From 9aabe3b4e5e9feb84da418e44244556fd0aa1933 Mon Sep 17 00:00:00 2001 From: Amin Roudaki Date: Sat, 19 Sep 2026 23:19:11 -0700 Subject: [PATCH 001/168] feat(bin): defer the wedge escalation for a lane parked at a supervisor-owed gate (#4974) * fix(watch): recheck a gate awaiting a human instead of wedge-escalating it A lane whose validation run is parked at a gate waiting on a human decision is correctly quiet, but nothing in its status line says so: the evidence is the pipeline's own gate state rather than anything the worker wrote. The wedge timer read that silence as a suspected wedge and climbed the escalation ladder for as long as the wait lasted, and each escalation cost a supervising turn. The landed declared-wait consult does not reach it, because a live ordinary crewmate never reports a declared pause, and raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for every lane by the same amount. The threshold now reads a second, independent record when the status line accounts for nothing: whether the crew's current state is a gate whose answer is owed by a human. That is minted only from the gate's own findings table, by a row whose `action` column is exactly `ask-user`, located by position out of the table header the way nm_gate_step_row already reads its row - never searched for over the run payload, where a finding's free-text description or a branch name satisfies a search just as well. A gate awaiting the CREWMATE's own answer keeps the unchanged escalation schedule, reason and demand-deep-inspection wording, because a crewmate that goes quiet before answering its own gate is exactly the wedge the ladder exists to catch. Each kind of wait now carries the human it is on, the action that clears it, and whether that human is the captain as data alongside the verdict, rather than as wording chosen per branch where the recheck is written, so the deferral cannot word one kind of wait as another and a new kind cannot ship without deciding all of them. A parked gate has no written record of when its wait began, so its recheck publishes no wait age at all rather than one read from the quiet window this deferral resets on every pass, which would report the same small number for a gate of any age. Like every other captain-facing recheck here it is absorbed in silence while the away-posture record exists, arming no throttle, so the recheck is owed in full the moment the record is archived. The consult runs only in the at-threshold branch that was about to escalate, beside the worktree walk already there, and only for lanes whose status line explained nothing. Closes #3055 * no-mistakes(review): require an unanswered decision before deferring a parked gate * no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records * test(watch): pass the pane hash wedge_timer_check now takes Upstream gave wedge_timer_check a sixth argument for its dead-record probe. The malformed-wait-record rounds drive the real function directly, so they pass one, and stub fm_backend_agent_state to a live agent so the probe that runs after a refused deferral keeps the unchanged ladder rather than reading a backend the child shell has none of. * no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate * no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling * feat(watch): make the parked-gate wait deferral opt-in The wedge timer deferring a lane parked at a validation gate is new supervision behaviour rather than a restored one, and it decides which lanes give up the escalation ladder, so it now ships as a default-off per-home option instead of changing every home on upgrade. config/wedge-defer-parked-gate arms it. The flag is read before the decision fold, so an unconfigured home spends no fold or current-state read, writes no record, and keeps the unchanged escalation schedule, reasons and demand-deep-inspection wording; a test counts the reader calls in both directions to pin that. It is not inherited by secondmate homes: each home supervises its own crew and owns that trade separately, the same reason config/turnend-churn-absorb is home-local. The away-posture absorb returns to leaving the idle timer alone, which it had restarted only because the costly consult could reach it. A parked-gate wait is owed to the supervisor rather than the captain, so it never enters that branch, and the recheck owed on return is again owed in full the moment the record is archived. * test(watch): pin that the away-silenced hold leaves the idle timer alone The absorb no longer restarts the timer, so the recheck owed on return is owed in full rather than a cadence into the return. Nothing asserted that, so a restart could be reintroduced silently. * no-mistakes(review): document away-silence rationale, pin captured gate component * no-mistakes(test): anchor gate row scan to the braced findings header * no-mistakes(document): pin same-block gate row invariant in crew-state comment --- AGENTS.md | 1 + bin/fm-classify-lib.sh | 78 ++++++ bin/fm-crew-state.sh | 85 ++++++- bin/fm-dod-lib.sh | 8 + bin/fm-watch.sh | 291 +++++++++++++++++------ docs/architecture.md | 34 ++- docs/configuration.md | 17 +- tests/fm-crew-state.test.sh | 197 ++++++++++++++++ tests/fm-watch-triage.test.sh | 433 +++++++++++++++++++++++++++++++++- tests/wake-helpers.sh | 4 + 10 files changed, 1048 insertions(+), 100 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index dd6062f9b6d..619beea3103 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -82,6 +82,7 @@ config/stow-pass-horizon optional presence flag opting this home in to /stow's config/herdr-presentation-spaces optional "off" opt-out from, or "on" opt-in to, Herdr's default-on disposable single-task visual projection, which is unconfigured-default-on only at or above a Herdr version floor; LOCAL, gitignored; inherited by secondmate homes; see docs/herdr-backend.md "Presentation spaces" config/trace-context optional presence flag enabling default-off native W3C trace-context propagation to spawned agents; LOCAL, gitignored; inherited by secondmate homes; see docs/configuration.md "Trace context propagation" and docs/trace-context.md config/turnend-churn-absorb optional presence flag opting this home into the default-off absorb of bare turn-end wakes on pane churn; LOCAL, gitignored, and not inherited; see docs/configuration.md "Turn-end pane-churn absorb" +config/wedge-defer-parked-gate optional presence flag opting this home into the default-off deferral of a wedge escalation for a lane parked at a validation gate awaiting the supervisor's own still-open decision; LOCAL, gitignored, and not inherited; see docs/configuration.md "Parked-gate wait deferral" config/cmux-socket-password optional cmux control-socket password; LOCAL, gitignored; read fresh on every cmux CLI call and passed through without ever overriding an operator's own ambient CMUX_SOCKET_PASSWORD when absent (docs/cmux-backend.md "Setup") config/wedge-alarm optional away-mode wedge-alarm active-alert directives; LOCAL, gitignored; absent means auto (macOS Notification Center when available); see docs/wedge-alarm.md config/watched-tools.json optional list of the tools this home depends on, read by the update check armed with bin/fm-tool-update-check.sh; LOCAL, gitignored, firstmate-maintained but human-editable, and NOT inherited by secondmate homes; see docs/configuration.md "Watched tool updates" diff --git a/bin/fm-classify-lib.sh b/bin/fm-classify-lib.sh index 0752e52f370..f737a7d96dd 100755 --- a/bin/fm-classify-lib.sh +++ b/bin/fm-classify-lib.sh @@ -621,6 +621,36 @@ EOF printf '%s\n' "$current" } +# 0 when the fold above still holds at least one decision OPENED by +# `needs-decision` - the status side's own record that a human was asked +# something and has not answered. A `blocked` record is deliberately not this: a +# blocker is an obstacle the crew reported, not an unanswered question, and a +# different action clears it. Whole-file and cursor-free on purpose: this answers +# a point-in-time question for a caller that holds no cursor and must not write +# one, so it reads status_open_decisions rather than the incremental fold. +# An unreadable, missing or symlinked status file folds to nothing and answers 1, +# which is the safe answer for every caller: no evidence, no exception. +# Given a , only a decision whose key is exactly `nm--` for +# a non-empty step counts - the key shape the brief mandates for a gate +# escalation - so an unrelated question left open earlier in the same task is +# never read as firstmate being told about THIS run's gate. +status_has_open_needs_decision() { # [] + local run=${2-} open line key verb + open=$(status_open_decisions "$1") + [ -n "$open" ] || return 1 + if [ $# -ge 2 ] && [ -z "$run" ]; then return 1; fi + while IFS= read -r line; do + key=${line%%$'\t'*} + verb=${line#*$'\t'}; verb=${verb%%$'\t'*} + [ "$verb" = needs-decision ] || continue + [ $# -ge 2 ] || return 0 + case "$key" in "nm-$run-"?*) return 0 ;; esac + done < has a record in a folded "\t\t" open set. _fm_open_set_has() { # case "$1" in @@ -1969,6 +1999,54 @@ crew_is_paused() { # [ "$(crew_absorb_class "$1")" = paused ] } +# The one spelling of the verdict component that says a parked gate's answer is +# owed by a HUMAN. bin/fm-crew-state.sh mints it (nm_gate_awaits_human_decision +# owns the derivation: the findings table's `action` column, read by position); +# crew_gate_awaits_human_decision below is its only consumer. +FM_GATE_HUMAN_DECISION='ask-user: authority decision' + +# 0 if crew 's authoritative current state is a no-mistakes gate whose answer +# is owed by a human rather than by the crewmate itself. +# +# `parked` alone cannot answer this: the gate's shape (awaiting_approval, +# fix_review, awaiting_agent) is reported parked in every case and does not by +# itself say who owes the answer; only a findings row whose `action` column is +# exactly `ask-user` does. A crewmate that goes quiet before answering its OWN +# gate is precisely the wedge the escalation ladder exists to catch, so only the +# minted component above - never the parked verdict, the gate name, or the +# finding text - admits a lane here. +# +# The whole component is compared for equality rather than searched for, so a +# gate name or a reconciliation note that happens to contain the words cannot +# mint it downstream either. +# On success it prints the reported run id, read from the line's whole +# `run: ` component, so the caller can bind the gate to the decision that +# names that run; a line carrying no run id is not evidence, since nothing could +# then tie a decision to this gate. +# Same cost and the same caveat as crew_absorb_class: one fm-crew-state.sh read, +# which may make a bounded no-mistakes call, so callers take it only where they +# already accept that cost. +crew_gate_awaits_human_decision() { # -> on stdout + local id=$1 line state src rest part human='' run='' + [ -n "$id" ] || return 1 + line=$("$FM_CREW_STATE_BIN" "$id" 2>/dev/null) || true + case "$line" in state:*) ;; *) return 1 ;; esac + state=${line#state: }; state=${state%% *} + [ "$state" = parked ] || return 1 + src=${line#*source: }; src=${src%% *} + [ "$src" = run-step ] || return 1 + rest="$line · " + while [ -n "$rest" ]; do + part=${rest%% · *} + rest=${rest#* · } + [ "$part" = "$FM_GATE_HUMAN_DECISION" ] && human=1 + case "$part" in "run: "?*) run=${part#run: } ;; esac + done + [ -n "$human" ] && [ -n "$run" ] || return 1 + case "$run" in *[[:space:]]*) return 1 ;; esac + printf '%s\n' "$run" +} + # Directories excluded from the worktree write probe below, and the depth it walks. # The excluded set is everything a supervisor read or a package manager can write # without the crew doing any work - .git first, so firstmate's own read-only git diff --git a/bin/fm-crew-state.sh b/bin/fm-crew-state.sh index 03f56c2fc88..f61e8d48653 100755 --- a/bin/fm-crew-state.sh +++ b/bin/fm-crew-state.sh @@ -491,6 +491,84 @@ nm_gate_findings_count() { case "$rest" in ''|*[!0-9]*) return 0 ;; esac printf '%s' "$rest" } +# 0 when the gate's own findings table holds at least one row whose `action` +# column is exactly `ask-user` - the pipeline's own record that this gate's +# answer is owed by a HUMAN, not by the crewmate (the gate's shape - +# awaiting_approval, fix_review, awaiting_agent - is reported parked in every +# case and does not by itself say who owes the answer; only a findings row whose +# `action` column is exactly `ask-user` does). +# +# Read POSITIONALLY, the way nm_gate_step_row above reads its row: locate the +# `findings[N]{...}` header, take the index of the `action` column from it, walk +# each of the N rows that follow to that index, and compare for EQUALITY. A +# substring search over the run payload cannot make this distinction - the +# trailing `description` column is free text that routinely quotes finding +# actions, and the payload also carries the branch name and step names, so a +# gate owed the crewmate's own answer would match just as readily as one owed a +# human. Column order is read from the header rather than assumed, so a table +# that grows a column keeps answering correctly. Both the header match and the +# row scan require the BRACE, so the count, the index and the rows all come from +# the same block: an earlier unbraced `findings[N]:` line from a resolved round +# must not supply the rows while the braced gate table supplies the index, which +# would read the wrong block's rows at the right block's offset +# (tests/fm-crew-state.test.sh's unbraced-precursor case pins it). +# +# Reading the index out of the header and then walking RAW COMMAS to it is only +# positional in name: the walk is sound only while every column before `action` +# is comma-free, and the producer does not quote commas inside `description` +# (tests/fm-crew-state.test.sh's own fixture proves it). A header ordering that +# puts free text before `action` would therefore let a row's description mint +# the marker - silently, with no error - which is the same class of hole the +# positional derivation exists to close, arriving by a different route. So the +# columns preceding `action` are checked against a WHITELIST of names this table +# is known to carry as short comma-free scalars, and anything else refuses: +# a whitelist rather than a blacklist of free-text names, because an unknown +# column must read as unsafe rather than as safe. When the table's shape is not +# provably safe the correct answer is the noisy one - a crewmate that went quiet +# before answering its own gate is the failure that must never be silenced. +# Residual bound, which no unquoted positional parse of this table escapes: a +# comma inside a whitelisted field's own value (a path with a comma in it, say) +# still shifts the walk. +nm_gate_awaits_human_decision() { + local header count cols idx i name field rows row rest + header=$(printf '%s\n' "$RUN_OUT" | grep -E '^[[:space:]]*findings\[[0-9]+\]\{[^}]*\}:' | head -1) + [ -n "$header" ] || return 1 + count=$(printf '%s' "$header" | sed -n 's/^[[:space:]]*findings\[\([0-9][0-9]*\)\].*/\1/p') + case "$count" in ''|*[!0-9]*) return 1 ;; esac + [ "$count" -gt 0 ] || return 1 + cols=$(printf '%s' "$header" | sed -n 's/^[^{]*{\([^}]*\)}.*/\1/p') + [ -n "$cols" ] || return 1 + idx=0 + i=0 + while [ -n "$cols" ]; do + i=$((i + 1)) + name=$(strip_quotes "$(trim "${cols%%,*}")") + if [ "$name" = action ]; then idx=$i; break; fi + case "$name" in + id|severity|file|line) ;; + *) return 1 ;; + esac + case "$cols" in *,*) cols=${cols#*,} ;; *) cols='' ;; esac + done + [ "$idx" -gt 0 ] || return 1 + rows=$(printf '%s\n' "$RUN_OUT" \ + | awk -v n="$count" 'f { print; if (++c >= n) exit; next } /^[[:space:]]*findings\[[0-9]+\]\{/ { f = 1 }') + while IFS= read -r row; do + case "$row" in *,*) ;; *) continue ;; esac + rest=$row + i=1 + while [ "$i" -lt "$idx" ]; do + case "$rest" in *,*) rest=${rest#*,} ;; *) rest=''; break ;; esac + i=$((i + 1)) + done + [ -n "$rest" ] || continue + field=$(strip_quotes "$(trim "${rest%%,*}")") + [ "$field" = ask-user ] && return 0 + done < ' } +# The `nm--` decision key this block mandates is load-bearing beyond +# the brief itself: the watcher binds an open `needs-decision` to the run a +# crew's current state reports by matching exactly that shape +# (wedge_wait_evidence in bin/fm-watch.sh, through +# status_has_open_needs_decision in bin/fm-classify-lib.sh), which is what buys +# a lane parked at a human-owed gate the long recheck cadence instead of a +# wedge escalation. A gate escalated under any other key still reads as a +# suspected wedge. fm_ask_user_escalation_block() { # local data=$1 id=$2 cat < triage_log "absorbed $label (worktree written since the idle window opened, idle ${age}s): $win" } +# One wait record, carrying every field a recheck needs to be correct. Emitting +# them together is the point: a recheck that names the wrong human, or asks for +# an action that does not clear the lane, points the reader away from the only +# person who can end the wait, so a new kind of evidence must not be able to +# reach wedge_defer_wait without deciding all of them. +# what the evidence IS, as the recheck names it. +# WHO the wait is on, in the recheck's own words. +# `captain` when that subject is the captain, `supervisor` when +# it is firstmate itself, `external` otherwise; this is what +# applies the away-posture rule below, which only `captain` takes. +# the one thing that clears the lane. +# the file whose mtime is when this wait started, or EMPTY when +# the wait has no written record. Empty is a real answer, not a +# degraded one: a gate the pipeline parked was never written down +# by the worker, so there is no honest age to publish and the +# deferral publishes none. +# The fields are joined with US (\037) rather than TAB because TAB is an IFS +# WHITESPACE character: consecutive tabs collapse under `read`, so a record with +# an empty middle field would not fail to parse, it would SHIFT every later field +# left into another field's position. US is not IFS whitespace, so consecutive +# delimiters yield genuinely empty fields and the record either parses as written +# or fails the deferral's guard. +wait_record() { # + printf '%s\037%s\037%s\037%s\037%s' "$1" "$2" "$3" "$4" "$5" +} + # The evidence that a quiet pane is a BOUNDED WAIT rather than a wedge suspect, # read at the one moment it decides anything: when an escalation is about to -# fire. The worker's own status line is that evidence - a declared `paused:` -# external wait, or a verified `captain-held` transfer. +# fire. Two records answer it, and they are independent: the worker's own status +# line - a declared `paused:` external wait, or a verified `captain-held` +# transfer - and, when that line explains nothing, the crew's authoritative +# current state. # # The generated brief promises that declaring one buys the long recheck cadence # instead of a wedge, and the wedge timer is reachable while that declaration @@ -937,75 +968,170 @@ wedge_defer_writing() { # # # A declared clearing time that has ALREADY passed (`paused: ... until `) is # not evidence: the wait the worker described is over, so it no longer explains -# the silence, and the pane keeps the unchanged schedule. -# Nothing here weakens detection for a pane with no declaration - it never runs -# for them beyond one status-line read, and their escalation schedule, reason and -# wording are untouched. -# WHICH verb declared it is printed, not just that one did, because the caller -# must not re-derive it: the two block on DIFFERENT humans - `paused:` on an -# external dependency the worker named, `captain-held:` on the captain themself - -# so a recheck that named the wrong one would point the reader away from the -# person who can clear it. -wedge_wait_evidence() { # -> `declared` or `held` on stdout - local task=$1 last until +# the silence, and the pane keeps the unchanged schedule. The records are read in +# this order rather than pooled because the routing already guarantees it is the +# right one: a pane whose last line is `paused:` or `captain-held:` reaches this +# timer only through pause_state_class answering `working`, so its crew state is +# a running step, never a parked gate. +# +# The second record is OFF unless the home creates config/wedge-defer-parked-gate, +# and that one guard is what makes an unconfigured home's behaviour identical to +# having no second record at all: it is read before the fold, so no fold or +# crew-state read is spent, no wait record exists to defer on, no recheck wording +# is reachable, and the lane keeps the unchanged escalation schedule, reason and +# demand-deep-inspection wording. Unlike the status line, which is the worker's +# own declaration about its own silence, this record is derived from a pipeline's +# gate state, so which lanes lose the ladder for it is a home's choice to make +# rather than a default every fleet inherits - the same reason +# config/turnend-churn-absorb gates its own widened absorb. +# +# The second record takes TWO signals, and needs both. The crew's authoritative +# current state must be a no-mistakes gate whose answer is owed by a HUMAN +# (crew_gate_awaits_human_decision in fm-classify-lib.sh, minted from the +# findings table's `action` column by position), AND the task's own decision fold +# must still hold an open `needs-decision` record whose key is `nm--` +# for the run that verdict reports. The gate's table alone says only that the +# answer is owed by a human; the open decision bound to that run is the positive +# evidence that firstmate was actually told about THIS gate and has not answered +# yet, which is what makes the lane's quiet a wait rather than a suspected wedge. +# An open decision under any other key - an unrelated question never closed - is +# not that evidence, and neither is a verdict that names no run. The wait is owed +# by firstmate, not the captain: ask-user findings are routed to firstmate, which +# decides most of them itself, and one it escalates becomes a captain-held +# transfer that the first record above already catches. So the away-posture +# silence does not apply to it: under away posture the supervision branch is the +# actor allowed to answer it, and it is rechecked on the long cadence throughout. +# The two signals come apart in both directions, and the ladder is kept in each: +# - the decision was ANSWERED and the crewmate has not yet relayed it with +# `axi respond`: the gate is still reported parked and still carries the +# ask-user row, but `fm-send --resolve-key` wrote the closing `resolved` line +# at answer time, so the fold is empty and what is outstanding is the +# crewmate's OWN next move; +# - the crewmate parked at a human-owed gate and went quiet before escalating +# it at all: nobody was ever told, so there is no wait to defer to. +# A `blocked` record does not count: a blocker is not an unanswered gate decision +# and a different action clears it. A gate awaiting the CREWMATE's own answer is +# deliberately NOT evidence either: a crewmate that goes quiet before answering +# its own gate is exactly the wedge this ladder exists to catch, so those keep +# the unchanged schedule, reason and demand-deep-inspection wording. +# Nothing here weakens detection for a pane with no wait at all - their +# escalation schedule, reason and wording are untouched, and every way this +# signal can come back empty (an unreadable status file, a fold with nothing +# open, a key convention nobody followed) escalates on the unchanged schedule +# rather than losing the ladder. The status-line and fold reads are file reads; +# the crew-state read is the costly one (it may make a bounded no-mistakes call), +# so it is taken only behind a first fold read that finds some open +# `needs-decision` at all, and only in the at-threshold branch - at most once per +# window per STALE_ESCALATE_SECS, never on an ordinary poll. +wedge_wait_evidence() { # -> one wait_record on stdout + local task=$1 last until statusf run [ -n "$task" ] || return 1 - last=$(last_status_line "$STATE/$task.status") + statusf="$STATE/$task.status" + last=$(last_status_line "$statusf") if status_is_captain_held "$last"; then - printf 'held' + wait_record 'captain-held' 'awaiting the captain - verified hold transfer' \ + captain 'answer the held decision or release the hold' "$statusf" return 0 fi - status_is_paused "$last" || return 1 - if until=$(status_paused_until "$last"); then - [ "$(date +%s)" -lt "$until" ] || return 1 + if status_is_paused "$last"; then + if until=$(status_paused_until "$last"); then + [ "$(date +%s)" -lt "$until" ] || return 1 + fi + wait_record 'declared wait' 'awaiting external' \ + external 'confirm the wait still holds' "$statusf" + return 0 + fi + [ -e "$CONFIG/wedge-defer-parked-gate" ] || return 1 + if status_has_open_needs_decision "$statusf" \ + && run=$(crew_gate_awaits_human_decision "$task") \ + && status_has_open_needs_decision "$statusf" "$run"; then + wait_record 'verified wait at a parked gate' "awaiting firstmate's ask-user decision" \ + supervisor "decide the gate's ask-user finding and relay the decision to the crewmate" '' + return 0 fi - printf 'declared' + return 1 } -# Defer ONE wedge escalation for a pane whose own declaration explains the quiet +# Defer ONE wedge escalation for a pane whose wait record explains the quiet # (wedge_wait_evidence above). Deliberately the same shape as # wedge_defer_writing: a DEFERRAL, not a cancellation, so the idle timer restarts # and the next window probes the evidence again - a wait that ends is escalating # again within one STALE_ESCALATE_SECS, which is why the worst-case detection # time for a pane that stops waiting does not move. -# How long the wait has held is read from the status file, which is when the -# worker wrote the line - anchored there rather than on a per-window marker for -# the same reason handle_paused_stale is: an idle pane churns its display (a -# clock, a token counter), and a marker this deferral kept touching would let -# that churn reset the cadence. -# The recheck names WHICH human the wait is on, for the same reason -# handle_paused_stale does: a hold is owed by the captain reading the recheck, so -# wording it as an external dependency points them away from the one action that -# clears it. -# A HOLD is not rechecked at all while the away-posture record exists: the one -# human who can answer it is away, the return brief already lists it, and every -# other captain-held path in this file absorbs it silently for that reason -# (handle_paused_stale, surface_nonterminal_stale, captain_call_stale_bound). -# That absorb arms no throttle, so the recheck is owed in full the moment the -# record is archived rather than starting a cadence nobody could act on. +# Every word of the recheck that could be wrong per kind of evidence - the human +# it names, the action it asks for, the age it publishes - is READ FROM THE +# RECORD rather than re-derived here, so this function cannot word one kind of +# wait as another. +# A wait with a written record is aged from that file, which is when the worker +# wrote the line - anchored there rather than on a per-window marker for the same +# reason handle_paused_stale is: an idle pane churns its display (a clock, a +# token counter), and a marker this deferral kept touching would let that churn +# reset the cadence. A wait with NO written record publishes no age at all: the +# quiet window is the only clock in hand and this deferral resets it on every +# pass, so a number read from it would never grow and would tell a supervisor +# that a day-old gate opened four minutes ago. The bounded re-surface still +# fires, governed by its own throttle instead of by a wait age. +# A CAPTAIN-facing wait is not rechecked at all while the away-posture record +# exists: the one human who can answer it is away, the return brief already lists +# it, and every other captain-facing path in this file absorbs it silently for +# that reason (handle_paused_stale, surface_nonterminal_stale, +# captain_call_stale_bound). That absorb arms no throttle and deliberately +# leaves the idle timer alone: a `captain` whom is minted only by the +# captain-held arm of wedge_wait_evidence, which returns before the +# wedge-defer-parked-gate flag test and therefore before any decision-fold or +# current-state read, so the only read that repeats under the away record is the +# one status-line read that predates this deferral. There is nothing costly to +# throttle there, so the recheck owed on return stays owed in full the moment the +# record is archived rather than starting a cadence nobody could act on. The +# costly parked-gate consult is owed to the supervisor instead, never silenced +# here, and its own deferral restarts the timer below. # The escalation counter is left alone, exactly as the write deferral leaves it: # this is not an escalation, and a later genuine one must keep the # demand-inspection history it had already earned. -wedge_defer_wait() { # - local win=$1 task=$2 since_file=$3 label=$4 age=$5 evidence=$6 key mtime wage min_age kind action waited - if [ "$evidence" = held ]; then - if afk_record_present; then - triage_log "absorbed $label (captain-held, never rechecked while the away-posture record exists): $win" - return 0 - fi - kind='captain-held, awaiting the captain - verified hold transfer' - action='answer the held decision or release the hold' - else - kind='declared wait, awaiting external' - action='confirm the wait still holds' +wedge_defer_wait() { # + local win=$1 since_file=$2 label=$3 age=$4 record=$5 + local kind subject whom action anchor key mtime wage min_age waited us ok + us=$(printf '\037') + IFS=$us read -r kind subject whom action anchor < < clear_write_tracking "$key" date +%s > "$since_file" resurface_absorbed "$win" "$STATE/.waiting-resurfaced-$key" "$wage" \ - "stale: $win (idle ${age}s${waited} - $kind, rechecked on a long cadence not a wedge; $action)" \ + "stale: $win (idle ${age}s${waited} - $kind, $subject, rechecked on a long cadence not a wedge; $action)" \ '' "$min_age" - triage_log "absorbed $label (the pane's own wait explains the quiet, idle ${age}s): $win" + triage_log "absorbed $label ($kind explains the quiet, idle ${age}s): $win" + return 0 } # Drop a window's write-deferral chain wherever its stale bookkeeping resets, so @@ -1048,7 +1175,7 @@ clear_write_tracking() { # # because the pane might still be working; this is terminal for as long as the # endpoint stays gone, because there is nothing left to re-probe on a cadence and a # repeat is exactly the noise it exists to stop. WHICH verdict fired is named for -# the same reason wedge_wait_evidence names its verb: the two ask the supervisor +# the same reason wedge_wait_evidence names its kind of wait: the two ask the supervisor # for different things. # # The marker is owned entirely by this function and records the verdict together @@ -1100,19 +1227,22 @@ wedge_dead_record() { # local win=$1 since_file=$2 label=$3 escalation_file=$4 task=$5 hash=$6 since age n reason evidence since=$(cat "$since_file" 2>/dev/null || true) @@ -1127,8 +1257,8 @@ wedge_timer_check() { # # the expected external wait. The caller has already confirmed liveness through # the busy verdict, so this exception does not suppress undeclared wedges or # alter the separate non-busy classification. handle_paused_stale keeps the -# exception bounded by re-surfacing it once per PAUSE_RESURFACE_SECS. Away mode -# remains daemon-owned and receives the undecorated wake identity for its own -# classification, which is why the declaration is read before the afk branch -# rather than after it. +# exception bounded by re-surfacing it once per PAUSE_RESURFACE_SECS. +# A pane that declared nothing falls through to the shared wedge timer, which, +# in a home that armed config/wedge-defer-parked-gate, applies the same rule to +# the one wait a busy pane cannot declare: a validation gate of its own awaiting +# a supervisor decision that is still open also takes the bounded recheck rather +# than the ladder, because who owes that answer does not depend on what the pane +# is rendering, and the recheck names that supervisor and the action that clears +# it. An unconfigured home keeps the unchanged ladder there. +# Away mode remains daemon-owned and receives the undecorated wake identity for +# its own classification, which is why the declaration is read before the afk +# branch rather than after it. busy_turn_bound_check() { # local win=$1 task=$2 h=$3 since_file=$4 escalation_file=$5 key statusf declared statusf="$STATE/$task.status" diff --git a/docs/architecture.md b/docs/architecture.md index 094b74df3dc..72c64aa55d7 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -9,7 +9,7 @@ firstmate's supervisor contract and routing index for conditional procedures is ## Event-driven supervision A zero-token bash watcher (`bin/fm-watch.sh`) sleeps on the fleet, classifies detected wakes in bash, and wakes the first mate only when something is actionable. -Actionable wakes include captain-relevant status signals, no-verb signals without positive evidence that their crew is still executing, authenticated check output such as PR merge polling or a Relay mention, stale panes whose crew is not provably working whether their status log looks terminal or non-terminal, provably-working stale panes that persist past `FM_STALE_ESCALATE_SECS` with neither a wait their own worker declared nor their own task worktree being written, declared external waits and attended captain-held transfers that remain declared past `FM_PAUSE_RESURFACE_SECS`, and heartbeat backstop hits. +Actionable wakes include captain-relevant status signals, no-verb signals without positive evidence that their crew is still executing, authenticated check output such as PR merge polling or a Relay mention, stale panes whose crew is not provably working whether their status log looks terminal or non-terminal, provably-working stale panes that persist past `FM_STALE_ESCALATE_SECS` with no wait their own worker declared, no writes to their own task worktree, and - in a home that armed `config/wedge-defer-parked-gate` - no validation gate of their own awaiting an unanswered supervisor decision, declared external waits and attended captain-held transfers that remain declared past `FM_PAUSE_RESURFACE_SECS`, and heartbeat backstop hits. For an ordinary crew task, a wait is read from both of its records: the status line a worker declared, and the backlog hold `bin/fm-captain-hold.sh` recorded once firstmate handed the work to the captain. So a delivered ordinary crew task whose last line stays `done: PR ...` bounds repeated alarms from new pane hashes to the `FM_PAUSE_RESURFACE_SECS` cadence for the length of the captain's decision. The first hash still alarms, each new hash inside that window is absorbed, and a new hash after the window re-surfaces the hold; a terminal pane hash that never changes stays inert after its first alarm exactly as it did before this bound. @@ -20,14 +20,31 @@ Repeated provably-working stale escalations on the same unchanged pane add an es In the same branch that is about to escalate, the pane's own account of its quiet is consulted first: the worker's declared `paused:` or verified `captain-held` status line. That declaration defers the escalation to the `FM_PAUSE_RESURFACE_SECS` recheck cadence instead, because a lane waiting on something it named is silent for a reason the escalation would misreport, and the ladder would otherwise climb for as long as the wait lasts. A declared clearing time (`paused: ... until `) that has already passed stops counting as that account, so a lane whose own wait is over, and a lane that never declared one, both keep the unchanged escalation schedule, reason and `demand-deep-inspection` wording. -Which verb declared it decides how the recheck is worded, because the two block on different people: a `paused:` declaration is owed by an external dependency the worker named and asks the reader to confirm the wait still holds, while a hold is owed by the captain reading the recheck and asks them to answer the held decision or release the hold. -Wording a hold as an external wait would point the captain away from the one action that clears it. -Both are aged from the status file, since that is when the worker wrote the line; anchoring on a per-window marker instead would let a churning display reset the cadence. -While the away-posture record exists a hold is not rechecked here at all, as on every other captain-held path: there is nobody to answer it and the return brief already lists it, so the pane is absorbed silently and no re-surface throttle is armed, leaving the recheck owed in full the moment the record is archived. -The consult costs one status-line read, taken in the same at-threshold branch as the worktree walk and never on an ordinary poll. +When the status line accounts for nothing and this home armed the default-off `config/wedge-defer-parked-gate` flag, one further record is read: whether the crew's own current state is a validation gate whose answer is owed to the supervisor, who has actually been asked and has not answered. +That record exists because the quiet of a parked lane is the pipeline's doing rather than anything the worker wrote down, so no status-line predicate can see it: the signal naming who owes that answer and what clears it lives here rather than in any line the worker could write. +Arming it is a per-home choice because every other wait here is the worker's own declaration about its own silence, while this one is derived from a pipeline's gate state, so which lanes give up the escalation ladder is a decision each home makes for itself. +A home that has not armed it reads no further record at all: the flag is tested before the fold, so no fold or current-state read is spent, no wait record exists to defer on, and every parked lane keeps the unchanged escalation schedule, reason and `demand-deep-inspection` wording. +Its first half is minted only from the gate's own findings table, by a row whose `action` column is exactly `ask-user`, read by position out of the table header rather than searched for over the run payload, where a finding's free-text description or a branch name would satisfy a search just as well. +Because the row is then split on raw commas and the producer does not quote commas inside free text, the derivation refuses outright - keeping the ladder - unless every column the header places before `action` is one of the short comma-free scalars this table is known to carry (`id`, `severity`, `file`, `line`), so a header that grows an unrecognised or free-text column ahead of `action` reads as unsafe rather than as safe. +That precision is what keeps the distinction the ladder depends on: the gate's shape - `awaiting_approval`, `fix_review`, `awaiting_agent` - is reported parked in every case and does not by itself say who owes the answer, only a findings row whose `action` column is exactly `ask-user` does, and a crewmate that goes quiet before answering its own gate is exactly the wedge this ladder exists to catch, so a gate with no such row keeps the unchanged escalation schedule, reason and `demand-deep-inspection` wording. +Its second half is the task's own decision fold still holding an open `needs-decision` record whose key is `nm--` for the run the current state reports, which is the positive evidence that firstmate was told about this gate rather than merely that someone owes it an answer. +An open decision under any other key, such as an unrelated question left open earlier in the same task, is not that evidence, and neither is a current state that names no run, so both keep the unchanged ladder. +That half is what keeps the ladder in the two cases where a parked supervisor-owed gate is really the crewmate's move: a decision that has already been answered, where `fm-send --resolve-key` closed it at answer time while the gate stays parked until the crewmate relays it, and a crewmate that parked at such a gate and went quiet before escalating it at all, where nobody was ever told. +A `blocked` record is not that evidence, since a blocker is an obstacle the crew reported rather than an unanswered question, and a different action clears it. +Every way the fold can come back empty, including an unreadable status file, leaves the unchanged escalation schedule in place rather than taking the ladder away. +Each kind of wait carries the human it is on and the action that clears it as data alongside the verdict, rather than as wording chosen per branch where the recheck is written, so a new kind of evidence cannot reach the deferral without deciding both. +The deferral refuses a record that does not carry all of them and escalates as it would have, because deferring on a half-filled record is what would print the wrong human or an action that clears nothing. +The three block on different people: a `paused:` declaration is owed by an external dependency the worker named and asks the reader to confirm the wait still holds, a hold is owed by the captain reading the recheck and asks them to answer the held decision or release the hold, and a parked gate is owed firstmate's `ask-user` decision and asks for that finding to be decided and relayed to the crewmate, because ask-user findings are routed to firstmate, which decides most of them itself, and one it escalates becomes a captain-held transfer that the hold record already covers. +Wording any of them as another would point the reader away from the one action that clears it. +A wait with a written record is aged from the status file, since that is when the worker wrote the line; anchoring on a per-window marker instead would let a churning display reset the cadence. +A parked gate has no such record - the worker never wrote the wait down - so its recheck publishes no wait age at all rather than one read from the quiet window, which this deferral resets on every pass and which would therefore report the same small number for a gate of any age. +While the away-posture record exists a hold is not rechecked here at all, as on every other captain-facing path: there is nobody to answer it and the return brief already lists it, so the pane is absorbed silently and no re-surface throttle is armed, leaving the recheck owed once the record is archived. +That absorb deliberately leaves the idle timer alone too, because only a captain-held hold ever reaches it and that verdict is reached before the armed-home flag is tested and so before any fold or current-state read: the one read repeating under the away record is the status-line read that predates this deferral, nothing costly enough to throttle, so the recheck owed on return stays owed in full the moment the record is archived rather than starting a cadence nobody could act on. +A parked gate is not silenced that way, because it is owed to the supervisor rather than the captain and under away posture the supervision branch is the actor allowed to answer it, so it keeps the long recheck cadence throughout. +The consult costs one status-line read, plus - only in an armed home - a status-log fold and then one current-state read for the lanes whose status line explained nothing and whose fold holds some open `needs-decision`, all taken in the same at-threshold branch as the worktree walk, so it is bounded to at most once per window per `FM_STALE_ESCALATE_SECS` and never runs on an ordinary poll. +The fold is read before the current state so a lane with no open decision never pays for the costlier read at all. A known bound: the recheck throttle is scoped to the pane hash, so the long cadence holds for a lane whose pane is genuinely static, while a lane whose display churns (a ticking clock, a token counter) drops the throttle with each new hash and is rechecked once per idle window instead. That lane still loses the escalation ladder and the `demand-deep-inspection` wording, which is the defect being fixed, but it is not the full delivery of a long cadence; the alternative, letting the throttle outlive the hash, trades this for a stale throttle surviving into an unrelated later episode and suppressing that episode's first recheck, which is the worse failure. -A lane that is quiet because its own validation run is parked at a gate awaiting a human decision is deliberately out of scope here and keeps the unchanged ladder: reading that state needs a signal that carries who the wait is on and what clears it, rather than one inferred from a parked verdict that also covers gates awaiting the crewmate itself. A pane holding a file newer than the start of its own quiet window, anywhere in the worktree recorded for that task, is deferred instead of escalated, because a crew writing source, then tests, then documentation behind a static pane is liveness that neither pane quietness nor the run step can show. That deferral re-surfaces on the same `FM_PAUSE_RESURFACE_SECS` cadence as a declared wait, with a reason naming the write evidence rather than a wedge, and it is bounded to one pruned, depth-bounded, wall-clock-bounded walk (`FM_WORKTREE_WRITE_PRUNE`, `FM_WORKTREE_WRITE_MAXDEPTH`, `FM_WORKTREE_WRITE_TIMEOUT`) taken only in the branch that was about to escalate, never on every poll. Every absence of write evidence, including a missing worktree record, a torn-down worktree, a walk that outlives its wall-clock bound on a hung mount, and a failed walk, leaves the existing escalation schedule untouched, so a crew that writes nothing still escalates exactly as before. @@ -39,7 +56,8 @@ The report decides nothing about the record's fate, because such a lane routinel The once-marker records the agent incarnation it was reported for - the task's per-incarnation busy gen (`state/.busy-gen`, minted by `bin/fm-busy-event.sh arm`, which changes exactly when the agent is replaced) - together with the verdict, so it re-arms when that endpoint reads live again and when the agent is replaced: a successor dying in the same window is reported again even when no threshold probe reads it alive in between and its dead display hashes identically to the one already reported. When no busy incarnation token is readable for the task (it was never armed, or its sidecar is unreadable), the marker falls back to keying on the pane hash: that keeps the once-per-display absorb for a record-less task rather than re-reporting on every threshold, at the residual cost that such a successor dying into a byte-identical dead display stays absorbed. A busy pane is otherwise exempt from staleness, but only until its last completed turn or explicit native-harness progress reaches `FM_BUSY_TURN_MAX_SECS` (`bin/fm-watch.sh` owns marker selection); past that bound it is routed through the same wedge escalation, with the identical reason, escalation count, worktree-write deferral, and `demand-deep-inspection` marker for a live agent and the same dead-record report when the endpoint is proven gone, for inspection only - never an automatic interrupt, signal, or restart. -A crew that declared an external wait (`paused:`) or a verified captain-held transfer is the one exception to that bound: its busy verdict supplies liveness while identifying the long-running foreground call as the declared wait, so it takes the bounded `FM_PAUSE_RESURFACE_SECS` recheck instead of a wedge escalation, except that a captain-held transfer is not rechecked while the away-posture record exists. +A crew that declared an external wait (`paused:`) or a verified captain-held transfer is the first exception to that bound: its busy verdict supplies liveness while identifying the long-running foreground call as the declared wait, so it takes the bounded `FM_PAUSE_RESURFACE_SECS` recheck instead of a wedge escalation, except that a captain-held transfer is not rechecked while the away-posture record exists. +In a home that armed `config/wedge-defer-parked-gate`, a crew whose own validation gate awaits the supervisor's still-open decision for that run is the second, reached through the shared wedge timer rather than the declaration branch, because who owes that answer does not depend on what the pane is rendering; it takes the same bounded recheck, including while the away-posture record exists. Lifting the declaration restores the unchanged busy-pane wedge path, while a pane that is no longer busy returns to the existing idle declared-wait classification. While the legacy daemon flag is active, a busy pane that crosses the bound under a declared external wait is handed to the daemon as the plain wake identity instead of taking that recheck in the watcher, because the daemon owns triage there and a wake already decorated as a possible wedge would override the daemon's own declared-wait verdict; an undeclared busy pane past the bound still takes the wedge escalation. That handoff is keyed on the declaration itself (the status log's signature) rather than on the pane capture, so a harness footer that ticks on every poll wakes the daemon once per declaration instead of once per poll, and it clears the wedge timer, escalation count, and worktree-write deferral exactly as the normal-mode absorber does, so an undeclared busy phase's timer does not resume when the declaration lifts. diff --git a/docs/configuration.md b/docs/configuration.md index 808ee716baa..8da010d8417 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -219,6 +219,15 @@ The bound is required rather than cosmetic because churn and pane staleness read The flag is a home-local supervision-noise preference and is not inherited by secondmate homes, which run their own crew mix. [`architecture.md`](architecture.md) owns the triage contract and `bin/fm-watch.sh`'s `signal_turnend_panes_churned` owns the exact evidence and fail-closed boundaries. +## Parked-gate wait deferral (config/wedge-defer-parked-gate) + +The optional local, gitignored `config/wedge-defer-parked-gate` presence flag opts this home into a default-off second form of wait evidence in the watcher's wedge timer. +With it present, a provably-working pane about to escalate is also deferred to the `FM_PAUSE_RESURFACE_SECS` recheck cadence when its crew's own current state is a validation gate whose answer is owed to the supervisor and whose decision for that run is still open, and the recheck names the supervisor and the action that clears the lane instead of reporting a suspected wedge. +It stays opt-in because the other evidence is the worker's own declaration about its own silence, while this is derived from a pipeline's gate state, so which lanes give up the escalation ladder for it is a home's choice. +With the flag absent the wedge timer spends no fold or current-state read for it, writes no record, and keeps the unchanged escalation schedule, reasons, and `demand-deep-inspection` wording. +The flag is a home-local supervision-noise preference and is not inherited by secondmate homes, which supervise their own crew and own that trade separately. +[`architecture.md`](architecture.md) owns the wait-evidence contract and which records may take the ladder away; `bin/fm-watch.sh`'s `wedge_wait_evidence` owns the exact derivation and its fail-closed boundaries. + ## Gate defaults (.no-mistakes.yaml) The tracked `.no-mistakes.yaml` sets `test.evidence.store_in_repo: true` and pins `commands.lint` to `bin/fm-lint.sh`, the same owner CI invokes. @@ -1089,7 +1098,7 @@ FM_CREW_STATE_NM_TIMEOUT=10 # seconds allowed per no-mistakes query inside fm- FM_TEARDOWN_NM_TIMEOUT=10 # seconds allowed per no-mistakes query or abort inside fm-teardown.sh FM_CREW_STATE_RUNS_LIMIT=200 # plain runs-ledger rows scanned for fallback attribution; does not change the CLI's AXI overview window (selection owner: bin/fm-nm-run-lib.sh) FM_TEARDOWN_NM_RUNS_LIMIT=200 # recent no-mistakes run rows scanned to prove an unresolved-head parked run belongs to teardown's task -FM_CREW_STATE_BIN=bin/fm-crew-state.sh # test override for the current-state reader used by working/paused watcher triage +FM_CREW_STATE_BIN=bin/fm-crew-state.sh # test override for the current-state reader used by watcher triage: the working/paused classification, and the wedge timer's parked-gate wait evidence FM_MAIL_USER= # mail-plane IMAP/SMTP login, from .env or environment (docs/configuration.md "Mail plane") FM_MAIL_PASS= # mail-plane IMAP/SMTP password FM_IMAP_HOST= # mail-plane IMAP server hostname @@ -1128,9 +1137,9 @@ FM_SIGNAL_GRACE=30 # seconds to coalesce nearby status and turn-end signals FM_TURNEND_CHURN_ABSORB_SECS=900 # longest one endpoint's bare turn-ends may be deferred on pane-churn evidence alone; only consulted when config/turnend-churn-absorb is present FM_CAPTAIN_RE='done:|needs-decision:|blocked:|failed:|PR ready|checks green|ready in branch|merged' # captain-relevant status regex; nonterminal progress verbs remain excluded even when their prose matches FM_CLASSIFY_PAUSED_VERB=paused # leading status verb for a declared external wait; excluded from FM_CAPTAIN_RE and distinct from blocked -FM_STALE_ESCALATE_SECS=240 # idle seconds before a provably-working stale pane escalates, unless that pane's own worker declared a wait that has not elapsed, which takes the FM_PAUSE_RESURFACE_SECS recheck below instead; stale panes whose crew is not provably working surface immediately unless admitted directly to the declared-wait cadence, while a live idle declared wait still surfaces once before that cadence bounds repeats; at that same escalation moment a recovery-grade agent-state probe (docs/architecture.md owns that dead-record contract) reports a pane whose endpoint is proven `dead` or `missing` once and stops re-escalating it while it stays that way -FM_BUSY_TURN_MAX_SECS=3600 # maximum age without a completed turn or explicit native-harness progress (bin/fm-watch.sh owns marker selection), before the same wedge escalation used for a provably-working non-busy stale takes over; inspection-only, never an automatic interrupt or restart; a declared external wait or attended verified captain-held transfer takes the FM_PAUSE_RESURFACE_SECS recheck below instead -FM_PAUSE_RESURFACE_SECS=14400 # four hours between bounded rechecks of a declared external wait or verified captain-held transfer, and between repeated new-hash stale alarms for an ordinary crew task with an open backlog captain call; a structured until time can make an external-wait recheck occur sooner but cannot extend this bound; this includes a live idle pane after its first inconclusive stale wake, a provably-working pane whose own unelapsed declared wait defers its FM_STALE_ESCALATE_SECS escalation, and a live busy pane past FM_BUSY_TURN_MAX_SECS, while the away-mode daemon uses the same setting and ages its window against the crew's own latest status line rather than pane busy state; a captain-held transfer is never rechecked while the away-posture record exists +FM_STALE_ESCALATE_SECS=240 # idle seconds before a provably-working stale pane escalates, unless that pane's own worker declared a wait that has not elapsed, or, where config/wedge-defer-parked-gate arms it, that pane's crew is parked at a validation gate awaiting the supervisor's decision on it that the crew raised under that run's key and nobody has answered yet, either of which takes the FM_PAUSE_RESURFACE_SECS recheck below instead; stale panes whose crew is not provably working surface immediately unless admitted directly to the declared-wait cadence, while a live idle declared wait still surfaces once before that cadence bounds repeats; at that same escalation moment a recovery-grade agent-state probe (docs/architecture.md owns that dead-record contract) reports a pane whose endpoint is proven `dead` or `missing` once and stops re-escalating it while it stays that way +FM_BUSY_TURN_MAX_SECS=3600 # maximum age without a completed turn or explicit native-harness progress (bin/fm-watch.sh owns marker selection), before the same wedge escalation used for a provably-working non-busy stale takes over; inspection-only, never an automatic interrupt or restart; a declared external wait, an attended verified captain-held transfer, or - where config/wedge-defer-parked-gate arms it - a validation gate of the crew's own awaiting the supervisor's still-unanswered decision takes the FM_PAUSE_RESURFACE_SECS recheck below instead +FM_PAUSE_RESURFACE_SECS=14400 # four hours between bounded rechecks of a declared external wait or verified captain-held transfer, and between repeated new-hash stale alarms for an ordinary crew task with an open backlog captain call; a structured until time can make an external-wait recheck occur sooner but cannot extend this bound; this includes a live idle pane after its first inconclusive stale wake, a provably-working pane whose own unelapsed declared wait or, where config/wedge-defer-parked-gate arms it, unanswered supervisor-owed validation gate defers its FM_STALE_ESCALATE_SECS escalation, and a live busy pane past FM_BUSY_TURN_MAX_SECS, while the away-mode daemon uses the same setting and ages its window against the crew's own latest status line rather than pane busy state; a captain-held transfer is never rechecked while the away-posture record exists, while an armed validation gate awaiting the supervisor's decision keeps this recheck in either posture FM_SECONDMATE_WAKE_STALL_SECS=180 # minimum interval with no change of the oldest actionable foreign wake-queue row (it advances as the mate drains, and a queue reprovisioned under the same task id starts a fresh interval at whatever sequence it restarts) before an endpoint-recorded local secondmate produces one durable parent wake-loop-stall notification for that no-progress episode; a mate that is provably inside an active turn (an exact busy verdict) does not escalate until that same no-progress interval reaches FM_BUSY_TURN_MAX_SECS above, declared external-wait pause rows are excluded, and zero or invalid values use 180 FM_WEDGE_DEMAND_INSPECT_COUNT=3 # consecutive provably-working stale escalations on the same unchanged pane before demand-deep-inspection is added FM_WORKTREE_WRITE_PRUNE='.git node_modules .venv venv __pycache__ .mypy_cache .pytest_cache .ruff_cache .tox target dist build .next .cache vendor' # directory names the wedge detector's task-worktree write probe skips; the default keeps .git out so a supervisor's own read-only git command can never look like crew progress; set it to the empty string to prune nothing, which widens the probe to the whole depth-bounded tree rather than disabling it diff --git a/tests/fm-crew-state.test.sh b/tests/fm-crew-state.test.sh index 2e247fb774e..7c4df16b96e 100755 --- a/tests/fm-crew-state.test.sh +++ b/tests/fm-crew-state.test.sh @@ -430,6 +430,95 @@ gate: review EOF } +# A gate owed the CREWMATE's own answer: every finding's `action` column is +# auto-fix. The free-text `description` column is where this repository's own +# review output routinely quotes finding actions, so one row spells the token out +# the way an enumeration does - surrounded by commas, in the exact shape a +# substring or unanchored-regex derivation would accept - and the branch name +# carries it too. Both are the counterexample: the ONLY thing that may mint the +# human-decision component is the `action` column read by position. +run_parked_crewmate_gate_with_ask_user_prose() { # + cat < + cat < + cat < + cat < cat </dev/null + fm_write_meta "$d/state/feat-au.meta" "window=fm:fm-feat-au" "worktree=$d/wt" "kind=ship" + printf 'needs-decision: review gate\n' > "$d/state/feat-au.status" + FM_FAKE_AXI_STATUS="$(run_parked fm/feat-au)" + out=$(run_crew_state "$d" feat-au) + assert_contains "$out" "state: parked" "an ask-user row still reports parked" + assert_contains "$out" " · ask-user: authority decision" \ + "an action column of ask-user mints the human-decision component" + + # The counterexample. Nothing here is owed a human: every action column is + # auto-fix. A description enumerating the action values, and a branch named + # after the same token, must not mint the component - a crewmate that goes + # quiet before answering its own gate has to keep the wedge ladder. + reset_fakes + d=$(new_case parked-ask-user-prose-only) + make_repo_on_branch "$d/wt" fm/ask-user-authority-fix + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-ap.meta" "window=fm:fm-feat-ap" "worktree=$d/wt" "kind=ship" + printf 'working: validation under way\n' > "$d/state/feat-ap.status" + FM_FAKE_AXI_STATUS="$(run_parked_crewmate_gate_with_ask_user_prose fm/ask-user-authority-fix)" + # Guard the counterexample against going vacuous: the payload this gate is read + # from must really contain the token in a position a substring or unanchored + # regex would accept, or the case below proves nothing. + assert_contains "$FM_FAKE_AXI_STATUS" ", ask-user," \ + "the counterexample payload must carry the token where a naive match accepts it" + assert_contains "$FM_FAKE_AXI_STATUS" "branch: fm/ask-user-authority-fix" \ + "the counterexample payload must also carry the token in its branch name" + out=$(run_crew_state "$d" feat-ap) + assert_contains "$out" "state: parked" "a crewmate-owed gate still reports parked" + assert_not_contains "$out" " · ask-user: authority decision" \ + "free text and a branch name must not mint the human-decision component" + + # Column order is read from the header, not assumed. + reset_fakes + d=$(new_case parked-ask-user-reordered) + make_repo_on_branch "$d/wt" fm/feat-ar + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-ar.meta" "window=fm:fm-feat-ar" "worktree=$d/wt" "kind=ship" + printf 'needs-decision: review gate\n' > "$d/state/feat-ar.status" + FM_FAKE_AXI_STATUS="$(run_parked_reordered_columns fm/feat-ar)" + out=$(run_crew_state "$d" feat-ar) + assert_contains "$out" " · ask-user: authority decision" \ + "the action column is located by header index, not by fixed position" + + # A header index alone is not enough, because the row is split on raw commas. + # With `description` ahead of `action` the comma walk lands inside free text, + # so a gate whose every action is auto-fix would mint the component. The table + # is not provably safe to walk, so the derivation must refuse and the crewmate + # must keep the wedge ladder. + reset_fakes + d=$(new_case parked-free-text-before-action) + make_repo_on_branch "$d/wt" fm/feat-af + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-af.meta" "window=fm:fm-feat-af" "worktree=$d/wt" "kind=ship" + printf 'needs-decision: review gate\n' > "$d/state/feat-af.status" + FM_FAKE_AXI_STATUS="$(run_parked_free_text_before_action fm/feat-af)" + # Non-vacuity: the payload must really carry the token at the comma offset the + # `action` index resolves to, or the case below proves nothing. + assert_contains "$FM_FAKE_AXI_STATUS" "findings[1]{id,severity,file,line,description,action}:" \ + "the fixture must really place free text before the action column" + assert_contains "$FM_FAKE_AXI_STATUS" ", ask-user," \ + "the fixture description must carry the token where the comma walk would accept it" + out=$(run_crew_state "$d" feat-af) + assert_contains "$out" "state: parked" "an unsafe findings header still reports parked" + assert_not_contains "$out" " · ask-user: authority decision" \ + "a findings header that puts free text before action must not mint the human-decision component" + + # The header and the rows must come from the SAME block. An earlier unbraced + # `findings[N]:` block ahead of the live gate's braced table would otherwise + # supply the rows while the braced header supplies the count and the `action` + # index, so the walk reads the wrong rows at the right index. Here that earlier + # block carries ask-user at exactly that offset while the live gate's only row + # is auto-fix: the crewmate owes this gate its own answer and must keep the + # wedge ladder. + reset_fakes + d=$(new_case parked-unbraced-findings-precursor) + make_repo_on_branch "$d/wt" fm/feat-ub + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-ub.meta" "window=fm:fm-feat-ub" "worktree=$d/wt" "kind=ship" + printf 'needs-decision: review gate\n' > "$d/state/feat-ub.status" + FM_FAKE_AXI_STATUS="$(run_parked_unbraced_findings_precursor fm/feat-ub)" + # Non-vacuity: the payload must really carry an unbraced findings block ahead + # of the braced one, with the token at the offset the walk would land on. + assert_contains "$FM_FAKE_AXI_STATUS" "findings[2]:" \ + "the fixture must really place an unbraced findings block before the gate's table" + assert_contains "$FM_FAKE_AXI_STATUS" ",ask-user," \ + "the earlier block must carry the token where the wrong-block walk would accept it" + out=$(run_crew_state "$d" feat-ub) + assert_contains "$out" "state: parked" "an unbraced findings precursor still reports parked" + assert_not_contains "$out" " · ask-user: authority decision" \ + "rows from an earlier unbraced findings block must not mint the human-decision component" + pass "the parked human-decision component is derived from the findings table's action column" +} + test_scalar_gate_parked_not_superseded() { reset_fakes local d; d=$(new_case parked-scalar-gate) @@ -4393,9 +4585,13 @@ test_captured_axi_status_shapes() { assert_contains "$out" '01NEW' "captured $shape preserves the selected identity" if [ "$shape" = parked ]; then assert_contains "$out" 'parked at test: 1 finding(s)' 'the captured gate retains its actual step and finding count' + assert_contains "$out" ' · ask-user: authority decision' \ + 'the captured gate mints the human-decision component from the real column layout' toolbin=$(make_no_python_toolbin "$d") out=$(PATH="$d/fakebin:$toolbin" FM_STATE_OVERRIDE="$d/state" "$CREW_STATE" competing) assert_contains "$out" 'parked at test: 1 finding(s)' 'a complete captured gate remains readable without Python' + assert_contains "$out" ' · ask-user: authority decision' \ + 'the captured gate mints the human-decision component without Python' assert_contains "$out" '01NEW' 'the captured gate retains its id without Python' fi pass "captured AXI $shape status replays through crew-state" @@ -4497,6 +4693,7 @@ test_single_owner_terminal_declaration_supersedes_stale_decision test_latest_status_preserves_legacy_completions test_latest_status_subshell_work_does_not_grow_with_history test_genuine_parked_not_superseded +test_parked_human_decision_comes_from_the_action_column test_scalar_gate_parked_not_superseded test_gate_block_parked_not_superseded test_ci_ready_done_log_beats_monitoring_run diff --git a/tests/fm-watch-triage.test.sh b/tests/fm-watch-triage.test.sh index 13c8648e53a..8a94ebdf22f 100755 --- a/tests/fm-watch-triage.test.sh +++ b/tests/fm-watch-triage.test.sh @@ -2595,6 +2595,7 @@ test_live_paused_until_controls_recheck_time() { wedge_threshold_round() { # local state=$1 fakebin=$2 out=$3 capture=$4 window=$5 verdict=$6 mode=$7 pid cycles=0 PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture" \ + FM_CONFIG_OVERRIDE="$(dirname "$state")/config" \ FM_FAKE_TMUX_CURRENT_COMMAND="${FM_TEST_PANE_COMMAND-grok}" \ FM_FAKE_TMUX_WINDOWS="${FM_TEST_TMUX_WINDOWS-}" FM_FAKE_CREW_STATE="$verdict" \ FM_WATCH_HANDLING_SUCCESSOR=1 \ @@ -2615,18 +2616,24 @@ wedge_threshold_round() { # backdates -# the status file so a case can put the bounded recheck cadence in or out of reach. -wedge_threshold_fixture() { # - local name=$1 line=$2 age=$3 dir state statusf window key text back +# A lane already stably stale at its recorded hash - exactly where +# wedge_timer_check owns the pane. is the WHOLE log, so a case can +# supply the multi-line history a decision fold actually reads; +# backdates the file so a case can put the bounded recheck cadence in or out of +# reach. , when given, pre-arms this key's wedge timer at that +# age: a log whose last line is captain-relevant (a `needs-decision:` escalation +# is) routes through the overridden-terminal-status branch, which reaches +# wedge_timer_check only for a hash whose timer is already running, so a case on +# that path must arm it rather than assume the plain non-terminal route. +wedge_threshold_fixture() { # [] + local name=$1 log=$2 age=$3 timer=${4-} dir state statusf window key text back dir=$(make_case "$name"); state="$dir/state" window="test:fm-wedge" statusf="$state/wedge.status" text='waiting at the gate' printf '%s' "$text" > "$dir/pane.txt" printf 'window=%s\nkind=ship\nharness=grok\nbackend=tmux\n' "$window" > "$state/wedge.meta" - printf '%s\n' "$line" > "$statusf" + printf '%s\n' "$log" > "$statusf" back=$(( $(date +%s) - age )) set_mtime "$back" "$statusf" printf '%s' "$(seen_sig "$statusf")" > "$state/.seen-wedge_status" @@ -2637,9 +2644,20 @@ wedge_threshold_fixture() { # # first sight: the suppressor holds this exact hash, so every further poll goes # straight to the wedge timer. printf '%s' "$(hash_text "$text")" > "$state/.stale-$key" + if [ -n "$timer" ]; then + printf '%s\n' "$(( $(date +%s) - timer ))" > "$state/.stale-since-$key" + fi + # An UNCONFIGURED home: the config dir exists and is empty, so every case here + # starts with the parked-gate wait evidence off and has to arm it deliberately. + mkdir -p "$dir/config" printf '%s\n' "$dir" } +# Arm the opt-in parked-gate wait evidence for a fixture built above. +arm_parked_gate() { # + : > "$1/config/wedge-defer-parked-gate" +} + wedge_stale_wakes() { # awk -F '\t' -v w="$2" '$3 == "stale" && $4 == w { n++ } END { print n + 0 }' \ "$1/.wake-queue" 2>/dev/null || echo 0 @@ -2739,7 +2757,7 @@ test_wedge_threshold_defers_to_a_declared_wait_under_a_working_verdict() { # confirm points them away from the only action that ends the wait. The sibling # absorber makes exactly this distinction, and a lane routed here must not lose it. test_wedge_threshold_recheck_names_the_captain_for_a_held_lane() { - local dir state fakebin out capture window key n + local dir state fakebin out capture window key n armed_timer local working='state: working · source: run-step · ci running' dir=$(wedge_threshold_fixture captain-held-wait \ @@ -2783,9 +2801,12 @@ test_wedge_threshold_recheck_names_the_captain_for_a_held_lane() { # at once rather than waiting out a cadence that started while the captain was # away. Same fixture and same age as the attended leg above, which is what makes # the difference attributable to the record alone. + # The idle timer is pre-armed well past the threshold, so every round below + # reaches the absorb with the same timer value and a restart would be visible. dir=$(wedge_threshold_fixture captain-held-away \ - 'captain-held: which retention window wins' 2000) + 'captain-held: which retention window wins' 2000 2000) state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out"; capture="$dir/pane.txt" + armed_timer=$(cat "$state/.stale-since-$key") write_away_record "$state" n=1 while [ "$n" -le 3 ]; do @@ -2803,8 +2824,12 @@ test_wedge_threshold_recheck_names_the_captain_for_a_held_lane() { || fail "an away-silenced hold counted $(cat "$state/.wedge-escalations-$key") wedge escalation(s)" grep -F 'never rechecked while the away-posture record exists' "$state/.watch-triage.log" >/dev/null \ || fail "the away-silenced hold was not recorded in the triage log: $(cat "$state/.watch-triage.log")" + [ "$(cat "$state/.stale-since-$key")" = "$armed_timer" ] \ + || fail "an away-silenced hold restarted the idle timer, so part of the away window would be spent against the cadence the recheck owed on return uses" - # And the recheck returns once the captain is back, so the hold is not lost. + # And the recheck is owed in full the moment the captain is back: the absorb + # above leaves the idle timer alone, so no part of the away window is spent + # against the cadence the hold is rechecked on. archive_away_record "$state" : > "$out" FM_TEST_PAUSE_RESURFACE=240 wedge_threshold_round "$state" "$fakebin" "$out" "$capture" "$window" "$working" exit \ @@ -2815,6 +2840,392 @@ test_wedge_threshold_recheck_names_the_captain_for_a_held_lane() { pass "a captain-held lane is rechecked as a hold on the captain, never as an external wait, and never at all while the captain is away" } +# --- the wedge threshold reads the crew's own parked-gate state -------------- +# Upstream kunchenguid/firstmate#3055: a lane parked at a validation gate that is +# waiting on a HUMAN is correctly quiet, but nothing in the status LINE says so - +# the evidence is the pipeline's gate state, not anything the worker wrote. One +# such lane reached 671 consecutive escalations on a single home. Neither landed +# mitigation covers it: a declared `paused:` does nothing because a live ordinary +# crewmate's absorb class never reads paused, and raising the threshold delays +# genuine wedge detection for every lane equally. +# +# The distinction that makes this safe is between the two gates the crew state +# both reports as `parked`: one owed a HUMAN, and one owed the CREWMATE's own +# answer. Only the first may go quiet - a crewmate that wedges before answering +# its own gate is exactly the failure this ladder exists to catch - so both +# directions are pinned here, and the crewmate direction is written so that a +# consumer which merely searched the verdict for the token would fail it. +# The second half of that evidence - that the human was actually asked and has +# not answered - is pinned in the test below this one. +test_wedge_threshold_defers_to_a_parked_gate_awaiting_a_human() { + local dir state fakebin out capture window key n queued + # The gate's own findings table said a human owes this answer, so + # bin/fm-crew-state.sh minted the human-decision component (its derivation from + # the `action` column by position is pinned in tests/fm-crew-state.test.sh). + local human='state: parked · source: run-step · parked at awaiting_approval: 2 finding(s) · ask-user: authority decision · run: 01RUNGATE' + # The same gate with no run component: nothing can tie a decision to it. + local runless='state: parked · source: run-step · parked at awaiting_approval: 2 finding(s) · ask-user: authority decision' + # The same shape owed the crewmate itself. The gate name is free text carried + # out of the run payload, so this one spells the whole marker inside it: a + # consumer that searched the verdict for those words instead of comparing a + # whole component for equality would read this lane as human-owed and take its + # ladder away. + local crewmate='state: parked · source: run-step · parked at fix_review (ask-user: authority decision follow-up): 2 finding(s) · run: 01RUNGATE' + + window="test:fm-wedge"; key=$(printf '%s' "$window" | tr ':/.' '___') + + # The log every case here shares: the crew escalated the gate's question and + # nobody has answered it yet, so its decision fold still holds one open + # `needs-decision`. That is the record of who was TOLD; the crew-state verdict + # above is the record of who OWES the answer, and the deferral needs both. + # The trailing `working:` note is what a crew appends next and does not close a + # decision, so it leaves the fold open while keeping the LAST line + # non-captain-relevant - the plain route into the wedge timer these cases want. + # The file is backdated well past the recheck cadence, and it is still not the + # record of when this wait began, so nothing about the recheck may be computed + # from its mtime. + local escalated='needs-decision [key=nm-01RUNGATE-review]: the gate raised an authority question +working: still parked at that gate' + # An open decision too, but under a key that names no run: an unrelated + # question raised earlier in the same task and never closed. It says nothing + # about whether anyone was told about THIS gate. + local unrelated='needs-decision [key=earlier-question]: which changelog section fits +working: still parked at that gate' + + dir=$(wedge_threshold_fixture parked-gate-human "$escalated" 2000) + arm_parked_gate "$dir" + state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out"; capture="$dir/pane.txt" + wedge_threshold_round "$state" "$fakebin" "$out" "$capture" "$window" "$human" exit \ + || fail "a gate awaiting a human was never rechecked at the threshold: $(cat "$out")" + grep -F 'verified wait at a parked gate' "$out" >/dev/null \ + || fail "the parked-gate recheck did not name its evidence: $(cat "$out")" + grep -F "awaiting firstmate's ask-user decision" "$out" >/dev/null \ + || fail "the parked-gate recheck did not name firstmate as the one the wait is on: $(cat "$out")" + grep -F "decide the gate's ask-user finding and relay the decision to the crewmate" "$out" >/dev/null \ + || fail "the parked-gate recheck did not name the action that clears the lane: $(cat "$out")" + grep -F 'awaiting the captain' "$out" >/dev/null \ + && fail "the parked-gate recheck named the captain for a decision firstmate owns: $(cat "$out")" + grep -F 'confirm the wait still holds' "$out" >/dev/null \ + && fail "a parked gate borrowed the external-wait action, which does not clear it: $(cat "$out")" + grep -F 'possible wedge' "$out" >/dev/null \ + && fail "a gate awaiting a human was reported as a possible wedge: $(cat "$out")" + # No wait age is published, because no record of when this wait began exists: + # the status file is an unrelated line, and the idle window this deferral + # resets every pass would report the same small number forever. + grep -E ', waiting [0-9]+s' "$out" >/dev/null \ + && fail "the parked-gate recheck published a wait age it has no record for: $(cat "$out")" + ack_stopped_cycle "$state" || fail "could not acknowledge the parked-gate recheck" + + # Long cadence, not a ladder: every further threshold inside the cadence is + # absorbed whole, with no escalation counted and nothing queued. + queued=$(wedge_stale_wakes "$state" "$window") + n=1 + while [ "$n" -le 3 ]; do + wedge_threshold_round "$state" "$fakebin" "$out" "$capture" "$window" "$human" absorb \ + || fail "a gate awaiting a human wedge-escalated at threshold $n: $(cat "$out")" + n=$((n + 1)) + done + [ "$(wedge_stale_wakes "$state" "$window")" -eq "$queued" ] \ + || fail "a gate awaiting a human queued a further wake inside its recheck cadence: $(cat "$state/.wake-queue")" + [ ! -e "$state/.wedge-escalations-$key" ] \ + || fail "a gate awaiting a human counted $(cat "$state/.wedge-escalations-$key") wedge escalation(s)" + + # The other direction, and the whole reason the distinction is drawn: a gate + # the crewmate itself must answer keeps the unchanged schedule, reason and + # demand-deep-inspection wording. + dir=$(wedge_threshold_fixture parked-gate-crewmate "$escalated" 2000) + arm_parked_gate "$dir" + state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out"; capture="$dir/pane.txt" + n=1 + while [ "$n" -le 3 ]; do + wedge_threshold_round "$state" "$fakebin" "$out" "$capture" "$window" "$crewmate" exit \ + || fail "a gate awaiting the crewmate stopped escalating at threshold $n: $(cat "$out")" + ack_stopped_cycle "$state" || fail "could not acknowledge crewmate-gate escalation $n" + grep -F "possible wedge, escalation $n" "$out" >/dev/null \ + || fail "a gate awaiting the crewmate did not reach escalation $n: $(cat "$out")" + n=$((n + 1)) + done + grep -F 'demand-deep-inspection: same pane has wedge-escalated 3 times in a row' "$out" >/dev/null \ + || fail "a gate awaiting the crewmate lost the demand-deep-inspection wording: $(cat "$out")" + grep -F 'verified wait at a parked gate' "$out" >/dev/null \ + && fail "a gate awaiting the crewmate was deferred as a wait on a human: $(cat "$out")" + + # The wait is owed by firstmate, not the captain, so the captain-away silence + # does not apply: under away posture the supervision branch is the actor + # allowed to answer it, and it keeps the long recheck cadence throughout. + dir=$(wedge_threshold_fixture parked-gate-away "$escalated" 2000) + arm_parked_gate "$dir" + state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out"; capture="$dir/pane.txt" + write_away_record "$state" + wedge_threshold_round "$state" "$fakebin" "$out" "$capture" "$window" "$human" exit \ + || fail "a parked gate owed firstmate's decision was silenced while the away-posture record existed: $(cat "$out")" + grep -F "awaiting firstmate's ask-user decision" "$out" >/dev/null \ + || fail "the away-posture parked-gate recheck did not name firstmate: $(cat "$out")" + grep -F 'possible wedge' "$out" >/dev/null \ + && fail "an away-posture parked gate was reported as a possible wedge: $(cat "$out")" + grep -F 'never rechecked while the away-posture record exists' "$state/.watch-triage.log" >/dev/null \ + && fail "a parked gate owed firstmate took the captain-away silence: $(cat "$state/.watch-triage.log")" + ack_stopped_cycle "$state" || fail "could not acknowledge the away-posture parked-gate recheck" + queued=$(wedge_stale_wakes "$state" "$window") + n=1 + while [ "$n" -le 3 ]; do + wedge_threshold_round "$state" "$fakebin" "$out" "$capture" "$window" "$human" absorb \ + || fail "an away-posture parked gate wedge-escalated at threshold $n: $(cat "$out")" + n=$((n + 1)) + done + [ "$(wedge_stale_wakes "$state" "$window")" -eq "$queued" ] \ + || fail "an away-posture parked gate queued a further wake inside its recheck cadence: $(cat "$state/.wake-queue")" + [ ! -e "$state/.wedge-escalations-$key" ] \ + || fail "an away-posture parked gate counted $(cat "$state/.wedge-escalations-$key") wedge escalation(s)" + + # An open decision under an unrelated key does not bind to this gate, so the + # lane keeps the unchanged ladder: nothing says anyone was told about it. + dir=$(wedge_threshold_fixture parked-gate-unrelated-key "$unrelated" 2000) + arm_parked_gate "$dir" + state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out"; capture="$dir/pane.txt" + n=1 + while [ "$n" -le 3 ]; do + wedge_threshold_round "$state" "$fakebin" "$out" "$capture" "$window" "$human" exit \ + || fail "a gate with only an unrelated open decision stopped escalating at threshold $n: $(cat "$out")" + ack_stopped_cycle "$state" || fail "could not acknowledge unrelated-key escalation $n" + grep -F "possible wedge, escalation $n" "$out" >/dev/null \ + || fail "a gate with only an unrelated open decision did not reach escalation $n: $(cat "$out")" + n=$((n + 1)) + done + grep -F 'demand-deep-inspection: same pane has wedge-escalated 3 times in a row' "$out" >/dev/null \ + || fail "a gate with only an unrelated open decision lost the demand-deep-inspection wording: $(cat "$out")" + grep -F 'verified wait at a parked gate' "$out" >/dev/null \ + && fail "an unrelated open decision was read as this gate's wait: $(cat "$out")" + + # A verdict naming no run cannot be bound to any decision, so it keeps the + # ladder even with the run-shaped key open. + dir=$(wedge_threshold_fixture parked-gate-runless "$escalated" 2000) + arm_parked_gate "$dir" + state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out"; capture="$dir/pane.txt" + wedge_threshold_round "$state" "$fakebin" "$out" "$capture" "$window" "$runless" exit \ + || fail "a runless human-owed gate never escalated: $(cat "$out")" + ack_stopped_cycle "$state" || fail "could not acknowledge the runless-gate escalation" + grep -F 'possible wedge, escalation 1' "$out" >/dev/null \ + || fail "a runless human-owed gate did not take the unchanged ladder: $(cat "$out")" + pass "a gate awaiting firstmate's decision for its own run is rechecked on the long cadence in either posture, while a crewmate-owed gate, an unrelated open decision and a runless verdict keep the unchanged ladder" +} + +# --- an unconfigured home behaves exactly as it did before this evidence ----- +# The parked-gate record is the one wait here that is not the worker's own +# declaration about its own silence: it is derived from a pipeline's gate state, +# so a home decides for itself whether a lane may give up the escalation ladder +# for it. Absent `config/wedge-defer-parked-gate` the lane this whole file +# otherwise defers - human-owed gate, open decision keyed to that run, every +# signal the armed cases assert on - must escalate on the unchanged schedule +# with the unchanged reason and demand-deep-inspection wording, and the evidence +# arm must not even be reached: no current-state read is spent and no recheck +# throttle is written. The fixture is byte-identical to the armed case above +# except for the flag, so the difference is attributable to the flag alone. +test_wedge_threshold_parked_gate_is_off_until_armed() { + local dir state fakebin out capture window key n unarmed_probes armed_probes + local human='state: parked · source: run-step · parked at awaiting_approval: 2 finding(s) · ask-user: authority decision · run: 01RUNGATE' + local escalated='needs-decision [key=nm-01RUNGATE-review]: the gate raised an authority question +working: still parked at that gate' + window="test:fm-wedge"; key=$(printf '%s' "$window" | tr ':/.' '___') + + dir=$(wedge_threshold_fixture parked-gate-unarmed "$escalated" 2000) + state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out"; capture="$dir/pane.txt" + [ ! -e "$dir/config/wedge-defer-parked-gate" ] \ + || fail "the unarmed fixture armed the flag, so it proves nothing" + export FM_FAKE_CREW_STATE_LOG="$dir/crew-state.calls" + : > "$FM_FAKE_CREW_STATE_LOG" + n=1 + while [ "$n" -le 3 ]; do + wedge_threshold_round "$state" "$fakebin" "$out" "$capture" "$window" "$human" exit \ + || fail "an unarmed home stopped escalating a parked gate at threshold $n: $(cat "$out")" + ack_stopped_cycle "$state" || fail "could not acknowledge unarmed-gate escalation $n" + grep -F "possible wedge, escalation $n" "$out" >/dev/null \ + || fail "an unarmed home did not reach escalation $n: $(cat "$out")" + n=$((n + 1)) + done + grep -F 'demand-deep-inspection: same pane has wedge-escalated 3 times in a row' "$out" >/dev/null \ + || fail "an unarmed home lost the demand-deep-inspection wording: $(cat "$out")" + grep -F 'verified wait at a parked gate' "$out" >/dev/null \ + && fail "an unarmed home deferred a parked gate: $(cat "$out")" + [ ! -e "$state/.waiting-resurfaced-$key" ] \ + || fail "an unarmed home wrote the parked-gate recheck throttle" + unarmed_probes=$(wc -l < "$FM_FAKE_CREW_STATE_LOG" | tr -d ' ') + unset FM_FAKE_CREW_STATE_LOG + + [ "$unarmed_probes" -eq 0 ] \ + || fail "an unarmed home spent $unarmed_probes current-state read(s) on a parked gate over three thresholds" + + # The same fixture with only the flag added, counted the same way, so the + # zero above is the flag's doing rather than a fixture that could never have + # reached the reader: one armed threshold must spend a read. A guard placed + # after the consult instead of before it would make both counts nonzero. + dir=$(wedge_threshold_fixture parked-gate-armed-probe-count "$escalated" 2000) + arm_parked_gate "$dir" + state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out"; capture="$dir/pane.txt" + export FM_FAKE_CREW_STATE_LOG="$dir/crew-state.calls" + : > "$FM_FAKE_CREW_STATE_LOG" + wedge_threshold_round "$state" "$fakebin" "$out" "$capture" "$window" "$human" exit \ + || fail "the armed control was never rechecked: $(cat "$out")" + ack_stopped_cycle "$state" || fail "could not acknowledge the armed control recheck" + armed_probes=$(wc -l < "$FM_FAKE_CREW_STATE_LOG" | tr -d ' ') + unset FM_FAKE_CREW_STATE_LOG + [ "$armed_probes" -gt 0 ] \ + || fail "the armed control spent no current-state read, so the probe count proves nothing" + pass "with config/wedge-defer-parked-gate absent a parked gate keeps the unchanged ladder, wording and reads" +} + +# --- a parked human-owed gate also needs the human to still owe an answer ---- +# The gate's findings table says who the answer is owed BY. It does not say the +# human was ever asked, and it does not stop saying `ask-user` once they answer: +# the run stays parked, and the row stays in the table, until the CREWMATE relays +# the decision with `axi respond`. So a lane that is quiet because the crewmate +# wedged before relaying an answer it already has would read exactly like a lane +# waiting on firstmate - and would lose the ladder for the one failure the +# ladder exists to catch. +# The task's own decision fold is the record that closes that hole, because it is +# written at ANSWER time rather than at relay time: `fm-send --resolve-key` +# appends the closing `resolved` line the moment the decision is answered. An open +# `needs-decision` therefore means the human was told and has not answered; its +# absence means the outstanding move belongs to the crewmate, or that nobody was +# ever told at all. Each of those keeps the unchanged schedule below. +test_wedge_threshold_parked_gate_needs_an_unanswered_decision() { + local dir state fakebin out capture window key n + local human='state: parked · source: run-step · parked at awaiting_approval: 2 finding(s) · ask-user: authority decision · run: 01RUNGATE' + window="test:fm-wedge"; key=$(printf '%s' "$window" | tr ':/.' '___') + + # Answered, not yet relayed. The gate verdict is byte-identical to the one the + # test above defers on; only the closing `resolved` line differs, and the + # `resolved:` verb is not captain-relevant, so this lane takes the same plain + # non-terminal route into the wedge timer as that one. + dir=$(wedge_threshold_fixture parked-gate-decided \ + 'needs-decision [key=nm-01RUNGATE-review]: the gate raised an authority question +resolved [key=nm-01RUNGATE-review]: firstmate chose the second fix' 2000) + arm_parked_gate "$dir" + state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out"; capture="$dir/pane.txt" + n=1 + while [ "$n" -le 3 ]; do + wedge_threshold_round "$state" "$fakebin" "$out" "$capture" "$window" "$human" exit \ + || fail "a decided-but-unrelayed gate stopped escalating at threshold $n: $(cat "$out")" + ack_stopped_cycle "$state" || fail "could not acknowledge decided-gate escalation $n" + grep -F "possible wedge, escalation $n" "$out" >/dev/null \ + || fail "a decided-but-unrelayed gate did not reach escalation $n: $(cat "$out")" + n=$((n + 1)) + done + grep -F 'demand-deep-inspection: same pane has wedge-escalated 3 times in a row' "$out" >/dev/null \ + || fail "a decided-but-unrelayed gate lost the demand-deep-inspection wording: $(cat "$out")" + grep -F 'verified wait at a parked gate' "$out" >/dev/null \ + && fail "a gate whose decision was already answered was deferred as a wait on the captain: $(cat "$out")" + + # Parked at a human-owed gate, quiet, and the crewmate never escalated it: no + # human has been told, so there is no wait to defer to. + dir=$(wedge_threshold_fixture parked-gate-unescalated 'working: validation under way' 2000) + arm_parked_gate "$dir" + state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out"; capture="$dir/pane.txt" + wedge_threshold_round "$state" "$fakebin" "$out" "$capture" "$window" "$human" exit \ + || fail "a human-owed gate nobody was told about never escalated: $(cat "$out")" + ack_stopped_cycle "$state" || fail "could not acknowledge the unescalated-gate escalation" + grep -F 'possible wedge, escalation 1' "$out" >/dev/null \ + || fail "a human-owed gate nobody was told about did not take the unchanged ladder: $(cat "$out")" + + # An open `blocked` record is not an unanswered question: it is an obstacle the + # crew reported, and a different action clears it. A `blocked:` last line is + # captain-relevant, so this lane reaches the wedge timer through the + # overridden-terminal-status branch instead, which only ever sees a hash whose + # timer is already running - hence the fixture's fourth argument. + dir=$(wedge_threshold_fixture parked-gate-blocked \ + 'blocked [key=nm-01RUNGATE-review]: the fixture cannot reach its dependency' 2000 600) + arm_parked_gate "$dir" + state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out"; capture="$dir/pane.txt" + wedge_threshold_round "$state" "$fakebin" "$out" "$capture" "$window" "$human" exit \ + || fail "a human-owed gate with only a blocker open never escalated: $(cat "$out")" + ack_stopped_cycle "$state" || fail "could not acknowledge the blocked-gate escalation" + grep -F 'possible wedge, escalation 1' "$out" >/dev/null \ + || fail "an open blocker was accepted as an unanswered gate decision: $(cat "$out")" + pass "a parked human-owed gate is deferred only while its decision is still open, so an answered-but-unrelayed gate, an unescalated one, and one holding only a blocker all keep the unchanged ladder" +} + +# --- a wait record that does not carry every field is refused ---------------- +# wait_record joins its five fields with US and wedge_defer_wait parses them with +# `IFS= read`, so consecutive delimiters yield genuinely EMPTY fields and no +# field can shift left into another's position. That is what makes the deferral's +# guard able to enforce the whole contract rather than a position-specific slice +# of it: each field the recheck prints must be present, and a record carrying +# more than its four delimiters is refused too, since `read` puts any surplus +# into the final variable. Deferring on a record that is not what it claims is +# what takes the ladder away, so every one of these must fall back to the +# escalation the caller was about to make instead. +# No shipped evidence producer can emit a malformed record, which is precisely +# the invariant under test, so this loads the real bin/fm-watch.sh through its +# own source guard in a child shell (the entry tests/fm-supervision-events.test.sh +# uses) and drives the real wedge_timer_check. The assertion is on the durable +# wake queue the watcher actually wrote. + +# One wedge_timer_check round against a malformed record. is the +# body of a wedge_wait_evidence override, so a case supplies exactly the record +# under test. Publishes the state directory it ran in as MALFORMED_STATE rather +# than on stdout, because fail() exits the shell it runs in and a command +# substitution would swallow a setup failure here. +run_malformed_wait_record_round() { # + local name=$1 body=$2 dir state out + dir=$(make_case "$name"); state="$dir/state" + printf 'working: validation under way\n' > "$state/wedge.status" + printf '%s\n' "$(( $(date +%s) - 600 ))" > "$state/.stale-since-test_fm-wedge" + + out="$dir/defer.out" + FM_STATE_OVERRIDE="$state" FM_STALE_ESCALATE_SECS=1 FM_PAUSE_RESURFACE_SECS=999 \ + FM_WEDGE_DEMAND_INSPECT_COUNT=3 \ + bash -c ' + # shellcheck disable=SC1090,SC1091 + . "$1" + wake() { :; } + # A live agent, so the dead-record probe that runs after a refused + # deferral keeps the unchanged ladder rather than reading a backend this + # child shell has none of. + fm_backend_agent_state() { printf alive; } + eval "wedge_wait_evidence() { $2 ; }" + wedge_timer_check "test:fm-wedge" "$FM_STATE_OVERRIDE/.stale-since-test_fm-wedge" \ + "non-terminal stale" "$FM_STATE_OVERRIDE/.wedge-escalations-test_fm-wedge" wedge \ + malformed-record-pane + ' _ "$WATCH" "$body" > "$out" 2>&1 \ + || fail "the wedge timer failed on a malformed wait record ($name): $(cat "$out")" + MALFORMED_STATE=$state +} + +assert_malformed_record_kept_the_ladder() { # + local state=$1 what=$2 + grep -F 'possible wedge, escalation 1' "$state/.wake-queue" >/dev/null \ + || fail "$what did not keep the unchanged ladder: $(cat "$state/.wake-queue" 2>/dev/null)" + grep -F 'rechecked on a long cadence not a wedge' "$state/.wake-queue" >/dev/null \ + && fail "$what was deferred on a record that is not what it claims: $(cat "$state/.wake-queue")" + [ "$(cat "$state/.wedge-escalations-test_fm-wedge" 2>/dev/null || echo 0)" -eq 1 ] \ + || fail "$what did not count its escalation" +} + +test_wedge_defer_refuses_a_half_filled_wait_record() { + # An empty subject - the field whose loss used to shift the prose action into + # `whom` and print an action that clears nothing. + run_malformed_wait_record_round malformed-wait-record \ + 'wait_record "declared wait" "" external "confirm the wait still holds" ""' + assert_malformed_record_kept_the_ladder "$MALFORMED_STATE" "a wait record with no subject" + + # An empty ACTION with a non-empty anchor. Under the old TAB join this parsed + # as a valid record: the doubled tab collapsed, the anchor path slid into + # `action`, and the recheck published a status-file path as the one thing that + # clears the lane while silently losing the wait-age anchor. + run_malformed_wait_record_round malformed-wait-record-no-action \ + "wait_record 'declared wait' 'awaiting external' external '' '$TMP_ROOT/anchor.status'" + assert_malformed_record_kept_the_ladder "$MALFORMED_STATE" "a wait record with no action" + + # A record carrying a surplus delimiter: `read` puts everything past the last + # field into `anchor`, so the fields after the extra one are not the fields + # they are read as. + run_malformed_wait_record_round malformed-wait-record-surplus \ + 'printf "%s\\037%s\\037%s\\037%s\\037%s\\037%s" "declared wait" "awaiting external" external "confirm the wait still holds" "" extra' + assert_malformed_record_kept_the_ladder "$MALFORMED_STATE" "a wait record with a surplus field" + + pass "a wait record missing a field the recheck must print, or carrying one it must not, is refused and the lane escalates exactly as it would have" +} + # --- a record whose agent is GONE reports once, instead of alarming forever --- # Observed on a live fleet: two finished lanes reached 226 and 203 CONSECUTIVE @@ -5616,6 +6027,10 @@ test_live_declared_wait_churn_honors_the_resurface_throttle test_live_paused_until_controls_recheck_time test_wedge_threshold_defers_to_a_declared_wait_under_a_working_verdict test_wedge_threshold_recheck_names_the_captain_for_a_held_lane +test_wedge_threshold_defers_to_a_parked_gate_awaiting_a_human +test_wedge_threshold_parked_gate_needs_an_unanswered_decision +test_wedge_threshold_parked_gate_is_off_until_armed +test_wedge_defer_refuses_a_half_filled_wait_record test_open_captain_call_bounds_stale_churn test_stale_churn_without_a_captain_call_still_alarms test_failed_wake_append_does_not_arm_the_captain_hold_throttle diff --git a/tests/wake-helpers.sh b/tests/wake-helpers.sh index 8e7106d8230..f8c7b05e2e1 100644 --- a/tests/wake-helpers.sh +++ b/tests/wake-helpers.sh @@ -113,12 +113,16 @@ SH # A per-id override FM_FAKE_CREW_STATE_ wins; otherwise the shared # FM_FAKE_CREW_STATE; otherwise an unknown verdict (NOT provably working), the # safe default so a test that forgets to set one surfaces rather than absorbs. +# Exporting FM_FAKE_CREW_STATE_LOG appends one line per call, so a test that +# asserts how many current-state reads a path spends - the reads are the costly +# half of watcher triage - can count them instead of inferring them. make_fake_crew_state() { # local fakebin=$1 cat > "$fakebin/fm-crew-state.sh" <<'SH' #!/usr/bin/env bash set -u id=${1:-} +[ -z "${FM_FAKE_CREW_STATE_LOG:-}" ] || printf '%s\n' "$id" >> "$FM_FAKE_CREW_STATE_LOG" key=$(printf '%s' "$id" | tr -c 'A-Za-z0-9' '_') var="FM_FAKE_CREW_STATE_$key" val=${!var:-${FM_FAKE_CREW_STATE:-}} From 1b1b3cd4c3130303803cf1c93a5f7127e0245bb7 Mon Sep 17 00:00:00 2001 From: Jon Roosevelt Date: Sun, 20 Sep 2026 02:19:16 -0400 Subject: [PATCH 002/168] fix(bin): reclaim a task whose herdr endpoint was destroyed (#5007) * fix(control): let the owning seat reclaim a task whose endpoint is gone A destroyed pane or workspace made `missing` a terminal state. Relaunch accepted only `dead` and said to stop the agent first; exit refused `missing` and said to reconcile the task first; there is no reconcile verb. Each command named the other as its prerequisite, so a task whose terminal went away could not be reclaimed by anything, and a no-mistakes approval it was parked on had no seat left to answer it. `missing` is agent-free a fortiori: there is no endpoint, so there is no agent in it. Widen the existing guards rather than add a verb. - fm-spawn --relaunch accepts a positively proven `missing` and creates one fresh endpoint in the recorded worktree; the record it already republishes rebinds the task to it. A `dead` endpoint is still adopted in place. - fm-control exit reports `endpoint-gone` instead of dying, so the relaunch transaction's stop step no longer dead-ends, and re-resolves the endpoint from the record before verifying the replacement. The duplicate-agent refusal is untouched: both verdicts come from the same recovery-grade classifier, which claims `missing` only from positive absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The backends' own create paths refuse a live same-labeled endpoint as a second independent guard. The worktree, its branch, commits, uncommitted changes, armed poll and registration, record rows, and status log are all untouched - a reclaim is a recovery, never a teardown. A secondmate is excluded: its gone-endpoint recovery already has one owner in the session-start liveness sweep, so relaunch refuses and names it rather than becoming a second path to the same outcome. Tests reproduce both halves of the deadlock, the reclaim succeeding, unlanded work surviving it, and the refusals that still hold. * no-mistakes(review): prove endpoint absence per backend before reclaim rebinds * no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session * no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly * no-mistakes(review): stop refusals and docs asserting unestablished causes * no-mistakes(review): stop herdr fixture helper losing tmp-root registration * no-mistakes(review): document workspace drift and absence-probe server residue * no-mistakes(review): correct rebind limitation to its one reachable case * no-mistakes(review): stop claiming reclaim leaves instructions untouched * no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage * no-mistakes(rebase): read the staged launch file in the herdr fixture Rebasing onto main picked up #4994, which stages a long worker launch command into a script and delivers the short `. ''` line instead of the literal command. The tmux fake and tests/fixtures.sh were updated for that; the herdr fake this branch adds was written before it and still keyed "an agent now exists on this pane" off the literal `encode launch-brief` text, so after the rebase it never marked the rebound pane live and the reclaim's alive-wait read `dead`. Dereference the staged file first, exactly as the tmux fake above does. Test-fixture only; no production path changes. Co-Authored-By: Claude Opus 5 (1M context) * no-mistakes(document): note reclaim placement in herdr and scripts inventories --------- Co-authored-by: Claude Opus 5 (1M context) --- .../skills/stuck-crewmate-recovery/SKILL.md | 5 + bin/backends/herdr.sh | 41 +- bin/fm-control-lib.sh | 74 ++- bin/fm-control.sh | 98 +++- bin/fm-spawn.sh | 180 +++++- docs/agent-control.md | 68 ++- docs/herdr-backend.md | 1 + docs/scripts.md | 2 +- tests/fm-control-relaunch.test.sh | 536 +++++++++++++++++- tests/fm-control.test.sh | 18 +- 10 files changed, 978 insertions(+), 45 deletions(-) diff --git a/.agents/skills/stuck-crewmate-recovery/SKILL.md b/.agents/skills/stuck-crewmate-recovery/SKILL.md index c5209051a44..9004c3872bd 100644 --- a/.agents/skills/stuck-crewmate-recovery/SKILL.md +++ b/.agents/skills/stuck-crewmate-recovery/SKILL.md @@ -37,6 +37,11 @@ Do not sweep another home's endpoints or infer ownership from a matching window Before relaunch, prove that no live agent still owns the recorded task and that the existing worktree remains available. Preserve its uncommitted changes and commits, keep the same task identity, and resume or relaunch the recorded harness in that existing worktree with the same brief plus a concise progress note. +A HERDR endpoint that is not merely idle but destroyed - a pane or workspace removed in Herdr churn - is recovered by that same relaunch, which creates one fresh endpoint in the existing worktree and rebinds the task's record to it; nothing special is needed, and the worktree is untouched ([`docs/agent-control.md`](../../../docs/agent-control.md) "Reclaiming a task whose endpoint is gone"). +That relaunch proves the endpoint is destroyed before it rebinds, so a Herdr server that was merely stopped is adopted back rather than duplicated. +On tmux there is no reclaim: a task record carries no socket identity for its endpoint, so a `missing` window cannot be told apart from one on a tmux server this seat cannot address, and both `exit` and `relaunch` refuse. +Do not work around either refusal by respawning - it means a live agent may still hold that worktree. +That reclaim is the owning home's operation only, and a secondmate is the one exception: recover it through `bin/fm-spawn.sh --secondmate` as above. Do not use a fresh generic spawn while the recorded worktree is unaccounted for, because allocating another worktree can split one task across two copies. If the worktree or ownership cannot be reconciled safely, leave all state intact and report the task failed or blocked with the conflicting evidence. diff --git a/bin/backends/herdr.sh b/bin/backends/herdr.sh index dee97f7866c..b836b77201e 100644 --- a/bin/backends/herdr.sh +++ b/bin/backends/herdr.sh @@ -2011,10 +2011,19 @@ fm_backend_herdr_workspace_ensure() { # [ is passed straight through to # fm_backend_herdr_workspace_ensure, which owns its meaning. -fm_backend_herdr_container_ensure() { # [] - local cwd=${1:-$PWD} relationship=${2:-launcher-home} session label status +# +# is optional and DEFAULTS to fm_backend_herdr_session, so every +# ordinary spawn keeps resolving the ambient session exactly as before. A +# RECOVERY passes the session its record already names, because a task must not +# be relocated onto whatever server the recovering seat happens to sit on. It is +# threaded as a parameter rather than by shadowing HERDR_SESSION on purpose: +# fm_backend_herdr_launcher_identity compares the launcher's own ambient session +# against this one, and shadowing would make that half of its cross-session +# guard compare the pinned value with itself and pass vacuously. +fm_backend_herdr_container_ensure() { # [] [] + local cwd=${1:-$PWD} relationship=${2:-launcher-home} session=${3:-} label status fm_backend_herdr_version_check || return 1 - session=$(fm_backend_herdr_session) + [ -n "$session" ] || session=$(fm_backend_herdr_session) fm_backend_herdr_server_ensure "$session" || return 1 fm_backend_herdr_workspace_ensure "$session" "$cwd" "$relationship" >/dev/null && status=0 || status=$? # A 3 already reported the exact placement it refused to guess at; adding the @@ -2353,6 +2362,32 @@ fm_backend_herdr_agent_state() { # esac } +# fm_backend_herdr_endpoint_absence_recheck: re-read with its own +# session's server running, and print the resulting fm_backend_agent_state +# verdict. For a recovery that is about to RE-CREATE an endpoint, this is the +# read that decides whether there is anything to re-create at all. +# +# fm_backend_herdr_agent_state maps a positively STOPPED session server to +# `missing` (issue #4091), which is correct for "no agent is running" but is +# NOT evidence the endpoint was destroyed: stopping and restarting a named +# Herdr server preserves workspace, tab, pane, and label ids (docs/herdr-backend.md +# "Restart and liveness behavior") - only the harness processes and their +# registrations die. So `missing` there means unreachable right now, and a +# caller that rebound on it would abandon a pane that was about to come back. +# +# Only the RECORDED session's server is ensured, never a workspace or tab, so +# this creates nothing: a merely-stopped server comes back and the recorded +# pane classifies `dead` (adoptable), a genuinely destroyed pane still reads +# `missing`, a returning agent reads `alive`, and a server that will not start +# is `unreadable` - unreachable, which refuses, rather than absence. +fm_backend_herdr_endpoint_absence_recheck() { # + local target=$1 + fm_backend_herdr_parse_target "$target" || { printf 'unreadable'; return 0; } + fm_backend_herdr_server_ensure "$FM_BACKEND_HERDR_SESSION" >/dev/null 2>&1 \ + || { printf 'unreadable'; return 0; } + fm_backend_herdr_agent_state "$target" +} + # Backward-compatible three-state view for callers that only need a yes/no # agent verdict. The detailed state contract is owned by fm_backend_agent_state. fm_backend_herdr_agent_alive() { # diff --git a/bin/fm-control-lib.sh b/bin/fm-control-lib.sh index 7bb4d580ec6..6e6be0d5c3a 100644 --- a/bin/fm-control-lib.sh +++ b/bin/fm-control-lib.sh @@ -12,9 +12,12 @@ # verbs addressed to an exact task id, with the per-harness mechanics owned # here rather than improvised per harness in agent prose. # -# This file owns three capability tables plus their pure artifact-path tables -# and nothing else. It has no side effects, runs no backend command, and reads -# no state, so it can be sourced by a test as a pure contract: +# This file owns three capability tables plus their pure artifact-path tables, +# and ONE named exception to that purity - fm_control_endpoint_absence_verdict, +# the single owner of the per-backend endpoint-absence proof, which does run +# backend reads. Everything else has no side effects, runs no backend command, +# and reads no state, so sourcing this file is still free and the tables can be +# read by a test as a pure contract: # # 1. Verb allowlist. There is no arbitrary-text and no generic raw-key entry # point on the control plane; a caller either names an allowlisted verb or @@ -218,6 +221,71 @@ fm_control_backend_state_verified() { # return 1 } +# fm_control_endpoint_absence_verdict: the ONE owner of the per-backend proof +# that an endpoint reading `missing` is actually GONE rather than merely +# unreachable from this seat. Call it only for a `missing` raw state. +# +# Prints "\t" - always exactly one TAB, so a caller splits +# unambiguously with ${raw%%$'\t'*} and ${raw#*$'\t'}. The reason is empty +# except on `unproven`, where it is the concrete sentence the caller's refusal +# message embeds. It is returned on stdout rather than set in a variable +# because every caller reads this through a command substitution, where an +# assignment made here could never reach them. +# +# The verdicts: +# gone - absence is PROVEN. There is no endpoint and therefore no agent. +# dead - the endpoint is there after all and holds no agent. +# alive - the endpoint is there and an agent is running in it. +# unproven - neither could be established; the caller must refuse. +# +# fm_backend_agent_state's `missing` conflates "the endpoint was DESTROYED" +# with "the endpoint is UNREACHABLE from here right now". An unreachable +# endpoint can still hold a live agent on the task's worktree, so every caller +# that would act on absence - `exit` claiming the agent stopped, `relaunch` +# re-creating the endpoint - must come through here rather than trusting the +# raw verdict. +# +# Whether absence is provable AT ALL is a property of the backend, not of the +# reading: +# herdr CAN prove it. Every read goes through fm_backend_herdr_cli, which +# passes `--session `, so the recheck starts and reads the session +# the RECORD names, through that session's own socket. The answer is about +# the task's endpoint and nothing else. +# tmux CANNOT. `list-windows -a` describes only the server the CURRENT +# process addresses (its TMUX_TMPDIR/socket), and a task's record does not +# carry the endpoint's socket identity - so a different but running server +# would answer "not anywhere" about a window it was never able to see. +# There is no read available here that closes that gap, so tmux always +# returns `unproven` and both verbs refuse. tmux is left exactly as +# deadlocked as it was before this change - no worse - but deliberately. +# +# Both control-plane callers share this one implementation so the proof cannot +# drift into two answers for the same endpoint. +fm_control_endpoint_absence_verdict() { # + local backend=${1-} target=${2-} + fm_backend_source "$backend" \ + || { printf 'unproven\tbackend %s could not be loaded to prove anything about that endpoint' "'$backend'"; return 0; } + case "$backend" in + tmux) + printf 'unproven\ttmux absence cannot be proven from a task record: the record does not carry the endpoint'"'"'s socket identity, and a server-wide window inventory only describes the tmux server this process addresses, so a window absent from it may still be alive on another' + ;; + herdr) + # Start the RECORDED session's server (only the server - nothing is + # created) and re-read the recorded pane. A pane that comes back with the + # server was never destroyed. + case "$(fm_backend_herdr_endpoint_absence_recheck "$target")" in + dead) printf 'dead\t' ;; + alive) printf 'alive\t' ;; + missing) printf 'gone\t' ;; + *) printf 'unproven\tthe recorded herdr session'"'"'s server could not be started, or its pane could not be classified once it was running' ;; + esac + ;; + *) + printf 'unproven\tbackend %s has no recovery-grade classifier, so absence cannot be proven on it at all' "'$backend'" + ;; + esac +} + # The per-task wiring artifacts a harness leaves behind, so a relaunch that # changes harness (or re-arms the same one with a fresh busy generation) can # clear the previous incarnation's wiring instead of leaving a stale hook diff --git a/bin/fm-control.sh b/bin/fm-control.sh index 4e1852c358d..1b73aa644ac 100755 --- a/bin/fm-control.sh +++ b/bin/fm-control.sh @@ -30,11 +30,35 @@ # every uncommitted change. Interrupts first when the task reads # busy, then submits the harness's exit command. Postcondition: # the backend's recovery-grade classifier reports the agent gone. -# Already-stopped is success (idempotent). +# Already-stopped is success (idempotent). An endpoint that reads +# `missing` is put through the control plane's per-backend absence +# proof (fm_control_endpoint_absence_verdict) before anything is +# claimed about it, because `missing` also covers an endpoint that +# is merely unreachable from this seat. That proof exists only on +# HERDR, whose reads are scoped to the session the record names: +# proven gone reports `endpoint-gone` rather than +# `already-stopped`, because the endpoint this verb normally +# preserves did not survive; a pane that turns out to be there and +# idle is the ordinary `already-stopped`; one whose agent is back +# takes the ordinary interrupt-then-exit path. A tmux `missing` +# always REFUSES: a task record carries no socket identity for its +# endpoint, so this verb cannot tell a destroyed window from one on +# a tmux server it cannot address, and it will not claim a stop it +# cannot see. # relaunch Transactionally replace the running agent with a new one, in the -# SAME endpoint and SAME worktree, on the same or a newly chosen +# SAME worktree - and the same endpoint whenever that endpoint +# still exists - on the same or a newly chosen # harness/model/effort - so switching harness is one ordinary use -# of this verb. An explicit `default` model or effort clears that +# of this verb. When the recorded endpoint is instead proven gone - +# a Herdr pane or workspace destroyed in churn - the launch owner +# re-creates one in that worktree, in the herdr session the record +# names, and the task's record rebinds to it; that is how a task +# whose terminal was destroyed is reclaimed by the home that owns +# it, rather than being stranded with a parked approval nobody can +# answer. Reclaim is HERDR-ONLY for the reason `exit` gives above: +# a tmux `missing` cannot be proven absent from a task record, so +# it refuses. +# An explicit `default` model or effort clears that # axis for the replacement. With no explicit axis, a secondmate # re-resolves its durable config/secondmate-harness pin (harness # plus its optional model and effort tokens) exactly as any other @@ -447,9 +471,9 @@ retire_busy_incarnation() { } # do_exit: stop the running agent, preserving endpoint and worktree. Prints -# `already-stopped` or `stopped`. +# `already-stopped`, `endpoint-gone`, or `stopped`. do_exit() { - local state cmd verdict composer_state cancel interrupt_result=not-needed + local state cmd verdict composer_state cancel absence interrupt_result=not-needed require_state_verified_backend exit state=$(agent_state) case "$state" in @@ -458,7 +482,40 @@ do_exit() { return 0 ;; alive) ;; - missing) die "task $ID's recorded endpoint is gone, so there is no agent to stop; reconcile the task before any further control action" ;; + missing) + # `missing` on its own is not a finding about the endpoint: it conflates + # "destroyed" with "unreachable from this seat". Route it through the + # control plane's one absence proof - the same one the relaunch gate uses + # - and report what that proof actually established, never more. + absence=$(fm_control_endpoint_absence_verdict "$BACKEND" "$T") + case "${absence%%$'\t'*}" in + gone) + # Proven gone, so the agent that lived in it went with it: exit's + # postcondition already holds and there is nothing to send. Its own + # outcome rather than `already-stopped`, because the endpoint this + # verb normally preserves did not survive. The worktree and every + # uncommitted change are untouched, and `relaunch` re-creates the + # endpoint from here. + printf 'endpoint-gone' + return 0 + ;; + dead) + # The endpoint was only unreachable and is there after all, holding + # no agent - a herdr pane whose session server was merely stopped is + # the common case. Nothing is gone, so this is the ordinary + # already-stopped outcome. + printf 'already-stopped' + return 0 + ;; + alive) + # The agent came back with its endpoint. Fall through to the ordinary + # alive path: interrupt if busy, then the harness's exit command. + ;; + *) + die "task $ID's endpoint $T reads 'missing', but ${absence#*$'\t'}; exit will not claim an agent stopped at an address it cannot trust, nor send lifecycle input to one" + ;; + esac + ;; *) die "task $ID's endpoint reads '$state' rather than a positively classified state; refusing to send a lifecycle command into an unattributed endpoint" ;; esac # A busy agent is interrupted first before the exit command is submitted. @@ -596,8 +653,16 @@ relaunch_rollback() { echo "error: $ID's agent stopped but relaunch did not reach replacement launch; no agent is running, and its work plus progress note are preserved at $WT" >&2 ;; *) - journal_write "failed:$RELAUNCH_PHASE" "rollback=none-agent-state-$state" || true - echo "error: relaunch of $ID failed while stopping the old agent and its state is '$state'; the durable record and progress note were retained for recovery" >&2 + # The old agent was NOT proven stopped, so no replacement is coming + # and the agent that may still be reading these instructions is the + # original one. The note exists to brief a replacement; leaving it in + # a possibly-live agent's brief would be an unrequested edit to a + # running task. Restore byte-exact, exactly as the alive case does. + if [ -n "$RELAUNCH_BRIEF" ] && [ -f "$BRIEF_PRIOR" ]; then + cp -p "$BRIEF_PRIOR" "$RELAUNCH_BRIEF" 2>/dev/null || true + fi + journal_write "failed:$RELAUNCH_PHASE" "rollback=instructions-restored-agent-state-$state" || true + echo "error: relaunch of $ID failed while stopping the old agent and its state is '$state', so it was not proven stopped; its original instructions were restored and the durable record was retained for recovery" >&2 ;; esac ;; @@ -851,6 +916,23 @@ do_relaunch() { if FM_CONTROL_RELAUNCH_TX="$RELAUNCH_TX" \ "$SCRIPT_DIR/fm-spawn.sh" "${spawn_args[@]}" >/dev/null; then RELAUNCH_META_PUBLISHED=1 + # $T was resolved from the record before the launch. When the recorded + # endpoint was gone, the launch owner created a fresh one and republished + # the record pointing at it, so every postcondition below must be read from + # the endpoint the task now HAS, not the one it had. Re-resolving through + # the same shared validation is what makes that safe: a record that no + # longer passes it refuses here rather than leaving this transaction + # polling an address nothing owns. + # stdout is dropped (it is only the resolved target), but the refusal on + # stderr names the exact row that failed - and in this one branch the record + # was just rewritten by the launch owner, so that row is the whole + # diagnostic. Let it through rather than dying with nothing to act on. + if fm_backend_validate_task_endpoint "$META" "$ID" >/dev/null \ + && [ -n "$FM_BACKEND_VALIDATED_TARGET" ]; then + T=$FM_BACKEND_VALIDATED_TARGET + else + die "the replacement agent for $ID was launched, but task $ID's republished record no longer passes endpoint validation (the refusal above names the row), so this transaction cannot say which endpoint to confirm it on; reconcile $META before any further control action" + fi else [ "$(fm_meta_get "$META" control_relaunch_tx)" != "$RELAUNCH_TX" ] \ || RELAUNCH_META_PUBLISHED=1 diff --git a/bin/fm-spawn.sh b/bin/fm-spawn.sh index 5c77dc23cb6..8bc3b25b0fd 100755 --- a/bin/fm-spawn.sh +++ b/bin/fm-spawn.sh @@ -30,7 +30,8 @@ # secondmate's charter. # fm-spawn.sh --relaunch [--harness ] [--model ] [--effort ] # --relaunch launches a replacement agent for an EXISTING task into that -# task's own recorded endpoint and worktree instead of creating either. It is +# task's own recorded worktree, reusing its recorded endpoint when that +# endpoint still exists, instead of creating either from scratch. It is # the launch half of the control plane (bin/fm-control.sh relaunch), which # owns the checkpoint, the progress note, stopping the previous agent, and the # transaction; call fm-control rather than this flag directly unless you are @@ -42,7 +43,20 @@ # ordinary relaunch. It refuses unless the recorded endpoint is positively # agent-free on a backend with a recovery-grade agent-state classifier (tmux # or herdr), and clears the previous harness's per-task wiring before arming -# the new incarnation. The replacement still never starts outside the copy +# the new incarnation. Two verdicts are agent-free: a `dead` endpoint is +# ADOPTED as-is, while an endpoint PROVEN gone is RE-CREATED in the recorded +# worktree and the republished record rebinds the task to it. That proof is +# its own step, because a backend's `missing` also covers an endpoint that is +# merely unreachable from here - and it is only available on HERDR, which must +# still read the recorded pane as gone once that session's server is running +# again. A tmux `missing` always refuses: a task record carries no socket +# identity for its endpoint, so no read here can tell a destroyed window from +# one on a tmux server this process cannot address. An endpoint that turns out +# to have survived refuses too. The worktree is reused untouched either way; a +# rebind is a recovery, never a teardown. Only a crewmate or scout rebinds: a +# secondmate whose endpoint is gone is respawned by its own owner +# (`--secondmate`, driven by the session-start liveness sweep). +# The replacement still never starts outside the copy # holding the work: a Herdr shell that has drifted out of the recorded # worktree is told once to return, and only a shell that will not go refuses. # --harness is the explicit per-spawn harness/profile adapter. The old @@ -1521,6 +1535,9 @@ RAW_LAUNCH=0 # validation teardown uses, so a malformed, ambiguous, or foreign record # refuses here exactly as it refuses there. RELAUNCH_PRIOR_HARNESS= +# 1 when the recorded endpoint is authoritatively gone and this relaunch must +# create a fresh one for the task rather than adopt its recorded address. +RELAUNCH_REBIND=0 if [ "$RELAUNCH" -eq 1 ]; then [ "${#POS[@]}" -eq 1 ] || { echo "error: --relaunch takes the task id only; its project or home comes from the task's own record" >&2 @@ -1554,14 +1571,69 @@ if [ "$RELAUNCH" -eq 1 ]; then echo "error: backend '$BACKEND' has no recovery-grade agent-state classifier, so a relaunch cannot prove the previous agent exited; refusing rather than risking two agents in one endpoint" >&2 exit 1 } + # Two states are agent-free, and both license a relaunch: + # dead - the endpoint exists and confidently holds no agent. The + # endpoint is ADOPTED, so the task keeps its exact address. + # missing - the endpoint itself is gone. There is no endpoint AND therefore + # no agent, so a relaunch cannot adopt it: it CREATES a fresh + # endpoint in the recorded worktree and the published record + # rebinds to it. + # `missing` is NOT one state, and that is what the duplicate-agent argument + # turns on. fm_backend_agent_state's per-backend `missing` conflates "the + # endpoint was DESTROYED" with "the endpoint is UNREACHABLE from here right + # now", and an unreachable endpoint can still hold the live agent this + # relaunch would duplicate. So absence is PROVEN before it may rebind, never + # inferred from a failed read - and only HERDR can prove it: + # herdr - the recorded session's server is started, and the recorded pane is + # RE-READ through that session's own socket. `dead` means the pane + # survived the restart and is adopted after all; `alive` means the + # agent came back and refuses; only a second `missing` proves the + # pane itself did not survive. + # tmux - REFUSES, always. A task record carries no socket identity for its + # endpoint, and a server-wide inventory describes only the server + # this process addresses, so no read available here can tell "gone" + # from "on a server I cannot see". A tmux `missing` therefore stays + # as deadlocked as it was before this change - deliberately, and + # with the reason stated rather than guessed past. + # Every transient or self-contradicting read stays `unreadable`/`ambiguous` + # and refuses as it always did (bin/fm-backend.sh's fm_backend_agent_state + # owns that vocabulary). The proof itself lives in one place for the whole + # control plane - fm_control_endpoint_absence_verdict - so `exit` and + # `relaunch` cannot reach two different answers about one endpoint. RELAUNCH_STATE=$(fm_backend_agent_state "$BACKEND" "$RELAUNCH_TARGET") - [ "$RELAUNCH_STATE" = dead ] || { - echo "error: task $ID's endpoint reads '$RELAUNCH_STATE'; a relaunch requires a positively agent-free endpoint (stop the agent first with bin/fm-control.sh $ID exit)" >&2 - exit 1 - } + if [ "$RELAUNCH_STATE" = missing ]; then + RELAUNCH_ABSENCE=$(fm_control_endpoint_absence_verdict "$BACKEND" "$RELAUNCH_TARGET") + case "${RELAUNCH_ABSENCE%%$'\t'*}" in + gone) RELAUNCH_STATE=missing ;; + dead) RELAUNCH_STATE=dead ;; + alive) RELAUNCH_STATE=alive ;; + *) + echo "error: task $ID's recorded endpoint $RELAUNCH_TARGET reads 'missing', but ${RELAUNCH_ABSENCE#*$'\t'}. An endpoint that cannot be proven absent may still hold a live agent on this task's worktree; refusing rather than launching a second agent into it" >&2 + exit 1 + ;; + esac + fi + case "$RELAUNCH_STATE" in + dead) ;; + missing) RELAUNCH_REBIND=1 ;; + *) + echo "error: task $ID's endpoint reads '$RELAUNCH_STATE'; a relaunch requires a positively agent-free endpoint (stop the agent first with bin/fm-control.sh $ID exit)" >&2 + exit 1 + ;; + esac RELAUNCH_PRIOR_HARNESS=$(fm_meta_get "$RELAUNCH_META" harness) KIND=$(fm_meta_get "$RELAUNCH_META" kind) [ -n "$KIND" ] || KIND=ship + # A secondmate whose endpoint is gone already has ONE owner for that + # recovery: the session-start liveness sweep respawns it with + # `fm-spawn.sh --secondmate`, which stands its home's own workspace back + # up (bin/fm-bootstrap.sh; the secondmate-provisioning skill). Rebinding one + # here as well would be a second path to the same outcome, so this refuses + # and names the one that owns it. + if [ "$RELAUNCH_REBIND" -eq 1 ] && [ "$KIND" = secondmate ]; then + echo "error: secondmate $ID's recorded endpoint is gone; its recovery is owned by the secondmate respawn path, not by relaunch (run bin/fm-spawn.sh $ID --secondmate, or let the session-start liveness sweep do it)" >&2 + exit 1 + fi MODE=$(fm_meta_get "$RELAUNCH_META" mode) YOLO=$(fm_meta_get "$RELAUNCH_META" yolo) RELAUNCH_WT=$(fm_meta_get "$RELAUNCH_META" worktree) @@ -1580,6 +1652,12 @@ if [ "$RELAUNCH" -eq 1 ]; then } fi if [ "$BACKEND" = herdr ]; then + # fm-spawn uses HERDR_PANE_ID for the TASK's pane, while the herdr adapter + # reads that SAME name as the pane THIS process is itself running in + # (fm_backend_herdr_launcher_identity). The record is about to overwrite it, + # so keep what herdr actually injected: a rebind still has to prove its own + # launcher identity, and a task's recorded pane is not it. + RELAUNCH_LAUNCHER_PANE_ID=${HERDR_PANE_ID:-} HERDR_SES=$(fm_meta_get "$RELAUNCH_META" herdr_session) HERDR_WORKSPACE_ID=$(fm_meta_get "$RELAUNCH_META" herdr_workspace_id) HERDR_TAB_ID=$(fm_meta_get "$RELAUNCH_META" herdr_tab_id) @@ -3042,16 +3120,92 @@ fi W="fm-$ID" if [ "$RELAUNCH" -eq 1 ]; then - # Adopt the recorded endpoint instead of creating one. This is what keeps a - # relaunch a REPLACEMENT rather than a second copy of the task: no new - # terminal, no second worktree, and every uncommitted change left exactly - # where the previous agent left it. - T=$RELAUNCH_TARGET # A secondmate's home already resolved WT above through the same validation a # fresh secondmate spawn uses; every other kind takes the recorded worktree. + # Either way the worktree is REUSED, never re-created: its branch, commits and + # uncommitted changes are exactly as the previous agent left them, and nothing + # below may touch them. [ "$KIND" = secondmate ] || WT=$RELAUNCH_WT - WT_TARGET=$T - SES=${T%%:*} + if [ "$RELAUNCH_REBIND" -eq 0 ]; then + # Adopt the recorded endpoint instead of creating one. This is what keeps a + # relaunch a REPLACEMENT rather than a second copy of the task: no new + # terminal, no second worktree, and every uncommitted change left exactly + # where the previous agent left it. + T=$RELAUNCH_TARGET + WT_TARGET=$T + SES=${T%%:*} + else + # The recorded endpoint is authoritatively gone, so there is nothing to + # adopt: create ONE fresh endpoint for the same task, opened directly in the + # recorded worktree. The record published below writes window= (and herdr's + # ids) from these values, which is the whole rebind - the task id, brief, + # worktree, armed poll and status log are untouched. + # + # Herdr is the ONLY backend that reaches here: the gate above rebinds only + # on a PROVEN-gone endpoint, and absence is provable only on herdr, whose + # every read is scoped to the session the record names + # (fm_control_endpoint_absence_verdict owns that argument). tmux and every + # secondmate were already refused, so there is no dispatch left to make. + # + # This deliberately uses the FLAT container shape rather than Herdr's + # presentation projection: projection is a presentation-only layout that is + # never endpoint or ownership authority, and flat is already the documented + # fallback for every recovery it cannot bind exactly + # (docs/herdr-backend.md "Presentation spaces"). + # + # KNOWN LIMITATION (bead fm-herdr-rebind-leak-20260913): the tab minted + # below is registered with no abort cleanup, so a later refusal leaves that + # pane behind and a retry mints another. Documented in + # docs/agent-control.md rather than fixed here, because the remedy is + # machinery the ordinary flat spawn path does not have either. + # + # Re-create the tab under the RECORDED herdr session. Without the explicit + # session the container would resolve from the AMBIENT one + # (${HERDR_SESSION:-default}), so reclaiming a task recorded on a named + # session from a seat that is not in it would silently relocate the task + # onto another herdr server - an identity change, published as a + # self-consistent but wrong record. + HERDR_REBIND_SES=${RELAUNCH_TARGET%%:*} + HERDR_CONTAINER_RAW=$(HERDR_PANE_ID="$RELAUNCH_LAUNCHER_PANE_ID" \ + fm_backend_herdr_container_ensure "$PROJ_ABS" launcher-home "$HERDR_REBIND_SES") || { + # container_ensure returns 1 for several unrelated reasons - a failed + # version check, a server that will not start, an ambiguous workspace + # label, a cross-session launcher identity, a failed workspace create - + # and each already printed its own accurate message. Add only what this + # layer actually knows, and name the session mismatch solely when there + # IS one, rather than asserting a cause this condition cannot establish. + # + # A seat with NO herdr pane never reaches the cross-session guard at all: + # fm_backend_herdr_launcher_identity returns 2 for it and the placement + # falls back to the recorded session's labeled container, which is what + # makes a plain ssh or cron reclaim work. Its ambient session still reads + # `default` (fm_backend_herdr_session's fallback), so the inequality alone + # would fire for EVERY named-session task reclaimed from a plain shell and + # send the operator chasing a session mismatch that was never the cause. + HERDR_AMBIENT_SES=$(fm_backend_herdr_session) + if [ -n "$RELAUNCH_LAUNCHER_PANE_ID" ] && [ "$HERDR_AMBIENT_SES" != "$HERDR_REBIND_SES" ]; then + echo "error: task $ID's endpoint could not be re-created in its recorded herdr session '$HERDR_REBIND_SES'; this seat is running in herdr session '$HERDR_AMBIENT_SES', and a reclaim never moves a task to another session" >&2 + else + echo "error: task $ID's endpoint could not be re-created in its recorded herdr session '$HERDR_REBIND_SES'; see the refusal above for what failed" >&2 + fi + exit 1 + } + CONTAINER=${HERDR_CONTAINER_RAW%%$'\t'*} + HERDR_SEEDED_DEFAULT_TAB_ID=${HERDR_CONTAINER_RAW#*$'\t'} + HERDR_SES=${CONTAINER%%:*} + HERDR_WORKSPACE_ID=${CONTAINER#*:} + HERDR_TASK_IDS=$(fm_backend_herdr_create_task "$CONTAINER" "$W" "$WT" "$HERDR_SEEDED_DEFAULT_TAB_ID") || exit 1 + read -r HERDR_TAB_ID HERDR_PANE_ID <&2 + exit 1 + fi + T="$HERDR_SES:$HERDR_PANE_ID" + SES=$HERDR_SES + WT_TARGET=$T + fi else case "$BACKEND" in tmux) diff --git a/docs/agent-control.md b/docs/agent-control.md index a2a4b1e49d8..19a0e4ad543 100644 --- a/docs/agent-control.md +++ b/docs/agent-control.md @@ -13,7 +13,7 @@ The failure repeated across harnesses and homes, and the workaround (remember to ## What the control plane owns -`bin/fm-control-lib.sh` is the single executable owner of three capability tables, with no side effects, so it can be read as a contract: +`bin/fm-control-lib.sh` is the single executable owner of three capability tables, which have no side effects, so they can be read as a contract: - The **verb allowlist**: `interrupt`, `exit`, `relaunch`. There is no arbitrary-text and no generic raw-key entry point. @@ -23,6 +23,8 @@ The failure repeated across harnesses and homes, and the workaround (remember to `bin/fm-send.sh`'s `--key` path reads the composer-clear table from this owner too, rather than keeping a second copy of it. - **Per-backend capability**: which named keys a runtime backend can deliver, and whether it has a recovery-grade agent-state classifier able to prove an agent stopped. +The one thing this file owns that is not a pure table is the [endpoint-absence proof](#reclaiming-a-task-whose-endpoint-is-gone) below, which does run backend reads; sourcing the file is still free. + A recorded `harness=` is not always an exact adapter name: a task launched from a raw command records that command's basename instead. `fm_control_harness_family` is the one place that prefix rule is stated, and an unrecognized value resolves to no adapter rather than being guessed into one. @@ -31,8 +33,8 @@ A recorded `harness=` is not always an exact adapter name: a task launched from | Verb | Effect | Postcondition | | --- | --- | --- | | `interrupt` | Deliver the harness's verified interrupt sequence while leaving the agent running. | Delivery succeeds while the endpoint still exists and the agent is still alive where the backend can classify that; cancellation is confirmed only from an adapter-owned acknowledgement and otherwise reports `cancel=unconfirmed`. | -| `exit` | Stop the agent, preserving the endpoint, the worktree, and every uncommitted change. | The backend's recovery-grade classifier reports the agent gone. Already-stopped is idempotent success. | -| `relaunch` | Replace the running agent with a new one in the same endpoint and worktree, on the exact recorded adapter or an explicitly chosen harness, model, and effort. | The new agent is alive on the recorded endpoint, and the durable record names the harness that is actually running. | +| `exit` | Stop the agent, preserving the endpoint, the worktree, and every uncommitted change. | The backend's recovery-grade classifier reports the agent gone. Already-stopped is idempotent success. An endpoint reading `missing` goes through the same [absence proof](#reclaiming-a-task-whose-endpoint-is-gone) the reclaim uses before anything is claimed about it, and only Herdr can supply one: proven gone reports `endpoint-gone` (the agent went with it, and the endpoint this verb normally preserves did not survive), a pane that turns out to be there and idle is the ordinary `already-stopped`, one whose agent is back takes the ordinary interrupt-then-exit path. A tmux `missing` always refuses rather than claim a stop it cannot see. | +| `relaunch` | Replace the running agent with a new one in the same worktree - and the same endpoint whenever that endpoint still exists - on the exact recorded adapter or an explicitly chosen harness, model, and effort. | The new agent is alive on the endpoint the task's record now names, and that record names the harness that is actually running. | An exit that delivers lifecycle input but cannot prove the agent stopped fails with `exit=unconfirmed`, reports the observed agent state and any interrupt cancellation claim, and never claims that nothing changed. Interrupt never rewrites busy state as proof of its own success. @@ -71,10 +73,63 @@ It is not deterministic across the verified adapters: codex, grok, and gemini re A ship or scout relaunch requires `--note`, because the replacement inherits the local copy but none of the conversation; the note is appended to the instructions it reads. A secondmate relaunch does not require one and never rewrites its standing charter. 4. **Stop the old agent** through the `exit` verb, with its postcondition. -5. **Launch the replacement** through its single owner, `bin/fm-spawn.sh --relaunch`, which adopts the recorded endpoint and worktree instead of creating either, clears the previous harness's per-task wiring, and arms a fresh busy generation. +5. **Launch the replacement** through its single owner, `bin/fm-spawn.sh --relaunch`, which reuses the recorded worktree instead of creating one, adopts the recorded endpoint when it still exists, clears the previous harness's per-task wiring, and arms a fresh busy generation. + When the recorded endpoint is proven gone rather than merely idle or unreachable - which only Herdr can establish - the launch owner creates one fresh endpoint in that same worktree and the republished record rebinds the task to it - see [Reclaiming a task whose endpoint is gone](#reclaiming-a-task-whose-endpoint-is-gone). Switching harness is therefore one ordinary relaunch rather than a separate mechanism. +### Reclaiming a task whose endpoint is gone + +A Herdr pane or workspace can be destroyed out from under a live task by churn or a session restart. +The task's worktree, branch, commits, and uncommitted changes all survive that; only its terminal does not. + +**Reclaim is Herdr-only.** On tmux, both verbs refuse a `missing` endpoint, leaving it exactly as deadlocked as it was before this mechanism existed - deliberately, and with the reason stated rather than guessed past. + +Two endpoint verdicts are agent-free, and both license a relaunch: + +- `dead` - the endpoint exists and confidently holds no agent. It is **adopted**, so the task keeps its exact recorded address. +- gone, **proven** - there is no endpoint and therefore no agent, and it cannot be adopted, so the launch owner **creates one fresh endpoint in the recorded worktree** and the republished record rebinds the task to it. + +That proof is its own step, because the classifier's `missing` is not one state: it conflates *the endpoint was destroyed* with *the endpoint is unreachable from here right now*. +An unreachable endpoint can still hold the live agent a rebind would duplicate, so absence is proven and never inferred from a failed read - and whether it is provable at all is a property of the backend: + +- **Herdr can prove it.** Every read goes through the adapter's `--session ` CLI, so the recheck starts and reads the session the *record* names, through that session's own socket. + It starts that server (only the server: no workspace and no tab are created) and **re-reads the recorded pane**. + `dead` means the pane survived the restart and is adopted after all, with no second tab; `alive` means the agent came back and refuses; only a second `missing` proves the pane itself did not survive ([`docs/herdr-backend.md`](herdr-backend.md) "Restart and liveness behavior"). + That server start is a real side effect, and the parenthetical above does not cover it: when the recorded session's server no longer exists at all, the probe stands a fresh empty one up in order to ask, and nothing afterwards uses it. + So in that state `exit` - which otherwise reads as a read-only inspection - leaves an idle herdr server behind. +- **tmux cannot.** `list-windows -a` describes only the tmux server the *current process* addresses (its `TMUX_TMPDIR`/socket), and a task record carries no socket identity for its endpoint. + A different but running server would answer "not anywhere" about a window it was never able to see, so a server-wide read cannot tell a destroyed window from one on a server this process cannot address. + There is no read available that closes that gap, so tmux always refuses - for a renamed session, a moved window, a foreign socket, and a dead server alike. + +Every transient or self-contradicting read stays `unreadable` or `ambiguous` and still refuses, so a momentary backend failure can never be mistaken for absence. + +That proof has one owner for the whole control plane (`fm_control_endpoint_absence_verdict` in `bin/fm-control-lib.sh`), so `exit` and `relaunch` cannot reach two different answers about one endpoint. +`exit` reports what the proof established and nothing more - see its row in the verb table above. + +What a reclaim is not: + +- It is **not a teardown**. The worktree is reused exactly as the previous agent left it; nothing unlanded is ever discarded, and the ordinary `--note` requirement still applies. +- It does **not** change the task's identity. The task id, its armed poll and registration, and its status log are untouched; only the endpoint binding in the record moves. + Its instructions are the one exception, and only in the way an ordinary relaunch already changes them: a ship or scout reclaim appends the required `--note` under a `## Progress note ()` heading in `data//brief.md`, so re-read that brief rather than assuming it is byte-identical - a reclaim that failed and was retried leaves one block per attempt. + A secondmate's standing charter is never rewritten. +- It is **not** a peer seat's operation. `fm-control` resolves an exact task id against **this** home's `state/`, so only the home that owns the task can reclaim it. +- It does **not** cover a secondmate. A secondmate whose endpoint is gone already has one owner for that recovery - `bin/fm-spawn.sh --secondmate`, driven by the session-start liveness sweep - so relaunch refuses and names it rather than becoming a second path to the same outcome. + +The re-created tab is opened in the herdr session the record names, never in whichever session the recovering seat happens to sit in - relocating a task onto another herdr server would be an identity change published as a self-consistent but wrong record. +A seat that *claims* a herdr launcher pane belonging to a different session is refused rather than allowed to place the endpoint somewhere else, so reclaim such a task from a seat in the recorded session. +A seat with no herdr launcher pane at all - a plain ssh or cron shell, which is the ordinary way an operator reclaims - is not refused: placement falls back to the recorded session's labeled container, so the tab still lands in the session the record names. +The reclaim pins the recorded **session** but not the **workspace**: the container follows the reclaiming seat, so a reclaim run from a seat inside the recorded session places the new tab in *that seat's* workspace rather than the recorded `herdr_workspace_id`, even when the recorded workspace still exists and only the pane was destroyed. +The record is republished consistently and no work is lost, but the task's `herdr_workspace_id` moves with it. +The pane id necessarily changes (the pane did not survive), and the record follows it. +A Herdr reclaim deliberately uses the flat container shape rather than presentation projection: projection is a presentation-only layout that is never endpoint or ownership authority, and flat is already the documented fallback for every recovery it cannot bind exactly ([`docs/herdr-backend.md`](herdr-backend.md)). + +**Known limitation - a refusal before the record is republished leaves a stray husk pane** (follow-up bead `fm-herdr-rebind-leak-20260913`). +The rebind registers no abort cleanup, so a refusal in the window between the new tab being created and the record being republished leaves that pane behind while the record still names the old, gone one. +The stray pane holds a bare shell - the harness is not delivered until after publication - so the next reclaim cleans up after it: the re-created tab carries the same `fm-` label, `tab create` finds it, classifies it a husk, and closes and replaces it. +That self-heals only when the retry resolves the *same* workspace, which the placement rule above does not guarantee. +The worktree and the task's records are unaffected either way. + ### Failure and rollback - A refusal **before** the agent is stopped leaves the durable record and the instructions byte-identical. @@ -102,7 +157,8 @@ Switching harness is therefore one ordinary relaunch rather than a separate mech - An ambiguous or unreadable endpoint state refuses. Only a positively classified state acts. - `exit`'s composer-empty check, above, is itself a fail-closed boundary that `relaunch` inherits by stopping the old agent through `exit`. -- `fm-spawn --relaunch` independently refuses unless the recorded endpoint is positively agent-free, so a replacement can never join a live agent. +- `fm-spawn --relaunch` independently refuses unless the endpoint is positively agent-free - either a `dead` endpoint that survives, or a Herdr endpoint proven gone by the absence proof above - so a replacement can never join a live agent. + An `alive`, `ambiguous`, or `unreadable` verdict all refuse, and so does any endpoint whose absence is not provable, which on tmux is every `missing`; absence is claimed only from positive evidence of it. It also requires the shell to be in the recorded worktree: tmux refuses immediately when it is not, while Herdr sends one `cd` to the recorded path and refuses unless a subsequent path read confirms the move. ## Capability matrix @@ -123,5 +179,5 @@ The empirical basis for each adapter's value is the `harness-adapters` skill's v ## Verification - `tests/fm-control.test.sh` - the adapter contract for its verified-harness lane (adapters outside the lane pin their control mechanics in their own harness suites), the backend capability matrix, exact-id scoping, the closed verb list, the busy, idle, dead, and idempotent lifecycle cases, and marker non-regression, all against a stubbed session provider. -- `tests/fm-control-relaunch.test.sh` - the relaunch transaction: identity preservation, harness switching, the progress note, checkpoint refusals, and rollback after a failed launch. +- `tests/fm-control-relaunch.test.sh` - the relaunch transaction: identity preservation, harness switching, the progress note, checkpoint refusals, rollback after a failed launch, and the endpoint-absence proof both verbs share - the Herdr reclaim of a destroyed endpoint, and tmux refusing one it cannot prove absent. - `tests/fm-control-herdr-smoke.test.sh` - the second state-verified backend against the real herdr binary, on an isolated throwaway lab session. diff --git a/docs/herdr-backend.md b/docs/herdr-backend.md index 1199bd8142d..ee944fdf1b1 100644 --- a/docs/herdr-backend.md +++ b/docs/herdr-backend.md @@ -72,6 +72,7 @@ That path needs the home label to identify exactly one workspace: two workspaces Avoid naming a personal workspace `firstmate` or `2ndmate-` for that reason, and because the adapter cannot distinguish that label collision from its own container. An older secondmate workspace using `firstmate-` is not migrated automatically; rename it manually before expecting new tasks or recovery to use it. Recovery and list-live still scan the first workspace matching the home label, because they address panes they already recorded rather than choosing where new work goes. +The one recovery that does place new work is the control plane's reclaim of a destroyed endpoint, which mints a replacement tab through this section's ordinary placement rules while pinning the herdr session the task's record names ([`agent-control.md`](agent-control.md) "Reclaiming a task whose endpoint is gone"). Existing task operations use recorded endpoint ids and do not move a live task when labels change. The per-home workspace is reused while it has task tabs. diff --git a/docs/scripts.md b/docs/scripts.md index ef45d68dfa2..a03683f16df 100644 --- a/docs/scripts.md +++ b/docs/scripts.md @@ -118,7 +118,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-lease.sh` | Claim, release, inspect, and sweep per-task supervision leases | | `fm-lease-lib.sh` | One owner of the supervision lease contract and the main-only role-partition guards | | `fm-control.sh` | Agent lifecycle control plane: allowlisted `interrupt`, `exit`, and transactional `relaunch` verbs for an exact task id ([agent-control.md](agent-control.md)) | -| `fm-control-lib.sh` | One executable owner of the control-plane verb allowlist, per-harness interrupt/exit mechanics, and per-backend capability | +| `fm-control-lib.sh` | One executable owner of the control-plane verb allowlist, per-harness interrupt/exit mechanics, per-backend capability, and the endpoint-absence proof both `exit` and `relaunch` read | | `fm-busy-lib.sh` | Single owner of the semantic busy-state contract: verdicts, source attribution, and per-harness sources | | `fm-busy-event.sh` | The only writer of a task's semantic busy-state record and native-harness progress marker; arms an incarnation and applies lifecycle events | | `fm-tmux-lib.sh` | Shared tmux pane primitives for composer capture, verified submit, and the submit-time busy check | diff --git a/tests/fm-control-relaunch.test.sh b/tests/fm-control-relaunch.test.sh index ae3da17d95f..5631488cebd 100755 --- a/tests/fm-control-relaunch.test.sh +++ b/tests/fm-control-relaunch.test.sh @@ -121,7 +121,53 @@ case "${1:-}" in printf '╭────╮\n│ │\n╰────╯\n' fi exit 0 ;; - list-windows) [ -f "$D/windows" ] && cat "$D/windows"; exit 0 ;; + list-windows) + # The three shapes real tmux answers a per-session inventory with. The + # first two are DEFINITIVE and classify `missing`; the third is not and + # classifies `unreadable`. + if [ -f "$D/server-dead" ]; then + echo 'no server running on /tmp/tmux-1000/default' >&2 + exit 1 + fi + if [ -f "$D/session-missing" ]; then + echo "can't find session: $(cat "$D/session-name")" >&2 + exit 1 + fi + if [ -f "$D/inventory-broken" ]; then + echo 'lost server' >&2 + exit 1 + fi + [ -f "$D/windows" ] && cat "$D/windows"; exit 0 ;; + new-session) + # Nothing in the relaunch path may ever create a session; recording the + # call is how a refusal test proves that. + shift + ses= + while [ $# -gt 0 ]; do + case "$1" in + -s) ses=${2:-}; shift 2 ;; + *) shift ;; + esac + done + printf '%s\n' "$ses" >> "$D/created-sessions" + exit 0 ;; + new-window) + # Model the one thing an endpoint re-creation depends on: the window now + # appears in the session inventory, so the very next agent-state read stops + # answering `missing`. Echo a stable window id the way the real -P -F does. + shift + name= + while [ $# -gt 0 ]; do + case "$1" in + -n) name=${2:-}; shift 2 ;; + -c|-t) shift 2 ;; + *) shift ;; + esac + done + printf '%s\n' "$name" >> "$D/windows" + printf '%s\n' "$name" >> "$D/created-windows" + printf '@9\n' + exit 0 ;; esac exit 0 SH @@ -143,13 +189,14 @@ new_case() { printf 'claude' > "$dir/fake/command" printf 'claude' > "$dir/fake/becomes" printf '%s\n' "fm-$id" > "$dir/fake/windows" + printf '%s' fmses > "$dir/fake/session-name" make_tmux_stub "$dir" printf '%s\n' "$dir" } -# add_ship_task [harness] +# add_ship_task [harness] [session] add_ship_task() { - local dir=$1 id=$2 harness=${3:-claude} + local dir=$1 id=$2 harness=${3:-claude} ses=${4:-fmses} local home="$dir/home" proj="$dir/proj" wt="$dir/wt" fm_git_worktree "$proj" "$wt" "task-$id" mkdir -p "$home/data/$id" @@ -162,7 +209,7 @@ Exercise relaunch behavior for $id. Preserve the task while replacing its agent process. EOF { - echo "window=fmses:fm-$id" + echo "window=$ses:fm-$id" echo "endpoint_task_id=$id" echo "worktree=$wt" echo "project=$proj" @@ -175,6 +222,7 @@ EOF echo "effort=default" } > "$home/state/$id.meta" printf '%s\n' "fm-$id" > "$dir/fake/windows" + printf '%s' "$ses" > "$dir/fake/session-name" printf '%s' "$wt" > "$dir/fake/cwd" TASK_TMPS+=("/tmp/fm-$id") } @@ -185,7 +233,9 @@ run_control() { # # store (bin/fm-claude-trust.sh), and a relaunch reaches it through fm-control.sh, so this runs against a throwaway HOME; # without it this suite would write the developer's real ~/.claude.json. mkdir -p "$dir/user-home" - env PATH="$dir/fakebin:$PATH" FM_HOME="$dir/home" FM_FAKE_DIR="$dir/fake" \ + env -u HERDR_ENV -u HERDR_PANE_ID -u HERDR_SESSION -u HERDR_SOCKET_PATH \ + -u HERDR_TAB_ID -u HERDR_WORKSPACE_ID \ + PATH="$dir/fakebin:$PATH" FM_HOME="$dir/home" FM_FAKE_DIR="$dir/fake" \ HOME="$dir/user-home" CLAUDE_CONFIG_DIR='' \ FM_SPAWN_NO_GUARD=1 GROK_HOME="$dir/grokhome" \ FM_CONTROL_POLL=0.01 FM_CONTROL_EXIT_WAIT=0.05 FM_CONTROL_LAUNCH_WAIT=0.05 \ @@ -205,7 +255,9 @@ run_spawn() { # # store (bin/fm-claude-trust.sh), so it runs against a throwaway HOME; # without it this suite would write the developer's real ~/.claude.json. mkdir -p "$dir/user-home" - env PATH="$dir/fakebin:$PATH" FM_HOME="$dir/home" FM_FAKE_DIR="$dir/fake" \ + env -u HERDR_ENV -u HERDR_PANE_ID -u HERDR_SESSION -u HERDR_SOCKET_PATH \ + -u HERDR_TAB_ID -u HERDR_WORKSPACE_ID \ + PATH="$dir/fakebin:$PATH" FM_HOME="$dir/home" FM_FAKE_DIR="$dir/fake" \ HOME="$dir/user-home" CLAUDE_CONFIG_DIR='' \ FM_SPAWN_NO_GUARD=1 GROK_HOME="$dir/grokhome" \ "$SPAWN" "$@" 2>&1 @@ -1649,6 +1701,467 @@ test_spawn_relaunch_refuses_a_pane_outside_the_worktree() { pass "fm-spawn --relaunch: refuses to start a replacement outside the copy holding its work" } +# --- 7. reclaiming a task whose endpoint is gone ---------------------------- +# +# Before this, `missing` was a terminal state: fm-spawn --relaunch accepted only +# `dead` and told the caller to stop the agent first, while fm-control exit +# refused `missing` outright and told the caller to reconcile the task first - +# and there is no reconcile verb. Each command named the other as its +# prerequisite, so a task whose pane or workspace was destroyed could not be +# reclaimed by anything, and any no-mistakes approval it was parked on had no +# seat left to answer it. + +# strand_endpoint : make a tmux endpoint read `missing` the way +# a destroyed window does - a successful session inventory that omits the exact +# window. +strand_endpoint() { # + : > "$1/fake/windows" +} + +# Every tmux `missing` refuses on BOTH verbs, whatever produced it. tmux is the +# one verified backend whose absence cannot be proven from a task record: the +# record carries no socket identity for the endpoint, and any inventory +# describes only the server this process happens to address. So a window that +# is merely on a server this seat cannot reach is indistinguishable from one +# that was destroyed, and neither verb will guess. +assert_tmux_missing_refuses() { # + local dir=$1 id=$2 what=$3 out rc brief_before + + out=$(run_spawn "$dir" "$id" --relaunch --harness claude); rc=$? + expect_code 1 "$rc" "relaunch must refuse a tmux endpoint whose absence cannot be proven ($what)"$'\n'"$out" + assert_absent "$dir/fake/created-windows" "a refused relaunch must not create a window ($what)" + assert_absent "$dir/fake/created-sessions" "a refused relaunch must not create a session ($what)" + [ ! -s "$dir/fake/literal" ] || fail "a refused relaunch must send nothing into any pane ($what)" + + brief_before=$(cat "$dir/home/data/$id/brief.md") + out=$(run_control "$dir" "$id" exit); rc=$? + expect_code 1 "$rc" "exit must refuse a tmux endpoint whose absence cannot be proven ($what)"$'\n'"$out" + assert_not_contains "$out" "endpoint-gone" \ + "exit must not report a stop it cannot see ($what)" + [ ! -s "$dir/fake/literal" ] || fail "a refused exit must send nothing into any pane ($what)" + + out=$(run_control "$dir" "$id" relaunch --note "this note must never reach a live agent"); rc=$? + expect_code 1 "$rc" "the relaunch transaction must fail closed ($what)"$'\n'"$out" + [ "$(cat "$dir/home/data/$id/brief.md")" = "$brief_before" ] \ + || fail "a refused relaunch edited instructions an agent that may still be running is reading ($what)" + assert_absent "$dir/fake/created-windows" "a refused transaction must not create a window ($what)" + assert_absent "$dir/fake/created-sessions" "a refused transaction must not create a session ($what)" + [ ! -s "$dir/fake/literal" ] || fail "a refused transaction must launch nothing ($what)" +} + +test_tmux_refuses_a_window_missing_from_its_session() { + local dir + dir=$(new_case tmux-gone rl60) + add_ship_task "$dir" rl60 claude + strand_endpoint "$dir" rl60 + assert_tmux_missing_refuses "$dir" rl60 "window absent from a readable session inventory" + pass "tmux: a window absent from its session refuses both verbs rather than being assumed gone" +} + +test_tmux_refuses_a_session_that_cannot_be_found() { + local dir + dir=$(new_case tmux-nosession rl61) + add_ship_task "$dir" rl61 claude + # Real tmux's answer to a renamed session, and to a different + # TMUX_TMPDIR/socket: definitive about the SESSION, silent about whether the + # window and its agent survived elsewhere. + : > "$dir/fake/session-missing" + assert_tmux_missing_refuses "$dir" rl61 "recorded session not found" + pass "tmux: an unfindable session refuses both verbs, so a live agent is never duplicated" +} + +test_tmux_refuses_when_the_server_is_gone() { + local dir + dir=$(new_case tmux-noserver rl62) + add_ship_task "$dir" rl62 claude + # No server on the socket this process addresses. Another server may still be + # running the task's window, and the record cannot say which socket is its. + : > "$dir/fake/server-dead" + assert_tmux_missing_refuses "$dir" rl62 "no tmux server on this socket" + pass "tmux: a dead server on this socket refuses both verbs rather than proving absence" +} + +test_reclaim_refuses_an_unreadable_endpoint() { + local dir out rc + dir=$(new_case gone-unreadable rl63) + add_ship_task "$dir" rl63 claude + # The inventory itself fails non-definitively. That is not evidence of + # absence, and reading it as one is exactly how two agents end up in one + # endpoint. + : > "$dir/fake/inventory-broken" + + out=$(run_spawn "$dir" rl63 --relaunch --harness claude); rc=$? + expect_code 1 "$rc" "an unreadable endpoint must still refuse" + assert_contains "$out" "positively agent-free endpoint" \ + "only a POSITIVELY proven agent-free endpoint may be relaunched into" + assert_absent "$dir/fake/created-windows" \ + "a refused relaunch must not create an endpoint" + [ ! -s "$dir/fake/literal" ] || fail "a refused relaunch must launch nothing" + pass "reclaim: an unclassifiable endpoint is still refused, so two agents cannot share one" +} + +# --- herdr: a stopped server is not a destroyed endpoint -------------------- +# +# Stopping and restarting a named Herdr server preserves workspace, tab, pane +# and label ids; only the harness processes and their registrations die +# (docs/herdr-backend.md "Restart and liveness behavior"). The recovery-grade +# classifier still reads a stopped server as `missing`, so a reclaim that +# believed that verdict would abandon a pane that was about to come back and +# open a second tab beside it. +# +# Canned/stateful fake only - never a real herdr session. +make_herdr_stub() { # + local fb="$1/fakebin" + mkdir -p "$fb" + # The herdr server-ensure poll must actually wait between reads, so this case + # keeps the real sleep rather than the tmux cases' instant stub. + rm -f "$fb/sleep" + cat > "$fb/herdr" <<'SH' +#!/usr/bin/env bash +set -u +D=$FM_FAKE_DIR +printf '%s\n' "$*" >> "$D/herdr-log" +if [ "${1:-}" = status ] && [ "${2:-}" = --json ]; then + if [ -f "$D/herdr-stopped" ]; then + printf '{"client":{"version":"0.9.0","protocol":22},"server":{"running":false}}\n' + else + printf '{"client":{"version":"0.9.0","protocol":22},"server":{"running":true}}\n' + fi + exit 0 +fi +if [ "${1:-}" = server ]; then + rm -f "$D/herdr-stopped" + exit 0 +fi +if [ -f "$D/herdr-stopped" ]; then + # Every operational call against a stopped server fails at the transport, + # with no JSON body to classify. + echo 'error: could not connect to the herdr server' >&2 + exit 1 +fi +case "${1:-} ${2:-}" in + 'pane get') + if [ "${3:-}" = "$(cat "$D/herdr-pane")" ]; then + printf '{"result":{"pane":{"pane_id":"%s","foreground_cwd":"%s"}}}\n' \ + "${3:-}" "$(cat "$D/cwd")" + else + # Only the pane this case says survived can be read back. Any other pane + # id is structurally gone, which is herdr's `pane_not_found`. + printf '{"error":{"code":"pane_not_found"}}\n' + fi + exit 0 ;; + 'agent get') + if [ -f "$D/herdr-agent-live" ]; then + # The agent came back with its server. Nothing here is reclaimable. + printf '{"result":{"agent":{"agent_status":"idle"}}}\n' + else + # A pane that comes back holding no agent is the adoptable state. + printf '{"error":{"code":"agent_not_found"}}\n' + fi + exit 0 ;; + 'pane process-info') + # Only asked for once an agent IS registered, to prove it at process level. + printf '{"result":{"type":"pane_process_info","process_info":{"pane_id":"%s","shell_pid":4242,"foreground_processes":[{"pid":4243,"name":"claude","argv":["claude"],"cmdline":"claude"}]}}}\n' \ + "$(cat "$D/herdr-pane")" + exit 0 ;; + 'pane send-text') + # Mirrors the tmux fake's `becomes`: delivering the launch brief is what + # makes an agent exist on this pane, so the control plane's alive-wait can + # observe the replacement come up. A launch arrives as a short line sourcing + # the staged launch file rather than the literal command, so read that file + # back before deciding what was delivered - exactly as the tmux fake above + # and tests/fixtures.sh do. + payload=${4:-} + case "$payload" in + ". '"*"'") staged=${payload#". '"}; staged=${staged%"'"}; [ ! -f "$staged" ] || payload=$(cat "$staged") ;; + esac + case "$payload" in + *'encode launch-brief'*) : > "$D/herdr-agent-live" ;; + esac + exit 0 ;; + 'workspace list') + printf '{"result":{"workspaces":[]}}\n' + exit 0 ;; + 'workspace create') + if [ -f "$D/herdr-workspace-create-fails" ]; then + echo 'error: workspace create failed' >&2 + exit 1 + fi + printf '{"result":{"workspace":{"workspace_id":"wsnew"},"tab":{"tab_id":"seedtab"}}}\n' + exit 0 ;; + 'tab list') + printf '{"result":{"tabs":[]}}\n' + exit 0 ;; + 'tab create') + # The re-created endpoint. Recording it lets a case prove the pane the + # record ends up naming is the one this call minted. + printf '%s\n' "$*" >> "$D/herdr-created-tabs" + printf '{"result":{"tab":{"tab_id":"tabnew"},"root_pane":{"pane_id":"%%9"}}}\n' + # From here on the new pane is the one that reads back. + printf '%s' '%9' > "$D/herdr-pane" + exit 0 ;; +esac +exit 0 +SH + chmod +x "$fb/herdr" +} + +# add_herdr_ship_task [session] [surviving-pane]: a ship task +# recorded on the herdr backend, with its server stopped so its endpoint +# classifies `missing`. is the pane id the fake will answer for +# once that server is back; default is the recorded one (it survived the +# restart). Pass a different id to model a pane that genuinely did not. +add_herdr_ship_task() { # [session] [surviving-pane] + local dir=$1 id=$2 ses=${3:-fmlab} survivor=${4:-'%7'} + local home="$dir/home" proj="$dir/proj" wt="$dir/wt" + fm_git_worktree "$proj" "$wt" "task-$id" + mkdir -p "$home/data/$id" + cat > "$home/data/$id/brief.md" < "$home/state/$id.meta" + printf '%s' "$wt" > "$dir/fake/cwd" + printf '%s' "$survivor" > "$dir/fake/herdr-pane" + : > "$dir/fake/herdr-log" + : > "$dir/fake/herdr-stopped" + TASK_TMPS+=("/tmp/fm-$id") +} + +# Sets HERDR_CASE_DIR rather than echoing it, so callers invoke it as a plain +# statement. A `dir=$(herdr_case_or_skip ...)` would run add_herdr_ship_task in +# a command-substitution subshell, where its TASK_TMPS registration would +# mutate a discarded copy and the EXIT trap would never remove the +# out-of-tmproot /tmp/fm- root the spawn creates. +HERDR_CASE_DIR= +herdr_case_or_skip() { # [session] [surviving-pane] + HERDR_CASE_DIR= + command -v jq >/dev/null 2>&1 || return 1 + HERDR_CASE_DIR=$(new_case "$1" "$2") + add_herdr_ship_task "$HERDR_CASE_DIR" "$2" "${3:-fmlab}" "${4:-%7}" + make_herdr_stub "$HERDR_CASE_DIR" + return 0 +} + +test_herdr_reclaim_adopts_a_pane_that_outlived_its_server() { + local dir out rc=0 log stray + herdr_case_or_skip gone-herdr rl68 || { + echo "skip - herdr reclaim needs jq (the herdr adapter parses JSON with it)" + return 0 + } + dir=$HERDR_CASE_DIR + + out=$(run_spawn "$dir" rl68 --relaunch --harness claude) || rc=$? + log=$(cat "$dir/fake/herdr-log") + expect_code 0 "$rc" "a pane that outlived its stopped server is adoptable"$'\n'"$out"$'\n'"$log" + + assert_contains "$log" "server --session fmlab" \ + "the reclaim must bring the RECORDED session's server back before deciding anything" + assert_contains "$log" "agent get %7 --session fmlab" \ + "the reclaim must re-read the recorded pane once its server is running" + assert_not_contains "$log" "workspace create" \ + "adopting a preserved pane must not create a workspace" + assert_not_contains "$log" "tab create" \ + "adopting a preserved pane must not open a second tab beside it" + # Every call belongs to the session the record names. A rebind resolves its + # container from the ambient session instead, which is how the preserved pane + # ends up orphaned in a workspace nothing points at. + stray=$(printf '%s\n' "$log" | grep -v -- '--session fmlab$' | grep -v '^status --json$' || true) + [ -z "$stray" ] || fail "a herdr reclaim touched a session the record does not name: $stray" + assert_contains "$out" "window=fmlab:%7" "the reclaim should report the adopted endpoint" + [ "$(meta_field "$dir" rl68 herdr_pane_id)" = '%7' ] \ + || fail "the adopted record's pane id changed, got $(meta_field "$dir" rl68 herdr_pane_id)" + [ "$(meta_field "$dir" rl68 herdr_tab_id)" = tab1 ] \ + || fail "the adopted record's tab id changed, got $(meta_field "$dir" rl68 herdr_tab_id)" + [ "$(meta_field "$dir" rl68 window)" = 'fmlab:%7' ] \ + || fail "the adopted record's endpoint moved, got $(meta_field "$dir" rl68 window)" + assert_contains "$log" "pane send-text %7 " \ + "the replacement's launch brief must be delivered into the adopted pane" + pass "reclaim: a herdr pane that outlived its stopped server is adopted, never orphaned beside a new tab" +} + +test_herdr_exit_reports_already_stopped_when_the_pane_outlived_its_server() { + local dir out rc=0 + herdr_case_or_skip gone-herdr-exit rl72 || { + echo "skip - herdr exit needs jq (the herdr adapter parses JSON with it)" + return 0 + } + dir=$HERDR_CASE_DIR + + out=$(run_control "$dir" rl72 exit) || rc=$? + expect_code 0 "$rc" "a pane that outlived its stopped server holds no agent, which is success"$'\n'"$out" + assert_contains "$out" "already-stopped" \ + "the endpoint is there and idle, which is the ordinary already-stopped outcome" + assert_not_contains "$out" "endpoint-gone" \ + "a pane that survived its server's restart was never gone" + [ "$(meta_field "$dir" rl72 window)" = 'fmlab:%7' ] \ + || fail "exit must leave the recorded endpoint exactly as it found it" + pass "fm-control exit: a herdr pane that outlived its stopped server is already-stopped, not gone" +} + +test_herdr_rebind_stays_in_the_recorded_session() { + local dir out rc=0 log + # The record names session `fmlab`; this seat has no ambient HERDR_SESSION, so + # the adapter's own default is `default`. The recorded pane does NOT come back + # with the server, so this reclaim really does rebind - and the rebind must + # land in `fmlab`, never in `default`. + herdr_case_or_skip gone-herdr-pin rl73 fmlab '%none' || { + echo "skip - herdr rebind needs jq (the herdr adapter parses JSON with it)" + return 0 + } + dir=$HERDR_CASE_DIR + + out=$(run_spawn "$dir" rl73 --relaunch --harness claude) || rc=$? + log=$(cat "$dir/fake/herdr-log") + expect_code 0 "$rc" "a herdr pane that did not survive its server should be rebound"$'\n'"$out"$'\n'"$log" + + assert_contains "$log" "tab create" "a destroyed pane must be replaced by a fresh tab" + [ -z "$(grep -v -- '--session fmlab$' <<<"$log" | grep -v '^status --json$' || true)" ] \ + || fail "the rebind used a herdr session the record does not name: $log" + [ "$(meta_field "$dir" rl73 herdr_session)" = fmlab ] \ + || fail "the rebound record left its recorded herdr session, got $(meta_field "$dir" rl73 herdr_session)" + [ "$(meta_field "$dir" rl73 window)" = 'fmlab:%9' ] \ + || fail "the rebound endpoint should be the new pane in the recorded session, got $(meta_field "$dir" rl73 window)" + [ "$(meta_field "$dir" rl73 herdr_pane_id)" = '%9' ] \ + || fail "the rebound record should name the pane the reclaim minted, got $(meta_field "$dir" rl73 herdr_pane_id)" + pass "reclaim: a herdr rebind is created in the session the record names, never the ambient one" +} + +test_herdr_reclaim_refuses_an_agent_that_came_back() { + local dir out rc log + herdr_case_or_skip gone-herdr-alive rl74 || { + echo "skip - herdr reclaim needs jq (the herdr adapter parses JSON with it)" + return 0 + } + dir=$HERDR_CASE_DIR + # The server was stopped, so the first read says `missing` - but starting it + # brings the pane AND its agent back. A rebind here would put a second agent + # in this task's worktree, which is the whole reason absence is re-proven. + : > "$dir/fake/herdr-agent-live" + + out=$(run_spawn "$dir" rl74 --relaunch --harness claude); rc=$? + log=$(cat "$dir/fake/herdr-log") + expect_code 1 "$rc" "a returning agent must refuse, never be duplicated"$'\n'"$out"$'\n'"$log" + assert_contains "$out" "alive" "the refusal should name the state it actually read" + assert_not_contains "$log" "tab create" "a refused reclaim must not mint a second tab" + assert_not_contains "$log" "workspace create" "a refused reclaim must not create a workspace" + [ "$(meta_field "$dir" rl74 herdr_pane_id)" = '%7' ] \ + || fail "a refused reclaim rewrote the record's pane id" + pass "reclaim: a herdr agent that came back with its server refuses, so one worktree keeps one agent" +} + +test_herdr_reclaim_keeps_the_task_whole() { + local dir out rc=0 head_before + herdr_case_or_skip gone-herdr-work rl75 fmlab '%none' || { + echo "skip - herdr reclaim needs jq (the herdr adapter parses JSON with it)" + return 0 + } + dir=$HERDR_CASE_DIR + printf 'landed on the branch\n' > "$dir/wt/committed.txt" + git -C "$dir/wt" add committed.txt + git -C "$dir/wt" -c user.email=t@example.com -c user.name=t commit -qm "work in progress" + head_before=$(git -C "$dir/wt" rev-parse HEAD) + printf 'never committed\n' > "$dir/wt/dirty.txt" + + # A reclaim rebinds the ENDPOINT and nothing else. Everything that identifies + # the task must come through untouched: a record row the reclaim does not + # own, the armed watcher check and the private binding that authorizes it, + # and the status log the supervisor reads. + printf '%s\n' "pr=https://example.invalid/pr/7" >> "$dir/home/state/rl75.meta" + printf '%s\n' '#!/usr/bin/env bash' 'exit 0' > "$dir/home/state/rl75.check.sh" + chmod 0700 "$dir/home/state/rl75.check.sh" + FM_HOME="$dir/home" "$ROOT/bin/fm-check-register.sh" rl75 >/dev/null \ + || fail "could not arm a custom check for the reclaim fixture" + printf 'working: parked on an approval nobody can answer\n' >> "$dir/home/state/rl75.status" + + out=$(run_control "$dir" rl75 relaunch --note "the pane was destroyed; pick the work back up") || rc=$? + expect_code 0 "$rc" "the owning seat should be able to reclaim a task whose pane is gone"$'\n'"$out" + + [ "$(git -C "$dir/wt" rev-parse HEAD)" = "$head_before" ] \ + || fail "a reclaim moved the worktree's HEAD" + [ "$(git -C "$dir/wt" rev-parse --abbrev-ref HEAD)" = "task-rl75" ] \ + || fail "a reclaim changed the worktree's branch" + assert_contains "$(cat "$dir/wt/dirty.txt")" "never committed" \ + "a reclaim destroyed or rewrote an uncommitted change" + assert_present "$dir/wt/committed.txt" "a reclaim destroyed committed work" + + [ "$(meta_field "$dir" rl75 worktree)" = "$dir/wt" ] \ + || fail "a reclaim must keep the recorded worktree" + [ "$(meta_field "$dir" rl75 pr)" = "https://example.invalid/pr/7" ] \ + || fail "a reclaim dropped a record row it does not own" + assert_present "$dir/home/state/rl75.check.sh" "a reclaim retired the task's armed check" + assert_present "$dir/home/state/rl75.check-trust" "a reclaim broke the armed check's registration" + assert_contains "$(cat "$dir/home/state/rl75.status")" "parked on an approval nobody can answer" \ + "a reclaim truncated the status log" + assert_contains "$(cat "$dir/home/data/rl75/brief.md")" "the pane was destroyed" \ + "the replacement must inherit the progress note" + [ "$(journal_field "$dir" rl75 exit_result)" = endpoint-gone ] \ + || fail "the transaction should record that the endpoint was already gone" + pass "reclaim: a herdr reclaim rebinds the endpoint and leaves the whole rest of the task alone" +} + +test_herdr_rebind_failure_from_a_plain_shell_names_the_real_cause() { + local dir out rc + # No HERDR_* env at all, which is how an operator reclaims from ssh or cron. + # The adapter's ambient session then reads `default` while the record names + # `fmlab`, but the cross-session launcher guard was never consulted - this + # seat claims no launcher pane, so placement fell back to the recorded + # session's labeled container and the container failed for its own reason. + herdr_case_or_skip gone-herdr-plain rl77 fmlab '%none' || { + echo "skip - herdr reclaim needs jq (the herdr adapter parses JSON with it)" + return 0 + } + dir=$HERDR_CASE_DIR + : > "$dir/fake/herdr-workspace-create-fails" + + out=$(run_spawn "$dir" rl77 --relaunch --harness claude); rc=$? + expect_code 1 "$rc" "a container that cannot be ensured must refuse"$'\n'"$out" + assert_contains "$out" "fmlab" "the refusal should name the session the reclaim was targeting" + assert_not_contains "$out" "this seat is running in herdr session" \ + "a seat with no launcher pane never hit the cross-session guard, so the refusal must not blame one" + assert_not_contains "$out" "a reclaim never moves a task to another session" \ + "the operator must not be sent to re-run from another seat when that would not help" + pass "reclaim: a rebind refused from a plain shell reports the real cause, not a fabricated session mismatch" +} + +test_herdr_reclaim_of_a_secondmate_names_its_own_owner() { + local dir out rc + herdr_case_or_skip gone-herdr-secondmate rl76 fmlab '%none' || { + echo "skip - herdr reclaim needs jq (the herdr adapter parses JSON with it)" + return 0 + } + dir=$HERDR_CASE_DIR + printf '%s\n' "kind=secondmate" "home=$dir/wt" >> "$dir/home/state/rl76.meta" + + out=$(run_spawn "$dir" rl76 --relaunch --harness claude); rc=$? + expect_code 1 "$rc" "a secondmate reclaim belongs to the secondmate respawn path" + assert_contains "$out" "--secondmate" "the refusal should name the path that owns this recovery" + assert_not_contains "$(cat "$dir/fake/herdr-log")" "tab create" \ + "the refusal must happen before any endpoint is created" + pass "reclaim: a herdr secondmate whose endpoint is gone is sent to its own respawn owner" +} + test_relaunch_reverifies_an_already_in_flight_item_instead_of_rewriting_it() { local dir out rc=0 command -v tasks-axi >/dev/null 2>&1 || { @@ -1738,5 +2251,16 @@ test_spawn_relaunch_refuses_a_pending_authoritative_close test_spawn_relaunch_refuses_contradicting_flags test_spawn_relaunch_refuses_an_unrecorded_task test_spawn_relaunch_refuses_a_pane_outside_the_worktree +test_tmux_refuses_a_window_missing_from_its_session +test_tmux_refuses_a_session_that_cannot_be_found +test_tmux_refuses_when_the_server_is_gone +test_reclaim_refuses_an_unreadable_endpoint +test_herdr_reclaim_adopts_a_pane_that_outlived_its_server +test_herdr_exit_reports_already_stopped_when_the_pane_outlived_its_server +test_herdr_rebind_stays_in_the_recorded_session +test_herdr_reclaim_refuses_an_agent_that_came_back +test_herdr_reclaim_keeps_the_task_whole +test_herdr_reclaim_of_a_secondmate_names_its_own_owner +test_herdr_rebind_failure_from_a_plain_shell_names_the_real_cause test_relaunch_reverifies_an_already_in_flight_item_instead_of_rewriting_it test_relaunch_moves_a_drifted_item_back_in_flight diff --git a/tests/fm-control.test.sh b/tests/fm-control.test.sh index 5c00d6cb04e..67105bcfd99 100755 --- a/tests/fm-control.test.sh +++ b/tests/fm-control.test.sh @@ -634,15 +634,23 @@ test_already_stopped_exit_is_idempotent() { pass "fm-control exit: an already-stopped agent is idempotent success with no bytes sent" } -test_missing_endpoint_refuses() { +test_missing_tmux_endpoint_refuses_rather_than_claiming_a_stop() { local dir out rc dir=$(new_case gone) add_task "$dir" t1 claude : > "$dir/fake/windows" out=$(run_control "$dir" t1 exit); rc=$? - expect_code 1 "$rc" "a missing endpoint should refuse" - assert_contains "$out" "recorded endpoint is gone" "the refusal should name the missing endpoint" - pass "fm-control exit: a vanished endpoint refuses instead of silently succeeding" + # `missing` on tmux is not a finding about the endpoint. A task record carries + # no socket identity for it, and any inventory describes only the tmux server + # this process addresses, so a window that is merely on a server this seat + # cannot reach is indistinguishable from one that was destroyed. exit refuses + # rather than claim a stop it cannot see, and sends nothing to an address it + # cannot trust. Reclaim of a destroyed endpoint is Herdr-only + # (docs/agent-control.md "Reclaiming a task whose endpoint is gone"). + expect_code 1 "$rc" "a tmux endpoint whose absence cannot be proven must refuse" + assert_not_contains "$out" "endpoint-gone" "exit must not report a stop it could not prove" + [ -z "$(literals "$dir")" ] || fail "nothing may be sent into an endpoint exit cannot trust" + pass "fm-control exit: an unprovable tmux endpoint refuses instead of claiming the agent stopped" } test_interrupt_refuses_when_no_agent_runs() { @@ -900,7 +908,7 @@ test_verb_allowlist_is_closed test_resume_is_refused_with_its_reason test_relaunch_only_flags_are_rejected_on_other_verbs test_already_stopped_exit_is_idempotent -test_missing_endpoint_refuses +test_missing_tmux_endpoint_refuses_rather_than_claiming_a_stop test_interrupt_refuses_when_no_agent_runs test_ambiguous_endpoint_refuses test_busy_agent_is_interrupted_before_the_exit_command From a09090d13ef24ce3cf71d171ade119896d6db301 Mon Sep 17 00:00:00 2001 From: Tiago Date: Sun, 20 Sep 2026 03:26:49 -0300 Subject: [PATCH 003/168] feat(bin): stamp status events with their emission time (#3764) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * test(status): reproduce missing event emission time * wip(status): preserve optional event emission time * test(status): document indirect clock stub invocation * no-mistakes(review): Preserve historical status bytes during reply recovery * no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies * no-mistakes(review): Preserve captain regex overrides for timestamped status events * no-mistakes(document): Clarify status event timing and publication contracts * no-mistakes(lint): Quote literal done to satisfy ShellCheck * no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed * no-mistakes(test): Preserve terminal notifications with malformed timestamp tags * no-mistakes(test): Stamp Rovo spawn failures with emission time * no-mistakes(document): Verify status event documentation * no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests * no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed * no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed * no-mistakes(review): Stamp remote escalations at call sites, drop new flag * no-mistakes(review): Accept stamped escalation and close lines in test assertions * no-mistakes(review): Restore reserved-key answered-note guard for stamped closes * test(status): accept optional emission time in PR-provenance assertions The #4148 provenance test landed on main with exact unstamped greps. Parent-channel lines from this branch carry [at=], so strip only that tag before the same exact match. No production change. * no-mistakes(review): Accept stamped ready signal in PR fallback scrape * no-mistakes(review): Drop relay flag, stamp parent events at call sites * no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture * no-mistakes(review): Accept optional stamp in live cmux drift guard * no-mistakes(review): Restore original test invocation order in two suites * no-mistakes(review): Strip only well-formed numeric status time tags * no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc * no-mistakes(review): Stamp agy spawn-failure status lines with event time * fix(bin): normalize status event times in-shell and freeze the budget test clock Two paths made a status event's emission time cost more than it should. The captain-relevance fallback piped every line through awk to drop a well-formed `[at=]` tag before matching, so a supervisor sweep paid a fork per line just to prepare a regex match. Shell parameter expansion does the same strip with no fork, and the retry-dedup scan now reuses that one helper instead of carrying a second copy of the rule in awk. The copies had already drifted: the shell side stripped tags from lines with no colon, which the awk rule left whole, so a colonless line could be mistaken for one already recorded. One definition, checked against the awk rule it replaces over the edge cases and a 4000-line fuzz. tests/fm-contributions.test.sh froze its fixture clock only in exhaust mode. In hang mode the poll set DEADLINE to the real now plus a one-second budget, and when the second ticked before the first forge call the loop broke without ever calling gh: forge/calls was never written and the assertion failed reading a missing file. Freezing the clock in both modes removes the dependence on wall time; the bounded call is still cut by the real timeout, so the observation the test asserts still starts. Emission time stays optional on new status records, and legacy or malformed lines keep an unknown age. * no-mistakes(review): Stamp ask-user escalation line and fix Kimi status assertion * no-mistakes(document): Drop stale unstamped done-line spelling from watcher docs * test: fold emission-time snapshot coverage into the fixture case Drop the incidental ci.yml 18-to-19 count hunk so the PR no longer touches workflows. Keep every emission-time assertion by folding it into test_fixture_snapshot_json. * no-mistakes(review): replace brief date substitution with epoch placeholder; drop emitted_at_epoch * no-mistakes(review): align untimed normalizer with epoch parser; tolerate placeholder stamp in PR scrape * no-mistakes(review): strip undelimited at-tags; correct brief stamp header * no-mistakes(review): normalize stamps at both captain-regex sites; restore mtime freshness * no-mistakes(review): strip colon-bearing stamps for relevance; fix headers and test oracles * no-mistakes(review): narrow escalation match to stamp tolerance; pin note verb * no-mistakes(review): read note and key past colon-bearing stamps * test(status): keep inactive reconcile assertions stamp-tolerant These two oracles were made stamp-tolerant while resolving one of the branch's merges from main. The rebase drops merge commits, so that adaptation was lost and both assertions went back to matching an exact substring that a stamped line no longer contains: the tag lands before the colon, so "failed [key=k]: ..." is now "failed [key=k] [at=N]: ...". Strip a well-formed tag before matching, as the branch's other oracles do. Co-Authored-By: Claude Opus 5 (1M context) * no-mistakes(review): unstamp fold colon tests; reserve stamp width in cap * no-mistakes(document): correct stale unstamped status-line spellings in docs * no-mistakes(document): quote brief-test literals for lint; correct stamp-helper contract comments * no-mistakes(ci): rename subshell-local epoch in delivery-race stub The serialization test overrides fm_pending_reply_mark_delivered inside a (..) subshell. Its `epoch` local collided with the same name in status_line_at_epoch/status_stamp_line, which this branch added and this suite now calls at top level, so ShellCheck 0.11.0 reported SC2030 and failed Lint 2. The stub already prefixes its other locals with `pending_` for the same reason; `epoch` was the leftover. Co-Authored-By: Claude Opus 5 (1M context) --------- Co-authored-by: Claude Opus 5 (1M context) --- AGENTS.md | 4 +- bin/fm-branch-prompt.sh | 2 +- bin/fm-brief.sh | 39 +-- bin/fm-classify-lib.sh | 187 +++++++++++++-- bin/fm-dod-lib.sh | 10 +- bin/fm-fleet-snapshot.sh | 35 ++- bin/fm-inactive-reconcile.sh | 10 +- bin/fm-merge-outcome-lib.sh | 5 +- bin/fm-parent-channel-lib.sh | 16 +- bin/fm-pending-reply-lib.sh | 9 +- bin/fm-procevent-remote-reply.sh | 20 +- bin/fm-secondmate-report.sh | 7 +- bin/fm-send.sh | 16 +- bin/fm-spawn.sh | 8 +- bin/fm-wake-lib.sh | 14 +- bin/fm-watch.sh | 2 +- docs/architecture.md | 2 +- docs/captain-hold-lifecycle.md | 2 +- docs/configuration.md | 3 +- docs/secondmate-parent-channel.md | 4 +- .../verification/secondmate-parent-channel.md | 2 + tests/fm-agy-harness.test.sh | 2 +- tests/fm-bearings-snapshot.test.sh | 8 +- tests/fm-branch-supervision.test.sh | 2 +- tests/fm-brief.test.sh | 61 ++++- tests/fm-captain-hold-lifecycle.test.sh | 26 +- tests/fm-classify-corr-token.test.sh | 224 ++++++++++++++++++ .../fm-cmux-claude-composer-live-e2e.test.sh | 25 +- tests/fm-contributions.test.sh | 4 +- tests/fm-fleet-snapshot-view.test.sh | 60 ++++- tests/fm-inactive-reconcile.test.sh | 53 +++-- tests/fm-kimi-harness.test.sh | 4 +- tests/fm-pending-reply.test.sh | 56 ++++- tests/fm-pr-check-security.test.sh | 6 +- tests/fm-pr-merge.test.sh | 8 +- tests/fm-remote-backlog-handoff.test.sh | 4 +- tests/fm-remote-reply.test.sh | 38 ++- tests/fm-rovo-harness.test.sh | 17 +- tests/fm-send-remote-delivery.test.sh | 2 +- tests/fm-send-resolve-key.test.sh | 47 +++- tests/fm-tangle-guard.test.sh | 3 +- tests/fm-task-delivery.test.sh | 3 +- tests/fm-teardown.test.sh | 6 +- tests/fm-wake-queue.test.sh | 4 +- 44 files changed, 861 insertions(+), 199 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 619beea3103..29ced552794 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -98,7 +98,7 @@ data/ personal fleet records; LOCAL, gitignored as a whole /report.md scout task deliverable, written by the crewmate; survives teardown projects/ cloned repos; gitignored; read-only except under hard rule 1's concrete captain-approved project operation exception state/ runtime records and signals; gitignored - .status appended by crewmates: ": " wake-event lines, not current-state truth + .status append-only wake events, not current-state truth; bin/fm-classify-lib.sh owns their syntax .turn-ended touched by turn-end hooks .progress touched for observed native-harness activity inside one Pi turn; bin/fm-busy-event.sh owns its generation binding and bin/fm-watch.sh reads it beside turn-ended for the busy-age bound only, never as a completed turn .busy-state .busy-gen semantic busy-state record (one line, atomically replaced) and its per-incarnation gen sidecar; bin/fm-busy-event.sh is the only writer and bin/fm-busy-lib.sh owns the record format and classification; arming again replaces the previous incarnation so late events carrying its gen are rejected as stale; removed by retire and teardown @@ -389,7 +389,7 @@ The worker reports the PR when CI first becomes green rather than waiting for me ### PR ready, landing, and teardown -For PR-based ship tasks, the ready signal depends on mode: `no-mistakes` reports `done: PR checks green` after CI is green, while `direct-PR` reports `done: PR ` after opening the PR. +For PR-based ship tasks, the ready signal depends on mode: `no-mistakes` reports `done [at=]: PR checks green` after CI is green, while `direct-PR` reports `done [at=]: PR ` after opening the PR. Run `bin/fm-pr-check.sh ` with the URL copied from that ready signal - it records `pr=` and the forge's `pr_head=` when available in the task's meta and arms the watcher's merge poll. Tell the captain the PR's full `https://...` URL copied from the worker's ready line or the task's `pr=` metadata, a concise outcome summary, and the no-mistakes risk level when applicable. A captain instruction to merge is explicit authority; `yolo` is the only standing routine merge authority. diff --git a/bin/fm-branch-prompt.sh b/bin/fm-branch-prompt.sh index 4bc5d883e4b..acf213ccff2 100755 --- a/bin/fm-branch-prompt.sh +++ b/bin/fm-branch-prompt.sh @@ -78,7 +78,7 @@ Write summaries in the captain's outcome language - the project, the fix, the PR # PR identity: copy or abstain -A PR URL you pass to a tool or write into a summary is copied verbatim from the task's `done: PR ` status line or its `pr=` metadata field. +A PR URL you pass to a tool or write into a summary is copied verbatim from the task's `done [at=]: PR ` status line or its `pr=` metadata field. Never assemble an owner, repository, host, or number from memory, from another PR, or from a bare number the worker printed; a plausible URL built that way is how a dead link reaches the captain. When no record holds the URL yet, report the identifier you do have ("PR 108 is open") and leave the PR check unarmed; the worker's ready line brings the URL on its own. diff --git a/bin/fm-brief.sh b/bin/fm-brief.sh index 264126f6d99..75d6717442c 100755 --- a/bin/fm-brief.sh +++ b/bin/fm-brief.sh @@ -60,6 +60,10 @@ # declared-external-wait verb (FM_CLASSIFY_PAUSED_VERB, default "paused") from # "blocked:": pause for a known external wait expected to clear on its own, # blocked when firstmate must act. +# Emission-time syntax and legacy unknown-time handling are owned by +# bin/fm-classify-lib.sh; each scaffold renders the stamp as a literal +# placeholder the worker replaces with a numeric Unix time as it appends, so a +# scaffold never emits a substitution a file-write tool would copy through. # Every scaffold also carries the steering-inbox receive-and-ack section: # process state/.inbox/*.msg in order and acknowledge each by moving it to # handled/ (record, doorbell, and ladder owned by bin/fm-task-inbox-lib.sh). @@ -284,8 +288,9 @@ $INBOX_SECTION # Escalation to main firstmate Handle routine work yourself. Report only true captain-relevant outcomes or a declared external wait by appending one line: - \`echo "{state}: {one short line}" >> $STATUS_FILE\` + \`echo "{state} [at=]: {one short line}" >> $STATUS_FILE\` States: working, needs-decision, blocked, $PAUSED_VERB, done, failed. +Substitute \`\` with the current Unix time in seconds - run \`date +%s\` and write the number it printed; a stamp that is not plain digits records no time at all. Use \`$PAUSED_VERB: {why}\` (distinct from \`blocked:\`) only when your domain is deliberately idling on a known external wait you expect to clear on its own, naming when it clears with \`until \` (UTC) when you know; use \`blocked:\` when you are stuck and need firstmate to act. Use this only for material phase changes, a captain decision, a real blocker, a failure, work ready for review, or work you landed. Work you landed includes a merge you performed yourself under standing merge authority and one the captain merged on the forge: under that authority nothing is ever \"ready for review\", so a landed merge that goes unreported reaches the captain as silence. @@ -294,9 +299,9 @@ A marked request requires one correlated answer after the work; it does not requ Never append \`working:\` merely to acknowledge receipt or announce that a marked request has started. When a routed-work phase has a supervisor-actionable material change worth reporting under the rule above, give that reported phase a stable key. If its first reportable event is \`working [key=]: {material phase}\`, use the same key on its later \`$PAUSED_VERB\`, \`done\`, \`failed\`, \`needs-decision\`, or \`blocked\` event so the earlier working phase is superseded. -When a keyed phase ends without another reportable state, append \`resolved [key=]: {why it is no longer active}\`. +When a keyed phase ends without another reportable state, append \`resolved [key=] [at=]: {why it is no longer active}\`. \`resolved\` separately closes an escalated decision or blocker, and only a \`resolved\` line carrying that decision's exact key closes it: a later \`done\` or \`working\` event never does, even when the answer is what started that work. -The main firstmate's answer normally writes that closing line at answer time; when a blocker or wait clears WITHOUT an answer from the main firstmate, append \`resolved: {how it cleared}\` yourself (keyed with \`[key=]\` if you opened it with one) as your domain resumes. +The main firstmate's answer normally writes that closing line at answer time; when a blocker or wait clears WITHOUT an answer from the main firstmate, append \`resolved [at=]: {how it cleared}\` yourself (keyed with \`[key=]\` if you opened it with one) as your domain resumes. Routine internal supervision, heartbeats, retries, and crewmate churn stay inside your own home and must not touch that status file. # Definition of done @@ -304,7 +309,7 @@ You are persistent by default. Do not exit just because your queue is empty. On startup and restart, run normal firstmate bootstrap and recovery through \`bin/fm-session-start.sh\` for your own home, but only to RECONCILE work that is already yours: in-flight crewmates, tracked backlog items, and durable watches recorded in this home. When you have no assigned or in-flight work after that reconciliation, go idle and wait silently for the main firstmate to route you a task. An empty queue is a healthy resting state, not a cue to invent work: never spawn a survey, audit, or any self-directed "find work" task on your own initiative. -If this charter cannot be carried out, append \`blocked: {why}\` or \`failed: {why}\` to the main status file and stop. +If this charter cannot be carried out, append \`blocked [at=]: {why}\` or \`failed [at=]: {why}\` to the main status file and stop. EOF if [ "$SECONDMATE_CHARTER" = "{TASK}" ]; then echo "scaffolded: $BRIEF (secondmate charter; replace {TASK})" @@ -382,8 +387,9 @@ The report is the only thing that survives, so anything worth keeping must be in 2. Stay inside this worktree; the only files you may write outside it are the report and the status file below. 3. Use gh-axi for GitHub operations and chrome-devtools-axi for browser operations. 4. Report status by appending one line: - \`echo "{state}: {one short line}" >> $STATUS_FILE\` + \`echo "{state} [at=]: {one short line}" >> $STATUS_FILE\` States: working, needs-decision, blocked, $PAUSED_VERB, done, failed. + Substitute \`\` with the current Unix time in seconds - run \`date +%s\` and write the number it printed; a stamp that is not plain digits records no time at all. Each append wakes firstmate, so report sparingly: only phase changes a supervisor would act on and the needs-decision/blocked/paused/done/failed states. No step-by-step FYI progress lines; firstmate reads your pane for that. @@ -396,17 +402,17 @@ The report is the only thing that survives, so anything worth keeping must be in treating it as a possible wedge. When you know when the wait clears, say so in the line with \`until \` (UTC) and firstmate rechecks at that time instead. Use \`blocked:\` when you are stuck and need help. -5. If you hit the same obstacle twice, append \`blocked: {why}\` and stop; firstmate will help. +5. If you hit the same obstacle twice, append \`blocked [at=]: {why}\` and stop; firstmate will help. 6. If a decision belongs to a human (product choices, destructive actions), - append \`needs-decision: {summary of options}\` and stop. Firstmate will reply with the decision. + append \`needs-decision [at=]: {summary of options}\` and stop. Firstmate will reply with the decision. A decision or blocker you opened stays open until a \`resolved\` line carrying its exact key lands; a later \`done:\` or \`working:\` line never closes it, even when the answer is what started that work. - Firstmate's reply normally writes that closing line at answer time; when a blocker or wait clears WITHOUT a firstmate reply, append \`resolved: {how it cleared}\` yourself (same \`[key=]\` if you opened it with one) as you resume. + Firstmate's reply normally writes that closing line at answer time; when a blocker or wait clears WITHOUT a firstmate reply, append \`resolved [at=]: {how it cleared}\` yourself (same \`[key=]\` if you opened it with one) as you resume. 7. Never stop, restart, or update the shared \`no-mistakes\` daemon - it is one instance serving every lane/home, so restarting it kills other lanes' in-flight pipeline runs; only firstmate manages the daemon. Before you append \`blocked:\` about the pipeline, run \`no-mistakes daemon status\` and \`no-mistakes axi status\`. If the daemon socket refuses connections or is missing, append - \`blocked: {the daemon error}\` and stop even when the local run record still says running or + \`blocked [at=]: {the daemon error}\` and stop even when the local run record still says running or fixing, because that record can be stale after the daemon exits. A run record failed with a daemon error is also a real block. Only after ruling out socket refusal, if the run is still running or fixing, reattach and keep @@ -421,7 +427,7 @@ Write your findings to \`$DATA/$ID/report.md\`. The report must stand alone: what you did, what you found, the evidence (commands run, output, file:line references), and what you recommend. $LAVISH_LINE Before reporting done, read and follow \`$FM_ROOT/.agents/skills/captain-hold-lifecycle/SKILL.md\` and pass its shared completion gate for the report and any visual review. -When the report is complete, append \`done: {one-line conclusion}\` to the status file and stop. +When the report is complete, append \`done [at=]: {one-line conclusion}\` to the status file and stop. If your findings reveal work that should ship (e.g. you reproduced a bug and the fix is clear), say so in the report; firstmate may promote this task in place, and you would then receive mode-specific ship instructions as a follow-up message. EOF echo "scaffolded: $BRIEF (scout; replace {TASK} and {FIRSTMATE_SPEC})" @@ -460,7 +466,7 @@ You are in a disposable git worktree of $REPO, at a detached HEAD on a clean def **Verify isolation before anything else.** Run \`pwd -P\` and \`git rev-parse --show-toplevel\`; both must resolve to the disposable task worktree you were launched in, such as a treehouse pool path or an Orca-managed worktree, not the primary checkout firstmate operates from. The path check is authoritative: \`git rev-parse --git-dir\` and \`git rev-parse --git-common-dir\` can help inspect the repo, but they do not prove you are outside the primary checkout. -If the top-level path is the primary checkout or not the worktree you were launched in, STOP - do not branch or commit here - append \`blocked: launched in primary checkout, not an isolated worktree\` to the status file and stop. +If the top-level path is the primary checkout or not the worktree you were launched in, STOP - do not branch or commit here - append \`blocked [at=]: launched in primary checkout, not an isolated worktree\` to the status file and stop. 1. First action: create your branch: \`git checkout -b fm/$ID\`$SETUP2 @@ -469,8 +475,9 @@ $RULE1 2. Stay inside this worktree; modify nothing outside it. 3. Use gh-axi for GitHub operations and chrome-devtools-axi for browser operations. 4. Report status by appending one line: - \`echo "{state}: {one short line}" >> $STATUS_FILE\` + \`echo "{state} [at=]: {one short line}" >> $STATUS_FILE\` States: working, needs-decision, blocked, $PAUSED_VERB, done, failed. + Substitute \`\` with the current Unix time in seconds - run \`date +%s\` and write the number it printed; a stamp that is not plain digits records no time at all. Each append wakes firstmate, so report sparingly: only phase changes a supervisor would act on (setup done, bug reproduced, fix implemented, validation passed) and the needs-decision/blocked/paused/done/failed states. No step-by-step FYI progress lines; @@ -484,18 +491,18 @@ $RULE1 known external wait you expect to clear on its own ($CREWMATE_PAUSE_WAIT_EXAMPLES): firstmate then leaves your idle pane alone and rechecks it on a long cadence instead of treating it as a possible wedge. Use \`blocked:\` when you are stuck and need help. -5. If you hit the same obstacle twice, append \`blocked: {why}\` and stop; firstmate will help. +5. If you hit the same obstacle twice, append \`blocked [at=]: {why}\` and stop; firstmate will help. 6. If a decision belongs above the implementation worker (product choices, destructive actions), - append \`needs-decision: {summary of options}\` and stop. Firstmate will reply with the decision. + append \`needs-decision [at=]: {summary of options}\` and stop. Firstmate will reply with the decision. $ASK_USER_BLOCK A decision or blocker you opened stays open until a \`resolved\` line carrying its exact key lands; a later \`done:\` or \`working:\` line never closes it, even when the answer is what started that work. - Firstmate's reply normally writes that closing line at answer time; when a blocker or wait clears WITHOUT a firstmate reply, append \`resolved: {how it cleared}\` yourself (same \`[key=]\` if you opened it with one) as you resume. + Firstmate's reply normally writes that closing line at answer time; when a blocker or wait clears WITHOUT a firstmate reply, append \`resolved [at=]: {how it cleared}\` yourself (same \`[key=]\` if you opened it with one) as you resume. 7. Never stop, restart, or update the shared \`no-mistakes\` daemon - it is one instance serving every lane/home, so restarting it kills other lanes' in-flight pipeline runs; only firstmate manages the daemon. Before you append \`blocked:\` about the pipeline, run \`no-mistakes daemon status\` and \`no-mistakes axi status\`. If the daemon socket refuses connections or is missing, append - \`blocked: {the daemon error}\` and stop even when the local run record still says running or + \`blocked [at=]: {the daemon error}\` and stop even when the local run record still says running or fixing, because that record can be stale after the daemon exits. A run record failed with a daemon error is also a real block. Only after ruling out socket refusal, if the run is still running or fixing, reattach and keep diff --git a/bin/fm-classify-lib.sh b/bin/fm-classify-lib.sh index f737a7d96dd..cc56ed3e06a 100755 --- a/bin/fm-classify-lib.sh +++ b/bin/fm-classify-lib.sh @@ -160,7 +160,7 @@ last_status_line() { # [] # A bare legacy free-text line counts as an event only when a captain token leads # it, so continuation prose that merely mentions one cannot hide a declaration. _fm_status_event_scan() { - local line last='' prev='' fallback='' verb legacy_re + local line last='' prev='' fallback='' verb legacy_re unstamped legacy_re="^[[:space:]]*(${FM_CAPTAIN_RE:-$FM_CLASSIFY_CAPTAIN_RE_DEFAULT})" while IFS= read -r line || [ -n "$line" ]; do case "$line" in *[![:space:]]*) fallback=$line ;; *) continue ;; esac @@ -170,7 +170,8 @@ _fm_status_event_scan() { "${FM_CLASSIFY_PAUSED_VERB:-$FM_CLASSIFY_PAUSED_VERB_DEFAULT}"|\ "${FM_CLASSIFY_RESOLVE_VERB:-$FM_CLASSIFY_RESOLVE_VERB_DEFAULT}"|\ "${FM_CLASSIFY_CAPTAIN_HELD_VERB:-$FM_CLASSIFY_CAPTAIN_HELD_VERB_DEFAULT}") prev=$last; last=$line ;; - *) _fm_classify_matches "$line" "$legacy_re" && { prev=$last; last=$line; } ;; + *) _fm_status_unstamped "$line" unstamped + _fm_classify_matches "$unstamped" "$legacy_re" && { prev=$last; last=$line; } ;; esac done printf '%s\n%s\n' "$prev" "${last:-$fallback}" @@ -205,8 +206,12 @@ status_is_terminal_verb() { # (working, resolved, captain-held) and paused never match from free-text prose; # only lines without those leading verbs may still match free-text tokens for # legacy bare lines such as "merged" or "PR ready". +# Regex matching ignores any emission-time tag before the first colon - here and +# in the shared event scan, the module's two FM_CAPTAIN_RE sites - so an override +# keeps matching a stamped event however the worker spelled the stamp; other +# metadata and note text remain intact, as do the stored and surfaced event bytes. status_is_captain_relevant() { - local line=$1 verb + local line=$1 verb unstamped [ -n "$line" ] || return 1 status_line_verb "$line" verb case "$verb" in @@ -219,7 +224,8 @@ status_is_captain_relevant() { done|needs-decision|blocked|failed) return 0 ;; esac fi - _fm_classify_matches "$line" "${FM_CAPTAIN_RE:-$FM_CLASSIFY_CAPTAIN_RE_DEFAULT}" + _fm_status_unstamped "$line" unstamped + _fm_classify_matches "$unstamped" "${FM_CAPTAIN_RE:-$FM_CLASSIFY_CAPTAIN_RE_DEFAULT}" } # 0 if a status line's leading verb is the pause verb (paused: ). A pure @@ -275,6 +281,135 @@ status_paused_until() { # -> epoch on stdout fm_utc_iso_to_epoch "$token" } +# --- optional event emission time ------------------------------------------- +# New writers may append "[at=]" before the first colon, alongside key +# and corr tags in any order. Epoch is UTC Unix seconds: canonical unsigned +# decimal, at most 12 digits (bounded for safe shell arithmetic). For example: +# resolved [key=api-shape] [at=1788576000]: answered: use REST +# No colons appear inside this field, so existing verb/key/note readers retain +# their grammar. Missing, malformed, or duplicate time fields mean UNKNOWN time; +# never infer emission time from file mtime, a wake, or observation time. Relays +# preserve source tags and leave legacy source events unstamped. Time describes +# event history only and must never decide current state or decision closure. +# This parser owns that grammar; every reader below is a thin adapter over it, +# so no second spelling of "well-formed" can drift against this one. +# Internals carry a reserved prefix: bash locals are dynamically scoped, so a +# plain name here would shadow the caller's out-var of the same name. +_fm_status_at_epoch() { # -> 0 and the epoch when known + local __fm_at_head __fm_at_value __fm_at_rest + printf -v "$2" '%s' '' + case "$1" in *:*) __fm_at_head=${1%%:*} ;; *) return 1 ;; esac + case "$__fm_at_head" in *\[at=*\]*) ;; *) return 1 ;; esac + __fm_at_rest=${__fm_at_head#*\[at=} + __fm_at_value=${__fm_at_rest%%\]*} + case "${__fm_at_rest#*\]}" in *\[at=*) return 1 ;; esac + case "$__fm_at_value" in ''|*[!0-9]*|0[0-9]*) return 1 ;; esac + [ "${#__fm_at_value}" -le 12 ] || return 1 + printf -v "$2" '%s' "$__fm_at_value" +} + +status_line_at_epoch() { # -> epoch; nonzero when unknown + local epoch + _fm_status_at_epoch "$1" epoch || return 1 + printf '%s' "$epoch" +} + +# Stamp only a newly emitted event. Preserve an existing tag, even malformed, +# and preserve the event itself if the clock cannot be read. Never use this to +# timestamp a copied historical line. +status_stamp_line() { # -> line (without newline) + local head epoch + case "$1" in + *:*) head=${1%%:*} ;; + *) printf '%s' "$1"; return 0 ;; + esac + case "$head" in *\[at=*) printf '%s' "$1"; return 0 ;; esac + if epoch=$(date +%s); then + printf '%s [at=%s]:%s' "$head" "$epoch" "${1#*:}" + else + printf '%s' "$1" + fi +} + +# Characters status_stamp_line would insert into a line it stamps: the space, +# the "[at=" and "]" delimiters, and the clock's own digit width. A writer that +# caps a status line BEFORE the append stamps it must subtract this from its +# cap, or the bytes actually appended overrun the cap that writer enforces and +# every capped rendering downstream loses that much real note text. Zero when +# the clock cannot be read, because then nothing is stamped either. +status_stamp_width() { # -> characters a stamp adds to a line + local epoch tag + epoch=$(date +%s) || { printf 0; return 0; } + case "$epoch" in ''|*[!0-9]*) printf 0; return 0 ;; esac + tag=" [at=$epoch]" + printf '%s' "${#tag}" +} + +# Strip the one well-formed time tag _fm_status_at_epoch accepts, for readers +# that need a stamped line as the exact bytes it carried before stamping: +# retry-dedup identity here, and the pending-reply escalation match in +# bin/fm-pending-reply-lib.sh, which compares against its own literal spellings. +# Every other [at=...] byte run - malformed, duplicate, or outside the canonical +# bounds - is ordinary line bytes here, never a time tag, so a retry of it stays +# a distinct event. A reader that instead asks where the HEAD ends owns a more +# tolerant rule in _fm_status_unstamped below and must route through that one; +# do not route such a reader through this one. It reads the grammar from that +# single parser rather than a second spelling of it, and a sweep that normalizes +# a line at a time never pays a fork for the match it prepares. +_fm_status_untimed() { # -> line without a time tag + local __fm_untimed_epoch __fm_untimed_head __fm_untimed_tag __fm_untimed_before + if _fm_status_at_epoch "$1" __fm_untimed_epoch; then + __fm_untimed_head=${1%%:*} + __fm_untimed_tag="[at=$__fm_untimed_epoch]" + __fm_untimed_before=${__fm_untimed_head%%"$__fm_untimed_tag"*} + printf -v "$2" '%s%s:%s' "${__fm_untimed_before% }" \ + "${__fm_untimed_head#*"$__fm_untimed_tag"}" "${1#*:}" + return 0 + fi + printf -v "$2" '%s' "$1" +} + +# Strip every time-tag-shaped run a worker could have written as the stamp, +# however malformed its value. This is the shared head-boundary rule for every +# reader that asks where a line's head ends rather than what its stamp means: +# captain-relevance, the event scan, and the note, key, and decision-fold +# readers. A tag is metadata a worker appended, so it must never decide whether +# a terminal event reaches its supervisor, which note or key that event carries, +# or whether a decision opens or closes - not when the worker left the brief's +# placeholder unsubstituted, and not when they wrote a readable time +# whose colons swallow the head/note separator. +# A run is the stamp only while nothing before it holds a colon; once one does, +# the head has ended and every later [at=...] is note text the override may +# legitimately match on, so scanning stops there. The caller's own bytes are +# untouched: this writes a throwaway copy used for matching only. +_fm_status_unstamped() { # -> line with its stamp removed + local __fm_unstamped_rest=$1 __fm_unstamped_keep='' __fm_unstamped_before + while :; do + case "$__fm_unstamped_rest" in *\[at=*\]*) ;; *) break ;; esac + __fm_unstamped_before=${__fm_unstamped_rest%%\[at=*} + case "$__fm_unstamped_before" in *:*) break ;; esac + __fm_unstamped_keep=$__fm_unstamped_keep${__fm_unstamped_before% } + __fm_unstamped_rest=${__fm_unstamped_rest#*\[at=} + __fm_unstamped_rest=${__fm_unstamped_rest#*\]} + done + printf -v "$2" '%s' "$__fm_unstamped_keep$__fm_unstamped_rest" +} + +# Retry deduplication ignores only a well-formed optional numeric time tag; +# all other bytes, including correlation metadata, still identify the event. +# Both sides normalize through _fm_status_untimed, so a stamped retry of an +# already-recorded event can never read as a new one. +status_event_recorded() { # + local wanted line untimed + [ -f "$1" ] || return 1 + _fm_status_untimed "$2" wanted + while IFS= read -r line || [ -n "$line" ]; do + _fm_status_untimed "$line" untimed + [ "$untimed" != "$wanted" ] || return 0 + done < "$1" + return 1 +} + # --- durable keyed decisions ------------------------------------------------ # # The status stream is an append-only EVENT log. Reading it last-event-wins @@ -428,16 +563,21 @@ _fm_decision_slug_ok() { # *) return 0 ;; esac } +# Both readers below locate the head/note separator on an unstamped copy, so a +# worker-written stamp cannot move it: a readable time like [at=10:30] carries +# colons that would otherwise end the head mid-tag and hand the caller a note +# and a key sliced out of the timestamp. The line's own bytes are never altered. status_line_note() { # -> text after the first colon, trimmed - local n k - case "$1" in - *:*) n=${1#*:}; n=${n#"${n%%[![:space:]]*}"} ;; - *) printf '%s' "$1"; return 0 ;; + local n k unstamped + _fm_status_unstamped "$1" unstamped + case "$unstamped" in + *:*) n=${unstamped#*:}; n=${n#"${n%%[![:space:]]*}"} ;; + *) printf '%s' "$unstamped"; return 0 ;; esac # A note-head token that states this line's key (no before-colon token, valid # slug) is key metadata, not note text: strip it so both stated-key positions # yield the same note. - if ! _fm_key_before_colon "$1" && k=$(_fm_key_at_note_head "$1") \ + if ! _fm_key_before_colon "$unstamped" && k=$(_fm_key_at_note_head "$unstamped") \ && _fm_decision_slug_ok "$k"; then n=${n#"[key=$k]"} n=${n#"${n%%[![:space:]]*}"} @@ -445,13 +585,14 @@ status_line_note() { # -> text after the first colon, trimmed printf '%s' "$n" } _fm_decision_key() { # -> key slug, or "default" when no token - local k - if _fm_key_before_colon "$1"; then - k=${1%%:*} + local k unstamped + _fm_status_unstamped "$1" unstamped + if _fm_key_before_colon "$unstamped"; then + k=${unstamped%%:*} k=${k#*\[key=} k=${k%%\]*} else - k=$(_fm_key_at_note_head "$1") || { printf 'default'; return 0; } + k=$(_fm_key_at_note_head "$unstamped") || { printf 'default'; return 0; } fi _fm_decision_slug_ok "$k" || return 1 printf '%s' "$k" @@ -535,7 +676,15 @@ _fm_status_kind() { } _fm_decision_fold_line() { # - local open=$1 line=$2 resolve=$3 held=$4 kind=$5 verb key note + local open=$1 line=$2 resolve=$3 held=$4 kind=$5 verb key note unstamped + # Both colon tests below ask where the head ends, the same question the note + # and key readers ask, so they read the same unstamped copy those readers do. + # A worker-written time tag must never decide whether a decision opens or + # closes: a readable [at=10:30] carries colons that would otherwise make bare + # prose look like a transition, or make a keyless line open a phantom + # decision no later line could close. The stored and surfaced bytes stay the + # caller's own. + _fm_status_unstamped "$line" unstamped # Declaration guard. A transition's verb ends at a colon, or - in the colonless # form _fm_decision_key still accepts below - at a complete "[key=...]" token. # A line holding neither is continuation prose, a bare word, or blank, and can @@ -543,12 +692,12 @@ _fm_decision_fold_line() { # # 8: a colonless line without a complete "[key=...]" token is no longer a # transition at all, so a cursor holding a phantom decision that bare prose # opened - which no later line could close - is discarded. +# 9: the two colon tests read the line with its time tag stripped, so a +# malformed worker stamp whose colons used to pose as the head/note separator +# no longer opens or closes anything; cursors folded under that reading are +# discarded. # Version 4 was already spent on the bracketed-tag parser change above, and a # cursor persisted under that reading predates this one, so it must still be # discarded and rebuilt from byte 0 under the new reading. -FM_OPEN_DECISIONS_FOLD_VERSION=8 +FM_OPEN_DECISIONS_FOLD_VERSION=9 # Portable device:inode identity for the rotation/recreation check below. _fm_open_decisions_file_ident() { # -> strongest available identity diff --git a/bin/fm-dod-lib.sh b/bin/fm-dod-lib.sh index b58a5a00d23..db70a186f88 100755 --- a/bin/fm-dod-lib.sh +++ b/bin/fm-dod-lib.sh @@ -235,7 +235,7 @@ fm_ask_user_escalation_block() { # local data=$1 id=$2 cat <-findings.txt\`, then report the gate with - \`needs-decision [key=nm--]: ask-user findings=,,... file=$data/$id/nm--findings.txt\` + \`needs-decision [at=] [key=nm--]: ask-user findings=,,... file=$data/$id/nm--findings.txt\` naming every ask-user finding id from that gate. The status line only points at the file; it never restates or summarizes a finding's content. EOF } @@ -249,7 +249,7 @@ fm_dod_block() { # Delivery contract: mode=direct-PR This task ships **direct-PR**: you raise the PR yourself, without the no-mistakes pipeline. The task is complete only when committed on your branch. -When it is implemented and committed, push your branch and open a PR with \`gh-axi\`, then append \`done: PR {url}\` to the status file and stop. +When it is implemented and committed, push your branch and open a PR with \`gh-axi\`, then append \`done [at=]: PR {url}\` to the status file and stop. Do NOT run /no-mistakes. The configured merge authority decides whether to merge the PR; firstmate relays the outcome. EOF ;; @@ -260,7 +260,7 @@ Delivery contract: mode=local-only This task ships **local-only**: no remote, no PR, no pipeline. The task is complete only when committed on your branch \`fm/$id\`. Do NOT push, do NOT open a PR, do NOT merge. Keep your branch a clean fast-forward onto the current default branch - if \`main\` has advanced, rebase onto it so the eventual merge stays a fast-forward. -When it is implemented and committed, append \`done: ready in branch fm/$id\` to the status file and stop. +When it is implemented and committed, append \`done [at=]: ready in branch fm/$id\` to the status file and stop. The configured merge authority approves the ready branch, then firstmate merges it into local \`main\` through the guarded fast-forward path. EOF ;; @@ -269,7 +269,7 @@ EOF # Definition of done Delivery contract: mode=no-mistakes The task is complete only when committed on your branch. -When you believe it is complete, append \`done: {summary}\` to the status file and stop. +When you believe it is complete, append \`done [at=]: {summary}\` to the status file and stop. Firstmate will then instruct you to run /no-mistakes to validate and ship a PR. You drive no-mistakes by responding to its gates, not by implementing fixes. @@ -297,7 +297,7 @@ Two firstmate-specific rules layer on top of that guidance: - NEVER pass \`--yes\` (or \`-y\`) to \`no-mistakes axi run\` or \`no-mistakes axi respond\`. It is banned fleet-wide. It auto-resolves every gate including ask-user findings with no escalation, and answering your own ask-user finding is a hard rule violation. -After /no-mistakes reports CI green (the CI-ready return point - do not wait for it to keep monitoring in the background until merge), append \`done: PR {url} checks green\` and stop. You are finished. +After /no-mistakes reports CI green (the CI-ready return point - do not wait for it to keep monitoring in the background until merge), append \`done [at=]: PR {url} checks green\` and stop. You are finished. EOF ;; *) diff --git a/bin/fm-fleet-snapshot.sh b/bin/fm-fleet-snapshot.sh index 296159ce04e..b2273996170 100755 --- a/bin/fm-fleet-snapshot.sh +++ b/bin/fm-fleet-snapshot.sh @@ -56,7 +56,10 @@ # an explicit unknown value because their endpoint liveness belongs to # supervision rather than this snapshot path. # paths.status_log.last_event is historical wake-event data only, never -# current state. +# current state. age_seconds is null when the emission time is unknown; +# fm-classify-lib.sh owns the optional emission-time field, and only the +# age derived from it is published here. A future event time leaves that age +# unknown rather than clamped to zero. # hints.open_decisions is the keyed open-decision set returned by # fm-classify-lib.sh's authoritative status_open_decisions fold and reconciled # against current_state; hints.pending_decision and hints.blocked_event are @@ -77,6 +80,10 @@ # each home with explicit provenance, freshness, endpoint evidence, and unknown # failure reasons. Parent status and bounded terminal evidence are historical, # untrusted supplements only and never override readable structured-home facts. +# parent_event carries age_seconds from the task's paths.status_log.last_event +# above. An unreadable-home fallback reports freshness.age_seconds from the +# observed status file's mtime instead: freshness is how fresh this snapshot's +# own observation is, never when a worker emitted the event. # Each structured-home record carries active_children, decisions_open, holds, # queued, landed, endpoints, counts, and omitted. provenance.summary_source # distinguishes "local-ledger", "remote-ledger", and "remote-ledger-cache"; @@ -348,20 +355,25 @@ crew_state_json() { # [] [] } status_event_json() { # [] - local log=$1 path=${2:-$1} present=0 raw='' verb='' note='' + local log=$1 path=${2:-$1} present=0 raw='' verb='' note='' epoch=null age=null if [ -f "$log" ]; then present=1 raw=$(last_nonempty_line "$log" || true) verb=$(status_line_verb "$raw") note=$(status_line_note "$raw") + epoch=$(status_line_at_epoch "$raw") || epoch=null + if [ "$epoch" != null ] && [ "$epoch" -le "$SNAPSHOT_EPOCH" ]; then + age=$((SNAPSHOT_EPOCH - epoch)) + fi fi jq -n \ --arg path "$path" \ --arg raw "$raw" \ --arg verb "$verb" \ --arg note "$note" \ + --argjson age "$age" \ --argjson present "$(bool_json "$present")" \ - '{path:$path,present:$present,kind:"event_history",last_event:{state:$verb,note:$note,raw:$raw}}' + '{path:$path,present:$present,kind:"event_history",last_event:{state:$verb,note:$note,raw:$raw,age_seconds:$age}}' } first_pr_url_in_file() { # @@ -1701,7 +1713,7 @@ parent_evidence_reconciliation_json() { # secondmate_current_json() { # local tasks_file=$1 output_file=$2 registry_file union_file records_file rows total_registered total shown truncated - local row id home host remote registered registry_error task sampled_spawn_gen status_file status_observation_file event_raw event_note event_epoch event_age + local row id home host remote registered registry_error task sampled_spawn_gen status_file status_observation_file event_raw event_note event_age observed_epoch observed_age local activity_scan activities decisions reconciliation provenance freshness reason summary_file summary_sampled summary_valid summary_invalidity state terminal terminal_contradiction contradiction local summary_source summary_age summary_observed summary_freshness cache_path collection_status collection_slot summary_index=0 local seen_homes='' @@ -1756,11 +1768,12 @@ secondmate_current_json() { # activity_scan=$(bounded_parent_activities_json "$status_observation_file") activities=$(printf '%s' "$activity_scan" | jq -c '.records') decisions=$(printf '%s' "$task" | jq -c '.hints.open_decisions // []') - event_epoch=$(file_mtime_epoch "$status_observation_file") - event_age=null - if [ -n "$event_epoch" ]; then - event_age=$((SNAPSHOT_EPOCH - event_epoch)) - [ "$event_age" -lt 0 ] && event_age=0 + event_age=$(printf '%s' "$task" | jq -r '.paths.status_log.last_event.age_seconds // "null"') + observed_epoch=$(file_mtime_epoch "$status_observation_file") + observed_age=null + if [ -n "$observed_epoch" ]; then + observed_age=$((SNAPSHOT_EPOCH - observed_epoch)) + [ "$observed_age" -lt 0 ] && observed_age=0 fi reason=$registry_error @@ -1893,7 +1906,7 @@ secondmate_current_json() { # --arg id "$id" --arg home "$home" --arg host "$host" --argjson remote "$remote" --arg reason "$reason" --arg observed "$SNAPSHOT_NOW" \ --arg spawn_gen "$sampled_spawn_gen" \ --arg provenance "$provenance" --arg freshness "$freshness" --arg event_raw "$event_raw" --arg event_note "$event_note" \ - --argjson registered "$registered" --argjson event_age "$event_age" --argjson activities "$activities" --argjson activity_scan "$activity_scan" \ + --argjson registered "$registered" --argjson event_age "$event_age" --argjson observed_age "$observed_age" --argjson activities "$activities" --argjson activity_scan "$activity_scan" \ --argjson decisions "$decisions" --argjson terminal "$terminal" --slurpfile summary "$summary_file" --argjson summary_sampled "$summary_sampled" ' ($summary[0]) as $summary | @@ -1902,7 +1915,7 @@ secondmate_current_json() { # current:{state:"unknown",reason:(if $summary_sampled then "structured home state invalid: " + ($summary.reason // "unknown reason") else $reason end)},invalidity:null, reconcile_inventory:(if $summary_sampled then $summary.invalidity else null end), provenance:{selected:$provenance,structured_home:($home | if . == "" then null else . end),parent_event_role:"fallback-only-not-current"}, - freshness:{status:$freshness,observed_at:$observed,age_seconds:$event_age}, + freshness:{status:$freshness,observed_at:$observed,age_seconds:$observed_age}, active_children:[],decisions_open:[],holds:[],queued:[],landed:[],endpoints:[],counts:{active_children:0,decisions_open:0,holds:0,queued:0,landed:0,endpoints:0},omitted:[], parent_event:{raw:$event_raw,note:$event_note,age_seconds:$event_age,open_activities:$activities,open_decisions:$decisions,activity_scan:$activity_scan}, terminal_evidence:$terminal,contradiction:false}' >> "$records_file" || return 1 diff --git a/bin/fm-inactive-reconcile.sh b/bin/fm-inactive-reconcile.sh index 5cbaf9e63d2..dc2308821d5 100755 --- a/bin/fm-inactive-reconcile.sh +++ b/bin/fm-inactive-reconcile.sh @@ -12,7 +12,7 @@ # first runs the LEDGER-FIRST parent delivery: a direct child whose status # ledger ends in a whole `done:` or `failed:` line has stated its own outcome, # so that line is published on the parent channel at once through -# bin/fm-parent-channel-lib.sh as +# bin/fm-parent-channel-lib.sh from this unstamped payload: # [key=child-outcome---]: child : [pr=] [mode=] [yolo=] [report=data//report.md] # carrying the child's recorded PR, delivery mode, merge posture, and scout # report pointer, without consulting fm-crew-state.sh and without waiting for @@ -313,8 +313,10 @@ meta_incarnation() { # # The task's delivered PR. Recorded meta pr= is the only authoritative source; # the fallback scrape accepts only a preferred terminal line in a mode's -# ready-signal shape (`done: PR ` or `done: PR checks green`), so a -# PR a worker merely mentioned in prose is never claimed as the delivery. +# ready-signal shape (`done: PR ` or `done: PR checks green`, +# optionally carrying an emission-time tag this scrape steps over without +# reading), so a PR a worker merely mentioned in prose is never claimed as the +# delivery. # A scout never delivers a PR, so it never carries one. pr_for_task() { # [preferred-line] local meta=$1 preferred=${2:-} value @@ -322,7 +324,7 @@ pr_for_task() { # [preferred-line] value=$(meta_field "$meta" pr) if [ -z "$value" ] && [ -n "$preferred" ]; then value=$(printf '%s\n' "$preferred" \ - | sed -nE 's|^done: PR (https?://[^[:space:])"]+/pull/[0-9]+)( checks green)?$|\1|p' \ + | sed -nE 's|^done( \[at=[^]]*\])?: PR (https?://[^[:space:])"]+/pull/[0-9]+)( checks green)?$|\2|p' \ | head -1 || true) fi clean_field "$value" diff --git a/bin/fm-merge-outcome-lib.sh b/bin/fm-merge-outcome-lib.sh index db279351145..bf11266213b 100755 --- a/bin/fm-merge-outcome-lib.sh +++ b/bin/fm-merge-outcome-lib.sh @@ -9,8 +9,7 @@ # # The destination is the home's role, never the caller's choice: # - a secondmate home reports upward on its parent channel, resolved and -# appended through bin/fm-parent-channel-lib.sh in the same -# " [key=]: " shape the charter contract defines; +# appended through bin/fm-parent-channel-lib.sh under its channel contract; # - a main home reports to the captain through the durable wake queue. # A poll observed in a secondmate home also receives a local durable wake after # the upward write, so the mate can handle its own poll observation. @@ -97,7 +96,7 @@ fm_merge_outcome_report() { # [autho fi if [ -n "$destination" ]; then - fm_parent_channel_append_once "$destination" "$line" || status=1 + fm_parent_channel_append_once "$destination" "$(status_stamp_line "$line")" || status=1 fi if [ "$status" -eq 0 ] && { [ "$origin" = poll ] || [ -z "$destination" ]; }; then fm_wake_append check "merged-$id-$FM_PR_URL" \ diff --git a/bin/fm-parent-channel-lib.sh b/bin/fm-parent-channel-lib.sh index 8b1feccd80c..f44c1eab449 100644 --- a/bin/fm-parent-channel-lib.sh +++ b/bin/fm-parent-channel-lib.sh @@ -40,10 +40,9 @@ # The parent watcher classifies lines there exactly as it classifies any # crewmate's status stream, so a captain-relevant line becomes a parent wake. # -# Lines follow the charter's " [key=]: " shape and are -# appended at most once by exact content, so a retried publication cannot -# duplicate a delivered event. An existing destination must be a regular, -# non-symlinked file; a missing one is created with its directory. +# Line syntax and retry equivalence are owned by fm-classify-lib.sh. +# An existing destination must be a regular, non-symlinked file; a missing one +# is created with its directory. # # Return codes, shared by every entry point that resolves the channel: # 0 resolved, or appended / already present @@ -59,6 +58,8 @@ _FM_PARENT_CHANNEL_LIB_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" # shellcheck source=bin/fm-secondmate-parent-lib.sh . "$_FM_PARENT_CHANNEL_LIB_DIR/fm-secondmate-parent-lib.sh" +# shellcheck source=bin/fm-classify-lib.sh +. "$_FM_PARENT_CHANNEL_LIB_DIR/fm-classify-lib.sh" # shellcheck disable=SC2034 # Output globals read by sourcing callers. FM_PARENT_CHANNEL_ID= @@ -128,7 +129,8 @@ fm_parent_channel_clean_note() { # printf '%s' "$1" | LC_ALL=C tr '\t\r\n' ' ' | cut -c1-1200 } -# Append to unless that exact line is already there. +# Append once, using fm-classify-lib.sh's retry contract. Time-insensitive: +# the caller declaring a new event is the one that stamps it. fm_parent_channel_append_once() { # local path=$1 line=$2 if [ -e "$path" ] || [ -L "$path" ]; then @@ -136,7 +138,7 @@ fm_parent_channel_append_once() { # else mkdir -p "$(dirname "$path")" || return 1 fi - if grep -Fqx -- "$line" "$path" 2>/dev/null; then + if status_event_recorded "$path" "$line"; then return 0 fi printf '%s\n' "$line" >> "$path" @@ -147,5 +149,5 @@ fm_parent_channel_report() { # local home=$1 state=$2 line=$3 destination rc=0 destination=$(fm_parent_channel_destination "$home" "$state") || rc=$? [ "$rc" -eq 0 ] || return "$rc" - fm_parent_channel_append_once "$destination" "$line" || return 4 + fm_parent_channel_append_once "$destination" "$(status_stamp_line "$line")" || return 4 } diff --git a/bin/fm-pending-reply-lib.sh b/bin/fm-pending-reply-lib.sh index 79283ba0941..93456d58717 100755 --- a/bin/fm-pending-reply-lib.sh +++ b/bin/fm-pending-reply-lib.sh @@ -1098,15 +1098,16 @@ fm_pending_reply_escalation_payload() { # # that exact escalation remains open. If an unrelated decision has since taken # over that key, the close is withheld so the unrelated decision is not cleared. fm_pending_reply_escalation_line() { # - local status_file=$1 rec=$2 corr=$3 line found='' kind payload own_key + local status_file=$1 rec=$2 corr=$3 line found='' kind payload own_key untimed [ -f "$status_file" ] || return 0 [ "$(fm_pending_reply_get "$rec" corr_id)" = "$corr" ] || return 0 own_key=$(fm_pending_reply_escalation_key "$corr") while IFS= read -r line || [ -n "$line" ]; do [ "$(status_line_verb "$line")" = blocked ] || continue + _fm_status_untimed "$line" untimed for kind in missed delivery-unknown recovery-delivery; do payload=$(fm_pending_reply_escalation_payload "$rec" "$kind") || continue - case "$line" in + case "$untimed" in "blocked [key=$own_key]: $payload"|"blocked: $payload") found=$line; break ;; "blocked [key=$own_key]: $payload "*|"blocked: $payload "*) found=$line; break ;; esac @@ -1260,8 +1261,8 @@ _fm_pending_reply_maybe_escalate_locked() { # [ -n "$parent_status" ] || return 1 mkdir -p "$(dirname "$parent_status")" 2>/dev/null || return 1 line="blocked [key=$(fm_pending_reply_escalation_key "$corr")]: $payload" - if ! grep -Fqx "$line" "$parent_status" 2>/dev/null; then - printf '%s\n' "$line" >> "$parent_status" 2>/dev/null || return 1 + if ! status_event_recorded "$parent_status" "$line"; then + printf '%s\n' "$(status_stamp_line "$line")" >> "$parent_status" 2>/dev/null || return 1 fi now=$(fm_pending_reply_now) fm_pending_reply_set "$rec" escalated_epoch "$now" || return 1 diff --git a/bin/fm-procevent-remote-reply.sh b/bin/fm-procevent-remote-reply.sh index 222a54c0809..b6615ab79d7 100755 --- a/bin/fm-procevent-remote-reply.sh +++ b/bin/fm-procevent-remote-reply.sh @@ -407,8 +407,10 @@ normalize_payload() { # } # Adapter-authored escalations and notes use exact-byte append suppression. -# Mirrored payload lines use their pre-rewrite source identity in -# stage_mirror_lines instead, because delivery state can change between replays. +# Their callers first apply fm-classify-lib.sh's retry contract and stamp only +# the line they append. Mirrored payload lines keep their source time (or its +# absence) and use their pre-rewrite source identity in stage_mirror_lines +# instead, because delivery state can change between replays. # Returns 0 appended, 1 already present, 2 the write itself failed. append_status_once() { # grep -Fqx -- "$2" "$1" 2>/dev/null && return 1 @@ -528,7 +530,11 @@ cmd_ingest() { if [ "$class" = continuity-broken ]; then line="blocked [key=remote-reply-continuity-$id]: remote reply continuity broke for $id ($reason)" append_rc=0 - append_status_once "$status_file" "$line" || append_rc=$? + if status_event_recorded "$status_file" "$line"; then + append_rc=1 + else + append_status_once "$status_file" "$(status_stamp_line "$line")" || append_rc=$? + fi [ "$append_rc" -ne 2 ] || { fm_lock_release "$lock"; die "cannot append continuity escalation"; } fm_lock_release "$lock" printf 'continuity-broken: %s (%s)\n' "$id" "$reason" @@ -587,9 +593,13 @@ EOF # fold, so it cannot stand open the way a keyed block did. while IFS=$'\t' read -r doc reason || [ -n "$doc" ]; do [ -n "$doc" ] || continue + line="note: remote document did not transfer for $id: $doc - $reason" append_rc=0 - append_status_once "$status_file" "note: remote document did not transfer for $id: $doc - $reason" \ - || append_rc=$? + if status_event_recorded "$status_file" "$line"; then + append_rc=1 + else + append_status_once "$status_file" "$(status_stamp_line "$line")" || append_rc=$? + fi [ "$append_rc" -ne 2 ] || { fm_lock_release "$lock"; die "cannot append remote document note"; } [ "$append_rc" -ne 0 ] || appended=$((appended + 1)) done <> "$DESTINATION" + printf -v line '%s [%s]: %s (%s via-helper)' "$VERB" "$token" "$NOTE" "$DOC_PATH" else - printf '%s [%s]: %s (via-helper)\n' "$VERB" "$token" "$DOC_PATH" >> "$DESTINATION" + printf -v line '%s [%s]: %s (via-helper)' "$VERB" "$token" "$DOC_PATH" fi else NOTE=$* - printf '%s [%s]: %s (via-helper)\n' "$VERB" "$token" "$NOTE" >> "$DESTINATION" + printf -v line '%s [%s]: %s (via-helper)' "$VERB" "$token" "$NOTE" fi +printf '%s\n' "$(status_stamp_line "$line")" >> "$DESTINATION" diff --git a/bin/fm-send.sh b/bin/fm-send.sh index 5e42f354213..e672b963823 100755 --- a/bin/fm-send.sh +++ b/bin/fm-send.sh @@ -156,8 +156,9 @@ # blocked: record in the target task's state/.status. fm-send itself # appends the closing resolved line to that status file, so the captain-facing # OPEN DECISIONS record closes at answer time and never depends on the busy -# worker writing a matching resolved line. Ordinary keys close with -# "resolved [key=]: answered: ". A reserved key +# worker writing a matching resolved line. For ordinary keys the payload is +# "resolved [key=]: answered: " before the emission-time +# handling owned by bin/fm-classify-lib.sh. A reserved key # (pending-reply-* today; bin/fm-classify-lib.sh's reserved-key guard) is # closed with the owning library's vocabulary note # (fm_pending_reply_close_note_for_key / fm_pending_reply_resolved_note), so @@ -559,6 +560,7 @@ RESOLVE_STATUS_FILE= # longer owns also keeps the common path free of any backlog read. RESOLVE_STATUS_KEYS= RESOLVE_HOLD_KEYS= +RESOLVE_CLOSE_MAX=$FM_LINE_CAP_DEFAULT # Resolve a --resolve-key key that the status log no longer owns to the # captain-held task that carries it: the key as a task id itself (the collapsed @@ -661,6 +663,12 @@ if [ -n "$RESOLVE_KEYS" ]; then fi # Refuse before send when a named status-log key cannot actually close: a # reserved key with an answered: note is a silent no-op in the fold. + # The cap bounds the line that is actually APPENDED, and the self-announced + # append stamps each line with its emission time. Reserve that stamp's width + # here so the probe below measures the same bytes the writer will produce and + # the close record stays inside the cap this refusal cites. + RESOLVE_CLOSE_MAX=$((FM_LINE_CAP_DEFAULT - $(status_stamp_width))) + [ "$RESOLVE_CLOSE_MAX" -ge 0 ] || RESOLVE_CLOSE_MAX=0 resolve_excerpt=$(printf '%s' "$*" | tr '\n\r\t' ' ' | LC_ALL=C tr -d '\000-\037\177') for k in $RESOLVE_STATUS_KEYS; do probe=$(fm_send_resolve_close_note "$k" "$resolve_excerpt") @@ -669,7 +677,7 @@ if [ -n "$RESOLVE_KEYS" ]; then exit 1 fi probe_line="resolved [key=$k]: $probe" - fm_cap_line_var "$probe_line" + fm_cap_line_var "$probe_line" "$RESOLVE_CLOSE_MAX" probe_key=$(_fm_decision_key "$FM_LINE_CAP_LINE") || probe_key= if [ "$(status_line_verb "$FM_LINE_CAP_LINE")" != resolved ] || [ "$probe_key" != "$k" ]; then echo "error: --resolve-key cannot close a decision key of length ${#k}: its ${#probe_line}-character close record exceeds the $FM_LINE_CAP_DEFAULT-character status-line cap, and truncation would remove the structural key delimiter. Refusing rather than writing an ineffective close; nothing was sent." >&2 @@ -694,7 +702,7 @@ fm_send_close_resolved_keys() { # note=$(printf '%s' "$note" | tr '\n\r\t' ' ' | LC_ALL=C tr -d '\000-\037\177') for k in $RESOLVE_STATUS_KEYS; do close_note=$(fm_send_resolve_close_note "$k" "$note") - fm_cap_line_var "resolved [key=$k]: $close_note" + fm_cap_line_var "resolved [key=$k]: $close_note" "$RESOLVE_CLOSE_MAX" close_lines+=("$FM_LINE_CAP_LINE") done [ "${#close_lines[@]}" -gt 0 ] || return 0 diff --git a/bin/fm-spawn.sh b/bin/fm-spawn.sh index 8bc3b25b0fd..b1b8608531d 100755 --- a/bin/fm-spawn.sh +++ b/bin/fm-spawn.sh @@ -509,6 +509,8 @@ fi . "$SCRIPT_DIR/fm-ff-lib.sh" # shellcheck source=bin/fm-wake-lib.sh . "$SCRIPT_DIR/fm-wake-lib.sh" +# shellcheck source=bin/fm-classify-lib.sh +. "$SCRIPT_DIR/fm-classify-lib.sh" fm_backlog_directory_present "$STATE" "state directory" || { echo "error: spawn refused: $FM_BACKLOG_TRANSITION_ERROR" >&2 exit 1 @@ -3636,7 +3638,7 @@ kimi_wait_for_delivery() { } kimi_spawn_fail() { # - printf 'failed: %s\n' "$1" >>"$STATE/$ID.status" + printf '%s\n' "$(status_stamp_line "failed: $1")" >>"$STATE/$ID.status" echo "error: $1; inspect window $T" >&2 } @@ -3704,7 +3706,7 @@ rovo_wait_for_delivery() { } rovo_spawn_fail() { # - printf 'failed: %s\n' "$1" >>"$STATE/$ID.status" + printf '%s\n' "$(status_stamp_line "failed: $1")" >>"$STATE/$ID.status" echo "error: $1; inspect window $T" >&2 rovo_endpoint_cleanup } @@ -3777,7 +3779,7 @@ agy_wait_for_working() { } agy_spawn_fail() { # - printf 'failed: %s\n' "$1" >> "$STATE/$ID.status" + printf '%s\n' "$(status_stamp_line "failed: $1")" >>"$STATE/$ID.status" echo "error: $1; inspect window $T" >&2 rovo_endpoint_cleanup } diff --git a/bin/fm-wake-lib.sh b/bin/fm-wake-lib.sh index d59672d9da4..6a7590cafe3 100755 --- a/bin/fm-wake-lib.sh +++ b/bin/fm-wake-lib.sh @@ -2185,23 +2185,31 @@ fm_wake_status_mark_current() { # # normally. # A later, different line from any other writer grows the size past the marker # and wakes as before: task identity alone can never suppress new content. +# Each line is stamped with its emission time on the way in (status_stamp_line, +# bin/fm-classify-lib.sh), so the appended bytes are the stamped ones, not the +# caller's: a caller that caps a line first must reserve status_stamp_width, +# and one that suppresses a repeat must ask status_event_recorded rather than +# compare exact bytes. # Returns 0 appended and self-announced, 1 appended but left for the watcher # (the safe direction), 2 the append itself failed. fm_wake_status_append_self_announced() { # ... local state=$1 file=$2 line appended=0 pre_size='' pre_ident='' post_size post_ident classified folded lag span_rc=0 - local LC_ALL=C + local LC_ALL=C stamped=() shift 2 _fm_wake_require_classify || return 1 + for line in "$@"; do + stamped+=("$(status_stamp_line "$line")") + done if [ -e "$file" ]; then pre_size=$(_fm_status_file_size "$file") || pre_size='' pre_ident=$(_fm_open_decisions_file_ident "$file") || pre_ident='' fi - printf '%s\n' "$@" >> "$file" || return 2 + printf '%s\n' "${stamped[@]}" >> "$file" || return 2 post_size=$(_fm_status_file_size "$file") || return 1 post_ident=$(_fm_open_decisions_file_ident "$file") || return 1 case "$pre_size$post_size" in ''|*[!0-9]*) return 1 ;; esac [ -n "$pre_ident" ] && [ "$post_ident" = "$pre_ident" ] || return 1 - for line in "$@"; do appended=$((appended + ${#line} + 1)); done + for line in "${stamped[@]}"; do appended=$((appended + ${#line} + 1)); done [ "$post_size" -eq $((pre_size + appended)) ] || return 1 classified=$(fm_wake_signal_seen_size "$state" "$file") if [ "$classified" != "$pre_size" ]; then diff --git a/bin/fm-watch.sh b/bin/fm-watch.sh index 8f9ec65fbbb..31cf64aa43f 100755 --- a/bin/fm-watch.sh +++ b/bin/fm-watch.sh @@ -1516,7 +1516,7 @@ pause_state_class() { # # the only record when the worker itself is waiting. It is not the only record # there is: once firstmate hands work to the captain, the wait is written into the # BACKLOG by bin/fm-captain-hold.sh, and the worker's last line stays whatever it -# was - routinely `done: PR ...` after a delivery, which no line predicate can +# was - routinely `done` after a PR delivery, which no line predicate can # read as a wait. An alarm bounded only by the line therefore re-fires for the # captain's whole thinking time, on exactly the work they already have in hand. # diff --git a/docs/architecture.md b/docs/architecture.md index 72c64aa55d7..7af90ab2e00 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -11,7 +11,7 @@ firstmate's supervisor contract and routing index for conditional procedures is A zero-token bash watcher (`bin/fm-watch.sh`) sleeps on the fleet, classifies detected wakes in bash, and wakes the first mate only when something is actionable. Actionable wakes include captain-relevant status signals, no-verb signals without positive evidence that their crew is still executing, authenticated check output such as PR merge polling or a Relay mention, stale panes whose crew is not provably working whether their status log looks terminal or non-terminal, provably-working stale panes that persist past `FM_STALE_ESCALATE_SECS` with no wait their own worker declared, no writes to their own task worktree, and - in a home that armed `config/wedge-defer-parked-gate` - no validation gate of their own awaiting an unanswered supervisor decision, declared external waits and attended captain-held transfers that remain declared past `FM_PAUSE_RESURFACE_SECS`, and heartbeat backstop hits. For an ordinary crew task, a wait is read from both of its records: the status line a worker declared, and the backlog hold `bin/fm-captain-hold.sh` recorded once firstmate handed the work to the captain. -So a delivered ordinary crew task whose last line stays `done: PR ...` bounds repeated alarms from new pane hashes to the `FM_PAUSE_RESURFACE_SECS` cadence for the length of the captain's decision. +So a delivered ordinary crew task whose last line stays a `done` PR-ready line bounds repeated alarms from new pane hashes to the `FM_PAUSE_RESURFACE_SECS` cadence for the length of the captain's decision. The first hash still alarms, each new hash inside that window is absorbed, and a new hash after the window re-surfaces the hold; a terminal pane hash that never changes stays inert after its first alarm exactly as it did before this bound. The throttle is scoped to both the current captain-call lifecycle and the status-log state, so releasing and re-holding the same task without a status append starts a fresh window whose first new hash alarms. A secondmate reaches the stale path only for a wait declared in its status line, so a hold recorded only in the backlog while its last line is `working:` or `done:` is outside this guard. diff --git a/docs/captain-hold-lifecycle.md b/docs/captain-hold-lifecycle.md index f97ba727218..bd2f08018fc 100644 --- a/docs/captain-hold-lifecycle.md +++ b/docs/captain-hold-lifecycle.md @@ -25,7 +25,7 @@ A hold whose `--until` date has passed keeps those annotations while tasks-axi r The `complete` subcommand unions the reviewed captain-held task ids into `decision_keys=` and appends `decisions_reviewed=1` while originating task metadata is live. A post-teardown visual review can complete against the surviving report and durable tasks without recreating volatile task metadata. It accepts `--none` as an explicit semantic inventory result, refused while the origin still has a lifecycle-open keyed status decision, and verifies every listed task against tasks-axi before recording completion. -With a non-empty inventory it appends a `captain-held [key=]: tracked by ` transfer event for every still-open keyed status decision, which `bin/fm-classify-lib.sh` recognizes as closing the live status copy without claiming that the captain has answered it. +With a non-empty inventory it appends a `captain-held [key=]` transfer event naming the reviewed inventory for every still-open keyed status decision, which `bin/fm-classify-lib.sh` recognizes as closing the live status copy without claiming that the captain has answered it. Scout teardown calls the read-only `verify` subcommand after checking for the report and before removing any source state. `verify` requires the recorded attestation, requires every recorded inventory entry to still be durable (actively captain-held, or carrying a recorded answer), and fails on any keyed status decision that opened after the last `complete`, which makes re-running `complete` the repair. diff --git a/docs/configuration.md b/docs/configuration.md index 8da010d8417..e21c13f8799 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -17,7 +17,8 @@ Untracked files and directories whose names begin with `scratchpad` are also git `bin/fm-spawn.sh` owns the base task-metadata fields it emits, while the runtime-backend section below owns backend-specific fields and selector interpretation. `bin/fm-contributions.sh` owns durable published-contribution records under each task, observation bounds, equivalent triage-label configuration, and the authenticated contribution check. -The producing PR and Relay helpers own the fields they append, `bin/fm-classify-lib.sh` owns status-event vocabulary, and `bin/fm-crew-state.sh` owns current-state reconciliation. +The producing PR and Relay helpers own the fields they append, [`bin/fm-classify-lib.sh`](../bin/fm-classify-lib.sh) owns status-event vocabulary, optional emission-time syntax, and legacy unknown-time handling, and `bin/fm-crew-state.sh` owns current-state reconciliation. +The [`bin/fm-fleet-snapshot.sh` header](../bin/fm-fleet-snapshot.sh) owns the snapshot's event-time and age fields, including secondmate parent-event projections. Wake, watcher, away-mode, and Relay-specific state mechanics remain with their named scripts and reference sections rather than being duplicated into one exhaustive state tree here. `bin/fm-session-start.sh`'s header is the single owner of session-start ordering, composed commands, digest contents, and the digest's startup mechanism. diff --git a/docs/secondmate-parent-channel.md b/docs/secondmate-parent-channel.md index a9e682c9945..e99e037e668 100644 --- a/docs/secondmate-parent-channel.md +++ b/docs/secondmate-parent-channel.md @@ -22,7 +22,7 @@ Every captain-facing outcome that leaves durable evidence in the mate home is pu | Outcome | Durable evidence in the mate home | Published by | |---|---|---| -| Ship child PR ready | the child's `done: PR ...` line; `pr=` in the child's record once registered | `bin/fm-inactive-reconcile.sh` on the next poll with the child's line; `bin/fm-pr-check.sh` at registration with the canonical URL | +| Ship child PR ready | the child's `done:` PR ready line, whose accepted spellings the publisher below owns; `pr=` in the child's record once registered | `bin/fm-inactive-reconcile.sh` on the next poll with the child's line; `bin/fm-pr-check.sh` at registration with the canonical URL | | Scout child findings | the child's `done:` line plus `data//report.md` | `bin/fm-inactive-reconcile.sh` on the next poll, with the report pointer | | Child failed | the child's `failed:` line | `bin/fm-inactive-reconcile.sh` on the next poll | | Child decision escalated to the captain | the task held for the captain in the mate backlog | `bin/fm-captain-hold.sh hold`, and its answer by `answer` | @@ -33,7 +33,7 @@ Every captain-facing outcome that leaves durable evidence in the mate home is pu | An outcome that exists only in the mate's reasoning | none | the charter and the `AGENTS.md` carve-outs only | The ledger delivery reads files only: it calls no harness, no forge, and no current-state reader, so it is identical for every harness and runtime backend. -Each delivery is keyed with the first eight hexadecimal characters of its receipt fingerprint and appended at most once by exact line, and the ledger path reuses the inactive scan's per-fingerprint receipts, so a replayed poll or restart cannot deliver an event twice while a genuinely new terminal event is delivered again. +Each delivery is keyed with the first eight hexadecimal characters of its receipt fingerprint and uses the shared append contract above, and the ledger path reuses the inactive scan's per-fingerprint receipts, so a replayed poll or restart cannot deliver an event twice while a genuinely new terminal event is delivered again. A duplicate line is harmless and a missed one is not, so the mate may still append its own judgement about a delivered outcome, and the parent reads the script's line as the fact and the mate's line as commentary. For marked replies, the report helper accepts no caller-selected destination and uses the channel resolver for both local and remote homes; its script header owns the exact invocation contract. The pending-reply guard may restate only the correlated line from a local mate's `state/.status` onto the parent channel, which repairs the common parent-home versus mate-home mixup without accepting arbitrary mate-home sightings as acknowledgement. diff --git a/docs/verification/secondmate-parent-channel.md b/docs/verification/secondmate-parent-channel.md index 4cd541c8028..7a3cfcdcfbb 100644 --- a/docs/verification/secondmate-parent-channel.md +++ b/docs/verification/secondmate-parent-channel.md @@ -2,6 +2,8 @@ Maintainer-verification record for the guarantee in [`secondmate-parent-channel.md`](../secondmate-parent-channel.md): a captain-facing outcome recorded inside a secondmate home reaches the parent channel without the mate model writing it. Refresh it by rerunning the fixture below after changing any publisher named in `bin/fm-parent-channel-lib.sh`. +This run predates emission-time stamping, so each published line below is the payload without its stamp: a rerun now writes the same bytes with an `[at=]` tag closing the head, as in `done [key=child-outcome-child-done-05b032a1] [at=]: child ...`. +[`bin/fm-classify-lib.sh`](../../bin/fm-classify-lib.sh) owns that tag's syntax; nothing this record proves about delivery depends on it. ## What was run diff --git a/tests/fm-agy-harness.test.sh b/tests/fm-agy-harness.test.sh index 4de94772c81..1ccf3b4ba10 100755 --- a/tests/fm-agy-harness.test.sh +++ b/tests/fm-agy-harness.test.sh @@ -815,7 +815,7 @@ test_agy_unregistered_path_without_a_dialog_fails_the_spawn() { || fail "the gate must not send Enter into a pane that shows no dialog" assert_contains "$(cat "$CASE_DIR/tmux-calls.log")" "kill-window" \ "a failed agy readiness gate left its launched endpoint running" - assert_grep 'failed: agy never showed its folder-trust dialog' "$HOME_DIR/state/$id.status" \ + assert_grep 'failed: agy never showed its folder-trust dialog' <(sed -E 's/ \[at=[0-9]+\]//' "$HOME_DIR/state/$id.status") \ "a failed agy readiness gate did not record the failure in the task status" pass "fm-spawn: a busy verdict on an unregistered path without a dialog fails and closes the endpoint" } diff --git a/tests/fm-bearings-snapshot.test.sh b/tests/fm-bearings-snapshot.test.sh index ecdde88a82c..9aefd7b7578 100755 --- a/tests/fm-bearings-snapshot.test.sh +++ b/tests/fm-bearings-snapshot.test.sh @@ -380,6 +380,8 @@ test_domain_alpha_stale_parent_event_does_not_become_current_work() { .secondmate_current.records[] | select(.id == "domain-alpha") | .provenance.selected == "structured-home" and .freshness.status == "fresh" + and .parent_event.age_seconds == null + and (.parent_event | has("emitted_at_epoch") | not) and .terminal_evidence.provenance == "parent-direct-report-terminal" and .terminal_evidence.trust == "untrusted-supplement" and .terminal_evidence.captured == true @@ -429,7 +431,11 @@ SH and .parent_event.activity_scan.available == true ' >/dev/null || fail "GNU stat fixture corrupted the authoritative secondmate summary: $canonical" assert_contains "$(cat "$stat_log")" '-c %a' "GNU registry mode must use stat -c" - assert_contains "$(cat "$stat_log")" '-c %Y' "GNU parent-event mtime must use stat -c" + assert_contains "$(cat "$stat_log")" '-c %Y' "GNU status-observation mtime must use stat -c" + printf '%s' "$canonical" | jq -e ' + .secondmate_current.records[] | select(.id == "domain-alpha") + | .parent_event.age_seconds == null and (.parent_event | has("emitted_at_epoch") | not) + ' >/dev/null || fail "legacy event acquired an age from GNU stat" assert_contains "$(cat "$stat_log")" '-c %s' "GNU parent-event size must use stat -c" if grep -q '^-f ' "$stat_log"; then fail "GNU snapshot invoked BSD stat -f before its GNU file reads: $(cat "$stat_log")" diff --git a/tests/fm-branch-supervision.test.sh b/tests/fm-branch-supervision.test.sh index 6dd79583271..3a19310026e 100644 --- a/tests/fm-branch-supervision.test.sh +++ b/tests/fm-branch-supervision.test.sh @@ -54,7 +54,7 @@ test_branch_prompt_is_byte_stable_and_above_cache_floor() { *) fail "branch prompt lost the requested-result, progress-routine, or routine-silence rules" ;; esac case "$out_a" in - *"# PR identity: copy or abstain"*"copied verbatim from the task's \`done: PR \` status line or its \`pr=\` metadata field"*"Never assemble an owner, repository, host, or number"*"report the identifier you do have"*) ;; + *"# PR identity: copy or abstain"*"copied verbatim from the task's \`done [at=]: PR \` status line or its \`pr=\` metadata field"*"Never assemble an owner, repository, host, or number"*"report the identifier you do have"*) ;; *) fail "branch prompt lost the copy-or-abstain PR identity rule" ;; esac pass "branch prompt is byte-stable across homes, cwd, timezone, and time, above the cache floor" diff --git a/tests/fm-brief.test.sh b/tests/fm-brief.test.sh index 88f4e566ff0..03da3f01453 100755 --- a/tests/fm-brief.test.sh +++ b/tests/fm-brief.test.sh @@ -395,7 +395,7 @@ test_ask_user_escalation_format() { assert_grep "write only the ask-user findings, verbatim and unparaphrased (id, severity, file, line, description, authority)" "$brief" \ "ship rule 6 must limit the verbatim axi slice to ask-user findings" # shellcheck disable=SC2016 # single quotes are deliberate: backticks and the key/findings/file tokens must stay literal - assert_grep 'needs-decision [key=nm--]: ask-user findings=,,... file='"$home/data/$id/nm--findings.txt" "$brief" \ + assert_grep 'needs-decision [at=] [key=nm--]: ask-user findings=,,... file='"$home/data/$id/nm--findings.txt" "$brief" \ "ship rule 6 must render the exact needs-decision ask-user status line" assert_grep "$home/data/$id/nm--findings.txt" "$brief" \ "ship rule 6 must point the snapshot file under this task's own data directory" @@ -765,16 +765,16 @@ test_herdr_lab_contract_applies_to_scouts_but_not_secondmates() { } test_pause_verb_override_renders_all_brief_scaffolds() { - local home kind id brief + local home kind id brief append now epoch templates template line signals home="$TMP_ROOT/pause-verb-home" mkdir -p "$home/data" - for kind in ship scout secondmate; do - id="brief-pause-verb-$kind" + for kind in ship:no-mistakes ship:direct-PR ship:local-only scout secondmate; do + id="brief-pause-verb-${kind//:/-}" case "$kind" in - ship) + ship:*) FM_HOME="$home" FM_CLASSIFY_PAUSED_VERB=awaiting \ - "$ROOT/bin/fm-brief.sh" "$id" firstmate --mode no-mistakes >/dev/null 2>&1 + "$ROOT/bin/fm-brief.sh" "$id" firstmate --mode "${kind#ship:}" >/dev/null 2>&1 ;; scout) FM_HOME="$home" FM_CLASSIFY_PAUSED_VERB=awaiting \ @@ -786,6 +786,55 @@ test_pause_verb_override_renders_all_brief_scaffolds() { ;; esac brief="$home/data/$id/brief.md" + # Fill the scaffold's generated status-append command the way a worker does + # and run it. The stamp must be a value the worker supplies, so the command + # may not carry an unevaluated substitution that a file-write tool would + # copy through verbatim. + # shellcheck disable=SC2016 # Match literal backticks in the generated interface. + append=$(sed -n '/`echo "{state}/s/.*`\(echo .*\)`.*/\1/p' "$brief") + now=$(date +%s) + append=${append//\{state\}/done} + append=${append//\{one short line\}/test event} + append=${append///$now} + case "$append" in + *"\$("*) fail "$kind scaffold left an unevaluated command in its status-append line" ;; + esac + mkdir -p "$home/state" + bash -c "$append" || fail "generated status command failed" + epoch=$(bash -c '. "$1"; status_line_at_epoch "$(cat "$2")"' _ \ + "$ROOT/bin/fm-classify-lib.sh" "$home/state/$id.status") + [ "$epoch" = "$now" ] || fail "$kind scaffold did not record the worker's event time" + # Every status signal the brief instructs a worker to append is a template + # the worker fills in and writes verbatim, with or without a shell, not only + # rule 4's echo: substitute each one's named placeholders and read the stamp + # back. Extracting by "append" as well as by the stamp means dropping a stamp + # from any instruction fails here rather than shrinking the set. + templates=$(grep -o -e "append \`[^\`]*: [^\`]*\`" \ + -e "\`[^\`]*\[at=\][^\`]*\`" "$brief" \ + | sed 's/^append //' | tr -d '`' | sort -u) + signals=0 + while IFS= read -r template; do + [ -n "$template" ] || continue + case "$template" in + 'echo "'*) template=${template#echo \"}; template=${template%%\" >>*} ;; + esac + case "$template" in + *"\$("*) fail "$kind signal embeds an unevaluated command: $template" ;; + esac + now=$(date +%s) + line=${template//\{state\}/done} + line=${line///$now} + line=$(printf '%s' "$line" \ + | sed -e 's/{[^}]*}/one short line/g' -e 's/<[^>]*>/slug/g') + epoch=$(bash -c '. "$1"; status_line_at_epoch "$2"' _ \ + "$ROOT/bin/fm-classify-lib.sh" "$line") + [ "$epoch" = "$now" ] || fail "$kind signal carries no worker-written stamp: $template" + signals=$((signals + 1)) + done </dev/null \ + sed -E 's/ \[at=[0-9]+\]//' "$home/state/$id.status" | grep -F 'captain-held [key=route]: tracked by sample-route-call' >/dev/null \ || fail "the transfer line does not name the tracking inventory" before=$(shasum -a 256 "$home/data/backlog.md" | awk '{print $1}') @@ -1353,7 +1353,7 @@ EOF run_teardown "$mate" "$origin" >/dev/null 2> "$mate/teardown.err" \ || fail "secondmate investigation teardown failed: $(cat "$mate/teardown.err")" tasks_in "$mate" "done" "$origin" --report "data/$origin/report.md" --keep 0 >/dev/null - grep -Eq "^done \\[key=child-outcome-$origin-done-[0-9a-f]{8}\\]: child $origin done: report and visual review complete mode=scout report=data/$origin/report.md$" \ + grep -Eq "^done \\[key=child-outcome-$origin-done-[0-9a-f]{8}\\] \\[at=[0-9]+\\]: child $origin done: report and visual review complete mode=scout report=data/$origin/report.md$" \ "$parent/state/sample-mate.status" \ || fail "the scout's final line did not reach the parent at teardown" @@ -1399,13 +1399,13 @@ EOF run_captain "$mate" hold quoted-record-call --reason "quoted record choice pending" \ --origin quoted-origin >/dev/null || fail "quoted-record hold failed" assert_grep 'needs-decision [key=captain-hold-quoted-record-call-1]: captain hold quoted-record-call: quoted record choice pending' \ - "$channel" "body prose was incorrectly counted as a resolution record" + <(sed -E 's/ \[at=[0-9]+\]//' "$channel") "body prose was incorrectly counted as a resolution record" run_captain "$mate" hold mate-call --title "Choose the mate release" \ --reason "release choice pending" --repo sample >/dev/null \ || fail "mate hold failed" assert_grep 'needs-decision [key=captain-hold-mate-call-1]: captain hold mate-call: release choice pending' \ - "$channel" "the mate's hold did not reach the parent channel" + <(sed -E 's/ \[at=[0-9]+\]//' "$channel") "the mate's hold did not reach the parent channel" run_captain "$mate" hold mate-call --reason "release choice pending" >/dev/null \ || fail "repeated mate hold failed" [ "$(grep -c 'captain-hold-mate-call-1' "$channel")" = 1 ] \ @@ -1415,17 +1415,17 @@ EOF run_captain "$mate" answer mate-call --decision-file "$decision" --release >/dev/null \ || fail "mate release answer failed" assert_grep 'resolved [key=captain-hold-mate-call-1]: captain hold mate-call: released' \ - "$channel" "the released answer did not close the parent decision" + <(sed -E 's/ \[at=[0-9]+\]//' "$channel") "the released answer did not close the parent decision" run_captain "$mate" hold mate-call --reason "second release choice" >/dev/null \ || fail "re-hold after release failed" assert_grep 'needs-decision [key=captain-hold-mate-call-2]: captain hold mate-call: second release choice' \ - "$channel" "a re-held task did not open a distinct parent decision" + <(sed -E 's/ \[at=[0-9]+\]//' "$channel") "a re-held task did not open a distinct parent decision" printf 'ship it\n' > "$decision" run_captain "$mate" answer mate-call --decision-file "$decision" >/dev/null \ || fail "mate close answer failed" assert_grep 'resolved [key=captain-hold-mate-call-2]: captain hold mate-call: answered' \ - "$channel" "the closing answer did not close the second parent decision" + <(sed -E 's/ \[at=[0-9]+\]//' "$channel") "the closing answer did not close the second parent decision" run_captain "$mate" answer mate-call --decision-file "$decision" >/dev/null \ || fail "idempotent answer retry failed" [ "$(grep -c 'captain-hold-mate-call-2' "$channel")" = 2 ] \ @@ -1492,13 +1492,15 @@ test_secondmate_reconcile_publishes_before_request_retirement() { assert_contains "$show" "Resolution mode: reconciled" \ "request retirement failure lost the reconciled resolution mode" [ -f "$request" ] || fail "the request retired despite its forced retirement failure" - [ "$(grep -c 'resolved \[key=captain-hold-reconcile-channel-call-1\]: captain hold reconcile-channel-call: reconciled' "$channel")" -eq 1 ] \ + [ "$(grep -c 'resolved \[key=captain-hold-reconcile-channel-call-1\]: captain hold reconcile-channel-call: reconciled' \ + <(sed -E 's/ \[at=[0-9]+\]//' "$channel"))" -eq 1 ] \ || fail "the parent resolution was not published before retirement failed: $(cat "$channel")" run_captain "$mate" reconcile close reconcile-channel-call --evidence-file "$evidence" >/dev/null \ || fail "the closed reconciliation could not finish publication and retirement" [ ! -e "$request" ] || fail "the retry did not retire the published reconcile request" - [ "$(grep -c 'resolved \[key=captain-hold-reconcile-channel-call-1\]: captain hold reconcile-channel-call: reconciled' "$channel")" -eq 1 ] \ + [ "$(grep -c 'resolved \[key=captain-hold-reconcile-channel-call-1\]: captain hold reconcile-channel-call: reconciled' \ + <(sed -E 's/ \[at=[0-9]+\]//' "$channel"))" -eq 1 ] \ || fail "the reconciliation retry duplicated or changed its parent resolution: $(cat "$channel")" tasks_in "$mate" add answer-channel-call "Answer the mate call" --kind ship --repo sample >/dev/null \ || fail "could not create the normal-answer channel call" @@ -1518,12 +1520,14 @@ test_secondmate_reconcile_publishes_before_request_retirement() { show=$(tasks_in "$mate" show answer-channel-call --full) assert_contains "$show" "state: done" "request retirement failure reversed the captain answer" [ -f "$request" ] || fail "the normal-answer retry trigger retired after its forced failure" - [ "$(grep -c 'resolved \[key=captain-hold-answer-channel-call-1\]: captain hold answer-channel-call: answered' "$channel")" -eq 1 ] \ + [ "$(grep -c 'resolved \[key=captain-hold-answer-channel-call-1\]: captain hold answer-channel-call: answered' \ + <(sed -E 's/ \[at=[0-9]+\]//' "$channel"))" -eq 1 ] \ || fail "the normal answer did not publish before retirement failed: $(cat "$channel")" run_captain "$mate" answer answer-channel-call --decision-file "$mate/answer.txt" >/dev/null \ || fail "the normal-answer retry could not finish request retirement" [ ! -e "$request" ] || fail "the normal-answer retry left its request pending" - [ "$(grep -c 'resolved \[key=captain-hold-answer-channel-call-1\]: captain hold answer-channel-call: answered' "$channel")" -eq 1 ] \ + [ "$(grep -c 'resolved \[key=captain-hold-answer-channel-call-1\]: captain hold answer-channel-call: answered' \ + <(sed -E 's/ \[at=[0-9]+\]//' "$channel"))" -eq 1 ] \ || fail "the normal-answer retry duplicated its parent resolution: $(cat "$channel")" pass "secondmate resolutions publish before retiring durable retry triggers" } diff --git a/tests/fm-classify-corr-token.test.sh b/tests/fm-classify-corr-token.test.sh index b9bcc80a2a3..f82dcf94898 100755 --- a/tests/fm-classify-corr-token.test.sh +++ b/tests/fm-classify-corr-token.test.sh @@ -524,6 +524,7 @@ EOF FM_HOME="$mate" "$REPORT" "done" "$corr" "audit clean" \ || fail "$REPORT failed writing a correlated report" helper_line=$(tail -1 "$state/pinned.status") + status_line_at_epoch "$helper_line" >/dev/null || fail "report helper emitted no time" verb=$(status_line_verb "$helper_line") [ "$verb" = "done" ] \ || fail "the classifier did not read through the helper's own line '$helper_line' (verb=[$verb])" @@ -533,6 +534,7 @@ EOF FM_HOME="$mate" "$REPORT" --doc needs-decision "$corr" data/x/report.md "see the report" \ || fail "$REPORT failed writing a correlated doc-pointer report" helper_line=$(tail -1 "$state/pinned.status") + status_line_at_epoch "$helper_line" >/dev/null || fail "doc report helper emitted no time" verb=$(status_line_verb "$helper_line") [ "$verb" = needs-decision ] \ || fail "the classifier did not read through the helper's doc line '$helper_line' (verb=[$verb])" @@ -540,6 +542,228 @@ EOF pass "both real correlation-token writers produce lines this classifier reads through" } +test_optional_event_time() { + local line stamped epoch before after dir + before=$(date +%s) + line="needs-decision corr=$CORR [key=timed]: choose: A or B" + stamped=$(status_stamp_line "$line") || fail "status writer could not stamp an event" + after=$(date +%s) + epoch=$(status_line_at_epoch "$stamped") || fail "new event has no emission time" + [ "$epoch" -ge "$before" ] && [ "$epoch" -le "$after" ] || fail "event time is not append time" + [ "$(status_line_verb "$stamped")" = needs-decision ] || fail "time changed verb" + [ "$(_fm_decision_key "$stamped")" = timed ] || fail "time changed key" + [ "$(status_line_note "$stamped")" = 'choose: A or B' ] || fail "time changed note" + [ "$(status_stamp_line "$stamped")" = "$stamped" ] || fail "restamping changed emission time" + line='done [at=1700000000]: old event' + [ "$(status_stamp_line "$line")" = "$line" ] || fail "writer replaced an old emission time" + ( + # shellcheck disable=SC2329 # status_stamp_line invokes this clock stub indirectly. + date() { return 1; } + [ "$(status_stamp_line 'done: clock unavailable')" = 'done: clock unavailable' ] + ) || fail "clock failure lost the event" + for line in 'done: legacy' 'done: [at=1700000000] prose' \ + 'done [at=]: empty' "done [at=\$(date +%s)]: literal substitution" \ + 'done [at=]: unsubstituted placeholder' \ + 'done [at=bad]: malformed' 'done [at=17:00]: malformed colon' 'done [at=-1]: negative' \ + 'done [at=01700000000]: noncanonical' 'done [at=99999999999999999999]: overflow' \ + 'done [at=1] [at=2]: ambiguous'; do + if status_line_at_epoch "$line" >/dev/null; then fail "invented time for $line"; fi + done + # A readable time a worker wrote instead of epoch seconds carries colons that + # must not move the head/note separator, in either metadata order. + for line in "needs-decision [key=api-shape] [at=10:30]: choose: A or B" \ + "needs-decision [at=10:30] [key=api-shape]: choose: A or B" \ + "needs-decision [key=api-shape] [at=2026-09-20T14:03:00Z]: choose: A or B"; do + [ "$(_fm_decision_key "$line")" = api-shape ] \ + || fail "a colon-bearing time hid the decision key: [$(_fm_decision_key "$line")] from $line" + [ "$(status_line_note "$line")" = 'choose: A or B' ] \ + || fail "a colon-bearing time garbled the note: [$(status_line_note "$line")] from $line" + done + for line in "done [at=1700000000] [corr=$CORR]: finished" \ + "done [corr=$CORR] [at=1700000000]: finished" \ + "done[at=1700000000] [corr=$CORR]: finished"; do + [ "$(status_line_at_epoch "$line")" = 1700000000 ] || fail "metadata order changed time" + done + dir=$(make_case event-time) + # The real parent publisher deduplicates a retry against both timed and + # legacy records without rewriting the first event's time. + . "$ROOT/bin/fm-parent-channel-lib.sh" + line="done [corr=$CORR]: path: C:\\notes" + printf '%s\n' "$(status_stamp_line "$line")" > "$dir/state/retry.status" + stamped=$(cat "$dir/state/retry.status") + fm_parent_channel_append_once "$dir/state/retry.status" "$line" || fail "parent retry failed" + [ "$(cat "$dir/state/retry.status")" = "$stamped" ] || fail "retry duplicated or restamped event" + printf '%s\n' "$line" > "$dir/state/legacy.status" + fm_parent_channel_append_once "$dir/state/legacy.status" "$line" || fail "legacy retry failed" + [ "$(cat "$dir/state/legacy.status")" = "$line" ] || fail "legacy retry acquired an invented time" + fm_parent_channel_append_once "$dir/state/retry.status" "done [corr=$CORR2]: path: C:\\notes" + [ "$(wc -l < "$dir/state/retry.status")" -eq 2 ] || fail "dedup discarded different correlation" + fm_parent_channel_append_once "$dir/state/retry.status" 'done: prose [at=1]' + fm_parent_channel_append_once "$dir/state/retry.status" 'done: prose [at=2]' + [ "$(wc -l < "$dir/state/retry.status")" -eq 4 ] || fail "dedup stripped a time mention from prose" + # A malformed time tag is ordinary event bytes, so it identifies the event: + # the unstamped line is a DIFFERENT event, while re-appending the same bytes + # is still a retry. + for line in 'done [at=17:00]: shipped' 'done [at=]: shipped' 'done [at=bad]: shipped' \ + 'done [at=1] [at=2]: shipped' 'done [at=01700000000]: shipped' \ + 'done [at=99999999999999999999]: shipped'; do + printf '%s\n' "$line" > "$dir/state/malformed.status" + fm_parent_channel_append_once "$dir/state/malformed.status" 'done: shipped' \ + || fail "append after a malformed time failed" + [ "$(wc -l < "$dir/state/malformed.status")" -eq 2 ] \ + || fail "dedup stripped a malformed time tag: $line" + fm_parent_channel_append_once "$dir/state/malformed.status" "$line" \ + || fail "malformed time retry failed" + [ "$(head -1 "$dir/state/malformed.status")" = "$line" ] \ + && [ "$(wc -l < "$dir/state/malformed.status")" -eq 2 ] \ + || fail "retry duplicated or rewrote malformed time: $line" + done + # A well-formed numeric tag still strips, in either metadata order. + for line in "done [at=1700000000] [corr=$CORR2]: stamped" \ + "done [corr=$CORR2] [at=1700000000]: stamped" \ + "done[at=1700000000] [corr=$CORR2]: stamped"; do + printf '%s\n' "$line" > "$dir/state/timed.status" + fm_parent_channel_append_once "$dir/state/timed.status" "done [corr=$CORR2]: stamped" \ + || fail "numeric time retry failed" + [ "$(cat "$dir/state/timed.status")" = "$line" ] \ + || fail "dedup did not ignore a well-formed numeric time: $line" + done + stamped=$(status_stamp_line "needs-decision corr=$CORR [key=timed]: choose: A or B") + printf '%s\n' "$stamped" 'working [at=1700000000]: unrelated progress' > "$dir/state/task.status" + [ -n "$(status_open_decisions "$dir/state/task.status")" ] || fail "time cleared an open decision" + printf '%s\n' 'resolved [at=1700000001] [key=timed]: answered' >> "$dir/state/task.status" + [ -z "$(status_open_decisions "$dir/state/task.status")" ] || fail "timed resolution did not close decision" + pass "optional event time preserves parsing and legacy unknown time" +} + +test_captain_override_ignores_event_time() { + local dir verb line event + local FM_CAPTAIN_RE='done:|needs-decision:|blocked:|failed:' + dir=$(make_case captain-override-time) + for verb in 'done' needs-decision blocked failed; do + for line in "$verb: audit complete" "$verb [at=1700000000]: audit complete" \ + "${verb}[at=1700000000]: audit complete"; do + status_is_captain_relevant "$line" || fail "override missed actionable event: $line" + printf '%s\n' "$line" > "$dir/state/task.status" + event=$(status_span_first_actionable "$dir/state/task.status" 0) \ + || fail "override hid actionable status span: $line" + [ "$event" = "$line" ] || fail "classification changed surfaced event bytes: $event" + [ "$(cat "$dir/state/task.status")" = "$line" ] || fail "classification rewrote stored event" + done + done + FM_CAPTAIN_RE='done:' + for line in 'blocked: waiting' 'blocked [at=1700000000]: waiting'; do + status_is_captain_relevant "$line" && fail "override admitted excluded event: $line" + printf '%s\n' "$line" > "$dir/state/task.status" + status_span_has_actionable "$dir/state/task.status" 0 \ + && fail "override surfaced excluded event: $line" + done + for verb in working paused resolved captain-held; do + for line in "$verb: done: mentioned" "$verb [at=1700000000]: done: mentioned"; do + status_is_captain_relevant "$line" && fail "override bypassed nonterminal suppression: $line" + done + done + FM_CAPTAIN_RE='^custom-verb: audit complete$' + for line in 'custom-verb: audit complete' 'custom-verb [at=1700000000]: audit complete' \ + 'custom-verb [at=]: audit complete'; do + status_is_captain_relevant "$line" || fail "timestamp broke custom verb override: $line" + printf '%s\n' 'working: started' "$line" > "$dir/state/task.status" + [ "$(last_status_line "$dir/state/task.status")" = "$line" ] \ + || fail "event scan skipped the stamped custom-verb event: $line" + done + FM_CAPTAIN_RE="^done \\[corr=$CORR\\]: literal \\[at=1700000000\\]$" + for line in "done [corr=$CORR]: literal [at=1700000000]" \ + "done [at=1700000000] [corr=$CORR]: literal [at=1700000000]" \ + "done [corr=$CORR] [at=1700000000]: literal [at=1700000000]"; do + status_is_captain_relevant "$line" || fail "normalization changed correlation metadata or note: $line" + done + pass "captain regex overrides preserve timed and legacy relevance and event bytes" +} + +# A malformed time tag is never read as a time: relevance, verb, and note all see +# the same ordinary bytes, so a FM_CAPTAIN_RE override matching ":" does not +# find a separator the line does not have, while the terminal-verb default still +# surfaces the event. +test_malformed_event_time_is_ordinary_bytes() { + local dir verb line event + dir=$(make_case malformed-event-time) + for verb in 'done' needs-decision blocked failed; do + for line in "$verb [at=]: audit complete" "$verb [at=bad]: audit complete" \ + "$verb [at=17:00]: audit complete" "$verb [at=bad] [at=17:00]: audit complete" \ + "$verb [at=2026-09-20T14:03:00Z]: audit complete" "$verb [at=10:30]: audit complete" \ + "$verb [at=\$(date +%s)]: audit complete" \ + "$verb [at=]: audit complete" \ + "$verb [at=1] [at=2]: audit complete" \ + "$verb [at=01700000000]: audit complete" \ + "$verb [at=99999999999999999999]: audit complete"; do + if status_line_at_epoch "$line" >/dev/null; then fail "invented time for $line"; fi + [ "$(status_line_verb "$line")" = "$verb" ] || fail "malformed time changed verb: $line" + [ "$(status_line_note "$line")" = 'audit complete' ] \ + || fail "malformed time garbled the note: [$(status_line_note "$line")] from $line" + [ "$(_fm_decision_key "$line")" = default ] \ + || fail "malformed time invented a decision key: [$(_fm_decision_key "$line")] from $line" + status_is_captain_relevant "$line" \ + || fail "default vocabulary lost an actionable event: $line" + printf '%s\n' "$line" > "$dir/state/task.status" + event=$(status_span_first_actionable "$dir/state/task.status" 0) \ + || fail "default vocabulary hid actionable status span: $line" + [ "$event" = "$line" ] || fail "classification changed surfaced event bytes: $event" + # A tag the worker spelled wrong is still a tag, so it must not decide + # whether the supervisor sees a terminal event - including a readable + # timestamp whose colons would otherwise swallow the head/note separator. + ( + FM_CAPTAIN_RE='done:|needs-decision:|blocked:|failed:' + status_is_captain_relevant "$line" || exit 1 + exit 0 + ) || fail "override lost a terminal event to a malformed tag: $line" + printf '%s\n' "$line" > "$dir/state/scan.status" + ( + FM_CAPTAIN_RE='done:|needs-decision:|blocked:|failed:' + event=$(last_status_line "$dir/state/scan.status") + [ "$event" = "$line" ] || exit 1 + ) || fail "event scan lost a terminal event to a malformed tag: $line" + done + done + pass "malformed event times stay ordinary line bytes without hiding the event" +} + +# The decision fold reads the head/note separator on the same unstamped copy the +# note and key readers use, so a worker's mis-spelled time tag cannot decide +# whether a captain's decision survives. Without that, a readable "[at=17:00]" +# hands the fold a colon it never wrote: a colonless terminal line closes every +# open decision, and a colonless declaration opens a phantom one no later line +# can close. +test_malformed_event_time_never_moves_the_decision_fold() { + local dir status tag + dir=$(make_case fold-malformed-event-time) + status="$dir/state/task.status" + printf 'kind=ship\n' > "$dir/state/task.meta" + for tag in '[at=17:00]' '[at=10:30]' '[at=2026-09-20T14:03:00Z]' '[at=]' '[at=bad]'; do + printf '%s\n%s\n' \ + 'needs-decision [key=api-shape] [at=1700000000]: REST or gRPC?' \ + "done $tag finished the audit" > "$status" + case "$(status_open_decisions "$status")" in + 'api-shape'$'\t''needs-decision'$'\t''REST or gRPC?') : ;; + *) fail "malformed tag $tag closed an open decision: [$(status_open_decisions "$status")]" ;; + esac + printf '%s\n' "needs-decision $tag which base branch" > "$status" + [ -z "$(status_open_decisions "$status")" ] \ + || fail "malformed tag $tag opened a phantom decision: [$(status_open_decisions "$status")]" + done + # The real separator still closes, so the tolerance above did not disarm the + # terminal rule itself. + printf '%s\n%s\n' \ + 'needs-decision [key=api-shape] [at=1700000000]: REST or gRPC?' \ + 'done [at=1700000001]: finished the audit' > "$status" + [ -z "$(status_open_decisions "$status")" ] \ + || fail "a well-formed terminal event stopped closing the decision" + pass "malformed event times never open or close a decision" +} + +test_captain_override_ignores_event_time +test_malformed_event_time_is_ordinary_bytes +test_malformed_event_time_never_moves_the_decision_fold +test_optional_event_time test_tokened_opener_opens_and_tokened_closer_closes test_token_is_read_through_in_every_position_it_is_written_in test_untokened_pair_is_unchanged diff --git a/tests/fm-cmux-claude-composer-live-e2e.test.sh b/tests/fm-cmux-claude-composer-live-e2e.test.sh index 439d9335e97..e1670fbaea3 100755 --- a/tests/fm-cmux-claude-composer-live-e2e.test.sh +++ b/tests/fm-cmux-claude-composer-live-e2e.test.sh @@ -15,12 +15,17 @@ SPAWNED=0 fail() { printf 'not ok - %s\n' "$1" >&2; exit 1; } pass() { printf 'ok - %s\n' "$1"; } +# The scout brief instructs the optional `[at=]` stamp on every append, +# and a live worker may place it before or after a key. Match these events with +# the stamp removed instead of pinning one spelling. +untimed_status() { sed -E 's/ \[at=[0-9]+\]//g' "$1" 2>/dev/null; } + cleanup() { [ "$SPAWNED" -eq 0 ] || { mkdir -p "$LAB/data/$TASK" : > "$LAB/data/$TASK/report.md" - if grep -q '^needs-decision \[key=probe-decision\]' "$LAB/state/$TASK.status" 2>/dev/null \ - && ! grep -q '^resolved \[key=probe-decision\]' "$LAB/state/$TASK.status" 2>/dev/null; then + if untimed_status "$LAB/state/$TASK.status" | grep -q '^needs-decision \[key=probe-decision\]' \ + && ! untimed_status "$LAB/state/$TASK.status" | grep -q '^resolved \[key=probe-decision\]'; then printf '%s\n' 'resolved [key=probe-decision]: live guard cleanup' >> "$LAB/state/$TASK.status" fi FM_HOME="$LAB" "$ROOT/bin/fm-decision-hold.sh" complete "$TASK" --none >/dev/null 2>&1 || true @@ -55,9 +60,9 @@ brief = Path(sys.argv[1]) status = sys.argv[2] brief.write_text(brief.read_text().replace("{TASK}", f'''Run a cmux communication probe. -Immediately append `working: cmux composer probe ready` to `{status}`. -Then append exactly `needs-decision [key=probe-decision]: awaiting codeword` to that file and stop to wait for a firstmate message. -When you receive a firstmate message containing `ALBATROSS`, append `done: received ALBATROSS` to that status file and stop. +Immediately append `working [at=]: cmux composer probe ready` to `{status}`, substituting `` as rule 4 instructs. +Then append exactly `needs-decision [at=] [key=probe-decision]: awaiting codeword` to that file and stop to wait for a firstmate message. +When you receive a firstmate message containing `ALBATROSS`, append `done [at=]: received ALBATROSS` to that status file and stop. Do not change project files or make a commit.''')) PY @@ -78,10 +83,10 @@ for _ in $(seq 1 45); do case "$CAPTURE" in *'Yes, I trust this folder'*) FM_HOME="$LAB" "$ROOT/bin/fm-send.sh" "$TASK" --key Enter || fail "could not accept Claude's folder-trust prompt" ;; esac - grep -q '^needs-decision \[key=probe-decision\]' "$STATUS" 2>/dev/null && break + untimed_status "$STATUS" | grep -q '^needs-decision \[key=probe-decision\]' && break sleep 2 done -grep -q '^needs-decision \[key=probe-decision\]' "$STATUS" 2>/dev/null \ +untimed_status "$STATUS" | grep -q '^needs-decision \[key=probe-decision\]' \ || fail "Claude $(claude --version) did not reach the communication decision" COMPOSER=$(fm_backend_cmux_composer_state "$TARGET" "$TASK") @@ -91,12 +96,12 @@ pass "cmux classifies the real Claude borderless composer as empty" FM_SEND_SETTLE=0 FM_HOME="$LAB" "$ROOT/bin/fm-send.sh" "$TASK" --resolve-key probe-decision ALBATROSS \ || fail "cmux did not confirm the real Claude steer" for _ in $(seq 1 30); do - grep -q '^done: received ALBATROSS' "$STATUS" 2>/dev/null && break + untimed_status "$STATUS" | grep -q '^done: received ALBATROSS' && break sleep 2 done -grep -q '^resolved \[key=probe-decision\]: answered: ALBATROSS' "$STATUS" \ +untimed_status "$STATUS" | grep -q '^resolved \[key=probe-decision\]: answered: ALBATROSS' \ || fail "confirmed cmux delivery did not close the keyed decision" -grep -q '^done: received ALBATROSS' "$STATUS" \ +untimed_status "$STATUS" | grep -q '^done: received ALBATROSS' \ || fail "the real Claude worker did not complete after the confirmed steer" CAPTURE=$(fm_backend_cmux_capture "$TARGET" 200 "$TASK") diff --git a/tests/fm-contributions.test.sh b/tests/fm-contributions.test.sh index 24e1d3b9909..d4c5abf0f1e 100755 --- a/tests/fm-contributions.test.sh +++ b/tests/fm-contributions.test.sh @@ -584,7 +584,9 @@ test_budget_exhaustion_keeps_prior_record() { # exhaust|hang wrap_forge "$home" mutate_record "$home" delivery '.records[0].checked_at="2026-09-15T08:00:00Z"' cp "$home/data/delivery/contributions.json" "$home/prior.json" - if [ "$mode" = exhaust ]; then /bin/date +%s > "$home/forge/clock"; fi + # Both modes freeze the clock: an unfrozen one can tick past a one-second + # budget before the first forge call, so nothing is ever observed. + /bin/date +%s > "$home/forge/clock" printf '%s\n' "$mode" > "$home/forge/fault" out=$(with_home "$home" env FM_CONTRIBUTIONS_BUDGET=1 "$ROOT/bin/fm-contributions.sh" poll) \ || fail "poll failed when its budget ran out ($mode)" diff --git a/tests/fm-fleet-snapshot-view.test.sh b/tests/fm-fleet-snapshot-view.test.sh index 4ee0e9bf219..1238568f31f 100755 --- a/tests/fm-fleet-snapshot-view.test.sh +++ b/tests/fm-fleet-snapshot-view.test.sh @@ -181,6 +181,11 @@ test_fixture_snapshot_json() { and .endpoint.agent_alive == "alive" and (.actions.watch | contains("do not routinely fm-peek")) ' >/dev/null || fail "secondmate return-channel guidance missing" + printf '%s' "$out" | jq -e ' + .tasks[] | select(.id == "secondmate-task") + | .paths.status_log.last_event + | has("age_seconds") and .age_seconds == null + ' >/dev/null || fail "legacy event must have an explicit unknown age" printf '%s' "$out" | jq -e ' .tasks[] | select(.id == "cmux-task") | .backend == "cmux" @@ -194,7 +199,60 @@ test_fixture_snapshot_json() { .backlog.records[] | select(.id == "done-task") | .state == "done" and .pr_url == "https://github.com/kunchenguid/firstmate/pull/7" ' >/dev/null || fail "done backlog PR row missing" - pass "fixture snapshot covers task rows, backlog rows, pointers, and stable ordering" + + local line expected_age before after emitted epoch observed + printf 'secondmate-task\n' > "$home/secondmate-home/.fm-secondmate-home" + printf 'schema=fm-secondmate-parent.v1\nroute=local\nparent_home=%s\n' "$home" \ + > "$home/secondmate-home/.fm-secondmate-parent" + before=$(date +%s) + FM_HOME="$home/secondmate-home" "$ROOT/bin/fm-secondmate-report.sh" \ + 'done' 0123456789abcdef 'audit complete' || fail "parent report failed" + after=$(date +%s) + emitted=$(tail -1 "$home/state/secondmate-task.status") + # shellcheck source=bin/fm-classify-lib.sh + . "$ROOT/bin/fm-classify-lib.sh" + epoch=$(status_line_at_epoch "$emitted") || fail "new parent report has unknown time" + [ "$epoch" -ge "$before" ] && [ "$epoch" -le "$after" ] \ + || fail "parent report did not record emission time" + for line in "$emitted" 'working: legacy' 'working [at=1700000000]: timed' \ + 'working [at=1700000200]: future' 'working [at=oops]: malformed'; do + printf '%s\n\n' "$line" > "$home/state/secondmate-task.status" + # Deliberately unrelated file age must never substitute for event age. + touch -t 202001010000 "$home/state/secondmate-task.status" + expected_age=null; observed=1700000100 + case "$line" in + "$emitted") expected_age=100; observed=$((epoch + 100)) ;; + *1700000000*) expected_age=100 ;; + esac + out=$(PATH="$fakebin:$PATH" FM_HOME="$home" FM_SNAPSHOT_NOW_EPOCH=$observed "$SNAPSHOT" --json) + printf '%s' "$out" | jq -e --argjson age "$expected_age" ' + .tasks[] | select(.id == "secondmate-task") + | .paths.status_log.last_event + | has("age_seconds") and .age_seconds == $age + and (has("emitted_at_epoch") | not) + ' >/dev/null || fail "event age came from something other than the record: $line" + # parent_event age is the emission age; freshness is how old this snapshot's + # own observation of the file is, so the 2020 mtime must show up there and + # only there. + printf '%s' "$out" | jq -e --argjson age "$expected_age" ' + .secondmate_current.records[] | select(.id == "secondmate-task") + | .current.state == "unknown" + and .parent_event.age_seconds == $age + and (.parent_event | has("emitted_at_epoch") | not) + and (.freshness.age_seconds | type) == "number" + and .freshness.age_seconds > 100000000 + ' >/dev/null || fail "fallback confused event age, observation freshness, and current state: $line" + if [ "${FM_TEST_EVIDENCE:-0}" = 1 ]; then + printf '$ touch -t 202001010000 %s\n' "$home/state/secondmate-task.status" + printf '$ FM_HOME=%s FM_SNAPSHOT_NOW_EPOCH=%s bin/fm-fleet-snapshot.sh --json\n' "$home" "$observed" + printf '%s' "$out" | jq '{ + last_event: (.tasks[] | select(.id == "secondmate-task") | .paths.status_log.last_event), + secondmate: (.secondmate_current.records[] | select(.id == "secondmate-task") + | {current, parent_event, freshness}) + }' + fi + done + pass "fixture snapshot covers task rows, backlog rows, pointers, stable ordering, and emission-time event age" } # R1 owner contract: main_inventory discloses orphan in-flight and unstructured diff --git a/tests/fm-inactive-reconcile.test.sh b/tests/fm-inactive-reconcile.test.sh index 0c57b55691e..720a75c7082 100755 --- a/tests/fm-inactive-reconcile.test.sh +++ b/tests/fm-inactive-reconcile.test.sh @@ -177,7 +177,7 @@ test_local_secondmate_delivers_terminal_ledger_line() { FM_FAKE_CREW_STATE='unknown' run_reconcile "$MATE" key=$(reported_outcome_key "$MATE" child 'done') || fail "ledger receipt did not retain its collision-resistant key" expected="done [key=$key]: child child done: PR https://example.test/owner/repo/pull/1 checks green pr=https://example.test/owner/repo/pull/1 mode=no-mistakes yolo=off" - grep -Fxq "$expected" "$MAIN/state/mate.status" \ + sed -E 's/ \[at=[0-9]+\]//' "$MAIN/state/mate.status" | grep -Fxq "$expected" \ || fail "secondmate did not deliver the child's ledger line on a plain poll: $(cat "$MAIN/state/mate.status" 2>/dev/null)" [ "$(outcome_count "$MATE" reported)" = 1 ] || fail "ledger delivery receipt was not durable" FM_FAKE_CREW_STATE='unknown' run_reconcile "$MATE" @@ -217,7 +217,8 @@ SH run_report "$MATE" child key=$(reported_outcome_key "$MATE" child "$terminal") \ || fail "$terminal with trailing prose arriving $timing state read was not owned by the ledger" - grep -Fq "$terminal [key=$key]: child child $terminal: validation finished" "$MAIN/state/mate.status" \ + sed -E 's/ \[at=[0-9]+\]//' "$MAIN/state/mate.status" \ + | grep -Fq "$terminal [key=$key]: child child $terminal: validation finished" \ || fail "$terminal with trailing prose arriving $timing state read was lost: $(cat "$MAIN/state/mate.status" 2>/dev/null)" [ "$(wc -l < "$MAIN/state/mate.status" | tr -d ' ')" = 1 ] \ || fail "$terminal with trailing prose arriving $timing state read was delivered twice" @@ -237,7 +238,8 @@ test_secondmate_unterminated_prose_reports_run_outcome() { printf 'Still going' >> "$MATE/state/child.status" age "$MATE/state/child.status" FM_FAKE_CREW_STATE='failed' run_reconcile "$MATE" --startup - grep -Fq "failed [key=inactive-outcome-mate-child-failed]: inactive terminal child=child" "$MAIN/state/mate.status" \ + sed -E 's/ \[at=[0-9]+\]//' "$MAIN/state/mate.status" \ + | grep -Fq "failed [key=inactive-outcome-mate-child-failed]: inactive terminal child=child" \ || fail "an unterminated prose line withheld a proven failure: $(cat "$MAIN/state/mate.status" 2>/dev/null)" [ "$(outcome_count "$MATE" reported)" = 1 ] || fail "the fallback report did not retain its receipt" age "$MATE/state/child.status" @@ -297,16 +299,16 @@ test_secondmate_ledger_delivery_carries_report_and_failure() { scout_key=$(reported_outcome_key "$MATE" scout 'done') || fail "scout receipt key missing" boom_key=$(reported_outcome_key "$MATE" boom failed) || fail "failed receipt key missing" replaced_key=$(reported_outcome_key "$MATE" replaced-pr 'done') || fail "replacement PR receipt key missing" - grep -Fxq "done [key=$scout_key]: child scout done: report written pr=https://example.test/owner/repo/pull/1 mode=no-mistakes yolo=off report=data/scout/report.md" \ - "$MAIN/state/mate.status" || fail "scout delivery lost its report pointer: $(cat "$MAIN/state/mate.status")" - grep -Fxq "failed [key=$boom_key]: child boom failed: build broke pr=https://example.test/owner/repo/pull/1 mode=no-mistakes yolo=off" \ - "$MAIN/state/mate.status" || fail "failed line was not delivered under the failed verb: $(cat "$MAIN/state/mate.status")" - grep -Fxq "done [key=$replaced_key]: child replaced-pr done: PR https://example.test/owner/repo/pull/22 pr=https://example.test/owner/repo/pull/22 mode=no-mistakes yolo=off" \ - "$MAIN/state/mate.status" || fail "ledger fallback did not prefer the terminal ready line PR: $(cat "$MAIN/state/mate.status")" + sed -E 's/ \[at=[0-9]+\]//' "$MAIN/state/mate.status" | grep -Fxq "done [key=$scout_key]: child scout done: report written pr=https://example.test/owner/repo/pull/1 mode=no-mistakes yolo=off report=data/scout/report.md" \ + || fail "scout delivery lost its report pointer: $(cat "$MAIN/state/mate.status")" + sed -E 's/ \[at=[0-9]+\]//' "$MAIN/state/mate.status" | grep -Fxq "failed [key=$boom_key]: child boom failed: build broke pr=https://example.test/owner/repo/pull/1 mode=no-mistakes yolo=off" \ + || fail "failed line was not delivered under the failed verb: $(cat "$MAIN/state/mate.status")" + sed -E 's/ \[at=[0-9]+\]//' "$MAIN/state/mate.status" | grep -Fxq "done [key=$replaced_key]: child replaced-pr done: PR https://example.test/owner/repo/pull/22 pr=https://example.test/owner/repo/pull/22 mode=no-mistakes yolo=off" \ + || fail "ledger fallback did not prefer the terminal ready line PR: $(cat "$MAIN/state/mate.status")" printf 'working: retrying\ndone: fixed on retry\n' >> "$MATE/state/boom.status" FM_FAKE_CREW_STATE='unknown' run_reconcile "$MATE" boom_key=$(reported_outcome_key "$MATE" boom 'done') || fail "recovered receipt key missing" - grep -Fq "done [key=$boom_key]: child boom done: fixed on retry" "$MAIN/state/mate.status" \ + sed -E 's/ \[at=[0-9]+\]//' "$MAIN/state/mate.status" | grep -Fq "done [key=$boom_key]: child boom done: fixed on retry" \ || fail "a new terminal line after recovery was not delivered" [ "$(grep -c 'child-outcome-boom-' "$MAIN/state/mate.status")" = 2 ] \ || fail "recovery delivered the wrong number of lines: $(cat "$MAIN/state/mate.status")" @@ -317,12 +319,14 @@ test_secondmate_ledger_delivery_carries_report_and_failure() { # task's delivered PR: without a recorded PR, only a terminal line in the # ready-signal shape carries one, and a scout never carries one at all. test_pr_field_requires_recorded_pr_or_ready_signal_line() { - local id prose_key ready_key scout_key + local id prose_key ready_key stamped_key placeholder_key scout_key make_world pr-provenance; bind_secondmate local write_child "$MATE" prose $'working: context in https://example.test/other/repo/pull/33\ndone: cleanup finished' write_child "$MATE" ready 'done: PR https://example.test/owner/repo/pull/44 checks green' + write_child "$MATE" stamped 'done [at=1788576000]: PR https://example.test/owner/repo/pull/66 checks green' + write_child "$MATE" placeholder 'done [at=]: PR https://example.test/owner/repo/pull/77 checks green' write_child "$MATE" lookout 'done: PR https://example.test/owner/repo/pull/55' - for id in prose ready; do + for id in prose ready stamped placeholder; do awk '$0 !~ /^pr=/' "$MATE/state/$id.meta" > "$MATE/state/$id.meta.tmp" mv "$MATE/state/$id.meta.tmp" "$MATE/state/$id.meta" done @@ -332,17 +336,21 @@ test_pr_field_requires_recorded_pr_or_ready_signal_line() { FM_FAKE_CREW_STATE='unknown' run_reconcile "$MATE" prose_key=$(reported_outcome_key "$MATE" prose 'done') || fail "prose receipt key missing" ready_key=$(reported_outcome_key "$MATE" ready 'done') || fail "ready receipt key missing" + stamped_key=$(reported_outcome_key "$MATE" stamped 'done') || fail "stamped ready receipt key missing" + placeholder_key=$(reported_outcome_key "$MATE" placeholder 'done') \ + || fail "unsubstituted-stamp ready receipt key missing" scout_key=$(reported_outcome_key "$MATE" lookout 'done') || fail "scout receipt key missing" - grep -Fxq "done [key=$prose_key]: child prose done: cleanup finished mode=no-mistakes yolo=off" \ - "$MAIN/state/mate.status" \ + sed -E 's/ \[at=[0-9]+\]//' "$MAIN/state/mate.status" | grep -Fxq "done [key=$prose_key]: child prose done: cleanup finished mode=no-mistakes yolo=off" \ || fail "a PR mentioned only in prose was claimed as the delivery: $(cat "$MAIN/state/mate.status")" - grep -Fxq "done [key=$ready_key]: child ready done: PR https://example.test/owner/repo/pull/44 checks green pr=https://example.test/owner/repo/pull/44 mode=no-mistakes yolo=off" \ - "$MAIN/state/mate.status" \ + sed -E 's/ \[at=[0-9]+\]//' "$MAIN/state/mate.status" | grep -Fxq "done [key=$ready_key]: child ready done: PR https://example.test/owner/repo/pull/44 checks green pr=https://example.test/owner/repo/pull/44 mode=no-mistakes yolo=off" \ || fail "a ready-signal terminal line did not carry its PR: $(cat "$MAIN/state/mate.status")" - grep -Fxq "done [key=$scout_key]: child lookout done: PR https://example.test/owner/repo/pull/55 mode=no-mistakes yolo=off" \ - "$MAIN/state/mate.status" \ + sed -E 's/ \[at=[0-9]+\]//' "$MAIN/state/mate.status" | grep -Fxq "done [key=$stamped_key]: child stamped done: PR https://example.test/owner/repo/pull/66 checks green pr=https://example.test/owner/repo/pull/66 mode=no-mistakes yolo=off" \ + || fail "a stamped ready-signal terminal line did not carry its PR: $(cat "$MAIN/state/mate.status")" + sed -E 's/ \[at=[0-9]+\]//' "$MAIN/state/mate.status" | grep -Fxq "done [key=$placeholder_key]: child placeholder done: PR https://example.test/owner/repo/pull/77 checks green pr=https://example.test/owner/repo/pull/77 mode=no-mistakes yolo=off" \ + || fail "a ready-signal line whose stamp was left unsubstituted lost its PR: $(cat "$MAIN/state/mate.status")" + sed -E 's/ \[at=[0-9]+\]//' "$MAIN/state/mate.status" | grep -Fxq "done [key=$scout_key]: child lookout done: PR https://example.test/owner/repo/pull/55 mode=no-mistakes yolo=off" \ || fail "a scout's ready-looking line carried a PR claim: $(cat "$MAIN/state/mate.status")" - pass "pr= requires the recorded PR or a ready-signal terminal line, and never a scout" + pass "pr= requires the recorded PR or a ready-signal terminal line, whatever its stamp, and never a scout" } # If a terminal ledger line lands while the authoritative state read is in @@ -440,7 +448,7 @@ test_secondmate_partial_ledger_line_waits_for_newline() { printf 'ten\n' >> "$MATE/state/child.status" FM_FAKE_CREW_STATE='unknown' run_reconcile "$MATE" key=$(reported_outcome_key "$MATE" child 'done') || fail "completed ledger receipt key missing" - grep -Fq "done [key=$key]: child child done: half written" "$MAIN/state/mate.status" \ + sed -E 's/ \[at=[0-9]+\]//' "$MAIN/state/mate.status" | grep -Fq "done [key=$key]: child child done: half written" \ || fail "the completed line was not delivered once its newline landed" FM_FAKE_CREW_STATE='done' run_reconcile "$MATE" --startup [ "$(wc -l < "$MAIN/state/mate.status" | tr -d ' ')" = 1 ] \ @@ -468,7 +476,7 @@ test_report_subcommand_delivers_and_refuses() { write_child "$MATE" child 'done: final word' run_report "$MATE" child || fail "report refused a deliverable ledger line" key=$(reported_outcome_key "$MATE" child 'done') || fail "report receipt key missing" - grep -Fq "done [key=$key]: child child done: final word" "$MAIN/state/mate.status" \ + sed -E 's/ \[at=[0-9]+\]//' "$MAIN/state/mate.status" | grep -Fq "done [key=$key]: child child done: final word" \ || fail "report did not deliver the child's final line" run_report "$MATE" child || fail "report did not treat an already delivered line as owed nothing" write_child "$MATE" quiet 'working: nothing terminal' @@ -762,8 +770,7 @@ test_watcher_poll_delivers_child_ledger_line_to_parent() { done reap "$pid" key=$(reported_outcome_key "$MATE" child 'done') || fail "watcher ledger receipt key missing" - grep -Fxq "done [key=$key]: child child done: PR https://example.test/owner/repo/pull/1 checks green pr=https://example.test/owner/repo/pull/1 mode=no-mistakes yolo=off" \ - "$MAIN/state/mate.status" \ + sed -E 's/ \[at=[0-9]+\]//' "$MAIN/state/mate.status" | grep -Fxq "done [key=$key]: child child done: PR https://example.test/owner/repo/pull/1 checks green pr=https://example.test/owner/repo/pull/1 mode=no-mistakes yolo=off" \ || fail "the watcher poll did not deliver the child's ledger line to the parent: $(cat "$MAIN/state/mate.status" 2>/dev/null; cat "$WORLD/mate-watch.out")" [ ! -s "$WORLD/forge.log" ] || fail "ledger delivery invoked a forge command" pass "the real watcher poll delivers a child's terminal ledger line to the parent channel" diff --git a/tests/fm-kimi-harness.test.sh b/tests/fm-kimi-harness.test.sh index 6f2aeae7f15..3a35a723b5c 100755 --- a/tests/fm-kimi-harness.test.sh +++ b/tests/fm-kimi-harness.test.sh @@ -704,7 +704,7 @@ test_kimi_unconfirmed_delivery_fails_loudly() { [ "$rc" -ne 0 ] || fail "an unconfirmed kimi delivery should fail" assert_contains "$out" "kimi brief pointer delivery was not confirmed" \ "unconfirmed kimi delivery lacked a loud diagnostic" - assert_grep 'failed: kimi brief pointer delivery was not confirmed' "$HOME_DIR/state/$id.status" \ + assert_grep 'failed: kimi brief pointer delivery was not confirmed' <(sed -E 's/ \[at=[0-9]+\]//' "$HOME_DIR/state/$id.status") \ "unconfirmed kimi delivery did not leave a supervisor-visible failure" pass "fm-spawn: kimi treats a silent pointer drop as a failed spawn" } @@ -934,7 +934,7 @@ test_kimi_stuck_trust_dialog_fails_before_delivery() { [ "$(wc -l < "$CASE_DIR/trust-enter.log" | tr -d ' ')" -gt 1 ] \ || fail "stuck Kimi trust dialog was not re-answered while it stayed on screen" [ ! -s "$CASE_DIR/pointer.log" ] || fail "Kimi pointer was sent through a stuck trust dialog" - assert_grep 'failed: kimi trust dialog did not clear' "$HOME_DIR/state/$id.status" \ + assert_grep 'failed: kimi trust dialog did not clear' <(sed -E 's/ \[at=[0-9]+\]//' "$HOME_DIR/state/$id.status") \ "stuck Kimi trust dialog did not leave a supervisor-visible failure" pass "fm-spawn: a Kimi trust dialog must visibly clear before brief delivery" } diff --git a/tests/fm-pending-reply.test.sh b/tests/fm-pending-reply.test.sh index 1c1353050db..cd31fbaf552 100755 --- a/tests/fm-pending-reply.test.sh +++ b/tests/fm-pending-reply.test.sh @@ -314,7 +314,7 @@ test_second_missed_turn_escalates_once_and_stays_durable() { [ "$(phase_of "$state" "$corr")" = escalated ] || fail "phase should be escalated" status_line=$(tail -1 "$state/hibit.status") case "$status_line" in - "blocked [key=pending-reply-$corr]:"*pending-reply-missed:*pending-reply-id=$corr*) : ;; + "blocked [key=pending-reply-$corr]"*pending-reply-missed:*pending-reply-id=$corr*) : ;; *) fail "parent status should carry one blocked missed-report line"$'\n'"$status_line" ;; esac [ ! -s "$state/.wake-queue" ] || fail "direct escalation must not enqueue a duplicate check wake" @@ -324,7 +324,7 @@ test_second_missed_turn_escalates_once_and_stays_durable() { : fi [ "$(phase_of "$state" "$corr")" = escalated ] || fail "phase must stay escalated" - escalations=$(grep -Fc "blocked [key=pending-reply-$corr]:" "$state/hibit.status") + escalations=$(grep -Fc "blocked [key=pending-reply-$corr]" "$state/hibit.status") [ "$escalations" = 1 ] || fail "missed recovery should publish one escalation, got $escalations" # Durable record retained (never silently expired). rec=$(fm_pending_reply_path "$state" "$corr") @@ -409,7 +409,7 @@ test_escalation_publication_failure_retries() { rmdir "$target" fm_pending_reply_maybe_escalate "$state" "$corr" || fail "escalation retry should succeed" [ "$(phase_of "$state" "$corr")" = escalated ] || fail "successful retry should commit escalation" - escalations=$(grep -Fc "blocked [key=pending-reply-$corr]:" "$target") + escalations=$(grep -Fc "blocked [key=pending-reply-$corr]" "$target") [ "$escalations" = 1 ] || fail "successful retry should publish exactly once, got $escalations" pass "failed escalation publication remains retryable and publishes once" } @@ -429,7 +429,7 @@ test_legacy_escalation_closes_default_decision() { printf 'done [corr=%s]: delayed legacy reply\n' "$corr" >> "$state/hibit.status" fm_pending_reply_try_resolve "$state" "$corr" || fail "legacy reply should resolve its record" - [ "$(grep -Fc "resolved [key=default]: pending-reply-resolved: task=hibit pending-reply-id=$corr" "$state/hibit.status")" -eq 1 ] \ + [ "$(sed -E 's/ \[at=[0-9]+\]//' "$state/hibit.status" | grep -Fc "resolved [key=default]: pending-reply-resolved: task=hibit pending-reply-id=$corr")" -eq 1 ] \ || fail "legacy escalation did not append one guarded default-key resolution" open=$(status_open_decisions "$state/hibit.status") [ -z "$open" ] || fail "resolved legacy escalation remained open: $open" @@ -454,7 +454,7 @@ test_legacy_escalation_does_not_close_taken_default_decision() { printf 'done [corr=%s]: delayed legacy reply\n' "$corr" >> "$state/hibit.status" fm_pending_reply_try_resolve "$state" "$corr" || fail "legacy reply should resolve its record" - if grep -Fq 'resolved [key=default]: pending-reply-resolved:' "$state/hibit.status"; then + if grep -Fq 'resolved [key=default]' "$state/hibit.status"; then fail "legacy escalation emitted an unsafe default-key resolution" fi fm_pending_reply_tick "$state" || fail "legacy close retry failed" @@ -486,7 +486,7 @@ test_foreign_blocker_is_not_selected_as_escalation() { "pending-reply closure cleared the foreign release decision" assert_not_contains "$open" "pending-reply-$corr" \ "genuine keyed escalation remained open" - assert_no_grep 'resolved [key=release]: pending-reply-resolved:' "$state/hibit.status" \ + assert_no_grep 'resolved [key=release]' "$state/hibit.status" \ "foreign release decision was selected as the pending-reply escalation" [ -n "$(fm_pending_reply_get "$rec" escalation_closed_epoch)" ] \ || fail "genuine keyed escalation closure was not recorded" @@ -657,7 +657,7 @@ test_delivery_confirmation_fallback_reconciles() { || fail "delivery uncertainty should use its distinct escalation" fm_pending_reply_tick_one "$state" "$prepared_corr" unknown \ || fail "repeated delivery-unknown tick should be inert" - escalations=$(grep -Fc "blocked [key=pending-reply-$prepared_corr]:" "$state/hibit.status") + escalations=$(grep -Fc "blocked [key=pending-reply-$prepared_corr]" "$state/hibit.status") [ "$escalations" = 1 ] \ || fail "delivery-unknown escalation should publish once, got $escalations" printf 'done [corr=%s]: late report proves delivery\n' "$prepared_corr" >> "$state/hibit.status" @@ -666,7 +666,7 @@ test_delivery_confirmation_fallback_reconciles() { || fail "late report should resolve escalated delivery-unknown" [ "$(fm_pending_reply_get "$prepared_rec" delivered_epoch)" = 5760 ] \ || fail "late report should provide delivery evidence" - escalations=$(grep -Fc "blocked [key=pending-reply-$prepared_corr]:" "$state/hibit.status") + escalations=$(grep -Fc "blocked [key=pending-reply-$prepared_corr]" "$state/hibit.status") [ "$escalations" = 1 ] || fail "late report must not re-escalate delivery-unknown" fm_pending_reply_tick "$state" || fail "resolved late report should remain idempotent" [ "$(phase_of "$state" "$prepared_corr")" = resolved ] \ @@ -708,12 +708,12 @@ test_delivery_confirmation_serializes_with_reconciliation() { entered="$home/mark-delivered.entered" release="$home/mark-delivered.release" fm_pending_reply_mark_delivered() { - local pending_state=$1 pending_corr=$2 epoch=$3 pending_rec phase + local pending_state=$1 pending_corr=$2 pending_epoch=$3 pending_rec phase printf '%s\n' "${BASHPID:-$$}" >> "$calls" : > "$entered" while [ ! -e "$release" ]; do /bin/sleep 0.01; done pending_rec=$(fm_pending_reply_path "$pending_state" "$pending_corr") - fm_pending_reply_set "$pending_rec" delivered_epoch "$epoch" || return 1 + fm_pending_reply_set "$pending_rec" delivered_epoch "$pending_epoch" || return 1 phase=$(fm_pending_reply_get "$pending_rec" phase) [ "$phase" != delivery_unknown ] \ || fm_pending_reply_set "$pending_rec" phase awaiting_report @@ -1259,7 +1259,7 @@ test_mirrored_remote_reply_never_triggers_a_repost() { } test_same_basename_self_home_corr_resolves_on_tick() { - local home state sm_home corr rec parent_status hook_log + local home state sm_home corr rec parent_status hook_log fb out home=$(setup_parent same-basename-repair) state="$home/state" sm_home=$(bind_local_mate "$home" mate) @@ -1297,6 +1297,9 @@ test_same_basename_self_home_corr_resolves_on_tick() { || fail "resolved_epoch must be set after the restatement copy" grep -Fq "corr=$corr" "$parent_status" \ || fail "parent channel must receive the restated corr= line" + if status_line_at_epoch "$(tail -1 "$parent_status")" >/dev/null; then + fail "a relayed copy must not acquire an emission time: $(cat "$parent_status")" + fi if grep -Fq pending-reply-missed "$parent_status"; then fail "same-basename self-home corr must not escalate as pending-reply-missed" fi @@ -1307,6 +1310,31 @@ test_same_basename_self_home_corr_resolves_on_tick() { "$(fm_pending_reply_get "$rec" wrong_home_first_sighting)")" = \ "$sm_home/state/mate.status:1" ] \ || fail "first wrong-home sighting must display the readable mate-home path and line" + fm_pending_reply_restatement_copy_same_basename "$state" "$corr" "$sm_home" \ + || fail "repeated restatement copy should succeed" + fm_parent_channel_report "$sm_home" "$sm_home/state" "$(cat "$sm_home/state/mate.status")" \ + || fail "publication retry of a recovered reply should succeed" + cmp -s "$sm_home/state/mate.status" "$parent_status" \ + || fail "recovery and retries must preserve the legacy reply bytes without duplicates" + fm_write_secondmate_meta "$state/mate.meta" "$sm_home" + fb=$(make_stubs "$home") + out=$(PATH="$fb:$PATH" FM_HOME="$home" "$ROOT/bin/fm-fleet-snapshot.sh" --json) \ + || fail "snapshot of the recovered reply should succeed" + printf '%s' "$out" | jq -e ' + .tasks[] | select(.id == "mate") | .paths.status_log.last_event + | has("age_seconds") and .age_seconds == null + ' >/dev/null || fail "recovered legacy reply must retain an unknown age" + printf '%s' "$out" | jq -e ' + .secondmate_current.records[] | select(.id == "mate") | .parent_event + | has("age_seconds") and .age_seconds == null + ' >/dev/null || fail "secondmate summary must retain the recovered reply's unknown age" + fm_parent_channel_report "$sm_home" "$sm_home/state" 'done: new report' \ + || fail "new publication should succeed" + status_line_at_epoch "$(tail -1 "$parent_status")" >/dev/null \ + || fail "new publication must still receive an emission time" + fm_parent_channel_report "$sm_home" "$sm_home/state" 'done: new report' \ + || fail "new publication retry should succeed" + [ "$(wc -l < "$parent_status")" -eq 2 ] || fail "new publication retry must not duplicate the event" unset FM_PENDING_REPLY_SEND_HOOK pass "same-basename self-home corr= is restated onto the parent channel and resolves" } @@ -1330,7 +1358,7 @@ test_same_basename_reply_resolves_after_recovery_failure() { rec=$(fm_pending_reply_path "$state" "$corr") parent_status=$(fm_pending_reply_get "$rec" parent_status) fm_write_secondmate_meta "$state/mate.meta" "$sm_home" - printf 'done [corr=%s]: answer landed after recovery failure\n' "$corr" \ + printf 'done [corr=%s] [at=11000]: answer landed after recovery failure\n' "$corr" \ > "$sm_home/state/mate.status" fm_pending_reply_tick "$state" @@ -1338,6 +1366,8 @@ test_same_basename_reply_resolves_after_recovery_failure() { || fail "late same-basename reply must resolve before recovery failure escalation" grep -Fq "corr=$corr" "$parent_status" \ || fail "late reply must be restated onto the parent channel" + cmp -s "$sm_home/state/mate.status" "$parent_status" \ + || fail "recovery must preserve the reply's original emission time" if grep -Fq pending-reply-recovery-delivery "$parent_status"; then fail "authorized late reply must prevent recovery delivery escalation" fi @@ -1521,7 +1551,7 @@ test_escalated_undelivered_correlation_stays_retryable() { fm_pending_reply_maybe_escalate "$state" "$corr" || fail "delivery-unknown escalation should fire" [ "$(phase_of "$state" "$corr")" = escalated ] || fail "phase should be escalated" [ -z "$(fm_pending_reply_get "$rec" delivered_epoch)" ] || fail "escalation must not invent delivery" - [ "$(grep -cF "blocked [key=pending-reply-$corr]:" "$state/hibit.status")" = 1 ] \ + [ "$(grep -cF "blocked [key=pending-reply-$corr]" "$state/hibit.status")" = 1 ] \ || fail "delivery-unknown escalation should publish once" fm_pending_reply_corr_reusable "$state" "$corr" hibit \ || fail "an escalated undelivered correlation must stay reusable by its owner" diff --git a/tests/fm-pr-check-security.test.sh b/tests/fm-pr-check-security.test.sh index 62a948f8447..364c2d8fba5 100755 --- a/tests/fm-pr-check-security.test.sh +++ b/tests/fm-pr-check-security.test.sh @@ -1619,7 +1619,7 @@ test_merged_poll_retries_a_failed_upward_report() { set -e [ "$rc" -eq 0 ] || fail "merged-poll-upward-retry: post-recovery retry failed: $(cat "$dir/watch-3.err")" fi - assert_grep "done [key=merged-task-a]: merged task-a $url" "$replies" \ + assert_grep "done [key=merged-task-a]: merged task-a $url" <(sed -E 's/ \[at=[0-9]+\]//' "$replies") \ "merged-poll-upward-retry: repaired binding did not receive the retry" assert_poll_absent "$state" task-a pass "a failed upward merge report keeps its poll armed for repair and retry" @@ -1648,7 +1648,7 @@ test_self_merge_and_poll_publish_one_outcome() { set -e [ "$rc" -eq 0 ] \ || fail "merge-outcome-committed: watcher failed: $(cat "$dir/watch.err")" - [ "$(grep -c -F "done [key=merged-task-a]: merged task-a $url" "$replies")" -eq 1 ] \ + [ "$(sed -E 's/ \[at=[0-9]+\]//' "$replies" | grep -c -F "done [key=merged-task-a]: merged task-a $url")" -eq 1 ] \ || fail "merge-outcome-committed: self and poll reports produced duplicate merge outcomes" assert_no_grep "check: $state/task-a.check.sh: merged" "$state/.wake-queue" \ "merge-outcome-committed: absorbed poll published a second outcome" @@ -1726,7 +1726,7 @@ test_merged_poll_reports_upward_from_a_secondmate_home_once() { check:*task-a.check.sh:*merged) ;; *) fail "merged-poll-upward: the poll's own row was lost: $(cat "$dir/watch-1.out")" ;; esac - assert_grep "done [key=merged-task-a]: merged task-a $url" "$replies" \ + assert_grep "done [key=merged-task-a]: merged task-a $url" <(sed -E 's/ \[at=[0-9]+\]//' "$replies") \ "merged-poll-upward: a merge this home did not perform was never reported upward" [ "$(grep -c -F "$url" "$replies")" -eq 1 ] \ || fail "merged-poll-upward: one detected merge produced more than one upward line" diff --git a/tests/fm-pr-merge.test.sh b/tests/fm-pr-merge.test.sh index a22690455ce..9104bbb730f 100755 --- a/tests/fm-pr-merge.test.sh +++ b/tests/fm-pr-merge.test.sh @@ -1886,13 +1886,13 @@ test_secondmate_merge_reports_upward_once() { FM_TEST_HOME="$case_dir/home" run_pr_merge "$case_dir" task-x1 "$url" \ >"$case_dir/stdout" 2>"$case_dir/stderr" || fail "secondmate-merge-reports: merge failed" - assert_grep "done [key=merged-task-x1]: merged task-x1 $url" "$replies" \ + assert_grep "done [key=merged-task-x1]: merged task-x1 $url" <(sed -E 's/ \[at=[0-9]+\]//' "$replies") \ "secondmate-merge-reports: the landed PR was not reported upward" [ "$(grep -c 'merged-task-x1' "$replies")" -eq 1 ] \ || fail "secondmate-merge-reports: one merge produced more than one upward merge line" # The merge path registers the PR first, and that registration publishes the # child's ready line on the same channel from fm-pr-check itself. - assert_grep "done [key=child-pr-task-x1]: child task-x1 PR ready: $url" "$replies" \ + assert_grep "done [key=child-pr-task-x1]: child task-x1 PR ready: $url" <(sed -E 's/ \[at=[0-9]+\]//' "$replies") \ "secondmate-merge-reports: the registration's ready line was not reported upward" # The same merge again: the forge accepts it in this fixture, so only the @@ -1918,7 +1918,7 @@ test_secondmate_merge_reports_on_the_local_route() { FM_TEST_HOME="$case_dir/home" run_pr_merge "$case_dir" task-x1 "$url" \ >"$case_dir/stdout" 2>"$case_dir/stderr" || fail "secondmate-merge-local: merge failed" - assert_grep "done [key=merged-task-x1]: merged task-x1 $url" "$parent_status" \ + assert_grep "done [key=merged-task-x1]: merged task-x1 $url" <(sed -E 's/ \[at=[0-9]+\]//' "$parent_status") \ "secondmate-merge-local: the landed PR did not reach the parent home's channel" [ ! -e "$case_dir/state/parent-replies.status" ] \ || fail "secondmate-merge-local: a local-route report also wrote the remote reply channel" @@ -1978,7 +1978,7 @@ test_gitlab_merge_reports_upward() { >"$case_dir/stdout" 2>"$case_dir/stderr" || fail "gitlab-merge-reports: merge failed" assert_grep "done [key=merged-task-x1]: merged task-x1 $url" \ - "$case_dir/state/parent-replies.status" \ + <(sed -E 's/ \[at=[0-9]+\]//' "$case_dir/state/parent-replies.status") \ "gitlab-merge-reports: a landed merge request was not reported upward" pass "a landed GitLab merge request is reported upward on the same channel" } diff --git a/tests/fm-remote-backlog-handoff.test.sh b/tests/fm-remote-backlog-handoff.test.sh index fdbb8e806e8..f05c1ad14fe 100755 --- a/tests/fm-remote-backlog-handoff.test.sh +++ b/tests/fm-remote-backlog-handoff.test.sh @@ -380,7 +380,7 @@ bash -c '. "$1"; fm_pending_reply_tick "$2"' _ "$ROOT/bin/fm-pending-reply-lib.s || fail "watcher tick did not escalate the undelivered wake, got $(grep '^phase=' "$escalated_rec")" [ -z "$(grep '^delivered_epoch=' "$escalated_rec" | cut -d= -f2-)" ] \ || fail "escalation must not invent a delivery for the undelivered wake" -[ "$(grep -cF "blocked [key=pending-reply-$escalated_corr]:" "$PARENT/state/ios.status")" -eq 1 ] \ +[ "$(grep -cF "blocked [key=pending-reply-$escalated_corr]" "$PARENT/state/ios.status")" -eq 1 ] \ || fail "undelivered wake escalation was not published exactly once" set +e handoff_env "$ROOT/bin/fm-backlog-handoff.sh" --resume-pending > "$TMP_ROOT/wake-escalated-resume.out" 2>&1 @@ -398,7 +398,7 @@ assert_absent "$PARENT/state/.backlog-handoff-ios.wake-pending" "escalated wake || fail "successful wake retry did not confirm delivery on the same correlation" [ "$(grep '^phase=' "$escalated_rec" | cut -d= -f2-)" = awaiting_report ] \ || fail "delivered wake retry did not return the correlation to awaiting its report" -[ "$(grep -cF "blocked [key=pending-reply-$escalated_corr]:" "$PARENT/state/ios.status")" -eq 1 ] \ +[ "$(grep -cF "blocked [key=pending-reply-$escalated_corr]" "$PARENT/state/ios.status")" -eq 1 ] \ || fail "wake retry duplicated the published escalation" write_backlog '- [ ] after-escalated - next handoff flows once the escalated wake is retried (repo: alpha)' handoff_env "$ROOT/bin/fm-backlog-handoff.sh" ios after-escalated >/dev/null \ diff --git a/tests/fm-remote-reply.test.sh b/tests/fm-remote-reply.test.sh index 216ff8e223e..49cd55ad37b 100755 --- a/tests/fm-remote-reply.test.sh +++ b/tests/fm-remote-reply.test.sh @@ -95,7 +95,7 @@ assert_contains "$out" "armed: $SID offset=0" "remote reply source was not armed remote_env "$ROOT/bin/fm-procevent.sh" start "$SID" > "$TMP_ROOT/start-one.out" 2>&1 & RUNNER=$! wait_for "$CLAIMS/$SID.claim" || fail "process-event runner never claimed the remote reply source" -printf 'done [corr=0123456789abcdef]: build verified report=data/reply/report.md\n' \ +printf 'done [corr=0123456789abcdef] [at=1700000000]: build verified report=data/reply/report.md\n' \ >> "$REMOTE/state/parent-replies.status" wait "$RUNNER" || fail "remote reply source failed to capture its first delta" RESULT=$(find "$PARENT/state/procevent-inbox" -name "$SID.1.result" -print -quit 2>/dev/null) @@ -160,6 +160,8 @@ cmp -s "$SOURCE_AFTER" "$REMOTE/state/parent-replies.status" \ || fail "handling consumed or rewrote the remote append-only log" expected_offset=$(LC_ALL=C wc -c < "$REMOTE/state/parent-replies.status" | tr -d ' ') assert_grep "offset=$expected_offset" "$PARENT/state/remote-replies/ios.cursor" "reply cursor did not advance to the committed delta" +assert_grep 'done [corr=0123456789abcdef] [at=1700000000]: build verified' "$PARENT/state/ios.status" \ + "relay replaced the source event time with observation time" pass "ingest appends one validated line, fetches its document, and advances the cursor" out=$(remote_env "$ADAPTER" handle ios 1 "$RESULT") @@ -198,6 +200,8 @@ assert_contains "$out" 'ingested: ios appended=0' "earlier generation did not re assert_contains "$out" 'handled: remote-reply-ios 2' "earlier generation remained unacknowledged after later cursor advancement" [ "$(grep -cF 'working [corr=1111111111111111]' "$PARENT/state/ios.status")" -eq 1 ] \ || fail "earlier generation replay duplicated its parent status" +grep -Fxq 'working [corr=1111111111111111]: second generation' "$PARENT/state/ios.status" \ + || fail "relay invented an emission time for a legacy source event" pass "later generations cannot invalidate an unacknowledged ingested result" # The channel mirrors the remote mate's content-bearing status lines at most once @@ -216,6 +220,8 @@ fm_pending_reply_mark_delivered "$PARENT/state" "$PENDING_CORR" \ { printf 'working [key=version-audit]: family --version audit complete (data/reply/prose-only.md)\n' printf 'needs-decision [key=rough-cut-version]: implement --version or retire the tool\n' + printf 'needs-decision [at=1700000000]: which base branch?\n' + printf 'needs-decision [at=1700086400]: which base branch?\n' printf 'done [corr=%s]: release chain audited\n' "$PENDING_CORR" } >> "$REMOTE/state/parent-replies.status" remote_env "$ROOT/bin/fm-procevent.sh" start "$SID" >/dev/null \ @@ -226,6 +232,10 @@ remote_env "$ADAPTER" handle ios 4 "$RESULT_FOUR" > "$TMP_ROOT/handle-mirror.out assert_grep 'working [key=version-audit]' "$PARENT/state/ios.status" "an uncorrelated progress line never reached the parent stream" assert_grep 'needs-decision [key=rough-cut-version]' "$PARENT/state/ios.status" "a newly raised remote decision never reached the parent stream" assert_grep "done [corr=$PENDING_CORR]" "$PARENT/state/ios.status" "the correlated answer sharing the delta was lost" +for epoch in 1700000000 1700086400; do + grep -Fxq "needs-decision [at=$epoch]: which base branch?" "$PARENT/state/ios.status" \ + || fail "relay discarded a distinct event with identical text and a different time" +done mirror_offset=$(LC_ALL=C wc -c < "$REMOTE/state/parent-replies.status" | tr -d ' ') assert_grep "offset=$mirror_offset" "$PARENT/state/remote-replies/ios.cursor" \ "the cursor did not advance past an uncorrelated line" @@ -258,6 +268,16 @@ remote_env "$ADAPTER" handle ios 4 "$RESULT_FOUR" >/dev/null 2>&1 || true assert_grep "offset=$mirror_offset" "$PARENT/state/remote-replies/ios.cursor" \ "replaying the mirrored delta moved the cursor" pass "a replayed mirrored delta is idempotent in both the stream and the cursor" +[ "$(grep -Fc ': which base branch?' "$PARENT/state/ios.status")" -eq 2 ] \ + || fail "replaying a delta duplicated distinct timed requests" +if [ "${FM_TEST_EVIDENCE:-0}" = 1 ]; then + printf 'Remote source status records:\n' + cat "$REMOTE/state/parent-replies.status" + printf '\nParent status after handling and replaying generation 4:\n' + cat "$PARENT/state/ios.status" + printf '\nCommitted remote cursor:\n' + cat "$PARENT/state/remote-replies/ios.cursor" +fi # Bytes crossing a machine boundary are normalized, never dropped: a control # character cannot make the parent's status file unsafe and cannot stop the @@ -425,7 +445,7 @@ cmp -s "$REMOTE/data/remote-secondmates/nested/data/reply/report.md" \ || fail "a nested remote report this mate holds was not relayed" assert_grep 'nested report=data/remote-secondmates/ios/data/remote-secondmates/nested/data/reply/report.md foreign report=data/remote-secondmates/other/data/reply/report.md' "$PARENT/state/ios.status" \ "the nested pointer was not rewritten or the undeliverable foreign pointer was changed" -assert_grep 'note: remote document did not transfer for ios: data/remote-secondmates/other/data/reply/report.md - ' "$PARENT/state/ios.status" \ +assert_grep 'note: remote document did not transfer for ios: data/remote-secondmates/other/data/reply/report.md - ' <(sed -E 's/ \[at=[0-9]+\]//' "$PARENT/state/ios.status") \ "an undeliverable foreign pointer left no note" assert_no_document_decision "an undeliverable foreign pointer raised a document decision" mirrored_cursor_is_current "an undeliverable foreign pointer prevented the cursor from advancing" @@ -461,8 +481,10 @@ pass "the reported incident raises no standing decision and still delivers the r # A structured offer the reader cannot deliver fails open with its own reason. # Offered again twice in one delta, the unchanged note is not repeated. mirror_lines 'reply [corr=4444444444444444]: dispatched a scout report=data/reply/never-written.md' -assert_grep 'note: remote document did not transfer for ios: data/reply/never-written.md - file is not a non-symlink regular file' "$PARENT/state/ios.status" \ +assert_grep 'note: remote document did not transfer for ios: data/reply/never-written.md - file is not a non-symlink regular file' <(sed -E 's/ \[at=[0-9]+\]//' "$PARENT/state/ios.status") \ "an undeliverable structured offer left no note carrying the reader's reason" +status_line_at_epoch "$(grep -E '^note( \[at=[0-9]+\])?: remote document did not transfer for ios: data/reply/never-written\.md' "$PARENT/state/ios.status")" >/dev/null \ + || fail "new remote document note has unknown emission time" assert_grep 'dispatched a scout report=data/reply/never-written.md' "$PARENT/state/ios.status" \ "an undeliverable offer's line was not mirrored with its own pointer intact" assert_no_document_decision "an undeliverable structured offer raised a document decision" @@ -470,7 +492,7 @@ mirrored_cursor_is_current "an undeliverable structured offer held the cursor ba mirror_lines \ 'reply [corr=4444444444444444]: still writing report=data/reply/never-written.md' \ 'reply [corr=4444444444444444]: same, report=data/reply/never-written.md' -[ "$(grep -cF 'note: remote document did not transfer for ios: data/reply/never-written.md' "$PARENT/state/ios.status")" -eq 1 ] \ +[ "$(sed -E 's/ \[at=[0-9]+\]//' "$PARENT/state/ios.status" | grep -cF 'note: remote document did not transfer for ios: data/reply/never-written.md')" -eq 1 ] \ || fail "re-offering the same undeliverable document repeated its note" assert_no_document_decision "re-offering an undeliverable document raised a document decision" pass "an undeliverable structured offer fails open with one note and never a decision" @@ -479,7 +501,7 @@ pass "an undeliverable structured offer fails open with one note and never a dec # reader bounds document size, and that refusal is visible by its own reason. head -c 300000 /dev/zero | tr '\0' 'x' > "$REMOTE/data/reply/big.md" mirror_lines 'done [key=big-report]: oversize deliverable report=data/reply/big.md' -assert_grep 'note: remote document did not transfer for ios: data/reply/big.md - file exceeds max-bytes' "$PARENT/state/ios.status" \ +assert_grep 'note: remote document did not transfer for ios: data/reply/big.md - file exceeds max-bytes' <(sed -E 's/ \[at=[0-9]+\]//' "$PARENT/state/ios.status") \ "an oversize document's refusal did not surface with its reason" assert_absent "$PARENT/data/remote-secondmates/ios/data/reply/big.md" \ "a refused oversize document was stored locally anyway" @@ -780,6 +802,12 @@ assert_absent "$PARENT/state/procevent/$SID.source" "continuity break was re-arm remote_env "$ADAPTER" ingest ios "$RESULT_TWELVE" >/dev/null 2>&1 || true [ "$(grep -cF 'blocked [key=remote-reply-continuity-ios]' "$PARENT/state/ios.status")" -eq 1 ] \ || fail "continuity replay duplicated the escalation" +status_line_at_epoch "$(grep -F 'blocked [key=remote-reply-continuity-ios]' "$PARENT/state/ios.status")" >/dev/null \ + || fail "new continuity escalation has unknown emission time" +if [ "${FM_TEST_EVIDENCE:-0}" = 1 ]; then + printf '\nNew continuity escalation after ingest retry:\n' + grep -F 'blocked [key=remote-reply-continuity-ios]' "$PARENT/state/ios.status" +fi pass "truncation is detected, escalated once, and not silently rebased" rm -f "$PARENT/state/procevent-inbox/$SID.$GEN.handled" diff --git a/tests/fm-rovo-harness.test.sh b/tests/fm-rovo-harness.test.sh index cfa6476f40c..dd405872df9 100644 --- a/tests/fm-rovo-harness.test.sh +++ b/tests/fm-rovo-harness.test.sh @@ -4,6 +4,8 @@ set -u # shellcheck source=tests/lib.sh . "$(dirname "${BASH_SOURCE[0]}")/lib.sh" +# shellcheck source=bin/fm-classify-lib.sh +. "$ROOT/bin/fm-classify-lib.sh" # bin/fm-harness.sh checks verified ENV markers before ancestry, but that # ordering settles the marker layer only: a structural (comm-strength) @@ -279,7 +281,7 @@ test_rovo_effort_high_sets_config_override() { } test_rovo_readiness_gate_precedes_pointer() { - local id rec out rc + local id rec out rc line id="rovo-not-ready-z3-$$" rec=$(make_spawn_case not-ready "$id") read_spawn_record "$rec" @@ -289,11 +291,18 @@ test_rovo_readiness_gate_precedes_pointer() { [ "$rc" -ne 0 ] || fail "rovo spawn without a ready signal should fail" assert_contains "$out" "rovo did not show a verified ready signal" \ "rovo readiness failure lacked a loud diagnostic" - assert_grep 'failed: rovo did not show a verified ready signal' "$HOME_DIR/state/$id.status" \ + line=$(cat "$HOME_DIR/state/$id.status") + [ "$(status_line_verb "$line")" = failed ] || fail "rovo readiness failure lost its failed verb" + assert_contains "$(status_line_note "$line")" 'rovo did not show a verified ready signal' \ "rovo readiness failure did not leave a supervisor-visible failure" [ ! -s "$CASE_DIR/pointer.log" ] || fail "rovo pointer was sent before an observable ready signal" grep -q "kill-window.*fm-$id" "$CASE_DIR/tmux-calls.log" \ || fail "a failed rovo readiness gate must tear down the exact endpoint it created instead of leaking an orphaned --yolo process" + status_line_at_epoch "$line" >/dev/null \ + || fail "new rovo spawn failure has unknown emission time: $line" + if [ "${FM_TEST_EVIDENCE:-0}" = 1 ]; then + printf 'Rovo readiness failure CLI output:\n%s\nPersisted status:\n%s\n' "$out" "$line" + fi pass "fm-spawn: rovo never sends the brief pointer before an observable ready signal, and tears down the created endpoint on failure" } @@ -311,7 +320,9 @@ test_rovo_unconfirmed_delivery_fails_loudly() { [ -n "$pointer" ] || fail "rovo never typed the pointer before the delivery gate" assert_contains "$out" "rovo brief pointer delivery was not confirmed" \ "unconfirmed rovo delivery lacked a loud diagnostic" - assert_grep 'failed: rovo brief pointer delivery was not confirmed' "$HOME_DIR/state/$id.status" \ + [ "$(status_line_verb "$(cat "$HOME_DIR/state/$id.status")")" = failed ] \ + || fail "unconfirmed rovo delivery lost its failed verb" + assert_contains "$(status_line_note "$(cat "$HOME_DIR/state/$id.status")")" 'rovo brief pointer delivery was not confirmed' \ "unconfirmed rovo delivery did not leave a supervisor-visible failure" grep -q "kill-window.*fm-$id" "$CASE_DIR/tmux-calls.log" \ || fail "an unconfirmed rovo delivery must tear down the exact endpoint it created instead of leaking an orphaned --yolo process" diff --git a/tests/fm-send-remote-delivery.test.sh b/tests/fm-send-remote-delivery.test.sh index fb52e7b21e8..8e37db6d529 100755 --- a/tests/fm-send-remote-delivery.test.sh +++ b/tests/fm-send-remote-delivery.test.sh @@ -516,7 +516,7 @@ test_remote_resolve_key_closes_at_enqueue() { send_env "$fb" "$home" "$ssh_log" \ "$SEND" rsm --resolve-key upgrade-window "the weekend, freeze Friday" >/dev/null 2>&1 || rc=$? expect_code 0 "$rc" "a durably recorded remote answer must exit 0" - grep -F 'resolved [key=upgrade-window]: answered: the weekend, freeze Friday' "$home/state/rsm.status" >/dev/null \ + sed -E 's/ \[at=[0-9]+\]//' "$home/state/rsm.status" | grep -qF 'resolved [key=upgrade-window]: answered: the weekend, freeze Friday' \ || fail "a recorded remote answer must close the decision at enqueue: $(cat "$home/state/rsm.status")" out=$(drain_out "$home") if printf '%s' "$out" | grep -F 'OPEN DECISIONS' >/dev/null; then diff --git a/tests/fm-send-resolve-key.test.sh b/tests/fm-send-resolve-key.test.sh index 9c57b0ea1a3..3141367c1c7 100755 --- a/tests/fm-send-resolve-key.test.sh +++ b/tests/fm-send-resolve-key.test.sh @@ -135,7 +135,7 @@ test_answer_send_closes_open_decision() { grep -qF "go with REST" "$home/state/t1.inbox/001.msg" \ || fail "the answer text should reach the worker's durable inbox record" assert_contains "$(cat "$log")" "Firstmate instruction waiting" "the doorbell should be rung for the answer" - grep -F 'resolved [key=api-shape]: answered: go with REST' "$home/state/t1.status" >/dev/null \ + sed -E 's/ \[at=[0-9]+\]//' "$home/state/t1.status" | grep -qF 'resolved [key=api-shape]: answered: go with REST' \ || fail "fm-send did not append the closing resolved line:"$'\n'"$(cat "$home/state/t1.status")" # The drain folded the worker's `working:` line but never listed it, so the # close must leave the file for the watcher instead of marking it seen. @@ -171,7 +171,7 @@ test_answer_close_is_self_announced() { run_send "$fb" "$home" "$log" t9 --resolve-key port-choice "use 9090"; rc=$? expect_code 0 "$rc" "the answer send should succeed" - grep -F 'resolved [key=port-choice]: answered: use 9090' "$home/state/t9.status" >/dev/null \ + sed -E 's/ \[at=[0-9]+\]//' "$home/state/t9.status" | grep -qF 'resolved [key=port-choice]: answered: use 9090' \ || fail "the closing resolved line is missing" FM_STATE_OVERRIDE="$home/state" bash -c ' . "$1"; fm_wake_signal_seen_current "$2" "$3" @@ -206,7 +206,7 @@ test_colon_first_key_position_is_answerable() { run_send "$fb" "$home" "$log" t8 --resolve-key seam-max-bound "cap it at 4"; rc=$? expect_code 0 "$rc" "answering a colon-first stated key should succeed, not refuse as unknown" - grep -F 'resolved [key=seam-max-bound]: answered: cap it at 4' "$home/state/t8.status" >/dev/null \ + sed -E 's/ \[at=[0-9]+\]//' "$home/state/t8.status" | grep -qF 'resolved [key=seam-max-bound]: answered: cap it at 4' \ || fail "the closing resolved line is missing:"$'\n'"$(cat "$home/state/t8.status")" out=$(drain_out "$home") @@ -308,7 +308,7 @@ test_failed_ring_still_closes_at_enqueue() { expect_code 0 "$rc" "a failed doorbell must not fail the durably enqueued answer" grep -qF 'token is in the vault now' "$home/state/t5.inbox/001.msg" \ || fail "the answer must be durably recorded despite the failed ring" - grep -F 'resolved [key=creds]' "$home/state/t5.status" >/dev/null \ + sed -E 's/ \[at=[0-9]+\]//' "$home/state/t5.status" | grep -qF 'resolved [key=creds]: answered: token is in the vault now' \ || fail "the enqueued answer must close the decision at answer time: $(cat "$home/state/t5.status")" out=$(drain_out "$home") if printf '%s' "$out" | grep -F '[key=creds]' >/dev/null; then @@ -470,7 +470,7 @@ test_remote_secondmate_answer_closes_locally() { expect_code 0 "$rc" "a remote secondmate answer send should succeed" assert_grep 'fm-remote-entrypoint.sh' "$ssh_log" \ "the answer message should cross the remote transport" - grep -F 'resolved [key=upgrade-window]: answered: the weekend, freeze Friday' "$home/state/rsm.status" >/dev/null \ + sed -E 's/ \[at=[0-9]+\]//' "$home/state/rsm.status" | grep -qF 'resolved [key=upgrade-window]: answered: the weekend, freeze Friday' \ || fail "the remote answer did not close the local ledger: $(cat "$home/state/rsm.status")" out=$(drain_out "$home") if printf '%s' "$out" | grep -F 'OPEN DECISIONS' >/dev/null; then @@ -505,7 +505,7 @@ test_remote_reply_corr_tag_does_not_block_resolve_key() { FM_SSH_BIN="$fb/fake-ssh" FM_SSH_LOG="$ssh_log" FM_FAKE_SSH_RC=0 \ "$SEND" rsm --resolve-key loan-installment-cadence-amount "monthly" >/dev/null 2>&1; rc=$? expect_code 0 "$rc" "answering a corr-tagged remote decision should succeed, not refuse as unknown" - grep -F 'resolved [key=loan-installment-cadence-amount]: answered: monthly' "$home/state/rsm.status" >/dev/null \ + sed -E 's/ \[at=[0-9]+\]//' "$home/state/rsm.status" | grep -qF 'resolved [key=loan-installment-cadence-amount]: answered: monthly' \ || fail "the closing resolved line is missing:"$'\n'"$(cat "$home/state/rsm.status")" out=$(drain_out "$home") @@ -606,7 +606,7 @@ test_reserved_pending_reply_key_closes_through_resolve_key() { grep -F "pending-reply-resolved: task=mate pending-reply-id=$corr via=operator-resolve-key" \ "$home/state/mate.status" >/dev/null \ || fail "the operator close did not write the owning library's close note:"$'\n'"$(cat "$home/state/mate.status")" - if grep -E "resolved \[key=$key\]: answered:" "$home/state/mate.status" >/dev/null; then + if grep -E "resolved \[key=$key\]( \[at=[0-9]+\])?: answered:" "$home/state/mate.status" >/dev/null; then fail "the operator close still wrote a bare answered: note that the fold ignores:"$'\n'"$(cat "$home/state/mate.status")" fi @@ -705,6 +705,32 @@ test_long_decision_key_refuses_before_send() { pass "fm-send --resolve-key: an overlong decision key refuses before sending" } +# The cap bounds the line that is actually APPENDED. The self-announced append +# stamps each close with its emission time, so a cap measured before the stamp +# lets the stored line overrun it and every 220-capped rendering downstream +# silently loses that much real note text. +test_stamped_close_line_stays_within_the_status_line_cap() { + local dir fb log home rc answer line + dir="$TMP_ROOT/cap-with-stamp"; mkdir -p "$dir" + fb=$(make_stubs "$dir"); log="$dir/send.log" + home=$(setup_home cap-with-stamp) + fm_write_meta "$home/state/t1.meta" "window=sess:fm-t1" "kind=ship" + printf 'needs-decision [key=api-shape]: REST or gRPC\n' > "$home/state/t1.status" + answer=$(printf 'x%.0s' {1..400}) + + run_send "$fb" "$home" "$log" t1 --resolve-key api-shape "$answer"; rc=$? + expect_code 0 "$rc" "answering with an over-long note should succeed, not refuse" + line=$(grep -F 'resolved [key=api-shape]' "$home/state/t1.status") \ + || fail "the closing resolved line is missing:"$'\n'"$(cat "$home/state/t1.status")" + case "$line" in + *' [at='*']: '*) : ;; + *) fail "the appended close carries no emission stamp: $line" ;; + esac + [ "${#line}" -le 220 ] \ + || fail "the appended close is ${#line} characters, past the 220-character cap: $line" + pass "fm-send --resolve-key: a stamped close line stays inside the status-line cap" +} + test_failed_close_recovery_command_is_shell_safe() { local dir fb log home err marker answer rc diagnostic manual out dir="$TMP_ROOT/manual-close"; mkdir -p "$dir" @@ -792,7 +818,7 @@ test_decision_answer_partition_relocates_under_the_record() { # ordinary lease guard alone. FM_SUPERVISION_ACTOR=branch run_send "$fb" "$home" "$log" t1 --resolve-key token "refreshed the token; resume"; rc=$? expect_code 0 "$rc" "an attended branch resolving a blocker is ordinary steering" - grep -qF 'resolved [key=token]: answered: refreshed the token; resume' "$home/state/t1.status" \ + sed -E 's/ \[at=[0-9]+\]//' "$home/state/t1.status" | grep -qF 'resolved [key=token]: answered: refreshed the token; resume' \ || fail "the branch's blocker answer did not close the key:"$'\n'"$(cat "$home/state/t1.status")" grep -qF "refreshed the token; resume" "$home/state/t1.inbox/001.msg" \ || fail "the branch's blocker answer did not reach the worker's inbox" @@ -804,7 +830,7 @@ test_decision_answer_partition_relocates_under_the_record() { FM_SUPERVISION_ACTOR=branch "$SEND" t1 --resolve-key api-shape "go with REST" 2>&1); rc=$? expect_code 0 "$rc" "under the away-posture record the branch's decision answer must be sent: $out" assert_contains "$out" "main is parked" "the relocation did not announce itself" - grep -qF 'resolved [key=api-shape]: answered: go with REST' "$home/state/t1.status" \ + sed -E 's/ \[at=[0-9]+\]//' "$home/state/t1.status" | grep -qF 'resolved [key=api-shape]: answered: go with REST' \ || fail "the relocated answer did not close the decision:"$'\n'"$(cat "$home/state/t1.status")" grep -qF "go with REST" "$home/state/t1.inbox/002.msg" \ || fail "the relocated answer did not reach the worker's inbox" @@ -818,7 +844,7 @@ test_decision_answer_partition_relocates_under_the_record() { printf 'needs-decision [key=db]: postgres or sqlite\n' >> "$home/state/t1.status" run_send "$fb" "$home" "$log" t1 --resolve-key db "postgres"; rc=$? expect_code 0 "$rc" "main answering a decision attended is unaffected by the partition" - grep -qF 'resolved [key=db]: answered: postgres' "$home/state/t1.status" \ + sed -E 's/ \[at=[0-9]+\]//' "$home/state/t1.status" | grep -qF 'resolved [key=db]: answered: postgres' \ || fail "main's attended decision answer did not close the key" pass "fm-send --resolve-key: a decision answer refuses the attended branch before sending, a blocked: key stays steering, and the away-posture record relocates the answer" } @@ -842,6 +868,7 @@ test_reserved_pending_reply_key_closes_through_resolve_key test_unrelated_writer_cannot_close_or_hijack_reserved_key test_unclosable_reserved_key_refuses_before_send test_long_decision_key_refuses_before_send +test_stamped_close_line_stays_within_the_status_line_cap test_failed_close_recovery_command_is_shell_safe test_remote_reserved_pending_reply_key_closes_locally test_decision_answer_partition_relocates_under_the_record diff --git a/tests/fm-tangle-guard.test.sh b/tests/fm-tangle-guard.test.sh index d59864e0dae..6f80e3078e0 100755 --- a/tests/fm-tangle-guard.test.sh +++ b/tests/fm-tangle-guard.test.sh @@ -131,7 +131,8 @@ test_brief_assertion_precedes_branch() { FM_HOME="$home" "$ROOT/bin/fm-brief.sh" tangle-brief-cc3 alpha --mode no-mistakes >/dev/null 2>&1 brief="$home/data/tangle-brief-cc3/brief.md" assert_present "$brief" "brief was not scaffolded" - assert_grep "blocked: launched in primary checkout, not an isolated worktree" "$brief" \ + # shellcheck disable=SC2016 # The generated instruction keeps the stamp literal. + assert_grep 'blocked [at=]: launched in primary checkout, not an isolated worktree' "$brief" \ "brief is missing the isolation blocked-status contract" assert_grep "The path check is authoritative" "$brief" \ "brief must make the path check authoritative" diff --git a/tests/fm-task-delivery.test.sh b/tests/fm-task-delivery.test.sh index 7457a5d87b5..51dbf4bd583 100755 --- a/tests/fm-task-delivery.test.sh +++ b/tests/fm-task-delivery.test.sh @@ -374,7 +374,8 @@ STUB "promoted no-mistakes worker did not receive the ask-user escalation rule" assert_grep "write only the ask-user findings, verbatim and unparaphrased (id, severity, file, line, description, authority)" "$payload" \ "promoted no-mistakes worker did not receive the ask-user-only snapshot contract" - assert_grep 'needs-decision [key=nm--]: ask-user findings=,,... file='"$home/data/promote-dod-no-mistakes/nm--findings.txt" "$payload" \ + # shellcheck disable=SC2016 # single quotes are deliberate: the placeholders must stay literal + assert_grep 'needs-decision [at=] [key=nm--]: ask-user findings=,,... file='"$home/data/promote-dod-no-mistakes/nm--findings.txt" "$payload" \ "promoted no-mistakes worker did not receive the structured escalation event" assert_grep "NEVER pass \`--yes\` (or \`-y\`)" "$payload" \ "promoted no-mistakes worker did not receive the --yes prohibition" diff --git a/tests/fm-teardown.test.sh b/tests/fm-teardown.test.sh index 43fa7df5543..04f7f096231 100755 --- a/tests/fm-teardown.test.sh +++ b/tests/fm-teardown.test.sh @@ -1852,7 +1852,7 @@ test_secondmate_pr_registration_publishes_ready_line() { PATH="$case_dir/fakebin:$PATH" "$PR_CHECK" task-x1 "$url" > "$case_dir/pr-check.out" 2> "$case_dir/pr-check.err" \ || fail "mate-pr-ready: fm-pr-check failed: $(cat "$case_dir/pr-check.err")" grep -q '^armed:' "$case_dir/pr-check.out" || fail "mate-pr-ready: poll was not armed" - assert_grep "done [key=child-pr-task-x1]: child task-x1 PR ready: $url mode=no-mistakes" "$channel" \ + assert_grep "done [key=child-pr-task-x1]: child task-x1 PR ready: $url mode=no-mistakes" <(sed -E 's/ \[at=[0-9]+\]//' "$channel") \ "mate-pr-ready: the ready line did not reach the parent channel" ! grep -q '^actionable:' "$case_dir/pr-check.err" \ || fail "mate-pr-ready: registration reported a channel problem: $(cat "$case_dir/pr-check.err")" @@ -1897,7 +1897,7 @@ test_secondmate_home_teardown_delivers_final_line_or_refuses() { rc=$? set -e expect_code 0 "$rc" "mate-teardown-delivers: teardown should succeed: $(cat "$case_dir/stderr")" - grep -Eq '^done \[key=child-outcome-task-x1-done-[0-9a-f]{8}\]: child task-x1 done: PR https://github.com/example/repo/pull/9 checks green pr=https://github.com/example/repo/pull/9 mode=local-only$' "$channel" \ + sed -E 's/ \[at=[0-9]+\]//' "$channel" | grep -Eq '^done \[key=child-outcome-task-x1-done-[0-9a-f]{8}\]: child task-x1 done: PR https://github.com/example/repo/pull/9 checks green pr=https://github.com/example/repo/pull/9 mode=local-only$' \ || fail "mate-teardown-delivers: the final ledger line did not reach the parent: $(cat "$channel" 2>/dev/null)" [ ! -e "$case_dir/state/task-x1.meta" ] || fail "mate-teardown-delivers: teardown left the task record" @@ -1940,7 +1940,7 @@ test_secondmate_home_teardown_delivers_final_line_or_refuses() { rc=$? set -e expect_code 0 "$rc" "mate-teardown-refuses: rerun after repair should succeed: $(cat "$case_dir/stderr2")" - grep -Eq '^done \[key=child-outcome-task-x1-done-[0-9a-f]{8}\]: child task-x1 done: PR https://github.com/example/repo/pull/9 checks green' "$channel" \ + sed -E 's/ \[at=[0-9]+\]//' "$channel" | grep -Eq '^done \[key=child-outcome-task-x1-done-[0-9a-f]{8}\]: child task-x1 done: PR https://github.com/example/repo/pull/9 checks green' \ || fail "mate-teardown-refuses: the rerun did not deliver the final line" [ ! -e "$case_dir/state/task-x1.meta" ] || fail "mate-teardown-refuses: rerun left the task record" pass "a secondmate home's teardown delivers the child's final line or refuses until it can" diff --git a/tests/fm-wake-queue.test.sh b/tests/fm-wake-queue.test.sh index 14cb8cc8d89..7924fd5b1c0 100755 --- a/tests/fm-wake-queue.test.sh +++ b/tests/fm-wake-queue.test.sh @@ -1605,7 +1605,7 @@ test_self_announced_append_guards() { run_wake_lib fm_wake_status_append_self_announced "$state" "$status" \ 'resolved [key=k1]: answered: closed by this home' \ || fail "self-announced append on an announced file was not suppressed (rc=$?)" - grep -Fq 'resolved [key=k1]: answered: closed by this home' "$status" \ + sed -E 's/ \[at=[0-9]+\]//' "$status" | grep -Fq 'resolved [key=k1]: answered: closed by this home' \ || fail "the suppressed close was not appended" run_wake_lib fm_wake_signal_seen_current "$state" "$status" \ || fail "the self-announced close left unannounced bytes behind" @@ -1621,7 +1621,7 @@ test_self_announced_append_guards() { run_wake_lib fm_wake_status_append_self_announced "$state" "$status" \ 'resolved [key=k1]: answered: second close' || rc=$? [ "$rc" -eq 1 ] || fail "a close over pending foreign bytes did not fail toward waking (rc=$rc)" - grep -Fq 'resolved [key=k1]: answered: second close' "$status" \ + sed -E 's/ \[at=[0-9]+\]//' "$status" | grep -Fq 'resolved [key=k1]: answered: second close' \ || fail "the fail-toward-waking close was not appended" run_wake_lib fm_wake_signal_seen_current "$state" "$status" \ && fail "a close over pending foreign bytes swallowed the pending wake" From c443d8c2596a5a2acaa1780fcd09a2e17072a27a Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Sun, 20 Sep 2026 15:09:54 -0700 Subject: [PATCH 004/168] fix(bin): unify Lavish host and disconnect handling (#5060) * fix: ship clean Lavish host fixes * no-mistakes(review): Fix Lavish classifications and fail-closed host loading * no-mistakes(review): Restore Lavish host state across retries and launches * no-mistakes(review): Preserve destination Lavish host when configuration is absent * no-mistakes(document): Document Lavish status and host guarantees --- .agents/skills/process-event-sources/SKILL.md | 11 +- .../skills/stuck-crewmate-recovery/SKILL.md | 2 + AGENTS.md | 3 +- bin/fm-brief.sh | 2 +- bin/fm-config-inherit-lib.sh | 4 +- bin/fm-procevent-lavish.sh | 95 ++++++++++++--- bin/fm-spawn.sh | 27 ++++- docs/configuration.md | 16 ++- docs/verification/process-event-sources.md | 13 ++- tests/fm-brief.test.sh | 2 +- tests/fm-procevent.test.sh | 108 +++++++++++++++++- tests/fm-spawn-dispatch-profile.test.sh | 46 ++++++++ 12 files changed, 291 insertions(+), 38 deletions(-) diff --git a/.agents/skills/process-event-sources/SKILL.md b/.agents/skills/process-event-sources/SKILL.md index a18e7b0eb3f..9219ddca1b7 100644 --- a/.agents/skills/process-event-sources/SKILL.md +++ b/.agents/skills/process-event-sources/SKILL.md @@ -27,12 +27,14 @@ Firstmate registers a source, keeps working, and is woken when that process comp ## Arming a source Use the adapter, not the generic runner, for a real source. -For a Lavish review artifact firstmate owns (a live investigating scout should host its own loop): +For a Lavish review artifact firstmate owns: ```sh bin/fm-procevent-lavish.sh arm ``` +Never arm a board that a live task hosts; follow the crew-hosted Lavish board contract in [`docs/configuration.md`](../../../docs/configuration.md#crew-hosted-lavish-review-boards). + Registering a source is not the same fact as listening to it: arming records the source, and a separate runner still has to pick it up. After arming by hand, confirm `bin/fm-procevent.sh list` reports that source as `live`, and run `bin/fm-procevent.sh reconcile` when it does not. Reconcile reports every launch that did not prove it took its claim within the confirm window as `failed=` and exits non-zero, so a source that cannot be started says so instead of looking armed, and it wakes you once per failure episode about it because the watcher discards that count; `start` does not fix that - if the source stays unowned, run `start` attached to read the runner's refusal, then check the source command and adapter binary the registration names, and if a later reconcile finds the source owned the episode closes on its own. @@ -105,11 +107,12 @@ Two rules the commands cannot enforce for you: ``` This call is atomically deduplicated by the exact source and sequence: it prints `handled: ` only the first time and `already-handled: ` on every repeat, so a paired effect gated on that distinction is never authorized twice. Reading the event line or the result file is not handling - only this call durably retires the wake, so call it every time, including on a repeat wake for a sequence you already acted on. : Ask the adapter what the result means rather than parsing it yourself. - `bin/fm-procevent.sh classify ` routes through the immutable built-in or extension identity captured with that result; for Lavish, its existing direct command returns `feedback`, `ended`, `waiting`, `missing`, or `unknown`. - Consume a Lavish capture with `bin/fm-procevent-lavish.sh read ` rather than grepping the raw file: that command reports declared and presented item counts plus a completeness verdict, enumerates every captured queued item while retaining supplied element identity, and surfaces a `tag=message` session-ending message as its own field. + `bin/fm-procevent.sh classify ` routes through the immutable built-in or extension identity captured with that result; for Lavish, its existing direct command returns `feedback`, `ended`, `waiting`, `disconnected`, `missing`, or `unknown`. + Consume a Lavish capture with `bin/fm-procevent-lavish.sh read ` rather than grepping the raw file: that command reports declared and presented item counts plus a completeness verdict, enumerates every captured queued item while retaining supplied element identity, and surfaces a `tag=message` freeform message as its own field, labeling it as session-ending only when the session ended. `answers` remains the keyed-choice extractor and never treats freeform prose as a decision key. A `feedback` result can still be the last one a review ever produces, so never assume another wake is coming just because the state is not `ended`. -: A routine no-op an adapter positively identifies never becomes a wake at all - it is recorded as handled and stays silent, so you never see it. For Lavish that is exactly an ended session carrying nothing: a board the captain closed without saying anything. A board close carrying a real answer, and every other result, still wakes you unchanged. Never read the absence of a wake as proof a review is still open; ask the source, not the queue. +The crew-hosted recovery ordering and interim polling rule are owned by the [crew-hosted Lavish board contract](../../../docs/configuration.md#crew-hosted-lavish-review-boards); `bin/fm-brief.sh` emits its interim instruction at the point of use. +: A routine no-op an adapter positively identifies never becomes a wake at all - it is recorded as handled and stays silent, so you never see it. For Lavish that is an ended session carrying nothing, or `browser_disconnected` (classified `disconnected`): a closed review window that still has an open session. A board close carrying a real answer, and every other result, still wakes you unchanged. Never read the absence of a wake as proof a review is still open; ask the source, not the queue. : A Lavish wake whose source id matches `bin/fm-procevent-lavish.sh source-id "$(bin/fm-bearings-board.sh path)"` is a bearings board result; load the `bearings` skill's board-wake handling regardless of which answer kinds the result contains. : A `when` wake carries the watch's one terminal captured outcome and may be re-announced until handled: `bin/fm-procevent-when.sh classify ` returns `fired` (relay the success and its output); `action-failed` (relay the captured error and decide recovery); `condition-error`, `never-true`, or `rejected` (the watch stopped safely without acting - report why and decide whether to re-arm); or `ambiguous` (the action was claimed but its outcome was never captured - verify its effect manually before anything else). Every `when` outcome is terminal and the action is never retried automatically, so after handling and the generic acknowledgement above, run `bin/fm-procevent-when.sh retire ` to clean the watch's private records before any re-arm. : A `quota` wake carries one terminal quota-check outcome: `bin/fm-procevent-quota.sh classify ` returns `low`, `exhausted`, `error`, or `unknown`. Report the provider and captured quota state, decide whether the active work should continue or move, then use the generic acknowledgement above. Re-arm explicitly if continued monitoring is needed. diff --git a/.agents/skills/stuck-crewmate-recovery/SKILL.md b/.agents/skills/stuck-crewmate-recovery/SKILL.md index 9004c3872bd..3d7ac5e1d66 100644 --- a/.agents/skills/stuck-crewmate-recovery/SKILL.md +++ b/.agents/skills/stuck-crewmate-recovery/SKILL.md @@ -14,6 +14,8 @@ metadata: Use this playbook when the session-start digest reports an ordinary direct report's endpoint dead or its metadata has no window, or when a direct report is stale, looping, repeatedly confused, asking a question its brief already answers, unresponsive, or when a steer failed to land. +Follow the crew-hosted Lavish board contract in [`docs/configuration.md`](../../../docs/configuration.md#crew-hosted-lavish-review-boards) when recovering a worker that hosts a board. + Interrupt, stop, and relaunch a worker through `bin/fm-control.sh interrupt|exit|relaunch`, which resolves the recorded runtime itself, verifies each action, and never tears down or discards anything ([`docs/agent-control.md`](../../../docs/agent-control.md)). That plane covers workers running in this home; a remotely placed secondmate is refused by name and reconciled through `secondmate-provisioning` instead. Load `harness-adapters` before a resume command or a harness-specific skill invocation, and whenever the adapter's own quirks matter. diff --git a/AGENTS.md b/AGENTS.md index 29ced552794..4c64dfe9a5d 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -81,6 +81,7 @@ config/startup-memory-budget primary-authoritative per-home startup-memory b config/stow-pass-horizon optional presence flag opting this home in to /stow's default-off pass-count decay horizon; LOCAL, gitignored, and not inherited; see docs/configuration.md "Stow pass horizon" config/herdr-presentation-spaces optional "off" opt-out from, or "on" opt-in to, Herdr's default-on disposable single-task visual projection, which is unconfigured-default-on only at or above a Herdr version floor; LOCAL, gitignored; inherited by secondmate homes; see docs/herdr-backend.md "Presentation spaces" config/trace-context optional presence flag enabling default-off native W3C trace-context propagation to spawned agents; LOCAL, gitignored; inherited by secondmate homes; see docs/configuration.md "Trace context propagation" and docs/trace-context.md +config/lavish-axi-host optional one-line per-machine Lavish server address; LOCAL, gitignored, inherited by secondmate homes, and exported into every worker launch; the adapter reads it before each board call; see docs/configuration.md "Lavish server address" config/turnend-churn-absorb optional presence flag opting this home into the default-off absorb of bare turn-end wakes on pane churn; LOCAL, gitignored, and not inherited; see docs/configuration.md "Turn-end pane-churn absorb" config/wedge-defer-parked-gate optional presence flag opting this home into the default-off deferral of a wedge escalation for a lane parked at a validation gate awaiting the supervisor's own still-open decision; LOCAL, gitignored, and not inherited; see docs/configuration.md "Parked-gate wait deferral" config/cmux-socket-password optional cmux control-socket password; LOCAL, gitignored; read fresh on every cmux CLI call and passed through without ever overriding an operator's own ambient CMUX_SOCKET_PASSWORD when absent (docs/cmux-backend.md "Setup") @@ -409,7 +410,7 @@ Retire one only on an explicit captain or main-firstmate decision, after loading A completed scout must leave a self-contained report before its scratch worktree can be discarded; read and relay its findings, record the report as the Done artifact, and re-evaluate the queue. A report may recommend implementation but does not authorize it. Before treating the investigation or any visual review as complete, load `captain-hold-lifecycle`; teardown enforces that shared completion gate. -When a scout's deliverable is a visual artifact the captain will iterate on, prefer keeping that scout alive to host its own Lavish loop rather than tearing it down and mediating from firstmate, so the scout keeps its investigation context and the captain iterates in one continuous session. +When a scout's deliverable is a visual artifact the captain will iterate on, keep it alive and follow the crew-hosted Lavish board contract in `docs/configuration.md` rather than arming or polling the board from firstmate. When implementation is separately authorized, promote the existing scout through `bin/fm-promote.sh` rather than creating a duplicate task. The promoted worker must inventory scratch state, return to a clean default-branch base, carry over only intended fix changes, create the ship branch, and follow the project's selected delivery path while leaving scratch commits and debug edits behind and turning a reproduced bug into the regression test. diff --git a/bin/fm-brief.sh b/bin/fm-brief.sh index 75d6717442c..891f4961628 100755 --- a/bin/fm-brief.sh +++ b/bin/fm-brief.sh @@ -365,7 +365,7 @@ TASK_SECTION=${TASK_SECTION%$'\n'} if [ "$KIND" = scout ]; then if "$SCRIPT_DIR/fm-bootstrap.sh" lavish-compatible >/dev/null 2>&1; then - LAVISH_LINE='If your deliverable is a visual artifact the captain will review and iterate on, you may host the Lavish review loop yourself (poll, revise, re-serve, staying alive) instead of handing it back to firstmate.' + LAVISH_LINE='If your deliverable is a visual artifact the captain will review and iterate on, use the lavish-axi rule: keep the poll in the foreground, or use your harness-native tracked background job; never use a bare &, nohup, disown, or redirected fire-and-forget polling; post needs-decision [key=board-review] with the live board URL, and stop at session_ended.' else LAVISH_LINE='Lavish is unavailable (lavish-axi is missing or below its supported version floor), so deliver your findings as a text report without Lavish, even for a visual deliverable.' fi diff --git a/bin/fm-config-inherit-lib.sh b/bin/fm-config-inherit-lib.sh index 79ff10605c2..f95d3647143 100644 --- a/bin/fm-config-inherit-lib.sh +++ b/bin/fm-config-inherit-lib.sh @@ -15,6 +15,8 @@ # "off" preferences propagate as files. Primary # config/trace-context is copied at the launch convergence point as part of the # default-off W3C trace-context setup, while live convergence leaves it unchanged. +# Primary config/lavish-axi-host carries the one per-machine Lavish server address +# to every worker so a worker never starts a second server on another interface. # The primary passes its frozen home-session decision into a newly launched # Secondmate; see docs/trace-context.md. # Primary config/claude-permission-mode is a captain-wide safety preference @@ -66,7 +68,7 @@ FM_SHARED_CAPTAIN_MODE="444" # The declared inheritable set (space-separated, config-dir-relative item paths). # Extend here to inherit more of the primary's local config; override via the # environment only in tests. Items must not contain whitespace. -FM_INHERITABLE_CONFIG="${FM_INHERITABLE_CONFIG:-crew-dispatch.json crew-harness backlog-backend backend herdr-presentation-spaces startup-memory-budget trace-context launch-env-allowlist claude-permission-mode}" +FM_INHERITABLE_CONFIG="${FM_INHERITABLE_CONFIG:-crew-dispatch.json crew-harness backlog-backend backend herdr-presentation-spaces startup-memory-budget trace-context launch-env-allowlist claude-permission-mode lavish-axi-host}" # Items whose value is a home-SESSION enablement decision rather than durable # local configuration. They are inherited at the launch convergence point, where diff --git a/bin/fm-procevent-lavish.sh b/bin/fm-procevent-lavish.sh index ece166e7d37..7d24b6d2534 100755 --- a/bin/fm-procevent-lavish.sh +++ b/bin/fm-procevent-lavish.sh @@ -14,13 +14,14 @@ # fm-procevent-lavish.sh poll # # classify Print the lifecycle state a handler should act on: feedback, ended, -# waiting, missing, or unknown. +# waiting, disconnected, missing, or unknown. # read Print a structured presentation of one already-captured result so a # handler consumes every queued item without grepping the raw file. # It is read-only over the capture: it does not arm, poll, or change -# what Lavish delivered. The session-ending freeform message -# (tag=message) is its own labeled field, printed first and distinct -# from per-element annotations. Declared and presented item counts, +# what Lavish delivered. The freeform message (tag=message) is its +# own labeled field, printed first and distinct from per-element +# annotations; it is labeled SESSION-ENDING MESSAGE only when the +# session ended. Declared and presented item counts, # plus a completeness verdict, follow before all annotations so a # partial read is obvious. Each annotation retains its element uid, # selector, tag, and text. A non-choice freeform comment (`prompt`) @@ -47,9 +48,10 @@ # Closing a review surface that carried nothing is the single most common Lavish # result: the captain reads a board, says nothing, and closes it. Announcing that # put a wake in front of the handler whose entire content was that nothing -# happened. `silent` therefore holds one narrow, positively-determined shape - +# happened. `silent` therefore holds two narrow, positively-determined shapes - # a session this adapter classifies `ended` that carries no queued content block -# at all - and every other result stays announced. +# at all, or `browser_disconnected`, which carries no answer while the session +# remains open - and every other result stays announced. # # Deliberately narrow, in both directions. A `Send & End` close carrying the # captain's actual answer arrives as `status: feedback` with `session_ended`, so @@ -66,6 +68,13 @@ # and how to read a completed result. Ownership, durable capture, publication, # and restart recovery all belong to bin/fm-procevent.sh. # +# The published poll vocabulary includes feedback, ended, waiting, and +# browser_disconnected. A waiting result from this no-timeout poll means a +# second poller was present; it is not a normal idle round. browser_disconnected +# means the session remains open and is handled as a silent reconnect wait. +# The poll reads config/lavish-axi-host from FM_HOME before every lavish-axi +# invocation so firstmate and workers reach the same server. +# # `answers` is this adapter's half of the generic keyed-answer contract in # bin/fm-procevent.sh. It reports what the captain actually chose, as # `\t\t