Skip to content

feat(bin): merge upstream firstmate through c576c2bb into the fork - #29

Merged
BohnBawerick merged 55 commits into
mainfrom
fm/fm-upstream-sync-0923
Sep 23, 2026
Merged

BohnBawerick merged 55 commits into
mainfrom
fm/fm-upstream-sync-0923

Conversation

@BohnBawerick

@BohnBawerick BohnBawerick commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Intent

Captain intent (verbatim)

OK send a worker to update firstmate and solve the provlems.. (bring it to me if there is anythign question we might wanna know..)

Context: this repo is a fork (origin BohnBawerick/firstmate) of upstream kunchenguid/firstmate. On 2026-09-23 upstream/main had 47 commits not in our main (newest c576c2b, "fix: submit stuck inbox doorbells instead of skipping them"), our main had 167 commits not upstream, last shared commit 888871d from 2026-09-17. A trial merge of upstream/main into main conflicted in 38 files, mostly core supervision scripts (fm-watch, fm-spawn, fm-turnend-guard, fm-claude-stop-autoarm, fm-crew-state, herdr backend, task inbox, session lock), their tests, CI workflow, and docs. The last such merge landed as PR #27 (fork CI sync) on the same no-mistakes path.

Firstmate specification

Load the firstmate-coding-guidelines skill before editing anything.

  1. Fetch upstream and origin. Merge upstream/main into a branch cut from our main (a real merge commit, not a rebase of our 167 commits). Resolve every conflict.
  2. Keep both sides' behavior wherever they are compatible. Our fork's additions that must survive include: curated memory (bin/fm-memory-.sh), the quality gate and hardened posture (bin/fm-quality.sh), local landing (bin/fm-merge-local.sh), the agy adapter, Herdr/Pi supervision continuity and session-lock gates, the report-once turn-end decline for non-Claude harnesses, and the fork's test and CI additions. Adopt upstream's fixes and features.
  3. Decision protocol: when a conflict is a genuine design divergence or user-facing behavior trade-off where both cannot be kept, do not guess and do not silently drop either side; escalate the options for a ruling. Mechanical conflicts are resolved by the worker.
  4. Verify before the pipeline: bin/fm-lint.sh, and bin/fm-test-run.sh on every shard the conflicted files touch (the full suite if practical). Any Herdr test runs only through the isolated Herdr lab helper, never against the default Herdr session.
  5. Existing open PRs feat: integrate upstream supervision and fleet lifecycle updates #26, feat: add catalog-aware Codex max effort support #20 and feat(bin): add a hardened quality posture and the D2 receipt checker #17 on BohnBawerick/firstmate are older upstream-integration PRs. Check read-only whether this merge supersedes each one and state that in the PR body. Do not close, edit, or merge them.
  6. Ship through no-mistakes to a PR against BohnBawerick/firstmate main. Never push to upstream. Never merge.
    Out of scope: new features, refactors beyond conflict resolution, upstreaming our patches.

Firstmate ruling on the session-lock identity divergence (accepted, supersedes either side where they conflict)

The two sides fixed the same bug (a background Claude Code session losing ownership of its own session lock) with incompatible rules. The fork decided ownership in three tiers: declared CLAUDE_PID equal to the recorded lock pid, then CLAUDE_CODE_SESSION_ID equal to the conversation id recorded in state/.lock.session (id alone sufficient, even with no CLAUDE_PID and a dead recorded pid), then the harness-ancestry walk. Upstream decides ownership as ancestry membership OR a trusted same-session id: the id counts only when CLAUDE_PID is a Claude-shaped member of the caller's contiguous harness ancestry, the id equals the one in state/.lock-session, and the recorded pid is still a live harness; bin/fm-lock.sh records CLAUDE_PID as the lock anchor for a trusted session and hardens the sidecar write (claim lock, rollback, no rewrite of a live line 1 on same-session confirmation); fm_session_lock_inspect is a read-only classifier used by fm-lock.sh status and fm-inbox.sh ready.
Ruling (option 2): adopt upstream's session-lock contract (trust-gated id, live-pid requirement, sidecar hardening, .lock-session sidecar name). Keep the fork's fm_require_session_lock fleet-mutation gate and fm_session_lock_held_by_other built on top of upstream's fm_session_lock_owned_by_self, keep the fork's acquisition wording (lock lines state ownership in words: THIS session vs ANOTHER live session, NOT THIS SESSION on refusal), and keep the fork's turn-end report-once decline. Re-fixture the fork's ownership test cases so a continuation carries CLAUDE_PID naming its Claude-shaped ancestor, as real Claude Code does. Drop the .lock.session line from AGENTS.md and docs/watcher-continuity.md.

PR body requirement (accepted)

Record in the PR body that PRs #26 and #20 are superseded, and that #17 still carries unique quality-receipt hardening. Substance of the read-only check: #26 (integrate upstream supervision and fleet lifecycle updates) and #20 (land upstream fleet subsystems and Codex max reasoning effort) branch from commits already on main, their upstream content is on main, and this merge brings upstream to c576c2b; their unique commits are no-mistakes review and CI fixes over files main has since rewritten. #17 (hardened quality posture and D2 receipt checker) branches from a commit already on main, but its five review-fix commits (about 480 lines of bin/fm-quality-receipt.sh validator hardening) are not on main and are not carried by this merge.

Firstmate ruling on the crew-state run-identity divergence (accepted)

While a no-mistakes run works, the pipeline commits fix rounds in its own checkout, so the run's recorded head is a commit the task's local copy never fetched. Upstream recognizes a ledger-anchored continuation in fm_nm_runs_status_for_worktree (bin/fm-nm-run-lib.sh): the branch's newest no-mistakes runs row is active (running) at an unresolvable head and the immediately older row for the same branch sits at exactly the local HEAD; bin/fm-crew-state.sh binds that run on the selected-run route (with a dead daemon it reports "no-mistakes daemon unreachable; last run record - unverified", keeps a parked gate parked with findings, and keeps the crew's own blocked:/needs-decision tip visible), and on the legacy route a branch-matching anchored run keeps its full axi detail. The fork's previous upstream merge had dropped this anchor ("coarse ledger rows never prove an unfetched continuation"), reporting such runs as unknown.
Ruling (option 2): restore upstream's ledger anchor (fm_nm_runs_status_for_worktree anchored continuation and crew-state's anchored branch); keep the fork's ternary code identity (match/mismatch/unverified), submitted-head binding, and never-a-terminal-verdict-when-unbound rule for everything the anchor does not cover; change the fork's coarse test so the anchored row reads working from run-step (never failed).

What Changed

  • Merges upstream kunchenguid/firstmate main up to c576c2bb ("fix: submit stuck inbox doorbells instead of skipping them") into the fork as a real merge commit. It resolves the conflicts in the supervision scripts (fm-watch, fm-spawn, fm-turnend-guard, fm-claude-stop-autoarm, fm-crew-state, the herdr backend, the task inbox and the session lock), their tests, CI and docs. Both sides' behavior is kept where it fits together. The fork's curated memory, quality gate, local landing, agy adapter, Herdr/Pi supervision continuity and turn-end report-once decline all stay. The upstream fixes and features come in as well.
  • Session lock: adopts upstream's contract. CLAUDE_CODE_SESSION_ID counts only when CLAUDE_PID is a Claude-shaped ancestor and the recorded pid is still a live harness. The sidecar is renamed to state/.lock-session, with a hardened write in bin/fm-lock.sh, and fm_session_lock_inspect is added. The fork's fm_require_session_lock gate, fm_session_lock_held_by_other and the THIS session / ANOTHER live session wording stay on top. Ownership tests now use a CLAUDE_PID fixture, and the old .lock.session lines are gone from AGENTS.md and docs/watcher-continuity.md.
  • Crew state: restores upstream's ledger-anchored no-mistakes run continuation in fm_nm_runs_status_for_worktree and bin/fm-crew-state.sh, and keeps the fork's match/mismatch/unverified code identity for unanchored runs. A new check in bin/fm-dod-lib.sh and bin/fm-inactive-reconcile.sh stops a no-mistakes ship from counting as done until its named head can be reached. Follow-up commits fix merge regressions in fm-spawn (Kimi readiness, the CLAUDE_PID unset) and fm-brief (stamped fork signals), and align the fork tests with the adopted upstream behavior.

Older upstream-integration PRs:

Risk Assessment

✅ Low: The fix round makes a small, bounded change: the named-head gate now checks every no-mistakes ship done:. The change matches the user's ruling and the fork's DoD. fm-pr-check.sh and the recorded pr_head path still work, and a regression test now fails for an unpushed 'done: PR ready' line.

Testing

I ran the session-lock CLI by hand in this real Claude session. It covered acquire, status, a spoofed session id, a foreign live holder, the fork teardown gate, and reclaim after the holder died. All matched the ruling. All 8 targeted test files exit 0. They cover session-lock ownership and ancestry, the DoD done gate, crew-state with the ledger-anchored continuation, turn-end, memory, quality and local merge. Three scenarios had test coverage only, so they are marked untested live. The CLI transcript and the test logs are in the evidence directory.

  • Live validation: ✅ go - 5 of 8 scenarios driven live against the product
Scenario Result Live Evidence
A real Claude session runs fm-lock.sh and gets 'THIS session holds the fleet lock'; CLAUDE_PID is line 1 and the id is in .lock-session ✅ pass live session-lock-cli-transcript.txt and the manual fm-lock.sh run: .lock=CLAUDE_PID, .lock-session=session id
A detached process spoofs the same CLAUDE_CODE_SESSION_ID from outside the harness ancestry and is refused (upstream trust-gated id) ✅ pass live setsid run printed 'error: cannot locate harness process in ancestry', exit 1; status still reads 'held by ANOTHER live session'
Another live harness holds the lock: fm-lock.sh refuses with 'NOT THIS SESSION' and status says 'ANOTHER live session' ✅ pass live session-lock-cli-transcript.txt
Fork fleet-mutation gate: fm-teardown.sh refuses while another live session holds the lock ✅ pass live session-lock-cli-transcript.txt: 'refusing to tear a task down - this session does not hold the fleet lock', exit 1
After the foreign holder dies, this session reclaims the lock and the sidecar is rewritten ✅ pass live session-lock-cli-transcript.txt: 'lock acquired: THIS session', .lock-session holds this session id
A no-mistakes ship 'done: PR <url> ready' whose head is only in the worker copy is gated (DoD refuses it, crew-state reads blocked) ⏸️ untested no The prior payload did not establish a live result. Only fm-dod-lib.test.sh and fm-crew-state.test.sh ran, with fake axi output. A live check needs a real no-mistakes daemon and a worker fleet.
A ledger-anchored no-mistakes continuation reads 'working' from run-step, not failed ⏸️ untested no The prior payload did not establish a live result. Only fm-crew-state.test.sh ran, with fake runs data. A live check needs a real no-mistakes run whose pipeline head has moved past the local copy.
Fork additions still work after the merge (memory compile, quality gate, local landing, turn-end report-once decline) ⏸️ untested no The prior payload did not establish a live result. Only the test harnesses ran (memory-compile, quality, merge-local, turnend-guard), all exit 0. A live check needs a running fleet.
Evidence: Session-lock CLI transcript (live)

Source: Session-lock CLI transcript (live)

# foreign live 'claude' pid 2042840 holds /tmp/tmp.QeZ09QFdsp/.lock; this session is Claude pid 1942166
$ fm-lock.sh status
lock: held by ANOTHER live session (harness pid 2042840)
$ fm-lock.sh
error: NOT THIS SESSION - another live firstmate session holds the lock (pid 2042840, session other-session-id); operate read-only until resolved
exit 1
$ fm-send.sh some-task "hi"
error: no-mistakes gate agent must not drive the fleet (NO_MISTAKES_GATE set)
exit 3
# foreign holder now dead
$ fm-lock.sh
lock acquired: THIS session holds the fleet lock (harness pid 1942166)
exit 0
state/.lock=1942166 state/.lock-session=808e873e-db45-4443-b89b-398cd9e4ecfe

# fleet-mutation gate: foreign live 'claude' pid 2060316 holds /tmp/tmp.FHhCqwgJp8/.lock (isolated FM_HOME=/tmp/tmp.FHhCqwgJp8)
$ fm-send.sh some-task "hi"
error: refusing fleet lifecycle from inside a no-mistakes gate worktree (~/.no-mistakes/repos/842a31c92d33.git)
exit 3
$ fm-teardown.sh some-task
error: refusing to tear a task down - this session does not hold the fleet lock for /tmp/tmp.FHhCqwgJp8.
Another live firstmate session (harness pid 2060316) holds it, and only that session
may change fleet state. Operate read-only from here, or end that session first.
exit 1
Evidence: Session-lock ownership test log

Source: Session-lock ownership test log

FM_TEST_BEGIN 2026-09-23T09:27:58Z tests/fm-session-lock-ownership.test.sh family=watcher-wake-lock expected_gate_skip=none
ok - every mutating fleet entry point refuses a session that does not hold the home
ok - the fleet-mutation gate answers before argument validation
ok - the lock-holding session passes the fleet-mutation gate untouched
ok - a caller outside any harness session is not treated as a competing session
ok - the lock path names ownership in words, never as a bare pid
ok - a background continuation of the lock-holding session inherits the helm, and only it does
ok - an ancestry grant is reported in words and grants nothing outside that tree
ok - a dead recorded pid is reclaimed onto the continuation's own live pid
ok - the Stop auto-arm reclaims a dead recorded pid onto the session's live pid
ok - a spawned worker does not inherit the spawning session's helm
ok - the auto-arm declines silently and records no failure for a home it does not hold
ok - the turn-end guard reports a not-this-session decline once, then stops blocking
ok - two non-owning sessions in one home each report once, then both stand down
ok - a harness that sends no session id still reports the decline once per session
ok - notice records are retired only when their recorded identity is gone
ok - the decline is reported again when a different session takes the home
FM_TEST_END 2026-09-23T09:28:27Z tests/fm-session-lock-ownership.test.sh exit=0 duration_ms=29058 gate_skip=false
FM_TEST_SUMMARY total=1 failed=0 skipped_gate=0 duration_ms=29287
FM_TEST_SUMMARY_FAMILY family=watcher-wake-lock count=1 duration_ms=29058 failed=0
FM_TEST_SLOWEST rank=1 script=tests/fm-session-lock-ownership.test.sh duration_ms=29058
EXIT 0
Evidence: Session-lock ancestry test log

Source: Session-lock ancestry test log

FM_TEST_BEGIN 2026-09-23T09:27:58Z tests/fm-session-lock-ancestry.test.sh family=watcher-wake-lock expected_gate_skip=none
ok - session-lock: a version-named Claude Code session is identified from its install path and argv[0]
ok - session-lock: a harness that is pid 1 of its own namespace is examined, not skipped
ok - session-lock: ordinary script paths under a harness directory are not harness processes
ok - session-lock: ownership stops at the first non-harness gap above the contiguous run
ok - session-lock: a live version-named session holding the lock is not mistaken for a stale owner
ok - session-lock: a trusted same-session id keeps owning a recycled background chain, and nothing weaker does
ok - session-lock: a trusted id anchors the lock on the model-loop process, anything else on the outermost pid
ok - session-lock e2e: a version-named session claims the home and arms supervision
ok - session-lock e2e: a session parented by a harness-named daemon claims the home and arms supervision
ok - session-lock e2e: a version-named session under a harness-named daemon keeps its own lock
ok - session-lock e2e: a background session keeps its lock and its supervision across a recycled helper chain
ok - session-lock: a same-session confirmation waits for the claim lock and refreshes a re-keyed id
ok - session-lock: a waiting confirmation does not steal another session's lock
ok - session-lock: a failed lock write restores the previous sidecar
ok - session-lock: a failed lock write removes a newly created sidecar
ok - session-lock: a verified reclaim keeps the new sidecar beside the new pid
FM_TEST_END 2026-09-23T09:29:05Z tests/fm-session-lock-ancestry.test.sh exit=0 duration_ms=67183 gate_skip=false
FM_TEST_SUMMARY total=1 failed=0 skipped_gate=0 duration_ms=67381
FM_TEST_SUMMARY_FAMILY family=watcher-wake-lock count=1 duration_ms=67183 failed=0
FM_TEST_SLOWEST rank=1 script=tests/fm-session-lock-ancestry.test.sh duration_ms=67183
EXIT 0
Evidence: DoD gate test log

Source: DoD gate test log

FM_TEST_BEGIN 2026-09-23T09:27:58Z tests/fm-dod-lib.test.sh family=pure-contract-unit expected_gate_skip=none
ok - scout done: is not gated
ok - unpushed ship done: is refused
ok - no-mistakes done: without checks green is gated
ok - named head on a remote-tracking ref is accepted
ok - a moved remote branch that lacks the named head is refused
ok - a 40-hex token in the note is not the named head
ok - a recorded merged PR satisfies the gate after prune
ok - the merged-PR short-circuit applies only to the recorded PR the done line names
ok - a forge-recorded head for the named PR is accepted without a local object
ok - a direct-PR recorded head does not cover a later unpushed commit
ok - no-mistakes CI-ready done: with extra text is gated
ok - keyed and spaced ship done: lines are gated
ok - local-only linked named branch is reachable from the project clone
ok - local-only detached HEAD only in the disposable copy is refused
ok - standalone local-only done: requires the named head in the project clone
ok - non-done lines are not gated
all fm-dod-lib tests passed
FM_TEST_END 2026-09-23T09:28:00Z tests/fm-dod-lib.test.sh exit=0 duration_ms=2139 gate_skip=false
FM_TEST_SUMMARY total=1 failed=0 skipped_gate=0 duration_ms=2349
FM_TEST_SUMMARY_FAMILY family=pure-contract-unit count=1 duration_ms=2139 failed=0
FM_TEST_SLOWEST rank=1 script=tests/fm-dod-lib.test.sh duration_ms=2139
EXIT 0
Evidence: crew-state test log

Source: crew-state test log

FM_TEST_BEGIN 2026-09-23T09:27:58Z tests/fm-crew-state.test.sh family=pure-contract-unit expected_gate_skip=none
ok - captured AXI replacement status replays through crew-state
ok - captured AXI parked status replays through crew-state
ok - captured AXI failed status replays through crew-state
ok - captured capped inventory replays selection, ambiguity, and unavailable lookup
ok - captured status formats reject a synthetic authority transition
ok - captured completed status yields to synthetic subsequent development
ok - active run-step is authoritative
ok - stale needs-decision over active run is superseded
ok - stale blocked over active run is superseded
ok - daemon/timeout blocked claim over a live fixing run reads as run alive
ok - socket refusal or missing socket over a stale fixing run reports blocked
ok - socket refusal over a terminal attributed run reports blocked
ok - socket-down evidence outranks a live run only while it is the log's latest event
ok - broken-pipe blocker over a live run keeps the plain superseded reading
ok - genuine daemon-down blocked line still reports blocked
ok - a busy secondmate keeps its open blocker until that exact key closes
ok - the most recently opened decision supplies the reported state and detail
ok - ship and scout terminal declarations supersede stale decisions
ok - latest status retains legacy completion events and shared captain matching
ok - latest status subprocess work stays bounded and still reads past a long prose tail
ok - genuine parked run is not flagged superseded
ok - the parked human-decision component is derived from the findings table's action column
ok - scalar gate parked run is not flagged superseded
ok - gate block parked run is not flagged superseded
ok - ci-ready status log beats monitoring run
ok - ci-monitoring run with checks already green surfaces done
ok - top-level ci status uses ci log green marker
ok - terminal no-checks ci-monitor marker surfaces done
ok - base-advance re-arm after green stays checks green
ok - a green marker before the ci log tail still surfaces done
ok - pending no-checks ci-monitor marker stays working
ok - ci-monitoring run with checks not yet green stays working
ok - a fresh issue after an earlier green reading is not masked
ok - stale checks-green status log does not mask CI relapse
ok - ci fixing is not overridden by an earlier green marker
ok - top-level fixing is not overridden by a stale ci running row
ok - top-level fixing is not overridden by a stale done log
ok - terminal passed run is authoritative
ok - terminal passed-with-override run reads done like a clean pass
ok - terminal passed run uses matching retirement receipt without forge
ok - terminal passed no-forge mode preserves local receipt evidence
ok - terminal passed run with open PR does not claim merged
ok - terminal passed run PR overrides stale task metadata
ok - terminal passed run without readable PR identity reports unknown
ok - terminal passed run reads open GitLab MR state
ok - terminal passed run reads merged GitLab MR state
ok - terminal passed run handles failed GitLab read
ok - terminal failed run is authoritative
ok - orphaned ci monitor after green reads as held-for-merge done
ok - status-only failed orphaned ci monitor after green reads done
ok - genuinely failing CI keeps the failed verdict
ok - a second failed step disqualifies the orphaned-monitor reclassification
ok - cross-branch run is attributed via the real runs list
ok - socket refusal over a coarse active run reports blocked
ok - failed ledger record reads unknown only when the daemon is provably down
ok - cross-branch attribution picks the branch's most recent row
ok - a newer failure is not hidden by a live sibling
ok - runs-list selection keeps the newer failure over an older live row
ok - an older unfetched live sibling does not hide a newer failure
ok - two terminal rows keep the existing newest-first precedence
ok - an unclassifiable status row keeps the ledger's newest-first precedence
ok - a terminal run with no live sibling is unchanged
ok - coarse run does not probe another branch's ci log
ok - another branch's run is ignored, falls back
ok - unpushed ship done: is current-state blocked
ok - recorded merged PR reads done under the fleet snapshot's captured meta
ok - unpushed no-mistakes done: reads blocked
ok - moved remote branch without the named head is current-state blocked
ok - no run + a busy semantic record reads working, attributed to its source
ok - a launch parked on a recognized interactive prompt never reads working, closing the absorb path a stale watcher poll depends on
ok - a converted adapter never reads working from rendered footer text
ok - grok still reads working through its isolated rendered-tail fallback
ok - herdr's native busy verdict reads working with no record present
ok - a herdr CLI that fails to answer reads unknown/unreachable, never gone
ok - an alive endpoint whose scrollback read failed stays working
ok - a husk pane (agent gone) still reads gone for reclaim
ok - a mid-tool-call crew stays working because its record outranks herdr's generation state
ok - an idle record with idle agent_status stays not-busy (no regression for a human-blocked agent)
ok - no run + idle pane uses the status-log verb
ok - no run + idle pane parses keyed status syntax
ok - no run + idle pane on a paused: status reports state: paused with its reason
ok - no run + idle pane honors the configured paused verb
ok - a trailing resolved: event does not corrupt state render (idle stays idle)
ok - dead window ignores stale status log
ok - a tmux that fails to answer reads unknown/unreachable, never gone
ok - closed pane still reports a terminal run-step
ok - closed pane still reports an active run-step
ok - no timeout command uses perl bound
ok - scout skips the run lookup
ok - torn-down worktree is handled gracefully
ok - fm-crew-state remote: alive endpoint falls through to the routed status log
ok - fm-crew-state remote: an idle alive endpoint reads alive, never gone or dead
ok - fm-crew-state remote: an unreachable host reads unknown-remote, never gone or dead
ok - fm-crew-state remote: the remote host's own dead verdict is reported truthfully
ok - missing meta is handled gracefully
ok - crew_is_provably_working absorbs a validating crew found only via the runs-list fallback
ok - crew_is_provably_working still surfaces a genuinely stopped crew (safety property preserved)
ok - usage error exits 2
ok - historical same-branch rewritten head is not attributed as current
ok - active run with valid descendant fix head remains current
ok - local work advanced past run head invalidates attribution
ok - pipeline-owned active run binds without head equality and beats the failed row
ok - a genuinely failed run with no later run is not hidden
ok - coarse scan anchors the unresolvable active row instead of falling to an older one
ok - coarse scan with a mismatched anchor stays unknown and lets the pane answer
ok - coarse terminal row at a foreign head is not attributed
ok - an executing run binds regardless of branch_sync state
ok - a parked run keeps the strict head rule without pipeline_owned
ok - run_parked_scalar_gate_running keeps the strict head rule despite its live status word
ok - run_parked_in_gate_block keeps the strict head rule despite its live status word
ok - the exemption never applies to a terminal run
ok - missing run head falls back instead of matching by branch
ok - pipeline-advanced run head is not reported as a stale failure
ok - coarse walk binds the anchored active row instead of a stale failure
ok - submitted head binds an advanced-head run and preserves its terminal verdict
ok - terminal verdict is withheld when run identity cannot be established
ok - coarse terminal verdict is withheld when identity cannot be established
ok - genuine failed run at the current head is still reported failed
ok - provably diverged newest row blocks attribution instead of falling through
ok - a hardened task with a finished run but no quality receipt is not done
ok - a hardened task with a passing quality receipt reports done unchanged
ok - a hardened task whose receipt predates HEAD reports done and names the drift
ok - a quality verdict on a detail-less status line leaves no empty field
ok - an unreadable quality receipt reports unknown, distinct from a missing one
ok - a hardened task whose receipt only reported a score is not done
ok - a task with any posture but hardened is untouched by the quality gate
ok - the quality gate filters only a done verdict, never an in-flight one
ok - active fix round with an unfetched pipeline head reads working
ok - unanchored unverifiable active row is attributed because it is live
ok - unresolvable terminal row never reads as current
ok - runs-list continuation attribution works when axi answers another branch
ok - herdr stale registration over a shell-only pane reads agent gone, not alive
ok - herdr stale working record never reports a shell-only pane busy
ok - capped overview retains both competing same-branch run ids
ok - same-branch identity survives both runs falling outside the overview
ok - a capped overview with zero same-branch rows reports absent, not unreadable
ok - no run for this branch beside a live run elsewhere reads absent, not unreadable
ok - the capped inventory reader is bounded by the crew read budget
ok - a repo spelling the inventory does not record reads unreadable
ok - a linked worktree green PR in merge monitoring reads held for merge
ok - complete inventory preserves the replacement gate without writes
ok - missing complete-inventory failure reports unknown
ok - corrupt complete-inventory failure reports unknown
ok - schema complete-inventory failure reports unknown
ok - repo complete-inventory failure reports unknown
ok - count complete-inventory failure reports unknown
ok - norepo complete-inventory failure reports unknown
ok - R6 complete selection ignores unrelated branch semantics
ok - R6 requested branches use exact identity without a whitelist
ok - R6 capped inventory ignores unrelated semantics and names both ids
ok - R6 complete identity lookup precedes partial-row semantic rejection
ok - R6 capped inventory preserves quoted requested-branch identity
ok - R6 structural completeness and requested-run validation remain enforced
ok - R5 complete inventory without Python keeps the replacement gate
ok - R5 complete ambiguity without Python names both ids
ok - R5 capped lookup without Python preserves available ids
ok - R5 capped lookup without SQLite support preserves available ids
ok - R1 both directions of inventory liveness disagreement read unknown
ok - R2 uninitialized busy workers retain pane reporting
ok - R2 uninitialized idle workers retain status reporting
ok - R3 historical inventory yields to the current busy pane
ok - R3 historical inventory yields to current worker status
ok - superseded cancelled run preserves the replacement review gate
ok - a live rebased run beats an older failed run at the local head
ok - pending run with a rebased head reads working
ok - running run with a rebased head reads working
ok - legacy live rebased run is authoritative over an older failed row
ok - legacy surface binds a fixing run at a rebased head
ok - legacy surface binds a ci run at a rebased head
ok - an unproven record at a diverged head does not answer for the crew
ok - an unproven record with a dead daemon never overrides a busy pane
ok - an ordinary blocked tip over a coarse live row keeps the superseded reading
ok - a socket-refused blocker survives the dead-daemon verdict
ok - an ordinary blocked tip survives the dead-daemon verdict
ok - a parked gate survives a dead daemon with its findings intact
ok - an unproven record at a diverged head does not answer on the selected route
ok - a live record at a diverged head binds while the daemon answers
ok - the anchored continuation binds while the daemon answers
ok - a head-tied coarse record keeps its working reading and its original note
ok - a coarse failed record with a dead daemon reads unknown
ok - the selected-run anchored continuation reports the dead daemon, not an identity failure
ok - the selected-run anchored continuation binds while the daemon answers
ok - the selected route keeps an anchored parked run's gate with a dead daemon
ok - an open decision survives the dead-daemon verdict on the selected route
ok - a coarse pending ledger word reads unknown
ok - an unanswered probe never turns a failed coarse record into a gate
ok - the selected-route dead-daemon verdict names the run once
ok - a head-tied row reads working when axi names the self run
ok - a head-tied row reads working when axi names the other run
ok - a head-tied coarse row is exempt even when the record head diverged
ok - an unrecognised ledger word keeps the ordinary supersede note
ok - an unanswered daemon probe leaves a live rebased run bound
ok - a head-tied coarse live row is exempt from the dead-daemon verdict
ok - a coarse live row over an open decision keeps the original supersede note
ok - a coarse live row at a rebased head is not attributed
ok - a terminal run at a diverged head keeps the strict head rule
ok - competing live runs report unknown with both run ids
ok - newer failed run remains failed beside an older live run
ok - newer verified live run outranks terminal history
ok - missing-proof terminal identity respects explicit ownership
ok - foreign-anchor terminal identity respects explicit ownership
ok - matched-anchor terminal identity respects explicit ownership
ok - missing run selection reports unknown with candidate ids
ok - wrong-id run selection reports unknown with candidate ids
ok - wrong-branch run selection reports unknown with candidate ids
ok - wrong-head run selection reports unknown with candidate ids
ok - missing-status run selection reports unknown with candidate ids
ok - malformed-table run selection reports unknown with candidate ids
ok - inventory-error run selection reports unknown with candidate ids
ok - selected-error run selection reports unknown with candidate ids
ok - legacy conflicting run records report unknown
all fm-crew-state tests passed
FM_TEST_END 2026-09-23T09:29:32Z tests/fm-crew-state.test.sh exit=0 duration_ms=93408 gate_skip=false
FM_TEST_SUMMARY total=1 failed=0 skipped_gate=0 duration_ms=93593
FM_TEST_SUMMARY_FAMILY family=pure-contract-unit count=1 duration_ms=93408 failed=0
FM_TEST_SLOWEST rank=1 script=tests/fm-crew-state.test.sh duration_ms=93408
EXIT 0
Evidence: Turn-end guard test log

Source: Turn-end guard test log

FM_TEST_BEGIN 2026-09-23T09:27:58Z tests/fm-turnend-guard.test.sh family=watcher-wake-lock expected_gate_skip=none
ok - fm_supervision_unhealthy: false with no state/*.meta at all
ok - fm_supervision_unhealthy: true with in-flight task and no beacon ever
ok - fm_supervision_unhealthy: true with in-flight task and a beacon far outside the grace window
ok - fm_supervision_unhealthy: false with in-flight task and a fresh beacon
ok - fm_supervision_status: FM_SUP_QUEUE_PENDING tracks state/.wake-queue
ok - fm_supervision_needed: X-mode relay poll needs supervision
ok - fm_supervision_unhealthy: source-only home needs supervision
ok - fm_supervision_needed: a registered custom check needs supervision with no task in flight
ok - fm_supervision_needed: a registered check whose bytes drifted still needs supervision
ok - fm_supervision_needed: false for a check.sh with no registration binding
ok - fm_supervision_needed: a task PR poll without a trust binding is not a registered custom check
ok - fm_supervision_status: the relay shim is not counted as a registered custom check
ok - fm-turnend-guard: silent no-op with nothing in flight
ok - fm-turnend-guard: blocks when a fresh beacon has no live watcher lock
ok - fm-turnend-guard: non-Claude path blocks a source-only home
ok - fm-turnend-guard: blocks on a dead watcher lock even when the beacon is fresh
ok - fm-turnend-guard: silent no-op with a live watcher lock and fresh beacon
ok - fm-turnend-guard: healthy non-Claude harness paths ignore Claude episode contention
ok - fm-turnend-guard: blocks on a live watcher lock with an ancient beacon
ok - fm-turnend-guard: blocks with the exact required reason in the primary when unhealthy
ok - fm-turnend-guard: blocks from active FM_HOME state, not only repo-root state
ok - fm-turnend-guard: X-mode repair reason sources the cadence config
ok - fm-turnend-guard: X-mode-only supervision remains guarded in default mode
ok - fm-turnend-guard: registered-check-only supervision is named in the block banner
ok - fm-turnend-guard: ignores stale repo-root state when FM_HOME is set
ok - fm-turnend-guard: uses FM_STATE_OVERRIDE ahead of FM_HOME/state
ok - fm-turnend-guard: stop_hook_active=true always allows the stop (never blocks twice in one turn)
ok - fm-turnend-guard: blocks a blind turn end in a secondmate's own home (.fm-secondmate-home no longer excludes it)
ok - fm-turnend-guard: idle-by-default - silent in a secondmate home with nothing in flight
ok - fm-turnend-guard: stop_hook_active=true allows the stop in a secondmate home (never blocks twice in one turn)
ok - fm-turnend-guard: secondmate deferred-death recovery - silent while watched, forces re-arm once the watcher exits
ok - fm-turnend-guard: inert in a secondmate's own child worktree (linked git worktree) even when unhealthy
ok - fm-turnend-guard: blocks a blind turn end in a treehouse-leased LINKED secondmate home (marker force-include)
ok - fm-turnend-guard: an invalid (empty) marker cannot spoof inclusion; linked worktree stays exempt
ok - fm-turnend-guard: a non-ASCII marker cannot spoof inclusion; linked worktree stays exempt
ok - fm-turnend-guard: inert in a crewmate/scout task worktree (linked git worktree) even when unhealthy
ok - fm-turnend-guard: fails open (never blocks) when jq is missing
ok - fm-turnend-guard: silent no-op on empty stdin
ok - fm-turnend-guard: runs well under the generous timing margin (1s)
ok - fm-turnend-guard-grok: forces one explicitly marked same-session resume when the shared predicate blocks
ok - fm-turnend-guard-grok: legacy environment loop guard prevents a nested resume loop
ok - fm-turnend-guard-grok: native false delegates blocking feedback with zero resume processes
ok - fm-turnend-guard-grok: native true remains bounded and starts no resume process
ok - fm-turnend-guard-grok: both spellings are typed and camelCase has deterministic precedence
ok - fm-turnend-guard-grok: malformed, invalidly typed, and missing-prerequisite payloads start neither path
ok - fm-turnend-guard-grok: missing jq and no-supervision-needed stops stay silent and bounded
ok - tracked .claude/settings.json entries: 5 inert under grok, the documented subagent exception still armed, all live under Claude
ok - .codex/hooks.json: Stop hook uses hook process root when payload cwd is outside
ok - .codex/hooks.json: Stop hook ignores nested git root guard scripts
ok - .opencode primary plugin: guard path is anchored to worktree, not directory
ok - .pi primary extension: no-tool and multi-tool runs each inject exactly one guard follow-up
ok - .pi primary extension: delivery failure resets the logical-run latch
ok - fm-turnend-guard --claude: re-blocks a loop-guarded stop while unhealthy and unclaimed (incident regression)
ok - fm-turnend-guard --claude: X-mode-only homes re-block when auto-arm recovery is absent
ok - fm-turnend-guard --claude: a live arming epoch advances once and repeated observation is idempotent
ok - fm-turnend-guard --claude: repeated failed-to-arming races make bounded monotonic progress
ok - fm-turnend-guard --claude: terminal owner boundary excludes a concurrent start without deadlock
ok - fm-turnend-guard --claude: fresh rewake epoch prevents a duplicate continuation for the same event
ok - fm-turnend-guard --claude: an abandoned auto-arm claim no longer allows a blind stop (incident regression)
ok - fm-turnend-guard --claude: a claim whose pid was reused stops counting as recovery even while its entry reads arming
ok - fm-turnend-guard --claude: a hung owner frozen at arming with no watcher beat no longer allows a blind stop
ok - fm-turnend-guard --claude: a live open generation claim owns recovery with no lock held
ok - fm-turnend-guard --claude: a stuck generation claim no longer allows a blind stop
ok - fm-turnend-guard --claude: the terminal path clears an abandoned claim instead of stepping aside silently
ok - fm-turnend-guard --claude: fresh failed epochs preserve and advance monotonic fail-open progression
ok - fm-turnend-guard --claude: integrated fresh failures reach one bounded fail-open, stop continuation, and reset on recovery
ok - fm-turnend-guard --claude: an inert auto-arm's frozen epoch reaches one bounded fail-open and resets on recovery
ok - fm-turnend-guard --claude: a frozen unverified epoch spends the budget yet still blocks
ok - fm-turnend-guard --claude: reset contention preserves all episode state until retry
ok - fm-turnend-guard --claude: concurrent auto-arm and guard resets are idempotent and deadlock-free
ok - fm-turnend-guard --claude: stale rewake epoch does not allow a blind stop
ok - fm-turnend-guard --claude: budget exhaustion alone cannot permit a blind stop
ok - fm-turnend-guard --claude: verified fail-open is loud, bounded, attended, and non-repeating
ok - fm-turnend-guard --claude: fail-open requires both exhausted retries and consumed notice
ok - fm-turnend-guard --claude: away ownership excludes the Stop-autoarm fail-open
ok - fm-turnend-guard: away mode is quiet with a live daemon and no watcher process
ok - fm-turnend-guard --claude: away mode is quiet with a live daemon and no watcher process
ok - fm-turnend-guard: away mode accepts the watcher beat as the daemon's tick
ok - fm-turnend-guard: away mode tolerates the window before the daemon's first tick
ok - fm-turnend-guard: away mode still blocks when the daemon pid is dead
ok - fm-turnend-guard: away mode blocks when the daemon pid was recycled
ok - fm-turnend-guard: away mode still blocks when the daemon stopped ticking
ok - fm-turnend-guard --claude: away mode blocks a dead daemon and never blames the auto-arm
ok - fm-turnend-guard: away mode's identity-less fallback matches the daemon command only
ok - fm-turnend-guard: a live daemon outside away mode changes nothing
ok - fm-turnend-guard --claude: positive watcher recovery resets failure episode state
ok - fm-turnend-guard --claude: bounded claim wait avoids a token-consuming forced continuation
ok - fm-turnend-guard --claude: secondmate home re-blocks unclaimed and allows auto-arm-claimed stops
ok - fm-turnend-guard: a live away-mode daemon satisfies supervision with no watcher holding the lock
ok - fm-turnend-guard: away-mode daemon ownership survives a leftover dead watcher lock
ok - fm-turnend-guard: away mode with no daemon and no watcher still blocks
ok - fm-turnend-guard: away mode blocks on a dead away-mode daemon
ok - fm-turnend-guard: away mode blocks on a pid-reused away-mode daemon lock
ok - fm-turnend-guard: away-mode daemon ownership never substitutes for a fresh beacon
ok - fm-turnend-guard: a daemon lock proves nothing while away mode is off
ok - fm-turnend-guard: away-mode beacon freshness uses the configured daemon tick grace, not the default tick budget
ok - fm-turnend-guard: a dead away-mode daemon still blocks under the configured daemon tick grace
ok - fm-turnend-guard: the configured daemon tick grace is bounded, not unlimited
ok - fm-turnend-guard: with away mode off, the poll-derived grace never applies
FM_TEST_END 2026-09-23T09:29:09Z tests/fm-turnend-guard.test.sh exit=0 duration_ms=71139 gate_skip=false
FM_TEST_SUMMARY total=1 failed=0 skipped_gate=0 duration_ms=71366
FM_TEST_SUMMARY_FAMILY family=watcher-wake-lock count=1 duration_ms=71139 failed=0
FM_TEST_SLOWEST rank=1 script=tests/fm-turnend-guard.test.sh duration_ms=71139
EXIT 0

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 1 issue found → auto-fixed ✅
  • ⚠️ AGENTS.md:411 - The merge kept two opposite no-mistakes flows. Upstream's line in AGENTS.md:411 says "In no-mistakes mode the earlier done [at=&lt;epoch&gt;]: {summary} is the pipeline handoff and is not gated". The bin/fm-dod-lib.sh header (lines 13-15) and fm_dod_should_gate_ship_done (line 379) say the same. The fork's flow, still in AGENTS.md:376-379 and in the rendered DoD ("Do not append done: until there is a PR", "Firstmate does not send a start trigger"), has no pre-validation done at all. Because of this, fm_dod_should_gate_ship_done gates a no-mistakes done: only when its note contains both "PR" and "checks green". Example: a fork worker appends done [at=...]: PR https://... ready. That line skips the named-head reachability gate, so fm-crew-state.sh and fm-pr-check.sh can accept a head that exists only in the worker copy. Fixing the docs alone is mechanical. Gating every no-mistakes ship done, which matches the fork's flow, changes the behavior of upstream's gate. So the remedy needs a ruling: (a) gate every no-mistakes done and drop the handoff wording (recommended), or (b) keep upstream's CI-ready-only gate and only drop the contradicting AGENTS.md line.

🔧 Fix applied.
✅ Re-checked - no issues remain.

🔧 **Test** - 1 issue found → auto-fixed ✅
  • ⚠️ tests/fm-session-lock-ancestry.test.sh:760 - This upstream test expected an orphaned pty-host to be adopted by pid 1. On hosts with a subreaper (here systemd --user, pid 1298) it failed with 'the pty-host was not reparented to init'. That is a host assumption, not a product bug. I fixed it: the check now accepts any live new parent that is not the dead daemon. The file passes on this host. The fix is in the working tree and not yet committed.
  • Live validation: ✅ go - 6 of 8 scenarios driven live against the product
Scenario Result Live Evidence
No-mistakes worker appends done: PR &lt;url&gt; ready with an unpushed head; crew-state reads blocked, then done after the push ✅ pass live dod-gate-transcript.txt: before 739a510 read done, after 7df5f1f read blocked with the named-head reason, and read done after the push
Firstmate acquires the session lock from a live Claude session; the lock names CLAUDE_PID, the .lock-session sidecar holds the id, and the wording says THIS session ✅ pass live session-lock-transcript.txt steps 1-2 and status
Another live Claude-shaped session holds the lock; acquire is refused with NOT THIS SESSION and status says ANOTHER live session ✅ pass live session-lock-transcript.txt step 3
Fleet mutation (fm-send.sh) is refused while another session holds the lock (fork gate kept) ✅ pass live session-lock-transcript.txt step 3b: 'refusing to steer a worker - this session does not hold the fleet lock'
Adversarial: a caller copies the holder's session id but its CLAUDE_PID is outside its ancestry; ownership is refused (upstream trust gate) ✅ pass live session-lock-transcript.txt step 4, exit=1
The holder dies; the session reclaims the lock ✅ pass live session-lock-transcript.txt step 5
Ledger-anchored no-mistakes continuation reads working in crew-state, never failed ⏸️ untested no The prior payload did not establish a live result. Only tests/fm-crew-state.test.sh with a stubbed no-mistakes CLI covered it. A live check needs real no-mistakes runs rows with an active run at an…
Fork additions still work after the merge: memory, quality gate, local landing, agy adapter, turn-end decline, inbox doorbell ⏸️ untested no The prior payload did not establish a live result. Only the script-level tests in targeted-tests.log covered it. A live check needs a Herdr or tmux fleet session with real agent panes, and this gate a…
  • drive-dod-gate.sh &lt;root&gt; &lt;dir&gt;: real bin/fm-crew-state.sh on a real git origin and worker clone, run on 739a510 (before) and 7df5f1f (after)
  • drive-session-lock.sh &lt;root&gt; &lt;dir&gt;: real bin/fm-lock.sh, fm-lock.sh status and bin/fm-send.sh, run from this live Claude Code session with its real CLAUDE_PID and CLAUDE_CODE_SESSION_ID
  • bin/fm-test-run.sh tests/fm-dod-lib.test.sh tests/fm-crew-state.test.sh tests/fm-session-lock-ownership.test.sh tests/fm-session-lock-ancestry.test.sh tests/fm-turnend-guard.test.sh tests/fm-send-inbox.test.sh tests/fm-merge-local.test.sh tests/fm-memory-verify.test.sh tests/fm-quality.test.sh tests/fm-agy-harness.test.sh tests/fm-bearings-snapshot.test.sh tests/fm-fleet-snapshot-view.test.sh
  • bin/fm-test-run.sh tests/fm-session-lock-ancestry.test.sh (re-run after the test fix)

🔧 Fix applied.
✅ Re-checked - no issues remain.

  • Live validation: ✅ go - 5 of 8 scenarios driven live against the product
Scenario Result Live Evidence
A real Claude session runs fm-lock.sh and gets 'THIS session holds the fleet lock'; CLAUDE_PID is line 1 and the id is in .lock-session ✅ pass live session-lock-cli-transcript.txt and the manual fm-lock.sh run: .lock=CLAUDE_PID, .lock-session=session id
A detached process spoofs the same CLAUDE_CODE_SESSION_ID from outside the harness ancestry and is refused (upstream trust-gated id) ✅ pass live setsid run printed 'error: cannot locate harness process in ancestry', exit 1; status still reads 'held by ANOTHER live session'
Another live harness holds the lock: fm-lock.sh refuses with 'NOT THIS SESSION' and status says 'ANOTHER live session' ✅ pass live session-lock-cli-transcript.txt
Fork fleet-mutation gate: fm-teardown.sh refuses while another live session holds the lock ✅ pass live session-lock-cli-transcript.txt: 'refusing to tear a task down - this session does not hold the fleet lock', exit 1
After the foreign holder dies, this session reclaims the lock and the sidecar is rewritten ✅ pass live session-lock-cli-transcript.txt: 'lock acquired: THIS session', .lock-session holds this session id
A no-mistakes ship 'done: PR <url> ready' whose head is only in the worker copy is gated (DoD refuses it, crew-state reads blocked) ⏸️ untested no The prior payload did not establish a live result. Only fm-dod-lib.test.sh and fm-crew-state.test.sh ran, with fake axi output. A live check needs a real no-mistakes daemon and a worker fleet.
A ledger-anchored no-mistakes continuation reads 'working' from run-step, not failed ⏸️ untested no The prior payload did not establish a live result. Only fm-crew-state.test.sh ran, with fake runs data. A live check needs a real no-mistakes run whose pipeline head has moved past the local copy.
Fork additions still work after the merge (memory compile, quality gate, local landing, turn-end report-once decline) ⏸️ untested no The prior payload did not establish a live result. Only the test harnesses ran (memory-compile, quality, merge-local, turnend-guard), all exit 0. A live check needs a running fleet.
  • bin/fm-lock.sh and bin/fm-lock.sh status with FM_STATE_OVERRIDE set to a temp dir, in this live Claude session
  • setsid -f env CLAUDE_PID=... CLAUDE_CODE_SESSION_ID=&lt;same id&gt; bin/fm-lock.sh (a spoofed id from outside the harness ancestry)
  • Foreign lock holder: a copied sleep binary named claude holds .lock; ran fm-lock.sh status, fm-lock.sh, fm-teardown.sh some-task, then killed the holder and ran fm-lock.sh again to reclaim
  • bin/fm-test-run.sh tests/fm-session-lock-ownership.test.sh
  • bin/fm-test-run.sh tests/fm-session-lock-ancestry.test.sh
  • bin/fm-test-run.sh tests/fm-dod-lib.test.sh
  • bin/fm-test-run.sh tests/fm-crew-state.test.sh
  • bin/fm-test-run.sh tests/fm-turnend-guard.test.sh
  • bin/fm-test-run.sh tests/fm-memory-compile.test.sh
  • bin/fm-test-run.sh tests/fm-quality.test.sh
  • bin/fm-test-run.sh tests/fm-merge-local.test.sh
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

kunchenguid and others added 30 commits September 17, 2026 19:24
)

* Improve CI reliability and rebalance full-coverage validation

* no-mistakes(document): Clarify lint partition documentation
…guid#4799)

* Handle Kimi workspace trust dialog

* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers

* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures

* no-mistakes(review): Read visible pane for Kimi trust and ready gates

* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate

* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection

* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca
…er (kunchenguid#4775)

* fix(bin): report a record whose agent is gone once instead of escalating forever

The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.

Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.

fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that was about to escalate, read it once and treat
only its two recovery-grade verdicts - dead (endpoint present, no agent in it)
and missing (endpoint authoritatively absent) - as proof, reporting that record
once and not re-escalating it while it stays that way. Every other verdict,
including alive, ambiguous, unreadable, unverified, and a read that failed
outright, keeps the identical schedule, reason, and escalation count, so a
genuinely wedged live agent is unaffected. The probe costs at most one backend
read per window per threshold, the same budget the declared-wait consult and the
worktree write probe already take.

The report decides nothing about the record's fate: both lanes still held
unlanded work and teardown refusing them was correct, so retiring, relaunching,
or cleaning up stays with the supervisor. The once-only marker is owned entirely
by that function and is dropped by the same read the moment the endpoint stops
reading gone, so a replacement launched into the same window escalates normally
and its own later death is reported again.

Related, and not closed by this: kunchenguid#4412, kunchenguid#4482, kunchenguid#4316.

Tests drive the real watcher against a record whose endpoint does not exist and
pin both directions: dead and missing report once and never advance the count
across later thresholds, while alive, ambiguous, and unreadable endpoints keep
escalating with the identical reason and a climbing count.

* fix(bin): bind the once-only dead report to the pane it reported

Review of the parent commit found a reachable sequence where a later death in
the same window lost its promised report. The marker was keyed on the verdict
string alone and dropped only when a threshold probe read a non-gone verdict,
but probes run only at thresholds: a replacement launched into the same window
that dies without ever being probed alive - it crashes at startup, or works and
then crashes - was absorbed by the previous death's marker. The pane's first
sight yielded only the generic stale wake and every later threshold matched the
stale marker, so the second death never got the detailed once-report that both
the function's own comment and docs/architecture.md promise.

Record the verdict together with the pane hash it was reported for, and absorb a
repeat only while both still match. A replacement churns the pane, which resets
the stale suppressor, wedge timer, and escalation count while no reset site
touches this marker, so the pane half is what tells the second death apart from
the first. The live-probe drop stays as it was.

Clearing the marker at those reset sites instead would re-open unbounded
re-alarming for a dead pane whose display ever ticks, which is the exact defect
the parent commit exists to close.

The noise bound is unchanged: an unchanged dead pane still absorbs on every
later threshold and never advances the escalation count, and every verdict short
of proof still escalates exactly as before.

* no-mistakes(review): Key the dead-record once-marker on the busy incarnation token

* no-mistakes(document): Document dead-record escalation cap in stale-pane config entry

* no-mistakes(document): Add busy-state inventory line to AGENTS.md

* no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path
…id#4854)

Captain holds have no due semantics and are a hold kind, not a Beads issue
type. The create path now waives due.required and maps to native type task.

Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(bin): launch every spawned agent with the compact adviser disabled

Every crewmate, scout, and secondmate Firstmate launches now starts with
COMPACT_ADVISER_DISABLE=1, on a fresh spawn and on a relaunch alike, so an
unattended session never activates the compact adviser.
The value is unconditional: no configuration file gates it and there is no
override, unlike the trace carrier beside it.

Three carriers deliver it, because no single one covers every launch shape.
The pane shell receives an export beside GOTMPDIR, so the agent's own children
inherit it too.
The launch command carries an explicit assignment, prepended outermost so it
wins over any ambient value the pane already held.
The cleared launch environment sets it again at the `env -i` boundary and keeps
COMPACT_ADVISER_DISABLE in the fixed operational floor, which is what preserves
the switch when config/launch-env-allowlist empties the environment, and what
delivers it on a remote host that never had the value.

bin/fm-control.sh relaunch, the bootstrap secondmate relaunch, and the remote
secondmate transport all rebuild their launch through bin/fm-spawn.sh, so they
inherit the same floor.
The captain's own primary session is untouched.

The two new suites drive the real spawn and then execute the launch command the
pane actually received, with the harness replaced by a probe that prints its own
environment, rather than matching script text.
They cover ship and secondmate launches with the allowlist absent and enabled,
the pane export and its ordering, fm-control.sh relaunch, and the full parent to
remote-host chain.

* no-mistakes(review): Export compact-adviser disable across compound launches

* no-mistakes(document): Document spawned-agent compact-adviser environment guarantee
…henguid#4894)

* fix(bin): let a background Claude session keep owning its session lock

Session-lock ownership was decided by process ancestry alone. Under an
unattended Claude session the model loop runs in a transient bg-spare
bridged to the front-end by a shared daemon; when that bridge is
recycled the contiguous claude-named ancestry from a hook to the
recorded owner breaks while the owner pid stays alive, so the Stop
auto-arm stood down as a foreign live owner, the turn-end guard ended
every turn with its read-only diagnostic, and fm-lock.sh refused - a
self-sustaining outage until restart.

Ownership is now ancestry membership OR a trusted same-session id,
never id-first:

- fm-session-lock-lib.sh accepts CLAUDE_CODE_SESSION_ID only when
  CLAUDE_PID is a Claude-shaped member of the current contiguous run,
  compares it against the id recorded in state/.lock-session, and
  requires the recorded pid to still be a live harness. No id, no
  sidecar, an untrusted id, a different id, or a dead recorded pid
  leaves the ancestry verdict unchanged. Ids are never read from ps
  argv.
- fm-lock.sh accepts a same-session holder at both refusal sites,
  writes, refreshes, and clears the sidecar only under its claim lock
  (including the early already-mine exit, skipped only while the
  deferred startup sweep leases that lock), keeps it byte-identical
  across a same-session confirmation, records CLAUDE_PID on lock line 1
  for a session with a trusted id so a shared daemon or front-end that
  outlives the session never keeps a dead session's lock alive, never
  rewrites a live line 1 on a same-session confirmation, and names the
  recorded id in the live-owner refusal.
- The .lock line-1 format is unchanged, so every reader that takes the
  whole first line as the pid keeps working; the guard's foreign-owner
  exit is unchanged and inherits the fix through the shared predicate.

Tests: the ancestry suite drives the ancestry and id signals apart in a
deterministic process table (asserting the divergence) and runs a real
orphaned front-end/daemon/pty-host/spare tree through six phases with
the real lock, auto-arm, and guard scripts; the foreign-owner repro
keeps its negative control and adds a same-id positive control.

Disclosure: no live unattended Claude background session ran on the
verifying machine. The topology is documented by the real process
listings in kunchenguid#3902, kunchenguid#2314, kunchenguid#3398, and kunchenguid#4066; coverage is the structural
predicate plus the executable fixtures, not a live pass.

Residual: bin/fm-sessionstart-nudge.sh keeps its own private ancestry
walk (it only decides whether to print a nudge) and may nudge on a
resume in the recycled case.

Out of scope, deliberately: no structured lock format, no guard budget
changes, no daemon-identity rejection, no fork lineage.

* no-mistakes(review): Wait for claim lock; revert failed sidecars

* no-mistakes(review): Revalidate ownership after wait; restore sidecars

* no-mistakes(review): Roll back sidecar by publication phase

* no-mistakes(review): Restore sidecar only if lock line is unchanged

* no-mistakes(review): Trust session ids without a spelling allowlist

* no-mistakes(review): Disarm sidecar rollback before backup cleanup

* no-mistakes(document): Updated session-lock ownership documentation
* feat: park main under the away posture on Pi

While the away-posture record exists on a Pi primary, the supervision branch
takes every actionable wake, no processing turn opens on main, captain rows
accumulate for the return brief, and main's standing authority relocates to
the branch through the existing guarded scripts.

- lib/fm-branch-dispatch.ts: read the record at every routing decision; while
  it exists claim check, decision-owned, and heartbeat rows too, keeping the
  two broken-queue vetoes; expose checkSeqs so a claimed check row lifts task
  scoping.
- fm-primary-pi-watch.ts: offer every actionable row under the record; a
  declined wake and every watcher-failure alarm still reach main.
- fm-branch-supervision.ts: drop the legacy .afk decline; append a fixed
  POSTURE: AWAY tail carrying the record's read-back verbatim per wake; open no
  processing request while the record exists, re-checked immediately before a
  request would open and at every run boundary; present the accumulated rows
  at the first run boundary after archive.
- fm-lease-lib.sh: fm_lease_forbid_branch passes the branch for opted-in
  actions only while fm-afk-contract.sh validate succeeds on a confirmed live
  record; PR merge, fresh spawn, and decision answer opt in, local landing
  never does.
- fm-send.sh: a --resolve-key naming an open needs-decision or captain-held
  task is a decision answer and meets the partition; blocked: keys stay
  steering.
- fm-spawn.sh: enforce the record's spend cap for a fresh ordinary spawn by
  either actor; relaunches and secondmates exempt.
- fm-branch-prompt.sh: fixed Postures section and the verbatim
  ask-user-authority policy; the prefix stays byte-stable.
- fm-afk-return.sh: count what the away session handled from the store.
- docs, afk skill, AGENTS.md stub: main parked on Pi, green merge gate
  absolute while away.
- tests: watcher and branch extension suites, fleet-record, merge, and
  decision-answer suites cover the relocation, the vetoes, the tail, the
  parked processing turn, the cancellation, the re-presentation, and the
  spend cap; dated live-guard evidence recorded.

* no-mistakes(review): Refuse branch merge after preflight archive race

* no-mistakes(review): Fix away wake, spawn, and processing races

* no-mistakes(review): Suppress parked processing; narrow away-only rejection

* no-mistakes(review): Abort dedicated processing; gate branch spawn once

* no-mistakes(review): Stamp away-only on the dispatch offer

* no-mistakes(review): Treat invalid away records as spend-cap absence

* no-mistakes(review): Drop spawn test hook; abort processing-opened runs

* no-mistakes(review): Bind abort to opening prompt; cap-read absence

* no-mistakes(review): Limit away branch spawn to queued work only

* no-mistakes(document): Correct AFK posture documentation
* ci: simplify CI job timeouts to a three-tier policy

Replace the scattered per-job timeout values (10m parallel, 25m lint, 30m
serial, 10m macOS) with three readable tiers, each a hang tripwire with
headroom rather than a packing estimate:

- fast (5m): coverage guard, repo invariants, timing aggregate
- normal (30m, one shared budget): lint partitions, portable parallel
  shards, portable serial shards, macOS stock Bash
- heavy (Herdr only): 20m step tripwire on the family run so always()
  cleanup still runs, under a 75m job-level last-resort backstop

The workflow's header comment states the policy and points at
docs/fm-test-portable-shards.md "Timeouts", which now owns it, and each
job names its tier beside timeout-minutes. tests/fm-ci-workflow.test.sh
asserts the policy against the parsed workflow instead of the old
per-job minute values: every job joins exactly one tier, exactly three
distinct job-level values exist, the fast tier stays within 5-10
minutes, the normal budget stays at least double the modeled parallel
lane sum reported by fm-test-run.sh --check-coverage, and the Herdr step
tripwire stays below its job backstop with an always() cleanup after it.

Concurrency supersession, shard counts, lane membership, and fail-fast
settings are unchanged.

* no-mistakes(review): Decouple the normal timeout from packing estimates

* no-mistakes(review): Assert Herdr teardown follows the family run

* no-mistakes(review): Pin Herdr family-run timeout to 20 minutes

* no-mistakes(review): Ignore comments when identifying Herdr steps

* no-mistakes(review): Identify Herdr steps by declarative ids

* no-mistakes(document): Clarify authoritative three-tier timeout policy
…nchenguid#4895)

* fix(bin): keep supervisor status closes from waking the same home

A drain that already folded OPEN DECISIONS has presented those bytes even
when the watcher has no matching seen marker. Treat that fold, and the
presentation cursor, as known so the bookkeeping close stays quiet while
later worker lines still signal.

* no-mistakes(review): Keep folded worker failures waking past supervisor closes

* no-mistakes(review): Wake on unlisted folded worker lines; batch multi-key closes

* no-mistakes(review): Stop folded worker resolved lines from counting as already read

* no-mistakes(document): Correct self-announced close marker contract in docs
* Stop steering operators away from Herdr

* no-mistakes(review): Neutralize remaining Herdr opt-out documentation wording
…enguid#4973)

* fix(bin): treat a live no-mistakes run as current after rebase

A running run on the task's branch is authoritative regardless of head.
Matching only the local head made a rebased in-flight run look failed.

* no-mistakes(review): restrict coarse live-any-head to foreign-branch answers

* no-mistakes(review): reject gate-parked runs from the executing predicate

* no-mistakes(review): hoist gate-marker patterns into single run-lib owner

* no-mistakes(review): require live daemon for head-free run binding

* no-mistakes(review): require answered daemon-down before unbinding live runs

* no-mistakes(review): extend daemon guard to anchored continuation routes

* no-mistakes(review): delete live-any-head; restore dead-daemon verdict

* no-mistakes(review): keep parked gates parked; name dead daemon everywhere

* no-mistakes(review): set dead-daemon verdict instead of emitting early

* no-mistakes(review): align selected route with legacy dead-daemon handling

* no-mistakes(review): drop unproven-record binds; narrow coarse gate reading

* no-mistakes(review): narrow header, drop vestigial guard, retarget tests

* no-mistakes(review): revert coarse gate override; require answered-down probe

* no-mistakes(review): cache one daemon probe; stop duplicating run id

* no-mistakes(review): restrict coarse dead-daemon verdict to moved-off rows

* no-mistakes(review): delete coarse dead-daemon extension and gate note

* no-mistakes(review): delete remaining coarse dead-daemon block and stale docs

* no-mistakes(document): document rebase-safe live-run bind and unverified-record verdict
…4994)

* fix(bin): stage the launch command in a private file and type a short source line

A long launch line typed while the fresh pane shell is still busy waits in the
terminal's canonical line buffer, which drops input past about 1,024 bytes on
macOS, so the pane was left at an unfinished command with no agent running.
fm-spawn now writes the assembled command to the task's own temp root under
umask 077 and types only a short line that sources it.

Refs kunchenguid#4559

* fix(bin): keep the per-task temp root private before staging the launch command

The root lives at a predictable path under /tmp and now holds the whole launch
command. Create it with mode 0700, refuse one that already exists as anything but
a directory owned by this user that nobody else can write, and tighten an owned
one, so no other local user can plant or swap the staged file.

Refs kunchenguid#4559

* fix(bin): enforce private staged launch file mode

* test(spawn): cover long staged Claude launches

* no-mistakes(review): Namespace launch files and prove truncation staging

* no-mistakes(review): Use immutable per-spawn launch filenames

* no-mistakes(document): Document staged launch delivery safeguards

* no-mistakes(ci): Updated eight behavior tests/fakes to execute or inspect immutable staged launch files instead of expecting inline launch commands. This restores Muse, secondmate lifecycle/restart, remote trace/parent binding, compact-adviser, and Orca coverage. All affected tests, dispatch-profile regression, fixture tests, syntax checks, ShellCheck, and git diff checks pass

---------

Co-authored-by: Vytautas Stankus <svycka@gmail.com>
* Add isolated Herdr runbook to test instructions

* no-mistakes(review): Drop substring matching from test.instructions contract

* no-mistakes(review): Assert commands.test key absence in YAML

* Drop unit-first sentence and instructions contract test

Captain-scoped follow-up on the Herdr-lab test.instructions ship:
keep the lab safety runbook only, and leave the no-mistakes contract
test focused on commands.test absence.
…uid#4873) (kunchenguid#5001)

* docs(vision): accept vendor-semantics and 9k contract-ceiling amendments (kunchenguid#4873)

Replace the pixels-of-today's-UI rule with a quarantined, version-pinned
surface-adapter exception recorded as standing debt. Cap the always-loaded
contract at 9,000 words and require prune-or-trigger before a crossing change
lands.

Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

* docs(vision): restore accepted three-sentence vendor-semantics form (kunchenguid#4873)

Replace the compressed paraphrase with the issue's accepted wording:
a named quarantined version-pinned adapter, expected to break, recorded
as standing debt that never hardens into a shared contract.

Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
…or-owed gate (kunchenguid#4974)

* fix(watch): recheck a gate awaiting a human instead of wedge-escalating it

A lane whose validation run is parked at a gate waiting on a human
decision is correctly quiet, but nothing in its status line says so: the
evidence is the pipeline's own gate state rather than anything the worker
wrote. The wedge timer read that silence as a suspected wedge and climbed
the escalation ladder for as long as the wait lasted, and each escalation
cost a supervising turn. The landed declared-wait consult does not reach
it, because a live ordinary crewmate never reports a declared pause, and
raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for
every lane by the same amount.

The threshold now reads a second, independent record when the status line
accounts for nothing: whether the crew's current state is a gate whose
answer is owed by a human. That is minted only from the gate's own
findings table, by a row whose `action` column is exactly `ask-user`,
located by position out of the table header the way nm_gate_step_row
already reads its row - never searched for over the run payload, where a
finding's free-text description or a branch name satisfies a search just
as well. A gate awaiting the CREWMATE's own answer keeps the unchanged
escalation schedule, reason and demand-deep-inspection wording, because a
crewmate that goes quiet before answering its own gate is exactly the
wedge the ladder exists to catch.

Each kind of wait now carries the human it is on, the action that clears
it, and whether that human is the captain as data alongside the verdict,
rather than as wording chosen per branch where the recheck is written, so
the deferral cannot word one kind of wait as another and a new kind
cannot ship without deciding all of them. A parked gate has no written
record of when its wait began, so its recheck publishes no wait age at
all rather than one read from the quiet window this deferral resets on
every pass, which would report the same small number for a gate of any
age. Like every other captain-facing recheck here it is absorbed in
silence while the away-posture record exists, arming no throttle, so the
recheck is owed in full the moment the record is archived.

The consult runs only in the at-threshold branch that was about to
escalate, beside the worktree walk already there, and only for lanes
whose status line explained nothing.

Closes kunchenguid#3055

* no-mistakes(review): require an unanswered decision before deferring a parked gate

* no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records

* test(watch): pass the pane hash wedge_timer_check now takes

Upstream gave wedge_timer_check a sixth <pane-hash> argument for its
dead-record probe. The malformed-wait-record rounds drive the real function
directly, so they pass one, and stub fm_backend_agent_state to a live agent so
the probe that runs after a refused deferral keeps the unchanged ladder rather
than reading a backend the child shell has none of.

* no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate

* no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling

* feat(watch): make the parked-gate wait deferral opt-in

The wedge timer deferring a lane parked at a validation gate is new
supervision behaviour rather than a restored one, and it decides which
lanes give up the escalation ladder, so it now ships as a default-off
per-home option instead of changing every home on upgrade.

config/wedge-defer-parked-gate arms it. The flag is read before the
decision fold, so an unconfigured home spends no fold or current-state
read, writes no record, and keeps the unchanged escalation schedule,
reasons and demand-deep-inspection wording; a test counts the reader
calls in both directions to pin that.

It is not inherited by secondmate homes: each home supervises its own
crew and owns that trade separately, the same reason
config/turnend-churn-absorb is home-local.

The away-posture absorb returns to leaving the idle timer alone, which
it had restarted only because the costly consult could reach it. A
parked-gate wait is owed to the supervisor rather than the captain, so
it never enters that branch, and the recheck owed on return is again
owed in full the moment the record is archived.

* test(watch): pin that the away-silenced hold leaves the idle timer alone

The absorb no longer restarts the timer, so the recheck owed on return is
owed in full rather than a cadence into the return. Nothing asserted
that, so a restart could be reintroduced silently.

* no-mistakes(review): document away-silence rationale, pin captured gate component

* no-mistakes(test): anchor gate row scan to the braced findings header

* no-mistakes(document): pin same-block gate row invariant in crew-state comment
…uid#5007)

* fix(control): let the owning seat reclaim a task whose endpoint is gone

A destroyed pane or workspace made `missing` a terminal state. Relaunch
accepted only `dead` and said to stop the agent first; exit refused
`missing` and said to reconcile the task first; there is no reconcile
verb. Each command named the other as its prerequisite, so a task whose
terminal went away could not be reclaimed by anything, and a no-mistakes
approval it was parked on had no seat left to answer it.

`missing` is agent-free a fortiori: there is no endpoint, so there is no
agent in it. Widen the existing guards rather than add a verb.

- fm-spawn --relaunch accepts a positively proven `missing` and creates
  one fresh endpoint in the recorded worktree; the record it already
  republishes rebinds the task to it. A `dead` endpoint is still adopted
  in place.
- fm-control exit reports `endpoint-gone` instead of dying, so the
  relaunch transaction's stop step no longer dead-ends, and re-resolves
  the endpoint from the record before verifying the replacement.

The duplicate-agent refusal is untouched: both verdicts come from the
same recovery-grade classifier, which claims `missing` only from positive
absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The
backends' own create paths refuse a live same-labeled endpoint as a
second independent guard. The worktree, its branch, commits, uncommitted
changes, armed poll and registration, record rows, and status log are all
untouched - a reclaim is a recovery, never a teardown.

A secondmate is excluded: its gone-endpoint recovery already has one
owner in the session-start liveness sweep, so relaunch refuses and names
it rather than becoming a second path to the same outcome.

Tests reproduce both halves of the deadlock, the reclaim succeeding,
unlanded work surviving it, and the refusals that still hold.

* no-mistakes(review): prove endpoint absence per backend before reclaim rebinds

* no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session

* no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly

* no-mistakes(review): stop refusals and docs asserting unestablished causes

* no-mistakes(review): stop herdr fixture helper losing tmp-root registration

* no-mistakes(review): document workspace drift and absence-probe server residue

* no-mistakes(review): correct rebind limitation to its one reachable case

* no-mistakes(review): stop claiming reclaim leaves instructions untouched

* no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage

* no-mistakes(rebase): read the staged launch file in the herdr fixture

Rebasing onto main picked up kunchenguid#4994, which stages a long worker launch
command into a script and delivers the short `. '<path>'` line instead of
the literal command. The tmux fake and tests/fixtures.sh were updated for
that; the herdr fake this branch adds was written before it and still
keyed "an agent now exists on this pane" off the literal
`encode launch-brief` text, so after the rebase it never marked the
rebound pane live and the reclaim's alive-wait read `dead`.

Dereference the staged file first, exactly as the tmux fake above does.
Test-fixture only; no production path changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(document): note reclaim placement in herdr and scripts inventories

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…3764)

* test(status): reproduce missing event emission time

* wip(status): preserve optional event emission time

* test(status): document indirect clock stub invocation

* no-mistakes(review): Preserve historical status bytes during reply recovery

* no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies

* no-mistakes(review): Preserve captain regex overrides for timestamped status events

* no-mistakes(document): Clarify status event timing and publication contracts

* no-mistakes(lint): Quote literal done to satisfy ShellCheck

* no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed

* no-mistakes(test): Preserve terminal notifications with malformed timestamp tags

* no-mistakes(test): Stamp Rovo spawn failures with emission time

* no-mistakes(document): Verify status event documentation

* no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests

* no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed

* no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed

* no-mistakes(review): Stamp remote escalations at call sites, drop new flag

* no-mistakes(review): Accept stamped escalation and close lines in test assertions

* no-mistakes(review): Restore reserved-key answered-note guard for stamped closes

* test(status): accept optional emission time in PR-provenance assertions

The kunchenguid#4148 provenance test landed on main with exact unstamped greps.
Parent-channel lines from this branch carry [at=<epoch>], so strip only
that tag before the same exact match. No production change.

* no-mistakes(review): Accept stamped ready signal in PR fallback scrape

* no-mistakes(review): Drop relay flag, stamp parent events at call sites

* no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture

* no-mistakes(review): Accept optional stamp in live cmux drift guard

* no-mistakes(review): Restore original test invocation order in two suites

* no-mistakes(review): Strip only well-formed numeric status time tags

* no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc

* no-mistakes(review): Stamp agy spawn-failure status lines with event time

* fix(bin): normalize status event times in-shell and freeze the budget test clock

Two paths made a status event's emission time cost more than it should.

The captain-relevance fallback piped every line through awk to drop a
well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a
fork per line just to prepare a regex match. Shell parameter expansion does the
same strip with no fork, and the retry-dedup scan now reuses that one helper
instead of carrying a second copy of the rule in awk. The copies had already
drifted: the shell side stripped tags from lines with no colon, which the awk
rule left whole, so a colonless line could be mistaken for one already
recorded. One definition, checked against the awk rule it replaces over the
edge cases and a 4000-line fuzz.

tests/fm-contributions.test.sh froze its fixture clock only in exhaust mode. In
hang mode the poll set DEADLINE to the real now plus a one-second budget, and
when the second ticked before the first forge call the loop broke without ever
calling gh: forge/calls was never written and the assertion failed reading a
missing file. Freezing the clock in both modes removes the dependence on wall
time; the bounded call is still cut by the real timeout, so the observation the
test asserts still starts.

Emission time stays optional on new status records, and legacy or malformed
lines keep an unknown age.

* no-mistakes(review): Stamp ask-user escalation line and fix Kimi status assertion

* no-mistakes(document): Drop stale unstamped done-line spelling from watcher docs

* test: fold emission-time snapshot coverage into the fixture case

Drop the incidental ci.yml 18-to-19 count hunk so the PR no longer
touches workflows. Keep every emission-time assertion by folding it
into test_fixture_snapshot_json.

* no-mistakes(review): replace brief date substitution with epoch placeholder; drop emitted_at_epoch

* no-mistakes(review): align untimed normalizer with epoch parser; tolerate placeholder stamp in PR scrape

* no-mistakes(review): strip undelimited at-tags; correct brief stamp header

* no-mistakes(review): normalize stamps at both captain-regex sites; restore mtime freshness

* no-mistakes(review): strip colon-bearing stamps for relevance; fix headers and test oracles

* no-mistakes(review): narrow escalation match to stamp tolerance; pin note verb

* no-mistakes(review): read note and key past colon-bearing stamps

* test(status): keep inactive reconcile assertions stamp-tolerant

These two oracles were made stamp-tolerant while resolving one of the
branch's merges from main. The rebase drops merge commits, so that
adaptation was lost and both assertions went back to matching an exact
substring that a stamped line no longer contains: the tag lands before
the colon, so "failed [key=k]: ..." is now "failed [key=k] [at=N]: ...".
Strip a well-formed tag before matching, as the branch's other oracles do.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): unstamp fold colon tests; reserve stamp width in cap

* no-mistakes(document): correct stale unstamped status-line spellings in docs

* no-mistakes(document): quote brief-test literals for lint; correct stamp-helper contract comments

* no-mistakes(ci): rename subshell-local epoch in delivery-race stub

The serialization test overrides fm_pending_reply_mark_delivered inside a
(..) subshell. Its `epoch` local collided with the same name in
status_line_at_epoch/status_stamp_line, which this branch added and this
suite now calls at top level, so ShellCheck 0.11.0 reported SC2030 and
failed Lint 2. The stub already prefixes its other locals with `pending_`
for the same reason; `epoch` was the leftover.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: ship clean Lavish host fixes

* no-mistakes(review): Fix Lavish classifications and fail-closed host loading

* no-mistakes(review): Restore Lavish host state across retries and launches

* no-mistakes(review): Preserve destination Lavish host when configuration is absent

* no-mistakes(document): Document Lavish status and host guarantees
…#5076)

* feat(afk): make the captain's away words the whole mandate

Retire the clause fields, verb list, never-set scan, refused records, and
the per-task merge-grant list from the away-posture record. The record is
now version 2: the captain's words verbatim plus expected return, spend
cap, and reach line; a version 1 record still validates, reads, and
archives so a live away window is never broken by the upgrade.

The supervision branch reads the words at the tail of every wake and acts
on them by its own judgment through the guarded scripts under standing
authority, never by analogy, holding for the return on doubt, and opens
each such outcome summary with "per your away instructions:" so the
return brief can render the words beside the session's account. While the
record exists any green merge runs under away authority (ledger tag
"away"); red merges, --allow-red, asynchronous and queued merges, and
local-only landing stay refused. The branch may file a backlog item the
words explicitly call for before dispatching it under the spend cap.

Tests drive fm-afk-contract.sh, fm-afk-launch.sh, fm-afk-return.sh, and
fm-pr-merge.sh as commands: version 2 written, version 1 read, retired
flags and subcommands refused by name, green merges landing under the
record, red and waived-red refused, the record lock still closing the
authority-read window, and the Pi away tail carrying the words.

* no-mistakes(review): carry the away read-back to the session verbatim

* no-mistakes(review): match the exact away-action marker in the return brief

* no-mistakes(review): refuse a words block truncated by a damaged line

* no-mistakes(document): Refresh away-role contract documentation
…unchenguid#5049)

* fix(bin): render the remote charter's steering-inbox path host-local

A freshly provisioned remote secondmate read a parent-home absolute
steering-inbox path in its charter - a location that exists on no route -
and spent its first turn discovering the gap and filing a blocked
decision for what was a render defect. The seed's remote-copy rewrite now
maps the inbox to the route's host-local parent-route inbox, exactly as
it already maps the reply-log path, so every mention - bare path, listing,
and handled/ acknowledgement - lands host-local.

Both rewrites also become plain assignments, because a quoted substitution
nested inside a double-quoted printf argument leaks literal quotes into
the replacement text on stock macOS bash. The lifecycle suite pins the
corrected render both directions against the real seed, provisioning,
and delivery route, sharing one fixture value between the render truth
and the delivery truth.

Closes kunchenguid#5012

* no-mistakes(document): document remote charter's host-local steering inbox
)

* feat(procevent): route worker-owned Lavish rounds

* no-mistakes(review): drop duplicate artifact field from task-owned registration

* no-mistakes(review): post worker reply once, fix ring label, keep re-arm atomic

* no-mistakes(review): keep worker board owned until terminal round acknowledged

* no-mistakes(review): refuse every retirement of an open worker-owned round

* no-mistakes(review): use real lavish reply flag, isolate reply generations

* no-mistakes(review): drop .posted marker for best-effort reply posting

* no-mistakes(review): consume staged reply after listener setup, refuse orphaned captures

* no-mistakes(review): require a reachable owner, redeliver open rounds, roll back failed re-arms

* no-mistakes(review): re-arm only to acknowledge an open round

* no-mistakes(review): conclude only a still-open terminal round

* no-mistakes(review): record the acknowledgement before retiring the board

* no-mistakes(review): retain the registration across a conclude, qualify terminal docs

* no-mistakes(document): Document worker-owned Lavish round lifecycle
…unchenguid#5107)

* fix(bin): reserve contribution observation budget

* no-mistakes(review): Strengthen slow-read regression test to exceed the poll budget
…ness JSON (kunchenguid#5103)

* feat(bin): add idempotent inbox orders, receipts, replies, and readiness

Let a caller supply a request id when publishing a captain inbox note so a
retry returns the original note instead of creating a second one, including
across the crash window between save and wake announcement. Separate saved
from announced so a failed wake is repairable without enqueueing again.
Add bounded receipts JSON with omission disclosure, a durable primary reply
against a note id, and a read-only readiness projection that can say
unknown instead of inferring liveness from a lock file.

* no-mistakes(review): fix(bin): honest inbox announce, reply cursor, and readiness verdict

* fix(bin): resolve ready from lock-holder ancestry; drop lock status --json

Remove the extra JSON surface from fm-lock.sh so its human status still
always exits zero. Have the readiness projection classify the inspected
home from the lock-holder pid via fm-harness.sh ancestry, with an explicit
FM_SUPERVISION_MODEL still winning and an unknown model when there is no
holder. Prove the yes path when that ancestry names a known harness.

* no-mistakes(review): Harden inbox announce, receipts reads, and reply sequence cursor

* no-mistakes(document): Note read-only lock inspection in scripts inventory

* no-mistakes(lint): Pass missing id argument to malformed-reply test printf

---------

Co-authored-by: cliflacata-svg <304148223+cliflacata-svg@users.noreply.github.com>
…ending text (kunchenguid#5118)

* fix(composer): stop a harness footer row from reading as a composer holding text

A harness draws its own furniture below the composer - a user statusLine, a
permission-mode hint - and the cursorless "bottom-most shape wins" rule looks
exactly there. `→` (U+2192) is Cursor's prompt glyph but ordinary text
everywhere else, so a statusLine opening with `→` was selected as a bare
composer, swallowed the hint row beneath it as wrapped input, and answered
`pending` on a visibly empty pane. `fm_task_inbox_ring` defers on exactly that
verdict, and `bin/fm-watch.sh`'s re-ring calls the same function, so the first
doorbell and every retry were skipped and the worker never saw the steer.

Measured live on 2026-09-20: three of five Claude Code 2.1.236 worker panes on
Herdr 0.8.0 had genuinely empty composers and every one of them was refused.

A separator pair that closed over a bare agent-glyph row is a proven composer
container, so the contiguous non-blank rows below its closing rule are that
composer's footer and are no longer composer candidates. The demotion is bounded
by all three of its own preconditions: a blank row ends the zone, a pair that
closed over no glyph row demotes nothing, and a shape with no separator pair at
all (Cursor's half-block rules) is untouched. Real unsubmitted text in that same
composer, including a stray SGR mouse report left by a click in the pane, still
reads `pending`.

Pinned by two portable regressions and by a new cursorless arm on the live
composer-matrix guard, which re-reads each harness's already-proven-idle pane
the way every non-tmux backend reads it and fails naming the harness and
version when that read is `pending`.

* no-mistakes(review): make composer footer-zone demotion shape-independent

* no-mistakes(review): make footer-zone demotion refuse-only and drop rescan

* no-mistakes(lint): quote probe-absent sentinel to clear ShellCheck SC2100

---------

Co-authored-by: Koen Muller <koen@catapult.nl>
…5115)

Co-authored-by: guanchengh-lgtm <271917158+guanchengh-lgtm@users.noreply.github.com>
… an unreadable runs table (kunchenguid#5114)

* fix(bin): stop misreading a no-run branch as an unreadable runs table

Defect: when `no-mistakes axi status`'s overview is truncated (a task's
own branch has zero rows among the shown ones), fm_nm_select_run's
Python fallback derived the repo identity for its direct SQLite query
from a `repo: <path>` line it expected in the overview text. The real
CLI never emits that line, truncated or not (see the genuine capture at
tests/captures/no-mistakes-v1.70.1/overview.toon, which has only
`count:`/`runs[...]:`), so the lookup always failed and reported
"unreadable runs table" for a task that simply has no run on its
branch. On a fleet with many concurrent runs, every idle-branch task
hits the truncated-overview path routinely, so this fired every few
minutes and drowned genuine unreadable/blocked verdicts in noise.

Fix: derive the repo identity from the task worktree path instead,
which is exactly the value `no-mistakes` records as a repo's
`working_path` (confirmed against the existing capped-overview test
fixtures, which already register repos by worktree path). A worktree
path that is not absolute cannot be matched and still reads as
unreadable rather than being guessed at. Also raise the reader's
SQLite busy timeout from 1s to 30s so ordinary lock contention on a
busy fleet cannot masquerade as an unreadable database.

Safety: every other verdict byte-for-byte unchanged - the repo lookup
still requires exactly one matching row (a genuinely corrupt or
mismatched repos table still reports unreadable, per the existing
`repo` failure-mode test), the branch query and row validation are
untouched, and a zero-row result for the branch still flows through
the same recursive re-parse that already turns an empty `runs[0]{...}`
table into `absent`. Added a regression test
(test_capped_overview_without_repo_line_and_no_runs_reports_absent)
that reproduces the real overview shape - capped, zero rows for the
task's branch, no `repo: ` line - and asserts the crew state falls
through to the pane/busy verdict instead of reporting unknown or
"unreadable". Full fm-crew-state.test.sh suite passes unchanged
otherwise.

* fix: recovered same-branch inventory awk misreads empty result as unreadable

fm_nm_select_run's deep SQLite reader rebuilds a `count:`/`runs[...]:`
overview and re-runs it through the same awk selection pass. When that
rebuilt inventory has zero rows for the branch, the row-matching loop never
executes, so its counters (`seen`) stay at awk's uninitialized empty string
while `expected` and `shown` are plain strings parsed from the header text.
Comparing an uninitialized value against a non-numeric string uses string
comparison, so "" != "0" is true, and the END block takes the "unreadable
runs table" branch instead of falling through to the correct "absent"
verdict for a branch with genuinely zero runs.

Coerce the affected END comparisons with `+0` so they are always numeric,
matching seen/expected/shown/total regardless of whether awk classified
them as strings or numeric strings. A truncated or genuinely malformed
inventory still differs numerically and still reports unreadable.

* no-mistakes(review): bound capped-overview inventory reader and canonicalize worktree lookup

* no-mistakes(review): match recorded repo path first, tolerate duplicate spellings

* no-mistakes(review): revert repo lookup to exact working_path match

* no-mistakes(document): note state-db inventory read under crew-state nm timeout
…ort (kunchenguid#5141)

* fix(bin): require a non-draft pull request before a PR-based done report

A PR-based ship could report done, and merge monitoring could be armed, while the pull request was still a draft. A draft cannot be merged, so the poll waited for an event that could not occur and nobody was asked to merge.

The PR-based definitions of done now require reading the pull request back from the forge and confirming it is not a draft, and a lane that deliberately holds a draft declares a wait instead of done.
bin/fm-pr-check.sh refuses to arm merge monitoring on a draft, naming the draft state, and treats an unreadable draft state as before.
The draft reading now lives in bin/fm-pr-lib.sh and bin/fm-pr-merge.sh uses it, with its refusal to merge a draft unchanged.

Closes kunchenguid#4757

* fix(review): Skip arm-time draft refusal when fm-pr-merge records metadata
* fix(bin): accept quota-axi schema 6 snapshots keyed by provider + accountKey

quota-axi 0.1.47 emits schemaVersion 6 once a provider expands to more
than one account: every provider row carries an accountKey and one
provider id may appear on several rows. fm_quota_json_valid accepted
only schema 5 with unique provider ids, so fm-dispatch-resolve.sh,
fm-quota-choose.sh, and fm-procevent-quota.sh all rejected the live
snapshot and quota-informed dispatch was dead against the current tool.

- bin/fm-quota-axi-lib.sh: the validator accepts schema 6 with
  accountKey required on every row and uniqueness on
  provider + accountKey; schema 5 keeps its exact rules. FM_QUOTA_ROW_JQ
  is the one join every consumer uses: schema 5 binds by provider alone,
  schema 6 binds to the row keyed by the candidate's Pi lane, else the
  provider's default row, else no row (unmeasured, never blocked, never
  by position or summed across accounts).
- bin/fm-quota-choose.sh: accepts schema 6 JSON and the TOON accountKey
  column, and joins through the shared function.
- bin/fm-dispatch-resolve.sh and bin/fm-procevent-quota.sh: join through
  the shared function; an expanded provider with no row for the
  candidate's account is reported as such.
- tests: schema 6 fixtures shaped like the real snapshot, each paired
  with a schema 5 case on the same path; every new case fails on the
  previous scripts and passes now.
- docs: the two sentences naming the row join describe the schema 6 key.

* no-mistakes(review): Fix native Codex quota and expanded provider watches

* no-mistakes(review): Align native Codex account matching across dispatch paths

* no-mistakes(document): Align quota documentation with account-aware snapshots

* no-mistakes(document): Align quota dispatch documentation with account matching

* fix(bin): keep CI lint and the quota watch test portable

- bin/fm-quota-axi-lib.sh: FM_QUOTA_ROW_JQ is read only by the scripts
  that source this library, so full-mode ShellCheck reported SC2034 on
  the assignment; mark it alongside the existing SC2016 disable.
- tests/fm-procevent-quota.test.sh: the schema 6 provider-watch
  assertions used rg, which CI runners do not install, so the case
  failed with 'rg: command not found' rather than on behavior; use grep
  like the rest of the file.

* no-mistakes(document): Documented schema-version account-row compatibility
* test: repair Claude live auto-arm regression

* no-mistakes(review): Assert SessionStart digest completeness within its hook_response event

* no-mistakes(document): Consolidate Claude live verification references
Roll the shared require-no-mistakes action to the tagged v1.80.1 SHA and grant pull-requests: read so the check can read PR bodies.
sdivanl and others added 25 commits September 21, 2026 19:30
…nchenguid#5174)

* fix: preserve Pi watcher ownership across session replacement

* no-mistakes(document): Scope Pi predecessor retention away from omp

* no-mistakes(ci): Diagnosed all three failing checks; only one was code-caused. (ci-3, genuine) Stock macOS Bash snapshot compatibility: `tests/fm-pi-watch-extension.test.sh` failed the macOS Bash 3.2 `bash -n` parse sweep with `line 4265: unexpected EOF while looking for matching '`. I built GNU Bash 3.2.0 from source locally and reproduced it. Root cause: the PR added a comment containing an apostrophe (`// Replacement shutdown deliberately retains module 2's established arm until`) inside a quoted here-document (`<<'EOF'`) nested inside a `$(...)` command substitution. Bash 3.2 has a parser bug (fixed in later bash) where an unmatched single quote inside such a here-doc body is treated as opening a shell quote and never closed, aborting the whole file parse. The base commit parses cleanly under Bash 3.2, confirming this PR introduced the break. Minimal fix: reworded the comment to remove the apostrophe (`... retains the established module-2 arm until`), preserving meaning. Verified `bin/fm-lint.sh --list-files` (the 6 changed shell files) now all pass `/tmp/bash-3.2/bash -n`; Bash 5 also parses. (ci-1, infrastructure) Behavior portable serial 8: GitHub API shows the `Run portable serial shard 8` step conclusion=success; only `Upload portable serial shard 8 timing artifact` failed with `Failed to FinalizeArtifact ... (403) Forbidden`. This is a transient artifact-service/cancellation failure, not a test or code failure. No change. (ci-2, infrastructure) Lint 1: fetched the job log via the GitHub API; it ends with `##[error]The runner has received a shutdown signal...` then exit 143. The step was cancelled mid-run, not a ShellCheck finding. Independently ran `bin/fm-lint.sh --partition 1of2 --telemetry ...` locally with pinned ShellCheck 0.11.0 and actionlint 1.7.12: exited rc=0 (no findings). No change. The only code change is the apostrophe removal in tests/fm-pi-watch-extension.test.sh; no other files modified
…d#5236)

* fix(bin): retire windowless leftovers and stop claiming a Pi daemon teardown

Catch-up correctly refuses while a leftover task record has no status file.
Cleanup used to deadlock on those same records when they also had no spawn_gen and no window, so they lingered and wedged every later away-mode return. Teardown now treats a windowless leftover as a missing-endpoint legacy record, and stop reports that no daemon terminal was running when none was launched.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Narrow windowless teardown exception to tmux legacy leftovers

* no-mistakes(review): Validate windowless leftover identity via shared endpoint validator

* no-mistakes(review): Refuse windowless leftovers carrying other backends' endpoint identity

* no-mistakes(document): Clarify windowless teardown retry documentation

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…rted (kunchenguid#5250)

* fix: surface parked launch prompts as not started

* no-mistakes(document): docs: record launch-prompt busy backstop classification

* no-mistakes(document): docs: align tail40 and rendered-text comments with launch-prompt backstop
* feat(afk): make /afk itself the go with a same-turn record write

Collapse the propose-then-confirm away entry into one 'enter' step that
writes state/.afk-contract immediately and prints the announcement and
read-back after the record exists, never asking for a go. The retired
propose, confirm, and --proposal inputs are refused by name, and a stale
proposal left by an older version is removed rather than promoted.
Refresh and replace semantics, verbatim words, the single writer, the
never-set, and per-harness launch behavior are unchanged.

* no-mistakes(document): Refresh away-entry documentation evidence
…nguid#5294)

* fix(bin): map passed-with-override to done instead of unknown

no-mistakes' axi status emits outcome: passed-with-override for a run
that finished with an explicitly approved Test or CI exception. Both
bin/fm-crew-state.sh's outcome resolver and bin/fm-teardown.sh's
pre-teardown terminal-run check only matched the literal passed and
checks-passed tokens, so this outcome fell through to unknown/parked
and a finished worker awaiting merge kept getting re-alerted as stale,
while an abort race during teardown could also leave a finished run
misreported as still parked.

Map passed-with-override to the same done/terminal handling as a
clean passed in both places.

* fix(document): Replace stale outcome mapping with authoritative pointer

* fix(ci): Fixed a pre-existing mock-clock race in tests/fm-contributions.test.sh by advancing time only during the serial issue read. Reproduced the exact CI failure before fixing it. Forced-race replay, all 38 contribution scenarios, scoped ShellCheck, Bash syntax, and diff checks pass. Only the test fixture changed; CI rerun remains with the outer executor
* fix: close landed workers from supervision in both postures and at return

During the 2026-09-22 away window every exemption worker whose pull request
had merged was left sitting for nine hours. The supervision branch received
the stale wake, the merge-landed check, and the hourly inactive-outcome row
for each of them, ran the recovery playbook, found nothing to recover, and
reported "no further action". The branch prompt granted ordinary teardown of
a confirmed-landed task without ever naming the moment or the command, and
the playbook has no landed exit, so the stale path ended at "nothing to
recover". The return brief then listed only blockers, decisions, and the
latest five routine outcomes, so the landed workers stayed invisible after
the captain came back.

- bin/fm-branch-prompt.sh: name the merge-landed wake, and any later stale,
  inactive-outcome, or heartbeat row on a done task with a merged PR, as the
  moment to claim the lease and run bin/fm-teardown.sh with no flags; a
  refusal is reported, never forced or worked around. Add teardown to the
  handling tool list.
- stuck-crewmate-recovery: a landed worker is not a recovery case; point at
  the ordinary teardown owner for each actor.
- bin/fm-afk-return.sh: render a "Landed, cleanup due" section from durable
  records only (a live task record whose recorded PR carries the
  merge-notification marker), between could-not-fix and handled, without
  holding the gate; the afk skill's return step closes each listed task
  through ordinary teardown once the check clears.
- tests: pin the prompt rule in fm-branch-supervision and the brief section
  in fm-afk-return through the real marker writer.

* no-mistakes(document): Document landed-task cleanup ownership
* fix(bin): surface a green no-mistakes PR still in ci merge monitoring

A green PR could sit unreported because neither the worker nor the
supervisor could observe checks-green while the ci step kept monitoring
for the merge.

Supervisor read: fm_nm_select_run's capped-overview inventory reader looked
the repository up by the task worktree path, but no-mistakes registers a
repository once by its main clone path and resolves every linked worktree
to it, so on every task copy of a busy repo the lookup matched no row and
each read reported "complete same-branch run inventory unreadable". Key the
lookup on the overview's own top-level `repo:` line, which every axi
release emits as the resolved working_path.

Even with a readable run, the ci-log classifier treated "base branch
advanced ..., re-arming CI monitor timeout" as not-ready. The monitor logs
a checks state only when it changes and a base advance does not clear
readiness, so a green PR read as still validating for as long as main kept
advancing. Stop treating that line as a marker, matching no-mistakes' own
ci-log parser, and name the run's PR URL in the held-for-merge reading so
the existing inactive-outcome path can act on it without a worker report.

Worker contract: `axi status` never reports checks-passed while the ci
step monitors for merge, so the definition of done no longer makes a
status poll the wait for the next gate or outcome; the drive call's own
return is the green signal, reattached with `no-mistakes axi run` after a
bounded return.

* no-mistakes(review): read the full ci log when checking checks-green

* no-mistakes(review): correct stale ci log tail wording in docs

* no-mistakes(document): Document checks-green supervisor fallback
* fix: derive Lavish polling server from its board session

* no-mistakes(document): Document session-derived Lavish polling

* no-mistakes(document): Correct Lavish routing verification claims
… vanish (kunchenguid#4900)

* fix(bin): ignore vanished state scratch files on secondmate relaunch

Relaunch refused when find(1) exited non-zero while listing a secondmate
home's state directory. A live watcher can delete scratch files between
readdir and processing, which is not evidence that child *.meta records
are unreadable.

Prove the directory is listable from its mode and keep the existing
readable-meta loop as the child-record guarantee. Fixes kunchenguid#4765.

* no-mistakes(review): Skip chmod-000 unlistable-state relaunch test when running as root
…d#4907)

* fix(bin): treat home-owned status closes as already read

Self-announced bookkeeping appends now record their exact byte ranges.
Later drains and signal scans skip those ranges, so two distinct
--resolve-key answers after an OPEN DECISIONS fold do not each wake the
supervisor. Worker-authored lines outside that ledger still signal.

* no-mistakes(review): Keep owned closes in unread status; lock ledger writes

* no-mistakes(review): Drop fold-lag wake suppression so folded worker decisions still wake

* no-mistakes(review): Require real owned growth before ledger marks status seen

* no-mistakes(document): Clarify home-appends ledger scope versus UNREAD STATUS

* no-mistakes(review): Restore fold-lag path, drop owned-range filters, fix test

* no-mistakes(review): Align ledger docs and scope ledger to wake path only

* no-mistakes(review): Restore stranded historical-annotation test comment to its function

* no-mistakes(review): Retire the home-appends lock alongside its ledger

* no-mistakes(document): Note ledger's lock-helper dependency in classify library

* no-mistakes(review): Append-and-coalesce home-appends ledger; fix stamped-line assertions

* no-mistakes(review): Drop redundant empty-span branch; make owned test pin ledger

* no-mistakes(document): Document covers' ascending-order dependency on home-appends ledger

* no-mistakes(document): Note owned-append skip in watcher signal-scan comment
…nguid#5350)

* chore(bin): raise tasks-axi, quota-axi, and lavish-axi floors to latest

Raise the minimum versions to tasks-axi 0.2.6, quota-axi 0.1.50, and
lavish-axi 0.1.77, pin CI's tasks-axi install to 0.2.6, and move the
floor-boundary test fixtures to the new versions.

tasks-axi 0.2.6 makes a failed relation deliverable for a promised-final
expecting pr-merged, so add the regression test: a bound work that ends
failed reports its honest outcome text through fm-public-followup-emit.sh,
consume marks the commitment ready, and deliver posts that text exactly
once.

Also make two hang-guard tests in fm-backlog-atomicity portable to hosts
without coreutils timeout, and stop an installed herdr from leaking into
the secondmate-liveness husk classifier test.

* no-mistakes(review): drop out-of-scope bounded_run hang-guard helper from atomicity test

* no-mistakes(review): pin quota-axi floor at 0.1.49 across fixtures

* no-mistakes(document): Document failed public-followup delivery behavior

* no-mistakes(ci): Updated quota-axi floor and all 0.1.49 fixtures to 0.1.51, corrected bootstrap boundaries to 0.1.51/0.1.52/0.1.50, and bumped the bearings lavish-axi stub to 0.1.77. Bearings, quota procevent, quota chooser, startup budget, and bootstrap floor coverage passed; the full bootstrap suite exceeded the 240-second local command limit after relevant checks passed. git diff --check passed
…rker copy (kunchenguid#4878)

* fix(bin): refuse ship done: when the named head lives only in the worker copy

A ship done: is not current-state done until that exact commit is reachable
outside the disposable copy. The check tests the named head, not whether
some branch moved.

* fix(bin): gate CI-ready ship done: on named-head reachability, not handoff

Keep no-mistakes' first done: as the pipeline handoff, apply the same shared
check when registering a PR and when a secondmate publishes ledger-first,
treat a recorded merged PR as landed after prune, and name the PR head
instead of scanning free-text SHAs.

* no-mistakes(review): Bind named-head gate to recorded PR and forge heads

* no-mistakes(review): Gate direct-PR forge heads and keep pending ledger deliveries

* no-mistakes(review): Align worker done wording, test mapping, pending-retry test

* no-mistakes(test): Raise watcher test time limit to stop load flake

* no-mistakes(document): Restore ledger-path fact and name named-head gate coverage

* ci: re-attest named-head ship-done gate for a fresh serial-3 verdict

* no-mistakes(review): Simplify local-only gate, gate keyed done lines, document recovery

* no-mistakes(document): Name fm-crew-state among named-head gate callers
…all alarm (kunchenguid#5204)

* fix(bin): ring a proven-idle secondmate before a wake-loop stall alarm

A leftover foreign-queue row on an idle, alive, ring-safe mate is still drainable in that home. Ring once, reset the observation interval, and keep the parent alarm for unknown, busy, or still-frozen rows.

* no-mistakes(review): Mark drain steer with from-firstmate fire-and-forget carrier
…unchenguid#5335)

The re-arm recovery cases judged "the watcher stayed live instead of
surfacing recovery" with fixed budgets below what a real stale-lock
recovery costs on a contended host: the arm's default 10s confirmation
deadline, a start helper that returned after about 4s whether or not the
arm had confirmed its watcher, and an 80-poll exit wait.
A changed-suite run beside other suites starves the recovery's many
short-lived processes while this suite's sleeping poll loops keep their
pace, so a watcher still surfacing its recovery read as one that stayed
live (issue kunchenguid#3793).
The original 0.25s window after confirmation was widened to 80 polls in
kunchenguid#3837, which left the same race at a larger size.

Following the CONTRIBUTING.md fixture-budget rule, the re-arm helper now
gives the arm an explicit 30s confirmation budget and waits for its
confirmation or exit within a ceiling that outlasts it, and every wait on
a re-armed watcher uses one named iteration-counted ceiling that outlasts
the same budget.
A passing case returns as soon as the arm reports or exits, and a watcher
that never surfaces its recovery still fails.

A new case delays every mktemp and readlink the re-armed watcher runs
after it publishes its beacon, so its first poll and exit take about 13s
on any host.
It fails with the reported symptom on the previous budgets and passes now.
No bin/ change.
* fix(bin): let one TERM always stop the watcher on bash 5.2

Bash 5.2 runs a pending trap from the parser entry of the next command
substitution it expands, where the trap body is parsed as the inside of
that substitution and fails ("trap: line 2: unexpected EOF while looking
for matching `)'") or is dropped silently, consuming the signal. The
watcher's `trap 'exit 1' HUP INT TERM` could therefore ignore a TERM and
keep polling while its stopper waited: the triage suite's reap waited
forever (CI jobs cancelled at 30 minutes), and the arm's signal path and
the away-mode daemon's shutdown wait for the watcher the same way.
Bash 5.3 fixed the parser; 5.2 is the stock bash on Ubuntu 24.04.

HUP and TERM now keep bash's native fatal-signal handling, which runs the
EXIT trap (watcher_cleanup) and exits on bash 3.2, 5.2, and 5.3. INT keeps
its trap because bash ignores a direct SIGINT while a child runs. The
check-spawn deferral window no longer contains a command substitution.

The triage suite's reap is now bounded and fails the case within 10s with
process evidence instead of hanging the job, and a new regression test
proves TERM stops a watcher blocked inside a poll's pane capture and still
releases its lock and records an acknowledgeable stop.

* no-mistakes(document): Clarify watcher stop-signal documentation
…id#5374)

* fix(bin): submit our own stuck doorbell instead of skipping every later ring

* no-mistakes(review): Confirm and retry Enter once on stuck-doorbell submit

* no-mistakes(document): Clarify doorbell retry and pending-composer documentation
Brings the 47 upstream commits since 888871d into the fork and resolves
the 38 conflicted files, keeping both sides wherever they are compatible.

Session-lock identity follows upstream's contract: ownership is harness
ancestry or a trusted Claude session id (CLAUDE_PID must be a
Claude-shaped member of the caller's ancestry, the id must match
state/.lock-session, and the recorded pid must be alive), with upstream's
sidecar hardening and lock inspection. The fork keeps its fleet-mutation
gate (fm_require_session_lock, fm_session_lock_held_by_other) on top of
that verdict, its ownership wording on every lock line, and the
report-once turn-end decline for non-Claude harnesses. The fork's
.lock.session sidecar and conversation-id-only inheritance are retired.

Other resolutions: Kimi trust dialog handles both the 0.36 and 2.0 shapes;
run attribution keeps the fork's head-identity verdict and adds upstream's
executing-run route; the definition of done keeps worker-started
validation and adds status stamps, the non-draft PR check, and the
pushed-head gate; CI adopts upstream's three-tier timeouts; the watcher
wedge timer keeps the fork's busy recheck and adds upstream's dead-record
probe and parked-gate deferral.
- fm-spawn: drop a leftover call to the removed kimi_trust_dialog_is_showing
  from the merged Kimi readiness loop, and clear CLAUDE_PID and
  CLAUDE_CODE_SESSION_ID with an unset statement instead of an env -u prefix,
  so a compound raw launch (cd <dir> && <agent>) still runs.
- fm-brief: stamp the fork-only worker signals (default-branch check,
  secondmate captain decision, dreamer brief) with [at=<epoch>] like every
  other scaffold signal.
- tests: align fork tests with the adopted upstream behavior (executing-run
  binding, stamped DoD lines, the lint partition matrix and nine serial
  shards, upstream's wait_for_exit diagnostics), plant the Kimi shared temp
  root after the fixture claims it, re-fixture the session-start helm case
  on harness ancestry, and give the dead-record wedge tests a verdict that is
  not provably working, since the fork suppresses wedge escalation for one.
The previous fork merge dropped upstream's ledger anchor, so a run whose head
the task copy never fetched stayed "identity unverified" even when the
pipeline's own ledger proved it was this task's continuation. During a daemon
outage that hid a parked gate, an open decision, or the crew's own blocker.

- fm_nm_runs_status_for_worktree recognizes the anchored continuation again:
  the branch's newest ledger row is active and the row immediately before it
  ended at exactly the local head.
- crew-state binds that run on the selected route, naming a dead daemon
  instead of an identity failure, and keeps a branch-matching anchored run's
  full detail on the legacy route.
- The fork's ternary identity, submitted-head binding, executing-run route,
  and no-terminal-verdict-when-unbound rule stay for everything else.
- Teardown still never aborts a run on ledger rows alone.
The fork's Herdr self-deadlock test still used the retired propose and
confirm steps, which upstream removed with the wait-for-go gate.
… home spent 3 current-state read(s) on a parked gate over three thresholds"): this is a real merge clash that fails every time. I reproduced it locally 3 times. Upstream's test expects 0 crew-state reads when config/wedge-defer-parked-gate is absent. The fork keeps its semantic busy recheck in wedge_timer_check (commit e85be9e; crew_is_provably_working runs before each non-busy wedge escalation). That recheck spends exactly one read per threshold, so the unarmed count is 3. Fix: tests/fm-watch-triage.test.sh now expects exactly 3 reads, one busy recheck per threshold. The comments now explain why. The test still catches the regression it guards: a flag guard placed after the parked-gate consult would add a consult read to every threshold. Production code is unchanged, so both behaviors survive. The single test passes locally. bin/fm-lint.sh is clean. When I returned this result, the full triage file had passed 87 tests with 0 failures and was still running. Serial 8 (tests/fm-bearings-board-render.test.sh, "source ... is not listening after reconcile"): this is not caused by this PR. bin/fm-bearings-board.sh and that test are identical to upstream c576c2b, except for the stub's lavish-axi version string. The fork's fm-procevent differences touch only the extension-adapter paths, not the lavish register/reconcile launch. The whole test file passed 3 times normally, 3 times with every CPU core saturated, and 8 times as parallel copies. The failure is a CI-only slow-runner timing hiccup: the detached listener did not claim within the 3s confirm plus 5s await window. No change was made for it
…est.sh, "settled history starved the active contribution"): this is a real merge clash. The failure log shows "bin/fm-contributions.sh: line 191: DEADLINE - : syntax error", so `date +%s` returned an empty string. The fork's test test_settled_history_does_not_starve_open_contributions has a fake `date` that reads and then rewrites $FORGE/clock with `>`. Upstream's fm-contributions.sh now runs six forge reads at the same time, and each one calls `date +%s`. One call can truncate the clock file while another reads it, and the reader gets an empty value. I reproduced the race directly: 6 parallel calls, 300 rounds, old stub gave 124 bad reads, new stub gave 0. Fix (test only): every write to the fake clock now goes to a temp file and is moved into place with `mv`, which replaces it in one step. This covers the fork's ticking date stub and the gh stub's reserve/exhaust/fail-late advances (a new small `advance` helper in the gh stub). Production code is unchanged. The full tests/fm-contributions.test.sh passed 3 times in a row locally. bin/fm-lint.sh (changed-file mode) and shellcheck on the file are clean. Lint 1: this is not caused by this PR. The job got SIGTERM (exit 143) after 7m54s. That is well under its 30-minute timeout, and the log has no lint finding. Lint 2 passed in the same run. Lint 1 passed on the previous run (8m19s), and the only change since then was one test file. This looks like the runner being shut down. I started a full local `bin/fm-lint.sh --partition 1of2` run. This machine was too slow to finish it: after about 20 minutes it was still in the ShellCheck extended analysis, so I stopped it. No change was made for Lint 1
@BohnBawerick
BohnBawerick merged commit 71bd89c into main Sep 23, 2026
19 checks passed
@BohnBawerick
BohnBawerick deleted the fm/fm-upstream-sync-0923 branch October 4, 2026 02:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.