Skip to content

Upstream catch-up: merge upstream/main (251 commits) with 163 conflicts resolved - #19

Merged
jorguez96 merged 253 commits into
mainfrom
fm/fm-fork-sync-2026-09-24
Sep 25, 2026
Merged

jorguez96 merged 253 commits into
mainfrom
fm/fm-fork-sync-2026-09-24

Conversation

@jorguez96

@jorguez96 jorguez96 commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

Bring the fork's main up to date with upstream's main: 251 commits behind, reconciled as a merge that keeps every local customization intact.

Method

bin/fm-upstream-sync.sh sync --dry-run (owned mechanics, opened nothing) reported 163 conflicted files.
Direct 3-way merge against the recorded merge-base over-conflicts: the prior catch-up (#14) landed as a squash, so the recorded base predates 165 already-absorbed upstream commits.
The squash point was re-anchored exactly (the 165th commit after the old base: 3eb5b633), proven by a 42-file diff that matches the fork's known customization footprint.
Every conflicted file was then classified against that anchor.

What synced cleanly

  • 44 upstream-added and 92 upstream-modified files merged with no conflicts (298 files changed total, +37303/-4635).
  • 146 conflicted files took upstream verbatim: their fork content was byte-identical to the squash anchor, so the fork never customized them (115 both-modified, 31 both-added-that-were-upstream-additions).
  • 13 conflicted files auto-combined on the anchor base: fork customization and upstream evolution in disjoint regions (skills, ci.yml, .gitignore, README.md, fm-fleet-snapshot.sh, fm-session-start.sh, fm-spawn.sh, documentation-audiences.json, 3 test files).

Hand-resolved (4 files)

  • bin/fm-brief.sh: fork's user-invoked-skill section and upstream's new shared-infra rule occupy the same anchor; kept both heredoc definitions, fork's first. Emission sites had merged cleanly.
  • bin/fm-test-run.sh: conflicting serial weight tables (CI-measured durations, not sizes) merged as union - upstream's refreshed measurements plus the 2 fork-only tests (fm-upstream-sync, fm-herdr-attached-viewer-live-e2e).
  • docs/configuration.md: took upstream's supervision-branch paragraph. The fork's sentences described pre-squash behavior that contradicts the merged implementation and the owner doc docs/pi-supervision-branch.md (legacy .afk flag semantics, away-posture authority relocation).
  • AGENTS.md (5 hunks): kept the layout pointer (new knobs all live in docs/configuration.md); blended the digest section (fork pointer, upstream bootstrap wording, added devin to verified harnesses - merged code supports it); restored the quota-array procedure the skill says AGENTS.md owns; kept the validation condensation (substance owned by no-mistakes-validate, ask-user-authority, task-landing, fm-dod-lib, fm-crew-state); restored the escalation final-message rule (section 9 owns escalation style, S1 had it, no other owner).

Verification

  • bin/fm-lint.sh exit 0 over the full sync footprint; bin/fm-doc-audience-check.sh ok; shellcheck clean on both hand-edited scripts; bash -n clean on boot scripts with help output sane.
  • Tests green: fm-brief, fm-upstream-sync (9/9), fm-spawn-dispatch-profile, fm-test-run.sh --check-coverage ok (both fork tests scheduled).
  • Two failures are environmental, proven by running the same tests on pure upstream/main in this container: fm-session-lock-ancestry (pty reparenting needs a real init) and fm-ci-workflow (only the ruby requirement; no ruby here).
  • Full fm-lint.sh --partition timed out locally after 20 minutes; CI runs it.
  • One new blank-line-at-EOF whitespace nit in upstream's bin/fm-brief-heading-lib.sh left verbatim.

Deferred / for the captain's review

  • Observation for future auto-syncs: upstream's own tests/fm-startup-network.test.sh contains the fixture GITHUB_TOKEN=ghp_supersecretvalue, which trips fm-upstream-sync.sh's secret pattern. Even a conflict-free future sync will refuse to land until that fixture is allowlisted or the pattern narrowed. Left untouched as upstream content.
  • The AGENTS.md escalation restoration (hunk 5) reverses the fork's earlier condensation; merge word is the captain's, so adjust there if the shorter form is preferred.
  • Merge authority stays with the captain: this PR is not merged, nothing was pushed to main, upstream was never touched.

Base reconciliation (2026-09-25, commit c6ec2a2)

Out of scope (untouched): no-mistakes pipeline, secondmate homes.

kunchenguid and others added 30 commits September 1, 2026 19:10
)

* fix(bin): resolve process-event state roots before validating them

The process-event module validated the caller's spelling of a home's state
root instead of the directory it operates on: it required the supplied path
to equal its own lexical normalization, which rejects any path reached
through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks,
so an operator home under either could never claim a source. Reconcile still
reported the runner started, while the detached runner died writing "cannot
claim source" to the discarded stderr, and the source silently never fired.

Resolve the state root to its physical directory once, then apply the
existing private-directory validation to that resolved directory and derive
every path, recorded claim identity, and later confinement check from it.
This keeps the confinement contract for the directory actually operated on
rather than only for callers that already spelled it physically, and removes
the window where an ancestor symlink could be repointed between check and
use. Homes already spelled physically behave identically.

This was the single cause of both deterministic macOS failures in
tests/fm-procevent.test.sh ("reconcile never claimed the registered source")
and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not
produce an outcome"). The new case pins the behavior with an explicit
symlinked-ancestor home, so it fails without the fix on any platform rather
than only where the temp root happens to be a symlink.

* fix(bin): pin the external capture staging boundary to its physical path

The extension capture path pinned its registry staging boundary by comparing
`pwd -P` against the caller-spelled registry directory, so a home reached
through a symlinked ancestor still refused to start an extension-backed
source after the state root itself resolved correctly. That left such a home
half working: built-in sources ran while external ones failed.

The staging preparer now prints the physical registry directory it validated,
matching the inbox and reservation preparers beside it, and the start path
pins on that returned path. The new end-to-end case drives the shipped
file-signal package from a symlinked home spelling.

* no-mistakes(review): Propagate canonical process-event state roots

* no-mistakes(review): Propagate canonical state to process-event adapters

* no-mistakes(document): Document physical process-event state roots
…kunchenguid#3312)

* fix(pi): persist captain outcomes visibly

* no-mistakes(review): Recover captain outcomes after cold-start lock acquisition

* no-mistakes(document): Document cold-start captain-outcome recovery

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* fix(pi): process captain outcomes through a sequence-keyed turn

PR kunchenguid#3312 made every captain-facing supervision outcome a durable, exact-once
visible transcript entry with the read cursor advancing only after that entry
exists. That is the display half of the delivery contract. Left alone it turns
a probabilistic silent loss into a deterministic one: the captain sees an
anchor line, and firstmate never acts, because nothing opens a turn and
nothing records whether main ever processed the outcome.

The 2026-08-31 timeline showed the two shapes this must survive on the
previous hidden-turn path: seven delivered decision outcomes each answered by
an empty assistant message (cursor advanced, no retry, unanswered for close
to three hours), and two answered by an unrelated prior reply. Both happened
because delivery advanced the cursor at enqueue and accepted whatever the
next assistant message was.

Add the processing half on top of the persistence half:

- bin/fm-branch-outcome.sh keeps a processed marker separate from the read
  cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It
  only advances through an explicit sequence-bound acknowledgement, never
  past the read cursor and never backwards; an absent marker reads as zero
  and `processed-init` migrates delivered history once so an upgraded home
  is not re-presented its past.
- After the visible entry for a captain outcome exists, the extension hands
  every still-unprocessed captain row to main as one hidden, typed
  `fm-branch-process` request listing each `[seq N] task: summary`, opening
  exactly one main turn. Main closes it only by calling the new
  `fm_branch_processed` tool with the highest sequence listed. An unrelated,
  empty, or paraphrased answer leaves the sequence open, and the same request
  is presented again at the end of the next main run and at session start.
  The first two presentations of a sequence set open a turn of their own;
  after that the request rides the captain's next prompt so an ignored
  request cannot loop, and a session replacement resets that budget.
  Routine outcomes stay turn-free.
- The regressions cover exactly those incident shapes against the real store
  scripts: an empty answer and an unrelated prior answer neither advance the
  marker nor stop re-presentation, the acknowledgement is refused beyond the
  read cursor and outside lock ownership, a partial acknowledgement keeps the
  newer sequence open, and kunchenguid#3312's own assertions now forbid an unkeyed turn
  rather than any turn. The store suite pins the marker's bounds and the
  migration; the real-SDK guard for appendEntry persistence and model
  exclusion is unchanged.

Docs move the protocol from "no model turn" to "one sequence-keyed processing
turn closed only by its acknowledgement", and the verification record carries
the dated run against Pi 0.84.4.

* no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements

* no-mistakes(review): Harden outcome state validation and request pacing

* no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores

* no-mistakes(review): Validate canonical mark-read cursor state

* no-mistakes(review): Guard cursor advancement against corrupt processed state

* no-mistakes(review): Bind acknowledgements to active processing requests

* no-mistakes(review): Reset pacing when processing sequence membership changes

* no-mistakes(review): Enforce silent outcome invariants at storage boundary

* no-mistakes(document): Document hardened captain outcome processing contracts

---------

Co-authored-by: kunchenguid <kun@kunchenguid.com>
…3481)

* feat: bound Bearings remote ledger collection

* no-mistakes(review): Clarify default remote-ledger collection behavior

* no-mistakes(review): Detach reconcile delivery from watcher loop

* no-mistakes(review): Enforce bounded snapshot and request captures

* no-mistakes(review): Bound legacy summary capture before parsing

* no-mistakes(review): Bound primary remote ledger captures

* no-mistakes(document): Correct snapshot and reconcile documentation

* no-mistakes(lint): Fix ShellCheck quoting in bounded collector

* no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks

* test: await reconcile request retirement

* no-mistakes(review): Avoid empty reconcile queue process churn

* no-mistakes(review): Read ledger summaries from immutable snapshots

* no-mistakes(review): Reject multi-document home ledger streams

* no-mistakes(review): Coalesce durable reconcile requests per target

* no-mistakes(review): Unify reconcile keys and reject snapshot streams

* no-mistakes(review): Key reconcile requests by stable target ID

* no-mistakes(document): Document per-target reconcile request coalescing

* no-mistakes(lint): Remove unused snapshot summary file variable

* no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass

* no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks
* fix(ci): rebalance the portable serial shards on measured durations

The "Behavior portable serial 3" shard ran 17-20 minutes against its
20-minute job cap and intermittently timed out seconds after a passing
test, on branches and on main alike.

Shards are packed longest-processing-time from per-script duration hints,
and those hints were last measured on 2026-08-21 at 116 scripts. The lane
has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had
no hint at all and fell back to the 20 s default, and several existing
hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured,
fm-public-followup 36 s vs 197 s). The partition therefore looked
perfectly balanced in hint space, 734.6 s per shard, while really running
11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the
tests asserted, stayed normal throughout and hid it.

Refresh the hints from the timing artifacts of three green runs, taking
the slowest measurement of each script so the balance holds on a slow
runner, and split the lane across five shards instead of four. Replayed
against those runs' real per-script durations the worst shard is now
12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's
wall clock drops from ~20 to ~12.5 minutes.

Bound the drift that caused this rather than relying on the hints being
refreshed by hand: the coverage guard now reports the unmeasured share as
serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT,
which leaves room for newly added tests while making a stale table fail
the guard instead of silently pushing one shard into its cap.

No test changes what it asserts and no test stops running; only the
partition across shards changes.

* no-mistakes(document): Clarify conservative shard timing aggregate
…uid#3491)

* fix(pi): fall back after settled branch errors

* no-mistakes(review): Detect provider errors across prompt compaction

* no-mistakes(review): Preserve in-flight branch state across selection changes
* fix(pi): recover supervision branch after cooldown

* no-mistakes(review): Defer branch recovery until prompt settlement

* no-mistakes(document): Clarify supervision cooldown recovery contract
* refactor: remove legacy remote summary reads

* no-mistakes(document): Document ledger-only snapshot reads

* no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass

* no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean

* no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux
…henguid#3498)

* fix(pi): rearm watcher after session replacement

* no-mistakes(review): Queue actionable closes across Pi session replacement

* no-mistakes(review): Stop replacement arm when handoff persistence fails

* no-mistakes(review): Preserve actionable wakes through branch and late child races

* no-mistakes(review): Surface late handoff failures without crashing Pi

* no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens

* no-mistakes(review): Retry stale deliveries and release settled claims

* no-mistakes(review): Distinguish branch settlement and retry handoff cleanup

* no-mistakes(review): Deduplicate persistent handoff cleanup alerts

* no-mistakes(review): Acknowledge watcher follow-ups only when consumed

* no-mistakes(review): Persist idle follow-ups until agent consumption

* no-mistakes(review): Preserve pending outcomes when handoff persistence fails

* no-mistakes(review): Arm replacement before awaiting prior delivery settlement

* no-mistakes(review): Adopt pending handoffs after lock reclamation

* no-mistakes(review): Prevent stale generations from adopting replacement handoffs

* no-mistakes(review): Scope replacement handoffs by watcher state

* no-mistakes(document): Clarify replacement handoff documentation

* no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks

* no-mistakes(review): Update branch settlement tests and preserve chunked outcomes

* no-mistakes(document): Document watcher-owned replacement handoffs

* no-mistakes(document): Verify replacement handoff documentation

* test(pi): cover watcher-owned branch fallback

* no-mistakes(document): Refresh watcher-owned fallback documentation
…d#3495)

* fix(bin): resurface terminal statuses lost after branch handling

* test(watch): canonicalize process-event fixture homes

* no-mistakes(review): Index branch outcomes by causal status position

* no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses

* no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics

* no-mistakes(review): Keep unclassifiable oversized statuses silent

* no-mistakes(document): Document lost-wake outcome backstop

* no-mistakes(document): Update outcome backstop documentation

* no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally

* no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes

* no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift

* no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state
…id#3503)

* fix(bin): deliver typed terminal results from remote work homes

A public commitment whose work is bound to a REMOTE secondmate home could
never receive its typed terminal result. `fm-public-followup.sh brief`
printed an emit command carrying this home's own absolute path and this
checkout's own script path, neither of which exists on the machine the
worker runs on, so the worker had nothing it could write to that the
owning home would ever read - and `consume` kept finding nothing while
the promise stayed open.

The brief is now route-aware: for a remote work home it prints that
route's own code root and home with `--stage-in`, so the typed event is
staged in the home where the work actually runs, and the closing
paragraph names the owning home as the one on the other machine instead
of pointing at the path above it. The owning home collects those staged
results over the same SSH route it reaches that secondmate on, because
the transport only runs outbound: `consume` pulls them into its own
inbox and reconciles them exactly as it reconciles a local report.
Collection is non-destructive until the result is durably held, so a
dropped connection cannot lose a terminal result, and a route that could
not be reached is named in `consume`'s output with the promise left open
rather than reported as an empty inbox.

A local work home is untouched: the brief still prints `--home` with this
home and this checkout's script, and the event still lands directly in
this home's typed terminal-result inbox.

This is the emit-side counterpart of the retire/clear fix in kunchenguid#3479 and
reuses the remote-route resolution that landed with it. Reconciling a
loop bound to a remote route now reaches that route, so the existing
remote cases drive `consume` through the same faked transport their
other steps already use.

* no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes

* no-mistakes(review): Fail collection when remote outbox is unreadable

* no-mistakes(review): Surface reassigned remote routes during empty collection

* no-mistakes(review): Fail remote collection on invalid registrations

* no-mistakes(review): Reject unsafe registration entries during remote collection

* no-mistakes(review): Restore healthy empty remote collection behavior

* no-mistakes(review): Skip remote collection for delivered registrations

* no-mistakes(review): Skip delivered registrations before route validation

* no-mistakes(document): Document remote follow-up collection semantics
…#3504)

* fix(bin): exclude secondmates from home-summary child inventory

kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed.

* no-mistakes(review): Cover terminal secondmate in-flight exclusion

* no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds
* fix(bin): self-heal status-outcome indexes on every drain

Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault.

* no-mistakes(review): Guard held-lock initialization and fail marker writes

* no-mistakes(document): Document cross-harness outcome-index self-healing
…nchenguid#3505)

* fix(bearings): keep active children underway beside a captain hold

Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work.

* no-mistakes(review): Preserve Underway repos and disclose child truncation

* no-mistakes(review): Fall back to task project for Underway repos

* no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean
…enguid#3513)

* fix(pi): settle watcher delivery on Pi accepting the follow-up

A follow-up queued while main is streaming joins the running run without
ever raising before_agent_start, so waiting on that event before clearing
the successor pipeline (kunchenguid#3498) stalled every later actionable close: no
successor started, no wake was delivered or offered to the branch, and the
turn-end guard woke main to re-arm by hand after every close.

The pipeline now settles once Pi accepts the follow-up. Consumption is
observed at before_agent_start for an idle main and at the user
message_start for a streaming main, and decides only what a replacement
session (/new, /resume, /fork, reload) replays. An exhausted restoration
delivers its typed failure without launching an arm past the retry bound,
which the stall had hidden. The replacement-coordinator map is typed so the
strict no-emit typecheck passes again.

Tests: the doubles no longer raise before_agent_start for a streaming send,
a portable regression drives two actionable closes while main streams and
proves the successor chain plus consumption-scoped replay, and a
credential-free real-SDK probe pins Pi's event contract for both the
streaming and the idle follow-up.

Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a

* fix(pi): retry a verified successor that fails during wake delivery

A verified successor can exit while the wake it was started for is still
being delivered, most plausibly during a branch turn that holds the
settlement for minutes. Its failure close arrived while the pipeline's
single-flight guard was set, so the close handler skipped the retry, and
the pipeline's end no longer launched an arm, which left the live
generation with no watcher and no retry timer.

The close handler now records that failure when the child had reported
readiness and was not retired by the restoration itself, and the pipeline
runs the ordinary bounded, lock-checked retry for it once the delivery
settles. A restoration started for a later pending supersedes it, and an
exhausted restoration still hands repair to main without a further arm.

The regression holds a branch settlement open while the verified
successor exits with a failure and proves one retry watcher starts after
the settlement releases, none while it is held.

Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a
* fix(bin): bound repeat stale wakes for a parked but live worker

A worker parked on a declared wait - `paused:` for an external or pipeline
wait, or a verified `captain-held` transfer - kept waking firstmate far inside
FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one
captain-held worker and dozens across a day on a pipeline wait, and reported
upstream as four wakes in 75 minutes against a 3600s window.

pause_state_class deliberately answers `none` for a still-live agent even under
a declared wait, so a worker genuinely waiting on a decision is never silenced.
That classification is correct and is left alone; it routes every parked but
live worker through surface_nonterminal_stale on first sight of each distinct
stale hash, and an idle parked pane still churns its hash on a clock or a token
counter without changing what is being waited on.

Two places let that churn re-alarm:

- surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was
  declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should
  have suppressed it. The throttle was never read on this path and was advanced
  by the wake it should have prevented.
- The hash-change path cleared that throttle through clear_pause_tracking
  whenever the classification came back `none`, so each tick also bought the same
  declared wait a fresh window. Fixing only the first site changes nothing.

Read the throttle before anything is queued and advance it only on a wake that
really fires, and on the hash-change path reset only the per-hash bookkeeping
while the declaration still stands, via a clear_stale_hash_tracking split so
neither half of clear_pause_tracking is duplicated. The throttle is keyed to the
declaration, not to the pane.

First sight still wakes, so an inconclusive state is still inspected, and the
window's end still re-surfaces once, so a forgotten wait cannot rot invisibly -
noise traded for a bounded cadence, never for silence. The wake identity stays
the plain `stale: <win>` the away-mode handoff depends on.

Tests cover both observed forms and were confirmed to fail against three
deliberate breaks: each site reverted on its own, and a re-surface that never
fires again.

* fix(document): Clarify declared-wait wake cadence documentation

* fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor

* fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed
…er (kunchenguid#3567)

* fix(turnend): accept the away-mode daemon as the supervision owner

While state/.afk exists the away-mode daemon owns supervision and runs
bin/fm-watch.sh one-shot: the watcher exits on every wake and the daemon
starts its replacement. The turn-end guard tested for a live watcher
process holding the watch lock at that instant, so a turn boundary that
landed in the hand-off blocked with "TURN WOULD END BLIND" while
supervision was completely healthy, costing a full handling turn each
time.

Reproduced with the real daemon wrapping the real watcher and the real
guard sampling the same home: 6 of 40 samples blocked, every one of them
with the daemon alive and the beacon 2-3 seconds old, and a new watcher
pid on each cycle. After the fix the same reproduction blocks 0 of 40,
and killing the daemon and its watcher (away mode still on, beacon still
fresh) blocks again.

The guard now accepts a live, identity-matched daemon holding this home
as proof of supervision while away mode is active. The identity match is
the same discipline the watcher lock uses, so a recycled pid or a lock
left by a killed daemon proves nothing. The fresh-beacon half of the
predicate is unchanged: a daemon that stops restarting its watcher still
blocks once the beacon passes grace, a home with no supervisor blocks
exactly as before, and with away mode off the strict watcher predicate is
untouched.

The predicate reads only durable state, so it behaves identically for
every primary harness and runtime backend.

* no-mistakes(document): clarify away-mode daemon supervision proof and test coverage

* no-mistakes(document): generalize stale turn-end predicate summary in architecture.md
…unchenguid#3582)

* fix(backlog): omit markdown file for beads probes

* no-mistakes(document): Narrow backlog addressing doc to mutations for backend-aware probes

* no-mistakes(ci): Fixed the Greptile P2 review comment (the only failing check) on tests/fm-backlog-atomicity.test.sh. The comment correctly noted that an exported TASKS_AXI_BACKEND environment variable would inherit into the spawned scripts and, because fm_tasks_axi_backend gives it top precedence, override each test case's .tasks.toml backend fixture — making the backend-specific argv assertions fail for environmental reasons. Fix: unset TASKS_AXI_BACKEND in the test harness right after sourcing tests/lib.sh, with a comment explaining why, so every case deterministically exercises its declared backend (4 lines added; no production code touched). Verified: reproduced the leak before the fix (TASKS_AXI_BACKEND=beads made the markdown dispatch case fail with 'beads show failed', exactly the reported failure mode); after the fix the full suite passes (0 failures, exit 0) both with and without TASKS_AXI_BACKEND=beads exported. The added lines are shellcheck-clean (the only shellcheck note, SC1091 on the lib.sh source line, pre-exists this change)
…chenguid#3589)

The supervision branch's verdict rule escalated every outcome that
answered a captain request, so "the work started" and "still working"
notes reached the captain with nothing to look at. The rule now keeps a
finished result of requested work captain-facing, even when healthy, and
treats start or still-working updates that bring no new artifact,
finding, or decision as routine. The captain list for review-ready PRs,
ask-user findings, exhausted blockers, credentials, and destructive or
security-sensitive cases is unchanged, as are the unsolicited-routine,
silent-fleet-review, and doubt-chooses-captain rules.

The fm_branch_report tool description and the two docs that restated the
old unconditional rule now point at the prompt's "Verdict: routine or
captain" section as the one owner instead of carrying a second copy.
* fix(bin): never close a captain call during cleanup

A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` closes a held row silently, and nothing in
teardown asked whether the row was the captain's own call.

bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when
the task is still an open captain call, 1 when it is not, 2 when that
cannot be established. It reads the row through the transition library's
backend-aware probe, so it addresses the same backlog teardown does; the
script's other commands now address the configured data directory the
same way instead of FM_HOME, which also fixes captain holds in a home
with a relocated data directory.

Teardown asks `open` before any destructive step and refuses on 2. On 0
only the close changes: after cleanup and still under the task's own
lock, the row gets one "Deliverable of the finished work" line at the end
of its body and returns to Queued through `tasks-axi reopen`, keeping its
hold, so it lands in Captain's Call instead of reading as work under way.
--force does not lift this: it authorizes discarding unlanded work, never
the captain's question. The deliverable goes into the body because
`tasks-axi update --report` rewrites the title of a row that is not Done.

The crash window reuses the pending-close record teardown already stages:
a `mode=retain` line makes the existing replay record the deliverable and
reopen instead of closing, with the same validator, stale-generation
check, cleanup-incomplete marking, and non-blocking bootstrap lock as an
ordinary close. A retained row the captain answered first simply retires
the record. No parallel record type, recovery command, or second bootstrap
loop is introduced.

Regressions run the real executables: the captain-held scout survives
cleanup queued, held, with its deliverable and on the board, only
`answer` closes it, --force keeps it open, and an ordinary scout still
closes with its report; an interrupted cleanup leaves the row untouched
and the next session start retains it; a relocated backlog keeps the
retention in its one configured file; and a ship row whose hold cannot be
read refuses cleanup before anything destructive.

Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np

* no-mistakes(review): Serialize captain holds and fix backend-aware listing

* no-mistakes(document): Update captain-call retention documentation

* no-mistakes(document): Fix relocated captain-hold backlog diagnostics
…uid#3592)

* fix(bin): deliver every secondmate outcome on the parent channel from the recording scripts

A secondmate's captain-facing outcomes could miss: the mate model addressed
the captain in its own unread chat instead of appending to the parent
channel, and a PR-ready report, a finding, a decision, a blocker, and a
failure all depended on that one remembered append. Make delivery
structural, so the parent channel never depends on the model:

- bin/fm-parent-channel-lib.sh is the one owner of channel resolution and
  exact-line append-once; the merge outcome path and the inactive-outcome
  scan now publish through it instead of two private copies.
- bin/fm-inactive-reconcile.sh gains a ledger-first path that runs on every
  watcher poll in a secondmate home: a direct child's whole terminal done or
  failed line is delivered at once with its note, recorded PR, mode, merge
  posture, and scout report pointer, keyed and receipted so it is delivered
  once, and the inactive path yields to it. `report <task-id>` runs the same
  delivery for a caller holding the child's meta lock.
- bin/fm-pr-check.sh publishes the PR-ready line with the canonical URL at
  registration.
- bin/fm-captain-hold.sh publishes a hold and its answer, keyed by task id
  and resolution-record count, with no new persisted state.
- bin/fm-teardown.sh delivers the child's final line before removing its
  record and refuses, retaining every record, while the channel cannot be
  written.
- The charter opens with the parent-channel rule and confines the mate's own
  appends to judgement; AGENTS.md carries the carve-out at the persona
  address rule and the escalation list.

docs/secondmate-parent-channel.md records the design and its coverage, and
docs/verification/secondmate-parent-channel.md records the live run with real
tmux panes and both real watchers delivering every line with no model.
Supersedes kunchenguid#3569.

* no-mistakes(review): Fix parent outcome retries and reconciliation locking

* no-mistakes(review): Prevent busy children from starving ledger delivery

* no-mistakes(review): Correct ledger metadata and hold occurrence handling

* no-mistakes(review): Disambiguate ledger outcomes and normalize hold reasons

* no-mistakes(review): Close ledger races and preserve teardown records

* no-mistakes(document): Correct parent-channel receipt and scanner documentation

* no-mistakes(lint): Quote done arguments for ShellCheck compliance

* no-mistakes(ci): Fixed both CI failures. Updated GOTMP teardown fixtures for the new final-outcome reporter and isolated them from host tmux state. Updated the PR security assertion to distinguish the accepted PR-ready line from duplicate merge outcomes. Verified with both failing test suites, bash syntax checks, and git diff checks

* no-mistakes(ci): Fixed Greptile’s duplicate-delivery race in bin/fm-inactive-reconcile.sh. Ledger events now claim matching already-delivered inactive receipts using the prior status fingerprint, preventing duplicate parent reports while preserving later same-state completions. Added behavioral regression coverage. Verified inactive-reconcile tests, project lint, documentation audience checks, syntax, and diff checks. Teardown tests passed relevant cases before the documented pre-existing herdr-preflight-missing-adapter failure
* fix(bin): sync remote second-mate homes to the parent primary commit

Session start and remote launch pointed a remote second-mate home at whatever
Firstmate copy its own host kept, so a home that had already advanced past that
copy refused as a non-fast-forward and every other home stopped at the host's
older commit while the primary ran ahead.

The parent now resolves ITS primary default-branch commit with the existing
helper and hands that commit to the host on both paths. Because a remote home
is a standalone clone, the host imports that one commit before advancing -
already present, else from that host's Firstmate copy without moving it, else
from the home's own origin - and then runs the SAME ff_target guards a local
home gets, so dirty, diverged, feature-branch, and unresolvable targets skip
untouched and the ancestry rules keep one owner. An unimportable target now
names /updatefirstmate instead of failing opaquely, and a host still running an
older Firstmate copy is reported the same way rather than echoing a bare
refusal.

The host-local launch leg no longer re-runs its own secondmate sync, so the
spawn it drives cannot re-target that host's copy after the parent has already
converged the home.

/updatefirstmate is unchanged: it still refreshes the remote code root from that
host's origin and then syncs the home to that refreshed copy, which is what the
sync call with no target commit means.

* no-mistakes(document): Document primary-targeted remote secondmate synchronization
)

* fix(bin): split brief task into captain intent and firstmate spec

Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs.

* fix(bin): stop task-subsection copies at the next heading

Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body.

* no-mistakes(review): Validate brief content and preserve nested specifications

* no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies

* no-mistakes(review): Ignore fenced subsection headings during brief validation

* no-mistakes(review): Preserve captain intent across scout promotion

* no-mistakes(review): Enforce safe intent boundaries for legacy promotions

* no-mistakes(review): Allow marked legacy intent and reject empty promotions

* no-mistakes(review): Scope task parsing and overlay legacy intent contracts

* no-mistakes(review): Overlay current intent contract for all no-mistakes spawns

* no-mistakes(review): Preserve later captain clarifications in intent overlays

* no-mistakes(document): Document brief intent enforcement and ownership

* no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed

* no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint
…guid#3600)

* fix(pi): start a new supervision branch conversation per main session

The supervision branch reopened one recorded conversation forever, so
every main session start reloaded the current generated prompt and then
weeks of accumulated thread, where a superseded rule could still outweigh
today's.

The branch conversation is now scoped to one main session: the session
generation owns the recorded conversation, so a cold start, /new,
/resume, /fork, or a reload always builds a new one, while a rebuild
inside one session (a model or effort change) still continues that
session's own conversation.

The dialog mirror re-anchors with it. Its durable cursor records what the
previous branch conversation received, so a /resume or reload - which
keeps main's own session file - would otherwise leave the new branch
blind to dialog main itself still has. The reset is bounded by the
current main session, and the cursor keeps advancing incrementally within
it. The durable outcome store and its processed marker are untouched, so
unacknowledged captain-facing outcomes still re-present on the new main
session.

* no-mistakes(document): Document fresh Pi supervision conversations

* no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks
* feat(update): restart second mates whose instructions changed

/updatefirstmate pulled new bytes onto disk and then asked each advanced
second mate to re-read them. A running agent holds AGENTS.md and every
loaded skill frozen from launch and no verified harness offers a reload,
so that steer could not reach a loaded skill at all and left the mate
holding two contradictory copies of its own job description.

An eligible mate is now restarted instead, in the same home and endpoint,
through the existing transactional relaunch. The restart is gated on the
mate first writing down the open work it holds only in conversation - the
open-record half of /stow, never its memory sweeps - so an unregistered
captain call is flushed before the conversation is spent. Anything that
leaves the reload unprovable falls back to the old re-read message and is
reported as exactly that, never as a clean reload.

Remote mates take the same path: fm-remote-secondmate-control.sh gains a
relaunch verb whose host-local leg runs that same control plane, since the
mate is an ordinary local secondmate from its host's point of view. The
primary resolves the profile and passes it explicitly, because
config/secondmate-harness is not inherited and the file on that host
belongs to a different home.

fm-update.sh now splits its advanced live mates into a restart set and a
nudge residual, and both sets require a changed instruction surface, which
also closes the over-nudge against the session-start sweep. Restart is
stricter still: a bin/-only advance reloads itself on the next call, so it
never costs a conversation.

Colocated tests cover the gating, the persist-then-restart order, the
task-subset persist request, each unsafe fallback, the remote hop, and the
remote sync's new instruction-surface report.

* no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting

* no-mistakes(review): Parallelize relaunches and classify replacement incarnations

* no-mistakes(review): Gate restart actions on live agent state

* no-mistakes(review): Handle failed restart workers without hanging

* no-mistakes(review): Nudge legacy remotes and preserve persist recovery

* no-mistakes(review): Document one-time secondmate restart rollout

* no-mistakes(review): Honor arrived replies and refresh remote profiles

* no-mistakes(review): Revert remote parent profile reconciliation

* no-mistakes(review): Reset remote profile defaults and honor published results

* no-mistakes(review): Preserve fallback nudges for unverifiable secondmates

* no-mistakes(document): Document second-mate restart update flow

* no-mistakes(lint): Fix ShellCheck warnings in restart scripts
…id#3644)

* perf(tests): route gate verification through the bounded concurrent runner

Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.

Three changes, each measured:

- `.no-mistakes.yaml` pins `commands.test` to
  `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
  already owns changed-file selection, bounded concurrency, the refusal of
  unproven scripts, and a generous automatic per-script bound, so the gate's
  baseline is neither a serial chain nor a guessed timeout. It stays
  intent-targeted - the Test step still runs its evidence agent on top - and
  excludes the live-Herdr family the required Herdr lane owns.

- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
  automatic scheduler and automatic bound that `--changed` gets. Naming several
  subjects is how a verification round asks for exactly those scripts. The
  curated selections are untouched: `--lane` still composes CI shards whose
  serial lane must stay serial, `--family` is what the required Herdr lane runs,
  and `--all` stays a deliberate complete regression.

- `pr-forge` is admitted to the concurrent-safe family registry on two
  consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
  and records `secondmate` and `session-bootstrap` as refused with the exact
  script and reason each failed on, so the refusals are actionable rather than
  silent.

Measured on this host, 0 failures on both sides:

  verification round, 4 scripts   448s chained -> 231s through the runner (-48%)
  pr-forge family                 409.2s at 1 worker -> 237.9s at 4 (1.72x)
  watcher-wake-lock family        1311.1s at 1 worker -> 539.3s at 4 (2.43x)

A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.

* no-mistakes(review): Separate concurrent runs by isolation proof family

* no-mistakes(review): Limit automatic timeouts to changed-file validation

* no-mistakes(document): Clarify validation concurrency documentation
* fix: copy PR URLs from records or abstain, never assemble them

Supervision reported a plausible but dead PR link three times because its
prompt demanded a full https:// URL at a moment when only a PR number was
observable, so the model assembled an owner/repository from memory, and the PR
check then accepted that URL and wrote it into the task record, after which the
model kept defending its own tool-endorsed guess over the worker's real link.

Three changes close that chain without any live forge lookup, so private
forges are treated exactly like public ones:

- bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy
  or abstain" section requires a URL to be copied verbatim from a durable
  record (the done: PR <url> status line, pr= metadata, or the backlog note),
  forbids assembling owner, repository, host, or number from memory, and has
  the branch report only the identifier it actually holds when no record names
  the URL yet, leaving the PR check unarmed until the worker's ready line
  arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for
  main in place of the bare full-URL mandate.

- Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full
  https:// URL wherever a PR is mentioned - status line, terminal, or summary -
  never a bare "PR 108", so the link is in view as early as the number is.

- bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that
  the task's own done lines contradict, printing both spellings; a log naming
  no URL still records the argument as before. fm_pr_status_ready_urls in
  bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches
  bin/fm-pr-merge.sh, so nothing merges under a contradicted URL.

Tests cover the offline refusal with zero side effects, the recorded spelling
being accepted, markdown-wrapped and punctuated URLs, working lines not
counting, the merge wrapper propagation, a self-hosted merge request with no
forge call, the prompt carrying the rule, and the brief carrying the worker
rule.

* no-mistakes(review): Remove stale PR URL enforcement

* no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check`
…uid#3661)

* fix(bin): disable Claude's feedback-draft flow for fleet-launched agents

Scope --settings '{"feedbackDrafts":"off"}' to every Firstmate-launched
Claude crewmate and secondmate, so /bug and /feedback never queue or
submit a bug report on the captain's behalf. feedbackDrafts is the
documented settings key (Claude Code changelog 2.1.247); the
per-launch CLI flag never touches the captain's global settings.json.

Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3

* no-mistakes(review): Prevent managed settings from re-enabling Claude feedback drafts

* no-mistakes(document): Fix Claude feedback documentation formatting

* fix(bin): layer both feedback-draft controls for defense in depth

The prior --settings-only fix can be overridden by a managed Claude
settings policy (feedbackDrafts precedence). Keep CLAUDE_CODE_SEND_FEEDBACK=0
alongside --settings '{"feedbackDrafts":"off"}': either control alone
disables the SendFeedback tool, so a managed override of one still
leaves the other in force.

Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3

* no-mistakes(document): Document Claude feedback-draft suppression ownership
…guid#3662)

* perf(tests): admit three more families to concurrent validation

The three families that `docs/fm-test-isolation-proof.md` recorded as refused
were not refused for concurrency. Each blocker was a test that decided a
property by wall clock, or a script filed where it cannot run. Fixing those
three things admits all three families and recovers 28.6 minutes of local
validation with no assertion removed or weakened.

- `tests/fm-backlog-handoff.test.sh` injected its pre-move crash by killing the
  handoff, sleeping a fixed second, then delegating the move to the real
  binary. Nothing ever killed the fake, so on a host slow enough for the case's
  next assertions to take longer than a second, the orphan woke and completed
  the very move the case requires left undone, and recovery then failed with
  `Task "pre-move-crash" not found in this backlog`. Watching the two backlogs
  during the injected crash showed exactly that, the item moving one second
  after the crash. All four crash injections in the file now go through a new
  `fm_fake_crash_injector` shim that signals the target and returns only once
  it is observably gone, and the pre-move fake never delegates the move at all.

- `tests/fm-session-start.test.sh` proved the startup digest does not block on
  a slow current-state read by timing the whole digest against a fixed
  eight-second sleep, which a loaded host exceeds without the property being
  violated. It now holds that read open until the case releases it and asserts,
  the moment the digest returns, that the read has not finished. A digest that
  waited would wait indefinitely rather than for an interval a slow host can
  out-run, so the assertion is stronger than the bound it replaces. Its scan
  budget moves to the maximum, because the old value left two seconds of margin
  over the fixed sleep and measured the host rather than the deadline that
  `tests/fm-inactive-reconcile.test.sh` owns.

- `fm-backend-herdr-focus-flash-e2e` was filed in the family map's catch-all,
  which put it in the portable serial lane, where Linux CI gate-skips it: that
  real-Herdr regression was running nowhere. It moves to `real-herdr-gated` and
  the required Herdr lane. `fm-claude-stop-autoarm-live-e2e` gate-skips on its
  opt-in variable and moves to `live-harness-optin`.

The 28 remaining ungrouped scripts become an enumerated `standalone` family
instead of admitting `unclassified` itself. `unclassified` is the family map's
`*)` arm, so admitting it would silently grant concurrency to every test added
afterwards, which is exactly the population with no proof. A new test still
lands in `unclassified` and stays serial, and `tests/fm-test-run.test.sh`
covers that split behaviorally.

Each family passes two consecutive four-worker proofs with zero failures. On
the production runner, `secondmate` goes 1233.1s to 453.4s, `session-bootstrap`
756.4s to 286.4s, and `standalone` 724.6s to 261.1s: 2.71x overall and 1713.2s
recovered. The whole suite runs 177 scripts in 52.6 minutes of wall clock
against 121 minutes of summed script time.

* no-mistakes(document): Refresh concurrent validation and shard documentation

* no-mistakes(ci): Fixed the real-Herdr focus-flash E2E race exposed by reclassification. Part C now starts its persistent child atomically via `pane run` and verifies stable child identity through Herdr’s public `process-info` interface, avoiding the racy send-text/send-keys sequence and platform-specific `ps` matching. Verified with bash syntax checking, ShellCheck, git diff checks, and the complete E2E test on Herdr 0.8.2
* feat(brief): structure no-mistakes ask-user escalation as event + snapshot file

Crewmates escalating a no-mistakes ask-user gate now report one status
event naming every finding id plus a snapshot file holding the gate's
axi finding records verbatim (id, severity, file, line, description,
authority), using the same shape even for a single finding. The status
line never paraphrases. The format is defined once in fm-dod-lib.sh and
rendered into both the scout and ship rule 6 in fm-brief.sh, so a
promoted scout - whose rule 6 fm-promote.sh preserves unchanged - gets
the identical contract as a freshly-spawned no-mistakes ship worker.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpiWaDerbYavTLPPtEjQei

* no-mistakes(review): Preserve ask-user escalation output contract

* no-mistakes(review): Align escalation format test expectation

* no-mistakes(review): Scope ask-user escalation instructions correctly

* no-mistakes(review): Remove ask-user from generic decision rules

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* fix(bin): require a self-sufficient no-mistakes intent

A no-mistakes worker's --intent is only as useful as the string it
passes. PR kunchenguid#3604 shipped with an intent that was only "do 1, 2, 3, 7
from the report": the real contract lived in a private scout report and
never reached --intent, so nobody holding that string plus the codebase
could have derived the specification.

This is pure instruction at the contract's one owner; no spawn-side or
promotion-side check is added.

- bin/fm-dod-lib.sh: the generated no-mistakes Definition of done now
  states that the --intent string must be self-sufficient (the string
  plus the codebase reconstructs roughly the same specification) and
  tells the worker to write the substance of any report, decision, or
  PR the captain's intent refers to into --intent rather than the
  pointer, while Firstmate build instructions and the worker's own
  decisions still stay out. The spawn-time overlay points back at that
  rule so its "supersedes" wording cannot cancel it, and the header's
  owner statement carries the rule.
- AGENTS.md section 11 and bin/fm-brief.sh's header ask Firstmate to
  include the substance of referenced material when filling
  ## Captain's intent, and section 11 points at the owner of the rule.
- tests/fm-brief.test.sh and tests/fm-task-delivery.test.sh assert the
  rendered brief and launch contract carry the rule.

Claude-Session: https://claude.ai/code/session_01YMhEe42q7BAAoN6RxNuzim

* no-mistakes(document): Replace incident-specific intent test commentary
blackxwhite88 and others added 29 commits September 23, 2026 00:08
…kunchenguid#5338)

* fix(test): repair tmux liveness and calm follow-up loaded_off regressions

Both self-tests fail on untouched main on a host whose coreutils are a
multicall binary and whose Chrome has no pre-warmed profile, and each failure
masks the other's file.

tests/fm-tmux-agent-liveness.test.sh - the stand-in harness processes were
symlinks to the host's `sleep`. A single-purpose `sleep` runs happily under
another name, but a multicall coreutils binary (uutils or busybox) resolves its
applet from argv[0]: `claude-link -> sleep` invoked under the harness name runs
the wrong applet and exits immediately, so no foreground process exists and
every positive case reads not-alive ("last verdict for liveness:agent was
missing (expected alive); title=sh comms=[sh ]"). Build a dedicated spinner as
the stand-in target, exactly the way the version-string case already builds its
executable, and require the fallback target to demonstrably survive the rename
before using it. Every assertion is untouched; the stand-in identity signal is
unchanged (the kernel still records the symlink name as the executable
identity).

tests/fm-calm-pi-extension.test.sh - render_export_dom pinned a brand-new
`--user-data-dir` per attempt. On Google Chrome for Testing 151.0.7922.34 that
pristine profile makes Chrome's first-run initialization never complete: the
browser and its renderers start, but --dump-dom never returns, so all three
bounded attempts end exit=0 timed_out=yes bytes=0 and the DOM assertions never
run ("could not render calm-mode HTML export DOM"). Chrome's own profile
creation under a fresh HOME renders the same document in about a second, so the
helper now gives Chrome a private per-attempt HOME instead of the explicit
profile flag. Each attempt still gets an isolated profile, and every DOM
assertion is unchanged.

Root-cause evidence: a pristine --user-data-dir with `--headless=new
--dump-dom` had not returned after 150s, while the same command with an empty
HOME and no --user-data-dir returned the full DOM in ~1s, and reusing an
already-populated profile also returned it in ~1s. The render failure masked
the rest of the file: with it repaired, the Pi follow-up loaded_off case passes
unmodified against an installed @earendil-works/pi-coding-agent package.

These two failures block downstream validation of every lane on hosts with
multicall coreutils or a fresh Chrome profile.

Verification:
- timeout 300 bash tests/fm-tmux-agent-liveness.test.sh -> exit 0, 16 assertions ok
- timeout 700 bash tests/fm-calm-pi-extension.test.sh -> exit 0, 13 assertions ok,
  including the Pi operational follow-up loaded_off case
- bash -n and shellcheck clean on both touched files
- rest of tests/: bin/fm-test-run.sh --all bounded by timeout 900 completed 17 files with 0 failures (fm-afk-contract.test.sh through fm-backend-herdr-launcher-workspace-e2e.test.sh), then the bound cut off the 18th (fm-backend-herdr-presentation-e2e.test.sh, a real-herdr-gated lab test) with no failure recorded

* fix(test): give wake-queue observation checkpoints the alerting ceiling

tests/fm-wake-queue.test.sh's secondmate stall case runs bounded foreground
watcher checkpoints whose job is to record an observation, with the alerting
checkpoint that follows asserting the stall. A checkpoint's exit publishes a
downtime marker, and the next checkpoint consumes it only by reaching the end of
the watcher's poll loop, where the recovery surfacing runs after the stall tick;
the observation itself is recorded by that same stall tick. On a loaded host a
1s ceiling sits under the cost of that iteration (which includes a pane capture
in the active-turn gate), so the observation was never recorded, the downtime
marker stayed pending, and the alerting checkpoint surfaced
`check: rearm-resurface` instead of the stall it asserts:

  not ok - a foreign queue with no progress did not alert: check: rearm-resurface
  not ok - a frozen reprovisioned queue generation was hidden: check: rearm-resurface

Give the observation checkpoints that feed a later alert the same 4s ceiling the
file already documents for alerting checkpoints. The ceiling is only a bound - a
checkpoint still returns on its first actionable wake - so no assertion is
weakened, and the quiet windows get longer, not shorter.

* no-mistakes(document): docs: correct export-DOM Chrome render root cause

* no-mistakes(review): Isolate Chrome profile on macOS, dedupe tmux CC_BIN lookup

* chore: re-trigger fork workflow approval for triage

---------

Co-authored-by: Captain <blackxwhite88@users.noreply.github.com>
Co-authored-by: kunchenguid <kunchenguid@users.noreply.github.com>
…chenguid#5383)

* fix(bin): classify a status span without re-folding the whole log

A watcher poll could take minutes, so its liveness beacon aged past the
guard's 300s grace and the Stop auto-arm reported the watcher down. On the
main home, cycles ended with beacon_age 91-235s while healthy and 534-706s
while the laptop was CPU-starved.

Cause: whenever a newly appended status span held a keyed needs-decision
or blocked line, status_span_first_actionable_record re-read and re-folded
the ENTIRE log to decide whether that opening was still live, forking
several subshells per line. On a remote second mate's mirrored parent
channel (1.2MB, ~2300 lines) that is 13-20k subshells, about 17s per log
per classification when idle, paid by every signal and heartbeat scan.

Nothing regressed recently: subshell counts per classification were
20,272 from kunchenguid#3268 (2026-08-29, which introduced the whole-log fold) and
13,188 from kunchenguid#3753 onward through HEAD. The cost grew with log size, since
parent-channel logs only grow.

Fix: fold only the captured span. An accepted opening does not depend on
earlier lines and only later lines close or supersede it, and every later
line lies inside the span, so the span fold names the same live openings
at a cost bounded by the span. Old and new classification outputs are
byte-identical across 51 span offsets of real-shaped secondmate and ship
logs.

A real-watcher regression test records every read the classification
makes through the span-reader seam and asserts none reaches before the
classified offset; it fails on the old code (5,157 bytes read from
offset 0 to classify an 84-byte span).

* no-mistakes(document): Clarify span classification and watcher regression coverage
kunchenguid#5362 and kunchenguid#4878) (kunchenguid#5381)

* test: fix watcher timing flakes in fm-pr-check-security

The bounded watcher's hang guard now counts only the watcher's own time: a
case marks the intervals where it holds the watcher on injected work or makes
it wait on concurrent work, and those no longer count against its budget. The
budget itself stays at main's sixty seconds. The helper also stops forcing a
one-second per-check timeout, which killed a correct merged poll whenever that
poll took longer than a second, so the watcher only retried it or exited on a
later check's wake without the merge.

The concurrent-publication case pauses the guard while its arming is in
flight, and its task now sorts before the contributions observer the arming
also registers, so the watcher stops on the poll under test before running
that unrelated fleet snapshot. The case also prints the watcher's stderr when
it fails.

The replacement case pauses the guard while the re-arm runs inside the
watcher, runs that injected arming with the fixture root every other arming
here uses, and waits on the replacement merge's process instead of a
two-second cap. Merged-poll runs retire the contributions observer before the
watcher starts, since no case here exercises it.

The returned-descendant case no longer races a four-second sleep or a TERM
landing at an arbitrary point in the watcher's idle loop: its descendant holds
until killed, and a second check in the same cycle witnesses that it was
drained and stops the watcher.

* no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed

* Revert "no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed"

This reverts commit 6a59859.
…enguid#5385)

* feat: record task.pr_ready in the fleet ledger when a task PR is registered

* feat: record worker status lines in the fleet ledger as they are written

* no-mistakes(review): Keep worker status append failures and pass the resolved config to the ledger

* no-mistakes(review): Resolve relative config override before embedding in worker command

* no-mistakes(document): Clarify fleet ledger status capture timing
…unchenguid#5386)

* test: synchronize foreign secondmate stall legs on the watcher's recorded observation

Each leg of test_secondmate_foreign_queue_stall_tracks_progress_and_alerts_once
ran the watcher under a 1s or 4s wall-clock checkpoint, but every later leg
depends on the progress observation the previous leg's watcher recorded. Under
load the watcher was killed before its first stall tick, the observation was
never written, and the next leg treated its own sighting as the first one, so
the stall alert never fired.

Run the watcher directly and end each leg on its observable outcome: the
progress marker recording the expected observation, or the watcher's own first
wake. Also move a comment orphaned above this test back to the drain liveness
test it describes.

* no-mistakes(review): Wait for full stall reset before stopping watcher leg
kunchenguid#5391)

The listener resolves its server from that store before it polls. Without a session for this board, it exits in the gap after the build has already sampled a live claim.
…#5390)

* fix(bin): prune a torn-down task's wake rows at teardown

Prune pending durable wake rows (.wake-queue) for a task when it is torn
down, clearing stale wakes for its target window, signal wakes for its status
or turn-ended files, and task-specific check wakes.

Fixes kunchenguid#3419.
Adjacent to kunchenguid#5252.

- bin/fm-wake-lib.sh: add fm_wake_queue_prune_task
- bin/fm-teardown.sh: call fm_wake_queue_prune_task in cleanup_firstmate_home_children and main teardown
- tests/fm-wake-queue.test.sh: add test_wake_queue_prune_task

* no-mistakes(document): docs: note teardown prunes a task's wake rows

---------

Co-authored-by: Captain <blackxwhite88@users.noreply.github.com>
…ixture readiness (kunchenguid#5392)

* Make portable tests match resolved host paths

Summary:
- Match macOS full Node command paths by basename in the Gemini behavior test.
- Mirror symlink-resolved Nix PATH behavior and give the loaded-host race bounded headroom.

Testing:
- bin/fm-lint.sh
- bin/fm-test-run.sh tests/fm-on.test.sh tests/fm-gemini-harness.test.sh tests/fm-procevent.test.sh

Related:
- None

* no-mistakes(review): Mirror production PATH helper rules per directory group in tests

* no-mistakes(test): Wait for orphan runner start marker instead of fixed sleep

* no-mistakes(document): Clarify gemini ancestry test comment for versioned node comm

---------

Co-authored-by: Sandeep Salwan <salwansa@amazon.com>
…ns (kunchenguid#5389)

The sibling secondmate stall cases in tests/fm-wake-queue.test.sh now wait for the watcher's recorded observation instead of a one-second wall-clock checkpoint, so they can neither fail nor pass vacuously under load. Deterministic proof with a 5s watcher launch delay: before the fix 4 cases passed vacuously and 6 failed; after it all 10 pass on the recorded observation.

Also includes a CI flake fix from validation: fm_control_harness_supported in bin/fm-control-lib.sh finishes reading the harness allowlist before returning, removing intermittent broken-pipe diagnostics. Behavior is unchanged.
… tail (kunchenguid#5336)

* fix(bin): refuse a Herdr submit that would send only a message tail

A long typed payload can sit in the composer as a suffix, or as a paste placeholder plus a remainder, and the following Enter was still reported as delivered. Prove the selected composer holds the payload before Enter, and report failure when it does not.

* no-mistakes(review): Scope Herdr payload proof to Claude, clear composer on refusal

* no-mistakes(test): Clear refused Herdr composer drafts one wrapped row per press

* no-mistakes(test): Accept Claude's multi-line paste placeholder in Herdr submit proof

* no-mistakes(review): Accept Claude read-back that drops U+2063 in Herdr proof

* no-mistakes(document): Document Herdr proof ignoring U+2063 operational mark

* no-mistakes(ci): I made a one-line test change. The failing check comes from a timing race in an existing test that this PR doesn't touch. **What failed:** `tests/fm-procevent.test.sh` failed at "the superseded paced runner invoked its stale command" (line ~3313). The PR only changes the Herdr files and their tests, and the same shard passed on main at the base commit. **Why it can fail:** the fixture starts a second runner with a 3-second launch floor (the minimum wait since the source's last launch). That runner sleeps for the rest of the floor and only then checks whether its registration was replaced (`fm_procevent_launch_floor_wait` in `bin/fm-procevent-lib.sh`). The test then waits for the claim and re-registers the source. If that takes longer than about 3 seconds after the first launch, the old runner wakes up, finds its registration still current, and runs the stale command. That produces the second log line the test reports. The CI shard was slow (this one test took 160 s). **Fix:** in `tests/fm-procevent.test.sh` I raised the superseded runner's floor from 3 to 15 seconds and added a comment explaining why. The floor now outlasts the fixture setup even on a loaded runner. Nothing else changed: the first launch and the later fresh-registration start still use a 3-second floor, and no product code changed. **Verification:** - The full test file can't give a reliable result on this machine (load average about 64 on 8 cores). It failed earlier, at the reconcile assertion around line 1680, before it reached this section. - I ran the changed section by itself (file setup plus the pacing-race block) five times with the fix: all passed, in about 9-13 s each. - The original code also passed five out of five, so the race didn't reproduce locally. The diagnosis rests on the code path and the CI log. - I haven't seen the full file or the CI shard pass with the fix yet

* no-mistakes(test): Accept Claude folder-trust prompt via down+enter in live e2e

* no-mistakes(document): Note unreadable Claude composer refusal in Herdr docs
…unchenguid#5427)

Speaking as Kun's firstmate: squash-merging — opt-in (forge=gerrit registry-gated; default project-mode stdout restored to two words), attestation MATCH, CI+NM green, safe review, MERGEABLE.
…uid#5358)

* feat(bin): add an opt-in per-home worker account pin

A home that mixes work and personal accounts for one runner had no way to
say which account its workers launch on: Claude workers inherited whatever
CLAUDE_CONFIG_DIR the supervising process had, Pi workers the pane's ambient
root, and an ambient API key outranked both, with no signal at launch.

config/claude-account and config/pi-account now pin that choice per home.
With neither file every launch is unchanged. With one, every launch of that
runner from the home (ship, scout, local secondmate, raw Claude command, and
relaunch) runs under the declared root, and the spawn refuses before any
endpoint exists when the file is malformed or the runner's own check
(claude auth status, pi auth check with a model-listing fallback) says the
pinned account is not signed in. The check runs in a cleared environment so
an ambient credential cannot answer for an empty root. A pinned Claude launch
sheds the environment credentials Claude ranks above a stored login; a pinned
Pi launch needs an explicit <provider>/<id> model for a declared provider and
also carries --provider. The chosen account is printed on the spawned line and
recorded in the task record, and relaunch checks the pin before stopping the
running agent.

* test(secondmate): give the concurrent config-push wait room for a slow host

test_config_reread_serializes_concurrent_pushes waited about two seconds for
the first fm-config-push.sh to reach its first send-keys. On a slower host
that push takes four to five seconds, so the test failed on main before the
push ever got there. The loop still leaves as soon as the marker appears, so
the larger bound costs nothing where the push is fast.

* no-mistakes(review): Refuse raw Claude account overrides under a pin
…ts (kunchenguid#5470)

* fix(bin): keep the Herdr lab session option before a -- delimiter

fm-herdr-lab.sh run appended --session <lab> after every argument, so a
command with a passthrough delimiter such as agent start ... -- <agent args>
handed the session flag to the agent and Herdr routed the call by the
caller's ambient socket instead of the lab.
The helper now inserts --session <lab> immediately before the first --
delimiter and keeps the trailing form otherwise.

* no-mistakes(document): Clarify Herdr lab session option placement
* feat(bin): guard the partition, harness pin, and bounded exec for a non-Pi supervision host

Lease liveness is now the pure record test in every calling context, so an
unmarked main honors a live branch lease held by a separate process, and a
lease file engages the guard's claim serialization for any caller; a home
with no lease files still takes no lock.

bin/fm-harness.sh honors FM_SUPERVISION_PRIMARY_HARNESS while
FM_SUPERVISION_ACTOR=branch, so a supervision branch running under another
harness resolves own, crew, and secondmate to the primary's harness.

fm_tasks_axi's watchdog moves into bin/fm-timeout-lib.sh as fm_exec_timed with
a separate grace: the perl watchdog is preferred, runs the command in its own
process group against wall-clock deadlines, forwards TERM/INT/HUP, and reaps
the group, so a descendant holding captured output can no longer keep the
caller waiting past the bound on a host without timeout.

The Claude Stop auto-arm header records that Claude drops the exit 2 of a hook
it terminated at the configured timeout, re-measured on Claude Code 2.1.281.

* fix(bin): state that fm_exec_timed cannot reach a descendant in its own process group

Live runs of real Claude and Pi engine turns under the bound showed both CLIs
start every tool command in a process group of its own, so those processes end
through the engine's own TERM handling rather than the group signal or reap.
Also clears the new timeout test's ShellCheck findings.

* no-mistakes(document): Clarify cross-harness lease documentation
…henguid#2648)

* feat(bin): make the ship-branch prefix configurable per project

fm-brief.sh hardcoded every generated ship branch to fm/<task-id>, which
leaks that firstmate produced the branch/PR - unwanted for a third-party
public repo that does not use this tooling.

Add an optional --branch-prefix flag to fm-brief.sh (default "fm/", so
existing installs are unaffected) and teach fm-project-mode.sh - the
registry's single-owner parser - to resolve a project's optional
"branch=<prefix>" data/projects.md annotation via a new --branch-prefix
query, order-independent with the existing mode/+yolo tokens. Firstmate
resolves the override at task intake and passes it explicitly, mirroring
how --mode already works; fm-brief.sh itself never reads the registry.

An empty override resolves to a bare "<task-id>" branch rather than a
leading slash. All five previously hardcoded fm/$ID sites (branch
creation, never-push rule text, definition-of-done text, and the status
message) now render the resolved prefix consistently.

* no-mistakes(review): Wire branch-prefix intake in AGENTS.md; fix fm-merge-local.sh hardcoded fm/ prefix

* no-mistakes(document): docs: document configurable ship-branch prefix in architecture.md

* no-mistakes(review): Persist immutable branch contracts

* no-mistakes(document): Document configurable ship branch prefixes

* no-mistakes(lint): Captain: fix ShellCheck test warnings

* fix(bin): map bearings PR rows to their recorded ship branch (kunchenguid#1887)

fm-bearings-snapshot.sh keyed a PR back to its task by string-matching
the headRefName against the fm/ prefix, so any project whose branch
prefix was overridden (e.g. via kunchenguid#2648's branch=<prefix> registry
annotation) had its PRs silently drop to task "-" in the bearings
view, exactly the third fm/-assumption issue kunchenguid#1887 named alongside
fm-merge-local.sh and fm-bearings-snapshot.sh itself.

fm-fleet-snapshot.sh now surfaces each task's recorded branch=
metadata field in its JSON task rows, and fm-bearings-snapshot.sh
cross-references a PR's headRefName against those recorded branches
before falling back to the legacy fm/ prefix heuristic, so a custom
branch prefix maps a PR back to its real task.

Adds a regression test proving a PR opened against a fix/<task-id>
branch resolves to that task instead of "-"; confirmed it fails on
the prior startswith("fm/") logic and passes with this change.
ShellCheck clean; full fm-bearings-snapshot.test.sh and
fm-fleet-snapshot-view.test.sh suites pass.

* fix(ci): align lint arithmetic-looking assignment and stale Bearings snapshot count

- Quote the --branch-prefix want_value assignment in fm-brief.sh, fm-promote.sh,
  and fm-spawn.sh so ShellCheck SC2100 no longer misreads the plain string
  'branch-prefix' as arithmetic shorthand.
- Bump the Stock macOS Bash snapshot job's hardcoded Bearings test-count
  assertion from 59 to 60: this PR added a Bearings test, so the count was
  stale, not the feature.

* no-mistakes(review): fix(bin): honor recorded ship branch in relaunch and review-diff

* no-mistakes(document): docs: complete branch-prefix flag in brief and promote headers

* fix(lint): quote branch-prefix parser token; drop unused BRANCH_Q after rebase

* no-mistakes(review): Restore %q branch escaping in promotion instructions with regression test

* no-mistakes(document): document recorded ship branch and prefix flag

fm-review-diff.sh's header is the owner of its branch-resolution
contract; it still described only the legacy local-branch behavior
after the change made review-diff honor state/<id>.meta's recorded
ship branch. README's feature bullet enumerates the registry's
optional flags and was missing the new branch=<prefix> override.

* no-mistakes(lint): Silence SC2016 on intentional single-quoted sed expression

* no-mistakes(review): address branch-prefix review findings in DoD and project-mode

* no-mistakes(test): branch-prefix suites pass under tasks-axi 0.2.6; environment-only failure

* no-mistakes(document): purge stale fm/ branch naming from docs and headers

* fix(test): assert the merged epoch status wording in the branch-prefix override test

The rebase resolution of tests/fm-brief.test.sh kept the branch's
pre-merge \`done: ready in branch ...\` assertion while the merged
fm-dod-lib.sh (carrying main's epoch-stamped status line) renders
\`done [at=<epoch>]: ready in branch ...\`. Align the assertion so the
override-consistency test matches the behavior it verifies.

* no-mistakes(review): Address remaining branch-prefix findings in four bin scripts

* no-mistakes(test): skip real-tasks-axi tests below the repo's 0.2.6 floor

* no-mistakes(document): document spawn's branch-prefix registry deviation notice
…d per-rule confidence floors (kunchenguid#5478)

* feat(bin): send dispatch resolver only the brief's task sections

* Sent Jev only the scaffolded Captain's intent and Firstmate spec
  sections, falling back to the whole brief when neither heading is
  present, so the identical setup, rules, and definition-of-done
  boilerplate no longer reads as a signal about the task
* Added an optional per-rule min_confidence that replaces the global 0.6
  floor for that rule; a picked rule below its own floor falls to the most
  probable other option that clears its floor, or returns ambiguous
* Kept files with no declared floor on the exact previous behavior and
  kept the model blind to the new field
* Recorded the live old-versus-new comparison over scaffolded fixtures

* no-mistakes(review): share brief heading parser, add kind line, fix floors

* no-mistakes(test): stop sending ship delivery mode to jev, keep scout tag

* no-mistakes(document): docs: list shared brief heading lib in scripts inventory
* feat(bin): supervision host core behind config/supervision-host

Add the supervision host (bin/fm-supervision-host.sh): beside a Claude
primary it owns the watcher cycle for the Stop auto-arm and, while the
away-posture record exists, hands each wake to a bounded headless Claude
engine session that runs the supervision branch's contract - the same
generated prompt, row eligibility, wake grant, per-actor drain, outcome
store, leases, and away relocation the Pi branch uses. Attended wakes pass
straight to main. Every path that cannot finish a wake hands it to main
with a supervision-host line; the park ends itself before the Stop hook
timeout with a cycle-boundary wake.

- bin/fm-supervision-engine-lib.sh: opt-in parse, verified engines
  (claude, default sonnet), one bounded engine turn, and a reap of engine
  tool processes that sit in their own process groups.
- bin/fm-branch-report.sh: the command twin of fm_branch_report, scoped
  to the tasks the current host turn claimed.
- bin/fm-branch-dispatch.mjs: command entry to the Pi dispatch module, so
  eligibility and the wake prompt have one owner.
- bin/fm-claude-stop-autoarm.sh runs the host in the arm's place when
  config/supervision-host exists; nothing changes without the file.
- bin/fm-watch-arm.sh --stop: home-scoped stop without a re-arm.
- bin/fm-lease-lib.sh: an opted-in home takes the lease-command lock for
  unmarked main too, closing the first-claim race; the refusal tells the
  caller to leave the lease alone and retry.
- /afk launches no away daemon on an opted-in Claude home; /quiet still
  does. Session start renders the host's main-side protocol there.

* fix(bin): relay a host turn's outcomes when the captain returns mid-turn, and log per-turn engine cost

Live validation found two supervision host gaps. A captain who returns while
an engine turn is running gets a return brief rendered before that turn's
outcomes exist, so the host now hands the close to main with those outcomes.
Claude reports a resumed conversation's running cost, so the engine lib now
derives each turn's cost from the total the host records, and the host log
records every close's destination.

* docs(verification): record the supervision host's live evidence

The dated live results behind docs/supervision-host.md: the Claude engine's
live guard, the away-wake cases against real workers, the engine's cost
reporting, and the flag-off before-and-after regression.

* docs: describe the supervision host ledger as covering every close

* no-mistakes(review): Harden supervision host ownership, boundary, ack, and late outcomes

* no-mistakes(review): Recheck park boundary just before starting an engine turn

* no-mistakes(review): Cap park boundary, deliver all host lines, reject incomplete results

* no-mistakes(document): Correct supervision host documentation and stale pointers
…t-in, and delivery (kunchenguid#5506)

Attestation MATCH; contract-class restore; CI/NM green. Squash-merged by Kun's firstmate.
…kunchenguid#5528)

* fix(bin): bound the startup-network worker's lock waits by its budget

Fixes kunchenguid#5377

The deferred startup network worker bounded its sweeps with a stage budget but
took the publish lock and the fleet-lock lease with an unbounded wait, so a live
holder of that lock kept the detached worker alive for hours past its timeout
with its output discarded at the end. Every wait now goes through the bounded
acquire and shares the remaining stage or delivery budget; a lock a live process
still holds at the deadline ends the worker with a failed record naming the
holder and the rerun command, and a wake so the result surfaces.

* no-mistakes(review): propagate publish exit code from cmd_run terminal paths
…dth (kunchenguid#5517)

* fix(bin): keep the ps fallback identity independent of terminal width

fm_pid_identity's portable fallback read the command column at the
ambient COLUMNS width, so an identity recorded from a wide shell never
matched the one recomputed inside a narrow hook and the continuity guard
denied every fleet command. Pass -ww so the column is never cut.

Fixes kunchenguid#799

* no-mistakes(ci): Fixed CI failure in Behavior portable serial 4. Root cause: the -ww flag added in commit ac7ab5d to fm_pid_identity (bin/fm-wake-lib.sh) shifted the ps argv so $1 became -ww instead of -p, breaking the positional fake-ps fixtures in tests/fm-procevent.test.sh (lines 2736, 3571) which then fell through to real ps and failed the fm-procevent test. Fix (already applied in the worktree, matching the authoritative user instruction exactly): replaced -ww with a COLUMNS=10000 environment pin so the call is COLUMNS=10000 LC_ALL=C ps -p "$pid" -o lstart= -o command=, mirroring fm_pending_reply_pid_identity in bin/fm-pending-reply-lib.sh:982. argv is back to -p PID -o lstart= -o command=, so the fixtures match again with no fixture edits. Comments above the call in bin/fm-wake-lib.sh and in test_pid_identity_is_terminal_width_invariant (tests/fm-watcher-lock.test.sh) now describe the COLUMNS pin instead of -ww; the regression test still asserts narrow-vs-wide byte equality and the full command. Verified: the terminal-width-invariant regression test passes. The only local not-ok results were flaky, run-varying timing tests (procevent launch/claim confirmation, listener reparenting) that differ each run and are unrelated to the ps argv change
…thorized intent (kunchenguid#5526)

Fixes kunchenguid#3608

When a scout is promoted to a ship, the captain's authorized intent is
extracted from a legacy `# Task` body by matching `Captain:` and `[captain]`
lines anywhere in the body, including inside fenced code blocks and indented
examples, while the heading reader already tracks fences. A fenced `Captain:`
example therefore passed the provenance gate and became the ship contract's
intent while the real ask was dropped.

Make the captain-words extractor fence-aware like the heading reader: a line
inside a ``` or ~~~ fenced block, or indented four spaces or a tab, is never a
marked line. The promotion and spawn callers need no change. The regression
test covers both the extractor and the promotion provenance gate refusing a
brief whose only Captain lines are fenced or indented examples.
…riefs (kunchenguid#2868)

* fix(bin): forbid administering the shared worktree pool in crewmate briefs

A crewmate ran a `git worktree remove` loop over the treehouse pool its own
worktree came from, destroying five worktrees - four belonging to tasks that
were running mid-pipeline. The generated brief's rule 2, "stay inside this
worktree; modify nothing outside it", is a rule about files: removing a
worktree is administration of shared state, not an edit outside a directory,
so the sentence never reached the act. The worker satisfied its brief
completely.

Rule 7 already named one piece of shared infrastructure - the no-mistakes
daemon, one instance serving every lane - with the reason stated plainly. The
worktree pool is the same class of thing and was unnamed.

Fold the pool into that existing rule rather than adding a second warning:
state the constraint around the act (create, remove, return, prune, move,
reassign a worktree or pool slot; write into a sibling slot), keep concrete
commands as examples rather than as the definition so no single provider is
pinned, and give the prohibition a real exit through `blocked:`.

The rule is emitted from one shared string interpolated into both crewmate
scaffolds, so the ship and scout copies cannot drift apart. The secondmate
charter deliberately omits it: that home runs its own fleet and legitimately
allocates and returns slots for its own crewmates.

Contract text only; no runtime enforcement layer.

* no-mistakes(document): Distill pool-safety comment rationale

* no-mistakes(review): align pool-rule test grep patterns with emitted [at=<epoch>] text
… a dispatch record (kunchenguid#5524)

* fix(bin): refuse tasks-axi add --start so In flight always has a dispatch record

Fixes kunchenguid#4753

Dispatch (bin/fm-spawn.sh) is the only path that moves a backlog row to
In flight, because it creates the task record, status file, and inbox
that go with the row. A row hand-placed there through the wrapper's
`add --start` had none of those, and nothing later noticed, so the
live-task count included work nobody was doing. The wrapper now refuses
`add --start` (exit 2) and names the dispatch path; plain `add` and
`start <id>` pass through unchanged, and the lifecycle transitions
address tasks-axi directly so dispatch is unaffected.

The issue's other half, a reconcile sweep in bin/fm-inactive-reconcile.sh
that notices an In flight row with no task record, is left as is; this
change closes the only path that creates such a row.

* no-mistakes(review): refuse create --start alias, not just add --start

* no-mistakes(review): reword add --start guard docs to drop only-path overclaim

* no-mistakes(review): scope add/create --start guard docs, drop universal claim
…#5503)

* feat(bin): run the supervision host beside the other non-Pi primaries while away

Cursor's stop-hook park, the OpenCode plugin, the omp watch extension, Grok's
model-owned background arm, and Codex's foreground checkpoint now run
bin/fm-supervision-host.sh in the watcher arm's place when the home opted in
with config/supervision-host, so the host's Claude engine takes away-posture
wakes beside those primaries exactly as it does beside Claude. Without the
file nothing changes.

- The host streams its first cycle's status line, accepts --restart and the
  owner's predecessor arm for its first cycle, and prints each exit in one
  write, so owners that wait for arm readiness and restart their own
  successor (OpenCode, omp) keep their handling handoff.
- Codex's checkpoint passes its bound to the host as the park boundary,
  raises it to FM_CODEX_WATCH_CHECKPOINT_AWAY (3600 s) while the away record
  exists, and lets an engine turn that starts before the bound finish after
  it (FM_SUPERVISION_HOST_PARK_LIMIT).
- /afk launches no away daemon on an opted-in home of those harnesses and
  says so at entry when the file selects no engine for that primary.
- Session start renders the host protocol for each arm owner, and Grok's
  arm command becomes the host.

* fix(bin): keep the watcher-down banner away from the supervision branch actor

A supervision host's engine turn runs guarded commands after its successor
watcher cycle may already have closed on a newer wake, so the guard showed it
the watcher-down banner with the primary's repair line. Under a Codex primary
pin that line is the checkpoint, and a live Codex lab run showed the away
session running it mid-turn (the nested host stood down on its ownership
check). The branch actor never owns watcher continuity, so the banner, its
reminder, and the episode state now leave that actor out, as the queued-wake
warning already does.

The lint telemetry fixture counts bin/fm-afk-launch.sh's source directives,
which the host engine note raised from four to five.

* fix(bin): queue away-session outcomes recorded after the return for main

A Cursor park superseded by the captain's return stops its host as the
engine turn ends, so the host's own handoff of that turn's outcomes was
never printed and the outcomes never reached main. The report surface now
queues every outcome it records after the away record is gone as a durable
check wake; the return owner archives the record before it reads the store,
so each outcome is in the return brief, queued, or both. A host stopped
mid-turn also removes its turn's result and error files.

The stream test now acknowledges its first close and accepts a restarted
cycle that closes on its resurface before the arm confirms it.

* fix(bin): clear a hard-killed host's turn at the next activation

A Cursor park superseded mid-turn can kill its host outright, which runs no
cleanup, so the turn's result, error, and descendant files stayed behind and
any tool process the engine started was left running. The next host's
activation now reaps the descendants that turn recorded and removes its
files. The host suite also registers its homes in a file, because
make_home runs in a command substitution, so its cleanup now stops every
host a case leaves running.

* fix(bin): leave rows that arrive after main's drain unclaimed at its acknowledgement

Main's acknowledgement re-claimed every unreserved queued row, including one
that arrived after the drain above the acknowledged cutoff. That row stayed
main's without ever being shown to it, so while away the supervision host
refused every later wake that included it and handed each back to main until
main drained again. The acknowledgement now claims only unreserved rows at or
below its cutoff.

* docs: name the killed turn's engine and files in the host's failure direction

* docs: record live supervision host runs on the non-Pi primaries

* no-mistakes(review): Replay host-only supervision boundaries across omp session replacement

* no-mistakes(review): Deliver omp supervision-host wakes only at the host's close

* no-mistakes(document): Correct supervision host documentation for non-Pi primaries

* no-mistakes(ci): Fixed the CI failure by naming FM_CODEX_WATCH_CHECKPOINT_AWAY in the rendered Codex host instructions. The focused instruction and checkpoint suites pass
)

* feat(bin): auto-relaunch dead persistent secondmates during ordinary supervision

A persistent secondmate whose primary agent exits mid-session previously
stayed down until the next session-start liveness sweep. Extract the
sweep's probe/classify/relaunch mechanics into a shared library and drive
the same contract from a cadence-gated watcher tick, so a positively dead
or missing endpoint is relaunched through the guarded spawn path within a
poll cycle instead of an hour later.

Only the recovery-grade `dead` and `missing` verdicts authorize relaunch;
ambiguous, unreadable, unverified, and unreachable-remote reads stay
fail-closed and a remote route is never replaced by a local endpoint.
Each relaunch emits exactly one `check` wake and appends to a durable
per-mate ledger; a mate exceeding the bounded attempt budget is parked
behind a marker until a live probe rearms it. A per-mate liveness lock
serializes the tick against a concurrent session-start sweep.

* no-mistakes(review): Fail closed on relaunch ledger errors; clear state on remote teardown

* no-mistakes(review): Share ledger read guard; retire relaunch state under liveness lock

* no-mistakes(review): Lazy-load wake lib; live rearm restores full relaunch budget

* no-mistakes(review): Finish liveness tick for every mate before waking once

* no-mistakes(review): Keep liveness tick scanning past per-mate errors, then wake

* no-mistakes(review): Wake only on queued rows; teardown holds liveness lock

* no-mistakes(review): Queue liveness outcome wake before releasing mate lock

* no-mistakes(document): Update secondmate liveness documentation for mid-session recovery

* no-mistakes(lint): Fix empty assignments flagged by ShellCheck

* no-mistakes(ci): Added ShellCheck analysis boundaries for the shared liveness library in both callers and marked its result globals as intentional library outputs. Changed-file lint passed; full CI partitions were not run locally

* no-mistakes(ci): Fixed Lint 2 by removing an unused test variable in tests/fm-wake-queue.test.sh. ShellCheck, bash syntax, and the full wake-queue test script pass

* no-mistakes(ci): Fixed the CI wake-queue fixture: stall-only watcher legs now seed the liveness cadence marker, preventing the new endpoint probe from interfering with their assertions. The full wake-queue test, ShellCheck, and diff checks pass locally
…cycle ends (kunchenguid#5550)

* fix(bin): start a successor when the Claude Stop-hook arm's attached cycle ends

Fixes kunchenguid#2381

When the Claude Stop hook's foreground arm attached to a peer watcher cycle
and that cycle ended, the arm reported the delivered wake and the hook exited
2 without starting a successor, so the handling turn ran with no watcher.
Pi, omp, and OpenCode start the next arm before delivering the wake and pass
the closed arm's pid as FM_WATCH_PREDECESSOR_ARM_PID; the Claude hook never
passed that predecessor identity.

The hook now runs its arm as a tracked child it waits on, so it holds that
arm's pid, and after any actionable close starts one handling-successor
bin/fm-watch-arm.sh with the closed arm's pid as FM_WATCH_PREDECESSOR_ARM_PID.
The successor is launched the one way a process outlives a Claude hook's
exit-2 rewake (nohup, detached stdio, own process group, the shape
bin/fm-startup-network.sh already uses); the hook waits for its status line
and adds one banner line when no live watcher was confirmed, never withholding
the wake. The supervision-host path is unchanged, as is the arm wrapper.

The regression test drives the real hook against an arm fixture whose attached
peer cycle ends: it fails on the previous tip because no successor starts, and
now asserts the successor names the closed arm as its predecessor and outlives
the rewake. A second case pins the unconfirmed-successor banner line.
docs/watcher-continuity.md no longer records the Claude asymmetry.

* no-mistakes(ci): Serial-4 failure was a real regression: tests/fm-session-lock-ancestry.test.sh asserts exact cumulative arm-invocation counts while driving the real fm-claude-stop-autoarm.sh hook against a stubbed fm-watch-arm.sh. This PR makes the hook start a handling successor after an actionable close, so every owned actionable phase now records TWO arm invocations (foreground arm + successor) instead of one, breaking "healthy chain: expected 1 arm(s), got 2". Fixed by updating the cumulative expectations to match the new behavior: owned phases 1/2/6 -> 2/4/6, foreign carry phases 3/4/5 -> 4, plus a comment explaining the +2-per-owned-phase model. Verified: phase-1 (the CI failure point) now passes on every run, syntax checks clean, and sibling arm-count tests (fm-claude-stop-autoarm.test.sh, fm-cursor-primary, fm-turnend-guard) pass unchanged. The only remaining local not-ok is a WSL-only environmental artifact (orphan reparents to a subreaper, not PID 1) that passes on the CI runner. Parallel-1 failure is an unrelated flake: its 11 tests (fm-lint, fm-pr-merge, fm-test-run, fm-cd-pretool-check, fm-pi-primary-types, fm-grok-harness, fm-composer-lib, fm-review-diff, fm-tmux-submit-busy, fm-composer-ghost, fm-brief) do not include fm-session-lock-ancestry and none reads any file this PR touches; all pass locally. It should clear on CI re-run. Made the smallest root-cause fix (one test file, 6 count updates + a clarifying comment). Validation of the branch continues through the no-mistakes pipeline, which owns re-running CI
…ning (kunchenguid#5566)

* fix(bin): report a Lavish source armed only after its listener is running

Registration alone was treated as ready, so arm could succeed before anything was collecting from the board.

* no-mistakes(review): Guard Lavish arm launches, keep retire refusals, report live prior listener

* no-mistakes(review): Keep polling through window before reporting a still-live prior listener

* test: wait for a capture's claim to drop before the next arm

The result is stored before the runner exits, so a re-arm in that gap was meeting a live claim.

* no-mistakes(document): Record Lavish arm readiness evidence in verification doc

* no-mistakes(ci): Both failures were caused by this PR, and both are fixed with test-only edits. Lint 2 (ShellCheck SC2034): this branch removed the only use of `reply_id` (a `start "$reply_id"` call) from tests/fm-procevent.test.sh, which left the assignment at line 1450 unused. I deleted that assignment. It was the only `reply_id` in the file. ShellCheck is now clean on both test files. Behavior portable serial 4: the failing test was tests/fm-bearings-board.test.sh, in the check "registration consumed its answer before the any-origin binding existed". I reproduced it locally: the hold was still `state: queued` when the test checked it. - What must hold: the test's check that the hold is closed must run after the listener has captured the answer. - Why it broke: the test used a stand-in adapter that ran `fm-procevent.sh start` in the foreground after `arm`, so capture finished before build returned. On this branch, `arm` starts the listener itself in the background, so the real listener captures the answer and closes the hold a moment after build returns. - Fix: removed the now-redundant stand-in adapter, the copied runtime directory, and its extra environment variables. The test now runs the real build through the existing `run_board` helper and waits up to about 10s for the hold to reach `state: done`. The checks that follow are unchanged: `Resolution mode: answered` and the any-origin binding. - Other tests: this was the only test in the file that stood in for the adapter this way. The shard's other pure-contract-unit test (tests/fm-trace-context-lib.test.sh) passed unchanged. Verification: - tests/fm-bearings-board.test.sh passed 3 times in a row via bin/fm-test-run.sh, all 18 checks, about 53s per run. - tests/fm-procevent.test.sh was not rerun, because the lint fix only removed an unused assignment
…ts resolved

Reconciles upstream/main into the fork, keeping every local customization intact.
Resolution method is documented in the PR body.
Merge origin/main into the 251-commit upstream reconciliation.
Keeps the landed #18 content where they overlap: the treehouse
--no-fetch spawn entry, the v2.3.0 pin, the tangle-guard expectation,
the CONTRIBUTING direct-PR wording, and the no-mistakes-required
workflow removal (upstream's later edits to that workflow are dropped
with the file, per the fork's direct-PR posture).
Upstream-synced content is otherwise untouched.
@jorguez96
jorguez96 merged commit bcf80bf into main Sep 25, 2026
18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.