Skip to content

feat: sync upstream supervision and harness support - #32

Merged
BohnBawerick merged 240 commits into
mainfrom
fm/fm-upstream-sync-1003
Oct 4, 2026
Merged

BohnBawerick merged 240 commits into
mainfrom
fm/fm-upstream-sync-1003

Conversation

@BohnBawerick

@BohnBawerick BohnBawerick commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

Intent

[captain] so could we run an update on all the axis stack and on firstmate and so on? /updatefirstmate and /stow
[captain] I think we were going thorught our fork update for the firstmate when the connection stoped.. so could you look into that one. it was running a no0mistakes thing

Context for the captain's words (Firstmate-authored, 3 Oct 2026): this repo is a fork (origin BohnBawerick/firstmate) of upstream kunchenguid/firstmate. Origin had nothing new, so the firstmate update means bringing in upstream. upstream/main (newest 1f3e769) had 203 commits not in our main; our main had 189 not upstream; the recorded merge base was still 888871d. The last upstream sync, PR #29, brought in upstream through c576c2b but landed as a squash, so git did not see c576c2b as merged. Recording c576c2b as already merged first (a merge commit with strategy ours, since its content is already in our main) drops a trial merge from 118 conflicted files to 56. Upstream has 156 commits after c576c2b. The connection stopped because the machine restarted while the update was mid-task, before its no-mistakes run had started; this run resumes that branch.

The captain's rulings from that earlier sync still stand:

  • lock-identity-contract, option 2: adopt upstream's session-lock identity contract (trusted session-id gate, live recorded pid, upstream's bin/fm-session-lock-lib.sh and bin/fm-lock.sh flow) and keep the fork's fleet-mutation gate fm_require_session_lock on top of it, with the fork's lock wording.
  • crew-state-ledger-anchor, option 2: adopt upstream's ledger-anchored continuation rule in bin/fm-nm-run-lib.sh and bin/fm-crew-state.sh (an active run at an unresolvable head is bound when the immediately older ledger row for the same branch sits at the local HEAD), keeping the fork's ternary identity for everything the anchor does not cover.

Firstmate specification (not the captain's literal words):

  1. On a branch cut from our main, first record c576c2b as merged with git merge -s ours c576c2bb (its content landed by squash in PR 29, verified by a tree comparison), then merge upstream/main (1f3e769) with a real merge commit, not a rebase. Resolve every conflict.
  2. Keep both sides' behavior wherever they are compatible. Fork additions that must survive: curated memory (bin/fm-memory-.sh), the quality gate and hardened posture (bin/fm-quality.sh), local landing (bin/fm-merge-local.sh), the agy adapter, Herdr/Pi supervision continuity and session-lock gates, the report-once turn-end decline for non-Claude harnesses, catalog-aware Codex max effort (bin/fm-codex-catalog-lib.sh and its callers), Lavish list-form feedback parsing and truncation guards in bin/fm-procevent-lavish.sh, remote secondmates, and the fork's tests and CI additions. Adopt upstream's fixes and features. Where an earlier ruling already settled a cluster, apply it again.
  3. A conflict that is a genuine design divergence or a user-facing behavior trade-off where both cannot be kept is not guessed and neither side is silently dropped: it goes to the captain for a ruling. Mechanical conflicts are resolved directly.
  4. Lint (bin/fm-lint.sh) and the behavior tests covering the conflicted files must pass, apart from failures that reproduce identically on unmodified upstream on the same host.
  5. Open PRs feat: add catalog-aware Codex max effort support #20, feat: integrate upstream supervision and fleet lifecycle updates #26, fix(bin): preserve Lavish list-form feedback #30 and fix: harden supervision feedback and worker launches #31 on BohnBawerick/firstmate stay open by design (their work already landed on our local main). Do not close, edit or merge them. The PR body says whether this merge supersedes feat: integrate upstream supervision and fleet lifecycle updates #26: it supersedes that PR's purpose (its upstream content is already on main and this branch brings upstream to 1f3e769) but does not carry that PR's own review and CI fix commits.
  6. Ship as a PR against BohnBawerick/firstmate main. Never push to upstream. Never merge.

Out of scope: new features, refactors beyond conflict resolution, upstreaming the fork's patches.

What Changed

  • Merge upstream Firstmate through 1f3e7696, adding the supervision host, expanded AFK and quiet handling, automatic secondmate recovery, and watcher lifecycle fixes across supported harnesses.
  • Add Gerrit and Devin support, configurable ship branch prefixes, an optional fleet activity ledger, live supervision labs, catalog-validated Codex effort selection, and structured Lavish feedback parsing.
  • Supersede feat: integrate upstream supervision and fleet lifecycle updates #26's purpose because its upstream content is already on main and this branch advances the fork through 1f3e7696; this branch does not include feat: integrate upstream supervision and fleet lifecycle updates #26's own review and CI fix commits.

Bundled work, earlier rulings, and follow-up

  • Commits from local main. The branch was cut from the fork's local main, which at the time held 24 commits not yet on origin/main. Those commits carried the work of feat: add catalog-aware Codex max effort support #20 (catalog-aware Codex max effort), fix(bin): preserve Lavish list-form feedback #30 (Lavish list-form feedback), and fix: harden supervision feedback and worker launches #31 (supervision feedback and worker launch hardening). origin/main has since been advanced to that same commit (1726136), so those three PRs now show as merged and this PR no longer adds those commits on top of its base.
  • feat: integrate upstream supervision and fleet lifecycle updates #26. This merge supersedes the purpose of feat: integrate upstream supervision and fleet lifecycle updates #26: its upstream content is already on main, and this branch brings upstream to 1f3e7696. It does not carry feat: integrate upstream supervision and fleet lifecycle updates #26's own review and CI fix commits. feat: integrate upstream supervision and fleet lifecycle updates #26 is left open and untouched.
  • Earlier rulings applied again. The review flagged two divergences as unresolved; both were already settled and the merge applies them:
    • No-mistakes handoff (dod-handoff-contradiction, option a): every no-mistakes ship done: is gated, the worker starts its own validation, and there is no pre-validation done:.
    • Closed contributions (fork commit 48b9e5e, kept in the last sync): only a merged contribution is final; a closed one is still re-read so a reopen is seen.
  • Follow-up, not fixed here. The review found that a raw launch command starting with a wrapper (env, exec, bash -c) skips the worker-account pin checks. bin/fm-worker-account-lib.sh and the raw-launch parse in bin/fm-spawn.sh are identical to upstream 1f3e7696, so the merge did not introduce it. It is tracked as separate follow-up work.
  • CI fixes made by the pipeline. Test-only changes, plus one product fix: an already-handled remote-reply replay no longer re-arms the source registration, which could retire the live listener (bin/fm-procevent-remote-reply.sh). tests/fm-supervision-host.test.sh was split in two (tests/fm-supervision-host-lifecycle.test.sh, shared tests/fm-supervision-host-fixture.sh) so each half fits the 30-minute job limit.

Risk Assessment

🚨 High: The change contains two unresolved semantic conflicts that the stated intent reserves for the captain, plus a reachable account-pin bypass in supported raw launches.

Testing

Inspected the target history and Herdr runbook, drove the captain's lock and ledger rulings through real Firstmate commands in disposable environments, ran the focused post-merge reconciliation tests, and reproduced then fixed the away-daemon deadlock in a guarded non-default Herdr lab. Every command passed, no lab remained running, and no worktree files changed.

  • Live validation: ✅ go - 3 of 3 scenarios driven live against the product
Scenario Result Live Evidence
A trusted session keeps its lock across helper-process ancestry changes, while another or untrusted session is refused ✅ pass live session-lock-ownership.log and session-lock-ancestry.log show real Firstmate lock entry points preserving the trusted session across a recycled process chain, refusing weaker identities, and recla…
An active validation run at an unfetched pipeline head remains bound only when the immediately older same-branch ledger row anchors it to local HEAD ✅ pass live crew-state.log shows the real crew-state command binding an unresolvable active row through the immediately older same-branch anchor, while mismatched anchors remain unknown and terminal rows retain…
Away supervision delivers to an idle Claude pane through a separate daemon terminal while a fleet worker remains busy, and rejects unsafe target selection ✅ pass live herdr-afk-self-deadlock-e2e.log records the old self-target deadlock, unsafe-target refusals, failed-run cleanup, separate daemon-terminal launch, and successful digest delivery while the fleet work…
Evidence: Session-lock ownership

Source: Session-lock ownership

ok - every mutating fleet entry point refuses a session that does not hold the home
ok - the fleet-mutation gate answers before argument validation
ok - the lock-holding session passes the fleet-mutation gate untouched
ok - a caller outside any harness session is not treated as a competing session
ok - the lock path names ownership in words, never as a bare pid
ok - a background continuation of the lock-holding session inherits the helm, and only it does
ok - an ancestry grant is reported in words and grants nothing outside that tree
ok - a dead recorded pid is reclaimed onto the continuation's own live pid
ok - the Stop auto-arm reclaims a dead recorded pid onto the session's live pid
ok - a spawned worker does not inherit the spawning session's helm
ok - the auto-arm declines silently and records no failure for a home it does not hold
ok - the turn-end guard reports a not-this-session decline once, then stops blocking
ok - two non-owning sessions in one home each report once, then both stand down
ok - a harness that sends no session id still reports the decline once per session
ok - notice records are retired only when their recorded identity is gone
ok - the decline is reported again when a different session takes the home
Evidence: Session-lock ancestry

Source: Session-lock ancestry

ok - session-lock: a version-named Claude Code session is identified from its install path and argv[0]
ok - session-lock: a harness that is pid 1 of its own namespace is examined, not skipped
ok - session-lock: ordinary script paths under a harness directory are not harness processes
ok - session-lock: ownership stops at the first non-harness gap above the contiguous run
ok - session-lock: a live version-named session holding the lock is not mistaken for a stale owner
ok - session-lock: a trusted same-session id keeps owning a recycled background chain, and nothing weaker does
ok - session-lock: a trusted id anchors the lock on the model-loop process, anything else on the outermost pid
ok - session-lock e2e: a version-named session claims the home and arms supervision
ok - session-lock e2e: a session parented by a harness-named daemon claims the home and arms supervision
ok - session-lock e2e: a version-named session under a harness-named daemon keeps its own lock
ok - session-lock e2e: a background session keeps its lock and its supervision across a recycled helper chain
ok - session-lock: a same-session confirmation waits for the claim lock and refreshes a re-keyed id
ok - session-lock: a waiting confirmation does not steal another session's lock
ok - session-lock: a failed lock write restores the previous sidecar
ok - session-lock: a failed lock write removes a newly created sidecar
ok - session-lock: a verified reclaim keeps the new sidecar beside the new pid
Evidence: Crew-state ledger anchoring

Source: Crew-state ledger anchoring

ok - captured AXI replacement status replays through crew-state
ok - captured AXI parked status replays through crew-state
ok - captured AXI failed status replays through crew-state
ok - captured capped inventory replays selection, ambiguity, and unavailable lookup
ok - captured status formats reject a synthetic authority transition
ok - captured completed status yields to synthetic subsequent development
ok - captured cancelled run leaves fleet inventory unverified without a failure contradiction
ok - failed-outcome/github/open: terminal delivery uses current disposition
ok - failed-outcome/github/merged: terminal delivery uses current disposition
ok - failed-outcome/github/closed: terminal delivery uses current disposition
ok - failed-outcome/github/unreadable: terminal delivery uses current disposition
ok - failed-outcome/github/skipped: terminal delivery uses current disposition
ok - failed-outcome/github/no-identity: terminal delivery uses current disposition
ok - failed-outcome/gitlab/open: terminal delivery uses current disposition
ok - failed-outcome/gitlab/merged: terminal delivery uses current disposition
ok - failed-outcome/gitlab/closed: terminal delivery uses current disposition
ok - failed-outcome/gitlab/unreadable: terminal delivery uses current disposition
ok - failed-outcome/gitlab/skipped: terminal delivery uses current disposition
ok - failed-outcome/gitlab/no-identity: terminal delivery uses current disposition
ok - failed-outcome/gerrit/open: terminal delivery uses current disposition
ok - failed-outcome/gerrit/merged: terminal delivery uses current disposition
ok - failed-outcome/gerrit/closed: terminal delivery uses current disposition
ok - failed-outcome/gerrit/unreadable: terminal delivery uses current disposition
ok - failed-outcome/gerrit/skipped: terminal delivery uses current disposition
ok - failed-outcome/gerrit/no-identity: terminal delivery uses current disposition
ok - failed-status/github/open: terminal delivery uses current disposition
ok - failed-status/github/merged: terminal delivery uses current disposition
ok - failed-status/github/closed: terminal delivery uses current disposition
ok - failed-status/github/unreadable: terminal delivery uses current disposition
ok - failed-status/github/skipped: terminal delivery uses current disposition
ok - failed-status/github/no-identity: terminal delivery uses current disposition
ok - failed-status/gitlab/open: terminal delivery uses current disposition
ok - failed-status/gitlab/merged: terminal delivery uses current disposition
ok - failed-status/gitlab/closed: terminal delivery uses current disposition
ok - failed-status/gitlab/unreadable: terminal delivery uses current disposition
ok - failed-status/gitlab/skipped: terminal delivery uses current disposition
ok - failed-status/gitlab/no-identity: terminal delivery uses current disposition
ok - failed-status/gerrit/open: terminal delivery uses current disposition
ok - failed-status/gerrit/merged: terminal delivery uses current disposition
ok - failed-status/gerrit/closed: terminal delivery uses current disposition
ok - failed-status/gerrit/unreadable: terminal delivery uses current disposition
ok - failed-status/gerrit/skipped: terminal delivery uses current disposition
ok - failed-status/gerrit/no-identity: terminal delivery uses current disposition
ok - cancelled-outcome/github/open: terminal delivery uses current disposition
ok - cancelled-outcome/github/merged: terminal delivery uses current disposition
ok - cancelled-outcome/github/closed: terminal delivery uses current disposition
ok - cancelled-outcome/github/unreadable: terminal delivery uses current disposition
ok - cancelled-outcome/github/skipped: terminal delivery uses current disposition
ok - cancelled-outcome/github/no-identity: terminal delivery uses current disposition
ok - cancelled-outcome/gitlab/open: terminal delivery uses current disposition
ok - cancelled-outcome/gitlab/merged: terminal delivery uses current disposition
ok - cancelled-outcome/gitlab/closed: terminal delivery uses current disposition
ok - cancelled-outcome/gitlab/unreadable: terminal delivery uses current disposition
ok - cancelled-outcome/gitlab/skipped: terminal delivery uses current disposition
ok - cancelled-outcome/gitlab/no-identity: terminal delivery uses current disposition
ok - cancelled-outcome/gerrit/open: terminal delivery uses current disposition
ok - cancelled-outcome/gerrit/merged: terminal delivery uses current disposition
ok - cancelled-outcome/gerrit/closed: terminal delivery uses current disposition
ok - cancelled-outcome/gerrit/unreadable: terminal delivery uses current disposition
ok - cancelled-outcome/gerrit/skipped: terminal delivery uses current disposition
ok - cancelled-outcome/gerrit/no-identity: terminal delivery uses current disposition
ok - cancelled-status/github/open: terminal delivery uses current disposition
ok - cancelled-status/github/merged: terminal delivery uses current disposition
ok - cancelled-status/github/closed: terminal delivery uses current disposition
ok - cancelled-status/github/unreadable: terminal delivery uses current disposition
ok - cancelled-status/github/skipped: terminal delivery uses current disposition
ok - cancelled-status/github/no-identity: terminal delivery uses current disposition
ok - cancelled-status/gitlab/open: terminal delivery uses current disposition
ok - cancelled-status/gitlab/merged: terminal delivery uses current disposition
ok - cancelled-status/gitlab/closed: terminal delivery uses current disposition
ok - cancelled-status/gitlab/unreadable: terminal delivery uses current disposition
ok - cancelled-status/gitlab/skipped: terminal delivery uses current disposition
ok - cancelled-status/gitlab/no-identity: terminal delivery uses current disposition
ok - cancelled-status/gerrit/open: terminal delivery uses current disposition
ok - cancelled-status/gerrit/merged: terminal delivery uses current disposition
ok - cancelled-status/gerrit/closed: terminal delivery uses current disposition
ok - cancelled-status/gerrit/unreadable: terminal delivery uses current disposition
ok - cancelled-status/gerrit/skipped: terminal delivery uses current disposition
ok - cancelled-status/gerrit/no-identity: terminal delivery uses current disposition
ok - outcome: cancellation without delivery carries no verdict
ok - status: cancellation without delivery carries no verdict
ok - selected: cancellation without delivery carries no verdict
ok - coarse: cancellation without delivery carries no verdict
ok - no-ci-log: cancellation without delivery carries no verdict
ok - red-ci: cancellation without delivery carries no verdict
ok - cancelled-test: cancellation without delivery carries no verdict
ok - skipped-test: cancellation without delivery carries no verdict
ok - synthetic cancelled run leaves fleet inventory unverified without a failure contradiction
ok - cancelled-outcome: terminal delivery reports only observed evidence
ok - cancelled-status: terminal delivery reports only observed evidence
ok - skipped-rebase: terminal delivery reports only observed evidence
ok - cancelled-skipped-rebase: terminal delivery reports only observed evidence
ok - passed: terminal delivery reports only observed evidence
ok - active run-step is authoritative
ok - stale needs-decision over active run is superseded
ok - stale blocked over active run is superseded
ok - daemon/timeout blocked claim over a live fixing run reads as run alive
ok - socket refusal or missing socket over a stale fixing run reports blocked
ok - socket refusal over a terminal attributed run reports blocked
ok - socket-down evidence outranks a live run only while it is the log's latest event
ok - broken-pipe blocker over a live run keeps the plain superseded reading
ok - genuine daemon-down blocked line still reports blocked
ok - a busy secondmate keeps its open blocker until that exact key closes
ok - the most recently opened decision supplies the reported state and detail
ok - ship and scout terminal declarations supersede stale decisions
ok - latest status retains legacy completion events and shared captain matching
ok - latest status subprocess work stays bounded an

... [4823 bytes truncated] ...

 reads alive, never gone or dead
ok - fm-crew-state remote: an unreachable host reads unknown-remote, never gone or dead
ok - fm-crew-state remote: the remote host's own dead verdict is reported truthfully
ok - missing meta is handled gracefully
ok - crew_is_provably_working absorbs a validating crew found only via the runs-list fallback
ok - crew_is_provably_working still surfaces a genuinely stopped crew (safety property preserved)
ok - usage error exits 2
ok - historical same-branch rewritten head is not attributed as current
ok - active run with valid descendant fix head remains current
ok - local work advanced past run head invalidates attribution
ok - pipeline-owned active run binds without head equality and beats the failed row
ok - a genuinely failed run with no later run is not hidden
ok - coarse scan anchors the unresolvable active row instead of falling to an older one
ok - coarse scan with a mismatched anchor stays unknown and lets the pane answer
ok - coarse terminal row at a foreign head is not attributed
ok - an executing run binds regardless of branch_sync state
ok - a parked run keeps the strict head rule without pipeline_owned
ok - run_parked_scalar_gate_running keeps the strict head rule despite its live status word
ok - run_parked_in_gate_block keeps the strict head rule despite its live status word
ok - the exemption never applies to a terminal run
ok - missing run head falls back instead of matching by branch
ok - pipeline-advanced run head is not reported as a stale failure
ok - coarse walk binds the anchored active row instead of a stale failure
ok - submitted head binds an advanced-head run and preserves its terminal verdict
ok - terminal verdict is withheld when run identity cannot be established
ok - coarse terminal verdict is withheld when identity cannot be established
ok - genuine failed run at the current head is still reported failed
ok - provably diverged newest row blocks attribution instead of falling through
ok - a hardened task with a finished run but no quality receipt is not done
ok - a hardened task with a passing quality receipt reports done unchanged
ok - a hardened task whose receipt predates HEAD reports done and names the drift
ok - a quality verdict on a detail-less status line leaves no empty field
ok - an unreadable quality receipt reports unknown, distinct from a missing one
ok - a hardened task whose receipt only reported a score is not done
ok - a task with any posture but hardened is untouched by the quality gate
ok - the quality gate filters only a done verdict, never an in-flight one
ok - active fix round with an unfetched pipeline head reads working
ok - unanchored unverifiable active row is attributed because it is live
ok - unresolvable terminal row never reads as current
ok - runs-list continuation attribution works when axi answers another branch
ok - herdr stale registration over a shell-only pane reads agent gone, not alive
ok - herdr stale working record never reports a shell-only pane busy
ok - capped overview retains both competing same-branch run ids
ok - same-branch identity survives both runs falling outside the overview
ok - a capped overview with zero same-branch rows reports absent, not unreadable
ok - no run for this branch beside a live run elsewhere reads absent, not unreadable
ok - the capped inventory reader is bounded by the crew read budget
ok - a repo spelling the inventory does not record reads unreadable
ok - a linked worktree green PR in merge monitoring reads held for merge
ok - complete inventory preserves the replacement gate without writes
ok - missing complete-inventory failure reports unknown
ok - corrupt complete-inventory failure reports unknown
ok - schema complete-inventory failure reports unknown
ok - repo complete-inventory failure reports unknown
ok - count complete-inventory failure reports unknown
ok - norepo complete-inventory failure reports unknown
ok - R6 complete selection ignores unrelated branch semantics
ok - R6 requested branches use exact identity without a whitelist
ok - R6 capped inventory ignores unrelated semantics and names both ids
ok - R6 complete identity lookup precedes partial-row semantic rejection
ok - R6 capped inventory preserves quoted requested-branch identity
ok - R6 structural completeness and requested-run validation remain enforced
ok - R5 complete inventory without Python keeps the replacement gate
ok - R5 complete ambiguity without Python names both ids
ok - R5 capped lookup without Python preserves available ids
ok - R5 capped lookup without SQLite support preserves available ids
ok - R1 both directions of inventory liveness disagreement read unknown
ok - R2 uninitialized busy workers retain pane reporting
ok - R2 uninitialized idle workers retain status reporting
ok - R3 historical inventory yields to the current busy pane
ok - R3 historical inventory yields to current worker status
ok - superseded cancelled run preserves the replacement review gate
ok - a live rebased run beats an older failed run at the local head
ok - pending run with a rebased head reads working
ok - running run with a rebased head reads working
ok - legacy live rebased run is authoritative over an older failed row
ok - legacy surface binds a fixing run at a rebased head
ok - legacy surface binds a ci run at a rebased head
ok - an unproven record at a diverged head does not answer for the crew
ok - an unproven record with a dead daemon never overrides a busy pane
ok - an ordinary blocked tip over a coarse live row keeps the superseded reading
ok - a socket-refused blocker survives the dead-daemon verdict
ok - an ordinary blocked tip survives the dead-daemon verdict
ok - a parked gate survives a dead daemon with its findings intact
ok - an unproven record at a diverged head does not answer on the selected route
ok - a live record at a diverged head binds while the daemon answers
ok - the anchored continuation binds while the daemon answers
ok - a head-tied coarse record keeps its working reading and its original note
ok - a coarse failed record with a dead daemon reads unknown
ok - the selected-run anchored continuation reports the dead daemon, not an identity failure
ok - the selected-run anchored continuation binds while the daemon answers
ok - the selected route keeps an anchored parked run's gate with a dead daemon
ok - an open decision survives the dead-daemon verdict on the selected route
ok - a coarse pending ledger word reads unknown
ok - an unanswered probe never turns a failed coarse record into a gate
ok - the selected-route dead-daemon verdict names the run once
ok - a head-tied row reads working when axi names the self run
ok - a head-tied row reads working when axi names the other run
ok - a head-tied coarse row is exempt even when the record head diverged
ok - an unrecognised ledger word keeps the ordinary supersede note
ok - an unanswered daemon probe leaves a live rebased run bound
ok - a head-tied coarse live row is exempt from the dead-daemon verdict
ok - a coarse live row over an open decision keeps the original supersede note
ok - a coarse live row at a rebased head is not attributed
ok - a terminal run at a diverged head keeps the strict head rule
ok - competing live runs report unknown with both run ids
ok - newer failed run remains failed beside an older live run
ok - newer verified live run outranks terminal history
ok - missing-proof terminal identity respects explicit ownership
ok - foreign-anchor terminal identity respects explicit ownership
ok - matched-anchor terminal identity respects explicit ownership
ok - missing run selection reports unknown with candidate ids
ok - wrong-id run selection reports unknown with candidate ids
ok - wrong-branch run selection reports unknown with candidate ids
ok - wrong-head run selection reports unknown with candidate ids
ok - missing-status run selection reports unknown with candidate ids
ok - malformed-table run selection reports unknown with candidate ids
ok - inventory-error run selection reports unknown with candidate ids
ok - selected-error run selection reports unknown with candidate ids
ok - legacy conflicting run records report unknown
all fm-crew-state tests passed
Evidence: Herdr away-supervision live drive

Source: Herdr away-supervision live drive

ok - reproduced: daemon hosted in its target pane defers forever on Herdr busy state
ok - failed-run cleanup removes the lab session and its created workspace
fm-afk-launch: Claude + Herdr native launch redirected to a non-visible daemon terminal
fm-afk-launch: daemon launched in non-visible herdr workspace w2 (pane fm-lab-fm-afk-inject-se-2658260-21431:w2:p1), supervising fm-lab-fm-afk-inject-se-2658260-21431:w1:p1
ok - fixed: digest delivered to idle Claude pane while a Herdr fleet worker stayed busy
Evidence: Supervision-host reconciliation

Source: Supervision-host reconciliation

ok - host+hook: a successor closed before exit_to_main does not suppress the branch-outcome rewake
ok - host+hook: a successor that closes during a held engine turn does not suppress the branch-outcome rewake
ok - host+hook: a refused hand-back becomes a delivered failure notice
ok - host+hook: a refused hand-back on an announced handling marker becomes a delivered failure notice
ok - host: parked child-exit sampling uses ordinary half-second sleeps
ok - report surface: only the branch actor's current turn may report, and only on the tasks its wake names
ok - report surface: visible late outcomes queue a relay, while silent outcomes remain stored without a wake or note
ok - dispatch entry: the host reads branch eligibility, the offer rule, and the wake prompt from the Pi branch's own owner
ok - drain: BRANCH OUTCOMES runs on a Claude home by default and on another primary with the file, never with off, and never on Pi
ok - drain: captain outcomes come first, and routine overflow collapses into a count one drain clears
ok - drain: repeated captain outcomes collapse per task, and the byte cap presents only the run its acknowledgement covers
ok - drain: a long away window costs one short drain, captain outcomes collapsed per task and routine overflow counted, and nothing from it is shown again
ok - drain: the BRANCH OUTCOMES budgets count bytes, cutting multibyte summaries by whole characters in any locale
ok - drain: branch outcomes stay unread when a projection of the store fails
ok - drain: branch outcomes stay unread and the drain fails when jq is missing
ok - drain: branch outcomes stay unread when the drain cannot print them
ok - drain: a legacy backlog is presented with each outcome's age and a check-first instruction, never adopted
ok - drain: a keyed decision survives acknowledgement through a newer outcome for its task
ok - drain: an outcome carried across a switch off Pi comes back with its age, never adopted
ok - drain: an outcome nothing has shown is presented until acknowledged, and a repeated acknowledgement changes nothing
ok - drain: an outcome the host drain presented but main never acknowledged stays unprocessed across a switch to Pi
ok - drain: an outcome the host drain presented but main never acknowledged survives an outcome index repair
ok - host: an attended wake the branch may take is handled on the engine, and its routine outcome never wakes main
ok - host: an attended captain outcome wakes main once and stays in its drain until main acknowledges it
ok - host: a captain outcome recorded after the captain left waits for the return, then reaches main's drain
ok - host: a quiet record without its daemon is a present captain, so outcomes and decisions reach main
ok - host: an attended decision close stays main's exactly as the plain arm delivers it
ok - host: an off written while the host is parked sends the next attended close to main, naming the opt-out
ok - host: a main-only pass-through leaves the successor watcher running and the close undelivered for main
ok - host: an attended close whose main session cannot be identified reaches main and runs no engine turn
ok - host: a decision close accepted away whose turn starts attended still reaches main unchanged
ok - host: an attended close whose task turns main-only before its turn still reaches main unchanged
ok - host+hook: an attended main-only pass-through rewakes main and keeps its successor watcher
ok - host+hook: a captain outcome beside a quiet record rewakes the present captain with no away note
ok - host+hook: a Claude home without config/supervision-host runs the host at the default engine, and an off file restores the plain arm
ok - host+hook: a close that turns main-only at its turn rewakes main and keeps its successor watcher
ok - host+hook: failed at-turn downtime write notifies main despite a healthy successor
ok - host+hook: a successor close that lands during main's turn is delivered at the next turn end
ok - host+hook: the next park takes over the cycle a main-only pass-through left, so one arm owns it
ok - host+hook: a park stopped mid take-over leaves the take-over to the next park
ok - host+hook: a successor that cannot be recorded is stopped, and main's next turn end arms a fresh cycle
ok - host: a primary with no verified dialog mirror keeps every attended close on main, and its away posture still runs
ok - host: each wake carries the captain's dialog since the last wake, without operational input
ok - host: the dialog mirror, its feed, and the wake file are owner-only
ok - host: dialog a turn never completed with its report (a park boundary, a stopped turn, no report) reaches the next turn
ok - host: an attended wake whose mirror is missing, cannot be read (at the feed or at the prompt), or holds a malformed entry reaches main before any engine turn, and the cursor stays put
ok - host: an away wake is handled on the engine through the branch contract and never reaches main
ok - host: an engine turn that records no outcome hands its durable wake to main
ok - host: a captain return during an engine turn hands that turn's outcomes to main
ok - host: silent outcomes are excluded from both captain-return handoff paths
ok - host: an early visible outcome survives more than 1,000 same-turn receipts
ok - host: an outcome lookup failure forces a visible main handoff
ok - host: an outcome recorded after the return reaches main even when its host dies at the turn's end
ok - host: a killed predecessor's engine is reaped and the next host removes its turn files
ok - host: a turn that reports but leaves its granted rows queued hands the wake to main
ok - host: a captain return during a failed engine turn still hands that turn's outcomes to main
ok - host: an engine turn whose result is incomplete hands its wake to main
ok - host: two engine errors latch the session, main keeps every away wake in the cooldown, a failed probe doubles it up to its cap, and a report recovers it silently
ok - host: an attended close in a latched session reaches main as the arm printed it and leaves the latch as it was, and a home that opted out with off never reads it
ok - host: attended, two engine errors latch the session, main keeps every close unchanged in the cooldown, a failed probe doubles it up to its cap, and a routine probe's recovery stays in the ledger, off main
ok - host: an engine turn is bounded, and tool processes outside its process group are reaped
ok - host: a restarted host stops, by recorded identity, the cycle a killed predecessor left running
ok - host: the park ends itself with a boundary wake and a stopped watcher
ok - host: waiting closes cannot carry the park past its boundary
ok - host: a close whose margin runs out while the successor starts reaches main at the boundary without a turn
ok - host: the park's test clock is inert without the test marker
ok - host: a park at or beyond the Stop-hook registration falls back to the default boundary
ok - host: an owner's later park limit lets a turn outlive the boundary, and no other limit does
ok - host: the first cycle's status streams once, --restart replaces a stale watcher, and an owner predecessor makes a handling successor
ok - host: an unchanged held outcome reaches the captain once across cadences, and a later decision on the task still does
ok - host: a home naming an unverified engine hands every away wake to main with the reason
ok - host: a host outside the session-lock owner stands down without arming
ok - host: a host under a superseded auto-arm generation stands down without touching the owner
Evidence: Task-delivery reconciliation

Source: Task-delivery reconciliation

ok - fm-spawn/fm-promote: authorized intent preserves exact words and refuses operator-address lines
ok - fm-spawn: every legacy worker receives scoped role instructions without changing project or primary instructions
ok - fm-spawn: a ship spawn requires a valid explicit mode and yolo before anything is created
ok - fm-spawn: scout and secondmate spawns refuse ship delivery flags
ok - fm-spawn: the brief's recorded mode and the spawn's explicit mode must agree
ok - fm-spawn: a rigor downgrade against the registered posture is announced, never blocked
ok - fm-spawn: a scout spawn resolves no delivery posture from the registry
ok - fm-promote: promotion requires the delivery contract and records it exactly once
ok - fm-promote: a hardened standing posture is announced on promotion, never blocked
ok - fm-project-mode: the conditional policy is accepted, mapped for mechanical callers, and readable raw
ok - fm-project-mode: --quality reads +hardened from any bracket position and defaults to standard
ok - fm-project-mode: +hardened rides a flat mode and is refused on the conditional policy
ok - fm-project-mode: the two-word stdout its three callers parse is unchanged by the quality posture
ok - fm-spawn: --quality is closed-set validated on ship spawns and refused everywhere else
ok - fm-spawn: the brief's recorded quality and the spawn's quality must agree in both directions
ok - fm-spawn: a ship task records quality= and base_sha= additively, and a scout records neither
ok - fm-spawn: a standard task under a hardened standing posture is announced, never blocked
ok - fm-promote: a symlinked task record is refused and its target is left untouched
ok - fm-promote: a promoted worker receives the same mode-specific delivery contract a briefed one does
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
●  WATCHER DOWN - SUPERVISION IS OFF
●  1 task(s) in flight, but no watcher has a fresh beacon (last beat: never, grace 300s).
●  Trust the emitted supervision protocol for this harness; do not use shell & for watcher repair.
●  This is a supervision warning only; the guarded operation WILL still run.
●  repair missing watcher supervision according to the session-start block for this harness; do not use shell &.
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
ok - fm-promote: a selected branch prefix reaches both worker instructions and durable task state
Switched to a new branch '$(touch${IFS}/tmp/fm-task-delivery.h7pQpc/promote-branch-shell-safe-marker)/promote-branch-safe-e3'
ok - fm-promote: ref-format-valid shell metacharacters stay literal in promotion branch commands
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
●  WATCHER DOWN - SUPERVISION IS OFF
●  1 task(s) in flight, but no watcher has a fresh beacon (last beat: never, grace 300s).
●  Trust the emitted supervision protocol for this harness; do not use shell & for watcher repair.
●  This is a supervision warning only; the guarded operation WILL still run.
●  repair missing watcher supervision according to the session-start block for this harness; do not use shell &.
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
ok - fm-merge-local: a registry change cannot redirect an in-flight local-only task
ok - fm-project-mode: the registry lookup matches a whole multi-word name, not just its first token
ok - fm-project-mode: the forge binds from its own token and is reported only through --forge
ok - fm-project-mode: only a malformed forge binding refuses; every other token keeps its old tolerance
ok - forge=gerrit: yolo is refused with its reason, never silently dropped
ok - forge=gerrit: no-mistakes runs with its forge steps skipped, recovers its fixes, then publishes one change
ok - forge=gerrit: direct-PR publishes one squashed change and a stack is refused with its reason
ok - fm-spawn: a registered forge must reach the worker's brief
ok - fm-spawn: the brief must carry the spawn's selected ship branch, and the selection is validated before anything is created
ok - fm-spawn: a ship branch that deviates from the registered prefix is announced, never blocked
ok - fm-spawn: a registry forge token the parser refuses stops the launch, reason included
ok - fm-promote: a promoted worker receives the project's registered forge contract with no flag to remember
ok - fm-spawn/fm-promote: leftover Task placeholders are refused until both subsections are filled
ok - fm-project-mode: --branch-prefix resolves order-independently and defaults to the legacy fm/ prefix
# all fm-task-delivery tests passed

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

⚠️ **Rebase** - 1 warning
  • ⚠️ .agents/skills/afk/SKILL.md - branch carries 24 commit(s) that exist on your local main branch but were never pushed to origin/main; these may be unintended bundled work (proposed PR changes 345 file(s)):
  • 1726136 no-mistakes(ci): Fixed all three CI failures. Updated the Calm export test for Pi 1.0's renderer API while retaining older compatibility, and increased two test-only process startup windows to avoid loaded-runner flakes. All three focused tests pass; timing-sensitive tests passed five repeated runs each. ShellCheck, actionlint, Bash syntax, and git diff checks pass
  • a873559 no-mistakes(review): Harden pending-reply hash fallback under nounset
  • 536aa1b fix(bin): stop cap-cut forge reads recording observation unavailable
  • 72dbec0 no-mistakes(ci): Fixed both CI failures. Updated the Lavish text-range assertion to match nested path metadata output and added fm-codex-catalog-lib.sh to the synthetic remote-root fixture. Both affected test files pass locally. Bash syntax and git diff checks also pass
  • d393b4c no-mistakes(document): Document Codex catalog validation and fallback
  • 68b48e7 no-mistakes(review): Validate Codex catalog schema before enabling max
  • 5bf3d87 no-mistakes(document): Document Codex catalog helper
  • fce2fe2 no-mistakes(review): Reject truncated Lavish list items as incomplete
  • ca58502 no-mistakes(review): Relay Codex max warnings through secondmate restarts
  • 112c121 no-mistakes(document): Document catalog validation and feedback metadata
  • 5298637 no-mistakes(review): Relay Codex max downgrade warnings through recovery
  • 25ab3ef no-mistakes(review): Preserve Lavish metadata and harden Codex catalog validation
  • 4299989 Validate Codex max effort from catalog
  • 61985d8 fix(pi): support Pi 1.0 rendering contracts
  • 48b50ae no-mistakes(document): Document catalog-based Codex max effort
  • 10d8297 no-mistakes(document): Align Codex and Lavish documentation
  • 4ea5d84 no-mistakes(review): Reject malformed Lavish lists and isolate Codex catalogs
  • c25c34e Pass supported Codex max effort to workers
  • 8a87193 no-mistakes(document): Document Lavish result decoding guarantees
  • 2a37594 Preserve Lavish UTF-8 feedback
  • e66a604 no-mistakes(ci): Updated the malformed-capture regression to require read's nonzero incomplete verdict while retaining output assertions. tests/fm-procevent.test.sh, tests/fm-bearings-board-render.test.sh, bin/fm-lint.sh, and git diff --check pass. Serial 8 was a transient runner failure and passed six local runs
  • c1e01b4 no-mistakes(document): Document Lavish list-form feedback parsing
  • 9147025 no-mistakes(review): Parse table rows beyond declared counts
  • e9d8cc2 Handle Lavish list-form feedback

Confirm these commits belong in this PR before approving, or manually separate the intended work onto origin/main before gating.

⚠️ **Review** - 3 errors
  • 🚨 bin/fm-worker-account-lib.sh:230 - Raw launch wrappers bypass both worker-account pins. For --harness "env CLAUDE_CONFIG_DIR=/other claude --print raw", and the analogous Pi command, the parser records HARNESS=env at bin/fm-spawn.sh:2208. This call then resolves no pin, so the Claude and Pi refusal branches at lines 236 and 251 and the launch prefix at bin/fm-spawn.sh:5264 never run. Wrappers such as command, exec, and bash -c have the same path. A durable remedy needs authorization: constrain the supported raw-command grammar so the effective runner and assignments can be proven, or remove raw launches from pinned-account support. Matching individual wrappers would leave the bypass reachable.
  • 🚨 bin/fm-dod-lib.sh:392 - The merge chose one of two incompatible no-mistakes handoff protocols without the captain ruling required by intent criterion 3. Current generated briefs tell workers to start validation themselves and withhold done: at lines 392-394 and 450-452; line 510 gates every no-mistakes done:, and .agents/skills/validation-supervision/SKILL.md:11-13 says Firstmate never sends the start trigger. Upstream parent 1f3e769 instead has the worker emit a pre-validation done: and stop until Firstmate triggers validation, and exempts that handoff from the named-head gate. Ask the captain which protocol should govern, then align every listed site.
  • 🚨 bin/fm-contributions.sh:350 - The merge chose the fork's meaning of a closed contribution without the captain ruling required by intent criterion 3. Upstream parent 1f3e769 treats both merged and closed as final, while fork commit 48b9e5e deliberately keeps closed contributions refreshable so reopening is detected. The merged result retains merged-only predicates here and at lines 382 and 394, with the same rule documented at line 61. A closed issue therefore remains in every poll and can later reopen; under upstream semantics it settles permanently. Ask the captain whether reopening must be observed or closed must be terminal, then apply that decision consistently to all predicates and documentation.
✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 3 of 3 scenarios driven live against the product
Scenario Result Live Evidence
A trusted session keeps its lock across helper-process ancestry changes, while another or untrusted session is refused ✅ pass live session-lock-ownership.log and session-lock-ancestry.log show real Firstmate lock entry points preserving the trusted session across a recycled process chain, refusing weaker identities, and recla…
An active validation run at an unfetched pipeline head remains bound only when the immediately older same-branch ledger row anchors it to local HEAD ✅ pass live crew-state.log shows the real crew-state command binding an unresolvable active row through the immediately older same-branch anchor, while mismatched anchors remain unknown and terminal rows retain…
Away supervision delivers to an idle Claude pane through a separate daemon terminal while a fleet worker remains busy, and rejects unsafe target selection ✅ pass live herdr-afk-self-deadlock-e2e.log records the old self-target deadlock, unsafe-target refusals, failed-run cleanup, separate daemon-terminal launch, and successful digest delivery while the fleet work…
  • bash tests/fm-session-lock-ownership.test.sh
  • bash tests/fm-session-lock-ancestry.test.sh
  • bash tests/fm-crew-state.test.sh
  • bash tests/fm-watch-triage.test.sh
  • bash tests/fm-control-relaunch.test.sh
  • bash tests/fm-remote-reply.test.sh
  • bash tests/fm-supervision-host.test.sh
  • bash tests/fm-brief.test.sh
  • bash tests/fm-claude-trust.test.sh
  • bash tests/fm-spawn-dispatch-profile.test.sh
  • bash tests/fm-task-delivery.test.sh
  • bash tests/fm-remote-secondmate-lifecycle-e2e.test.sh
  • bash tests/fm-secondmate-sync.test.sh
  • bash tests/fm-sessionstart-nudge.test.sh
  • bash tests/fm-teardown.test.sh
  • bash tests/fm-afk-inject-self-deadlock-e2e.test.sh
  • Preflight and postflight with bin/fm-herdr-lab.sh run fm-lab-test-inspect session list --json
  • Final git status --short --branch cleanup check
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

kunchenguid and others added 30 commits September 17, 2026 19:24
)

* Improve CI reliability and rebalance full-coverage validation

* no-mistakes(document): Clarify lint partition documentation
…guid#4799)

* Handle Kimi workspace trust dialog

* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers

* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures

* no-mistakes(review): Read visible pane for Kimi trust and ready gates

* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate

* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection

* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca
…er (kunchenguid#4775)

* fix(bin): report a record whose agent is gone once instead of escalating forever

The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.

Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.

fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that was about to escalate, read it once and treat
only its two recovery-grade verdicts - dead (endpoint present, no agent in it)
and missing (endpoint authoritatively absent) - as proof, reporting that record
once and not re-escalating it while it stays that way. Every other verdict,
including alive, ambiguous, unreadable, unverified, and a read that failed
outright, keeps the identical schedule, reason, and escalation count, so a
genuinely wedged live agent is unaffected. The probe costs at most one backend
read per window per threshold, the same budget the declared-wait consult and the
worktree write probe already take.

The report decides nothing about the record's fate: both lanes still held
unlanded work and teardown refusing them was correct, so retiring, relaunching,
or cleaning up stays with the supervisor. The once-only marker is owned entirely
by that function and is dropped by the same read the moment the endpoint stops
reading gone, so a replacement launched into the same window escalates normally
and its own later death is reported again.

Related, and not closed by this: kunchenguid#4412, kunchenguid#4482, kunchenguid#4316.

Tests drive the real watcher against a record whose endpoint does not exist and
pin both directions: dead and missing report once and never advance the count
across later thresholds, while alive, ambiguous, and unreadable endpoints keep
escalating with the identical reason and a climbing count.

* fix(bin): bind the once-only dead report to the pane it reported

Review of the parent commit found a reachable sequence where a later death in
the same window lost its promised report. The marker was keyed on the verdict
string alone and dropped only when a threshold probe read a non-gone verdict,
but probes run only at thresholds: a replacement launched into the same window
that dies without ever being probed alive - it crashes at startup, or works and
then crashes - was absorbed by the previous death's marker. The pane's first
sight yielded only the generic stale wake and every later threshold matched the
stale marker, so the second death never got the detailed once-report that both
the function's own comment and docs/architecture.md promise.

Record the verdict together with the pane hash it was reported for, and absorb a
repeat only while both still match. A replacement churns the pane, which resets
the stale suppressor, wedge timer, and escalation count while no reset site
touches this marker, so the pane half is what tells the second death apart from
the first. The live-probe drop stays as it was.

Clearing the marker at those reset sites instead would re-open unbounded
re-alarming for a dead pane whose display ever ticks, which is the exact defect
the parent commit exists to close.

The noise bound is unchanged: an unchanged dead pane still absorbs on every
later threshold and never advances the escalation count, and every verdict short
of proof still escalates exactly as before.

* no-mistakes(review): Key the dead-record once-marker on the busy incarnation token

* no-mistakes(document): Document dead-record escalation cap in stale-pane config entry

* no-mistakes(document): Add busy-state inventory line to AGENTS.md

* no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path
…id#4854)

Captain holds have no due semantics and are a hold kind, not a Beads issue
type. The create path now waives due.required and maps to native type task.

Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(bin): launch every spawned agent with the compact adviser disabled

Every crewmate, scout, and secondmate Firstmate launches now starts with
COMPACT_ADVISER_DISABLE=1, on a fresh spawn and on a relaunch alike, so an
unattended session never activates the compact adviser.
The value is unconditional: no configuration file gates it and there is no
override, unlike the trace carrier beside it.

Three carriers deliver it, because no single one covers every launch shape.
The pane shell receives an export beside GOTMPDIR, so the agent's own children
inherit it too.
The launch command carries an explicit assignment, prepended outermost so it
wins over any ambient value the pane already held.
The cleared launch environment sets it again at the `env -i` boundary and keeps
COMPACT_ADVISER_DISABLE in the fixed operational floor, which is what preserves
the switch when config/launch-env-allowlist empties the environment, and what
delivers it on a remote host that never had the value.

bin/fm-control.sh relaunch, the bootstrap secondmate relaunch, and the remote
secondmate transport all rebuild their launch through bin/fm-spawn.sh, so they
inherit the same floor.
The captain's own primary session is untouched.

The two new suites drive the real spawn and then execute the launch command the
pane actually received, with the harness replaced by a probe that prints its own
environment, rather than matching script text.
They cover ship and secondmate launches with the allowlist absent and enabled,
the pane export and its ordering, fm-control.sh relaunch, and the full parent to
remote-host chain.

* no-mistakes(review): Export compact-adviser disable across compound launches

* no-mistakes(document): Document spawned-agent compact-adviser environment guarantee
…henguid#4894)

* fix(bin): let a background Claude session keep owning its session lock

Session-lock ownership was decided by process ancestry alone. Under an
unattended Claude session the model loop runs in a transient bg-spare
bridged to the front-end by a shared daemon; when that bridge is
recycled the contiguous claude-named ancestry from a hook to the
recorded owner breaks while the owner pid stays alive, so the Stop
auto-arm stood down as a foreign live owner, the turn-end guard ended
every turn with its read-only diagnostic, and fm-lock.sh refused - a
self-sustaining outage until restart.

Ownership is now ancestry membership OR a trusted same-session id,
never id-first:

- fm-session-lock-lib.sh accepts CLAUDE_CODE_SESSION_ID only when
  CLAUDE_PID is a Claude-shaped member of the current contiguous run,
  compares it against the id recorded in state/.lock-session, and
  requires the recorded pid to still be a live harness. No id, no
  sidecar, an untrusted id, a different id, or a dead recorded pid
  leaves the ancestry verdict unchanged. Ids are never read from ps
  argv.
- fm-lock.sh accepts a same-session holder at both refusal sites,
  writes, refreshes, and clears the sidecar only under its claim lock
  (including the early already-mine exit, skipped only while the
  deferred startup sweep leases that lock), keeps it byte-identical
  across a same-session confirmation, records CLAUDE_PID on lock line 1
  for a session with a trusted id so a shared daemon or front-end that
  outlives the session never keeps a dead session's lock alive, never
  rewrites a live line 1 on a same-session confirmation, and names the
  recorded id in the live-owner refusal.
- The .lock line-1 format is unchanged, so every reader that takes the
  whole first line as the pid keeps working; the guard's foreign-owner
  exit is unchanged and inherits the fix through the shared predicate.

Tests: the ancestry suite drives the ancestry and id signals apart in a
deterministic process table (asserting the divergence) and runs a real
orphaned front-end/daemon/pty-host/spare tree through six phases with
the real lock, auto-arm, and guard scripts; the foreign-owner repro
keeps its negative control and adds a same-id positive control.

Disclosure: no live unattended Claude background session ran on the
verifying machine. The topology is documented by the real process
listings in kunchenguid#3902, kunchenguid#2314, kunchenguid#3398, and kunchenguid#4066; coverage is the structural
predicate plus the executable fixtures, not a live pass.

Residual: bin/fm-sessionstart-nudge.sh keeps its own private ancestry
walk (it only decides whether to print a nudge) and may nudge on a
resume in the recycled case.

Out of scope, deliberately: no structured lock format, no guard budget
changes, no daemon-identity rejection, no fork lineage.

* no-mistakes(review): Wait for claim lock; revert failed sidecars

* no-mistakes(review): Revalidate ownership after wait; restore sidecars

* no-mistakes(review): Roll back sidecar by publication phase

* no-mistakes(review): Restore sidecar only if lock line is unchanged

* no-mistakes(review): Trust session ids without a spelling allowlist

* no-mistakes(review): Disarm sidecar rollback before backup cleanup

* no-mistakes(document): Updated session-lock ownership documentation
* feat: park main under the away posture on Pi

While the away-posture record exists on a Pi primary, the supervision branch
takes every actionable wake, no processing turn opens on main, captain rows
accumulate for the return brief, and main's standing authority relocates to
the branch through the existing guarded scripts.

- lib/fm-branch-dispatch.ts: read the record at every routing decision; while
  it exists claim check, decision-owned, and heartbeat rows too, keeping the
  two broken-queue vetoes; expose checkSeqs so a claimed check row lifts task
  scoping.
- fm-primary-pi-watch.ts: offer every actionable row under the record; a
  declined wake and every watcher-failure alarm still reach main.
- fm-branch-supervision.ts: drop the legacy .afk decline; append a fixed
  POSTURE: AWAY tail carrying the record's read-back verbatim per wake; open no
  processing request while the record exists, re-checked immediately before a
  request would open and at every run boundary; present the accumulated rows
  at the first run boundary after archive.
- fm-lease-lib.sh: fm_lease_forbid_branch passes the branch for opted-in
  actions only while fm-afk-contract.sh validate succeeds on a confirmed live
  record; PR merge, fresh spawn, and decision answer opt in, local landing
  never does.
- fm-send.sh: a --resolve-key naming an open needs-decision or captain-held
  task is a decision answer and meets the partition; blocked: keys stay
  steering.
- fm-spawn.sh: enforce the record's spend cap for a fresh ordinary spawn by
  either actor; relaunches and secondmates exempt.
- fm-branch-prompt.sh: fixed Postures section and the verbatim
  ask-user-authority policy; the prefix stays byte-stable.
- fm-afk-return.sh: count what the away session handled from the store.
- docs, afk skill, AGENTS.md stub: main parked on Pi, green merge gate
  absolute while away.
- tests: watcher and branch extension suites, fleet-record, merge, and
  decision-answer suites cover the relocation, the vetoes, the tail, the
  parked processing turn, the cancellation, the re-presentation, and the
  spend cap; dated live-guard evidence recorded.

* no-mistakes(review): Refuse branch merge after preflight archive race

* no-mistakes(review): Fix away wake, spawn, and processing races

* no-mistakes(review): Suppress parked processing; narrow away-only rejection

* no-mistakes(review): Abort dedicated processing; gate branch spawn once

* no-mistakes(review): Stamp away-only on the dispatch offer

* no-mistakes(review): Treat invalid away records as spend-cap absence

* no-mistakes(review): Drop spawn test hook; abort processing-opened runs

* no-mistakes(review): Bind abort to opening prompt; cap-read absence

* no-mistakes(review): Limit away branch spawn to queued work only

* no-mistakes(document): Correct AFK posture documentation
* ci: simplify CI job timeouts to a three-tier policy

Replace the scattered per-job timeout values (10m parallel, 25m lint, 30m
serial, 10m macOS) with three readable tiers, each a hang tripwire with
headroom rather than a packing estimate:

- fast (5m): coverage guard, repo invariants, timing aggregate
- normal (30m, one shared budget): lint partitions, portable parallel
  shards, portable serial shards, macOS stock Bash
- heavy (Herdr only): 20m step tripwire on the family run so always()
  cleanup still runs, under a 75m job-level last-resort backstop

The workflow's header comment states the policy and points at
docs/fm-test-portable-shards.md "Timeouts", which now owns it, and each
job names its tier beside timeout-minutes. tests/fm-ci-workflow.test.sh
asserts the policy against the parsed workflow instead of the old
per-job minute values: every job joins exactly one tier, exactly three
distinct job-level values exist, the fast tier stays within 5-10
minutes, the normal budget stays at least double the modeled parallel
lane sum reported by fm-test-run.sh --check-coverage, and the Herdr step
tripwire stays below its job backstop with an always() cleanup after it.

Concurrency supersession, shard counts, lane membership, and fail-fast
settings are unchanged.

* no-mistakes(review): Decouple the normal timeout from packing estimates

* no-mistakes(review): Assert Herdr teardown follows the family run

* no-mistakes(review): Pin Herdr family-run timeout to 20 minutes

* no-mistakes(review): Ignore comments when identifying Herdr steps

* no-mistakes(review): Identify Herdr steps by declarative ids

* no-mistakes(document): Clarify authoritative three-tier timeout policy
…nchenguid#4895)

* fix(bin): keep supervisor status closes from waking the same home

A drain that already folded OPEN DECISIONS has presented those bytes even
when the watcher has no matching seen marker. Treat that fold, and the
presentation cursor, as known so the bookkeeping close stays quiet while
later worker lines still signal.

* no-mistakes(review): Keep folded worker failures waking past supervisor closes

* no-mistakes(review): Wake on unlisted folded worker lines; batch multi-key closes

* no-mistakes(review): Stop folded worker resolved lines from counting as already read

* no-mistakes(document): Correct self-announced close marker contract in docs
* Stop steering operators away from Herdr

* no-mistakes(review): Neutralize remaining Herdr opt-out documentation wording
…enguid#4973)

* fix(bin): treat a live no-mistakes run as current after rebase

A running run on the task's branch is authoritative regardless of head.
Matching only the local head made a rebased in-flight run look failed.

* no-mistakes(review): restrict coarse live-any-head to foreign-branch answers

* no-mistakes(review): reject gate-parked runs from the executing predicate

* no-mistakes(review): hoist gate-marker patterns into single run-lib owner

* no-mistakes(review): require live daemon for head-free run binding

* no-mistakes(review): require answered daemon-down before unbinding live runs

* no-mistakes(review): extend daemon guard to anchored continuation routes

* no-mistakes(review): delete live-any-head; restore dead-daemon verdict

* no-mistakes(review): keep parked gates parked; name dead daemon everywhere

* no-mistakes(review): set dead-daemon verdict instead of emitting early

* no-mistakes(review): align selected route with legacy dead-daemon handling

* no-mistakes(review): drop unproven-record binds; narrow coarse gate reading

* no-mistakes(review): narrow header, drop vestigial guard, retarget tests

* no-mistakes(review): revert coarse gate override; require answered-down probe

* no-mistakes(review): cache one daemon probe; stop duplicating run id

* no-mistakes(review): restrict coarse dead-daemon verdict to moved-off rows

* no-mistakes(review): delete coarse dead-daemon extension and gate note

* no-mistakes(review): delete remaining coarse dead-daemon block and stale docs

* no-mistakes(document): document rebase-safe live-run bind and unverified-record verdict
…4994)

* fix(bin): stage the launch command in a private file and type a short source line

A long launch line typed while the fresh pane shell is still busy waits in the
terminal's canonical line buffer, which drops input past about 1,024 bytes on
macOS, so the pane was left at an unfinished command with no agent running.
fm-spawn now writes the assembled command to the task's own temp root under
umask 077 and types only a short line that sources it.

Refs kunchenguid#4559

* fix(bin): keep the per-task temp root private before staging the launch command

The root lives at a predictable path under /tmp and now holds the whole launch
command. Create it with mode 0700, refuse one that already exists as anything but
a directory owned by this user that nobody else can write, and tighten an owned
one, so no other local user can plant or swap the staged file.

Refs kunchenguid#4559

* fix(bin): enforce private staged launch file mode

* test(spawn): cover long staged Claude launches

* no-mistakes(review): Namespace launch files and prove truncation staging

* no-mistakes(review): Use immutable per-spawn launch filenames

* no-mistakes(document): Document staged launch delivery safeguards

* no-mistakes(ci): Updated eight behavior tests/fakes to execute or inspect immutable staged launch files instead of expecting inline launch commands. This restores Muse, secondmate lifecycle/restart, remote trace/parent binding, compact-adviser, and Orca coverage. All affected tests, dispatch-profile regression, fixture tests, syntax checks, ShellCheck, and git diff checks pass

---------

Co-authored-by: Vytautas Stankus <svycka@gmail.com>
* Add isolated Herdr runbook to test instructions

* no-mistakes(review): Drop substring matching from test.instructions contract

* no-mistakes(review): Assert commands.test key absence in YAML

* Drop unit-first sentence and instructions contract test

Captain-scoped follow-up on the Herdr-lab test.instructions ship:
keep the lab safety runbook only, and leave the no-mistakes contract
test focused on commands.test absence.
…uid#4873) (kunchenguid#5001)

* docs(vision): accept vendor-semantics and 9k contract-ceiling amendments (kunchenguid#4873)

Replace the pixels-of-today's-UI rule with a quarantined, version-pinned
surface-adapter exception recorded as standing debt. Cap the always-loaded
contract at 9,000 words and require prune-or-trigger before a crossing change
lands.

Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

* docs(vision): restore accepted three-sentence vendor-semantics form (kunchenguid#4873)

Replace the compressed paraphrase with the issue's accepted wording:
a named quarantined version-pinned adapter, expected to break, recorded
as standing debt that never hardens into a shared contract.

Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
…or-owed gate (kunchenguid#4974)

* fix(watch): recheck a gate awaiting a human instead of wedge-escalating it

A lane whose validation run is parked at a gate waiting on a human
decision is correctly quiet, but nothing in its status line says so: the
evidence is the pipeline's own gate state rather than anything the worker
wrote. The wedge timer read that silence as a suspected wedge and climbed
the escalation ladder for as long as the wait lasted, and each escalation
cost a supervising turn. The landed declared-wait consult does not reach
it, because a live ordinary crewmate never reports a declared pause, and
raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for
every lane by the same amount.

The threshold now reads a second, independent record when the status line
accounts for nothing: whether the crew's current state is a gate whose
answer is owed by a human. That is minted only from the gate's own
findings table, by a row whose `action` column is exactly `ask-user`,
located by position out of the table header the way nm_gate_step_row
already reads its row - never searched for over the run payload, where a
finding's free-text description or a branch name satisfies a search just
as well. A gate awaiting the CREWMATE's own answer keeps the unchanged
escalation schedule, reason and demand-deep-inspection wording, because a
crewmate that goes quiet before answering its own gate is exactly the
wedge the ladder exists to catch.

Each kind of wait now carries the human it is on, the action that clears
it, and whether that human is the captain as data alongside the verdict,
rather than as wording chosen per branch where the recheck is written, so
the deferral cannot word one kind of wait as another and a new kind
cannot ship without deciding all of them. A parked gate has no written
record of when its wait began, so its recheck publishes no wait age at
all rather than one read from the quiet window this deferral resets on
every pass, which would report the same small number for a gate of any
age. Like every other captain-facing recheck here it is absorbed in
silence while the away-posture record exists, arming no throttle, so the
recheck is owed in full the moment the record is archived.

The consult runs only in the at-threshold branch that was about to
escalate, beside the worktree walk already there, and only for lanes
whose status line explained nothing.

Closes kunchenguid#3055

* no-mistakes(review): require an unanswered decision before deferring a parked gate

* no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records

* test(watch): pass the pane hash wedge_timer_check now takes

Upstream gave wedge_timer_check a sixth <pane-hash> argument for its
dead-record probe. The malformed-wait-record rounds drive the real function
directly, so they pass one, and stub fm_backend_agent_state to a live agent so
the probe that runs after a refused deferral keeps the unchanged ladder rather
than reading a backend the child shell has none of.

* no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate

* no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling

* feat(watch): make the parked-gate wait deferral opt-in

The wedge timer deferring a lane parked at a validation gate is new
supervision behaviour rather than a restored one, and it decides which
lanes give up the escalation ladder, so it now ships as a default-off
per-home option instead of changing every home on upgrade.

config/wedge-defer-parked-gate arms it. The flag is read before the
decision fold, so an unconfigured home spends no fold or current-state
read, writes no record, and keeps the unchanged escalation schedule,
reasons and demand-deep-inspection wording; a test counts the reader
calls in both directions to pin that.

It is not inherited by secondmate homes: each home supervises its own
crew and owns that trade separately, the same reason
config/turnend-churn-absorb is home-local.

The away-posture absorb returns to leaving the idle timer alone, which
it had restarted only because the costly consult could reach it. A
parked-gate wait is owed to the supervisor rather than the captain, so
it never enters that branch, and the recheck owed on return is again
owed in full the moment the record is archived.

* test(watch): pin that the away-silenced hold leaves the idle timer alone

The absorb no longer restarts the timer, so the recheck owed on return is
owed in full rather than a cadence into the return. Nothing asserted
that, so a restart could be reintroduced silently.

* no-mistakes(review): document away-silence rationale, pin captured gate component

* no-mistakes(test): anchor gate row scan to the braced findings header

* no-mistakes(document): pin same-block gate row invariant in crew-state comment
…uid#5007)

* fix(control): let the owning seat reclaim a task whose endpoint is gone

A destroyed pane or workspace made `missing` a terminal state. Relaunch
accepted only `dead` and said to stop the agent first; exit refused
`missing` and said to reconcile the task first; there is no reconcile
verb. Each command named the other as its prerequisite, so a task whose
terminal went away could not be reclaimed by anything, and a no-mistakes
approval it was parked on had no seat left to answer it.

`missing` is agent-free a fortiori: there is no endpoint, so there is no
agent in it. Widen the existing guards rather than add a verb.

- fm-spawn --relaunch accepts a positively proven `missing` and creates
  one fresh endpoint in the recorded worktree; the record it already
  republishes rebinds the task to it. A `dead` endpoint is still adopted
  in place.
- fm-control exit reports `endpoint-gone` instead of dying, so the
  relaunch transaction's stop step no longer dead-ends, and re-resolves
  the endpoint from the record before verifying the replacement.

The duplicate-agent refusal is untouched: both verdicts come from the
same recovery-grade classifier, which claims `missing` only from positive
absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The
backends' own create paths refuse a live same-labeled endpoint as a
second independent guard. The worktree, its branch, commits, uncommitted
changes, armed poll and registration, record rows, and status log are all
untouched - a reclaim is a recovery, never a teardown.

A secondmate is excluded: its gone-endpoint recovery already has one
owner in the session-start liveness sweep, so relaunch refuses and names
it rather than becoming a second path to the same outcome.

Tests reproduce both halves of the deadlock, the reclaim succeeding,
unlanded work surviving it, and the refusals that still hold.

* no-mistakes(review): prove endpoint absence per backend before reclaim rebinds

* no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session

* no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly

* no-mistakes(review): stop refusals and docs asserting unestablished causes

* no-mistakes(review): stop herdr fixture helper losing tmp-root registration

* no-mistakes(review): document workspace drift and absence-probe server residue

* no-mistakes(review): correct rebind limitation to its one reachable case

* no-mistakes(review): stop claiming reclaim leaves instructions untouched

* no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage

* no-mistakes(rebase): read the staged launch file in the herdr fixture

Rebasing onto main picked up kunchenguid#4994, which stages a long worker launch
command into a script and delivers the short `. '<path>'` line instead of
the literal command. The tmux fake and tests/fixtures.sh were updated for
that; the herdr fake this branch adds was written before it and still
keyed "an agent now exists on this pane" off the literal
`encode launch-brief` text, so after the rebase it never marked the
rebound pane live and the reclaim's alive-wait read `dead`.

Dereference the staged file first, exactly as the tmux fake above does.
Test-fixture only; no production path changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(document): note reclaim placement in herdr and scripts inventories

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…3764)

* test(status): reproduce missing event emission time

* wip(status): preserve optional event emission time

* test(status): document indirect clock stub invocation

* no-mistakes(review): Preserve historical status bytes during reply recovery

* no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies

* no-mistakes(review): Preserve captain regex overrides for timestamped status events

* no-mistakes(document): Clarify status event timing and publication contracts

* no-mistakes(lint): Quote literal done to satisfy ShellCheck

* no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed

* no-mistakes(test): Preserve terminal notifications with malformed timestamp tags

* no-mistakes(test): Stamp Rovo spawn failures with emission time

* no-mistakes(document): Verify status event documentation

* no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests

* no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed

* no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed

* no-mistakes(review): Stamp remote escalations at call sites, drop new flag

* no-mistakes(review): Accept stamped escalation and close lines in test assertions

* no-mistakes(review): Restore reserved-key answered-note guard for stamped closes

* test(status): accept optional emission time in PR-provenance assertions

The kunchenguid#4148 provenance test landed on main with exact unstamped greps.
Parent-channel lines from this branch carry [at=<epoch>], so strip only
that tag before the same exact match. No production change.

* no-mistakes(review): Accept stamped ready signal in PR fallback scrape

* no-mistakes(review): Drop relay flag, stamp parent events at call sites

* no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture

* no-mistakes(review): Accept optional stamp in live cmux drift guard

* no-mistakes(review): Restore original test invocation order in two suites

* no-mistakes(review): Strip only well-formed numeric status time tags

* no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc

* no-mistakes(review): Stamp agy spawn-failure status lines with event time

* fix(bin): normalize status event times in-shell and freeze the budget test clock

Two paths made a status event's emission time cost more than it should.

The captain-relevance fallback piped every line through awk to drop a
well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a
fork per line just to prepare a regex match. Shell parameter expansion does the
same strip with no fork, and the retry-dedup scan now reuses that one helper
instead of carrying a second copy of the rule in awk. The copies had already
drifted: the shell side stripped tags from lines with no colon, which the awk
rule left whole, so a colonless line could be mistaken for one already
recorded. One definition, checked against the awk rule it replaces over the
edge cases and a 4000-line fuzz.

tests/fm-contributions.test.sh froze its fixture clock only in exhaust mode. In
hang mode the poll set DEADLINE to the real now plus a one-second budget, and
when the second ticked before the first forge call the loop broke without ever
calling gh: forge/calls was never written and the assertion failed reading a
missing file. Freezing the clock in both modes removes the dependence on wall
time; the bounded call is still cut by the real timeout, so the observation the
test asserts still starts.

Emission time stays optional on new status records, and legacy or malformed
lines keep an unknown age.

* no-mistakes(review): Stamp ask-user escalation line and fix Kimi status assertion

* no-mistakes(document): Drop stale unstamped done-line spelling from watcher docs

* test: fold emission-time snapshot coverage into the fixture case

Drop the incidental ci.yml 18-to-19 count hunk so the PR no longer
touches workflows. Keep every emission-time assertion by folding it
into test_fixture_snapshot_json.

* no-mistakes(review): replace brief date substitution with epoch placeholder; drop emitted_at_epoch

* no-mistakes(review): align untimed normalizer with epoch parser; tolerate placeholder stamp in PR scrape

* no-mistakes(review): strip undelimited at-tags; correct brief stamp header

* no-mistakes(review): normalize stamps at both captain-regex sites; restore mtime freshness

* no-mistakes(review): strip colon-bearing stamps for relevance; fix headers and test oracles

* no-mistakes(review): narrow escalation match to stamp tolerance; pin note verb

* no-mistakes(review): read note and key past colon-bearing stamps

* test(status): keep inactive reconcile assertions stamp-tolerant

These two oracles were made stamp-tolerant while resolving one of the
branch's merges from main. The rebase drops merge commits, so that
adaptation was lost and both assertions went back to matching an exact
substring that a stamped line no longer contains: the tag lands before
the colon, so "failed [key=k]: ..." is now "failed [key=k] [at=N]: ...".
Strip a well-formed tag before matching, as the branch's other oracles do.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): unstamp fold colon tests; reserve stamp width in cap

* no-mistakes(document): correct stale unstamped status-line spellings in docs

* no-mistakes(document): quote brief-test literals for lint; correct stamp-helper contract comments

* no-mistakes(ci): rename subshell-local epoch in delivery-race stub

The serialization test overrides fm_pending_reply_mark_delivered inside a
(..) subshell. Its `epoch` local collided with the same name in
status_line_at_epoch/status_stamp_line, which this branch added and this
suite now calls at top level, so ShellCheck 0.11.0 reported SC2030 and
failed Lint 2. The stub already prefixes its other locals with `pending_`
for the same reason; `epoch` was the leftover.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: ship clean Lavish host fixes

* no-mistakes(review): Fix Lavish classifications and fail-closed host loading

* no-mistakes(review): Restore Lavish host state across retries and launches

* no-mistakes(review): Preserve destination Lavish host when configuration is absent

* no-mistakes(document): Document Lavish status and host guarantees
…#5076)

* feat(afk): make the captain's away words the whole mandate

Retire the clause fields, verb list, never-set scan, refused records, and
the per-task merge-grant list from the away-posture record. The record is
now version 2: the captain's words verbatim plus expected return, spend
cap, and reach line; a version 1 record still validates, reads, and
archives so a live away window is never broken by the upgrade.

The supervision branch reads the words at the tail of every wake and acts
on them by its own judgment through the guarded scripts under standing
authority, never by analogy, holding for the return on doubt, and opens
each such outcome summary with "per your away instructions:" so the
return brief can render the words beside the session's account. While the
record exists any green merge runs under away authority (ledger tag
"away"); red merges, --allow-red, asynchronous and queued merges, and
local-only landing stay refused. The branch may file a backlog item the
words explicitly call for before dispatching it under the spend cap.

Tests drive fm-afk-contract.sh, fm-afk-launch.sh, fm-afk-return.sh, and
fm-pr-merge.sh as commands: version 2 written, version 1 read, retired
flags and subcommands refused by name, green merges landing under the
record, red and waived-red refused, the record lock still closing the
authority-read window, and the Pi away tail carrying the words.

* no-mistakes(review): carry the away read-back to the session verbatim

* no-mistakes(review): match the exact away-action marker in the return brief

* no-mistakes(review): refuse a words block truncated by a damaged line

* no-mistakes(document): Refresh away-role contract documentation
…unchenguid#5049)

* fix(bin): render the remote charter's steering-inbox path host-local

A freshly provisioned remote secondmate read a parent-home absolute
steering-inbox path in its charter - a location that exists on no route -
and spent its first turn discovering the gap and filing a blocked
decision for what was a render defect. The seed's remote-copy rewrite now
maps the inbox to the route's host-local parent-route inbox, exactly as
it already maps the reply-log path, so every mention - bare path, listing,
and handled/ acknowledgement - lands host-local.

Both rewrites also become plain assignments, because a quoted substitution
nested inside a double-quoted printf argument leaks literal quotes into
the replacement text on stock macOS bash. The lifecycle suite pins the
corrected render both directions against the real seed, provisioning,
and delivery route, sharing one fixture value between the render truth
and the delivery truth.

Closes kunchenguid#5012

* no-mistakes(document): document remote charter's host-local steering inbox
)

* feat(procevent): route worker-owned Lavish rounds

* no-mistakes(review): drop duplicate artifact field from task-owned registration

* no-mistakes(review): post worker reply once, fix ring label, keep re-arm atomic

* no-mistakes(review): keep worker board owned until terminal round acknowledged

* no-mistakes(review): refuse every retirement of an open worker-owned round

* no-mistakes(review): use real lavish reply flag, isolate reply generations

* no-mistakes(review): drop .posted marker for best-effort reply posting

* no-mistakes(review): consume staged reply after listener setup, refuse orphaned captures

* no-mistakes(review): require a reachable owner, redeliver open rounds, roll back failed re-arms

* no-mistakes(review): re-arm only to acknowledge an open round

* no-mistakes(review): conclude only a still-open terminal round

* no-mistakes(review): record the acknowledgement before retiring the board

* no-mistakes(review): retain the registration across a conclude, qualify terminal docs

* no-mistakes(document): Document worker-owned Lavish round lifecycle
…unchenguid#5107)

* fix(bin): reserve contribution observation budget

* no-mistakes(review): Strengthen slow-read regression test to exceed the poll budget
…ness JSON (kunchenguid#5103)

* feat(bin): add idempotent inbox orders, receipts, replies, and readiness

Let a caller supply a request id when publishing a captain inbox note so a
retry returns the original note instead of creating a second one, including
across the crash window between save and wake announcement. Separate saved
from announced so a failed wake is repairable without enqueueing again.
Add bounded receipts JSON with omission disclosure, a durable primary reply
against a note id, and a read-only readiness projection that can say
unknown instead of inferring liveness from a lock file.

* no-mistakes(review): fix(bin): honest inbox announce, reply cursor, and readiness verdict

* fix(bin): resolve ready from lock-holder ancestry; drop lock status --json

Remove the extra JSON surface from fm-lock.sh so its human status still
always exits zero. Have the readiness projection classify the inspected
home from the lock-holder pid via fm-harness.sh ancestry, with an explicit
FM_SUPERVISION_MODEL still winning and an unknown model when there is no
holder. Prove the yes path when that ancestry names a known harness.

* no-mistakes(review): Harden inbox announce, receipts reads, and reply sequence cursor

* no-mistakes(document): Note read-only lock inspection in scripts inventory

* no-mistakes(lint): Pass missing id argument to malformed-reply test printf

---------

Co-authored-by: cliflacata-svg <304148223+cliflacata-svg@users.noreply.github.com>
…ending text (kunchenguid#5118)

* fix(composer): stop a harness footer row from reading as a composer holding text

A harness draws its own furniture below the composer - a user statusLine, a
permission-mode hint - and the cursorless "bottom-most shape wins" rule looks
exactly there. `→` (U+2192) is Cursor's prompt glyph but ordinary text
everywhere else, so a statusLine opening with `→` was selected as a bare
composer, swallowed the hint row beneath it as wrapped input, and answered
`pending` on a visibly empty pane. `fm_task_inbox_ring` defers on exactly that
verdict, and `bin/fm-watch.sh`'s re-ring calls the same function, so the first
doorbell and every retry were skipped and the worker never saw the steer.

Measured live on 2026-09-20: three of five Claude Code 2.1.236 worker panes on
Herdr 0.8.0 had genuinely empty composers and every one of them was refused.

A separator pair that closed over a bare agent-glyph row is a proven composer
container, so the contiguous non-blank rows below its closing rule are that
composer's footer and are no longer composer candidates. The demotion is bounded
by all three of its own preconditions: a blank row ends the zone, a pair that
closed over no glyph row demotes nothing, and a shape with no separator pair at
all (Cursor's half-block rules) is untouched. Real unsubmitted text in that same
composer, including a stray SGR mouse report left by a click in the pane, still
reads `pending`.

Pinned by two portable regressions and by a new cursorless arm on the live
composer-matrix guard, which re-reads each harness's already-proven-idle pane
the way every non-tmux backend reads it and fails naming the harness and
version when that read is `pending`.

* no-mistakes(review): make composer footer-zone demotion shape-independent

* no-mistakes(review): make footer-zone demotion refuse-only and drop rescan

* no-mistakes(lint): quote probe-absent sentinel to clear ShellCheck SC2100

---------

Co-authored-by: Koen Muller <koen@catapult.nl>
…5115)

Co-authored-by: guanchengh-lgtm <271917158+guanchengh-lgtm@users.noreply.github.com>
… an unreadable runs table (kunchenguid#5114)

* fix(bin): stop misreading a no-run branch as an unreadable runs table

Defect: when `no-mistakes axi status`'s overview is truncated (a task's
own branch has zero rows among the shown ones), fm_nm_select_run's
Python fallback derived the repo identity for its direct SQLite query
from a `repo: <path>` line it expected in the overview text. The real
CLI never emits that line, truncated or not (see the genuine capture at
tests/captures/no-mistakes-v1.70.1/overview.toon, which has only
`count:`/`runs[...]:`), so the lookup always failed and reported
"unreadable runs table" for a task that simply has no run on its
branch. On a fleet with many concurrent runs, every idle-branch task
hits the truncated-overview path routinely, so this fired every few
minutes and drowned genuine unreadable/blocked verdicts in noise.

Fix: derive the repo identity from the task worktree path instead,
which is exactly the value `no-mistakes` records as a repo's
`working_path` (confirmed against the existing capped-overview test
fixtures, which already register repos by worktree path). A worktree
path that is not absolute cannot be matched and still reads as
unreadable rather than being guessed at. Also raise the reader's
SQLite busy timeout from 1s to 30s so ordinary lock contention on a
busy fleet cannot masquerade as an unreadable database.

Safety: every other verdict byte-for-byte unchanged - the repo lookup
still requires exactly one matching row (a genuinely corrupt or
mismatched repos table still reports unreadable, per the existing
`repo` failure-mode test), the branch query and row validation are
untouched, and a zero-row result for the branch still flows through
the same recursive re-parse that already turns an empty `runs[0]{...}`
table into `absent`. Added a regression test
(test_capped_overview_without_repo_line_and_no_runs_reports_absent)
that reproduces the real overview shape - capped, zero rows for the
task's branch, no `repo: ` line - and asserts the crew state falls
through to the pane/busy verdict instead of reporting unknown or
"unreadable". Full fm-crew-state.test.sh suite passes unchanged
otherwise.

* fix: recovered same-branch inventory awk misreads empty result as unreadable

fm_nm_select_run's deep SQLite reader rebuilds a `count:`/`runs[...]:`
overview and re-runs it through the same awk selection pass. When that
rebuilt inventory has zero rows for the branch, the row-matching loop never
executes, so its counters (`seen`) stay at awk's uninitialized empty string
while `expected` and `shown` are plain strings parsed from the header text.
Comparing an uninitialized value against a non-numeric string uses string
comparison, so "" != "0" is true, and the END block takes the "unreadable
runs table" branch instead of falling through to the correct "absent"
verdict for a branch with genuinely zero runs.

Coerce the affected END comparisons with `+0` so they are always numeric,
matching seen/expected/shown/total regardless of whether awk classified
them as strings or numeric strings. A truncated or genuinely malformed
inventory still differs numerically and still reports unreadable.

* no-mistakes(review): bound capped-overview inventory reader and canonicalize worktree lookup

* no-mistakes(review): match recorded repo path first, tolerate duplicate spellings

* no-mistakes(review): revert repo lookup to exact working_path match

* no-mistakes(document): note state-db inventory read under crew-state nm timeout
…ort (kunchenguid#5141)

* fix(bin): require a non-draft pull request before a PR-based done report

A PR-based ship could report done, and merge monitoring could be armed, while the pull request was still a draft. A draft cannot be merged, so the poll waited for an event that could not occur and nobody was asked to merge.

The PR-based definitions of done now require reading the pull request back from the forge and confirming it is not a draft, and a lane that deliberately holds a draft declares a wait instead of done.
bin/fm-pr-check.sh refuses to arm merge monitoring on a draft, naming the draft state, and treats an unreadable draft state as before.
The draft reading now lives in bin/fm-pr-lib.sh and bin/fm-pr-merge.sh uses it, with its refusal to merge a draft unchanged.

Closes kunchenguid#4757

* fix(review): Skip arm-time draft refusal when fm-pr-merge records metadata
* fix(bin): accept quota-axi schema 6 snapshots keyed by provider + accountKey

quota-axi 0.1.47 emits schemaVersion 6 once a provider expands to more
than one account: every provider row carries an accountKey and one
provider id may appear on several rows. fm_quota_json_valid accepted
only schema 5 with unique provider ids, so fm-dispatch-resolve.sh,
fm-quota-choose.sh, and fm-procevent-quota.sh all rejected the live
snapshot and quota-informed dispatch was dead against the current tool.

- bin/fm-quota-axi-lib.sh: the validator accepts schema 6 with
  accountKey required on every row and uniqueness on
  provider + accountKey; schema 5 keeps its exact rules. FM_QUOTA_ROW_JQ
  is the one join every consumer uses: schema 5 binds by provider alone,
  schema 6 binds to the row keyed by the candidate's Pi lane, else the
  provider's default row, else no row (unmeasured, never blocked, never
  by position or summed across accounts).
- bin/fm-quota-choose.sh: accepts schema 6 JSON and the TOON accountKey
  column, and joins through the shared function.
- bin/fm-dispatch-resolve.sh and bin/fm-procevent-quota.sh: join through
  the shared function; an expanded provider with no row for the
  candidate's account is reported as such.
- tests: schema 6 fixtures shaped like the real snapshot, each paired
  with a schema 5 case on the same path; every new case fails on the
  previous scripts and passes now.
- docs: the two sentences naming the row join describe the schema 6 key.

* no-mistakes(review): Fix native Codex quota and expanded provider watches

* no-mistakes(review): Align native Codex account matching across dispatch paths

* no-mistakes(document): Align quota documentation with account-aware snapshots

* no-mistakes(document): Align quota dispatch documentation with account matching

* fix(bin): keep CI lint and the quota watch test portable

- bin/fm-quota-axi-lib.sh: FM_QUOTA_ROW_JQ is read only by the scripts
  that source this library, so full-mode ShellCheck reported SC2034 on
  the assignment; mark it alongside the existing SC2016 disable.
- tests/fm-procevent-quota.test.sh: the schema 6 provider-watch
  assertions used rg, which CI runners do not install, so the case
  failed with 'rg: command not found' rather than on behavior; use grep
  like the rest of the file.

* no-mistakes(document): Documented schema-version account-row compatibility
* test: repair Claude live auto-arm regression

* no-mistakes(review): Assert SessionStart digest completeness within its hook_response event

* no-mistakes(document): Consolidate Claude live verification references
Roll the shared require-no-mistakes action to the tagged v1.80.1 SHA and grant pull-requests: read so the check can read PR bodies.
BohnBawerick and others added 19 commits October 3, 2026 10:31
…e assertion to match nested path metadata output and added fm-codex-catalog-lib.sh to the synthetic remote-root fixture. Both affected test files pass locally. Bash syntax and git diff checks also pass
…es (kunchenguid#5863)

* fix(pi): silence unacknowledged processing retry replies

Suppress autonomous processing prose before persistence and during streaming while retaining tool calls, signed reasoning, and retryable outcomes. Restore ordinary output after acknowledgement or a user message.

Fixes kunchenguid#4954

* no-mistakes(review): Silence only processing retries, keep first presentation visible

* fix(pi): preserve differing processing retry replies
A forge read killed by its five-second cap was recorded as forge
unavailable, staled the observation, and rang a false supervision wake
whenever a healthy read ran slow beside its five parallel siblings.
Treat every deadline-cut read, budget deadline or per-read cap, as
unmeasured: the prior record is kept and the URL is observed first on
the next poll. A genuine forge failure still records the error and
wakes once per episode.

publish_pending hashed its wake key with shasum, which this Arch host
carries only under /usr/bin/core_perl; a watcher PATH without it minted
empty wake keys and broke wake dedup. Hash through shasum or sha256sum
and fail with a named error when neither exists. The same sha256sum
fallback is folded into the pending-reply correlation-id fallback.

Regression tests cover both: a slow parallel read keeps the prior
record without a wake, and a PATH without core_perl still publishes a
64-hex wake key exactly once.
… test for Pi 1.0's renderer API while retaining older compatibility, and increased two test-only process startup windows to avoid loaded-runner flakes. All three focused tests pass; timing-sensitive tests passed five repeated runs each. ShellCheck, actionlint, Bash syntax, and git diff checks pass
Its content landed by squash in #29 (71bd89c), whose tree equals the PR head 0ef444c, a real merge of c576c2b.
Recording the ancestry lets later upstream merges use c576c2b as the base.
Merges upstream/main (156 commits after c576c2b) with a real merge commit.
Keeps the fork's curated memory, quality posture, local landing, agy adapter,
session-lock gates, report-once turn-end decline, catalog-aware Codex max
effort, Lavish list-form parsing, remote secondmates, and the
worker-starts-its-own-validation contract, and adopts upstream's forge and
ship-branch-prefix support, Devin adapter, supervision host, skill-routed
AGENTS.md, secondmate liveness library, and contribution poll fairness.
The fork's fleet-mutation gate refuses a drain from a session that does not
hold the lock, so the host fixture records its session beside the lock and
main-side drains and acknowledgements declare that same session.
Also rename the turn-end guard doc's away-daemon heading to the anchor other
sections link to.
… sync

Keeps the fix's behavior: a forge read cut by the budget or the per-read cap is
unmeasured, never unavailable. Upstream carries the same rule and additionally
moves on to the next URL, which is the form kept. Two upstream contribution
tests are adapted to the fork's rules that only a merged contribution is final
and that a settled owner's unnotified pending signal replays once.
…am merge

The lone-separator rule now spares a bare glyph only inside its titled rules
or when its opening rule was clipped off the top of the capture, so the fork's
clipped idle Claude and upstream's mismatched titled rule both read correctly.
The watcher retraction test uses a verdict that is not provably working,
because a provably working crew suppresses a wedge escalation.
…efix

Launch-parsing helpers strip the session-identity unset statement before
evaluating the command, so they split the agent command and never run it.
Teardown names the session-lock library in its required-source check.
…ly listener

fm-control-relaunch: the Herdr exited-agent fixture reports the pane shell as
the one foreground process and answers the stat-bearing ps read, the shape
the fork's departed-Pi proof requires.
fm-supervision-host: main's drain helper runs as the lock-holding session, so
the session-lock gate lets it acknowledge the watcher downtime.
fm-remote-reply: the helper-report section awaits its capture through the
kept-alive listener instead of waiting for start to return.
…iers

A Claude home now runs the supervision host by default and launches no away
daemon, so the case opts its home out to keep pinning the daemon's Herdr
topology. A Claude primary receives the digest as a durable record plus a
doorbell line, so the case reads the record the submitted doorbell names.
…riable to avoid ShellCheck SC2031 false positives across both launch paths. Increased the shared supervision-host registration-case deadline from 25 to 60 seconds because CI completed successfully after about 34 seconds under load. Verified with repository lint, both full affected behavior suites, Bash syntax checks, and git diff checks
…25-second supervision-host wait to 60 seconds. Product behavior and deadlines are unchanged. The full 65-case supervision-host suite, repository lint, Bash syntax, and git diff checks pass
…s now accept the documented bounded lock refusal, retry after contention clears, and retain all endpoint, husk-removal, focus, and teardown assertions. The full real-Herdr suite, repository lint, Bash syntax, and git diff checks pass
…ture clock for both generated-check budget paths, removing the one-second boundary race. The full contribution suite passed in 151.5s, the 65-case supervision-host suite passed in 719.7s, repository lint passed, and git diff checks passed. Serial 2 needs no separate code change: origin/main's latest green run predates this test, while upstream green runs varied from 737.3s to 1527.6s; the local merged suite matched the faster upstream run, so the 1792.8s cancellation was runner-load variance rather than a merge-caused hang
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.