feat: sync upstream supervision and fleet capabilities - #31
Merged
Merged
Conversation
* fix(bin): derive passed PR state from PR record A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists. For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim. Fixes kunchenguid#4607 * no-mistakes(review): Add bounded GitLab merge-request state reads * no-mistakes(review): Preserve network-free inactive crew-state scans * no-mistakes(document): Document PR record readers in shared library
kunchenguid#4627) * fix: restore published contribution follow-up (Fixes kunchenguid#4469) * fix(review): Fix contribution freshness and merge actor routing * fix(review): Restore issue triage and scope contribution follow-up * fix(test): test: assert one wake per contribution signal * fix(document): Document contribution follow-up * fix: restore truthful terminal delivery evidence * fix(review): Disclose unsupported contributions and deduplicate watcher wakes * fix(review): Preserve unmeasured unsupported contributions across Bearings * fix(review): Deduplicate shared contribution wakes and isolate diagnostics * fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower
…nguid#4658) * fix(bin): make a remote-reply document gap self-clearing and re-attemptable A remote mate's undelivered document raised a keyed `blocked` decision that nothing could ever resolve, and any `data/*.md` substring in any mirrored line was an unconditional fetch instruction. A mate announcing a report it had not written yet therefore manufactured a permanent, factually false blocker, and its own explanation of the false alarm manufactured more. The reader has no permanence vocabulary: a report still being written refuses exactly like a path that will never exist. So an undelivered document is now a durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`, re-attempted on the next delta and on the channel's own quiet poll, and retired with a matching `resolved` line naming the local copy once it arrives. The cursor still advances and no delta stalls on one bad pointer. Only a structured `report=data/....md` pointer now offers a document, so a path merely mentioned in prose - including one under another home's mirror tree, which is provably not that mate's to serve - is never fetched. Offers are deduplicated across the whole delta, the escalation names each missing document once and carries the reader's own reason instead of discarding it, and a strictly increasing notice ordinal keeps a later escalation from being swallowed as duplicate bytes. A mirrored line still lands once whichever pointer form it was first written under. * no-mistakes(review): Require structured pointer token boundaries * no-mistakes(review): Unify boundary-safe pointer extraction and rewriting * fix(bin): identify a mirrored line independently of its delivery state Two defects in the boundary-safe pointer work. The at-most-once check compared only the all-remote and all-local renderings of a line, so it could not recognize a mixed one. A line offering two documents where only the first was deliverable mirrored as local-plus-remote; once the second arrived, a cursor-loss whole-log recapture rendered the same line all-local, matched neither alternate, and mirrored a second time. A line's identity is now the canonical form every boundary-valid pointer would take once delivered, derived by the same parser that does extraction and rewriting, so it no longer depends on which documents happened to be deliverable at the time. The pointer map was passed to awk through the process environment. A delta may carry up to the configured 1 MiB bound, and an expanded map of delivered pointers can exceed the platform's exec argument limit, so awk would fail to start; because no caller checked, the empty result would have been appended as blank lines while the cursor advanced past dropped status content. The map now travels in a file, and every call site checks the exit status and stops the ingest rather than committing a delta it could not render. Both passes now run once per stream instead of twice per line. * no-mistakes(review): Abort ingest when document pointer extraction fails * no-mistakes(review): Exclude structured cross-home pointers from document transfer * fix(bin): fail open on an undeliverable remote document instead of tracking it Narrow the remote-reply document fix to the scope the diagnosis actually requires, as decided after measuring a simpler alternative. A document the reader cannot deliver now fails open. The mate's line is mirrored with its own pointer, the cursor advances, and one unkeyed note carries the reader's reason. A note never enters the open-decision fold, so it cannot stand open the way the original keyed block did - which removes the never-clearing false blocker by construction rather than by resolving it. That makes the durable self-clearing obligation unnecessary, so it goes: the per-mate pending-documents record, its notice ordinal and resolved announcements, and the poll-side retry. Canonical line identity goes too, and with it a way to silently drop a genuine status line; mirroring is back to at-most-once on exact bytes. The cross-home exclusion goes as well: under fail-open a cross-home report= either fails harmlessly or is a nested remote report this mate genuinely holds, which is now relayed again. Kept: fetching only on a structured report= pointer, the boundary-correct parser, the file-based rewrite map, and checked extraction and rewrite exit status. The parser now scans behind a sentinel byte so a rejected candidate can no longer give the text right after it a false leading boundary. The reported incident is covered end to end: a report path announced in prose before it exists raises no decision, and the report still arrives through the ledger publisher's structured offer once written. * no-mistakes(review): Preserve source-line identity across remote reply replays * no-mistakes(document): Document remote reply transfer and replay semantics * no-mistakes(lint): Fix staging truncation lint checks
* Preserve substantive Calm mid-turn text * no-mistakes(review): Distinguish newline-preserved replies from short narration * no-mistakes(document): Document Calm mid-turn preservation boundaries * no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`
…#4656) * fix(bin): re-record PR poll identity after a volume device renumber (Fixes kunchenguid#4260) A volume remount can renumber the state filesystem's st_dev while every inode and byte stays the same; APFS does this across a reboot. A poll registration records its sidecar and check as device:inode, so every poll armed before the remount failed strict validation and the watcher refused all of them as unauthenticated state checks until each was re-armed by hand. There are two device comparisons. fm_pr_private_file_valid compares a live file's device with the state directory's device read in the same invocation: it refuses a file that is not on the state directory's own filesystem and already survives a renumber, so it is unchanged. The registration's recorded identity versus the live identity (from kunchenguid#556, reused by the kunchenguid#932 retirement receipt) binds the registration to the exact files published in its own transaction; its device part is what breaks. When strict capture fails, the watcher now proves the device is the only difference: every other artifact check passes (template bytes, both hashes, private mode, single link, live device, metadata), both recorded identities name one device, and each recorded inode equals its live inode. Only then, under the task's control lock, does it rewrite the two identity lines, repeating the whole proof and comparing the registration's file identity and bytes just before the rename, and then capture strictly again. A swapped, altered, re-moded, relinked, split-device, or foreign-device artifact still fails a proof and is still refused, and a pending retirement receipt blocks the rewrite. Reproduction: on macOS a poll armed on an APFS disk image that was detached and re-attached behind another image moved st_dev 16777239 -> 16777243 with inodes, bytes, mode, and link count unchanged; the real watcher refused it on main and reports its merge with this change. The portable regression test rewrites a real registration's recorded device and drives the watcher. Not changed here: the status presentation cursor keys rows by its own device:inode identity in bin/fm-classify-lib.sh, a different helper that needs its own fix; a retirement receipt left by a reboot between its publication and removal still names the old device and stays refused; custom check trust binds only a content hash and is unaffected. * fix(review): Serialize PR poll publication writers * fix(review): Bound PR poll publication lock scope
…llow-up to kunchenguid#4627) (kunchenguid#4661) A budget that expires partway through an observation no longer records an error or prints the unavailable wake; the URL keeps its prior record and is observed first next poll. forge() flags budget exhaustion at the point it refuses, or when a read is killed at the budget's own deadline, so a genuine forge failure still records the error and wakes. Each distinct URL is now observed once per poll and applied to every owning task.
…kunchenguid#4680) * fix(bin): clear parent pending-replies on local secondmate retirement Local secondmate teardown left resolved parent pending-reply records behind after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced retirement while any reply for that id is still unresolved, and delete every matching record plus its delivery confirmation after a successful local or remote retirement, matching the remote cleanup path. * no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup * no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung * no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern * no-mistakes(review): Pending-replies Basename und corr_id abgleichen * no-mistakes(document): Clarify forced retirement pending-reply cleanup --------- Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>
kunchenguid#4677) * fix(bin): accept Orca's composite worktree id at teardown Teardown refused every Orca-backed task because the endpoint validator checked orca_worktree_id with the simple-atom rule meant for tmux-style window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca returns that id as `<orca id>::<absolute worktree path>`, so the colon and slashes in every real value made validation fail and finished Orca tasks could never be cleaned up. Validate the field as the composite it is: both halves of the first `::` split present, the path half absolute, and no embedded newline, carriage return, or tab. The terminal field keeps the atom check, which is correct for it, and no other backend's validation changes. The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca never returns, which is why the suite passed a check the real value fails. They now carry the composite form, so the tests exercise the real value. * no-mistakes(document): name Orca's repo id in the composite worktree id * no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points
* feat(bin): add opt-in typed dispatch resolution through typesafe.ai Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or scout profile from a written brief with typesafe.ai's System One model: one Choice question over the rules' `when` texts, then the confidence floor, the rule's `approval` and `floor`, each profile's `provider` and `floor`, one quota-axi snapshot, and the spendPriority argmax all in code. It is off unless TYPESAFE_API_KEY is in the environment or the home's gitignored .env; off means one stderr line, exit 0, and no network call, so firstmate dispatches exactly as before. The key reaches curl on a file descriptor, never argv. Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates the four new optional dispatch fields. Document the schema, the operator contract, the AGENTS.md intake step, and the live and benchmark evidence. * no-mistakes(review): Harden typed dispatch resolution and quota bounds * no-mistakes(review): Validate dispatch floors and ranking evidence * no-mistakes(review): Tighten dispatch response and floor evidence * no-mistakes(review): Neutralize none matching and resolve defaults locally * no-mistakes(review): Preserve providerless profiles outside typed resolution * no-mistakes(review): Validate response usage and reject duplicate profiles * no-mistakes(review): Escalate unverifiable floors and validate probabilities * no-mistakes(review): Validate probability mass and unknown profile floors * no-mistakes(review): Simplify resolver interface and preserve fallback routing * no-mistakes(review): Fix constants and rank partial quota evidence * no-mistakes(review): Add authoritative provider mapping and enforce explicit providers * no-mistakes(review): Declare provider for documented Pi profile * no-mistakes(review): Validate provider identifiers and support Gemini dispatch * no-mistakes(review): Strictly anchor provider identifiers * no-mistakes(review): Validate selectors and preserve fallback candidate evidence * no-mistakes(review): Gate typed validation and harden resolver evidence * no-mistakes(review): Preserve opt-in routing and harden candidate evidence * no-mistakes(review): Prioritize known exhaustion over quota uncertainty * no-mistakes(review): Isolate API secrets and preserve no-key diagnostics * no-mistakes(review): Fallback safely when dispatch rules are absent * no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets * no-mistakes(document): Document typed dispatch safety and fallback behavior
…n decisions aren't lost (kunchenguid#3753) * test: reproduce buried status declarations in shared readers * fix: share status event reads and preserve open blockers * fix: retain terminal scout and ship status declarations * no-mistakes(review): Fix status chronology, legacy completions, and reader performance * no-mistakes(review): Share terminal decision reconciliation across fleet snapshots * no-mistakes(review): Unify terminal supersession across cached folds and consumers * no-mistakes(review): Filter per-key status history while preserving terminal chronology * no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells * no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses * no-mistakes(document): Document latest-event status read and kind-scoped fold cursor * no-mistakes(lint): Quote literal done in test for-lists for SC1010 * ci: expect 19 snapshot/fleet-view tests This branch adds a fleet-snapshot regression, so the stock macOS Bash lane's hardcoded guard of 18 'ok - ' lines fails on the new count. Bump the guard and its message to 19. * no-mistakes(review): Restore multiline child outcome reporting * no-mistakes(review): Select ledger terminal events through bounded shared reader * no-mistakes(review): Report newest open decision instead of preferring blocked * no-mistakes(review): Require colon before ship/scout terminal supersession in fold * no-mistakes(review): Gate socket-down override on latest event; drop lock matrix * no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions * no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold * no-mistakes(test): Update fleet-view expectations to newest-open-decision rule * no-mistakes(document): Align status-read docs with fold-resolved crew state * no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers * no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree * test: fold terminal-cleanup snapshot coverage into the completed-scout case Keep the ship/scout/secondmate supersession assertions without adding a nineteenth top-level fleet-view test, so CI can stay at the upstream suite count. * no-mistakes(document): Clarify socket-down override expiry in architecture doc * ci: retrigger flaky contribution check
…nchenguid#4689) * fix(spawn): launch codex crewmates with codex's hook layer disabled A freshly launched Codex worker never reached its instructions. Codex stopped it on an interactive "Hooks need review" modal whose selection sits on "Review hooks", which is neither trusting nor declining. Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow navigation, so the selection cannot be moved, and pre-accepting the prompt by writing Codex's own trust store would record an operator consent that was never given. The hooks are the machine's own ~/.codex/hooks.json plus any project's .codex/hooks.json. A crewmate needs neither: its turn-end signal is the -c notify= program on the same launch, and Firstmate's project hooks are primary-session infrastructure that stands down in a child worktree. Crewmate and scout launches now pass --disable hooks. That is the opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted hooks; disabling the feature runs none of them and leaves the operator's ~/.codex untouched. An unknown feature name is a hard Codex error, so a release that drops the flag fails the launch loudly instead of silently restoring the modal. A secondmate is a primary in its own home and keeps the project hooks its turn-end guard and session-start digest ride on. Verified on codex-cli 0.151.0: the modal is gone and the turn-end notification still lands. This unblocks the second review that every finished pull request is supposed to get. Fixes kunchenguid#4673 * no-mistakes(review): Fix contradictory hook count in Codex verification record
…d#4669, Fixes kunchenguid#4670) (kunchenguid#4710) * fix(bin): settle terminal contributions and wake once per read-failure episode A contribution whose last good observation is merged or closed is final: poll no longer re-reads it, projection keeps it fresh, and a stale error recorded beside it is cleared once. A genuine forge-read failure on an open contribution still records its error on every cycle but prints the unavailable wake only when it starts a failure episode; a successful read ends the episode. Open PRs linked from done tasks keep being observed. The false unavailable beside a complete observation was budget exhaustion mid-observation, already fixed by kunchenguid#4661. * fix(review): Settle terminal contribution owners * fix(review): Deduplicate shared contribution failure episodes * fix(test): Preserve settled terminal contribution records
* fix(crew-state): select authoritative validation runs by identity Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row. Refs: kunchenguid#3215 * fix(review): Resolve same-branch run identities beyond capped history * fix(review): Fix run-selection compatibility, races, and worker-state fallbacks * fix(review): Limit run validation to the requested branch * fix(test): Anchor AXI fixtures and document remaining live evidence gaps * fix(document): Clarify run selection documentation and capture ownership * fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation
* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work MAIN answered a supervision-branch outcome for completed captain-requested work (implementation done, PR ready for review and merge approval) with "Captain, shipshape.", reading section 9's no-action reply as covering it and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no captain-facing response is owed". Section 9 now limits the shipshape reply to true no-ops (idle re-read, empty heartbeat, consequence-free acknowledgement) and requires a short outcome response naming what finished and what word is needed whenever requested work finishes or a result needs the captain's word, even when a transcript entry already shows the substance. The Pi protocol's re-emit rule now says it bounds repetition only, and carries a worked example of the ready-for-review outcome whose correct processing turn a shipshape reply fails. No executable contract evaluates the content of MAIN's captain-facing reply, so the regression is the protocol example in the owner doc rather than a text-match test. * no-mistakes(document): Clarify captain-facing outcomes versus no-ops * docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line The document step condensed the Pi protocol's re-emit rule and dropped the worked example of a finished, ready-for-review outcome whose correct processing turn a "Captain, shipshape." reply fails. That example is the contract's regression: no executable contract evaluates the content of MAIN's captain-facing reply, so the owner doc's example is the test case. Restore it directly under the re-emit rule, prefixed as a regression example that is kept verbatim and never condensed or summarized away. * no-mistakes(review): Clarify captain outcome and decision-word requirements * no-mistakes(document): Clarify captain-facing completion outcomes * docs(pi): require the PR URL in the visible captain-facing outcome reply Captain review on the regression example: drop the sample reply string and say only that the ready-for-review outcome requires relaying a captain-facing outcome response, not just "Captain, shipshape.". Fold in the visible-PR-handoff failure seen this session: after the branch outcome reporting this fix green, MAIN's visible reply was only "Awaiting your merge call." with no PR URL, leaning on the dim anchor. Section 9's URL rule now also covers a review or merge ask and names the visible reply as where the URL goes, sourced from the ready status, pr= metadata, or the supervision branch's summary and never left to a transcript entry. The Pi protocol adds the same-way failure and places the captain-facing text in the final visible assistant reply after the fm_branch_processed call, because Calm hides assistant text emitted in the same step as a tool call as a working note. Investigation verdict, evidence in the PR comment: no recent PR caused the handoff failure; Pi has hidden same-step pre-tool assistant text since kunchenguid#2339 (2026-08-13), kunchenguid#4655 changed only the Claude Code mod, and kunchenguid#4658 touched only remote report transfer. * no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs * no-mistakes(document): Clarify captain-facing supervision outcomes * docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule The consolidated section 9 URL rule narrowed its trigger to a review or merge ask, dropping the "whenever a PR is mentioned" catch-all from kunchenguid#3648 that keeps every PR URL copied from a durable record and never assembled from memory. Restore that trigger as a union with the review or merge ask so the one consolidated rule covers both.
* Fix foreign-owner turn-end supervision loop * no-mistakes(review): Scope foreign-owner safe exit to Claude guard * no-mistakes(document): Document Claude foreign-owner safe exit
…orb (kunchenguid#4778) Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty indexed array as an unbound variable and aborts the shell. In signal_turnend_panes_churned() the missing_keys loop was reachable with an empty array whenever every churned key already held a fresh .churn-since-* marker (a second churning turn-end inside an open deferral window), so each watcher cycle died about half a minute in and supervision restarted endlessly. The created_keys rollback loops had the same latent crash on their error paths. Audit of bin/ for the same pattern found one more confirmed-reachable case: remote_handoff's noncanonical-body scan iterates to_move, which is empty when a retried remote handoff finds every key already staged in the outbox. All other "${arr[@]}" sites are either count-guarded, guaranteed non-empty by construction, or unreachable while empty. Guard the three reachable expansions with the repo's existing "${arr[@]+...}" idiom. New regression test drives a real watcher through the all-marked churn path; the macos-stock-bash CI lane runs it under real /bin/bash 3.2 via FM_TEST_ONLY.
… lock. (kunchenguid#4783) The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it. Co-authored-by: Cursor <cursoragent@cursor.com>
* docs: require complete final responses across harnesses * no-mistakes(document): Document complete final replies for Grok Bot * docs: point Grok replies to the shared contract owner * no-mistakes(review): Clarify final recap without batching decision asks
* fix(calm): preserve substantive Pi mid-turn text * no-mistakes(review): Preserve substantive Pi Calm text per block * no-mistakes(test): Cover shared Calm preservation boundaries behaviorally * no-mistakes(document): Consolidate Calm preservation documentation
…guid#4799) * Handle Kimi workspace trust dialog * no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers * no-mistakes(review): Gate Kimi ready on any trust marker and clean captures * no-mistakes(review): Read visible pane for Kimi trust and ready gates * no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate * no-mistakes(review): Harden Kimi viewport capture and trust dialog detection * no-mistakes(document): Document Kimi spawn refusal on cmux and Orca
…er (kunchenguid#4775) * fix(bin): report a record whose agent is gone once instead of escalating forever The wedge escalation path never asked whether there was still an agent to be wedged. A wedge is something stuck that might recover, so re-alarming it earns its cost; an agent that is gone never moves again, its pane never churns, the idle timer never resets, and the escalate path clears its own timer and re-arms with nothing bounding the count. Observed on a live fleet: two finished lanes reached 226 and 203 consecutive escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400 notifications a day from two lanes with no agent running at all. On one, fm-control.sh exit answered already-stopped and fm-crew-state.sh read "failed - run failed". Closing the Herdr pane did not stop it either: with the pane genuinely gone and herdr pane read returning pane_not_found, the count kept climbing, because the poll is driven by the record's window= line rather than by the pane. The cost is not the repetition but that it drowns the alarms that matter. fm_backend_agent_state already separates a thinking agent from a gone one at process level. In the branch that was about to escalate, read it once and treat only its two recovery-grade verdicts - dead (endpoint present, no agent in it) and missing (endpoint authoritatively absent) - as proof, reporting that record once and not re-escalating it while it stays that way. Every other verdict, including alive, ambiguous, unreadable, unverified, and a read that failed outright, keeps the identical schedule, reason, and escalation count, so a genuinely wedged live agent is unaffected. The probe costs at most one backend read per window per threshold, the same budget the declared-wait consult and the worktree write probe already take. The report decides nothing about the record's fate: both lanes still held unlanded work and teardown refusing them was correct, so retiring, relaunching, or cleaning up stays with the supervisor. The once-only marker is owned entirely by that function and is dropped by the same read the moment the endpoint stops reading gone, so a replacement launched into the same window escalates normally and its own later death is reported again. Related, and not closed by this: kunchenguid#4412, kunchenguid#4482, kunchenguid#4316. Tests drive the real watcher against a record whose endpoint does not exist and pin both directions: dead and missing report once and never advance the count across later thresholds, while alive, ambiguous, and unreadable endpoints keep escalating with the identical reason and a climbing count. * fix(bin): bind the once-only dead report to the pane it reported Review of the parent commit found a reachable sequence where a later death in the same window lost its promised report. The marker was keyed on the verdict string alone and dropped only when a threshold probe read a non-gone verdict, but probes run only at thresholds: a replacement launched into the same window that dies without ever being probed alive - it crashes at startup, or works and then crashes - was absorbed by the previous death's marker. The pane's first sight yielded only the generic stale wake and every later threshold matched the stale marker, so the second death never got the detailed once-report that both the function's own comment and docs/architecture.md promise. Record the verdict together with the pane hash it was reported for, and absorb a repeat only while both still match. A replacement churns the pane, which resets the stale suppressor, wedge timer, and escalation count while no reset site touches this marker, so the pane half is what tells the second death apart from the first. The live-probe drop stays as it was. Clearing the marker at those reset sites instead would re-open unbounded re-alarming for a dead pane whose display ever ticks, which is the exact defect the parent commit exists to close. The noise bound is unchanged: an unchanged dead pane still absorbs on every later threshold and never advances the escalation count, and every verdict short of proof still escalates exactly as before. * no-mistakes(review): Key the dead-record once-marker on the busy incarnation token * no-mistakes(document): Document dead-record escalation cap in stale-pane config entry * no-mistakes(document): Add busy-state inventory line to AGENTS.md * no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path
…id#4854) Captain holds have no due semantics and are a hold kind, not a Beads issue type. The create path now waives due.required and maps to native type task. Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(bin): launch every spawned agent with the compact adviser disabled Every crewmate, scout, and secondmate Firstmate launches now starts with COMPACT_ADVISER_DISABLE=1, on a fresh spawn and on a relaunch alike, so an unattended session never activates the compact adviser. The value is unconditional: no configuration file gates it and there is no override, unlike the trace carrier beside it. Three carriers deliver it, because no single one covers every launch shape. The pane shell receives an export beside GOTMPDIR, so the agent's own children inherit it too. The launch command carries an explicit assignment, prepended outermost so it wins over any ambient value the pane already held. The cleared launch environment sets it again at the `env -i` boundary and keeps COMPACT_ADVISER_DISABLE in the fixed operational floor, which is what preserves the switch when config/launch-env-allowlist empties the environment, and what delivers it on a remote host that never had the value. bin/fm-control.sh relaunch, the bootstrap secondmate relaunch, and the remote secondmate transport all rebuild their launch through bin/fm-spawn.sh, so they inherit the same floor. The captain's own primary session is untouched. The two new suites drive the real spawn and then execute the launch command the pane actually received, with the harness replaced by a probe that prints its own environment, rather than matching script text. They cover ship and secondmate launches with the allowlist absent and enabled, the pane export and its ordering, fm-control.sh relaunch, and the full parent to remote-host chain. * no-mistakes(review): Export compact-adviser disable across compound launches * no-mistakes(document): Document spawned-agent compact-adviser environment guarantee
…henguid#4894) * fix(bin): let a background Claude session keep owning its session lock Session-lock ownership was decided by process ancestry alone. Under an unattended Claude session the model loop runs in a transient bg-spare bridged to the front-end by a shared daemon; when that bridge is recycled the contiguous claude-named ancestry from a hook to the recorded owner breaks while the owner pid stays alive, so the Stop auto-arm stood down as a foreign live owner, the turn-end guard ended every turn with its read-only diagnostic, and fm-lock.sh refused - a self-sustaining outage until restart. Ownership is now ancestry membership OR a trusted same-session id, never id-first: - fm-session-lock-lib.sh accepts CLAUDE_CODE_SESSION_ID only when CLAUDE_PID is a Claude-shaped member of the current contiguous run, compares it against the id recorded in state/.lock-session, and requires the recorded pid to still be a live harness. No id, no sidecar, an untrusted id, a different id, or a dead recorded pid leaves the ancestry verdict unchanged. Ids are never read from ps argv. - fm-lock.sh accepts a same-session holder at both refusal sites, writes, refreshes, and clears the sidecar only under its claim lock (including the early already-mine exit, skipped only while the deferred startup sweep leases that lock), keeps it byte-identical across a same-session confirmation, records CLAUDE_PID on lock line 1 for a session with a trusted id so a shared daemon or front-end that outlives the session never keeps a dead session's lock alive, never rewrites a live line 1 on a same-session confirmation, and names the recorded id in the live-owner refusal. - The .lock line-1 format is unchanged, so every reader that takes the whole first line as the pid keeps working; the guard's foreign-owner exit is unchanged and inherits the fix through the shared predicate. Tests: the ancestry suite drives the ancestry and id signals apart in a deterministic process table (asserting the divergence) and runs a real orphaned front-end/daemon/pty-host/spare tree through six phases with the real lock, auto-arm, and guard scripts; the foreign-owner repro keeps its negative control and adds a same-id positive control. Disclosure: no live unattended Claude background session ran on the verifying machine. The topology is documented by the real process listings in kunchenguid#3902, kunchenguid#2314, kunchenguid#3398, and kunchenguid#4066; coverage is the structural predicate plus the executable fixtures, not a live pass. Residual: bin/fm-sessionstart-nudge.sh keeps its own private ancestry walk (it only decides whether to print a nudge) and may nudge on a resume in the recycled case. Out of scope, deliberately: no structured lock format, no guard budget changes, no daemon-identity rejection, no fork lineage. * no-mistakes(review): Wait for claim lock; revert failed sidecars * no-mistakes(review): Revalidate ownership after wait; restore sidecars * no-mistakes(review): Roll back sidecar by publication phase * no-mistakes(review): Restore sidecar only if lock line is unchanged * no-mistakes(review): Trust session ids without a spelling allowlist * no-mistakes(review): Disarm sidecar rollback before backup cleanup * no-mistakes(document): Updated session-lock ownership documentation
* feat: park main under the away posture on Pi While the away-posture record exists on a Pi primary, the supervision branch takes every actionable wake, no processing turn opens on main, captain rows accumulate for the return brief, and main's standing authority relocates to the branch through the existing guarded scripts. - lib/fm-branch-dispatch.ts: read the record at every routing decision; while it exists claim check, decision-owned, and heartbeat rows too, keeping the two broken-queue vetoes; expose checkSeqs so a claimed check row lifts task scoping. - fm-primary-pi-watch.ts: offer every actionable row under the record; a declined wake and every watcher-failure alarm still reach main. - fm-branch-supervision.ts: drop the legacy .afk decline; append a fixed POSTURE: AWAY tail carrying the record's read-back verbatim per wake; open no processing request while the record exists, re-checked immediately before a request would open and at every run boundary; present the accumulated rows at the first run boundary after archive. - fm-lease-lib.sh: fm_lease_forbid_branch passes the branch for opted-in actions only while fm-afk-contract.sh validate succeeds on a confirmed live record; PR merge, fresh spawn, and decision answer opt in, local landing never does. - fm-send.sh: a --resolve-key naming an open needs-decision or captain-held task is a decision answer and meets the partition; blocked: keys stay steering. - fm-spawn.sh: enforce the record's spend cap for a fresh ordinary spawn by either actor; relaunches and secondmates exempt. - fm-branch-prompt.sh: fixed Postures section and the verbatim ask-user-authority policy; the prefix stays byte-stable. - fm-afk-return.sh: count what the away session handled from the store. - docs, afk skill, AGENTS.md stub: main parked on Pi, green merge gate absolute while away. - tests: watcher and branch extension suites, fleet-record, merge, and decision-answer suites cover the relocation, the vetoes, the tail, the parked processing turn, the cancellation, the re-presentation, and the spend cap; dated live-guard evidence recorded. * no-mistakes(review): Refuse branch merge after preflight archive race * no-mistakes(review): Fix away wake, spawn, and processing races * no-mistakes(review): Suppress parked processing; narrow away-only rejection * no-mistakes(review): Abort dedicated processing; gate branch spawn once * no-mistakes(review): Stamp away-only on the dispatch offer * no-mistakes(review): Treat invalid away records as spend-cap absence * no-mistakes(review): Drop spawn test hook; abort processing-opened runs * no-mistakes(review): Bind abort to opening prompt; cap-read absence * no-mistakes(review): Limit away branch spawn to queued work only * no-mistakes(document): Correct AFK posture documentation
* ci: simplify CI job timeouts to a three-tier policy Replace the scattered per-job timeout values (10m parallel, 25m lint, 30m serial, 10m macOS) with three readable tiers, each a hang tripwire with headroom rather than a packing estimate: - fast (5m): coverage guard, repo invariants, timing aggregate - normal (30m, one shared budget): lint partitions, portable parallel shards, portable serial shards, macOS stock Bash - heavy (Herdr only): 20m step tripwire on the family run so always() cleanup still runs, under a 75m job-level last-resort backstop The workflow's header comment states the policy and points at docs/fm-test-portable-shards.md "Timeouts", which now owns it, and each job names its tier beside timeout-minutes. tests/fm-ci-workflow.test.sh asserts the policy against the parsed workflow instead of the old per-job minute values: every job joins exactly one tier, exactly three distinct job-level values exist, the fast tier stays within 5-10 minutes, the normal budget stays at least double the modeled parallel lane sum reported by fm-test-run.sh --check-coverage, and the Herdr step tripwire stays below its job backstop with an always() cleanup after it. Concurrency supersession, shard counts, lane membership, and fail-fast settings are unchanged. * no-mistakes(review): Decouple the normal timeout from packing estimates * no-mistakes(review): Assert Herdr teardown follows the family run * no-mistakes(review): Pin Herdr family-run timeout to 20 minutes * no-mistakes(review): Ignore comments when identifying Herdr steps * no-mistakes(review): Identify Herdr steps by declarative ids * no-mistakes(document): Clarify authoritative three-tier timeout policy
…nchenguid#4895) * fix(bin): keep supervisor status closes from waking the same home A drain that already folded OPEN DECISIONS has presented those bytes even when the watcher has no matching seen marker. Treat that fold, and the presentation cursor, as known so the bookkeeping close stays quiet while later worker lines still signal. * no-mistakes(review): Keep folded worker failures waking past supervisor closes * no-mistakes(review): Wake on unlisted folded worker lines; batch multi-key closes * no-mistakes(review): Stop folded worker resolved lines from counting as already read * no-mistakes(document): Correct self-announced close marker contract in docs
* Stop steering operators away from Herdr * no-mistakes(review): Neutralize remaining Herdr opt-out documentation wording
…enguid#4973) * fix(bin): treat a live no-mistakes run as current after rebase A running run on the task's branch is authoritative regardless of head. Matching only the local head made a rebased in-flight run look failed. * no-mistakes(review): restrict coarse live-any-head to foreign-branch answers * no-mistakes(review): reject gate-parked runs from the executing predicate * no-mistakes(review): hoist gate-marker patterns into single run-lib owner * no-mistakes(review): require live daemon for head-free run binding * no-mistakes(review): require answered daemon-down before unbinding live runs * no-mistakes(review): extend daemon guard to anchored continuation routes * no-mistakes(review): delete live-any-head; restore dead-daemon verdict * no-mistakes(review): keep parked gates parked; name dead daemon everywhere * no-mistakes(review): set dead-daemon verdict instead of emitting early * no-mistakes(review): align selected route with legacy dead-daemon handling * no-mistakes(review): drop unproven-record binds; narrow coarse gate reading * no-mistakes(review): narrow header, drop vestigial guard, retarget tests * no-mistakes(review): revert coarse gate override; require answered-down probe * no-mistakes(review): cache one daemon probe; stop duplicating run id * no-mistakes(review): restrict coarse dead-daemon verdict to moved-off rows * no-mistakes(review): delete coarse dead-daemon extension and gate note * no-mistakes(review): delete remaining coarse dead-daemon block and stale docs * no-mistakes(document): document rebase-safe live-run bind and unverified-record verdict
…chenguid#5928) * feat: make /quiet a statement where the attended supervision host runs On a home that opted into the supervision host, quiet mode is what the attended host already does, so /quiet now enters nothing there instead of launching the quiet daemon and writing a record that would park a present captain's main. - bin/fm-afk-launch.sh quiet-check says quiet mode needs nothing where the attended host runs, or that the session is paused while its broken-session latch holds; a quiet enter refuses there before writing anything. - Where the home opted in but the attended host lacks a part (engine, tools, verified mirror writer, identifiable main session, valid mirror), quiet-check names it and quiet mode falls back to the daemon. - Under a live away record on that home, quiet-check and a quiet enter refuse and name the record, so the return runs first, whatever state/.afk says. - A quiet enter records mode: quiet in the posture record, so start and start-native launch the quiet daemon without FM_AFK_MODE, and the away refusal wording fires only for away. - bin/fm-host-mirror.sh check validates the dialog mirror read-only and exits 1 on a missing, unreadable, or invalid mirror. - The quiet and afk skills and the supervision-host docs describe the new behavior; homes without the opt-in and Pi homes keep the daemon path. * no-mistakes(review): Archive the quiet record when a quiet daemon start fails * no-mistakes(document): Clarify quiet-mode documentation and remove stale duplicates
…uid#5884) * fix(bin): grant Claude workers their task-channel dirs via --add-dir Since Claude Code 2.1.257, a file-tool read (Read/Glob/Grep, and an Edit's mandatory prior read) of a path outside the working directories parks --permission-mode auto panes on a one-time interactive question, and a "Block" answer lands permissions.blockReadsOutsideWorkingDirectories in user settings, refusing the same reads even under bypass. Firstmate launches Claude with no --add-dir, so a secondmate's parent-home steering inbox and a ship or scout worker's launch record, steering inbox, brief dir, and code-root .agents/skills were all outside: workers wedged on the question the first time they read a steer. Every Claude launch, spawn and relaunch, in both permission modes, now grants exactly the task's channel directories: state/<id>.inbox for a secondmate (in the parent home), or state/operational-inbox, state/<id>.inbox, data/<id>, and the code root's .agents/skills for a ship or scout. Paths resolve to real paths and lazily created channel dirs are made before launch so the grant never names a not-yet-existing directory; the whole state/ is deliberately never granted. The grant keeps the bypass-mode launch argv changed on purpose: it also protects bypass workers against a machine-recorded Block answer. * no-mistakes(document): Consolidate Claude launch guidance in configuration reference * no-mistakes(document): Clarify Claude permission documentation reference
…nguid#4819) * fix(supervision): prevent idle recovery loops without stranding wakes * no-mistakes(review): Remove unused wake-append rollback helper
…id#5889) * fix(bin): stop the remote-job worker busy-polling an idle queue The serving loop slept 50ms between passes and re-ran state preparation (chmod on every queue directory), the heartbeat publish, and the stale sweep on every pass. It now blocks on a worker.wake FIFO that staging, cancellation, and lane exit nudge, keeps a short fast-poll window after activity, refreshes the heartbeat at most once a second, and runs the sweep (which re-applies the queue directories' 0700 modes) at startup and then on a bounded interval. Lane-owned records are no longer re-read every pass. Measured with a fork/execve-interposing counter on a --serve worker in a disposable HOME and queue, bash 3.2, 20-second windows (the counter slows the old loop to about 5 passes a second, so real-host rates were higher): idle worker 146 forks/s, 61 execs/s -> 4.8 forks/s, 3.1 execs/s one running long job 232 forks/s, 100 execs/s -> 15 forks/s, 11 execs/s Stage-to-result latency for a no-op job, idle and back to back, stayed at about 0.8-1.2s in both versions (dominated by job execution, not pickup). * perf(bin): drop per-cycle forks from watcher, drain, and lock helpers The watcher, drain, inactive-reconcile scan, and branch-outcome reads forked small external commands on every cycle where bash can do the same work. - fm-wake-lib.sh gains fm_dirname_to, fm_basename_to, and fm_epoch_seconds_to, exact stand-ins for $(dirname --), $(basename --), and $(date +%s); the clock uses printf %(%s)T on bash 4.2+ and still forks date exactly once on stock macOS bash 3.2. - fm_lock_abs_path, fm_wake_signal_seen_path, fm_path_age, the watcher's age_of and wedge timer, and the recovery-marker line count use them or plain reads instead of dirname/basename/tr/date/wc. - window_to_task reads a meta file once instead of two grep | tail -1 | cut -d= -f2- pipelines per file per call. - fm-classify-lib.sh reads uname -s once at source time instead of in every status stat helper. - Libraries sourced every cycle derive their own directory without forking dirname, including the backend adapter siblings a subshell re-sources on each probe. tests/fm-fork-free-helpers.test.sh pins each replacement against the command it replaces on edge-case inputs, under every available bash and both the C and a UTF-8 locale; CI's stock macOS bash lane runs it under /bin/bash 3.2. Measured with a fork/execve-interposing counter in a disposable home, one tmux crew task, FM_POLL=1 (forks and execs per watcher cycle, per run otherwise): watcher cycle bash 5.3 299/138 -> 199/66 bash 3.2 341/146 -> 224/80 drain bash 5.3 492/238 -> 430/200 bash 3.2 567/250 -> 491/212 inactive scan bash 5.3 27/14 -> 17/4 bash 3.2 37/14 -> 17/4 branch-outcome bash 5.3 40/21 -> 35/16 bash 3.2 48/24 -> 38/19 * test: note the interpreter-expanded version probe for shellcheck * no-mistakes(review): Fix worker idle bounds and fork-free contributions snapshot * no-mistakes(review): Coalesce buffered worker wake nudges into one wake * no-mistakes(review): Coalesce wake nudges via pending marker so publishers never block * no-mistakes(review): Claim wake nudges atomically via noclobber pending marker * no-mistakes(review): Release abandoned wake claims only after a 30-second bound * no-mistakes(review): Drop worker wake FIFO; load path helpers side-effect free * no-mistakes(document): Document remote worker polling and preemption cadence * no-mistakes(ci): Fixed both failing CI shards: isolated remote and teardown test fixtures now include fm-path-lib.sh, which fm-wake-lib.sh requires. The three affected tests, fm-lint.sh, and git diff --check pass locally
…uid#5941) * fix: keep a successor watcher and remote-reply listeners across the gaps that dropped them A main-only supervision pass-through exited without leaving a watcher, and each remote-reply poll released its claim until the next cycle, so short-lived listeners stayed down. * no-mistakes(document): Clarify listener and supervision continuity documentation * no-mistakes(ci): Fixed the three Greptile findings: failed ingestion leaves one durable capture, failed reads exit instead of relistening, and the disposable-checkout guard rejects state paths outside the marked lab. Added behavioral tests; the remote-reply and watcher-lock suites, shell syntax checks, and git diff checks passed * no-mistakes(ci): Fixed a race in the wake-queue interruption test: it now waits for the drain to own the lock and enter handling before signaling it. The wake-queue suite, shell syntax check, and diff check pass
…enguid#5925) * fix: date replayed branch outcomes and ask main to check current state first A captain outcome main never acknowledged is presented again, which after a harness or posture switch, or the first drain after the upgrade whose earlier presenter never advanced the read cursor, can be days after its situation settled. The replay read as fresh news, so a PR since merged looked ready. bin/fm-branch-outcome.sh now adds a "recordedAgo" age (minutes, hours, then days) to present and unprocessed rows, one owner of that wording for both presenters. The drain's BRANCH OUTCOMES captain lines and the Pi branch's processing request name that age and ask main to check the task's current state first; an outcome already settled needs only the acknowledgement, with nothing relayed to the captain. Nothing is adopted as processed, so a fresh home's first outcome is still presented until acknowledged. * no-mistakes(review): Absent processed marker reads 0; never adopt read cursor * no-mistakes(review): Require recordedAgo in Pi requests; report undated rows to main * no-mistakes(review): Keep recordedAgo on captain rows only in present output * no-mistakes(document): Correct cutover documentation and retire stale migration guidance * fix: keep settled branch outcomes out of main's reply to the captain A live Pi primary that took over a host-drain home received the carried-over outcomes dated and check-first, but its processing reply still told the captain about an outcome whose decision had since been answered. The request also claimed every outcome was already shown as an anchor entry in this transcript, which is false for an outcome carried over from before a restart or a switch of primary. The Pi processing request now says each outcome was recorded earlier and may already have been seen or handled, and that a settled outcome gets no captain-facing mention at all in the reply or any recap, not even that it is settled. The drain's BRANCH OUTCOMES header and the supervision docs state the same rule, and the tests check both delivered texts. * fix: scope main's outcome reply to what is still open Telling main what not to say about a settled outcome was not enough: in two live Pi trials the processing reply still told the captain that an answered decision was settled. Main now sorts the outcomes by current state first, and its reply to the captain covers only the still-open ones, written as if the settled ones had never been listed. With that framing three live Pi trials kept the settled outcome out of the reply and relayed the open one each time. The drain's BRANCH OUTCOMES header and the supervision docs use the same framing, and the tests check both delivered texts. * no-mistakes(review): Clarify that main acknowledges every presented captain outcome * no-mistakes(document): Clarify outcome cursor ownership across Pi and host * no-mistakes(ci): Fixed Pi replay by batching unprocessed captain outcomes oldest-first and acknowledging only through each batch. Verified a marker-less backlog over 1 MiB replays through all batches. A real-drain regression confirms an older keyed decision remains under OPEN DECISIONS after a newer branch row is acknowledged; the check-first instruction now names those decisions. Relevant targeted tests and branch-supervision tests passed; the full host suite timed out * no-mistakes(ci): Fixed the host drain’s check-first wording in bin/fm-wake-drain.sh; the CI fixture now passes. The full host suite passed the affected fixtures but timed out later. Syntax and diff checks passed * no-mistakes(ci): Fixed ci-1: abbreviated Pi outcome summaries now stay within 1,024 characters and include a row-specific full-outcome lookup command. The delivered instruction requires reading the full outcome before acting, relaying, or acknowledging it. The new extension-driver regression failed before the fix and passes now; the Pi and supervision-host suites pass * no-mistakes(ci): Corrected the batching sentence in docs/pi-supervision-branch.md. The cancelled CI check needs no code fix; its clean rerun passed. The Pi branch extension suite and git diff check passed
…guid#5961) * fix: wake an idle Claude primary for attended main-only hand-backs An attended main-only pass-through confirmed a handling handoff for the successor it leaves running, which flipped the recovery marker to handling. The Claude Stop hook only rewakes main while that marker reads downtime, so the close reached no one and an idle primary slept with wakes queued. The pass-through now leaves the marker at downtime, and a close that turns main-only at its turn hands the consumed handoff back to downtime before it reaches main. Regression tests drive the real Stop hook around the real host on both paths and for the successor's own later close, and a new opt-in live guard proves it against an idle interactive Claude primary with a pre-fix negative control. * no-mistakes(review): Assert live lab Stop-hook registration via parsed settings JSON * no-mistakes(document): Correct supervision hand-back documentation * no-mistakes(ci): Fixed the failed downtime-write path so the Stop hook notifies main instead of silently dropping the close. Corrected the live guard’s tracked-hook check and added the requested at-turn main-only scenario. The new regression failed before the fix and passed after it; the host suite, syntax checks, and diff check pass. The credentialed live guard was not run in this CI phase because it writes outside the worktree * no-mistakes(ci): Fixed the Stop hook’s retry ordering: a crashed host gets its bounded retry before a non-crash hand-back failure is reported. The Stop-hook suite passes, including the crash regression. The host suite passed the failed-marker-write regression but timed out before completing; syntax and diff checks pass
kunchenguid#5916) * Fix worker launches to enter recorded worktrees * no-mistakes(review): placeholder * no-mistakes(document): Update agent-control.md worktree-refusal note to match new universal cd+assert * no-mistakes(review): Add regression tests for Orca spawn/relaunch worktree carve-outs * Fix PR relaunch and prelaunch cwd verification * no-mistakes(document): Fix docs/agent-control.md: worktree cwd check is pre-launch, not post-launch * fix: slim worktree launch change onto upstream main
…ardown (kunchenguid#5997) * WIP: retire task-keyed watcher markers and orphan journals at teardown Re-applies old PR kunchenguid#5584 on current main: teardown retires the turn-ended .seen-* signature and an orphaned Herdr presentation journal whose workspace is already gone, and the wake-drain rotates its own dead scratch files. Not yet validated through no-mistakes. * no-mistakes(ci): Fixed the Greptile P1 finding in bin/backends/herdr.sh. fm_backend_herdr_projection_token_workspace_gone used `! ... jq -e ... 2>&1`, which swallowed a jq runtime error (thrown when a non-object workspace entry, e.g. a number before a live token-bearing workspace, hits `.label`) and flipped it to a "gone" verdict, causing teardown to delete a still-live v1 presentation journal. Invariant: a workspace-query error/ambiguity must never be read as token absence; only a cleanly-parsed list with no token-bearing label is "gone". Replaced the body with a single jq verdict (unknown/present/gone): a non-array list or any non-object/non-string-label entry yields "unknown", jq errors/empty output fall through `|| return 1` to unknown, and only "gone" returns 0. Sibling fm_backend_herdr_projection_endpoint_matches_journal already fails safe on jq error (empty match -> journal kept), so it needed no change, matching the author's scoping. Added test_teardown_retains_v1_journal_when_workspace_query_ambiguous driving real teardown with a malformed workspace-list entry, proving the journal is kept and no workspace close occurs. Verified the old logic returns GONE on that input (test fails before, passes after); full tests/fm-teardown.test.sh suite passes (exit 0) and shellcheck is clean. Marker-naming finding left untouched per explicit out-of-scope instruction
…unchenguid#6002) * fix(tests): disable Claude Code's auto-updater during live harness runs fm_live_gate let a live run proceed without ever setting DISABLE_AUTOUPDATER, so a live Claude test could let the real updater repoint ~/.local/bin/claude into a temporary directory and stop every Claude process on the machine from starting. Export DISABLE_AUTOUPDATER=1 on every path where the gate lets a live run proceed, and assert the export in tests/fm-live-gate.test.sh, including that it reaches a child process the same way a real harness pane would inherit it. * no-mistakes(ci): Greptile flagged that the PR's DISABLE_AUTOUPDATER inheritance test only checked a `bash -c` direct child, not the fm-spawn.sh launch path. The user chose to fix it with a regression on that path. In tests/fm-live-gate.test.sh I replaced the generic child test with test_disable_autoupdater_reaches_the_claude_pane_on_the_fm_spawn_launch_path: it drives the real fm-spawn claude launch through the spawn fixtures, captures the exact staged launch command, and runs it as a synthetic pane whose only `claude` is a stub recording the inherited DISABLE_AUTOUPDATER, asserting it saw 1. Switched the file to source fixtures.sh (pulls in lib.sh, guarded) for the spawn helpers. Verified it is a real guard: the stub records `1` when the ambient var is set and `unset` when absent, so it fails if fm-spawn ever scrubbed the variable (e.g. env -i or -u). This confirms fm-spawn's launch construction never references the name and passes it through via ordinary ambient inheritance with no allowlist. Full suite passes (12 tests ok), shellcheck clean. Note for the outer executor: I embedded the daemon caveat as a code comment in the test, but the finding also asks the PR body to state that a backend daemon already running before the gate exported the variable does not inherit it and fully covering that would need launcher support - that forge-side PR-body sentence is outside this CI phase's scope * no-mistakes(ci): Fixed Greptile finding ci-1. Root cause: fm-spawn.sh handed its launch command to an already-running backend daemon that never inherited the test process's exported DISABLE_AUTOUPDATER, so ambient inheritance dropped it and Claude's auto-updater could still run. Fix (bin/fm-spawn.sh): when DISABLE_AUTOUPDATER is set in the spawn's own environment, embed `export DISABLE_AUTOUPDATER=<value>;` into the LAUNCH command text (same idiom as the adjacent COMPACT_ADVISER_DISABLE export), so it survives a daemon-built pane, the env -i allowlist path, and relaunch alike; gated on presence so ordinary spawns are unchanged. Added regression test test_disable_autoupdater_survives_a_daemon_pane_that_never_inherited_it in tests/fm-live-gate.test.sh: stages a real claude launch with DISABLE_AUTOUPDATER set in the spawn env, then runs that exact command in a synthetic pane with `env -u DISABLE_AUTOUPDATER` and asserts the claude stub still recorded autoupdater=1. Verified the test fails (autoupdater=unset) without the fix and passes with it; the round-1 ambient test stays green either way. Full suite passes (13 ok); test file and isolated snippet shellcheck-clean (full fm-spawn.sh shellcheck kept getting terminated by the memory-constrained host, not by findings). Forge-side note for the outer executor: the PR-body caveat that fully covering the daemon case would need launcher support no longer applies to the Claude launch path and should be corrected
…cevent record (kunchenguid#6010) * fix(bin): ring the inbox doorbell only for a newly published procevent result publish_result rewrote a worker's captured Lavish round idempotently on every reconcile, unconditionally moved an already-acknowledged inbox record back out of handled/, and rang the doorbell every time - so an already-processed round rang the owning worker on every cycle. Snapshot the existing active and handled records before the idempotent write and ring, or move anything, only when the write actually created a fresh record; re-delivery of a still-open round is left to the inbox's own re-ring ladder. * no-mistakes(document): docs: reflect worker-board doorbell rings only on fresh inbox record * no-mistakes(ci): Fixed Greptile finding ci-1 in tests/fm-procevent.test.sh (test-only change). The redelivery regression previously moved the delivered note into handled/ before any repeated reconciles, so it only proved an acknowledged note stays quiet and would still pass if an unchanged active note rang every cycle. Per the user's instruction, I inserted (before the mv into handled/) five repeated `pe reconcile` runs with the note still in the active inbox and asserted the ring log holds exactly one line and 001.msg remains active; the existing acknowledged-note assertion after the move is kept unchanged. No product code changed. bash -n confirms syntax is valid; the block mirrors the already-passing post-move reconcile/ring-count assertion directly below it
…unchenguid#6032) * fix(bin): make the Claude Stop auto-arm refuse arguments before arming A model running bin/fm-claude-stop-autoarm.sh --help mid-turn armed a real supervision-host park owned by its short-lived tool process, leaving supervision down once that process exited. The Stop hook passes no arguments, so -h/--help now prints usage and any other argument is refused before anything is sourced, read, or armed. * no-mistakes(document): Clarify Claude Stop hook documentation for manual invocations * no-mistakes(ci): Updated the argument-run regression test to compare checksums of state files as well as entry names. The Stop auto-arm test suite passes, and git diff --check is clean * docs: restore the bin/ toolbelt intro's manual-use clause The document step dropped "interactive entrypoints work by hand too" from docs/scripts.md, which still holds for most bin/ scripts. * no-mistakes(document): Clarify Claude auto-arm manual-use guidance
…uid#6039) * feat(calm): show supervision sailboat and anchor notes on Claude Code The Calm mod follows a bounded display tail copy of the outcome store, which bin/fm-branch-outcome.sh append now refreshes, and the supervision host's latch, and appends one dim transcript line per visible routine outcome, captain outcome, and latch change, replaying unread and unprocessed outcomes at session start. It shows them whenever the mod is active, regardless of config/calm, and never marks anything read. * fix(calm): show each supervision note once per session on Claude Code Claude Code 2.1.283 stores ui.log lines in the session and restores them on --continue, so the mod records how far each session has followed the outcome store and a resume replays only newer outcomes. It also checks file existence before reads so absent files do not log debug errors. The live guard gains the supervision-notes scenario and the dated 2.1.283 record documents the observed behavior. * docs: name the Claude supervision note row as the engine draws it * no-mistakes(review): Seed outcome tail on present and anchor first tail on markers * no-mistakes(review): Seed outcome tail at session start; replay against start markers * no-mistakes(review): Bound outcome tail by bytes; reread recently changed files * no-mistakes(review): Skip store validation when outcome tail already exists * no-mistakes(document): Clarify bounded Claude supervision note replay * no-mistakes(ci): Fixed seed-tail to validate only a bounded suffix of complete store rows and write it through the existing byte- and row-limited tail writer. Added a regression test with malformed history outside that window and updated the script header. Outcome tests and shellcheck passed; the full session-start suite timed out after 240 seconds
…enguid#6033) * fix(bin): read a quiet-mode record as a present captain, never hold-for-return Daemon-backed quiet mode writes the away-posture record marked mode: quiet, but the entry announcement, read-back, and session-start digest rendered it as "hold-for-return only", and the spend cap and PR merge gate treated it as away. A present captain's requested actions could then be held for a return that was not coming. bin/fm-afk-contract.sh now owns which posture a record is (the mode subcommand, fm_afk_contract_mode, fm_afk_contract_away_present). A quiet record announces, reads back, and appears in the digest as a present captain holding nothing; merges under it stay attended and it binds no spend cap. An away record is unchanged, an /afk entry over quiet mode rewrites the record as away, and a quiet entry never turns a standing away record quiet. * no-mistakes(document): Clarify quiet-mode authority and remove stale away guidance * no-mistakes(ci): The CI failure came from a race in the supervision-host test: its restart fixture could observe a watcher left by the preceding cycle. The test now retires that watcher and waits for the fixture arm to report its own started cycle. The focused test passed three times; the full suite was attempted but stopped at a separate intermittent test failure * no-mistakes(ci): Fixed daemon refresh mode selection so an unset-mode refresh follows the posture record: /afk over a running quiet daemon now changes state/.afk to away, while a plain quiet refresh stays quiet. Added script-level regression coverage for start and start-native and corrected a quiet-refresh fixture. The launch test suite, syntax checks, and diff check passed * no-mistakes(ci): Herdr was blocked before tests ran by a GitHub HTTP 500 downloading pinned Treehouse; no code change was warranted for that check. Fixed the Lint 1 ShellCheck warning in tests/fm-afk-launch.test.sh by annotating the intentional background PID capture. The focused test suite, ShellCheck, syntax check, and diff check passed
…merge (kunchenguid#6053) * fix(bin): accept a task's next PR once fm-pr-merge confirms the bound one merged require_recorded_pr_identity now checks fm_pr_poll_merge_already_notified for the recorded pr= before refusing a different URL, so a task's later PR is accepted once its earlier PR's merge is confirmed, while it keeps refusing while the bound PR is still unmerged. * no-mistakes(document): docs(fm-pr-merge): note next-PR accepted after bound PR merges
…6064) * fix(bin): read a live quiet record as a present captain at the host and watcher A quiet record left without its daemon (a quiet start that never ran or was interrupted) was read as away by the supervision host, so it parked a present captain's main and held captain outcomes for a return that never comes, and the watcher and daemon silenced captain-held rechecks on record presence. The host's posture checks, the watcher's and daemon's captain-held silencing, and the host's outcome path (branch report, drain BRANCH OUTCOMES, relocated branch authority, the owners' away wake note, and the Codex checkpoint bound) now ask the record owner's away-or-quiet reading, so only an away record is away. A live away record keeps today's behavior. * no-mistakes(document): Correct quiet-record documentation and supervision guidance * no-mistakes(document): Clarify quiet-record posture and captain-held rechecks * no-mistakes(document): Clarify quiet-record posture in documentation
…kunchenguid#6043) * fix(bin): name an in-window engine latch in the return brief and drop the false handling GAP line The away return brief said nothing had failed after the supervision host latched on engine errors during the window, and printed a GAP: watcher downtime line whenever a wake was merely being handled or queued at return. The failures section now reads the host ledger and latch record and names the latch time, the window's engine-error count, and whether the session is still paused or recovered. An open recovery episode is reported as information, and as a gap only when a queued episode outlived the return grace or the marker cannot be read. * no-mistakes(review): Fix latch trip time, drop marker-age grace, bound error count * no-mistakes(review): Report paused latch without ledger trip row; bound errors * no-mistakes(review): Never report a failed probe's latch row as trip time * no-mistakes(review): Only a retained trip row marks a pre-window latch * no-mistakes(document): Clarify return-brief latch and watcher-gap documentation * no-mistakes(ci): Fixed Lint 1 by marking the shared cooldown constant as used by sourcing scripts. The repository lint command and diff check pass; the return test run was stopped by a 180-second timeout after its completed cases passed * no-mistakes(ci): Fixed the return brief so the trip time and error count come from the same initial latch row, and ledger rows before the current session’s lock boundary cannot affect its latch report. Added real-script regressions for both findings. The return test suite, repository lint, and diff check pass * no-mistakes(ci): Fixed the return brief’s restart cutoff so it retains in-window failures, prints one line per initial-trip row, and omits zero-error count wording. Added real-script restart regressions. The return test suite, ShellCheck, and diff check pass * no-mistakes(ci): Fixed the return brief so a recorded trip followed by recovery stays recovered, while a later pause with a lost trip append gets a separate “trip time unavailable” line. Added a real-script regression that failed before the fix. The return test suite, ShellCheck, syntax checks, and diff check pass * no-mistakes(ci): Fixed the false second latch during recovery. A real-script regression failed before the fix and passes now; the lost-second-trip test still passes. The return test suite, ShellCheck, syntax checks, and diff check pass
* fix(calm): name the Claude Code Calm plugin fm so supervision notes read "fm: " Claude Code labels every mod transcript line with the plugin name, so the notes rendered as "firstmate-calm: ⚓ ...". Rename the plugin to fm, update the live guard to assert the fm: label, and document the one-time replay for sessions resumed across the rename. * no-mistakes(document): Clarify Calm plugin rename in documentation
…#6037) * feat(bin): add fm-live-lab.sh, a one-command live supervision lab builder * fix(bin): exact lab windows, per-lab task ids, self-safe teardown * fix(bin): target lab windows by id, stop lab descendants, add readiness tests * fix(bin): keep Claude's auto-updater off in live labs; list fm-live-lab.sh * fix(bin): start the lab tmux server without user config * no-mistakes(review): Scope lab teardown to its store, root, and task ids * no-mistakes(review): Record selected user stores at up for check and down * no-mistakes(document): Clarify live lab documentation and remove stale narratives * no-mistakes(ci): Fixed the CI failure by checking for an existing lab root before looking up the harness executable. The affected behavioral test and shell syntax check pass; the refusal also works with Claude absent from PATH * no-mistakes(ci): Fixed all four Greptile findings: teardown signals only recorded lab processes and their descendants; the worker gate is in its granted task directory and its path is exposed; readiness uses current crew state; and mate and worker IDs use 12 nonce hex digits. The CLI behavior tests pass, as do shell syntax, ShellCheck, and diff checks. The Claude no-host path is unchanged * no-mistakes(ci): Fixed the CI test’s dependence on an installed Claude binary by supplying a test-local stub. The full fm-live-lab test, shell syntax check, and diff check pass * no-mistakes(ci): Fixed all three selected findings in bin/fm-live-lab.sh: down waits for recorded processes and escalates before cleanup, PID roots are checked against recorded start times, and Claude primary trust is rechecked after mate/worker readiness. Added behavioral tests in tests/fm-live-lab.test.sh. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass * no-mistakes(ci): Fixed the pre-primary settle wait, worker gate instructions, unused retry variable, and teardown PID revalidation in bin/fm-live-lab.sh. Added behavioral tests in tests/fm-live-lab.test.sh. Both requested commands pass: tests/fm-live-lab.test.sh and bin/fm-lint.sh * no-mistakes(ci): Fixed teardown to track pre-kill lab processes by PID and start time, including children orphaned when a root exits. Up now rejects an empty pane PID before calling ps. Added regression tests and a Linux-safe worker fixture. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass * no-mistakes(ci): Fixed teardown tracking for children spawned during shutdown and made the worker fixture verify its exact window with a Linux-available shell. Both requested checks pass. The lab test takes about 66 seconds locally, so the under-one-minute target remains unmet * no-mistakes(ci): Fixed ci-2 and ci-4 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. Teardown now tracks identity-checked members of captured lab process groups, including children orphaned during shutdown, without signaling the caller’s group or unrelated processes. Lint passed, and the lab test passed four times * no-mistakes(ci): Fixed teardown so an observed-empty process group is permanently dropped, preventing a reused group ID from signalling unrelated work. Added a ps-shim regression test. The lab test, lint, and diff checks pass * no-mistakes(ci): Fixed ci-1 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. The TERM-born-child fixture now waits until its handler is installed before calling down. Down sends SIGKILL to identity-valid survivors on every pass from pass 20 onward and includes survivor process details if it must refuse cleanup. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass locally; Linux CI remains to be verified * no-mistakes(ci): Fixed down’s teardown wait to require two empty identity-checked scans separated by 0.5 seconds, and removed the unused test loop variable without changing the TERM-born-child test. The lab test, lint, and diff check pass locally
Bring the fork up to date with kunchenguid/firstmate at b5fdf74. This commit keeps both parents so fork history and earlier pull-request references stay intact. Land it as a merge, not a squash. Still-needed fork work is re-applied on upstream's structure: cmux failed list-panes, Herdr in-process schema match, Claude session identity stripping, hosted Codex lock ownership, guarded Herdr legacy repair, tasks-axi handoff gating, Bitwarden ceremony, Automic Vault, project lifecycle posture, versioned Bearings contract, Orchestra refresh, authenticated PR poll across relaunch, captain-hold prune reconciliation, and active-only startup clone refresh. Superseded by upstream: review-diff always-HEAD (e31fd31) and the Pi Calm renderer baseline (cdd5462), except the watch-arm cleanup. keep-ai-trailers defaults off. Absent config/keep-ai-trailers inserts empty Claude attribution settings, so commits and pull requests do not gain an AI co-author trailer. Away-mode composer fixes d08e327 and 050a446 are ancestors of the upstream parent. No fork commit touched the composer classifier.
…c Vault launch substitution, updated staged-launch and terminal-signal tests, made Pi Calm rendering delegate to stock shells outside Calm mode, and supported Pi 0.99 hidden export messages. Verified affected tests against Pi 0.87.1 and 0.99.0, ran the full Claude Vault and Pi watcher tests, and passed fm-lint plus git diff checks. TypeScript typecheck was unavailable locally because tsc is not installed
… now waits for the queued post-reload repaint to restore the required 2-row gap instead of accepting Pi's intermediate 4-row frame. Verified with bash syntax checking, the focused test suite, `bin/fm-lint.sh`, and `git diff --check`. The live geometry case could not run locally because tmux is unavailable
…start boundary. Pre-start built-in tool rows are now invalidated after persisted Calm state is restored, covering all seven wrapped tools and both call/result rendering paths. Added behavioral lifecycle coverage. Focused suite, fm-lint, and git diff checks pass. The tmux E2E was unavailable locally because tmux is not installed; tsc was also unavailable
… Calm state before transcript rows are reconstructed. Disabling Calm now reloads the transcript so rows return to genuine Pi stock shells. Added behavioral coverage for all seven built-ins. Pi 0.99.1 focused tests, strict typecheck, branch rendering checks, fm-lint, bash syntax, and diff checks pass. The tmux geometry E2E could not run locally because tmux is unavailable
…r on or off, ensuring restored rows use the correct shell. Updated the geometry E2E to wait for the exact 2-row settled gap and added behavioral coverage for reloads in both directions. Focused Calm tests, Pi stock-shell tests, fm-lint, bash syntax, and diff checks pass. The tmux geometry E2E was skipped locally because pi or tmux is unavailable. Remote provisioning was not changed; serial 8 requires only a CI rerun
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Asked of 2026-09-28, about the running Firstmate: "As far as firstmate, are we behind from kunchenguid's main?" After hearing yes: "yes, queue the fork sync after the tool updates" (the tool updates finished 2026-09-29).
Context needed to read that ask: the running Firstmate is the fork https://github.com/rega10/firstmate (remote
origin); kunchenguid's main is https://github.com/kunchenguid/firstmate (remoteupstream). A "fork sync" means bringing the fork up to date with upstream main through one validated PR while keeping the fork's own captain-authorized work, as the two earlier fork syncs did (#26 and #27). At intake on 2026-09-29, origin/main 6266122 is 190 commits behind upstream/main b5fdf74 and 29 commits ahead of it (merge base b430bf5). Upstream's new work includes away-mode composer fixes (d08e327, footer rows; 050a446, full Herdr viewport) that are expected to fix the away-mode injection wedge where every away-mode escalation was deferred because the Claude composer read as pending; and a keep-ai-trailers switch whose default must be checked against this fork's standing rule that commits never carry an AI agent co-author trailer.What Changed
Risk Assessment
🚨 High: The Calm correction leaves the reported restored-row failure reachable during /reload, and its geometry test can inspect an unsettled frame.
Testing
Real Herdr injection, Claude composer classification, Pi 0.99 Calm layout, and Git trailer behavior passed. Disposable spawn and Vault checks passed after a test fixture correction. Merge ancestry was verified. The tmux-dependent Pi test cases skipped because tmux is absent, so the Pi UI was driven through a sized pseudo-terminal and its raw output replayed through xterm for visual and row-count evidence.
bash tests/fm-afk-inject-herdr-e2e.test.shexercised partial input, a swallowed Enter, normal delivery, and a persistently pending composer against an isolated real Herdr session.bash tests/fm-git-strip-ai-trailers.test.shbash tests/fm-claude-automic-vault.test.shand the passing rerun ofbash tests/fm-spawn-dispatch-profile.test.shEvidence: Claude composer classification
Source: Claude composer classification
Evidence: Actual Git commit attribution
Source: Actual Git commit attribution
Evidence: Fork merge ancestry
Source: Fork merge ancestry
Fork sync record
Upstream main merged here is
b5fdf74d654c9810517c3f473f2d4bc987af217a.The merge commit is
0deb804c72c69f47713d45238640e8664dd2bf0c, with parents6266122f54e01a4d73e7fd767bb9620d8bbbeb22(fork main) and that upstream commit.Land this pull request as a merge, not a squash, so both parents stay.
Pull request 27 was squashed and lost its upstream parent.
config/keep-ai-trailersdefaults off.The running home has no
config/keep-ai-trailersfile, and this change does not invert that branch.Absent config inserts empty Claude attribution settings, so commits and pull requests do not gain an AI co-author trailer.
Away-mode composer fixes
d08e327d(footer rows) and050a4464(full Herdr viewport) are ancestors of the upstream parent.No fork-only content commit edited the composer classifier.
The Calm presentation files that conflicted were taken as upstream.
After the first CI run failed on Pi 0.99, the Calm renderer was restored to the merge commit.
bin/fm-spawn.shandtests/fm-claude-automic-vault.test.shkeep the__CLAUDELAUNCH__placeholder..pi/extensions/fm-branch-supervision.ts,tests/fm-calm-pi-extension.test.sh,docs/calm.md, anddocs/calm-mode-feasibility.mdtake upstream pull request 6162 at51274f45.The review findings that asked to put the reload path back were not applied.
GitHub checks on
cb3f3b79785d3fe186e313856ad97d3f9f90b0a6then passed, including the Pi lanes.Fork commits
None of the unique fork commits were already upstream.
History rows are merges, not a separate capability.
--contract; CI expects 82 Bearings testsConflict resolutions
The first merge reported 56 content conflicts.
24 files had no fork delta since
fa930971and were taken as upstream:.claude/mods/firstmate-calm/hooks/register.ts,.claude/mods/firstmate-calm/lib/fm-calm-presentation.ts,bin/fm-classify-lib.sh,bin/fm-contributions.jq,bin/fm-contributions.sh,bin/fm-control-lib.sh,bin/fm-crew-state.sh,bin/fm-dispatch-resolve.sh,bin/fm-inactive-reconcile.sh,bin/fm-pr-check.sh,bin/fm-pr-lib.sh,bin/fm-procevent-remote-reply.sh,bin/fm-quota-axi-lib.sh,docs/calm.md,docs/remote-secondmates.md,docs/verification/dispatch-resolve.md,tests/fm-bootstrap.test.sh,tests/fm-calm-claude-mod.test.sh,tests/fm-classify-decision-key.test.sh,tests/fm-crew-state.test.sh,tests/fm-dispatch-resolve.test.sh,tests/fm-inactive-reconcile.test.sh,tests/fm-pr-check-security.test.sh, andtests/fm-remote-reply.test.sh.Four files re-merged cleanly:
bin/fm-bootstrap.sh,bin/fm-captain-hold.sh,docs/verification/runtime-backends.md, andtests/fm-secondmate-lifecycle-e2e.test.sh.The other 28 were resolved by hand:
.agents/skills/quota-array-dispatch/SKILL.mdAGENTS.md.github/workflows/ci.ymlbin/fm-config-inherit-lib.shkeep-ai-trailers;claude-automic-vaultstays remote-local-onlydocs/agent-control.mddocs/architecture.mddocs/fm-test-portable-shards.mdtests/fm-afk-contract.test.shtests/fm-watch-arm.test.shbin/fm-project-mode.shbranch=/forge=parsingbin/fm-review-diff.shtests/fm-review-diff.test.shbin/fm-session-lock-lib.shbin/fm-lock.shbin/fm-bearings-snapshot.shtask_for_branchbin/fm-fleet-snapshot.shbin/fm-spawn.shbin/fm-test-run.shdocs/captain-hold-lifecycle.mddocs/configuration.mddocs/scripts.mddocs/verification/supervision.mdtests/fm-bearings-snapshot.test.shtests/fm-contributions.test.shtests/fm-spawn-dispatch-profile.test.shtests/fm-task-delivery.test.shtests/fm-watch-checkpoint.test.shtests/fm-watch-triage.test.shenvso the assignment is a real prefixTwo follow-up fixes are in the same merge commit.
Fleet-sync fixtures add a task and then start it, because upstream refuses
add --start.The watch-triage background launch uses
envwithFM_SECONDMATE_LIVENESS_SECS=99999999.Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
.pi/extensions/fm-calm.ts:418- Pi rebuilds transcript rows before emitting session_start on /reload. A freshly loaded Calm module starts with presentation off: this change removed the load-time preference restore, then line 418 clears the invalidators for rows built before line 420 restores Calm. Those rows retain visible call and result components despite config/calm=on. The same transition must hold for all seven built-in wrappers at line 275, fm_watch_arm_pi at .pi/extensions/fm-primary-pi-watch.ts:1154, and fm_branch_outcomes and fm_branch_processed at .pi/extensions/fm-branch-supervision.ts:2219 and :2281. Restore the preference before row construction or invalidate every restored tool row after session_start publishes it.tests/fm-calm-pi-extension.test.sh:2811- The /reload wait now returns as soon as the final text reappears, including Pi's intermediate four-row frame, so the immediate two-row assertion can fail before the layout settles. The Calm-on wait at line 2898 likewise stops when tool text disappears rather than when the gap reaches two. Both waits need the recorded decision's exact settled-gap condition.✅ **Test** - passed
✅ No issues found.
bash tests/fm-afk-inject-herdr-e2e.test.shexercised partial input, a swallowed Enter, normal delivery, and a persistently pending composer against an isolated real Herdr session.bash tests/fm-git-strip-ai-trailers.test.shbash tests/fm-claude-automic-vault.test.shand the passing rerun ofbash tests/fm-spawn-dispatch-profile.test.shbash tests/fm-afk-inject-herdr-e2e.test.shbash tests/fm-git-strip-ai-trailers.test.shplus a disposable Git commit inspected throughgit log -1bash tests/fm-calm-pi-extension.test.shPi 0.99.0 in a sized 100x44 pseudo-terminal:/skill:ahoy,/reload,/calmoff,/calmon; raw output replayed through xtermClaude Code 2.1.284 in a disposable profile and sized pseudo-terminal: idle and typed-draft viewports passed tofm_composer_classify_screenbash tests/fm-claude-automic-vault.test.shbash tests/fm-spawn-dispatch-profile.test.shafter the pane-log test fixgit showandgit merge-base --is-ancestorfor both merge parents✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.