feat(bin): renew fork from upstream with typed dispatch resolution - #27
Merged
Merged
Conversation
…kunchenguid#4424) * fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (kunchenguid#42) * fix(pr-merge): read a plan-gated 403 on branch rules as no merge queue github_read_queue_method left status=unreadable for every failed rules read, including a 403 whose body is GitHub's own "Upgrade to GitHub Pro or make this repository public" message. A repository whose plan cannot expose branch rules cannot have a merge_queue rule either, so that specific 403 now resolves to status=none instead of unreadable - unblocking the away-merge grant on private repos without GitHub Pro. Any other failure (auth, rate limit, network, 404, unrelated 403) still reads as unreadable. * no-mistakes(document): Update stale away-merge queue-grant comment for plan-gated 403 --------- Co-authored-by: NewAiCoder <claude@theinbtw.com> * no-mistakes(review): Fix misleading away-queue-grant comment in fm-pr-merge and its test * no-mistakes(document): Update architecture.md for plan-gated-403 merge queue exception --------- Co-authored-by: NewAiCoder <claude@theinbtw.com>
…unchenguid#4246) * fix(tests): select readers of a changed top-level test fixture bin/fm-test-run.sh --changed recognised shared test helpers by an explicit list, tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh. A top-level tests/*-fixture.sh matched none of those, fell through to the tests/* catch-all, and was marked unmapped, so selection aborted with "no changed-test mapping for source path" and the run selected nothing at all. tests/herdr-client-pair-fixture.sh and tests/remote-herdr-fixture.sh are real shared fixtures with real consumers, so any branch touching one of them left a validation pipeline driving --changed with a hard abort rather than a narrowed selection. Extend the helper arm to tests/*-fixture.sh rather than routing it through the tests/fixtures/*/* arm. Both arms resolve consumers with the same reference scan, and that scan is what selects the right suites here: it finds exactly the tests that read the fixture. The fixtures/ arm adds only a directory-keying step, which has nothing to key on for a top-level file, so the helper arm is the same behaviour with no extra machinery. A tests/ path nothing reads still reaches the catch-all and still refuses loudly. Refs kunchenguid#4100 * no-mistakes(test): order nested fixtures arm before top-level fixture glob * no-mistakes(document): document tests/ shared-file mapping contract and arm order * no-mistakes(review): drop vacuous test phase, correct header claim, restore comment
… asked, not declined (kunchenguid#4387) * fix(bin): read Claude Code's default external-imports flags as never asked, not declined (kunchenguid#4378) fm-claude-trust.sh refused the whole trust registration whenever the project-root entry carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code writes that value only on an explicit "No, disable". Claude Code's default project entry carries Approved and WarningShown both false before the dialog is ever shown, so every such project refused every spawn. Only Approved === false with WarningShown === true — the pair the dialog writes on a decline — now counts as a decline. false/false behaves like an absent flag: trust is registered and no import consent is manufactured. New case test_project_root_entry_default_import_flags_are_not_a_decline fails on b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31, bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * no-mistakes(review): Correct harness doc's external-imports decline predicate --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…chenguid#4445) * fix(brief): keep operator address out of composed intent Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body. The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content. Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion. Fixes kunchenguid#3882 * no-mistakes(review): Refuse operator-address lines in Captain's intent body * no-mistakes(document): Document operator-address refusal in intent contract comments
…as a proven empty composer (kunchenguid#4455) * fix(composer): accept Grok title overhang * no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage
…ailure (kunchenguid#4474) * fix(bin): recover Claude auto-arm after timeout * no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list
* fix(spawn): establish Claude task channel authority * no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference
…or pending text (kunchenguid#4458) * fix: guard relaunch exit against pending input * no-mistakes(review): Verifying test run in progress * no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard * no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now
…unchenguid#4460) * fix: reconcile diverged secondmate updates * no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md * no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md
…d#4497) * fix(dispatch): support Codex Luna max effort * no-mistakes(review): use portable CODEX_HOME path in codex effort reference
kunchenguid#4498) * feat(calm): render smooth Unicode swell * feat(calm): make sails asymmetric * feat(calm): use quarter sail glyph * no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer * no-mistakes(document): docs: sync calm wave phase doc comment * no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。
kunchenguid#4491) * fix: supersede scout delivery brief on promotion * fix: preserve ship safety contract after promotion * no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch
…d stop cleanup dropping accents from a held body (kunchenguid#4471) * fix(bin): let captain holds work on hosts with an older JSON::PP Holding a task for the captain, and the cleanup that keeps a captain-held row open, both fail outright on any host whose JSON::PP defaults allow_nonref off - 2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`, but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older library rejects that whole value with "must be object or array". The consequence is fleet-wide on such a host, not one broken command: a worker there cannot formally record a decision for the captain at all. It can only mention the decision in passing in a status line, where it can be missed - which is how a real decision goes unrecorded. The hold reports that the task lost its hold-set stamp; the cleanup cannot return the row to Queued. Both call sites now ask for allow_nonref explicitly rather than inheriting whatever the installed library defaults to. The second one is worth naming: its `/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string case that fails, so the guard selects for the failing input rather than protecting against it. The regression case forces the older default back off for every perl the commands spawn, then drives both paths - holding a task that carries a body, and tearing down a captain-held row whose deliverable must still be appended. It also probes that the simulation genuinely rejects a bare scalar, so the case cannot pass vacuously on a lenient host. Each half was verified failing on its own unfixed call site with that site's real error message. Suites: fm-captain-hold-lifecycle 51 cases, fm-backlog-atomicity 99 cases, 0 failures. Verification limit: the mechanism is reproduced and tested, but neither fix is verified against a real JSON::PP 2.27202 host, because none is in the loop. This laptop runs 4.06, where the bug does not manifest. `bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a brace-delimited object before decoding, so allow_nonref never applies. * fix(bin): stop cleanup silently dropping accented characters from a held body Cleanup rewrites a captain-held row's body to append the finished work's deliverable, and the decoder it reads that body with printed decoded characters to a stream with no `:raw` layer. A character at or below U+00FF then came out as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the accent. `fm_backlog_retain` writes that body straight back through `--body-file`, and nothing reported an error - the character was simply gone from a row still waiting on the captain. The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus `utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already used. Review of the parent commit found this on one of the lines that commit already changed. It predates that change. The test asserts bytes rather than decoded strings, because comparing strings cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any character above U+00FF makes perl print the whole string as UTF-8, so one body carrying both an accent and an em dash passes even unfixed and proves nothing. Verified failing before the fix on the accented row, passing after. Suites: fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures. * no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc * no-mistakes(review): drop whole-file UTF-8 check from retained-body test * no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc * no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc
…furniture (kunchenguid#4532) * fix(composer): read codex 0.154's idle starfield and status footer as furniture codex-cli 0.154.0 animates a braille "starfield" around its idle composer: on the row above the bold `›` prompt row, on the `›` row behind the SGR-2 dim `Ask Codex to do anything` placeholder, and on the row below it, then draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`). The cells are truecolor greys on both sides of the ghost luminance ceiling, so the brighter ones survive ghost stripping, and the rows below the glyph carry no structural edge. The shared classifier selected the bare `›` shape, extended its wrap region over the two rows beneath the glyph, read the survivors and the footer as wrapped typed input, and answered `pending`; the steering doorbell defers on exactly that verdict, so no doorbell ever reached an idle codex 0.154 pane. bin/fm-composer-lib.sh now recognises that furniture by shape, declared once next to the idle placeholders and reached from the two wrap-region boundary points: - a row whose non-whitespace content is entirely braille cells (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it never counts as wrapped typed content and bounds a bare composer's wrap region; braille behind the glyph row's content is stripped before the emptiness decision when nothing else follows the glyph; a row mixing braille with other text stays typed content; - the codex status footer bounds the wrap region exactly as omp's status row does, anchored on the effort token, a spaced middle dot, and a `~` or `/` path cell, so a typed `fix · tests` stays composer input; - `^Ask Codex to do anything$` joins the verified idle-placeholder set; the ghost strip remains what proves that row empty, and the bare-row rule that bright placeholder text is real input is unchanged. Unchanged: the strict blank-row rule, the styled=0 degradation (a plain cmux/orca capture of this screen still reads `unknown`, never `pending`), FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape. tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte with the divergence (letters in place of the starfield read `pending`) and the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh is the default-on live guard (token-free, skips explicitly without codex or tmux) that launches the installed codex idle and asserts `empty` through both the tmux and the cursorless styled reads, naming codex --version on failure. docs/verification/runtime-backends.md records the dated Herdr evidence: `pending` before, `empty` after, on the captured screen. * no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry --------- Co-authored-by: Todd Billings <todd@usdvcapital.com>
* fix(bin): refuse empty text steers in fm-send A marked secondmate request sent with an empty message delivered only marker and correlation bytes and minted a pending-reply expectation the parent could never see resolved, stalling the fleet with no loud error (kunchenguid#4255). Fail closed on an empty or whitespace-only message on the text path, mirroring the existing --resolve-key refusal. * chore: retain ambient Pi-lens autoformat as its own commit Formatting-only edits produced by ambient Pi-lens autoformat during the msg-loss investigation, kept separate from the behavioural change in c23acba so the fix stays reviewable on its own. AGENTS.md is deliberately excluded: its only autoformat edit stripped the trailing space from the documented FM_OPERATIONAL_PREFIX value, which bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11 records as permanent compatibility. Documenting that constant without its trailing space makes the doc wrong about the contract, so that one line was restored rather than retained.
…chenguid#4554) On rose-pine-moon the two-color water (cyan crests over blue troughs) read as a pink stripe over aqua, the yellow left sail and mast clashed with the red right sail, and the hull carried a blue interior run. Every water cell is now blue so the swell reads through glyph height alone, and both sail halves, the mast, and the whole hull are one yellow run. Geometry, cadence, animation, direction flip, resize clamping, and the narrow fallback are unchanged. Update the unit and real-TUI color assertions to the new palette and the Calm docs that described the old one.
…chenguid#4270) * fix(watch): stop aging a second mate's active turn from its launch The parent watcher's second-mate wake-loop stall check exempts a mate that is demonstrably inside an active turn, but secondmate_in_active_turn asked busy_turn_over_age first and returned "not in a turn" whenever that said the bound was crossed. busy_turn_over_age ages from state/<task>.turn-ended, falling back to state/<task>.meta. A second mate's turns end in its own home, so the parent never gets a turn-ended mark for it and the fallback ages the mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago was therefore permanently "over age", the busy pane was never consulted, and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false wake-loop stall. The gate now bounds the busy exemption by <idle> - how long the queue's drain position has not moved - which is evidence this home actually holds. A busy mate stays exempt while the queue has been frozen for less than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so the bound that stops a busy pane from proving liveness forever is kept rather than removed. busy_turn_over_age is untouched; its remaining callers are the ordinary crew busy-pane bound. The regression pins the case that actually broke: a mate whose launch record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn must not escalate, while the same mate with its queue frozen past the bound still publishes exactly one notification. The existing coverage only exercised a freshly launched mate, which passes either way. Reaching that alert now costs a pane capture inside the gate, so the three checkpoints in this suite that assert an alert move from a 1s to a 4s bound - the value the neighbouring active-turn cases already use. The bound is a ceiling, not a wait: the checkpoint returns on the first actionable wake. On a loaded machine a 1s bound missed the alert repeatedly; at 4s it did not miss in 20 runs under the same load. * no-mistakes(review): scope the second-mate active-turn regression test's coverage claim * no-mistakes(document): fix stale second-mate active-turn comments in fm-watch
…unchenguid#4278) * feat(bin): add read-only PR blocker and reviewer-discovery commands Two focused, opt-in commands that read GitHub and never write to it. fm-pr-state.sh reports what still blocks one pull request from the author's side: a closed or merged state, draft state, unknown or conflicting mergeability, absent or failing required checks, and a blocking CHANGES_REQUESTED decision explained by each reviewer's latest verdict, marked STALE when it was left at a superseded head. A pull request that only awaits an approval is not reported as blocked, and advisory checks are omitted. Every reading is taken against one exact head; a push that lands mid-read invalidates the whole result rather than mixing two snapshots. fm-pr-reviewers.sh suggests reviewers from the most recent commits to the pull request's exact changed paths, counting each commit once, resolving handles through GitHub's own commit author.login mapping, and excluding the author and Bot accounts. Both stay read-only: no review request, no approval, no merge. Unresolved review-thread state is left unreported because the REST API does not expose it and unattended commands may not use GraphQL. Closes kunchenguid#3731 * no-mistakes(review): accept only PR URLs and stop at terminal state * no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate * no-mistakes(review): stop attributing readings to unverified heads * no-mistakes(review): narrow readiness contract to checks that have reported * no-mistakes(review): read the pull request once, drop the head guard * no-mistakes(document): scope pr-forge isolation proof to its measured members * no-mistakes(document): record uncovered pr-forge members and their pending proof * docs(isolation-proof): re-prove pr-forge at its full membership tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the pr-forge family in this branch, and script_allows_concurrency grants four workers by family membership alone, so both ran concurrently on a proof measured before they existed. Re-proved the family at all eight members: two consecutive runs, 0 failures, each begun with the one-minute load average below 6.0 so the result measures isolation rather than contention. A third run taken between them is disclosed rather than recorded, because it started while the previous run's workers were still decaying. The new durations are not comparable with the six-member measurement above them, so they are not presented as evidence about the two new members, and that record's 1.72x four-worker figure is left as a statement about its own run rather than restated as current. * no-mistakes(review): disclose gh error-text coupling at its matching site and tests
…uid#2752) * fix(bin): teach validation-round pauses in briefs * no-mistakes(document): Point classifier comments to authoritative pause examples
…guid#4510) * fix(teardown): refuse a cleanup whose endpoint close failed bin/fm-teardown.sh discarded both the exit status and the stderr of every fm_backend_kill call, so a close that genuinely failed was indistinguishable from one that succeeded. Teardown continued past it, deleted the task's durable records, returned its worktree, and reported the cleanup as completed. The deleted metadata is the only record of which endpoint belongs to the task, so such a close did not merely leave a stray session behind, it stranded one: nothing was left on disk naming it. The adapters could not carry that signal either. Driven against the real code, every backend arm returned 0 for a genuine failure exactly as it did for an already-exited endpoint, so there was nothing for the four call sites to propagate even once they stopped swallowing it. The tmux arm now resolves a close that did not succeed against the window's exact recorded identity, since kill-window fails the same way for a window that is gone and one that is still there. The Orca arm reports a close its missing CLI never attempted. Both stay silent for an endpoint that is already legitimately gone, and the remaining arms are unchanged: their close-command timing cannot be established without the real Zellij, Orca, and cmux binaries, and a gate that refused ordinary cleanup of an already-exited session would be worse than the defect. docs/verification/runtime-backends.md records what each backend can prove. A reported close failure now reaches teardown's existing retain-and-stop refusal before the records naming the endpoint are removed, matching where the Herdr confirmed-gone gates already sit for the same hazard, and the retained records let a rerun finish once the close works. * no-mistakes(review): refuse unreadable tmux close re-read; honor --force override * no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close * no-mistakes(document): document endpoint-close refusal in its backend and retirement owners * no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree
* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that brings Calm to Claude Code: the sailboat replaces the stock working row through a Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration, and canonically classified operational user rows draw at zero height. /calm is registered by the hooks module itself and toggles the same per-home config/calm preference the Pi extension uses, so one choice applies on either harness; rows redraw retroactively on toggle and stay hidden across claude --continue. The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS flag is on. Nothing sets that flag in any settings file, and the plugin carries no command file, skill, agent, or classic hook, so it is a complete no-op while the flag is off. The trusted project auto-loads it through an .agents/skills symlink, the only path Claude Code scans for project plugins. Extract the working-ship geometry, bounce track, cadences, and freeze/resume state into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module imports from outside the plugin folder) and have the Pi widget paint that core's frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify operational rows through a port of bin/fm-operational-input.sh's classify command guarded by a corpus parity test against the shell owner. Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering, Raster packing, policy, classifier parity), the mod's own claude plugin test suites behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op, the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272. Docs: record the version-scoped Claude Code evidence and the three bounded gaps in docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md, and make the shared preference, layout, and contributor notes harness-neutral. * no-mistakes(review): Preserve colliding final replies and strengthen parser parity * no-mistakes(review): Preserve final replies and strengthen canonical parity checks * no-mistakes(review): Require exact function-hooks opt-in before Calm activation * no-mistakes(review): Clarify Calm module loading and activation boundaries * no-mistakes(review): Reset Calm presentation state across session starts * no-mistakes(document): Refresh Calm session lifecycle documentation * feat(calm): paint the Claude Code working ship in Claude's own theme colors The captain picked the "Claude native" palette for the Claude Code mod's Raster: every water cell takes the spinner blue of the active theme family (#93a5ff dark, #5769f7 light) and the whole boat takes the Claude orange of the stock spinner (#d77757), one water color and one boat color. The family follows the `theme` setting's prefix, read at load through $.config.list and re-read on a config.set of that row, with `auto` and custom themes falling back to the dark set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte. Rename the shared sprite's color classes from hue names to `water` and `boat`, since each harness now maps them to its own colors; geometry, motion, cadence, and the activation gate are untouched. Tests cover both palettes' packing and the family rule under Node, and the plugin kit drives every theme value, a theme change mid-session, the Calm-off pass-through, and inertness of the menu read while the flag is off. The docs describe the Claude Code colors and record the guard passing on 2.1.273. * no-mistakes(review): Use light palette for unresolved Claude themes * no-mistakes(document): Refresh Claude Calm verification evidence
…kunchenguid#4586) * fix(watch): honour a declared wait before wedge-escalating a quiet pane wedge_timer_check escalated on elapsed idle time alone. Nothing asked whether the worker had already said why its pane was quiet, so a lane that declared a bounded external wait climbed the escalation ladder for as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every repeat carried demand-deep-inspection - which by its own wording forbids re-absorbing on the run-step or pane state, so the supervisor could not use the evidence that was there either. The generated brief promises that declaring `paused:` buys the long recheck cadence instead of a wedge, but the timer was still reachable while that declaration stood: a crew that declares a wait and then has an active run or busy pane attributed to it is handed to the timer as provably-working. The declaration is what the worker said about its own silence, so it now outranks a liveness verdict that only says something is running. The consult runs in the at-threshold branch that was about to escalate, beside the worktree walk already there, and costs one status-line read. Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS recheck the declared-wait absorber already uses, so the wait is still rechecked and cannot rot invisibly. Which verb declared it decides the wording, because the two block on different people: a `paused:` wait is owed by an external dependency and asks the reader to confirm it still holds, while a `captain-held:` transfer is owed by the captain reading the recheck and asks them to answer or release the hold. A hold is not rechecked at all while the away-posture record exists, as on every other captain-held path, and that absorb arms no throttle so the recheck is owed in full on return. A declared clearing time that has already passed stops counting, and a lane that never declared one keeps the identical escalation schedule, reason, count and demand-deep-inspection wording, so detection and its worst-case time are unchanged. The deferral restarts the idle timer rather than cancelling it, so a lane that stops waiting escalates again within one threshold. A lane quiet because its own validation run is parked at a gate awaiting a human decision is deliberately out of scope: reading that state needs a signal carrying who the wait is on and what clears it, rather than one inferred from a parked verdict that also covers gates awaiting the crewmate itself. Tests pin both directions for each case and were each confirmed to fail with the consult removed. * no-mistakes(document): docs: honour declared waits in stale-escalation docs
* fix(bin): derive passed PR state from PR record A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists. For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim. Fixes kunchenguid#4607 * no-mistakes(review): Add bounded GitLab merge-request state reads * no-mistakes(review): Preserve network-free inactive crew-state scans * no-mistakes(document): Document PR record readers in shared library
kunchenguid#4627) * fix: restore published contribution follow-up (Fixes kunchenguid#4469) * fix(review): Fix contribution freshness and merge actor routing * fix(review): Restore issue triage and scope contribution follow-up * fix(test): test: assert one wake per contribution signal * fix(document): Document contribution follow-up * fix: restore truthful terminal delivery evidence * fix(review): Disclose unsupported contributions and deduplicate watcher wakes * fix(review): Preserve unmeasured unsupported contributions across Bearings * fix(review): Deduplicate shared contribution wakes and isolate diagnostics * fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower
…nguid#4658) * fix(bin): make a remote-reply document gap self-clearing and re-attemptable A remote mate's undelivered document raised a keyed `blocked` decision that nothing could ever resolve, and any `data/*.md` substring in any mirrored line was an unconditional fetch instruction. A mate announcing a report it had not written yet therefore manufactured a permanent, factually false blocker, and its own explanation of the false alarm manufactured more. The reader has no permanence vocabulary: a report still being written refuses exactly like a path that will never exist. So an undelivered document is now a durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`, re-attempted on the next delta and on the channel's own quiet poll, and retired with a matching `resolved` line naming the local copy once it arrives. The cursor still advances and no delta stalls on one bad pointer. Only a structured `report=data/....md` pointer now offers a document, so a path merely mentioned in prose - including one under another home's mirror tree, which is provably not that mate's to serve - is never fetched. Offers are deduplicated across the whole delta, the escalation names each missing document once and carries the reader's own reason instead of discarding it, and a strictly increasing notice ordinal keeps a later escalation from being swallowed as duplicate bytes. A mirrored line still lands once whichever pointer form it was first written under. * no-mistakes(review): Require structured pointer token boundaries * no-mistakes(review): Unify boundary-safe pointer extraction and rewriting * fix(bin): identify a mirrored line independently of its delivery state Two defects in the boundary-safe pointer work. The at-most-once check compared only the all-remote and all-local renderings of a line, so it could not recognize a mixed one. A line offering two documents where only the first was deliverable mirrored as local-plus-remote; once the second arrived, a cursor-loss whole-log recapture rendered the same line all-local, matched neither alternate, and mirrored a second time. A line's identity is now the canonical form every boundary-valid pointer would take once delivered, derived by the same parser that does extraction and rewriting, so it no longer depends on which documents happened to be deliverable at the time. The pointer map was passed to awk through the process environment. A delta may carry up to the configured 1 MiB bound, and an expanded map of delivered pointers can exceed the platform's exec argument limit, so awk would fail to start; because no caller checked, the empty result would have been appended as blank lines while the cursor advanced past dropped status content. The map now travels in a file, and every call site checks the exit status and stops the ingest rather than committing a delta it could not render. Both passes now run once per stream instead of twice per line. * no-mistakes(review): Abort ingest when document pointer extraction fails * no-mistakes(review): Exclude structured cross-home pointers from document transfer * fix(bin): fail open on an undeliverable remote document instead of tracking it Narrow the remote-reply document fix to the scope the diagnosis actually requires, as decided after measuring a simpler alternative. A document the reader cannot deliver now fails open. The mate's line is mirrored with its own pointer, the cursor advances, and one unkeyed note carries the reader's reason. A note never enters the open-decision fold, so it cannot stand open the way the original keyed block did - which removes the never-clearing false blocker by construction rather than by resolving it. That makes the durable self-clearing obligation unnecessary, so it goes: the per-mate pending-documents record, its notice ordinal and resolved announcements, and the poll-side retry. Canonical line identity goes too, and with it a way to silently drop a genuine status line; mirroring is back to at-most-once on exact bytes. The cross-home exclusion goes as well: under fail-open a cross-home report= either fails harmlessly or is a nested remote report this mate genuinely holds, which is now relayed again. Kept: fetching only on a structured report= pointer, the boundary-correct parser, the file-based rewrite map, and checked extraction and rewrite exit status. The parser now scans behind a sentinel byte so a rejected candidate can no longer give the text right after it a false leading boundary. The reported incident is covered end to end: a report path announced in prose before it exists raises no decision, and the report still arrives through the ledger publisher's structured offer once written. * no-mistakes(review): Preserve source-line identity across remote reply replays * no-mistakes(document): Document remote reply transfer and replay semantics * no-mistakes(lint): Fix staging truncation lint checks
* Preserve substantive Calm mid-turn text * no-mistakes(review): Distinguish newline-preserved replies from short narration * no-mistakes(document): Document Calm mid-turn preservation boundaries * no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`
…#4656) * fix(bin): re-record PR poll identity after a volume device renumber (Fixes kunchenguid#4260) A volume remount can renumber the state filesystem's st_dev while every inode and byte stays the same; APFS does this across a reboot. A poll registration records its sidecar and check as device:inode, so every poll armed before the remount failed strict validation and the watcher refused all of them as unauthenticated state checks until each was re-armed by hand. There are two device comparisons. fm_pr_private_file_valid compares a live file's device with the state directory's device read in the same invocation: it refuses a file that is not on the state directory's own filesystem and already survives a renumber, so it is unchanged. The registration's recorded identity versus the live identity (from kunchenguid#556, reused by the kunchenguid#932 retirement receipt) binds the registration to the exact files published in its own transaction; its device part is what breaks. When strict capture fails, the watcher now proves the device is the only difference: every other artifact check passes (template bytes, both hashes, private mode, single link, live device, metadata), both recorded identities name one device, and each recorded inode equals its live inode. Only then, under the task's control lock, does it rewrite the two identity lines, repeating the whole proof and comparing the registration's file identity and bytes just before the rename, and then capture strictly again. A swapped, altered, re-moded, relinked, split-device, or foreign-device artifact still fails a proof and is still refused, and a pending retirement receipt blocks the rewrite. Reproduction: on macOS a poll armed on an APFS disk image that was detached and re-attached behind another image moved st_dev 16777239 -> 16777243 with inodes, bytes, mode, and link count unchanged; the real watcher refused it on main and reports its merge with this change. The portable regression test rewrites a real registration's recorded device and drives the watcher. Not changed here: the status presentation cursor keys rows by its own device:inode identity in bin/fm-classify-lib.sh, a different helper that needs its own fix; a retirement receipt left by a reboot between its publication and removal still names the old device and stays refused; custom check trust binds only a content hash and is unaffected. * fix(review): Serialize PR poll publication writers * fix(review): Bound PR poll publication lock scope
…llow-up to kunchenguid#4627) (kunchenguid#4661) A budget that expires partway through an observation no longer records an error or prints the unavailable wake; the URL keeps its prior record and is observed first next poll. forge() flags budget exhaustion at the point it refuses, or when a read is killed at the budget's own deadline, so a genuine forge failure still records the error and wakes. Each distinct URL is now observed once per poll and applied to every owning task.
…kunchenguid#4680) * fix(bin): clear parent pending-replies on local secondmate retirement Local secondmate teardown left resolved parent pending-reply records behind after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced retirement while any reply for that id is still unresolved, and delete every matching record plus its delivery confirmation after a successful local or remote retirement, matching the remote cleanup path. * no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup * no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung * no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern * no-mistakes(review): Pending-replies Basename und corr_id abgleichen * no-mistakes(document): Clarify forced retirement pending-reply cleanup --------- Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>
kunchenguid#4677) * fix(bin): accept Orca's composite worktree id at teardown Teardown refused every Orca-backed task because the endpoint validator checked orca_worktree_id with the simple-atom rule meant for tmux-style window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca returns that id as `<orca id>::<absolute worktree path>`, so the colon and slashes in every real value made validation fail and finished Orca tasks could never be cleaned up. Validate the field as the composite it is: both halves of the first `::` split present, the path half absolute, and no embedded newline, carriage return, or tab. The terminal field keeps the atom check, which is correct for it, and no other backend's validation changes. The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca never returns, which is why the suite passed a check the real value fails. They now carry the composite form, so the tests exercise the real value. * no-mistakes(document): name Orca's repo id in the composite worktree id * no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points
* feat(bin): add opt-in typed dispatch resolution through typesafe.ai Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or scout profile from a written brief with typesafe.ai's System One model: one Choice question over the rules' `when` texts, then the confidence floor, the rule's `approval` and `floor`, each profile's `provider` and `floor`, one quota-axi snapshot, and the spendPriority argmax all in code. It is off unless TYPESAFE_API_KEY is in the environment or the home's gitignored .env; off means one stderr line, exit 0, and no network call, so firstmate dispatches exactly as before. The key reaches curl on a file descriptor, never argv. Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates the four new optional dispatch fields. Document the schema, the operator contract, the AGENTS.md intake step, and the live and benchmark evidence. * no-mistakes(review): Harden typed dispatch resolution and quota bounds * no-mistakes(review): Validate dispatch floors and ranking evidence * no-mistakes(review): Tighten dispatch response and floor evidence * no-mistakes(review): Neutralize none matching and resolve defaults locally * no-mistakes(review): Preserve providerless profiles outside typed resolution * no-mistakes(review): Validate response usage and reject duplicate profiles * no-mistakes(review): Escalate unverifiable floors and validate probabilities * no-mistakes(review): Validate probability mass and unknown profile floors * no-mistakes(review): Simplify resolver interface and preserve fallback routing * no-mistakes(review): Fix constants and rank partial quota evidence * no-mistakes(review): Add authoritative provider mapping and enforce explicit providers * no-mistakes(review): Declare provider for documented Pi profile * no-mistakes(review): Validate provider identifiers and support Gemini dispatch * no-mistakes(review): Strictly anchor provider identifiers * no-mistakes(review): Validate selectors and preserve fallback candidate evidence * no-mistakes(review): Gate typed validation and harden resolver evidence * no-mistakes(review): Preserve opt-in routing and harden candidate evidence * no-mistakes(review): Prioritize known exhaustion over quota uncertainty * no-mistakes(review): Isolate API secrets and preserve no-key diagnostics * no-mistakes(review): Fallback safely when dispatch rules are absent * no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets * no-mistakes(document): Document typed dispatch safety and fallback behavior
…n decisions aren't lost (kunchenguid#3753) * test: reproduce buried status declarations in shared readers * fix: share status event reads and preserve open blockers * fix: retain terminal scout and ship status declarations * no-mistakes(review): Fix status chronology, legacy completions, and reader performance * no-mistakes(review): Share terminal decision reconciliation across fleet snapshots * no-mistakes(review): Unify terminal supersession across cached folds and consumers * no-mistakes(review): Filter per-key status history while preserving terminal chronology * no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells * no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses * no-mistakes(document): Document latest-event status read and kind-scoped fold cursor * no-mistakes(lint): Quote literal done in test for-lists for SC1010 * ci: expect 19 snapshot/fleet-view tests This branch adds a fleet-snapshot regression, so the stock macOS Bash lane's hardcoded guard of 18 'ok - ' lines fails on the new count. Bump the guard and its message to 19. * no-mistakes(review): Restore multiline child outcome reporting * no-mistakes(review): Select ledger terminal events through bounded shared reader * no-mistakes(review): Report newest open decision instead of preferring blocked * no-mistakes(review): Require colon before ship/scout terminal supersession in fold * no-mistakes(review): Gate socket-down override on latest event; drop lock matrix * no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions * no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold * no-mistakes(test): Update fleet-view expectations to newest-open-decision rule * no-mistakes(document): Align status-read docs with fold-resolved crew state * no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers * no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree * test: fold terminal-cleanup snapshot coverage into the completed-scout case Keep the ship/scout/secondmate supersession assertions without adding a nineteenth top-level fleet-view test, so CI can stay at the upstream suite count. * no-mistakes(document): Clarify socket-down override expiry in architecture doc * ci: retrigger flaky contribution check
…nchenguid#4689) * fix(spawn): launch codex crewmates with codex's hook layer disabled A freshly launched Codex worker never reached its instructions. Codex stopped it on an interactive "Hooks need review" modal whose selection sits on "Review hooks", which is neither trusting nor declining. Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow navigation, so the selection cannot be moved, and pre-accepting the prompt by writing Codex's own trust store would record an operator consent that was never given. The hooks are the machine's own ~/.codex/hooks.json plus any project's .codex/hooks.json. A crewmate needs neither: its turn-end signal is the -c notify= program on the same launch, and Firstmate's project hooks are primary-session infrastructure that stands down in a child worktree. Crewmate and scout launches now pass --disable hooks. That is the opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted hooks; disabling the feature runs none of them and leaves the operator's ~/.codex untouched. An unknown feature name is a hard Codex error, so a release that drops the flag fails the launch loudly instead of silently restoring the modal. A secondmate is a primary in its own home and keeps the project hooks its turn-end guard and session-start digest ride on. Verified on codex-cli 0.151.0: the modal is gone and the turn-end notification still lands. This unblocks the second review that every finished pull request is supposed to get. Fixes kunchenguid#4673 * no-mistakes(review): Fix contradictory hook count in Codex verification record
…d#4669, Fixes kunchenguid#4670) (kunchenguid#4710) * fix(bin): settle terminal contributions and wake once per read-failure episode A contribution whose last good observation is merged or closed is final: poll no longer re-reads it, projection keeps it fresh, and a stale error recorded beside it is cleared once. A genuine forge-read failure on an open contribution still records its error on every cycle but prints the unavailable wake only when it starts a failure episode; a successful read ends the episode. Open PRs linked from done tasks keep being observed. The false unavailable beside a complete observation was budget exhaustion mid-observation, already fixed by kunchenguid#4661. * fix(review): Settle terminal contribution owners * fix(review): Deduplicate shared contribution failure episodes * fix(test): Preserve settled terminal contribution records
* fix(crew-state): select authoritative validation runs by identity Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row. Refs: kunchenguid#3215 * fix(review): Resolve same-branch run identities beyond capped history * fix(review): Fix run-selection compatibility, races, and worker-state fallbacks * fix(review): Limit run validation to the requested branch * fix(test): Anchor AXI fixtures and document remaining live evidence gaps * fix(document): Clarify run selection documentation and capture ownership * fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation
* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work MAIN answered a supervision-branch outcome for completed captain-requested work (implementation done, PR ready for review and merge approval) with "Captain, shipshape.", reading section 9's no-action reply as covering it and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no captain-facing response is owed". Section 9 now limits the shipshape reply to true no-ops (idle re-read, empty heartbeat, consequence-free acknowledgement) and requires a short outcome response naming what finished and what word is needed whenever requested work finishes or a result needs the captain's word, even when a transcript entry already shows the substance. The Pi protocol's re-emit rule now says it bounds repetition only, and carries a worked example of the ready-for-review outcome whose correct processing turn a shipshape reply fails. No executable contract evaluates the content of MAIN's captain-facing reply, so the regression is the protocol example in the owner doc rather than a text-match test. * no-mistakes(document): Clarify captain-facing outcomes versus no-ops * docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line The document step condensed the Pi protocol's re-emit rule and dropped the worked example of a finished, ready-for-review outcome whose correct processing turn a "Captain, shipshape." reply fails. That example is the contract's regression: no executable contract evaluates the content of MAIN's captain-facing reply, so the owner doc's example is the test case. Restore it directly under the re-emit rule, prefixed as a regression example that is kept verbatim and never condensed or summarized away. * no-mistakes(review): Clarify captain outcome and decision-word requirements * no-mistakes(document): Clarify captain-facing completion outcomes * docs(pi): require the PR URL in the visible captain-facing outcome reply Captain review on the regression example: drop the sample reply string and say only that the ready-for-review outcome requires relaying a captain-facing outcome response, not just "Captain, shipshape.". Fold in the visible-PR-handoff failure seen this session: after the branch outcome reporting this fix green, MAIN's visible reply was only "Awaiting your merge call." with no PR URL, leaning on the dim anchor. Section 9's URL rule now also covers a review or merge ask and names the visible reply as where the URL goes, sourced from the ready status, pr= metadata, or the supervision branch's summary and never left to a transcript entry. The Pi protocol adds the same-way failure and places the captain-facing text in the final visible assistant reply after the fm_branch_processed call, because Calm hides assistant text emitted in the same step as a tool call as a working note. Investigation verdict, evidence in the PR comment: no recent PR caused the handoff failure; Pi has hidden same-step pre-tool assistant text since kunchenguid#2339 (2026-08-13), kunchenguid#4655 changed only the Claude Code mod, and kunchenguid#4658 touched only remote report transfer. * no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs * no-mistakes(document): Clarify captain-facing supervision outcomes * docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule The consolidated section 9 URL rule narrowed its trigger to a review or merge ask, dropping the "whenever a PR is mentioned" catch-all from kunchenguid#3648 that keeps every PR URL copied from a durable record and never assembled from memory. Restore that trigger as a union with the review or merge ask so the one consolidated rule covers both.
* Fix foreign-owner turn-end supervision loop * no-mistakes(review): Scope foreign-owner safe exit to Claude guard * no-mistakes(document): Document Claude foreign-owner safe exit
…orb (kunchenguid#4778) Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty indexed array as an unbound variable and aborts the shell. In signal_turnend_panes_churned() the missing_keys loop was reachable with an empty array whenever every churned key already held a fresh .churn-since-* marker (a second churning turn-end inside an open deferral window), so each watcher cycle died about half a minute in and supervision restarted endlessly. The created_keys rollback loops had the same latent crash on their error paths. Audit of bin/ for the same pattern found one more confirmed-reachable case: remote_handoff's noncanonical-body scan iterates to_move, which is empty when a retried remote handoff finds every key already staged in the outbox. All other "${arr[@]}" sites are either count-guarded, guaranteed non-empty by construction, or unreachable while empty. Guard the three reachable expansions with the repo's existing "${arr[@]+...}" idiom. New regression test drives a real watcher through the all-marked churn path; the macos-stock-bash CI lane runs it under real /bin/bash 3.2 via FM_TEST_ONLY.
… lock. (kunchenguid#4783) The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it. Co-authored-by: Cursor <cursoragent@cursor.com>
* docs: require complete final responses across harnesses * no-mistakes(document): Document complete final replies for Grok Bot * docs: point Grok replies to the shared contract owner * no-mistakes(review): Clarify final recap without batching decision asks
* fix(calm): preserve substantive Pi mid-turn text * no-mistakes(review): Preserve substantive Pi Calm text per block * no-mistakes(test): Cover shared Calm preservation boundaries behaviorally * no-mistakes(document): Consolidate Calm preservation documentation
…guid#4799) * Handle Kimi workspace trust dialog * no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers * no-mistakes(review): Gate Kimi ready on any trust marker and clean captures * no-mistakes(review): Read visible pane for Kimi trust and ready gates * no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate * no-mistakes(review): Harden Kimi viewport capture and trust dialog detection * no-mistakes(document): Document Kimi spawn refusal on cmux and Orca
…er (kunchenguid#4775) * fix(bin): report a record whose agent is gone once instead of escalating forever The wedge escalation path never asked whether there was still an agent to be wedged. A wedge is something stuck that might recover, so re-alarming it earns its cost; an agent that is gone never moves again, its pane never churns, the idle timer never resets, and the escalate path clears its own timer and re-arms with nothing bounding the count. Observed on a live fleet: two finished lanes reached 226 and 203 consecutive escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400 notifications a day from two lanes with no agent running at all. On one, fm-control.sh exit answered already-stopped and fm-crew-state.sh read "failed - run failed". Closing the Herdr pane did not stop it either: with the pane genuinely gone and herdr pane read returning pane_not_found, the count kept climbing, because the poll is driven by the record's window= line rather than by the pane. The cost is not the repetition but that it drowns the alarms that matter. fm_backend_agent_state already separates a thinking agent from a gone one at process level. In the branch that was about to escalate, read it once and treat only its two recovery-grade verdicts - dead (endpoint present, no agent in it) and missing (endpoint authoritatively absent) - as proof, reporting that record once and not re-escalating it while it stays that way. Every other verdict, including alive, ambiguous, unreadable, unverified, and a read that failed outright, keeps the identical schedule, reason, and escalation count, so a genuinely wedged live agent is unaffected. The probe costs at most one backend read per window per threshold, the same budget the declared-wait consult and the worktree write probe already take. The report decides nothing about the record's fate: both lanes still held unlanded work and teardown refusing them was correct, so retiring, relaunching, or cleaning up stays with the supervisor. The once-only marker is owned entirely by that function and is dropped by the same read the moment the endpoint stops reading gone, so a replacement launched into the same window escalates normally and its own later death is reported again. Related, and not closed by this: kunchenguid#4412, kunchenguid#4482, kunchenguid#4316. Tests drive the real watcher against a record whose endpoint does not exist and pin both directions: dead and missing report once and never advance the count across later thresholds, while alive, ambiguous, and unreadable endpoints keep escalating with the identical reason and a climbing count. * fix(bin): bind the once-only dead report to the pane it reported Review of the parent commit found a reachable sequence where a later death in the same window lost its promised report. The marker was keyed on the verdict string alone and dropped only when a threshold probe read a non-gone verdict, but probes run only at thresholds: a replacement launched into the same window that dies without ever being probed alive - it crashes at startup, or works and then crashes - was absorbed by the previous death's marker. The pane's first sight yielded only the generic stale wake and every later threshold matched the stale marker, so the second death never got the detailed once-report that both the function's own comment and docs/architecture.md promise. Record the verdict together with the pane hash it was reported for, and absorb a repeat only while both still match. A replacement churns the pane, which resets the stale suppressor, wedge timer, and escalation count while no reset site touches this marker, so the pane half is what tells the second death apart from the first. The live-probe drop stays as it was. Clearing the marker at those reset sites instead would re-open unbounded re-alarming for a dead pane whose display ever ticks, which is the exact defect the parent commit exists to close. The noise bound is unchanged: an unchanged dead pane still absorbs on every later threshold and never advances the escalation count, and every verdict short of proof still escalates exactly as before. * no-mistakes(review): Key the dead-record once-marker on the busy incarnation token * no-mistakes(document): Document dead-record escalation cap in stale-pane config entry * no-mistakes(document): Add busy-state inventory line to AGENTS.md * no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path
…id#4854) Captain holds have no due semantics and are a hold kind, not a Beads issue type. The create path now waives due.required and maps to native type task. Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(bin): launch every spawned agent with the compact adviser disabled Every crewmate, scout, and secondmate Firstmate launches now starts with COMPACT_ADVISER_DISABLE=1, on a fresh spawn and on a relaunch alike, so an unattended session never activates the compact adviser. The value is unconditional: no configuration file gates it and there is no override, unlike the trace carrier beside it. Three carriers deliver it, because no single one covers every launch shape. The pane shell receives an export beside GOTMPDIR, so the agent's own children inherit it too. The launch command carries an explicit assignment, prepended outermost so it wins over any ambient value the pane already held. The cleared launch environment sets it again at the `env -i` boundary and keeps COMPACT_ADVISER_DISABLE in the fixed operational floor, which is what preserves the switch when config/launch-env-allowlist empties the environment, and what delivers it on a remote host that never had the value. bin/fm-control.sh relaunch, the bootstrap secondmate relaunch, and the remote secondmate transport all rebuild their launch through bin/fm-spawn.sh, so they inherit the same floor. The captain's own primary session is untouched. The two new suites drive the real spawn and then execute the launch command the pane actually received, with the harness replaced by a probe that prints its own environment, rather than matching script text. They cover ship and secondmate launches with the allowlist absent and enabled, the pane export and its ordering, fm-control.sh relaunch, and the full parent to remote-host chain. * no-mistakes(review): Export compact-adviser disable across compound launches * no-mistakes(document): Document spawned-agent compact-adviser environment guarantee
…henguid#4894) * fix(bin): let a background Claude session keep owning its session lock Session-lock ownership was decided by process ancestry alone. Under an unattended Claude session the model loop runs in a transient bg-spare bridged to the front-end by a shared daemon; when that bridge is recycled the contiguous claude-named ancestry from a hook to the recorded owner breaks while the owner pid stays alive, so the Stop auto-arm stood down as a foreign live owner, the turn-end guard ended every turn with its read-only diagnostic, and fm-lock.sh refused - a self-sustaining outage until restart. Ownership is now ancestry membership OR a trusted same-session id, never id-first: - fm-session-lock-lib.sh accepts CLAUDE_CODE_SESSION_ID only when CLAUDE_PID is a Claude-shaped member of the current contiguous run, compares it against the id recorded in state/.lock-session, and requires the recorded pid to still be a live harness. No id, no sidecar, an untrusted id, a different id, or a dead recorded pid leaves the ancestry verdict unchanged. Ids are never read from ps argv. - fm-lock.sh accepts a same-session holder at both refusal sites, writes, refreshes, and clears the sidecar only under its claim lock (including the early already-mine exit, skipped only while the deferred startup sweep leases that lock), keeps it byte-identical across a same-session confirmation, records CLAUDE_PID on lock line 1 for a session with a trusted id so a shared daemon or front-end that outlives the session never keeps a dead session's lock alive, never rewrites a live line 1 on a same-session confirmation, and names the recorded id in the live-owner refusal. - The .lock line-1 format is unchanged, so every reader that takes the whole first line as the pid keeps working; the guard's foreign-owner exit is unchanged and inherits the fix through the shared predicate. Tests: the ancestry suite drives the ancestry and id signals apart in a deterministic process table (asserting the divergence) and runs a real orphaned front-end/daemon/pty-host/spare tree through six phases with the real lock, auto-arm, and guard scripts; the foreign-owner repro keeps its negative control and adds a same-id positive control. Disclosure: no live unattended Claude background session ran on the verifying machine. The topology is documented by the real process listings in kunchenguid#3902, kunchenguid#2314, kunchenguid#3398, and kunchenguid#4066; coverage is the structural predicate plus the executable fixtures, not a live pass. Residual: bin/fm-sessionstart-nudge.sh keeps its own private ancestry walk (it only decides whether to print a nudge) and may nudge on a resume in the recycled case. Out of scope, deliberately: no structured lock format, no guard budget changes, no daemon-identity rejection, no fork lineage. * no-mistakes(review): Wait for claim lock; revert failed sidecars * no-mistakes(review): Revalidate ownership after wait; restore sidecars * no-mistakes(review): Roll back sidecar by publication phase * no-mistakes(review): Restore sidecar only if lock line is unchanged * no-mistakes(review): Trust session ids without a spelling allowlist * no-mistakes(review): Disarm sidecar rollback before backup cleanup * no-mistakes(document): Updated session-lock ownership documentation
* feat: park main under the away posture on Pi While the away-posture record exists on a Pi primary, the supervision branch takes every actionable wake, no processing turn opens on main, captain rows accumulate for the return brief, and main's standing authority relocates to the branch through the existing guarded scripts. - lib/fm-branch-dispatch.ts: read the record at every routing decision; while it exists claim check, decision-owned, and heartbeat rows too, keeping the two broken-queue vetoes; expose checkSeqs so a claimed check row lifts task scoping. - fm-primary-pi-watch.ts: offer every actionable row under the record; a declined wake and every watcher-failure alarm still reach main. - fm-branch-supervision.ts: drop the legacy .afk decline; append a fixed POSTURE: AWAY tail carrying the record's read-back verbatim per wake; open no processing request while the record exists, re-checked immediately before a request would open and at every run boundary; present the accumulated rows at the first run boundary after archive. - fm-lease-lib.sh: fm_lease_forbid_branch passes the branch for opted-in actions only while fm-afk-contract.sh validate succeeds on a confirmed live record; PR merge, fresh spawn, and decision answer opt in, local landing never does. - fm-send.sh: a --resolve-key naming an open needs-decision or captain-held task is a decision answer and meets the partition; blocked: keys stay steering. - fm-spawn.sh: enforce the record's spend cap for a fresh ordinary spawn by either actor; relaunches and secondmates exempt. - fm-branch-prompt.sh: fixed Postures section and the verbatim ask-user-authority policy; the prefix stays byte-stable. - fm-afk-return.sh: count what the away session handled from the store. - docs, afk skill, AGENTS.md stub: main parked on Pi, green merge gate absolute while away. - tests: watcher and branch extension suites, fleet-record, merge, and decision-answer suites cover the relocation, the vetoes, the tail, the parked processing turn, the cancellation, the re-presentation, and the spend cap; dated live-guard evidence recorded. * no-mistakes(review): Refuse branch merge after preflight archive race * no-mistakes(review): Fix away wake, spawn, and processing races * no-mistakes(review): Suppress parked processing; narrow away-only rejection * no-mistakes(review): Abort dedicated processing; gate branch spawn once * no-mistakes(review): Stamp away-only on the dispatch offer * no-mistakes(review): Treat invalid away records as spend-cap absence * no-mistakes(review): Drop spawn test hook; abort processing-opened runs * no-mistakes(review): Bind abort to opening prompt; cap-read absence * no-mistakes(review): Limit away branch spawn to queued work only * no-mistakes(document): Correct AFK posture documentation
* ci: simplify CI job timeouts to a three-tier policy Replace the scattered per-job timeout values (10m parallel, 25m lint, 30m serial, 10m macOS) with three readable tiers, each a hang tripwire with headroom rather than a packing estimate: - fast (5m): coverage guard, repo invariants, timing aggregate - normal (30m, one shared budget): lint partitions, portable parallel shards, portable serial shards, macOS stock Bash - heavy (Herdr only): 20m step tripwire on the family run so always() cleanup still runs, under a 75m job-level last-resort backstop The workflow's header comment states the policy and points at docs/fm-test-portable-shards.md "Timeouts", which now owns it, and each job names its tier beside timeout-minutes. tests/fm-ci-workflow.test.sh asserts the policy against the parsed workflow instead of the old per-job minute values: every job joins exactly one tier, exactly three distinct job-level values exist, the fast tier stays within 5-10 minutes, the normal budget stays at least double the modeled parallel lane sum reported by fm-test-run.sh --check-coverage, and the Herdr step tripwire stays below its job backstop with an always() cleanup after it. Concurrency supersession, shard counts, lane membership, and fail-fast settings are unchanged. * no-mistakes(review): Decouple the normal timeout from packing estimates * no-mistakes(review): Assert Herdr teardown follows the family run * no-mistakes(review): Pin Herdr family-run timeout to 20 minutes * no-mistakes(review): Ignore comments when identifying Herdr steps * no-mistakes(review): Identify Herdr steps by declarative ids * no-mistakes(document): Clarify authoritative three-tier timeout policy
…wal-2026-09-19 Brings in 52 upstream commits, including typed dispatch resolution (bin/fm-dispatch-resolve.sh), which resolves one concrete crewmate or scout dispatch profile from a written brief and stays inert unless TYPESAFE_API_KEY is present. Six files conflicted; each resolution keeps both sides' intent: - .agents/skills/afk/SKILL.md: kept upstream's Pi main-parking wake rule and the fork's herdr carve-out that bars the native tracked-background tool. - README.md: kept upstream's reworded /bearings and /updatefirstmate rows and the fork's /clickup row. - bin/fm-promote.sh: kept upstream's brief.md relaunch append and its PROMOTION_SHIP_SPEC variable, with the fork's evidence rule carried into that spec as a trailing item so upstream's own numbering is untouched. - bin/fm-test-run.sh: took upstream's eight-member pr-forge pool and its refreshed serial-weight hints, re-adding the four fork-only script hints. - docs/documentation-audiences.json: kept both maintainer-verification entries. - docs/fm-test-portable-shards.md: took upstream's refreshed provenance and kept the fork's note recording where its own hints came from.
…sent means absent The suite took TYPESAFE_API_KEY from the caller's environment, so on the machine that actually has the credential configured the absent-key case resolved a profile instead of staying silent and aborted the run before the other seventeen cases. Scrub the variable, and the private name the tool re-homes it under, at the top of the file; every case that wants a key already passes its own. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
…s/fm-pr-description-guard.test.sh case "a readable description printed a diagnostic: fm-contributions: data directory unavailable / contributions: observation not armed". Root cause: the fork-local test's make_case() fake home created only home/state. The upstream commits brought in by this renewal added an unconditional `fm-contributions.sh arm` to PR registration (bin/fm-pr-check.sh:193); arm refuses a home whose data directory is absent, so fm-pr-check.sh degraded to a STDERR diagnostic, and the test asserts clean STDERR. The fixture, not the upstream code, was wrong — upstream's own tests/fm-pr-check-security.test.sh:129 creates home/data. Invariant: a fake firstmate home in tests must carry the directories a real home has (state/ and data/) so scripts take their production path. make_case() is the single place this file builds a home, so all cases are covered by one edit. Checked the sibling arm call site (bin/fm-bootstrap.sh:1662, already guarded on [ -d "$DATA" ]) and the other four fork-local test files (fm-clickup-contract, fm-download-lib, fm-evidence-artifacts, download-helpers) — none invoke fm-pr-check.sh or contributions, so no sibling site remains. Change: tests/fm-pr-description-guard.test.sh:153 — added "$dir/home/data" to the fixture mkdir -p, with a comment stating why. Verified: the test fails on this host before the edit with exactly the CI diagnostic; after it, `bin/fm-test-run.sh tests/fm-pr-description-guard.test.sh` reports FM_TEST_SUMMARY total=1 failed=0 with all 22 cases ok
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
The captain asked, on 2026-09-19:
"In a separate worker update my firstmate fork and lets setup typesafe api key for jev, because it should be introduced in the latest firstmate"
The ask has two halves. This task is the first half only: bring the fork up to date with upstream, so the typed dispatch resolution feature arrives in it. The second half, the credential itself, is a local gitignored .env value in the captain's home and is held separately. It is not yours to create, invent, name a value for, or reference in code.
What the captain is after, in their words, is the feature that upstream added: bin/fm-dispatch-resolve.sh resolves one concrete crewmate or scout dispatch profile from a written brief using typesafe.ai's System One model (jev-latest), gated on a TYPESAFE_API_KEY. It does not exist in the fork yet. It arrives with this renewal, along with fifty-one other upstream commits.
State established at intake, in this repository:
The captain holds merge authority. Do not merge.
What Changed
bin/fm-dispatch-resolve.sh, which resolves one crewmate or scout dispatch profile from a brief via typesafe.ai's System One model and stays entirely silent unless aTYPESAFE_API_KEYis present (environment, else$FM_HOME/.envthrough the newbin/fm-env-lib.sh); alongsidebin/fm-contributions.sh,bin/fm-pr-state.sh,bin/fm-pr-reviewers.sh, the flag-gated Claude Code Calm mode under.claude/mods/firstmate-calm/, and their suites, docs, and recorded captures.fm-test-run.sh --changednow maps root-level prose andtests/captures/*the way it already mappedtests/fixtures/*(and its header comment matches the new arms), andtests/fm-dispatch-resolve.test.shscrubs an inheritedTYPESAFE_API_KEYso the absent-key cases stay absent on a machine that has the credential configured.Conflict record
The intake prediction named twelve files. Three of those twelve actually conflicted;
nine merged cleanly. Three files conflicted that the prediction did not name. All
fifteen are accounted for below, and every clean auto-merge was checked line by line
that both sides' added lines survived.
Predicted, and conflicted (3)
.agents/skills/afk/SKILL.md— upstream added its Pi main-parking rule to the awaysupervision block; the fork had rewritten the same block so the native in-pane
tracked-background tool is never used on the herdr backend. Kept upstream's new rule
inside the fork's herdr carve-out.
bin/fm-promote.sh— upstream restructured the promotion handoff (the ship spec nowappends to
brief.mdand is hoisted intoPROMOTION_SHIP_SPEC); the fork had addedthe evidence rule as a numbered item in the old inline heredoc. Carried the fork's
evidence rule into upstream's
PROMOTION_SHIP_SPECas a trailing item so upstream'sown numbering and its cross-reference to the brief's rule 6 stay correct, and rewrote
the fork's now-stale rationale comment to describe upstream's shape.
bin/fm-test-run.sh— two hunks. Upstream grew thepr-forgeserial pool from sixscripts to eight (with a matching isolation proof) and refreshed the entire portable
serial-weight table from new CI runs. Took upstream's pool and table wholesale, then
re-added the four fork-only script hints upstream has no way to know about.
Predicted, but merged cleanly (9)
AGENTS.md— fork: evidence-artifacts delivery and the/clickupskill contract;upstream: Pi away-posture parking, lock ownership after helper recycling, dead-agent
reporting. Disjoint sections.
CONTRIBUTING.md— fork: gate test-command ownership and the tightened PR-compliancecontract; upstream: mail-check hardening and Calm mode.
bin/fm-afk-return.sh— fork: the skipped-wake-drain catch-up gate; upstream: Pimain-parking under the away posture.
bin/fm-brief.sh— fork: the evidence rule; upstream: validation-round pauses andkeeping operator-address labels out of the no-mistakes intent.
bin/fm-claude-trust.sh— fork: the symmetric trust-decline check; upstream: treatingClaude Code's default external-imports flags as never declined.
bin/fm-pr-check.sh— fork: refusing a description that delivers evidence as a localpath; upstream: preserving PR merge polls across volume remounts.
bin/fm-pr-merge.sh— fork: passing--no-description-checkso the local-path guardstays with the ready report rather than blocking an authorized merge; upstream:
plan-gated 403 branch rules with no merge queue.
bin/fm-supervise-daemon.sh— fork: stopping the away daemon from blocking its owndelivery; upstream: honouring a declared wait before wedge-escalating a quiet pane.
bin/fm-update.sh— fork: keeping a rebind-all failure as a warning only; upstream:reconciling a redundant secondmate divergence during updates.
Not predicted, but conflicted (3)
README.md— upstream reworded the/bearingsand/updatefirstmaterows while thefork inserted a
/clickuprow in the same table region. Kept all three.docs/documentation-audiences.json— both sides added amaintainer-verificationentry at the same position (upstream
dispatch-resolve.md, forkevidence-artifacts.md). Kept both, in path order.docs/fm-test-portable-shards.md— upstream rewrote the provenance paragraph for itsrefreshed hints. Kept upstream's text and preserved the fork's sentence recording
where its own four hints came from.
Risk Assessment
â^^E Low: The change is a routine upstream-renewal merge whose authored content is limited to conflict resolution in 24 files; every conflicted file's delta was verified against the pure upstream delta (10 identical, 2 adaptations checked against their cross-referencing invariants), the intent-required typed-dispatch feature landed byte-identical to upstream, and repo-wide sweeps found no lost test registrations, dangling helper calls, conflict markers, or parse errors.
Testing
Stood the renewed scripts up the way the captain will run them and drove nine end-user scenarios plus two adversarial ones. The headline feature works live: with the key configured, a written brief comes back as one concrete
--harness 'codex' --model 'gpt-5.5' --effort 'medium'profile from jev, and with no key the tool prints a single off line and makes no call. The round-1--changedrefusal is resolved â^@^T the documented contributor command now completes and selects the documentation suite for prose, the readers of a shared capture, and the readers of a shared test executable, while a new unread test helper still refuses rather than silently selecting nothing. One genuine failure appeared and was fixed here: the feature's own suite inherited the operator's TYPESAFE_API_KEY, so its absent-key case was not hermetic on the very machine where the key is being set up; scrubbing the variable makes the suite pass from a shell with or without the credential, and the key appears in no transcript. The runner's own suite leaves one red on this host, reproduced identically at the pure upstream tip, so it is inherited rather than introduced â^@^T reported as a warning, not a blocker. No screenshot applies: every surface here is a CLI, and the transcripts are the end-user output.bin/fm-dispatch-resolve.sh <brief> --project demowith example rules â^@^T real api.typesafe.ai call returned model jev-1.13.0 andprofile: --harness 'codex' --model 'gpt-5.5' --effort 'medium'(dispaâ^@¦env -u TYPESAFE_API_KEY FM_HOME=/tmp/fm-empty-home bin/fm-dispatch-resolve.sh /dev/nullâ^@^T single stderr line naming both lookup sites, exit 0, no curl (01-gate-off-without-key.txt)bin/fm-dispatch-resolve.sh <brief> --project demoin this worktree â^@^Tstatus: escalate reason: no rules to match, exit 0 (dispatch-resolve-live.txt, real-repo-no-rules.txt)bin/fm-test-run.sh tests/fm-dispatch-resolve.test.shfrom a shell exporting the real key: exit 1 before the fix, 18 ok / exit 0 after (dispatch-resolve-suite-hermeticity.md)bin/fm-test-run.sh --changedcommandbin/fm-test-run.sh --list --changedâ^@^T exit 0, 222 scripts selected, including tests/fm-documentation-audiences.test.sh; the GROK_BOT.md-only case that refused in round 1 now selects the documentatioâ^@¦ZZ_NEW_PROSE_9f3a.mdâ^F^R selects tests/fm-documentation-audiences.test.sh, exit 0 (changed-mapping-probes.md)fm-test-run: no changed-test mapping for source path: ..., exit 2 (changed-mapping-probes.md)env -u TYPESAFE_API_KEY bin/fm-test-run.sh tests/fm-dispatch-resolve.test.shâ^@^T 18 ok, exit 0, same as the key-present run (dispatch-resolve-suite-hermeticity.md)bin/fm-test-run.sh tests/fm-test-run.test.shâ^@^T 37 ok covering the prose, capture, and shared-executable assertions added by the fix (owning-suite-differential.md)not ok - the timed-out script left grandchild <pid> runningreproduces in a clean clone checked out at upstream tip 2bcb88c on this host, and bin/fm-timeout-lib.sh is byte-identical to bothâ^@¦Evidence: Resolver driven live against api.typesafe.ai â^@^T escalate with no rules, one concrete profile with rules
Source: Resolver driven live against api.typesafe.ai â^@^T escalate with no rules, one concrete profile with rules
Evidence: Feature suite hermeticity: fails on a machine exporting the real key, passes after the scrub
Source: Feature suite hermeticity: fails on a machine exporting the real key, passes after the scrub
Evidence: Documented --changed verification before and after the mapping fix, with the GROK_BOT.md-only reproduction
Source: Documented --changed verification before and after the mapping fix, with the GROK_BOT.md-only reproduction
Evidence: Mapper probes: new root prose, shared capture, shared executable, and an adversarial unread helper
Source: Mapper probes: new root prose, shared capture, shared executable, and an adversarial unread helper
Evidence: Runner suite differential â^@^T branch vs pure upstream tip on the same host
Source: Runner suite differential â^@^T branch vs pure upstream tip on the same host
Evidence: Resolver stays inert without a key
Source: Resolver stays inert without a key
Evidence: Credential never reaches stdout, stderr, or any file
Source: Credential never reaches stdout, stderr, or any file
Evidence: Evidence index for the whole renewal validation
Source: Evidence index for the whole renewal validation
Pipeline
Updates from git push no-mistakes
â^\^E **intent** - passed
â^^E No issues found.
â^\^E **Rebase** - passed
â^^E No issues found.
â^\^E **Review** - passed
â^^E No issues found.
â^Z ï¸^O **Test** - 1 warning
bin/fm-test-run.sh:1711-bin/fm-test-run.sh --changedâ^@^T the verification command CONTRIBUTING.md:110 hands contributors â^@^T refuses outright on this branch: it exits 2 withno changed-test mapping for source path: GROK_BOT.md. The renewal modifies GROK_BOT.md, and neither side of the merge maps that path to a test family (0 mapping lines in the fork base 599ea9e, the upstream tip 2bcb88c, and the merge a15ec5f alike), so the mapping gap is inherited from upstream rather than introduced by conflict resolution. Scope is narrow and it self-heals: CI never invokes --changed (all eight ci.yml call sites use --lane/--family/--check-coverage/--aggregate-json), and once the captain merges, origin/main already contains the new GROK_BOT.md so it leaves the changed set and the command works again. Left unfixed deliberately â^@^T adding a mapping arm means editing upstream's table inside a renewal merge, which is the captain's call, not this step's.bin/fm-test-run.sh --changedcommandenv -u TYPESAFE_API_KEY PATH=<curl removed> FM_HOME=<sandbox> bin/fm-dispatch-resolve.sh brief.md --project pagerâ^F^R 0 bytes on stdout, onedispatch-resolve: offline on stderr, exit 0 (01-gate-offâ^@¦TYPESAFE_API_KEY= bin/fm-dispatch-resolve.sh brief.mdâ^F^R samedispatch-resolve: off, 0 bytes stdout, exit 0 (01-gate-off-without-key.txt, CASE B)bin/fm-dispatch-resolve.sh does-not-exist.mdwith no key â^F^R gate fires first, exit 0, no error (01-gate-off-without-key.txt, CASE C)model: jev-1.13.0 latency_ms: 729,status: clear,rule_2 confidence 1.0, singleprofile: --harness 'cursor' --model 'cursor-grok-4.6-medium'(02-live-resolution-real-api.txt)approval: captainescalates instead of auto-resolvingstatus: escalate,reason: rule requires the captain's explicit approval before dispatch, and noprofile:line is emitted at all (03-captain-approval-escalates.txt).env, the tool is no longeroffâ^@^T it calls the API and reportsstatus: error / http 401 authentication_error, eâ^@¦.envâ^F^Rstatus: clearand a resolved profile, proving the environment value was used (04-dotenv-activation-and-precedence.txt, CASE B)grep -rlFfor the live key across all 11 evidence files and the entire worktree â^F^R 0 hits each; a 401 error is surfaced without echoing the key (05-credential-never-leaks.txt)bin/fm-dispatch-resolve.sh --helprenders the operator interface; with noconfig/crew-dispatch.jsonthe run returnsstatus: escalate / reason: no rules to match, exit 0 (real-repo-no-rules.txt)bin/fm-test-run.shover the 8 suites covering them plus the feature's own suite â^F^RFM_TEST_SUMMARY total=8 failed=0 skipped_gate=0(conflicted-surface-suites.txt)bin/fm-test-run.sh --changedcommandbin/fm-test-run.sh --list --changedexits 2 withno changed-test mapping for source path: GROK_BOT.md; the path is modified by this renewal and mapped by neither merge parent (changed-selection-reâ^@¦bin/fm-dispatch-resolve.sh brief.md --project pagerwithenv -u TYPESAFE_API_KEY, curl stripped from PATH, and a sandbox FM_HOME â^@^T gate-off pathbin/fm-dispatch-resolve.shwithTYPESAFE_API_KEY=exported empty â^@^T adversarial: empty must count as absentbin/fm-dispatch-resolve.sh missing-brief.mdwith no key â^@^T adversarial: gate must fire before file validationbin/fm-dispatch-resolve.sh brief.md --project pagerwith the live key â^@^T real call to typesafe.ai System One, profile resolvedbin/fm-dispatch-resolve.sh brief-hard.md --project platformagainst anapproval: captainrule â^@^T adversarial: must escalate, must emit noprofile:lineenv -u TYPESAFE_API_KEY bin/fm-dispatch-resolve.shwith only an invalidTYPESAFE_API_KEY=line in$FM_HOME/.envâ^@^T .env activationbin/fm-dispatch-resolve.shwith the real key in the environment and an invalid placeholder in.envâ^@^T environment precedencegrep -rlF "$TYPESAFE_API_KEY"over every evidence file and the entire worktree â^@^T 0 hits eachbin/fm-dispatch-resolve.sh --helpand a run in this repository as shipped (noconfig/crew-dispatch.json) â^@^T escalates withno rules to match, exit 0bin/fm-test-run.sh tests/fm-dispatch-resolve.test.sh tests/fm-brief.test.sh tests/fm-update.test.sh tests/fm-pr-merge.test.sh tests/fm-claude-trust.test.sh tests/fm-afk-return.test.sh tests/fm-afk-contract.test.sh tests/fm-pr-check-security.test.shâ^@^T 8 suites, 0 failuresbin/fm-test-run.sh --list --changedon the renewal branch â^@^T reproduced the exit-2 refusalgrep -n 'fm-test-run.sh' .github/workflows/*.ymlâ^@^T confirmed no CI job uses--changedð^_^T§ Fix applied.
1 warning still open:
tests/fm-test-run.test.sh:1547-bin/fm-test-run.sh tests/fm-test-run.test.shends with one red on this host:not ok - the timed-out script left grandchild <pid> running(37 ok, 1 not ok). It is not from this branch.bin/fm-timeout-lib.shis byte-identical to both merge parents, and a clean clone checked out at the pure upstream tip 2bcb88c fails the same case on the same machine (34 ok, 1 not ok). Upstream CI is Linux; this looks like a macOS reaping difference in the per-script timeout path rather than a renewal defect. Recorded so the captain is not surprised by a red line when they run the runner's own suite locally after merging; no action taken here because fixing it means changing upstream's timeout library inside a renewal merge.bin/fm-dispatch-resolve.sh <brief> --project demowith example rules â^@^T real api.typesafe.ai call returned model jev-1.13.0 andprofile: --harness 'codex' --model 'gpt-5.5' --effort 'medium'(dispaâ^@¦env -u TYPESAFE_API_KEY FM_HOME=/tmp/fm-empty-home bin/fm-dispatch-resolve.sh /dev/nullâ^@^T single stderr line naming both lookup sites, exit 0, no curl (01-gate-off-without-key.txt)bin/fm-dispatch-resolve.sh <brief> --project demoin this worktree â^@^Tstatus: escalate reason: no rules to match, exit 0 (dispatch-resolve-live.txt, real-repo-no-rules.txt)bin/fm-test-run.sh tests/fm-dispatch-resolve.test.shfrom a shell exporting the real key: exit 1 before the fix, 18 ok / exit 0 after (dispatch-resolve-suite-hermeticity.md)bin/fm-test-run.sh --changedcommandbin/fm-test-run.sh --list --changedâ^@^T exit 0, 222 scripts selected, including tests/fm-documentation-audiences.test.sh; the GROK_BOT.md-only case that refused in round 1 now selects the documentatioâ^@¦ZZ_NEW_PROSE_9f3a.mdâ^F^R selects tests/fm-documentation-audiences.test.sh, exit 0 (changed-mapping-probes.md)fm-test-run: no changed-test mapping for source path: ..., exit 2 (changed-mapping-probes.md)env -u TYPESAFE_API_KEY bin/fm-test-run.sh tests/fm-dispatch-resolve.test.shâ^@^T 18 ok, exit 0, same as the key-present run (dispatch-resolve-suite-hermeticity.md)bin/fm-test-run.sh tests/fm-test-run.test.shâ^@^T 37 ok covering the prose, capture, and shared-executable assertions added by the fix (owning-suite-differential.md)not ok - the timed-out script left grandchild <pid> runningreproduces in a clean clone checked out at upstream tip 2bcb88c on this host, and bin/fm-timeout-lib.sh is byte-identical to bothâ^@¦bin/fm-dispatch-resolve.sh /tmp/synthetic-brief.md --project demowith the configured key (real api.typesafe.ai call)FM_CONFIG_OVERRIDE=<tmp with docs/examples/crew-dispatch.json> bin/fm-dispatch-resolve.sh /tmp/synthetic-brief.md --project demoenv -u TYPESAFE_API_KEY FM_HOME=/tmp/fm-empty-home bin/fm-dispatch-resolve.sh /dev/nullbin/fm-test-run.sh tests/fm-dispatch-resolve.test.shfrom a shell exporting the real key (before fix: exit 1; after fix: 18 ok, exit 0)env -u TYPESAFE_API_KEY bin/fm-test-run.sh tests/fm-dispatch-resolve.test.shbin/fm-test-run.sh --list --changedon the renewal branch (the CONTRIBUTING.md:110 command)bin/fm-test-run.sh --list --changed --base HEADwith only GROK_BOT.md touched (round-1 blocker reproduction)bin/fm-test-run.sh --list --changed --base HEADwith a new root prose file, a touched tests/captures/no-mistakes-v1.70.1/overview.toon, and a touched tests/fm-turnend-foreign-owner-repro.pybin/fm-test-run.sh --list --changed --base HEADwith an unreferenced new tests/ helper (adversarial fail-closed probe)bin/fm-test-run.sh tests/fm-test-run.test.shin this worktree and in a clean clone checked out at upstream tip 2bcb88c (differential)verbatim sweep for the real key across every evidence file and the whole worktreeâ^Z ï¸^O **Document** - 1 info
docs/scripts.md:9- The bin/ index is one row per script, but ~42 tracked bin/ files have no row, including bin/fm-dispatch-resolve.sh â^@^T the typed dispatch resolver this renewal brings in. The gap is not caused by this change: the same 42 are missing at the pure upstream tip (2bcb88c), and upstream authored the resolver without adding its row. Its contract is documented in docs/configuration.md#typed-dispatch-resolution-env-typesafe_api_key and docs/verification/dispatch-resolve.md, so nothing is undocumented â^@^T only the index is incomplete. Left alone here deliberately: adding fork-local rows to an upstream-owned index creates conflict surface at every future renewal. Worth raising upstream (or a one-time fork decision) rather than patching inside this merge.â^\^E **Lint** - passed
â^^E No issues found.
â^\^E **Push** - passed
â^^E No issues found.