Skip to content

Reconcile canonical upstream through c5131a33 - #75

Merged
timbarreto merged 41 commits into
mainfrom
reconcile/upstream-2026-09-22-c5131a33
Sep 23, 2026
Merged

timbarreto merged 41 commits into
mainfrom
reconcile/upstream-2026-09-22-c5131a33

Conversation

@timbarreto

@timbarreto timbarreto commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner

CI repair update

Current PR head: 4159f4965393456b988cd22f8c294c53edeacbab. Frozen upstream remains c5131a33a1e35a42e34733a5334fcc4e0225a656; the original two-parent reconciliation f1ab3008c86e657fc2813bded88da55909d45a06 remains reachable.

Current exact-head CI: 25/25 successful; 0 pending, 0 failed, 0 absent. Observed at 2026-09-23T00:47:43.6098730Z.
Keep this PR open and unmerged. Use a merge commit, not squash or rebase merge; auto-merge remains disabled.

The original head finished with 18 successful and seven failed checks. Those failures included seven test suites and a terminated Lint 2 process. Every failing test was reproduced in an isolated Ubuntu environment; all seven now have focused passing evidence. Three affected legacy suites use the shared named-case registry without changing their full-suite order.

Repaired subject CI owner Focused pass (seconds) Root cause / correction
Azure ready identity Behavior portable serial 2 4.650 Normalize only the new timestamp field while requiring the exact provider URL and outcome key.
Azure interrupted publication Behavior portable serial 3 13.261 Normalize only report timestamps; retain notification-before-marker, exact-identity retry, and no-duplicate assertions.
Stopped child cleanup Behavior portable serial 3 13.975 Match the existing cleanup diagnostic while still requiring timeout 124 and proof the TERM-resistant child is reaped.
Restored endpoint refusal Behavior portable serial 4 4.458 Assert the shared absence-proof refusal and preserve byte-identical metadata/brief plus no lifecycle action.
Large mail poll Behavior portable serial 5 3.858 Install the existing complete private-path fixture dependency closure; retain real large-output draining and exact summaries.
Inbox lock readiness Behavior portable serial 7 19.980 Distinguish a live unverified owner from a dead owner in human output, retaining unknown readiness and the untouched lock.
Compact adviser relaunch Behavior portable serial 9 37.204 Remove the no-op sleep that instantly fired the real watchdog; use existing relaunch fixture bounds, leaving production deadlines unchanged.

Production runtime, native privacy, guarded recovery, and all production deadlines are unchanged by the initial repairs. The runtime-file edit is a ShellCheck source boundary only. The compact-adviser check passed its complete suite after removing its no-op sleep.

Lint failure and preserved coverage

Lint 2 failed twice on the unchanged original head with exit 143, no ShellCheck diagnostic, and no uploaded telemetry. Two bounded local partition attempts, including identical bytes on native Linux storage, timed out. Isolated full-rigor probes then showed the watcher root growing to 6,478,968 KiB RSS and the new live-test import to 7,968,876 KiB before their 90-second bounds.
The correction follows the existing canonical-owner boundary pattern: the watcher no longer re-expands the pending-reply owner's backend/classifier/wake graph, and the live test does not re-import production graphs for static analysis. All these owners remain independently linted canonical roots, with the same full source-aware flags, two partitions, concurrency cap, and coverage checks. No lint mode, production deadline, or job timeout was relaxed. Exact-head CI, not the failed local partition attempts, must establish the complete lint result.

Bounded repair evidence

The user explicitly approved a fresh 2400-second repair allowance and later two additional targeted-lint timeouts without adding any seconds or authorizing another full-partition retry. After that four-timeout cutoff was reached and the remote clone failure repeated in CI, the user approved using only the remaining approximately 13 minutes for that specific failure, with at most two additional timeouts. Charged total: 1913.892 seconds; 4 timeouts out of the approved maximum of 6. No time was added. The original reconciliation allowance remains exhausted and unchanged.
The local Ubuntu image's uutils date returned nanoseconds for a millisecond request. After observing that host mismatch, the already-installed GNU date was selected only in the session-local tool PATH. The final focused results use GNU date; earlier runner duration fields are not reliable and the controller's elapsedMs is the timing authority. Node and pinned linters were installed only after missing-prerequisite failures, into the session namespace, with verified checksums. No shared OS configuration was changed.
The seven fixed test files passed source-aware lint; the completed compact-adviser repair was relinted. The existing deterministic lint-owner contract also passed. Both complete canonical lint partitions passed at the first repair head after the source-boundary correction.

First complete repair run and targeted retry

The first complete run at 05c50fef3ae048042f0e08166e33bd9a846acb46 had 24 successful checks and one failure. Both canonical lint partitions passed, as did all seven originally failing behavior cases. Serial 4 instead encountered a different failure in the unchanged remote compact-adviser suite: Git reported ENOENT while copying an object into the fixture's newly cloned remote home, followed on the first attempt by a nonempty-directory cleanup error. That same remote suite passed on the original reconciliation head; its provisioning and transport code were unchanged by the first repair.
Only serial 4 was retried on the identical commit, without changing code, weakening assertions, or increasing a deadline. It repeated the same clone failure. Both complete logs are retained; this is not dismissed as a transient infrastructure problem. The current table uses GitHub's latest check run for each exact-head producer; superseded failed attempts remain recorded separately.
A complete tracked-only native Linux replay passed, as did a provisioning-only 12-iteration trace and a second 12-iteration trace making automatic fixture garbage collection eligible. Those minimized probes stop after provisioning and are not substitutes for the full launch assertions. The full fixture with failure-only diagnostics and its source-aware lint also passed locally. CI diagnostics capture the failed object and source/destination state before rollback; local passes do not establish the cause of the CI failure.
At diagnostic head 8a3915f8fd9e5313ab5973277e5c37783e27d187, the wrapped compact-adviser fixture passed, but the same clone-copy ENOENT appeared in the independent fm-remote-secondmate-trace-context.test.sh suite in serial 9. The fixture wrapper was therefore removed. Head d3f546331fe5677422eb8160820e90ad47882232 instrumented only the shared owner's failure and rollback paths, leaving the real Git invocation and all success-path timing untouched. This was one shared failure signature, not a reason to weaken unrelated assertions or retry the full matrix locally.
All 25 checks passed at shared-diagnostic head d3f546331fe5677422eb8160820e90ad47882232, so that run produced no clone-failure facts. A further controlled source-repack probe also passed rather than reproducing the failure. The intermittent clone issue therefore remains an unproven follow-up, not a claimed fix or a proven Git/GC root cause.
All temporary fixture and production diagnostics have been removed. The final repair tree is byte-identical to the initial ten-file repair commit 05c50fef3ae048042f0e08166e33bd9a846acb46; no speculative change to Git cloning, recovery, runtime deadlines, or native safety remains. The table below belongs only to the cleaned final head.
The first run of cleaned head 4159f4965393456b988cd22f8c294c53edeacbab again passed 24 checks and failed serial 9 on the same remote trace-context clone signature. Only that shard was retried without a code change. An additional local probe rotating redundant loose objects while retaining a valid packed source also passed; it did not establish a root cause. The intermittent finding remains documented for follow-up rather than being described as repaired by the original fixture/lint changes.

Repair commands, outcomes, and timing ledger

All test invocations used FM_LIVE=0 and cache=false. The existing validate-local.mjs controller owned the budget and process-group deadlines; only focused cases were selected.

Phase Allowance (seconds) Charged elapsed (seconds) Timeouts
preflight 2266 1.198 0
red 2264 49.066 0
lint-red 2215 481.747 1
green 1734 66.181 0
green-gnu 1621 76.714 0
compact-green 1544 40.328 0
lint-native 1501 481.604 1
lint-diagnostics 990 190.701 2
remote-red 799 1.496 0
remote-red-lf 798 86.578 0
remote-trace 711 57.317 0
remote-gc 654 57.551 0
remote-diagnostics-green 596 75.598 0
remote-shared-diagnostics 521 26.820 0
remote-repack-red 494 3.868 0
remote-rotation-red 490 4.506 0

Prerequisites charged 180.501 seconds; native export 2.118 seconds. A failed probe against an already-removed temporary export is conservatively charged another 30 seconds and was not a test pass.

red/azure-ready-identity: failed; 4061 ms; ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_azure_ready_identity_survives_reconciliation","tests/fm-inactive-reconcile.test.sh"]
red/azure-interrupted-publication: failed; 6464 ms; ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_azure_poll_interrupted_publication","tests/fm-pr-check-security.test.sh"]
red/stopped-child-cleanup: failed; 13575 ms; ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_wait_deadline_reaps_a_stopped_child","tests/fm-watcher-lock.test.sh"]
red/restored-endpoint-refusal: failed; 3349 ms; ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_missing_endpoint_restored_unsafe_state_refuses","tests/fm-control-recovery.test.sh"]
red/large-mail-poll: failed; 2047 ms; ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_large_poll_output_is_drained","tests/fm-mail-check.test.sh"]
red/inbox-lock-wording: failed; 15280 ms; ["bash","bin/fm-test-run.sh","tests/fm-inbox.test.sh"]
red/compact-adviser-relaunch: failed; 2755 ms; ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_relaunch_rebuilds_the_switch","tests/fm-spawn-compact-adviser-disable.test.sh"]
lint-red/lint-partition-two: timeout; 480210 ms; ["bash","bin/fm-lint.sh","--partition","2of2","--telemetry","$EVIDENCE/ci-lint-partition-two.tsv"]
green/azure-ready-identity: passed; 4673 ms; ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_azure_ready_identity_survives_reconciliation","tests/fm-inactive-reconcile.test.sh"]
green/azure-interrupted-publication: passed; 15280 ms; ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_azure_poll_interrupted_publication","tests/fm-pr-check-security.test.sh"]
green/stopped-child-cleanup: passed; 13875 ms; ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_wait_deadline_reaps_a_stopped_child","tests/fm-watcher-lock.test.sh"]
green/restored-endpoint-refusal: failed; 4666 ms; ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_missing_endpoint_restored_unsafe_state_refuses","tests/fm-control-recovery.test.sh"]
green/large-mail-poll: passed; 3955 ms; ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_large_poll_output_is_drained","tests/fm-mail-check.test.sh"]
green/inbox-lock-wording: passed; 19586 ms; ["bash","bin/fm-test-run.sh","tests/fm-inbox.test.sh"]
green/compact-adviser-relaunch: failed; 2749 ms; ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_relaunch_rebuilds_the_switch","tests/fm-spawn-compact-adviser-disable.test.sh"]
green-gnu/repair-syntax: passed; 252 ms; ["bash","-c","for file in \"$@\"; do bash -n \"$file\" || exit; done","_","tests/fm-control-recovery.test.sh","tests/fm-inactive-reconcile.test.sh","tests/fm-inbox.test.sh","tests/fm-mail-check.test.sh","tests/fm-pr-check-security.test.sh","tests/fm-spawn-compact-adviser-disable.test.sh","tests/fm-watcher-lock.test.sh"]
green-gnu/azure-ready-identity: passed; 4650 ms; ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_azure_ready_identity_survives_reconciliation","tests/fm-inactive-reconcile.test.sh"]
green-gnu/azure-interrupted-publication: passed; 13261 ms; ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_azure_poll_interrupted_publication","tests/fm-pr-check-security.test.sh"]
green-gnu/stopped-child-cleanup: passed; 13975 ms; ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_wait_deadline_reaps_a_stopped_child","tests/fm-watcher-lock.test.sh"]
green-gnu/restored-endpoint-refusal: passed; 4458 ms; ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_missing_endpoint_restored_unsafe_state_refuses","tests/fm-control-recovery.test.sh"]
green-gnu/large-mail-poll: passed; 3858 ms; ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_large_poll_output_is_drained","tests/fm-mail-check.test.sh"]
green-gnu/inbox-lock-wording: passed; 19980 ms; ["bash","bin/fm-test-run.sh","tests/fm-inbox.test.sh"]
green-gnu/compact-adviser-relaunch: failed; 2860 ms; ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_relaunch_rebuilds_the_switch","tests/fm-spawn-compact-adviser-disable.test.sh"]
green-gnu/lint-owner-contract: passed; 3157 ms; ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_jobs_are_deterministic_and_complete","tests/fm-lint.test.sh"]
green-gnu/repair-source-aware-lint: passed; 8264 ms; ["bash","bin/fm-lint.sh","tests/fm-control-recovery.test.sh","tests/fm-inactive-reconcile.test.sh","tests/fm-inbox.test.sh","tests/fm-mail-check.test.sh","tests/fm-pr-check-security.test.sh","tests/fm-spawn-compact-adviser-disable.test.sh","tests/fm-watcher-lock.test.sh"]
compact-green/compact-adviser-full: passed; 37204 ms; ["bash","bin/fm-test-run.sh","tests/fm-spawn-compact-adviser-disable.test.sh"]
compact-green/compact-adviser-source-aware-lint: passed; 1452 ms; ["bash","bin/fm-lint.sh","tests/fm-spawn-compact-adviser-disable.test.sh"]
lint-native/lint-partition-two-native: timeout; 480156 ms; ["bash","bin/fm-lint.sh","--partition","2of2","--telemetry","$EVIDENCE/ci-lint-partition-two-native.tsv"]
lint-diagnostics/lint-partition-inventory: passed; 7160 ms; ["bash","bin/fm-lint.sh","--partition","2of2","--list-files"]
lint-diagnostics/new-launch-prompt-root: timeout; 90792 ms; ["python3","$EVIDENCE/ci-lint-resource-probe.py","tests/fm-launch-prompt-signals-live-e2e.test.sh"]
lint-diagnostics/watch-production-root: timeout; 90726 ms; ["python3","$EVIDENCE/ci-lint-resource-probe.py","bin/fm-watch.sh"]
remote-red/remote-compact-adviser: failed; 144 ms; ["bash","$EVIDENCE/ci-remote-suite.sh"]
remote-red-lf/remote-compact-adviser: passed; 85080 ms; ["bash","$EVIDENCE/ci-remote-suite.sh"]
remote-trace/remote-provision-trace: passed; 55931 ms; ["bash","$EVIDENCE/ci-remote-suite.sh"]
remote-gc/remote-provision-auto-gc: passed; 56233 ms; ["bash","$EVIDENCE/ci-remote-suite.sh"]
remote-diagnostics-green/remote-compact-adviser: passed; 72551 ms; ["bash","$EVIDENCE/ci-remote-suite.sh"]
remote-diagnostics-green/remote-fixture-lint: passed; 1248 ms; ["bash","bin/fm-lint.sh","tests/fm-spawn-compact-adviser-disable-remote.test.sh"]
remote-shared-diagnostics/remote-failure-diagnostics: passed; 14473 ms; ["bash","$EVIDENCE/ci-remote-suite.sh"]
remote-shared-diagnostics/remote-fixture-lint: passed; 9258 ms; ["bash","bin/fm-lint.sh","bin/fm-remote-home-provision.sh","tests/fm-spawn-compact-adviser-disable-remote.test.sh"]
remote-repack-red/remote-source-repack: passed; 2549 ms; ["bash","-c",". tests/lib.sh; python3 \"$1\" \"$PWD/bin/fm-remote-home-provision.sh\"","_","$EVIDENCE/ci-clone-repack.py"]
remote-rotation-red/remote-source-object-rotation: passed; 3149 ms; ["bash","-c",". tests/lib.sh; python3 \"$1\" \"$PWD/bin/fm-remote-home-provision.sh\"","_","$EVIDENCE/ci-clone-repack.py"]

Current exact-head CI

Observed at 2026-09-23T00:47:43.6098730Z; head 4159f4965393456b988cd22f8c294c53edeacbab; GitHub mergeability: true.

Check Status Conclusion Producers
Lint 1 successful success 1
Lint 2 successful success 1
Test coverage guard successful success 1
Behavior portable parallel 1 successful success 1
Behavior portable parallel 2 successful success 1
Behavior portable serial 1 successful success 1
Behavior portable serial 2 successful success 1
Behavior portable serial 3 successful success 1
Behavior portable serial 4 successful success 1
Behavior portable serial 5 successful success 1
Behavior portable serial 6 successful success 1
Behavior portable serial 7 successful success 1
Behavior portable serial 8 successful success 1
Behavior portable serial 9 successful success 1
Behavior tests (Herdr) successful success 1
Behavior timing aggregate successful success 1
Stock macOS Bash snapshot compatibility successful success 1
Repo invariants successful success 1
Windows self-update entry point successful success 1
Windows reconciliation (core) successful success 1
Windows reconciliation (copilot-launch) successful success 1
Windows reconciliation (legacy-rollback) successful success 1
Windows reconciliation (pr-completion) successful success 1
Windows Copilot management successful success 1
Harness package compatibility successful success 1

Producer mismatches: 0; duplicate producers: 0; unexpected checks: 0.
Pending, absent, skipped, cancelled, or failed required checks are not passes. GitHub Actions owns the complete cross-platform matrix.

The original Windows-local Pi handoff failure and incomplete original local cases remain recorded below. Passing portable CI is not represented as a native Windows rerun. No live vendor session, fleet operation, PR merge, auto-merge, or local-main synchronization was performed.

Original reconciliation evidence and historical initial check snapshot

Status and merge method

Ordinary non-draft reconciliation PR; not yet merge-ready. Local readiness has unresolved outcomes listed below, and GitHub Actions owns the complete cross-platform matrix.
Use a merge commit, not squash or rebase merge, to preserve canonical upstream ancestry. Leave this PR unmerged until its exact-head checks and unresolved findings are settled; no auto-merge was enabled.

Frozen inputs and ancestry

  • Frozen fork / PR base: 395d209da567c8a706f0f17cc1f31a17d445846a.
  • Frozen canonical upstream: c5131a33a1e35a42e34733a5334fcc4e0225a656 (fetched once; no moving-upstream refresh).
  • Proven prior upstream: 888871de5cdf875ba4f4c0d231da6efdf7bad9a8.
  • Prior proof: newest valid trailer commit 460aca84d171cc64a1132fece6228910150d4456, reachable through fork merge 1c11c416fe4d70b475332795e7c698f3afb03dd6, corroborated by fork PR Reconcile upstream 888871de (September 17, 2026) #68. The prior SHA exists, is an ancestor of both frozen inputs, and equals their merge base.
  • Strategy: normal two-parent merge; 36 upstream commits, 169 canonical changed paths and 85 overlap paths. No reconstruction, synthetic anchor, or baseline worktree.
  • Original reconciliation head: f1ab3008c86e657fc2813bded88da55909d45a06; committed tree: 216df835715cf79c7c40aae05eec92034d56c458.
  • Verified parents, in order: 395d209da567c8a706f0f17cc1f31a17d445846a, c5131a33a1e35a42e34733a5334fcc4e0225a656. Git's trailer formatter returned the exact upstream SHA; the worktree is clean.

Firstmate-Upstream-SHA: c5131a3

Reconciliation decisions

  • Immutable, home-hashed and generation-bound launch files replace long inline delivery. Windows uses the existing literal PowerShell/Git-Bash transport, secures newly created native directories, validates reused directories freshly, and secures staged files before publication.
  • Kimi viewport-only trust readiness, timestamped launch failures, unconditional compact-adviser disabling, Lavish host propagation, and away-posture changes are retained. Signed Pi and other nonpilot paths stay with their existing owners.
  • Shared endpoint-absence proof feeds the fork's existing guarded recovery transaction. Direct spawn cannot bypass exact task-branch, task-held lease, competing-claim, metadata/session-lock and journal checks by creating another endpoint.
  • Pi generation handoff retains the fork's ownership token and shell-visible PID. Claude same-session sidecars compose with native ancestry and unverifiable-owner refusal.
  • Azure identity/observation, native private publication and rollback, generic forge URLs, bounded run selection, and in-process/batched status-reader cost guarantees remain intact.
  • Incoming test registrations and refreshed durations live in the existing catalogs. Slower fork measurements remain explicit overrides; the intentionally removed mandatory no-mistakes suite has no orphan duration. VISION.md now has explicit changed-source routing.
  • Two lint partitions and nine serial shards are retained without weakening analysis or granting concurrency through catalog metadata. The full slow lint inventory audit stays in its existing serial suite. Mandatory no-mistakes remains deleted and shared-template validation stays optional/context-selected.
  • Fixed merge-generated missing registry separators, the orphan duration, missing source routing, native symlink creation in the new Claude fixture, jq exposure in the draft fixture, and missing guarded-recovery fixture ports. The final two fixture repairs have the limitations below.

Local readiness and limits

The bounded controller consumed 2122.721 of 2400 seconds including preflight, inventory, failed attempts, retries and Git-for-Windows startup. Two 150-second timeouts triggered the shared circuit breaker; remaining local commands were deferred without resetting the allowance, increasing concurrency or widening production deadlines.
The four required gates passed during the bounded phase: Bash syntax, pinned context-selected lint (including actionlint), documentation audiences, and test coverage. Later fixture-only edits received source-aware lint; no production or catalog bytes changed after the final coverage pass. Gate evidence is not a claim that every final-tree behavior case passed.
Native directory ACL creation/revalidation, real PowerShell/Bash transport, Copilot and long launch delivery, Lavish forwarding, guarded dirty-copy recovery, direct-spawn refusal, native ancestry facts, Pi persistence-failure handoff, reader/process-cost contracts, catalog gate classes, source routing, copied private dependency layouts, draft refusal, and documentation behavior all have passing focused evidence.
Caching was disabled throughout. No interrupted, failed, skipped or merely preflighted result was reused as a pass. Earlier checks were retained only for unchanged dependency surfaces; affected fixture files were relinted.
Host probes observed Bash 5.3.15, Git 2.55.0.windows.5, Node 24.19.0, jq 1.8.2, ShellCheck 0.11.0, actionlint 1.7.12, Python 3.13.15 and Windows PowerShell 5.1.26100.9444. Ruby was absent. Relevant Git policies were inspected, including safe.bareRepository=explicit; signing, hooks and safety policies were not disabled.
No live Firstmate fleet, credentialed vendor session, no-mistakes pipeline, merge, auto-merge, or local-main synchronization was run. Historical live evidence in imported docs is not this run's live proof. All validation commands have exited.

Unresolved and deferred subjects

Subject Existing CI owner Result / remaining evidence
pi-predecessor-handoff Behavior portable serial 4 failed: Failed twice on Windows. At the fixture's fixed 1.2-second observation, only arm= and actionable-emitted were present; no predecessor-close/successor evidence had arrived. The serial retry reproduced it. No frozen-upstream differential was run, so this is unresolved, not a proven platform baseline.
trusted-claude-session Behavior portable serial 6 timeout-after-fixture-repair: The first run used Git-for-Windows ln's copied-file behavior instead of a real symlink. Repaired to fm_test_make_symlink; the complete repaired case reached the 150-second local bound. Its final outcome remains unverified.
guarded-recovery-endpoint Behavior portable serial 6 deferred-after-fixture-repair: The composed fixture lacked the guarded owner's session/socket inventory. Added that fact and a private, case-owned lock namespace through a copied adapter. Source-aware lint passed, but the two-timeout breaker deferred its behavioral rerun. The separate guarded-copy recovery and direct-spawn refusal cases passed.
lint-partition-contract Behavior portable parallel 1 timeout: The full named partition fixture exceeded 150 seconds. It was not retried with a larger limit. Actual pinned lint, actionlint, and selected source-aware lint passed.
parsed-ci-workflow Behavior portable serial 9 missing-prerequisite: Ruby is absent locally. Both preflight and execution explicitly deferred the parsed-YAML contract; actionlint is not claimed as a substitute for that behavior test.
Exact local commands, attempts, durations, and preflight accounting

The controller execution environment was FM_LIVE=0; deferred rows did not run. Commands are argv arrays; $EVIDENCE denotes the outside-repository evidence directory. The controller is skills/reconcile-firstmate-upstream/scripts/validate-local.mjs with the recorded absolute Git-for-Windows Bash, --max-timeouts 2 initially and 1 after the first timeout. Subsequent --budget-seconds values were only the unspent shared allowance.

Phase Allowance (s) Charged duration (s) Timeouts
preflight-gates 2400 5.600 0
gates 2394 204.440 0
retry-gates 2189 411.292 0
preflight-focused 1778 8.886 0
focused 1769 1204.806 1
retry-focused 564 287.697 1

Preflight-only phases executed only each plan's unique prerequisite argv/environment pair. Their aggregate durations include every probe; the controller does not separately publish successful probe durations. Failed probes and skips remain explicit. No test pass is inferred from preflight.

["bash","--version"]
["git","--version"]
["bash","-c","for tool in awk sed grep sort find perl mktemp date; do command -v \"$tool\" || exit 1; done"]
["shellcheck","--version"]
["actionlint","-version"]
["python3","--version"]
["node","--version"]
["jq","--version"]
["bash","-c","for tool in awk sed grep sort find perl mktemp date ps; do command -v \"$tool\" || exit 1; done"]
["powershell.exe","-NoProfile","-NonInteractive","-Command","$PSVersionTable.PSVersion.ToString()"]
["cygpath","--version"]
["ruby","--version"]
C1: ["bash","-c","git --version; git config --show-origin --get-regexp \"^(commit[.]gpgsign|core[.]hooksPath|safe[.]bareRepository)$\"; rc=$?; [ \"$rc\" -le 1 ]"]
  env: {"FM_LIVE":"0"}
C2: ["bash","bin/fm-test-run.sh","--list","--changed","--base","395d209da567c8a706f0f17cc1f31a17d445846a"]
  env: {"FM_LIVE":"0"}
C3: ["bash","-c","set -eu; files=$(bin/fm-lint.sh --list-files); while IFS= read -r script; do [ -z \"$script\" ] || /bin/bash -n \"$script\" || exit; done <<< \"$files\"; printf \"Shell syntax passed\\n\""]
  env: {"FM_LIVE":"0"}
C4: ["bash","bin/fm-lint.sh"]
  env: {"FM_LIVE":"0"}
C5: ["bash","bin/fm-lint.sh","bin/fm-spawn.sh","bin/fm-control-recovery-lib.sh","bin/fm-session-lock-lib.sh","bin/fm-classify-lib.sh","bin/fm-private-path-lib.sh"]
  env: {"FM_LIVE":"0"}
C6: ["bash","bin/fm-doc-audience-check.sh"]
  env: {"FM_LIVE":"0"}
C7: ["bash","bin/fm-test-run.sh","--check-coverage"]
  env: {"FM_LIVE":"0"}
C8: ["bash","-c","set -eu; for lane in \"$@\"; do printf \"FM_RECONCILE_LANE=%s\\n\" \"$lane\"; bin/fm-test-run.sh --list --lane \"$lane\"; done; printf \"FM_RECONCILE_LANE=herdr\\n\"; bin/fm-test-run.sh --list --family real-herdr-gated","_","portable-parallel-1","portable-parallel-2","portable-serial-1of9","portable-serial-2of9","portable-serial-3of9","portable-serial-4of9","portable-serial-5of9","portable-serial-6of9","portable-serial-7of9","portable-serial-8of9","portable-serial-9of9"]
  env: {"FM_LIVE":"0"}
C9: ["bash","bin/fm-test-run.sh","tests/fm-ci-workflow.test.sh"]
  env: {"FM_LIVE":"0"}
C10: ["bash","bin/fm-lint.sh","tests/fm-test-run.test.sh"]
  env: {"FM_LIVE":"0"}
C11: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_private_native_worker_directory_creation","tests/fm-private-path.test.sh"]
  env: {"FM_LIVE":"0"}
C12: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_native_transport_round_trip","tests/fm-platform-process.test.sh"]
  env: {"FM_LIVE":"0"}
C13: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_copilot_threads_model_effort_and_hooks","tests/fm-spawn-dispatch-profile.test.sh"]
  env: {"FM_LIVE":"0"}
C14: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_claude_long_launch_is_delivered_intact","tests/fm-spawn-dispatch-profile.test.sh"]
  env: {"FM_LIVE":"0"}
C15: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_lavish_server_address_is_exported_to_worker_launch","tests/fm-spawn-dispatch-profile.test.sh"]
  env: {"FM_LIVE":"0"}
C16: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_relaunch_missing_endpoint_preserves_copy","tests/fm-control-recovery.test.sh"]
  env: {"FM_LIVE":"0"}
C17: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_herdr_reclaim_keeps_the_task_whole","tests/fm-control-relaunch.test.sh"]
  env: {"FM_LIVE":"0"}
C18: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_spawn_relaunch_defers_missing_endpoint_to_control","tests/fm-control-relaunch.test.sh"]
  env: {"FM_LIVE":"0"}
C19: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_same_session_id_owns_a_recycled_background_chain","tests/fm-session-lock-ancestry.test.sh"]
  env: {"FM_LIVE":"0"}
C20: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_windows_native_ancestry_uses_verified_parent_rows","tests/fm-session-lock-ancestry.test.sh"]
  env: {"FM_LIVE":"0"}
C21: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_pi_actionable_output_waits_for_predecessor_close","tests/fm-pi-watch-extension.test.sh"]
  env: {"FM_LIVE":"0"}
C22: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_pi_replacement_persistence_failure_keeps_predecessor_until_successor","tests/fm-pi-watch-extension.test.sh"]
  env: {"FM_LIVE":"0"}
C23: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_snapshot_projection_bounds_json_tool_launches","tests/fm-startup-performance.test.sh"]
  env: {"FM_LIVE":"0"}
C24: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_status_text_destinations_preserve_legacy_capture_semantics","tests/fm-startup-performance.test.sh"]
  env: {"FM_LIVE":"0"}
C25: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_current_gen_path_stays_in_process","tests/fm-busy-state.test.sh"]
  env: {"FM_LIVE":"0"}
C26: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_launch_prompt_never_shortens_a_working_launch","tests/fm-busy-state.test.sh"]
  env: {"FM_LIVE":"0"}
C27: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_catalog_preserves_existing_gate_classes","tests/fm-test-catalog.test.sh"]
  env: {"FM_LIVE":"0"}
C28: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_reconciled_modules_select_their_consumers","tests/fm-test-run.test.sh"]
  env: {"FM_LIVE":"0"}
C29: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_harness_modules_select_all_consumers","tests/fm-test-run.test.sh"]
  env: {"FM_LIVE":"0"}
C30: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_process_modules_select_all_consumers","tests/fm-test-run.test.sh"]
  env: {"FM_LIVE":"0"}
C31: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_private_tracked_layouts","tests/fm-private-path.test.sh"]
  env: {"FM_LIVE":"0"}
C32: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_canonical_partitions_preserve_full_lint","tests/fm-lint.test.sh"]
  env: {"FM_LIVE":"0"}
C33: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_draft_pull_request_is_not_armed","tests/fm-pr-check-security.test.sh"]
  env: {"FM_LIVE":"0"}
C34: ["bash","bin/fm-test-run.sh","tests/fm-documentation-audiences.test.sh"]
  env: {"FM_LIVE":"0"}
C35: ["bash","bin/fm-lint.sh","tests/fm-session-lock-ancestry.test.sh","tests/fm-pi-watch-extension.test.sh","tests/fm-pr-check-security.test.sh","tests/fm-control-relaunch.test.sh"]
  env: {"FM_LIVE":"0"}
C36: ["bash","-c",". bin/fm-pr-lib.sh; value=$(fm_pr_json_draft_state '{\"isDraft\":true}'); printf \"draft value: %q\\n\" \"$value\"; [ \"$value\" = true ]"]
  env: {"FM_LIVE":"0"}
C37: ["bash","-c","FM_TEST_ONLY=\"$1\" exec bin/fm-test-run.sh \"$2\"","_","test_herdr_reclaim_keeps_the_task_whole","tests/fm-control-relaunch.test.sh"]
  env: {"FM_LIVE":"0","BASH_ENV":"$EVIDENCE/control-trace.sh","FM_RECONCILE_TRACE":"$EVIDENCE/recovery-retry.trace"}
Phase Subject Command Result Duration (s) / reason
gates git-fixture-policy C1 passed 0.865
gates changed-inventory C2 failed 1.826
gates shell-syntax C3 failed 8.457
gates lint C4 failed 132.408
gates source-aware-composed-lint C5 passed 49.532
gates documentation-audiences C6 passed 2.942
gates coverage C7 failed 1.254
gates lane-inventory C8 failed 1.823
retry-gates changed-inventory C2 failed 12.633
retry-gates shell-syntax C3 passed 13.586
retry-gates lint C4 passed 128.282
retry-gates coverage C7 passed 124.193
retry-gates lane-inventory C8 passed 128.000
preflight-focused parsed-ci-workflow C9 deferred prerequisite probe failed
focused changed-inventory C2 passed 25.747
focused routing-fix-lint C10 passed 7.175
focused native-launch-directory C11 passed 17.730
focused native-command-transport C12 passed 10.299
focused copilot-launch C13 passed 102.274
focused long-launch-delivery C14 passed 98.462
focused lavish-launch-environment C15 passed 50.294
focused guarded-recovery-copy C16 passed 42.800
focused guarded-recovery-endpoint C17 failed 42.647
focused direct-spawn-recovery-refusal C18 passed 29.026
focused trusted-claude-session C19 failed 110.506
focused native-session-ancestry C20 passed 37.385
focused pi-predecessor-handoff C21 failed 13.273
focused pi-predecessor-rollback C22 passed 16.188
focused status-snapshot-process-cost C23 passed 30.432
focused status-reader-compatibility C24 passed 8.879
focused busy-generation-process-cost C25 passed 9.149
focused launch-prompt-scope C26 passed 10.426
focused catalog-gate-classes C27 passed 8.234
focused reconciled-source-routing C28 passed 61.621
focused harness-source-routing C29 passed 54.476
focused process-source-routing C30 passed 68.324
focused private-dependency-layouts C31 passed 15.444
focused lint-partition-contract C32 timeout 150.954
focused draft-pr-refusal C33 failed 26.062
focused documentation-contract C34 passed 18.447
focused parsed-ci-workflow C9 deferred prerequisite probe failed
focused coverage-after-routing-fix C7 passed 126.821
retry-focused fixture-fix-lint C35 passed 33.901
retry-focused draft-parser-native-bytes C36 passed 3.436
retry-focused pi-predecessor-handoff C21 failed 16.161
retry-focused draft-pr-refusal C33 passed 68.456
retry-focused trusted-claude-session C19 timeout 150.917
retry-focused guarded-recovery-endpoint C37 deferred timeout circuit breaker

Initial syntax/lint failures were the two composed registry separators; the initial inventory/coverage failures were the orphan mandatory-pipeline duration, followed by missing VISION.md routing. Their repaired runs are separate rows, not overwritten history. The draft case's initial failure was its restricted PATH missing jq; after exposing its declared dependency, both native parser bytes and draft/ready/unreadable behavior passed.

CI inventory and selection

The resolved workflows expand to 25 automatic check names with one producer each (18 shared, seven fork). Both retain main push/PR triggers, read-only contents and workflow-qualified concurrency. The manual-only Windows Herdr experiment remains separate.
The shared real-Herdr job retains Herdr 0.7.4, protocol >=16, Treehouse and Pi 0.84.3. Portable Pi installation remains upstream's unpinned package with TypeScript 5.9.3; fork package compatibility pins Pi 0.84.3, OpenCode 1.18.23 and TypeScript 5.9.3. Native jobs use Node 24 and FM_LIVE=0. Serial lanes refuse parallel --jobs; catalog registration still grants no concurrency. Timing aggregation retains its dependencies on both parallel jobs, all nine serial jobs and Herdr, plus always-run artifact collection.

Expected check Sole producer
Lint 1 .github/workflows/ci.yml / lint
Lint 2 .github/workflows/ci.yml / lint
Test coverage guard .github/workflows/ci.yml / test-coverage
Behavior portable parallel 1 .github/workflows/ci.yml / tests-portable-parallel-1
Behavior portable parallel 2 .github/workflows/ci.yml / tests-portable-parallel-2
Behavior portable serial 1 .github/workflows/ci.yml / tests-portable-serial
Behavior portable serial 2 .github/workflows/ci.yml / tests-portable-serial
Behavior portable serial 3 .github/workflows/ci.yml / tests-portable-serial
Behavior portable serial 4 .github/workflows/ci.yml / tests-portable-serial
Behavior portable serial 5 .github/workflows/ci.yml / tests-portable-serial
Behavior portable serial 6 .github/workflows/ci.yml / tests-portable-serial
Behavior portable serial 7 .github/workflows/ci.yml / tests-portable-serial
Behavior portable serial 8 .github/workflows/ci.yml / tests-portable-serial
Behavior portable serial 9 .github/workflows/ci.yml / tests-portable-serial
Behavior tests (Herdr) .github/workflows/ci.yml / tests-herdr
Behavior timing aggregate .github/workflows/ci.yml / tests-timing-aggregate
Stock macOS Bash snapshot compatibility .github/workflows/ci.yml / macos-stock-bash
Repo invariants .github/workflows/ci.yml / invariants
Windows self-update entry point .github/workflows/fork-ci.yml / windows-update
Windows reconciliation (core) .github/workflows/fork-ci.yml / reconciliation-windows
Windows reconciliation (copilot-launch) .github/workflows/fork-ci.yml / reconciliation-windows
Windows reconciliation (legacy-rollback) .github/workflows/fork-ci.yml / reconciliation-windows
Windows reconciliation (pr-completion) .github/workflows/fork-ci.yml / reconciliation-windows
Windows Copilot management .github/workflows/fork-ci.yml / windows-management
Harness package compatibility .github/workflows/fork-ci.yml / harness-package-compatibility
All 246 selected scripts and their complete-suite CI dispositions

Selection command: bin/fm-test-run.sh --list --changed --base 395d209da567c8a706f0f17cc1f31a17d445846a.
Every script listed below is assigned to the existing producer heading for its complete suite, because the local budget covers selected composed-conflict cases rather than a second full matrix. The local command ledger shows supplemental named cases. The executable coverage guard proved 246 total = 24 parallel + 206 serial + 16 Herdr, complete/disjoint across nine serial shards.
Entries marked [live-capability], [herdr] or [optional-binary] retain that catalog gate. Assignment to a producer is not proof a live/platform gate was enabled. Native directory/process/path suites additionally run in Windows reconciliation (core), selected launch behavior in Windows reconciliation (copilot-launch), and the existing Windows management/package jobs retain their own exact scopes.

Behavior portable serial 4

tests/fm-afk-contract.test.sh
tests/fm-afk-pi-herdr-return-e2e.test.sh [live-capability]
tests/fm-afk-return.test.sh
tests/fm-bearings-board-lavish-live-e2e.test.sh [live-capability]
tests/fm-claude-session-lock-live-e2e.test.sh [live-capability]
tests/fm-control-recovery.test.sh
tests/fm-documentation-audiences.test.sh
tests/fm-fleet-sync.test.sh
tests/fm-herdr-pi-stale-registration-live-e2e.test.sh [live-capability]
tests/fm-home-summary-refresh.test.sh [optional-binary]
tests/fm-muse-signals-live-e2e.test.sh [live-capability]
tests/fm-nm-test-contract.test.sh
tests/fm-omp-harness.test.sh
tests/fm-peek-remote.test.sh
tests/fm-pi-watch-extension.test.sh
tests/fm-procevent.test.sh
tests/fm-quota-choose.test.sh
tests/fm-remote-reply.test.sh
tests/fm-spawn-compact-adviser-disable-remote.test.sh
tests/fm-spawn-worktree-settle.test.sh
tests/fm-startup-network.test.sh
tests/fm-task-delivery.test.sh
tests/fm-watch-checkpoint.test.sh

Behavior portable serial 7

tests/fm-afk-inject-e2e.test.sh
tests/fm-agy-harness.test.sh
tests/fm-backend-cmux-smoke.test.sh [optional-binary]
tests/fm-backend-cmux.test.sh [optional-binary]
tests/fm-backend-herdr-treehouse.test.sh
tests/fm-backend-herdr-windows-treehouse-live-e2e.test.sh
tests/fm-backlog-read-bound.test.sh
tests/fm-bearings-snapshot.test.sh [optional-binary]
tests/fm-busy-state.test.sh
tests/fm-extension-binding.test.sh
tests/fm-herdr-submit-confirm-live-e2e.test.sh [live-capability]
tests/fm-herdr-version-floor-live-e2e.test.sh [live-capability]
tests/fm-inbox.test.sh
tests/fm-pr-local-cost.test.sh
tests/fm-private-path.test.sh
tests/fm-rovo-harness.test.sh
tests/fm-rovo-signals-live-e2e.test.sh [live-capability]
tests/fm-secondmate-liveness.test.sh
tests/fm-spawn-dispatch-profile.test.sh
tests/fm-tasks-axi.test.sh
tests/fm-test-catalog.test.sh
tests/fm-wake-drain-open-decisions-cursor.test.sh
tests/fm-wake-drain-open-decisions.test.sh
tests/fm-wake-queue.test.sh
tests/fm-watch-recovery-loop.test.sh

Behavior tests (Herdr)

tests/fm-afk-inject-herdr-e2e.test.sh [herdr]
tests/fm-afk-launch.test.sh [herdr]
tests/fm-backend-autodetect-smoke.test.sh [herdr]
tests/fm-backend-herdr-agent-exit-shell-e2e.test.sh [herdr]
tests/fm-backend-herdr-eventwait-smoke.test.sh [herdr]
tests/fm-backend-herdr-focus-flash-e2e.test.sh [herdr]
tests/fm-backend-herdr-launcher-workspace-e2e.test.sh [herdr]
tests/fm-backend-herdr-presentation-e2e.test.sh [herdr]
tests/fm-backend-herdr-prune-safety-e2e.test.sh [herdr]
tests/fm-backend-herdr-respawn-idem-e2e.test.sh [herdr]
tests/fm-backend-herdr-smoke.test.sh [herdr]
tests/fm-backend-herdr-stale-active-tab-e2e.test.sh [herdr]
tests/fm-backend-herdr-workspace-per-home-e2e.test.sh [herdr]
tests/fm-control-herdr-smoke.test.sh [herdr]
tests/fm-herdr-attached-viewer-live-e2e.test.sh [herdr]
tests/fm-herdr-session-cleanup-e2e.test.sh [herdr]

Behavior portable serial 8

tests/fm-agy-signals-live-e2e.test.sh [live-capability]
tests/fm-backlog-handoff.test.sh
tests/fm-bearings-board-render.test.sh [optional-binary]
tests/fm-bootstrap-network-parallel.test.sh
tests/fm-calm-claude-mod.test.sh
tests/fm-claude-stop-autoarm-live-e2e.test.sh [live-capability]
tests/fm-claude-stop-autoarm.test.sh
tests/fm-copilot-primary-live-e2e.test.sh [live-capability]
tests/fm-cursor-primary-live-e2e.test.sh [live-capability]
tests/fm-pi-primary-live-e2e.test.sh [live-capability]
tests/fm-project-origin.test.sh
tests/fm-public-followup.test.sh
tests/fm-remote-job-orphan-reap.test.sh
tests/fm-remote-secondmate-parent-binding.test.sh
tests/fm-secondmate-harness.test.sh
tests/fm-send-agy-confirm.test.sh
tests/fm-send-remote-delivery.test.sh
tests/fm-sessionstart-nudge.test.sh
tests/fm-spawn-queue.test.sh
tests/fm-subagent-pretool-check.test.sh
tests/fm-supervision-events.test.sh
tests/fm-tangle-guard.test.sh
tests/fm-update.test.sh
tests/fm-vendor-auth-probe.test.sh

Behavior portable parallel 2

tests/fm-arm-pretool-check.test.sh
tests/fm-backend-herdr.test.sh
tests/fm-captain-hold-lifecycle.test.sh
tests/fm-crew-state.test.sh
tests/fm-ensure-agents-md.test.sh
tests/fm-herdr-lab.test.sh
tests/fm-send-popup-settle.test.sh
tests/fm-send-settle.test.sh
tests/fm-send-strict.test.sh
tests/fm-spawn-batch.test.sh
tests/fm-supervision-instructions.test.sh
tests/fm-transition-lib.test.sh
tests/fm-x-mode.test.sh

Behavior portable serial 9

tests/fm-ask-user-authority.test.sh
tests/fm-backend-zellij-smoke.test.sh [optional-binary]
tests/fm-busy-adapter-wiring.test.sh
tests/fm-ci-workflow.test.sh
tests/fm-classify-corr-token.test.sh
tests/fm-claude-trust.test.sh
tests/fm-composer-codex-idle-live-e2e.test.sh [live-capability]
tests/fm-copilot-management-live-e2e.test.sh [live-capability]
tests/fm-cursor-harness.test.sh
tests/fm-home-summary-request.test.sh
tests/fm-lint-inventory.test.sh
tests/fm-mail.test.sh
tests/fm-omp-primary-live-e2e.test.sh [live-capability]
tests/fm-pi-windows-shell-invocation.test.sh
tests/fm-procevent-stop-proof.test.sh
tests/fm-remote-backlog-handoff.test.sh
tests/fm-remote-secondmate-trace-context.test.sh
tests/fm-secondmate-restart.test.sh
tests/fm-send-secondmate-marker-herdr-e2e.test.sh [live-capability]
tests/fm-spawn-compact-adviser-disable.test.sh
tests/fm-stat-shadowing.test.sh
tests/fm-teardown.test.sh
tests/fm-test-fixture-cleanup.test.sh
tests/fm-test-fixtures.test.sh
tests/fm-wake-drain-unread-status.test.sh

Behavior portable serial 3

tests/fm-backend-orca.test.sh [optional-binary]
tests/fm-calm-claude-mod-live-e2e.test.sh [live-capability]
tests/fm-contributions.test.sh [optional-binary]
tests/fm-daemon.test.sh
tests/fm-gotmp.test.sh
tests/fm-grok-continuity-live-e2e.test.sh [live-capability]
tests/fm-harness-adapter-instructions-live-e2e.test.sh [live-capability]
tests/fm-harness-adapter-references.test.sh
tests/fm-harness-contract.test.sh
tests/fm-harness-precedence.test.sh
tests/fm-kimi-harness.test.sh
tests/fm-launch-status.test.sh
tests/fm-lint-workflows.test.sh
tests/fm-on.test.sh
tests/fm-pi-branch-extension.test.sh
tests/fm-pr-check-security.test.sh
tests/fm-secondmate-lifecycle-e2e.test.sh
tests/fm-secondmate-sync.test.sh
tests/fm-send-inbox.test.sh
tests/fm-test-isolation-proof.test.sh
tests/fm-trace-context-spawn.test.sh
tests/fm-wake-daemon-lifecycle-e2e.test.sh
tests/fm-watcher-lock.test.sh

Behavior portable serial 1

tests/fm-backend-tmux-smoke.test.sh
tests/fm-gemini-harness.test.sh
tests/fm-gitignore-config.test.sh
tests/fm-grok-stop-live-e2e.test.sh [live-capability]
tests/fm-live-gate.test.sh
tests/fm-lock-fast.test.sh
tests/fm-opencode-primary-live-e2e.test.sh [live-capability]
tests/fm-operational-input.test.sh
tests/fm-pr-state-live-e2e.test.sh [live-capability]
tests/fm-send-secondmate-marker.test.sh
tests/fm-stow-cascade.test.sh
tests/fm-turnend-foreign-owner-arm-fix.test.sh
tests/fm-watch-triage.test.sh

Behavior portable serial 2

tests/fm-backend-zellij.test.sh [optional-binary]
tests/fm-backend.test.sh
tests/fm-bootstrap.test.sh
tests/fm-check-unregister.test.sh
tests/fm-codex-continuity-live-e2e.test.sh [live-capability]
tests/fm-codex-hook-layer-live-e2e.test.sh [live-capability]
tests/fm-control.test.sh
tests/fm-copilot-harness.test.sh
tests/fm-dispatch-resolve.test.sh
tests/fm-herdr-session-cleanup.test.sh
tests/fm-inactive-reconcile.test.sh
tests/fm-path.test.sh
tests/fm-remote-doctor.test.sh
tests/fm-remote-entrypoint.test.sh
tests/fm-remote-secondmate-lifecycle-e2e.test.sh
tests/fm-secondmate-reconcile.test.sh
tests/fm-send-inbox-doorbell-live-e2e.test.sh [live-capability]
tests/fm-spawn-pool-base-freshen.test.sh
tests/fm-startup-performance.test.sh
tests/fm-task-inbox.test.sh
tests/fm-teardown-endpoint-safety.test.sh
tests/fm-tmux-agent-liveness.test.sh
tests/fm-tool-update-check.test.sh
tests/fm-trace-context-lib.test.sh

Behavior portable serial 5

tests/fm-backlog-atomicity.test.sh
tests/fm-calm-pi-extension.test.sh
tests/fm-classify-decision-key.test.sh
tests/fm-cmux-claude-composer-live-e2e.test.sh [live-capability]
tests/fm-copilot-hooks-live-e2e.test.sh [live-capability]
tests/fm-harness-liveness-drift-live-e2e.test.sh [live-capability]
tests/fm-herdr-unregistered-agent.test.sh
tests/fm-herdr-windows-liveness-live-e2e.test.sh [live-capability]
tests/fm-mail-check.test.sh
tests/fm-pending-reply.test.sh
tests/fm-pi-branch-live-e2e.test.sh [live-capability]
tests/fm-pi-branch-responsiveness-live-e2e.test.sh [live-capability]
tests/fm-pi-codex-native.test.sh [live-capability]
tests/fm-pr-reviewers.test.sh
tests/fm-procevent-quota.test.sh
tests/fm-reconcile-validation.test.sh
tests/fm-remote-herdr-guard.test.sh
tests/fm-remote-job.test.sh
tests/fm-remote-transport-lanes.test.sh
tests/fm-secondmate-safety.test.sh
tests/fm-sessionstart-hook-live-e2e.test.sh [live-capability]
tests/fm-startup-memory-budget.test.sh
tests/fm-update-windows.test.sh
tests/fm-voice-relay.test.sh
tests/fm-wake-drain-outcome-backstop.test.sh

Behavior portable serial 6

tests/fm-bearings-board.test.sh
tests/fm-branch-supervision.test.sh
tests/fm-calm-claude-mod-plugin.test.sh [live-capability]
tests/fm-composer-matrix-live-e2e.test.sh [live-capability]
tests/fm-control-relaunch.test.sh
tests/fm-cursor-primary.test.sh
tests/fm-fleet-snapshot-view.test.sh [optional-binary]
tests/fm-gate-refuse.test.sh
tests/fm-guard-stale-banner.test.sh
tests/fm-launch-prompt-signals-live-e2e.test.sh [live-capability]
tests/fm-muse-harness.test.sh
tests/fm-platform-process.test.sh
tests/fm-pr-state.test.sh
tests/fm-procevent-when.test.sh
tests/fm-quota-array-dispatch-live-e2e.test.sh [live-capability]
tests/fm-remote-job-wait.test.sh
tests/fm-send-resolve-key.test.sh
tests/fm-session-lock-ancestry.test.sh
tests/fm-session-start.test.sh
tests/fm-sessionstart-instruction-refresh-live-e2e.test.sh [live-capability]
tests/fm-shared-captain-inheritance.test.sh
tests/fm-turnend-guard.test.sh
tests/fm-watch-arm.test.sh
tests/herdr-workspace-move.test.sh

Behavior portable parallel 1

tests/fm-brief.test.sh
tests/fm-cd-pretool-check.test.sh
tests/fm-composer-ghost.test.sh
tests/fm-composer-lib.test.sh
tests/fm-grok-harness.test.sh
tests/fm-lint.test.sh
tests/fm-pi-primary-types.test.sh
tests/fm-pr-merge.test.sh
tests/fm-review-diff.test.sh
tests/fm-test-run.test.sh
tests/fm-tmux-submit-busy.test.sh

No-renames divergence and locality audit

Comparison against frozen upstream Modified upstream paths Additive fork paths Deletions Total
Frozen fork before reconciliation 270 73 6 349
Exact final head 189 73 2 264

The branch changes 176 manifest paths (172 modified, 4 added). 85 existing comparison paths become exactly upstream-equivalent; this mainly reflects adopting the upstream interval, not a new extraction. No moved implementation or rename is counted as a divergence reduction. The .github/workflows/no-mistakes-required.yml and tests/fm-no-mistakes-required.test.sh deletions remain intentional.
13 extracted/exception owners are byte-identical to the frozen fork, including both pilot adapters, process implementation/declarations/native helper, path helper, catalog loader, independent isolation proof, Azure poll and the remote doctor/entrypoint. The doctor's hash protocol is unchanged. All 47 non-conflicted overlap paths are retained in the source comparison evidence; conflict-free status was not treated as proof of correct composition.

git diff --no-renames --name-status c5131a33a1e35a42e34733a5334fcc4e0225a656 395d209da567c8a706f0f17cc1f31a17d445846a
git diff --no-renames --name-status c5131a33a1e35a42e34733a5334fcc4e0225a656 f1ab3008c86e657fc2813bded88da55909d45a06
git diff --no-renames --numstat 395d209da567c8a706f0f17cc1f31a17d445846a f1ab3008c86e657fc2813bded88da55909d45a06
git diff --no-renames --unified=0 c5131a33a1e35a42e34733a5334fcc4e0225a656 f1ab3008c86e657fc2813bded88da55909d45a06 -- bin/fm-spawn.sh bin/fm-control.sh bin/fm-control-recovery-lib.sh .pi/extensions/fm-primary-pi-watch.ts bin/fm-classify-lib.sh bin/fm-test-run.sh
Upstream-equivalent path inventory
.agents/skills/harness-adapters/references/harness/claude.md
.agents/skills/harness-adapters/references/harness/kimi.md
.agents/skills/process-event-sources/SKILL.md
.agents/skills/quota-array-dispatch/SKILL.md
.omp/extensions/fm-primary-omp-watch.ts
.pi/extensions/lib/fm-branch-dispatch.ts
VISION.md
bin/backends/cmux.sh
bin/backends/zellij.sh
bin/fm-afk-contract.sh
bin/fm-afk-launch.sh
bin/fm-afk-return.sh
bin/fm-branch-prompt.sh
bin/fm-composer-lib.sh
bin/fm-config-inherit-lib.sh
bin/fm-contributions.jq
bin/fm-dod-lib.sh
bin/fm-inbox.sh
bin/fm-lease-lib.sh
bin/fm-mail-check.sh
bin/fm-merge-authority-lib.sh
bin/fm-merge-local.sh
bin/fm-merge-outcome-lib.sh
bin/fm-parent-channel-lib.sh
bin/fm-pending-reply-lib.sh
bin/fm-procevent-lavish.sh
bin/fm-procevent-quota.sh
bin/fm-procevent-remote-reply.sh
bin/fm-quota-axi-lib.sh
bin/fm-remote-home-seed.sh
bin/fm-secondmate-report.sh
bin/fm-task-inbox-lib.sh
bin/fm-tmux-lib.sh
docs/captain-hold-lifecycle.md
docs/pi-supervision-branch.md
docs/secondmate-parent-channel.md
docs/supervision-protocols/pi.md
docs/verification/dispatch-auth.md
docs/verification/dispatch-resolve.md
docs/verification/process-event-sources.md
docs/verification/secondmate-parent-channel.md
docs/voice-relay.md
tests/fixtures.sh
tests/fm-afk-contract.test.sh
tests/fm-afk-launch.test.sh
tests/fm-afk-pi-herdr-return-e2e.test.sh
tests/fm-afk-return.test.sh
tests/fm-agy-harness.test.sh
tests/fm-backend-autodetect-smoke.test.sh
tests/fm-backend-orca.test.sh
tests/fm-backend.test.sh
tests/fm-bearings-snapshot.test.sh
tests/fm-branch-supervision.test.sh
tests/fm-ci-workflow.test.sh
tests/fm-classify-corr-token.test.sh
tests/fm-claude-stop-autoarm-live-e2e.test.sh
tests/fm-cmux-claude-composer-live-e2e.test.sh
tests/fm-composer-lib.test.sh
tests/fm-composer-matrix-live-e2e.test.sh
tests/fm-contributions.test.sh
tests/fm-dispatch-resolve.test.sh
tests/fm-guard-stale-banner.test.sh
tests/fm-inbox.test.sh
tests/fm-launch-prompt-signals-live-e2e.test.sh
tests/fm-muse-harness.test.sh
tests/fm-pending-reply.test.sh
tests/fm-pr-merge.test.sh
tests/fm-procevent-quota.test.sh
tests/fm-quota-choose.test.sh
tests/fm-remote-backlog-handoff.test.sh
tests/fm-remote-reply.test.sh
tests/fm-remote-secondmate-lifecycle-e2e.test.sh
tests/fm-remote-secondmate-parent-binding.test.sh
tests/fm-remote-secondmate-trace-context.test.sh
tests/fm-rovo-harness.test.sh
tests/fm-send-remote-delivery.test.sh
tests/fm-send-resolve-key.test.sh
tests/fm-spawn-compact-adviser-disable-remote.test.sh
tests/fm-spawn-compact-adviser-disable.test.sh
tests/fm-tangle-guard.test.sh
tests/fm-task-delivery.test.sh
tests/fm-task-inbox.test.sh
tests/fm-turnend-foreign-owner-repro.py
tests/fm-wake-queue.test.sh
tests/secondmate-helpers.sh
Retained owner Boundary / remaining reason Removal condition
bin/fm-harness-lib.sh and bin/harnesses/{copilot,pi}.sh Lifecycle callers retain profile resolution, generation, publication, delivery and rollback. Kimi and signed-Pi legacy branches remain separate. Equivalent closed upstream pilot capabilities and failure semantics across every caller.
bin/platform/process.mjs, declarations, native process helper and compatibility imports Pi owner records retain the fork ownership token and shell-visible PID while adopting upstream generation handoff. Equivalent process identity, native transport, and supported package import contracts.
bin/fm-private-path-lib.sh and bin/platform/windows-private-path.ps1 New native launch directories use the worker secure operation; reused directories are freshly validated. Callers still own staging and publication. Equivalent native ACL/ownership policy without added subprocesses or cached authorization.
bin/fm-control-recovery-lib.sh Shared absence proof feeds the existing lease/claim/task-branch transaction. Direct spawn cannot create a second unguarded recovery path. Upstream recovery preserves exact task-held leases, unique claims, both copies, approval binding and journals.
bin/fm-classify-lib.sh, bin/fm-fleet-snapshot.sh and existing hot readers Timestamp support composes with in-process readers, fresh file facts and batched JSON construction. Equivalent output, freshness, refusals and equal or better cost on matched native Windows inputs.
bin/fm-pr-poll.sh through bin/fm-pr-lib.sh Positive GitHub draft refusal does not replace Azure canonical identity, private publication, observation or replay. Equivalent supported Azure/native behavior without manufactured GitHub identity or broader merge authority.
tests/catalog/{core,fork}.tsv and bin/fm-test-catalog-lib.sh Incoming registrations, durations and nine shards use catalog metadata. The independent proof owner alone grants concurrency. VISION.md is explicitly routed. Equivalent upstream metadata loading, changed-reference selection and proof-owned admission.
.github/workflows/{ci,fork-ci}.yml Two shared lint partitions and nine serial shards coexist with seven fork checks; 25 unique automatic producers. Mandatory no-mistakes remains deleted. Equivalent required checks, gates, package pins, workflow-qualified concurrency and artifact dependencies.
bin/fm-remote-doctor.sh and bin/fm-remote-entrypoint.sh Both byte-identical to the frozen fork; no hash-pin update or legacy harness migration is needed. A separately approved migration preserving the integrity protocol and legacy contracts.

Locality evidence is behavioral: actual native ACL/transport calls, named harness/process changed-source routing, copied native dependency layouts, guarded recovery/refusal, and bounded hot-reader/JSON-launch assertions. The unresolved recovery and Pi outcomes remain limitations rather than directory-name or output-equality claims.
Raw NUL-delimited source/staging manifests, merge stages, complete source deltas, per-command logs, registrations/weights, expanded producer definitions and before/after inventories are preserved outside the checkout in the reconciliation session. No diagnostic or planning artifact is added to the repository.

Historical initial check snapshot

Observed at 2026-09-22T20:50:02.8814296Z; PR mergeability: true.

Check Status Conclusion Producers
Lint 1 pending pending 1
Lint 2 pending pending 1
Test coverage guard successful success 1
Behavior portable parallel 1 pending pending 1
Behavior portable parallel 2 pending pending 1
Behavior portable serial 1 pending pending 1
Behavior portable serial 2 pending pending 1
Behavior portable serial 3 pending pending 1
Behavior portable serial 4 pending pending 1
Behavior portable serial 5 pending pending 1
Behavior portable serial 6 pending pending 1
Behavior portable serial 7 pending pending 1
Behavior portable serial 8 pending pending 1
Behavior portable serial 9 pending pending 1
Behavior tests (Herdr) pending pending 1
Behavior timing aggregate absent - 0
Stock macOS Bash snapshot compatibility pending pending 1
Repo invariants successful success 1
Windows self-update entry point successful success 1
Windows reconciliation (core) pending pending 1
Windows reconciliation (copilot-launch) successful success 1
Windows reconciliation (legacy-rollback) successful success 1
Windows reconciliation (pr-completion) pending pending 1
Windows Copilot management pending pending 1
Harness package compatibility successful success 1

A pending, absent, skipped, cancelled or failed required check is not a pass. The PR remains unmerged; the unresolved local findings above still prevent a fully validated claim.

kunchenguid and others added 30 commits September 17, 2026 19:24
)

* Improve CI reliability and rebalance full-coverage validation

* no-mistakes(document): Clarify lint partition documentation
…guid#4799)

* Handle Kimi workspace trust dialog

* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers

* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures

* no-mistakes(review): Read visible pane for Kimi trust and ready gates

* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate

* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection

* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca
…er (kunchenguid#4775)

* fix(bin): report a record whose agent is gone once instead of escalating forever

The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.

Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.

fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that was about to escalate, read it once and treat
only its two recovery-grade verdicts - dead (endpoint present, no agent in it)
and missing (endpoint authoritatively absent) - as proof, reporting that record
once and not re-escalating it while it stays that way. Every other verdict,
including alive, ambiguous, unreadable, unverified, and a read that failed
outright, keeps the identical schedule, reason, and escalation count, so a
genuinely wedged live agent is unaffected. The probe costs at most one backend
read per window per threshold, the same budget the declared-wait consult and the
worktree write probe already take.

The report decides nothing about the record's fate: both lanes still held
unlanded work and teardown refusing them was correct, so retiring, relaunching,
or cleaning up stays with the supervisor. The once-only marker is owned entirely
by that function and is dropped by the same read the moment the endpoint stops
reading gone, so a replacement launched into the same window escalates normally
and its own later death is reported again.

Related, and not closed by this: kunchenguid#4412, kunchenguid#4482, kunchenguid#4316.

Tests drive the real watcher against a record whose endpoint does not exist and
pin both directions: dead and missing report once and never advance the count
across later thresholds, while alive, ambiguous, and unreadable endpoints keep
escalating with the identical reason and a climbing count.

* fix(bin): bind the once-only dead report to the pane it reported

Review of the parent commit found a reachable sequence where a later death in
the same window lost its promised report. The marker was keyed on the verdict
string alone and dropped only when a threshold probe read a non-gone verdict,
but probes run only at thresholds: a replacement launched into the same window
that dies without ever being probed alive - it crashes at startup, or works and
then crashes - was absorbed by the previous death's marker. The pane's first
sight yielded only the generic stale wake and every later threshold matched the
stale marker, so the second death never got the detailed once-report that both
the function's own comment and docs/architecture.md promise.

Record the verdict together with the pane hash it was reported for, and absorb a
repeat only while both still match. A replacement churns the pane, which resets
the stale suppressor, wedge timer, and escalation count while no reset site
touches this marker, so the pane half is what tells the second death apart from
the first. The live-probe drop stays as it was.

Clearing the marker at those reset sites instead would re-open unbounded
re-alarming for a dead pane whose display ever ticks, which is the exact defect
the parent commit exists to close.

The noise bound is unchanged: an unchanged dead pane still absorbs on every
later threshold and never advances the escalation count, and every verdict short
of proof still escalates exactly as before.

* no-mistakes(review): Key the dead-record once-marker on the busy incarnation token

* no-mistakes(document): Document dead-record escalation cap in stale-pane config entry

* no-mistakes(document): Add busy-state inventory line to AGENTS.md

* no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path
…id#4854)

Captain holds have no due semantics and are a hold kind, not a Beads issue
type. The create path now waives due.required and maps to native type task.

Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(bin): launch every spawned agent with the compact adviser disabled

Every crewmate, scout, and secondmate Firstmate launches now starts with
COMPACT_ADVISER_DISABLE=1, on a fresh spawn and on a relaunch alike, so an
unattended session never activates the compact adviser.
The value is unconditional: no configuration file gates it and there is no
override, unlike the trace carrier beside it.

Three carriers deliver it, because no single one covers every launch shape.
The pane shell receives an export beside GOTMPDIR, so the agent's own children
inherit it too.
The launch command carries an explicit assignment, prepended outermost so it
wins over any ambient value the pane already held.
The cleared launch environment sets it again at the `env -i` boundary and keeps
COMPACT_ADVISER_DISABLE in the fixed operational floor, which is what preserves
the switch when config/launch-env-allowlist empties the environment, and what
delivers it on a remote host that never had the value.

bin/fm-control.sh relaunch, the bootstrap secondmate relaunch, and the remote
secondmate transport all rebuild their launch through bin/fm-spawn.sh, so they
inherit the same floor.
The captain's own primary session is untouched.

The two new suites drive the real spawn and then execute the launch command the
pane actually received, with the harness replaced by a probe that prints its own
environment, rather than matching script text.
They cover ship and secondmate launches with the allowlist absent and enabled,
the pane export and its ordering, fm-control.sh relaunch, and the full parent to
remote-host chain.

* no-mistakes(review): Export compact-adviser disable across compound launches

* no-mistakes(document): Document spawned-agent compact-adviser environment guarantee
…henguid#4894)

* fix(bin): let a background Claude session keep owning its session lock

Session-lock ownership was decided by process ancestry alone. Under an
unattended Claude session the model loop runs in a transient bg-spare
bridged to the front-end by a shared daemon; when that bridge is
recycled the contiguous claude-named ancestry from a hook to the
recorded owner breaks while the owner pid stays alive, so the Stop
auto-arm stood down as a foreign live owner, the turn-end guard ended
every turn with its read-only diagnostic, and fm-lock.sh refused - a
self-sustaining outage until restart.

Ownership is now ancestry membership OR a trusted same-session id,
never id-first:

- fm-session-lock-lib.sh accepts CLAUDE_CODE_SESSION_ID only when
  CLAUDE_PID is a Claude-shaped member of the current contiguous run,
  compares it against the id recorded in state/.lock-session, and
  requires the recorded pid to still be a live harness. No id, no
  sidecar, an untrusted id, a different id, or a dead recorded pid
  leaves the ancestry verdict unchanged. Ids are never read from ps
  argv.
- fm-lock.sh accepts a same-session holder at both refusal sites,
  writes, refreshes, and clears the sidecar only under its claim lock
  (including the early already-mine exit, skipped only while the
  deferred startup sweep leases that lock), keeps it byte-identical
  across a same-session confirmation, records CLAUDE_PID on lock line 1
  for a session with a trusted id so a shared daemon or front-end that
  outlives the session never keeps a dead session's lock alive, never
  rewrites a live line 1 on a same-session confirmation, and names the
  recorded id in the live-owner refusal.
- The .lock line-1 format is unchanged, so every reader that takes the
  whole first line as the pid keeps working; the guard's foreign-owner
  exit is unchanged and inherits the fix through the shared predicate.

Tests: the ancestry suite drives the ancestry and id signals apart in a
deterministic process table (asserting the divergence) and runs a real
orphaned front-end/daemon/pty-host/spare tree through six phases with
the real lock, auto-arm, and guard scripts; the foreign-owner repro
keeps its negative control and adds a same-id positive control.

Disclosure: no live unattended Claude background session ran on the
verifying machine. The topology is documented by the real process
listings in kunchenguid#3902, kunchenguid#2314, kunchenguid#3398, and kunchenguid#4066; coverage is the structural
predicate plus the executable fixtures, not a live pass.

Residual: bin/fm-sessionstart-nudge.sh keeps its own private ancestry
walk (it only decides whether to print a nudge) and may nudge on a
resume in the recycled case.

Out of scope, deliberately: no structured lock format, no guard budget
changes, no daemon-identity rejection, no fork lineage.

* no-mistakes(review): Wait for claim lock; revert failed sidecars

* no-mistakes(review): Revalidate ownership after wait; restore sidecars

* no-mistakes(review): Roll back sidecar by publication phase

* no-mistakes(review): Restore sidecar only if lock line is unchanged

* no-mistakes(review): Trust session ids without a spelling allowlist

* no-mistakes(review): Disarm sidecar rollback before backup cleanup

* no-mistakes(document): Updated session-lock ownership documentation
* feat: park main under the away posture on Pi

While the away-posture record exists on a Pi primary, the supervision branch
takes every actionable wake, no processing turn opens on main, captain rows
accumulate for the return brief, and main's standing authority relocates to
the branch through the existing guarded scripts.

- lib/fm-branch-dispatch.ts: read the record at every routing decision; while
  it exists claim check, decision-owned, and heartbeat rows too, keeping the
  two broken-queue vetoes; expose checkSeqs so a claimed check row lifts task
  scoping.
- fm-primary-pi-watch.ts: offer every actionable row under the record; a
  declined wake and every watcher-failure alarm still reach main.
- fm-branch-supervision.ts: drop the legacy .afk decline; append a fixed
  POSTURE: AWAY tail carrying the record's read-back verbatim per wake; open no
  processing request while the record exists, re-checked immediately before a
  request would open and at every run boundary; present the accumulated rows
  at the first run boundary after archive.
- fm-lease-lib.sh: fm_lease_forbid_branch passes the branch for opted-in
  actions only while fm-afk-contract.sh validate succeeds on a confirmed live
  record; PR merge, fresh spawn, and decision answer opt in, local landing
  never does.
- fm-send.sh: a --resolve-key naming an open needs-decision or captain-held
  task is a decision answer and meets the partition; blocked: keys stay
  steering.
- fm-spawn.sh: enforce the record's spend cap for a fresh ordinary spawn by
  either actor; relaunches and secondmates exempt.
- fm-branch-prompt.sh: fixed Postures section and the verbatim
  ask-user-authority policy; the prefix stays byte-stable.
- fm-afk-return.sh: count what the away session handled from the store.
- docs, afk skill, AGENTS.md stub: main parked on Pi, green merge gate
  absolute while away.
- tests: watcher and branch extension suites, fleet-record, merge, and
  decision-answer suites cover the relocation, the vetoes, the tail, the
  parked processing turn, the cancellation, the re-presentation, and the
  spend cap; dated live-guard evidence recorded.

* no-mistakes(review): Refuse branch merge after preflight archive race

* no-mistakes(review): Fix away wake, spawn, and processing races

* no-mistakes(review): Suppress parked processing; narrow away-only rejection

* no-mistakes(review): Abort dedicated processing; gate branch spawn once

* no-mistakes(review): Stamp away-only on the dispatch offer

* no-mistakes(review): Treat invalid away records as spend-cap absence

* no-mistakes(review): Drop spawn test hook; abort processing-opened runs

* no-mistakes(review): Bind abort to opening prompt; cap-read absence

* no-mistakes(review): Limit away branch spawn to queued work only

* no-mistakes(document): Correct AFK posture documentation
* ci: simplify CI job timeouts to a three-tier policy

Replace the scattered per-job timeout values (10m parallel, 25m lint, 30m
serial, 10m macOS) with three readable tiers, each a hang tripwire with
headroom rather than a packing estimate:

- fast (5m): coverage guard, repo invariants, timing aggregate
- normal (30m, one shared budget): lint partitions, portable parallel
  shards, portable serial shards, macOS stock Bash
- heavy (Herdr only): 20m step tripwire on the family run so always()
  cleanup still runs, under a 75m job-level last-resort backstop

The workflow's header comment states the policy and points at
docs/fm-test-portable-shards.md "Timeouts", which now owns it, and each
job names its tier beside timeout-minutes. tests/fm-ci-workflow.test.sh
asserts the policy against the parsed workflow instead of the old
per-job minute values: every job joins exactly one tier, exactly three
distinct job-level values exist, the fast tier stays within 5-10
minutes, the normal budget stays at least double the modeled parallel
lane sum reported by fm-test-run.sh --check-coverage, and the Herdr step
tripwire stays below its job backstop with an always() cleanup after it.

Concurrency supersession, shard counts, lane membership, and fail-fast
settings are unchanged.

* no-mistakes(review): Decouple the normal timeout from packing estimates

* no-mistakes(review): Assert Herdr teardown follows the family run

* no-mistakes(review): Pin Herdr family-run timeout to 20 minutes

* no-mistakes(review): Ignore comments when identifying Herdr steps

* no-mistakes(review): Identify Herdr steps by declarative ids

* no-mistakes(document): Clarify authoritative three-tier timeout policy
…nchenguid#4895)

* fix(bin): keep supervisor status closes from waking the same home

A drain that already folded OPEN DECISIONS has presented those bytes even
when the watcher has no matching seen marker. Treat that fold, and the
presentation cursor, as known so the bookkeeping close stays quiet while
later worker lines still signal.

* no-mistakes(review): Keep folded worker failures waking past supervisor closes

* no-mistakes(review): Wake on unlisted folded worker lines; batch multi-key closes

* no-mistakes(review): Stop folded worker resolved lines from counting as already read

* no-mistakes(document): Correct self-announced close marker contract in docs
* Stop steering operators away from Herdr

* no-mistakes(review): Neutralize remaining Herdr opt-out documentation wording
…enguid#4973)

* fix(bin): treat a live no-mistakes run as current after rebase

A running run on the task's branch is authoritative regardless of head.
Matching only the local head made a rebased in-flight run look failed.

* no-mistakes(review): restrict coarse live-any-head to foreign-branch answers

* no-mistakes(review): reject gate-parked runs from the executing predicate

* no-mistakes(review): hoist gate-marker patterns into single run-lib owner

* no-mistakes(review): require live daemon for head-free run binding

* no-mistakes(review): require answered daemon-down before unbinding live runs

* no-mistakes(review): extend daemon guard to anchored continuation routes

* no-mistakes(review): delete live-any-head; restore dead-daemon verdict

* no-mistakes(review): keep parked gates parked; name dead daemon everywhere

* no-mistakes(review): set dead-daemon verdict instead of emitting early

* no-mistakes(review): align selected route with legacy dead-daemon handling

* no-mistakes(review): drop unproven-record binds; narrow coarse gate reading

* no-mistakes(review): narrow header, drop vestigial guard, retarget tests

* no-mistakes(review): revert coarse gate override; require answered-down probe

* no-mistakes(review): cache one daemon probe; stop duplicating run id

* no-mistakes(review): restrict coarse dead-daemon verdict to moved-off rows

* no-mistakes(review): delete coarse dead-daemon extension and gate note

* no-mistakes(review): delete remaining coarse dead-daemon block and stale docs

* no-mistakes(document): document rebase-safe live-run bind and unverified-record verdict
…4994)

* fix(bin): stage the launch command in a private file and type a short source line

A long launch line typed while the fresh pane shell is still busy waits in the
terminal's canonical line buffer, which drops input past about 1,024 bytes on
macOS, so the pane was left at an unfinished command with no agent running.
fm-spawn now writes the assembled command to the task's own temp root under
umask 077 and types only a short line that sources it.

Refs kunchenguid#4559

* fix(bin): keep the per-task temp root private before staging the launch command

The root lives at a predictable path under /tmp and now holds the whole launch
command. Create it with mode 0700, refuse one that already exists as anything but
a directory owned by this user that nobody else can write, and tighten an owned
one, so no other local user can plant or swap the staged file.

Refs kunchenguid#4559

* fix(bin): enforce private staged launch file mode

* test(spawn): cover long staged Claude launches

* no-mistakes(review): Namespace launch files and prove truncation staging

* no-mistakes(review): Use immutable per-spawn launch filenames

* no-mistakes(document): Document staged launch delivery safeguards

* no-mistakes(ci): Updated eight behavior tests/fakes to execute or inspect immutable staged launch files instead of expecting inline launch commands. This restores Muse, secondmate lifecycle/restart, remote trace/parent binding, compact-adviser, and Orca coverage. All affected tests, dispatch-profile regression, fixture tests, syntax checks, ShellCheck, and git diff checks pass

---------

Co-authored-by: Vytautas Stankus <svycka@gmail.com>
* Add isolated Herdr runbook to test instructions

* no-mistakes(review): Drop substring matching from test.instructions contract

* no-mistakes(review): Assert commands.test key absence in YAML

* Drop unit-first sentence and instructions contract test

Captain-scoped follow-up on the Herdr-lab test.instructions ship:
keep the lab safety runbook only, and leave the no-mistakes contract
test focused on commands.test absence.
…uid#4873) (kunchenguid#5001)

* docs(vision): accept vendor-semantics and 9k contract-ceiling amendments (kunchenguid#4873)

Replace the pixels-of-today's-UI rule with a quarantined, version-pinned
surface-adapter exception recorded as standing debt. Cap the always-loaded
contract at 9,000 words and require prune-or-trigger before a crossing change
lands.

Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

* docs(vision): restore accepted three-sentence vendor-semantics form (kunchenguid#4873)

Replace the compressed paraphrase with the issue's accepted wording:
a named quarantined version-pinned adapter, expected to break, recorded
as standing debt that never hardens into a shared contract.

Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
…or-owed gate (kunchenguid#4974)

* fix(watch): recheck a gate awaiting a human instead of wedge-escalating it

A lane whose validation run is parked at a gate waiting on a human
decision is correctly quiet, but nothing in its status line says so: the
evidence is the pipeline's own gate state rather than anything the worker
wrote. The wedge timer read that silence as a suspected wedge and climbed
the escalation ladder for as long as the wait lasted, and each escalation
cost a supervising turn. The landed declared-wait consult does not reach
it, because a live ordinary crewmate never reports a declared pause, and
raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for
every lane by the same amount.

The threshold now reads a second, independent record when the status line
accounts for nothing: whether the crew's current state is a gate whose
answer is owed by a human. That is minted only from the gate's own
findings table, by a row whose `action` column is exactly `ask-user`,
located by position out of the table header the way nm_gate_step_row
already reads its row - never searched for over the run payload, where a
finding's free-text description or a branch name satisfies a search just
as well. A gate awaiting the CREWMATE's own answer keeps the unchanged
escalation schedule, reason and demand-deep-inspection wording, because a
crewmate that goes quiet before answering its own gate is exactly the
wedge the ladder exists to catch.

Each kind of wait now carries the human it is on, the action that clears
it, and whether that human is the captain as data alongside the verdict,
rather than as wording chosen per branch where the recheck is written, so
the deferral cannot word one kind of wait as another and a new kind
cannot ship without deciding all of them. A parked gate has no written
record of when its wait began, so its recheck publishes no wait age at
all rather than one read from the quiet window this deferral resets on
every pass, which would report the same small number for a gate of any
age. Like every other captain-facing recheck here it is absorbed in
silence while the away-posture record exists, arming no throttle, so the
recheck is owed in full the moment the record is archived.

The consult runs only in the at-threshold branch that was about to
escalate, beside the worktree walk already there, and only for lanes
whose status line explained nothing.

Closes kunchenguid#3055

* no-mistakes(review): require an unanswered decision before deferring a parked gate

* no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records

* test(watch): pass the pane hash wedge_timer_check now takes

Upstream gave wedge_timer_check a sixth <pane-hash> argument for its
dead-record probe. The malformed-wait-record rounds drive the real function
directly, so they pass one, and stub fm_backend_agent_state to a live agent so
the probe that runs after a refused deferral keeps the unchanged ladder rather
than reading a backend the child shell has none of.

* no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate

* no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling

* feat(watch): make the parked-gate wait deferral opt-in

The wedge timer deferring a lane parked at a validation gate is new
supervision behaviour rather than a restored one, and it decides which
lanes give up the escalation ladder, so it now ships as a default-off
per-home option instead of changing every home on upgrade.

config/wedge-defer-parked-gate arms it. The flag is read before the
decision fold, so an unconfigured home spends no fold or current-state
read, writes no record, and keeps the unchanged escalation schedule,
reasons and demand-deep-inspection wording; a test counts the reader
calls in both directions to pin that.

It is not inherited by secondmate homes: each home supervises its own
crew and owns that trade separately, the same reason
config/turnend-churn-absorb is home-local.

The away-posture absorb returns to leaving the idle timer alone, which
it had restarted only because the costly consult could reach it. A
parked-gate wait is owed to the supervisor rather than the captain, so
it never enters that branch, and the recheck owed on return is again
owed in full the moment the record is archived.

* test(watch): pin that the away-silenced hold leaves the idle timer alone

The absorb no longer restarts the timer, so the recheck owed on return is
owed in full rather than a cadence into the return. Nothing asserted
that, so a restart could be reintroduced silently.

* no-mistakes(review): document away-silence rationale, pin captured gate component

* no-mistakes(test): anchor gate row scan to the braced findings header

* no-mistakes(document): pin same-block gate row invariant in crew-state comment
…uid#5007)

* fix(control): let the owning seat reclaim a task whose endpoint is gone

A destroyed pane or workspace made `missing` a terminal state. Relaunch
accepted only `dead` and said to stop the agent first; exit refused
`missing` and said to reconcile the task first; there is no reconcile
verb. Each command named the other as its prerequisite, so a task whose
terminal went away could not be reclaimed by anything, and a no-mistakes
approval it was parked on had no seat left to answer it.

`missing` is agent-free a fortiori: there is no endpoint, so there is no
agent in it. Widen the existing guards rather than add a verb.

- fm-spawn --relaunch accepts a positively proven `missing` and creates
  one fresh endpoint in the recorded worktree; the record it already
  republishes rebinds the task to it. A `dead` endpoint is still adopted
  in place.
- fm-control exit reports `endpoint-gone` instead of dying, so the
  relaunch transaction's stop step no longer dead-ends, and re-resolves
  the endpoint from the record before verifying the replacement.

The duplicate-agent refusal is untouched: both verdicts come from the
same recovery-grade classifier, which claims `missing` only from positive
absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The
backends' own create paths refuse a live same-labeled endpoint as a
second independent guard. The worktree, its branch, commits, uncommitted
changes, armed poll and registration, record rows, and status log are all
untouched - a reclaim is a recovery, never a teardown.

A secondmate is excluded: its gone-endpoint recovery already has one
owner in the session-start liveness sweep, so relaunch refuses and names
it rather than becoming a second path to the same outcome.

Tests reproduce both halves of the deadlock, the reclaim succeeding,
unlanded work surviving it, and the refusals that still hold.

* no-mistakes(review): prove endpoint absence per backend before reclaim rebinds

* no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session

* no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly

* no-mistakes(review): stop refusals and docs asserting unestablished causes

* no-mistakes(review): stop herdr fixture helper losing tmp-root registration

* no-mistakes(review): document workspace drift and absence-probe server residue

* no-mistakes(review): correct rebind limitation to its one reachable case

* no-mistakes(review): stop claiming reclaim leaves instructions untouched

* no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage

* no-mistakes(rebase): read the staged launch file in the herdr fixture

Rebasing onto main picked up kunchenguid#4994, which stages a long worker launch
command into a script and delivers the short `. '<path>'` line instead of
the literal command. The tmux fake and tests/fixtures.sh were updated for
that; the herdr fake this branch adds was written before it and still
keyed "an agent now exists on this pane" off the literal
`encode launch-brief` text, so after the rebase it never marked the
rebound pane live and the reclaim's alive-wait read `dead`.

Dereference the staged file first, exactly as the tmux fake above does.
Test-fixture only; no production path changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(document): note reclaim placement in herdr and scripts inventories

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…3764)

* test(status): reproduce missing event emission time

* wip(status): preserve optional event emission time

* test(status): document indirect clock stub invocation

* no-mistakes(review): Preserve historical status bytes during reply recovery

* no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies

* no-mistakes(review): Preserve captain regex overrides for timestamped status events

* no-mistakes(document): Clarify status event timing and publication contracts

* no-mistakes(lint): Quote literal done to satisfy ShellCheck

* no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed

* no-mistakes(test): Preserve terminal notifications with malformed timestamp tags

* no-mistakes(test): Stamp Rovo spawn failures with emission time

* no-mistakes(document): Verify status event documentation

* no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests

* no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed

* no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed

* no-mistakes(review): Stamp remote escalations at call sites, drop new flag

* no-mistakes(review): Accept stamped escalation and close lines in test assertions

* no-mistakes(review): Restore reserved-key answered-note guard for stamped closes

* test(status): accept optional emission time in PR-provenance assertions

The kunchenguid#4148 provenance test landed on main with exact unstamped greps.
Parent-channel lines from this branch carry [at=<epoch>], so strip only
that tag before the same exact match. No production change.

* no-mistakes(review): Accept stamped ready signal in PR fallback scrape

* no-mistakes(review): Drop relay flag, stamp parent events at call sites

* no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture

* no-mistakes(review): Accept optional stamp in live cmux drift guard

* no-mistakes(review): Restore original test invocation order in two suites

* no-mistakes(review): Strip only well-formed numeric status time tags

* no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc

* no-mistakes(review): Stamp agy spawn-failure status lines with event time

* fix(bin): normalize status event times in-shell and freeze the budget test clock

Two paths made a status event's emission time cost more than it should.

The captain-relevance fallback piped every line through awk to drop a
well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a
fork per line just to prepare a regex match. Shell parameter expansion does the
same strip with no fork, and the retry-dedup scan now reuses that one helper
instead of carrying a second copy of the rule in awk. The copies had already
drifted: the shell side stripped tags from lines with no colon, which the awk
rule left whole, so a colonless line could be mistaken for one already
recorded. One definition, checked against the awk rule it replaces over the
edge cases and a 4000-line fuzz.

tests/fm-contributions.test.sh froze its fixture clock only in exhaust mode. In
hang mode the poll set DEADLINE to the real now plus a one-second budget, and
when the second ticked before the first forge call the loop broke without ever
calling gh: forge/calls was never written and the assertion failed reading a
missing file. Freezing the clock in both modes removes the dependence on wall
time; the bounded call is still cut by the real timeout, so the observation the
test asserts still starts.

Emission time stays optional on new status records, and legacy or malformed
lines keep an unknown age.

* no-mistakes(review): Stamp ask-user escalation line and fix Kimi status assertion

* no-mistakes(document): Drop stale unstamped done-line spelling from watcher docs

* test: fold emission-time snapshot coverage into the fixture case

Drop the incidental ci.yml 18-to-19 count hunk so the PR no longer
touches workflows. Keep every emission-time assertion by folding it
into test_fixture_snapshot_json.

* no-mistakes(review): replace brief date substitution with epoch placeholder; drop emitted_at_epoch

* no-mistakes(review): align untimed normalizer with epoch parser; tolerate placeholder stamp in PR scrape

* no-mistakes(review): strip undelimited at-tags; correct brief stamp header

* no-mistakes(review): normalize stamps at both captain-regex sites; restore mtime freshness

* no-mistakes(review): strip colon-bearing stamps for relevance; fix headers and test oracles

* no-mistakes(review): narrow escalation match to stamp tolerance; pin note verb

* no-mistakes(review): read note and key past colon-bearing stamps

* test(status): keep inactive reconcile assertions stamp-tolerant

These two oracles were made stamp-tolerant while resolving one of the
branch's merges from main. The rebase drops merge commits, so that
adaptation was lost and both assertions went back to matching an exact
substring that a stamped line no longer contains: the tag lands before
the colon, so "failed [key=k]: ..." is now "failed [key=k] [at=N]: ...".
Strip a well-formed tag before matching, as the branch's other oracles do.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): unstamp fold colon tests; reserve stamp width in cap

* no-mistakes(document): correct stale unstamped status-line spellings in docs

* no-mistakes(document): quote brief-test literals for lint; correct stamp-helper contract comments

* no-mistakes(ci): rename subshell-local epoch in delivery-race stub

The serialization test overrides fm_pending_reply_mark_delivered inside a
(..) subshell. Its `epoch` local collided with the same name in
status_line_at_epoch/status_stamp_line, which this branch added and this
suite now calls at top level, so ShellCheck 0.11.0 reported SC2030 and
failed Lint 2. The stub already prefixes its other locals with `pending_`
for the same reason; `epoch` was the leftover.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: ship clean Lavish host fixes

* no-mistakes(review): Fix Lavish classifications and fail-closed host loading

* no-mistakes(review): Restore Lavish host state across retries and launches

* no-mistakes(review): Preserve destination Lavish host when configuration is absent

* no-mistakes(document): Document Lavish status and host guarantees
…#5076)

* feat(afk): make the captain's away words the whole mandate

Retire the clause fields, verb list, never-set scan, refused records, and
the per-task merge-grant list from the away-posture record. The record is
now version 2: the captain's words verbatim plus expected return, spend
cap, and reach line; a version 1 record still validates, reads, and
archives so a live away window is never broken by the upgrade.

The supervision branch reads the words at the tail of every wake and acts
on them by its own judgment through the guarded scripts under standing
authority, never by analogy, holding for the return on doubt, and opens
each such outcome summary with "per your away instructions:" so the
return brief can render the words beside the session's account. While the
record exists any green merge runs under away authority (ledger tag
"away"); red merges, --allow-red, asynchronous and queued merges, and
local-only landing stay refused. The branch may file a backlog item the
words explicitly call for before dispatching it under the spend cap.

Tests drive fm-afk-contract.sh, fm-afk-launch.sh, fm-afk-return.sh, and
fm-pr-merge.sh as commands: version 2 written, version 1 read, retired
flags and subcommands refused by name, green merges landing under the
record, red and waived-red refused, the record lock still closing the
authority-read window, and the Pi away tail carrying the words.

* no-mistakes(review): carry the away read-back to the session verbatim

* no-mistakes(review): match the exact away-action marker in the return brief

* no-mistakes(review): refuse a words block truncated by a damaged line

* no-mistakes(document): Refresh away-role contract documentation
…unchenguid#5049)

* fix(bin): render the remote charter's steering-inbox path host-local

A freshly provisioned remote secondmate read a parent-home absolute
steering-inbox path in its charter - a location that exists on no route -
and spent its first turn discovering the gap and filing a blocked
decision for what was a render defect. The seed's remote-copy rewrite now
maps the inbox to the route's host-local parent-route inbox, exactly as
it already maps the reply-log path, so every mention - bare path, listing,
and handled/ acknowledgement - lands host-local.

Both rewrites also become plain assignments, because a quoted substitution
nested inside a double-quoted printf argument leaks literal quotes into
the replacement text on stock macOS bash. The lifecycle suite pins the
corrected render both directions against the real seed, provisioning,
and delivery route, sharing one fixture value between the render truth
and the delivery truth.

Closes kunchenguid#5012

* no-mistakes(document): document remote charter's host-local steering inbox
)

* feat(procevent): route worker-owned Lavish rounds

* no-mistakes(review): drop duplicate artifact field from task-owned registration

* no-mistakes(review): post worker reply once, fix ring label, keep re-arm atomic

* no-mistakes(review): keep worker board owned until terminal round acknowledged

* no-mistakes(review): refuse every retirement of an open worker-owned round

* no-mistakes(review): use real lavish reply flag, isolate reply generations

* no-mistakes(review): drop .posted marker for best-effort reply posting

* no-mistakes(review): consume staged reply after listener setup, refuse orphaned captures

* no-mistakes(review): require a reachable owner, redeliver open rounds, roll back failed re-arms

* no-mistakes(review): re-arm only to acknowledge an open round

* no-mistakes(review): conclude only a still-open terminal round

* no-mistakes(review): record the acknowledgement before retiring the board

* no-mistakes(review): retain the registration across a conclude, qualify terminal docs

* no-mistakes(document): Document worker-owned Lavish round lifecycle
…unchenguid#5107)

* fix(bin): reserve contribution observation budget

* no-mistakes(review): Strengthen slow-read regression test to exceed the poll budget
…ness JSON (kunchenguid#5103)

* feat(bin): add idempotent inbox orders, receipts, replies, and readiness

Let a caller supply a request id when publishing a captain inbox note so a
retry returns the original note instead of creating a second one, including
across the crash window between save and wake announcement. Separate saved
from announced so a failed wake is repairable without enqueueing again.
Add bounded receipts JSON with omission disclosure, a durable primary reply
against a note id, and a read-only readiness projection that can say
unknown instead of inferring liveness from a lock file.

* no-mistakes(review): fix(bin): honest inbox announce, reply cursor, and readiness verdict

* fix(bin): resolve ready from lock-holder ancestry; drop lock status --json

Remove the extra JSON surface from fm-lock.sh so its human status still
always exits zero. Have the readiness projection classify the inspected
home from the lock-holder pid via fm-harness.sh ancestry, with an explicit
FM_SUPERVISION_MODEL still winning and an unknown model when there is no
holder. Prove the yes path when that ancestry names a known harness.

* no-mistakes(review): Harden inbox announce, receipts reads, and reply sequence cursor

* no-mistakes(document): Note read-only lock inspection in scripts inventory

* no-mistakes(lint): Pass missing id argument to malformed-reply test printf

---------

Co-authored-by: cliflacata-svg <304148223+cliflacata-svg@users.noreply.github.com>
…ending text (kunchenguid#5118)

* fix(composer): stop a harness footer row from reading as a composer holding text

A harness draws its own furniture below the composer - a user statusLine, a
permission-mode hint - and the cursorless "bottom-most shape wins" rule looks
exactly there. `→` (U+2192) is Cursor's prompt glyph but ordinary text
everywhere else, so a statusLine opening with `→` was selected as a bare
composer, swallowed the hint row beneath it as wrapped input, and answered
`pending` on a visibly empty pane. `fm_task_inbox_ring` defers on exactly that
verdict, and `bin/fm-watch.sh`'s re-ring calls the same function, so the first
doorbell and every retry were skipped and the worker never saw the steer.

Measured live on 2026-09-20: three of five Claude Code 2.1.236 worker panes on
Herdr 0.8.0 had genuinely empty composers and every one of them was refused.

A separator pair that closed over a bare agent-glyph row is a proven composer
container, so the contiguous non-blank rows below its closing rule are that
composer's footer and are no longer composer candidates. The demotion is bounded
by all three of its own preconditions: a blank row ends the zone, a pair that
closed over no glyph row demotes nothing, and a shape with no separator pair at
all (Cursor's half-block rules) is untouched. Real unsubmitted text in that same
composer, including a stray SGR mouse report left by a click in the pane, still
reads `pending`.

Pinned by two portable regressions and by a new cursorless arm on the live
composer-matrix guard, which re-reads each harness's already-proven-idle pane
the way every non-tmux backend reads it and fails naming the harness and
version when that read is `pending`.

* no-mistakes(review): make composer footer-zone demotion shape-independent

* no-mistakes(review): make footer-zone demotion refuse-only and drop rescan

* no-mistakes(lint): quote probe-absent sentinel to clear ShellCheck SC2100

---------

Co-authored-by: Koen Muller <koen@catapult.nl>
…5115)

Co-authored-by: guanchengh-lgtm <271917158+guanchengh-lgtm@users.noreply.github.com>
… an unreadable runs table (kunchenguid#5114)

* fix(bin): stop misreading a no-run branch as an unreadable runs table

Defect: when `no-mistakes axi status`'s overview is truncated (a task's
own branch has zero rows among the shown ones), fm_nm_select_run's
Python fallback derived the repo identity for its direct SQLite query
from a `repo: <path>` line it expected in the overview text. The real
CLI never emits that line, truncated or not (see the genuine capture at
tests/captures/no-mistakes-v1.70.1/overview.toon, which has only
`count:`/`runs[...]:`), so the lookup always failed and reported
"unreadable runs table" for a task that simply has no run on its
branch. On a fleet with many concurrent runs, every idle-branch task
hits the truncated-overview path routinely, so this fired every few
minutes and drowned genuine unreadable/blocked verdicts in noise.

Fix: derive the repo identity from the task worktree path instead,
which is exactly the value `no-mistakes` records as a repo's
`working_path` (confirmed against the existing capped-overview test
fixtures, which already register repos by worktree path). A worktree
path that is not absolute cannot be matched and still reads as
unreadable rather than being guessed at. Also raise the reader's
SQLite busy timeout from 1s to 30s so ordinary lock contention on a
busy fleet cannot masquerade as an unreadable database.

Safety: every other verdict byte-for-byte unchanged - the repo lookup
still requires exactly one matching row (a genuinely corrupt or
mismatched repos table still reports unreadable, per the existing
`repo` failure-mode test), the branch query and row validation are
untouched, and a zero-row result for the branch still flows through
the same recursive re-parse that already turns an empty `runs[0]{...}`
table into `absent`. Added a regression test
(test_capped_overview_without_repo_line_and_no_runs_reports_absent)
that reproduces the real overview shape - capped, zero rows for the
task's branch, no `repo: ` line - and asserts the crew state falls
through to the pane/busy verdict instead of reporting unknown or
"unreadable". Full fm-crew-state.test.sh suite passes unchanged
otherwise.

* fix: recovered same-branch inventory awk misreads empty result as unreadable

fm_nm_select_run's deep SQLite reader rebuilds a `count:`/`runs[...]:`
overview and re-runs it through the same awk selection pass. When that
rebuilt inventory has zero rows for the branch, the row-matching loop never
executes, so its counters (`seen`) stay at awk's uninitialized empty string
while `expected` and `shown` are plain strings parsed from the header text.
Comparing an uninitialized value against a non-numeric string uses string
comparison, so "" != "0" is true, and the END block takes the "unreadable
runs table" branch instead of falling through to the correct "absent"
verdict for a branch with genuinely zero runs.

Coerce the affected END comparisons with `+0` so they are always numeric,
matching seen/expected/shown/total regardless of whether awk classified
them as strings or numeric strings. A truncated or genuinely malformed
inventory still differs numerically and still reports unreadable.

* no-mistakes(review): bound capped-overview inventory reader and canonicalize worktree lookup

* no-mistakes(review): match recorded repo path first, tolerate duplicate spellings

* no-mistakes(review): revert repo lookup to exact working_path match

* no-mistakes(document): note state-db inventory read under crew-state nm timeout
…ort (kunchenguid#5141)

* fix(bin): require a non-draft pull request before a PR-based done report

A PR-based ship could report done, and merge monitoring could be armed, while the pull request was still a draft. A draft cannot be merged, so the poll waited for an event that could not occur and nobody was asked to merge.

The PR-based definitions of done now require reading the pull request back from the forge and confirming it is not a draft, and a lane that deliberately holds a draft declares a wait instead of done.
bin/fm-pr-check.sh refuses to arm merge monitoring on a draft, naming the draft state, and treats an unreadable draft state as before.
The draft reading now lives in bin/fm-pr-lib.sh and bin/fm-pr-merge.sh uses it, with its refusal to merge a draft unchanged.

Closes kunchenguid#4757

* fix(review): Skip arm-time draft refusal when fm-pr-merge records metadata
* fix(bin): accept quota-axi schema 6 snapshots keyed by provider + accountKey

quota-axi 0.1.47 emits schemaVersion 6 once a provider expands to more
than one account: every provider row carries an accountKey and one
provider id may appear on several rows. fm_quota_json_valid accepted
only schema 5 with unique provider ids, so fm-dispatch-resolve.sh,
fm-quota-choose.sh, and fm-procevent-quota.sh all rejected the live
snapshot and quota-informed dispatch was dead against the current tool.

- bin/fm-quota-axi-lib.sh: the validator accepts schema 6 with
  accountKey required on every row and uniqueness on
  provider + accountKey; schema 5 keeps its exact rules. FM_QUOTA_ROW_JQ
  is the one join every consumer uses: schema 5 binds by provider alone,
  schema 6 binds to the row keyed by the candidate's Pi lane, else the
  provider's default row, else no row (unmeasured, never blocked, never
  by position or summed across accounts).
- bin/fm-quota-choose.sh: accepts schema 6 JSON and the TOON accountKey
  column, and joins through the shared function.
- bin/fm-dispatch-resolve.sh and bin/fm-procevent-quota.sh: join through
  the shared function; an expanded provider with no row for the
  candidate's account is reported as such.
- tests: schema 6 fixtures shaped like the real snapshot, each paired
  with a schema 5 case on the same path; every new case fails on the
  previous scripts and passes now.
- docs: the two sentences naming the row join describe the schema 6 key.

* no-mistakes(review): Fix native Codex quota and expanded provider watches

* no-mistakes(review): Align native Codex account matching across dispatch paths

* no-mistakes(document): Align quota documentation with account-aware snapshots

* no-mistakes(document): Align quota dispatch documentation with account matching

* fix(bin): keep CI lint and the quota watch test portable

- bin/fm-quota-axi-lib.sh: FM_QUOTA_ROW_JQ is read only by the scripts
  that source this library, so full-mode ShellCheck reported SC2034 on
  the assignment; mark it alongside the existing SC2016 disable.
- tests/fm-procevent-quota.test.sh: the schema 6 provider-watch
  assertions used rg, which CI runners do not install, so the case
  failed with 'rg: command not found' rather than on behavior; use grep
  like the rest of the file.

* no-mistakes(document): Documented schema-version account-row compatibility
* test: repair Claude live auto-arm regression

* no-mistakes(review): Assert SessionStart digest completeness within its hook_response event

* no-mistakes(document): Consolidate Claude live verification references
Roll the shared require-no-mistakes action to the tagged v1.80.1 SHA and grant pull-requests: read so the check can read PR bodies.
sdivanl and others added 11 commits September 21, 2026 19:30
…nchenguid#5174)

* fix: preserve Pi watcher ownership across session replacement

* no-mistakes(document): Scope Pi predecessor retention away from omp

* no-mistakes(ci): Diagnosed all three failing checks; only one was code-caused. (ci-3, genuine) Stock macOS Bash snapshot compatibility: `tests/fm-pi-watch-extension.test.sh` failed the macOS Bash 3.2 `bash -n` parse sweep with `line 4265: unexpected EOF while looking for matching '`. I built GNU Bash 3.2.0 from source locally and reproduced it. Root cause: the PR added a comment containing an apostrophe (`// Replacement shutdown deliberately retains module 2's established arm until`) inside a quoted here-document (`<<'EOF'`) nested inside a `$(...)` command substitution. Bash 3.2 has a parser bug (fixed in later bash) where an unmatched single quote inside such a here-doc body is treated as opening a shell quote and never closed, aborting the whole file parse. The base commit parses cleanly under Bash 3.2, confirming this PR introduced the break. Minimal fix: reworded the comment to remove the apostrophe (`... retains the established module-2 arm until`), preserving meaning. Verified `bin/fm-lint.sh --list-files` (the 6 changed shell files) now all pass `/tmp/bash-3.2/bash -n`; Bash 5 also parses. (ci-1, infrastructure) Behavior portable serial 8: GitHub API shows the `Run portable serial shard 8` step conclusion=success; only `Upload portable serial shard 8 timing artifact` failed with `Failed to FinalizeArtifact ... (403) Forbidden`. This is a transient artifact-service/cancellation failure, not a test or code failure. No change. (ci-2, infrastructure) Lint 1: fetched the job log via the GitHub API; it ends with `##[error]The runner has received a shutdown signal...` then exit 143. The step was cancelled mid-run, not a ShellCheck finding. Independently ran `bin/fm-lint.sh --partition 1of2 --telemetry ...` locally with pinned ShellCheck 0.11.0 and actionlint 1.7.12: exited rc=0 (no findings). No change. The only code change is the apostrophe removal in tests/fm-pi-watch-extension.test.sh; no other files modified
…d#5236)

* fix(bin): retire windowless leftovers and stop claiming a Pi daemon teardown

Catch-up correctly refuses while a leftover task record has no status file.
Cleanup used to deadlock on those same records when they also had no spawn_gen and no window, so they lingered and wedged every later away-mode return. Teardown now treats a windowless leftover as a missing-endpoint legacy record, and stop reports that no daemon terminal was running when none was launched.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Narrow windowless teardown exception to tmux legacy leftovers

* no-mistakes(review): Validate windowless leftover identity via shared endpoint validator

* no-mistakes(review): Refuse windowless leftovers carrying other backends' endpoint identity

* no-mistakes(document): Clarify windowless teardown retry documentation

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…rted (kunchenguid#5250)

* fix: surface parked launch prompts as not started

* no-mistakes(document): docs: record launch-prompt busy backstop classification

* no-mistakes(document): docs: align tail40 and rendered-text comments with launch-prompt backstop
* feat(afk): make /afk itself the go with a same-turn record write

Collapse the propose-then-confirm away entry into one 'enter' step that
writes state/.afk-contract immediately and prints the announcement and
read-back after the record exists, never asking for a go. The retired
propose, confirm, and --proposal inputs are refused by name, and a stale
proposal left by an older version is removed rather than promoted.
Refresh and replace semantics, verbatim words, the single writer, the
never-set, and per-harness launch behavior are unchanged.

* no-mistakes(document): Refresh away-entry documentation evidence
…nguid#5294)

* fix(bin): map passed-with-override to done instead of unknown

no-mistakes' axi status emits outcome: passed-with-override for a run
that finished with an explicitly approved Test or CI exception. Both
bin/fm-crew-state.sh's outcome resolver and bin/fm-teardown.sh's
pre-teardown terminal-run check only matched the literal passed and
checks-passed tokens, so this outcome fell through to unknown/parked
and a finished worker awaiting merge kept getting re-alerted as stale,
while an abort race during teardown could also leave a finished run
misreported as still parked.

Map passed-with-override to the same done/terminal handling as a
clean passed in both places.

* fix(document): Replace stale outcome mapping with authoritative pointer

* fix(ci): Fixed a pre-existing mock-clock race in tests/fm-contributions.test.sh by advancing time only during the serial issue read. Reproduced the exact CI failure before fixing it. Forced-race replay, all 38 contribution scenarios, scoped ShellCheck, Bash syntax, and diff checks pass. Only the test fixture changed; CI rerun remains with the outer executor
Merge the frozen September 22 upstream snapshot while retaining the fork's
native identity, private publication, Azure, recovery, and module boundaries.
Compose immutable launch delivery and shared endpoint-absence proof with the
existing Windows transport and guarded recovery transaction.
Carry incoming test registrations and timing hints through the catalogs,
preserving independent concurrency admission and optional delivery validation.

Firstmate-Upstream-SHA: c5131a3
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Update Azure status assertions for timestamped reports, preserve unknown
lock ownership and guarded endpoint-refusal semantics, install the complete
mail-check fixture dependency set, and retain the bounded cleanup verdict.

The compact-adviser relaunch fixture replaced sleep with a no-op, causing
the real watchdog to expire immediately. Use real sleep and the established
relaunch fixture bounds without changing any production deadline.

Register the affected legacy suites with the shared named-case runner while
preserving their full-suite order.

Watcher source-aware lint reproduces unbounded memory growth through its
import graph. Apply the existing separately analyzed canonical-owner boundary
to pending replies and the new live test's production imports, retaining
full lint partitions, source-aware owners, and all analysis flags.

Firstmate-Upstream-SHA: c5131a3
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Capture source object inventories, the failing object and destination directory, and Git trace events before rollback erases the evidence. The complete fixture and source-aware lint pass locally, while the unchanged CI failure has repeated twice. This diagnostic does not change production runtime or relax assertions.

Firstmate-Upstream-SHA: c5131a3
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The instrumented compact-adviser suite passed, but the same clone-copy failure appeared in the independent trace-context suite. Remove the fixture Git wrapper and capture object-store, destination, and rollback identity evidence only after the original Git command fails. Preserve the existing failure status and rollback contract.

Firstmate-Upstream-SHA: c5131a3
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
All 25 checks passed on the shared diagnostic head. Neither repeated local provisioning, eligible fixture auto-maintenance, nor a controlled source repack reproduced the intermittent CI clone failure. Retain that uncertainty in PR evidence rather than shipping a speculative runtime change. Restore the exact ten-file repair tree and remove all temporary diagnostic code.

Firstmate-Upstream-SHA: c5131a3
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@timbarreto
timbarreto merged commit b13de80 into main Sep 23, 2026
42 of 43 checks passed
@timbarreto
timbarreto deleted the reconcile/upstream-2026-09-22-c5131a33 branch September 23, 2026 00:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.