Skip to content

Reconcile upstream b805823a with fork module boundaries - #76

Merged
timbarreto merged 70 commits into
mainfrom
reconcile/upstream-2026-09-25-b805823a-b382ab77
Sep 25, 2026
Merged

timbarreto merged 70 commits into
mainfrom
reconcile/upstream-2026-09-25-b805823a-b382ab77

Conversation

@timbarreto

@timbarreto timbarreto commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

Frozen snapshot and landing requirement

Use a merge commit to land this PR. Do not squash or rebase it: upstream ancestry is part of the reconciliation contract. This is an ordinary, non-draft PR to leave unmerged; auto-merge is not enabled.

Identity SHA
Frozen fork / actual PR base b13de801f35b53757a8c61ea18d8c577b07a7aeb
Frozen canonical upstream b805823a6beef7f59b2a10fe50ae6054c39cf571
Proven prior upstream c5131a33a1e35a42e34733a5334fcc4e0225a656
Reconciliation head 85c1d779065e7447727aecf5af2c15a1870444e5
Reconciled tree 860e6a757cc9da66aee4013d0f74678c85075ded

Firstmate-Upstream-SHA: b805823

The prior point is proven by the parsed trailer on 4159f4965393456b988cd22f8c294c53edeacbab, retained head of PR #75, whose merge is the frozen fork base. It is an ancestor of both frozen inputs and their merge base. The normal two-parent reconciliation merge is b4a3eb287c9442656e2009344bac6d31a1e7a0e0; subsequent commits repair CI dependencies and integration wiring. This is not reconstruction and no ancestry-only anchor was needed. Canonical upstream was fetched once and has remained frozen.

Work is isolated on reconcile/upstream-2026-09-25-b805823a-b382ab77 in a separate worktree. Original local main remains clean at the frozen fork SHA, with its recorded index SHA-256 unchanged. No live Firstmate fleet, session-start, credentialed vendor prompt, or no-mistakes execution was used.

Conflict and ownership decisions

Resolved 55 conflicted paths from both histories. Per-hunk choices and source manifests are retained in external session evidence.

  • Devin worker/scout support, guarded double-Escape, and devin-hook coexist with Copilot and the closed Copilot/Pi adapter seam; secondmate support is not broadened.
  • New pilot interrupt-safety queries live in each pilot adapter, with interface assertions. Shared nonpilot and hash-pinned doctor behavior remains legacy.
  • Supervision-host opt-in composes with OpenCode's explicit Bash launch, per-child owner tokens, and native Windows graceful/forced retirement. Copilot stays on its asynchronous shell protocol; Pi keeps its native branch extension.
  • The shared native process owner recognizes an owned supervision host as well as an owned watch arm. Native fixtures assert that each tokenized tree can be retired without touching the other or a foreign tree; their execution remains CI-owned.
  • New Devin configs secure the empty temporary file through the shared native private-path owner before copying user settings; POSIX keeps umask 077. Config failures preserve the prior valid publication. The existing Windows reconciliation core job picks up the new native success and pre-payload refusal case in fm-private-path.test.sh.
  • Shared secondmate liveness is actually sourced by both bootstrap and watcher. Copilot is included in the verified local recovery allowlist.
  • Decision-key parsing preserves the fork's caller-owned destination argument; the upstream keyless sentinel is the third argument. Declared-wait callers use that contract.
  • Fleet snapshots add immutable branch metadata to the existing one-pass reader, not a new per-field process. Lint exposed and fixed the missing local declaration for branch.
  • Startup network harvest stays lock-free, while upstream's bounded publication wait and worker wake behavior are retained.
  • Azure private registration, observed live source heads, backlog link retention, and notification ordering coexist with Gerrit. Neither unsupported forge gains merge authority.
  • Spawn retains Windows command transport, resolved pilot executable, launch-status generations, rollback, and lifecycle ownership while adding account pins, branch prefixes, immutable branch metadata, and Devin configuration.
  • Conflicted test registries retain fork cases in order and append upstream cases. The new worker-account and host suites use the existing named-case helper without changing default case ordering.
  • Documentation keeps native Windows privacy, cleanup, presentation, Copilot, and publication safety facts while adopting upstream headings and host/Gerrit material.
  • Initial exact-head CI exposed tasks-axi 0.2.5 below the imported runtime floor 0.2.6. The remaining shared portable and fork Windows backlog-fixture pins now match 0.2.6; a workflow contract compares every such pin against the runtime owner.
  • The subsequent CI failure came from a new completion caller retaining meta_value after the fork extracted batched metadata reads. Mode and project now join that same batch, restoring the named-head refusal without extra processes.
  • The AFK return test's copied PR reader lacked the private-path closure. The existing fixture installer now supplies it, and the existing 23 cases retain their order through the supported named-case registry.

The two whole-file fork selections were deliberate: bin/fm-test-run.sh had only incoming inline metadata changes, migrated to tests/catalog/core.tsv; the fork's recursive Pi type fixture already carries new upstream extension files while preserving its repository-relative dependency closure.

Coupled verification surface

  • Shared CI adopts the incoming stock-macOS tasks-axi 0.2.6 update, Bearings assertion count 60, and two-assertion adapter-file regression. Exact-head CI then demonstrated that the remaining portable jobs and fork Windows backlog-fixture jobs also require tasks-axi 0.2.6; those four stale 0.2.5 pins are corrected. That pin is the fork workflow's only change. The workflow contract now compares every tasks-axi installation with the runtime minimum.
  • Eleven incoming test registrations plus routes/duration hints are in the core catalog. Fleet-ledger fixtures are classified as backend-dispatch instead of inheriting upstream's missing registration.
  • The catalog loader, runner algorithms, fork prior-value overrides, and independent isolation-proof owner are unchanged. Registration does not grant concurrency admission.
  • Pi 0.84.3, OpenCode 1.18.23, TypeScript 5.9.3, Herdr 0.7.4/protocol >=16, and Node 24 in native/package jobs are retained. Shared Linux portable jobs retain their existing Pi installation and compiler prerequisites. Timing/artifact dependencies and optional live gates are preserved.
  • The Windows Herdr experiment remains manual-only (workflow_dispatch), outside the automatic check set.

Initial bounded local evidence — not a full validation pass

The controller had a 2,400-second shared allowance and a two-timeout circuit breaker. It stopped after 903.639 seconds and two timeouts. The remaining allowance was not reset and the circuit breaker was not bypassed. No full local suite, dependency install, baseline differential, or retry-until-green was run. No task-owned validation process was observed remaining after completion.

Host: Windows with explicit C:\Program Files\Git\bin\bash.exe, Bash 5.3.15(2)-release, Git 2.55.0.windows.5, Node 24.19.0, Python 3.13.15, jq 1.8.2, Perl 5.42.3, ShellCheck 0.11.0, actionlint 1.7.12. All probed prerequisites were available; there was no missing-tool or live-capability pass.

Command Result Seconds CI owner
bash -c 'set -e; files=$(bin/fm-lint.sh --list-files); while IFS= read -r script; do /bin/bash -n "$script" || exit; done <<< "$files"' passed before subsequent bounded shell/test edits; final refresh deferred 66.977 Lint 1; Lint 2; Stock macOS Bash snapshot compatibility
bash bin/fm-lint.sh failed: two findings fixed afterward; rerun deferred 363.292 Lint 1; Lint 2
bash bin/fm-doc-audience-check.sh passed 8.948 Repo invariants
bash bin/fm-test-run.sh --check-coverage timeout 185.565 Test coverage guard
bash bin/fm-test-run.sh --list --changed --base b13de801f35b53757a8c61ea18d8c577b07a7aeb passed 65.121 Test coverage guard; Behavior portable parallel 1/2; Behavior portable serial 1-9; Behavior tests (Herdr)
bash -c 'set -e; node --check .opencode/plugins/fm-primary-watch-arm.js; node --check bin/fm-branch-dispatch.mjs; node --check tests/fm-platform-process.test.mjs' passed 2.106 Harness package compatibility; Behavior portable serial 1-9
bash bin/fm-test-run.sh --jobs 1 tests/fm-test-fixtures.test.sh timeout 181.701 Behavior portable serial 1-9
bash bin/fm-test-run.sh --jobs 1 tests/fm-harness-contract.test.sh tests/fm-test-catalog.test.sh tests/fm-supervision-instructions.test.sh tests/fm-classify-decision-key.test.sh tests/fm-pr-local-cost.test.sh deferred not run Behavior portable parallel 1/2; Behavior portable serial 1-9; Windows Copilot management
bash bin/fm-lint.sh bin/fm-control-lib.sh bin/fm-harness-lib.sh bin/harnesses/copilot.sh bin/harnesses/pi.sh bin/fm-devin-config.sh bin/fm-secondmate-liveness-lib.sh deferred not run Lint 1; Lint 2

Prerequisite probes: bash --version, git --version, shellcheck --version, actionlint -version, python3 --version, node --version, jq --version, and perl -e 'print "$^V\n"'. Their execution plus controller overhead accounts for the other 29.929 seconds. Test commands used FM_LIVE=0. No successful-result cache was enabled or reused.

The lint findings were a missing local declaration for the snapshot's batched branch field and the new Devin POSIX permission assertion's use of ls. They were fixed without adding metadata subprocesses or weakening permissions. One incoming trailing blank line at EOF was also removed for the whitespace gate. The session-only staging auditor initially assumed every .sh was executable; that assumption was corrected to compare inherited Git modes, without changing repository modes.

Both timeouts remain unclassified unresolved observations, not claimed platform baseline failures: the coverage guard reached its 180-second command deadline, and the Git-fixture policy suite reached its 180-second deadline before reporting case completion. No matching frozen-upstream differential was run. The mandatory local gate set therefore did not finish green.

After lint, the new host/account suites gained the existing named-case registry (default order preserved), host malformed-token assertions were added, and the already CI-owned private-path suite gained native Devin publication/pre-payload-refusal coverage. These final edits were not rerun locally after the circuit breaker.

Specific final-tree ownership: Lint 1/Lint 2 cover the complete source-aware roots and final shell edits; Repo invariants owns docs; Test coverage guard owns executed partition proof; Windows reconciliation (core) exercises the native host process tree and Devin private config via its existing process/private-path suites; Harness package compatibility owns real installed Pi/OpenCode layouts; Windows Copilot management and the other reconciliation subjects own retained native transport/rollback/Azure cases. GitHub Actions owns the complete cross-platform matrix.

No new live vendor support is certified. Account-pin files retain upstream's strict POSIX-absolute/ordinary grammar rather than inventing broader native account-root support during this merge.

CI observations and user-approved targeted follow-up

Initial head b4a3eb287c9442656e2009344bac6d31a1e7a0e0 exposed stale tasks-axi 0.2.5 pins against the imported 0.2.6 runtime floor (Fork CI 36166249695). After correcting those pins, head b9fa47b2799c0a90a14ad64a53d134401c09ea34 finished with 23 successful checks and two failures: portable parallel 2 and serial 7 (CI run 36167657667). Earlier-head passes are not final-head evidence.

The remaining failures were two integration defects: a new crew-state completion path still called the removed meta_value reader, and the copied AFK return fixture omitted the PR reader's private-path dependencies. Mode/project now join the existing one-pass metadata read, without new subprocesses. The fixture uses the existing dependency installer and retains all 23 default cases through the shared named-case registry. No completion guard, production deadline, or privacy check was weakened.

The user explicitly approved up to 600 additional command seconds and two additional timeouts. This follow-up used 485.804 seconds and 1 timeout; the original allowance was not reset. Crew-state reproduced the exact CI failure locally; the pre-fix AFK case hit its 120-second local bound, while its exact assertion/missing-dependency failure is preserved from CI. Both repaired named cases then passed, as did changed-shell syntax and source-aware local fast lint. Tool versions were already probed; cheap presence checks were reused, never test-result caches.

Follow-up command Result Seconds
FM_TEST_ONLY=test_moved_remote_branch_without_named_head_is_blocked bin/fm-test-run.sh --jobs 1 tests/fm-crew-state.test.sh passed 45.044
FM_TEST_ONLY=test_return_brief_lists_landed_work_awaiting_cleanup bin/fm-test-run.sh --jobs 1 tests/fm-afk-return.test.sh passed 124.105
set -e; bash -n bin/fm-crew-state.sh; bash -n tests/fm-afk-return.test.sh passed 1.371
bash bin/fm-lint.sh --fast bin/fm-crew-state.sh tests/fm-afk-return.test.sh passed 15.879

The same two named cases before repair returned failure (69.586 seconds) and timeout (128.848 seconds), respectively. The latter is not claimed as a proven platform baseline issue. Full-dataflow lint and broad lanes remain CI-owned; the final-head matrix must finish before merge readiness is claimed.

Disposition for all 257 selected scripts (257 root-level scripts total)

The changed-inventory command completed. Exact per-script CI owners below are derived from the resolved catalogs, unchanged literal parallel lists, and unchanged deterministic serial LPT assignment. This static routing audit is not a substitute for the timed-out executable coverage guard. Live-harness rows cover their gate on automatic CI, not credentialed live proof.

Every basename below is under tests/; broad suite execution is delegated to the named CI owner. The selected crew-state/AFK case outcomes are recorded above, and the earlier fixture-policy timeout is marked. [gate-only] means automatic CI does not provide live proof. Herdr runs only in its isolated backend job.

Behavior portable parallel 1

fm-brief.test.sh, fm-cd-pretool-check.test.sh, fm-composer-ghost.test.sh, fm-composer-lib.test.sh, fm-grok-harness.test.sh, fm-lint.test.sh, fm-pi-primary-types.test.sh, fm-pr-merge.test.sh, fm-review-diff.test.sh, fm-test-run.test.sh, fm-tmux-submit-busy.test.sh

Behavior portable parallel 2

fm-arm-pretool-check.test.sh, fm-backend-herdr.test.sh, fm-captain-hold-lifecycle.test.sh, fm-crew-state.test.sh, fm-ensure-agents-md.test.sh, fm-herdr-lab.test.sh, fm-send-popup-settle.test.sh, fm-send-settle.test.sh, fm-send-strict.test.sh, fm-spawn-batch.test.sh, fm-supervision-instructions.test.sh, fm-transition-lib.test.sh, fm-x-mode.test.sh

Behavior portable serial 1

fm-busy-state.test.sh, fm-calm-claude-mod-plugin.test.sh [gate-only], fm-check-unregister.test.sh, fm-gemini-harness.test.sh, fm-gotmp.test.sh, fm-grok-stop-live-e2e.test.sh [gate-only], fm-guard-stale-banner.test.sh, fm-lock-fast.test.sh, fm-opencode-primary-live-e2e.test.sh [gate-only], fm-pr-state-live-e2e.test.sh [gate-only], fm-send-agy-confirm.test.sh, fm-send-secondmate-marker-herdr-e2e.test.sh [gate-only], fm-shared-captain-inheritance.test.sh, fm-wake-daemon-lifecycle-e2e.test.sh, fm-watch-triage.test.sh

Behavior portable serial 2

fm-afk-contract.test.sh, fm-afk-pi-herdr-return-e2e.test.sh [gate-only], fm-backend-zellij.test.sh, fm-bootstrap.test.sh, fm-claude-session-lock-live-e2e.test.sh [gate-only], fm-codex-continuity-live-e2e.test.sh [gate-only], fm-composer-codex-idle-live-e2e.test.sh [gate-only], fm-cursor-harness.test.sh, fm-gate-refuse.test.sh, fm-harness-liveness-drift-live-e2e.test.sh [gate-only], fm-herdr-pi-stale-registration-live-e2e.test.sh [gate-only], fm-herdr-unregistered-agent.test.sh, fm-inactive-reconcile.test.sh, fm-launch-prompt-signals-live-e2e.test.sh [gate-only], fm-live-gate.test.sh, fm-mail.test.sh, fm-muse-harness.test.sh, fm-pi-windows-shell-invocation.test.sh, fm-procevent-quota.test.sh, fm-remote-entrypoint.test.sh, fm-remote-secondmate-lifecycle-e2e.test.sh, fm-secondmate-reconcile.test.sh, fm-spawn-pool-base-freshen.test.sh, fm-voice-relay.test.sh, fm-worker-account.test.sh

Behavior portable serial 3

fm-backend.test.sh, fm-calm-claude-mod-live-e2e.test.sh [gate-only], fm-contributions.test.sh, fm-control.test.sh, fm-devin-signals-live-e2e.test.sh [gate-only], fm-fleet-snapshot-view.test.sh, fm-grok-continuity-live-e2e.test.sh [gate-only], fm-harness-adapter-instructions-live-e2e.test.sh [gate-only], fm-lint-workflows.test.sh, fm-muse-signals-live-e2e.test.sh [gate-only], fm-peek-remote.test.sh, fm-pi-branch-extension.test.sh, fm-pr-check-security.test.sh, fm-procevent-stop-proof.test.sh, fm-remote-herdr-guard.test.sh, fm-rovo-harness.test.sh, fm-secondmate-sync.test.sh, fm-send-remote-delivery.test.sh, fm-spawn-compact-adviser-disable.test.sh, fm-supervision-host-live-e2e.test.sh [gate-only], fm-test-fixture-cleanup.test.sh, fm-test-isolation-proof.test.sh, fm-trace-context-spawn.test.sh, fm-wake-drain-open-decisions.test.sh, fm-watcher-lock.test.sh

Behavior portable serial 4

fm-backend-orca.test.sh, fm-claude-stop-autoarm-live-e2e.test.sh [gate-only], fm-cursor-primary-live-e2e.test.sh [gate-only], fm-daemon.test.sh, fm-devin-harness.test.sh, fm-extension-binding.test.sh, fm-harness-precedence.test.sh, fm-home-summary-refresh.test.sh, fm-home-summary-request.test.sh, fm-omp-harness.test.sh, fm-operational-input.test.sh, fm-pi-primary-live-e2e.test.sh [gate-only], fm-pi-watch-extension.test.sh, fm-platform-process.test.sh, fm-procevent.test.sh, fm-quota-array-dispatch-live-e2e.test.sh [gate-only], fm-remote-reply.test.sh, fm-secondmate-liveness.test.sh, fm-send-secondmate-marker.test.sh, fm-session-lock-ancestry.test.sh, fm-spawn-compact-adviser-disable-remote.test.sh, fm-startup-network.test.sh, fm-subagent-pretool-check.test.sh, fm-update.test.sh, fm-wake-drain-open-decisions-cursor.test.sh

Behavior portable serial 5

fm-agy-signals-live-e2e.test.sh [gate-only], fm-ask-user-authority.test.sh, fm-backlog-atomicity.test.sh, fm-backlog-read-bound.test.sh, fm-bearings-board-render.test.sh, fm-branch-supervision.test.sh, fm-calm-pi-extension.test.sh, fm-codex-hook-layer-live-e2e.test.sh [gate-only], fm-copilot-harness.test.sh, fm-copilot-management-live-e2e.test.sh [gate-only], fm-dod-lib.test.sh, fm-launch-status.test.sh, fm-on.test.sh, fm-remote-job.test.sh, fm-remote-transport-lanes.test.sh, fm-secondmate-safety.test.sh, fm-send-inbox-doorbell-live-e2e.test.sh [gate-only], fm-send-resolve-key.test.sh, fm-supervision-events.test.sh, fm-test-catalog.test.sh, fm-tmux-agent-liveness.test.sh, fm-turnend-guard.test.sh, fm-wake-drain-outcome-backstop.test.sh, fm-watch-checkpoint.test.sh

Behavior portable serial 6

fm-afk-inject-e2e.test.sh, fm-bearings-board.test.sh, fm-calm-claude-mod.test.sh, fm-control-recovery.test.sh, fm-control-relaunch.test.sh, fm-copilot-hooks-live-e2e.test.sh [gate-only], fm-cursor-primary.test.sh, fm-dispatch-resolve.test.sh, fm-forge-detect.test.sh, fm-gitignore-config.test.sh, fm-harness-adapter-references.test.sh, fm-herdr-windows-liveness-live-e2e.test.sh [gate-only], fm-pending-reply.test.sh, fm-pi-codex-native.test.sh [gate-only], fm-pr-state.test.sh, fm-remote-doctor.test.sh, fm-remote-job-orphan-reap.test.sh, fm-remote-job-wait.test.sh, fm-session-start.test.sh, fm-spawn-worktree-settle.test.sh, fm-stat-shadowing.test.sh, fm-supervision-host.test.sh, fm-task-delivery.test.sh, fm-tool-update-check.test.sh, fm-update-windows.test.sh, fm-watch-arm.test.sh

Behavior portable serial 7

fm-afk-return.test.sh, fm-agy-harness.test.sh, fm-backend-cmux.test.sh, fm-backend-herdr-windows-treehouse-live-e2e.test.sh, fm-backend-zellij-smoke.test.sh, fm-bearings-snapshot.test.sh, fm-cmux-claude-composer-live-e2e.test.sh [gate-only], fm-fleet-ledger.test.sh, fm-fleet-sync.test.sh, fm-herdr-session-cleanup.test.sh, fm-kimi-harness.test.sh, fm-omp-primary-live-e2e.test.sh [gate-only], fm-pi-branch-live-e2e.test.sh [gate-only], fm-pi-branch-responsiveness-live-e2e.test.sh [gate-only], fm-pr-reviewers.test.sh, fm-private-path.test.sh, fm-quota-choose.test.sh, fm-send-inbox.test.sh, fm-sessionstart-hook-live-e2e.test.sh [gate-only], fm-spawn-dispatch-profile.test.sh, fm-spawn-queue.test.sh, fm-startup-performance.test.sh, fm-timeout-lib.test.sh, fm-turnend-foreign-owner-arm-fix.test.sh, fm-wake-queue.test.sh, fm-watch-recovery-loop.test.sh

Behavior portable serial 8

fm-backlog-handoff.test.sh, fm-bearings-board-lavish-live-e2e.test.sh [gate-only], fm-bootstrap-network-parallel.test.sh, fm-ci-workflow.test.sh, fm-classify-decision-key.test.sh, fm-claude-stop-autoarm.test.sh, fm-claude-trust.test.sh, fm-composer-matrix-live-e2e.test.sh [gate-only], fm-copilot-primary-live-e2e.test.sh [gate-only], fm-documentation-audiences.test.sh, fm-path.test.sh, fm-pr-local-cost.test.sh, fm-procevent-when.test.sh, fm-project-origin.test.sh, fm-public-followup.test.sh, fm-secondmate-harness.test.sh, fm-sessionstart-instruction-refresh-live-e2e.test.sh [gate-only], fm-sessionstart-nudge.test.sh, fm-tangle-guard.test.sh, fm-task-inbox.test.sh, fm-tasks-axi.test.sh, fm-teardown-endpoint-safety.test.sh, fm-vendor-auth-probe.test.sh, fm-wake-drain-unread-status.test.sh, herdr-workspace-move.test.sh

Behavior portable serial 9

fm-backend-cmux-smoke.test.sh, fm-backend-herdr-treehouse.test.sh, fm-backend-tmux-smoke.test.sh, fm-busy-adapter-wiring.test.sh, fm-calm-pi-queue-retention-live-e2e.test.sh [gate-only], fm-classify-corr-token.test.sh, fm-harness-contract.test.sh, fm-herdr-submit-confirm-live-e2e.test.sh [gate-only], fm-herdr-version-floor-live-e2e.test.sh [gate-only], fm-inbox.test.sh, fm-lint-inventory.test.sh, fm-mail-check.test.sh, fm-nm-test-contract.test.sh, fm-reconcile-validation.test.sh, fm-remote-backlog-handoff.test.sh, fm-remote-secondmate-parent-binding.test.sh, fm-remote-secondmate-trace-context.test.sh, fm-rovo-signals-live-e2e.test.sh [gate-only], fm-secondmate-lifecycle-e2e.test.sh, fm-secondmate-restart.test.sh, fm-startup-memory-budget.test.sh, fm-stow-cascade.test.sh, fm-teardown.test.sh, fm-test-fixtures.test.sh [local timeout], fm-trace-context-lib.test.sh, fm-worker-account-live-e2e.test.sh [gate-only]

Behavior tests (Herdr)

fm-afk-inject-herdr-e2e.test.sh, fm-afk-launch.test.sh, fm-backend-autodetect-smoke.test.sh, fm-backend-herdr-agent-exit-shell-e2e.test.sh, fm-backend-herdr-eventwait-smoke.test.sh, fm-backend-herdr-focus-flash-e2e.test.sh, fm-backend-herdr-launcher-workspace-e2e.test.sh, fm-backend-herdr-presentation-e2e.test.sh, fm-backend-herdr-prune-safety-e2e.test.sh, fm-backend-herdr-respawn-idem-e2e.test.sh, fm-backend-herdr-smoke.test.sh, fm-backend-herdr-stale-active-tab-e2e.test.sh, fm-backend-herdr-workspace-per-home-e2e.test.sh, fm-control-herdr-smoke.test.sh, fm-herdr-attached-viewer-live-e2e.test.sh, fm-herdr-session-cleanup-e2e.test.sh

No-renames divergence and locality audit

All comparisons use frozen upstream b805823a6beef7f59b2a10fe50ae6054c39cf571, actual base b13de801f35b53757a8c61ea18d8c577b07a7aeb, and head 85c1d779065e7447727aecf5af2c15a1870444e5. Commands:

git diff --no-renames --name-status <frozen-upstream> <actual-base>
git diff --no-renames --name-status <frozen-upstream> <final-head>
git diff --no-renames --numstat <actual-base> <final-head>
git diff --no-renames --unified=0 <frozen-upstream> <final-head> -- <retained-caller>
Comparison Paths Modified upstream paths Additions Deletions Inserted/deleted lines
Upstream → frozen base 399 294 73 32 +29765/-33304
Upstream → final head 276 201 73 2 +27470/-5802
Actual base → final head 246 216 30 0 +27980/-2773

124 paths return to exact frozen-upstream equivalence. These are actual identical paths/blobs, not renamed implementation credited as removed divergence. No new implementation relocation is introduced. The retained fork additions and compatibility imports remain explicit; full path inventories and binary/full branch diffs are retained in session evidence.

Retained integration Implementation owner and reason Removal condition
Copilot/Pi bin/fm-harness-lib.sh; bin/harnesses/copilot.sh; bin/harnesses/pi.sh. Detection, spawn, bootstrap, control, busy, teardown retain lifecycle boundaries. New interrupt facts delegate to the adapters. Equivalent upstream pilot registry and lifecycle-compatible capabilities.
Native process and transport bin/platform/process.mjs; bin/platform/windows-process.ps1; bin/fm-platform-process-lib.sh. OpenCode host and ordinary arm use distinct exact-child tokens; existing public wrappers and Windows transport remain. Equivalent upstream native identity, owned retirement, and import contracts.
Native private paths bin/fm-private-path-lib.sh; bin/platform/windows-private-path.ps1. PR, X, worker, Herdr callers retain transaction policy. Devin secures the empty temporary file before copying user settings. Equivalent upstream native privacy without weaker publication or rollback.
Startup and Windows cost bin/fm-backend.sh; bin/fm-classify-lib.sh; bin/fm-wake-lib.sh; bin/fm-home-summary-refresh.sh; bin/fm-pr-lib.sh. In-process metadata and decision readers, lock-free harvest, batch ACLs, request-driven summary publication, and native event transport remain. Matched output, freshness, safety, and equal-or-better Windows cost evidence.
Forge bin/fm-pr-poll.sh; bin/fm-pr-lib.sh. Azure identity-bound observation and private publication compose with Gerrit; neither gains merge authority. Equivalent upstream identity, private registration, notification, and landed-work proofs.
Test metadata tests/catalog/core.tsv; tests/catalog/fork.tsv; bin/fm-test-catalog-lib.sh. New upstream registrations/routes/hints go in core; runner algorithms, fork overrides, and independent proof admissions are retained. Equivalent upstream metadata, selection, and proof-owned admission.
Legacy/integrity bin/fm-spawn.sh; bin/fm-remote-doctor.sh. Nonpilots and pi-signed remain legacy; remote doctor stays hash-pinned and unchanged. Separately scoped compatibility migration or integrity-protocol change.

Complete before/after path inventories and the 124 exact upstream-equivalent paths are retained outside the repository in divergence-before.name-status.tsv, divergence-after.name-status.tsv, and divergence-audit.json. No source evidence was discarded to shorten this PR body.

Expected exact-head automatic checks

The resolved workflows expand to 25 unique checks, with one intended producer per name. Live check status is reported separately from PR creation; pending, failed, absent, cancelled, or skipped required lanes do not establish merge readiness.

Check Workflow/job producer
Lint 1 .github/workflows/ci.yml / lint
Lint 2 .github/workflows/ci.yml / lint
Test coverage guard .github/workflows/ci.yml / test-coverage
Behavior portable parallel 1 .github/workflows/ci.yml / tests-portable-parallel-1
Behavior portable parallel 2 .github/workflows/ci.yml / tests-portable-parallel-2
Behavior portable serial 1 .github/workflows/ci.yml / tests-portable-serial
Behavior portable serial 2 .github/workflows/ci.yml / tests-portable-serial
Behavior portable serial 3 .github/workflows/ci.yml / tests-portable-serial
Behavior portable serial 4 .github/workflows/ci.yml / tests-portable-serial
Behavior portable serial 5 .github/workflows/ci.yml / tests-portable-serial
Behavior portable serial 6 .github/workflows/ci.yml / tests-portable-serial
Behavior portable serial 7 .github/workflows/ci.yml / tests-portable-serial
Behavior portable serial 8 .github/workflows/ci.yml / tests-portable-serial
Behavior portable serial 9 .github/workflows/ci.yml / tests-portable-serial
Behavior tests (Herdr) .github/workflows/ci.yml / tests-herdr
Behavior timing aggregate .github/workflows/ci.yml / tests-timing-aggregate
Stock macOS Bash snapshot compatibility .github/workflows/ci.yml / macos-stock-bash
Repo invariants .github/workflows/ci.yml / invariants
Windows self-update entry point .github/workflows/fork-ci.yml / windows-update
Windows reconciliation (core) .github/workflows/fork-ci.yml / reconciliation-windows
Windows reconciliation (copilot-launch) .github/workflows/fork-ci.yml / reconciliation-windows
Windows reconciliation (legacy-rollback) .github/workflows/fork-ci.yml / reconciliation-windows
Windows reconciliation (pr-completion) .github/workflows/fork-ci.yml / reconciliation-windows
Windows Copilot management .github/workflows/fork-ci.yml / windows-management
Harness package compatibility .github/workflows/fork-ci.yml / harness-package-compatibility

The timing aggregate retains dependencies on both parallel jobs, every serial shard, and Herdr; lane timing artifacts and always-run cleanup remain unchanged. The final head must clear these checks and the unresolved local observations before this PR is considered validated.

kunchenguid and others added 30 commits September 22, 2026 12:02
* fix: close landed workers from supervision in both postures and at return

During the 2026-09-22 away window every exemption worker whose pull request
had merged was left sitting for nine hours. The supervision branch received
the stale wake, the merge-landed check, and the hourly inactive-outcome row
for each of them, ran the recovery playbook, found nothing to recover, and
reported "no further action". The branch prompt granted ordinary teardown of
a confirmed-landed task without ever naming the moment or the command, and
the playbook has no landed exit, so the stale path ended at "nothing to
recover". The return brief then listed only blockers, decisions, and the
latest five routine outcomes, so the landed workers stayed invisible after
the captain came back.

- bin/fm-branch-prompt.sh: name the merge-landed wake, and any later stale,
  inactive-outcome, or heartbeat row on a done task with a merged PR, as the
  moment to claim the lease and run bin/fm-teardown.sh with no flags; a
  refusal is reported, never forced or worked around. Add teardown to the
  handling tool list.
- stuck-crewmate-recovery: a landed worker is not a recovery case; point at
  the ordinary teardown owner for each actor.
- bin/fm-afk-return.sh: render a "Landed, cleanup due" section from durable
  records only (a live task record whose recorded PR carries the
  merge-notification marker), between could-not-fix and handled, without
  holding the gate; the afk skill's return step closes each listed task
  through ordinary teardown once the check clears.
- tests: pin the prompt rule in fm-branch-supervision and the brief section
  in fm-afk-return through the real marker writer.

* no-mistakes(document): Document landed-task cleanup ownership
* fix(bin): surface a green no-mistakes PR still in ci merge monitoring

A green PR could sit unreported because neither the worker nor the
supervisor could observe checks-green while the ci step kept monitoring
for the merge.

Supervisor read: fm_nm_select_run's capped-overview inventory reader looked
the repository up by the task worktree path, but no-mistakes registers a
repository once by its main clone path and resolves every linked worktree
to it, so on every task copy of a busy repo the lookup matched no row and
each read reported "complete same-branch run inventory unreadable". Key the
lookup on the overview's own top-level `repo:` line, which every axi
release emits as the resolved working_path.

Even with a readable run, the ci-log classifier treated "base branch
advanced ..., re-arming CI monitor timeout" as not-ready. The monitor logs
a checks state only when it changes and a base advance does not clear
readiness, so a green PR read as still validating for as long as main kept
advancing. Stop treating that line as a marker, matching no-mistakes' own
ci-log parser, and name the run's PR URL in the held-for-merge reading so
the existing inactive-outcome path can act on it without a worker report.

Worker contract: `axi status` never reports checks-passed while the ci
step monitors for merge, so the definition of done no longer makes a
status poll the wait for the next gate or outcome; the drive call's own
return is the green signal, reattached with `no-mistakes axi run` after a
bounded return.

* no-mistakes(review): read the full ci log when checking checks-green

* no-mistakes(review): correct stale ci log tail wording in docs

* no-mistakes(document): Document checks-green supervisor fallback
* fix: derive Lavish polling server from its board session

* no-mistakes(document): Document session-derived Lavish polling

* no-mistakes(document): Correct Lavish routing verification claims
… vanish (kunchenguid#4900)

* fix(bin): ignore vanished state scratch files on secondmate relaunch

Relaunch refused when find(1) exited non-zero while listing a secondmate
home's state directory. A live watcher can delete scratch files between
readdir and processing, which is not evidence that child *.meta records
are unreadable.

Prove the directory is listable from its mode and keep the existing
readable-meta loop as the child-record guarantee. Fixes kunchenguid#4765.

* no-mistakes(review): Skip chmod-000 unlistable-state relaunch test when running as root
…d#4907)

* fix(bin): treat home-owned status closes as already read

Self-announced bookkeeping appends now record their exact byte ranges.
Later drains and signal scans skip those ranges, so two distinct
--resolve-key answers after an OPEN DECISIONS fold do not each wake the
supervisor. Worker-authored lines outside that ledger still signal.

* no-mistakes(review): Keep owned closes in unread status; lock ledger writes

* no-mistakes(review): Drop fold-lag wake suppression so folded worker decisions still wake

* no-mistakes(review): Require real owned growth before ledger marks status seen

* no-mistakes(document): Clarify home-appends ledger scope versus UNREAD STATUS

* no-mistakes(review): Restore fold-lag path, drop owned-range filters, fix test

* no-mistakes(review): Align ledger docs and scope ledger to wake path only

* no-mistakes(review): Restore stranded historical-annotation test comment to its function

* no-mistakes(review): Retire the home-appends lock alongside its ledger

* no-mistakes(document): Note ledger's lock-helper dependency in classify library

* no-mistakes(review): Append-and-coalesce home-appends ledger; fix stamped-line assertions

* no-mistakes(review): Drop redundant empty-span branch; make owned test pin ledger

* no-mistakes(document): Document covers' ascending-order dependency on home-appends ledger

* no-mistakes(document): Note owned-append skip in watcher signal-scan comment
…nguid#5350)

* chore(bin): raise tasks-axi, quota-axi, and lavish-axi floors to latest

Raise the minimum versions to tasks-axi 0.2.6, quota-axi 0.1.50, and
lavish-axi 0.1.77, pin CI's tasks-axi install to 0.2.6, and move the
floor-boundary test fixtures to the new versions.

tasks-axi 0.2.6 makes a failed relation deliverable for a promised-final
expecting pr-merged, so add the regression test: a bound work that ends
failed reports its honest outcome text through fm-public-followup-emit.sh,
consume marks the commitment ready, and deliver posts that text exactly
once.

Also make two hang-guard tests in fm-backlog-atomicity portable to hosts
without coreutils timeout, and stop an installed herdr from leaking into
the secondmate-liveness husk classifier test.

* no-mistakes(review): drop out-of-scope bounded_run hang-guard helper from atomicity test

* no-mistakes(review): pin quota-axi floor at 0.1.49 across fixtures

* no-mistakes(document): Document failed public-followup delivery behavior

* no-mistakes(ci): Updated quota-axi floor and all 0.1.49 fixtures to 0.1.51, corrected bootstrap boundaries to 0.1.51/0.1.52/0.1.50, and bumped the bearings lavish-axi stub to 0.1.77. Bearings, quota procevent, quota chooser, startup budget, and bootstrap floor coverage passed; the full bootstrap suite exceeded the 240-second local command limit after relevant checks passed. git diff --check passed
…rker copy (kunchenguid#4878)

* fix(bin): refuse ship done: when the named head lives only in the worker copy

A ship done: is not current-state done until that exact commit is reachable
outside the disposable copy. The check tests the named head, not whether
some branch moved.

* fix(bin): gate CI-ready ship done: on named-head reachability, not handoff

Keep no-mistakes' first done: as the pipeline handoff, apply the same shared
check when registering a PR and when a secondmate publishes ledger-first,
treat a recorded merged PR as landed after prune, and name the PR head
instead of scanning free-text SHAs.

* no-mistakes(review): Bind named-head gate to recorded PR and forge heads

* no-mistakes(review): Gate direct-PR forge heads and keep pending ledger deliveries

* no-mistakes(review): Align worker done wording, test mapping, pending-retry test

* no-mistakes(test): Raise watcher test time limit to stop load flake

* no-mistakes(document): Restore ledger-path fact and name named-head gate coverage

* ci: re-attest named-head ship-done gate for a fresh serial-3 verdict

* no-mistakes(review): Simplify local-only gate, gate keyed done lines, document recovery

* no-mistakes(document): Name fm-crew-state among named-head gate callers
…all alarm (kunchenguid#5204)

* fix(bin): ring a proven-idle secondmate before a wake-loop stall alarm

A leftover foreign-queue row on an idle, alive, ring-safe mate is still drainable in that home. Ring once, reset the observation interval, and keep the parent alarm for unknown, busy, or still-frozen rows.

* no-mistakes(review): Mark drain steer with from-firstmate fire-and-forget carrier
…unchenguid#5335)

The re-arm recovery cases judged "the watcher stayed live instead of
surfacing recovery" with fixed budgets below what a real stale-lock
recovery costs on a contended host: the arm's default 10s confirmation
deadline, a start helper that returned after about 4s whether or not the
arm had confirmed its watcher, and an 80-poll exit wait.
A changed-suite run beside other suites starves the recovery's many
short-lived processes while this suite's sleeping poll loops keep their
pace, so a watcher still surfacing its recovery read as one that stayed
live (issue kunchenguid#3793).
The original 0.25s window after confirmation was widened to 80 polls in
kunchenguid#3837, which left the same race at a larger size.

Following the CONTRIBUTING.md fixture-budget rule, the re-arm helper now
gives the arm an explicit 30s confirmation budget and waits for its
confirmation or exit within a ceiling that outlasts it, and every wait on
a re-armed watcher uses one named iteration-counted ceiling that outlasts
the same budget.
A passing case returns as soon as the arm reports or exits, and a watcher
that never surfaces its recovery still fails.

A new case delays every mktemp and readlink the re-armed watcher runs
after it publishes its beacon, so its first poll and exit take about 13s
on any host.
It fails with the reported symptom on the previous budgets and passes now.
No bin/ change.
* fix(bin): let one TERM always stop the watcher on bash 5.2

Bash 5.2 runs a pending trap from the parser entry of the next command
substitution it expands, where the trap body is parsed as the inside of
that substitution and fails ("trap: line 2: unexpected EOF while looking
for matching `)'") or is dropped silently, consuming the signal. The
watcher's `trap 'exit 1' HUP INT TERM` could therefore ignore a TERM and
keep polling while its stopper waited: the triage suite's reap waited
forever (CI jobs cancelled at 30 minutes), and the arm's signal path and
the away-mode daemon's shutdown wait for the watcher the same way.
Bash 5.3 fixed the parser; 5.2 is the stock bash on Ubuntu 24.04.

HUP and TERM now keep bash's native fatal-signal handling, which runs the
EXIT trap (watcher_cleanup) and exits on bash 3.2, 5.2, and 5.3. INT keeps
its trap because bash ignores a direct SIGINT while a child runs. The
check-spawn deferral window no longer contains a command substitution.

The triage suite's reap is now bounded and fails the case within 10s with
process evidence instead of hanging the job, and a new regression test
proves TERM stops a watcher blocked inside a poll's pane capture and still
releases its lock and records an acknowledgeable stop.

* no-mistakes(document): Clarify watcher stop-signal documentation
…id#5374)

* fix(bin): submit our own stuck doorbell instead of skipping every later ring

* no-mistakes(review): Confirm and retry Enter once on stuck-doorbell submit

* no-mistakes(document): Clarify doorbell retry and pending-composer documentation
* feat(bin): add the opt-in fleet activity ledger

Homes that create config/fleet-ledger get an append-only JSONL file,
state/fleet-ledger.jsonl, recording task.dispatched, task.status,
task.merged, and task.cleaned_up so outside tools can follow a fleet.
With the flag absent each producer does one file test and nothing else.
docs/fleet-ledger.md owns the record contract and its documented limits.

* no-mistakes(review): Record task.status text verbatim after the first colon

* no-mistakes(document): Clarify fleet ledger status and setup documentation

* no-mistakes(ci): Fixed a timing race in tests/fm-pi-branch-extension.test.sh: the replacement-wake test now waits for the prompt to start before releasing it. The focused test passed twice, and git diff --check passed
…nchenguid#5352)

* fix(bin): format, validate, and surface public-followup deliverables

brief pre-fills report_path=data/<work-id>/report.md and states the accepted
format of every value it cannot know instead of a bare <value> placeholder.
fm-public-followup-emit.sh refuses a deliverable tasks-axi would refuse, in
both the direct and staged destinations, naming the key, value, and format.
consume records the specific deliverable, outcome, or missing key behind a
tasks-axi refusal, and each refusal wakes the owning home once through the
existing relay poll.

* no-mistakes(review): refuse emits missing a required deliverable in both destinations

* no-mistakes(review): require promised deliverables and keep rejections recoverable

* no-mistakes(review): mirror tasks-axi's canonical pull request URL rule

* no-mistakes(review): keep a rejection wake whose line cannot be read

* no-mistakes(review): key emit-time rules on the promise, not the outcome

* no-mistakes(review): bound deliverable keys and values as tasks-axi does

* no-mistakes(review): state rejection wakes as at-least-once and pin it

* no-mistakes(review): enforce the promised contract tasks-axi holds at emit

* no-mistakes(review): stop inferring a staged promise from its outcome

* no-mistakes(document): Refresh public follow-up documentation

* no-mistakes(ci): Fixed both CI flakes. Watcher cleanup is now installed before singleton acquisition, preventing timeout races from leaving stale locks while preserving recovery-failure evidence. Bearings render fixtures now publish a valid isolated Lavish session store and retire each listener after rendering, eliminating false unowned-source races. Verified with checkpoint stress, fm-watch-checkpoint, fm-watcher-lock, repeated fm-bearings-board-render runs, project lint, syntax checks, and git diff checks

* Revert unrelated CI auto-fix edits to the watcher and bearings board test

The CI step's automatic repair changed bin/fm-watch.sh and
tests/fm-bearings-board-render.test.sh to chase two intermittent CI
failures that also occur on main and are not part of this change. Restore
both files so this branch carries only the public-followup deliverable fix.

* no-mistakes(review): Refuse a repeated --deliverable key at emit argument parsing

* no-mistakes(document): Clarify public-followup validation and rejection-wake documentation
* Add verified Devin CLI worker adapter

* no-mistakes(review): Drop Devin resolver refusal and launch marker

* no-mistakes(review): Verify devin in bootstrap, fold kind rule, update docs

* no-mistakes(document): Document Devin sidecar, resume, and worker-only facts

* no-mistakes(document): Document Devin interrupt, liveness anchor, composer signals

* fix(control): never pair Devin interrupt presses on an idle agent

A fast double Escape on an idle Devin opens its /revert picker, where Enter
reverts file changes. fm-control now sends the second press only after the
first renders Devin's 'esc again to interrupt' armed hint, never sooner than
0.5 s, closes a revert picker a mistimed press opened with one Escape, and
refuses to type the exit command while that picker is open. An unarmed
interrupt reports cancel=not-running and leaves the busy record untouched.

* fix(devin): disable Claude hook import and commit attribution for workers

The per-task Devin config now forces read_config_from.claude=false, so a
worker no longer runs the user's or project's Claude Code hooks (including
Herdr's Claude agent-state hook), and attribution=false, so Devin adds no
Co-Authored-By trailer or Generated-with line to commits and PRs.

* test(devin): extend live guard and record Herdr and revert-picker evidence

The credentialed live guard now fails if an imported Claude Code hook runs,
if the worker's commit carries Devin attribution, if an idle interrupt sends
more than one press or opens the revert picker, or if an open picker lets
exit through or is closed with a revert. The Devin reference, agent-control
doc, and verification records carry the 2026-09-22 tmux and Herdr lab results,
including the Herdr exit refusal.

* no-mistakes(document): Correct Devin documentation links and lifecycle guidance

---------

Co-authored-by: Denis Beliaev <battler73@yandex.ru>
…id#5322)

fm-crew-state classifies the no-mistakes outcome 'passed-with-skips' as
unknown, so a finished worker awaiting merge is re-alerted as stale. The
same blind spot lets fm-teardown's pre-teardown terminal-run check refuse
a legitimate abort race that lands on this outcome.

Map passed-with-skips to done in crew-state resolution, keeping the
skipped publication/CI verification visible in the detail rather than
reporting a clean pass, and recognize it as terminal during teardown.
…nguid#5382)

* fix: refuse missing backend adapter before source

* no-mistakes(review): Gate backend precheck under stock Bash

* no-mistakes(document): Clarify adapter precheck docs

* no-mistakes(lint): Suppress intentional child Bash ShellCheck warning
…kunchenguid#5338)

* fix(test): repair tmux liveness and calm follow-up loaded_off regressions

Both self-tests fail on untouched main on a host whose coreutils are a
multicall binary and whose Chrome has no pre-warmed profile, and each failure
masks the other's file.

tests/fm-tmux-agent-liveness.test.sh - the stand-in harness processes were
symlinks to the host's `sleep`. A single-purpose `sleep` runs happily under
another name, but a multicall coreutils binary (uutils or busybox) resolves its
applet from argv[0]: `claude-link -> sleep` invoked under the harness name runs
the wrong applet and exits immediately, so no foreground process exists and
every positive case reads not-alive ("last verdict for liveness:agent was
missing (expected alive); title=sh comms=[sh ]"). Build a dedicated spinner as
the stand-in target, exactly the way the version-string case already builds its
executable, and require the fallback target to demonstrably survive the rename
before using it. Every assertion is untouched; the stand-in identity signal is
unchanged (the kernel still records the symlink name as the executable
identity).

tests/fm-calm-pi-extension.test.sh - render_export_dom pinned a brand-new
`--user-data-dir` per attempt. On Google Chrome for Testing 151.0.7922.34 that
pristine profile makes Chrome's first-run initialization never complete: the
browser and its renderers start, but --dump-dom never returns, so all three
bounded attempts end exit=0 timed_out=yes bytes=0 and the DOM assertions never
run ("could not render calm-mode HTML export DOM"). Chrome's own profile
creation under a fresh HOME renders the same document in about a second, so the
helper now gives Chrome a private per-attempt HOME instead of the explicit
profile flag. Each attempt still gets an isolated profile, and every DOM
assertion is unchanged.

Root-cause evidence: a pristine --user-data-dir with `--headless=new
--dump-dom` had not returned after 150s, while the same command with an empty
HOME and no --user-data-dir returned the full DOM in ~1s, and reusing an
already-populated profile also returned it in ~1s. The render failure masked
the rest of the file: with it repaired, the Pi follow-up loaded_off case passes
unmodified against an installed @earendil-works/pi-coding-agent package.

These two failures block downstream validation of every lane on hosts with
multicall coreutils or a fresh Chrome profile.

Verification:
- timeout 300 bash tests/fm-tmux-agent-liveness.test.sh -> exit 0, 16 assertions ok
- timeout 700 bash tests/fm-calm-pi-extension.test.sh -> exit 0, 13 assertions ok,
  including the Pi operational follow-up loaded_off case
- bash -n and shellcheck clean on both touched files
- rest of tests/: bin/fm-test-run.sh --all bounded by timeout 900 completed 17 files with 0 failures (fm-afk-contract.test.sh through fm-backend-herdr-launcher-workspace-e2e.test.sh), then the bound cut off the 18th (fm-backend-herdr-presentation-e2e.test.sh, a real-herdr-gated lab test) with no failure recorded

* fix(test): give wake-queue observation checkpoints the alerting ceiling

tests/fm-wake-queue.test.sh's secondmate stall case runs bounded foreground
watcher checkpoints whose job is to record an observation, with the alerting
checkpoint that follows asserting the stall. A checkpoint's exit publishes a
downtime marker, and the next checkpoint consumes it only by reaching the end of
the watcher's poll loop, where the recovery surfacing runs after the stall tick;
the observation itself is recorded by that same stall tick. On a loaded host a
1s ceiling sits under the cost of that iteration (which includes a pane capture
in the active-turn gate), so the observation was never recorded, the downtime
marker stayed pending, and the alerting checkpoint surfaced
`check: rearm-resurface` instead of the stall it asserts:

  not ok - a foreign queue with no progress did not alert: check: rearm-resurface
  not ok - a frozen reprovisioned queue generation was hidden: check: rearm-resurface

Give the observation checkpoints that feed a later alert the same 4s ceiling the
file already documents for alerting checkpoints. The ceiling is only a bound - a
checkpoint still returns on its first actionable wake - so no assertion is
weakened, and the quiet windows get longer, not shorter.

* no-mistakes(document): docs: correct export-DOM Chrome render root cause

* no-mistakes(review): Isolate Chrome profile on macOS, dedupe tmux CC_BIN lookup

* chore: re-trigger fork workflow approval for triage

---------

Co-authored-by: Captain <blackxwhite88@users.noreply.github.com>
Co-authored-by: kunchenguid <kunchenguid@users.noreply.github.com>
…chenguid#5383)

* fix(bin): classify a status span without re-folding the whole log

A watcher poll could take minutes, so its liveness beacon aged past the
guard's 300s grace and the Stop auto-arm reported the watcher down. On the
main home, cycles ended with beacon_age 91-235s while healthy and 534-706s
while the laptop was CPU-starved.

Cause: whenever a newly appended status span held a keyed needs-decision
or blocked line, status_span_first_actionable_record re-read and re-folded
the ENTIRE log to decide whether that opening was still live, forking
several subshells per line. On a remote second mate's mirrored parent
channel (1.2MB, ~2300 lines) that is 13-20k subshells, about 17s per log
per classification when idle, paid by every signal and heartbeat scan.

Nothing regressed recently: subshell counts per classification were
20,272 from kunchenguid#3268 (2026-08-29, which introduced the whole-log fold) and
13,188 from kunchenguid#3753 onward through HEAD. The cost grew with log size, since
parent-channel logs only grow.

Fix: fold only the captured span. An accepted opening does not depend on
earlier lines and only later lines close or supersede it, and every later
line lies inside the span, so the span fold names the same live openings
at a cost bounded by the span. Old and new classification outputs are
byte-identical across 51 span offsets of real-shaped secondmate and ship
logs.

A real-watcher regression test records every read the classification
makes through the span-reader seam and asserts none reaches before the
classified offset; it fails on the old code (5,157 bytes read from
offset 0 to classify an 84-byte span).

* no-mistakes(document): Clarify span classification and watcher regression coverage
kunchenguid#5362 and kunchenguid#4878) (kunchenguid#5381)

* test: fix watcher timing flakes in fm-pr-check-security

The bounded watcher's hang guard now counts only the watcher's own time: a
case marks the intervals where it holds the watcher on injected work or makes
it wait on concurrent work, and those no longer count against its budget. The
budget itself stays at main's sixty seconds. The helper also stops forcing a
one-second per-check timeout, which killed a correct merged poll whenever that
poll took longer than a second, so the watcher only retried it or exited on a
later check's wake without the merge.

The concurrent-publication case pauses the guard while its arming is in
flight, and its task now sorts before the contributions observer the arming
also registers, so the watcher stops on the poll under test before running
that unrelated fleet snapshot. The case also prints the watcher's stderr when
it fails.

The replacement case pauses the guard while the re-arm runs inside the
watcher, runs that injected arming with the fixture root every other arming
here uses, and waits on the replacement merge's process instead of a
two-second cap. Merged-poll runs retire the contributions observer before the
watcher starts, since no case here exercises it.

The returned-descendant case no longer races a four-second sleep or a TERM
landing at an arbitrary point in the watcher's idle loop: its descendant holds
until killed, and a second check in the same cycle witnesses that it was
drained and stops the watcher.

* no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed

* Revert "no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed"

This reverts commit 6a59859.
…enguid#5385)

* feat: record task.pr_ready in the fleet ledger when a task PR is registered

* feat: record worker status lines in the fleet ledger as they are written

* no-mistakes(review): Keep worker status append failures and pass the resolved config to the ledger

* no-mistakes(review): Resolve relative config override before embedding in worker command

* no-mistakes(document): Clarify fleet ledger status capture timing
…unchenguid#5386)

* test: synchronize foreign secondmate stall legs on the watcher's recorded observation

Each leg of test_secondmate_foreign_queue_stall_tracks_progress_and_alerts_once
ran the watcher under a 1s or 4s wall-clock checkpoint, but every later leg
depends on the progress observation the previous leg's watcher recorded. Under
load the watcher was killed before its first stall tick, the observation was
never written, and the next leg treated its own sighting as the first one, so
the stall alert never fired.

Run the watcher directly and end each leg on its observable outcome: the
progress marker recording the expected observation, or the watcher's own first
wake. Also move a comment orphaned above this test back to the drain liveness
test it describes.

* no-mistakes(review): Wait for full stall reset before stopping watcher leg
kunchenguid#5391)

The listener resolves its server from that store before it polls. Without a session for this board, it exits in the gap after the build has already sampled a live claim.
…#5390)

* fix(bin): prune a torn-down task's wake rows at teardown

Prune pending durable wake rows (.wake-queue) for a task when it is torn
down, clearing stale wakes for its target window, signal wakes for its status
or turn-ended files, and task-specific check wakes.

Fixes kunchenguid#3419.
Adjacent to kunchenguid#5252.

- bin/fm-wake-lib.sh: add fm_wake_queue_prune_task
- bin/fm-teardown.sh: call fm_wake_queue_prune_task in cleanup_firstmate_home_children and main teardown
- tests/fm-wake-queue.test.sh: add test_wake_queue_prune_task

* no-mistakes(document): docs: note teardown prunes a task's wake rows

---------

Co-authored-by: Captain <blackxwhite88@users.noreply.github.com>
…ixture readiness (kunchenguid#5392)

* Make portable tests match resolved host paths

Summary:
- Match macOS full Node command paths by basename in the Gemini behavior test.
- Mirror symlink-resolved Nix PATH behavior and give the loaded-host race bounded headroom.

Testing:
- bin/fm-lint.sh
- bin/fm-test-run.sh tests/fm-on.test.sh tests/fm-gemini-harness.test.sh tests/fm-procevent.test.sh

Related:
- None

* no-mistakes(review): Mirror production PATH helper rules per directory group in tests

* no-mistakes(test): Wait for orphan runner start marker instead of fixed sleep

* no-mistakes(document): Clarify gemini ancestry test comment for versioned node comm

---------

Co-authored-by: Sandeep Salwan <salwansa@amazon.com>
…ns (kunchenguid#5389)

The sibling secondmate stall cases in tests/fm-wake-queue.test.sh now wait for the watcher's recorded observation instead of a one-second wall-clock checkpoint, so they can neither fail nor pass vacuously under load. Deterministic proof with a 5s watcher launch delay: before the fix 4 cases passed vacuously and 6 failed; after it all 10 pass on the recorded observation.

Also includes a CI flake fix from validation: fm_control_harness_supported in bin/fm-control-lib.sh finishes reading the harness allowlist before returning, removing intermittent broken-pipe diagnostics. Behavior is unchanged.
… tail (kunchenguid#5336)

* fix(bin): refuse a Herdr submit that would send only a message tail

A long typed payload can sit in the composer as a suffix, or as a paste placeholder plus a remainder, and the following Enter was still reported as delivered. Prove the selected composer holds the payload before Enter, and report failure when it does not.

* no-mistakes(review): Scope Herdr payload proof to Claude, clear composer on refusal

* no-mistakes(test): Clear refused Herdr composer drafts one wrapped row per press

* no-mistakes(test): Accept Claude's multi-line paste placeholder in Herdr submit proof

* no-mistakes(review): Accept Claude read-back that drops U+2063 in Herdr proof

* no-mistakes(document): Document Herdr proof ignoring U+2063 operational mark

* no-mistakes(ci): I made a one-line test change. The failing check comes from a timing race in an existing test that this PR doesn't touch. **What failed:** `tests/fm-procevent.test.sh` failed at "the superseded paced runner invoked its stale command" (line ~3313). The PR only changes the Herdr files and their tests, and the same shard passed on main at the base commit. **Why it can fail:** the fixture starts a second runner with a 3-second launch floor (the minimum wait since the source's last launch). That runner sleeps for the rest of the floor and only then checks whether its registration was replaced (`fm_procevent_launch_floor_wait` in `bin/fm-procevent-lib.sh`). The test then waits for the claim and re-registers the source. If that takes longer than about 3 seconds after the first launch, the old runner wakes up, finds its registration still current, and runs the stale command. That produces the second log line the test reports. The CI shard was slow (this one test took 160 s). **Fix:** in `tests/fm-procevent.test.sh` I raised the superseded runner's floor from 3 to 15 seconds and added a comment explaining why. The floor now outlasts the fixture setup even on a loaded runner. Nothing else changed: the first launch and the later fresh-registration start still use a 3-second floor, and no product code changed. **Verification:** - The full test file can't give a reliable result on this machine (load average about 64 on 8 cores). It failed earlier, at the reconcile assertion around line 1680, before it reached this section. - I ran the changed section by itself (file setup plus the pacing-race block) five times with the fix: all passed, in about 9-13 s each. - The original code also passed five out of five, so the race didn't reproduce locally. The diagnosis rests on the code path and the CI log. - I haven't seen the full file or the CI shard pass with the fix yet

* no-mistakes(test): Accept Claude folder-trust prompt via down+enter in live e2e

* no-mistakes(document): Note unreadable Claude composer refusal in Herdr docs
…unchenguid#5427)

Speaking as Kun's firstmate: squash-merging — opt-in (forge=gerrit registry-gated; default project-mode stdout restored to two words), attestation MATCH, CI+NM green, safe review, MERGEABLE.
…uid#5358)

* feat(bin): add an opt-in per-home worker account pin

A home that mixes work and personal accounts for one runner had no way to
say which account its workers launch on: Claude workers inherited whatever
CLAUDE_CONFIG_DIR the supervising process had, Pi workers the pane's ambient
root, and an ambient API key outranked both, with no signal at launch.

config/claude-account and config/pi-account now pin that choice per home.
With neither file every launch is unchanged. With one, every launch of that
runner from the home (ship, scout, local secondmate, raw Claude command, and
relaunch) runs under the declared root, and the spawn refuses before any
endpoint exists when the file is malformed or the runner's own check
(claude auth status, pi auth check with a model-listing fallback) says the
pinned account is not signed in. The check runs in a cleared environment so
an ambient credential cannot answer for an empty root. A pinned Claude launch
sheds the environment credentials Claude ranks above a stored login; a pinned
Pi launch needs an explicit <provider>/<id> model for a declared provider and
also carries --provider. The chosen account is printed on the spawned line and
recorded in the task record, and relaunch checks the pin before stopping the
running agent.

* test(secondmate): give the concurrent config-push wait room for a slow host

test_config_reread_serializes_concurrent_pushes waited about two seconds for
the first fm-config-push.sh to reach its first send-keys. On a slower host
that push takes four to five seconds, so the test failed on main before the
push ever got there. The loop still leaves as soon as the marker appears, so
the larger bound costs nothing where the push is fast.

* no-mistakes(review): Refuse raw Claude account overrides under a pin
…ts (kunchenguid#5470)

* fix(bin): keep the Herdr lab session option before a -- delimiter

fm-herdr-lab.sh run appended --session <lab> after every argument, so a
command with a passthrough delimiter such as agent start ... -- <agent args>
handed the session flag to the agent and Herdr routed the call by the
caller's ambient socket instead of the lab.
The helper now inserts --session <lab> immediately before the first --
delimiter and keeps the trailing form otherwise.

* no-mistakes(document): Clarify Herdr lab session option placement
* feat(bin): guard the partition, harness pin, and bounded exec for a non-Pi supervision host

Lease liveness is now the pure record test in every calling context, so an
unmarked main honors a live branch lease held by a separate process, and a
lease file engages the guard's claim serialization for any caller; a home
with no lease files still takes no lock.

bin/fm-harness.sh honors FM_SUPERVISION_PRIMARY_HARNESS while
FM_SUPERVISION_ACTOR=branch, so a supervision branch running under another
harness resolves own, crew, and secondmate to the primary's harness.

fm_tasks_axi's watchdog moves into bin/fm-timeout-lib.sh as fm_exec_timed with
a separate grace: the perl watchdog is preferred, runs the command in its own
process group against wall-clock deadlines, forwards TERM/INT/HUP, and reaps
the group, so a descendant holding captured output can no longer keep the
caller waiting past the bound on a host without timeout.

The Claude Stop auto-arm header records that Claude drops the exit 2 of a hook
it terminated at the configured timeout, re-measured on Claude Code 2.1.281.

* fix(bin): state that fm_exec_timed cannot reach a descendant in its own process group

Live runs of real Claude and Pi engine turns under the bound showed both CLIs
start every tool command in a process group of its own, so those processes end
through the engine's own TERM handling rather than the group signal or reap.
Also clears the new timeout test's ShellCheck findings.

* no-mistakes(document): Clarify cross-harness lease documentation
tiago-peixoto and others added 24 commits September 24, 2026 19:21
…unchenguid#5544)

* fix(bin): terminate a remote job worker that lost ownership when it receives TERM

A serving worker whose lock directory is gone can no longer quarantine shutdown, and resuming service publishes a false ready heartbeat. Exit after stopping only that worker's own command tree, without removing a replacement owner's lock.

* no-mistakes(review): Check worker lock ownership before publishing shutdown quarantine

* no-mistakes(document): Correct worker shutdown comment on replacement-owned lock

* fix(bin): keep an ousted remote job worker off the replacement quarantine

Shutdown can lose the lock after the first ownership check and before it
writes or clears quarantine. Bind both operations to the directory object
this process still owns so a replacement's quarantine stays untouched.

* no-mistakes(review): Make ousted-worker shutdown test reliably reach quarantine clear

* no-mistakes(document): Reattach worker_shutdown doc comment to its function

* no-mistakes(ci): Fixed the failing check (Behavior portable serial 7) with a test-only change to the stall test in tests/fm-remote-job.test.sh. Product code is unchanged; no other test changed. Cause: after the decoy dies, both workers run the same check-exists, read, delete sequence on the job records. On the CI runner the replacement deleted a record between the ousted worker's check and its read. The ousted worker exited 125, and because the file runs under set -e the unguarded `wait` ended the test with 125. The exit trap then killed the replacement, which produced the "Killed" line. Reproduction: a temporary 0.3 s delay between the check and the read, applied to the ousted worker only, made the committed test fail exactly as in CI (exit 125 and the "Killed" line). The new test passed with the same delay. The delay is reverted, along with a similar debug hook that the timed-out attempt had left in bin/fm-remote-job-worker.sh. Test changes: - The replacement is frozen (and confirmed stopped) before the decoy is killed and resumed only after the ousted worker exits, so only one worker touches the job records at a time. - The ousted worker is stopped only once its quarantine exists and its lane is reaped, which places it inside its stop loop. - Every fixed poll loop is now a wait on a named condition with a 30 s deadline and an explicit failure message. Exit detection also handles zombies. - The exit trap kills and waits for the decoy and both workers on every path. - A non-zero exit from the ousted worker now fails with its exit code and stderr instead of silently ending the file. The test still proves that the resumed ousted worker exits 0 and leaves the replacement's lock, quarantine contents and quarantine inode unchanged. Verification: the full test file passed four times on its own and three times under nice -n 10 with four busy-loop CPU hogs; bin/fm-lint.sh passes. Changes are not committed

* no-mistakes(ci): I fixed the failing check (Behavior portable serial 7) by changing only the stall test in tests/fm-remote-job.test.sh. Product code is unchanged. **What failed:** "an ousted worker in shutdown leaves the replacement quarantine untouched" failed on CI with the ousted worker exiting 125 ("could not stop the active command tree"). **Why:** during shutdown, the worker retries the still-running decoy command group a fixed 100 times, 0.01 s apart, then gives up and exits 125. The test tried to freeze the worker partway through those retries by sending SIGSTOP from outside. On a slow runner the retries ran out before the stop arrived, so the worker had already given up. The invariant is that the test must hold the ousted worker inside that retry loop until the replacement owns the lock. That was the only place the test depended on timing. The other waits already watch for a named state change with a 30 s deadline. **Fix:** - The ousted worker now starts with a small `sleep` wrapper at the front of its PATH, and the SIGSTOP race is gone. - The wrapper only holds a `sleep` called directly by that worker's own process (it checks its parent pid against a hold file) while its quarantine file exists. - The only such `sleep` is the first retry in the shutdown stop loop, so the worker waits there as long as needed. - The wrapper writes a marker when it starts holding. The test waits for that marker, then hands the lock to the replacement, freezes the replacement, and kills the decoy. - The test releases the worker by deleting the hold file. Deleting the whole temp directory also releases it, so a failed run cannot leave the wrapper looping. - A process leak: the test overwrites the job's command-group record with the decoy, so no worker ever stopped the job's real command. `fm-hold-job.sh` and its `sleep 30` stayed running for up to 30 s after the test. The test now records that group before overwriting it and kills it at the end of the test and in the exit cleanup. - The test still asserts the same things: the ousted worker exits 0, and the replacement's lock, quarantine contents and quarantine inode are unchanged. **Verification:** - The full file passed twice on its own, twice under `nice -n 10` with six busy-loop CPU hogs, and twice more after the leak fix. - `pgrep` found no leftover processes afterwards. - With the worker from just before the fix commit (cf45cb6^), the test still fails with "the ousted worker wrote or cleared the replacement quarantine during shutdown", so it still proves the fix. - `bin/fm-lint.sh` passes. - I did not reproduce the CI failure locally. The cause comes from the fixed retry limit and the CI error message. The changes are not committed

* no-mistakes(ci): I changed only the stall test ("an ousted worker in shutdown leaves the replacement quarantine untouched") in tests/fm-remote-job.test.sh. Product code is unchanged, and so is every other test. **Invariant:** the pid written to the job's group record must be a process-group leader whose group dies when that one process is killed. Otherwise the worker's bounded stop loop never sees the group die, gives up, and exits 125 ("could not stop the active command tree") before it reaches the lost-ownership exit. The decoy is the only place in this test that depends on this. **Fix:** - The decoy used to be `set -m; sleep 30 &`. It now starts as `perl -MPOSIX=setsid -e 'setsid() >= 0 or exit 1; exec @argv' sleep 30 &`, which gets its own session and group without shell job control. tests/fm-procevent.test.sh already uses the same idiom. - The test now waits, with the file's usual 30 s deadline and a named failure, until `ps -o pgid=` of the decoy equals its pid before writing it into the group record. This way the worker can never read the record before `setsid` has run. - The existing steps are unchanged: the test kills the decoy, reaps it with `wait` before releasing the hold file, and the exit trap still kills and reaps the decoy and both workers. - The assertions are unchanged: the ousted worker exits 0, and the replacement's lock pid, quarantine text and quarantine inode stay the same. **Cleanup:** I reverted a debug `printf` hook that the timed-out previous attempt had left in bin/fm-remote-job-worker.sh, and deleted its untracked `.tmp-repro/` directory. Neither was committed. **Verification:** - The full tests/fm-remote-job.test.sh passed twice normally and once under `setsid -w` with stdin from /dev/null (no controlling terminal). - `bin/fm-lint.sh` passes. - No leftover `sleep 30` processes afterwards. **Not reproduced:** I could not reproduce the CI failure locally. On this host `set -m` made the decoy its own group leader even without a controlling terminal, so the cause on the runner is not confirmed. The change removes the test's reliance on shell job control, as the user asked. Changes are not committed

* no-mistakes(ci): I changed only the stall test ("an ousted worker in shutdown leaves the replacement quarantine untouched") in tests/fm-remote-job.test.sh. Product code is unchanged, and so is every other test. **Invariant:** the group record the ousted worker checks in its stop loop must stay the job's own command group, and the test must stop that group before it releases the hold. Otherwise the bounded retry keeps seeing a live group, gives up, and exits 125 ("could not stop the active command tree") before it reaches the lost-ownership exit. The test overwrote this record in one place (the decoy) and stopped the group in one place (killing the decoy); both are changed. **Fix:** - I removed the setsid decoy and the overwrite of `.claim/group`. The record keeps the job's real command group, which the test still saves as `STALL_JOB_GROUP`. - The two-line `group_start` stays. It is still needed: without it the worker kills the real group on its first pass, before the replacement takes over, so the hold would never matter. - The `sleep` wrapper that holds the worker at its first stop-loop retry is unchanged. - After the replacement owns the lock, its quarantine is planted and it is frozen, the test runs `kill -KILL -- -$STALL_JOB_GROUP`. It then waits, with the file's usual 30 s deadline and a named failure, until `kill -0` on the group fails. Only then does it remove the hold file. The worker therefore always sees its own command already stopped and never races its retry budget. - The exit trap still kills the saved command group if the test fails. It can't `wait` on that group because the group is not a child of the test shell. The decoy variable and its cleanup entry are gone. - The assertions are unchanged: the ousted worker exits 0, and the replacement's lock pid, quarantine text and quarantine inode stay the same. **Verification:** - The full tests/fm-remote-job.test.sh passed twice normally. - It passed once under `setsid -w` with stdin from /dev/null (no controlling terminal). - It passed once under `nice -n 10` with six busy-loop CPU hogs. - With the worker from before the fix (cf45cb6^), the test still fails with "the ousted worker wrote or cleared the replacement quarantine during shutdown", so it still proves the fix. - No `fm-hold-job` or `sleep 30` processes were left afterwards. - `bin/fm-lint.sh` passes. **Not reproduced:** I couldn't reproduce the CI failure locally; the decoy version also passed on this host. So I can't confirm why the decoy group stayed alive on the runner. The new wait turns any leftover live group into a clear named failure instead of an exit 125. The changes are not committed

* fix(bin): keep a dead command group dead on bash 5.2

A bare return inside the liveness check drops the failing kill status when the check runs in a conditional, so shutdown keeps treating a stopped group as alive and exits 125.

* no-mistakes(review): Use bash 3.2 fd syntax and fix trap return comments
…henguid#5589)

* docs: make configuration settings easier to find and understand

* no-mistakes(review): Restore dropped qualifiers and fix misplaced config doc labels

* no-mistakes(review): Restore three dropped qualifiers in configuration reference
)

* fix: bound worker edits of project AGENTS.md/CLAUDE.md to factual corrections

These files are loaded into every agent session of a project, so additions
should be a deliberate human choice rather than automated task output. The
ship brief's project-memory section and AGENTS.md section 6 previously invited
workers to record durable knowledge, which let project AGENTS.md files accrete
detail the codebase or README already carries. Workers now edit only to fix
factually wrong content - including content their own change made wrong - and
fm-ensure-agents-md.sh runs only alongside such a correction. Stow no longer
routes project-memory additions through ship tasks, and the generated skeleton
no longer invites discovery-driven additions.

* no-mistakes(review): Stop running fm-ensure-agents-md.sh on memory-file corrections

* no-mistakes(document): Clarify manual project-memory initialization and remove duplicate guidance
…enguid#5635)

* fix(bin): let gate agents drive lifecycle against marked lab homes

Part 2 of the kunchenguid#5615 split. A no-mistakes gate agent runs inside a
checkout carrying the fleet-captain identity, so fm-gate-refuse-lib
refuses fleet mutation on the gate signal. That refusal was absolute,
which kept gate validation from ever exercising the real lifecycle.

Stamp a disposable lab FM_HOME with a .fm-lab-home marker file that
only bin/fm-lab-home.sh writes, and only onto a fresh empty dir, so
no call path can mark a populated real home. fm_refuse_if_gate_agent
then permits lifecycle only when FM_HOME carries the marker and is
driven through its stock layout - any FM_*_OVERRIDE relocation stays
refused so part of the "lab" cannot be split back onto the real fleet.
The threat model is a confused agent touching the real fleet, not
deliberate forgery, so the marker is a plain token file rather than a
bound record. FM_GATE_REFUSE_BYPASS is unchanged: it still serves the
test harness, which cannot mark hundreds of temp homes.

Teardown's slot-ownership scan compared state-dir paths textually
while fm_firstmate_root_home canonicalizes, so a lab home under a
symlinked TMPDIR scanned its own record twice and self-collided;
compare file identity (-ef) instead.

* no-mistakes(review): Refuse unlistable lab homes and hardlinked slot records

* no-mistakes(review): Mint lab markers only on verified-empty fresh dirs

* no-mistakes(document): Clarify lab-home gate documentation and comment contracts

* no-mistakes(document): Clarify lab-home gate documentation and remove stale claims

* no-mistakes(document): Clarify gate lab-home documentation and boundary wording
…ad of refusing every re-arm (kunchenguid#5594)

* fix(bin): replace a watcher whose beacon stalls past a hard bound instead of refusing every re-arm

A fleet watcher that is alive but whose liveness beacon has gone stale could
never be replaced: every re-arm was refused because the lock holder was a live
pid, and the holder was never evicted because it was not dead. Add
FM_WATCHER_STALL_BOUND (default 3x the stale grace): below it the refusal is
unchanged; at or past it the arm re-verifies the holder against the lock's
recorded identity, sends TERM, waits boundedly, and takes the lock the normal
way, ledgering a stalled-holder-replaced row. A holder that survives TERM keeps
the old refusal.

Fixes kunchenguid#4400

* no-mistakes(test): poll for replacement message to fix watcher-lock test flake

* no-mistakes(document): document FM_WATCHER_STALL_BOUND in config inventory
…n can keep them (kunchenguid#5563)

* fix(pi): hide queued Firstmate notifications under Calm only when the session can keep them

Calm now keeps authenticated Firstmate operational inputs out of Pi's queued-message
listing, but only after proving the live session exposes every member needed to keep
them across Escape. A session missing any of them keeps stock rows and Escape and shows
one generic warning. Escape and the dequeue key return only captain-authored messages to
the editor and re-queue hidden notifications in order; after an abort that kept any in
Pi's agent queue, the adapter starts the delivery turn itself because Pi 0.87.1 does not
continue an aborted run. Compaction-held notifications stay with Pi's compaction flush and
never start or announce a turn.

Fixes kunchenguid#1588

* docs(calm): record Pi 0.87.1 queued-row retention verification

* no-mistakes(review): Deliver kept Calm notifications after tree-navigation aborts too

* no-mistakes(review): Defer Calm notification turn until tree navigation finishes

* no-mistakes(lint): Silence SC2016 for literal JavaScript in queue-retention e2e test
…uid#5548)

* fix(bin): refuse teardown when a required source disappears

A missing sibling was sourced after cleanup had started, so Bash 3.2
exited 0 from the EXIT trap and Bash 5 continued and reported success.

* no-mistakes(review): Remove unused FM_TEST_ONLY hook from teardown tests

* no-mistakes(review): Check task backend sources before any teardown cleanup

* test(gotmp): give teardown fixtures every tmux adapter sibling

Teardown now refuses when a sibling the recorded backend's adapter sources
is missing, so the fake bin must carry fm-session-lock-lib.sh,
fm-agent-process-lib.sh and fm-gemini-lib.sh.
Restructure the supervision host doc's prose into shorter sections, lists,
and tables without changing documented behavior. Every original heading,
anchor, identifier, number, quoted string, and link target is preserved.
* docs: make herdr-backend easier to read

Restructure the Herdr backend doc's prose into shorter sections, lists, numbered procedures, and tables without changing documented behavior. Every original heading, anchor, fenced code block, link target, and documented fact is kept.

* no-mistakes(document): Restore composer-proof reason and complete Herdr topic table
Restructure the prose into sections, lists, and tables without changing
documented behavior. Every original heading and anchor, inline-code span,
link target, number, and quoted string is kept, and each sentence sits on
its own line. Adds a topic navigation table and short subsections under
the existing headings.
* docs: make watcher-continuity easier to read

Restructure the prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, identifier, link target, and number is kept.

* no-mistakes(review): Fix actor and supervision-host scope in watcher-continuity doc

* no-mistakes(review): Make readiness TERM and retry conditional on unready successor
* docs: make sessionstart-nudge easier to read

Restructure the prose into sections, lists, and tables without changing documented behavior. Every original heading, inline-code span, link target, number, and fact is preserved, and a harness-to-tier table now sits near the top.

* no-mistakes(review): Drop helm glossary line and dedupe exit-code lead-in
* docs: make captain-hold-lifecycle easier to read

Restructure the captain-hold lifecycle prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, identifier, number, quoted string, and link target is kept.

* no-mistakes(review): Fix verification record subjects and grouping headings

* no-mistakes(review): Clarify task-body read-back cases belong to the suite
* docs: make remote-secondmates easier to read

Restructure the remote second mates prose into sections, lists, numbered procedures, and tables without changing documented behavior. Every original heading, anchor, fenced code block, identifier, link target, and qualifier is preserved.

* no-mistakes(review): Merge remote-home table cell into one sentence

* no-mistakes(review): Tighten readiness lead-in, restore causal link, fix dangling reference
…nguid#5554)

* fix(bin): bound the away digest and log why a delivery failed

The away daemon joined every buffered escalation into one unbounded
digest. A start-up catch-all span can exceed what one transport argument
carries (tmux rejects the send-keys command; Linux refuses to exec any
argument above 131,071 bytes, which is how herdr receives it), so the
initial send failed on every housekeeping pass and was logged as an
unconfirmed Enter with text possibly in the composer.

escalate_flush now builds the injected digest under a fixed byte budget:
each event is cut at a UTF-8 boundary with an omitted-bytes marker, the
joined events stop with a "+K more event(s)" tail, and a bounded digest
names a state/.subsuper-digests/ file that keeps every buffered event
verbatim. The buffer itself is untouched, so the return catch-up stays
complete.

The tmux submit core and the herdr literal send now replay the
transport's stderr on failure, and inject_msg logs the failing stage
(initial send versus Enter confirmation) with the byte count and that
stderr. The wedge alarm line and marker carry the last failure reason.

Fixes kunchenguid#4382

* no-mistakes(review): Drop digest pruning; label send-failed as send-or-Enter stage

* no-mistakes(review): Keep digest full text once submit ran; reuse on retry

* no-mistakes(lint): Count digest files with find instead of ls

---------

Co-authored-by: firstmate-oss <firstmate@kunchenguid.local>
…henguid#5638)

* feat(tests): add FM_TEST_SEAM launch seam and gate lab-primary recipe

Part 1 of the kunchenguid#5615 split: the pieces that let the no-mistakes pipeline
live-validate firstmate changes, without the gate-refusal rescoping.

- bin/fm-afk-launch.sh: FM_TEST_HARNESS pins the detected harness only
  alongside the FM_TEST_SEAM=1 marker test suites set, so a leaked variable
  in a real primary's environment stays inert and unknown tokens fall
  through to real detection.
- tests/lib.sh: export FM_TEST_SEAM=1 for every suite.
- .no-mistakes.yaml: per-harness recipe for running a real fixture primary
  from a gate run - a plain mktemp lab FM_HOME on a private tmux socket,
  with FM_GATE_REFUSE_BYPASS=1 scoped to it and NO_MISTAKES_GATE scrubbed.
- tests/fm-wake-queue.test.sh: stop the owned watcher fixture with KILL and
  clear its lifecycle state so the next leg starts clean; TERM could leave
  bash waiting in a child on some runners.
- tests/fm-remote-secondmate-lifecycle-e2e.test.sh: wait for the liveness
  lock holder's post-acquire marker instead of the lock dir, which is
  published before the claim finishes.

* no-mistakes(review): Scrub lab home overrides and require FM_TEST_SEAM separately

* no-mistakes(document): Clarify test seam and disposable lab bypass documentation

* no-mistakes(document): Clarify lab isolation and test-seam documentation

* no-mistakes(ci): Fixed the CI failure: test cleanup killed the remote worker child but left its supervisor able to restart it during fixture removal. Cleanup now stops the worker tree. The lifecycle test passed locally; ShellCheck and diff checks passed
… home is gone (kunchenguid#5552)

* fix(bin): refuse watchers from disposable checkouts and exit when the home is gone

Fixes kunchenguid#321
Fixes kunchenguid#4760

A watcher armed from a disposable no-mistakes validation checkout under
.no-mistakes/worktrees/ outlived the validation step and kept writing the
real home's state, and a running watcher never noticed when its home,
state directory, or code root disappeared. The arm now refuses from such
a checkout with the typed failure line, the watcher checks once per poll
that its home, state directory (or its own lock holder record), and bin
directory still exist and exits with a logged reason scoped to itself,
and the shared test helpers reap every watcher a suite armed for a
temporary home through the home-scoped stop.

* no-mistakes(lint): fix SC1007 by assigning empty string in watch-arm test

* no-mistakes(ci): Found and fixed a genuine, reproducible hang introduced by this branch's test-watcher reaper, which is what killed both CI checks (serial-2 cancelled at the 30-min cap; Lint 2 exit 143 = the suite's own TERM-trap code). Root cause: test_drain_asserts_watcher_liveness (tests/fm-wake-queue.test.sh) fabricates a .watch.lock whose pid is the test runner's own $$ with the runner's real identity, to make the drain believe a live watcher exists. The new make_case tracking registers that state dir for reaping, so at fm_test_cleanup the new fm_test_reap_watchers drives fm-watch-arm.sh --stop; its identity check matches (the fixture recorded the runner's identity) and it kill -TERMs the test runner. tests/lib.sh:231 is `trap 'fm_test_cleanup; exit 143' TERM`, so the TERM re-enters cleanup -> reap -> kills $$ again -> infinite loop until the runner cap. I reproduced this locally: the suite ran all tests then looped forever in cleanup spawning fm-watch-arm.sh --stop against a lock naming its own PID. Fix (tests/lib.sh, +5 lines): in fm_test_reap_watchers, skip any tracked lock whose pid equals our own $$ before driving --stop. This is the single shared reap boundary; seven $$-self-lock fixtures across four test files are all covered by the one guard, and real armed watchers (pid != $$) are still reaped. Invariant: the test reaper must only signal real armed watcher processes, never the test runner itself. Verified locally: tests/fm-wake-queue.test.sh -> EXIT 0 (63 ok, no hang); tests/fm-watch-arm.test.sh -> EXIT 0 (21 ok, including test_reaper_stops_a_tracked_watcher, confirming the guard does not over-skip). Lint 2's exit 143 was the same shard/cap signature; a fresh CI run on this new commit will re-evaluate it

---------

Co-authored-by: firstmate-oss <firstmate@kunchenguid.local>
…m as silence (kunchenguid#5588)

* fix(bin): surface an unrecognized status prefix instead of dropping it

A parked or holding declaration, and a verb whose correlation token did not parse, never became an event, so the supervisor still saw the earlier line.

* no-mistakes(review): Require verb-shaped unrecognized status prefixes, add continuation tests

* no-mistakes(document): Document unrecognized status prefix escalation in afk skill

* no-mistakes(ci): I reproduced the "Behavior portable serial 6" failure locally and fixed it by changing the test data in one test. No product code changed. **What failed:** `tests/fm-session-start.test.sh`, in `test_orphan_status_logs_are_printed`, with "matched status log was printed 2 times". **Why:** the test writes status lines with made-up prefixes, `matched: surfaced once` and `orphan: step N`. The test only uses them as placeholder text. It checks that the session-start digest prints each task's status tail exactly once. This PR (kunchenguid#4763) deliberately makes an unrecognized one-word lowercase prefix a status event. So those lines now surface as captain-relevant events, and the wake queue's STATUS OUTCOME BACKSTOP section prints them a second time. The code under review is behaving as the issue asks. Only the test's placeholder data had become meaningful. **Rule the test depends on:** its status lines must not be captain-relevant, so the digest is the only place they are printed. Both lines in this test broke that rule. The orphan line would have failed the same count check right after the matched line did. **Fix:** in that test only, I switched both lines to the recognized, non-captain verb `working:`: `working: surfaced once` and `working: orphan step 1..6`. I updated the matching assertions and counts to use the new text. What the test checks is unchanged: orphan logs are labelled, the tail is bounded, the log path is printed, and each tail appears once. **Verification:** before the fix, the test failed locally the same way as in CI. After it, `bash tests/fm-session-start.test.sh` reports "all assertions passed

* no-mistakes(review): Detect unrecognized prefixes on unstamped lines; share verb list

---------

Co-authored-by: Kun's firstmate <kunchenguid+firstmate@users.noreply.github.com>
…henguid#5658)

Fixes kunchenguid#5295

Session start now reports a remote inheritance failure using the
push's own error line instead of the first unchanged item that
happened to print before it, and the shared captain preferences
header check now names the first required phrase it did not find,
on both the local and remote inheritance paths.
…yloads (kunchenguid#5657)

* fix(bin): stand down the Claude Stop auto-arm on pi-code-delivered payloads

pi-code loads the tracked Claude settings but has no asyncRewake, so it
awaits every Stop hook; without a stand-down the auto-arm runs
synchronously inside Pi's turn end and holds it open for the declared
multi-hour timeout. Stand down when the payload's transcript_path
contains a /.pi/ path component, the same discriminator the closed-but-
unmerged fix in kunchenguid#3352 used, with an explicit string-type check on the
jq filter.

Fixes kunchenguid#3343

* no-mistakes(document): document pi-code stand-down in harness integrations reference
…kunchenguid#5659)

* fix(bin): match whole multi-word project names in the registry lookup

bin/fm-project-mode.sh matched a registered project name against only the
first whitespace-delimited token of a registry row, so a name containing a
space never matched, silently defaulting the project to no-mistakes off
instead of its declared posture.

The lookup now matches the whole registered name against the raw line text,
so a name is compared literally (never as a regex) and a name that is a
leading prefix of another registered name still resolves to its own row.

* no-mistakes(document): docs already accurate for multiword registry name match

* chore: drop accidental empty err file

Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
…d#5546)

* fix(bin): classify the stdin program of `bash -s` with operands in the arm policy

With -s, sh/bash/zsh read the program from stdin even when operands follow;
the operands are only positional parameters. The arm policy treated the first
operand as a script path, so heredoc and here-string payloads were never
classified and a hidden bin/fm-watch.sh execution was allowed.

A protected path in the operand position still fails closed as before.

Fixes kunchenguid#1489

* no-mistakes(document): Clarify stdin shell operand documentation

* no-mistakes(ci): Captain, fixed `shellInvocation` so `bash -- -s` treats `-s` as a script name, and updated R21 to test the exact command. The targeted policy suite, lint, documentation check, and diff check pass. Both hosted workflows show `action_required` before any jobs ran; that external approval state remains unresolved

* fix(bin): keep main's handling of words after a leading `--`

Revert the pipeline CI-step change that made the first word after a leading
`--` always a script. It turned forms that main denies today into allow
(for example `bash -- -c 'bin/fm-watch.sh'`), which is outside kunchenguid#1489 and
loosens a fail-closed policy. `--` after `-s` still ends option parsing.
Compose Devin, supervision-host, worker-account, branch, Gerrit, liveness,
and ledger changes with the extracted pilot, native Windows, Azure,
startup-cost, private-path, and test-catalog boundaries.

Firstmate-Upstream-SHA: b805823
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Use tasks-axi 0.2.6 in the portable and Windows backlog-fixture jobs.
Keep the installed versions at or above the runtime owner through the
existing workflow contract suite.

Firstmate-Upstream-SHA: b805823
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@timbarreto

Copy link
Copy Markdown
Owner Author

Exact-head CI snapshot (2026-09-25T17:32:34.780Z)

Head: b9fa47b2799c0a90a14ad64a53d134401c09ea34; base: b13de801f35b53757a8c61ea18d8c577b07a7aeb. The PR is open, non-draft, unmerged, and GitHub reports it mergeable. Auto-merge remains disabled. Mergeable is not merge-ready.

Of 25 expected checks: 5 successful, 19 pending, 0 failed, 1 absent at this snapshot. All 24 observed checks have one producer from the intended exact-head workflow; no duplicate producers were observed.

Check State Exact producer run / check ID
Lint 1 pending: in_progress 36167657667 / 108179541323
Lint 2 pending: in_progress 36167657667 / 108179541148
Test coverage guard successful: success 36167657667 / 108179540851
Behavior portable parallel 1 pending: in_progress 36167657667 / 108179541255
Behavior portable parallel 2 pending: in_progress 36167657667 / 108179541162
Behavior portable serial 1 pending: queued 36167657667 / 108179541429
Behavior portable serial 2 pending: queued 36167657667 / 108179541365
Behavior portable serial 3 pending: in_progress 36167657667 / 108179541447
Behavior portable serial 4 pending: in_progress 36167657667 / 108179541377
Behavior portable serial 5 pending: in_progress 36167657667 / 108179541019
Behavior portable serial 6 pending: in_progress 36167657667 / 108179541364
Behavior portable serial 7 pending: in_progress 36167657667 / 108179541095
Behavior portable serial 8 pending: in_progress 36167657667 / 108179541006
Behavior portable serial 9 pending: in_progress 36167657667 / 108179541141
Behavior tests (Herdr) pending: in_progress 36167657667 / 108179541114
Behavior timing aggregate absent: waiting for lane dependencies .github/workflows/ci.yml / tests-timing-aggregate
Stock macOS Bash snapshot compatibility pending: in_progress 36167657667 / 108179541143
Repo invariants successful: success 36167657667 / 108179541286
Windows self-update entry point successful: success 36167657663 / 108179221342
Windows reconciliation (core) pending: in_progress 36167657663 / 108179221212
Windows reconciliation (copilot-launch) successful: success 36167657663 / 108179221339
Windows reconciliation (legacy-rollback) successful: success 36167657663 / 108179221391
Windows reconciliation (pr-completion) pending: in_progress 36167657663 / 108179221137
Windows Copilot management pending: in_progress 36167657663 / 108179221113
Harness package compatibility pending: in_progress 36167657663 / 108179220797

The timing aggregate has an existing producer but does not start until its lane dependencies settle; absence is not a pass. Cancelled, skipped, or failed required lanes would not be accepted as validation.

The initial head's dependency-floor failure and correction are documented in the PR body. The corrected head already passes the previously failing Windows legacy-rollback job. Azure downstream completion and both portable parallel lanes still require their own final-head results; the pin fix is not a claim that all earlier failures are resolved.

Local execution stopped at the documented two-timeout circuit breaker; no local budget reset or additional local suite was run. This PR remains not fully validated until the required final-head matrix and unresolved observations are cleared. Preserve the requested merge-commit landing strategy and leave the PR unmerged.

Read mode and project through the existing batched metadata owner instead
of the removed meta_value helper. The empty values bypassed the named-head
completion guard and classified unpreserved work as done.

Install the private-path dependency closure for the copied AFK PR reader.
Expose its existing cases through the shared named-case registry while
preserving the default order.

Firstmate-Upstream-SHA: b805823
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@timbarreto

Copy link
Copy Markdown
Owner Author

Exact-head CI snapshot (2026-09-25T18:32:28.160Z)

Head: 85c1d779065e7447727aecf5af2c15a1870444e5; base: b13de801f35b53757a8c61ea18d8c577b07a7aeb. The PR is open, non-draft, unmerged, and GitHub reports it mergeable. Auto-merge remains disabled. Mergeable is not merge-ready.

Of 25 expected checks: 6 successful, 18 pending, 0 failed, 1 absent at this snapshot. All 24 observed checks have one producer from the intended exact-head workflow; no duplicate producers were observed.

Check State Exact producer run / check ID
Lint 1 pending: in_progress 36173901648 / 108199751776
Lint 2 pending: in_progress 36173901648 / 108199751576
Test coverage guard successful: success 36173901648 / 108199751669
Behavior portable parallel 1 pending: in_progress 36173901648 / 108199751338
Behavior portable parallel 2 pending: in_progress 36173901648 / 108199751763
Behavior portable serial 1 pending: in_progress 36173901648 / 108199752156
Behavior portable serial 2 pending: in_progress 36173901648 / 108199752078
Behavior portable serial 3 pending: in_progress 36173901648 / 108199752089
Behavior portable serial 4 pending: in_progress 36173901648 / 108199752039
Behavior portable serial 5 pending: in_progress 36173901648 / 108199752074
Behavior portable serial 6 pending: in_progress 36173901648 / 108199751975
Behavior portable serial 7 pending: in_progress 36173901648 / 108199752108
Behavior portable serial 8 pending: in_progress 36173901648 / 108199752166
Behavior portable serial 9 pending: in_progress 36173901648 / 108199752066
Behavior tests (Herdr) pending: in_progress 36173901648 / 108199751709
Behavior timing aggregate absent: waiting for lane dependencies .github/workflows/ci.yml / tests-timing-aggregate
Stock macOS Bash snapshot compatibility pending: in_progress 36173901648 / 108199751936
Repo invariants successful: success 36173901648 / 108199751928
Windows self-update entry point successful: success 36173901843 / 108199751536
Windows reconciliation (core) pending: in_progress 36173901843 / 108199752256
Windows reconciliation (copilot-launch) successful: success 36173901843 / 108199752013
Windows reconciliation (legacy-rollback) successful: success 36173901843 / 108199752231
Windows reconciliation (pr-completion) pending: in_progress 36173901843 / 108199752207
Windows Copilot management pending: in_progress 36173901843 / 108199751878
Harness package compatibility successful: success 36173901843 / 108199751911

The timing aggregate has an existing producer but does not start until its lane dependencies settle; absence is not a pass. Cancelled, skipped, or failed required lanes would not be accepted as validation.

The previous head finished with 23 successful checks and failures in portable parallel 2 and serial 7. Both exact failing cases now pass locally after the metadata-reader and copied-dependency repairs described in the PR body. Those targeted results do not replace the final-head CI matrix.

The user-approved focused follow-up used 485.804 seconds and 1 timeout within its 600-second/two-timeout cap. Both named regressions, changed-shell syntax, and source-aware fast lint passed; full dataflow and broad suites remain CI-owned. The original budget was not reset. This PR remains not fully validated until the required final-head matrix and unresolved observations are cleared. Preserve merge-commit landing and leave the PR unmerged.

@timbarreto
timbarreto merged commit 22517d8 into main Sep 25, 2026
25 checks passed
@timbarreto
timbarreto deleted the reconcile/upstream-2026-09-25-b805823a-b382ab77 branch September 25, 2026 18:48
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.