Skip to content

Reconcile upstream e0d269e (September 11, 2026) - #47

Merged
timbarreto merged 6 commits into
mainfrom
reconcile/upstream-2026-09-11-e0d269e
Sep 11, 2026
Merged

timbarreto merged 6 commits into
mainfrom
reconcile/upstream-2026-09-11-e0d269e

Conversation

@timbarreto

@timbarreto timbarreto commented Sep 11, 2026 •

Copy link
Copy Markdown
Owner

Frozen reconciliation

Normal merge of four canonical commits, preserving the fork's architecture and compatibility behavior. Use a merge commit, not squash/rebase. Merge commits are enabled; this ordinary non-draft PR stays unmerged, with auto-merge disabled. GitHub Actions owns the complete cross-platform matrix; this is not yet a merge-readiness claim.

Identity SHA
Frozen fork / actual PR base ff61baecb7df70222a1bf50271105a84878f69a5
Proven prior upstream 4768e98d469b7216569acf6e7a5cc696ecd15c2e
Frozen upstream e0d269e07318a5a80069ab6193b4a0af4c077a61
Upstream merge commit bdf50f25d5020c8944d3d60984cca5e926369ef3
Published head 22e143cbe049052dc8c4afca1c5afb8cb7de37a7
Reviewed / committed tree d70975fc67292372190f542be82c7d7d2f9144ae

Firstmate-Upstream-SHA: e0d269e

Upstream was fetched exactly once. The prior upstream is both frozen heads' merge base. Prior merge e855c91e49ca1003d2668f9fbffdb5d6722f0701 has parents 31bb375d24cb4dcc8a3402ce7f15d7b2e02c7826 and that prior upstream; retained head 7def92706f2d302d5dce6493f7234edaad3ca173 and PR #46 preserve its trailer. PR #46 merged September 11, 2026 at 01:43:08 UTC, producing the frozen fork. Merge commit bdf50f25d5020c8944d3d60984cca5e926369ef3 has the frozen fork and upstream as its two parents. The published head adds the CI correction below without rewriting that merge. Git's trailer parser, committed tree, exact scope, clean worktree, and unchanged local main were verified. No reconstruction or ancestry anchor was needed.

Integration decisions

Incoming: kunchenguid#4151 (lane timing/packing), kunchenguid#4148 (delivered-PR provenance), kunchenguid#4212 (confirmed process-event launches and notifications), and kunchenguid#4191 (Herdr process-level liveness).

Conflict Decision
bin/backends/tmux.sh Share upstream's classifier; retain exact copilot / copilot.exe recognition and dependency-error propagation in the shared owner and both consumers. No duplicate classifier.
bin/fm-test-run.sh Keep catalogs, batched serial scheduling, Windows handling, and independent proof admission. All 24 upstream parallel hints become separate parallel-duration records; upstream memberships and metrics remain. Only explicit parallel lanes use parallel weights.
tests/fm-test-run.test.sh Keep the named registry and fork cases, register upstream cases, and use complete catalog fixtures with deliberately opposed serial/parallel weights.
docs/tmux-backend.md Retain supported Copilot, signed Pi, OMP, and other identities while naming the shared owner.

Loader, override/cache tests, fixture projection, routes, and docs changed together. Shared-classifier routing covers both backends, Orca, live identity gates, and the native Treehouse contract through an exact fork override. Empty hint streams preserve serial fallback and accurate parallel coverage; missing parallel hints do not acquire a new coverage-guard failure policy.

Process-event starts require bounded confirmation; failed launches and stranded claims get distinct deduplicated notifications. Generation checks, ambiguous-group refusal, native privacy, and the timeout owner remain. Delivered PRs require recorded metadata or an exact terminal ready line, never arbitrary prose or scout attribution. Herdr stale-agent permits pane-reusing recovery, not closing, replacement, or duplicate spawning; unreadable liveness evidence refuses. Earlier direct-child Pi documentation now explicitly excludes the newer nested-shell case, preserving both dated measurements and anchors.

Verification ownership

Shared CI changed only comments versus the fork: executable jobs, permissions, triggers, matrices, timeouts, prerequisites, and artifacts are unchanged. Fork CI, proof admission, and timeout code are byte-identical. Both parallel membership functions and all 24 hints match frozen upstream; every prior core/fork serial hint is unchanged.

Retained: two parallel lanes, five serial shards, Pi setup in parallel lane 1, shared portable Pi installation policy, TypeScript 5.9.3, tasks-axi 0.2.5; real Herdr 0.7.4 / protocol 16, Treehouse 2.0.1, Pi 0.84.3; fork package Pi 0.84.3, OpenCode 1.18.23, TypeScript 5.9.3. Both workflows retain main push/PR triggers, read-only permissions, and workflow-qualified concurrency. Portable/Herdr timing uploads still feed the dependent aggregate. The Windows Herdr experiment remains manual-only.

The full Herdr response shape was measured upstream on 0.9.0; older-client subcommand presence is not server-response proof. Exact-head real-Herdr CI must establish pinned compatibility. Native transport fakes do not prove full Windows Herdr liveness. Existing live/optional gates remain limitations, not empirical passes. No live Firstmate session, vendor prompt, fleet action, or operational-record change occurred.

Local evidence

Recorded shell: C:\Program Files\Git\bin\bash.exe, never ambient WSL Bash; FM_LIVE=0 throughout. Probes passed: Git 2.55.0.windows.5, Bash 5.3.15, awk 5.4.1, cygpath 3.6.10, Node 24.19.0, ShellCheck 0.11.0, actionlint 1.7.12, jq 1.8.2, Python 3.13.15, Perl 5.42.3. No installs, disabled signing/hooks, or relaxation of safe.bareRepository=explicit.

Initial snapshot: All four repository gates and 13 focused behavioral invocations passed. Two POSIX-only Herdr cases were deferred before execution because ps -axo pid=,ppid=,comm= >/dev/null is unsupported here; Linux parallel 2 owns both. Before publication there was no executed test failure, timeout, interruption, retry, differential run, or cached pass. Only explanatory runtime-backend prose changed after those behavioral runs; the docs gate and complete lane/routing inventory were refreshed. Post-publication CI exposed the separate routing regression documented below; those failed reproductions are not counted as passes.

One 2,400-second budget: 815428 ms charged / 1584 s remaining, zero timeouts; caching disabled throughout. Invocation totals: preflight 5375 ms, validation-01 169595 ms, validation-02 373873 ms, validation-03 50801 ms, validation-ci-01 14213 ms, validation-ci-02 11000 ms, validation-ci-03 190571 ms. Probe/controller overhead is included.

Probes were the listed tools' --version commands (Perl used -e to print $^V). Read-only Git policy probes checked commit.gpgsign, core.hooksPath, and safe.bareRepository (initially 903 ms, repeated without changing policy for CI diagnosis). Preflight-only exited 75 without running tests; validation-02 exited 75 solely for the two prerequisite deferrals. Other initial validation invocations exited 0. CI diagnosis invocations 01 and 02 exited 1 on the reproduced assertion; invocation 03 exited 0. No validator is left running.

Gate recipes: S is the syntax sweep below; I is bin/fm-test-run.sh --list --changed --base ff61baecb7df70222a1bf50271105a84878f69a5, plus --list-lanes and --list --lane for both parallel, all five numbered serial, and real-Herdr lanes. I-short is the changed inventory plus assertions that Treehouse and the new stale-registration script are selected. Outputs were saved outside the checkout.

set -e
inventory=$(bin/fm-lint.sh --list-files)
while IFS= read -r script; do
  [ -z "$script" ] || bash -n "$script" || exit
done <<< "$inventory"

X1 is source-aware bin/fm-lint.sh with these explicit roots: bin/fm-agent-process-lib.sh bin/fm-test-catalog-lib.sh bin/fm-test-run.sh bin/backends/tmux.sh bin/backends/herdr.sh tests/fm-test-catalog.test.sh tests/fm-harness-contract.test.sh. X2 adds tests/fm-test-run.test.sh tests/fm-backend-herdr.test.sh.

Passed gate / command First ms Second ms Final docs-only refresh ms
S 3634 2680 unchanged inputs
bin/fm-lint.sh 41110 41162 unchanged inputs
bin/fm-doc-audience-check.sh 1742 1626 2151
bin/fm-test-run.sh --check-coverage 48962 46048 unchanged inputs
X1 / X2 17888 58797 unchanged inputs
I / I-short / I 50635 8943 48628

Initial behavior commands: bin/fm-test-run.sh --jobs 1 tests/<script>.test.sh; set FM_TEST_ONLY=<case> for named rows, omit it for FULL. Every executed row below passed with gate_skip=false; the two deferred rows were not executed. Common Bash path and FM_LIVE=0 are given above. These results remain attributed to their original snapshots, not relabeled as final-head CI passes.

Script Case ms / result
fm-test-catalog FULL 19290 passed
fm-harness-contract test_shared_backend_process_identity_keeps_pilot_boundaries 4266 passed
fm-harness-contract test_missing_or_broken_adapter_refuses 4742 passed
fm-test-run test_list_scheduled_proven_isolated_uses_serial_weights 6193 passed
fm-test-run test_list_scheduled_non_lane_selections_use_serial_weights 10173 passed
fm-test-run test_empty_duration_hints_preserve_serial_fallback_and_parallel_coverage 53075 passed
fm-test-run test_portable_parallel_lanes_stay_duration_balanced 46347 passed
fm-test-run test_fork_workflow_selects_its_contracts 12011 passed
fm-backend-herdr-treehouse FULL 4547 passed
fm-backend-herdr test_registered_agent_with_a_live_foreground_process_stays_alive 9703 passed
fm-backend-herdr test_registered_copilot_foreground_process_stays_alive 16342 passed
fm-backend-herdr test_registered_agent_with_an_unreadable_process_view_is_unknown 23021 passed
fm-backend-herdr test_projection_reclaim_rollback_refuses_a_stale_registration 3863 passed
fm-backend-herdr test_stale_registration_over_a_shell_only_pane_is_agent_free NOT RUN: POSIX ps; CI parallel 2
fm-backend-herdr test_busy_state_never_reports_a_shell_only_pane_busy NOT RUN: POSIX ps; CI parallel 2

Coverage: 210 = 24 parallel + 171 serial + 15 Herdr. All parallel members have hints; max packed sum 414299 ms, imbalance 30 ms, with the named case enforcing at most 5%. Estimates are not CI wall-time/headroom proof. Docs: 103 surfaces / 422 local links.

Post-publication CI correction

Behavior portable parallel 1, job 103337702426 on initial head bdf50f25d5020c8944d3d60984cca5e926369ef3, failed because changing herdr-client-pair-fixture.sh omitted fm-secondmate-safety.test.sh. The new curated Herdr route returned before the fork's reference expansion. Both references and the secondmate registration were intact.

Follow-up 22e143cbe049052dc8c4afca1c5afb8cb7de37a7 keeps the upstream mapping and adds reference-derived consumer families for shared fixtures/helpers. It changes only the runner, its existing named regression, and two fork guides within the same 46-path scope. The regression also protects curated backend coverage, unrelated-family exclusion, curated fixtures with no references, and refusal of fixtures with neither mapping nor consumers. No timeout, dependency, package pin, lane membership, or workflow change was needed.

For the named rows, the exact recipe is FM_TEST_ONLY=<case> bin/fm-test-run.sh --jobs 1 tests/fm-test-run.test.sh, under the recorded Bash with FM_LIVE=0. The original case failed on the initial committed tree; the strengthened case failed on test-only tree d61262c63a3721adf462d0b085c4643845bc8d2f before the production fix. All post-fix rows use committed tree d70975fc67292372190f542be82c7d7d2f9144ae.

Case / command Result ms
test_changed_shared_fixtures_select_consumers (original reproduction) failed: exact CI assertion 10168
test_changed_shared_fixtures_select_consumers (strengthened regression, one serial retry) failed: same assertion 9838
test_changed_shared_fixtures_select_consumers (fixed) passed 14398
test_changed_reference_scan_batches_test_files passed 9796
test_changed_dependency_selection_and_unmapped_failure passed 51741
test_changed_bin_reference_selects_per_script_not_per_family passed 8585
S passed 2713
bin/fm-lint.sh passed 41552
bin/fm-doc-audience-check.sh passed 2276
bin/fm-test-run.sh --check-coverage passed 46604
I-short passed; identical 175-script selection 9196

The failure reproduced on both Linux CI and local Windows; no platform-baseline claim or frozen-upstream differential was needed. No debug instrumentation was added. The original normal merge remains reachable. Exact-head CI was restarted by the ordinary push; old-head successes are not proof for this head.

No-renames audit

Comparison M A D Total + lines - lines
Prior upstream -> frozen fork 149 52 3 204 14718 2986
Frozen upstream -> fork (before) 163 52 5 220 14960 5551
Frozen upstream -> head (after) 150 52 3 205 14911 3044
PR base -> head 44 2 0 46 2720 262

Scope: 38 incoming paths + 8 explicitly recorded integration paths = 46. The initial merge's intended/staged/committed NUL path sets match exactly. The follow-up staged exactly four paths within that manifest; the final PR still has the identical 46-path scope. All modes match the fork or new upstream file. No broad renormalization. The two PR additions are upstream files, not additional fork modules. The before-comparison's two extra D entries are those not-yet-adopted upstream additions.

A smaller count is not behavior proof: persistent divergence against the prior sync grows 204 -> 205 paths, including the shared-classifier compatibility hunk; additive fork paths stay 52. Inherited deletions stay .claude/settings.json, .github/workflows/no-mistakes-required.yml, and tests/fm-no-mistakes-required.test.sh. No rename manufactures equivalence.

Complete before/after inventories are retained outside the checkout and reproducible with:

git diff --no-renames --name-status e0d269e07318a5a80069ab6193b4a0af4c077a61 ff61baecb7df70222a1bf50271105a84878f69a5
git diff --no-renames --name-status e0d269e07318a5a80069ab6193b4a0af4c077a61 22e143cbe049052dc8c4afca1c5afb8cb7de37a7
git diff --no-renames --numstat ff61baecb7df70222a1bf50271105a84878f69a5 22e143cbe049052dc8c4afca1c5afb8cb7de37a7

Exactly upstream-equivalent now (15, including the new live guard): .agents/skills/process-event-sources/SKILL.md, bin/fm-backend.sh, bin/fm-inactive-reconcile.sh, docs/verification/process-event-sources.md, docs/verification/rovo.md, tests/fm-control-herdr-smoke.test.sh, tests/fm-crew-state.test.sh, tests/fm-cursor-harness.test.sh, tests/fm-herdr-pi-stale-registration-live-e2e.test.sh, tests/fm-omp-harness.test.sh, tests/fm-tmux-agent-liveness.test.sh, tests/fm-watch-arm.test.sh, tests/fm-watch-triage.test.sh, tests/herdr-client-pair-fixture.sh, tests/remote-herdr-fixture.sh.

Classifier code moved to the shared owner, and parallel timing data to catalogs; no duplicate classifier or inline timing table remains. Scheduling, proof admission, and backend transactions intentionally remain separate. Retained exceptions/removal conditions (full ownership in docs/fork/architecture.md):

Retained boundary Remove only when
Pilot adapter calls / explicit load errors Upstream supplies equivalent capabilities and failure semantics across lifecycle callers.
Pi/OpenCode compatibility imports Supported consumers share an equivalent public import contract.
Native process / ACL / transport modules and caller policy Upstream preserves privacy, ownership, rollback and subprocess bounds.
Catalog loader / runner / proof Upstream supplies equivalent metadata, reference routing and independent admission.
Lint discovery / fixture closure Upstream discovers and distributes the full owning module closure.
Fork and shared CI integration Every required native/package/portable subject, gate and single producer remains covered.
Legacy nonpilots / hash-pinned remote doctor A separately reviewed migration preserves those contracts and integrity protocol.
Classified documentation owners Upstream preserves supported guidance, anchors and safety facts.

All 175 selected scripts and their CI owners

Names below expand to tests/<name>.test.sh. Default local disposition: full regression deferred to the named CI job. FULL = full script passed locally; CASE = only tabled named cases passed. LIVE, HERDR, OPTIONAL identify unchanged capability gates, not empirical passes. Procevent/delivered-PR full suites belong to serial 3; other watcher, stop-proof and lifecycle subjects are explicitly assigned below. Real mutations/cleanup need the Herdr producer; stock Bash needs macOS; native Windows and installed packages need Fork CI.

Complete per-script routing (175)

Behavior portable parallel 1 (9)

fm-brief
fm-cd-pretool-check
fm-composer-ghost
fm-composer-lib
fm-grok-harness
fm-lint
fm-pi-primary-types
fm-test-run [CASE]
fm-tmux-submit-busy

Behavior portable parallel 2 (12)

fm-arm-pretool-check
fm-backend-herdr [CASE]
fm-captain-hold-lifecycle
fm-crew-state
fm-ensure-agents-md
fm-herdr-lab
fm-send-popup-settle
fm-send-settle
fm-send-strict
fm-spawn-batch
fm-supervision-instructions
fm-transition-lib

Behavior portable serial 1 (27)

fm-afk-pi-herdr-return-e2e [LIVE]
fm-branch-supervision
fm-codex-continuity-live-e2e [LIVE]
fm-copilot-harness
fm-copilot-hooks-live-e2e [LIVE]
fm-cursor-harness
fm-cursor-primary
fm-documentation-audiences
fm-herdr-pi-stale-registration-live-e2e [LIVE]
fm-herdr-version-floor-live-e2e [LIVE]
fm-on
fm-peek-remote
fm-pending-reply
fm-pi-branch-live-e2e [LIVE]
fm-private-path
fm-procevent-when
fm-quota-choose
fm-remote-backlog-handoff
fm-secondmate-liveness
fm-secondmate-sync
fm-send-secondmate-marker
fm-spawn-worktree-settle
fm-task-delivery
fm-test-fixtures
fm-turnend-guard
fm-watch-checkpoint
fm-watch-triage

Behavior portable serial 2 (29)

fm-backend-orca [OPTIONAL]
fm-busy-adapter-wiring
fm-claude-stop-autoarm-live-e2e [LIVE]
fm-control-relaunch
fm-extension-binding
fm-harness-contract [CASE]
fm-harness-liveness-drift-live-e2e [LIVE]
fm-herdr-submit-confirm-live-e2e [LIVE]
fm-live-gate
fm-omp-harness
fm-operational-input
fm-pi-branch-extension
fm-pi-branch-responsiveness-live-e2e [LIVE]
fm-platform-process
fm-procevent-quota
fm-quota-array-dispatch-live-e2e [LIVE]
fm-remote-entrypoint
fm-remote-job-orphan-reap
fm-remote-secondmate-lifecycle-e2e
fm-send-remote-delivery
fm-send-secondmate-marker-herdr-e2e [LIVE]
fm-sessionstart-hook-live-e2e [LIVE]
fm-shared-captain-inheritance
fm-startup-memory-budget
fm-supervision-events
fm-vendor-auth-probe
fm-wake-drain-unread-status
fm-wake-queue
fm-watch-recovery-loop

Behavior portable serial 3 (30)

fm-bearings-board-lavish-live-e2e [LIVE]
fm-busy-state
fm-classify-corr-token
fm-claude-stop-autoarm
fm-composer-matrix-live-e2e [LIVE]
fm-control
fm-gitignore-config
fm-inactive-reconcile
fm-kimi-harness
fm-lint-workflows
fm-mail-check
fm-opencode-primary-live-e2e [LIVE]
fm-pi-codex-native [LIVE]
fm-pi-primary-live-e2e [LIVE]
fm-pi-watch-extension
fm-procevent
fm-project-origin
fm-public-followup
fm-reconcile-validation
fm-remote-doctor
fm-remote-herdr-guard
fm-remote-secondmate-trace-context
fm-rovo-signals-live-e2e [LIVE]
fm-sessionstart-instruction-refresh-live-e2e [LIVE]
fm-teardown-endpoint-safety
fm-test-fixture-cleanup
fm-voice-relay
fm-wake-daemon-lifecycle-e2e
fm-wake-drain-outcome-backstop
fm-watcher-lock

Behavior portable serial 4 (25)

fm-backend-herdr-treehouse [FULL]
fm-backend-tmux-smoke
fm-bearings-board
fm-calm-pi-extension
fm-claude-trust
fm-cursor-primary-live-e2e [LIVE]
fm-daemon
fm-grok-continuity-live-e2e [LIVE]
fm-herdr-session-cleanup
fm-lint-inventory
fm-mail
fm-muse-harness
fm-muse-signals-live-e2e [LIVE]
fm-procevent-stop-proof
fm-remote-reply
fm-secondmate-lifecycle-e2e
fm-secondmate-reconcile
fm-secondmate-safety
fm-send-resolve-key
fm-session-lock-ancestry
fm-stow-cascade
fm-subagent-pretool-check
fm-test-isolation-proof
fm-wake-drain-open-decisions-cursor
fm-watch-arm

Behavior portable serial 5 (28)

fm-ask-user-authority
fm-backend
fm-backlog-handoff
fm-classify-decision-key
fm-cmux-claude-composer-live-e2e [LIVE]
fm-grok-stop-live-e2e [LIVE]
fm-guard-stale-banner
fm-harness-adapter-instructions-live-e2e [LIVE]
fm-harness-adapter-references
fm-omp-primary-live-e2e [LIVE]
fm-remote-job-wait
fm-remote-job
fm-remote-secondmate-parent-binding
fm-remote-transport-lanes
fm-rovo-harness
fm-secondmate-harness
fm-secondmate-restart
fm-send-inbox-doorbell-live-e2e [LIVE]
fm-send-inbox
fm-spawn-dispatch-profile
fm-spawn-pool-base-freshen
fm-task-inbox
fm-test-catalog [FULL]
fm-tmux-agent-liveness
fm-tool-update-check
fm-trace-context-lib
fm-trace-context-spawn
fm-wake-drain-open-decisions

Behavior tests (Herdr) (15)

fm-afk-inject-herdr-e2e [HERDR]
fm-afk-launch [HERDR]
fm-backend-autodetect-smoke [HERDR]
fm-backend-herdr-agent-exit-shell-e2e [HERDR]
fm-backend-herdr-eventwait-smoke [HERDR]
fm-backend-herdr-focus-flash-e2e [HERDR]
fm-backend-herdr-launcher-workspace-e2e [HERDR]
fm-backend-herdr-presentation-e2e [HERDR]
fm-backend-herdr-prune-safety-e2e [HERDR]
fm-backend-herdr-respawn-idem-e2e [HERDR]
fm-backend-herdr-smoke [HERDR]
fm-backend-herdr-stale-active-tab-e2e [HERDR]
fm-backend-herdr-workspace-per-home-e2e [HERDR]
fm-control-herdr-smoke [HERDR]
fm-herdr-session-cleanup-e2e [HERDR]

Exact-head CI inventory

Each expected name has one verified producer. The post-publication CI comment records successful/pending/failed/absent states for this exact head; skipped or cancelled required lanes are not proof. The manual Windows Herdr experiment is separate.

Expected check Workflow / job
Lint CI / lint
Test coverage guard CI / test-coverage
Behavior portable parallel 1 CI / tests-portable-parallel-1
Behavior portable parallel 2 CI / tests-portable-parallel-2
Behavior portable serial 1 CI / tests-portable-serial / shard=1
Behavior portable serial 2 CI / tests-portable-serial / shard=2
Behavior portable serial 3 CI / tests-portable-serial / shard=3
Behavior portable serial 4 CI / tests-portable-serial / shard=4
Behavior portable serial 5 CI / tests-portable-serial / shard=5
Behavior tests (Herdr) CI / tests-herdr
Behavior timing aggregate CI / tests-timing-aggregate
Stock macOS Bash snapshot compatibility CI / macos-stock-bash
Repo invariants CI / invariants
Windows self-update entry point Fork CI / windows-update
Windows reconciliation (core) Fork CI / reconciliation-windows / subject=core
Windows reconciliation (copilot-launch) Fork CI / reconciliation-windows / subject=copilot-launch
Windows reconciliation (legacy-rollback) Fork CI / reconciliation-windows / subject=legacy-rollback
Harness package compatibility Fork CI / harness-package-compatibility

mremond and others added 6 commits September 10, 2026 23:55
…nchenguid#4151)

* ci: rebalance the portable parallel lanes on measured runner durations

Both portable parallel lanes are capped at 10 minutes. Lane 1 was cancelled at
that cap on every request raised on 2026-09-10 while lane 2 finished in about
3.5 minutes, so no request could go green.

CONTRACT CLASS: RESTORE.
The workflow already promises two duration-balanced lanes and the shard
documentation already claims a measured wall; this re-establishes both against
what the lanes now cost, and changes no lane count, no cap, and no scope of what
runs. The counter-argument, so nobody has to take that on trust: two pieces here
are genuinely new rather than restored, and either could be argued to make this
a NEW-behavior change. `--list-scheduled` now ranks a parallel lane on measured
durations where it previously handed every parallel script the serial default
weight and returned an alphabetical order; and `--check-coverage` gains three
reported fields. I classify the change RESTORE because both exist only to make
the already-promised property checkable, but they are named here rather than
folded into the restoration.

=== PART 1: THE TOTAL, AND HOW IT WAS OBTAINED ===

This section stands on its own. It establishes what the parallel set costs. It
derives no packing; Part 2 does that, from this number.

THE TOTAL: 828568 ms, about 13 min 49 s of serial work across the 24 scripts.
Lane 1 held 624299 ms of it and lane 2 held 204269 ms, a 3.06:1 split.

HOW IT WAS OBTAINED. The difficulty was that lane 1 had never finished, so its
duration did not exist as a recorded figure anywhere and no timing artifact was
expected for it. It turned out to be recoverable from the real lane without
estimating, by two routes, across six CI runs on 2026-09-10 (34459949083,
34460760299, 34462530836, 34462758357, 34466966385, 34470382458):

  - Run 34462758357's lane-1 job finished its suite 18 s BEFORE the wall and
    uploaded a complete fm-test-timing-portable-parallel-1 artifact carrying all
    11 scripts, FM_TEST_SUMMARY total=11 failed=0 duration_ms=598225. The
    upload step is if: always(), so the cancellation did not suppress it. This
    is one full, untruncated lane-1 measurement.
  - The five other lane-1 jobs were cancelled mid-suite, but each logs every
    script that had already finished as an FM_TEST_END duration_ms= marker.
    Those per-script records are complete measurements of completed scripts;
    only the script in flight at cancellation is lost, and it differs by run.

Lane 2 completed in all six runs, so its scripts come from the six uploaded
fm-test-timing-portable-parallel-2 artifacts.

Every one of the 24 scripts therefore carries at least one untruncated
measurement: 20 of them measured in all six runs, two in three or four runs, and
two (fm-brief, fm-transition-lib, the tail of lane 1) in the single complete run.
Each hint is the SLOWEST value that script reached, so the total is an upper
envelope rather than an average. NO FIGURE IN IT IS DERIVED FROM A TRUNCATED
LANE, and no lower bound was ever extrapolated into a total.

THE ENVIRONMENT, AND WHETHER IT TRANSFERS. Every hint is a serial run of the
real portable parallel lane on a GitHub ubuntu-latest runner, produced by the
lane's own CI job. It transfers because it is not a proxy for the lane; it is
the lane. Nothing in the total came from this machine or from any harness of
mine.

That mattered, and here is what it would have cost. A same-day macOS
cross-check of the same scripts ran 1.7x to 5.0x slower with the ratio varying
per script (fm-test-run 157420 ms against 92944 ms, fm-x-mode 67217 ms against
31870 ms, fm-composer-ghost 10521 ms against 2120 ms). Local timings therefore
do not scale the lane, they REORDER it, so a packing derived from them would
have balanced the wrong thing while looking clean.

WHAT IT REPLACES, which is the root cause. The lanes were packed from the
2026-08-20 concurrent isolation proof: 24 candidates across four LOCAL workers.
That record answers whether the candidates are isolation-safe, not how long a
SERIAL CI lane runs, so it was structurally incapable of representing lane wall
clock even when it was fresh. It was also never refreshed while the set grew
about 3.2x. Both the wrong instrument and the staleness are fixed here: the
hints now come from the lane itself and carry their run ids and date.

=== PART 2: THE SPLIT DERIVED FROM THAT TOTAL ===

Longest-processing-time assignment over those hints gives 414269 ms and
414299 ms, 30 ms apart, against 624299/204269 before.

tests/fm-pi-primary-types.test.sh stays in lane 1 because that is the job which
installs the Pi package, so ci.yml needs no step changes.

=== PART 3: DOES THE MARGIN SURVIVE MACHINE VARIANCE ===

Stated explicitly, because 6.90 min against a 10 min cap is 69% of cap before
any variance is applied, and the cap covers the whole job rather than the suite.

  worst lane, script time                         414299 ms   6.90 min
  job overhead, measured on the real lane             ~18 s   (see below)
  expected healthy job                            ~432300 ms  7.21 min
  x1.29 on the script time, plus overhead         ~552400 ms  9.21 min
  cap                                             600000 ms  10.00 min
  room left after the multiplication                ~47.6 s   7.9% of cap

The 1.29x is the runner variance measured today on the SIBLING SERIAL lane, as
supplied; it is not this lane's own figure. This lane family does have its own,
and it is tighter: the six full lane-2 sums today span 192939 ms to 203451 ms,
a spread of 1.054x. At that figure the worst lane lands near 7.58 min with about
2.4 min of room. I have used the LARGER, borrowed 1.29x for the verdict rather
than the tighter one this lane actually shows, and note that the hints are
already per-script maxima, so 1.29x on top is conservative twice over.

THE MARGIN SURVIVES THE MULTIPLICATION, so this proceeds rather than stopping.
The 18 s overhead is measured, not assumed: in run 34462758357 the lane-1 job
ran 10 min 16 s against a 598.2 s suite, and lane 2 ran 3 min 21 s against a
192.9 s suite, a ~10 s difference that matches lane 1's extra Pi package install.

The cap is unchanged, the lane count is unchanged, and nothing in the serial
lane, its shard count, its guard or its hint table is touched.

=== PART 4: THE RECORDED FACT ===

The workflow comment no longer restates the shard wall as a literal, which is
how "~1 min of serial sum" survived a 10x change without announcing it. It now
points at bin/fm-test-run.sh --check-coverage, which prints parallel_max_ms,
parallel_imbalance_ms and parallel_unhinted derived from the hint table, so the
current number is computed on demand. The shard documentation carries the dated
run ids, which route it was taken by, and the local cross-check that shows why
local numbers are not admissible as hints.

Two regressions pin what rotted: lane membership must be stored
longest-measured-first, and the lanes must be fully hinted and packed within 5%
of each other. Both were run against the old composition and both fail on it
(420030 ms imbalance against a 624299 ms worst lane). The ordering assertion they
replace named a specific script by hand and had itself gone stale.

=== PART 5: NAMED AND LEFT, OUTSIDE THIS REBALANCE ===

tests/fm-captain-hold-lifecycle.test.sh alone is 296481 ms, 36% of the whole
set, so it is the floor of any two-lane split: no repacking can put a lane below
it. After this rebalance the cap is about 1.45x the healthy lane where the
sibling serial lane keeps roughly 2x.

Nothing refuses a stale parallel hint the way PORTABLE_SERIAL_MAX_UNHINTED_PERCENT
bounds the serial lane. parallel_unhinted is reported, not enforced, which is
what let this drift for three weeks unnoticed.

* fix(review): Restrict parallel scheduling hints to portable parallel lanes

* fix(document): Clarify parallel lane scheduling and timing evidence
… PR (kunchenguid#4148)

pr_for_task fell back to scraping the whole status log with tail -1, so
any PR URL a worker ever mentioned in prose - including a scout citing
someone else's PR - became the task's delivered PR in the parent-channel
terminal report. Recorded meta pr= is now the only authoritative source,
the fallback scrape accepts only a preferred terminal line in a mode's
ready-signal shape (done: PR <url> or done: PR <url> checks green), and
a scout never carries pr= at all.
…claims instead of counting a dead drop as started (kunchenguid#4212)

* fix(procevent): stop a dead runner owning a source and reconcile reporting it

The captain answered ten calls on a bearings board, the board accepted
them, and nothing collected them. He had to answer all ten again in chat.
A surface that presents as armed while being a dead drop is worse than one
that visibly fails, because the answers looked recorded.

Two independent defects, reproduced together in an isolated home where
reconcile reports started=1 on every run while ownership never moves and
no runner ever attaches.

1. reconcile counted a launch it never verified. detach_runner is
   fire-and-forget and discards the child's stderr, so a runner that died
   before it could claim was counted exactly like one that is listening.
   Launches are now confirmed - the source observed owned, or its runner
   record moved - before being reported as started; the rest are reported
   as failed= with a non-zero exit. The runner-record clause is what keeps
   a fast-completing source from being reported as a failure when it
   finished between two polls. One bounded window covers a whole cycle's
   launches, so a home full of broken sources costs the same wait as one.

2. A claim whose whole generation is provably gone could be refused
   forever. Reclaiming it ran cleanups over that dead generation's own
   leftovers, and any failure vetoed the claim - permanently, because none
   of those conditions clears on its own. Every one of those leftovers is
   keyed by the dead generation's claim token and a replacement always
   claims a fresh one, so none can collide with what replaces it.
   fm_procevent_claim_capture_reservation_reclaim_locked already said this
   for the reservation record; the staging file and the shape check on the
   registry directory recorded to hold it now take the same rule. Removing
   the claim record itself stays a hard precondition: two owners is the one
   outcome worse than none.

Two smaller repairs to the same "registered is not listening" confusion:

- `list` reported OWNER=none for a source nothing can claim. A reused PID
  whose process group survives reaches that state through the stale branch
  rather than the leaderless one, so it read as an idle source waiting to
  be started - the reassuring answer this surface gave while a board
  collected nothing. It now reports the orphaned state it shares.
- reconcile relaunched into that same unclaimable state on every cycle,
  spawning a runner that could only die on the claim. docs/configuration.md
  already promised it preserves such a claim without starting a
  replacement; the code now does that and reports it as uncertain.

This is NOT a third instance of today's two lock-identity defects
(4e1bf9aa and its replayed predecessor). Those were wrong liveness
predicates: a reused PID read as a live holder, then an exec'd holder read
as dead. Here the predicate is right - the code correctly proves the owner
dead and refuses the claim anyway, on a condition unrelated to liveness.

Regression coverage, each failing on the parent commit for its own reason:
- tests/fm-procevent.test.sh: a source that cannot start is reported as
  failed rather than started; a dead generation whose leftovers cannot be
  tidied no longer keeps owning its source (the parent reports a start
  while nothing ever runs); the existing reused-PID fixture now also
  asserts the orphaned listing and that no doomed relaunch is reported.
- tests/fm-captain-hold-lifecycle.test.sh: a board answer reaches the
  keyed-answer intake through the runner end to end - durable capture, the
  wake, and the closed task carrying the captain's selection. This one
  passes on the parent, because that chain was never what broke.

fm-procevent 100, fm-bearings-board 18, fm-captain-hold-lifecycle 50,
fm-procevent-when 13 and fm-procevent-quota 18 pass; bin/fm-lint.sh and
bin/fm-doc-audience-check.sh clean. tests/fm-extension-binding.test.sh has
two failures identical on the parent commit (EACCES on package install in
this sandbox) and unrelated to this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gxgshn5jkWJ3GEYWy7vTG

* no-mistakes(review): confirm reconcile launches on durable launch stamps

* no-mistakes(review): announce stranded sources and refuse bad confirm windows

* no-mistakes(review): announce leaderless strands, bound confirm window, fix recovery docs

* no-mistakes(review): announce unconfirmed launches once per episode, qualify start reclaim

* no-mistakes(review): nonce launch-failed keys, refuse bad window at arm

* no-mistakes(review): state only observed launch outcome, shorten episode nonce

* no-mistakes(test): assert launch-failed headline not re-delivered, allow recovery wake

* no-mistakes(document): docs: cover strand and launch-failure wakes in skill trigger and verification record

* no-mistakes(lint): restructure SC2015 chain into explicit if-block

* test(watch-triage): fix two timing-exposed defects the pipeline found

Both surfaced in the no-mistakes test step on this branch, each failing one
full run of tests/fm-watch-triage.test.sh; neither was accepted as a flake to
retry past.

1. The new launch-failed delivery test assumed an already-surfaced key never
   wakes the watcher again. That is false: a fresh watcher legitimately
   re-surfaces any unacknowledged queue row through its downtime-recovery
   path ("check: rearm-resurface"), so the assertion failed whenever a
   re-arm landed between its two checks. The pipeline's own fix tolerated any
   wake lacking the repeated key's headline; this tightens it to exactly one
   tolerated reason, by its exact line, with a failure message that names the
   expectation so a reworded path reads as "the tolerated recovery path
   changed" rather than as a mystery - and so nobody restores the strict
   silence check. The positive assertion (a fresh-suffix key is delivered
   under its own headline) is unchanged.

2. seed_captured_procevent_result retired its source in the gap between the
   runner publishing its wake and releasing its claim, so retire read the
   exiting runner's ownership as uncertain and refused ("cannot confirm
   runner identity"). The fixture and retire path pre-date this branch; the
   confirm window returns reconcile closer to the moment of capture, which
   made the gap easier to hit. The fixture now waits, bounded, for the claim
   release the publish promises, with the reason at the wait.

Verified on this head with tasks-axi on PATH: fm-watch-triage 113/113 with
no skips, fm-procevent 106/106, fm-captain-hold-lifecycle 50/50,
fm-watch-arm 15/15, fm-bearings-board 18/18, fm-procevent-when 13/13,
fm-procevent-quota 18/18; bin/fm-lint.sh and bin/fm-doc-audience-check.sh
exit 0. First attempt, no retries.

* no-mistakes(document): docs: route stranded and launch-failed wakes in skill handling

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…gistration (kunchenguid#4191)

* fix(herdr): verify agent registrations at process level before trusting them

Herdr keeps a Pi registration (`agent get` -> agent=pi, agent_status=idle)
after the Pi process has exited to a plain shell whenever a nested interactive
shell sits under the pane's top shell, which is the crew shape `treehouse get`
leaves behind. The pane classifier trusted that registration alone, so
`fm-control.sh <id> relaunch`, `fm-spawn.sh --relaunch`, and the crew-state
recovery read all treated a shell-only pane as a live agent and refused
recovery for as long as the record lived.

The Herdr adapter now reads `pane process-info` plus the real process table
through a shared harness-process classifier (bin/fm-agent-process-lib.sh,
moved verbatim out of the tmux adapter so both backends mean the same thing by
agent, shell, and other) before a registered agent counts as live. A
registration over a shell-only pane is the new explicit `stale-agent` pane
state, which the recovery-grade read maps to `dead`; husk detection, reclaim,
presentation recovery, and session cleanup keep refusing it, so recovery reuses
the pane and nothing gains close authority. A working record is verified the
same way before the native busy verdict reports busy, so the recovery
classifier never reports a shell-only pane as working. An unreadable process
view reads unknown, trusting neither the registration nor its absence.

Reproduced and measured on Herdr 0.9.0 with Pi 0.85.1 in an isolated lab; the
new default-on live guard tests/fm-herdr-pi-stale-registration-live-e2e.test.sh
exercises the real stale record, tests/fm-control-herdr-smoke.test.sh proves
exit and relaunch through the control plane, and the portable suites pin the
classifier over real processes.

Fixes kunchenguid#4115. Duplicates: kunchenguid#3639, kunchenguid#3487, kunchenguid#2908, kunchenguid#3545.

* no-mistakes(review): settle transient prompt helpers before trusting herdr process state

* no-mistakes(review): drop stray codegraph file; read spaced comm whole in descendant walk

* no-mistakes(review): untrack stray .codegraph/.gitignore

* no-mistakes(review): untrack codegraph file; make spaced-path walk test discriminating

* no-mistakes(review): untrack stray .codegraph/.gitignore

* no-mistakes(review): untrack stray .codegraph/.gitignore re-added by fix round

* no-mistakes(review): untrack stray .codegraph/.gitignore

* no-mistakes(review): untrack codegraph file, drop dead control case, record process-info floor

* no-mistakes(review): refuse stale-agent on fresh herdr spawn preflight

Documented non-goal: fresh-spawn, reclaim, and presentation-recovery auto-recovery for a stale-agent pane is a separate design change, out of scope here, to be proposed upstream as its own issue if wanted.

* no-mistakes(test): Fix herdr flake: don't misread transient empty foreground as unreadable

* no-mistakes(document): Add fm-agent-process-lib.sh to scripts inventory

* no-mistakes(fix): update remote herdr fixture to the real pane process-info shape

The shared remote-secondmate herdr fixture still returned the old flat
process-info body ({"result":{"process":{"name":...}}}). The process-level
liveness classifier added for kunchenguid#4115 requires the real
{"result":{"type":"pane_process_info","process_info":{...foreground_processes}}}
shape and treated the old body as unreadable, so an already-launched remote
endpoint's agent-state read failed and any relaunch attempt against it died
with "remote endpoint state is unreadable; refusing duplicate launch"
instead of reaching the state it was actually exercising
(tests/fm-remote-secondmate-parent-binding.test.sh,
tests/fm-remote-secondmate-lifecycle-e2e.test.sh).

* no-mistakes(review): test: add empty-foreground regression test for herdr flake fix

* no-mistakes(document): docs: register new stale-registration live-e2e test in herdr entry points
Merge the frozen canonical snapshot while preserving the fork's Copilot/Pi
interfaces, native process and private-path owners, catalog runner, and CI.
Share backend process recognition and keep parallel-lane timing records
separate from serial scheduling weights and concurrency admission.

Firstmate-Upstream-SHA: e0d269e
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The upstream Herdr pairing-fixture mapping returned before the fork's
reference expansion, omitting secondmate coverage. Keep curated families
and add reference-derived shared-fixture consumers without hard-coding
another family or dropping the upstream backend mapping.

Extend the named regression to retain unrelated-family exclusion,
unreferenced curated mappings, and unmapped-fixture refusal. Document the
additive routing contract. The reported CI assertion was reproduced twice
locally before the fix; the strengthened case, three adjacent routing
cases, and all four repository gates pass afterward.

Firstmate-Upstream-SHA: e0d269e
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@timbarreto

timbarreto commented Sep 11, 2026 •

Copy link
Copy Markdown
Owner Author

CI after fixture-routing fix

Observed 2026-09-11T17:14:38.979Z, exact head 22e143cbe049052dc8c4afca1c5afb8cb7de37a7, base ff61baecb7df70222a1bf50271105a84878f69a5.

The original Behavior portable parallel 1 failure is fixed and that lane is now successful on this head. Curated fixture mappings now preserve reference-derived consumers; the existing and strengthened assertion failed before the fix and passed afterward. The normal upstream merge remains reachable.

16 successful, 1 pending, 0 failed, 1 absent across 18 expected checks. No previous-head success is reused.

Expected check State Exact-head evidence
Lint successful success
Test coverage guard successful success
Behavior portable parallel 1 successful success
Behavior portable parallel 2 successful success
Behavior portable serial 1 successful success
Behavior portable serial 2 successful success
Behavior portable serial 3 successful success
Behavior portable serial 4 pending in_progress
Behavior portable serial 5 successful success
Behavior tests (Herdr) successful success
Behavior timing aggregate absent Not yet reported; depends on portable/Herdr jobs
Stock macOS Bash snapshot compatibility successful success
Repo invariants successful success
Windows self-update entry point successful success
Windows reconciliation (core) successful success
Windows reconciliation (copilot-launch) successful success
Windows reconciliation (legacy-rollback) successful success
Harness package compatibility successful success

The timing aggregate is not yet reported and depends on the portable/Herdr producers. Skipped or cancelled required checks would not count as passes. Live/optional/manual coverage remains limited as described in the PR body.

GitHub reports mergeable=true, mergeable_state=unstable. The PR is open, non-draft, unmerged, and has auto-merge disabled. Not yet merge-ready while required checks remain incomplete. Use a merge commit, not squash/rebase.

Local evidence: 815428 ms charged to the shared 2400-second budget, 1584 seconds remaining, zero timeouts. The four post-fix named cases and all four local gates passed; the 175-script selection, 46-path scope, and executable modes were verified. Full cross-platform evidence remains owned by GitHub Actions.

@timbarreto
timbarreto merged commit 502bb35 into main Sep 11, 2026
18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants