Reconcile upstream e0d269e (September 11, 2026) - #47
Conversation
…nchenguid#4151) * ci: rebalance the portable parallel lanes on measured runner durations Both portable parallel lanes are capped at 10 minutes. Lane 1 was cancelled at that cap on every request raised on 2026-09-10 while lane 2 finished in about 3.5 minutes, so no request could go green. CONTRACT CLASS: RESTORE. The workflow already promises two duration-balanced lanes and the shard documentation already claims a measured wall; this re-establishes both against what the lanes now cost, and changes no lane count, no cap, and no scope of what runs. The counter-argument, so nobody has to take that on trust: two pieces here are genuinely new rather than restored, and either could be argued to make this a NEW-behavior change. `--list-scheduled` now ranks a parallel lane on measured durations where it previously handed every parallel script the serial default weight and returned an alphabetical order; and `--check-coverage` gains three reported fields. I classify the change RESTORE because both exist only to make the already-promised property checkable, but they are named here rather than folded into the restoration. === PART 1: THE TOTAL, AND HOW IT WAS OBTAINED === This section stands on its own. It establishes what the parallel set costs. It derives no packing; Part 2 does that, from this number. THE TOTAL: 828568 ms, about 13 min 49 s of serial work across the 24 scripts. Lane 1 held 624299 ms of it and lane 2 held 204269 ms, a 3.06:1 split. HOW IT WAS OBTAINED. The difficulty was that lane 1 had never finished, so its duration did not exist as a recorded figure anywhere and no timing artifact was expected for it. It turned out to be recoverable from the real lane without estimating, by two routes, across six CI runs on 2026-09-10 (34459949083, 34460760299, 34462530836, 34462758357, 34466966385, 34470382458): - Run 34462758357's lane-1 job finished its suite 18 s BEFORE the wall and uploaded a complete fm-test-timing-portable-parallel-1 artifact carrying all 11 scripts, FM_TEST_SUMMARY total=11 failed=0 duration_ms=598225. The upload step is if: always(), so the cancellation did not suppress it. This is one full, untruncated lane-1 measurement. - The five other lane-1 jobs were cancelled mid-suite, but each logs every script that had already finished as an FM_TEST_END duration_ms= marker. Those per-script records are complete measurements of completed scripts; only the script in flight at cancellation is lost, and it differs by run. Lane 2 completed in all six runs, so its scripts come from the six uploaded fm-test-timing-portable-parallel-2 artifacts. Every one of the 24 scripts therefore carries at least one untruncated measurement: 20 of them measured in all six runs, two in three or four runs, and two (fm-brief, fm-transition-lib, the tail of lane 1) in the single complete run. Each hint is the SLOWEST value that script reached, so the total is an upper envelope rather than an average. NO FIGURE IN IT IS DERIVED FROM A TRUNCATED LANE, and no lower bound was ever extrapolated into a total. THE ENVIRONMENT, AND WHETHER IT TRANSFERS. Every hint is a serial run of the real portable parallel lane on a GitHub ubuntu-latest runner, produced by the lane's own CI job. It transfers because it is not a proxy for the lane; it is the lane. Nothing in the total came from this machine or from any harness of mine. That mattered, and here is what it would have cost. A same-day macOS cross-check of the same scripts ran 1.7x to 5.0x slower with the ratio varying per script (fm-test-run 157420 ms against 92944 ms, fm-x-mode 67217 ms against 31870 ms, fm-composer-ghost 10521 ms against 2120 ms). Local timings therefore do not scale the lane, they REORDER it, so a packing derived from them would have balanced the wrong thing while looking clean. WHAT IT REPLACES, which is the root cause. The lanes were packed from the 2026-08-20 concurrent isolation proof: 24 candidates across four LOCAL workers. That record answers whether the candidates are isolation-safe, not how long a SERIAL CI lane runs, so it was structurally incapable of representing lane wall clock even when it was fresh. It was also never refreshed while the set grew about 3.2x. Both the wrong instrument and the staleness are fixed here: the hints now come from the lane itself and carry their run ids and date. === PART 2: THE SPLIT DERIVED FROM THAT TOTAL === Longest-processing-time assignment over those hints gives 414269 ms and 414299 ms, 30 ms apart, against 624299/204269 before. tests/fm-pi-primary-types.test.sh stays in lane 1 because that is the job which installs the Pi package, so ci.yml needs no step changes. === PART 3: DOES THE MARGIN SURVIVE MACHINE VARIANCE === Stated explicitly, because 6.90 min against a 10 min cap is 69% of cap before any variance is applied, and the cap covers the whole job rather than the suite. worst lane, script time 414299 ms 6.90 min job overhead, measured on the real lane ~18 s (see below) expected healthy job ~432300 ms 7.21 min x1.29 on the script time, plus overhead ~552400 ms 9.21 min cap 600000 ms 10.00 min room left after the multiplication ~47.6 s 7.9% of cap The 1.29x is the runner variance measured today on the SIBLING SERIAL lane, as supplied; it is not this lane's own figure. This lane family does have its own, and it is tighter: the six full lane-2 sums today span 192939 ms to 203451 ms, a spread of 1.054x. At that figure the worst lane lands near 7.58 min with about 2.4 min of room. I have used the LARGER, borrowed 1.29x for the verdict rather than the tighter one this lane actually shows, and note that the hints are already per-script maxima, so 1.29x on top is conservative twice over. THE MARGIN SURVIVES THE MULTIPLICATION, so this proceeds rather than stopping. The 18 s overhead is measured, not assumed: in run 34462758357 the lane-1 job ran 10 min 16 s against a 598.2 s suite, and lane 2 ran 3 min 21 s against a 192.9 s suite, a ~10 s difference that matches lane 1's extra Pi package install. The cap is unchanged, the lane count is unchanged, and nothing in the serial lane, its shard count, its guard or its hint table is touched. === PART 4: THE RECORDED FACT === The workflow comment no longer restates the shard wall as a literal, which is how "~1 min of serial sum" survived a 10x change without announcing it. It now points at bin/fm-test-run.sh --check-coverage, which prints parallel_max_ms, parallel_imbalance_ms and parallel_unhinted derived from the hint table, so the current number is computed on demand. The shard documentation carries the dated run ids, which route it was taken by, and the local cross-check that shows why local numbers are not admissible as hints. Two regressions pin what rotted: lane membership must be stored longest-measured-first, and the lanes must be fully hinted and packed within 5% of each other. Both were run against the old composition and both fail on it (420030 ms imbalance against a 624299 ms worst lane). The ordering assertion they replace named a specific script by hand and had itself gone stale. === PART 5: NAMED AND LEFT, OUTSIDE THIS REBALANCE === tests/fm-captain-hold-lifecycle.test.sh alone is 296481 ms, 36% of the whole set, so it is the floor of any two-lane split: no repacking can put a lane below it. After this rebalance the cap is about 1.45x the healthy lane where the sibling serial lane keeps roughly 2x. Nothing refuses a stale parallel hint the way PORTABLE_SERIAL_MAX_UNHINTED_PERCENT bounds the serial lane. parallel_unhinted is reported, not enforced, which is what let this drift for three weeks unnoticed. * fix(review): Restrict parallel scheduling hints to portable parallel lanes * fix(document): Clarify parallel lane scheduling and timing evidence
… PR (kunchenguid#4148) pr_for_task fell back to scraping the whole status log with tail -1, so any PR URL a worker ever mentioned in prose - including a scout citing someone else's PR - became the task's delivered PR in the parent-channel terminal report. Recorded meta pr= is now the only authoritative source, the fallback scrape accepts only a preferred terminal line in a mode's ready-signal shape (done: PR <url> or done: PR <url> checks green), and a scout never carries pr= at all.
…claims instead of counting a dead drop as started (kunchenguid#4212) * fix(procevent): stop a dead runner owning a source and reconcile reporting it The captain answered ten calls on a bearings board, the board accepted them, and nothing collected them. He had to answer all ten again in chat. A surface that presents as armed while being a dead drop is worse than one that visibly fails, because the answers looked recorded. Two independent defects, reproduced together in an isolated home where reconcile reports started=1 on every run while ownership never moves and no runner ever attaches. 1. reconcile counted a launch it never verified. detach_runner is fire-and-forget and discards the child's stderr, so a runner that died before it could claim was counted exactly like one that is listening. Launches are now confirmed - the source observed owned, or its runner record moved - before being reported as started; the rest are reported as failed= with a non-zero exit. The runner-record clause is what keeps a fast-completing source from being reported as a failure when it finished between two polls. One bounded window covers a whole cycle's launches, so a home full of broken sources costs the same wait as one. 2. A claim whose whole generation is provably gone could be refused forever. Reclaiming it ran cleanups over that dead generation's own leftovers, and any failure vetoed the claim - permanently, because none of those conditions clears on its own. Every one of those leftovers is keyed by the dead generation's claim token and a replacement always claims a fresh one, so none can collide with what replaces it. fm_procevent_claim_capture_reservation_reclaim_locked already said this for the reservation record; the staging file and the shape check on the registry directory recorded to hold it now take the same rule. Removing the claim record itself stays a hard precondition: two owners is the one outcome worse than none. Two smaller repairs to the same "registered is not listening" confusion: - `list` reported OWNER=none for a source nothing can claim. A reused PID whose process group survives reaches that state through the stale branch rather than the leaderless one, so it read as an idle source waiting to be started - the reassuring answer this surface gave while a board collected nothing. It now reports the orphaned state it shares. - reconcile relaunched into that same unclaimable state on every cycle, spawning a runner that could only die on the claim. docs/configuration.md already promised it preserves such a claim without starting a replacement; the code now does that and reports it as uncertain. This is NOT a third instance of today's two lock-identity defects (4e1bf9aa and its replayed predecessor). Those were wrong liveness predicates: a reused PID read as a live holder, then an exec'd holder read as dead. Here the predicate is right - the code correctly proves the owner dead and refuses the claim anyway, on a condition unrelated to liveness. Regression coverage, each failing on the parent commit for its own reason: - tests/fm-procevent.test.sh: a source that cannot start is reported as failed rather than started; a dead generation whose leftovers cannot be tidied no longer keeps owning its source (the parent reports a start while nothing ever runs); the existing reused-PID fixture now also asserts the orphaned listing and that no doomed relaunch is reported. - tests/fm-captain-hold-lifecycle.test.sh: a board answer reaches the keyed-answer intake through the runner end to end - durable capture, the wake, and the closed task carrying the captain's selection. This one passes on the parent, because that chain was never what broke. fm-procevent 100, fm-bearings-board 18, fm-captain-hold-lifecycle 50, fm-procevent-when 13 and fm-procevent-quota 18 pass; bin/fm-lint.sh and bin/fm-doc-audience-check.sh clean. tests/fm-extension-binding.test.sh has two failures identical on the parent commit (EACCES on package install in this sandbox) and unrelated to this change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gxgshn5jkWJ3GEYWy7vTG * no-mistakes(review): confirm reconcile launches on durable launch stamps * no-mistakes(review): announce stranded sources and refuse bad confirm windows * no-mistakes(review): announce leaderless strands, bound confirm window, fix recovery docs * no-mistakes(review): announce unconfirmed launches once per episode, qualify start reclaim * no-mistakes(review): nonce launch-failed keys, refuse bad window at arm * no-mistakes(review): state only observed launch outcome, shorten episode nonce * no-mistakes(test): assert launch-failed headline not re-delivered, allow recovery wake * no-mistakes(document): docs: cover strand and launch-failure wakes in skill trigger and verification record * no-mistakes(lint): restructure SC2015 chain into explicit if-block * test(watch-triage): fix two timing-exposed defects the pipeline found Both surfaced in the no-mistakes test step on this branch, each failing one full run of tests/fm-watch-triage.test.sh; neither was accepted as a flake to retry past. 1. The new launch-failed delivery test assumed an already-surfaced key never wakes the watcher again. That is false: a fresh watcher legitimately re-surfaces any unacknowledged queue row through its downtime-recovery path ("check: rearm-resurface"), so the assertion failed whenever a re-arm landed between its two checks. The pipeline's own fix tolerated any wake lacking the repeated key's headline; this tightens it to exactly one tolerated reason, by its exact line, with a failure message that names the expectation so a reworded path reads as "the tolerated recovery path changed" rather than as a mystery - and so nobody restores the strict silence check. The positive assertion (a fresh-suffix key is delivered under its own headline) is unchanged. 2. seed_captured_procevent_result retired its source in the gap between the runner publishing its wake and releasing its claim, so retire read the exiting runner's ownership as uncertain and refused ("cannot confirm runner identity"). The fixture and retire path pre-date this branch; the confirm window returns reconcile closer to the moment of capture, which made the gap easier to hit. The fixture now waits, bounded, for the claim release the publish promises, with the reason at the wait. Verified on this head with tasks-axi on PATH: fm-watch-triage 113/113 with no skips, fm-procevent 106/106, fm-captain-hold-lifecycle 50/50, fm-watch-arm 15/15, fm-bearings-board 18/18, fm-procevent-when 13/13, fm-procevent-quota 18/18; bin/fm-lint.sh and bin/fm-doc-audience-check.sh exit 0. First attempt, no retries. * no-mistakes(document): docs: route stranded and launch-failed wakes in skill handling --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…gistration (kunchenguid#4191) * fix(herdr): verify agent registrations at process level before trusting them Herdr keeps a Pi registration (`agent get` -> agent=pi, agent_status=idle) after the Pi process has exited to a plain shell whenever a nested interactive shell sits under the pane's top shell, which is the crew shape `treehouse get` leaves behind. The pane classifier trusted that registration alone, so `fm-control.sh <id> relaunch`, `fm-spawn.sh --relaunch`, and the crew-state recovery read all treated a shell-only pane as a live agent and refused recovery for as long as the record lived. The Herdr adapter now reads `pane process-info` plus the real process table through a shared harness-process classifier (bin/fm-agent-process-lib.sh, moved verbatim out of the tmux adapter so both backends mean the same thing by agent, shell, and other) before a registered agent counts as live. A registration over a shell-only pane is the new explicit `stale-agent` pane state, which the recovery-grade read maps to `dead`; husk detection, reclaim, presentation recovery, and session cleanup keep refusing it, so recovery reuses the pane and nothing gains close authority. A working record is verified the same way before the native busy verdict reports busy, so the recovery classifier never reports a shell-only pane as working. An unreadable process view reads unknown, trusting neither the registration nor its absence. Reproduced and measured on Herdr 0.9.0 with Pi 0.85.1 in an isolated lab; the new default-on live guard tests/fm-herdr-pi-stale-registration-live-e2e.test.sh exercises the real stale record, tests/fm-control-herdr-smoke.test.sh proves exit and relaunch through the control plane, and the portable suites pin the classifier over real processes. Fixes kunchenguid#4115. Duplicates: kunchenguid#3639, kunchenguid#3487, kunchenguid#2908, kunchenguid#3545. * no-mistakes(review): settle transient prompt helpers before trusting herdr process state * no-mistakes(review): drop stray codegraph file; read spaced comm whole in descendant walk * no-mistakes(review): untrack stray .codegraph/.gitignore * no-mistakes(review): untrack codegraph file; make spaced-path walk test discriminating * no-mistakes(review): untrack stray .codegraph/.gitignore * no-mistakes(review): untrack stray .codegraph/.gitignore re-added by fix round * no-mistakes(review): untrack stray .codegraph/.gitignore * no-mistakes(review): untrack codegraph file, drop dead control case, record process-info floor * no-mistakes(review): refuse stale-agent on fresh herdr spawn preflight Documented non-goal: fresh-spawn, reclaim, and presentation-recovery auto-recovery for a stale-agent pane is a separate design change, out of scope here, to be proposed upstream as its own issue if wanted. * no-mistakes(test): Fix herdr flake: don't misread transient empty foreground as unreadable * no-mistakes(document): Add fm-agent-process-lib.sh to scripts inventory * no-mistakes(fix): update remote herdr fixture to the real pane process-info shape The shared remote-secondmate herdr fixture still returned the old flat process-info body ({"result":{"process":{"name":...}}}). The process-level liveness classifier added for kunchenguid#4115 requires the real {"result":{"type":"pane_process_info","process_info":{...foreground_processes}}} shape and treated the old body as unreadable, so an already-launched remote endpoint's agent-state read failed and any relaunch attempt against it died with "remote endpoint state is unreadable; refusing duplicate launch" instead of reaching the state it was actually exercising (tests/fm-remote-secondmate-parent-binding.test.sh, tests/fm-remote-secondmate-lifecycle-e2e.test.sh). * no-mistakes(review): test: add empty-foreground regression test for herdr flake fix * no-mistakes(document): docs: register new stale-registration live-e2e test in herdr entry points
Merge the frozen canonical snapshot while preserving the fork's Copilot/Pi interfaces, native process and private-path owners, catalog runner, and CI. Share backend process recognition and keep parallel-lane timing records separate from serial scheduling weights and concurrency admission. Firstmate-Upstream-SHA: e0d269e Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The upstream Herdr pairing-fixture mapping returned before the fork's reference expansion, omitting secondmate coverage. Keep curated families and add reference-derived shared-fixture consumers without hard-coding another family or dropping the upstream backend mapping. Extend the named regression to retain unrelated-family exclusion, unreferenced curated mappings, and unmapped-fixture refusal. Document the additive routing contract. The reported CI assertion was reproduced twice locally before the fix; the strengthened case, three adjacent routing cases, and all four repository gates pass afterward. Firstmate-Upstream-SHA: e0d269e Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CI after fixture-routing fixObserved 2026-09-11T17:14:38.979Z, exact head The original 16 successful, 1 pending, 0 failed, 1 absent across 18 expected checks. No previous-head success is reused.
The timing aggregate is not yet reported and depends on the portable/Herdr producers. Skipped or cancelled required checks would not count as passes. Live/optional/manual coverage remains limited as described in the PR body. GitHub reports mergeable= Local evidence: 815428 ms charged to the shared 2400-second budget, 1584 seconds remaining, zero timeouts. The four post-fix named cases and all four local gates passed; the 175-script selection, 46-path scope, and executable modes were verified. Full cross-platform evidence remains owned by GitHub Actions. |
Frozen reconciliation
Normal merge of four canonical commits, preserving the fork's architecture and compatibility behavior. Use a merge commit, not squash/rebase. Merge commits are enabled; this ordinary non-draft PR stays unmerged, with auto-merge disabled. GitHub Actions owns the complete cross-platform matrix; this is not yet a merge-readiness claim.
ff61baecb7df70222a1bf50271105a84878f69a54768e98d469b7216569acf6e7a5cc696ecd15c2ee0d269e07318a5a80069ab6193b4a0af4c077a61bdf50f25d5020c8944d3d60984cca5e926369ef322e143cbe049052dc8c4afca1c5afb8cb7de37a7d70975fc67292372190f542be82c7d7d2f9144aeFirstmate-Upstream-SHA: e0d269e
Upstream was fetched exactly once. The prior upstream is both frozen heads' merge base. Prior merge
e855c91e49ca1003d2668f9fbffdb5d6722f0701has parents31bb375d24cb4dcc8a3402ce7f15d7b2e02c7826and that prior upstream; retained head7def92706f2d302d5dce6493f7234edaad3ca173and PR #46 preserve its trailer. PR #46 merged September 11, 2026 at 01:43:08 UTC, producing the frozen fork. Merge commitbdf50f25d5020c8944d3d60984cca5e926369ef3has the frozen fork and upstream as its two parents. The published head adds the CI correction below without rewriting that merge. Git's trailer parser, committed tree, exact scope, clean worktree, and unchanged localmainwere verified. No reconstruction or ancestry anchor was needed.Integration decisions
Incoming: kunchenguid#4151 (lane timing/packing), kunchenguid#4148 (delivered-PR provenance), kunchenguid#4212 (confirmed process-event launches and notifications), and kunchenguid#4191 (Herdr process-level liveness).
bin/backends/tmux.shcopilot/copilot.exerecognition and dependency-error propagation in the shared owner and both consumers. No duplicate classifier.bin/fm-test-run.shparallel-durationrecords; upstream memberships and metrics remain. Only explicit parallel lanes use parallel weights.tests/fm-test-run.test.shdocs/tmux-backend.mdLoader, override/cache tests, fixture projection, routes, and docs changed together. Shared-classifier routing covers both backends, Orca, live identity gates, and the native Treehouse contract through an exact fork override. Empty hint streams preserve serial fallback and accurate parallel coverage; missing parallel hints do not acquire a new coverage-guard failure policy.
Process-event starts require bounded confirmation; failed launches and stranded claims get distinct deduplicated notifications. Generation checks, ambiguous-group refusal, native privacy, and the timeout owner remain. Delivered PRs require recorded metadata or an exact terminal ready line, never arbitrary prose or scout attribution. Herdr
stale-agentpermits pane-reusing recovery, not closing, replacement, or duplicate spawning; unreadable liveness evidence refuses. Earlier direct-child Pi documentation now explicitly excludes the newer nested-shell case, preserving both dated measurements and anchors.Verification ownership
Shared CI changed only comments versus the fork: executable jobs, permissions, triggers, matrices, timeouts, prerequisites, and artifacts are unchanged. Fork CI, proof admission, and timeout code are byte-identical. Both parallel membership functions and all 24 hints match frozen upstream; every prior core/fork serial hint is unchanged.
Retained: two parallel lanes, five serial shards, Pi setup in parallel lane 1, shared portable Pi installation policy, TypeScript
5.9.3, tasks-axi0.2.5; real Herdr0.7.4/ protocol 16, Treehouse2.0.1, Pi0.84.3; fork package Pi0.84.3, OpenCode1.18.23, TypeScript5.9.3. Both workflows retain main push/PR triggers, read-only permissions, and workflow-qualified concurrency. Portable/Herdr timing uploads still feed the dependent aggregate. The Windows Herdr experiment remains manual-only.The full Herdr response shape was measured upstream on 0.9.0; older-client subcommand presence is not server-response proof. Exact-head real-Herdr CI must establish pinned compatibility. Native transport fakes do not prove full Windows Herdr liveness. Existing live/optional gates remain limitations, not empirical passes. No live Firstmate session, vendor prompt, fleet action, or operational-record change occurred.
Local evidence
Recorded shell:
C:\Program Files\Git\bin\bash.exe, never ambient WSL Bash;FM_LIVE=0throughout. Probes passed: Git2.55.0.windows.5, Bash5.3.15, awk5.4.1, cygpath3.6.10, Node24.19.0, ShellCheck0.11.0, actionlint1.7.12, jq1.8.2, Python3.13.15, Perl5.42.3. No installs, disabled signing/hooks, or relaxation ofsafe.bareRepository=explicit.Initial snapshot: All four repository gates and 13 focused behavioral invocations passed. Two POSIX-only Herdr cases were deferred before execution because
ps -axo pid=,ppid=,comm= >/dev/nullis unsupported here; Linux parallel 2 owns both. Before publication there was no executed test failure, timeout, interruption, retry, differential run, or cached pass. Only explanatory runtime-backend prose changed after those behavioral runs; the docs gate and complete lane/routing inventory were refreshed. Post-publication CI exposed the separate routing regression documented below; those failed reproductions are not counted as passes.One 2,400-second budget: 815428 ms charged / 1584 s remaining, zero timeouts; caching disabled throughout. Invocation totals: preflight 5375 ms, validation-01 169595 ms, validation-02 373873 ms, validation-03 50801 ms, validation-ci-01 14213 ms, validation-ci-02 11000 ms, validation-ci-03 190571 ms. Probe/controller overhead is included.
Probes were the listed tools'
--versioncommands (Perl used-eto print$^V). Read-only Git policy probes checkedcommit.gpgsign,core.hooksPath, andsafe.bareRepository(initially 903 ms, repeated without changing policy for CI diagnosis). Preflight-only exited 75 without running tests; validation-02 exited 75 solely for the two prerequisite deferrals. Other initial validation invocations exited 0. CI diagnosis invocations 01 and 02 exited 1 on the reproduced assertion; invocation 03 exited 0. No validator is left running.Gate recipes:
Sis the syntax sweep below;Iisbin/fm-test-run.sh --list --changed --base ff61baecb7df70222a1bf50271105a84878f69a5, plus--list-lanesand--list --lanefor both parallel, all five numbered serial, and real-Herdr lanes.I-shortis the changed inventory plus assertions that Treehouse and the new stale-registration script are selected. Outputs were saved outside the checkout.X1is source-awarebin/fm-lint.shwith these explicit roots:bin/fm-agent-process-lib.sh bin/fm-test-catalog-lib.sh bin/fm-test-run.sh bin/backends/tmux.sh bin/backends/herdr.sh tests/fm-test-catalog.test.sh tests/fm-harness-contract.test.sh.X2addstests/fm-test-run.test.sh tests/fm-backend-herdr.test.sh.Initial behavior commands:
bin/fm-test-run.sh --jobs 1 tests/<script>.test.sh; setFM_TEST_ONLY=<case>for named rows, omit it forFULL. Every executed row below passed withgate_skip=false; the two deferred rows were not executed. Common Bash path andFM_LIVE=0are given above. These results remain attributed to their original snapshots, not relabeled as final-head CI passes.Coverage: 210 = 24 parallel + 171 serial + 15 Herdr. All parallel members have hints; max packed sum 414299 ms, imbalance 30 ms, with the named case enforcing at most 5%. Estimates are not CI wall-time/headroom proof. Docs: 103 surfaces / 422 local links.
Post-publication CI correction
Behavior portable parallel 1, job103337702426on initial headbdf50f25d5020c8944d3d60984cca5e926369ef3, failed because changingherdr-client-pair-fixture.shomittedfm-secondmate-safety.test.sh. The new curated Herdr route returned before the fork's reference expansion. Both references and the secondmate registration were intact.Follow-up
22e143cbe049052dc8c4afca1c5afb8cb7de37a7keeps the upstream mapping and adds reference-derived consumer families for shared fixtures/helpers. It changes only the runner, its existing named regression, and two fork guides within the same 46-path scope. The regression also protects curated backend coverage, unrelated-family exclusion, curated fixtures with no references, and refusal of fixtures with neither mapping nor consumers. No timeout, dependency, package pin, lane membership, or workflow change was needed.For the named rows, the exact recipe is
FM_TEST_ONLY=<case> bin/fm-test-run.sh --jobs 1 tests/fm-test-run.test.sh, under the recorded Bash withFM_LIVE=0. The original case failed on the initial committed tree; the strengthened case failed on test-only treed61262c63a3721adf462d0b085c4643845bc8d2fbefore the production fix. All post-fix rows use committed treed70975fc67292372190f542be82c7d7d2f9144ae.The failure reproduced on both Linux CI and local Windows; no platform-baseline claim or frozen-upstream differential was needed. No debug instrumentation was added. The original normal merge remains reachable. Exact-head CI was restarted by the ordinary push; old-head successes are not proof for this head.
No-renames audit
Scope: 38 incoming paths + 8 explicitly recorded integration paths = 46. The initial merge's intended/staged/committed NUL path sets match exactly. The follow-up staged exactly four paths within that manifest; the final PR still has the identical 46-path scope. All modes match the fork or new upstream file. No broad renormalization. The two PR additions are upstream files, not additional fork modules. The before-comparison's two extra D entries are those not-yet-adopted upstream additions.
A smaller count is not behavior proof: persistent divergence against the prior sync grows 204 -> 205 paths, including the shared-classifier compatibility hunk; additive fork paths stay 52. Inherited deletions stay
.claude/settings.json,.github/workflows/no-mistakes-required.yml, andtests/fm-no-mistakes-required.test.sh. No rename manufactures equivalence.Complete before/after inventories are retained outside the checkout and reproducible with:
Exactly upstream-equivalent now (15, including the new live guard):
.agents/skills/process-event-sources/SKILL.md,bin/fm-backend.sh,bin/fm-inactive-reconcile.sh,docs/verification/process-event-sources.md,docs/verification/rovo.md,tests/fm-control-herdr-smoke.test.sh,tests/fm-crew-state.test.sh,tests/fm-cursor-harness.test.sh,tests/fm-herdr-pi-stale-registration-live-e2e.test.sh,tests/fm-omp-harness.test.sh,tests/fm-tmux-agent-liveness.test.sh,tests/fm-watch-arm.test.sh,tests/fm-watch-triage.test.sh,tests/herdr-client-pair-fixture.sh,tests/remote-herdr-fixture.sh.Classifier code moved to the shared owner, and parallel timing data to catalogs; no duplicate classifier or inline timing table remains. Scheduling, proof admission, and backend transactions intentionally remain separate. Retained exceptions/removal conditions (full ownership in
docs/fork/architecture.md):All 175 selected scripts and their CI owners
Names below expand to
tests/<name>.test.sh. Default local disposition: full regression deferred to the named CI job.FULL= full script passed locally;CASE= only tabled named cases passed.LIVE,HERDR,OPTIONALidentify unchanged capability gates, not empirical passes. Procevent/delivered-PR full suites belong to serial 3; other watcher, stop-proof and lifecycle subjects are explicitly assigned below. Real mutations/cleanup need the Herdr producer; stock Bash needs macOS; native Windows and installed packages need Fork CI.Complete per-script routing (175)
Behavior portable parallel 1 (9)
Behavior portable parallel 2 (12)
Behavior portable serial 1 (27)
Behavior portable serial 2 (29)
Behavior portable serial 3 (30)
Behavior portable serial 4 (25)
Behavior portable serial 5 (28)
Behavior tests (Herdr) (15)
Exact-head CI inventory
Each expected name has one verified producer. The post-publication CI comment records successful/pending/failed/absent states for this exact head; skipped or cancelled required lanes are not proof. The manual Windows Herdr experiment is separate.