Reconcile upstream d499323 (September 14, 2026) - #52
Conversation
* Bind GitHub merges to a live green head and require an away-task grant. A GitHub merge now re-reads the pull request and passes --match-head-commit, so a red or moved head cannot land the way GitLab already refused. While an away record exists, only yolo or a named grant may merge, so hold-for-return cannot ship an ungated PR. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(review): Harden away merge authorization and grant parsing * no-mistakes(review): Restrict fallback outcomes to proved GitHub merges * no-mistakes(document): Refresh merge safety documentation * no-mistakes(ci): Fixed all three CI failures by updating legacy GitHub merge fixtures for live-head verification/direct gh merges and removing a process-event runner cleanup race. Verified fm-pr-check-security, fm-captain-hold-lifecycle, and fm-watch-triage pass locally; shell syntax and git diff checks also pass * no-mistakes(document): Document attended red-check exception --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…m epoch (kunchenguid#4221) The --claude guard's re-block budget charged the auto-arm ledger epoch, not the re-block: `budget_account_current_epoch` advanced the session count only when `state/.claude-autoarm-epoch` named a different generation than the previous accounting. The epoch advances only inside the auto-arm hook's generation claim, so a hook kept inert before that claim - a session lock held by a live harness outside its ancestry, a hook that never fires, or an identity or write failure ahead of `fm_autoarm_claim_next` - left the ledger frozen at its last outcome and the count frozen with it. Reproduced in a fixture: twelve consecutive Stops re-blocked with the count at 0 and the attended fail-open never fired, leaving only Claude's silent 8-block override, the blind end the bounded alarm exists to prevent. The budget now charges a re-block against an epoch the previous re-block already charged, while still charging each epoch at most once per Stop so the wait loop's repeated observations of one fresh terminal outcome and the same invocation's block decision cannot double count. The advancing-epoch progression is unchanged: three re-blocks, then one attended fail-open for a verified failure episode, and a frozen epoch now follows the same shape. Budget exhaustion without a verified failure still blocks, by the existing contract, and positive watcher recovery still clears the whole episode. Regression coverage drives the real auto-arm hook against a foreign session lock holder, asserts the ledger itself stays frozen, and fails before the fix in both the verified and unverified shapes; the existing unverified budget test now proves its budget actually ran out.
…enguid#4242) * feat(herdr): attach a real foreground viewer so the live-client teardown cases can be driven PR kunchenguid#4131 gated the Herdr active-tab close refusal on a live foreground client instead of the persisted `.focused` pointer, but only its two detached scenarios could be validated live. Every pseudo-terminal the runner built started at a zero-sized window grid, so Herdr registered no foreground client and `terminal title clear` kept answering `no_foreground_client`, leaving the four attached-client scenarios untested. That was a harness limit, not a product one. Add `fm-herdr-lab.sh viewer start|stop <session>`, backed by `bin/fm-herdr-lab-viewer.py`. The launcher sets the pty window size on the master fd BEFORE the fork, so the TUI cannot read the grid until it is already non-zero, and scrubs the inherited `HERDR_*` variables so Herdr's nested-viewer refusal does not fire when the helper runs inside one of its own panes. Attach and detach are both confirmed against the session's own foreground-client reason rather than assumed from a signal. The viewer inherits the lab's isolation contract: it attaches only to a session carrying this lab's ownership tripwire, never to `default`, and it signals only the processes it recorded, so a client someone else attached is never touched. Teardown now refuses while an owned viewer is still attached. Turn the reproduction into the regression with `tests/fm-herdr-attached-viewer-live-e2e.test.sh`, which drives kunchenguid#4131's scenarios 3, 4, 5, and 7 live against real Herdr and asserts the close refusal fires. Scenarios 4 and 5 need a focus change at one exact product boundary, so a PATH shim performs the real `tab focus` when the close helper issues its planning `pane get`. Removing either half of the recipe from the launcher makes the guard fail with the same `no_foreground_client` symptom kunchenguid#4131 reported. * test(herdr): fail loudly when an attached-viewer fixture cannot be created The fixture helpers run inside command substitutions, where fail() exits only the subshell and leaves the script running with empty ids. Return non-zero instead and carry the message at each call site. * fix(herdr): stop the viewer launcher's kill timer from raising on an exited child The SIGALRM escalation called os.kill unguarded, so a viewer that exited during the grace window turned an ordinary shutdown into a traceback inside the signal handler. * docs: list the lab viewer's pty engine in the bin toolbelt * no-mistakes(review): Harden Herdr viewer ownership and live CI coverage * no-mistakes(review): Validate viewer startup timeout and process ownership * no-mistakes(document): Document Herdr viewer safety contracts * no-mistakes(review): Fix viewer timeout to two seconds * no-mistakes(review): Cancel timed-out viewers and fix PTY grid * no-mistakes(review): Serialize viewer transitions and verify process parentage * no-mistakes(review): Harden viewer ownership locks and deduplicate CI * no-mistakes(review): Release interrupted locks and preserve viewer escalation * no-mistakes(review): Remove viewer locks and cancel interrupted launches * no-mistakes(review): Close viewer launch signal races * no-mistakes(document): Document attached Herdr viewer regression
…unchenguid#3766) * fix(bootstrap): allow nonvisual work without Lavish * no-mistakes(review): Gate scout brief Lavish line on bootstrap version floor
…kunchenguid#3825) * fix(tests): isolate fixture Git configuration from host preferences Ignore global and system Git configuration in the shared test library, which all four fixture helper entry points source. Keep local config, command-line overrides and explicitly supplied test config usable without changing the caller's environment or real project signing preferences. Exercise global and system signing inputs through all four helpers, real fixture and child commits, explicit signing overrides, unchanged input files, and signing refusal outside fixture subprocesses. Verification evidence for issue kunchenguid#3770: On pristine upstream f09de8a, all 12 reported suites failed and each logged "No secret key" using a private GIT_CONFIG_GLOBAL containing commit.gpgsign=true and gpg.format=openpgp, GIT_CONFIG_NOSYSTEM=1, and an empty private GNUPGHOME (GIT_CONFIG_COUNT and GIT_CONFIG_PARAMETERS unset). With this change, all 12 pass in the identical environment through bin/fm-test-run.sh --per-script-timeout-secs 900: fm-backlog-atomicity, fm-bootstrap-network-parallel, fm-bootstrap, fm-crew-state, fm-fleet-sync, fm-gate-refuse, fm-grok-harness, fm-session-start, fm-sessionstart-nudge, fm-tangle-guard, fm-test-run, and fm-update (all tests/<name>.test.sh). The new fm-test-fixtures regression failed before the library change and passes after it. Canonical bin/fm-lint.sh passes. Additional verification exposed fm-teardown's herdr-preflight-missing-adapter assertion on both this branch and an unchanged f09de8a archive with signing neutralized. That pre-existing failure needs separate disposition; it is not repaired or skipped here. The separately owned Muse and composer fixture defects remain untouched. Fixes kunchenguid#3770 * no-mistakes(review): Complete fixture Git isolation and scope config assertions * no-mistakes(review): Share Git isolation across standalone fixture entry points * no-mistakes(review): Map git-config helper changes to lib.sh dependents * no-mistakes(review): Select fixture-isolation regression on runner change; halve config matrix * no-mistakes(review): Scope fixture-isolation regression selection to the runner alone * no-mistakes(document): Give fixture Git isolation helper its owning header * no-mistakes(document): Record fixture Git-isolation coverage in fixtures suite header * no-mistakes(review): Fix linked-worktree fixtures and remove redundant Git isolation * no-mistakes(document): Correct stale runner-selection documentation * no-mistakes(document): Clarify family antecedent in isolation-proof runner evidence
… in auto mode (kunchenguid#4239) * feat(spawn): add config/claude-permission-mode to launch Claude workers in auto mode Every Claude worker launched with --dangerously-skip-permissions, and a captain who refuses bypass mode had no way to select Claude Code's classifier-reviewed auto mode instead. A new one-token local config, config/claude-permission-mode, selects the permission flag for every Claude launch: absent or `bypass` keeps today's launch byte-for-byte, `auto` swaps in --permission-mode auto, and any other value refuses the spawn before any endpoint, worktree, or record exists and names the accepted values. fm-spawn resolves the file on every spawn and relaunch, threads the flag through the Claude launch template for crewmates, scouts, and secondmates alike, and records claude_permission_mode=auto in the task meta only under auto so the default meta stays unchanged; a relaunch re-resolves rather than preserving the line. The file is a captain-wide safety preference, so it joins the inherited local material pushed into secondmate homes. The Claude adapter reference records the verified auto launch shape on Claude Code 2.1.269 and that it never meets the once-per-machine bypass confirmation dialog; docs/configuration.md owns the schema. * no-mistakes(review): drop unread claude_permission_mode meta line and its assertions
… untouched (kunchenguid#4243) * fix(teardown): refuse to return a Treehouse pool slot reassigned to another task A pool slot is reused across tasks, so a finished task's worktree= line can name a slot a different, live task now holds. Teardown already refused when a second task record named the same live path, but that scan cannot prove the record it is tearing down is the current owner: the task that took the slot next may leave no record the scan can reach - its own worker may have exited and its record been cleaned up, or it may live in a home this machine does not register. Teardown then killed every process under the path, hard-reset it and returned it, and its unlanded-work refusal never fired because it was inspecting a directory that no longer belonged to the task being torn down (observed 2026-09-07). Treehouse's own state file cannot answer the ownership question. It records a slot's owner as a live process lease (owner_pid plus owner_started_at, with `treehouse status` reporting in-use from the processes actually running under the path), which names no task and is released by the very event that makes a record stale - the worker exiting. An unleased slot therefore reads identical whether it is still this task's or has since been handed on, and a slot whose new holder has also exited but left uncommitted work reads as free. So the identity source is Firstmate's own claim, not Treehouse's lease. fm-spawn writes that claim - the task id - into the slot at the moment it takes it, under the same project lock that allocates the slot, and fm-teardown drops it only after the slot is genuinely returned. It lives at <pool>/<slot>/.fm-slot-owner, a sibling of the repo checkout rather than a file inside it, so claiming a slot can never dirty the copy the landed-work checks inspect. A claim naming another task, or one that cannot be read, refuses; --force does not lift either refusal, because --force authorizes discarding this task's unlanded work, never another task's live work. A slot that cannot be claimed refuses the spawn instead. An absent claim proceeds on exactly the record-scan protection it had before: slots taken before claims existed, and slots already returned, carry none, and refusing those would strand every task in flight across this change on no evidence at all. The refusal is deliberately all-or-nothing rather than partially completing the task's own cleanup. state/<id>.meta is the only durable record naming the worktree and endpoint, so removing it would destroy the evidence needed to reconcile which record is wrong, and its removal is one step with the backlog transition. Nothing is stranded: clearing the stale worktree= line leaves a record with no slot to release, which then tears down normally, and the refusal names that remedy. Repairing the previous claimant's stale worktree= line at spawn time is left for separate work. It would have the new owner write another task's record - the same class of cross-task mutation this bug is - and would need that record's own meta lock; with the claim in place teardown refuses on evidence rather than depending on the stale pointer having been scrubbed. For the same reason the relaunch path writes no claim: it holds no allocation lock, and a record whose worktree= is already stale would stamp the wrong task's claim onto a live sibling's slot. The regression reproduces the reuse sequence with only one discoverable record, including a clean, fully landed ship copy torn down without --force - the shape of the real incident, which the previous code returned to the pool - and fails against the previous code; the existing two-record, cross-home, own-slot and no-claim cases still pass unchanged. This builds ON upstream b028e8b (kunchenguid#3837), which is already in this branch's base (origin/main 40c50ea) and owns the record-exclusivity scan. Nothing here replaces that scan; the claim is the positive proof it cannot supply. Claude-Session: https://claude.ai/code/session_01JTBmuqKugaPUj7k9TXQwFS * no-mistakes(review): teardown leaves reassigned slot; spawn abort drops claim * no-mistakes(review): narrow Treehouse lease evidence; gate abort claim release on lock * no-mistakes(review): pin spawn-side slot claim; narrow abort-release header * no-mistakes(document): docs: point slot-claim rationale at fm-wake-lib owner
…nguid#4247) The reviewer treats Captain's intent as acceptance criteria, so a widened ask there drives over-built work; the spec should carry only what the ask requires.
* feat(bearings): name the Underway rows and order Charted Next newest filed first The fleet board's Underway rows led with the run status alone, so a scan told the captain where a pipeline stood but never which task the row was, and Charted Next rendered in backlog order rather than by when work was filed. The snapshot now projects the durable task name onto every in_flight row - from this home's backlog title, and from a secondmate home's own ledger for an active child - and the durable filed date onto every gate. The board's Underway row leads with that name and keeps the run status on its second line, and Charted Next renders newest filed first, with rows carrying no comparable date keeping their payload order after every dated row. The payload validator requires an explicit name marker on every Underway row and refuses a filed value that is not an ISO date, so the board can never sort on garbage or invent a label. * no-mistakes(review): Fix Bearings labels, bounds, and filed validation * no-mistakes(review): Fix Bearings identifiers and eligible queue bounds * no-mistakes(document): Document Bearings labels and newest-first bounds * no-mistakes(ci): Updated the stock macOS Bash CI expectation from 56 to 59 Bearings tests. Verified the suite under /bin/bash 3.2: all 59 tests pass. git diff --check also passes
…nchenguid#4248) * fix(bearings): report the away-return catch-up instead of refusing A captain returning from away and asking for bearings got zero bytes and an error: fm-bearings-snapshot.sh ran the away-return guard with `|| exit $?` before reading any fleet state, so the mere existence of the catch-up gate killed every bearings mode (and /ahoy with them). Bearings now consults that guard rather than obeying it. fm-afk-return.sh separates its two refusal branches by exit status, so an ACTIVE away window still refuses exactly as before - the right answer there is to run the return first - while return catch-up (exit 4) lets collection and projection proceed and is disclosed as one action-free `(return-catchup)` gate row, following the existing `(main-inventory)` precedent. It stays out of decisions_open: these blockers are firstmate-actionable, not the captain's own call, and the per-task blockers already project as their own Underway rows. The guard's refusal text also stops promising a blocker list it cannot produce: a gate retained for a lifecycle reason alone now names that retention reason, and bearings carries the same reason in the gate row's title. Reporting is not ordinary work. AGENTS.md already scopes the return hold to work rather than reporting, so only the /afk and bearings skills needed the correction. * no-mistakes(document): Refresh away-return Bearings verification * no-mistakes(review): Reserve catch-up gate outside Bearings truncation * no-mistakes(review): Preserve filed dates in catch-up gate output * no-mistakes(document): Document reserved catch-up gate projection
…forked code-root copy (kunchenguid#4223) * fix(backlog): address the home's backlog from any directory and detect a forked code-root copy A home outside the code root forks its queue: the tracked .tasks.toml names data/backlog.md relative to tasks-axi's working directory, so a bare tasks-axi call from the code root writes the code root's data/ while session start, spawn, and teardown use $FM_HOME/data. Linking the code-root copy into the home does not hold, because tasks-axi 0.2.4 writes by renaming a temp file over its target and rename(2) replaces a symlink: add, start, hold, and done from the code root each turn the link back into a regular file. The archive path is resolved against the working directory too, even with --file. bin/fm-tasks-axi.sh runs tasks-axi against this home's backlog from any directory, using the lifecycle transitions' existing addressing (run from the data directory's parent, pin <data>/backlog.md through TASKS_AXI_FILE). It keeps relative --to/--*-file arguments meaning the caller's paths, and refuses a caller --file, an unresolvable home, and a symlinked home backlog. The fm-send hold lookup, fm-public-followup, and the fm-decision-hold shim, which relied on cwd discovery, now go through it with an explicit FM_HOME and a cleared data override, so they keep addressing exactly $FM_HOME/data and an ambient TASKS_AXI_FILE cannot divert them; every agent-facing backlog command names it instead of bare tasks-axi. Bootstrap gains a detect-only BACKLOG_RECONCILE check, also run read-only: when the home's data directory is not the code root's, a code-root data/backlog.md or data/done-archive.md that is not the home's own file is reported as a fork, with the merge procedure in bootstrap-diagnostics. * test(teardown): assert the completion hint names bin/fm-tasks-axi.sh ready The completion hint now points at the home-addressed command instead of a bare tasks-axi call, so the dependency-cleared follow-up assertion checks for that command. * no-mistakes(test): clear ambient tasks-axi env in tests/lib.sh * no-mistakes(document): drop bare tasks-axi example from cd-guard doc * no-mistakes(lint): replace ls -A decoy listing with find for SC2012 * no-mistakes: apply CI fixes * revert: keep the compliance gate unchanged; the synchronize race is filed separately
* fix(spawn): pre-register Claude workspace trust for secondmate homes A claude --secondmate launch skipped workspace-trust registration entirely, so a standalone-clone secondmate home (an explicit ~/fm-homes/<id> path) had no store entry and its pane wedged on the "Is this a project you trust?" dialog before it read its charter. The step was gated on the task kind rather than on the harness, so the spawn's fail-closed guard had nothing to run against and reported a launch that could never start work. fm-claude-trust.sh gains a secondmate-home mode. A secondmate home is a whole firstmate instance, produced either as a leased worktree or as a standalone clone, so the linked-worktree test cannot decide it and the seed is the evidence instead: the .fm-secondmate-home marker must be a regular file this user owns naming exactly the id being spawned, the home must hold AGENTS.md and bin/, and each operational directory must resolve inside the home. That is the set fm-home-seed.sh writes and fm-spawn.sh's own home validation re-checks, so nothing wider than a home a secondmate spawn would launch into can earn home-level trust. The worktree path is unchanged, and still refuses a home. fm-spawn.sh now runs the registration for every claude launch and keeps refusing the spawn when it fails, rather than launching an agent that would wedge. * no-mistakes(document): Correct Claude secondmate trust guidance
* fix(pr-merge): judge each required check by its current run When the base branch advances, GitHub cancels a pull request's in-flight run and re-triggers it. The cancelled run stays in statusCheckRollup beside the passing re-run, so the rollup can hold several runs of one check name at the same head while GitHub itself reports the pull request CLEAN. github_checks_not_green judged every run independently, so that superseded failure refused a genuinely mergeable pull request and pushed the operator toward a needless --allow-red. Group the rollup by the reported name and judge each check by its current run. Supersession is proven, never assumed: a name leaves the red set only when every one of its non-green runs is strictly older than one of its green runs, dated by the forge's own settled timestamp - a check run's completedAt once its status is COMPLETED, or a status context's createdAt - and only in the whole-second UTC form GitHub emits, which is the one spelling that orders correctly as plain text. A run with no such timestamp is never superseded, so a still-running, queued or undated run keeps its check red, and a name with no green run at all stays red. An unnamed entry is grouped alone so two unrelated unnamed checks are never treated as one. Every comparison is one-directional: it can only clear a failure a later success provably replaced, and never clears a check whose current run failed, is pending, or is missing. No other guard moves - the pull request must still be open, undrafted, mergeable, conflict-free and head-bound, and --allow-red still waives exactly its named check with every other check green. Live reproduction: PR kunchenguid#4224 read CLEAN with an old FAILURE and a newer SUCCESS for one check name and was refused; it now verifies, while kunchenguid#4208 and kunchenguid#4210, whose latest runs failed, still refuse. * no-mistakes(review): Use check-run start times for safe supersession * no-mistakes(document): Clarify GitHub check-rollup documentation
…guid#4266) * fix(merge): persist the merge authority on poll-detected merge outcomes The merge ledger tags a merge with the authority that permitted it while the away-posture record existed, but only the direct attended merge in bin/fm-pr-merge.sh recorded it. A merge the forge queued, or one the merge poll detected after the fact, published an untagged row, so exactly the merges no agent watched were the least auditable. bin/fm-merge-authority-lib.sh now owns that answer, read from the same structured sources the merge gate already used: the task's recorded yolo posture and the away-posture record's mechanical grant list, never prose. bin/fm-pr-merge.sh keeps its own refusal wording and gates on that answer; bin/fm-watch.sh only records it on the row its poll publishes, so reading the authority never becomes a second path to a merge. An unresolved answer records an untagged row rather than dropping the outcome or inventing an authority. * no-mistakes(review): Persist canonical merge authority for queued poll outcomes * no-mistakes(review): Harden merge authority persistence against lifecycle races * no-mistakes(review): Serialize poll authority publication with teardown * no-mistakes(document): Clarify persisted merge authority lifecycle * no-mistakes(ci): Added targeted SC2034 suppressions for the two public result assignments in bin/fm-merge-authority-lib.sh. Verified successfully with `CI=true bin/fm-lint.sh`
…4281) The 2026-09-12 Actions starvation incident found firstmate CI with no concurrency deduplication, so every superseded PR head kept its full 13-job fan-out, and four jobs with no timeout at all. Add per-PR supersession keyed on the PR number for pull_request events and on the unique run id for push events, cancelling only pull_request runs, so a new PR head replaces its own in-flight CI while every main push keeps its own group and is never cancelled. Add hang tripwires to the four previously unbounded jobs: 25 minutes for lint (measured at 14-16 minutes) and 5 minutes each for the coverage guard, the timing aggregate, and the repo invariants. Measured lane bounds are unchanged. tests/fm-ci-workflow.test.sh resolves the workflow's concurrency expressions against simulated pull_request and push contexts and holds every job's finite timeout.
…henguid#4288) Every other make_hold_home caller in this file skips when tasks-axi is absent; this test was the one unguarded call, so hosts without tasks-axi hard-fail the fixture build instead of skipping.
…nnot blind a session start (kunchenguid#4027) * fix(bin): bound each backlog row read so one wedged backend cannot blind a session start bin/fm-bootstrap.sh's reconcile and close-replay sweeps read the backlog backend once per item through fm_backlog_row_show, and that read was unbounded. A single wedged `tasks-axi show` therefore consumed the whole FM_SESSION_START_TIMEOUT and truncated the digest before the wake queue, supervision instructions, fleet state, and context sections ever printed, leaving the fleet unsupervised with no live watcher. The harm was a blind startup, not a slow one. Bound the read with the existing shared timeout primitive (bin/fm-timeout-lib.sh), so a wedged backend degrades to a loud partial reconcile: the sweep's existing BACKLOG_RECONCILE diagnostic names the item it could not read and the loop continues to the next one. The first bound hit also latches FM_BACKLOG_ROW_SHOW_WEDGED, so a sweep over many items pays one bound rather than one per item and still names every item it skipped, which is what keeps the digest whole on a home carrying a large fleet. The bound holds regardless of any particular tasks-axi install, so it does not depend on the 0.2.5 `show` hang being resolved separately. * fix(bin): set the wedged-backend latch where it survives, and prove it The latch added with the read bound was inert. fm_backlog_row_show runs inside a command substitution in both of its status-capturing callers, so the subshell read the inherited value correctly but its write died with the subshell. Every item still paid a full bound and reported `exceeded`, never `skipped`, which left the large-fleet case the latch existed to cover completely uncovered. Move the write to the two callers that capture the read's status and own the surviving shell, and leave fm_backlog_row_show reading the latch only. Correct the comments that claimed an ownership the function never had. The test that was supposed to cover this asserted only that the second read finished under a generous ceiling, which is true whether or not the latch works. Assert instead that a latched read is strictly faster than one bound and that it reports its own item as skipped, so an inert latch fails the test. * test: cover every item the wedged-backend latch skips The latch assertion exercised a single skipped item, so "every skipped item is still named" was inferred rather than tested. Probe three items instead and assert each skipped one names itself and costs less than a bound. Verified as a real guard by removing both latch writes: the suite then fails on the first skipped item instead of passing. * no-mistakes(review): distinguish backlog read-bound hits from absent rows * no-mistakes(review): preserve read-bound status through the captain verify gates * no-mistakes(review): Preserve backlog read-bound hits through resolve_entry and reconcile instead of spending them as absent rows * no-mistakes(review): Preserve backlog read-bound 124 through migrated-prefix scan and remaining task_show call sites * no-mistakes(document): Document bounded backlog row reads and FM_BACKLOG_ROW_TIMEOUT_SECS * no-mistakes(ci): Fixed all four failing CI checks with one root-cause fix plus one test-heredity fix. (1) bin/fm-captain-hold.sh: task_show carries the row in TASK_SHOW_OUTPUT and emits no stdout, but four call sites still used the stale command-substitution convention show=$(task_show ...), leaving show empty: task_show_or_fail (every captain hold failed with 'did not retain its hold-set stamp' - broke fm-captain-hold-lifecycle in parallel 1 and fm-bearings-board in serial 3), resolve_migrated_entry (migrated-prefix resolution could never match), reconcile-requests (existing rows were refused as absent), and command_open --identity (printed a constant '#0' identity, so fm-watch-triage's re-held captain call inherited the previous call's silence in serial 1). This is also the Greptile P1. Fixed by invoking task_show in the current shell and reading show=$TASK_SHOW_OUTPUT, the convention the other eight call sites already use; read-bound hits still stop loudly by name. (2) tests/fm-backlog-read-bound.test.sh (serial 4, unclassified family): the new e2e half implicitly relied on the author's process tree containing a harness process so fm-lock.sh would grant the fleet lock; on CI runners the lock is refused, the reconcile sweep is skipped, and the final BACKLOG_RECONCILE assertion fails. Reproduced by simulating a CI ancestry via a ps shim, fixed by pinning the lock evidence with the established fake-ps harness fixture pattern from tests/fm-session-start.test.sh. Verified: shellcheck clean; parallel-1, serial-3, and serial-4 lanes fully green locally (failed=0); serial-1 lane green except fm-gemini-harness, which fails only under local Node v26 (comm=node-MainThread); CI's default Node 22 reports comm=node, the branch that test passes on, so it is not a CI failure * no-mistakes(document): Verified bounded backlog read docs accurate across branch
…guid#4285) * fix(merge): serialize the away-authority check with a synchronous merge bin/fm-pr-merge.sh read the away-posture record for merge authority (the per-task merge grant and the yolo/away-grant decision) and handed the merge to the forge afterwards. An archive at the captain's return or a grant revoked by a replacement record could land in between, so a merge could proceed on away authority that no longer held. The away record now carries a cross-subsystem lock, built on the existing bounded lock primitive rather than a new lock format: the record-mutating subcommands hold it across their mutation, and the merge holds it across both its authority read and the forge command. Because a queued or auto merge returns before the pull request lands, and would therefore outlive the lock, an away merge is now refused whenever it could land asynchronously: a requested --auto, a base branch whose merge-queue state does not prove an immediate merge, and GitLab's asynchronous flags and configuration. What remains permitted while away is the synchronous merge that lands inside the lock. This closes the common away-record/merge race against a live lock owner. It does not make the merge atomic in every case, and two narrow races are accepted and documented at their sites rather than hidden, both confused-agent-grade in the sense bin/fm-lease-lib.sh already uses: - A merge-queue rule change or a PR base change in the window between the queue-free preflight and the forge call can still enqueue the merge, which can then land after its grant lapses. - Killing the lock-owning shell while its gh or glab child is still running lets stale-owner recovery reclaim the lock and the record be archived or replaced, after which the orphaned child can complete the merge on lapsed authority. Closing either one needs landing verification or an ownership handoff, which is deliberately out of scope here. No existing gate is relaxed. The lock is taken after the live green-at-head verify and the captain-hold check, the in-lock authority read is unchanged, and a lock that cannot be taken refuses the merge rather than proceeding unlocked. The away grant stays a structured field; no prose is parsed. * no-mistakes(review): Fix GitHub rollup fixture base branch * no-mistakes(document): Document atomic away-authority merge locking * no-mistakes(ci): Updated two executable GitHub API fixtures to include the required baseRefName. Both previously failing test suites now pass: fm-captain-hold-lifecycle.test.sh and fm-pr-check-security.test.sh. git diff --check also passes
…unchenguid#4200) * feat(agy): verify Antigravity CLI as third worker/scout adapter Detection by anchored ancestry in fm-harness.sh (no marker of its own); bootstrap harness and effort validation; launch template with model and effort mapping plus reachable-catalog model validation; rendered-tail busy fallback in fm-busy-lib.sh with delivery footer in fm-composer-lib.sh; control mechanics with crewmate/scout-only refusal; tmux liveness naming; router entry with concise adapter reference; dated verification record; portable regression plus opt-in live drift guard. Verified live on agy 1.2.0: supervised spawn, durable steering, same-copy relaunch, and exit, with Herdr-native busy agreement. * no-mistakes(review): bound agy model probe, gate trust dialog, narrow busy signature * no-mistakes(review): pre-register agy workspace trust, make readiness gate strict * no-mistakes(review): Close Orca terminal on gate failure; isolate live-guard HOME; tighten agy matching * no-mistakes(document): Document agy adapter in stale harness enumerations * no-mistakes(review): Clamp non-positive FM_AGY_MODELS_TIMEOUT to the default bound * no-mistakes(document): Fix stale test-shard snapshots after agy lane additions * no-mistakes(ci): Fixed ci-3 (tests/fm-agy-harness.test.sh:519). Root cause: the agy spawn fixture's default base PATH (/usr/bin:/bin:/usr/sbin:/sbin) omits node's directory, but the spawn drives the real bin/fm-agy-trust.sh (which hard-requires node to record trust) and the fixture's fake tmux trust lookup (node -e) under that PATH. On the ubuntu-latest CI runner node lives in the toolcache (/usr/local/bin), so trust pre-registration failed on portable serial 2; on typical Arch hosts node is in /usr/bin, masking the defect. Fix (smallest, following the existing tests/fm-kimi-harness.test.sh precedent of carrying the interpreter's resolved directory): resolve node from the invoking environment (failing the test with 'test needs node' if absent, as kimi does for python3) and prepend its directory to the fixture's default base PATH; the FM_TEST_BASE_PATH override contract is untouched. Verified locally: (1) pre-fix reproduction with a CI-shaped base PATH (system bins minus node) produced exactly the reported failure — 'node is required to record workspace trust and was not found on PATH' plus the fake tmux 'node: command not found'; (2) post-fix, all 29 tests in the file pass both with node available only via a leading non-standard dir in the base PATH (CI's shape) and with the default base PATH on this host. bash -n clean; ShellCheck is not installed in this worktree (previously recorded as environmental) * no-mistakes(test): Give agy typed sends a longer submit-confirm budget * no-mistakes(document): Document agy send budget, trust gate, and control coverage * no-mistakes(document): Document agy busy fallback inventory and send-timing evidence
…uid#4337) * feat(afk): add quiet supervision mode for a present captain Adds a first-class quiet supervision mode alongside /afk for kunchenguid#2356: the same away-mode daemon, injection, busy/composer guards, classification policy, and reliability properties, but the captain staying present and chatting no longer exits it - only an explicit /quiet off does. state/.afk's first line now declares its mode (away, the default, or quiet); fm_afk_mode() in bin/fm-wake-lib.sh is the single reader, falling back to away for missing/empty/unreadable/unrecognized content (including the legacy bare-epoch-timestamp format written before mode existed) so nothing regresses. fm_afk_flag_write() preserves the on-disk mode on a bare refresh (no explicit mode given) rather than defaulting to away, which is what keeps the daemon's own redundant terminal-side re-write from silently resetting a captain's quiet mode back to away underneath them. New .agents/skills/quiet/SKILL.md is a thin wrapper cross-referencing /afk for every shared mechanism, per the one-owner rule. AGENTS.md gains the state/.afk table entry and section 8's exit-trigger line. bin/fm-supervision-instructions.sh, bin/fm-session-start.sh, and bin/fm-guard.sh's stale-watcher banner all become mode-aware so a quiet-mode captain is never misdirected to /afk in captain-facing text. Closes kunchenguid#2356 * no-mistakes(review): Fix AFK epoch parsing and quiet-mode digest wording for two-line flag * no-mistakes(document): Fix turnend-guard.md daemon-ownership contract for quiet mode --------- Co-authored-by: NewAiCoder <claude@theinbtw.com> Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…kunchenguid#3578) * fix(bin): let verified harness ancestry outrank retained markers (#3) * fix(bin): let a structural harness ancestor outrank a retained marker bin/fm-harness.sh treated a verified environment marker as unconditionally authoritative, so a Codex session started from an environment that had retained CLAUDECODE=1 detected as claude. Session start then emitted Claude's Stop-owned supervision protocol to a Codex primary, and every turn end was blocked for missing Claude recovery. The defect is the precedence boundary, not any one harness. codex, opencode, kimi, and muse publish no identity marker at all, so with markers winning outright any retained CLAUDECODE renamed them; the Cursor-before-Claude ordering was a point patch on the same class of problem, and the launch-time marker clearing only ever covered sessions fm-spawn started. Markers and ancestry are now separate evidence layers that detect_own arbitrates: - no ancestry match, or no marker: the single available layer answers, unchanged; - same harness family: the marker's finer verdict stands, so a launch-selected pi-signed is not flattened to pi by an ancestry walk that can only see the shared launcher name; - different harness with a structural (command-name) ancestor: ancestry wins, because only ancestry proves who owns the process tree; - different harness with only a bare-interpreter script-path match: the marker wins, since a harness-shaped path in some node process's arguments is weaker evidence than a harness publishing its own identity. The correction is symmetric: a retained CURSOR_AGENT no longer renames a claude worker nested under cursor either. Adds fm-harness.sh ancestry [<pid>], ancestry evidence with no marker layer, so a real harness process can be asked what the walk makes of it. tests/fm-harness-precedence.test.sh is the portable regression, built from real renamed processes with no harness installed. Every case drives the two layers apart and asserts each alone as well as the combination, so no case can pass vacuously; it also pins Codex's real two-process install topology, since the fix depends on the native binary being what a tool subprocess meets first. The opt-in drift guard gains the matching live half: each installed harness's real running process must still be identified by the ancestry walk, and it fails naming the harness and version when a release changes that name. Documentation follows the corrected contract in the script header, the harness-adapters detection section, the codex, opencode, kimi, and cursor references, and a dated verification record. * fix(tests): drop the unused argument pass-through in the shim-topology helper bin/fm-lint.sh refused the branch: run_shim declared a `[ancestry]` argument and forwarded "$@", but every call site that varies the environment or passes the ancestry subcommand invokes the shim entry point directly, so the helper is only ever called with no arguments (ShellCheck SC2120/SC2119). Behavior is unchanged: with no arguments "$@" expanded to nothing. * fix(bin): examine the top of the process chain instead of assuming init harness_ancestry stopped as soon as the next pid was 1, on the assumption that pid 1 is always init and can never be a harness. Inside a PID namespace that assumption inverts: the harness itself is pid 1, so the walk never examined the one process that proves who owns the tree, reported no ancestry at all, and handed the verdict straight back to a retained marker. A real Codex session under `codex sandbox`, holding CLAUDECODE=1 and CLAUDE_CODE_ENTRYPOINT=cli, is exactly that shape: it resolved claude and rendered Claude's Stop-owned supervision protocol even with the marker-vs-ancestry precedence boundary in place. The same probe now resolves codex and renders the Codex foreground checkpoint. A host's real pid 1 (init, systemd, launchd) matches no harness name, so examining it costs one ps call and can introduce no false positive; the walk still stops once that top process has been read, and a non-numeric or zero ppid still ends it. tests/fm-harness-precedence.test.sh pins the namespace shape with a fake ps that reports every process as bash with ppid 1 and pid 1 as the harness. The case asserts the marker still answers alone when pid 1 is host-shaped, so it cannot pass vacuously, and it fails against the previous stop condition. * docs(verification): record the real-Codex retained-marker evidence The existing record proved the precedence boundary with the portable regression and recorded each installed harness's process name behind the ancestry walk, but it had no evidence from a real Codex process actually holding a retained Claude marker, which is the failure the boundary exists for. Adds the dated before/after result from codex-cli 0.152.0 under `codex sandbox`, with the exact command and the decisive verdict and rendered protocol on each side, and records the second boundary that shape exposed: the walk must examine the top of the process chain, because inside a PID namespace the harness is pid 1. Refreshes the portable regression's observed output for the case it gained. * no-mistakes(review): blind ancestry in marker-pinned harness tests * no-mistakes(review): blind ancestry in the Pi guard-routing test * no-mistakes(review): classify precedence suite, dedupe ps stub, soften claims * no-mistakes(review): model the spawn-and-wait Codex shim topology * no-mistakes(document): correct stale muse marker-clearing detection claims * no-mistakes: apply CI fixes * fix(bin): examine the top of the chain in the lock and nudge walks too The pid-1 defect corrected in bin/fm-harness.sh survived unchanged in the two other harness-ancestry walks, on the exact topology the branch verified against a real Codex process. bin/fm-session-lock-lib.sh's fm_harness_ancestry_pids stopped as soon as the next pid was 1, so a firstmate whose harness is pid 1 of its own PID namespace could not find that harness at all and did not recognize its own session lock. bin/fm-sessionstart-nudge.sh carried the same stop plus a blanket rejection of a lock pid of 1, so the same session was told to run session start again on every turn. Both walks now compare the top process before stopping, matching the shape used in bin/fm-harness.sh. For the lock walk this is safe because fm_harness_process_matches rejects a host's real pid 1. For the nudge, `kill -0` still gates the lock pid, and on a host an unprivileged `kill -0 1` fails, so a lock file that wrongly names pid 1 leaves the hook silent rather than acting on init. Each walk gains one regression case. The lock case drives a deterministic process table whose pid 1 is the harness and asserts a host-shaped pid 1 still finds nothing, so it cannot pass vacuously. The nudge case needs a real PID namespace, because the builtin `kill -0` gate cannot be reached through a fake ps, and it first proves the same fixture nudges with no lock present; it skips explicitly where unprivileged namespaces are unavailable. * no-mistakes(review): assert comm-strength detection from subprocess vantage in drift guard * fix(bin): verify the live harness guard at the strength the guarantee needs The marker-versus-ancestry boundary this branch ships is a strength claim: detect_own hands an args-strength verdict straight back to a retained foreign marker, so a harness is only protected where the ancestry walk reaches it at comm strength. The installed-harness drift guard probed the pane process alone. Under an interpreter shim the pane process IS the shim, whose own script path is args strength, while the native binary that carries comm strength is its child. The guard therefore observed args for Codex, passed, and would have kept passing if a release stopped spawning that native child at all, while real sessions silently regressed to the original bug. fm-harness.sh gains `ancestry-subtree`, which asks the walk from the pane process and every descendant of it, the vantage a tool subprocess actually occupies. The guard now requires comm strength somewhere in that set and requires every vantage to name the same harness. This supersedes the preceding commit's in-guard leaf walk, which reached the same vantage but left the logic inside the test file, where CI could not pin it and nothing else could reuse it. A harness-dependent check needs both halves: `tests/fm-harness-precedence.test.sh` now carries a portable case proving the subtree probe reaches a strength the top-of-session probe cannot, mutation checked twice, once against the pre-change script and once by disabling descendant enumeration. The subtree walk also avoids depending on tty and process-group semantics that differ between Linux and macOS. Verified live: codex-cli 0.152.0 reports [args codex;comm codex] and Claude Code 2.1.257 reports [comm claude]. * no-mistakes(review): narrow drift guard to the upward vantage path * no-mistakes(review): judge only comm-strength vantages in drift guard * no-mistakes(document): drop duplicated rationale in detection precedence evidence * no-mistakes(review): fix pid-1 nudge case vacuity and descent no-arg expansion * no-mistakes(document): drop branch-relative phrasing in detection precedence evidence * no-mistakes(review): guard remaining empty positional expansions in fm-harness * no-mistakes(document): scope cursor marker-ordering claim to the marker layer * no-mistakes(review): Prefer comm-strength leaves in equal-depth descent ties * no-mistakes(document): Document comm-strength descent tie-break --------- * no-mistakes(review): Blind ancestry in stale gemini/rovo marker-precedence tests * no-mistakes(document): Add missing equal-depth-tie test line to precedence evidence transcript * no-mistakes(review): Fix stale/vacuous agy precedence test, add agy to precedence suite and docs * no-mistakes(document): Fix stale kimi.md marker doc missed by ancestry-precedence fix --------- Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…rkers (kunchenguid#3944) Claude Code's external-imports check (hasClaudeMdExternalIncludesApproved) reads only the canonical git-root project entry in ~/.claude.json, which its own worktree-to-primary-checkout canonicalization means is never the task worktree fm-claude-trust.sh registered. The trust dialog kept working previously only because its check has an ancestor-walk fallback that happens to reach the worktree entry; the external-imports check has no such fallback. Verified by disassembling the installed claude binary and reproducing in an isolated three-way tmux launch: identical flags registered only at the worktree key still showed the external-imports dialog, and registering them at the primary checkout key suppressed both dialogs. fm-claude-trust.sh now registers all three flags on both the worktree entry and the primary-checkout entry in one atomic write, and refuses when the <project> argument is not itself a primary checkout (its own write target would then be wrong). Extends the harness-adapters Claude reference and the trust test suite. Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…nguid#4355) The marker lifecycle (fm-wake-lib.sh _fm_recovery_marker_ack) leaves state/.watcher-down behind in an acked:* state after a downtime episode is handled. health_snapshot's presence check reported that as an open gap on every later return, so a handled episode kept surfacing as a false GAP forever.
…kunchenguid#4361) * fix(update): rebind fm-procevent-when watches after a self-update A self-update fast-forwards bin/ in place, changing an armed watch's action executable bytes with no tampering involved. The watch's trust binding was hashed at arm time, so the very next fire was refused as not matching the registered binding and the watch died silently. Add fm-procevent-when.sh rebind-all: it re-hashes and republishes the trust binding for every watch whose action executable lives under FM_ROOT, using the same spec/trust validation as an ordinary fire, and leaves any watch whose action lives outside FM_ROOT untouched. Wire it into fm-update.sh right after a successful fast-forward, for both the primary home and any local secondmate home that advances. * no-mistakes(review): Canonicalize FM_ROOT for rebind-all's containment check * no-mistakes(document): Document fm-update.sh's automatic watch rebind and its verification evidence * no-mistakes(lint): fix(tests): double-quote printf scripts to satisfy shellcheck SC2016 * no-mistakes(review): Reload trust binding from disk before firing to reach live pollers * no-mistakes(review): Lock the fire-time trust reload against rebind_one's publish race * no-mistakes(document): Document rebind-all's self-update guarantee and its two review-round test rows --------- Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…kunchenguid#4424) * fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#42) * fix(pr-merge): read a plan-gated 403 on branch rules as no merge queue github_read_queue_method left status=unreadable for every failed rules read, including a 403 whose body is GitHub's own "Upgrade to GitHub Pro or make this repository public" message. A repository whose plan cannot expose branch rules cannot have a merge_queue rule either, so that specific 403 now resolves to status=none instead of unreadable - unblocking the away-merge grant on private repos without GitHub Pro. Any other failure (auth, rate limit, network, 404, unrelated 403) still reads as unreadable. * no-mistakes(document): Update stale away-merge queue-grant comment for plan-gated 403 --------- Co-authored-by: NewAiCoder <claude@theinbtw.com> * no-mistakes(review): Fix misleading away-queue-grant comment in fm-pr-merge and its test * no-mistakes(document): Update architecture.md for plan-gated-403 merge queue exception --------- Co-authored-by: NewAiCoder <claude@theinbtw.com>
…unchenguid#4246) * fix(tests): select readers of a changed top-level test fixture bin/fm-test-run.sh --changed recognised shared test helpers by an explicit list, tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh. A top-level tests/*-fixture.sh matched none of those, fell through to the tests/* catch-all, and was marked unmapped, so selection aborted with "no changed-test mapping for source path" and the run selected nothing at all. tests/herdr-client-pair-fixture.sh and tests/remote-herdr-fixture.sh are real shared fixtures with real consumers, so any branch touching one of them left a validation pipeline driving --changed with a hard abort rather than a narrowed selection. Extend the helper arm to tests/*-fixture.sh rather than routing it through the tests/fixtures/*/* arm. Both arms resolve consumers with the same reference scan, and that scan is what selects the right suites here: it finds exactly the tests that read the fixture. The fixtures/ arm adds only a directory-keying step, which has nothing to key on for a top-level file, so the helper arm is the same behaviour with no extra machinery. A tests/ path nothing reads still reaches the catch-all and still refuses loudly. Refs kunchenguid#4100 * no-mistakes(test): order nested fixtures arm before top-level fixture glob * no-mistakes(document): document tests/ shared-file mapping contract and arm order * no-mistakes(review): drop vacuous test phase, correct header claim, restore comment
… asked, not declined (kunchenguid#4387) * fix(bin): read Claude Code's default external-imports flags as never asked, not declined (kunchenguid#4378) fm-claude-trust.sh refused the whole trust registration whenever the project-root entry carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code writes that value only on an explicit "No, disable". Claude Code's default project entry carries Approved and WarningShown both false before the dialog is ever shown, so every such project refused every spawn. Only Approved === false with WarningShown === true — the pair the dialog writes on a decline — now counts as a decline. false/false behaves like an absent flag: trust is registered and no import consent is manufactured. New case test_project_root_entry_default_import_flags_are_not_a_decline fails on b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31, bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * no-mistakes(review): Correct harness doc's external-imports decline predicate --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…chenguid#4445) * fix(brief): keep operator address out of composed intent Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body. The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content. Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion. Fixes kunchenguid#3882 * no-mistakes(review): Refuse operator-address lines in Captain's intent body * no-mistakes(document): Document operator-address refusal in intent contract comments
…as a proven empty composer (kunchenguid#4455) * fix(composer): accept Grok title overhang * no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage
…ailure (kunchenguid#4474) * fix(bin): recover Claude auto-arm after timeout * no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list
* fix(spawn): establish Claude task channel authority * no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference
…or pending text (kunchenguid#4458) * fix: guard relaunch exit against pending input * no-mistakes(review): Verifying test run in progress * no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard * no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now
…unchenguid#4460) * fix: reconcile diverged secondmate updates * no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md * no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md
Merge the frozen canonical snapshot while preserving the fork's pilot interfaces, Windows identity and transport, Azure observation, catalog routing, and independent native/package CI producers. Compose ancestry precedence with pilot identity and POSIX namespace PID 1 without losing the Windows parent bridge. Keep reassigned-slot ownership guards ahead of destructive cleanup at the fork's existing lifecycle site. Transfer upstream registrations and fixture Git isolation through the catalog seam, including the new Herdr viewer's changed-source route. Firstmate-Upstream-SHA: d499323
Exact-head CI handoffHead: Current observation: 3 successful, 15 pending, 0 failed, 1 absent of 19 expected checks.
Workflows: Not yet merge-ready. GitHub Actions owns complete regression and cross-platform/package evidence; required skipped/cancelled lanes and historical-head successes are not proof. Please merge with a merge commit, after readiness is established; do not squash/rebase this reconciliation. |
The fork returns status 3 without output for an already-active target, leaving deferred-result reporting to lifecycle callers. The incoming live viewer test still expected upstream's warning at that earlier boundary. Require the exact quiet result for direct close and seeded pruning, retain pane-preservation assertions, and also verify that focus stays unchanged. Keep late-focus warning, successful switch-away, and detached-close cases intact. No production guard, dependency pin, or timeout changes. Firstmate-Upstream-SHA: d499323
PR 52: attached-viewer assertion reconciliationFrozen upstream: d499323 Observed failureThe completed Herdr job ran 16 scripts, with exactly one failure. The preceding nonzero-status assertion passed; the pane-preservation assertion had not yet run. DiagnosisRanked hypotheses: (1) the fork intentionally defers silently, (2) a pinned-version foreground API mismatch, (3) a lost protection. The late-on scenario still reaches the warning-bearing mutation checkpoint, so its diagnostic assertion must remain. Fix and validation routingChange only tests/fm-herdr-attached-viewer-live-e2e.test.sh, already in the 158-path reconciliation manifest. The red-capable signal is the existing real-Herdr script on the Linux CI producer, not a local mock or source-text assertion. |
Attached-viewer correction: exact-head CI resultHead: Current observation: 12 successful, 6 pending, 0 failed, 1 absent of 19 expected checks.
Workflows: Not yet merge-ready. GitHub Actions owns complete regression and cross-platform/package evidence; required skipped/cancelled lanes and historical-head successes are not proof. Please merge with a merge commit, after readiness is established; do not squash/rebase this reconciliation. |
Create the fixture repository with explicit test identity and model one stable live session owner through the platform's proc-root test seam. Bind all operational directories to the temporary home, verify actual lock acquisition, and require the specific bounded-read diagnostic. Expose the two existing halves through the shared named-case registry. Keep Linux's 2-second read and 30-second ceiling; allow measured native Windows setup overhead with 10/120-second fixture limits while preserving strict latch and timeout assertions. Production deadlines are unchanged. Firstmate-Upstream-SHA: d499323
Backlog fixture correction and resumed validationCurrent head: CI diagnosisOn The incoming fixture faked ChangesOnly Local Windows execution then exposed fixture timing assumptions: a latched read took 2 seconds, and full protected startup took 65 seconds. Local validation after explicit authorization to continue without circuit breakersAll commands use the recorded
Earlier cancellation timeout resolvedThe unchanged named cancellation assertion also passed on Windows: FM_TEST_STUB_MAX_BLOCK_SECONDS=900 \
FM_TEST_ONLY=test_timed_out_provision_cancels_late_launch \
bin/fm-test-run.sh --jobs 1 tests/fm-herdr-lab.test.shThe controller allowed 420 seconds for this command; the fake server's existing lifetime knob kept it alive until the cancellation being tested, rather than expiring first. Total cumulative local validation, including all probes and failures: 2297298 ms (38m17.298s); caching remains disabled. The full exact-head CI matrix is running; this report does not relabel an older head's checks as current passes. |
Final exact-head CI: all 19 automatic checks passedHead: Current observation: 19 successful, 0 pending, 0 failed, 0 absent of 19 expected checks.
Workflows: Both reported CI failures are resolved and verified on this exact head. Local continuation also passed the full backlog fixture on Windows, all four repository gates, source-aware backlog lint, and identical 220-script routing. GitHub Actions supplies the complete automatic cross-platform/package matrix. Optional credentialed vendor checks remain gated, and the manual-only Windows Herdr spike was not dispatched; those are not new live-vendor guarantees. |
Frozen upstream reconciliation
Normal merge of 35 canonical commits, retaining the fork's interfaces and platform behavior.
Use a merge commit, not squash or rebase. Merge commits are enabled; leave this ordinary, non-draft PR unmerged with auto-merge disabled.
All 19 automatic CI checks now pass on the published head, with one expected producer each.
GitHub Actions owns the complete cross-platform matrix; optional credentialed vendor checks remain gated.
f96ed25d1ba1a866fc5ffe5bfb5862f4466ce4aae0d269e07318a5a80069ab6193b4a0af4c077a61d49932335485e2098f13fe1cafaa3077d4410f6cb087c8dd287a0172cd07f175623f9fc2704dc64bd9b73af6f291577a8e15dd1ff7f54c68d8761038a188a6f54b954ff8ac788b36a75c4b73d163dbe2bf563df3c2f8fad5479e8d176d87537094d2744fFirstmate-Upstream-SHA: d499323
Upstream was fetched exactly once.
The newest valid reachable trailer, on
22e143cbe049052dc8c4afca1c5afb8cb7de37a7, names the prior upstream above; that object exists and is an ancestor of both frozen inputs.Prior merge
bdf50f25d5020c8944d3d60984cca5e926369ef3has it as its second parent, and merged PR #47 records the same snapshot and retained head.This proves synchronization independently of the matching merge base.
The new merge's parents are the frozen fork and upstream: no reconstruction, synthetic test worktree or ancestry anchor was needed.
Git's trailer parser, tree, modes, exact path manifest and clean worktree were verified.
An origin-fetch race with another process's primary-checkout fast-forward settled at the frozen fork SHA.
This run used a separate clean worktree at that SHA, without reset, stash, discard or primary-checkout changes.
Existing local rerere settings (
enabled=true,autoupdate=false) and global settings were preserved; no reused resolution was blindly staged.Conflict and ownership decisions
All 28 conflicted paths were resolved against base/fork/upstream versions and originating history.
Primary sources: canonical PRs kunchenguid#3578 (ancestry), kunchenguid#4200 (AGY), kunchenguid#4199 (live-head merge gates), kunchenguid#4243 (reassigned slots), kunchenguid#3825 (Git fixtures), kunchenguid#4246 (fixture routing), kunchenguid#4281 (CI); fork PRs #43 (modules), #48 (native Claude), #50 (Copilot transport), #51 (Azure completion).
bin/fm-harness.sh,bin/fm-agent-process-lib.shbin/fm-session-lock-lib.sh,bin/fm-sessionstart-nudge.shbin/fm-bootstrap.sh,bin/fm-control-lib.sh,bin/fm-spawn.shbin/fm-pr-merge.shghfor--match-head-commit;gh-axiremains the outcome fallback and this maintainer's forge client.bin/fm-teardown.shbin/fm-test-run.sh.github/workflows/ci.ymltests/fm-control-relaunch.test.sh,tests/fm-herdr-lab.test.sh,tests/fm-pr-check-security.test.sh,tests/fm-session-lock-ancestry.test.sh,tests/fm-sessionstart-nudge.test.sh,tests/fm-spawn-dispatch-profile.test.sh,tests/fm-test-run.test.sh,tests/fm-x-mode.test.sh.agents/skills/harness-adapters/SKILL.md,AGENTS.md,CONTRIBUTING.md,docs/agent-control.md,docs/architecture.md,docs/tmux-backend.md,docs/trace-context.mddocs/fm-test-portable-shards.md,docs/verification/runtime-backends.mdUnconflicted quiet supervision, bounded backlog reads, Claude trust/import consent, task-first launch identity, merge-authority persistence, update-time watch rebinding, redundant secondmate divergence handling and composer/Stop safeguards remain with their lifecycle owners.
Canonical process/private-path implementations, Pi/OpenCode wrappers, both pilot implementations, timeout library, remote doctor/integrity entrypoint and concurrency proof owner are byte-identical to the frozen fork.
No live Firstmate home, fleet operation, vendor prompt, credentials or repository settings were modified.
Catalog composition:
bin/fm-herdr-lab-viewer.py. A named regression failed before its explicit Herdr catalog mapping and passed afterward, as did the complete inventory. Unknown-source refusal remains.CI corrections
b6476dba893dbc906810240f205cc62dd83aa2d7corrects the attached-viewer test to require silent status 3 for already-active targets, preserving panes and focus; all six live-viewer cases and the 16-script Herdr lane passed on that head.d9b73af6f291577a8e15dd1ff7f54c68d8761038fixes the backlog-bound startup fixture's Git identity and proc-root/session-owner model, rejects read-only false positives, and requires the actual bound diagnostic.Its original Linux limits remain 2/30 seconds; Windows-only fixture limits of 10/120 accommodate measured launch/ACL overhead, with strict latch and timeout assertions retained.
Both follow-ups change tests only; production code, production deadlines, package pins and workflows are unchanged.
The continuation record contains the failure evidence, every resumed command/outcome/timing, fixture PATH preparation, and validation authorization.
Verification ownership and limits
The resolved workflows expand to 19 automatic checks, one producer each (13 shared, six fork).
Both retain main push/PR triggers and read-only permissions; groups supersede the same PR, preserve each main push and never cancel the other workflow.
The Windows Herdr experiment remains manual and was not dispatched.
Shared CI retains two parallel lanes, five serial shards, real Herdr, timing aggregation, stock Bash and invariants.
Portable jobs retain tasks-axi
0.2.5, TypeScript5.9.3, the existing public Pi install policy and missing-typecheck-prerequisite refusal.Herdr retains
0.7.4/ protocol 16, Treehouse2.0.1, Pi0.84.3, 75-minute job/20-minute suite bounds and isolated cleanup.Fork package compatibility retains Pi
0.84.3, OpenCode1.18.23, TypeScript5.9.3, Node 24 and no vendor credentials.Windows keeps core, Copilot launch, legacy rollback and PR completion, including tasks-axi/native review/cleanup and the existing PR-completion 15-minute/600-second bounds.
Portable/Herdr uploads still feed the dependent
always()aggregate; action versions remain checkout v6, setup-node v7, upload-artifact v7 and download-artifact v8.The earlier Windows cancellation timeouts are resolved: after the user authorized continuation without circuit breakers, the unchanged named assertion passed in 134422 ms with a 420-second external window and the existing fake-server lifetime knob set to 900 seconds.
It proved exact-PID cancellation and no late launch without changing any production deadline or assertion; no platform-baseline incompatibility is claimed.
Missing prerequisite: Ruby was absent; the workflow-contract suite was deferred before execution to
CI / Behavior portable serial 1.Live/optional gates remain limitations; upstream's AGY PR also records an inconclusive full live-spawn scenario.
Both reported CI regressions and the earlier local cancellation-window issue are resolved; the complete exact-head automatic matrix is green.
Local evidence
Interpreter:
C:\Program Files\Git\usr\bin\bash.exe, never WSL;FM_LIVE=0; no dependency installs.Observed: Git
2.55.0.windows.5, Bash5.3.15, Node24.19.0, jq1.8.2, Perl5.42.3, cygpath3.6.10, ShellCheck0.11.0, actionlint1.7.12.Read-only Git policy probes did not disable host safeguards; the fixture regression proved explicit configuration and outside-fixture signing remain authoritative.
MSYS and scoped native Windows inspection found no remaining test process.
Local passes: four repository gates and source-aware backlog lint were refreshed on the current tree; the full backlog-bound fixture and the continued Herdr cancellation case passed.
The initial tree additionally passed 20 focused behavior invocations and source-aware owner lint; those remain attributed to that tree.
Every resumed command and its failed predecessors is in the linked continuation record.
Inventory/coverage: 220 selected = 24 parallel + 180 serial + 16 Herdr, five complete serial shards, zero missing parallel hints, 20 unhinted serial scripts.
The full selection was routed, never run locally. Docs: 106 surfaces / 444 local links.
Cumulative local ledger: 2,297,298 ms (38m17.298s), including all probes, failures and resumed checks.
The user explicitly removed the earlier circuit-breaker stop; commands retain individual cancellation deadlines.
Caching remains disabled; interrupted, failed, skipped and preflight-only results are not passes.
Initial commands are retained below; continuation commands and per-invocation totals are in the linked record.
Plans, results, logs, side versions, manifests and no-renames inventories are outside the repository.
Exact local commands and outcomes
All commands below ran through the recorded Bash and the bounded controller.
S,XandIexpand to the following exact recipes:S:
X:
I:
Preflight commands (unique probes; shared probes are memoized only within an invocation):
T(subject, case)below expands toFM_TEST_ONLY=<case> bin/fm-test-run.sh --jobs 1 tests/<subject>.test.sh;FULLomitsFM_TEST_ONLY.Sbin/fm-lint.shbin/fm-doc-audience-check.shbin/fm-test-run.sh --check-coverageXIT(fm-test-run, test_herdr_viewer_selects_its_consumers)IT(fm-test-run, test_herdr_viewer_selects_its_consumers)T(fm-test-catalog, FULL)T(fm-harness-contract, FULL)T(fm-test-run, test_changed_dependency_selection_and_unmapped_failure)T(fm-test-run, test_changed_shared_fixture_selects_its_readers)T(fm-test-run, test_changed_shared_fixtures_select_consumers)T(fm-test-run, test_harness_modules_select_all_consumers)T(fm-test-run, test_fork_workflow_selects_its_contracts)T(fm-test-run, test_changed_runner_surfaces_select_their_family)T(fm-test-fixtures, FULL)T(fm-session-lock-ancestry, test_windows_native_ancestry_uses_verified_parent_rows)T(fm-session-lock-ancestry, test_windows_orphaned_claude_uses_verified_pid_handoff)T(fm-session-lock-ancestry, test_native_windows_claude_session_acquires_lock)T(fm-session-lock-ancestry, test_harness_at_namespace_pid1_is_examined)T(fm-spawn-dispatch-profile, test_copilot_threads_model_effort_and_hooks)T(fm-spawn-dispatch-profile, test_claude_permission_mode_auto_swaps_only_the_permission_flag)T(fm-spawn-dispatch-profile, test_claude_task_launch_carries_control_channel_authority)T(fm-control-relaunch, test_relaunch_refuses_before_exit_when_the_composer_holds_pending_text)T(fm-control-relaunch, test_relaunch_refuses_before_exit_when_the_composer_state_is_unproven)T(fm-pr-check-security, test_azure_merge_is_explicitly_unsupported)T(fm-herdr-lab, test_timed_out_provision_cancels_late_launch)T(fm-ci-workflow, FULL)Sbin/fm-lint.shbin/fm-doc-audience-check.shbin/fm-test-run.sh --check-coverageT(fm-herdr-lab, test_timed_out_provision_cancels_late_launch)No-renames divergence and locality audit
The 152-path source manifest plus seven explicit integration paths resolves to 158 changed paths:
docs/fm-test-portable-shards.mddeliberately returns to unchanged fork bytes.The intended, staged and committed NUL-delimited path sets match, and all file modes retain the corresponding fork/upstream source mode.
The final path-manifest SHA-256 is
51ca177e82e4abda20de2d482300aac85c7ff9d636a40555ab7de4ff66a06253.Of 152 incoming paths, 76 are now byte-equivalent to frozen upstream: 62 existing paths plus 14 incoming additions.
The other 76 compose retained fork behavior or the reconciliation fixes, including whitespace repair and the attached-viewer assertion correction.
All 16 PR additions originate upstream; this run adds no new fork-only module.
Persistent fork divergence grows from 221 paths at the proven prior snapshot to 225 against the new snapshot; the apparent 301 -> 225 decrease is not a claim that fork behavior disappeared.
The 55 additive fork paths remain 55, and inherited deletions remain the mandatory no-mistakes workflow and its policy test.
No rename detection was used to manufacture a smaller count.
Reproduce the complete inventories and caller hunks:
fm-pr-poll.sh, registration infm-pr-lib.sh, downstream backlog/review/cleanupNew detection orchestration remains in its caller and uses the existing adapter implementation; metadata moves to catalogs, not back into inline tables.
Existing signed-Pi duplication and static integrity facts remain deliberate exceptions, not accidentally retained extraction work.
Native process/private-path code was not duplicated or moved; its compatibility calls and lifecycle authority remain distinct.
The complete 76-path equivalence inventory is retained with the outside-repository audit and reproducible from the literal comparisons above.
Complete selected-script routing
Each entry expands to
tests/<name>.test.sh.Full regression is assigned to its existing Linux producer;
LOCAL FULLandLOCAL CASESrefer only to the command outcomes above.Native Windows and pinned-package contracts additionally remain with fork CI, while live/optional gates are not converted into passes.
All 220 selected scripts and their primary CI owners
CI / Behavior portable parallel 1 (11)
CI / Behavior portable parallel 2 (13)
CI / Behavior portable serial 1 (34)
CI / Behavior portable serial 2 (37)
CI / Behavior portable serial 3 (36)
CI / Behavior portable serial 4 (36)
CI / Behavior portable serial 5 (37)
CI / Behavior tests (Herdr) (16)
Expected exact-head checks
The live-check follow-up records successful, pending, failed or absent status for this exact head.
Skipped/cancelled required checks and checks for older heads do not establish merge readiness.
linttest-coveragetests-portable-parallel-1tests-portable-parallel-2tests-portable-serialtests-portable-serialtests-portable-serialtests-portable-serialtests-portable-serialtests-herdrtests-timing-aggregatemacos-stock-bashinvariantswindows-updatereconciliation-windowsreconciliation-windowsreconciliation-windowsreconciliation-windowsharness-package-compatibilityManual-only:
Windows Herdr automation spike / measure(Measure Herdr automation primitives), not included in the 19 automatic producers and not invoked.