diff --git a/.agents/skills/afk/SKILL.md b/.agents/skills/afk/SKILL.md index a00beb35038..8d3a6e7a10e 100644 --- a/.agents/skills/afk/SKILL.md +++ b/.agents/skills/afk/SKILL.md @@ -27,7 +27,10 @@ Hold-for-return is the default and the only reach profile this release records: Write only clauses the words actually support; a wish with no object or no stated precondition is not a clause. Plain `/afk` with no words has no clauses. 2. **Propose and read back.** - Run `bin/fm-afk-launch.sh propose --words-file [--action --object --when [--stop ]]... [--expected-return ] [--spend ]` (or `--words `), and relay its read-back to the captain in `AGENTS.md` section 9 language: the accepted clauses as a numbered list, every refused clause with the part it is missing, the expected return, the spend cap, and the one-sentence reach announcement. + Run `bin/fm-afk-launch.sh propose --words-file [--action --object --when [--stop ]]... [--expected-return ] [--spend ] [--grant ]...` (or `--words `), and relay its read-back to the captain in `AGENTS.md` section 9 language: the accepted clauses as a numbered list, every refused clause with the part it is missing, the expected return, the spend cap, any merge-when-green task ids, and the one-sentence reach announcement. + When the captain names task ids that may merge while green, pass `--grant ` for each named id. + Never infer task ids from clause prose, object text, or the away words. + Red-check exceptions stay in the words or clause `when` text and are not executed. A refused clause does not fail the proposal; the captain can restate it or leave it refused. Exit 3 only means a clause was refused; the proposal stands. 3. **Confirm on the captain's go.** @@ -65,17 +68,23 @@ No `/back` is needed. The first genuine message is the return signal: The gate keeps every open `blocked:` event until that blocker's own resolution is proven: remediate each immediately through the normal lifecycle, or explicitly reclassify it with a durable reason and close its decision key with `resolved [key=...]`, then run `bin/fm-afk-return.sh check`. Captain-verdict outcomes are listed under "waiting on you", but do not exempt open blockers because per-blocker provenance is deferred to phase 4. Once the record is archived, resume full per-wake responsiveness through the emitted primary-harness supervision protocol while blocker handling proceeds, so the gate never creates a blind wait. - Do not answer a Bearings request or perform any other ordinary captain work until the check exits successfully. + A Bearings request may be answered while the gate is open, and the digest surfaces the catch-up state as a Charted Next `(return-catchup)` warning row naming what still holds it. + Acting on the fleet - dispatching, steering, merging, or any other ordinary captain work - still waits until the check exits successfully. - A message **with** the current operational prefix (`FM_OPERATIONAL_PREFIX`, U+2063 INVISIBLE SEPARATOR followed by `FIRSTMATE_OP: `), or a legacy bare `FM_INJECT_MARK` daemon escalation -> stay away and process it. - Re-invoking `/afk` while already away -> stay away (refresh); this does **not** trigger an exit. Bias ambiguous cases toward exit: a present captain beats token savings, and a false exit is self-correcting (the captain re-runs `/afk`). +When the captain wants this same token-saving supervision while staying present and chatting - ordinary messages should NOT exit it - that is `/quiet` (kunchenguid/firstmate#2356), not `/afk`. ## Orthogonal to approval authority afk changes how the captain is informed and what happens at a captain-owned decision point, **not who approves what**. "Away" never means "approves more" or "approves less." A PR ready for merge keeps the merge authority from `AGENTS.md` section 7, and a needs-decision finding keeps the `ask-user-authority` policy; anything requiring the captain still waits for the captain's explicit word. +While the away-posture record exists, a merge proceeds only when that task's recorded yolo posture is on or its id is in the record's merge-grant list; otherwise it is held for the captain's return. +A merge grant never releases a captain hold, and it expires when the away record is archived. +`--allow-red` remains attended-only and is refused while the record exists. +A merge under away authority must be synchronous; `fm-pr-merge.sh` refuses auto-merge and any GitHub queue state that cannot prove an immediate merge while the record exists. A mandate clause is the captain's explicit instruction given before leaving, recorded with its named object and condition; a clause is never inferred, never applied by analogy, and expires at return. Forbidden, destructive, irreversible, and security-sensitive actions are never pre-authorizable regardless of clause text, and no recorded clause is authority by itself. This release records clauses and does not execute them. diff --git a/.agents/skills/bearings/SKILL.md b/.agents/skills/bearings/SKILL.md index 12c08aa915f..907145c03d3 100644 --- a/.agents/skills/bearings/SKILL.md +++ b/.agents/skills/bearings/SKILL.md @@ -54,6 +54,8 @@ Board answers are acted on later under the normal authority rules; this skill's The `(main-inventory)` gate is an action-free integrity warning rather than queued work. Render it under Charted Next with the related `omitted` disclosure, never invent an Underway row from backlog-only state, and never move it into Captain's Call. The same holds for a secondmate home whose current state is unavailable, and for a readable home whose `invalidity` reports a backlog-vs-metadata mismatch: the mismatch is a repair notice about that home's own books, not a reason to drop its separately projected decisions, queued, landed, or live work. + The `(return-catchup)` gate is the same shape: an action-free notice that an away-return catch-up is still open, naming the blockers left to clear or the reason the catch-up was retained. + Render it under Charted Next like any other warning row: reporting is not ordinary work, while acting on the fleet still waits for `bin/fm-afk-return.sh check` (`/afk`). 2. **Record a later reconcile notification for any home whose own books disagree.** When the snapshot reports a secondmate home whose `invalidity` is `orphan_in_flight`, `unowned_current`, or `terminal_in_flight`, that home's backlog and its own task metadata disagree and only that home may fix it. @@ -99,8 +101,13 @@ Compose the payload from the same snapshot with the same ranking judgment as the - Decision cards carry agent-authored copy: a short noun-phrase title, one-line `about` and `decide` context rows, and option labels with hints, with the recommended option marked. - Card `type` (decision, merge, credential) is your composing judgment from the row's content; no backlog field types a card for you. - When the card's task is a captain-gated WORK item (the answer should free it to proceed rather than complete it), set the card's `close: "release"` so the answer lifts the hold instead of closing the task; question-shaped items omit it. -- A Charted Next row's optional `kind` separates work from alarms: omit it (or set `"queued"`) for real queued work, and set `"warning"` on every action-free fleet-integrity notice - the `(main-inventory)` gate, an unavailable secondmate home, and an inventory-mismatch repair notice. The board badges a warning row `needs repair` instead of `waiting` and leaves it out of the Charted Next count, so those rows never read as dispatchable queued work. +- A Charted Next row's optional `kind` separates work from alarms: omit it (or set `"queued"`) for real queued work, and set `"warning"` on every action-free fleet-integrity notice - the `(main-inventory)` gate, the `(return-catchup)` gate, an unavailable secondmate home, and an inventory-mismatch repair notice. The board badges a warning row `needs repair` instead of `waiting` and leaves it out of the Charted Next count, so those rows never read as dispatchable queued work. - `charted_more` counts omitted queued rows only, while `charted_warning_more` counts omitted warning rows only; keep both counts separate whenever the board payload truncates Charted Next. +- Every Underway row copies the task-identifying `in_flight.name` from the snapshot into an explicit `name` field, which the board leads with while keeping the run status on its second line. + The snapshot command's header owns its durable-title-or-id normalization; never replace the projected label with run status or invent another label. +- Every Charted Next row copies the snapshot gate's durable filed date into `filed`, and the board orders the section by it, newest filed first. + Follow `bin/fm-bearings-board.sh`'s payload contract for the accepted format. + Omit it or pass null for a row with no durable filed date - the main-inventory or return-catchup warning, an unavailable secondmate home, or a queued row filed before dates were recorded - and the board keeps those rows in payload order after every dated row. - Every Captain's Call item and every Underway, Recently Landed, and Charted Next row carries an explicit `repo` field. Fill it from the snapshot and task records wherever known; use null or an empty string only as the deliberate genuinely-no-repo marker, in which case the template may show the internal id. Ids otherwise stay in the payload only as the routing channel, and composed reasons name blockers in plain words. Run `build` once after composing the payload. diff --git a/.agents/skills/bearings/assets/board-template.html b/.agents/skills/bearings/assets/board-template.html index 614bef7426b..987db2d8c03 100644 --- a/.agents/skills/bearings/assets/board-template.html +++ b/.agents/skills/bearings/assets/board-template.html @@ -439,6 +439,17 @@ so every count of queued work excludes them. */ function isWarning(t) { return t && t.kind === "warning"; } function chartedQueued(rows) { return (rows || []).filter(function (t) { return !isWarning(t); }); } + /* Charted Next reads newest filed first, so the most recently filed upcoming + work is at the top. ISO filed dates compare as text; a row with no + comparable date keeps its payload order after every dated row. */ + function chartedOrder(rows) { + var dated = [], undated = []; + (rows || []).forEach(function (t) { + if (t && typeof t.filed === "string" && t.filed) dated.push(t); else undated.push(t); + }); + dated.sort(function (a, b) { return a.filed < b.filed ? 1 : (a.filed > b.filed ? -1 : 0); }); + return dated.concat(undated); + } var chartedMoreQueued = data.charted_more || 0; var chartedMoreWarnings = data.charted_warning_more || 0; function utf8ByteLength(text) { return new TextEncoder().encode(text).length; } @@ -617,10 +628,13 @@ var row = el("div", "bb-row"); row.appendChild(badge(t.state === "working" ? "online" : "info", t.state)); var main = el("div", "bb-row__main"); - main.appendChild(el("div", "bb-row__title", t.doing)); + /* the snapshot's durable name-or-id label leads the row so a scan says + WHICH task this is; the run status keeps its place on the second line */ + main.appendChild(el("div", "bb-row__title", t.name)); /* captain-facing rows name the repo; the internal task id shows only when no repo is known */ - main.appendChild(el("div", "bb-row__sub", t.kind + " · " + (t.repo || t.id))); + main.appendChild(el("div", "bb-row__sub", + t.doing + " · " + t.kind + " · " + (t.repo || t.id))); row.appendChild(main); uw.appendChild(row); }); @@ -664,7 +678,7 @@ if (!chartedQueued(data.charted).length && !chartedMoreQueued) { ch.appendChild(el("div", "bb-empty", "Nothing is queued.")); } - data.charted.forEach(function (t) { + chartedOrder(data.charted).forEach(function (t) { var row = el("div", "bb-row"); if (t.dispatchable && !isWarning(t)) { anyPickable = true; diff --git a/.agents/skills/bootstrap-diagnostics/SKILL.md b/.agents/skills/bootstrap-diagnostics/SKILL.md index e4dde858ea8..d6772de62d3 100644 --- a/.agents/skills/bootstrap-diagnostics/SKILL.md +++ b/.agents/skills/bootstrap-diagnostics/SKILL.md @@ -2,7 +2,7 @@ name: bootstrap-diagnostics description: >- Agent-only handling playbook for session-start bootstrap diagnostics. - Use whenever the session-start digest's bootstrap or network-checks section prints an actionable diagnostic line - MISSING, MISSING_MANUAL, BACKEND_INVALID, NEEDS_GH_AUTH, TANGLE, VAULT_DRIFT, UPSTREAM, GBRAIN_SERVING_CREDENTIAL, GBRAIN_PIN, GBRAIN_CAPTURE, STARTUP_MEMORY_BUDGET, CREW_DISPATCH (invalid or backend mismatch), FLEET_SYNC, BOARD_SWEEP, NETWORK_CHECKS, HOME_SUMMARY, BACKLOG_RECONCILE, ENDPOINT_BINDING_MIGRATION, RUN_ATTRIBUTION, SECONDMATE_SYNC, SECONDMATE_LIVENESS, SECONDMATE_HANDOFF, NUDGE_SECONDMATES, USAGE_STORE, or FMX - or reports that an interrupted backlog cleanup may have left an endpoint or local copy, or when a standalone bin/fm-bootstrap.sh or bin/fm-startup-network.sh run prints one of those lines. + Use whenever the session-start digest's bootstrap or network-checks section prints an actionable diagnostic line - MISSING, MISSING_MANUAL, PRESENTATION_UNAVAILABLE, BACKEND_INVALID, NEEDS_GH_AUTH, TANGLE, VAULT_DRIFT, UPSTREAM, GBRAIN_SERVING_CREDENTIAL, GBRAIN_PIN, GBRAIN_CAPTURE, STARTUP_MEMORY_BUDGET, CREW_DISPATCH (invalid or backend mismatch), FLEET_SYNC, BOARD_SWEEP, NETWORK_CHECKS, HOME_SUMMARY, BACKLOG_RECONCILE, ENDPOINT_BINDING_MIGRATION, RUN_ATTRIBUTION, SECONDMATE_SYNC, SECONDMATE_LIVENESS, SECONDMATE_HANDOFF, NUDGE_SECONDMATES, USAGE_STORE, or FMX - or reports that an interrupted backlog cleanup may have left an endpoint or local copy, or when a standalone bin/fm-bootstrap.sh or bin/fm-startup-network.sh run prints one of those lines. A silent bootstrap section, or any other BOOTSTRAP_INFO fact, means no skill load. user-invocable: false metadata: @@ -19,9 +19,12 @@ When any diagnostic needs captain attention, report the plain consequence and re - `MISSING: (install: )` - list the missing tools to the captain with a one-line purpose each plus the printed install commands, wait for consent (one approval may cover the list), then run `bin/fm-bootstrap.sh install `. For `treehouse`, this also covers an installed version whose `treehouse get` lacks `--lease`; treat it as an upgrade request. For `no-mistakes`, this also covers an installed version older than 1.46.0, because this repo's PR gate requires structured pipeline attestation that older builds do not write. - For any axi-family tool - `gh-axi`, `lavish-axi`, `tasks-axi`, `quota-axi` - an installed version below its floor is a plain upgrade request; [`bin/fm-bootstrap.sh`](../../../bin/fm-bootstrap.sh) owns the floor policy, and never argue the floor down to whatever the home happens to have installed. + For essential axi-family tools - `gh-axi`, `tasks-axi`, `quota-axi` - an installed version below its floor is a plain upgrade request; [`bin/fm-bootstrap.sh`](../../../bin/fm-bootstrap.sh) owns the floor policy, and never argue the floor down to whatever the home happens to have installed. For `tasks-axi`, this additionally covers an installed build that fails the separate feature probe (`bin/fm-tasks-axi-lib.sh` owns the definition); `config/backlog-backend=manual` only suppresses the verbose `BOOTSTRAP_INFO: tasks-axi available` fact, not this missing-tool report. For `quota-axi`, bootstrap requires it because firstmate reads its current output directly before resolving every crew-dispatch profile array; without it, report the missing requirement and do not choose around an unexamined candidate. +- `PRESENTATION_UNAVAILABLE: lavish-axi ...` - explain that visual presentation is unavailable and continue nonvisual work with plain-text decisions and reports; do not hold unrelated dispatch for installation consent. + Do not use Lavish until it satisfies the floor owned by `bin/fm-bootstrap.sh`; when visual work needs it, request consent for the printed install or upgrade command, then rerun bootstrap to confirm compatibility before using it. + Scout briefs check the same floor when scaffolded and ask for a text report instead of a Lavish loop, so scaffold a visual scout only after that rerun confirms compatibility. - `MISSING_MANUAL: (instructions: )` - tell the captain why the tool is required and give them the printed instructions URL, but do not pass the tool to `bin/fm-bootstrap.sh install`; wait for the captain to complete the manual installation, then rerun session start to confirm the dependency is present. - `BACKEND_INVALID: (known: )` - the resolved runtime backend has no verified dependency or lifecycle contract, so do not dispatch work until the invalid `FM_BACKEND` or `config/backend` value is corrected to one of the listed backends. - `NEEDS_GH_AUTH` - ask the captain to run `! gh auth login` (interactive; you cannot run it for them). @@ -88,6 +91,9 @@ When any diagnostic needs captain attention, report the plain consequence and re Treat the task's run state as unreadable rather than absent, and reconstruct any urgent supervision decision from the recorded worktree's git state plus `no-mistakes axi status` before allowing branch edits or commits. Never hand-write `branch=` from the checked-out branch, an `fm/` naming match, a run listing, or a work item; those are the same unproven inferences the attribution guard refuses. A future locked startup may converge a GitHub task only when its recorded PR URL and recorded PR-head SHA still match the forge's current PR head, while every other task remains diagnosed until cleanup. +- `BACKLOG_RECONCILE: code-root is not this home's ; ...` - a tasks-axi write addressed the code root instead of this home, so the queue has already forked and either copy may hold rows the other lacks; [`docs/configuration.md`](../../../docs/configuration.md) ("Backlog backend") owns why. + Neither copy is a safe winner: union-merge them into this home's file by task id, resolve each conflicting id to its most recent real transition, check this home's archive before treating a missing Done row as lost, and verify the merged id set equals the union of both inputs before installing it. + Then move the code-root file aside rather than deleting it, tell the captain which rows were recovered, and run every later backlog command through `bin/fm-tasks-axi.sh`; re-linking the code-root copy is never the fix, because the next cwd-relative tasks-axi write replaces the link again. - `SECONDMATE_SYNC: secondmate : skipped: ` - secondmate convergence left a live home on its existing checkout because the home was dirty, diverged, unsafe, on the wrong branch, missing its placement-specific target commit, unreachable, or otherwise not fast-forwardable, or because inherited local-material propagation failed; bootstrap continued, but inspect the reason because the secondmate's tracked instructions, inherited settings, or shared captain preferences may be stale after a primary update. - `SECONDMATE_LIVENESS: secondmate : skipped: |respawn failed after : ` - the session-start liveness sweep could not guarantee that the registered secondmate is running a real agent process. Investigate the reason because that secondmate is not guaranteed live. diff --git a/.agents/skills/firstmate-coding-guidelines/SKILL.md b/.agents/skills/firstmate-coding-guidelines/SKILL.md index a595b37f48f..77e979855bc 100644 --- a/.agents/skills/firstmate-coding-guidelines/SKILL.md +++ b/.agents/skills/firstmate-coding-guidelines/SKILL.md @@ -56,7 +56,7 @@ That is the trigger condition for loading the skill, plus any safety-critical fa Everything else - the procedure, the mechanism, the surrounding detail - moves out completely. Do not leave a partial restatement behind "just in case". A partial copy is exactly the duplication the one-owner rule forbids. -The model to copy is `AGENTS.md` section 8's "Away-mode stub": it keeps only the marker format, the ownership-transfer rule, and the exit condition inline, and points everything else at the `/afk` skill. +The model to copy is `AGENTS.md` section 8's "Away-mode and quiet-mode stub": it keeps only the marker format, the ownership-transfer rule, and the exit condition inline, and points everything else at the `/afk` and `/quiet` skills. ## Size discipline diff --git a/.agents/skills/fmx-respond/SKILL.md b/.agents/skills/fmx-respond/SKILL.md index 4ffb0d17a04..2e7504c4c95 100644 --- a/.agents/skills/fmx-respond/SKILL.md +++ b/.agents/skills/fmx-respond/SKILL.md @@ -257,7 +257,7 @@ So treat second-mate-routed Relay work as a promised final by construction: the **When you promise a final (including every Relay request whose work is routed to a second mate):** -1. Create the typed obligation with `tasks-axi public-followup add` and bind the work with `bind-work`, keeping the public-safe summary and the opaque thread binding in the obligation and the full request context where the poll already put it. +1. Create the typed obligation with `bin/fm-tasks-axi.sh public-followup add` and bind the work with its `bind-work`, keeping the public-safe summary and the opaque thread binding in the obligation and the full request context where the poll already put it. When the public ask plainly implies follow-on work ("look into X and fix it"), register the promised-final against the outcome and deliver any interim report as a separate `--purpose milestone` obligation on the same thread. An ask that genuinely terminates at a report stays `report-ready`; do not invent a ship commitment for work the captain has not authorized. 2. Register it with `bin/fm-public-followup.sh register --relation --work-home > --work-id --generation `. diff --git a/.agents/skills/harness-adapters/SKILL.md b/.agents/skills/harness-adapters/SKILL.md index 1cd341faa98..7a741a1e3de 100644 --- a/.agents/skills/harness-adapters/SKILL.md +++ b/.agents/skills/harness-adapters/SKILL.md @@ -40,7 +40,8 @@ Agy is verified only for crewmate and scout work on the Herdr backend, never a p ## Detection -`../../../bin/fm-harness.sh` prints firstmate's own harness from verified environment markers, then process ancestry. +`../../../bin/fm-harness.sh` prints firstmate's own harness from verified environment markers and process ancestry, and owns how they combine. +A marker names its harness, but a structural ancestor of a different harness outranks it, because a marker is ordinary environment state a child or a multiplexer can retain while ancestry is what proves who owns the process tree. Only `FM_PI_HARNESS=pi-signed` at the launch boundary together with `PI_CODING_AGENT=true` selects Pi-signed; shared unmarked launcher ancestry remains Pi. omp publishes no marker of its own; `FM_OMP_HARNESS=omp` is Firstmate's launch marker and the anchored process name `omp` is its ancestry evidence, as `references/harness/omp.md` records. `../../../bin/fm-spawn.sh` owns worker marker establishment, while the README launch command owns the signed-primary boundary. diff --git a/.agents/skills/harness-adapters/references/common/control-and-recovery.md b/.agents/skills/harness-adapters/references/common/control-and-recovery.md index c16a78bb8a8..4223b63b895 100644 --- a/.agents/skills/harness-adapters/references/common/control-and-recovery.md +++ b/.agents/skills/harness-adapters/references/common/control-and-recovery.md @@ -17,15 +17,12 @@ Select only its documented trust choice from the active Firstmate home, binding No observed dialog proves only that launch. Each supported harness handles its folder-trust gate differently, and the tool reference owns the detail. -Claude gates a fresh worktree and cannot be answered by key, so the spawn pre-registers the path in Claude's own store. +For Claude, load `references/harness/claude.md`; its workspace-trust section owns the non-key-answerable gate and spawn-time pre-registration for every spawn kind. +agy gates every fresh worktree too; the spawn pre-registers it in agy's own store the same way, and a strict post-launch gate answers any dialog that still renders before the spawn reports success. Cursor suppresses its dialog with launch-time `--trust`, and Muse suppresses its own with `--yolo`. Grok dodges its gate instead of granting trust, because its project picker appears only outside a project and the spawn starts in the isolated git root. Pi gates the fresh-worktree case too, but unlike Claude its dialog is answered with Enter, and `references/harness/pi.md` owns that recipe and where the decision persists. Codex shows a directory-trust dialog on the first run for a repository root. -A Claude secondmate is deliberately not pre-registered, because `../../../bin/fm-spawn.sh` runs its per-harness pre-launch setup only for non-secondmate kinds, so the registration is never invoked for one. -That kind guard is the whole exclusion, because a treehouse-leased secondmate home is itself a linked worktree that the scope test would accept, and only a plain-clone home would be refused as a primary checkout. -The consequence is that a claude secondmate whose home Claude has never trusted meets the workspace-trust dialog itself, and firstmate cannot answer it any more than it can for a crewmate. -This is rarely seen because a secondmate home is persistent and reused, so its trust decision is made once and survives, unlike a per-task worktree that is new every time. Use the tool's exact skill form, or natural language only when no separate command is verified or the form remains uncertain. A successful send or key return is not proof of submission; require the tool-specific postcondition. diff --git a/.agents/skills/harness-adapters/references/harness/agy.md b/.agents/skills/harness-adapters/references/harness/agy.md index c52d17bdf2f..21026e92601 100644 --- a/.agents/skills/harness-adapters/references/harness/agy.md +++ b/.agents/skills/harness-adapters/references/harness/agy.md @@ -1,50 +1,56 @@ -# Antigravity CLI (agy) +# Antigravity CLI -Verified 2026-07-23 with agy 1.1.5 on herdr 0.7.4. -`agy` (Antigravity CLI, Gemini) is a captain-approved divergence of Firstmate's tracked surface, so crewmates can use the captain's paid Gemini subscription. -Upstream carries no agy support at all, so every mechanism here is fork-local; `../../../docs/fork-divergence.md` owns that ledger entry. -The full rationale is in `data/captain.md` and the empirical evidence is `data/cursor-agy-verify/report.md`. -Cross-harness provider and credential identity is owned by `references/common/model-and-effort.md`. +Antigravity's `agy` TUI, verified end to end on 2026-09-10 with agy 1.2.0 on Linux through the Herdr backend. +Verified as a CREWMATE and SCOUT adapter on Herdr only; `../../../../../bin/fm-spawn.sh` refuses a secondmate launch on it because `../../../../../docs/supervision-protocols/` carries no agy wake protocol. +`../../../../../docs/verification/agy.md` owns how every fact below was established and what is still unproven. ## Operating facts | Fact | Value | |---|---| -| Launch | The initial prompt stays interactive and tool use is auto-approved by the verified template in `../../../bin/fm-launch-lib.sh`. | -| Busy state | Native herdr `agent_status == working`, generic and with no screen scrape; `../../../bin/fm-busy-lib.sh` owns the identity-gated contract. | -| Turn end | Watcher-side identity-gated, debounced native-idle wake notification; no hook installed and no repo or new global hook file written. | -| Composer state | `unknown`, the safe default, with no override. | -| Model | `--model `. | -| Effort | `--effort `, verified on agy 1.1.5, whose `--help` advertises only these three levels while omitting `xhigh` and `max`. | - -## Crew-only and Herdr-only boundary - -agy is CREW-ONLY and HERDR-ONLY: never a primary runtime, never a secondmate launcher, and never on any non-herdr backend. -`../../../bin/fm-spawn.sh` refuses a `--secondmate` agy spawn and refuses agy on a non-herdr backend, both before any backend or worktree work; `../../../tests/fm-agy-adapter.test.sh` covers the refusals. -tmux is deliberately out of scope: it has no native agent detection, so the liveness, turn-end, and composer signals below would all be absent there. - -The raw-launch escape hatch must never be used for `agy`; use the sanctioned `--harness` path so the required gates and supervision apply. -`../../../bin/fm-launch-lib.sh`'s `fm_launch_raw_restricted_harness` and `fm_launch_write_raw_guard` comments own the early classifier, exec-time PATH-shim invariant, accepted same-user removal residual, and security-boundary rationale. -`../../../tests/fm-launch-lib.test.sh` covers each bypass class, and `../../../tests/fm-agy-adapter.test.sh` covers real-spawn guard installation. - -## Turn end and composer - -agy installs no turn-end hook or status writer; the watcher instead converts its identity-gated, debounced herdr-native idle into the shared `state/.turn-ended` wake notification without treating it as current-state truth or relaxing the event-stream policy. -`../../../bin/fm-transition-lib.sh`'s `fm_transition_native_completion` comment owns the native-identity gate, debounce state machine, and re-arm behavior, while `../../../bin/fm-watch.sh` owns its poll-loop integration. -After the cursor adoption that mechanism serves agy alone. -No repo `.agents/hooks.json` is ever written for agy, and no new shared global hook file is added. - -Composer classification stays `unknown` for agy and no override is added. -agy's prompt shape is Pi's "separated" shape but native identity reports `agy`, not `pi`, so the Pi separated-shape gate in `../../../bin/backends/herdr.sh` correctly rejects it. -A generic bare-glyph "empty" rule must NOT be added: agy's prompt glyph is literally `>`, identical to a dead bash shell, so a generic rule would be a dead-shell send hazard, and any future override must be native-identity-gated exactly like the Pi gate. -The only cost of `unknown` is that the away-mode escalation injector defers rather than injects into an agy pane, which is a minor functional gap and never a safety hole. -`../../../docs/herdr-backend.md` under "Current transport behavior" owns agy prerequisites and atomic-prompt delivery semantics. -Treat `verdict=unverifiable` as possibly accepted and do not blindly resend it. - -## Workspace trust - -agy workspace trust is the one extra launch step. -An interactive agy launch gates on a per-workspace trust modal that `--dangerously-skip-permissions` does NOT cover, and trust is an EXACT-path entry, never a prefix, in agy's SINGLE global settings file `~/.gemini/antigravity-cli/settings.json` under `trustedWorkspaces`. -So the spawn pre-seeds the exact crew-worktree path before launch and teardown removes only a Firstmate-owned entry. -`../../../bin/fm-agy-trust-lib.sh`'s header and function comments own the created-versus-preexisting signal, ownership-and-liveness lock, atomic mutation, abort rollback, teardown ordering, retry-evidence preservation, and refusal behavior. -An agy spawn aborts if that trust write does not land, because an unseeded launch would wedge on the modal. +| Binary | Absolute `agy` from `PATH`, refused if absent; a Go-compiled single binary, so the live process name is exactly `agy` with `argv[0]=agy`. | +| Launch | `agy --prompt-interactive "" --model --effort --dangerously-skip-permissions`, with the resolved absolute binary; the brief auto-submits with no extra Enter. The spawn pre-registers the worktree in agy's trust store first, then waits for a busy turn (answering the folder-trust dialog if it renders anyway) before reporting success. | +| Busy state | No hook or plugin writer, so nothing is armed and no record is seeded; on Herdr only native `working` status classifies busy, and unavailable native evidence stays unknown. | +| Rendered tail | Busy status row carries `esc to cancel` on the left; the idle row shows `? for shortcuts` instead. The `Generating...` word beside the braille spinner is free-floating output and is not a signal. | +| Turn end | No hook is installed; the watcher emits identity-gated, debounced native-idle notifications through `bin/fm-transition-lib.sh`. | +| Exit | `/quit`, one Enter; the process exits. | +| Interrupt | Single `Escape`, which prints the Interrupted row and leaves an idle composer with no repollution, so no clear key follows. | +| Skill | No verified slash-skill form; use natural language. | +| Autonomy | `--dangerously-skip-permissions` auto-approves tool calls for the run. | +| Marker | None; a live TUI carries no `AGY_*` or `ANTIGRAVITY_*` variable. | +| Resume | `--continue` and `--conversation` exist but carry no verified pane-resume contract; use deterministic relaunch. | +| Model | `--model ` with the bare catalog id from `agy models` (for example `gemini-3.8-flash-high`); `bin/fm-spawn.sh` refuses a requested id a reachable listing omits. The listing is a remote fetch, so the probe runs stdin-detached under the shared hard bound and an unreachable or hung listing launches unvalidated with a notice. | +| Effort | `--effort low\|medium\|high`; `xhigh` and `max` stay in task metadata under the record-and-omit contract. | +| Composer | Borderless bare `>` row, which the shared classifier reads as `unknown` under the dead-shell rule, never `empty`; steering confirms delivery through native agent-state and the delivery footer instead, the cursor precedent. | + +## Fork boundaries and trust ownership + +The fork retains the crew-only and Herdr-only spawn gates, raw-launch refusal, atomic prompt delivery, and ownership-aware workspace trust cleanup recorded in [`docs/fork-divergence.md`](../../../../../docs/fork-divergence.md). +`bin/fm-agy-trust-lib.sh` owns trust mutation, created-versus-preexisting ownership, abort rollback, and cleanup retries. +A failed trust write refuses the spawn; successful registration is followed by upstream's readiness gate before dispatch reports success. +`bin/fm-transition-lib.sh` owns native completion notification, and `bin/fm-busy-lib.sh` requires matching native identity before accepting a busy verdict over an untrusted record. +The shared composer classifier keeps a bare `>` unknown, never empty. +`docs/herdr-backend.md` owns atomic-prompt delivery semantics; an `unverifiable` verdict may already have been accepted and must not be blindly resent. + +## Credential precondition + +A verified agy worker ran under a signed-in Google account with no key export and no dialog. +The unauthenticated failure mode was not observed, so treat any auth prompt or refusal as a credential blocker under `../../../../../AGENTS.md` section 9, fix the environment, and retire the endpoint rather than typing into it. + +## Detection + +Detected by ancestry alone: `../../../../../bin/fm-harness.sh` matches the anchored process name `agy`, never `*agy*`. +No environment marker is promoted: `AGENT=1` observed on a live TUI is an inherited launcher value, not an agy identity, and agy does not clear an inherited `CLAUDECODE` - but a structural agy ancestor now outranks that retained marker, which `../../../../../bin/fm-harness.sh` decides without depending on the spawn's own launch-boundary marker clearing. +agy is deliberately absent from the session-lock name vocabulary in `../../../../../bin/fm-session-lock-lib.sh`, where muse, gemini, and rovo are also absent: a crewmate-only adapter must never own a home session lock. + +## Worker busy state and turn end + +`../../../../../bin/fm-spawn.sh` arms no busy generation for agy and writes no sidecar, exactly because no writer could ever clear a seeded record. +The shared rendered-tail helper remains available for standalone non-Herdr observations, but does not authorize a supervised Herdr busy verdict. +Teardown removes only the workspace trust entry this task created, using the durable ownership marker; preexisting trust remains untouched. + +## Primary integration + +Unsupported and unverified. +`../../../../../docs/supervision-protocols/` carries no agy protocol, no turn-end guard adapter exists for it, and this adapter verified only the crewmate-side launch, busy state, interrupt, and exit. +`references/common/primary-hooks.md`'s unsupported-boundary rule applies: never invent a wake protocol from a similar TUI. diff --git a/.agents/skills/harness-adapters/references/harness/claude.md b/.agents/skills/harness-adapters/references/harness/claude.md index c9b9834f33b..f29720ca411 100644 --- a/.agents/skills/harness-adapters/references/harness/claude.md +++ b/.agents/skills/harness-adapters/references/harness/claude.md @@ -12,22 +12,33 @@ Busy hooks verified 2026-07-28 on Claude Code 2.1.220. | Skill | `/`, for example `/no-mistakes`. | | Model | `--model `; discover through the interactive `/model` picker, with alias or full-name shape documented by `claude --help`. | | Effort | `--effort `, verified on 2.1.196. | +| Permissions | `--dangerously-skip-permissions` by default, or `--permission-mode auto` when `config/claude-permission-mode` is `auto`; the `auto` shape verified on 2.1.269, and `../../../../../docs/configuration.md` "Claude permission mode" owns the file. | ## Workspace trust -Claude gates a folder it has never seen behind an interactive workspace-trust dialog, so every fresh task worktree would hit it. -`--dangerously-skip-permissions` does not cover that gate: `claude --help` records that the dialog is skipped only in non-interactive mode, through `-p` or a non-TTY stdout, and a crewmate pane is interactive. -A ship or scout spawn therefore pre-registers the worktree before launch, and the dialog does not appear. -`../../../bin/fm-claude-trust.sh` records `hasTrustDialogAccepted` for that worktree path in `${CLAUDE_CONFIG_DIR:-$HOME}/.claude.json`, and `../../../bin/fm-spawn.sh` refuses the spawn when the write fails rather than launching a worker that would wedge. +Claude gates a folder it has never seen behind an interactive workspace-trust dialog (titled "Quick safety check: Is this a project you created or one you trust?"), so every fresh task worktree would hit it, and so would every secondmate home no operator has opened by hand. +`--dangerously-skip-permissions` does not cover that gate: `claude --help` records that the dialog is skipped only in non-interactive mode, through `-p` or a non-TTY stdout, and a spawned pane is interactive. +Every claude spawn therefore pre-registers the directory its pane starts in before launch, and the dialog does not appear: the task worktree for a ship or scout, and the home itself for a `--secondmate` spawn, in either seeded shape (a leased worktree or a standalone clone). -Never try to answer the trust dialog with a key. -Firstmate's key plane carries only Enter, Escape, and C-c with no arrow navigation, so it cannot move a dialog's selection at all, and the observed rendering starts on `No, exit`, which means a sent Enter ends the session instead of accepting. -A visible trust dialog means pre-registration did not take effect, so inspect the store and the spawn's error output rather than sending keys. +A second, separate dialog - "Allow external CLAUDE.md file imports?" - renders whenever a loaded CLAUDE.md chain reaches outside the project tree, which every crewmate's does through the captain's own `~/.claude/CLAUDE.md` importing `~/.claude/RTK.md`. +`--setting-sources project,local` (the minimal worker tool surface) does not suppress it either, and it gates the pane exactly like the trust dialog: cursor on "No, disable external imports", no way to move the selection from firstmate's steering plane. -The once-per-machine bypass-permissions confirmation is a separate dialog, scoped to the machine rather than the path, and pre-registration does not address it. +`../../../bin/fm-claude-trust.sh` records `hasTrustDialogAccepted` for both the worktree and its primary checkout in `${CLAUDE_CONFIG_DIR:-$HOME}/.claude.json` for a ship or scout spawn; a secondmate spawn registers only its own home entry, since a secondmate home has no separate primary-checkout entry to carry import consent forward from. +For a ship or scout spawn, the external-imports flags (`hasClaudeMdExternalIncludesApproved`, `hasClaudeMdExternalIncludesWarningShown`) are carried forward alongside the trust flag only when the primary checkout's project entry already carries an explicit `hasClaudeMdExternalIncludesApproved===true` from a prior interactive session - the common first-spawn case is a project claude has never been asked about, so those two flags are left unwritten and the import dialog still renders, even though trust registers normally. +When the project entry instead already carries an explicit decline (`===false`), the whole registration refuses - including the trust flag - rather than manufacture consent the human never gave, so that spawn wedges on the trust dialog before it would even reach the import one. +The why-two-entries mechanism and the consent-gating logic live in the script's own header comment, which is the one owner for that contract; the fact worth repeating here is that `../../../bin/fm-spawn.sh` refuses the spawn when the trust flag fails to land, rather than launching a worker that would wedge on that dialog. + +Never try to answer either dialog with a key. +Firstmate's key plane carries only Enter, Escape, and C-c with no arrow navigation, so it cannot move a dialog's selection at all, and both dialogs render with the cursor on their declining option, which means a sent Enter ends the session instead of accepting. +A visible trust dialog means pre-registration did not take effect (or the project entry already carries an explicit decline) - inspect the store and the spawn's error output rather than sending keys. +A visible external-imports dialog is expected, not a failure signal, whenever the project entry has no prior explicit approval on record - the common first-spawn case; `fm-control.sh interrupt` delivers Escape, which dismisses whichever of the two is on screen without answering it, and is the safe way to clear a wedged pane for inspection. + +The once-per-machine bypass-permissions confirmation is a third, separate dialog, scoped to the machine rather than the path, and pre-registration does not address it. Never send Enter to that one either: it was observed rendering in the same shape as the trust dialog, with the selection on `No, exit` and the footer `Enter to confirm . Esc to cancel`, so Enter ends the session rather than accepting. Firstmate cannot move a selection with Enter, Escape, and C-c alone, so it cannot accept this dialog at all, and an operator accepts it once per machine instead. Inspect the pane to identify which dialog is on screen, and report it rather than answering it. +A launch under `config/claude-permission-mode=auto` never meets the bypass confirmation, because it does not request bypass mode: on 2.1.269 `claude --permission-mode auto` reached the composer directly with the footer `⏵⏵ auto mode on (shift+tab to cycle)`, so a captain who refuses the bypass dialog selects `auto` there instead of accepting it. +The workspace-trust dialog is unaffected by the permission mode and still needs the pre-registration above. ## Composer ghost diff --git a/.agents/skills/harness-adapters/references/harness/codex.md b/.agents/skills/harness-adapters/references/harness/codex.md index 5fb95b8e494..368afadddf9 100644 --- a/.agents/skills/harness-adapters/references/harness/codex.md +++ b/.agents/skills/harness-adapters/references/harness/codex.md @@ -14,6 +14,7 @@ Verified on 2026-06-11 with codex-cli 0.139.0 unless a fact gives a newer versio | Model flag | `--model `. | | Effort flag | `-c 'model_reasoning_effort=""'`, verified on codex-cli 0.142.1 whose installed schema contains `model_reasoning_effort`, active config uses it, and bundled catalog advertises only these four values while omitting `max`. | | Model discovery | Open the current interactive session's `/model` picker. | +| Marker | None; identity comes from ancestry, and `../../../bin/fm-harness.sh` is what keeps a retained foreign `CLAUDECODE` from renaming it. Verified on 2026-09-01 with codex-cli 0.152.0: the pane process is the `node` npm shim and the native `codex` binary runs as its foreground child, so a tool subprocess reaches the native name directly while the shim itself is identified from its script path. | A directory trust dialog appears on the first run for a repository root: "Do you trust the contents of this directory?" Accept it with Enter and verify the instructions begin processing. diff --git a/.agents/skills/harness-adapters/references/harness/cursor.md b/.agents/skills/harness-adapters/references/harness/cursor.md index 3048a0a8347..0bdede0f20a 100644 --- a/.agents/skills/harness-adapters/references/harness/cursor.md +++ b/.agents/skills/harness-adapters/references/harness/cursor.md @@ -28,6 +28,7 @@ The slash popup consumes the first Enter; that Enter closes it and a genuine sec Cursor does not clear inherited `CLAUDECODE`, so a Cursor worker under Claude carries both markers. `../../../bin/fm-harness.sh` tests Cursor first, and launch also clears foreign markers. Both remain necessary: sanitization covers Firstmate launches, ordering covers hand-started sessions. +That ordering settles the marker layer only, and a nearer Claude ancestor still outranks a retained Cursor marker. Cursor is a bundled Node script, so tmux can report bare `node` while `ps -o comm=` carries its install path. Bare `node` matches nothing; `../../../bin/fm-cursor-lib.sh` proves identity from Cursor's name or install tree in path or argv zero. diff --git a/.agents/skills/harness-adapters/references/harness/kimi.md b/.agents/skills/harness-adapters/references/harness/kimi.md index 8b61d813e48..8799f8bcfcd 100644 --- a/.agents/skills/harness-adapters/references/harness/kimi.md +++ b/.agents/skills/harness-adapters/references/harness/kimi.md @@ -16,7 +16,7 @@ Verified on 2026-07-25 with Kimi Code CLI 0.29.1. | Autonomy | `--auto`; `-y` and `--yolo` are weaker and are not used. | | Trust dialog | None observed on a clean first launch in a fresh pooled worktree. | | Slash submission | One Enter submits, with no popup swallow or settle hazard. | -| Environment marker | None; detection uses process ancestry command name `kimi`. | +| Environment marker | None; identity comes from process ancestry command name `kimi`, which `../../../bin/fm-harness.sh` keeps a retained foreign marker from overriding. | | Composer | Bordered box with a bare `>` prompt glyph and no observed ghost or placeholder text. | | Effort | No verified reasoning-effort flag; `references/common/model-and-effort.md` owns unsupported-value handling. | diff --git a/.agents/skills/harness-adapters/references/harness/muse.md b/.agents/skills/harness-adapters/references/harness/muse.md index a0a9df3e004..d00258e1f45 100644 --- a/.agents/skills/harness-adapters/references/harness/muse.md +++ b/.agents/skills/harness-adapters/references/harness/muse.md @@ -17,7 +17,7 @@ The router owns Muse's task-kind boundary. | Resume | `muse resume --last` or `muse resume `; bare `muse resume` opens a picker. | | Autonomy | `--yolo` disables approval and sandbox and trusts the workspace. | | Trust | Dialog `Do you trust this workspace?`, choice `1 Trust and continue` preselected for Enter; `--yolo` suppresses it, which fresh task paths require. | -| Marker | None; detect anchored `muse-bin-*` ancestry after clearing foreign primary markers, while `MUSE_CURRENT_SESSION_LOG` is a path rather than identity and its export to tools is unverified. | +| Marker | None; identity comes from anchored `muse-bin-*` ancestry, which `../../../bin/fm-harness.sh` keeps a retained foreign marker from overriding, while `MUSE_CURRENT_SESSION_LOG` is a path rather than identity and its export to tools is unverified. | | Composer | Bordered `⟩`, truecolor `38;2;90;160;255`, luminance about 149.9 and narrowly above ghost threshold 128; typed text is `38;2;204;211;219`, about 209.8, with no observed placeholder or ghost. | | Effort | `--reasoning-effort`, default `high`, accepts `none\|minimal\|low\|medium\|high\|xhigh\|ultra`; shared values expose low through xhigh, explicit captain `max` maps to `ultra`, and `none` or `minimal` remain unreachable. | diff --git a/.agents/skills/harness-adapters/references/harness/opencode.md b/.agents/skills/harness-adapters/references/harness/opencode.md index 8b37a8d35ad..ca9ff18b3f5 100644 --- a/.agents/skills/harness-adapters/references/harness/opencode.md +++ b/.agents/skills/harness-adapters/references/harness/opencode.md @@ -15,6 +15,7 @@ Verified on 2026-06-11 across versions 1.15.7 through 1.17.6, with busy-queue be | Effort flag | None for Firstmate's interactive `opencode --prompt` launch verified on 1.17.6; `opencode run` has `--variant`, but that is not this path. | | Model discovery | Run `opencode models [provider]` to list available provider/model identifiers. | | Trust dialog | None. | +| Marker | None; OpenCode publishes no identity marker, so `../../../bin/fm-harness.sh` identifies it from process ancestry. | OpenCode can auto-upgrade in the background, and the running TUI can exit mid-task. That behavior was observed live during an upgrade from 1.15.7 to 1.17.3. diff --git a/.agents/skills/process-event-sources/SKILL.md b/.agents/skills/process-event-sources/SKILL.md index a3a6cf8a5f7..050ba1a025d 100644 --- a/.agents/skills/process-event-sources/SKILL.md +++ b/.agents/skills/process-event-sources/SKILL.md @@ -3,8 +3,10 @@ name: process-event-sources description: >- Agent-only procedure for registered process-to-event sources and their wakes. Use before arming a long-polling source firstmate owns, before registering a - deterministic condition->action watch, and on any - `procevent ` check wake. + deterministic condition->action watch, on any + `procevent ` check wake, and on any + `process-event source stranded` or `process-event source failed to start` + check wake. Owns the arming commands, the condition->action eligibility boundary, the durable result read, which wakes must be routed to their adapter instead of acknowledged generically, the handled acknowledgement contract, the one-owner @@ -17,7 +19,7 @@ metadata: # process-event-sources -Load this before arming a long-polling source, before registering a deterministic condition->action watch, and whenever a `check:` wake carries `procevent `. +Load this before arming a long-polling source, before registering a deterministic condition->action watch, whenever a `check:` wake carries `procevent `, and whenever the watcher headlines a `process-event source stranded` or `process-event source failed to start` wake. The runner exists so a blocking external process never holds firstmate's conversational turn. Firstmate registers a source, keeps working, and is woken when that process completes. @@ -31,6 +33,14 @@ For a Lavish review artifact firstmate owns (a live investigating scout should h bin/fm-procevent-lavish.sh arm ``` +Registering a source is not the same fact as listening to it: arming records the source, and a separate runner still has to pick it up. +After arming by hand, confirm `bin/fm-procevent.sh list` reports that source as `live`, and run `bin/fm-procevent.sh reconcile` when it does not. +Reconcile reports every launch that did not prove it took its claim within the confirm window as `failed=` and exits non-zero, so a source that cannot be started says so instead of looking armed, and it wakes you once per failure episode about it because the watcher discards that count; `start` does not fix that - if the source stays unowned, run `start` attached to read the runner's refusal, then check the source command and adapter binary the registration names, and if a later reconcile finds the source owned the episode closes on its own. +A source `list` reports as `orphaned` is one reconcile will not relaunch, because something may still be polling it; reconcile wakes you once about it, and that wake's payload says which of two recoveries applies. +If the claim's recorded pid is alive under a different identity, `bin/fm-procevent.sh start ` takes the source back once you have checked nothing is still polling it - provided the dead generation's reservation records can still be tidied; otherwise it refuses with `cannot claim source`. +If the runner itself died and its process group survives, `start` reports `already owned` and takes nothing back: verify whether the dead runner's polling child is still attached to the source, and once that group is empty the next reconcile reclaims the source on its own. +Nothing signals that group automatically. + When a source carries captain answers to captain-held tasks, bind it BEFORE arming it, so it can never produce an answer that has nowhere to go: ```sh @@ -108,6 +118,10 @@ Two rules the commands cannot enforce for you: : Never append a raw result to a task's status history; that log is a bounded event record, not a payload channel. : A source whose adapter returns a terminal verdict for the captured result has already retired itself, so an ended review needs no cleanup from you and produces no further wake. Retire any other finished source with the adapter's `retire`, which accepts an artifact path or its source id, and stays retirable after the artifact is deleted: a gone path resolves to the id it registered while the file existed. Retire by source id is always safe to repeat, and so is retire by a path that still resolves. Retire by a *gone* path repeats safely once a result has been captured for that source - which an `artifact-missing` source always has - because the captured result is the durable proof the id was registered here; a source that never produced one leaves nothing behind after its registration is removed, so a second retire by that gone path refuses with guidance rather than silently no-op'ing a path that may never have been armed. When in doubt, re-retire by source id. For an `artifact-missing` source the id is in the wake text itself (`procevent lavish `) or `bin/fm-procevent.sh list`. Retirement stops future completions; it is independent of acknowledging a result already captured, which only `handled` does. +`process-event source stranded` or `process-event source failed to start` (queue keys `procevent::stranded:` and `procevent::launch-failed:-`) +: Nothing was captured: the source named in the payload is registered but nothing is confirmed to be collecting from it. There is no result file to read and no `handled` call to make; the ordinary drain acknowledgement consumes the row. +: The payload says which shape it is and what clears it. Follow it exactly as the arming section above describes - a `start` is named only for the reused-pid strand, a leaderless group is a human check and reclaims itself once its group is empty, and a launch that never proved its claim closes its own episode if a later cycle finds the source owned. + ## What the runner guarantees, exactly Supported by tests: diff --git a/.agents/skills/quiet/SKILL.md b/.agents/skills/quiet/SKILL.md new file mode 100644 index 00000000000..dc6560a3748 --- /dev/null +++ b/.agents/skills/quiet/SKILL.md @@ -0,0 +1,50 @@ +--- +name: quiet +description: >- + Enter quiet supervision mode when the captain invokes /quiet or asks for quiet mode, quiet-while-present, or fewer routine wake turns while they stay in the session. + It sets the same durable away/quiet-mode flag as /afk, in `quiet` mode, so the sub-supervisor daemon self-handles routine wakes and escalates captain-relevant events exactly as away mode does, but ordinary captain chat does NOT exit it - only an explicit `/quiet off` does. +user-invocable: true +metadata: + internal: true +--- + +# quiet + +Quiet supervision mode (kunchenguid/firstmate#2356): the same token-saving daemon tradeoff as `/afk`, made explicit for a captain who is staying, watching the session, and does not want to exit the mode just by chatting. + +This skill is a thin wrapper. +Every mechanism below - the daemon, its injection, its busy/composer guards, its classification policy, its reliability properties - is owned once by the `afk` skill and is IDENTICAL in quiet mode; nothing here restates it. +Quiet uses an attended entry and explicit exit without an away-posture record; the launch and return scripts own those lifecycle differences. + +## What it does + +1. **Enter the shared lifecycle through `bin/fm-afk-launch.sh` with `FM_AFK_MODE=quiet`.** + Use the `afk` skill's existing terminal-backed or harness-native daemon launch procedure with `FM_AFK_MODE=quiet` on `start` or `start-native`. + Do not propose or confirm an away-posture record: quiet is attended and grants no away authority. + Finish an existing away return and its catch-up gate before quiet entry. + Refreshing an existing quiet daemon without an explicit mode preserves quiet. + The one daemon continues to own supervision, including Pi and OMP extension standby. + +2. **Acknowledge** in `AGENTS.md` section 9 language: "Captain, quiet mode is active; I will batch routine updates and surface only decisions, failures, credentials, or review-ready work - ordinary chat will not exit this, say `/quiet off` when you want normal per-wake responses back." + +## How to exit quiet mode + +Unlike `/afk`, ordinary chat is never the exit signal - that is the entire point of this mode (AGENTS.md section 8's away-mode stub, quiet branch). + +- Only an explicit `/quiet off` or a plain request to leave quiet mode exits it. + Run `bin/fm-afk-return.sh quiet-off` to stop the existing daemon and clear the quiet flag through the shared teardown owner. + Quiet exit renders no away return brief, archives no fictional away record, and creates no away catch-up gate. + The ordinary away return path remains required when an actual away record exists. +- A marked daemon escalation, or a message beginning `/quiet` while already in quiet mode (refresh, not exit) -> stay in quiet mode and process it, the same two carve-outs `/afk` documents for away mode. +- Every other message while in quiet mode is simply answered as ordinary work; the flag and daemon are left untouched. + +## Orthogonal to approval authority + +Identical to `/afk`: quiet mode changes how aggressively firstmate surfaces things, never who approves what. +A PR ready for merge keeps the merge authority from `AGENTS.md` section 7, and a needs-decision finding keeps the `ask-user-authority` policy. + +## Must not hide a decision or a failure + +Per the issue's own author triage: quiet mode is presentation only. +Progress, retries, and internal mechanics stay below deck exactly as in away mode, but review-ready work, findings, decisions, failures, and credentials escalate every time, through the same classification policy `/afk` owns. +Quiet mode is opt-in and never the unconsented default; only an explicit `/quiet` invocation enters it. diff --git a/.agents/skills/stow/SKILL.md b/.agents/skills/stow/SKILL.md index ccb793f4897..85e6c957dad 100644 --- a/.agents/skills/stow/SKILL.md +++ b/.agents/skills/stow/SKILL.md @@ -282,7 +282,7 @@ This sequence owns every user-owned local skill destination creation, whether in Autonomously relocate it only by adding it to an already-existing allowed JIT note, by applying the local destination creation contract when Destinations allows the one bootstrap or requires a validated weight or fit split, or by routing it through a project's established delivery path to its existing owning `AGENTS.md`, then confirming that destination holds the quoted entry before removing the memory entry. A destination that needs creation is not live until the allowed bootstrap or validated split completes the local destination creation contract; uncompleted project delivery or any other future work still cannot count as relief, so continue with the next archival or eviction rung instead of leaving an over-budget proposal pending. 2. Propose pinned relocation only. - For a pinned candidate, append a `proposed-offload` section with the same fields to the completion receipt, create or refresh one durable backlog item with `tasks-axi add`, `tasks-axi show --full`, and `tasks-axi update --body-file ` as appropriate, then hold it through `bin/fm-captain-hold.sh hold`. + For a pinned candidate, append a `proposed-offload` section with the same fields to the completion receipt, create or refresh one durable backlog item with `bin/fm-tasks-axi.sh add`, `bin/fm-tasks-axi.sh show --full`, and `bin/fm-tasks-axi.sh update --body-file ` as appropriate, then hold it through `bin/fm-captain-hold.sh hold`. Preserve each candidate's approval state in that item, and require explicit plain-chat approval for that named item before any migration. If the captain never answers, nothing migrates and the held item persists, but it is never treated as budget relief. 3. Migrate an approved pinned candidate outside this pass. @@ -309,7 +309,7 @@ This sequence owns every user-owned local skill destination creation, whether in - Project-intrinsic knowledge never goes directly into a project's `AGENTS.md`. Route it through a normal ship task so a crewmate records it with `bin/fm-ensure-agents-md.sh` and the project's delivery path. - Knowledge general to every Firstmate user belongs in this repo's shared tracked material through the normal branch, no-mistakes, PR, and captain-merge path. - - For task-scoped notes, inspect the item with `tasks-axi show --full`, classify the change as new, duplicate, superseding, or obsolete, then use a considered replacement body through `tasks-axi update --body-file `. + - For task-scoped notes, inspect the item with `bin/fm-tasks-axi.sh show --full`, classify the change as new, duplicate, superseding, or obsolete, then use a considered replacement body through `bin/fm-tasks-axi.sh update --body-file `. Use `--archive-body` when recoverability matters. Never append. - File each undone next step as a queued backlog item with a genuine `blocked-by` dependency when applicable. diff --git a/.agents/skills/sync-upstream/SKILL.md b/.agents/skills/sync-upstream/SKILL.md index 27c1365e586..6ded28772ba 100644 --- a/.agents/skills/sync-upstream/SKILL.md +++ b/.agents/skills/sync-upstream/SKILL.md @@ -69,8 +69,12 @@ A sync PR is expected to carry the red `PR must be raised via no-mistakes` check The round deliberately uses `direct-PR` because no-mistakes rebases onto `origin/main`, which would replay and linearize a merge-only branch. Every other required check must pass, the PR must be mergeable, and the body must carry the applicability table and fork-survival evidence before it is reported ready. -Register the ready PR through the normal task lifecycle and stop for the configured merge authority. -When that authority later approves landing, the normal merge handler must use: +Firstmate has standing authority to land upstream-sync rounds after its review of tests, lint, fork-preservation evidence and passing substantive CI. +The worker still reports the ready PR through the normal task lifecycle and stops without merging. +After that review, Firstmate records the exact round review declaration in the task metadata as specified by `github_verified_upstream_sync` in [`bin/fm-pr-merge.sh`](../../../bin/fm-pr-merge.sh). +That guard owns the narrow expected-policy-check exception and its identity and ancestry proofs; any new head needs a fresh review declaration. +Every other merge gate, including the existing away-authority gate, remains applicable. +The landing handler must use: ```sh bin/fm-pr-merge.sh -- --merge diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index ab6442f625c..99605b31a80 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -9,10 +9,26 @@ on: permissions: contents: read +# Per-PR supersession: a new push to the same PR replaces that PR's in-flight +# CI instead of letting superseded heads keep 13 jobs of hosted-runner work. +# The group uses the PR number for pull_request events, so every run of one PR +# shares a group, and falls back to the unique run id for push events, so each +# main push gets its own group and is never cancelled. Cancellation is likewise +# limited to pull_request events. Evidence and rationale: the September 12 +# Actions starvation report, section 3 "Workflow mechanics to ship first". +# The compliance workflow deliberately keeps its own event-specific groups; do +# not collapse it onto this simpler shape. +concurrency: + group: ci-${{ github.workflow }}-${{ github.event_name }}-${{ github.event.pull_request.number || github.run_id }} + cancel-in-progress: ${{ github.event_name == 'pull_request' }} + jobs: lint: name: Lint runs-on: ubuntu-latest + # Hang tripwire only: lint executions measured at 14-16 minutes in the + # September 12 starvation report, so this leaves deliberate margin. + timeout-minutes: 25 steps: - uses: actions/checkout@v6 - name: Install pinned ShellCheck @@ -37,6 +53,8 @@ jobs: test-coverage: name: Test coverage guard runs-on: ubuntu-latest + # Hang tripwire: the coverage guard is a seconds-long local computation. + timeout-minutes: 5 steps: - uses: actions/checkout@v6 - name: Prove complete regression partition @@ -44,11 +62,18 @@ jobs: # Two duration-balanced portable parallel shards of the Phase 2 proven-isolated # set only. Composition owner: bin/fm-test-run.sh (docs/fm-test-portable-shards.md). + # Two admitted workers overlap this isolated work inside each existing job. + # docs/fm-test-isolation-proof.md owns the concurrency evidence. tests-portable-parallel-1: name: Behavior portable parallel 1 runs-on: ubuntu-latest - # Measured shard wall is ~7 min of serial sum on proven scripts; this cap - # is a hang tripwire with margin, not the expected healthy end of the lane. + # This cap is intended as a hang tripwire, but the previous lane 1 reached + # it; the former "~1 min of serial sum" estimate no longer applies. + # Compare it with the derived hints from fm-test-run.sh --check-coverage + # and completed job timings, allowing for setup and runner-speed spread. + # A packed hint sum is not a measured job wall time or proof of headroom. + # Evidence and refresh procedure: docs/fm-test-portable-shards.md. + # Changes to this cap or the lane count require a separate scope decision. timeout-minutes: 10 steps: - uses: actions/checkout@v6 @@ -78,7 +103,7 @@ jobs: run: | set -eu mkdir -p "$RUNNER_TEMP/fm-test" - bin/fm-test-run.sh --lane portable-parallel-1 \ + bin/fm-test-run.sh --lane portable-parallel-1 --jobs 2 \ --fail-on-gate-skip 'Pi extension typecheck prerequisite not found' \ --json "$RUNNER_TEMP/fm-test/fm-test-timing-portable-parallel-1.json" - name: Upload shard 1 timing artifact @@ -92,6 +117,7 @@ jobs: tests-portable-parallel-2: name: Behavior portable parallel 2 runs-on: ubuntu-latest + # Same timeout rationale as portable parallel shard 1 above. timeout-minutes: 10 steps: - uses: actions/checkout@v6 @@ -116,7 +142,7 @@ jobs: run: | set -eu mkdir -p "$RUNNER_TEMP/fm-test" - bin/fm-test-run.sh --lane portable-parallel-2 \ + bin/fm-test-run.sh --lane portable-parallel-2 --jobs 2 \ --json "$RUNNER_TEMP/fm-test/fm-test-timing-portable-parallel-2.json" - name: Upload shard 2 timing artifact if: always() @@ -337,6 +363,8 @@ jobs: tests-timing-aggregate: name: Behavior timing aggregate runs-on: ubuntu-latest + # Hang tripwire: aggregation is seconds of work over lane artifacts. + timeout-minutes: 5 needs: - tests-portable-parallel-1 - tests-portable-parallel-2 @@ -421,8 +449,8 @@ jobs: bearings_output=$(/bin/bash tests/fm-bearings-snapshot.test.sh) printf '%s\n' "$bearings_output" bearings_count=$(printf '%s\n' "$bearings_output" | grep -c '^ok - ') - [ "$bearings_count" -ge 56 ] || { - echo "::error::expected at least 56 Bearings tests, got $bearings_count" + [ "$bearings_count" -ge 59 ] || { + echo "::error::expected at least 59 Bearings tests, got $bearings_count" exit 1 } @@ -440,6 +468,8 @@ jobs: invariants: name: Repo invariants runs-on: ubuntu-latest + # Hang tripwire: the invariant checks are seconds-long file comparisons. + timeout-minutes: 5 steps: - uses: actions/checkout@v6 - name: Compatibility pointers must stay intact diff --git a/AGENTS.md b/AGENTS.md index c65d043390e..411e5bafd23 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -97,6 +97,7 @@ config/gbrain-local.json this home's own brain-root override and OAuth client i config/gbrain-secrets/ one brain credential per file; LOCAL, gitignored, must be mode 0600, and deliberately NOT inherited (docs/configuration.md "Brain scoping") config/wedge-alarm optional away-mode wedge-alarm active-alert directives; LOCAL, gitignored (docs/wedge-alarm.md) config/watched-tools.json optional list of the tools this home depends on, read by the update check armed with bin/fm-tool-update-check.sh; LOCAL, gitignored, firstmate-maintained but human-editable, and NOT inherited by secondmate homes (docs/configuration.md "Watched tool updates") +config/claude-permission-mode optional Claude worker permission posture; LOCAL, gitignored; inherited by secondmate homes (docs/configuration.md "Claude permission mode") config/x-mode.env generated Relay watcher cadence; LOCAL, gitignored; source before arming watcher when present data/ personal fleet records; LOCAL, gitignored as a whole backlog.md task queue, dependencies, history @@ -139,6 +140,7 @@ state/ runtime records and signals; gitignored .pr-poll private validated data sidecar for the byte-static PR merge poll .pr-poll-registration private transactional provenance record binding the task, canonical metadata identity, sidecar, and static poll publication (bin/fm-pr-lib.sh) .pr-poll-retirement private identity-bound crash-recovery receipt for one exact validated merged result; removed after its poll artifacts retire (bin/fm-pr-lib.sh) + .merge-authority private canonical-PR-bound authority persisted after firstmate's forge merge request is accepted and consumed by a later merged poll; bin/fm-merge-authority-lib.sh owns its format and lifecycle .pr-poll-merge-notified canonical PR identity of the last merge outcome delivered for this task; bin/fm-pr-lib.sh owns the marker format and identity mechanics, while bin/fm-merge-outcome-lib.sh owns locked publication, duplicate suppression, and replacement branch-outcomes.jsonl .branch-outcomes-cursor .branch-outcomes-processed ..branch-outcome-index .branch-outcome-index-ready Pi supervision-branch durable outcome store, its read cursor, main's processed marker, bounded latest per-task status-coverage caches, and their recovery marker; bin/fm-branch-outcome.sh owns the formats branch-session/ .branch-session .branch-mirror-cursor the branch's per-main-session conversations, the pointer to the current one, and the dialog-mirror cursor; extension-owned (docs/pi-supervision-branch.md) @@ -170,9 +172,9 @@ state/ runtime records and signals; gitignored .watcher-down private generation-bound recovery state coupling watcher downtime, durable wake presentation, and post-handling acknowledgement; never touch ..open-decisions-cursor per-task byte cursor and folded open-decision set bounding the OPEN DECISIONS scan's cost to new status-log appends; written only by fm-classify-lib.sh's status_open_decisions_incremental, removed by teardown, safe to delete (forces one full re-fold) .status-presentation-cursor .status-presentation-lock fleet-wide per-task status identity plus independent annotation and outcome-backstop byte offsets, with a serialization lock preventing already-presented lines from replaying while preserving delayed signal annotations; owned by fm-classify-lib.sh, with each task's row retired by teardown - .afk-contract the away-posture record: the captain's verbatim away words, expected return, reach profile, spend cap, and structured mandate clauses; written only by bin/fm-afk-contract.sh after the captain confirms the read-back, archived under afk-contracts/ at return; its presence IS the away posture in every harness + .afk-contract the away-posture record: the captain's verbatim away words, expected return, reach profile, spend cap, and structured mandate clauses; written only by bin/fm-afk-contract.sh after the captain confirms the read-back, archived under afk-contracts/ at return; its presence IS the away posture in every harness; its sibling .afk-contract.lock serializes actions authorized by the live record (contract: bin/fm-afk-contract.sh) afk-contracts/ archived away-posture records: one final record per away window keyed by entry time, plus any superseded mandates from that window - .afk durable away-mode daemon flag; present = sub-supervisor may inject escalations while Pi/OMP extension supervision stands by (set by the daemon entry, cleared on user return) + .afk durable away/quiet-mode daemon flag; its mode and exit semantics are owned by fm_classify_afk_mode in bin/fm-classify-lib.sh; Pi/OMP extension supervision stands by while the daemon owns the cycle .watch.lock .wake-queue.lock watcher singleton and queue serialization locks .claude-autoarm.lock .claude-autoarm-epoch .claude-autoarm-failure-notified .claude-autoarm-failure-alarmed .turnend-claude-blocks .turnend-claude-blocks.lock Claude Stop auto-arm single-flight, epoch, failure-episode, attended-alarm, guard-budget, and budget-lock records; never touch .cursor-park-owner .cursor-park-owner.lock .turnend-cursor-blocks Cursor stop-hook owner record, publication and commit lock, and bounded repair-nag budget; never touch @@ -219,8 +221,8 @@ The fleet-state digest's endpoint-liveness line is a fast presence check only, n A context file that does not exist prints an explicit `ABSENT` marker, never confused with an empty-but-present file. Bootstrap detects first, asks for consent, and installs only after the captain approves in the current session. -Do not dispatch until the required tools are present and GitHub authentication is good. -Use `gh-axi` for GitHub, `chrome-devtools-axi` for browser work, and `lavish-axi` for structured decisions or reports; consult current help rather than memorizing flags. +Do not dispatch until the essential launch tools are present and GitHub authentication is good; presentation availability follows `bootstrap-diagnostics` and does not block nonvisual work. +Use `gh-axi` for GitHub, `chrome-devtools-axi` for browser work, and compatible `lavish-axi` for visual decisions or reports; consult current help rather than memorizing flags. A silent bootstrap section needs no action; for any printed actionable diagnostic line, load `bootstrap-diagnostics` and follow its owner procedure. `BOOTSTRAP_INFO:` lines are completed no-action facts and do not require loading a skill. `secondmate-provisioning` owns startup secondmate sync, liveness, and inherited local-material convergence. @@ -263,7 +265,7 @@ For an ordinary direct report whose endpoint is dead or metadata has no window, For a dead secondmate direct report, load `secondmate-provisioning` and reconcile only that secondmate, never its whole child tree from the main home. Each secondmate reconciles work already in its own home and then idles; recovery never authorizes it to invent work. -If away mode is present, load `/afk` and let the daemon own supervision rather than arming another cycle; `docs/watcher-continuity.md` owns the Pi and OMP extension handoff. +If `state/.afk` is present, load `/afk` in away mode or `/quiet` in quiet mode (`bin/fm-classify-lib.sh`'s `fm_classify_afk_mode`); let the daemon own supervision rather than arming another cycle; `docs/watcher-continuity.md` owns the Pi and OMP extension handoff. Surface only captain-relevant decisions, review-ready PRs, failures, and credential needs; otherwise resume the emitted supervision protocol silently. A restart must be a non-event because durable state and live backend inventory, not conversation memory, are authoritative. @@ -379,8 +381,9 @@ The path's worker, automated gates, and captain approval remain authoritative: Delivery mode and `yolo` are orthogonal. `yolo` governs merge authority only: with it off, the captain approves every PR merge and every local-only landing; with it on, firstmate merges green, in-scope work itself. -Never merge a red PR under either setting; destructive, irreversible, and security-sensitive merges still escalate. -Without a current explicit captain instruction that states the concrete merge, that default stands, and standing `yolo` cannot authorize a red merge; section 1 owns when such an instruction overrides a Firstmate-written standing rule within its exact scope. +Never merge a red PR under either setting unless a current explicit captain instruction names the single GitHub check waived through `fm-pr-merge.sh --allow-red`; that attended-only waiver still requires every other check green. +Destructive, irreversible, and security-sensitive merges still escalate. +Without a current explicit captain instruction that states the concrete merge, the green default stands, and standing `yolo` cannot authorize a red merge; section 1 owns when such an instruction overrides a Firstmate-written standing rule within its exact scope. Load `ask-user-authority` before deciding any ask-user finding; the implementation worker never answers its own finding. Use `bin/fm-pr-merge.sh` for every task PR merge so merge metadata and any eligible explicitly recorded work-item lifecycle are handled and an unproved merge is refused instead of reported as landed, and use `bin/fm-merge-local.sh` for approved local-only landing; never call a lower-level merge command around their guards. After an autonomous merge, give the captain a one-line full-URL or local-main outcome. @@ -483,19 +486,20 @@ Queued wakes must be presented before other action and acknowledged only after h The spawn assertion and generated ship or design brief must both enforce that project work starts in an isolated disposable worktree, never the primary checkout. Harness-aware turn-end guards are structural backstops, not permission to omit the live cycle. -### Away-mode stub +### Away-mode and quiet-mode stub Invoke the `/afk` skill when the captain says `/afk`, says they are going afk, `state/.afk-contract` or `state/.afk` exists, an incoming message starts with `FM_INJECT_MARK`, or any `state/.subsuper-*` marker is involved. -The skill owns the daemon procedure; these safety facts remain inline: +Invoke the `/quiet` skill instead when the captain says `/quiet` or asks for quiet mode, or `state/.afk` already exists in quiet mode (`fm_classify_afk_mode` in `bin/fm-classify-lib.sh`). +The `/afk` skill owns the shared daemon procedure; `/quiet` owns its attended entry and explicit exit, and these safety facts remain inline for both: - Every current daemon injection uses the `away-supervisor` kind from `bin/fm-operational-input.sh` after `FM_OPERATIONAL_PREFIX` (U+2063 INVISIBLE SEPARATOR followed by `FIRSTMATE_OP: `), while the `/afk` skill owns legacy bare-marker compatibility. - `state/.afk-contract` is the away posture, written only after the captain confirms the read-back of their away words; entry announces hold-for-return only, and the record's clauses are recorded, not executed, in this release. - While `state/.afk` exists, the daemon owns supervision; do not arm a separate watcher. Pi and OMP extensions stand by while the daemon owns that cycle (`docs/watcher-continuity.md`). -- A marked message while away mode is active is internal escalation and does not exit away mode. -- A message beginning `/afk` refreshes away mode. -- Any other unmarked message means the captain returned; load `/afk`, run the return owner, and do not process that message as ordinary work until its durable catch-up gate clears. -- Away mode never expands approval authority for merges, ask-user findings, destructive actions, irreversible actions, or security-sensitive choices. +- A marked message while away or quiet mode is active is internal escalation and does not exit that mode. +- A message beginning `/afk` refreshes away mode; a message beginning `/quiet` refreshes quiet mode. +- Any other unmarked message means the captain returned in away mode (load `/afk`, run the return owner, and do not process that message as ordinary work until its durable catch-up gate clears), or, in quiet mode, is simply answered as ordinary work with the flag and daemon left untouched until an explicit `/quiet off`. +- Away and quiet mode never expand approval authority for merges, ask-user findings, destructive actions, irreversible actions, or security-sensitive choices. - Bias ambiguous input toward exit because a present captain takes precedence. ### Stuck-worker trigger @@ -553,14 +557,14 @@ Mention cost as a courtesy when unusually much work is running, but never block The configured `tasks-axi` backend is the durable queue; the tracked default is `data/backlog.md`. It tracks work items only, never agents; persistent secondmates never appear as backlog items. Work routed to a secondmate is recorded in that secondmate home's own backlog, not the main backlog. -A decision is simply a task held for the captain: create the task with `tasks-axi add` when needed, then always hold it through `bin/fm-captain-hold.sh hold --reason ""`, with `--until ` when the captain defers it. +A decision is simply a task held for the captain: create the task with `bin/fm-tasks-axi.sh add` when needed, then always hold it through `bin/fm-captain-hold.sh hold --reason ""`, with `--until ` when the captain defers it. When a main-side thread such as a pending captain decision or relay reminder is worth durable tracking, file it as its own work item and hold it through that wrapper. Captain calls discovered by investigations or visual reviews follow `captain-hold-lifecycle`, which owns their completion gate and recorded-answer rules. When the automatic transition gate applies, dispatch and completion move the item themselves - `bin/fm-spawn.sh` and `bin/fm-teardown.sh` own those transitions and refuse rather than report success without them - so what remains yours is filing the item before dispatch, recording decisions, and keeping notes current; `docs/configuration.md` owns gate applicability and the manual-backend exception. Re-evaluate queued work after every teardown and heartbeat, dispatching items only when dependencies and time gates have cleared. `.tasks.toml`, `docs/configuration.md`, and current `tasks-axi --help` own the backlog schema, compatibility, retention, and routine command syntax. -Use compatible `tasks-axi` when the configured backend selects it and the documented manual path otherwise; keep only the configured recent Done entries. +Use compatible `tasks-axi` when the configured backend selects it, always through `bin/fm-tasks-axi.sh` so the call reaches this home's backlog from any directory, and the documented manual path otherwise; keep only the configured recent Done entries. `secondmate-provisioning` and `bin/fm-backlog-handoff.sh` own cross-home handoff safety. Keep free-form notes free of temporary paths, moving versions, ephemeral identifiers, and copied state that will rot. @@ -571,7 +575,8 @@ Preserve durable structured identifiers, dependencies, and completion artifact l ## 11. Crewmate briefs `bin/fm-brief.sh` and its help own scaffold syntax, generated variants, status protocol, delivery-mode definitions of done, and exact safety mechanics. -Use its scaffold as the contract, then fill `## Captain's intent` (`{TASK}`) with the captain's own ask plus the context needed to read it, including the substance of any report, decision, or PR the ask refers to, and fill `## Firstmate spec` (`{FIRSTMATE_SPEC}`) with Firstmate's build instructions. +Use its scaffold as the contract, then fill `## Captain's intent` (`{TASK}`) with the captain's own ask and any boundary the captain stated, plus the context needed to read it, including the substance of any report, decision, or PR the ask refers to; never widen the ask there into a general goal or an enumerated coverage list, because the reviewer treats that subsection as acceptance criteria. +Fill `## Firstmate spec` (`{FIRSTMATE_SPEC}`) with only the build instructions that ask requires, naming what stays out of scope when the ask is narrow; a generalization, consistency sweep, or extra hardening the captain did not ask for is follow-up work to note, not scope to add. `bin/fm-dod-lib.sh` owns what a no-mistakes worker may pass as `--intent` and its rule that the string must be self-sufficient. Keep additions task-specific rather than repeating lifecycle instructions, and alter generated sections only when the task genuinely differs from the standard shape. Pass the task's own words with `--query` so the scaffold's one brain read searches for the right thing, and read the nearest-prior-work line it prints before dispatch: it claims proximity, never duplication, and never enters the brief. @@ -602,7 +607,7 @@ When the captain asks to check or update this fleet's toolchain ("check tool upd These skills are not captain-invocable; load them only at their precise triggers. -- `bootstrap-diagnostics` - load whenever the session-start digest's bootstrap or network-checks section prints an actionable diagnostic line (`MISSING:`, `MISSING_MANUAL:`, `BACKEND_INVALID:`, `NEEDS_GH_AUTH`, `TANGLE:`, `VAULT_DRIFT:`, `UPSTREAM:`, `GBRAIN_SERVING_CREDENTIAL:`, `GBRAIN_PIN:`, `GBRAIN_CAPTURE:`, `STARTUP_MEMORY_BUDGET:`, `CREW_DISPATCH:` (invalid or backend mismatch), `FLEET_SYNC:`, `BOARD_SWEEP:`, `NETWORK_CHECKS:`, `HOME_SUMMARY:`, `BACKLOG_RECONCILE:`, `ENDPOINT_BINDING_MIGRATION:`, `RUN_ATTRIBUTION:`, `SECONDMATE_SYNC:`, `SECONDMATE_LIVENESS:`, `SECONDMATE_HANDOFF:`, `NUDGE_SECONDMATES:`, `USAGE_STORE:`, or `FMX:`), or when `BOOTSTRAP_INFO:` says an interrupted backlog cleanup may have left an endpoint or local copy; silence and other `BOOTSTRAP_INFO:` facts need no load. +- `bootstrap-diagnostics` - load whenever the session-start digest's bootstrap or network-checks section prints an actionable diagnostic line (`MISSING:`, `MISSING_MANUAL:`, `PRESENTATION_UNAVAILABLE:`, `BACKEND_INVALID:`, `NEEDS_GH_AUTH`, `TANGLE:`, `VAULT_DRIFT:`, `UPSTREAM:`, `GBRAIN_SERVING_CREDENTIAL:`, `GBRAIN_PIN:`, `GBRAIN_CAPTURE:`, `STARTUP_MEMORY_BUDGET:`, `CREW_DISPATCH:` (invalid or backend mismatch), `FLEET_SYNC:`, `BOARD_SWEEP:`, `NETWORK_CHECKS:`, `HOME_SUMMARY:`, `BACKLOG_RECONCILE:`, `ENDPOINT_BINDING_MIGRATION:`, `RUN_ATTRIBUTION:`, `SECONDMATE_SYNC:`, `SECONDMATE_LIVENESS:`, `SECONDMATE_HANDOFF:`, `NUDGE_SECONDMATES:`, `USAGE_STORE:`, or `FMX:`), or when `BOOTSTRAP_INFO:` says an interrupted backlog cleanup may have left an endpoint or local copy; silence and other `BOOTSTRAP_INFO:` facts need no load. - `diagnostic-reasoning` - load before scoping a reported bug and before acting on a diagnostic report. - `design-profile` - load before scaffolding, dispatching, answering, completing, or cleaning up a one-conversation design task whose tracked deliverable is a short ADR. - `ask-user-authority` - load before deciding any ask-user finding, regardless of the project's `yolo` posture. @@ -615,7 +620,7 @@ These skills are not captain-invocable; load them only at their precise triggers - `secondmate-provisioning` - load before creating, seeding, validating, launching, handing backlog to, recovering, pushing inherited local material into, or retiring a secondmate home, and before editing `data/secondmates.md`. - `work-item-visibility` - load at intake before scaffolding a PR-based ship or design brief that carries a work item, and before posting any milestone the lifecycle scripts do not post themselves. - `captain-hold-lifecycle` - load before treating an investigation or visual review as complete, before ending a visual review that exposed a captain decision, when recording or routing the captain's answer, and on any `RECORD DIVERGENCE` line from the wake drain. -- `process-event-sources` - load before arming a long-polling source, before registering a deterministic condition->action watch (do X as soon as Y is true), and on any `procevent ` check wake. +- `process-event-sources` - load before arming a long-polling source, before registering a deterministic condition->action watch (do X as soon as Y is true), on any `procevent ` check wake, and on any `process-event source stranded` or `process-event source failed to start` check wake. Never run a registered source's blocking command yourself in a conversational turn. - `fmx-respond` - load on an `x-mention ` `check:` wake to handle the mention, on an `x-mode-error ...` `check:` wake to report the Relay configuration blocker, on a `public-followup ...` `check:` wake or a startup-surfaced public commitment, and on any milestone or terminal wake for a Relay-linked task before posting its completion follow-up; relevant only when Relay is on. - `firstmate-codexapp` - load before coordinating a visible Codex Desktop thread, evaluating a Codex App backend request, or reconciling Codex Desktop host-tool smoke evidence for Firstmate work. diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 4732df85ba7..697622a37a8 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -59,7 +59,7 @@ See the [no-mistakes quick start](https://kunchenguid.github.io/no-mistakes/star It pins one exact shellcheck version and one exact actionlint version and refuses to run under any other. Print the shellcheck pin with `bin/fm-lint.sh --required-version` and the actionlint pin with `bin/fm-lint-workflows.sh --required-version`. Use `bin/fm-install-shellcheck.sh` and `bin/fm-install-actionlint.sh` to install those exact builds locally; each installer's header owns its destination usage and supported platforms. -- Harness-adapter ownership spans detection in `bin/fm-harness.sh`, launch-command construction in `bin/fm-launch-lib.sh`, the surrounding gates and hook mechanics in `bin/fm-spawn.sh`, spawn-time Claude workspace-trust pre-registration in `bin/fm-claude-trust.sh`, semantic busy sources and trust gates in `bin/fm-busy-lib.sh`, delivery-only rendered guards in `bin/fm-composer-lib.sh`, cleanup in `bin/fm-teardown.sh`, and facts in the skill tree rooted at `.agents/skills/harness-adapters/SKILL.md`; the `firstmate-coding-guidelines` skill owns the validation policy for checks that depend on those harnesses. +- Harness-adapter ownership spans detection in `bin/fm-harness.sh`, launch-command construction in `bin/fm-launch-lib.sh`, the surrounding gates and hook mechanics in `bin/fm-spawn.sh`, spawn-time Claude workspace-trust and external-CLAUDE.md-import pre-approval in `bin/fm-claude-trust.sh`, semantic busy sources and trust gates in `bin/fm-busy-lib.sh`, delivery-only rendered guards in `bin/fm-composer-lib.sh`, cleanup in `bin/fm-teardown.sh`, and facts in the skill tree rooted at `.agents/skills/harness-adapters/SKILL.md`; the `firstmate-coding-guidelines` skill owns the validation policy for checks that depend on those harnesses. - Changes to runtime session backends (`bin/fm-backend.sh`, `bin/backends/`, and the scripts that dispatch through them) keep current setup and limits in the relevant backend guide and active empirical evidence in [`docs/verification/runtime-backends.md`](docs/verification/runtime-backends.md). - [`docs/documentation-audiences.md`](docs/documentation-audiences.md) and its machine-consumed inventory own prose classification. - For the contract governing dated verification entries, see [`docs/verification/README.md`](docs/verification/README.md). @@ -119,13 +119,14 @@ Every standard-mode sweep above also arms an automatic per-script bound, so a hu Portable shard balance evidence lives in `docs/fm-test-portable-shards.md`. Family selection is the ordinary local path; `--all` is deliberate full regression only. CI owns broad regression across required portable parallel shards, the portable serial lane's separate-runner shards, the Herdr lane, lint, invariants, the coverage guard, and stock macOS Bash compatibility in [`.github/workflows/ci.yml`](.github/workflows/ci.yml). +Pushing a new head to a pull request cancels that pull request's still-running CI so only the current head is validated; pushes to `main` are never cancelled, and the workflow owns that contract and its rationale. Use `bin/fm-test-run.sh --list-lanes` for exact lane names and `--help` for `--jobs` rules and required gate-skip flags when reproducing a lane locally. Leave the `sleep 0.1` cadence in the suites' bounded condition waits alone. Those sleeps look like recoverable overhead - `fm-watch-triage.test.sh` alone issues about 1,900 of them, each paying a flat ~100ms scheduler wake-up penalty on macOS - but they are not overhead added to the clock; they are how a test waits for a subject that only moves on `fm-watch.sh`'s own one-second `FM_POLL` cadence. Sampling less often does not remove that wait, it only delays detection: raising the interval to 0.5s and charging each sample proportionally measured `fm-watch-triage.test.sh` at 435s and 440s against 390s and 393s for the unchanged script, back to back on 2026-09-03, because each of its ~40 poll-cycle waits and ~73 process-exit waits paid up to half a second more. Some of those loops are also catching a transient rather than waiting for a settled condition, so a coarser sample can step over the state they assert on. Discover tests by listing `tests/*.test.sh`: each is a self-contained bash script named `.test.sh`, and its header comment describes what it covers, so pass one to `bin/fm-test-run.sh` to focus on a subject with canonical timing output. -Shared test helpers live in `tests/lib.sh` (reporters, temp roots, git fixtures), `tests/fixtures.sh` (fake toolchain and spawn-world builders), `tests/wake-helpers.sh`, and `tests/secondmate-helpers.sh`. +Shared test helpers live in `tests/lib.sh` (reporters, temp roots, git fixtures), `tests/fixtures.sh` (fake toolchain and spawn-world builders), `tests/wake-helpers.sh`, `tests/secondmate-helpers.sh`, and `tests/git-config-helpers.sh` (fixture Git isolation from the host's global and system configuration, already sourced by `tests/lib.sh` and `tests/herdr-test-safety.sh`; a suite that sources neither must source it itself before its first Git operation so a direct invocation stays isolated). Source those instead of copying a fake toolchain into a new suite. A fixture may shorten a production timeout to keep a failure path prompt, but never below what the real work inside that window costs on a loaded machine: a fork, an exec, a lock acquisition, a beacon publication, or a first-poll check. Where a case's assertion is not about the timeout itself, give that window headroom over the measured loaded cost, and bound the test's own waiting with iteration-counted poll loops, which stretch under load where a wall-clock budget does not. diff --git a/README.md b/README.md index 86333cfeff0..4645e53899e 100644 --- a/README.md +++ b/README.md @@ -186,6 +186,7 @@ Claude and grok use the slash form shown here; codex uses the same names with `$ | Skill | What it does | | ------------------ | -------------------------------------------------------------------------------------------------------------------------------------------- | | `/afk` | Enter away-mode supervision: the sub-supervisor self-handles routine notifications in bash, escalates captain-relevant events and bounded declared-external-wait rechecks as batched digests, and actively alerts if delivery gets stuck while you step away | +| `/quiet` | Enter quiet supervision mode: the same token-saving sub-supervisor tradeoff as `/afk`, for a captain who is staying and chatting - ordinary messages do not exit it, only an explicit `/quiet off` does | | `/ahoy` | Recap visible session events since the prior real captain message plus visibly unanswered captain decisions, then guide the captain through any open decisions one at a time in agent-judged impact order; fall back to Bearings when invoked as the session's first real captain message | | `/bearings` | Generate a concise four-section chat digest from bounded fleet state, including registered remote-home ledgers; use `/bearings file` to also replace today's dated report in `data/`, and add `include PRs` for live GitHub enrichment | | `/sync-upstream` | Check the fork's read-only upstream drift status or, when asked to sync, dispatch the next contiguous full-merge round as a reviewable fork PR without merging it | diff --git a/bin/backends/herdr.sh b/bin/backends/herdr.sh index 51550ba3355..e7c1858e3b3 100644 --- a/bin/backends/herdr.sh +++ b/bin/backends/herdr.sh @@ -86,6 +86,13 @@ FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" # shellcheck source=bin/fm-transition-lib.sh . "$FM_BACKEND_HERDR_ROOT/bin/fm-transition-lib.sh" +# Shared, backend-neutral harness-process identity (bin/fm-agent-process-lib.sh): +# the same agent|shell|other vocabulary the tmux adapter proves liveness with, +# so a Herdr registration is verified against the pane's real processes by the +# same rule (fm_backend_herdr_pane_process_state). +# shellcheck source=bin/fm-agent-process-lib.sh +. "$FM_BACKEND_HERDR_ROOT/bin/fm-agent-process-lib.sh" + FM_BACKEND_HERDR_MIN_PROTOCOL=14 FM_BACKEND_HERDR_MIN_AGENT_PROMPT_VERSION=0.7.5 # events.subscribe (the native pane.agent_status_changed push stream) and its @@ -1542,9 +1549,13 @@ fm_backend_herdr_pane_agent_free_proof() { # # fm_backend_herdr_pane_agent_free_sample: one tri-state instantaneous # observation for fm_backend_herdr_pane_agent_free_proof: 0 is agent-free, # 1 is retryable, and 2 is a conclusively live process-info foreground. -fm_backend_herdr_pane_agent_free_sample() { # +fm_backend_herdr_pane_agent_free_sample() { # [process-info-json] local session=$1 pane=$2 info shell_pid fg_pgid count fg_pid name argv0 shell_name ps_bin rows - info=$(fm_backend_herdr_cli "$session" pane process-info --pane "$pane" 2>/dev/null) || return 1 + if [ "$#" -ge 3 ]; then + info=$3 + else + info=$(fm_backend_herdr_cli "$session" pane process-info --pane "$pane" 2>/dev/null) || return 1 + fi printf '%s' "$info" | jq -e --arg pane "$pane" ' .result.type == "pane_process_info" and .result.process_info.pane_id == $pane @@ -2250,40 +2261,151 @@ fm_backend_herdr_explicit_close_pane_confirmed() { # [ "$presence" = dead ] } +# fm_backend_herdr_pane_process_state: what the operating system says is +# running in , as one of agent|shell|other|unreadable, from `pane +# process-info` plus the real process table. This is the process-level proof +# fm_backend_herdr_pane_agent_state demands before it lets a registration count +# as a live agent (issue #4115), built on the same shape the tmux adapter uses: +# the foreground process group is authoritative, read through the shared +# classifier in bin/fm-agent-process-lib.sh. +# +# agent - a foreground process has a verified harness identity through +# the shared classifier. +# shell - the exact process-info sample satisfies the terminal-wide +# absence proof owned by fm_backend_herdr_pane_agent_free_proof. +# other - readable foreground or process-tree evidence cannot prove +# absence. This includes background agents, unknown processes, +# and an empty foreground during exec handoff. Retry within the +# bounded settle window to let transient prompt helpers finish. +# unreadable - process-info failed or did not identify this pane and shell; +# an empty foreground naming a vanished shell also fails closed. +# +# Verified on Herdr 0.9.0 (docs/verification/runtime-backends.md "Stale agent +# registration"): process-info's `.name` is the kernel process name (`node` for +# Pi, `zsh` for a shell), `.argv0` the argv[0] basename (`pi`), and `.argv` / +# `.cmdline` the full command line, so Pi is identified by argv[0] exactly as +# the tmux probe identifies it from `ps`. +fm_backend_herdr_pane_process_state() { # + local attempt=0 max_attempts=${FM_BACKEND_HERDR_IDLE_SHELL_PROOF_POLLS:-10} verdict + while :; do + verdict=$(fm_backend_herdr_pane_process_state_sample "$1" "$2") + [ "$verdict" = other ] || break + attempt=$((attempt + 1)) + [ "$attempt" -lt "$max_attempts" ] || break + sleep 0.1 + done + printf '%s' "$verdict" +} + +# fm_backend_herdr_pane_process_state_sample: one instantaneous observation +# for fm_backend_herdr_pane_process_state, which owns the verdict contract and +# the settle retry. +fm_backend_herdr_pane_process_state_sample() { # + local session=$1 pane_id=$2 info shell_pid count i pid name argv0 args verdict + local others=0 + info=$(fm_backend_herdr_cli "$session" pane process-info --pane "$pane_id" 2>/dev/null) \ + || { printf 'unreadable'; return 0; } + printf '%s' "$info" | jq -e --arg pane "$pane_id" ' + .result.type == "pane_process_info" + and .result.process_info.pane_id == $pane + ' >/dev/null 2>&1 || { printf 'unreadable'; return 0; } + shell_pid=$(printf '%s' "$info" | jq -er \ + '.result.process_info.shell_pid | select(type == "number" and . > 1) | floor' 2>/dev/null) \ + || { printf 'unreadable'; return 0; } + count=$(printf '%s' "$info" | jq -er \ + '.result.process_info.foreground_processes | select(type == "array") | length' 2>/dev/null) \ + || { printf 'unreadable'; return 0; } + if [ "$count" -eq 0 ]; then + # An exec handoff may temporarily omit the foreground. Retry a real shell, + # but never infer absence from a provider naming a vanished process. + if ! "${FM_HERDR_PS_BIN:-ps}" -axo pid=,ppid=,stat=,comm=,tty= 2>/dev/null \ + | awk -v want="$shell_pid" '$1 == want { found=1 } END { exit !found }'; then + printf 'unreadable' + else + printf 'other' + fi + return 0 + fi + i=0 + while [ "$i" -lt "$count" ]; do + pid=$(printf '%s' "$info" | jq -r --argjson i "$i" \ + '.result.process_info.foreground_processes[$i].pid | select(type == "number") | floor' 2>/dev/null) + name=$(printf '%s' "$info" | jq -r --argjson i "$i" \ + '.result.process_info.foreground_processes[$i].name // empty' 2>/dev/null) + argv0=$(printf '%s' "$info" | jq -r --argjson i "$i" ' + .result.process_info.foreground_processes[$i] as $p + | (($p.argv // [])[0]) // $p.argv0 // empty' 2>/dev/null) + args=$(printf '%s' "$info" | jq -r --argjson i "$i" ' + .result.process_info.foreground_processes[$i] as $p + | $p.cmdline // (($p.argv // []) | join(" ")) // empty' 2>/dev/null) + verdict=$(fm_agent_process_classify "$name" "$argv0" "$args" "$pid") + case "$verdict" in + agent) printf 'agent'; return 0 ;; + shell) ;; + *) others=$((others + 1)) ;; + esac + i=$((i + 1)) + done + + # Positive harness identity is shared with upstream. Agent absence must also + # satisfy the fork's terminal-wide proof, including orphaned same-TTY jobs, + # suspended agents, and the nested Treehouse shell chain. Reuse this exact + # process-info sample so the proof and identity view cannot drift apart. + if [ "$others" -eq 0 ] && fm_backend_herdr_pane_agent_free_sample "$session" "$pane_id" "$info"; then + printf 'shell' + else + printf 'other' + fi +} + # fm_backend_herdr_pane_agent_state: classify in as one of -# dead|no-agent|live|unknown, purely from the JSON body of two read-only -# calls - never from process exit status, since a business-logic "not found" -# response is a normal, expected outcome here, not a call failure (real herdr -# 0.7.1 exits 1 for it; the canned-response test fakes exit 0; parsing only -# the JSON keeps this function correct against either). +# dead|no-agent|stale-agent|live|unknown, from the JSON body of two read-only +# calls plus, for a registered agent, the pane's process-level view - never +# from process exit status, since a business-logic "not found" response is a +# normal, expected outcome here, not a call failure (real herdr 0.7.1 exits 1 +# for it; the canned-response test fakes exit 0; parsing only the JSON keeps +# this function correct against either). # -# dead - `pane get` responds with error code pane_not_found: the pane -# itself is gone (closed, or its process died and herdr already -# reaped it - verified empirically: killing a pane's shell pid -# on a live server makes herdr immediately drop both the pane -# and its tab from `pane get`/`tab list`). -# no-agent - `pane get` succeeds (the pane structurally exists) but `agent -# get` responds with error code agent_not_found: nothing is -# registered in it - exactly what a herdr session-layout restore -# produces (verified empirically: `session stop` + fresh `herdr -# server` restart leaves the pane alive, agent_status "unknown", -# agent get -> agent_not_found - docs/herdr-backend.md "ID -# stability across a server restart"), and what a future -# `resume_agents_on_restore = false` restore would produce too -# (a plain shell, never an agent). -# live - `agent get` succeeds and reports a real agent_status (working, -# idle, done, or blocked - any registered value). An idle or -# blocked agent is still a genuine, still-registered agent, not -# a restored husk, so it is never a close-and-replace candidate. -# A registration can outlive its agent, though: the recovery-grade -# fm_backend_herdr_agent_state below owns the stale-registration -# cross-check layered on top of this raw verdict. -# unknown - anything else: an unparseable/unexpected response from either -# call, or a `pane get` success whose own echoed pane_id does not -# round-trip (guards against misreading a herdr response shape -# change as "the pane exists"). The caller must fail safe toward -# refusal here, never toward closing - this is the conservative -# backstop the husk check depends on. +# dead - `pane get` responds with error code pane_not_found: the pane +# itself is gone (closed, or its process died and herdr already +# reaped it - verified empirically: killing a pane's shell pid +# on a live server makes herdr immediately drop both the pane +# and its tab from `pane get`/`tab list`). +# no-agent - `pane get` succeeds (the pane structurally exists) but `agent +# get` responds with error code agent_not_found: nothing is +# registered in it - exactly what a herdr session-layout restore +# produces (verified empirically: `session stop` + fresh `herdr +# server` restart leaves the pane alive, agent_status "unknown", +# agent get -> agent_not_found - docs/herdr-backend.md "ID +# stability across a server restart"), and what a future +# `resume_agents_on_restore = false` restore would produce too +# (a plain shell, never an agent). +# stale-agent - `agent get` reports a registered agent_status (working, idle, +# done, or blocked) but fm_backend_herdr_pane_process_state +# proves the pane is shell-only: the registered agent's process +# has exited and Herdr kept its registration (issue #4115; +# Herdr does not release a Pi registration on TUI shutdown when +# a nested shell sits under the pane's top shell, the crew +# shape). This is the explicit agent-free reason: the pane is +# recoverable, and the record it carries is not evidence of a +# running agent. No registered status outranks the process +# view, because a killed mid-turn agent leaves `working` +# behind just as a quit one leaves `idle`. +# live - `agent get` succeeds with a registered agent_status and the +# process-level view is `agent` or `other`: a harness process +# is running, or something that is not a bare shell is, so the +# registration keeps its authority. An idle or blocked agent +# is still a genuine, still-registered agent, not a restored +# husk, so it is never a close-and-replace candidate. +# unknown - anything else: an unparseable/unexpected response from +# either call, a `pane get` success whose own echoed pane_id +# does not round-trip (guards against misreading a herdr +# response shape change as "the pane exists"), or a registered +# agent whose process-level view is unreadable - the +# registration alone is no longer trusted, and its absence is +# not claimed either. The caller must fail safe toward refusal +# here, never toward closing - this is the conservative +# backstop the husk check depends on. fm_backend_herdr_pane_agent_state() { # local session=$1 pane_id=$2 out code presence status presence=$(fm_backend_herdr_pane_presence_state "$session" "$pane_id") @@ -2302,16 +2424,23 @@ fm_backend_herdr_pane_agent_state() { # fi status=$(printf '%s' "$out" | jq -r '.result.agent.agent_status // empty' 2>/dev/null) case "$status" in - working|idle|done|blocked) printf 'live' ;; + working|idle|done|blocked) ;; + *) printf 'unknown'; return 0 ;; + esac + case "$(fm_backend_herdr_pane_process_state "$session" "$pane_id")" in + agent|other) printf 'live' ;; + shell) printf 'stale-agent' ;; *) printf 'unknown' ;; esac } # fm_backend_herdr_tab_is_husk: true (0) only for the two conservative husk # states (dead, no-agent) fm_backend_herdr_pane_agent_state can positively -# confirm; live and unknown both refuse (1), so an inconclusive read never -# licenses closing anything. Restored-layout recovery depends on this -# fail-safe-toward-refusal behavior. +# confirm; live, stale-agent, and unknown all refuse (1), so an inconclusive +# read never licenses closing anything, and a stale registration - agent-free +# for RECOVERY, which reuses the pane - still never licenses closing it, because +# the shell it holds may be a nested worktree shell. Restored-layout recovery +# depends on this fail-safe-toward-refusal behavior. fm_backend_herdr_tab_is_husk() { # case "$(fm_backend_herdr_pane_agent_state "$1" "$2")" in dead|no-agent) return 0 ;; @@ -2347,9 +2476,10 @@ fm_backend_herdr_server_running_state() { # # fm_backend_herdr_agent_state: recovery-grade state for the same session-start # sweep as the tmux classifier. It reuses the husk classifier rather than # creating a second Herdr state machine: a structurally gone pane is `missing`, -# a confirmed agent-less pane is `dead`, a registered agent is `alive` unless -# its registration is proven stale (below), and an unexpected or failed API -# read is `unreadable`. +# a confirmed agent-less pane is `dead` - whether nothing is registered or a +# registration lingers over a shell-only pane (stale-agent, issue #4115) - a +# registered agent with a live process is `alive`, and an unexpected or failed +# API read is `unreadable`. # # A registered agent is NOT taken at its word here, because herdr can hold a # permanently stale registration. For an agent whose lifecycle reporting is @@ -2363,23 +2493,12 @@ fm_backend_herdr_server_running_state() { # # recovery exists for. docs/verification/runtime-backends.md owns the active # versioned evidence. # -# So a live registration is cross-checked against the OS-level agent-free -# proof (fm_backend_herdr_pane_agent_free_proof): when the pane's entire -# process tree is provably nothing but recognized idle shells plus the -# treehouse worktree wrapper - the shape a task pane is left in when its agent -# dies - the registered agent's process does not exist, the registration is -# stale, and the endpoint classifies `dead` (agent-free, recovery licensed). -# This is deliberately a process-table fact, never a -# screen read: a rendered-prompt heuristic cannot be trusted here (agy's live -# prompt glyph is a bare `>`, and a themed shell prompt can itself end in -# `❯`), while any genuinely live agent - working, idle, suspended, or -# backgrounded as a shell job - puts a non-shell process in the pane's tree, -# fails the proof, and stays `alive`. Every inconclusive read (process info -# unavailable, unknown shell, unreadable process table) also fails the proof -# and stays `alive`, preserving the husk classifier's -# fail-safe-toward-refusal contract: only a positively proven agent-free pane -# ever unlocks recovery. - +# The shared process classifier uses fm_backend_herdr_pane_agent_free_sample +# for absence, preserving the terminal-wide proof owned by +# fm_backend_herdr_pane_agent_free_proof. Only positive absence unlocks recovery; +# a failed process-info read is unreadable, and an inconclusive readable sample +# remains alive. Neither state licenses removal or recovery. +# # One exception to that last case, and it is deliberately made HERE rather than # in the husk classifier: a read can fail because the recorded session's server # is not running at all, which is authoritative absence for every pane in that @@ -2398,15 +2517,8 @@ fm_backend_herdr_agent_state() { # fm_backend_herdr_parse_target "$target" || { printf 'unreadable'; return 0; } case "$(fm_backend_herdr_pane_agent_state "$FM_BACKEND_HERDR_SESSION" "$FM_BACKEND_HERDR_PANE")" in dead) printf 'missing' ;; - no-agent) printf 'dead' ;; - live) - if fm_backend_herdr_pane_agent_free_proof \ - "$FM_BACKEND_HERDR_SESSION" "$FM_BACKEND_HERDR_PANE" >/dev/null 2>&1; then - printf 'dead' - else - printf 'alive' - fi - ;; + no-agent|stale-agent) printf 'dead' ;; + live) printf 'alive' ;; *) case "$(fm_backend_herdr_server_running_state "$FM_BACKEND_HERDR_SESSION")" in stopped) printf 'missing' ;; @@ -2743,7 +2855,7 @@ fm_backend_herdr_projection_reclaim_rollback() { # case "$state" in dead) return 0 ;; no-agent) ;; - live|unknown) return 1 ;; + live|stale-agent|unknown) return 1 ;; esac fm_backend_herdr_projection_close_pane_focus_preserving "$session" "$new_pane" no-agent || return 1 [ "$(fm_backend_herdr_pane_agent_state "$session" "$new_pane")" = dead ] @@ -2795,7 +2907,7 @@ fm_backend_herdr_projection_reclaim_task() { # &2 return 2 ;; - live|unknown) + live|stale-agent|unknown) echo "error: exact herdr presentation pane for $id is $state; refusing duplicate launch" >&2 return 1 ;; @@ -2844,7 +2956,7 @@ fm_backend_herdr_projection_reclaim_task() { # &2 return 1 @@ -2867,7 +2979,7 @@ fm_backend_herdr_projection_reclaim_task() { # &2 return 1 ;; @@ -2949,7 +3061,7 @@ fm_backend_herdr_projection_recovery_allows_flat() { # &2 return 1 ;; @@ -3712,10 +3824,22 @@ fm_backend_herdr_agent_status_raw() { # # gets real semantics" per the design report. See # fm_backend_herdr_classify_agent_status for the status->busy/idle/unknown # mapping. +# +# A `busy` verdict is proven at process level before it is reported: a +# lingering `working` registration over a shell-only pane (an agent killed +# mid-turn, issue #4115) reads `unknown`, never busy, so the recovery classifier +# cannot report a shell-only pane as working. Only the busy case pays the extra +# process read; idle and unknown are never trusted as busy by any consumer. fm_backend_herdr_busy_state() { # + local verdict fm_backend_herdr_target_ready "$1" || { printf 'unknown'; return 0; } - fm_backend_herdr_classify_agent_status \ - "$(fm_backend_herdr_agent_status_raw "$FM_BACKEND_HERDR_SESSION" "$FM_BACKEND_HERDR_PANE")" + verdict=$(fm_backend_herdr_classify_agent_status \ + "$(fm_backend_herdr_agent_status_raw "$FM_BACKEND_HERDR_SESSION" "$FM_BACKEND_HERDR_PANE")") + if [ "$verdict" = busy ] \ + && [ "$(fm_backend_herdr_pane_process_state "$FM_BACKEND_HERDR_SESSION" "$FM_BACKEND_HERDR_PANE")" = shell ]; then + verdict=unknown + fi + printf '%s' "$verdict" } # fm_backend_herdr_wait_for_working: poll :'s NATIVE diff --git a/bin/backends/tmux.sh b/bin/backends/tmux.sh index 42a87fcc49d..4477eb97423 100644 --- a/bin/backends/tmux.sh +++ b/bin/backends/tmux.sh @@ -22,10 +22,8 @@ . "$FM_BACKEND_LIB_DIR/fm-tmux-lib.sh" # shellcheck source=bin/fm-session-lock-lib.sh . "$FM_BACKEND_LIB_DIR/fm-session-lock-lib.sh" -# shellcheck source=bin/fm-cursor-lib.sh -. "$FM_BACKEND_LIB_DIR/fm-cursor-lib.sh" -# shellcheck source=bin/fm-gemini-lib.sh -. "$FM_BACKEND_LIB_DIR/fm-gemini-lib.sh" +# shellcheck source=bin/fm-agent-process-lib.sh +. "$FM_BACKEND_LIB_DIR/fm-agent-process-lib.sh" # fm_backend_tmux_resolve_bare_selector: the live-window-listing fallback for a # selector that is neither an explicit target nor a task selector routed @@ -154,50 +152,10 @@ fm_backend_tmux_current_command() { # tmux display-message -p -t "$1" '#{pane_current_command}' 2>/dev/null } -# fm_backend_tmux_classify_process_name: the single owner of the process-name -# vocabulary shared by every liveness signal below - `agent` for a verified -# harness, `shell` for an idle login/interactive shell, `other` for anything -# else. Keeping one classifier means the two independent name sources can never -# drift into disagreeing about what a given name means. -fm_backend_tmux_classify_process_name() { # [argv0] -> agent|shell|other - local path=$1 argv0=${2:-} base - base=${path##*/} - base=${base#-} - case "$base" in - # muse is anchored rather than globbed like its neighbours: its installed - # binary is muse-bin- (the launcher execs it, so the version is the - # live process name and changes on every auto-update), and unlike `claude` or - # `codex` the substring `muse` is a common English fragment - a *muse* glob - # would classify musescore or amuse as a live agent pane. The install path - # cannot carry it either: ~/.local/bin/muse-bin- has no `muse` path - # COMPONENT, so the fm_harness_path_name fallback below never fires for it. - muse|muse-bin-*) printf 'agent' ;; - # omp (Oh My Pi) is anchored for the same reason as muse: its live process - # name is the bare word `omp` (verified, omp 18.1.11) and a glob would claim - # unrelated commands such as ompd or comp. - *claude*|*codex*|*opencode*|*grok*|*kimi*|*rovo*|pi|pi-signed|pi-launcher|Pi|omp) printf 'agent' ;; - zsh|bash|sh|dash|ash|ksh|mksh|tcsh|csh|fish) printf 'shell' ;; - *) - if fm_harness_path_name "$path" >/dev/null || fm_harness_path_name "$argv0" >/dev/null; then - printf 'agent' - # cursor-agent runs as a bundled node script, so tmux reports the pane - # command as a bare `node` that no name pattern above can own, and its - # other installed name is the far-too-generic `agent` (verified live on - # cursor-agent 2026.08.11-e8db854: #{pane_current_command} is `node` while - # `ps -o comm=` carries the cursor-agent install path). Identity therefore - # comes from the narrowed structural rule in bin/fm-cursor-lib.sh, which - # demands Cursor's own name or install tree in the path or argv[0]. An - # unrelated `node` or `agent` matches nothing here and stays `other`, - # which the callers above fold into `ambiguous` rather than `dead`, so a - # stranger's node pane is never reported as an agent-free pane. - elif fm_cursor_process_matches "${path:-$argv0}" '' "$argv0"; then - printf 'agent' - else - printf 'other' - fi - ;; - esac -} +# The process-name classifier every liveness signal below feeds +# (fm_agent_process_classify_name) is owned by bin/fm-agent-process-lib.sh, +# shared with the Herdr adapter so both backends mean the same thing by +# `agent`, `shell`, and `other`. # fm_backend_tmux_foreground_comms: the kernel-side names of every process in # 's pane tty foreground process group, one full value per line. @@ -330,7 +288,7 @@ fm_backend_tmux_agent_state() { # while IFS= read -r name; do [ -n "$name" ] || continue fg_seen=1 - case "$(fm_backend_tmux_classify_process_name "$name")" in + case "$(fm_agent_process_classify_name "$name")" in agent) printf 'alive'; return 0 ;; shell) fg_shell=1 ;; *) fg_other=1 ;; @@ -342,7 +300,7 @@ EOF argv0s=$(fm_backend_tmux_foreground_argv0s "$target") while IFS= read -r name; do [ -n "$name" ] || continue - if [ "$(fm_backend_tmux_classify_process_name '' "$name")" = agent ]; then + if [ "$(fm_agent_process_classify_name '' "$name")" = agent ]; then printf 'alive' return 0 fi @@ -379,7 +337,7 @@ EOF printf 'unreadable' return 0 } - if [ "$(fm_backend_tmux_classify_process_name "$comm")" = agent ]; then + if [ "$(fm_agent_process_classify_name "$comm")" = agent ]; then printf 'alive' return 0 fi @@ -398,7 +356,7 @@ EOF case "$comm" in '') printf 'unreadable'; return 0 ;; esac - case "$(fm_backend_tmux_classify_process_name "$comm")" in + case "$(fm_agent_process_classify_name "$comm")" in shell) printf 'dead' ;; *) printf 'ambiguous' ;; esac diff --git a/bin/fm-afk-contract.sh b/bin/fm-afk-contract.sh index 593b65e3a01..04f8197f9a6 100755 --- a/bin/fm-afk-contract.sh +++ b/bin/fm-afk-contract.sh @@ -21,6 +21,9 @@ # reach_channels: none # reach_announced: # spend_max_concurrent_workers: +# merge_grants: - | task ids that may merge while this record exists +# - (empty is `merge_grants: -`; a missing field on +# ... a pre-field v1 record reads as an empty list) # confirmed: # confirmed_epoch: # words: | or |- the captain's words, verbatim, never edited, @@ -82,12 +85,14 @@ # Usage: # fm-afk-contract.sh propose [--words-file | --words ] # [--action --object --when [--stop ]]... -# [--expected-return ] [--spend ] +# [--expected-return ] [--spend ] [--grant ]... # Compile and write the proposal, then print the read-back. Exit 0 with every # clause accepted, 3 when at least one clause was refused (the read-back names # the missing part), and 2 on a usage error. --words-file keeps the file's # bytes verbatim, trailing newlines included. A refused clause remains in the -# proposal so the captain can restate it before saying go. +# proposal so the captain can restate it before saying go. Repeatable --grant +# records captain-named task ids that may merge-when-green while the record +# exists; invalid or duplicate ids are a usage error, never a refused clause. # fm-afk-contract.sh confirm # Promote the proposal into the record with the confirmed timestamp and # print the entry announcement. A proposal is required when no confirmed @@ -103,12 +108,31 @@ # (`\\`, `\t`, `\r`, and `\n`) so every record remains one row per clause; # a literal `-` is `\x2d` to distinguish it from the empty-stop marker. # fm-afk-contract.sh refused [--proposal | --path ] TSV: id text missing +# fm-afk-contract.sh grants [--proposal | --path ] one task id per line # fm-afk-contract.sh archive move the record aside; print its path # fm-afk-contract.sh archived print that archived record's path # -# Sourceable: with the BASH_SOURCE guard, other scripts get the path and -# presence helpers (fm_afk_contract_path, fm_afk_contract_present, -# fm_afk_contract_proposal_path, fm_afk_contract_archive_dir) without running main. +# CROSS-SUBSYSTEM LOCK (state/.afk-contract.lock; this script is its one owner). +# This record is authority another subsystem reads and then ACTS on outside this +# script: bin/fm-pr-merge.sh reads the merge grants and afterwards hands a merge +# to the forge. A publication, replacement, or archive landing between that read +# and the forge handoff would land a merge on authority that no longer holds, so +# the two subsystems share one lock instead of each locking its own records: the +# record-mutating subcommands (confirm, archive) hold it across their mutation, +# and a reader that acts on the record holds it across both its read and that +# action (fm_afk_contract_lock_hold / fm_afk_contract_lock_release). The +# read-only subcommands never take it, so a holder can still read the record it +# locked. Neither side ever proceeds without it: the acquire is bounded, and a +# bound that is hit refuses and names the live holder rather than racing. That +# fixed bound is 120 seconds, sized so only a genuinely wedged holder trips it. +# A lock left by a killed process is reclaimed +# by the ordinary stale-owner recovery in bin/fm-wake-lib.sh, which owns the lock +# primitive itself. +# +# Sourceable: with the BASH_SOURCE guard, other scripts get the path, presence, +# and lock helpers (fm_afk_contract_path, fm_afk_contract_present, +# fm_afk_contract_proposal_path, fm_afk_contract_archive_dir, +# fm_afk_contract_lock_hold, fm_afk_contract_lock_release) without running main. set -u FM_AFK_CONTRACT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -123,6 +147,10 @@ FM_AFK_CONTRACT_VERSION=1 FM_AFK_CONTRACT_VERBS="merge land prerelease install rerun dispatch abort-run answer discard wake-me" FM_AFK_CONTRACT_REACH_ANNOUNCED='No phone channel is configured; anything that needs you waits for your return.' FM_AFK_CONTRACT_SPEND_DEFAULT=4 +# Generous against the longest legitimate holder, a merge waiting on the forge, +# so the bound only ever trips on something genuinely wedged. +_FM_AFK_CONTRACT_LOCK_TIMEOUT=120 +FM_AFK_CONTRACT_LOCK_HELD= fm_afk_contract_path() { # [state-dir] printf '%s/.afk-contract' "${1:-$FM_AFK_CONTRACT_STATE}" @@ -140,6 +168,54 @@ fm_afk_contract_present() { # [state-dir] [ -f "$(fm_afk_contract_path "${1:-$FM_AFK_CONTRACT_STATE}")" ] } +fm_afk_contract_lock_path() { # [state-dir] + printf '%s/.afk-contract.lock' "${1:-$FM_AFK_CONTRACT_STATE}" +} + +# Lazily reach the lock primitive. bin/fm-wake-lib.sh is a canonical lint root +# in its own right, so keep this an analysis boundary for the same reason +# bin/fm-lease-lib.sh's fm_lease_lock_helpers does. +fm_afk_contract_lock_helpers() { + command -v fm_lock_acquire_wait_bounded >/dev/null 2>&1 && return 0 + # shellcheck source=/dev/null + . "$FM_AFK_CONTRACT_DIR/fm-wake-lib.sh" +} + +# fm_afk_contract_lock_hold [state-dir]: take the cross-subsystem lock described +# in the header. The acquire is bounded so a wedged holder is refused instead of +# blocking a merge or a captain return forever, and returns 1 WITHOUT the lock so +# every caller refuses rather than proceeding unlocked. +fm_afk_contract_lock_hold() { # [state-dir] + local lock rc=0 STATE timeout + STATE=${1:-$FM_AFK_CONTRACT_STATE} + lock=$(fm_afk_contract_lock_path "$STATE") + timeout=${FM_TEST_AFK_CONTRACT_LOCK_TIMEOUT:-$_FM_AFK_CONTRACT_LOCK_TIMEOUT} + fm_afk_contract_lock_helpers || { + fm_afk_contract_log "could not load the lock primitive for $lock" + return 1 + } + fm_lock_acquire_wait_bounded "$lock" "$timeout" || rc=$? + if [ "$rc" -ne 0 ]; then + if [ "$rc" -eq 124 ] && [ -n "${FM_LOCK_HELD_PID:-}" ]; then + fm_afk_contract_log "the away-posture record is locked by live process $FM_LOCK_HELD_PID (an in-flight merge, or another change to this record); nothing was changed" + else + fm_afk_contract_log "could not take the away-posture record lock at $lock; nothing was changed" + fi + return 1 + fi + FM_AFK_CONTRACT_LOCK_HELD=$lock +} + +# Release the lock taken by fm_afk_contract_lock_hold. Idempotent, so callers can +# invoke it unconditionally from their own cleanup. +fm_afk_contract_lock_release() { + local lock=$FM_AFK_CONTRACT_LOCK_HELD + [ -n "$lock" ] || return 0 + FM_AFK_CONTRACT_LOCK_HELD= + fm_afk_contract_lock_helpers || return 1 + fm_lock_release "$lock" +} + fm_afk_contract_log() { printf 'fm-afk-contract: %s\n' "$*" >&2; } fm_afk_contract_usage() { @@ -162,6 +238,15 @@ fm_afk_contract_blank() { # [ -z "$(printf '%s' "$1" | tr -d '[:space:]')" ] } +# Same alphabet as fm_pr_task_id_valid / fm_task_id_path_safe in bin/fm-pr-lib.sh. +# Kept local so sourcing this file cannot reset that library's parse globals. +fm_afk_contract_grant_id_valid() { # + local LC_ALL=C id=${1-} + case "$id" in + ''|.*|*[!A-Za-z0-9._-]*) return 1 ;; + esac +} + fm_afk_contract_escape() { # local value=$1 value=${value//\\/\\\\} @@ -266,9 +351,10 @@ fm_afk_contract_validate_iso() { # # Compile every input into a record body on stdout (everything except the # confirmed fields). Inputs: WORDS (verbatim), the parallel clause field arrays -# CLAUSE_ACTIONS CLAUSE_OBJECTS CLAUSE_WHENS CLAUSE_STOPS, EXPECTED_RETURN, SPEND. +# CLAUSE_ACTIONS CLAUSE_OBJECTS CLAUSE_WHENS CLAUSE_STOPS, EXPECTED_RETURN, +# SPEND, MERGE_GRANTS. fm_afk_contract_render_body() { # - local entered=$1 entered_epoch=$2 ordinal=0 i as_given + local entered=$1 entered_epoch=$2 ordinal=0 i as_given grant local accepted_block="" refused_block="" i=0 while [ "$i" -lt "${#CLAUSE_ACTIONS[@]}" ]; do @@ -303,6 +389,14 @@ fm_afk_contract_render_body() { # printf 'reach_channels: none\n' printf 'reach_announced: %s\n' "$FM_AFK_CONTRACT_REACH_ANNOUNCED" printf 'spend_max_concurrent_workers: %s\n' "${SPEND:-$FM_AFK_CONTRACT_SPEND_DEFAULT}" + if [ "${#MERGE_GRANTS[@]}" -eq 0 ]; then + printf 'merge_grants: -\n' + else + printf 'merge_grants:\n' + for grant in "${MERGE_GRANTS[@]}"; do + printf ' - %s\n' "$grant" + done + fi if [ -n "$WORDS" ]; then local words_body=$WORDS words_indicator='|-' case "$words_body" in @@ -370,6 +464,52 @@ fm_afk_contract_read_words() { # ' "$path" } +# One granted task id per line. A missing merge_grants field is an empty list +# so a pre-field v1 record fails closed for non-yolo merges instead of skipping +# the grant check. A present but unreadable field fails rather than guessing. +fm_afk_contract_read_grants() { # + local path=$1 + [ -f "$path" ] || return 1 + awk -v record="$path" ' + function die(reason) { + printf "fm-afk-contract: record %s has an invalid merge_grants field: %s\n", record, reason > "/dev/stderr" + bad = 1 + exit 2 + } + function valid_id(value) { + if (value == "" || substr(value, 1, 1) == ".") return 0 + return value ~ /^[A-Za-z0-9._-]+$/ + } + /^merge_grants:/ { + if (found) die("the field is defined more than once") + found = 1 + if ($0 == "merge_grants: -") { empty = 1; next } + if ($0 == "merge_grants:") { inlist = 1; next } + die("the empty form is merge_grants: -") + } + inlist && /^ - / { + id = substr($0, 5) + if (!valid_id(id)) die("task id \"" id "\" is not a valid task id") + if (seen[id]++) die("task id \"" id "\" is listed more than once") + print id + count++ + next + } + inlist && /^[^ ]/ { + if (count == 0) die("the list form has no stored ids") + inlist = 0 + next + } + empty && /^[^ ]/ { empty = 0; next } + inlist || empty { die("a stored grant line is malformed") } + END { + if (bad) exit 2 + if (!found) exit 0 + if (inlist && count == 0) die("the list form has no stored ids") + } + ' "$path" +} + # TSV rows for a list section:
is clauses or refused. fm_afk_contract_read_list() { #
local path=$1 section=$2 @@ -466,6 +606,10 @@ fm_afk_contract_validate() { # words_header=$(sed -n '/^words: /{p;q;}' "$path") case "$words_header" in 'words: -'|'words: |'|'words: |-') ;; *) fm_afk_contract_log "record $path has no valid words field"; return 1 ;; esac fm_afk_contract_read_words "$path" >/dev/null || return 1 + fm_afk_contract_read_grants "$path" >/dev/null || { + fm_afk_contract_log "record $path has no valid merge_grants field" + return 1 + } if [ "$require_confirmed" -eq 1 ]; then confirmed=$(fm_afk_contract_read_field "$path" confirmed) fm_afk_contract_validate_iso "$confirmed" || { fm_afk_contract_log "record $path has no valid confirmed time"; return 1; } @@ -526,13 +670,22 @@ EOF # --- rendering -------------------------------------------------------------- fm_afk_contract_render_readback() { # - local path=$1 title=$2 words count id action object when stop text missing expected spend flag + local path=$1 title=$2 words count id action object when stop text missing expected spend flag grants grant_list expected=$(fm_afk_contract_read_field "$path" expected_return) spend=$(fm_afk_contract_read_field "$path" spend_max_concurrent_workers) + grants=$(fm_afk_contract_read_grants "$path") || return 1 + grant_list= + while IFS= read -r id; do + [ -n "$id" ] || continue + grant_list="${grant_list:+$grant_list, }$id" + done <<EOF +$grants +EOF printf '%s\n' "$title" printf ' entered: %s\n' "$(fm_afk_contract_read_field "$path" entered)" printf ' expected return: %s\n' "$( [ "$expected" = - ] && printf 'not given' || printf '%s' "$expected")" printf ' spend cap: %s concurrent workers\n' "$spend" + printf ' merge when green (task ids): %s\n' "${grant_list:-(none)}" printf ' reach: hold-for-return only. %s\n' "$(fm_afk_contract_read_field "$path" reach_announced)" words=$(fm_afk_contract_read_words "$path"; printf x) words=${words%x} @@ -601,10 +754,11 @@ fm_afk_contract_render_announcement() { # <path> # --- subcommands ------------------------------------------------------------ -fm_afk_contract_parse_inputs() { # <args...>; sets WORDS, the CLAUSE_* arrays, EXPECTED_RETURN, SPEND - local words_file='' open=-1 +fm_afk_contract_parse_inputs() { # <args...>; sets WORDS, the CLAUSE_* arrays, EXPECTED_RETURN, SPEND, MERGE_GRANTS + local words_file='' open=-1 grant WORDS=; EXPECTED_RETURN=-; SPEND=$FM_AFK_CONTRACT_SPEND_DEFAULT CLAUSE_ACTIONS=(); CLAUSE_OBJECTS=(); CLAUSE_WHENS=(); CLAUSE_STOPS=(); CLAUSE_STOP_GIVENS=() + MERGE_GRANTS=() while [ "$#" -gt 0 ]; do case "$1" in --words-file) @@ -642,6 +796,23 @@ fm_afk_contract_parse_inputs() { # <args...>; sets WORDS, the CLAUSE_* arrays, case "$2" in ''|*[!0-9]*|0) fm_afk_contract_log "--spend must be a positive integer, got '$2'"; return 2 ;; esac SPEND=$2 shift 2 ;; + --grant) + [ "$#" -gt 1 ] || { fm_afk_contract_log '--grant requires a task id'; return 2; } + fm_afk_contract_grant_id_valid "$2" || { + fm_afk_contract_log "--grant must be a valid task id, got '$2'" + return 2 + } + for grant in "${MERGE_GRANTS[@]+"${MERGE_GRANTS[@]}"}"; do + [ "$grant" != "$2" ] || { + fm_afk_contract_log "--grant lists '$2' more than once" + return 2 + } + done + MERGE_GRANTS+=("$2") + shift 2 ;; + --grant=*) + fm_afk_contract_log '--grant takes a separate task-id argument' + return 2 ;; *) fm_afk_contract_log "unknown option '$1'" return 2 ;; @@ -773,13 +944,28 @@ fm_afk_contract_select_path() { # <args...> -> prints the record path chosen by printf '%s' "$path" } +# The record-mutating subcommands run inside the cross-subsystem lock, so no +# publication, replacement, or archive can land between another subsystem's +# authority read and the action it takes on that authority. +fm_afk_contract_locked_cmd() { # <command> [args...] + local rc=0 + fm_afk_contract_lock_hold || return 1 + trap 'fm_afk_contract_lock_release || true' EXIT + "$@" || rc=$? + trap - EXIT + fm_afk_contract_lock_release || true + return "$rc" +} + fm_afk_contract_main() { local cmd=${1:-} path [ -n "$cmd" ] || { fm_afk_contract_usage >&2; return 2; } shift case "$cmd" in propose) fm_afk_contract_cmd_propose "$@" ;; - confirm) [ "$#" -eq 0 ] || { fm_afk_contract_usage >&2; return 2; }; fm_afk_contract_cmd_confirm ;; + confirm) + [ "$#" -eq 0 ] || { fm_afk_contract_usage >&2; return 2; } + fm_afk_contract_locked_cmd fm_afk_contract_cmd_confirm ;; readback) path=$(fm_afk_contract_select_path "$@") || { fm_afk_contract_usage >&2; return 2; } [ -f "$path" ] || { fm_afk_contract_log "no record at $path"; return 1; } @@ -812,7 +998,11 @@ fm_afk_contract_main() { refused) path=$(fm_afk_contract_select_path "$@") || { fm_afk_contract_usage >&2; return 2; } fm_afk_contract_read_list "$path" refused ;; - archive) fm_afk_contract_cmd_archive ;; + grants) + path=$(fm_afk_contract_select_path "$@") || { fm_afk_contract_usage >&2; return 2; } + [ -f "$path" ] || { fm_afk_contract_log "no record at $path"; return 1; } + fm_afk_contract_read_grants "$path" ;; + archive) fm_afk_contract_locked_cmd fm_afk_contract_cmd_archive ;; archived) [ "$#" -eq 1 ] || { fm_afk_contract_usage >&2; return 2; } path="$(fm_afk_contract_archive_dir)/$1.afk-contract" diff --git a/bin/fm-afk-launch.sh b/bin/fm-afk-launch.sh index dc8a619f169..46dfff5fece 100755 --- a/bin/fm-afk-launch.sh +++ b/bin/fm-afk-launch.sh @@ -14,7 +14,10 @@ # only: no phone channel exists). The record is the posture in every harness. # Every harness retains daemon-backed away supervision. Pi and OMP extensions # stand down while state/.afk exists; docs/watcher-continuity.md owns that handoff. -# `start` and `start-native` require the confirmed record before daemon launch. +# `start` and `start-native` require the confirmed record for away daemon launch. +# Explicit FM_AFK_MODE=quiet enters attended supervision without a posture record; +# an existing quiet flag preserves that mode on refresh. Active away state must +# finish its normal return before quiet entry. # `stop` (the return, driven by bin/fm-afk-return.sh) shuts the daemon down, # clears state/.afk last, and archives the record under state/afk-contracts/. # @@ -38,11 +41,14 @@ # fm-afk-launch.sh propose [--words-file <path> | --words <text>] # [--action <verb> --object <text> --when <text> [--stop <text>]]... # [--expected-return <UTC ISO 8601>] [--spend <n>] +# [--grant <task-id>]... # Record the captain's away words and mandate # clause fields into a proposal and print the # read-back. Exit 3 when a clause was refused (its # missing part is named in the read-back); the # proposal still records it as refused. +# Repeatable --grant records captain-named task +# ids that may merge-when-green while away. # fm-afk-launch.sh confirm Promote the required proposal and print the entry # announcement; daemon launch follows. # fm-afk-launch.sh start Capture the captain pane, then (unless the daemon @@ -68,6 +74,9 @@ # terminal (default bin/fm-afk-start.sh), so a topology test can run a harmless # placeholder instead of a real daemon. FM_SUPERVISOR_TARGET/FM_SUPERVISOR_BACKEND # override the captured captain pane/backend (an isolated lab pane in tests). +# FM_AFK_MODE (away|quiet, default away) declares which mode a `start` entry +# requests; leave it unset for a plain refresh of an already-running daemon +# so its current mode is preserved (bin/fm-afk-start.sh fm_afk_flag_write). set -u FM_AFK_LAUNCH_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -194,6 +203,16 @@ fm_afk_launch_catchup_pending() { fm_afk_launch_record_require() { local record + # Quiet is attended presentation, never an away grant or an implicit return. + if [ "${FM_AFK_MODE:-$(fm_afk_mode "$FM_AFK_LAUNCH_STATE")}" = quiet ]; then + if fm_afk_contract_present "$FM_AFK_LAUNCH_STATE" || { + [ -e "$FM_AFK_LAUNCH_STATE/.afk" ] && [ "$(fm_afk_mode "$FM_AFK_LAUNCH_STATE")" != quiet ]; + }; then + fm_afk_launch_log "finish the active away return before entering quiet mode" + return 1 + fi + return 0 + fi record=$(fm_afk_contract_path "$FM_AFK_LAUNCH_STATE") if ! fm_afk_contract_present "$FM_AFK_LAUNCH_STATE"; then fm_afk_launch_log "a confirmed away-posture record is required; run propose and confirm before starting the daemon" @@ -230,7 +249,11 @@ fm_afk_launch_record_write() { # <backend> <target> <extra> } fm_afk_launch_flag_write() { - fm_afk_flag_write "$FM_AFK_LAUNCH_STATE" + # FM_AFK_MODE is the ONE place a caller declares which mode this entry + # requests (away, the unset default, or quiet - kunchenguid/firstmate#2356); + # fm_afk_flag_write itself preserves the on-disk mode when it is unset, so + # a plain /afk refresh of an already-quiet daemon never resets it. + fm_afk_flag_write "$FM_AFK_LAUNCH_STATE" "${FM_AFK_MODE:-}" } # Read the recorded terminal into FM_AFK_REC_BACKEND/FM_AFK_REC_TARGET. The third diff --git a/bin/fm-afk-return.sh b/bin/fm-afk-return.sh index 923f1259eee..6dff521ea61 100755 --- a/bin/fm-afk-return.sh +++ b/bin/fm-afk-return.sh @@ -6,7 +6,10 @@ # fm-afk-return.sh Stop away mode, render the return brief, and open/check the gate. # fm-afk-return.sh begin Same as the default command. # fm-afk-return.sh check Re-render the brief and close the gate only after blockers resolve. -# fm-afk-return.sh guard Read-only refusal while away or catch-up is pending. +# fm-afk-return.sh guard Read-only consult: exit 3 while away mode is still +# active, exit 4 while return catch-up is pending. +# fm-afk-return.sh quiet-off Explicit attended quiet cleanup, with no away return. +# fm-afk-return.sh catchup-summary Read-only catch-up projection for a reporting surface. # # THE RETURN BRIEF (stdout, on begin and on every check) is rendered from durable # records, never from conversation memory: the archived away-posture record @@ -37,9 +40,12 @@ # so a crash between stopping, wake presentation, and blocker handling fails # closed. It retains the presented wake, buffered-escalation, wedge-marker, # health, and posture-record evidence until every live open blocker is closed -# and `check` succeeds. Repeated begin/check calls are idempotent. `guard` -# never mutates state and is suitable for ordinary read entrypoints such as -# fm-bearings-snapshot.sh. +# and `check` succeeds. Repeated begin/check calls are idempotent. `guard` and +# `catchup-summary` never mutate state and are suitable for ordinary read +# entrypoints such as fm-bearings-snapshot.sh. `guard` separates its two +# refusal branches by exit status so a reporting surface can keep refusing +# during an active away window while still rendering the catch-up posture as +# content; this file owns the gate format both branches read. set -u SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -59,7 +65,7 @@ RETURN_GRACE=${FM_GUARD_GRACE:-300} CONTRACT="$SCRIPT_DIR/fm-afk-contract.sh" usage() { - sed -n '2,9p' "${BASH_SOURCE[0]}" | sed 's/^# \{0,1\}//' + sed -n '2,11p' "${BASH_SOURCE[0]}" | sed 's/^# \{0,1\}//' } clean_field() { @@ -135,6 +141,9 @@ window_start_epoch() { fi if [ -z "$epoch" ] && [ -f "$STATE/.afk" ]; then flag=$(head -1 "$STATE/.afk" 2>/dev/null || true) + case "$flag" in + ''|*[!0-9]*) flag=$(sed -n '2p' "$STATE/.afk" 2>/dev/null || true) ;; + esac case "$flag" in ''|*[!0-9]*) ;; *) epoch=$flag ;; esac fi case "$epoch" in ''|*[!0-9]*) printf '' ;; *) printf '%s' "$epoch" ;; esac @@ -245,15 +254,58 @@ clear_delivery_artifacts() { "$STATE/.subsuper-inject-wedged" } +# The lifecycle retention reasons the gate kept, one per line, empty when the +# gate was retained for open blockers alone. +gate_retention_reasons() { # <file> + local file=$1 tag kind text + while IFS="$(printf '\t')" read -r tag kind text; do + [ "$tag" = evidence ] && [ "$kind" = lifecycle ] || continue + printf '%s\n' "$text" + done < "$file" +} + +gate_has_blockers() { # <file> + grep -q "^blocker$(printf '\t')" "$1" 2>/dev/null +} + +# Read-only catch-up projection for a reporting surface such as +# fm-bearings-snapshot.sh: one tab-separated line +# `<open-blocker-count><TAB><first-retention-reason>`, and exit 1 when no gate +# is open. The reason field is empty when open blockers alone hold the gate. +catchup_summary() { + local count reason + [ -e "$GATE" ] || return 1 + count=$(grep -c "^blocker$(printf '\t')" "$GATE" 2>/dev/null || true) + case "$count" in ''|*[!0-9]*) count=0 ;; esac + reason=$(gate_retention_reasons "$GATE" | head -1) + printf '%s\t%s\n' "$count" "$reason" +} + return_guard() { - if [ -e "$STATE/.afk" ] || fm_afk_contract_present "$STATE"; then + local reasons + if fm_afk_contract_present "$STATE" || { [ -e "$STATE/.afk" ] && [ "$(fm_classify_afk_mode "$STATE")" != quiet ]; }; then printf 'fm-afk-return: away mode is still active; run bin/fm-afk-return.sh before ordinary captain work\n' >&2 return 3 fi if [ -e "$GATE" ]; then - printf 'fm-afk-return: return catch-up is pending; remediate or durably reclassify every listed blocker, then run bin/fm-afk-return.sh check\n' >&2 - print_blockers "$GATE" >&2 - return 3 + if gate_has_blockers "$GATE"; then + printf 'fm-afk-return: return catch-up is pending; remediate or durably reclassify every listed blocker, then run bin/fm-afk-return.sh check\n' >&2 + print_blockers "$GATE" >&2 + else + # No blocker row exists, so naming "every listed blocker" would ask for + # something the gate does not list. Name the lifecycle retention reason + # that actually holds it instead. + printf 'fm-afk-return: return catch-up is pending with no open blocker; clear the retention reason below, then run bin/fm-afk-return.sh check\n' >&2 + reasons=$(gate_retention_reasons "$GATE") + if [ -n "$reasons" ]; then + printf '%s\n' "$reasons" | while IFS= read -r text; do + printf 'catch-up retained: %s\n' "$text" >&2 + done + else + printf 'catch-up retained: the durable gate recorded no retention reason\n' >&2 + fi + fi + return 4 fi return 0 } @@ -264,7 +316,18 @@ health_snapshot() { # <evidence-file> local evidence=$1 beat_age lines="" beat_age=$(fm_path_age "$STATE/.last-watcher-beat") if [ -e "$STATE/.watcher-down" ]; then - lines="GAP: watcher downtime was detected during the away window (recovery marker present)" + # The marker survives past its episode in an acked:* state + # (fm-wake-lib.sh _fm_recovery_marker_ack); only pending:* and + # announced:* mean the downtime is still open. A marker this read + # cannot parse is treated the same as an open gap, conservatively. + if fm_recovery_marker_snapshot "$STATE/.watcher-down"; then + case "$FM_RECOVERY_MARKER_TOKEN" in + acked:*) : ;; + *) lines="GAP: watcher downtime was detected during the away window (recovery marker present)" ;; + esac + else + lines="GAP: watcher downtime was detected during the away window (recovery marker present)" + fi fi if [ -e "$STATE/.afk" ] && ! fm_afk_daemon_owns_supervision "$STATE"; then lines="$lines @@ -648,8 +711,26 @@ EOF main() { local mode=${1:-begin} rc window_epoch contract_epoch case "$mode" in - begin|check) ;; + quiet-off) + # Explicit attended exit uses the same daemon teardown, without an away + # brief, archive, or catch-up gate. Never consume an actual away record. + if fm_afk_contract_present "$STATE" || [ -e "$GATE" ] || { + [ -e "$STATE/.afk" ] && [ "$(fm_classify_afk_mode "$STATE")" != quiet ]; + }; then + printf 'fm-afk-return: quiet-off cannot clear an away lifecycle; use the ordinary return path\n' >&2 + return 3 + fi + "$SCRIPT_DIR/fm-afk-launch.sh" stop + return + ;; + begin|check) + if [ -e "$STATE/.afk" ] && [ "$(fm_classify_afk_mode "$STATE")" = quiet ] && ! fm_afk_contract_present "$STATE"; then + printf 'fm-afk-return: quiet mode remains active; use quiet-off only for an explicit quiet exit\n' >&2 + return 3 + fi + ;; guard) return_guard; return ;; + catchup-summary) catchup_summary; return ;; -h|--help|help) usage; return 0 ;; *) usage >&2; return 2 ;; esac diff --git a/bin/fm-afk-start.sh b/bin/fm-afk-start.sh index e86c54f170a..e268d2d61e0 100755 --- a/bin/fm-afk-start.sh +++ b/bin/fm-afk-start.sh @@ -3,8 +3,8 @@ # foreground process when one is not already alive. # # Usage: fm-afk-start.sh -# Sets state/.afk unless FM_AFK_STATE_PREPARED=1, checks -# state/.supervise-daemon.lock, and: +# Sets state/.afk (mode preserved on refresh, see fm_afk_flag_write) unless +# FM_AFK_STATE_PREPARED=1, checks state/.supervise-daemon.lock, and: # - prints "afk: daemon already running pid=<pid>" then exits 0 when that # lock is held by a live daemon (a REFRESH: no stale-artifact clear); # - otherwise clears any prior away session's stale escalation artifacts @@ -110,12 +110,24 @@ daemon_lock_held_by_live_daemon() { daemon_pid_matches "$pid" "$owner" } -fm_afk_flag_write() { # <state-dir> - local state=$1 lock="$1/.cursor-park-owner.lock" pending attempt=0 status=1 +fm_afk_flag_write() { # <state-dir> [mode] + local state=$1 requested_mode=${2:-} lock="$1/.cursor-park-owner.lock" \ + pending attempt=0 status=1 mode mkdir -p "$state" || return 1 [ ! -d "$state/.afk" ] || return 1 + # An explicit mode is a caller's deliberate request (a fresh /afk or /quiet + # entry). Omitted means "just refresh" (an already-running daemon, or + # recovery re-entering generically) and PRESERVES whatever mode is already + # on disk via fm_afk_mode - which itself falls back to "away" when nothing + # is on disk yet, so a genuinely fresh unspecified entry still defaults + # away. This is what keeps a refresh from silently flipping a captain's + # quiet mode back to away underneath them (kunchenguid/firstmate#2356). + case "$requested_mode" in + away|quiet) mode=$requested_mode ;; + *) mode=$(fm_afk_mode "$state") ;; + esac pending=$(mktemp "$state/.afk.pending.XXXXXX") || return 1 - date '+%s' > "$pending" || { rm -f "$pending"; return 1; } + { printf '%s\n' "$mode"; date '+%s'; } > "$pending" || { rm -f "$pending"; return 1; } while [ "$attempt" -lt 50 ]; do attempt=$((attempt + 1)) if fm_lock_try_acquire "$lock"; then diff --git a/bin/fm-agent-process-lib.sh b/bin/fm-agent-process-lib.sh new file mode 100644 index 00000000000..dcf4b59ff4e --- /dev/null +++ b/bin/fm-agent-process-lib.sh @@ -0,0 +1,111 @@ +#!/usr/bin/env bash +# Backend-neutral harness-process identity. +# Sourced by bin/backends/tmux.sh and bin/backends/herdr.sh. This file is +# sourced by scripts and has no side effects on source. +# +# Why one owner: every runtime backend that proves an agent is alive does it by +# attributing operating-system processes - the pane's foreground process group +# on tmux, Herdr's `pane process-info` view plus the pane shell's descendants +# on Herdr - and the two must agree on what a given process name means, or a +# harness one backend recognizes silently reads as a dead pane on the other. +# The classifier moved here verbatim from the tmux adapter, where it was born; +# docs/tmux-backend.md "Agent liveness probe" owns the empirical basis for the +# names below, and tests/fm-tmux-agent-liveness.test.sh plus +# tests/fm-harness-liveness-drift-live-e2e.test.sh keep them honest. + +# shellcheck source=bin/fm-session-lock-lib.sh +. "$(dirname -- "${BASH_SOURCE[0]}")/fm-session-lock-lib.sh" +# shellcheck source=bin/fm-gemini-lib.sh +. "$(dirname -- "${BASH_SOURCE[0]}")/fm-gemini-lib.sh" + +# fm_agent_process_classify_name: the single owner of the process-name +# vocabulary shared by every liveness signal - `agent` for a verified harness, +# `shell` for an idle login/interactive shell, `other` for anything else. +# Keeping one classifier means independent name sources (a kernel process +# name, an argv[0], a rendered pane title) can never drift into disagreeing +# about what a given name means. +fm_agent_process_classify_name() { # <path> [argv0] -> agent|shell|other + local path=$1 argv0=${2:-} base + base=${path##*/} + base=${base#-} + case "$base" in + # muse is anchored rather than globbed like its neighbours: its installed + # binary is muse-bin-<version> (the launcher execs it, so the version is the + # live process name and changes on every auto-update), and unlike `claude` or + # `codex` the substring `muse` is a common English fragment - a *muse* glob + # would classify musescore or amuse as a live agent pane. The install path + # cannot carry it either: ~/.local/bin/muse-bin-<version> has no `muse` path + # COMPONENT, so the fm_harness_path_name fallback below never fires for it. + muse|muse-bin-*) printf 'agent' ;; + # omp (Oh My Pi) is anchored for the same reason as muse: its live process + # name is the bare word `omp` (verified, omp 18.1.11) and a glob would claim + # unrelated commands such as ompd or comp. + *claude*|*codex*|*opencode*|*grok*|*kimi*|*rovo*|pi|pi-signed|pi-launcher|Pi|omp) printf 'agent' ;; + # agy (Antigravity CLI) is anchored for the same reason as muse and omp: its + # live process name is the bare word `agy` (verified, agy 1.2.0: a Go-compiled + # single binary, comm=agy with argv[0]=agy), and a glob would claim + # unrelated commands containing that fragment. + agy) printf 'agent' ;; + zsh|bash|sh|dash|ash|ksh|mksh|tcsh|csh|fish) printf 'shell' ;; + *) + if fm_harness_path_name "$path" >/dev/null || fm_harness_path_name "$argv0" >/dev/null; then + printf 'agent' + # cursor-agent runs as a bundled node script, so tmux reports the pane + # command as a bare `node` that no name pattern above can own, and its + # other installed name is the far-too-generic `agent` (verified live on + # cursor-agent 2026.08.11-e8db854: #{pane_current_command} is `node` while + # `ps -o comm=` carries the cursor-agent install path). Identity therefore + # comes from the narrowed structural rule in bin/fm-cursor-lib.sh, which + # demands Cursor's own name or install tree in the path or argv[0]. An + # unrelated `node` or `agent` matches nothing here and stays `other`, + # which the callers fold into `ambiguous` rather than `dead`, so a + # stranger's node pane is never reported as an agent-free pane. + elif fm_cursor_process_matches "${path:-$argv0}" '' "$argv0"; then + printf 'agent' + else + printf 'other' + fi + ;; + esac +} + +# fm_agent_process_classify: one process, from every identity surface a +# backend can hand over, as agent|shell|other. Any single surface naming a +# verified harness carries `agent`, because a false negative is the one outcome +# that launches a duplicate agent onto a live worktree; `shell` needs every +# readable surface to agree the process is a shell; anything else is `other`. +# +# <name> the kernel process name (ps comm, or Herdr's process-info .name): +# on Linux the exec name, on macOS argv[0] truncated to 16 bytes. +# <argv0> argv[0] as the process reports it - a bare name or an install +# path, whichever the launcher used (empty when unknown). +# <args> the flattened command line, read only for the node-bundle +# harnesses whose identity sits in argv[1] (bin/fm-gemini-lib.sh). +# [pid] when given, lets the Gemini rule read argv boundaries from the +# live process instead of the flattened line. +fm_agent_process_classify() { # <name> <argv0> <args> [pid] -> agent|shell|other + local name=${1:-} argv0=${2:-} args=${3:-} pid=${4:-} by_name by_argv0 + by_name=$(fm_agent_process_classify_name "$name" "$argv0") + [ "$by_name" != agent ] || { printf 'agent'; return 0; } + if [ -n "$argv0" ]; then + # argv[0] is classified as a path in its own right, so a bare `pi` or a + # `-zsh` login name reads by basename and an install path by component. + by_argv0=$(fm_agent_process_classify_name "$argv0" "$argv0") + [ "$by_argv0" != agent ] || { printf 'agent'; return 0; } + else + by_argv0=$by_name + fi + if [ -n "$pid" ] && fm_gemini_pid_is_gemini "$pid"; then + printf 'agent' + return 0 + fi + if [ -n "$args" ] && fm_gemini_args_are_gemini "$args"; then + printf 'agent' + return 0 + fi + if [ "$by_name" = shell ] && [ "$by_argv0" = shell ]; then + printf 'shell' + else + printf 'other' + fi +} diff --git a/bin/fm-agy-trust-lib.sh b/bin/fm-agy-trust-lib.sh index 40f48ba546d..d4a94c2a7dd 100644 --- a/bin/fm-agy-trust-lib.sh +++ b/bin/fm-agy-trust-lib.sh @@ -22,7 +22,8 @@ # - LOCKED with the repository's ownership-and-liveness lock (fm_lock_*): a # stale lock is reclaimed only when its holder PID is provably dead, so a slow # mutator's LIVE lock is never stolen, and only the acquiring process releases. -# - ATOMIC - written to a sibling temp file and mv'd into place. +# - ATOMIC - written to a sibling temp file and mv'd into place after checking +# that an external writer has not changed the original bytes. # - FAIL-CLOSED - a settings file that is not valid JSON is left UNTOUCHED # rather than clobbered, and a missing file is created minimally only for add. # @@ -80,7 +81,7 @@ fm_agy_trust_add() { # <abs-path> # the caller (fm-spawn), not within this library. # shellcheck disable=SC2034 FM_AGY_TRUST_ADDED= - local path=$1 file lock tmp rc=0 already + local path=$1 file lock tmp rc=0 already before command -v jq >/dev/null 2>&1 || { echo "warning: jq unavailable; cannot add agy workspace trust for $path" >&2; return 1; } case "$path" in /*) : ;; @@ -91,8 +92,11 @@ fm_agy_trust_add() { # <abs-path> mkdir -p "$(dirname "$file")" 2>/dev/null || true lock="$file.fm-trust.lock" fm_agy_trust_lock_acquire "$lock" || { echo "warning: could not lock agy settings to add workspace trust for $path" >&2; return 1; } + before=absent if [ -e "$file" ]; then - if ! jq -e . "$file" >/dev/null 2>&1; then + [ -f "$file" ] && [ -O "$file" ] && [ -w "$file" ] && [ ! -L "$file" ] || { fm_agy_trust_lock_release "$lock"; return 1; } + before=$(cksum < "$file") || { fm_agy_trust_lock_release "$lock"; return 1; } + if ! jq -e 'type == "object" and ((.trustedWorkspaces // []) | type == "array")' "$file" >/dev/null 2>&1; then echo "warning: agy settings at $file is not valid JSON; leaving it untouched (workspace trust add skipped for $path)" >&2 fm_agy_trust_lock_release "$lock" return 1 @@ -111,6 +115,10 @@ fm_agy_trust_add() { # <abs-path> tmp=$(mktemp "$file.fm-trust.XXXXXX" 2>/dev/null) || { fm_agy_trust_lock_release "$lock"; return 1; } jq -n --arg p "$path" '{trustedWorkspaces: [$p]}' > "$tmp" 2>/dev/null || rc=1 fi + if [ "$before" != "$(if [ -e "$file" ]; then cksum < "$file"; else printf absent; fi)" ]; then + echo "warning: agy settings changed during trust mutation; preserving the external write" >&2 + rc=1 + fi if [ "$rc" -eq 0 ] && [ -s "$tmp" ]; then mv -f "$tmp" "$file" || rc=1 else @@ -147,7 +155,7 @@ fm_agy_trust_rollback() { # <abs-path> <marker> # contention, write failure) so the caller can treat it as an incomplete teardown # and retry. fm_agy_trust_remove() { # <abs-path> - local path=$1 file lock tmp rc=0 + local path=$1 file lock tmp rc=0 before command -v jq >/dev/null 2>&1 || { echo "warning: jq unavailable; cannot remove agy workspace trust for $path" >&2; return 1; } case "$path" in /*) : ;; @@ -158,6 +166,8 @@ fm_agy_trust_remove() { # <abs-path> [ -e "$file" ] || return 0 lock="$file.fm-trust.lock" fm_agy_trust_lock_acquire "$lock" || { echo "warning: could not lock agy settings to remove workspace trust for $path" >&2; return 1; } + [ -f "$file" ] && [ -O "$file" ] && [ -w "$file" ] && [ ! -L "$file" ] || { fm_agy_trust_lock_release "$lock"; return 1; } + before=$(cksum < "$file") || { fm_agy_trust_lock_release "$lock"; return 1; } if ! jq -e . "$file" >/dev/null 2>&1; then echo "warning: agy settings at $file is not valid JSON; leaving it untouched (workspace trust remove skipped for $path)" >&2 fm_agy_trust_lock_release "$lock" @@ -165,6 +175,10 @@ fm_agy_trust_remove() { # <abs-path> fi tmp=$(mktemp "$file.fm-trust.XXXXXX" 2>/dev/null) || { fm_agy_trust_lock_release "$lock"; return 1; } jq --arg p "$path" 'if (.trustedWorkspaces | type) == "array" then .trustedWorkspaces |= map(select(. != $p)) else . end' "$file" > "$tmp" 2>/dev/null || rc=1 + if [ "$before" != "$(if [ -e "$file" ]; then cksum < "$file"; else printf absent; fi)" ]; then + echo "warning: agy settings changed during trust mutation; preserving the external write" >&2 + rc=1 + fi if [ "$rc" -eq 0 ] && [ -s "$tmp" ]; then mv -f "$tmp" "$file" || rc=1 else diff --git a/bin/fm-agy-trust.sh b/bin/fm-agy-trust.sh new file mode 100755 index 00000000000..77a62ac5e19 --- /dev/null +++ b/bin/fm-agy-trust.sh @@ -0,0 +1,120 @@ +#!/usr/bin/env bash +# Pre-register Antigravity CLI's workspace trust for the isolated task worktree +# a ship/scout spawn is about to launch an agy crewmate into, so the worker +# reaches its brief in the worktree instead of parking on the folder-trust +# dialog and running its turn in agy's own scratch directory. +# +# Usage: fm-agy-trust.sh <worktree> <project> +# <worktree> the isolated task worktree this spawn launches into +# <project> the primary checkout that worktree belongs to +# Prints one line naming what it registered; refuses loudly on anything else. +# +# WHY THIS EXISTS. agy 1.2.0 gates a folder it has never seen behind +# "Do you trust the contents of this project?" and no launch flag suppresses +# it (`agy --help` lists none). Answering appends the folder to the +# `trustedWorkspaces` array of ${HOME}/.gemini/antigravity-cli/settings.json, +# and agy honours an entry written there ahead of launch: verified live under a +# throwaway HOME, a pre-registered folder launched straight into its turn while +# an unregistered sibling parked on the dialog (docs/verification/agy.md). agy +# compares the pane's LOGICAL working directory, not its resolved path (a +# symlinked cwd with only the real path registered still parked), so both the +# logical path and its resolved form are recorded when they differ. +# +# bin/fm-spawn.sh keeps a post-launch gate as the backstop: it answers the +# dialog if one renders anyway and never counts a busy turn as ready on a path +# that was neither pre-registered here nor answered there. +# +# THE SCOPE TEST IS THE SAFETY PROPERTY and mirrors bin/fm-claude-trust.sh: +# <worktree> must be a LINKED git worktree - its own git dir, sharing +# <project>'s common dir - whose top level is exactly the resolved argument. A +# primary checkout, a worktree of an unrelated repo, a subdirectory of a +# worktree, a plain directory, and a home directory are each refused with a +# non-zero exit, never a warning and never a silent skip. Only the launching +# user's own store is written, it must be a regular file this uid owns, every +# unrelated key and entry is preserved, and the replacement is atomic. +set -u +unset CDPATH \ + GIT_DIR GIT_WORK_TREE GIT_COMMON_DIR GIT_OBJECT_DIRECTORY GIT_INDEX_FILE \ + GIT_ALTERNATE_OBJECT_DIRECTORIES GIT_CEILING_DIRECTORIES GIT_NAMESPACE \ + GIT_DISCOVERY_ACROSS_FILESYSTEM GIT_CONFIG GIT_CONFIG_GLOBAL \ + GIT_CONFIG_SYSTEM GIT_CONFIG_NOSYSTEM GIT_CONFIG_COUNT + +[ "$#" -eq 2 ] || { echo "usage: fm-agy-trust.sh <worktree> <project>" >&2; exit 2; } +WT_ARG=$1 +PROJ_ARG=$2 + +refuse() { echo "error: refusing to pre-register agy trust: $1" >&2; exit 1; } + +real_dir() { (cd -P -- "$1" 2>/dev/null && pwd -P); } +logical_dir() { (cd -- "$1" 2>/dev/null && pwd -L); } +real_file() { node -e 'process.stdout.write(require("node:fs").realpathSync(process.argv[1]))' "$1" 2>/dev/null; } + +common_dir_of() { + local dir=$1 common + common=$(git -C "$dir" rev-parse --git-common-dir 2>/dev/null) || return 1 + (cd -P -- "$dir" && real_dir "$common") +} + +WT_REAL=$(real_dir "$WT_ARG") || true +[ -n "$WT_REAL" ] || refuse "worktree '$WT_ARG' is not an accessible directory" +WT_LOGICAL=$(logical_dir "$WT_ARG") || true +[ -n "$WT_LOGICAL" ] || WT_LOGICAL=$WT_REAL +PROJ_REAL=$(real_dir "$PROJ_ARG") || true +[ -n "$PROJ_REAL" ] || refuse "project '$PROJ_ARG' is not an accessible directory" + +[ -n "${HOME:-}" ] || refuse "HOME is not set, so agy's settings store cannot be located" +HOME_REAL=$(real_dir "$HOME") || true +[ -n "$HOME_REAL" ] || refuse "HOME '$HOME' is not an accessible directory" +[ "$WT_REAL" != "$HOME_REAL" ] || refuse "'$WT_REAL' is the home directory, not a task worktree" + +WT_TOP=$(git -C "$WT_REAL" rev-parse --show-toplevel 2>/dev/null) || true +[ -n "$WT_TOP" ] || refuse "'$WT_REAL' is not inside a git repository" +WT_TOP_REAL=$(real_dir "$WT_TOP") || true +[ "$WT_TOP_REAL" = "$WT_REAL" ] || refuse "'$WT_REAL' is not a worktree root (its root is '${WT_TOP_REAL:-unresolvable}')" + +WT_GIT_DIR=$(git -C "$WT_REAL" rev-parse --absolute-git-dir 2>/dev/null) || true +[ -n "$WT_GIT_DIR" ] || refuse "'$WT_REAL' has no resolvable git directory" +WT_GIT_DIR=$(real_dir "$WT_GIT_DIR") || true +[ -n "$WT_GIT_DIR" ] || refuse "'$WT_REAL' has an unresolvable git directory" +WT_COMMON=$(common_dir_of "$WT_REAL") || true +[ -n "$WT_COMMON" ] || refuse "'$WT_REAL' has no resolvable git common directory" +[ "$WT_GIT_DIR" != "$WT_COMMON" ] || refuse "'$WT_REAL' is a primary checkout, not an isolated worktree" + +PROJ_COMMON=$(common_dir_of "$PROJ_REAL") || true +[ -n "$PROJ_COMMON" ] || refuse "project '$PROJ_REAL' is not inside a git repository" +[ "$WT_COMMON" = "$PROJ_COMMON" ] || refuse "'$WT_REAL' is not a worktree of project '$PROJ_REAL'" + +command -v node >/dev/null 2>&1 || refuse "node is required to record workspace trust and was not found on PATH" + +STORE_DIR="$HOME_REAL/.gemini/antigravity-cli" +mkdir -p "$STORE_DIR" 2>/dev/null || true +STORE_DIR_REAL=$(real_dir "$STORE_DIR") || true +[ -n "$STORE_DIR_REAL" ] || refuse "agy settings directory '$STORE_DIR' does not exist and could not be created" +STORE="$STORE_DIR_REAL/settings.json" +if [ -L "$STORE" ]; then + STORE_REAL=$(real_file "$STORE") || true + [ -n "$STORE_REAL" ] || refuse "'$STORE' is a symlink whose target cannot be resolved" + STORE=$STORE_REAL +fi +if [ -e "$STORE" ]; then + [ -f "$STORE" ] || refuse "'$STORE' is not a regular file" + [ -O "$STORE" ] || refuse "'$STORE' is not owned by this user" + [ -w "$STORE" ] || refuse "'$STORE' is not writable" +fi + +# The shared library is the sole settings mutation owner, including locking, +# atomic replacement, external-write detection, and created/preexisting identity. +# This command owns the upstream linked-worktree scope check above. +# shellcheck source=bin/fm-agy-trust-lib.sh +. "$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/fm-agy-trust-lib.sh" +export FM_AGY_SETTINGS_OVERRIDE="$STORE" +fm_agy_trust_add "$WT_LOGICAL" || refuse "could not record trust for '$WT_LOGICAL' in '$STORE'" +if [ "$WT_LOGICAL" != "$WT_REAL" ]; then + fm_agy_trust_add "$WT_REAL" || refuse "could not record trust for '$WT_REAL' in '$STORE'" +fi + +if [ "$WT_LOGICAL" != "$WT_REAL" ]; then + echo "trusted: $WT_LOGICAL ($WT_REAL)" +else + echo "trusted: $WT_REAL" +fi diff --git a/bin/fm-backend.sh b/bin/fm-backend.sh index f6eada330b7..d4644f7bffb 100644 --- a/bin/fm-backend.sh +++ b/bin/fm-backend.sh @@ -927,14 +927,17 @@ fm_backend_target_exists() { # <backend> <target> [expected-label] [expected-ta # ambiguous - the endpoint exists but its process cannot be attributed. # unreadable - a target or inventory read failed or contradicted itself. # unverified - this backend has no recovery classifier. -# Only `dead` and `missing` license recovery. The tmux adapter requires a -# successful session inventory and returns `missing` only when it omits the -# exact window; the Herdr adapter reuses its husk classifier plus a -# stale-registration cross-check (fm_backend_herdr_agent_state owns that -# contract), and maps a positively stopped session server to `missing` only -# in this recovery-grade view. Zellij remains unverified because its secondmate ghost-tab and -# agent-process recovery path has not been empirically validated. Orca and cmux -# do not support secondmate spawns. +# Only `dead` and `missing` license recovery. Every `alive` is proven at +# process level through the shared classifier in bin/fm-agent-process-lib.sh, +# never from a registration or a rendered title alone. The tmux adapter +# requires a successful session inventory and returns `missing` only when it +# omits the exact window; the Herdr adapter reuses its strict husk classifier - +# which verifies a registered agent against `pane process-info` and the real +# process table, so a registration Herdr kept over a shell-only pane reads +# `dead` here (issue #4115) - then maps a positively stopped session server to +# `missing` only in this recovery-grade view. Zellij remains unverified because +# its secondmate ghost-tab and agent-process recovery path has not been +# empirically validated. Orca and cmux do not support secondmate spawns. fm_backend_agent_state() { # <backend> <target> local backend=$1 target=$2 fm_backend_source "$backend" || { printf 'unverified'; return 0; } diff --git a/bin/fm-backlog-handoff.sh b/bin/fm-backlog-handoff.sh index b40aae28758..dff23761c1f 100755 --- a/bin/fm-backlog-handoff.sh +++ b/bin/fm-backlog-handoff.sh @@ -317,7 +317,7 @@ warn_stale_public_commitments() { # <secondmate-id> <moved-key>... out=$("$SCRIPT_DIR/fm-public-followup.sh" guard-work main "$key" 2>/dev/null) || rc=$? [ "$rc" -ne 0 ] || continue [ -z "$out" ] || printf '%s\n' "$out" >&2 - printf 'warning: %s still owes a public reply bound to main/%s; rebind it to secondmate:%s (tasks-axi public-followup bind-work, then bin/fm-public-followup.sh register <obligation-id> --relation <relation-id> --work-home secondmate:%s --work-id %s --generation <n>) or the promised reply will be reconciled against work this home no longer owns.\n' \ + printf 'warning: %s still owes a public reply bound to main/%s; rebind it to secondmate:%s (bin/fm-tasks-axi.sh public-followup bind-work, then bin/fm-public-followup.sh register <obligation-id> --relation <relation-id> --work-home secondmate:%s --work-id %s --generation <n>) or the promised reply will be reconciled against work this home no longer owns.\n' \ "$key" "$key" "$id" "$id" "$key" >&2 done if fm_pf_relay_active "$FM_HOME" && fm_pf_has_delivered_open_loops "$STATE"; then diff --git a/bin/fm-backlog-transition-lib.sh b/bin/fm-backlog-transition-lib.sh index 4442ec3c64f..c116d016b0e 100644 --- a/bin/fm-backlog-transition-lib.sh +++ b/bin/fm-backlog-transition-lib.sh @@ -73,6 +73,17 @@ FM_BACKLOG_ROW_HOLD_KIND= # shellcheck disable=SC2034 # Output global, read by the sourcing caller. FM_BACKLOG_CLOSE_REPLAY_RESULT= +# Bounded execution is fm-timeout-lib.sh's alone; source it rather than +# re-deriving a deadline here. It is stateless, so the memoisation reason this +# library does not source fm-tasks-axi-lib.sh does not apply. +# shellcheck source=bin/fm-timeout-lib.sh disable=SC1091 +. "$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/fm-timeout-lib.sh" + +# Latched when a row read hits its bound. fm_backlog_row_show runs inside a +# command substitution, so the subshell can READ this latch but cannot set it; +# the callers that capture its status own the write. +FM_BACKLOG_ROW_SHOW_WEDGED=0 + # Emit each byte of a value as a decimal number, locale-independently. # Deliberately perl rather than od: the spawn and teardown lifecycle runs under a # curated PATH (tests/fm-teardown.test.sh make_path_without_lsof pins that set) @@ -382,22 +393,64 @@ fm_tasks_axi() { exit 127 } -# Print one row's `tasks-axi show` output (plus stderr); the exit status is -# tasks-axi's. Extra flags (such as --full) are passed through. +# Print one row's `tasks-axi show` output (plus stderr) from the addressing +# fm_backlog_tasks_axi_addressing resolved, with `--file` only for the markdown +# backend. Addressing or backend-resolution errors return before tasks-axi runs; +# otherwise its exit status is preserved. Extra flags (--full) pass through. +# +# Every read is bounded, because a wedged backend read here is what blinds a +# whole session start: bin/fm-bootstrap.sh's reconcile and close-replay sweeps +# call this once per item, and one unbounded read consumes the entire +# FM_SESSION_START_TIMEOUT and truncates the digest before the wake queue, +# supervision instructions, fleet state and context sections ever print. The +# bound turns that into a loud partial reconcile: the caller reports the item it +# could not read and moves to the next one. +# +# A per-item bound alone is not enough on a home carrying a large fleet, because +# N wedged items still cost N bounds and the digest is truncated anyway. So the +# first bound hit latches FM_BACKLOG_ROW_SHOW_WEDGED and every later read in the +# same sweep returns immediately, still naming its own item so nothing is +# silently skipped. This function only READS that latch: it runs inside a +# command substitution, and a write here would die with the subshell, so the +# callers that capture its status set it. The latch is deliberately +# process-wide because these scripts are short-lived and a backend that wedged +# once will wedge again within the same run. fm_backlog_row_show() { # <resolved-data-dir> <id> [flag...] - local data=$1 id=$2 addressing_status + local data=$1 id=$2 out status addressing_status secs=${FM_BACKLOG_ROW_TIMEOUT_SECS:-10} shift 2 + # A non-positive bound is not a bound (fm-timeout-lib.sh), and a padded zero + # such as 00 is still zero, so the digits test alone would let the very read + # this bound exists to prevent back in. Compare arithmetically, tolerating a + # value too large for the shell to compare at all. + case "$secs" in ''|*[!0-9]*) secs=10 ;; esac + [ "$secs" -gt 0 ] 2>/dev/null || secs=10 fm_backlog_tasks_axi_addressing "$data" addressing_status=$? if [ "$addressing_status" -ne 0 ]; then [ -z "${FM_BACKLOG_TRANSITION_ERROR:-}" ] || printf '%s\n' "$FM_BACKLOG_TRANSITION_ERROR" >&2 return "$addressing_status" fi + if [ "$FM_BACKLOG_ROW_SHOW_WEDGED" = 1 ]; then + printf 'tasks-axi show %s skipped: the backlog backend already exceeded its %ss read bound\n' "$id" "$secs" + return 124 + fi if [ -n "$FM_BACKLOG_AXI_FILE" ]; then - (cd "$FM_BACKLOG_AXI_ROOT" 2>/dev/null && fm_tasks_axi show "$id" "$@" --file "$FM_BACKLOG_AXI_FILE" 2>&1) + set -- "$@" --file "$FM_BACKLOG_AXI_FILE" + fi + # shellcheck disable=SC2016 # Expansion is deliberately deferred to the child shell. + out=$(fm_run_timed "$secs" bash -c 'cd "$1" 2>/dev/null || exit 1; shift; exec tasks-axi show "$@"' \ + _ "$FM_BACKLOG_AXI_ROOT" "$id" "$@" 2>&1) + status=$? + # A backend that wrote a header or a progress line before wedging leaves that + # fragment as the first output line, and every caller reads the first line as + # the failure reason. Whatever a timed-out read managed to emit is incomplete + # by definition, so the bound speaks for it instead. + if [ "$status" -eq 124 ]; then + printf 'tasks-axi show %s exceeded its %ss backlog read bound\n' "$id" "$secs" else - (cd "$FM_BACKLOG_AXI_ROOT" 2>/dev/null && fm_tasks_axi show "$id" "$@" 2>&1) + printf '%s\n' "$out" fi + return "$status" } fm_backlog_row_list() { # <resolved-data-dir> [flag...] @@ -436,6 +489,7 @@ fm_backlog_row_probe() { # <data-dir> <id> fi out=$(fm_backlog_row_show "$data" "$id") command_status=$? + [ "$command_status" -ne 124 ] || FM_BACKLOG_ROW_SHOW_WEDGED=1 if [ "$command_status" -ne 0 ]; then if printf '%s\n' "$out" | grep -q '^code: NOT_FOUND$'; then FM_BACKLOG_ROW_RESULT=not_found @@ -559,6 +613,7 @@ fm_backlog_retain() { # <data-dir> <id> [flag...] if [ -n "$deliverable" ]; then out=$(fm_backlog_row_show "$data" "$id" --full) command_status=$? + [ "$command_status" -ne 124 ] || FM_BACKLOG_ROW_SHOW_WEDGED=1 if [ "$command_status" -ne 0 ]; then FM_BACKLOG_TRANSITION_ERROR=$(printf '%s\n' "$out" | sed -n '1p') [ -n "$FM_BACKLOG_TRANSITION_ERROR" ] \ diff --git a/bin/fm-bearings-board.sh b/bin/fm-bearings-board.sh index b25ad5e9c10..2cb9506d721 100755 --- a/bin/fm-bearings-board.sh +++ b/bin/fm-bearings-board.sh @@ -72,6 +72,13 @@ # the template may display the routing id. Anything else refuses before the # existing board is touched. # +# Every Underway row likewise carries a non-empty `name`: the durable task name +# when known, otherwise its durable identifier. +# A Charted Next row MAY carry `filed`, the durable filed date (YYYY-MM-DD, or +# that date with a UTC timestamp) the template orders the section by, newest +# first; a row with no comparable date keeps its payload order after every dated +# row. Anything else in that field refuses rather than sorting on garbage. +# # The board path is stable - $FM_HOME/.lavish/bearings-board.html - so a # re-invocation rebuilds the same file in place, which keeps the same Lavish # session URL and the same canonical process-event source id. Injection escapes @@ -109,6 +116,17 @@ validate_payload() { # <data.json> def nonempty_string: type == "string" and length > 0; def slug($max): type == "string" and test("^[A-Za-z0-9._-]{1," + ($max | tostring) + "}$"); def repo_marker: has("repo") and (.repo == null or (.repo | type == "string")); + def name_marker: has("name") and (.name | nonempty_string); + def valid_filed: + . as $filed + | type == "string" + and test("^[0-9]{4}-[0-9]{2}-[0-9]{2}(T[0-9]{2}:[0-9]{2}:[0-9]{2}Z)?$") + and (if test("T") + then try ((fromdateiso8601 | strftime("%Y-%m-%dT%H:%M:%SZ")) == $filed) catch false + else try (((. + "T00:00:00Z") | fromdateiso8601 | strftime("%Y-%m-%d")) == $filed) catch false + end); + def optional_filed: + (has("filed") | not) or (.filed == null) or (.filed | valid_filed); def optional_string($name): (has($name) | not) or (.[$name] | type == "string"); def optional_https_url($name): (has($name) | not) @@ -152,7 +170,7 @@ validate_payload() { # <data.json> and ([.options[].value] | index("reconcile") == null) and (if .type == "merge" then (.risk | nonempty_string) else true end); def underway_item: - type == "object" and repo_marker and (.id | nonempty_string) + type == "object" and repo_marker and name_marker and (.id | nonempty_string) and (.state | nonempty_string) and (.doing | nonempty_string) and (.kind | nonempty_string); def landed_item: type == "object" and repo_marker and (.id | nonempty_string) @@ -164,6 +182,7 @@ validate_payload() { # <data.json> and (.title | nonempty_string) and (.reason | type == "string") and (.dispatchable | type == "boolean") and ((has("kind") | not) or (.kind == "queued" or .kind == "warning")) + and optional_filed and (if .kind == "warning" then .dispatchable == false else true end); type == "object" and (.schema == $schema) diff --git a/bin/fm-bearings-snapshot.sh b/bin/fm-bearings-snapshot.sh index f5b565bae85..8b5b7307a91 100755 --- a/bin/fm-bearings-snapshot.sh +++ b/bin/fm-bearings-snapshot.sh @@ -25,7 +25,10 @@ # decisions from report or visual-review prose or reimplements snapshot semantics. # Underway (in_flight) projects every main live worker plus every active child # from every readable secondmate ledger, independently of that home's -# bearings_state. A home classified captain_decision because it has an open +# bearings_state. Each row's name is the durable task title when nonblank and +# its durable task id otherwise, so renderers always receive a task-identifying +# label instead of having to substitute run status. A home classified +# captain_decision because it has an open # captain hold still contributes each working child as its own Underway row; # the home row on secondmates[] keeps the decision and gate classification. # Captain-hold placement follows the canonical snapshot's hold_bucket and @@ -41,12 +44,23 @@ # Aging is a projection safety net only; the durable # deferral remains re-holding with --until. # +# Ordinary Charted Next gates are ordered by durable filed date, newest first, +# before the FM_BEARINGS_GATES bound is applied. Gates without a comparable filed +# date keep their input order after dated gates. The synthetic (return-catchup) +# posture row is reserved ahead of that ordering and bound so it always surfaces. +# # Main-home inventory validity comes from the canonical snapshot's main_inventory # object (orphan structured in-flight without meta, unstructured current rows). # Bearings never invents Underway rows from backlog-only ids; it discloses those # gaps in omitted[] and, when invalid, a Charted Next gate line so the four-section # chat cannot claim an empty fleet while main current state is broken. # +# An open away-return catch-up is disclosed the same way, as a single action-free +# (return-catchup) gate row naming the blockers left to clear or the reason the +# catch-up was retained. Reporting is not ordinary captain work, so the gate never +# suppresses the digest; an ACTIVE away window still refuses, because the right +# answer there is to run the return first. bin/fm-afk-return.sh owns the gate. +# # The landed section merges this home's Done with the canonical snapshot's # secondmate_landed roll-up (fm-fleet-snapshot.sh), so merges a secondmate managed - # recorded in ITS OWN backlog, never the main one - are visible. It stays bounded by @@ -129,12 +143,14 @@ Default collection performs bounded concurrent remote-ledger reads for registere remote homes under one shared snapshot budget and may refresh the parent-side cache. --include-prs additionally performs live GitHub discovery and checks. -Default fields: schema, home, generated, prs, in_flight{id,kind,state,repo,doing}, +Default fields: schema, home, generated, prs, in_flight{id,kind,state,repo,name,doing}, secondmates{id,state,doing,provenance,freshness,age_seconds,contradiction,reason}, secondmate_reconcile{id,spawn_gen,host,kind,ids}, decisions_open{id,key,verb,summary,owner}, landed{id,what,artifact,owner}, - gates{id,title,blocked_by,reason,owner}, reports{id,path}, recorded_prs{id,url}, + gates{id,title,blocked_by,reason,owner,filed}, reports{id,path}, recorded_prs{id,url}, unhealthy_endpoints{...} (only when non-empty), omitted{surface,reveal}. +Default gates are selected newest filed first before their bound; undated gates + retain input order after dated gates. landed merges this home's Done with registered secondmate homes' Done, bounded by a per-home cap (FM_BEARINGS_LANDED_PER_HOME) and an overall cap (FM_BEARINGS_LANDED), with omitted[] disclosure. Default selection is balanced across deterministic home @@ -191,10 +207,29 @@ command -v jq >/dev/null 2>&1 || { echo "fm-bearings-snapshot: jq not found" >&2 # JSON documents that grow with the fleet are bound with --slurpfile and unwrapped # by `doc(...)`, never with --argjson; bin/fm-fleet-snapshot.sh owns the reason. -# The deterministic return-catch-up owner must clear before this or any other -# ordinary captain request proceeds. Bearings does not reproduce that policy; -# it only consults the shared read-only gate. -"$SCRIPT_DIR/fm-afk-return.sh" guard || exit $? +# The shared read-only away-return owner is consulted, not obeyed. An active +# away window still refuses here: the correct answer to a bearings request then +# is to run the return first. Return CATCH-UP is different - the captain is +# back and asking for the picture, so the catch-up posture is reported as +# content (a Charted Next gate row) and collection continues. bin/fm-afk-return.sh +# owns both the gate format and the branch distinction; bearings reproduces +# neither. Acting on the fleet still waits for its `check`. +RETURN_CATCHUP=null +GUARD_RC=0 +GUARD_ERR=$("$SCRIPT_DIR/fm-afk-return.sh" guard 2>&1 >/dev/null) || GUARD_RC=$? +if [ "$GUARD_RC" -ne 0 ] && [ "$GUARD_RC" -ne 4 ]; then + [ -z "$GUARD_ERR" ] || printf '%s\n' "$GUARD_ERR" >&2 + exit "$GUARD_RC" +fi +if [ "$GUARD_RC" -eq 4 ]; then + CATCHUP_LINE=$("$SCRIPT_DIR/fm-afk-return.sh" catchup-summary) || CATCHUP_LINE="" + CATCHUP_BLOCKERS=${CATCHUP_LINE%%$'\t'*} + case "$CATCHUP_BLOCKERS" in ''|*[!0-9]*) CATCHUP_BLOCKERS=0 ;; esac + CATCHUP_REASON="" + case "$CATCHUP_LINE" in *"$(printf '\t')"*) CATCHUP_REASON=${CATCHUP_LINE#*$'\t'} ;; esac + RETURN_CATCHUP=$(jq -n --argjson blockers "$CATCHUP_BLOCKERS" --arg reason "$CATCHUP_REASON" \ + '{pending:true,blockers:$blockers,reason:$reason}') +fi NOW=${FM_BEARINGS_NOW:-$(date -u +%Y-%m-%dT%H:%M:%SZ)} if [ "$ALL_LANDED" = 1 ] || [ "$ALL_SECONDMATES" = 1 ]; then @@ -338,6 +373,7 @@ MODEL=$(printf '%s' "$SNAP" | jq \ --argjson pr_repos_shown "$PR_REPOS_SHOWN" \ --argjson pr_rows_capped "$PR_ROWS_CAPPED" \ --argjson pr_rows_min_total "$PR_ROWS_MIN_TOTAL" \ + --argjson return_catchup "$RETURN_CATCHUP" \ --slurpfile candidate_prs_doc <(printf '%s' "$CANDIDATE_PRS") "$FM_LANDED_JQ_DEFS"' def doc($v): ($v[0] // error("fm-bearings-snapshot: empty jq payload")); def trunc($n): if . == null then null else @@ -389,7 +425,8 @@ MODEL=$(printf '%s' "$SNAP" | jq \ def as_gate($owner): {id, title:(.title | trunc(60)), blocked_by:((.unresolved_blocker_ids // []) | if length > 0 then join(",") else "-" end | trunc(120)), - reason:(hold_gate_reason | trunc(40)), owner:$owner}; + reason:(hold_gate_reason | trunc(40)), owner:$owner, + filed:((.since // null) | trunc(40))}; def round_robin_landed($n): . as $groups | [range(0; (($groups | map(length) | max) // 0)) as $i @@ -472,6 +509,8 @@ MODEL=$(printf '%s' "$SNAP" | jq \ | {id, kind, state: .current_state.state, repo:(.backlog.repo // .project // null), + name:((.backlog.title // "") as $name + | (if ($name | test("[^[:space:]]")) then $name else .id end) | trunc(70)), doing: ((.current_state.detail // "") as $d | (if $d != "" then $d else (.hints.last_event_text // "") end) | trunc(90)) } ] @@ -481,6 +520,9 @@ MODEL=$(printf '%s' "$SNAP" | jq \ kind:(.kind // "secondmate"), state:(.state // "working"), repo:(.repo // null), + name:((.name // "") as $name + | (if (($name | type) == "string" and ($name | test("[^[:space:]]"))) + then $name else ($m.id + "/" + .id) end) | trunc(70)), doing:((.doing // .state) | trunc(90))} ]) as $in_flight_all | ([ .backlog.records[] | . as $record @@ -511,12 +553,26 @@ MODEL=$(printf '%s' "$SNAP" | jq \ + [ (.secondmate_current.records // [])[] | .queued[]? | select(.hold_kind == "captain" and projected_deferred_hold) ] | length) as $decisions_marked_deferred + | (if ($return_catchup.pending // false) then + [{id:"(return-catchup)", + title:((if ($return_catchup.blockers // 0) > 0 then + "\($return_catchup.blockers) blocker(s) to clear before ordinary work" + elif (($return_catchup.reason // "") != "") then + ("catch-up retained: " + + ($return_catchup.reason | sub("[,;] *catch-up stays gated$"; ""))) + else "away-return catch-up is still open" end) | trunc(60)), + blocked_by:"-", + reason:"away-return catch-up", + owner:"(main)", + filed:null}] + else [] end) as $return_catchup_gate | ((if (.main_inventory.valid == false) then [{id:"(main-inventory)", title:((.main_inventory.reason // "main inventory invalid") | trunc(60)), blocked_by:"-", reason:"main inventory", - owner:"(main)"}] + owner:"(main)", + filed:null}] else [] end) + [ .backlog.records[] | . as $record @@ -537,7 +593,17 @@ MODEL=$(printf '%s' "$SNAP" | jq \ | select(($all_reports == 1) or (($rel_ids | index($r.id)) != null)) | {id, path} ]) as $reports_all | ([ .tasks[] | select(.kind != "secondmate" and .pr.url != null and .pr.source == "meta") | {id, url:.pr.url} ]) as $recorded_prs_all - | . as $snap + | def filed_epoch: + (.filed // null) as $filed + | if ($filed | type) != "string" then null + elif ($filed | test("T")) then try ($filed | fromdateiso8601) catch null + else try (($filed + "T00:00:00Z") | fromdateiso8601) catch null end; + def newest_filed_first: + to_entries + | sort_by((.value | filed_epoch) as $epoch + | if $epoch == null then [1, 0, .key] else [0, -$epoch, .key] end) + | map(.value); + . as $snap | { schema: "fm-bearings.v1", home: $home, @@ -551,7 +617,9 @@ MODEL=$(printf '%s' "$SNAP" | jq \ decisions_open: (if $all_decisions == 1 then $decisions_all else $decisions_all[:$decisions_n] end), landed: ($done | map({id, what:(.title | trunc(70)), artifact:(landed_artifact // "-"),owner:.home_id})), - gates: (if $all_queued == 1 then $gates_all else $gates_all[:$gates_n] end), + gates: ($return_catchup_gate + + ($gates_all | newest_filed_first + | if $all_queued == 1 then . else .[:$gates_n] end)), reports: (if $all_reports == 1 then $reports_all else $reports_all[:$reports_n] end), recorded_prs: (if $all_recorded_prs == 1 then $recorded_prs_all else $recorded_prs_all[:$recorded_prs_n] end) } diff --git a/bin/fm-bootstrap.sh b/bin/fm-bootstrap.sh index f318d3031e5..9be5b50b777 100755 --- a/bin/fm-bootstrap.sh +++ b/bin/fm-bootstrap.sh @@ -6,6 +6,7 @@ # exits 0. # Silent = all good. # Lines: "MISSING: <tool> (install: <command>)", +# "PRESENTATION_UNAVAILABLE: lavish-axi (requires >=<floor>; install: <command>) - nonvisual work may proceed with plain-text decisions and reports; install or upgrade before using Lavish", # "MISSING_MANUAL: <tool> (instructions: <url>)", "NEEDS_GH_AUTH", # "BACKEND_INVALID: <name> (known: <names>)", # "STARTUP_MEMORY_BUDGET: invalid config/startup-memory-budget - <reason>", @@ -18,6 +19,7 @@ # "ENDPOINT_BINDING_MIGRATION: task <id> (<backend>): <reason>", # "RUN_ATTRIBUTION: task <id>: legacy no-mistakes metadata has no proven branch=; any run is unattributable until task cleanup", # "BACKLOG_RECONCILE: <id>: <what this home could not reconcile>", +# "BACKLOG_RECONCILE: code-root <file> is not this home's <file>; ...", # "TANGLE: <remediation>", # "VAULT_DRIFT: <project>: <vault problem and remedy>", # "UPSTREAM: <fork drift or measurement failure>", @@ -73,14 +75,15 @@ # "treehouse get --lease" support. # no-mistakes is also MISSING when its installed version is older than # 1.46.0 (structured pipeline attestation floor; see CONTRIBUTING.md). -# The AXI-family floor policy is owned beside GH_AXI_MIN, -# LAVISH_AXI_MIN, and CHROME_DEVTOOLS_AXI_MIN below; the per-tool -# owners point there. An installed -# build below its floor reports MISSING like no-mistakes, so the operator -# is asked to upgrade rather than silently running an older tool. +# The AXI-family floor policy is owned beside GH_AXI_MIN, CHROME_DEVTOOLS_AXI_MIN, and +# LAVISH_AXI_MIN below; the per-tool owners point there. An installed +# essential build below its floor reports MISSING like no-mistakes. +# Missing or incompatible lavish-axi reports PRESENTATION_UNAVAILABLE: +# nonvisual dispatch continues with plain-text decisions and reports, +# but Lavish use still requires a compatible build at or above its floor. # tasks-axi feature probes remain a separate defense-in-depth check. -# tasks-axi and quota-axi are required bootstrap tools (same class as -# lavish-axi). A compatible tasks-axi default backend is silent. +# tasks-axi and quota-axi are essential bootstrap tools. +# A compatible tasks-axi default backend is silent. # quota-axi is required for the agent-owned dispatch-profile array # procedure in AGENTS.md section 4 and # .agents/skills/quota-array-dispatch/SKILL.md. @@ -150,6 +153,9 @@ # reads or writes another home; the fleet snapshot's classifier and # bin/fm-secondmate-reconcile.sh's nudge stay as backstops. Replayed # transitions and restored In-flight rows print BOOTSTRAP_INFO facts. +# The `code-root <file>` variant is a detect-only local check that runs +# even in a read-only session; detect_code_root_backlog_fork owns what +# it reports. # Set FM_BOOTSTRAP_DETECT_ONLY=1 to skip the eleven MUTATING sweeps # (endpoint-binding migration, run-attribution transition, # backlog_record_reconcile, secondmate_sync, @@ -202,6 +208,9 @@ # keeps detect-only meaning unlocked, exactly as before. # fm-bootstrap.sh install <tool>... # Install the named tools (only ones the captain approved). +# fm-bootstrap.sh lavish-compatible +# Exit 0 when lavish-axi meets LAVISH_AXI_MIN, 1 otherwise, printing +# nothing; bin/fm-brief.sh uses it to gate scout Lavish hosting. set -u SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -1207,7 +1216,7 @@ missing_tool_diagnostic() { # jq is universal because the read-only fleet snapshot, the crew-dispatch # profile reader, and the durable outcome manifest teardown publishes all parse # and emit JSON through it; without jq a home cannot archive a finished task. -COMMON_TOOLS="node git gh jq no-mistakes gh-axi chrome-devtools-axi lavish-axi tasks-axi quota-axi" +COMMON_TOOLS="node git gh jq no-mistakes gh-axi chrome-devtools-axi tasks-axi quota-axi" BACKEND=$(fm_backend_name) BACKEND_VALID=1 if ! BACKEND_TOOLS=$(fm_backend_required_tools "$BACKEND"); then @@ -1434,6 +1443,7 @@ crew_dispatch_validate() { elif $h == "claude" then (["low","medium","high","xhigh","max"] | index($e)) elif $h == "codex" then (["low","medium","high","xhigh"] | index($e)) elif $h == "grok" then (["low","medium","high"] | index($e)) + elif $h == "agy" then (["low","medium","high"] | index($e)) elif $h == "pi" or $h == "pi-signed" or $h == "omp" then (["low","medium","high","xhigh","max"] | index($e)) elif $h == "agy" then (["low","medium","high"] | index($e)) elif $h == "muse" then (["low","medium","high","xhigh","max"] | index($e)) @@ -1674,6 +1684,11 @@ startup_memory_budget_setup() { fi } +if [ "${1:-}" = "lavish-compatible" ]; then + tool_version_at_least lavish-axi "$LAVISH_AXI_MIN" + exit +fi + if [ "${1:-}" = "install" ]; then shift [ $# -gt 0 ] || { echo "usage: fm-bootstrap.sh install <tool>..." >&2; exit 1; } @@ -1781,8 +1796,8 @@ detect_local_tools() { if command -v chrome-devtools-axi >/dev/null 2>&1 && ! tool_version_at_least chrome-devtools-axi "$CHROME_DEVTOOLS_AXI_MIN"; then echo "MISSING: chrome-devtools-axi (install: $(install_cmd chrome-devtools-axi))" fi - if command -v lavish-axi >/dev/null 2>&1 && ! tool_version_at_least lavish-axi "$LAVISH_AXI_MIN"; then - echo "MISSING: lavish-axi (install: $(install_cmd lavish-axi))" + if ! tool_version_at_least lavish-axi "$LAVISH_AXI_MIN"; then + echo "PRESENTATION_UNAVAILABLE: lavish-axi (requires >=$LAVISH_AXI_MIN; install: $(install_cmd lavish-axi)) - nonvisual work may proceed with plain-text decisions and reports; install or upgrade before using Lavish" fi if command -v quota-axi >/dev/null 2>&1 && ! fm_quota_axi_compatible; then echo "MISSING: quota-axi (install: $(install_cmd quota-axi))" @@ -1823,9 +1838,27 @@ detect_local_config() { && ! fm_backlog_backend_manual "$CONFIG" && fm_tasks_axi_compatible; then echo "BOOTSTRAP_INFO: tasks-axi available" fi + detect_code_root_backlog_fork detect_home_summary_publication } +# Shadow-backlog check. When this home's data directory is not the code root's, +# a code-root data/backlog.md or data/done-archive.md that is not this home's +# own file is a queue a cwd-relative tasks-axi write has already forked; a link +# into the home does not survive such a write (docs/configuration.md "Backlog +# backend" owns why). Detect-only: neither copy is a safe winner, so nothing is +# merged here. +detect_code_root_backlog_fork() { + local name root_copy + [ "$FM_ROOT/data" -ef "$DATA" ] && return 0 + for name in backlog.md done-archive.md; do + root_copy="$FM_ROOT/data/$name" + [ -e "$root_copy" ] || [ -L "$root_copy" ] || continue + [ "$root_copy" -ef "$DATA/$name" ] && continue + echo "BACKLOG_RECONCILE: code-root $root_copy is not this home's $DATA/$name; tasks-axi wrote the code root instead of this home, so rows in it may be missing here - merge it into this home's copy and move it aside" + done +} + # This home's ledger publication is deliberately best-effort: every lifecycle # trigger calls it with --best-effort so a failure can never change the result # of a session start, a spawn, a teardown, or a watcher poll. That is correct, diff --git a/bin/fm-branch-prompt.sh b/bin/fm-branch-prompt.sh index c426f7ec8d2..0ed62dd0552 100755 --- a/bin/fm-branch-prompt.sh +++ b/bin/fm-branch-prompt.sh @@ -45,9 +45,9 @@ Handle it start to finish in one turn sequence: 1. Drain first: run `bin/fm-wake-drain.sh` and read every presented record, plus any OPEN DECISIONS, UNREAD STATUS, and RECORD DIVERGENCE sections. 2. For each task you are about to mutate, claim its lease first: `bin/fm-lease.sh claim <task>`. - Claim the reserved `backlog` lease around backlog writes (`bin/fm-lease.sh claim backlog`, then `tasks-axi ...`, then release). + Claim the reserved `backlog` lease around backlog writes (`bin/fm-lease.sh claim backlog`, then `bin/fm-tasks-axi.sh ...`, then release). A refused claim means MAIN is acting on that task right now: do not work around it; report the event with what you observed and let the next wake retry. -3. Handle with real tools: `bin/fm-crew-state.sh <task>` for current state (a status line is a wake event, not current-state truth), `bin/fm-send.sh` for a short steer, `bin/fm-control.sh <task> interrupt|exit|relaunch` for lifecycle, `bin/fm-pr-check.sh <task> <url>` when the task's ready status or `pr=` metadata names the PR's URL, `tasks-axi` for backlog moves. +3. Handle with real tools: `bin/fm-crew-state.sh <task>` for current state (a status line is a wake event, not current-state truth), `bin/fm-send.sh` for a short steer, `bin/fm-control.sh <task> interrupt|exit|relaunch` for lifecycle, `bin/fm-pr-check.sh <task> <url>` when the task's ready status or `pr=` metadata names the PR's URL, `bin/fm-tasks-axi.sh` for backlog moves. 4. Report: call the fm_branch_report tool exactly once per handled event, with the task id, the verdict, and a one-or-two-sentence summary; set silent true only for a fleet-wide heartbeat review that found literally nothing worth reporting. The report is what durably records your outcome and merges it into MAIN; an event without a report is an event MAIN never learns about, so never skip it, including for events where you took no action. 5. Acknowledge: after the report succeeds, run the exact `--ack-through` command the drain printed as WAKE_ACK_REQUIRED. diff --git a/bin/fm-brief.sh b/bin/fm-brief.sh index eae9dc91806..bdcf6f75801 100755 --- a/bin/fm-brief.sh +++ b/bin/fm-brief.sh @@ -58,6 +58,8 @@ # modified here. Plugin lifecycle is captain-owned outside this repository. # --scout writes the scout contract instead: the deliverable is a report at # data/<task-id>/report.md (no branch, no push, no PR) and the worktree is scratch. +# It offers the Lavish review loop only when `fm-bootstrap.sh lavish-compatible` +# confirms the supported lavish-axi floor; otherwise it asks for a text report. # --secondmate writes a persistent secondmate charter. The project list # is cloned into the secondmate home, while the natural-language scope # tells the main firstmate when to route work there; routine churn stays in its own home; @@ -963,6 +965,11 @@ EOF TASK_SECTION=${TASK_SECTION%$'\n'} if [ "$KIND" = scout ]; then +if "$SCRIPT_DIR/fm-bootstrap.sh" lavish-compatible >/dev/null 2>&1; then + LAVISH_LINE='If your deliverable is a visual artifact the captain will review and iterate on, you may host the Lavish review loop yourself (poll, revise, re-serve, staying alive) instead of handing it back to firstmate.' +else + LAVISH_LINE='Lavish is unavailable (lavish-axi is missing or below its supported version floor), so deliver your findings as a text report without Lavish, even for a visual deliverable.' +fi cat > "$BRIEF" <<EOF You are a crewmate: an autonomous worker agent managed by firstmate. Work on your own; do not wait for a human. @@ -1023,7 +1030,7 @@ $INBOX_SECTION ${BRAIN_SECTION}# Definition of done Write your findings to \`$DATA/$ID/report.md\`. The report must stand alone: what you did, what you found, the evidence (commands run, output, file:line references), and what you recommend. -If your deliverable is a visual artifact the captain will review and iterate on, you may host the Lavish review loop yourself (poll, revise, re-serve, staying alive) instead of handing it back to firstmate. +$LAVISH_LINE Before reporting done, read and follow \`$FM_ROOT/.agents/skills/captain-hold-lifecycle/SKILL.md\` and pass its shared completion gate for the report and any visual review. When the report is complete, append \`done: {one-line conclusion}\` to the status file and stop. If your findings reveal work that should ship (e.g. you reproduced a bug and the fix is clear), say so in the report; firstmate may promote this task in place, and you would then receive mode-specific ship instructions as a follow-up message. diff --git a/bin/fm-busy-lib.sh b/bin/fm-busy-lib.sh index b03165ddf0d..78c0a175c9a 100755 --- a/bin/fm-busy-lib.sh +++ b/bin/fm-busy-lib.sh @@ -42,7 +42,7 @@ # fm-interrupt the legacy Claude fm-send --key Escape idle event # fm-recovery a documented recovery reset after relaunch # Classifier-only sources (never written into a record): -# endpoint-gone, herdr-native, grok-regex, rovo-regex, muse-session-log, +# endpoint-gone, herdr-native, grok-regex, rovo-regex, agy-regex, muse-session-log, # cursor-transcript, missing, malformed, gen-mismatch, source-mismatch, # kimi-unverified, codex-unverified, capture-failed, no-target # @@ -53,17 +53,18 @@ # 3. a valid, gen-matching, source-trusted record -> its state and source # 4. no record at all: herdr's native busy verdict is trusted as busy # (generation state is sufficient for busy, not for idle), then the -# muse session-log and cursor transcript pull sources, then the Grok/Rovo -# temporary regex fallbacks classify a grok or rovo task from its -# rendered tail, then unknown missing +# muse session-log and cursor transcript pull sources, then the +# Grok/Rovo/AGY temporary regex fallbacks classify a grok, rovo, or agy +# task from its rendered tail, then unknown missing # 5. an untrusted agy record on Herdr may classify busy only when one native # sample reports both the matching identity and working status # 6. every other malformed, stale, or untrusted record -> unknown, never a # fallback -# Grok and Rovo are the ONLY rendered-text classifications that survive the -# redesign, because neither's structured lifecycle was credited-live-verified +# Grok, Rovo, and AGY are the ONLY rendered-text classifications that survive the +# redesign, because none of their structured lifecycles was credited-live-verified # in the approved audit (Rovo's clean ACP stopReason lives outside the TUI -# path firstmate drives, see references/harness/rovo.md); each is scoped to +# path firstmate drives, see references/harness/rovo.md; agy 1.2.0 exposes no +# hook surface at all, see references/harness/agy.md); each is scoped to # its own harness= and can never classify another adapter. The delivery # guards in bin/fm-composer-lib.sh match rendered footers for submit # acknowledgement and away-mode supervisor injection only; neither is a @@ -854,12 +855,27 @@ fm_busy_rovo_tail_busy() { | grep -qiE "${FM_BUSY_ROVO_REGEX:-Rovo is thinking}" } +# fm_busy_agy_tail_busy: the AGY-only temporary rendered-tail fallback. +# Consumes the tail on stdin; 0 when AGY's verified busy signature matches: +# the `esc to cancel` token in the status row the TUI pins to the bottom of +# the pane while a turn runs (verified live on agy 1.2.0; the idle status row +# shows `? for shortcuts` instead). The `Generating...` spinner word that +# renders beside it is deliberately NOT matched: it is a free-floating output +# line, so ordinary worker output echoing the word would classify an idle +# worker as busy. agy exposes no hook surface, so this fallback is the only +# pane-side source; it is never armed as a semantic writer +# (fm_busy_sources_for_harness trusts nothing for agy). +fm_busy_agy_tail_busy() { + grep -v '^[[:space:]]*$' | tail -12 \ + | grep -qiE 'esc[[:space:]]+to[[:space:]]+cancel' +} + # fm_busy_classify: semantic classification for a task whose endpoint the # caller has already established as present. Prints "<verdict> <source>": # busy|idle|unknown plus the producing source (see header). Never probes # process state. <tail40> is optional pre-captured plain output used only by -# the Grok arm; when absent the Grok arm captures through fm_backend_capture -# if available, else reports unknown capture-failed. +# the grok, rovo, and agy arms; when absent each captures through +# fm_backend_capture if available, else reports unknown capture-failed. fm_busy_classify() { # <backend> <target> <harness> <id> <state-dir> [tail40] local backend=$1 target=$2 harness=$3 id=$4 state=$5 tail40=${6-} local out rc r_state r_source native='' native_identity='' native_sample='' native_identity_required=0 log @@ -942,6 +958,10 @@ fm_busy_classify() { # <backend> <target> <harness> <id> <state-dir> [tail40] printf 'unknown source-mismatch' return 0 fi + if [ "$backend:$harness" = herdr:agy ]; then + printf 'unknown herdr-native' + return 0 + fi case "$harness" in muse*) # Semantic, on demand: fold this task's bound session log. An open run is @@ -1000,6 +1020,27 @@ fm_busy_classify() { # <backend> <target> <harness> <id> <state-dir> [tail40] fi return 0 ;; + agy) + if [ -z "$tail40" ]; then + if command -v fm_backend_capture >/dev/null 2>&1; then + tail40=$(fm_backend_capture "$backend" "$target" 40 2>/dev/null) || { + printf 'unknown capture-failed' + return 0 + } + else + printf 'unknown capture-failed' + return 0 + fi + fi + # Best-effort like rovo: a long turn can scroll the busy marker out of + # the captured tail, so its absence means "can't tell," never idle. + if printf '%s' "$tail40" | fm_busy_agy_tail_busy; then + printf 'busy agy-regex' + else + printf 'unknown agy-regex' + fi + return 0 + ;; esac printf 'unknown missing' } diff --git a/bin/fm-captain-hold.sh b/bin/fm-captain-hold.sh index a9ef07b132a..c3d3a98f67e 100755 --- a/bin/fm-captain-hold.sh +++ b/bin/fm-captain-hold.sh @@ -350,10 +350,38 @@ require_tasks_axi() { || fail "tasks-axi does not expose the captain-hold contract" } -task_show() { # <id> - local data +# Read one row into TASK_SHOW_OUTPUT; a non-zero return means the row is +# absent. A read that could not finish inside its bound is NOT absence, and +# every caller below would otherwise spend it as one - minting a duplicate task, +# skipping a keyed answer, or reporting a task that exists as missing. So the +# bound's own status stops the command instead, loudly and by name, and it +# leaves 124 intact rather than collapsing to fail's 1 so a caller running this +# inside a command substitution can still tell a wedged backend from a +# genuinely unknown id. +TASK_SHOW_OUTPUT= +task_show() { # <id>; sets TASK_SHOW_OUTPUT + local data status=0 reason data=$(fm_backlog_data_absolute "$DATA") || fail "data directory cannot be resolved: $DATA" - fm_backlog_row_show "$data" "$1" --full 2>/dev/null + TASK_SHOW_OUTPUT=$(fm_backlog_row_show "$data" "$1" --full 2>/dev/null) || status=$? + if [ "$status" -eq 124 ]; then + reason=${TASK_SHOW_OUTPUT%%$'\n'*} + printf 'fm-captain-hold: %s\n' \ + "${reason:-tasks-axi show $1 exceeded its backlog read bound}" >&2 + exit 124 + fi + return "$status" +} + +# Read one row into `show`, failing with <absence-message> only when the read +# genuinely failed; a read-bound hit (124) stops the command by name instead. +# task_show must be called in THIS shell, not inside a command substitution: +# it carries the row in TASK_SHOW_OUTPUT, which a subshell cannot hand back. +task_show_or_fail() { # <id> <absence-message>; sets show + task_show "$1" || { + [ "$?" -ne 124 ] || fail "the backlog backend exceeded its read bound reading $1" + fail "$2" + } + show=$TASK_SHOW_OUTPUT } show_field() { # <show-output> <field> @@ -388,7 +416,7 @@ show_field_value() { # <show-output> <field> origin_exists_here() { # <origin-id> [ -f "$STATE/$1.meta" ] && return 0 [ -f "$DATA/$1/report.md" ] && return 0 - task_show "$1" >/dev/null 2>&1 + task_show "$1" } list_has_key() { # <comma-list> <key> @@ -493,7 +521,8 @@ resolution_block() { # <mode> # surviving even when a date gate has expired) or a recorded captain answer. verify_hold_durable() { # <task-id> local id=$1 show state hold_kind body - show=$(task_show "$id") || fail "captain-held task $id is absent from this home's configured backlog (data directory $DATA)" + task_show "$id" || fail "captain-held task $id is absent from this home's configured backlog (data directory $DATA)" + show=$TASK_SHOW_OUTPUT state=$(show_field "$show" state) hold_kind=$(show_field_value "$show" hold_kind) body=$(show_field "$show" body) @@ -678,7 +707,13 @@ resolve_migrated_entry() { # <origin-or-empty> <entry> *-) prefixed="$prefix$candidate" ;; *) prefixed="$prefix-$candidate" ;; esac - show=$(task_show "$prefixed" 2>/dev/null) || continue + # Same shell rule as task_show_or_fail: the row is read out of + # TASK_SHOW_OUTPUT, so the read cannot sit inside a command substitution. + task_show "$prefixed" 2>/dev/null || { + [ "$?" -ne 124 ] || return 124 + continue + } + show=$TASK_SHOW_OUTPUT [ "$(show_field_value "$show" hold_kind)" = captain ] || continue prefixed_matches="${prefixed_matches}${prefixed_matches:+$NL_SEP}$prefixed" done @@ -699,13 +734,13 @@ resolve_migrated_entry() { # <origin-or-empty> <entry> # migrated-prefix, so a caller can record which evidence carried the attestation. resolve_entry() { # <origin-or-empty> <entry>; prints "<id> <how>" or fails local origin=$1 entry=$2 legacy migrated rc - if task_show "$entry" >/dev/null 2>&1; then + if task_show "$entry"; then printf '%s exact' "$entry" return 0 fi if [ -n "$origin" ] && [ "$origin" != "$BINDING_ANY" ]; then legacy=$(legacy_hold_id "$origin" "$entry") - if task_show "$legacy" >/dev/null 2>&1; then + if task_show "$legacy"; then printf '%s legacy' "$legacy" return 0 fi @@ -715,6 +750,7 @@ resolve_entry() { # <origin-or-empty> <entry>; prints "<id> <how>" or fails case "$rc" in 0) printf '%s' "$migrated"; return 0 ;; 2) return 2 ;; + 124) return 124 ;; esac if [ -n "$origin" ] && [ "$origin" != "$BINDING_ANY" ]; then legacy=$(legacy_hold_id "$origin" "$entry") @@ -763,6 +799,24 @@ write_hold_set_stamp() { # <task-id> <shown-body> <timestamp> <preserve-existin rm -f -- "$tmp" } +# Resolve one entry and verify the row it names is durably captain-held. A +# resolution failure that is not the read bound keeps resolve_entry's own +# status - its stderr already named the entry; 124 means the backend never +# answered, which is not the same as an unknown entry and must not be spent +# as absence. On success prints "<id> <how>" so the caller can keep the +# attestation evidence. +verify_entry_durable() { # <origin-or-empty> <entry>; prints "<id> <how>" + local origin=$1 entry=$2 resolved resolve_status=0 + resolved=$(resolve_entry "$origin" "$entry") || resolve_status=$? + if [ "$resolve_status" -ne 0 ]; then + [ "$resolve_status" -ne 124 ] \ + || fail "the backlog backend exceeded its read bound resolving $entry" + exit "$resolve_status" + fi + printf '%s\n' "$resolved" + verify_hold_durable "${resolved%% *}" +} + command_hold() { local id=${1:-} title='' reason='' repo='' origin='' until='' show state existing_title body='' hold_kind hold_set occurrence local existing_hold_kind='' existing_held='' preserve_hold_set=0 @@ -798,7 +852,8 @@ command_hold() { esac acquire_task_control_lock "$id" require_tasks_axi - if show=$(task_show "$id"); then + if task_show "$id"; then + show=$TASK_SHOW_OUTPUT state=$(show_field "$show" state) [ "$state" != "done" ] \ || fail "task $id is already closed; a new captain call needs its own task" @@ -833,9 +888,9 @@ command_hold() { # Publish the timestamp before the captain-hold annotation. A concurrent # snapshot may see the harmless stamp by itself, but can never see a newly # held task without the timestamp that defines this hold lifecycle's age. - show=$(task_show "$id") || fail "task $id disappeared before recording its hold-set stamp" + task_show_or_fail "$id" "task $id disappeared before recording its hold-set stamp" write_hold_set_stamp "$id" "$(show_field "$show" body)" "$hold_set" "$preserve_hold_set" - show=$(task_show "$id") || fail "task $id disappeared while recording its hold-set stamp" + task_show_or_fail "$id" "task $id disappeared while recording its hold-set stamp" [ -n "$(body_hold_set_timestamp "$(show_field_value "$show" body)")" ] \ || fail "task $id did not retain its hold-set stamp" if [ -n "$until" ]; then @@ -845,7 +900,8 @@ command_hold() { tasks_axi hold "$id" --reason "$reason" --kind captain >/dev/null \ || fail "could not hold task $id for the captain" fi - show=$(task_show "$id") || fail "task $id disappeared while holding it" + task_show "$id" || fail "task $id disappeared while holding it" + show=$TASK_SHOW_OUTPUT hold_kind=$(show_field_value "$show" hold_kind) [ "$hold_kind" = captain ] || fail "task $id did not retain its captain hold" occurrence=$(( $(resolution_record_count "$(show_field "$show" body)") + 1 )) @@ -922,7 +978,7 @@ close_answered() { # <task-id> <release-0-or-1> remove_interrupted_answer_stamp() { # <task-id> local id=$1 show body existing tmp - show=$(task_show "$id") || fail "task $id disappeared after closing" + task_show_or_fail "$id" "task $id disappeared after closing" body=$(decode_shown_value "$(show_field "$show" body)") \ || fail "could not decode the closed body for $id" existing=$(body_hold_set_timestamp "$body") @@ -958,7 +1014,8 @@ command_answer() { load_decision "$decision_file" acquire_task_control_lock "$id" require_tasks_axi - show=$(task_show "$id") || fail "captain-held task $id is absent from this home's configured backlog (data directory $DATA)" + task_show "$id" || fail "captain-held task $id is absent from this home's configured backlog (data directory $DATA)" + show=$TASK_SHOW_OUTPUT state=$(show_field "$show" state) hold_kind=$(show_field_value "$show" hold_kind) body=$(show_field "$show" body) @@ -994,7 +1051,8 @@ command_answer() { || fail "task $id was never held for the captain; nothing to record an answer on" write_resolution_record "$id" repaired "$body" remove_interrupted_answer_stamp "$id" - show=$(task_show "$id") || fail "task $id disappeared while recording the answer" + task_show "$id" || fail "task $id disappeared while recording the answer" + show=$TASK_SHOW_OUTPUT [ "$(show_field "$show" state)" = "done" ] || fail "recording the answer reopened closed task $id" body_has_resolution_record "$(show_field "$show" body)" \ || fail "captain-held task $id did not retain its durable resolution record" @@ -1031,7 +1089,8 @@ command_answer() { fail "could not close answered captain-held task $id" fi remove_interrupted_answer_stamp "$id" - show=$(task_show "$id") || fail "task $id disappeared after closing" + task_show "$id" || fail "task $id disappeared after closing" + show=$TASK_SHOW_OUTPUT body_has_resolution_record "$(show_field "$show" body)" \ || fail "captain-held task $id did not retain its durable resolution record" publish_parent_resolution_then_retire "$id" "$occurrence" "$outcome" @@ -1152,6 +1211,7 @@ sanitize_reconcile_provenance() { command_answers() { local origin='' source='' row rest key answer label mode id show state hold_kind body digest legacy_digest legacy_key local recorded_digest recorded_mode occurrence tmp err closed=0 skipped=0 reason release_flag tab=$'\t' + local resolve_rc while [ "$#" -gt 0 ]; do case "$1" in --source) shift; source=${1:-} ;; @@ -1212,6 +1272,12 @@ command_answers() { continue fi if [ "$resolve_rc" -ne 0 ]; then + # resolve_entry runs in a command substitution, so task_show's exit + # cannot stop this loop; only its status crosses back. 124 means the + # backend never answered, which is not the same as an unknown key and + # must not be spent as a skip. + [ "$resolve_rc" -ne 124 ] \ + || fail "the backlog backend exceeded its read bound resolving $key" printf 'skipped: %s (no captain-held task with that id)\n' "$key" skipped=$((skipped + 1)) continue @@ -1231,7 +1297,8 @@ command_answers() { if [ -n "$legacy_key" ]; then legacy_digest=$(sha256_text "$(legacy_keyed_decision_text "$source" "$legacy_key" "$answer" "$label")") fi - show=$(task_show "$id") || { printf 'skipped: %s (absent)\n' "$id"; skipped=$((skipped + 1)); continue; } + task_show "$id" || { printf 'skipped: %s (absent)\n' "$id"; skipped=$((skipped + 1)); continue; } + show=$TASK_SHOW_OUTPUT state=$(show_field "$show" state) hold_kind=$(show_field_value "$show" hold_kind) body=$(show_field "$show" body) @@ -1347,7 +1414,7 @@ publish_parent_resolution_then_retire() { # <task-id> <occurrence> <note> } command_reconcile_requests() { - local source_id='' source='' origin row id note provenance show created=0 skipped=0 tab=$'\t' + local source_id='' source='' origin row id note provenance show show_status=0 created=0 skipped=0 tab=$'\t' while [ "$#" -gt 0 ]; do case "$1" in --source-id) shift; source_id=${1:-} ;; @@ -1372,7 +1439,13 @@ command_reconcile_requests() { [ "${#id}" -le 128 ] \ || { printf 'refused: %s (task id is too long)\n' "$id"; skipped=$((skipped + 1)); continue; } acquire_task_control_lock "$id" - show=$(task_show "$id") || true + show_status=0 + show='' + task_show "$id" || show_status=$? + [ "$show_status" -ne 0 ] || show=$TASK_SHOW_OUTPUT + if [ "$show_status" -eq 124 ]; then + fail "the backlog backend exceeded its read bound reading $id" + fi if [ -z "$show" ]; then printf 'refused: %s (absent)\n' "$id" skipped=$((skipped + 1)) @@ -1447,7 +1520,7 @@ reconcile_close() { reconcile_request_read "$id" \ || fail "task $id has no pending board-created reconcile request" require_tasks_axi - show=$(task_show "$id") || fail "captain-held task $id is absent from this home's configured backlog (data directory $DATA)" + task_show_or_fail "$id" "captain-held task $id is absent from this home's configured backlog (data directory $DATA)" state=$(show_field "$show" state) hold_kind=$(show_field_value "$show" hold_kind) body=$(show_field "$show" body) @@ -1483,7 +1556,7 @@ reconcile_close() { fi close_answered "$id" 0 || fail "could not close reconciled captain-held task $id" remove_interrupted_answer_stamp "$id" - show=$(task_show "$id") || fail "task $id disappeared after closing" + task_show_or_fail "$id" "task $id disappeared after closing" body_has_resolution_record "$(show_field "$show" body)" \ || fail "captain-held task $id did not retain its durable resolution record" publish_parent_hold "$id" "$occurrence" resolved reconciled @@ -1519,7 +1592,7 @@ reconcile_note() { require_tasks_axi command_open "$id" \ || fail "task $id is not an open captain call; a note cannot keep a closed call open" - show=$(task_show "$id") || fail "captain-held task $id is absent from this home's configured backlog (data directory $DATA)" + task_show_or_fail "$id" "captain-held task $id is absent from this home's configured backlog (data directory $DATA)" body=$(decode_shown_value "$(show_field "$show" body)") \ || fail "could not decode the existing body for $id" note_digest=$(sha256_text "$note") @@ -1584,13 +1657,9 @@ command_complete() { if [ -n "$keys" ]; then while IFS= read -r entry; do [ -n "$entry" ] || continue - if ! resolved=$(resolve_entry "$origin" "$entry"); then - # resolve_entry has already refused on stderr naming the entry. - exit 1 - fi + resolved=$(verify_entry_durable "$origin" "$entry") || exit $? resolved_how=${resolved##* } resolved=${resolved%% *} - verify_hold_durable "$resolved" if [ "$resolved_how" = migrated-prefix ]; then attested_by_prefix="${attested_by_prefix}${attested_by_prefix:+ }$entry=$resolved" fi @@ -1648,11 +1717,7 @@ command_verify() { if [ -n "$keys" ]; then while IFS= read -r entry; do [ -n "$entry" ] || continue - if ! resolved=$(resolve_entry "$origin" "$entry"); then - # resolve_entry has already refused on stderr naming the entry. - exit 1 - fi - verify_hold_durable "${resolved%% *}" + verify_entry_durable "$origin" "$entry" >/dev/null done <<EOF $(printf '%s\n' "$keys" | tr ',' '\n') EOF @@ -1770,7 +1835,8 @@ command_diverged() { while IFS= read -r key; do list_has_line "$tokens" "$key" || continue [ "$(status_key_closing_verb "$f" "$key")" = "$resolve" ] || continue - show=$(task_show "$id") || continue + task_show "$id" || continue + show=$TASK_SHOW_OUTPUT [ "$(show_field "$show" state)" != "done" ] || continue [ "$(show_field_value "$show" hold_kind)" = captain ] || continue # The title is the only free-text field here, and the report is @@ -1838,10 +1904,11 @@ command_open() { # <task-id> [--identity] [--distinguish-absent] state=${FM_BACKLOG_ROW_STATE%% *} if [ "$state" != "done" ] && [ "$FM_BACKLOG_ROW_HOLD_KIND" = captain ]; then if [ "$identity" -eq 1 ]; then - show=$(task_show "$id") || { + task_show "$id" || { printf 'fm-captain-hold: captain call %s is open but its record could not be read\n' "$id" >&2 exit 2 } + show=$TASK_SHOW_OUTPUT shown_body=$(show_field "$show" body) printf '%s#%s\n' \ "$(body_hold_set_timestamp "$(decode_shown_value "$shown_body")")" \ diff --git a/bin/fm-classify-lib.sh b/bin/fm-classify-lib.sh index 5030356c4f7..c301e1a114d 100755 --- a/bin/fm-classify-lib.sh +++ b/bin/fm-classify-lib.sh @@ -2170,3 +2170,25 @@ stale_is_terminal() { # <window> <state> last=$(last_status_line "$state/$(window_to_task "$win" "$state").status") [ -n "$last" ] && status_is_captain_relevant "$last" } + +# fm_classify_afk_mode <state> +# The single owner of reading state/.afk's declared mode. Always prints +# exactly one of "away" or "quiet" and always succeeds - every caller gets a +# definitive answer, never an error to handle. Presence/liveness stays owned +# by fm_afk_daemon_owns_supervision and the raw `-e "$state/.afk"` checks +# throughout the tree; this is the mode of an ALREADY-present flag. +# "away" (today's return-on-any-unmarked-message behavior) is the safe +# default: missing, empty, unreadable, or unrecognized content, and the +# legacy bare-epoch-timestamp content written before mode existed, all read +# as "away". Only an exact first-line "quiet" ever reads as "quiet" - +# kunchenguid/firstmate#2356's standing captain-present quiet mode, entered +# only through /quiet and exited only through an explicit /quiet off +# (AGENTS.md section 8's away-mode stub). +fm_classify_afk_mode() { + local state=$1 mode + mode=$(head -n 1 "$state/.afk" 2>/dev/null) || { printf '%s\n' away; return 0; } + case "$mode" in + quiet) printf '%s\n' quiet ;; + *) printf '%s\n' away ;; + esac +} diff --git a/bin/fm-claude-trust.sh b/bin/fm-claude-trust.sh index 732a6b7c5c4..07762cf9a0a 100755 --- a/bin/fm-claude-trust.sh +++ b/bin/fm-claude-trust.sh @@ -1,32 +1,100 @@ #!/usr/bin/env bash -# Pre-register Claude Code's workspace trust for the isolated task worktree a -# ship/scout spawn is about to launch a claude crewmate into, so the worker -# reaches its brief instead of wedging on the trust dialog. +# Pre-register Claude Code's workspace trust for the directory a claude spawn is +# about to launch into - the isolated task worktree of a ship or scout crewmate, +# or the seeded home of a secondmate - so the agent reaches its brief or charter +# instead of wedging on the trust dialog. In worktree mode it also carries +# forward the external-CLAUDE.md-import approval, but only when the primary +# checkout already holds standing consent for it - see the consent-gating +# block below for why that dialog is otherwise left for the worker to wedge +# on rather than answered on the human's behalf. # # Usage: fm-claude-trust.sh <worktree> <project> +# fm-claude-trust.sh --secondmate-home <home> <id> # <worktree> the isolated task worktree this spawn launches into # <project> the primary checkout that worktree belongs to +# <home> the seeded secondmate home this spawn launches into +# <id> the secondmate id that home must already be marked for # Prints one line naming what it registered; refuses loudly on anything else. # # WHY THIS EXISTS. Claude Code gates a folder it has never seen behind an # interactive workspace-trust dialog, and --dangerously-skip-permissions does # NOT cover it: `claude --help` records that the dialog is skipped only in -# non-interactive mode (-p, or a non-TTY stdout), and a crewmate pane is -# interactive. Every fresh task worktree therefore hits it. The dialog renders +# non-interactive mode (-p, or a non-TTY stdout), and a spawned pane is +# interactive. Every fresh task worktree therefore hits it, and so does every +# secondmate home the operator has not opened by hand. The dialog renders # with the cursor on "No, exit" and firstmate's steering plane carries only # Enter, Escape and C-c with no arrow navigation, so firstmate cannot answer it -# and must not try - pressing Enter would select exit. The worker wedges before +# and must not try - pressing Enter would select exit. The agent wedges before # it ever reads the brief. Registering the trust before launch is the only -# control that reaches an interactive pane. +# control that reaches an interactive pane. The same reasoning covers Claude +# Code's separate "Allow external CLAUDE.md file imports?" dialog, which +# `--setting-sources project,local` (firstmate PR 10's minimal worker tool +# surface) stopped suppressing: it renders whenever a loaded CLAUDE.md chain +# reaches outside the project tree - which every crewmate's does, through the +# captain's own `~/.claude/CLAUDE.md` importing `~/.claude/RTK.md` - and it is +# gated the same fail-closed way as trust: cursor on "No, disable", no arrow +# navigation from firstmate's steering plane. Only worktree mode reaches this +# second dialog's flags: a secondmate home has no separate "project" entry to +# carry consent forward from, so its registration stays trust-only. +# +# TWO PROJECT-CONFIG ENTRIES IN WORKTREE MODE, NOT ONE. Registering both flags +# on the worktree entry alone (the original trust-only design) leaves the +# external-imports dialog showing. Verified 2026-09-06 by disassembling the +# installed `claude` binary and reproducing in an isolated three-way tmux +# launch: Claude Code's own trust check (`Rde`) reads the canonical +# project-root entry first and, failing that, falls back to an ancestor walk +# from the worktree upward that DOES reach the worktree's own entry - which is +# why the trust dialog kept working after PR 10. The external-imports check +# (`es`/`F1e`) has no such fallback: it reads ONLY the canonical project-root +# entry, and that root is never the worktree - Claude Code's own git-root +# canonicalization (`Fr`/`Se`) walks a linked worktree's `.git` file through +# its `commondir` pointer back to the PRIMARY CHECKOUT, exactly the <project> +# argument this script already receives for the worktree-mode scope test +# below. So the trust flag is registered on BOTH the worktree entry (for +# trust's ancestor-walk fallback and defense in depth) and the project entry +# (the trust check's first, canonical-shaped, look); the two external-imports +# flags land on those same two entries only when the project entry already +# carries standing consent (see the consent-gating block below) - the project +# entry is the only place the external-imports check ever looks. Registering +# the project entry is a write to the launching user's OWN Claude config +# store, keyed by a project PATH the scope test below has already verified is +# real - not a write to the project's tracked content, so hard rule 1 does not +# apply, same as the existing worktree-entry write. +# +# THAT SAME PROJECT ENTRY IS ALSO THE LAUNCHING HUMAN'S OWN INTERACTIVE +# CONFIG, though, so this registration must never overwrite a decision the +# human already made there. If the project entry already carries +# hasClaudeMdExternalIncludesApproved===false - Claude Code only ever writes +# that on an explicit "No, disable" answer - the whole registration refuses +# rather than flipping it, because doing so would grant every future +# interactive session in that checkout silent external-file inclusion the +# human declined, permanently and without being asked. The worktree entry is +# left unwritten too: the spawn wedges on the dialog, which is the honest +# outcome given a standing decline, not registered trust with a stripped +# consent record. # # THE SCOPE TEST IS THE SAFETY PROPERTY, and it is STRUCTURAL rather than a -# path policy. <worktree> must be a LINKED git worktree - its own git dir, +# path policy. Each mode has its own, because the two directories have entirely +# different shapes on disk. +# +# WORKTREE MODE. <worktree> must be a LINKED git worktree - its own git dir, # sharing <project>'s common dir - whose top level is exactly the resolved # argument. Git is the ground truth, so the argument is never trusted on its # own word: a primary checkout (git dir == common dir), a worktree of an # unrelated repo, a subdirectory of a worktree, a plain directory, and a home # directory are each refused. Refusal is a non-zero exit, never a warning and -# never a silent skip. +# never a silent skip. When <project> is itself a linked worktree (a +# secondmate home spawned from, rather than as, the primary checkout), +# refusing outright would wedge a relaunch that is otherwise perfectly valid: +# its own common dir already IS the primary checkout's own git dir (git's +# git-common-dir answer never changes by which worktree asks), so the +# checkout is derived structurally from it - its parent directory in the +# standard non-bare, non-GIT_DIR-overridden layout this script already +# requires elsewhere - and verified, never assumed: the candidate's own +# resolved git dir must equal that common dir, the same primary-checkout +# definition used throughout, or this refuses rather than guess. The +# consent-gated external-imports flags land on that resolved canonical +# checkout, never on the linked-worktree argument itself. # # The test is deliberately NOT a treehouse or orca path prefix. Treehouse's # root is configurable (--root, TREEHOUSE_ROOT, config, and a relative @@ -45,13 +113,43 @@ # opt-in guard family (FM_*_LIVE_E2E=1) and record the result in # docs/verification/runtime-backends.md, rather than assuming the shape here. # -# Only the launching user's own store is written: the projects entry for the -# worktree path in ${CLAUDE_CONFIG_DIR:-$HOME}/.claude.json, which must be a -# regular file this uid owns. Every unrelated key and project entry is -# preserved, and the replacement is atomic. fm-spawn.sh forwards CLAUDE_CONFIG_DIR -# onto the claude launch verbatim rather than resolving it, and the worker's pane -# starts in the task worktree, so only an absolute value names the same store on -# both sides; a relative one is refused below rather than guessed at. +# SECONDMATE-HOME MODE. A secondmate home is a whole firstmate instance rather +# than a task worktree, and bin/fm-home-seed.sh produces it in two shapes: a +# leased treehouse worktree (linked) and a standalone clone of the firstmate +# repo (a primary checkout). The worktree test above therefore cannot decide +# this case at all - it refuses the standalone clone as a primary checkout, +# which is why a claude secondmate in an explicit ~/fm-homes/<id> home met the +# dialog with nothing registered. Git shape is not the evidence here; THE SEED +# IS. The home must carry a .fm-secondmate-home marker that is a regular file +# this user owns, never a symlink, naming exactly the <id> passed; it must hold +# the firstmate instance files AGENTS.md and bin/; and each of its data, state, +# config and projects paths must resolve inside the home. That is the set +# bin/fm-home-seed.sh writes and bin/fm-spawn.sh's validate_firstmate_home_for_spawn +# re-checks before launch, so this accepts exactly the homes a secondmate spawn +# will launch into and nothing wider: a plain directory, a project checkout, an +# ordinary firstmate checkout, a home marked for a different secondmate, and a +# home whose operational directory escapes it are each refused. An ABSENT +# operational directory is accepted for the same reason the spawn accepts one - +# a test stricter than the spawn's own would move the wedge from the dialog to +# a refusal without making any unseeded directory less trusted. +# +# Home-level trust is broader than worktree trust, since the pane starts in the +# home and the secondmate works across it, so it is granted on that seed +# evidence alone and never on a caller's word about what a path is. It is +# trust-only: a secondmate home has no separate primary-checkout "project" +# argument to gate external-imports consent against, so the two import flags +# are never written there. +# +# Only the launching user's own store is written. In worktree mode: the +# projects entries for the worktree path and the resolved canonical project +# path in ${CLAUDE_CONFIG_DIR:-$HOME}/.claude.json, which must be a regular +# file this uid owns; every unrelated key and project entry is preserved, and +# both entries land in one atomic replacement. In secondmate-home mode: the +# single projects entry for the registered home path, same store, same atomic +# replacement. fm-spawn.sh forwards CLAUDE_CONFIG_DIR onto the claude launch +# verbatim rather than resolving it, and the pane starts in the registered +# directory, so only an absolute value names the same store on both sides; a +# relative one is refused below rather than guessed at. set -u # Path resolution here must answer from the filesystem, never from the caller's # environment, because the refusals below are the safety property. CDPATH would @@ -69,9 +167,36 @@ unset CDPATH \ GIT_DISCOVERY_ACROSS_FILESYSTEM GIT_CONFIG GIT_CONFIG_GLOBAL \ GIT_CONFIG_SYSTEM GIT_CONFIG_NOSYSTEM GIT_CONFIG_COUNT -[ "$#" -eq 2 ] || { echo "usage: fm-claude-trust.sh <worktree> <project>" >&2; exit 2; } -WT_ARG=$1 -PROJ_ARG=$2 +usage() { + echo "usage: fm-claude-trust.sh <worktree> <project>" >&2 + echo " fm-claude-trust.sh --secondmate-home <home> <id>" >&2 + exit 2 +} + +# MODE selects which structural scope test decides the argument, and SCOPE_NOUN +# names what the argument was expected to be so every shared refusal below reads +# correctly in both modes. +case "${1:-}" in + --secondmate-home) + [ "$#" -eq 3 ] || usage + MODE=secondmate-home + TARGET_ARG=$2 + SUB_ID=$3 + PROJ_ARG= + SCOPE_NOUN="secondmate home" + ;; + '' | -h | --help) + usage + ;; + *) + [ "$#" -eq 2 ] || usage + MODE=worktree + TARGET_ARG=$1 + SUB_ID= + PROJ_ARG=$2 + SCOPE_NOUN="task worktree" + ;; +esac refuse() { echo "error: refusing to pre-register Claude trust: $1" >&2; exit 1; } @@ -90,10 +215,12 @@ common_dir_of() { (cd -P -- "$dir" && real_dir "$common") } -WT_REAL=$(real_dir "$WT_ARG") || true -[ -n "$WT_REAL" ] || refuse "worktree '$WT_ARG' is not an accessible directory" -PROJ_REAL=$(real_dir "$PROJ_ARG") || true -[ -n "$PROJ_REAL" ] || refuse "project '$PROJ_ARG' is not an accessible directory" +TARGET_REAL=$(real_dir "$TARGET_ARG") || true +[ -n "$TARGET_REAL" ] || refuse "$SCOPE_NOUN '$TARGET_ARG' is not an accessible directory" +if [ "$MODE" = worktree ]; then + PROJ_REAL=$(real_dir "$PROJ_ARG") || true + [ -n "$PROJ_REAL" ] || refuse "project '$PROJ_ARG' is not an accessible directory" +fi CONFIG_DIR=${CLAUDE_CONFIG_DIR:-${HOME:-}} [ -n "$CONFIG_DIR" ] || refuse "neither CLAUDE_CONFIG_DIR nor HOME is set, so the store cannot be located" @@ -116,30 +243,95 @@ if [ -z "$CONFIG_DIR_REAL" ]; then fi [ -n "$CONFIG_DIR_REAL" ] || refuse "Claude config directory '$CONFIG_DIR' does not exist and could not be created" -# A home or config directory is never a task worktree. Checked explicitly so -# the refusal names the real reason instead of the git verdict behind it. -[ "$WT_REAL" != "$CONFIG_DIR_REAL" ] || refuse "'$WT_REAL' is the Claude config directory, not a task worktree" +# The filesystem root, a home directory, and the config directory are never +# something this registers, in either mode. Checked explicitly so the refusal +# names the real reason instead of the scope verdict behind it. +[ "$TARGET_REAL" != / ] || refuse "'/' is the filesystem root, not a $SCOPE_NOUN" +[ "$TARGET_REAL" != "$CONFIG_DIR_REAL" ] || refuse "'$TARGET_REAL' is the Claude config directory, not a $SCOPE_NOUN" if [ -n "${HOME:-}" ]; then HOME_REAL=$(real_dir "$HOME") || true - [ "$WT_REAL" != "${HOME_REAL:-}" ] || refuse "'$WT_REAL' is the home directory, not a task worktree" + [ "$TARGET_REAL" != "${HOME_REAL:-}" ] || refuse "'$TARGET_REAL' is the home directory, not a $SCOPE_NOUN" fi -WT_TOP=$(git -C "$WT_REAL" rev-parse --show-toplevel 2>/dev/null) || true -[ -n "$WT_TOP" ] || refuse "'$WT_REAL' is not inside a git repository" -WT_TOP_REAL=$(real_dir "$WT_TOP") || true -[ "$WT_TOP_REAL" = "$WT_REAL" ] || refuse "'$WT_REAL' is not a worktree root (its root is '${WT_TOP_REAL:-unresolvable}')" +if [ "$MODE" = worktree ]; then + WT_TOP=$(git -C "$TARGET_REAL" rev-parse --show-toplevel 2>/dev/null) || true + [ -n "$WT_TOP" ] || refuse "'$TARGET_REAL' is not inside a git repository" + WT_TOP_REAL=$(real_dir "$WT_TOP") || true + [ "$WT_TOP_REAL" = "$TARGET_REAL" ] || refuse "'$TARGET_REAL' is not a worktree root (its root is '${WT_TOP_REAL:-unresolvable}')" -WT_GIT_DIR=$(git -C "$WT_REAL" rev-parse --absolute-git-dir 2>/dev/null) || true -[ -n "$WT_GIT_DIR" ] || refuse "'$WT_REAL' has no resolvable git directory" -WT_GIT_DIR=$(real_dir "$WT_GIT_DIR") || true -[ -n "$WT_GIT_DIR" ] || refuse "'$WT_REAL' has an unresolvable git directory" -WT_COMMON=$(common_dir_of "$WT_REAL") || true -[ -n "$WT_COMMON" ] || refuse "'$WT_REAL' has no resolvable git common directory" -[ "$WT_GIT_DIR" != "$WT_COMMON" ] || refuse "'$WT_REAL' is a primary checkout, not an isolated worktree" + WT_GIT_DIR=$(git -C "$TARGET_REAL" rev-parse --absolute-git-dir 2>/dev/null) || true + [ -n "$WT_GIT_DIR" ] || refuse "'$TARGET_REAL' has no resolvable git directory" + WT_GIT_DIR=$(real_dir "$WT_GIT_DIR") || true + [ -n "$WT_GIT_DIR" ] || refuse "'$TARGET_REAL' has an unresolvable git directory" + WT_COMMON=$(common_dir_of "$TARGET_REAL") || true + [ -n "$WT_COMMON" ] || refuse "'$TARGET_REAL' has no resolvable git common directory" + [ "$WT_GIT_DIR" != "$WT_COMMON" ] || refuse "'$TARGET_REAL' is a primary checkout, not an isolated worktree" -PROJ_COMMON=$(common_dir_of "$PROJ_REAL") || true -[ -n "$PROJ_COMMON" ] || refuse "project '$PROJ_REAL' is not inside a git repository" -[ "$WT_COMMON" = "$PROJ_COMMON" ] || refuse "'$WT_REAL' is not a worktree of project '$PROJ_REAL'" + PROJ_COMMON=$(common_dir_of "$PROJ_REAL") || true + [ -n "$PROJ_COMMON" ] || refuse "project '$PROJ_REAL' is not inside a git repository" + [ "$WT_COMMON" = "$PROJ_COMMON" ] || refuse "'$TARGET_REAL' is not a worktree of project '$PROJ_REAL'" + + # The external-imports flags must land on the primary checkout - its own git + # dir equals the common dir - because that is exactly the path Claude Code's + # own git-root canonicalization collapses every linked worktree to. When + # <project> is itself a linked worktree (a secondmate home spawned from, + # rather than as, the primary checkout), refusing outright would wedge a + # relaunch that is otherwise perfectly valid: PROJ_COMMON already IS that + # primary checkout's own git dir (git's git-common-dir answer never changes + # by which worktree asks), so the checkout is derived structurally from it - + # its parent directory in the standard non-bare, non-GIT_DIR-overridden + # layout this script already requires elsewhere - and verified, never + # assumed: the candidate's own resolved git dir must equal PROJ_COMMON, the + # same primary-checkout definition used above, or this refuses rather than + # guess. + PROJ_GIT_DIR=$(git -C "$PROJ_REAL" rev-parse --absolute-git-dir 2>/dev/null) || true + [ -n "$PROJ_GIT_DIR" ] || refuse "project '$PROJ_REAL' has no resolvable git directory" + PROJ_GIT_DIR=$(real_dir "$PROJ_GIT_DIR") || true + [ -n "$PROJ_GIT_DIR" ] || refuse "project '$PROJ_REAL' has an unresolvable git directory" + if [ "$PROJ_GIT_DIR" = "$PROJ_COMMON" ]; then + PROJ_CANON=$PROJ_REAL + else + PROJ_CANON=$(real_dir "$(dirname -- "$PROJ_COMMON")") || true + [ -n "$PROJ_CANON" ] \ + || refuse "project '$PROJ_REAL' is a linked worktree whose primary checkout could not be resolved" + CANON_GIT_DIR=$(git -C "$PROJ_CANON" rev-parse --absolute-git-dir 2>/dev/null) || true + CANON_GIT_DIR=$(real_dir "${CANON_GIT_DIR:-}") || true + [ -n "$CANON_GIT_DIR" ] && [ "$CANON_GIT_DIR" = "$PROJ_COMMON" ] \ + || refuse "project '$PROJ_REAL' is a linked worktree whose primary checkout could not be resolved" + fi +else + # The seed evidence, in the order that names the most useful reason first: the + # marker decides whether this is a secondmate home at all, the id decides + # whose, and the instance files and operational directories decide whether it + # is the shape bin/fm-home-seed.sh leaves behind. The marker is the token the + # whole boundary rests on, so it is judged as a file rather than as a value: a + # symlink is refused outright rather than followed, because a link is a way to + # make some other file's bytes stand in for the seed, and a marker this user + # does not own was planted by someone else. + [ -n "$SUB_ID" ] || refuse "no secondmate id was supplied, so '$TARGET_REAL' cannot be matched against its seed marker" + SUB_MARKER="$TARGET_REAL/.fm-secondmate-home" + [ ! -L "$SUB_MARKER" ] || refuse "'$SUB_MARKER' is a symlink; a seeded secondmate home carries the marker as a regular file" + [ -f "$SUB_MARKER" ] || refuse "'$TARGET_REAL' carries no .fm-secondmate-home marker, so it is not a seeded secondmate home" + [ -O "$SUB_MARKER" ] || refuse "'$SUB_MARKER' is not owned by this user" + SUB_MARKER_ID=$(cat "$SUB_MARKER" 2>/dev/null) || true + [ "$SUB_MARKER_ID" = "$SUB_ID" ] || refuse "'$TARGET_REAL' is marked for secondmate '${SUB_MARKER_ID:-unknown}', not '$SUB_ID'" + [ -f "$TARGET_REAL/AGENTS.md" ] || refuse "'$TARGET_REAL' has no AGENTS.md, so it is not a firstmate home" + [ -d "$TARGET_REAL/bin" ] || refuse "'$TARGET_REAL' has no bin/, so it is not a firstmate home" + for sub_dir_name in data state config projects; do + sub_dir="$TARGET_REAL/$sub_dir_name" + if [ -L "$sub_dir" ] && [ ! -e "$sub_dir" ]; then + refuse "'$sub_dir' is a broken symlink, so this home's $sub_dir_name directory cannot be shown to stay inside it" + fi + [ -e "$sub_dir" ] || continue + [ -d "$sub_dir" ] || refuse "'$sub_dir' is not a directory, so '$TARGET_REAL' is not a seeded secondmate home" + sub_dir_real=$(real_dir "$sub_dir") || true + [ -n "$sub_dir_real" ] || refuse "'$sub_dir' cannot be resolved" + case "$sub_dir_real" in + "$TARGET_REAL"/*) ;; + *) refuse "'$sub_dir' resolves to '$sub_dir_real', outside the home, so '$TARGET_REAL' is not a safe secondmate home" ;; + esac + done +fi # The store write needs node, and a missing interpreter refuses like every other # failure here. Degrading instead would launch a worker straight into the dialog @@ -190,11 +382,44 @@ fi # attempts, and it must fail loudly rather than report a trust it did not leave. # ponytail: fingerprint-and-refuse, not a lock; flock is absent on macOS and # cannot stop a vendor session's own rewrite anyway. -if ! node - "$STORE" "$WT_REAL" <<'NODE' +# +# In worktree mode every flag lands on both the worktree entry and the project +# entry in the same read-modify-write attempt, so a single rename either +# records all of it or none of it - there is no state where the worktree entry +# is fresh and the project entry stale, or the other way round. In +# secondmate-home mode only the single home entry is written. +# +# The two external-imports flags (worktree mode only) are gated separately +# from the trust flag, because they are a CONSENT grant, not a pre-approval +# this script is allowed to manufacture. Claude Code only ever writes +# hasClaudeMdExternalIncludesApproved itself, on an explicit interactive +# answer; this script's own job is to keep a worker from wedging on a dialog, +# never to answer that dialog on the human's behalf. So the import flags land +# on the project entry - the only place the imports check ever reads (see the +# disassembly note above) - only when that entry ALREADY carries +# hasClaudeMdExternalIncludesApproved===true, i.e. the human already said yes +# at some point and this write is a same-value refresh, not new consent from +# an absent flag. When it is not already true (including plain absent, the +# common case for a project claude has never asked about), the import flags +# are left untouched on both entries: writing them to the worktree entry alone +# would be a pure no-op (the imports check never reads it) that only obscures +# the real state, so trust still registers normally but the import dialog is +# left exactly as undecided as it already was - the worker wedges on it, the +# same honest outcome as an explicit decline, rather than a spawn spending +# consent the human was never asked for. +TRUST_FLAG='hasTrustDialogAccepted' +IMPORT_FLAGS='["hasClaudeMdExternalIncludesApproved","hasClaudeMdExternalIncludesWarningShown"]' +if [ "$MODE" = worktree ]; then + WRITE_ARGS=("$STORE" "$MODE" "$TARGET_REAL" "$PROJ_CANON" "$TRUST_FLAG" "$IMPORT_FLAGS") +else + WRITE_ARGS=("$STORE" "$MODE" "$TARGET_REAL" "" "$TRUST_FLAG" "$IMPORT_FLAGS") +fi +if ! node - "${WRITE_ARGS[@]}" <<'NODE' const fs = require("node:fs"); const path = require("node:path"); const crypto = require("node:crypto"); -const [store, worktree] = process.argv.slice(2); +const [store, mode, target, project, trustFlag, importFlagsJson] = process.argv.slice(2); +const importFlags = JSON.parse(importFlagsJson); const readStore = () => { try { return fs.readFileSync(store); @@ -205,6 +430,32 @@ const readStore = () => { }; const fingerprint = (buf) => buf === null ? "absent" : crypto.createHash("sha256").update(buf).digest("hex"); +const setFlags = (projects, key, flags) => { + let entry = projects[key]; + if (entry === undefined || entry === null || typeof entry !== "object" || Array.isArray(entry)) { + entry = {}; + } + for (const flag of flags) entry[flag] = true; + projects[key] = entry; +}; +const flagsLanded = (projects, key, flags) => + flags.every((flag) => projects?.[key]?.[flag] === true); +// The project entry is the launching user's OWN interactive config, not a +// throwaway worktree, so a spawn must never silently reverse a decision the +// human already recorded there. hasClaudeMdExternalIncludesApproved===false +// is exactly that decision (Claude Code only ever writes it on an explicit +// "No, disable" answer); flipping it to true would grant every future +// interactive session in that checkout silent external-file inclusion the +// human declined. Refuse the whole registration instead of overriding it - +// the worktree entry is not written either, so the spawn wedges on the +// dialog rather than the human's consent being spent without being asked. +const declinedExternalImports = (projects, key) => + projects?.[key]?.hasClaudeMdExternalIncludesApproved === false; +// True only on an explicit prior "Yes, allow" answer - the sole state this +// script may treat as standing consent to refresh. Absent, or any other +// value, is NOT consent (see the block comment above this script's node call). +const approvedExternalImports = (projects, key) => + projects?.[key]?.hasClaudeMdExternalIncludesApproved === true; const attempt = () => { const original = readStore(); const before = fingerprint(original); @@ -223,12 +474,23 @@ const attempt = () => { if (projects === null || typeof projects !== "object" || Array.isArray(projects)) { throw new Error(`${store} has a non-object "projects" value`); } - let entry = projects[worktree]; - if (entry === undefined || entry === null || typeof entry !== "object" || Array.isArray(entry)) { - entry = {}; + let keys; + if (mode === "worktree") { + if (declinedExternalImports(projects, project)) { + throw new Error( + `project entry for ${project} in ${store} already declined external CLAUDE.md imports; refusing to override that consent`, + ); + } + const carryImportConsent = approvedExternalImports(projects, project); + const targetFlags = carryImportConsent ? [trustFlag, ...importFlags] : [trustFlag]; + const projectFlags = carryImportConsent ? [trustFlag, ...importFlags] : [trustFlag]; + setFlags(projects, target, targetFlags); + setFlags(projects, project, projectFlags); + keys = [[target, targetFlags], [project, projectFlags]]; + } else { + setFlags(projects, target, [trustFlag]); + keys = [[target, [trustFlag]]]; } - entry.hasTrustDialogAccepted = true; - projects[worktree] = entry; // Unpredictable name plus an exclusive create: the config directory may be // writable by another local account, and a predictable path could be // pre-created there as a symlink that a plain write would follow into some @@ -250,7 +512,8 @@ const attempt = () => { if (!renamed) fs.rmSync(tmp, { force: true }); } const back = JSON.parse(fs.readFileSync(store, "utf8")); - return back.projects?.[worktree]?.hasTrustDialogAccepted === true ? "recorded" : "dropped"; + const landed = keys.every(([key, flags]) => flagsLanded(back.projects, key, flags)); + return landed ? "recorded" : "dropped"; }; try { for (let i = 0; i < 3; i += 1) { @@ -265,11 +528,18 @@ try { console.error(`error: ${err.message}`); process.exit(1); } -console.error(`error: ${store} did not retain trust for ${worktree} after 3 attempts`); +console.error(`error: ${store} did not retain trust for ${target}${project ? ` and ${project}` : ""} after 3 attempts`); process.exit(1); NODE then - refuse "could not record trust for '$WT_REAL' in '$STORE'" + if [ "$MODE" = worktree ]; then + refuse "could not record trust for '$TARGET_REAL' and project '$PROJ_CANON' in '$STORE'" + else + refuse "could not record trust for '$TARGET_REAL' in '$STORE'" + fi fi -echo "trusted: $WT_REAL" +echo "trusted: $TARGET_REAL" +if [ "$MODE" = worktree ]; then + echo "trusted (project root): $PROJ_CANON" +fi diff --git a/bin/fm-composer-lib.sh b/bin/fm-composer-lib.sh index cdea8d98abd..058dadc7293 100644 --- a/bin/fm-composer-lib.sh +++ b/bin/fm-composer-lib.sh @@ -290,7 +290,7 @@ fm_composer_strip_ghost() { # Matching a footer to confirm a keystroke landed is a different question from # asking what a worker is doing, and the two must not be conflated. # Delivery-only rendered busy footers per harness. claude/codex: "esc to -# interrupt"; opencode: "esc interrupt"; pi: "Working..."; omp: "Working…"; grok: "Ctrl+c:cancel". +# interrupt"; opencode: "esc interrupt"; pi: "Working..."; omp: "Working…"; grok: "Ctrl+c:cancel"; agy: "esc to cancel". # Claude's current spinner has a rotating glyph and word, but every active-turn # line has an ellipsis followed by a parenthesized elapsed duration. Keep this # signature separate from the shared default because that shape is not generic @@ -311,7 +311,11 @@ fm_composer_strip_ghost() { # part of that union for the same reason the others are: without it a cursor # submit could never be acknowledged, because cursor parks its terminal cursor # outside its composer and the composer verdict is therefore always `unknown`. -FM_DELIVERY_BUSY_REGEX_DEFAULT='esc (to )?interrupt|Working(\.\.\.|…)|Ctrl\+c:cancel|ctrl\+c to stop' +# agy's `esc to cancel` is part of the union for the same reason: an explicit +# tmux agy endpoint reaches the submit core with no recorded harness, and its +# bare `>` composer verdict is `unknown`, so the busy footer is the only +# turn-started acknowledgement that path can read. +FM_DELIVERY_BUSY_REGEX_DEFAULT='esc (to )?interrupt|Working(\.\.\.|…)|Ctrl\+c:cancel|ctrl\+c to stop|esc[[:space:]]+to[[:space:]]+cancel' FM_DELIVERY_CLAUDE_BUSY_REGEX_DEFAULT='esc to interrupt|…[[:space:]]+\([0-9]+[smh]' FM_DELIVERY_CODEX_BUSY_REGEX_DEFAULT='esc to interrupt' FM_DELIVERY_OPENCODE_BUSY_REGEX_DEFAULT='esc interrupt' @@ -342,6 +346,14 @@ FM_DELIVERY_GROK_BUSY_REGEX_DEFAULT='Ctrl\+c:cancel' # injection. Cursor's recorded worker state comes from its transcript fold in # bin/fm-busy-lib.sh, never from this row. FM_DELIVERY_CURSOR_BUSY_REGEX_DEFAULT='ctrl\+c to stop' +# agy (Antigravity CLI) renders a pinned status row while a turn runs: the +# `esc to cancel` token on the left and the model cell on the right (verified +# live, agy 1.2.0; the idle row shows `? for shortcuts` instead). The +# `Generating...` spinner word beside it is a free-floating output line and is +# deliberately not matched, so echoed worker output cannot fake an +# acknowledgement. Delivery guard only; recorded worker state comes from the +# agy-regex fold in bin/fm-busy-lib.sh. +FM_DELIVERY_AGY_BUSY_REGEX_DEFAULT='esc[[:space:]]+to[[:space:]]+cancel' FM_DELIVERY_KIMI_BUSY_REGEX_DEFAULT='^[[:space:]]*(🌑|🌒|🌓|🌔|🌕|🌖|🌗|🌘)[[:space:]]+·[[:space:]]+' fm_busy_lines_match() { # [harness] @@ -357,6 +369,7 @@ fm_busy_lines_match() { # [harness] pi|pi-signed) regex=$FM_DELIVERY_PI_BUSY_REGEX_DEFAULT ;; omp) regex=$FM_DELIVERY_OMP_BUSY_REGEX_DEFAULT ;; grok) regex=$FM_DELIVERY_GROK_BUSY_REGEX_DEFAULT ;; + agy) regex=$FM_DELIVERY_AGY_BUSY_REGEX_DEFAULT ;; kimi) regex=$FM_DELIVERY_KIMI_BUSY_REGEX_DEFAULT ;; cursor) regex=$FM_DELIVERY_CURSOR_BUSY_REGEX_DEFAULT ;; '') regex=$FM_DELIVERY_BUSY_REGEX_DEFAULT ;; diff --git a/bin/fm-config-inherit-lib.sh b/bin/fm-config-inherit-lib.sh index de4e53630f5..c50e3a60cfe 100644 --- a/bin/fm-config-inherit-lib.sh +++ b/bin/fm-config-inherit-lib.sh @@ -20,6 +20,9 @@ # default-off W3C trace-context setup, while live convergence leaves it unchanged. # The primary passes its frozen home-session decision into a newly launched # Secondmate; see docs/trace-context.md. +# Primary config/claude-permission-mode is a captain-wide safety preference +# (bypass or auto for every claude launch), so it flows down too and a +# secondmate's own claude crewmates launch on the same permission posture. # It also pushes # the one primary-authoritative shared captain-preference file, # data/captain-shared.md, into each secondmate home's data/ as a read-only copy. @@ -76,7 +79,7 @@ FM_SHARED_CAPTAIN_MODE="444" # The declared inheritable set (space-separated, config-dir-relative item paths). # Extend here to inherit more of the primary's local config; override via the # environment only in tests. Items must not contain whitespace. -FM_INHERITABLE_CONFIG="${FM_INHERITABLE_CONFIG:-crew-dispatch.json crew-harness backlog-backend backend herdr-presentation-spaces startup-memory-budget trace-context gbrain.json project-board launch-env-allowlist}" +FM_INHERITABLE_CONFIG="${FM_INHERITABLE_CONFIG:-crew-dispatch.json crew-harness backlog-backend backend herdr-presentation-spaces startup-memory-budget trace-context gbrain.json project-board launch-env-allowlist claude-permission-mode}" # Items whose value is a home-SESSION enablement decision rather than durable # local configuration. They are inherited at the launch convergence point, where diff --git a/bin/fm-control-lib.sh b/bin/fm-control-lib.sh index 81352f09584..3dfd843f77b 100644 --- a/bin/fm-control-lib.sh +++ b/bin/fm-control-lib.sh @@ -76,13 +76,15 @@ fm_control_harness_supported() { # <harness> # harness= that way), which is why the spawn adapters match `claude*`, `muse*`, # and friends. This is the one place that prefix rule is stated. `pi` and # `pi-signed` are exact because a `pi*` prefix would swallow the signed adapter, -# `omp` is exact because an `omp*` prefix would claim unrelated commands, and an +# `omp` is exact because an `omp*` prefix would claim unrelated commands, `agy` +# is exact for the same reason on an even shorter name, and an # unrecognized value returns nonzero rather than being guessed into a family. fm_control_harness_family() { # <recorded-harness> case "${1-}" in pi) printf 'pi' ;; pi-signed) printf 'pi-signed' ;; omp) printf 'omp' ;; + agy) printf 'agy' ;; claude*) printf 'claude' ;; codex*) printf 'codex' ;; opencode*) printf 'opencode' ;; @@ -117,7 +119,9 @@ fm_control_harness_supports_kind() { # <harness> <kind> # gemini names its own key in the running turn's status row # (`(esc to cancel, <n>s)`), and a single Escape was verified to cancel it. # rovo cancels on a single Escape too, printing "Agent cancelled" (verified, -# 202609.1.2). omp (Oh My Pi) shares Pi's single Escape, empty composer +# 202609.1.2). agy cancels on a single Escape, printing the Interrupted row +# with an idle composer and no repollution (verified live, agy 1.2.0 through +# Herdr). omp (Oh My Pi) shares Pi's single Escape, empty composer # afterwards, and /quit exit (verified omp 18.1.2 in a PTY, re-verified 18.1.11 # through Herdr). fm_control_interrupt_key() { # <harness> @@ -182,8 +186,8 @@ fm_control_interrupt_ack_source() { # <harness> # The command that exits the agent from its own composer. fm_control_exit_command() { # <harness> case "${1-}" in - claude|opencode|grok|kimi|cursor|muse|rovo|agy) printf '/exit' ;; - codex|pi|pi-signed|omp|gemini) printf '/quit' ;; + claude|opencode|grok|kimi|cursor|muse|rovo) printf '/exit' ;; + codex|pi|pi-signed|omp|gemini|agy) printf '/quit' ;; *) return 1 ;; esac } diff --git a/bin/fm-crew-state.sh b/bin/fm-crew-state.sh index df12bed0b10..a60a3075c25 100755 --- a/bin/fm-crew-state.sh +++ b/bin/fm-crew-state.sh @@ -1366,8 +1366,10 @@ if ! pane_readable "$BACKEND_TARGET"; then # genuine server death - a socket-connection failure is NOT # covered by the unknown-never-death rule above). # dead - the endpoint exists but confidently has no agent (herdr's agent - # get answered agent_not_found; tmux's readable foreground process - # group is nothing but shells), still positive death evidence. + # get answered agent_not_found, or its registration lingers over a + # pane whose processes are nothing but shells - issue #4115; + # tmux's readable foreground process group is nothing but + # shells), still positive death evidence. # alive - the endpoint and its agent answered and only the heavy # scrollback read failed, so the live state is classified by the # normal flow below instead of being discarded. diff --git a/bin/fm-decision-hold.sh b/bin/fm-decision-hold.sh index c1a7a6c9f03..f5538deed16 100755 --- a/bin/fm-decision-hold.sh +++ b/bin/fm-decision-hold.sh @@ -59,7 +59,7 @@ compose() { # <origin> <key> } task_show() { - (cd "$FM_HOME" && tasks-axi show "$1" --full) 2>/dev/null + FM_HOME="$FM_HOME" FM_DATA_OVERRIDE='' "$SCRIPT_DIR/fm-tasks-axi.sh" show "$1" --full 2>/dev/null } show_field() { @@ -169,7 +169,7 @@ command_resolve() { for dep in $routed; do show=$(task_show "$dep") || fail "routed task $dep disappeared before routing" if list_has_key "$(normalized_blocked_by "$show")" "$id"; then - (cd "$FM_HOME" && tasks-axi unblock "$dep" --by "$id" >/dev/null) \ + FM_HOME="$FM_HOME" FM_DATA_OVERRIDE='' "$SCRIPT_DIR/fm-tasks-axi.sh" unblock "$dep" --by "$id" >/dev/null \ || fail "could not route the recorded decision to $dep" fi done diff --git a/bin/fm-fleet-snapshot.sh b/bin/fm-fleet-snapshot.sh index 5a32c500b32..df00e2c6e26 100755 --- a/bin/fm-fleet-snapshot.sh +++ b/bin/fm-fleet-snapshot.sh @@ -145,6 +145,10 @@ # reconcile_inventory independently of projection trust. # Actionable captain holds appear in decisions_open; every captain hold remains # in the bounded queued inventory with its structured classification metadata. +# Before that queued bound is applied, non-captain-actionable rows are selected +# ahead of captain-actionable rows so separately projected live decisions cannot +# crowd Charted-Next-eligible work out of the summary. Each group is ordered by +# filed date newest first, with undated rows stable at the end. # Structured-home input must declare the current home-summary and hold-classifier # schemas; a live ledger or cached copy missing either declaration or declaring # an unsupported version is unavailable even when it contains no captain holds. @@ -1489,6 +1493,16 @@ secondmate_home_summary_json() { # <backlog-json-file> <tasks-json-file> | def trunc($n): tostring | gsub("\\s+"; " ") | if length > $n then .[:$n] + "…" else . end; + def filed_epoch: + (.since // null) as $filed + | if ($filed | type) != "string" then null + elif ($filed | test("T")) then try ($filed | fromdateiso8601) catch null + else try (($filed + "T00:00:00Z") | fromdateiso8601) catch null end; + def newest_filed_first: + to_entries + | sort_by((.value | filed_epoch) as $epoch + | if $epoch == null then [1, 0, .key] else [0, -$epoch, .key] end) + | map(.value); ([ $backlog.records[]? | select((.state == "in_flight" or .state == "queued") and (.structured | not)) ]) as $unstructured_current | ([ $backlog.records[]? | select(.state == "in_flight" and .structured) ]) as $owned_in_flight @@ -1552,6 +1566,7 @@ secondmate_home_summary_json() { # <backlog-json-file> <tasks-json-file> | select(.id == $work.id and .current_state.state == "working") | {id,kind,state:.current_state.state, repo:(($work.repo // .project // null) | if . == null then null else trunc(120) end), + name:(($work.title // null) | if . == null then null else trunc(70) end), source:.current_state.source, doing:((.current_state.detail // "") | trunc(120))} ]) as $active_all | ([ $owned_in_flight[] as $work @@ -1626,7 +1641,11 @@ secondmate_home_summary_json() { # <backlog-json-file> <tasks-json-file> hold_age_days:(.hold_age_days // null), captain_actionable:(.captain_actionable // false), repo:((.repo // null) | if . == null then null else trunc(120) end), - kind:((.kind // null) | if . == null then null else trunc(40) end)}][:$queued_n]), + kind:((.kind // null) | if . == null then null else trunc(40) end), + since:((.since // null) | if . == null then null else trunc(40) end)}] + | ((map(select(.captain_actionable != true)) | newest_filed_first) + + (map(select(.captain_actionable == true)) | newest_filed_first)) + | .[:$queued_n]), landed:(if $landed_n == 0 then $landed_all else $landed_all[:$landed_n] end), endpoints:([$tasks[] | {id,state:.current_state.state,source:.current_state.source, endpoint:(.endpoint + {target:((.endpoint.target // null) | if . == null then null else trunc(240) end)})}][:$child_n]), diff --git a/bin/fm-guard.sh b/bin/fm-guard.sh index 6abc517c12e..b086526aae6 100755 --- a/bin/fm-guard.sh +++ b/bin/fm-guard.sh @@ -226,6 +226,7 @@ if [ "$watcher_healthy" = false ]; then fix=$("$SCRIPT_DIR/fm-supervision-instructions.sh" \ --read-only "$READ_ONLY" \ --afk "$afk" \ + --afk-mode "$(fm_afk_mode "$STATE")" \ --x-mode "$x_mode" \ --queue-pending "$queue_arg" \ --repair-line 2>/dev/null || printf '%s\n' 'Repair missing watcher supervision according to the session-start operating block.') diff --git a/bin/fm-harness.sh b/bin/fm-harness.sh index 96443cf60c9..7989643f1b6 100755 --- a/bin/fm-harness.sh +++ b/bin/fm-harness.sh @@ -1,6 +1,6 @@ #!/usr/bin/env bash # Detect the agent harness this process tree runs on. -# Usage: fm-harness.sh print own harness: claude|codex|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse|rovo|omp|unknown +# Usage: fm-harness.sh print own harness: claude|codex|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse|rovo|omp|agy|unknown # fm-harness.sh crew print the effective CREWMATE harness # (config/crew-harness; "default" resolves to own) # fm-harness.sh secondmate print the harness the PRIMARY uses to launch @@ -19,13 +19,42 @@ # codex-native/<id>. Other efforts retain # their adapter's existing policy. Native # Codex validates model support at startup. +# fm-harness.sh ancestry [<pid>] print "<strength> <harness>" for the nearest +# harness process at or above <pid> (default this +# process), or nothing when the walk finds none. +# Ancestry evidence only, with no marker layer, so +# a real harness process can be asked what the walk +# makes of it (tests/fm-harness-liveness-drift-live-e2e.test.sh). +# fm-harness.sh ancestry-descent [<pid>] [<leaf-pid>...] +# print each DISTINCT "<strength> <harness>" the walk +# reaches from the vantages on the UPWARD path +# between the deepest descendant of <pid> and <pid> +# itself, deepest first. Same evidence-only purpose +# as `ancestry`, asked from the vantage point a tool +# subprocess actually occupies rather than from the +# top of the session, which is the only place a +# harness behind an interpreter shim can be seen at +# comm strength. Optional <leaf-pid> values restrict +# which descendants may be chosen as the deepest one, +# so a caller that knows the terminal's foreground +# process group can keep a backgrounded process out +# of the selection. # config/secondmate-harness format: a single line "<harness> [<model>] [<effort>]", # whitespace-separated. A bare "<harness>" (today's format) behaves exactly as before: # harness only, no model/effort. Only the first non-empty, non-comment line is parsed. # Model/effort come ONLY from this file - config/crew-harness stays a bare adapter # name and is never parsed for a model. -# Detection layers: verified environment markers first, then process ancestry. -# Record each newly verified env marker here. +# Detection evidence and precedence: +# Markers - verified environment variables a harness publishes about itself. +# Cheap and unambiguous about WHICH harness set them, but they are +# ordinary environment state: a child inherits them, and a terminal +# multiplexer can replay a stale one into an unrelated session. +# Ancestry - the nearest harness process in this process's parent chain. This +# is the structural fact about who actually owns the process tree, +# so it is what settles a disagreement. +# detect_own is the single owner of how the two combine; harness_marker and +# harness_ancestry only report evidence. Record each newly verified env marker +# in harness_marker, and each newly verified command name in harness_ancestry. set -u SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -38,24 +67,18 @@ CONFIG="${FM_CONFIG_OVERRIDE:-$FM_HOME/config}" # shellcheck source=bin/fm-gemini-lib.sh . "$SCRIPT_DIR/fm-gemini-lib.sh" -detect_own() { - # Layer 1: environment markers for verified harnesses. - # Keep marker detection before ancestry detection as an explicit precedence rule. - # Claude, Pi, Grok, and Cursor set verified markers of their own; codex, - # opencode, Kimi, and Muse are markerless, so a foreign marker retained in a terminal - # multiplexer's stored environment can silently misidentify one of them before - # ancestry is consulted. This is a precedence hazard, not evidence that - # CLAUDECODE inheritance into a kimi child was observed; it was not observed. - # Cursor is checked BEFORE claude, deliberately. cursor-agent does NOT clear - # an inherited CLAUDECODE, so a cursor worker launched from a claude primary - # carries BOTH markers and whichever is tested first wins. Cursor's own - # markers are unambiguous when present, so ordering them first is what makes - # the verdict correct; bin/fm-spawn.sh additionally clears the foreign markers - # at the launch boundary. Both are kept: the launch sanitization only covers - # sessions fm-spawn started, while this ordering also covers a cursor session - # a human started by hand. Verified live on cursor-agent 2026.08.11-e8db854: - # CURSOR_INVOKED_AS=cursor-agent is set on the agent process itself, and - # CURSOR_AGENT=1 is set for the child/tool processes this script runs as. +# Print the harness named by a verified environment marker, or nothing when no +# marker is present. Markers only report what the environment CLAIMS; detect_own +# decides whether that claim survives contradicting ancestry. +harness_marker() { + # Cursor is tested BEFORE claude, deliberately. cursor-agent does NOT clear an + # inherited CLAUDECODE, so a cursor session started by hand from a claude + # primary carries BOTH markers and whichever is tested first wins. This + # ordering only settles the case where ancestry finds nothing to arbitrate + # with; a nearer claude ancestor still outranks both in detect_own. + # Verified live on cursor-agent 2026.08.11-e8db854: CURSOR_INVOKED_AS=cursor-agent + # is set on the agent process itself, and CURSOR_AGENT=1 is set for the + # child/tool processes this script runs as. [ "${CURSOR_AGENT:-}" = "1" ] && { echo cursor; return; } [ "${CURSOR_INVOKED_AS:-}" = "cursor-agent" ] && { echo cursor; return; } # Gemini is checked BEFORE claude for exactly cursor's reason above: the @@ -110,89 +133,21 @@ detect_own() { # identified, and any rule that must be RELIABLE under grok has to test the hook # markers too (see .claude/settings.json Stop entries, docs/turnend-guard.md). [ "${GROK_AGENT:-}" = "1" ] && { echo grok; return; } - # muse (Muse Code) publishes no harness-identity marker of its own. The only - # MUSE_* variable it is documented to hand a child is MUSE_CURRENT_SESSION_LOG, - # a per-session log PATH rather than an identity, and its export to tool - # subprocesses is unverified (verified: muse 0.1.0-R708.1), so muse is detected - # by ancestry alone below. Do NOT promote MUSE_CURRENT_SESSION_LOG to a marker - # without verifying it reaches children AND that it cannot survive in a - # multiplexer's stored environment, which is the precedence hazard above. - # Layer 2: walk the parent chain and match the command name. - local pid=$$ comm args argv0 - for _ in 1 2 3 4 5 6 7 8; do - comm=$(ps -o comm= -p "$pid" 2>/dev/null) || break - argv0=$(fm_cursor_argv0_for_pid "$pid" "$comm" 2>/dev/null || true) - if fm_cursor_process_matches "$comm" '' "$argv0"; then - echo cursor - return - fi - if fm_gemini_path_is_gemini "$comm"; then - echo gemini - return - fi - case "$(basename -- "$comm")" in - # gemini precedes claude here for the same precedence reason as the - # marker layer above, so a gemini worker under a claude primary is never - # read as claude. This arm covers a natively-named gemini binary only. - # It does NOT reach the currently installed CLI, which is a node bundle - # (~/.local/bin/gemini -> @google/gemini-cli/bundle/gemini.js): modern - # Node on Linux reports `comm` as MainThread rather than node (measured - # on Node v24.20.0), so neither this arm nor the node interpreter arm - # below matches a live gemini process. GEMINI_CLI above is therefore - # load-bearing for gemini rather than a fast path, which is why gemini - # is not offered as a primary or secondmate harness. Do NOT add - # MainThread to the interpreter arm to close this: that would make the - # args of EVERY node process searchable and let an unrelated node - # command carrying a harness name in its arguments claim an identity. - *claude*) echo claude; return ;; - *codex*) echo codex; return ;; - *opencode*) echo opencode; return ;; - *grok*) echo grok; return ;; - kimi) echo kimi; return ;; - rovo) echo rovo; return ;; - # muse's installed launcher ~/.local/bin/muse execs ~/.local/bin/muse-bin-<version> - # (verified in the published launcher, muse 0.1.0-R708.1), so the live process - # name carries the version and CHANGES on every auto-update. Match the stable - # prefix rather than any exact name. Deliberately anchored, never *muse*, so - # unrelated commands (musescore, amuse) cannot be misread as this harness. - muse|muse-bin-*) echo muse; return ;; - pi-signed) echo pi; return ;; - pi) echo pi; return ;; - # omp is a Bun-compiled single binary whose process name is exactly `omp` - # (verified, omp 18.1.11: `ps -o comm=` reports omp from both its `!` - # bash path and the model's bash tool). Anchored, never *omp*, so ompd, - # comp, and similar unrelated commands are not misread as this harness. - # It sits above the node*|python* interpreter fallback deliberately: the - # optional claude-bridge extension runs a nested executable literally - # named `claude` with its own node child, and that fallback's *claude* - # args glob would otherwise claim it if that subtree were ever walked. - omp) echo omp; return ;; - node*|python*) - # Bare interpreter: match the harness name in its script path. - args=$(ps -o args= -p "$pid" 2>/dev/null) - if fm_gemini_args_are_gemini "$args"; then - echo gemini - return - fi - case "$args" in - *claude*) echo claude; return ;; - *codex*) echo codex; return ;; - *opencode*) echo opencode; return ;; - *grok*) echo grok; return ;; - *" pi "*|*/pi) echo pi; return ;; - esac ;; - esac - pid=$(ps -o ppid= -p "$pid" 2>/dev/null | tr -d ' ') - if [ -z "$pid" ] || [ "$pid" -le 1 ]; then - break - fi - done - echo unknown + # codex, opencode, kimi, muse, and agy publish no harness-identity marker at all, so + # they are never named here and are identified by ancestry alone. That is the + # whole reason a foreign marker must not outrank ancestry: with markers winning + # unconditionally, any retained CLAUDECODE would silently rename one of them. + # muse's only documented child variable is MUSE_CURRENT_SESSION_LOG, a + # per-session log PATH rather than an identity, and its export to tool + # subprocesses is unverified (verified: muse 0.1.0-R708.1). Do NOT promote it + # to a marker without verifying it reaches children AND that it cannot survive + # in a multiplexer's stored environment. + return 0 } # True when an exact `omp` process sits within eight parents of this one. The -# same anchored match as the ancestry walk in detect_own, kept separate so the -# marker precedence above can demand real process evidence. +# same anchored match as the ancestry walk below, kept separate so the marker +# precedence above can demand real process evidence before trusting FM_OMP_HARNESS. ancestry_names_omp() { local pid=$$ comm for _ in 1 2 3 4 5 6 7 8; do @@ -204,6 +159,259 @@ ancestry_names_omp() { return 1 } +# Print "<strength> <harness>" when one process identifies a harness, or nothing. +# Strength records how the match was made: +# comm - the ancestor's own executable name identifies the harness. This is a +# structural fact about the running program, so it outranks a marker. +# args - a bare interpreter matched only because a harness name appears in the +# script path it was handed. This is the weakest inference in this file +# (any node process holding a harness-shaped path matches it), so it is +# used only when no marker is present. +harness_process_verdict() { # <pid> + local pid=$1 comm args argv0 + comm=$(ps -o comm= -p "$pid" 2>/dev/null) || return 0 + argv0=$(fm_cursor_argv0_for_pid "$pid" "$comm" 2>/dev/null || true) + if fm_cursor_process_matches "$comm" '' "$argv0"; then + echo "comm cursor" + return + fi + if fm_gemini_path_is_gemini "$comm"; then + echo "comm gemini" + return + fi + case "$(basename -- "$comm")" in + # gemini precedes claude here for the same precedence reason as the + # marker layer above, so a gemini worker under a claude primary is never + # read as claude. This arm covers a natively-named gemini binary only. + # It does NOT reach the currently installed CLI, which is a node bundle + # (~/.local/bin/gemini -> @google/gemini-cli/bundle/gemini.js): modern + # Node on Linux reports `comm` as MainThread rather than node (measured + # on Node v24.20.0), so neither this arm nor the node interpreter arm + # below matches a live gemini process. GEMINI_CLI above is therefore + # load-bearing for gemini rather than a fast path, which is why gemini + # is not offered as a primary or secondmate harness. Do NOT add + # MainThread to the interpreter arm to close this: that would make the + # args of EVERY node process searchable and let an unrelated node + # command carrying a harness name in its arguments claim an identity. + *claude*) echo "comm claude"; return ;; + *codex*) echo "comm codex"; return ;; + *opencode*) echo "comm opencode"; return ;; + *grok*) echo "comm grok"; return ;; + kimi) echo "comm kimi"; return ;; + rovo) echo "comm rovo"; return ;; + # muse's installed launcher ~/.local/bin/muse execs ~/.local/bin/muse-bin-<version> + # (verified in the published launcher, muse 0.1.0-R708.1), so the live process + # name carries the version and CHANGES on every auto-update. Match the stable + # prefix rather than any exact name. Deliberately anchored, never *muse*, so + # unrelated commands (musescore, amuse) cannot be misread as this harness. + muse|muse-bin-*) echo "comm muse"; return ;; + # Both Pi identities share this launcher name. Ancestry can only prove the + # FAMILY; only the launch-boundary marker selects the signed identity, which + # is why detect_own keeps a marker that agrees on the family. + pi-signed) echo "comm pi"; return ;; + pi) echo "comm pi"; return ;; + # omp is a Bun-compiled single binary whose process name is exactly `omp` + # (verified, omp 18.1.11: `ps -o comm=` reports omp from both its `!` + # bash path and the model's bash tool). Anchored, never *omp*, so ompd, + # comp, and similar unrelated commands are not misread as this harness. + # It sits above the node*|python* interpreter fallback deliberately: the + # optional claude-bridge extension runs a nested executable literally + # named `claude` with its own node child, and that fallback's *claude* + # args glob would otherwise claim it if that subtree were ever walked. + omp) echo "comm omp"; return ;; + # agy (Antigravity CLI) is a Go-compiled single binary whose process name + # is exactly `agy` (verified, agy 1.2.0: `ps -o comm=` reports agy and + # Herdr's process-info reports name agy with argv[0] agy). Anchored, never + # *agy*, so unrelated commands cannot be misread as this harness. agy + # publishes no harness-identity marker of its own (a live 1.2.0 TUI + # carries no AGY_* or ANTIGRAVITY_* variable; AGENT=1 seen there is an + # inherited launcher value, not an agy identity), so like muse it is + # detected by ancestry alone. + agy) echo "comm agy"; return ;; + node*|python*) + # Bare interpreter: match the harness name in its script path. + args=$(ps -o args= -p "$pid" 2>/dev/null) + if fm_gemini_args_are_gemini "$args"; then + echo "args gemini" + return + fi + case "$args" in + *claude*) echo "args claude"; return ;; + *codex*) echo "args codex"; return ;; + *opencode*) echo "args opencode"; return ;; + *grok*) echo "args grok"; return ;; + *" pi "*|*/pi) echo "args pi"; return ;; + esac ;; + esac +} + +# Print the verdict for the NEAREST harness process in the parent chain, or +# nothing when the walk finds none. The nearest match wins, so a worker nested +# inside another harness resolves to its own harness. +harness_ancestry() { # [<pid>] + local pid=${1:-$$} verdict + for _ in 1 2 3 4 5 6 7 8; do + verdict=$(harness_process_verdict "$pid") + [ -z "$verdict" ] || { echo "$verdict"; return; } + pid=$(ps -o ppid= -p "$pid" 2>/dev/null | tr -d ' ') + # Stop only once the walk has EXAMINED the top of the chain. Inside a PID + # namespace the harness itself is pid 1 - a container, or the `codex sandbox` + # this boundary was proven in - so breaking as soon as the next pid is 1 + # skips the one process that identifies the session and hands the verdict + # straight back to a retained marker. A host's real pid 1 (init, systemd, + # launchd) matches no harness name above, so examining it costs one ps call + # and can introduce no false positive. + case "$pid" in '' | *[!0-9]*) break ;; esac + [ "$pid" -ge 1 ] || break + done + return 0 +} + +# Print the pids on the UPWARD path between the deepest descendant of <root> and +# <root> itself, deepest first. Optional <eligible-leaf-pid> values restrict which +# descendants may be chosen as that deepest one; with none given every descendant +# is eligible. Bounded to the same eight levels harness_ancestry climbs, so a deep +# or pathological tree cannot make this walk unbounded. +process_descent_path() { # <root> [<eligible-leaf-pid>...] + local root=${1:-$$} eligible any hit pairs frontier next pid child parent verdict + local parents='' depth=0 best best_depth=0 best_strength='' hops=0 + case "$root" in '' | *[!0-9]*) return 0 ;; esac + shift 2>/dev/null || true + eligible=" ${*+$*} " + any=0 + [ "$#" -eq 0 ] && any=1 + pairs=$(ps -eo pid=,ppid= 2>/dev/null) || { printf '%s\n' "$root"; return 0; } + best=$root + frontier=$root + while [ -n "$frontier" ] && [ "$depth" -lt 8 ]; do + next= + for pid in $frontier; do + while read -r child parent; do + [ "$parent" = "$pid" ] || continue + [ "$child" != "$pid" ] || continue + parents="$parents $child:$pid" + next="$next $child" + if [ "$any" = 1 ]; then + hit=1 + else + case "$eligible" in + *" $child "*) hit=1 ;; + *) hit=0 ;; + esac + fi + if [ "$hit" = 1 ]; then + verdict=$(harness_process_verdict "$child") + if [ $((depth + 1)) -gt "$best_depth" ]; then + best=$child + best_depth=$((depth + 1)) + best_strength=${verdict%% *} + # At equal depth, prefer the leaf whose own executable reaches comm + # strength. Otherwise an earlier MCP interpreter carrying a foreign + # harness path can hide a native harness sibling purely through ps + # ordering. This repairs the chosen path's comm-strength guarantee; + # args-strength foreign verdicts remain excluded from cross-checking. + elif [ $((depth + 1)) -eq "$best_depth" ] \ + && [ "$best_strength" != comm ] && [ "${verdict%% *}" = comm ]; then + best=$child + best_strength='comm' + fi + fi + done <<EOF +$pairs +EOF + done + frontier=$next + depth=$((depth + 1)) + done + + pid=$best + while [ -n "$pid" ] && [ "$hops" -le 8 ]; do + printf '%s\n' "$pid" + [ "$pid" != "$root" ] || break + parent= + case "$parents" in + *" $pid:"*) + parent=${parents##*" $pid:"} + parent=${parent%% *} ;; + esac + pid=$parent + hops=$((hops + 1)) + done +} + +# Print each DISTINCT "<strength> <harness>" verdict harness_ancestry reaches from +# the vantages on the upward path between the deepest descendant of <root> and +# <root>, one per line, deepest first. +# +# Why a descent path and not <root> alone: detect_own always runs from a TOOL +# SUBPROCESS inside a session, never from the process at the top of it, and that +# difference decides whether a retained foreign marker can rename the session. A +# harness that ships as an interpreter shim spawning its native binary as a CHILD +# is only args strength when asked from the shim, and detect_own hands an +# args-strength verdict straight back to the marker; the native child is comm +# strength and outranks it. Asking from below is what puts the question at the +# vantage point a real session uses, so a guard built on this can assert the +# strength the shipped guarantee actually depends on +# (tests/fm-harness-liveness-drift-live-e2e.test.sh). +# +# Why the upward path and not the whole subtree: harness_ancestry only ever climbs, +# so a SIBLING branch is a vantage firstmate's own detection can never occupy. A +# harness-spawned MCP server running as `node <home>/.claude/mcp/<server>.js` matches +# *claude* on its script path in the bare-interpreter branch above and would report a +# foreign harness from a process no real tool subprocess ever asks from. +harness_ancestry_descent() { # <root> [<eligible-leaf-pid>...] + local pid verdict seen= + for pid in $(process_descent_path "$@"); do + verdict=$(harness_ancestry "$pid") + [ -n "$verdict" ] || continue + case "$seen" in *"|$verdict|"*) continue ;; esac + seen="$seen|$verdict|" + printf '%s\n' "$verdict" + done +} + +# Collapse a verdict to the harness FAMILY its evidence can actually prove, so a +# marker's more specific verdict and ancestry's coarser one are not read as a +# disagreement. Only Pi has two identities behind one launcher name. +harness_family() { + case "$1" in + pi-signed) printf 'pi\n' ;; + *) printf '%s\n' "$1" ;; + esac +} + +# Combine the two evidence layers. The precedence boundary, in one rule: a +# marker names its harness, but only ancestry proves which harness owns this +# process tree, so a structural (comm) ancestor of a DIFFERENT harness wins. +# - No ancestry match: the marker is the only evidence there is. +# - No marker: ancestry is the only evidence there is. +# - Same family: keep the marker's verdict, which is the more specific one +# (pi-signed, which ancestry can only see as pi). +# - Different harness, structural ancestor: ancestry wins. This is what stops +# an inherited or multiplexer-retained CLAUDECODE from renaming a markerless +# codex, opencode, kimi, or muse session, and symmetrically stops a retained +# CURSOR_AGENT from renaming a claude worker nested under cursor. +# - Different harness, interpreter-args ancestor only: the marker wins, because +# a harness-shaped path in some node process's arguments is weaker evidence +# than a harness publishing its own identity. +detect_own() { + local marker ancestry strength harness + marker=$(harness_marker) + ancestry=$(harness_ancestry) + if [ -z "$ancestry" ]; then + if [ -n "$marker" ]; then echo "$marker"; else echo unknown; fi + return + fi + strength=${ancestry%% *} + harness=${ancestry#* } + [ -n "$marker" ] || { echo "$harness"; return; } + if [ "$(harness_family "$marker")" = "$(harness_family "$harness")" ]; then + echo "$marker" + return + fi + if [ "$strength" = comm ]; then echo "$harness"; else echo "$marker"; fi +} + # Resolve the effective crewmate harness: config/crew-harness (a bare adapter # name) wins; absent or "default" mirrors firstmate's own harness. resolve_crew() { @@ -290,6 +498,23 @@ validate_native_effort() { case "${1:-}" in validate-native-effort) shift; validate_native_effort "$@" ;; + ancestry) + case "${2:-}" in + ''|*[!0-9]*) [ -z "${2:-}" ] || { echo "error: ancestry takes a numeric pid" >&2; exit 2; } ;; + esac + harness_ancestry "${2:-$$}" + ;; + ancestry-descent) + shift + for arg in ${1+"$@"}; do + case "$arg" in + ''|*[!0-9]*) echo "error: ancestry-descent takes numeric pids" >&2; exit 2 ;; + esac + done + descent_pid="${1:-$$}" + [ "$#" -eq 0 ] || shift + harness_ancestry_descent "$descent_pid" ${1+"$@"} + ;; crew) resolve_crew ;; secondmate) resolve_secondmate ;; secondmate-model) resolve_secondmate_model ;; diff --git a/bin/fm-herdr-lab-viewer.py b/bin/fm-herdr-lab-viewer.py new file mode 100755 index 00000000000..ce40e0b4a9e --- /dev/null +++ b/bin/fm-herdr-lab-viewer.py @@ -0,0 +1,204 @@ +#!/usr/bin/env python3 +"""Attach one real foreground Herdr viewer to a named lab session over a pty. + +bin/fm-herdr-lab.sh's ``viewer start`` is the only supported caller; run this +through that guard rather than directly, so the lab's ownership tripwire and +refuse-default checks still apply. + +Herdr registers a foreground client only when the attaching terminal reports a +usable window grid. A pty created by ``script`` or a bare ``pty.fork()`` from a +non-tty parent starts at 0x0, which makes Herdr report a zero-sized grid and +keeps ``client.window_title.clear`` answering ``no_foreground_client``. That is +why firstmate could not drive the attached-viewer teardown cases live before +this helper existed. The fix is ordering as much as sizing: the window size is +set on the master fd BEFORE the fork, so the TUI cannot read the grid until it +is already non-zero. + +The child also drops the inherited ``HERDR_*`` variables listed in +``SCRUBBED_ENV`` below. Herdr refuses to launch a nested viewer inside one of +its own panes, and this helper normally runs from exactly there. + +Usage: fm-herdr-lab-viewer.py <session> <pidfile> + +Exit status: + 0 the viewer ran and exited; + 2 the session or pidfile was invalid; + 3 the pty or the viewer process could not be created. +""" + +import errno +import fcntl +import os +import re +import signal +import struct +import subprocess +import sys +import termios + +# Herdr inherits these from the pane this helper runs in, and a nested viewer +# is refused outright. HERDR_SESSION is scrubbed with them so the explicit +# --session argument stays the viewer's only session source. +SCRUBBED_ENV = ( + "HERDR_ENV", + "HERDR_PANE_ID", + "HERDR_TAB_ID", + "HERDR_WORKSPACE_ID", + "HERDR_SOCKET_PATH", + "HERDR_BIN_PATH", + "HERDR_SESSION", +) + +SESSION_PATTERN = re.compile(r"\Afm-lab-[A-Za-z0-9][A-Za-z0-9_-]*\Z") +TERMINATE_GRACE_SECONDS = 5.0 +READ_CHUNK = 65536 +ROWS = 40 +COLS = 120 +TERMINATION_SIGNALS = (signal.SIGTERM, signal.SIGINT, signal.SIGHUP) + + +def _child(slave, master, session): + signal.pthread_sigmask(signal.SIG_UNBLOCK, TERMINATION_SIGNALS) + os.setsid() + try: + fcntl.ioctl(slave, termios.TIOCSCTTY, 0) + except OSError: + pass + for target in (0, 1, 2): + os.dup2(slave, target) + if slave > 2: + os.close(slave) + os.close(master) + env = {key: value for key, value in os.environ.items() if key not in SCRUBBED_ENV} + env.setdefault("TERM", "xterm-256color") + try: + os.execvpe("herdr", ["herdr", "--session", session], env) + except OSError: + pass + os._exit(127) + + +def _process_start(pid): + result = subprocess.run( + ["ps", "-p", str(pid), "-o", "lstart="], + check=True, + capture_output=True, + text=True, + env={**os.environ, "LC_ALL": "C"}, + ) + value = result.stdout.strip() + if not value: + raise RuntimeError("process start time unavailable") + return value + + +def _write_pidfile(path, launcher_pid, viewer_pid): + launcher_start = _process_start(launcher_pid) + viewer_start = _process_start(viewer_pid) + temporary = "%s.%d.tmp" % (path, launcher_pid) + with open(temporary, "w", encoding="utf-8") as handle: + handle.write("launcher_pid=%d\n" % launcher_pid) + handle.write("launcher_start=%s\n" % launcher_start) + handle.write("viewer_pid=%d\n" % viewer_pid) + handle.write("viewer_start=%s\n" % viewer_start) + os.rename(temporary, path) + + +def _drain(master): + while True: + try: + if not os.read(master, READ_CHUNK): + return + except OSError as error: + if error.errno == errno.EINTR: + continue + return + + +def main(argv): + if len(argv) != 3: + sys.stderr.write("fm-herdr-lab-viewer: usage: <session> <pidfile>\n") + return 2 + session, pidfile = argv[1:] + if session == "default" or not SESSION_PATTERN.match(session): + sys.stderr.write("fm-herdr-lab-viewer: refusing session %r\n" % session) + return 2 + if not os.path.isabs(pidfile): + sys.stderr.write("fm-herdr-lab-viewer: pidfile must be an absolute path\n") + return 2 + + try: + master, slave = os.openpty() + except OSError as error: + sys.stderr.write("fm-herdr-lab-viewer: could not create a pty: %s\n" % error) + return 3 + # Before the fork, so the TUI's first grid read already sees a real size. + fcntl.ioctl(master, termios.TIOCSWINSZ, struct.pack("HHHH", ROWS, COLS, 0, 0)) + + signal.pthread_sigmask(signal.SIG_BLOCK, TERMINATION_SIGNALS) + try: + viewer_pid = os.fork() + except OSError as error: + signal.pthread_sigmask(signal.SIG_UNBLOCK, TERMINATION_SIGNALS) + sys.stderr.write("fm-herdr-lab-viewer: could not fork the viewer: %s\n" % error) + return 3 + if viewer_pid == 0: + _child(slave, master, session) + + def _cancel_before_record(signum, _frame): + try: + os.kill(viewer_pid, signal.SIGKILL) + except OSError: + pass + os._exit(128 + signum) + + signal.signal(signal.SIGTERM, _cancel_before_record) + signal.signal(signal.SIGINT, _cancel_before_record) + signal.signal(signal.SIGHUP, _cancel_before_record) + signal.pthread_sigmask(signal.SIG_UNBLOCK, TERMINATION_SIGNALS) + + os.close(slave) + try: + _write_pidfile(pidfile, os.getpid(), viewer_pid) + except (OSError, RuntimeError, subprocess.SubprocessError) as error: + sys.stderr.write("fm-herdr-lab-viewer: could not record process identity: %s\n" % error) + try: + os.kill(viewer_pid, signal.SIGKILL) + except OSError: + pass + os.close(master) + try: + os.waitpid(viewer_pid, 0) + except OSError: + pass + return 3 + + def _signal_viewer(number): + # The viewer may already be gone; that is the outcome we wanted anyway. + try: + os.kill(viewer_pid, number) + except OSError: + pass + + def _terminate(_signum, _frame): + _signal_viewer(signal.SIGTERM) + signal.setitimer(signal.ITIMER_REAL, TERMINATE_GRACE_SECONDS) + + signal.signal(signal.SIGTERM, _terminate) + signal.signal(signal.SIGINT, _terminate) + signal.signal(signal.SIGHUP, _terminate) + signal.signal(signal.SIGALRM, lambda _s, _f: _signal_viewer(signal.SIGKILL)) + + _drain(master) + _terminate(None, None) + signal.setitimer(signal.ITIMER_REAL, TERMINATE_GRACE_SECONDS) + try: + _, status = os.waitpid(viewer_pid, 0) + except OSError: + status = 0 + signal.setitimer(signal.ITIMER_REAL, 0) + return 0 if os.WIFSIGNALED(status) else os.WEXITSTATUS(status) + + +if __name__ == "__main__": + sys.exit(main(sys.argv)) diff --git a/bin/fm-herdr-lab.sh b/bin/fm-herdr-lab.sh index f8ea014c6bc..d0aa633df55 100755 --- a/bin/fm-herdr-lab.sh +++ b/bin/fm-herdr-lab.sh @@ -7,6 +7,8 @@ # fm-herdr-lab.sh prepare <session> # fm-herdr-lab.sh provision <session> # fm-herdr-lab.sh run <session> <herdr arguments...> +# fm-herdr-lab.sh viewer start <session> +# fm-herdr-lab.sh viewer stop <session> # fm-herdr-lab.sh stop <session> # fm-herdr-lab.sh teardown <session> # @@ -23,6 +25,14 @@ # destructive call. # Provision records the running default session as a fleet-state tripwire and # teardown requires that record to be identical afterward. +# The viewer command attaches or detaches one real foreground Herdr client on +# an owned lab session over a fixed 40-row by 120-column pty; +# bin/fm-herdr-lab-viewer.py owns the pty mechanics. +# Start succeeds only when that session reports a foreground client and the +# recorded viewer process still matches its launch identity. +# Stop signals only identity-matched recorded processes and retains its +# ownership record until detach is confirmed or the session is stopped or +# absent; teardown refuses when that stop cannot be confirmed. set -u fm_herdr_lab_error() { @@ -153,6 +163,227 @@ fm_herdr_lab_cli() { # <session> <herdr arguments...> fm_herdr_lab_raw "$name" "$@" } +# --- foreground viewer ------------------------------------------------------ +# +# Herdr counts a client as the session's foreground viewer only once that +# client reports a usable window grid, so a zero-sized pty attaches nothing and +# leaves `terminal title clear` answering no_foreground_client. Attaching a +# real viewer is what lets a test drive the live-client teardown paths instead +# of only their detached halves. bin/fm-herdr-lab-viewer.py owns the pty and +# environment mechanics; the guards below own who may be attached to. +# Per-session locks are deliberately absent: generated fm-lab-<label>-$$-$RANDOM +# names have no caller that starts one viewer concurrently, so locks add risk. +# A subsecond interrupt window and SIGKILL residue are accepted in this isolated +# lab helper because teardown drops any stray viewer connection with the session. + +readonly fm_herdr_lab_viewer_timeout_seconds=5 +readonly fm_herdr_lab_viewer_launcher_grace_seconds=6 + +fm_herdr_lab_viewer_record_path() { # <session> + printf '%s/%s.viewer' "$(fm_herdr_lab_state_dir)" "$1" +} + +fm_herdr_lab_viewer_log_path() { # <session> + printf '%s/%s.viewer.log' "$(fm_herdr_lab_state_dir)" "$1" +} + +fm_herdr_lab_viewer_launcher_path() { + printf '%s/fm-herdr-lab-viewer.py' "$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +} + +# Prints the session's current foreground-client reason, or nothing when it +# cannot be read. +fm_herdr_lab_viewer_reason() { # <session> + local name=$1 out + out=$(fm_herdr_lab_cli "$name" terminal title clear 2>/dev/null) || return 1 + printf '%s' "$out" | jq -r '.result.reason // empty' 2>/dev/null +} + +fm_herdr_lab_process_start() { # <pid> + LC_ALL=C ps -p "$1" -o lstart= 2>/dev/null | sed 's/^[[:space:]]*//;s/[[:space:]]*$//' +} + +fm_herdr_lab_process_parent() { # <pid> + LC_ALL=C ps -p "$1" -o ppid= 2>/dev/null | sed 's/^[[:space:]]*//;s/[[:space:]]*$//' +} + +fm_herdr_lab_viewer_recorded_value() { # <session> <key> + local record value + record=$(fm_herdr_lab_viewer_record_path "$1") + [ -f "$record" ] || return 1 + value=$(sed -n "s/^$2=//p" "$record" | head -n 1) + [ -n "$value" ] || return 1 + printf '%s' "$value" +} + +fm_herdr_lab_viewer_owned_pair() { # <session> + local launcher_pid viewer_pid launcher_start viewer_start current_start parent_pid + launcher_pid=$(fm_herdr_lab_viewer_recorded_value "$1" launcher_pid) || return 1 + viewer_pid=$(fm_herdr_lab_viewer_recorded_value "$1" viewer_pid) || return 1 + case "$launcher_pid:$viewer_pid" in + *[!0-9:]*) return 1 ;; + esac + launcher_start=$(fm_herdr_lab_viewer_recorded_value "$1" launcher_start) || return 1 + viewer_start=$(fm_herdr_lab_viewer_recorded_value "$1" viewer_start) || return 1 + current_start=$(fm_herdr_lab_process_start "$launcher_pid") || return 1 + [ -n "$current_start" ] && [ "$current_start" = "$launcher_start" ] || return 1 + current_start=$(fm_herdr_lab_process_start "$viewer_pid") || return 1 + [ -n "$current_start" ] && [ "$current_start" = "$viewer_start" ] || return 1 + parent_pid=$(fm_herdr_lab_process_parent "$viewer_pid") || return 1 + [ "$parent_pid" = "$launcher_pid" ] || return 1 + printf '%s %s' "$launcher_pid" "$viewer_pid" +} + +fm_herdr_lab_viewer_owned_pid() { # <session> <launcher|viewer> + local pair + pair=$(fm_herdr_lab_viewer_owned_pair "$1") || return 1 + case "$2" in + launcher) printf '%s' "${pair%% *}" ;; + viewer) printf '%s' "${pair#* }" ;; + *) return 1 ;; + esac +} + +fm_herdr_lab_viewer_signal() { # <session> <launcher|viewer> <signal> + local pid + pid=$(fm_herdr_lab_viewer_owned_pid "$1" "$2") || return 0 + kill "-$3" "$pid" 2>/dev/null || true +} + +# True while this lab owns a viewer process that is still running. +fm_herdr_lab_viewer_owned_alive() { # <session> + fm_herdr_lab_viewer_owned_pair "$1" >/dev/null +} + +fm_herdr_lab_viewer_session_stopped_or_absent() { # <session> + local sessions running + sessions=$(fm_herdr_lab_session_list "$1" 2>/dev/null) || return 1 + running=$(printf '%s' "$sessions" | jq -r --arg name "$1" \ + '[.sessions[]? | select(.name == $name) | .running] | if length == 0 then "absent" elif length == 1 then .[0] else "ambiguous" end' \ + 2>/dev/null) || return 1 + [ "$running" = false ] || [ "$running" = absent ] +} + +fm_herdr_lab_viewer_start() { # <session> + local name=$1 record log launcher launcher_pid waited attempt reason pid interrupt_traps=0 timeout=$fm_herdr_lab_viewer_timeout_seconds + fm_herdr_lab_validate_name "$name" || return 1 + command -v herdr >/dev/null 2>&1 || { fm_herdr_lab_error "herdr is required"; return 1; } + command -v jq >/dev/null 2>&1 || { fm_herdr_lab_error "jq is required"; return 1; } + command -v python3 >/dev/null 2>&1 || { fm_herdr_lab_error "python3 is required for the lab viewer"; return 1; } + + [ -f "$(fm_herdr_lab_tripwire_path "$name")" ] || { + fm_herdr_lab_error "missing fleet-state tripwire for '$name'; refusing to attach a viewer to a session this lab does not own" + return 1 + } + fm_herdr_lab_refuse_if_default "$name" || return 1 + + record=$(fm_herdr_lab_viewer_record_path "$name") + if fm_herdr_lab_viewer_owned_alive "$name"; then + fm_herdr_lab_error "a lab viewer is already attached to '$name'; stop it before starting another" + return 1 + fi + rm -f "$record" + + launcher=$(fm_herdr_lab_viewer_launcher_path) + [ -f "$launcher" ] || { fm_herdr_lab_error "missing viewer launcher at $launcher"; return 1; } + log=$(fm_herdr_lab_viewer_log_path "$name") + mkdir -p "$(fm_herdr_lab_state_dir)" || return 1 + launcher_pid= + if [ "${BASH_SOURCE[0]}" = "$0" ]; then + interrupt_traps=1 + trap 'trap - INT TERM; [ -z "${launcher_pid:-}" ] || fm_herdr_lab_cancel_viewer_launcher "$launcher_pid"; exit 130' INT + trap 'trap - INT TERM; [ -z "${launcher_pid:-}" ] || fm_herdr_lab_cancel_viewer_launcher "$launcher_pid"; exit 143' TERM + fi + nohup python3 "$launcher" "$name" "$record" >"$log" 2>&1 & + launcher_pid=$! + + waited=0 + attempt=$((timeout * 5)) + while [ "$waited" -lt "$attempt" ]; do + reason=$(fm_herdr_lab_viewer_reason "$name") || reason= + if [ "$reason" = cleared ]; then + pid=$(fm_herdr_lab_viewer_owned_pid "$name" viewer) || pid= + if [ -n "$pid" ]; then + [ "$interrupt_traps" = 0 ] || trap - INT TERM + disown "$launcher_pid" 2>/dev/null || true + printf 'viewer attached to %s (pid %s)\n' "$name" "$pid" + return 0 + fi + fi + sleep 0.2 + waited=$((waited + 1)) + done + fm_herdr_lab_cancel_viewer_launcher "$launcher_pid" + [ "$interrupt_traps" = 0 ] || trap - INT TERM + fm_herdr_lab_error "lab viewer did not become the foreground client of '$name' within $timeout seconds (last reason: ${reason:-<unreadable>})" + [ ! -s "$log" ] || fm_herdr_lab_error "viewer log: $(tail -n 5 "$log" | tr '\n' ' ')" + fm_herdr_lab_viewer_stop "$name" >/dev/null 2>&1 || true + return 1 +} + +fm_herdr_lab_viewer_stop() { # <session> + local name=$1 record log role waited attempt reason timeout=$fm_herdr_lab_viewer_timeout_seconds + fm_herdr_lab_validate_name "$name" || return 1 + record=$(fm_herdr_lab_viewer_record_path "$name") + log=$(fm_herdr_lab_viewer_log_path "$name") + # An absent record means this lab owns no viewer. Any client attached in that + # case belongs to someone else and must never be signalled from here. + [ -f "$record" ] || return 0 + + for role in viewer launcher; do + fm_herdr_lab_viewer_signal "$name" "$role" TERM + done + waited=0 + while fm_herdr_lab_viewer_owned_alive "$name" && [ "$waited" -lt 50 ]; do + sleep 0.1 + waited=$((waited + 1)) + done + for role in viewer launcher; do + fm_herdr_lab_viewer_signal "$name" "$role" KILL + done + + waited=0 + attempt=$((timeout * 5)) + while [ "$waited" -lt "$attempt" ]; do + reason=$(fm_herdr_lab_viewer_reason "$name") || reason= + if [ "$reason" = no_foreground_client ] \ + || { [ -z "$reason" ] && fm_herdr_lab_viewer_session_stopped_or_absent "$name"; }; then + rm -f "$record" "$log" + return 0 + fi + sleep 0.2 + waited=$((waited + 1)) + done + fm_herdr_lab_error "lab viewer for '$name' did not detach within $timeout seconds (last reason: ${reason:-<unreadable>})" + return 1 +} + +fm_herdr_lab_viewer() { # <start|stop> <session> + case "${1:-}" in + start) fm_herdr_lab_viewer_start "$2" ;; + stop) fm_herdr_lab_viewer_stop "$2" ;; + *) + fm_herdr_lab_error "viewer takes 'start' or 'stop'" + return 2 + ;; + esac +} + +fm_herdr_lab_cancel_viewer_launcher() { # <pid> + local pid=$1 attempt=0 max_attempts=$((fm_herdr_lab_viewer_launcher_grace_seconds * 10)) + if kill -0 "$pid" 2>/dev/null; then + kill -TERM "$pid" 2>/dev/null || true + while kill -0 "$pid" 2>/dev/null && [ "$attempt" -lt "$max_attempts" ]; do + sleep 0.1 + attempt=$((attempt + 1)) + done + if kill -0 "$pid" 2>/dev/null; then + kill -KILL "$pid" 2>/dev/null || true + fi + fi + wait "$pid" 2>/dev/null || true +} + fm_herdr_lab_cancel_provision() { # <pid> local pid=$1 attempt=0 if kill -0 "$pid" 2>/dev/null; then @@ -261,6 +492,10 @@ fm_herdr_lab_teardown() { # <session> fm_herdr_lab_error "missing fleet-state tripwire for '$name'; refusing destructive calls" return 1 } + fm_herdr_lab_viewer_stop "$name" || { + fm_herdr_lab_error "refusing teardown of '$name' while this lab's viewer is still attached" + return 1 + } sessions=$(fm_herdr_lab_session_list "$name" 2>/dev/null) || { fm_herdr_lab_error "cannot list Herdr sessions before teardown" return 1 @@ -299,7 +534,7 @@ fm_herdr_lab_name() { # <label> } fm_herdr_lab_usage() { - sed -n '2,13p' "${BASH_SOURCE[0]}" | sed 's/^# \{0,1\}//' + sed -n '2,15p' "${BASH_SOURCE[0]}" | sed 's/^# \{0,1\}//' } fm_herdr_lab_main() { @@ -322,6 +557,10 @@ fm_herdr_lab_main() { shift fm_herdr_lab_cli "$@" ;; + viewer) + [ "$#" -eq 3 ] || { fm_herdr_lab_usage >&2; return 2; } + fm_herdr_lab_viewer "$2" "$3" + ;; stop) [ "$#" -eq 2 ] || { fm_herdr_lab_usage >&2; return 2; } fm_herdr_lab_stop "$2" diff --git a/bin/fm-inactive-reconcile.sh b/bin/fm-inactive-reconcile.sh index 9c30a9074be..5cf22755626 100755 --- a/bin/fm-inactive-reconcile.sh +++ b/bin/fm-inactive-reconcile.sh @@ -311,15 +311,19 @@ meta_incarnation() { # <meta> printf 'legacy-%s\n' "$(sha256_text "$identity")" } -pr_for_task() { # <meta> <status> [preferred-line] - local meta=$1 status=$2 preferred=${3:-} value +# The task's delivered PR. Recorded meta pr= is the only authoritative source; +# the fallback scrape accepts only a preferred terminal line in a mode's +# ready-signal shape (`done: PR <url>` or `done: PR <url> checks green`), so a +# PR a worker merely mentioned in prose is never claimed as the delivery. +# A scout never delivers a PR, so it never carries one. +pr_for_task() { # <meta> [preferred-line] + local meta=$1 preferred=${2:-} value + [ "$(meta_field "$meta" kind)" != scout ] || return 0 value=$(meta_field "$meta" pr) if [ -z "$value" ] && [ -n "$preferred" ]; then value=$(printf '%s\n' "$preferred" \ - | grep -Eo 'https?://[^[:space:])"]+/pull/[0-9]+' | head -1 || true) - fi - if [ -z "$value" ] && [ -f "$status" ]; then - value=$(grep -Eo 'https?://[^[:space:])"]+/pull/[0-9]+' "$status" 2>/dev/null | tail -1 || true) + | sed -nE 's|^done: PR (https?://[^[:space:])"]+/pull/[0-9]+)( checks green)?$|\1|p' \ + | head -1 || true) fi clean_field "$value" } @@ -398,7 +402,7 @@ report_child_ledger_locked() { # <id> <meta> status="$STATE/$id.status" last=$(child_terminal_ledger_line "$status") || return 0 state=$(status_line_verb "$last") - pr=$(pr_for_task "$meta" "$status" "$last") + pr=$(pr_for_task "$meta" "$last") incarnation=$(meta_incarnation "$meta") fingerprint=$(sha256_text "$incarnation|$id|$state|ledger|$last") previous=$(grep -v '^[[:space:]]*$' "$status" 2>/dev/null \ @@ -496,7 +500,7 @@ reconcile_direct_child_locked() { # <id> <meta> <secondmate-id-or-empty> <timeou 'state: failed '*) state='failed' ;; *) return 0 ;; esac - pr=$(pr_for_task "$meta" "$status") + pr=$(pr_for_task "$meta") incarnation=$(meta_incarnation "$meta") fingerprint=$(sha256_text "$incarnation|$id|$state|$pr|$(clean_field "$last")") if [ -n "$self" ]; then diff --git a/bin/fm-launch-lib.sh b/bin/fm-launch-lib.sh index 95f8afd96f6..3af72916fb9 100644 --- a/bin/fm-launch-lib.sh +++ b/bin/fm-launch-lib.sh @@ -113,6 +113,8 @@ fm_launch_render() { # <template> <model-flag> <effort-flag> <brief> <turnend> __PITURNEND__) out=$out$pi_turnend ;; __PIWATCH__) out=$out$pi_watch ;; __OPINPUT__) out=$out$op_input ;; + __AGYBIN__) out=$out'agy' ;; + __CLAUDEPERMFLAG__) out=$out'--dangerously-skip-permissions' ;; __PIBIN__|__PITUIMODE__|__CURSORBIN__|__WORKTREE__) out=$out$token ;; *) if [ "$allow_unresolved" = 1 ]; then @@ -276,7 +278,7 @@ fm_launch_template() { # Carry attribution-off with the per-launch settings because worker settings # sources may omit the user's scope. Keep both feedback controls alongside # it so managed settings cannot re-enable the model-drafted feedback tool. - claude) printf '%s' 'CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --dangerously-skip-permissions --settings '\''{"feedbackDrafts":"off","attribution":{"commit":"","pr":"","sessionUrl":false}}'\'' __MODELFLAG____EFFORTFLAG__"$(__OPINPUT__ encode launch-brief < __BRIEF__)"' ;; + claude) printf '%s' 'CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude __CLAUDEPERMFLAG__ --settings '\''{"feedbackDrafts":"off","attribution":{"commit":"","pr":"","sessionUrl":false}}'\'' __MODELFLAG____EFFORTFLAG__"$(__OPINPUT__ encode launch-brief < __BRIEF__)"' ;; codex) if [ "$kind" = secondmate ]; then printf '%s' 'codex __MODELFLAG____EFFORTFLAG__--dangerously-bypass-approvals-and-sandbox "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' @@ -338,7 +340,7 @@ fm_launch_template() { # before launch (bin/fm-agy-trust-lib.sh). --effort accepts only low|medium|high # (agy --help). Turn-end notification is the watcher's debounced native-idle detector, # so no launch-time hook is installed. - agy) printf '%s' 'agy --dangerously-skip-permissions __MODELFLAG____EFFORTFLAG__--prompt-interactive "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' ;; + agy) printf '%s' 'env -u CLAUDECODE -u PI_CODING_AGENT -u GROK_AGENT -u FM_PI_HARNESS __AGYBIN__ --dangerously-skip-permissions __MODELFLAG____EFFORTFLAG__--prompt-interactive "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' ;; # omp (Oh My Pi), a Pi fork. Same one-positional-brief, --model, --thinking, # and -e shape as Pi, verified on omp 18.1.11. The differences are all at # the launch boundary and documented in the header above: foreign markers diff --git a/bin/fm-merge-authority-lib.sh b/bin/fm-merge-authority-lib.sh new file mode 100755 index 00000000000..9dbbadda2b1 --- /dev/null +++ b/bin/fm-merge-authority-lib.sh @@ -0,0 +1,201 @@ +#!/usr/bin/env bash +# Durable ownership of the authority under which a task's merge was accepted. +# +# The away-posture record (state/.afk-contract) and the task's recorded yolo +# posture are resolved only at the merge gate. After a forge accepts the merge, +# bin/fm-pr-merge.sh persists that answer as: +# state/<task-id>.merge-authority +# fm-merge-authority-v1 +# <provider> +# <host> +# <path> +# <number> +# <authority> yolo | away-grant | attended +# The identity comes from the merge run's immutable canonical URL parse; +# persistence revalidates the task's current pr= metadata under its metadata +# and lifecycle locks and refuses a mismatch. The file is atomically published, +# mode 0600, single-link, and on the state filesystem. A poll consumes it only +# when all identity fields match its own validated snapshot. Missing, malformed, +# or mismatched state means external; it is never resolved again from a later +# away-posture record. +# +# Resolution authorizes nothing by itself. bin/fm-pr-merge.sh owns the merge +# gate and persists only after a forge command succeeds, before releasing the +# task lifecycle lock. After observing a landed merge, bin/fm-watch.sh acquires +# that same lock, revalidates the poll, publishes its durable outcome, and +# retires only the exact authority record it read. Teardown uses the same lock, +# so it cannot interleave with that consumption transaction, and removes any +# remaining record. +# +# Sourced by those scripts and by tests. No side effects on source beyond its +# sourced libraries. + +_FM_MERGE_AUTHORITY_LIB_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +# shellcheck source=bin/fm-pr-lib.sh +. "$_FM_MERGE_AUTHORITY_LIB_DIR/fm-pr-lib.sh" +# shellcheck source=bin/fm-afk-contract.sh +. "$_FM_MERGE_AUTHORITY_LIB_DIR/fm-afk-contract.sh" + +# shellcheck disable=SC2034 # Public results consumed by sourcing callers. +FM_MERGE_AUTHORITY= +# shellcheck disable=SC2034 # Public results consumed by sourcing callers. +FM_MERGE_AUTHORITY_REASON= +# shellcheck disable=SC2034 # Public results consumed by sourcing callers. +FM_MERGE_AUTHORITY_RECORD_IDENTITY= + +fm_merge_authority_resolve() { # <home> <state> <meta> <task-id> + local home=${1-} state=${2-} meta=${3-} id=${4-} + local yolo='' grants grant + FM_MERGE_AUTHORITY= + FM_MERGE_AUTHORITY_REASON='invalid' + [ -n "$home" ] && [ -n "$state" ] && [ -n "$meta" ] && [ -n "$id" ] || return 1 + + if ! fm_afk_contract_present "$state"; then + FM_MERGE_AUTHORITY='attended' + FM_MERGE_AUTHORITY_REASON='attended' + return 0 + fi + if ! FM_HOME="$home" FM_STATE_OVERRIDE="$state" \ + "$_FM_MERGE_AUTHORITY_LIB_DIR/fm-afk-contract.sh" validate >/dev/null 2>&1; then + FM_MERGE_AUTHORITY_REASON='record-unreadable' + return 1 + fi + if [ -f "$meta" ]; then + yolo=$(grep '^yolo=' "$meta" | tail -1 | cut -d= -f2- || true) + fi + if [ "$yolo" = on ]; then + FM_MERGE_AUTHORITY='yolo' + FM_MERGE_AUTHORITY_REASON='granted' + return 0 + fi + grants=$(FM_HOME="$home" FM_STATE_OVERRIDE="$state" \ + "$_FM_MERGE_AUTHORITY_LIB_DIR/fm-afk-contract.sh" grants 2>/dev/null) || { + FM_MERGE_AUTHORITY_REASON='grants-unreadable' + return 1 + } + while IFS= read -r grant; do + [ "$grant" = "$id" ] || continue + FM_MERGE_AUTHORITY='away-grant' + FM_MERGE_AUTHORITY_REASON='granted' + return 0 + done <<EOF +$grants +EOF + # shellcheck disable=SC2034 # Public results consumed by sourcing callers. + FM_MERGE_AUTHORITY_REASON='not-granted' + return 1 +} + +fm_merge_authority_record_matches() { # <record> <device> <provider> <host> <path> <number> + local record=$1 device=$2 expected_provider=$3 expected_host=$4 expected_path=$5 expected_number=$6 + local version provider host path number authority + fm_pr_private_file_valid "$record" 600 "$device" || return 1 + exec 8< "$record" || return 1 + IFS= read -r version <&8 || { exec 8<&-; return 1; } + IFS= read -r provider <&8 || { exec 8<&-; return 1; } + IFS= read -r host <&8 || { exec 8<&-; return 1; } + IFS= read -r path <&8 || { exec 8<&-; return 1; } + IFS= read -r number <&8 || { exec 8<&-; return 1; } + IFS= read -r authority <&8 || { exec 8<&-; return 1; } + if IFS= read -r _extra <&8; then + exec 8<&- + return 1 + fi + exec 8<&- + case "$authority" in yolo|away-grant|attended) ;; *) return 1 ;; esac + [ "$version" = fm-merge-authority-v1 ] \ + && [ "$provider" = "$expected_provider" ] \ + && [ "$host" = "$expected_host" ] \ + && [ "$path" = "$expected_path" ] \ + && [ "$number" = "$expected_number" ] || return 1 + FM_MERGE_AUTHORITY=$authority +} + +fm_merge_authority_persist() { # <state> <task-id> <meta> <provider> <host> <path> <number> <authority> + local state=$1 id=$2 meta=$3 provider=$4 host=$5 path=$6 number=$7 authority=$8 + local record tmp='' state_device lock status=0 + fm_pr_task_id_valid "$id" || return 1 + case "$authority" in yolo|away-grant|attended) ;; *) return 1 ;; esac + [ -d "$state" ] && [ ! -L "$state" ] || return 1 + state_device=$(fm_pr_file_device "$state") || return 1 + fm_pr_metadata_identity_parse "$meta" || return 1 + [ "$FM_PR_META_PROVIDER" = "$provider" ] \ + && [ "$FM_PR_META_HOST" = "$host" ] \ + && [ "$FM_PR_META_PATH" = "$path" ] \ + && [ "$FM_PR_META_NUMBER" = "$number" ] || return 1 + record="$state/$id.merge-authority" + lock="$record.lock" + fm_lock_acquire_wait "$lock" || return 1 + fm_pr_regular_destination_on_device_or_absent "$record" "$state_device" || status=1 + if [ "$status" -eq 0 ]; then + umask 077 + tmp=$(mktemp "$state/.fm-merge-authority.XXXXXX") || status=1 + fi + if [ "$status" -eq 0 ]; then + printf '%s\n%s\n%s\n%s\n%s\n%s\n' \ + fm-merge-authority-v1 "$provider" "$host" "$path" "$number" "$authority" > "$tmp" \ + || status=1 + fi + if [ "$status" -eq 0 ]; then + chmod 0600 "$tmp" \ + && fm_merge_authority_record_matches "$tmp" "$state_device" \ + "$provider" "$host" "$path" "$number" \ + && fm_pr_regular_destination_on_device_or_absent "$record" "$state_device" \ + && mv -f -- "$tmp" "$record" \ + && fm_merge_authority_record_matches "$record" "$state_device" \ + "$provider" "$host" "$path" "$number" \ + || status=1 + fi + [ "$status" -eq 0 ] || rm -f -- "$tmp" + fm_lock_release "$lock" || status=1 + return "$status" +} + +fm_merge_authority_read() { # <state> <task-id> <provider> <host> <path> <number> + local state=$1 id=$2 provider=$3 host=$4 path=$5 number=$6 + local record state_device lock status=0 + FM_MERGE_AUTHORITY='external' + FM_MERGE_AUTHORITY_RECORD_IDENTITY= + fm_pr_task_id_valid "$id" || return 1 + [ -d "$state" ] && [ ! -L "$state" ] || return 1 + state_device=$(fm_pr_file_device "$state") || return 1 + record="$state/$id.merge-authority" + lock="$record.lock" + fm_lock_acquire_wait "$lock" || return 1 + if fm_merge_authority_record_matches "$record" "$state_device" \ + "$provider" "$host" "$path" "$number"; then + # shellcheck disable=SC2034 # Public results consumed by sourcing callers. + FM_MERGE_AUTHORITY_RECORD_IDENTITY=$(fm_pr_file_identity "$record") || status=1 + else + FM_MERGE_AUTHORITY='external' + status=1 + fi + fm_lock_release "$lock" || status=1 + return "$status" +} + +fm_merge_authority_remove_if_matches() { # <state> <task-id> <provider> <host> <path> <number> <authority> <file-identity> + local state=$1 id=$2 provider=$3 host=$4 path=$5 number=$6 + local authority=$7 expected_file_identity=$8 record state_device lock current_file_identity status=0 + fm_pr_task_id_valid "$id" || return 1 + [ -d "$state" ] && [ ! -L "$state" ] || return 1 + state_device=$(fm_pr_file_device "$state") || return 1 + record="$state/$id.merge-authority" + lock="$record.lock" + fm_lock_acquire_wait "$lock" || return 1 + if [ -e "$record" ] || [ -L "$record" ]; then + if fm_merge_authority_record_matches "$record" "$state_device" \ + "$provider" "$host" "$path" "$number"; then + current_file_identity=$(fm_pr_file_identity "$record") || status=1 + if [ "$status" -eq 0 ] \ + && [ "$FM_MERGE_AUTHORITY" = "$authority" ] \ + && [ "$current_file_identity" = "$expected_file_identity" ]; then + rm -f -- "$record" || status=1 + fi + elif ! fm_pr_private_file_valid "$record" 600 "$state_device"; then + status=1 + fi + fi + fm_lock_release "$lock" || status=1 + return "$status" +} diff --git a/bin/fm-merge-outcome-lib.sh b/bin/fm-merge-outcome-lib.sh index ab0b96a6778..db279351145 100755 --- a/bin/fm-merge-outcome-lib.sh +++ b/bin/fm-merge-outcome-lib.sh @@ -35,27 +35,38 @@ _FM_MERGE_OUTCOME_LIB_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" # shellcheck disable=SC2034 # Public result consumed by sourcing callers. FM_MERGE_OUTCOME_ALREADY_RECORDED=false -# fm_merge_outcome_report <home> <state> <task-id> <pr-url> <origin> +# fm_merge_outcome_report <home> <state> <task-id> <pr-url> <origin> [authority] # # <origin> says who observed the merge, because that decides whether the # existing poll path also needs a local wake: # self - this home performed the merge. # poll - this home's merge poll detected the merge, so the canonical outcome # also wakes this home after any upward hop needed by a secondmate. +# Optional <authority> is yolo, away-grant, attended, or external. Yolo, +# away-grant, and external are appended to the ledger line; attended remains +# untagged. The merge entrypoint supplies its authority after forge acceptance, +# while the poll supplies the persisted identity-bound value or external when +# no matching record proves that this home authorized the merge. # # Returns 0 when the outcome is recorded (or already was), 2 on an invalid # request, 3 when this home's own role or parent binding cannot be read well # enough to say where the outcome belongs, and 1 on any other failure to # record. A caller that has already merged must report a non-zero return rather # than treat it as success: the merge landed and the record did not. -fm_merge_outcome_report() { # <home> <state> <task-id> <pr-url> <origin> +fm_merge_outcome_report() { # <home> <state> <task-id> <pr-url> <origin> [authority] local home=$1 state=$2 id=$3 url=$4 origin=$5 + local authority=${6-} suffix= local self_rc=0 destination='' line lock status=0 local provider host path number # shellcheck disable=SC2034 # Sourced wake helpers consume these scoped globals. local STATE FM_WAKE_QUEUE FM_WAKE_QUEUE_LOCK FM_MERGE_OUTCOME_ALREADY_RECORDED=false case "$origin" in self|poll) ;; *) return 2 ;; esac + case "$authority" in + yolo|away-grant|external) suffix=" $authority" ;; + attended|'') ;; + *) return 2 ;; + esac fm_pr_task_id_valid "$id" || return 2 fm_pr_url_parse "$url" || return 2 provider=$FM_PR_PROVIDER @@ -65,7 +76,7 @@ fm_merge_outcome_report() { # <home> <state> <task-id> <pr-url> <origin> [ -d "$state" ] && [ ! -L "$state" ] || return 1 if destination=$(fm_parent_channel_destination "$home" "$state"); then - line="done [key=merged-$id]: merged $id $FM_PR_URL" + line="done [key=merged-$id]: merged $id $FM_PR_URL$suffix" else self_rc=$? [ "$self_rc" -eq 1 ] || return 3 @@ -90,7 +101,7 @@ fm_merge_outcome_report() { # <home> <state> <task-id> <pr-url> <origin> fi if [ "$status" -eq 0 ] && { [ "$origin" = poll ] || [ -z "$destination" ]; }; then fm_wake_append check "merged-$id-$FM_PR_URL" \ - "check: merge landed: $id $FM_PR_URL" || status=1 + "check: merge landed: $id $FM_PR_URL$suffix" || status=1 fi if [ "$status" -eq 0 ]; then fm_pr_poll_merge_mark_notified "$state" "$id" \ diff --git a/bin/fm-pr-merge.sh b/bin/fm-pr-merge.sh index 3cf2cac6fad..13fee57563d 100755 --- a/bin/fm-pr-merge.sh +++ b/bin/fm-pr-merge.sh @@ -2,41 +2,46 @@ # Merge a task's PR or MR after recording pr= and any available pr_head= through # bin/fm-pr-check.sh, so teardown can verify landed work after squash merges. # The full canonical URL is parsed by bin/fm-pr-lib.sh. A GitHub pull request is -# addressed through gh-axi by the derived owner and repository; a GitLab merge +# addressed through gh by the derived owner and repository; a GitLab merge # request is addressed through glab by the project URL rebuilt from the parsed # host and path, so any instance works and no host is hardcoded. # # Merge method on GitHub defaults to --squash when the caller passes none of # --squash, --merge, --rebase, or --method after the optional -- separator. -# The gh-axi merge abstraction always performs the merge; the outcome read that -# follows it never becomes a prerequisite for reaching that abstraction. After -# gh-axi returns success, GitHub's live state is read back and accepted only -# when the pull request is merged or in the merge queue. gh's GraphQL API -# supplies that queue-aware read when gh is on PATH; when gh is absent or its -# read fails, gh-axi's own view still proves a landed merge, and every outcome -# it cannot prove refuses, reporting the single failed read when gh is absent -# and naming both failed reads when gh is present and its own read failed; -# neither degraded route accepts an outcome the gh-axi view cannot prove, so the -# same evidence yields the same verdict whether gh is absent or merely broken. +# A GitHub merge is refused unless every pre-merge condition holds, each read +# live at merge time rather than taken from recorded metadata: the pull request +# is open, not a draft, mergeable, free of conflicts, and every unwaived check +# is green at the exact current head commit, where github_checks_not_green below +# owns what makes a check green and judges each one by its current run. +# Every failing condition is reported, not +# just the first. The verified head is then passed to gh as +# --match-head-commit, so a push that lands between that read and the merge +# fails the merge instead of landing commits nothing verified. Reading that +# state needs gh and jq, and either one absent stops the merge before any +# state is recorded. github_verified_upstream_sync below owns the narrow +# reviewed upstream-sync policy exception. An attended --allow-red <check-name> may be passed once, +# with the name as a separate argument; it waives only checks with that exact +# name, still requires every other check green, and still binds the head. It is +# refused while the away-posture record exists, and it never +# applies on GitLab, where a merge already requires the head pipeline to have +# succeeded. After gh returns success, GitHub's live state is read back and +# accepted only when the pull request is merged or in the merge queue. gh's +# GraphQL API supplies that queue-aware read; when that read fails, gh-axi's +# own view still proves a landed merge, and every outcome it cannot prove +# refuses, reporting the failed gh read and naming both failed reads when the +# gh-axi view could not prove the outcome either. # If the pull request remains open and the base branch has an effective -# merge_queue rule, the refusal names the queue's configured merge method and -# the exact -- --auto --<method> retry flags, unless the caller already passed -# that method with --auto to a merge command that returned success, in which -# case it reports instead that the accepted request has not entered the queue -# and the queue state has to be re-checked. Reporting a queued request as -# upstream reports it is a captain decision of 2026-09-03 that RETIRED this -# fork's earlier refusal of deferred execution; see docs/fork-divergence.md. +# merge_queue rule, an attended refusal names the queue's configured merge +# method and exact --attended-override -- --auto --<method> retry flags. While +# the away-posture record exists, asynchronous merge requests are refused and +# queue retry flags are not offered because they would outlive away authority. +# An attended caller that already passed the configured method with --auto is +# told instead that the accepted request has not entered the queue and its queue +# state has to be re-checked. # No method is selected for the caller in any case. A rules response that names # no queue rule, one that could not be read, rules that disagree, and a method # this script does not recognise are four distinct outcomes and are reported # apart, because each one leaves the operator somewhere different. -# A caller-requested --auto that leaves the pull request neither merged nor -# queued is refused the same way and says auto-merge was armed with nothing -# landed or queued yet, or, when the merge command itself failed, that auto-merge -# was only requested; both are read from the caller's own arguments rather than -# from the forge's prose. The observed state is judged the same way whichever -# read produced it, and a refusal built on the gh-axi view says the merge queue -# could not be observed at all rather than implying an unqueued pull request. # Every refusal that follows a merge command which returned success quotes that # command's own output, marked as the forge's text and kept apart from this # script's verdict, including the refusal for an outcome that cannot be read; @@ -60,8 +65,8 @@ # setting, which the merge API applies, and imposing squash there would override # that convention rather than mirror the GitHub default. # Extra GitHub args are an ALLOW-LIST, not an unchanged passthrough: merge-method -# selectors, message arguments, post-execution branch cleanup, and head-binding -# arguments are admitted with a recorded reason each, and everything else is +# selectors, message arguments, and post-execution branch cleanup +# are admitted with a recorded reason each, and everything else is # refused by name. See assert_merge_args_allowed below. # # A GitLab merge is refused unless every pre-merge condition holds, each read @@ -79,13 +84,36 @@ # Before either forge merge, the task's existing per-task control lock # serializes the captain-hold check through the forge command. A still-held or # unreadable row refuses before that command, so a captain approval must be -# recorded as an `answer --release` before this entrypoint is invoked. The lock -# ends when the local forge command returns; docs/captain-hold-lifecycle.md owns -# the accepted asynchronous-landing and merge-to-cleanup residuals. +# recorded as an `answer --release` before this entrypoint is invoked. While +# state/.afk-contract exists, a merge for this task also proceeds only if its +# meta yolo=on or its id is in that record's merge-grant list; otherwise it is +# held for the captain return. An unreadable record refuses rather than being +# skipped. Neither posture releases a captain hold, and the grant lapses when +# the record is archived. +# The authority read and synchronous forge command share the away record's +# cross-subsystem lock, which bin/fm-afk-contract.sh owns, closing the common +# live-owner TOCTOU; failure to take it refuses before the forge call. Async and +# queued paths are refused while away. Two confused-agent-grade limitations are +# accepted rather than hidden: queue or base changes after GitHub's preflight can +# still enqueue, and killing this shell can orphan a forge child after stale-lock +# recovery. docs/architecture.md owns those away-merge limits, while +# docs/captain-hold-lifecycle.md owns the separate merge-to-cleanup residual. +# A failed forge command releases the lock after it returns. A successful one +# retains the lock until the accepted merge authority is persisted against the +# still-matching task metadata. # # Extra args must not include --repo or -R in any form, including a bundled # short-option cluster such as -yR, because the repository comes only from the -# URL, nor --sha on GitLab because the head comes only from the live read. +# URL, nor --sha or --match-head-commit because the head comes only from the +# live read. An existing task-meta pr= must equal the requested canonical URL; +# a task cannot be rebound here. Auto-merge (--auto), a protection bypass +# (--admin), and branch +# deletion (--delete-branch, -d, and GitLab's --remove-source-branch) are +# refused by default. --attended-override permits supported cleanup flags for +# an explicit instruction, but the fork allowlist always refuses --auto and +# --admin. It never skips live checks, away authority, or a task hold. +# +# Usage: fm-pr-merge.sh <task-id> <pr-url> [--attended-override] [--allow-red <check-name>] [-- <extra forge merge args>] # After a successful merge, an optional work item recorded in task metadata is # verified and, when it is open on a forge with a write adapter, closed with a # comment linking the merged PR. A work_item= record names the tracker the @@ -105,7 +133,6 @@ # (report_landed_after_failed_command), because a run that exits zero in silence # straight after the forge CLI printed its own error reads as an unexplained # success. -# Usage: fm-pr-merge.sh <task-id> <pr-url> [-- <extra forge merge args>] set -eu SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -128,6 +155,10 @@ esac . "$SCRIPT_DIR/fm-issue-lib.sh" # shellcheck source=bin/fm-forge-lib.sh . "$SCRIPT_DIR/fm-forge-lib.sh" +# shellcheck source=bin/fm-merge-authority-lib.sh +. "$SCRIPT_DIR/fm-merge-authority-lib.sh" +# shellcheck source=bin/fm-afk-contract.sh +. "$SCRIPT_DIR/fm-afk-contract.sh" if [ "$#" -lt 2 ]; then echo "error: invalid PR merge request" >&2 @@ -141,6 +172,8 @@ if ! fm_pr_task_id_valid "$ID" || ! fm_pr_url_parse "$RAW_URL"; then fi URL=$FM_PR_URL PROVIDER=$FM_PR_PROVIDER +PR_HOST=$FM_PR_HOST +PR_PATH=$FM_PR_PATH PR_OWNER=$FM_PR_OWNER PR_REPO=$FM_PR_REPO PR_NUMBER=$FM_PR_NUMBER @@ -148,7 +181,36 @@ PR_NUMBER=$FM_PR_NUMBER # rebuilt from the parsed identity rather than read from any ambient default. PROJECT_URL="https://$FM_PR_HOST/$FM_PR_PATH" shift 2 -[ "${1:-}" = "--" ] && shift +ATTENDED_OVERRIDE=false +ALLOW_RED=() +while [ "$#" -gt 0 ]; do + case "$1" in + --attended-override) + ATTENDED_OVERRIDE=true + shift + ;; + --attended-override=*) + echo "error: --attended-override takes no value" >&2 + exit 2 + ;; + --allow-red) + [ -n "${2:-}" ] || { echo "error: --allow-red requires a check name" >&2; exit 2; } + [ "${#ALLOW_RED[@]}" -eq 0 ] || { echo "error: --allow-red may be specified only once" >&2; exit 2; } + ALLOW_RED+=("$2") + shift 2 + ;; + --allow-red=*) + echo "error: --allow-red requires a separate check name argument" >&2 + exit 2 + ;; + --) shift; break ;; + *) break ;; + esac +done +if [ "${#ALLOW_RED[@]}" -gt 0 ] && [ "$PROVIDER" = gitlab ]; then + echo "error: --allow-red does not apply to GitLab, where a merge already requires the head pipeline to have succeeded" >&2 + exit 2 +fi # Forwarded arguments are an ALLOW-LIST, not a denylist. An unbounded passthrough # cannot be closed: each refused flag only reveals the next one. Every admitted @@ -157,89 +219,14 @@ shift 2 # it cannot defer execution. A flag whose safety justification cannot be written # honestly does not belong here. # -# WHICH ENTRIES ARE INERT ON GITHUB TODAY, disclosed for the same reason the -# --sha entry discloses it: an admitted flag must not read as a working control -# when the tool behind it has no such flag. As OBSERVED FROM `gh-axi pr merge -# --help` AT gh-axi 0.1.34 - a version-observed fact, not a permanent property of -# the tool - gh-axi accepts exactly --method <merge|squash|rebase>, --merge, -# --squash, --rebase, --auto, --delete-branch, --body <text>, --body-file <path> -# and --subject: NO short options at all, and no --remove-source-branch. So on -# GitHub -s, -m, -t, -b, -F, -d, -r and --remove-source-branch are inert exactly -# as --sha is - a forwarded one is REJECTED BY THE TOOL rather than doing -# anything. They stay admitted because the allow-list records what this script -# considers safe to forward, which is a judgement about deferral and not a claim -# about any one CLI's current spelling; the failure mode is a loud command -# rejection, never a silent one. Mirror-image on GitLab: --delete-branch is -# glab's wrong spelling there, where the same cleanup is --remove-source-branch. -# One consequence worth stating where it is, so it is not read as a bug: -# caller_has_merge_method recognises only the long spellings, so `-- -s` still -# gets --squash prepended. That is harmless while the short forms are inert, and -# is the thing to revisit first if gh-axi ever grows them. -# -# --squash --merge -s -m merge-method selectors: choose how the merge -# commit is formed, not when it happens. -# --method --method=* the same selection by name. -# --subject --body -t -b commit message text, applied to the commit the -# --body-file -F merge itself creates. -# -d --delete-branch GitHub branch cleanup, which runs AFTER the -# --remove-source-branch merge has executed and so cannot defer it. -# --sha --sha=* head-binding: constrains the mutation to the -# exact state this run verified. It cannot defer -# execution and makes the merge strictly -# narrower, so it is admitted on principle. -# INERT ON BOTH FORGES TODAY, and the entry says -# so rather than reading as a working control: -# on GitLab reject_head_overrides refuses a -# caller --sha earlier by design, because this -# script supplies its own verified --sha; on -# GitHub `gh-axi pr merge --help` lists its -# supported flags as --method, --merge, --squash, -# --rebase, --auto, --delete-branch, --body, -# --body-file and --subject, with no --sha, so a -# forwarded one would be rejected by the tool -# rather than binding anything. NOTHING ELSE -# BINDS THE HEAD ON GITHUB EITHER: this script -# carries no head check on that forge, and the -# entry says so rather than pointing at a check -# that would have to exist for it to be true. -# The BASE is bound - github_read_default_tip -# refuses a merge whose default-branch tip moved -# - and a base is not a head. Head binding waits -# on the deferred synchronous REST merge -# boundary, which is what could carry one. -# -# Deliberately NOT admitted, each for a stated reason: -# --auto requests deferred execution outright. -# --rebase -r REFUSED ON GITLAB ONLY. There it rewrites the -# source branch BEFORE the merge and leaves it -# rewritten even when the SHA-bound merge then -# fails, which is the mutation the no-auto-rebase -# rule prevents. On GitHub rebase merge replays -# onto the base without touching the source, so -# it is admitted as an ordinary merge-method -# selector. Do not re-collapse these into one -# rule: the rationale is forge-specific. -# --repo -R would retarget the mutation away from the -# identity this run validated. -# Refusing by name rather than silently dropping keeps a caller's mistake loud. -# -# WHY NO TOKEN IS EXEMPT FROM CLASSIFICATION, which is the one rule below that is -# about the TOOL rather than about deferral. This guard used to model the forge -# CLI as a POSITIONAL parser - a value-taking flag consumes the next word, -# whatever that word is - and that model was wrong in both directions before it -# was wrong here. The tool does not parse positionally: it SCANS the whole -# argument list for each flag it knows, so a token that LOOKS like a flag IS a -# flag to it, wherever it stands. OBSERVED AT gh-axi 0.1.34 and recorded as a -# version-observed fact about a third-party tool rather than a permanent -# property: dist/src/commands/pr.js calls takeBoolFlag(args, "--auto") before -# takeFlag(args, "--subject"), and dist/src/args.js's takeFlag returns undefined -# without erroring when its flag is left valueless, so `-- --subject --auto` -# satisfies a positional guard while gh-axi still arms deferred execution; the -# same file's rejectUnknownFlags inspects EVERY dash-leading token, so even one -# the tool would not act on is judged as a flag rather than carried as text. -# Hence a value-taking flag consumes the next word only when that word is not -# itself flag-shaped. A value that legitimately begins with a dash is passed with -# the --flag=<value> spelling, which is one token and cannot be mistaken for one. +# GitHub mutations use gh, so --method is normalized to a native method flag +# after validation. Message flags and supported post-merge cleanup flags retain +# their native meaning. Caller head overrides are rejected separately on both +# providers because this script supplies the verified head itself. +# Every dash-leading token is classified, including one in a value position, +# so a forbidden option cannot hide behind another option's missing value. +# GitLab --rebase rewrites the source before merging and remains refused; +# GitHub's rebase merge does not rewrite the source branch. assert_merge_args_allowed() { local arg # Detached values belong to the flag before them - "--sha abc123" is one @@ -383,7 +370,7 @@ reject_head_overrides() { local arg for arg in "$@"; do case "$arg" in - --sha|--sha=*) + --sha|--sha=*|--match-head-commit|--match-head-commit=*) echo "error: extra merge arguments must not override the head commit" >&2 return 1 ;; @@ -391,8 +378,49 @@ reject_head_overrides() { done } +reject_protected_forge_args() { + local arg + [ "$ATTENDED_OVERRIDE" = true ] && return 0 + for arg in "$@"; do + case "$arg" in + --auto|--auto=*|--admin|--admin=*|--delete-branch|--delete-branch=*|--remove-source-branch|--remove-source-branch=*) + printf 'error: extra merge arguments must not request auto-merge, a protection bypass, or branch deletion (%s); pass --attended-override only for an explicit captain instruction\n' "$arg" >&2 + return 1 + ;; + --*) ;; + # A single-dash argument is a short-option cluster. -d is gh's + # --delete-branch, and -yd carries it the same way -yR carries --repo. + -*d*) + printf 'error: extra merge arguments must not request auto-merge, a protection bypass, or branch deletion (%s); pass --attended-override only for an explicit captain instruction\n' "$arg" >&2 + return 1 + ;; + esac + done +} + reject_repo_overrides "$@" || exit 1 -[ "$PROVIDER" != gitlab ] || reject_head_overrides "$@" || exit 1 +reject_head_overrides "$@" || exit 1 +reject_protected_forge_args "$@" || exit 1 + +FM_PR_GITHUB_AUTO_REQUESTED=false +if [ "$PROVIDER" = github ] && caller_requested_auto_merge "$@"; then + FM_PR_GITHUB_AUTO_REQUESTED=true +fi +FM_PR_GITLAB_ASYNC_REQUESTED=false +if [ "$PROVIDER" = gitlab ]; then + for arg in "$@"; do + case "$arg" in + --auto-merge|--when-pipeline-succeeds) FM_PR_GITLAB_ASYNC_REQUESTED=true ;; + --auto-merge=*|--when-pipeline-succeeds=*) + case "${arg#*=}" in + [tT]|[tT][rR][uU][eE]|1) FM_PR_GITLAB_ASYNC_REQUESTED=true ;; + [fF]|[fF][aA][lL][sS][eE]|0) FM_PR_GITLAB_ASYNC_REQUESTED=false ;; + esac + ;; + esac + done +fi +FM_PR_AWAY_POSTURE=false fm_backlog_directory_present "$STATE" "state directory" || { echo "error: PR merge refused: $FM_BACKLOG_TRANSITION_ERROR" >&2 @@ -421,7 +449,12 @@ fi MERGE_EXPECTED_SPAWN_GEN=$FM_BACKLOG_META_SPAWN_GEN MERGE_CONTROL_LOCK= +MERGE_META_LOCK= +# Called by the EXIT trap. +# shellcheck disable=SC2329 merge_control_cleanup() { + [ -z "$MERGE_META_LOCK" ] || fm_lock_release "$MERGE_META_LOCK" || true + fm_afk_contract_lock_release || true [ -z "$MERGE_CONTROL_LOCK" ] || fm_lock_release "$MERGE_CONTROL_LOCK" || true } trap merge_control_cleanup EXIT @@ -450,6 +483,17 @@ if [ "$PROVIDER" = gitlab ]; then exit 1 fi fi +GITHUB_MISSING= +if [ "$PROVIDER" = github ]; then + command -v gh >/dev/null 2>&1 || GITHUB_MISSING="gh" + if ! command -v jq >/dev/null 2>&1; then + GITHUB_MISSING="${GITHUB_MISSING:+$GITHUB_MISSING and }jq" + fi + if [ -n "$GITHUB_MISSING" ]; then + echo "error: merging a GitHub pull request requires $GITHUB_MISSING on PATH" >&2 + exit 1 + fi +fi # The recorded head is read before bin/fm-pr-check.sh rewrites the metadata, # because that script re-records pr= and drops a pr_head= it cannot resolve. @@ -579,11 +623,12 @@ gitlab_refuse_if_behind() { # <because-clause> return 1 } +FM_PR_GITLAB_ASYNC_CONFIGURED=false gitlab_verify_mergeable() { local json fields line local total=0 named=0 refusals='' local state='' detail='' conflicts='' discussions='' - local live_head='' pipeline_sha='' pipeline_status='' + local live_head='' pipeline_sha='' pipeline_status='' async_configured='' # GITLAB_HOST is set to the same host the project URL already carries, so the # instance is taken from the parsed URL by both signals and never from the @@ -605,7 +650,8 @@ gitlab_verify_mergeable() { "discussions=" + (.blocking_discussions_resolved | tostring), "head=" + ((.sha // "") | tostring), "pipeline_sha=" + ((.head_pipeline.sha // "") | tostring), - "pipeline_status=" + ((.head_pipeline.status // "") | tostring) + "pipeline_status=" + ((.head_pipeline.status // "") | tostring), + "async_configured=" + (if .merge_when_pipeline_succeeds == true or (.merge_after != null) then "true" else "false" end) else error("merge request payload is not an object") end' 2>/dev/null); then @@ -622,6 +668,7 @@ gitlab_verify_mergeable() { head=*) live_head=${line#head=} ;; pipeline_sha=*) pipeline_sha=${line#pipeline_sha=} ;; pipeline_status=*) pipeline_status=${line#pipeline_status=} ;; + async_configured=*) async_configured=${line#async_configured=} ;; *) continue ;; esac named=$((named + 1)) @@ -631,7 +678,7 @@ FIELDS # Every field named exactly once and no unnamed line: a value carrying a # newline would split into a line no name matches, so it is refused here # rather than silently truncated into a value a check could accept. - if [ "$named" -ne 7 ] || [ "$total" -ne 7 ]; then + if [ "$named" -ne 8 ] || [ "$total" -ne 8 ]; then echo "error: could not read the GitLab merge request state before merging" >&2 return 1 fi @@ -674,13 +721,274 @@ FIELDS printf 'verified: %s is open and mergeable, with a successful pipeline at head %s\n' \ "$URL" "$live_head" >&2 FM_PR_MERGE_HEAD=$live_head + FM_PR_GITLAB_ASYNC_CONFIGURED=$async_configured } -# Read one live GitHub pull request view after gh-axi returns. The selected +# Every GitHub check that is not green in the given live pull-request JSON, one +# name per line. An entry is green when it is a status context whose state is +# SUCCESS, or a check run that completed with SUCCESS, NEUTRAL, or SKIPPED (so +# a pending check is not green either). Exits nonzero when the rollup cannot be +# read, so a malformed answer is a failed read and never an empty red set. +# +# The rollup can hold several runs of one check name at the same head, because +# GitHub cancels a pull request's in-flight run when the base branch advances +# and re-triggers it; the cancelled run stays in the rollup beside the passing +# re-run. A check is therefore judged by its current run rather than by any run +# that a later one superseded, which is what makes this agree with GitHub's own +# CLEAN mergeStateStatus instead of refusing a pull request GitHub considers +# mergeable. +# +# Supersession applies only among check runs with the same reported name. A +# name is dropped from the red set only when every non-green run is COMPLETED, +# has a whole-second UTC startedAt, and started strictly before a green run. +# Status contexts are never grouped or superseded, and every non-green one is +# reported independently. A still-running, queued, undated, or tied check run +# stays red. A name whose runs are all green needs no timestamp, while a name +# with no green run stays red. +# +# The reported name is also what --allow-red matches. An unnamed check run is +# grouped alone and can neither supersede nor be superseded, because unrelated +# unnamed checks must not be treated as one. +github_checks_not_green() { + local json=$1 + printf '%s' "$json" | jq -r ' + def settled_at: + if type == "string" and test("^[0-9]{4}-[0-9]{2}-[0-9]{2}T[0-9]{2}:[0-9]{2}:[0-9]{2}Z$") + then . else null end; + if (.statusCheckRollup | type) != "array" then error("no check rollup") else . end + | [ .statusCheckRollup + | to_entries[] + | .key as $i + | .value + | if .__typename == "CheckRun" then + { + kind: "check_run", + name: (.name // ""), + completed: (.status == "COMPLETED"), + ok: (.status == "COMPLETED" and (.conclusion == "SUCCESS" or .conclusion == "NEUTRAL" or .conclusion == "SKIPPED")), + at: (.startedAt | settled_at) + } + | . + {group: (if .name == "" then ["", $i] else [.name, -1] end)} + else + {kind: "status_context", name: (.context // ""), ok: (.state == "SUCCESS")} + end + ] + | . as $entries + | ( + ($entries[] + | select(.kind == "status_context" and (.ok | not)) + | .name + ), + ($entries + | [.[] | select(.kind == "check_run")] + | group_by(.group)[] + | { + name: .[0].name, + reds: [.[] | select(.ok | not)], + newest_green: ([.[] | select(.ok) | .at | select(. != null)] | max) + } + | select( + (.reds | length) > 0 + and ( + .newest_green == null + or any(.reds[]; (.completed | not) or .at == null) + or ([.reds[] | .at] | max) >= .newest_green + ) + ) + | .name + ) + ) + | if . == "" then "(unnamed check)" else . end + ' 2>/dev/null || return 1 +} + +# A review declaration is task metadata, not a general red-check waiver. +# Firstmate records exactly one upstream_sync_review=<base>:<target>:<head> +# after reviewing this round's tests, lint, fork survival and substantive CI. +# These full commit IDs pin the fork base, upstream endpoint and reviewed head. +# Only a direct-PR fm-upstream-sync-* task in HelloWorldSungin/firstmate can +# consume it, and only with --merge. The task worktree must prove the live +# branch/head, fork push target, upstream ancestry and exactly one two-parent +# merge from that base to that endpoint, followed only by linear fix commits. +# The base must still equal the live default tip. No fetch or metadata write +# happens here. Missing, duplicate, stale or unreadable proof leaves checks red. +# This exempts only completed FAILURE CheckRuns named exactly +# "PR must be raised via no-mistakes"; pending/cancelled checks and status +# contexts never qualify. All ordinary authority, hold and merge gates remain. +github_upstream_sync_task() { + [ "$PR_HOST/$PR_PATH" = github.com/HelloWorldSungin/firstmate ] || return 1 + case "$ID" in fm-upstream-sync-*) ;; *) return 1 ;; esac + [ "$(sed -n 's/^mode=//p' "$META")" = direct-PR ] +} + +github_verified_upstream_sync() { + local json=$1 live_head=$2 review wt base target reviewed rest + local origin upstream branch merges merge parents upstream_line + github_upstream_sync_task || return 1 + [ "$FM_PR_GITHUB_CALLER_METHOD" = merge ] || return 1 + [ "${#ALLOW_RED[@]}" -eq 0 ] || return 1 + review=$(sed -n 's/^upstream_sync_review=//p' "$META") + wt=$(sed -n 's/^worktree=//p' "$META") + [ -d "$wt" ] || return 1 + base=${review%%:*}; rest=${review#*:} + target=${rest%%:*}; reviewed=${rest#*:} + fm_pr_head_valid "$base" && fm_pr_head_valid "$target" \ + && fm_pr_head_valid "$reviewed" || return 1 + [ "$review" = "$base:$target:$reviewed" ] || return 1 + [ "$reviewed" = "$live_head" ] && [ "$base" = "$github_judged_default_tip" ] || return 1 + branch=$(git -C "$wt" symbolic-ref --quiet --short HEAD) || return 1 + [ "$branch" = "fm/$ID" ] || return 1 + [ "$(git -C "$wt" rev-parse --verify HEAD)" = "$reviewed" ] || return 1 + origin=$(git -C "$wt" remote get-url --push --all origin) || return 1 + case "$origin" in + https://github.com/HelloWorldSungin/firstmate.git|git@github.com:HelloWorldSungin/firstmate.git) ;; + *) return 1 ;; + esac + upstream=$(git -C "$wt" remote get-url upstream) || return 1 + case "$upstream" in + https://github.com/kunchenguid/firstmate.git|git@github.com:kunchenguid/firstmate.git) ;; + *) return 1 ;; + esac + upstream_line=$(git -C "$wt" rev-list --first-parent refs/remotes/upstream/main) || return 1 + printf '%s\n' "$upstream_line" | grep -qxF "$target" || return 1 + git -C "$wt" merge-base --is-ancestor "$base" "$reviewed" || return 1 + if git -C "$wt" merge-base --is-ancestor "$target" "$base"; then return 1; fi + # First-parent traversal excludes upstream's own merge commits. + merges=$(git -C "$wt" rev-list --first-parent --min-parents=2 "$base..$reviewed") || return 1 + [ -n "$merges" ] || return 1 + merge=$merges + parents=$(git -C "$wt" show -s --format=%P "$merge" 2>/dev/null) || return 1 + [ "$parents" = "$base $target" ] || return 1 + printf '%s' "$json" | jq -e --arg branch "$branch" ' + .headRefName == $branch + and .headRepository.nameWithOwner == "HelloWorldSungin/firstmate" + and ([.statusCheckRollup[] | select(.name == "PR must be raised via no-mistakes")] | length > 0) + and all(.statusCheckRollup[]; + if (.name == "PR must be raised via no-mistakes" or .context == "PR must be raised via no-mistakes") + then .__typename == "CheckRun" and .status == "COMPLETED" + and (.conclusion == "FAILURE" or .conclusion == "SUCCESS" or .conclusion == "NEUTRAL" or .conclusion == "SKIPPED") + else true end) + ' >/dev/null 2>&1 +} + +# Pre-merge conditions for a GitHub pull request, read from one live view. +# Sets FM_PR_MERGE_HEAD to the verified head on success. +github_verify_mergeable() { + local json fields line red name covered sync_policy=false + local total=0 named=0 refusals='' + local state='' draft='' mergeable='' merge_state='' live_head='' base='' + + if github_upstream_sync_task && [ "$FM_PR_GITHUB_CALLER_METHOD" != merge ]; then + echo 'error: upstream-sync tasks require explicit --merge to preserve upstream parentage' >&2 + return 1 + fi + if ! json=$(gh pr view "$URL" --json state,isDraft,mergeable,mergeStateStatus,headRefOid,baseRefName,statusCheckRollup,headRefName,headRepository 2>/dev/null) \ + || [ -z "$json" ]; then + echo "error: could not read the GitHub pull request state before merging" >&2 + return 1 + fi + if ! fields=$(printf '%s' "$json" | jq -r ' + if type == "object" then + "state=" + ((.state // "") | tostring), + "draft=" + (if (.isDraft | type) == "boolean" then (.isDraft | tostring) else "" end), + "mergeable=" + ((.mergeable // "") | tostring), + "merge_state=" + ((.mergeStateStatus // "") | tostring), + "head=" + ((.headRefOid // "") | tostring), + "base=" + ((.baseRefName // "") | tostring) + else + error("pull request payload is not an object") + end' 2>/dev/null); then + echo "error: could not read the GitHub pull request state before merging" >&2 + return 1 + fi + while IFS= read -r line; do + total=$((total + 1)) + case "$line" in + state=*) state=${line#state=} ;; + draft=*) draft=${line#draft=} ;; + mergeable=*) mergeable=${line#mergeable=} ;; + merge_state=*) merge_state=${line#merge_state=} ;; + head=*) live_head=${line#head=} ;; + base=*) base=${line#base=} ;; + *) continue ;; + esac + named=$((named + 1)) + done <<FIELDS +$fields +FIELDS + if [ "$named" -ne 6 ] || [ "$total" -ne 6 ] || [ -z "$base" ]; then + echo "error: could not read the GitHub pull request state before merging" >&2 + return 1 + fi + + if ! fm_pr_head_valid "$live_head"; then + echo "error: could not read the GitHub pull request head commit before merging" >&2 + return 1 + fi + if ! red=$(github_checks_not_green "$json"); then + echo "error: could not read the GitHub pull request state before merging" >&2 + return 1 + fi + + case "$state" in + [oO][pP][eE][nN]) ;; + *) + refusals="$refusals - state is \"${state:-unreadable}\", not open +" + ;; + esac + [ "$draft" = false ] \ + || refusals="$refusals - the pull request is a draft +" + [ "$mergeable" = MERGEABLE ] \ + || refusals="$refusals - mergeable is \"${mergeable:-unreadable}\", not MERGEABLE +" + [ "$merge_state" != DIRTY ] \ + || refusals="$refusals - mergeStateStatus is DIRTY (conflicts) +" + + if github_verified_upstream_sync "$json" "$live_head"; then + sync_policy=true + printf 'verified: reviewed upstream-sync graph at %s permits only the completed no-mistakes policy failure\n' "$live_head" >&2 + fi + uncovered='' + while IFS= read -r name; do + [ -n "$name" ] || continue + covered=0 + if [ "$sync_policy" = true ] && [ "$name" = 'PR must be raised via no-mistakes' ]; then + covered=1 + fi + if [ "${#ALLOW_RED[@]}" -gt 0 ]; then + for check in "${ALLOW_RED[@]}"; do + [ "$check" = "$name" ] && covered=1 + done + fi + [ "$covered" -eq 1 ] || { + refusals="$refusals - check '$name' is not green +" + uncovered="${uncovered:+$uncovered, }$name" + } + done <<EOF +$red +EOF + + if [ -n "$refusals" ]; then + printf 'error: refusing to merge %s\n' "$URL" >&2 + printf '%s' "$refusals" >&2 + [ -z "$uncovered" ] || printf 'error: these checks are not green: %s\n' "$uncovered" >&2 + return 1 + fi + printf 'verified: %s is open and mergeable, with every unwaived check green at head %s\n' \ + "$URL" "$live_head" >&2 + FM_PR_MERGE_HEAD=$live_head + FM_PR_GITHUB_BASE=$base +} + +# Read one live GitHub pull request view after gh returns. The selected # fields distinguish a landed pull request from a merge-queue entry and retain # the concrete state needed for a refusal. gh supplies the complete queue-aware -# view when available; gh-axi remains the degradation path that can prove a -# landed merge without making gh a prerequisite for the merge abstraction. +# view; if that post-merge read becomes unavailable, gh-axi is the degradation +# path that can prove only a landed merge. gh remains a pre-merge prerequisite. FM_PR_GITHUB_STATE= FM_PR_GITHUB_MERGED= FM_PR_GITHUB_QUEUED= @@ -974,7 +1282,102 @@ require_released_captain_hold() { esac } -FM_PR_GITHUB_AUTO_REQUESTED=false +FM_PR_MERGE_AUTHORITY= +# The gate on top of the shared authority read. bin/fm-merge-authority-lib.sh +# owns what the away-posture record and the task's recorded yolo posture say; +# this function owns what a merge run may do about it, so the answer the merge +# poll later tags its ledger row with is the same answer gated here. +require_away_merge_grant() { + FM_PR_MERGE_AUTHORITY= + if fm_merge_authority_resolve "$FM_HOME" "$STATE" "$META" "$ID"; then + FM_PR_MERGE_AUTHORITY=$FM_MERGE_AUTHORITY + return 0 + fi + case "$FM_MERGE_AUTHORITY_REASON" in + record-unreadable) + echo "error: PR merge refused - the away-posture record could not be read; nothing was merged" >&2 + ;; + grants-unreadable) + echo "error: PR merge refused - the away-posture record's grants could not be read; nothing was merged" >&2 + ;; + *) + echo "error: task $ID is held for the captain return" >&2 + ;; + esac + return 1 +} + +# Take the away record's own lock (bin/fm-afk-contract.sh owns it) so that +# record cannot be published, replaced, or archived between the authority read +# below and the forge command that acts on it. Refuses without the lock: a merge +# on authority nothing is holding still is exactly what this closes. This is the +# only path that holds both the per-task control lock and the away-record lock, +# and it always takes them in that order; the away-record side takes only its own +# lock, so the pair cannot deadlock. +hold_away_record_for_merge() { + fm_afk_contract_lock_hold "$STATE" && return 0 + echo "error: PR merge refused - the away-posture record could not be locked for the merge; nothing was merged" >&2 + return 1 +} + +require_current_away_authority() { + FM_PR_AWAY_POSTURE=false + if fm_afk_contract_present "$STATE"; then + FM_PR_AWAY_POSTURE=true + if [ "$PROVIDER" = github ] && [ "$FM_PR_GITHUB_AUTO_REQUESTED" = true ]; then + echo "error: --auto is attended-only; while the away-posture record exists only a synchronous merge may run under its authority lock" >&2 + return 2 + fi + if [ "$PROVIDER" = gitlab ] \ + && { [ "$FM_PR_GITLAB_ASYNC_REQUESTED" = true ] || [ "$FM_PR_GITLAB_ASYNC_CONFIGURED" = true ]; }; then + echo "error: GitLab auto-merge is attended-only; while the away-posture record exists only an immediate merge may run under its authority lock" >&2 + return 2 + fi + fi + require_away_merge_grant || return 1 + if [ "$FM_PR_AWAY_POSTURE" = true ] && [ "${#ALLOW_RED[@]}" -gt 0 ]; then + echo "error: --allow-red is attended-only; while the away-posture record exists the green check is absolute" >&2 + return 2 + fi +} + +persist_accepted_merge_authority() { + local status=0 + MERGE_META_LOCK=$(fm_meta_lock_path "$META") || return 1 + fm_lock_acquire_wait "$MERGE_META_LOCK" || return 1 + fm_merge_authority_persist "$STATE" "$ID" "$META" \ + "$PROVIDER" "$PR_HOST" "$PR_PATH" "$PR_NUMBER" "$FM_PR_MERGE_AUTHORITY" \ + || status=1 + fm_lock_release "$MERGE_META_LOCK" || status=1 + MERGE_META_LOCK= + if [ "$status" -eq 0 ]; then + return 0 + fi + printf 'actionable: the forge accepted the merge request for %s but its merge authority could not be persisted; the merge poll remains armed\n' \ + "$URL" >&2 + return 1 +} + +refuse_github_queue_while_away() { + [ "$FM_PR_AWAY_POSTURE" = true ] || return 0 + # Accepted confused-agent-grade limitation, as in bin/fm-lease-lib.sh, not an + # oversight: a queue rule or PR base change after this preflight can still + # enqueue the merge, which can land after its away grant lapses. + github_read_queue_method + [ "$FM_PR_GITHUB_QUEUE_STATUS" = none ] && return 0 + echo "error: GitHub merge refused while away because the base branch's merge-queue state does not prove an immediate merge; nothing was handed to the forge" >&2 + return 2 +} + +require_recorded_pr_identity() { + local existing + existing=$(grep '^pr=' "$META" | tail -1 | cut -d= -f2- || true) + [ -n "$existing" ] || return 0 + [ "$existing" = "$URL" ] && return 0 + echo "error: task $ID is bound to $existing, not $URL" >&2 + return 1 +} + FM_PR_GITHUB_MERGE_ACCEPTED=false FM_PR_GITHUB_CALLER_METHOD= @@ -1018,6 +1421,10 @@ github_caller_method_is() { github_report_queue_rules() { local queue_method methods_display + if [ "$FM_PR_AWAY_POSTURE" = true ]; then + printf 'error: the direct merge did not land while the away-posture record exists; merge-queue retry flags are unavailable because a queued merge would outlive its authority\n' >&2 + return 0 + fi github_read_queue_method case "$FM_PR_GITHUB_QUEUE_STATUS" in single) @@ -1040,7 +1447,7 @@ github_report_queue_rules() { printf 'error: this run refuses even though the request for %s was accepted with the exact flags base branch %s requires (--auto --%s): the pull request has still not entered the merge queue, so no landed or queued outcome is proven; re-check the pull request'"'"'s merge queue state before retrying\n' \ "$URL" "$FM_PR_GITHUB_BASE" "$queue_method" >&2 else - printf 'error: base branch %s requires the merge queue; retry with: %s %s %s -- --auto --%s\n' \ + printf 'error: base branch %s requires the merge queue; retry with: %s %s %s --attended-override -- --auto --%s\n' \ "$FM_PR_GITHUB_BASE" "$0" "$ID" "$URL" "$queue_method" >&2 fi ;; @@ -1083,8 +1490,12 @@ github_report_unmerged_outcome() { # one. Left whole rather than deleted, so a future upstream merge conflicts on # as little as possible. if [ "$FM_PR_GITHUB_QUEUE_OBSERVED" != true ]; then - printf 'error: the merge queue could not be observed for %s because the queue-aware read was unavailable, so a pull request already in the merge queue cannot be told apart from one that never entered it; re-check the pull request'"'"'s merge queue state before retrying\n' \ - "$URL" >&2 + if [ "$FM_PR_AWAY_POSTURE" = true ]; then + printf 'error: the synchronous merge did not land while the away-posture record exists; no asynchronous merge or queue retry is available under away authority\n' >&2 + else + printf 'error: the merge queue could not be observed for %s because the queue-aware read was unavailable, so a pull request already in the merge queue cannot be told apart from one that never entered it; re-check the pull request'"'"'s merge queue state before retrying\n' \ + "$URL" >&2 + fi return 0 fi github_report_queue_rules @@ -1268,8 +1679,17 @@ if [ "$PROVIDER" = github ]; then fi fi +away_status=0 +require_current_away_authority || away_status=$? +[ "$away_status" -eq 0 ] || exit "$away_status" +require_recorded_pr_identity || exit 1 record_pr_metadata || exit 1 +require_released_captain_hold || exit 1 +# Accepted confused-agent-grade limitation, as in bin/fm-lease-lib.sh, not an +# oversight: if this lock-owning shell dies while its gh or glab child lives, +# stale-owner recovery can release the record for archive or replacement and +# the orphaned forge child can still merge on the lapsed away authority. case "$PROVIDER" in github) merge_output= @@ -1277,10 +1697,31 @@ case "$PROVIDER" in if ! caller_has_merge_method "$@"; then merge_args=(--squash) fi - if caller_requested_auto_merge "$@"; then - FM_PR_GITHUB_AUTO_REQUESTED=true - fi + github_forward_args=() + while [ "$#" -gt 0 ]; do + case "$1" in + --method) github_method=$2; shift ;; + --method=*) github_method=${1#--method=} ;; + *) github_forward_args+=("$1"); shift; continue ;; + esac + case "$github_method" in + merge|squash|rebase) github_forward_args+=("--$github_method") ;; + *) printf 'error: unsupported merge method: %s\n' "$github_method" >&2; exit 1 ;; + esac + shift + done + set -- "${github_forward_args[@]+"${github_forward_args[@]}"}" FM_PR_GITHUB_CALLER_METHOD=$(caller_merge_method "$@") + if [ "$github_landed_observed" != true ]; then + github_verify_mergeable || exit 1 + fi + # The away record is locked first, so this last presence and authority read + # and the forge command below share one live-owner critical section. + hold_away_record_for_merge || exit 1 + away_status=0 + require_current_away_authority || away_status=$? + [ "$away_status" -eq 0 ] || exit "$away_status" + refuse_github_queue_while_away || exit 2 # THE DEFAULT TIP IS RE-READ IMMEDIATELY BEFORE THE MERGE CALL, and the merge # is refused if it moved. Refusing deferred execution is only half of the # immediate-execution guarantee; the other half is that the base was compared @@ -1318,13 +1759,22 @@ case "$PROVIDER" in fi require_released_captain_hold || exit 1 merge_status=0 - merge_output=$(gh-axi pr merge "$PR_NUMBER" --repo "$PR_OWNER/$PR_REPO" \ - "${merge_args[@]+"${merge_args[@]}"}" "$@" 2>&1) || merge_status=$? - merge_control_cleanup - MERGE_CONTROL_LOCK= + merge_output= + if [ "$github_landed_observed" != true ]; then + merge_output=$(gh pr merge "$PR_NUMBER" --repo "$PR_OWNER/$PR_REPO" \ + --match-head-commit "$FM_PR_MERGE_HEAD" \ + "${merge_args[@]+"${merge_args[@]}"}" "$@" 2>&1) || merge_status=$? + fi if [ "$merge_status" -eq 0 ]; then FM_PR_GITHUB_MERGE_ACCEPTED=true + persist_accepted_merge_authority || exit 1 + fm_afk_contract_lock_release || true + fm_lock_release "$MERGE_CONTROL_LOCK" || true + MERGE_CONTROL_LOCK= else + fm_afk_contract_lock_release || true + fm_lock_release "$MERGE_CONTROL_LOCK" || true + MERGE_CONTROL_LOCK= [ -z "$merge_output" ] || printf '%s\n' "$merge_output" >&2 if [ "$github_landed_observed" = true ] || github_read_outcome; then if [ "$FM_PR_GITHUB_MERGED" = true ]; then @@ -1376,20 +1826,25 @@ case "$PROVIDER" in # skips the interactive confirmation, which no supervised run can answer; # the conditions above are what authorize the merge. require_released_captain_hold || exit 1 - # Capture the command status rather than letting set -e abort here: if the - # forge LANDS the merge and a post-execution step or the response transport - # then fails, aborting would skip the confirmation read entirely and leave a - # merge that really happened with nothing recording it. The forge state read - # below is the authority; only a merge confirmed not to have landed fails. + # The away record is locked first, so this last presence and authority read + # and the forge command below share one live-owner critical section. + hold_away_record_for_merge || exit 1 + away_status=0 + require_current_away_authority || away_status=$? + [ "$away_status" -eq 0 ] || exit "$away_status" gitlab_merge_rc=0 + gitlab_merge_args=() + if [ "$FM_PR_AWAY_POSTURE" = true ]; then + gitlab_merge_args=(--auto-merge=false) + fi GITLAB_HOST="$FM_PR_HOST" glab mr merge "$PR_NUMBER" -R "$PROJECT_URL" \ - --sha "$FM_PR_MERGE_HEAD" --yes "$@" || gitlab_merge_rc=$? - merge_control_cleanup + --sha "$FM_PR_MERGE_HEAD" --yes "$@" "${gitlab_merge_args[@]+"${gitlab_merge_args[@]}"}" || gitlab_merge_rc=$? + if [ "$gitlab_merge_rc" -eq 0 ]; then + persist_accepted_merge_authority || exit 1 + fi + fm_afk_contract_lock_release || true + fm_lock_release "$MERGE_CONTROL_LOCK" || true MERGE_CONTROL_LOCK= - # Before the command status was captured, set -e guaranteed the merge command - # had succeeded whenever this ran, so "accepted the merge request" was always - # true. It is not any more, and a confirm that claims acceptance on a failed - # command contradicts the refusal printed immediately after it. gitlab_confirm_rc=0 if [ "$gitlab_merge_rc" -ne 0 ]; then FM_PR_GITLAB_ACCEPTANCE='rejected the merge command' @@ -1416,7 +1871,8 @@ esac # refused or failed merge above, and a queued forge merge exits without an # outcome while its existing poll remains armed. outcome_rc=0 -fm_merge_outcome_report "$FM_HOME" "$STATE" "$ID" "$URL" self || outcome_rc=$? +fm_merge_outcome_report "$FM_HOME" "$STATE" "$ID" "$URL" self \ + "${FM_PR_MERGE_AUTHORITY:-}" || outcome_rc=$? case "$outcome_rc" in 0) ;; 3) diff --git a/bin/fm-procevent-lib.sh b/bin/fm-procevent-lib.sh index 266365dd526..f5fce33dee1 100644 --- a/bin/fm-procevent-lib.sh +++ b/bin/fm-procevent-lib.sh @@ -211,22 +211,50 @@ fm_procevent_launch_floor_seconds() { printf '%s\n' "$value" } -fm_procevent_launch_floor_reset_locked() { # <state-root> <source-id> <registration-identity> +# How long reconcile waits for a runner it just detached to prove it took the +# source's claim. Confirmation reads durable evidence, so a healthy launch +# settles on the first poll and only a launch not yet proved spends the +# window. The default stays well below FM_POLL because bin/fm-watch.sh runs +# reconcile once per supervision cycle, and every launch of a cycle shares ONE +# window rather than taking a window each. +FM_PROCEVENT_LAUNCH_CONFIRM_DEFAULT_SECONDS=3 +FM_PROCEVENT_LAUNCH_CONFIRM_MIN_SECONDS=1 +FM_PROCEVENT_LAUNCH_CONFIRM_MAX_SECONDS=600 + +fm_procevent_launch_confirm_seconds() { + local value=${FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS-} + if [ -z "$value" ]; then + printf '%s\n' "$FM_PROCEVENT_LAUNCH_CONFIRM_DEFAULT_SECONDS" + return 0 + fi + case "$value" in ''|*[!0-9]*) return 1 ;; esac + [ "$value" -ge "$FM_PROCEVENT_LAUNCH_CONFIRM_MIN_SECONDS" ] || return 1 + [ "$value" -le "$FM_PROCEVENT_LAUNCH_CONFIRM_MAX_SECONDS" ] || return 1 + printf '%s\n' "$value" +} + +# The one place the launch-pacing stamp's name is constructed. Every writer, +# pruner and reader goes through here so the naming rule is stated once. +fm_procevent_launch_floor_stamp_path() { # <state-root> <source-id> <registration-identity> local reg identity case "$3" in *:*) ;; *) return 1 ;; esac case "$3" in ''|*[!0-9:]*) return 1 ;; esac + fm_procevent_source_id_valid "$2" || return 1 reg=$(fm_procevent_registry_dir "$1") || return 1 identity=${3//:/-} - rm -f -- "$reg/$2.$identity.last-launch" + printf '%s\n' "$reg/$2.$identity.last-launch" +} + +fm_procevent_launch_floor_reset_locked() { # <state-root> <source-id> <registration-identity> + local stamp + stamp=$(fm_procevent_launch_floor_stamp_path "$1" "$2" "$3") || return 1 + rm -f -- "$stamp" } fm_procevent_launch_floor_prune_locked() { # <state-root> <source-id> <registration-identity> - local reg identity keep stamp - case "$3" in *:*) ;; *) return 1 ;; esac - case "$3" in ''|*[!0-9:]*) return 1 ;; esac + local reg keep stamp + keep=$(fm_procevent_launch_floor_stamp_path "$1" "$2" "$3") || return 1 reg=$(fm_procevent_registry_dir "$1") || return 1 - identity=${3//:/-} - keep="$reg/$2.$identity.last-launch" for stamp in "$reg/$2".*.last-launch "$reg/$2.last-launch"; do [ "$stamp" = "$keep" ] && continue [ -e "$stamp" ] || [ -L "$stamp" ] || continue @@ -235,12 +263,9 @@ fm_procevent_launch_floor_prune_locked() { # <state-root> <source-id> <registra } fm_procevent_launch_floor_wait() { # <state-root> <source-id> <registration-identity> <seconds> - local state=$1 id=$2 expected=$3 floor=$4 reg stamp identity registration current_identity status=0 - case "$expected" in *:*) ;; *) return 1 ;; esac - case "$expected" in ''|*[!0-9:]*) return 1 ;; esac + local state=$1 id=$2 expected=$3 floor=$4 reg stamp registration current_identity status=0 + stamp=$(fm_procevent_launch_floor_stamp_path "$state" "$id" "$expected") || return 1 reg=$(fm_procevent_registry_dir "$state") || return 1 - identity=${expected//:/-} - stamp="$reg/$id.$identity.last-launch" [ ! -L "$stamp" ] || return 1 [ ! -e "$stamp" ] || [ -f "$stamp" ] || return 1 perl -MTime::HiRes=clock_gettime,sleep,CLOCK_MONOTONIC -e ' @@ -607,6 +632,29 @@ fm_procevent_claim_generation_gone_locked() { && ! fm_procevent_group_alive "${FM_PROCEVENT_CLAIM_PID:-}" } +# fm_procevent_claim_undisplaceable_locked <source-id> +# The single owner of "this stale claim is one no unattended caller may +# displace". True when a claim record is still present for the source and its +# generation is NOT provably gone. Call it only where +# fm_procevent_claim_state_locked has just returned 1, so the FM_PROCEVENT_CLAIM_* +# globals below describe this source: that same return also covers a source with +# no claim record at all, which leaves those globals holding whatever the +# previous load put there, so the record check has to travel with the generation +# check rather than being left to each caller. +# +# What the surviving process group means is why this refuses rather than +# relaunches. fm_procevent_group_alive probes the runner's OWN process group, +# and the runner leads that group with its polling source child inside it, so +# "the group still has members" can mean that child is still attached to the +# session the source collects from. Starting a replacement there puts a second +# destructive poller on one session, which drains and loses what the source was +# collecting. A source that needs a human beats a source that silently eats what +# it was supposed to deliver. +fm_procevent_claim_undisplaceable_locked() { # <source-id> + [ -e "$(fm_procevent_claim_path "$1")" ] || return 1 + ! fm_procevent_claim_generation_gone_locked +} + # Capture-reservation cleanup for a claim being reclaimed. # # Reservation records are keyed by CLAIM TOKEN, and every replacement claims a @@ -707,6 +755,26 @@ fm_procevent_claim_acquire_locked() { if [ "$status" -eq 0 ]; then fm_procevent_claim_capture_reservation_reclaim_locked || status=1 fi + # Every cleanup above tidies leftovers that belong to the DEAD + # generation - its staging file and its capture reservation, both keyed + # by ITS claim token - and a replacement always claims a fresh token, + # so nothing a failed tidy-up leaves behind can collide with the + # generation that replaces it. + # fm_procevent_claim_capture_reservation_reclaim_locked already states + # that rule for the reservation record; the staging file takes the same + # rule here, and so does the shape check on the registry directory + # recorded to hold it, which only decides whether that removal is safe + # to attempt. Once the stale owner and the + # independently absent process group prove the whole generation gone, + # the documented ownership promise is already granted, so a failed + # tidy-up may leave litter and nothing more. Vetoing the claim instead + # is what leaves a provably dead runner owning the source permanently, + # where no reconcile, no retire and no fresh arm can displace it. + if [ "$status" -ne 0 ] && fm_procevent_claim_generation_gone_locked; then + status=0 + fi + # Two owners is the one outcome worse than none: never proceed on a + # claim record that is still there. [ "$status" -ne 0 ] || rm -f -- "$claim" || status=1 else status=1 diff --git a/bin/fm-procevent-when.sh b/bin/fm-procevent-when.sh index c67539f27c9..76f11f7df68 100755 --- a/bin/fm-procevent-when.sh +++ b/bin/fm-procevent-when.sh @@ -10,6 +10,7 @@ # fm-procevent-when.sh terminal <result-file> # fm-procevent-when.sh source-id <name> # fm-procevent-when.sh retire <name> +# fm-procevent-when.sh rebind-all # fm-procevent-when.sh run <source-id> # # arm Bind a (condition, action) pair as process-event source @@ -49,6 +50,19 @@ # record, and fired marker. Idempotent. Captured results and their # handled acknowledgements are never touched. Warns when the action # had already fired without a captured outcome. +# rebind-all Refresh the trust binding of every registered watch whose action +# executable lives under this repo (FM_ROOT), re-hashing it against +# its CURRENT on-disk bytes. A self-update fast-forwards bin/ in +# place, which changes those bytes with no tampering involved; left +# alone, the next fire is refused as not matching the registered +# trust binding, and the watch dies silently. rebind-all is meant to +# run right after such an update. It still validates each watch's +# existing spec and trust chain exactly as an ordinary fire would +# (a watch already broken for some other reason is reported, not +# silently patched over), and it never touches an action executable +# outside FM_ROOT: rebinding follows this repo's own tracked +# update, never an arbitrary swapped action. Idempotent: a watch +# whose action bytes already match its binding is left alone. # run The blocking child the generic runner executes; never run it in a # conversational turn. It polls the condition on the registered # cadence, requires the stable count of consecutive trues, claims a @@ -78,6 +92,7 @@ SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" FM_ROOT="${FM_ROOT_OVERRIDE:-$(cd "$SCRIPT_DIR/.." && pwd)}" FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" +FM_ROOT_REAL=$(cd "$FM_ROOT" 2>/dev/null && pwd -P) || FM_ROOT_REAL=$FM_ROOT # shellcheck source=bin/fm-pr-lib.sh . "$SCRIPT_DIR/fm-pr-lib.sh" @@ -367,7 +382,7 @@ cmd_run() { emit_doc "$sid" rejected "cannot stage command output; nothing was executed" 0 '' '' exit 0 fi - trap 'rm -f -- "$out"' EXIT + trap 'fm_procevent_source_lock_release "$sid"; rm -f -- "$out"' EXIT while :; do now=$(date +%s) @@ -415,15 +430,29 @@ cmd_run() { exit 0 fi - # Revalidate the registered action bytes immediately before claiming the - # fire. A changed or unavailable executable must never be run. + # Reload the trust binding from disk immediately before claiming the fire, + # rather than trusting the value cached at spec_load time when this poll + # loop started: a rebind-all can run (e.g. after a self-update) while this + # process is still polling, and only a fresh read sees its rebound hash. + # rebind_one publishes the spec and trust files as two separate renames, so + # the lock brackets this reload exactly as it brackets that publish, + # keeping the reader from observing a torn intermediate state. local current_action_hash + if ! fm_procevent_source_lock_acquire "$sid"; then + emit_doc "$sid" rejected "refused without executing anything: cannot lock the watch source" "$polls" '' '' + exit 0 + fi + if ! spec_load "$sid"; then + emit_doc "$sid" rejected "refused without executing anything: $SPEC_ERROR" "$polls" '' '' + exit 0 + fi current_action_hash=$(fm_pr_sha256 "${ACT_ARGV[0]}") || current_action_hash= if [ "$current_action_hash" != "$SPEC_ACTION_SHA256" ]; then emit_doc "$sid" rejected \ "refused without executing the action: its bytes do not match the registered trust binding" "$polls" '' '' exit 0 fi + fm_procevent_source_lock_release "$sid" # Claim the fire durably and exclusively BEFORE the action, so no restart or # concurrent runner can ever run the action a second time. @@ -473,6 +502,110 @@ cmd_terminal() { [ "$(cmd_classify "$file")" != unknown ] } +# --- rebind-all --------------------------------------------------------------- + +# publish_spec <sid> <device> <action_hash>: write and hash-bind a spec from +# the SPEC_* scalars and COND_ARGV/ACT_ARGV a prior spec_load already +# populated, using the given action hash. Mirrors cmd_arm's write block; the +# only caller today is rebind_one, refreshing action_sha256 alone. +publish_spec() { + local sid=$1 device=$2 action_hash=$3 tmp trust_tmp hash + tmp=$(umask 077; mktemp "$WHEN_DIR/.spec.XXXXXX") || return 1 + { + printf 'fm-when-spec-v1\n' + printf 'armed=%s\n' "$SPEC_ARMED" + printf 'interval=%s\n' "$SPEC_INTERVAL" + printf 'stable=%s\n' "$SPEC_STABLE" + printf 'deadline=%s\n' "$SPEC_DEADLINE" + printf 'condition_timeout=%s\n' "$SPEC_CONDITION_TIMEOUT" + printf 'action_timeout=%s\n' "$SPEC_ACTION_TIMEOUT" + printf 'error_budget=%s\n' "$SPEC_ERROR_BUDGET" + printf 'action_sha256=%s\n' "$action_hash" + printf 'condition_argc=%s\n' "${#COND_ARGV[@]}" + printf 'action_argc=%s\n' "${#ACT_ARGV[@]}" + printf 'argv:\n' + printf '%s\n' "${COND_ARGV[@]}" + printf '%s\n' "${ACT_ARGV[@]}" + } > "$tmp" || { rm -f -- "$tmp"; return 1; } + chmod 0600 "$tmp" || { rm -f -- "$tmp"; return 1; } + hash=$(fm_pr_sha256 "$tmp") || { rm -f -- "$tmp"; return 1; } + trust_tmp=$(umask 077; mktemp "$WHEN_DIR/.trust.XXXXXX") || { rm -f -- "$tmp"; return 1; } + printf 'fm-when-trust-v1\n%s\n' "$hash" > "$trust_tmp" || { rm -f -- "$tmp" "$trust_tmp"; return 1; } + chmod 0600 "$trust_tmp" || { rm -f -- "$tmp" "$trust_tmp"; return 1; } + mv -f -- "$tmp" "$(spec_file "$sid")" || { rm -f -- "$tmp" "$trust_tmp"; return 1; } + mv -f -- "$trust_tmp" "$(trust_file "$sid")" || { rm -f -- "$(spec_file "$sid")" "$trust_tmp"; return 1; } + if ! fm_pr_private_file_valid "$(spec_file "$sid")" 600 "$device" \ + || ! fm_pr_private_file_valid "$(trust_file "$sid")" 600 "$device"; then + rm -f -- "$(spec_file "$sid")" "$(trust_file "$sid")" + return 1 + fi +} + +# rebind_one <source-id>: 0 = rebound, 1 = failed (reported to stderr), 2 = +# unchanged or the action lives outside FM_ROOT (skipped, not an error). +rebind_one() { + local sid=$1 action_path action_hash device + if ! fm_procevent_source_lock_acquire "$sid"; then + printf 'skip: %s (cannot lock)\n' "$sid" >&2 + return 1 + fi + if ! spec_load "$sid"; then + printf 'skip: %s (%s)\n' "$sid" "$SPEC_ERROR" >&2 + fm_procevent_source_lock_release "$sid" + return 1 + fi + if ! action_path=$(action_executable "${ACT_ARGV[0]}"); then + printf 'skip: %s (action executable is unavailable: %s)\n' "$sid" "${ACT_ARGV[0]}" >&2 + fm_procevent_source_lock_release "$sid" + return 1 + fi + case "$action_path" in + "$FM_ROOT_REAL"/*) ;; + *) fm_procevent_source_lock_release "$sid"; return 2 ;; + esac + if ! action_hash=$(fm_pr_sha256 "$action_path"); then + printf 'skip: %s (cannot hash the action executable)\n' "$sid" >&2 + fm_procevent_source_lock_release "$sid" + return 1 + fi + if [ "$action_hash" = "$SPEC_ACTION_SHA256" ]; then + fm_procevent_source_lock_release "$sid" + return 2 + fi + if ! device=$(fm_pr_file_device "$WHEN_DIR"); then + printf 'skip: %s (cannot inspect the watch directory)\n' "$sid" >&2 + fm_procevent_source_lock_release "$sid" + return 1 + fi + if ! publish_spec "$sid" "$device" "$action_hash"; then + printf 'skip: %s (could not publish the refreshed trust binding)\n' "$sid" >&2 + fm_procevent_source_lock_release "$sid" + return 1 + fi + fm_procevent_source_lock_release "$sid" + printf 'rebound: %s\n' "$sid" + return 0 +} + +cmd_rebind_all() { + [ "$#" -eq 0 ] || usage + local spec sid rebound=0 skipped=0 failed=0 rc + [ -d "$WHEN_DIR" ] || { printf 'no watches registered\n'; return 0; } + for spec in "$WHEN_DIR"/when-*.spec; do + [ -e "$spec" ] || continue + sid=$(basename "$spec" .spec) + rebind_one "$sid" + rc=$? + case "$rc" in + 0) rebound=$((rebound + 1)) ;; + 2) skipped=$((skipped + 1)) ;; + *) failed=$((failed + 1)) ;; + esac + done + printf 'rebind-all: %s rebound, %s unchanged or out of scope, %s failed\n' "$rebound" "$skipped" "$failed" + [ "$failed" -eq 0 ] +} + # --- retire ------------------------------------------------------------------ cmd_retire() { @@ -499,6 +632,7 @@ case "${1-}" in terminal) shift; cmd_terminal "$@" ;; source-id) shift; cmd_source_id "$@" ;; retire) shift; cmd_retire "$@" ;; + rebind-all) shift; cmd_rebind_all "$@" ;; ''|-h|--help|help) usage ;; *) die "unknown command: $1" ;; esac diff --git a/bin/fm-procevent.sh b/bin/fm-procevent.sh index 93361b64552..ee31dd8b3be 100755 --- a/bin/fm-procevent.sh +++ b/bin/fm-procevent.sh @@ -47,6 +47,29 @@ # start a runner for any registered source that has no live owner. # This is liveness repair only - it never discovers results by # polling the source, because the child blocks on the source itself. +# A start is REPORTED only once it is confirmed: starting a runner is +# detached and its errors reach no caller, so a source that cannot +# start would otherwise be counted exactly like one that is +# listening, and a wedged source would go on presenting as armed. +# Every launch is counted as `started` only after the source is +# observed owned or its launch-pacing stamp has moved, `failed` +# otherwise, and any failure also makes this command exit non-zero. +# One bounded window covers a whole cycle's launches +# (FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS; docs/configuration.md). +# A launch that fails to confirm is also announced as a durable +# `check` wake, once per failure episode - keyed by the registration +# identity it ran under and ended by a later launch of that source +# confirming - because the supervision cycle discards the `failed=` +# count. The launch itself is retried every cycle exactly as before. +# A source whose claim nothing may automatically displace is not +# relaunched at all; it is counted `uncertain` and announced once per +# stranded claim generation as a durable `check` wake, because the +# supervision cycle discards this command's own output and exit +# status. The wake names what clears that strand: the `start` +# command for a reused pid whose group survives, or the check a +# human makes for a group that lost its leader, which `start` +# reports as owned and which the next cycle reclaims on its own +# once that group is empty. # handled Durably and idempotently record that a captured result has been # fully handled: <source-id> <sequence>. Prints "handled: id seq" # the first time for that exact source-and-sequence generation and @@ -346,6 +369,8 @@ adapter_self_announcing() { # <adapter> source_file() { printf '%s/%s.source\n' "$REG" "$1"; } runner_file() { printf '%s/%s.runner\n' "$REG" "$1"; } staging_file() { printf '%s/.%s.%s.output\n' "$REG" "$1" "$2"; } +stranded_file() { printf '%s/.%s.stranded\n' "$REG" "$1"; } +launch_failed_file() { printf '%s/.%s.launch-failed\n' "$REG" "$1"; } # Let the source's own adapter apply and acknowledge one captured result. See # the header for why this exists and what each exit means. An already @@ -1195,8 +1220,107 @@ detach_runner() { # <source-id> isolate_runner detach "$1" } +# Announce a source whose claim no unattended caller may displace, once per +# stranded claim generation. +# +# The supervision cycle runs this command with its output and its exit status +# both discarded, so a strand that only shows up in `list` as `orphaned` and in +# this command's `uncertain=` count reaches nobody. A durable `check` wake does +# reach firstmate through the ordinary queue, and it carries what clears the +# strand so acting on it needs no hunt. The caller supplies that part, because +# the two strand shapes clear differently and naming the wrong recovery would +# send someone to a command that reports `already owned` and changes nothing. +# +# The marker records the claim generation that was reported, so the same strand +# never wakes twice while a genuinely new claim still does - an alarm that +# repeats every supervision cycle is as unusable as one nobody gets. It is +# written before the wake and removed again if the wake does not land, so a +# failed announcement retries instead of being silently marked as delivered. +report_stranded_source() { # <source-id> <claim-token> <why-and-recovery> + local id=$1 token=$2 detail=$3 + case "$token" in ''|*[!A-Za-z0-9._-]*) return 1 ;; esac + [ -n "$detail" ] || return 1 + announce_source_once "$(stranded_file "$id")" "$token" \ + "procevent:$id:stranded:$token" \ + "check: process-event source $id is registered but nothing can arm it: $detail" +} + +# Announce a launch that reconcile could not confirm, once per failure episode. +# +# A launch that never proves it took the claim - a runner that died before +# claiming on unreadable argv, a missing adapter binary or a guard that refused +# to start, or one merely too slow under load - is relaunched every supervision +# cycle and reported `failed=` to a stdout that cycle discards: armed in +# appearance, a dead drop in fact, which is the incident with a different cause. +# Confirmation observes only that no claim and no launch stamp appeared inside +# the window, so this says exactly that and no more about why. An episode is +# keyed by the registration identity the launch ran under and ends when a later +# cycle finds the source owned or a launch confirms, so a second failure inside +# one episode announces nothing, a slow runner that arms later closes its own +# episode without a retraction, and a source that recovers and then fails again +# announces a new one. Nothing here changes what reconcile does about the launch +# itself: it keeps relaunching exactly as before, and this only says so once. +# +# The queue key carries a nonce beyond the episode: the watcher remembers every +# key it has surfaced for good, so a key made of the registration identity alone +# would be surfaced for the first episode only and every later episode of the +# same registration would sit in the queue unannounced. The marker records the +# episode and that nonce together, and the episode alone decides whether to +# announce. +report_launch_failure() { # <source-id> <registration-identity> + local id=$1 identity=$2 episode nonce + case "$identity" in ''|*[!0-9:]*) episode=unreadable ;; *) episode=${identity//:/-} ;; esac + nonce="$RANDOM$RANDOM" + announce_source_once "$(launch_failed_file "$id")" "$episode" \ + "procevent:$id:launch-failed:$episode-$nonce" \ + "check: process-event source $id is registered but its launch did not prove it took the source's claim within FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS, so nothing is confirmed to be collecting from it; reconcile reports that as failed= and keeps launching it every supervision cycle. If it stays that way, check the source command and the adapter binary the registration names, and run an attached bin/fm-procevent.sh start $id to reproduce a refusal on its stderr - the detached launch discards it, and a hand-run reconcile only counts it as failed=. A later cycle that finds the source owned ends this episode on its own, so a runner that was merely slow to claim needs nothing from you." \ + "$episode $nonce" +} + +# Shared marker discipline for the announcements above: <marker> holds the +# generation last reported as its first field, written before the wake and +# removed again if the wake does not land, so a failed announcement retries +# instead of being marked delivered, and the same generation never announces +# twice. A caller may store more after that field (the launch-failure nonce); +# only the first field decides. +announce_source_once() { # <marker> <generation> <key> <payload> [marker-record] + local marker=$1 generation=$2 key=$3 payload=$4 record=${5:-$2} previous + previous=$(cat -- "$marker" 2>/dev/null || true) + [ "${previous%%[[:space:]]*}" != "$generation" ] || return 1 + (umask 077; printf '%s\n' "$record" > "$marker") || return 1 + if ! fm_wake_append check "$key" "$payload"; then + rm -f -- "$marker" + return 1 + fi + return 0 +} + +# The reused-pid strand: the recorded pid is alive under a different identity +# while the runner's process group still has members. The claim path does not +# consult the process group, so a deliberate `start` reclaims this - provided +# the dead generation's reservation records can still be tidied, because that +# tidy-up is only waived for a generation proven gone, and this one is not. +stranded_reused_pid_detail() { # <source-id> + printf '%s' "its claim names a dead runner whose process group still has members, so reconcile preserves that claim and starts no replacement. Check that nothing is still polling the source, then reclaim it with: bin/fm-procevent.sh start $1 - that reclaims it provided the dead generation's reservation records can still be tidied, and otherwise refuses with: cannot claim source" +} + +# The leaderless strand: the runner leader is gone and its group still has +# members. `start` reports this as owned and reclaims nothing, and nothing +# automatic signals that group, so the only honest recovery to name is the +# check a human makes; an empty group reads as gone on the next cycle. +stranded_leaderless_detail() { # <source-id> + printf '%s' "its runner died and its polling child may still be attached to the source's session, so reconcile preserves that claim and starts no replacement, and nothing automatic will touch that group. Verify whether anything is still polling $1; once that process group is empty, the next reconcile reclaims the source on its own." +} + cmd_reconcile() { - local rec id published started=0 stopped=0 uncertain=0 claim owner pid token identity claim_state stop_state + local rec id published started=0 stopped=0 uncertain=0 failed=0 claim owner pid token identity claim_state stop_state + local launch_identity launch_stamp launch_mark unconfirmed entry + local -a launched=() + # Rejected before anything is launched, and by name. A window this command + # cannot use makes every launch unconfirmable, so validating it later would + # report a fleet of perfectly healthy runners as `failed=` and blame nothing. + fm_procevent_launch_confirm_seconds >/dev/null \ + || die "FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS must be whole seconds from $FM_PROCEVENT_LAUNCH_CONFIRM_MIN_SECONDS to $FM_PROCEVENT_LAUNCH_CONFIRM_MAX_SECONDS" owner_lease_refresh published=$(publish_pending) @@ -1251,15 +1375,39 @@ cmd_reconcile() { if [ -f "$(source_file "$id")" ] && [ ! -L "$(source_file "$id")" ]; then fm_procevent_claim_state_locked "$id" claim_state=$? - if [ "$claim_state" -eq 1 ]; then + if [ "$claim_state" -eq 1 ] && fm_procevent_claim_undisplaceable_locked "$id"; then + # A stale claim whose process group still has members, which can mean + # the dead runner's polling child is still on the source's session + # (fm_procevent_claim_undisplaceable_locked owns that reasoning). + # Preserve the claim, start nothing, and say the cycle could not + # settle it, which is what this command already promises for the + # leaderless variant below. Only a deliberate `start` reclaims here, + # so report the strand durably rather than leaving it to whoever + # happens to run this command. + uncertain=$((uncertain + 1)) + report_stranded_source "$id" "$FM_PROCEVENT_CLAIM_TOKEN" \ + "$(stranded_reused_pid_detail "$id")" || true + elif [ "$claim_state" -eq 1 ]; then if ! cleanup_extension_registration_invocations_locked "$id"; then uncertain=$((uncertain + 1)) fm_procevent_source_lock_release "$id" continue fi + # Snapshot the launch-pacing stamp for the registration generation + # this launch will run under, while the source lock still keeps that + # registration from being replaced underneath it. The runner writes + # this stamp after it claims and before it runs the source command, + # and nothing removes it on the way out, so an advanced or newly + # appeared value is durable evidence the launch got going. + launch_identity=$(fm_pr_file_identity "$(source_file "$id")" 2>/dev/null) || launch_identity= + launch_mark= + if [ -n "$launch_identity" ] \ + && launch_stamp=$(fm_procevent_launch_floor_stamp_path "$STATE" "$id" "$launch_identity"); then + launch_mark=$(cat -- "$launch_stamp" 2>/dev/null || true) + fi fm_procevent_source_lock_release "$id" detach_runner "$id" - started=$((started + 1)) + launched+=("$id"$'\t'"$launch_identity"$'\t'"$launch_mark") continue elif [ "$claim_state" -eq 4 ]; then owner=$FM_PROCEVENT_CLAIM_HOME @@ -1277,15 +1425,119 @@ cmd_reconcile() { elif [ "$claim_state" -eq 3 ]; then # A leaderless group's generation is ambiguous under PID/PGID reuse, # so preserve its claim without signalling or starting a replacement. + # This is the ordinary crash shape, and `start` cannot clear it + # either, so it is announced the same way as the reused-pid strand + # above but naming what a human should check rather than a command. uncertain=$((uncertain + 1)) + report_stranded_source "$id" "$FM_PROCEVENT_CLAIM_TOKEN" \ + "$(stranded_leaderless_detail "$id")" || true elif [ "$claim_state" -eq 2 ]; then uncertain=$((uncertain + 1)) + elif [ "$claim_state" -eq 0 ]; then + # A live owner is the same evidence confirmation reads, however the + # runner was started, so it ends any launch-failure episode here. + rm -f -- "$(launch_failed_file "$id")" fi fi fm_procevent_source_lock_release "$id" done fi - printf 'reconciled: published=%s started=%s stopped=%s uncertain=%s\n' "$published" "$started" "$stopped" "$uncertain" + if [ "${#launched[@]}" -gt 0 ]; then + unconfirmed=$(confirm_launched_runners "${launched[@]}") \ + || unconfirmed=$(printf '%s\n' "${launched[@]}") + for entry in "${launched[@]}"; do + id=${entry%%$'\t'*} + launch_identity=${entry#*$'\t'} + launch_identity=${launch_identity%%$'\t'*} + if launch_entry_listed "$entry" "$unconfirmed"; then + failed=$((failed + 1)) + report_launch_failure "$id" "$launch_identity" || true + else + started=$((started + 1)) + rm -f -- "$(launch_failed_file "$id")" + fi + done + fi + printf 'reconciled: published=%s started=%s stopped=%s uncertain=%s failed=%s\n' \ + "$published" "$started" "$stopped" "$uncertain" "$failed" + [ "$failed" -eq 0 ] +} + +launch_entry_listed() { # <entry> <newline-separated entries> + local entry=$1 line + while IFS= read -r line; do + [ "$line" = "$entry" ] && return 0 + done <<< "$2" + return 1 +} + +# Bounded confirmation that every runner just detached actually took its +# source's claim, printing every launch entry that did not, one per line. +# +# detach_runner is fire-and-forget and discards the child's stderr, so before +# this every failure inside _start - a refused claim above all - was still +# counted and reported as a start. That made a source that CANNOT start +# indistinguishable from one that had, which is exactly how a wedged review +# board goes on presenting as armed while collecting nothing. +# +# Two signals confirm a launch, and each covers what the other cannot see: +# ownership covers the runner still blocked on its source, which is the only +# evidence such a runner ever shows; the launch-pacing stamp covers the runner +# that claimed, ran and exited between two polls, because the runner writes that +# stamp after claiming and before running the source command and nothing removes +# it on the way out - only registration replacement does, which also changes the +# snapshotted identity this reads under. A runner that dies BEFORE claiming +# reaches neither, and that is the case this confirmation exists to catch; a +# runner merely slow to claim looks the same inside the window, which is why +# the failure this reports is "not proved within the window" and nothing more. +# +# Every launch shares ONE window rather than taking a window each, so a whole +# fleet of failing sources costs a watcher cycle the same bounded wait as one. +confirm_launched_runners() { # <source-id><TAB><registration-identity><TAB><launch-stamp-before>... + local deadline window entry id rest identity before state stamp mark + local -a pending=("$@") remaining=() + window=$(fm_procevent_launch_confirm_seconds) || return 1 + # A zero-padded window is a valid value to its validator, which reads base 10; + # reading it as octal here would silently shorten the window or abort this + # subshell under `set -u` and report every launch as failed. + # SECONDS is an integer clock that can tick at any moment after this + # assignment, so a deadline of exactly SECONDS + window waits anywhere in + # [window - 1, window] and a healthy launch could be reported failed for + # losing a second it was promised. The extra second bounds the wait to + # [window, window + 1] instead: never less than configured. + deadline=$((SECONDS + 10#$window + 1)) + while :; do + remaining=() + for entry in "${pending[@]+"${pending[@]}"}"; do + id=${entry%%$'\t'*} + rest=${entry#*$'\t'} + identity=${rest%%$'\t'*} + before=${rest#*$'\t'} + state=1 + if fm_procevent_source_lock_try_acquire "$id"; then + fm_procevent_claim_state_locked "$id" + state=$? + fm_procevent_source_lock_release "$id" + fi + if [ "$state" -eq 0 ]; then + continue + fi + mark= + if [ -n "$identity" ] \ + && stamp=$(fm_procevent_launch_floor_stamp_path "$STATE" "$id" "$identity"); then + mark=$(cat -- "$stamp" 2>/dev/null || true) + fi + if [ -n "$mark" ] && [ "$mark" != "$before" ]; then + continue + fi + remaining+=("$entry") + done + pending=("${remaining[@]+"${remaining[@]}"}") + [ "${#pending[@]}" -gt 0 ] || break + [ "$SECONDS" -lt "$deadline" ] || break + sleep 0.05 + done + [ "${#pending[@]}" -eq 0 ] || printf '%s\n' "${pending[@]}" } # Stop a runner and the child it is blocked on. A runner started by reconcile is @@ -1489,6 +1741,8 @@ cmd_retire() { fi rm -f -- "$(source_file "$id")" rm -f -- "$(runner_file "$id")" + rm -f -- "$(stranded_file "$id")" + rm -f -- "$(launch_failed_file "$id")" fm_procevent_source_lock_release "$id" # A retired source produces no further answer, so drop any decision binding it # carried. Generic and idempotent: the binding owner is asked to forget this @@ -1653,7 +1907,7 @@ cmd_sweep_home() { } cmd_list() { - local rec id adapter owner pending + local rec id adapter owner pending claim_state owner_lease_refresh if ! fm_procevent_any_registered "$STATE"; then printf 'no sources registered\n' @@ -1666,7 +1920,23 @@ cmd_list() { adapter=$(read_adapter "$id" 2>/dev/null || echo '?') fm_procevent_source_lock_acquire "$id" || continue fm_procevent_claim_state_locked "$id" - case "$?" in 0) owner=live ;; 1) owner=none ;; 3) owner=orphaned ;; *) owner=uncertain ;; esac + claim_state=$? + # A stale claim whose process group still has members is exactly as + # undisplaceable as the leaderless group state 3 already reports, and a + # reused PID reaches it through state 1 rather than state 3. Reporting that + # as `none` reads like an idle source waiting to be started, which is the + # reassuring answer this whole surface gave while a board collected nothing. + case "$claim_state" in + 0) owner=live ;; + 1) + owner=none + if fm_procevent_claim_undisplaceable_locked "$id"; then + owner=orphaned + fi + ;; + 3) owner=orphaned ;; + *) owner=uncertain ;; + esac fm_procevent_source_lock_release "$id" pending=$(fm_procevent_pending "$STATE" | grep -c "/$id\." || true) printf '%-28s %-12s %-10s %s\n' "$id" "$adapter" "$owner" "$pending" diff --git a/bin/fm-public-followup.sh b/bin/fm-public-followup.sh index ea93e801dcc..ea5173902d5 100755 --- a/bin/fm-public-followup.sh +++ b/bin/fm-public-followup.sh @@ -208,9 +208,11 @@ require_tools() { command -v tasks-axi >/dev/null 2>&1 || die "tasks-axi is required" 1 } -# Every tasks-axi call runs from the home whose backlog owns the obligation, the -# same convention bin/fm-captain-hold.sh uses for typed backlog state. -tx() { (cd "$FM_HOME" && tasks-axi "$@"); } +# Every tasks-axi call addresses $FM_HOME/data, the home whose backlog owns the +# obligation, through bin/fm-tasks-axi.sh. An inherited FM_DATA_OVERRIDE is +# cleared because a caller such as a secondmate teardown names the parent home +# in FM_HOME while its own data override is still in the environment. +tx() { FM_HOME="$FM_HOME" FM_DATA_OVERRIDE='' "$SCRIPT_DIR/fm-tasks-axi.sh" "$@"; } # obligation_json <id>: the complete typed obligation payload on stdout, empty # when the backlog simply has no such public-followup item, and a non-zero exit @@ -285,7 +287,7 @@ cmd_register() { payload=$(obligation_json "$id") \ || die "could not read the backlog through tasks-axi" 1 [ -n "$payload" ] \ - || die "no public-followup obligation '$id' in this home's backlog; create it with tasks-axi public-followup add before registering" 1 + || die "no public-followup obligation '$id' in this home's backlog; create it with bin/fm-tasks-axi.sh public-followup add before registering" 1 # The relation must already be bound, so a registration can never describe a # binding tasks-axi does not have. @@ -293,7 +295,7 @@ cmd_register() { '(.public_followup.work_relations // []) | map(select(.relation_id == $r and .work_ref.home_id == $h and .work_ref.task_id == $w)) | length > 0' >/dev/null 2>&1 \ - || die "obligation '$id' has no bound relation '$relation' for $work_home/$work_id; run tasks-axi public-followup bind-work first" 1 + || die "obligation '$id' has no bound relation '$relation' for $work_home/$work_id; run bin/fm-tasks-axi.sh public-followup bind-work first" 1 [ -n "$platform" ] || platform=$(pf_field "$payload" '.public_followup.request.platform') [ -n "$request" ] || request=$(pf_field "$payload" '.public_followup.request.request_id') diff --git a/bin/fm-send.sh b/bin/fm-send.sh index 2c956a1d860..351eb8fd560 100755 --- a/bin/fm-send.sh +++ b/bin/fm-send.sh @@ -74,10 +74,11 @@ # so callers treat it as neither success nor safe-to-resend failure. # Submission dispatches through the target's recorded backend; the # tmux adapter shares its composer/submit core with the away-mode daemon via -# bin/fm-tmux-lib.sh. Tune with FM_SEND_RETRIES (default 3) / FM_SEND_SLEEP -# (0.4). Slash commands, and codex `$...` skill invocations resolved through -# harness meta, get a longer pre-Enter settle so completion popups do not -# swallow Enter. A remote secondmate target has no typed text plane at all: +# bin/fm-tmux-lib.sh. Tune with FM_SEND_RETRIES (default 3; agy typed targets +# default to 20 for agy's late busy render) / FM_SEND_SLEEP (0.4). Slash +# commands, and codex `$...` skill invocations resolved through harness meta, +# get a longer pre-Enter settle so completion popups do not swallow Enter. +# A remote secondmate target has no typed text plane at all: # every remote text steer rides the inbox (a marked secondmate request already # reaches the harness as marker-prefixed chat rather than a parser command, so # routing a remote "/..." or "$..." through the record changes nothing the @@ -575,7 +576,7 @@ fm_send_hold_resolved_id() { # <task-id> <decision-key> local show id state hold_kind command -v tasks-axi >/dev/null 2>&1 || return 1 for id in "$2" "$1-decision-$2"; do - show=$( (cd "$FM_HOME" && tasks-axi show "$id" --full) 2>/dev/null ) || continue + show=$(FM_HOME="$FM_HOME" FM_DATA_OVERRIDE='' "$SCRIPT_DIR/fm-tasks-axi.sh" show "$id" --full 2>/dev/null) || continue state=$(printf '%s\n' "$show" | sed -n 's/^ state: //p' | head -1) hold_kind=$(printf '%s\n' "$show" | sed -n 's/^ hold_kind: //p' | head -1) [ "$state" != "done" ] || continue @@ -1072,7 +1073,21 @@ else ;; *) settle=0.3 ;; esac - retries=${FM_SEND_RETRIES:-3} + # Per-harness submit-confirm budget. agy's bare `>` composer verdict is + # `unknown`, so a landed submit is acknowledged only by the idle-to-busy + # transition poll, and agy renders its verified busy footer well after the + # shared budget expires: ~1.5s after Enter for a short steer, ~4-5s for a + # realistic longer brief (live-measured, agy 1.2.1), against the shared + # default's 3 x 0.4s. With the shared default a typed steer to an agy + # endpoint was reported exit-1 non-delivery for a message that landed and + # ran, inviting a duplicate resend. agy typed targets get a longer default + # budget (~8s at the default cadence, twice the worst measured render); an + # explicit FM_SEND_RETRIES still wins, and every other harness keeps the + # shared 3-retry default untouched. + case "$TARGET_HARNESS" in + agy) retries=${FM_SEND_RETRIES:-20} ;; + *) retries=${FM_SEND_RETRIES:-3} ;; + esac sleep_s=${FM_SEND_SLEEP:-0.4} fm_send_reopen_scout_completion || exit 1 # Type once, submit, verify. Only exact empty confirms delivery; every other diff --git a/bin/fm-session-lock-lib.sh b/bin/fm-session-lock-lib.sh index c2a117b0b84..91c901f820b 100644 --- a/bin/fm-session-lock-lib.sh +++ b/bin/fm-session-lock-lib.sh @@ -122,7 +122,12 @@ fm_harness_ancestry_pids() { break fi pid=$(ps -o ppid= -p "$pid" 2>/dev/null | tr -d ' ') - [ -n "$pid" ] && [ "$pid" -gt 1 ] || break + # Examine the top of the chain before stopping. Inside a PID namespace the + # harness itself is pid 1, so stopping as soon as the next pid is 1 hides the + # very process this walk exists to find. A host's real pid 1 (init, systemd, + # launchd) is not harness-shaped, so fm_harness_process_matches rejects it. + case "$pid" in '' | *[!0-9]*) break ;; esac + [ "$pid" -ge 1 ] || break done [ "$printed" -eq 1 ] } diff --git a/bin/fm-session-start.sh b/bin/fm-session-start.sh index 996fb34563f..25acce13e92 100755 --- a/bin/fm-session-start.sh +++ b/bin/fm-session-start.sh @@ -158,7 +158,7 @@ # stay out of the startup digest; the same never-bound-a-held-or-blocked-row # rule applies, recognized there from the title line's own hold/blocked-by # markers. -# Full bodies are targeted follow-up only: `tasks-axi show <id> --full` when +# Full bodies are targeted follow-up only: `bin/fm-tasks-axi.sh show <id> --full` when # compatible tasks-axi is available, or `data/backlog.md` when the file body is # truly needed. # @@ -388,7 +388,7 @@ print_file_or_absent() { } print_backlog_pointer() { - printf 'Full task bodies remain available on demand: tasks-axi show <id> --full when compatible tasks-axi is available, or data/backlog.md.\n' + printf 'Full task bodies remain available on demand: bin/fm-tasks-axi.sh show <id> --full when compatible tasks-axi is available, or data/backlog.md.\n' } # A queued title line whose own text already marks it held or blocked. The @@ -455,8 +455,8 @@ strip_axi_help() { # and every other line it prints (its count, its public-followup line) passes # through untouched. Whatever is cut is disclosed exactly. print_ready_queued_bounded() { - local ready=$1 path=$2 - printf '%s\n' "$ready" | awk -v max="$QUEUED_LIMIT" -v path="$path" ' + local ready=$1 + printf '%s\n' "$ready" | awk -v max="$QUEUED_LIMIT" ' /^help\[/ { exit } /^ready\[/ { rows = 1; print; next } rows && /^[[:space:]]/ { @@ -469,7 +469,7 @@ print_ready_queued_bounded() { if (total > 0) { printf "(shown %d of %d ready queued item(s))\n", shown, total if (total > shown) { - printf "(%d more queued - tasks-axi ready --file %s)\n", total - shown, path + printf "(%d more queued - bin/fm-tasks-axi.sh ready)\n", total - shown } } } @@ -496,7 +496,7 @@ print_backlog_tasks_axi_compact() { printf '\nblocked queued:\n' printf '%s\n' "$blocked" | strip_axi_help printf '\nready queued (dispatchable now):\n' - print_ready_queued_bounded "$ready" "$path" + print_ready_queued_bounded "$ready" return 0 fi printf 'tasks-axi compact listing failed; falling back to title-line rendering.\n' @@ -751,6 +751,7 @@ fi stage supervision-instructions AFK_PRESENT=0 [ -e "$STATE/.afk" ] && AFK_PRESENT=1 +AFK_MODE=$(fm_afk_mode "$STATE") X_MODE_PRESENT=0 [ -f "$CONFIG/x-mode.env" ] && X_MODE_PRESENT=1 @@ -791,6 +792,7 @@ fi --harness "$PRIMARY_HARNESS" \ --read-only "$READ_ONLY" \ --afk "$AFK_PRESENT" \ + --afk-mode "$AFK_MODE" \ --x-mode "$X_MODE_PRESENT" # --- 5. read-once contract ------------------------------------------------- @@ -816,7 +818,7 @@ Go to a source directly only when: - an individual full status log is needed for older wake-event history, or a status line was capped and its tail matters (each task's full log path is printed with its tail), - - a full task body is needed (tasks-axi show <id> --full, or data/backlog.md), + - a full task body is needed (bin/fm-tasks-axi.sh show <id> --full, or data/backlog.md), - the backlog listing disclosed omitted queued items and this turn needs them, - the NETWORK CHECKS section reported its checks still IN PROGRESS and this turn needs their verdict (bin/fm-startup-network.sh report), @@ -881,12 +883,20 @@ if [ -f "$STATE/.afk-contract" ]; then printf 'present - away posture recorded at %s (hold-for-return only; bin/fm-afk-contract.sh readback for the mandate)' \ "$("$SCRIPT_DIR/fm-afk-contract.sh" field entered 2>/dev/null || printf unknown)" if [ -e "$STATE/.afk" ]; then - printf '; the away daemon owns the watcher.\n' + if [ "$AFK_MODE" = quiet ]; then + printf '; the quiet daemon owns the watcher.\n' + else + printf '; the away daemon owns the watcher.\n' + fi else printf '; no daemon runs, the ordinary supervision session continues.\n' fi elif [ -e "$STATE/.afk" ]; then - printf 'present - away-mode supervision is active; the daemon owns the watcher (legacy flag with no posture record).\n' + if [ "$AFK_MODE" = quiet ]; then + printf 'present - quiet-mode supervision is active; the daemon owns the watcher, only an explicit /quiet off exits it (attended mode without an away record).\n' + else + printf 'present - away-mode supervision is active; the daemon owns the watcher (legacy flag with no posture record).\n' + fi else printf 'absent\n' fi @@ -950,6 +960,14 @@ This session did not acquire the fleet lock. Stay read-only: do not arm, drain, spawn, steer, merge, or repair fleet state from here. Only a session with verified fleet-lock ownership may perform mutable follow-up. +EOF +elif [ "$AFK_PRESENT" -eq 1 ] && [ "$AFK_MODE" = quiet ]; then + cat <<'EOF' +Quiet mode is active. Follow the supervision operating instructions block +above: load /quiet and ensure the daemon is running, because the daemon owns +watcher supervision. Ordinary captain chat does not exit it; only an +explicit /quiet off does. + EOF elif [ "$AFK_PRESENT" -eq 1 ]; then cat <<'EOF' diff --git a/bin/fm-sessionstart-nudge.sh b/bin/fm-sessionstart-nudge.sh index fccf775dd95..a12aa3e4628 100755 --- a/bin/fm-sessionstart-nudge.sh +++ b/bin/fm-sessionstart-nudge.sh @@ -25,13 +25,22 @@ lock_is_in_ancestry() { [ -f "$STATE/.lock" ] || return 1 IFS= read -r lock_pid < "$STATE/.lock" 2>/dev/null || return 1 case "$lock_pid" in - ''|*[!0-9]*|1) return 1 ;; + # A lock pid of 1 is legitimate inside a PID namespace, where the harness + # holding the home lock IS pid 1, so it is no longer rejected outright; the + # liveness check below still gates it. On a host, a lock file that wrongly + # names pid 1 can now make this hook conclude the lock is already held and + # stay silent, which is the safe direction for a SessionStart hook whose only + # outputs are one nudge line or nothing. + ''|*[!0-9]*) return 1 ;; esac kill -0 "$lock_pid" 2>/dev/null || return 1 for _ in 1 2 3 4 5 6 7 8; do [ "$pid" = "$lock_pid" ] && return 0 pid=$(ps -o ppid= -p "$pid" 2>/dev/null | tr -d ' ') - [ -n "$pid" ] && [ "$pid" -gt 1 ] || return 1 + # Stop only after the top of the chain has been compared, for the same + # namespace reason as bin/fm-session-lock-lib.sh's walk. + case "$pid" in '' | *[!0-9]*) return 1 ;; esac + [ "$pid" -ge 1 ] || return 1 done return 1 } diff --git a/bin/fm-spawn.sh b/bin/fm-spawn.sh index 312fc394d43..7d848dc4068 100755 --- a/bin/fm-spawn.sh +++ b/bin/fm-spawn.sh @@ -124,7 +124,15 @@ # root Firstmate home's state directory before slot allocation and holds it through # task metadata publication. Teardown holds that same lock while proving and # returning a slot, so allocation cannot reuse a slot before its owner record -# is published. The local root is whatever bin/fm-wake-lib.sh's +# is published. Under that same lock it writes the slot's owner claim, which is +# what lets teardown leave a slot reassigned since untouched; bin/fm-wake-lib.sh +# owns the claim and bin/fm-teardown.sh owns what it protects. A slot that +# cannot be claimed refuses the spawn rather than launching a worker whose slot +# could later be released out from under its successor. A spawn that aborts +# while it still holds the allocation lock drops its own claim; an abort after +# metadata publication has released that lock leaves the claim in place, and +# the next spawn's claim replaces it. +# The local root is whatever bin/fm-wake-lib.sh's # fm_firstmate_root_home resolves, so a home seeded from another machine anchors # that lock itself rather than failing to resolve one; # contention refuses rather than waits. @@ -276,9 +284,21 @@ # This is an exec environment boundary, not a sandbox for the pane's startup # shell, credential files, same-user processes, or later shell initialization. # See docs/configuration.md for provider/Git setup and supported limits. +# Claude permission mode (config/claude-permission-mode): +# One token selecting the permission flag every claude launch (ship, scout, +# secondmate, and relaunch) carries. Absent or `bypass` keeps today's +# `--dangerously-skip-permissions`; `auto` launches with `--permission-mode +# auto` instead, Claude Code's classifier-reviewed mode, for a captain who +# refuses to run workers in bypass mode. Every other part of the claude launch +# is unchanged. The token is the file's whitespace-trimmed content; any other +# value, or an unreadable file, refuses the spawn before any endpoint, +# worktree, or record exists and names the accepted values. The file is read +# on every spawn and relaunch, so a change reaches the next launch without a +# restart, and it is inherited into secondmate homes (bin/fm-config-inherit-lib.sh). # Launch templates live in bin/fm-launch-lib.sh (fm_launch_template); placeholders replaced before launch: # __BRIEF__ absolute path to the worker-facing brief; design dispatches use # a per-dispatch copy carrying their pinned skill paths +# __CLAUDEPERMFLAG__ the permission flag selected by config/claude-permission-mode # __PIBIN__ quoted concrete Pi-family executable path resolved from PATH # __PITUIMODE__ optional --tui-mode regular when that executable advertises it # __TURNEND__ absolute path to state/<task-id>.turn-ended (for harnesses whose @@ -297,6 +317,7 @@ # __CURSORBIN__ resolved, cursor-verified executable for a cursor launch # __GEMINISETTINGS__ firstmate-owned per-task gemini settings file (busy-state hooks) # __ROVOBIN__ resolved, rovo-verified executable for a rovo launch +# __AGYBIN__ resolved, agy-verified executable for an agy launch # Verified per-harness turn-end hooks are installed automatically where enabled; some live outside the worktree. # Kimi uses one surgically installed Firstmate region in $HOME/.kimi-code/config.toml, # a firstmate-owned global hook and registry, and a gitignored per-task pointer. @@ -306,7 +327,7 @@ # plus a gitignored .fm-grok-turnend worktree pointer and a state token. # muse installs no hook at all - its plugin engine is off in the default build - so # it writes state/<id>.muse-session to bind the pane to muse's own session event -# log; muse and gemini are crewmate/scout only and are refused for --secondmate. +# log; muse, gemini, and agy are crewmate/scout only and are refused for --secondmate. # rovo installs no hook either - its eventHooks fire at tool granularity only, # never turn-end - so it carries no busy-source wiring at all and no turn-end # hook. A positional brief is dead-on-arrival (rovo loads, never works, and drops @@ -314,6 +335,9 @@ # only after a TUI readiness gate, then a delivery-confirmation gate - the same # launch-then-send shape as kimi. Its busy state is a screen-scrape fallback like # grok. rovo is crewmate/scout only and is refused for --secondmate, like muse. +# agy's crew-only Herdr lifecycle is owned by the harness-adapters agy reference. +# bin/fm-agy-trust-lib.sh owns fail-closed trust, rollback, and teardown ownership; +# the post-launch gate answers a residual dialog and confirms native working state. # cursor installs no per-task hook either: it writes state/<id>.cursor-session to # bind the pane to cursor's own conversation transcript (projects root, the exact # workspace path cursor records in .workspace-trusted, and the conversations that @@ -323,13 +347,13 @@ # park owns that home's supervision (docs/supervision-protocols/cursor.md). # Claude has a separate workspace-trust gate before its hook setup: before # any per-task state exists, and before its worktree .claude/settings.local.json -# hooks are written, a non-secondmate claude launch pre-registers the worktree in -# the launching user's own Claude trust store through bin/fm-claude-trust.sh, -# because Claude's interactive workspace-trust dialog gates a fresh worktree and -# firstmate cannot answer it. That helper's header owns the structural scope test -# and every refusal; a failed registration stops this spawn rather than launching -# a worker that would wedge on the dialog. A --secondmate launch never runs it, -# so a claude secondmate home keeps its own one-time trust decision. +# hooks are written, every claude launch pre-registers the directory the pane +# starts in - the task worktree, or the secondmate home for a --secondmate spawn - +# in the launching user's own Claude trust store through bin/fm-claude-trust.sh, +# because Claude's interactive workspace-trust dialog gates a folder it has never +# seen and firstmate cannot answer it. That helper's header owns the structural +# scope test for both shapes and every refusal; a failed registration stops this +# spawn rather than launching a worker that would wedge on the dialog. # Every claude launch also carries the attribution-off policy in its per-launch # --settings JSON, so a spawned worker never writes a Co-Authored-By trailer, # Claude-Session link, or generated-with line into a commit or PR body; @@ -438,6 +462,31 @@ if [ "$LAUNCH_ENV_ENABLED" = 1 ]; then exit 1 fi fi +# config/claude-permission-mode (header above): resolved once per spawn or +# relaunch, before any mutation, so a malformed file refuses instead of +# launching a worker on a permission posture the captain did not choose. +if ! CLAUDE_PERM_PRESENT=$(fm_config_source_present "$CONFIG/claude-permission-mode"); then + exit 1 +fi +CLAUDE_PERMISSION_MODE=bypass +if [ "$CLAUDE_PERM_PRESENT" = 1 ]; then + if [ ! -f "$CONFIG/claude-permission-mode" ] || [ ! -r "$CONFIG/claude-permission-mode" ]; then + echo "error: config/claude-permission-mode must be a readable regular file holding one of: bypass, auto" >&2 + exit 1 + fi + CLAUDE_PERMISSION_MODE=$(tr -d '[:space:]' < "$CONFIG/claude-permission-mode" || true) + case "$CLAUDE_PERMISSION_MODE" in + bypass|auto) ;; + *) + echo "error: config/claude-permission-mode holds '$CLAUDE_PERMISSION_MODE'; accepted values are: bypass (--dangerously-skip-permissions, the default when the file is absent), auto (--permission-mode auto)" >&2 + exit 1 + ;; + esac +fi +case "$CLAUDE_PERMISSION_MODE" in + auto) CLAUDE_PERM_FLAG='--permission-mode auto' ;; + *) CLAUDE_PERM_FLAG='--dangerously-skip-permissions' ;; +esac SUB_HOME_MARKER=".fm-secondmate-home" if [ -e "$STATE" ] || [ -L "$STATE" ]; then fm_backlog_directory_present "$STATE" "state directory" || { @@ -481,6 +530,8 @@ fm_backlog_directory_present "$STATE" "state directory" || { . "$SCRIPT_DIR/fm-trace-context-lib.sh" # shellcheck source=bin/fm-remote-readiness-lib.sh . "$SCRIPT_DIR/fm-remote-readiness-lib.sh" +# shellcheck source=bin/fm-timeout-lib.sh +. "$SCRIPT_DIR/fm-timeout-lib.sh" # Fail closed before any fleet mutation: a no-mistakes gate agent must never spawn # a direct report (see bin/fm-gate-refuse-lib.sh). fm_refuse_if_gate_agent @@ -924,6 +975,7 @@ SPAWN_TASK_SET_LOCK= SPAWN_TASK_SET_LOCK_HELD=0 SPAWN_TREEHOUSE_PROJECT_LOCK= SPAWN_TREEHOUSE_PROJECT_LOCK_HELD=0 +SPAWN_SLOT_CLAIMED=0 RELAUNCH_REPLACEMENT_PENDING=0 RELAUNCH_REPLACEMENT_BUSY_GEN= RELAUNCH_REPLACEMENT_HARNESS= @@ -1074,6 +1126,23 @@ spawn_abort_cleanup() { SPAWN_META_LOCK_HELD=0 fm_lock_release "$SPAWN_META_LOCK" || true fi + # A spawn that aborts after claiming its slot but before its record survives + # must not leave a claim naming a task no record describes. The release is a + # read-then-remove, so it runs only while the project lock that wrote the + # claim is still held (aborts before metadata publication); a later abort has + # already released that lock and leaves the claim for the next spawn's + # atomic replacement rather than racing it. The release itself never removes + # another task's claim. + if [ "$SPAWN_SLOT_CLAIMED" = 1 ] && [ -n "${WT:-}" ] \ + && [ ! -e "$STATE/$ID.meta" ] && [ ! -L "$STATE/$ID.meta" ] \ + && fm_treehouse_pool_slot "$PROJ_ABS" "$WT"; then + SPAWN_SLOT_CLAIMED=0 + if [ "$SPAWN_TREEHOUSE_PROJECT_LOCK_HELD" = 1 ]; then + fm_treehouse_slot_owner_release "$WT" "$ID" || true + else + echo "warning: leaving task $ID's slot claim on $WT in place; the Treehouse project lock is no longer held, so the next spawn's claim replaces it" >&2 + fi + fi if [ "$SPAWN_TREEHOUSE_PROJECT_LOCK_HELD" = 1 ]; then SPAWN_TREEHOUSE_PROJECT_LOCK_HELD=0 fm_lock_release "$SPAWN_TREEHOUSE_PROJECT_LOCK" || true @@ -1395,7 +1464,7 @@ if [ "$RELAUNCH" -eq 1 ]; then } elif [ "$KIND" = secondmate ]; then case "${POS[1]:-}" in - ''|claude|codex|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse|rovo|omp) + ''|claude|codex|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse|rovo|omp|agy) ARG3=${POS[1]:-} ;; *' '*) @@ -1467,6 +1536,37 @@ omp_model_validate() { # <omp-bin> <model> return 1 } +# agy pre-launch model validation. `agy models` (agy 1.2.0) prints one model per +# line as "<id>\t<label>" for the account's catalog only; model ids are bare +# (gemini-3.8-flash-high), never provider-prefixed. A requested model absent +# from a reachable listing is concrete unsupported evidence and refuses the +# spawn, so a stale id (the unlisted bare gemini-3.8-flash) fails loudly here +# instead of wedging a worker pane. The listing is a remote fetch that needs +# network and a signed-in account, so the probe runs under the shared hard +# bound (bin/fm-timeout-lib.sh) with stdin detached: a stalled fetch or a +# sign-in prompt can never block the spawn before any pane exists. An +# unreachable listing establishes nothing (harness-adapters +# model-and-effort.md) and launches unvalidated with a notice. +agy_model_validate() { # <agy-bin> <model> + local bin=$1 model=$2 listing rc=0 bound=${FM_AGY_MODELS_TIMEOUT:-15} + case "$bound" in ''|*[!0-9]*|0*) bound=15 ;; esac + [ -n "$model" ] && [ "$model" != default ] || return 0 + listing=$(fm_run_timed "$bound" "$bin" models 2>/dev/null < /dev/null) || rc=$? + if [ "$rc" -ne 0 ] || [ -z "$listing" ]; then + if [ "$rc" -eq 124 ]; then + echo "notice: 'agy models' did not answer within ${bound}s; launching with --model '$model' unvalidated" >&2 + else + echo "notice: 'agy models' listing is unreachable (exit $rc); launching with --model '$model' unvalidated" >&2 + fi + return 0 + fi + if printf '%s\n' "$listing" | awk '{print $1}' | grep -qxF -- "$model"; then + return 0 + fi + echo "error: agy model '$model' is not listed by 'agy models'; choose a listed id or omit --model" >&2 + return 1 +} + # The verified launch command per adapter lives in bin/fm-launch-lib.sh # (fm_launch_template), sourced above - including origin/main #909's __OPINPUT__ # operational-input encoding threaded into each template there. @@ -1477,7 +1577,6 @@ case "$ARG3" in *' '*) # raw launch command (unverified-adapter escape hatch) RAW_LAUNCH=1 LAUNCH=$ARG3 - RAW_LAUNCH=1 HARNESS="" for word in $LAUNCH; do case "$word" in [A-Za-z_]*=*) continue ;; *) HARNESS=$(basename "$word"); break ;; esac @@ -1532,7 +1631,7 @@ esac # agy is a CREW-ONLY, herdr-ONLY adapter (captain-approved divergence; # data/captain.md, verification data/cursor-agy-verify/report.md). Upstream carries -# no agy support, so this gate is fork-local. Enforce both halves before any +# a compatible agy adapter, but this scope gate remains fork-local. Enforce both halves before any # container or worktree work: agy can supervise a ship/design/scout crewmate only, # never a secondmate, and only on the herdr backend, whose native agent-state gives # liveness plus the debounced turn-boundary wake the watcher derives @@ -1555,7 +1654,7 @@ case "$HARNESS" in ;; esac -# muse and gemini are verified as CREWMATE/SCOUT adapters only. A secondmate is +# muse, gemini, and agy are verified as CREWMATE/SCOUT adapters only. A secondmate is # a firstmate instance, so it needs a primary supervision protocol. # gemini has none: docs/supervision-protocols/ carries no gemini wake protocol # and this task verified only crewmate-side launch, busy state, interrupt, and @@ -1565,7 +1664,9 @@ esac # asyncRewake handlers that firstmate's primary turn-end supervision is built on # (muse 0.1.0-R708.1). Refusing here keeps that gap loud instead of standing up a # secondmate whose supervision cycle could never be armed. -if [ "$KIND" = secondmate ] && { [ "$HARNESS" = muse ] || [ "$HARNESS" = gemini ]; }; then +# agy has none either: it exposes no hook surface for primary supervision and +# docs/supervision-protocols/ carries no agy wake protocol (agy 1.2.0). +if [ "$KIND" = secondmate ] && { [ "$HARNESS" = muse ] || [ "$HARNESS" = gemini ] || [ "$HARNESS" = agy ]; }; then echo "error: $HARNESS is a verified crewmate/scout adapter only and cannot run a secondmate; it has no primary supervision protocol. Select a harness verified for secondmates." >&2 exit 1 fi @@ -1621,6 +1722,12 @@ case "$HARNESS" in exit 1 } ;; + agy) + AGY_BIN=$(resolve_pi_executable agy) || { + echo "error: agy executable not found on PATH; install Antigravity CLI or select a different verified harness" >&2 + exit 1 + } + ;; esac # config/secondmate-harness may carry optional model/effort tokens alongside the @@ -1656,6 +1763,9 @@ fi if [ "$HARNESS" = omp ]; then omp_model_validate "$OMP_BIN" "$MODEL" || exit 1 fi +if [ "$HARNESS" = agy ]; then + agy_model_validate "$AGY_BIN" "$MODEL" || exit 1 +fi secondmate_registry_value() { secondmate_registry_field "$DATA/secondmates.md" "$1" "$2" @@ -2551,8 +2661,13 @@ herdr_projection_existing_meta_allows_flat() { # <meta> } old_state=$(fm_backend_herdr_pane_agent_state "$old_session" "$old_pane") case "$old_state" in + # A stale registration over a shell-only pane is agent-free for RECOVERY + # (--relaunch reuses the pane, issue #4115), but the duplicate-launch + # corridor keeps refusing it like every other non-husk state, so a fresh + # spawn is refused here consistently with the reclaim and presentation + # gates downstream. dead|no-agent) return 0 ;; - live|unknown) + live|stale-agent|unknown) echo "error: existing herdr endpoint for $ID is $old_state; refusing duplicate launch" >&2 return 1 ;; @@ -2580,7 +2695,7 @@ if fm_backlog_transition_applies "$CONFIG" "$DATA" "$KIND"; then if fm_backlog_row_probe "$DATA" "$ID"; then BACKLOG_ROW_STATE=$FM_BACKLOG_ROW_STATE elif [ "$FM_BACKLOG_ROW_RESULT" = not_found ]; then - echo "error: task $ID has no backlog item in this home, so dispatching it would leave a worker no record owns; add it first (tasks-axi add $ID '<title>' --kind $KIND) and re-run" >&2 + echo "error: task $ID has no backlog item in this home, so dispatching it would leave a worker no record owns; add it first (bin/fm-tasks-axi.sh add $ID '<title>' --kind $KIND) and re-run" >&2 exit 1 else echo "error: task $ID's backlog item could not be read before dispatch ($FM_BACKLOG_ROW_ERROR)" >&2 @@ -3026,19 +3141,79 @@ rovo_spawn_fail() { # <detail> rovo_endpoint_cleanup } -# No task record is ever published on this failure path, so nothing else -# (teardown, the watcher) will ever learn this endpoint exists to close it: -# without this, the already-launched --yolo rovo process keeps running as an -# orphaned autonomous agent outside task control. Mirrors fm-teardown.sh's own -# generic non-orca kill call; orca's worktree+terminal are owned by the -# separate ORCA_ABORT_CLEANUP trap path and are out of scope here. +# The launch-then-confirm gates run after the task record is published, when +# ORCA_ABORT_CLEANUP is already cleared and neither the abort trap nor a +# teardown owns this endpoint yet, so a gate failure must close the launched +# process here or it keeps running as an orphaned autonomous agent outside +# task control. Mirrors fm-teardown.sh's own generic kill call. On orca only +# the exact terminal is closed: that stops the CLI while its worktree stays +# for the record's own teardown, which owns worktree deletion. rovo_endpoint_cleanup() { - [ "$BACKEND" = orca ] && return 0 + if [ "$BACKEND" = orca ]; then + fm_backend_kill orca "$T" 2>/dev/null || true + return 0 + fi local tab_id= [ "$BACKEND" = zellij ] && tab_id=$ZELLIJ_TAB_ID fm_backend_kill "$BACKEND" "$T" "$tab_id" "fm-$ID" 2>/dev/null || true } +# agy carries its brief on the launch command, so it needs no delivery gate, +# but a worktree agy does not trust parks the TUI on the folder-trust dialog +# and an unanswered dialog sends the turn into agy's scratch directory instead +# of the worktree. The trust is pre-registered before launch +# (bin/fm-agy-trust-lib.sh), and this gate is the +# backstop in the rovo/kimi launch-then-confirm shape: answer the dialog once +# with the preselected safe default if it renders anyway, then require +# positive proof that the brief is being processed - the same verdict the +# supervisor reads (Herdr's native working state through fm_busy_classify) +# before the spawn reports success. +# The gate is strict about ordering because on Herdr the native working +# verdict is known to coexist with an unanswered dialog: a busy verdict counts +# only when the path was pre-registered or the dialog has been seen and +# answered; on an unregistered path it keeps polling for the dialog instead. +AGY_TRUST_DIALOG='Do you trust the contents of this project?' +AGY_TRUST_ANSWERED=0 + +agy_capture() { + fm_backend_capture "$BACKEND" "$T" 120 "$W" 2>/dev/null || true +} + +agy_pane_shows_trust_dialog() { # <plain-pane-capture> + printf '%s\n' "$1" | grep -Fq "$AGY_TRUST_DIALOG" +} + +agy_pane_is_working() { # <plain-pane-capture> + case "$(fm_busy_classify "$BACKEND" "$T" agy "$ID" "$STATE" "$1")" in + busy*) return 0 ;; + esac + return 1 +} + +agy_wait_for_working() { + local pane i=0 max=${FM_AGY_READY_POLLS:-60} interval=${FM_AGY_POLL_INTERVAL:-0.5} + while [ "$i" -lt "$max" ]; do + pane=$(agy_capture) + if agy_pane_shows_trust_dialog "$pane"; then + if [ "$AGY_TRUST_ANSWERED" -eq 0 ]; then + spawn_send_key "$T" Enter + AGY_TRUST_ANSWERED=1 + fi + elif [ "$AGY_TRUST_PREREGISTERED" -eq 1 ] || [ "$AGY_TRUST_ANSWERED" -eq 1 ]; then + agy_pane_is_working "$pane" && return 0 + fi + i=$((i + 1)) + [ "$i" -ge "$max" ] || sleep "$interval" + done + return 1 +} + +agy_spawn_fail() { # <detail> + printf 'failed: %s\n' "$1" >> "$STATE/$ID.status" + echo "error: $1; inspect window $T" >&2 + rovo_endpoint_cleanup +} + if [ "$RELAUNCH" -eq 1 ]; then # No worktree is acquired: the recorded one is reused as-is. What must be # proven instead is that the adopted endpoint's shell is actually sitting in @@ -3133,42 +3308,67 @@ elif [ "$KIND" != secondmate ] && [ "$BACKEND" != orca ]; then fi validate_spawn_worktree "treehouse get" "$T" + + # Claim the pool slot for this task. The interactive `treehouse get` sent to + # the pane above records only a process lease (Treehouse's durable + # `get --lease --lease-holder`, which bin/fm-home-seed.sh uses for secondmate + # homes, is not this path), so Treehouse cannot say which task a slot belongs + # to once that task's worker exits - and that is exactly when the slot is + # handed on and this task's worktree= line goes stale. The claim is what lets + # bin/fm-teardown.sh leave a slot that has since been reassigned untouched, so + # a slot that cannot be claimed is refused here, at the cheapest point, rather + # than launching a worker whose slot teardown could later release out from + # under its successor. + # Written under the Treehouse project lock held from before slot allocation + # through metadata publication, so no other spawn or return sees a half-claim. + if fm_treehouse_pool_slot "$PROJ_ABS" "$WT"; then + if ! fm_treehouse_slot_owner_claim "$WT" "$ID" "$FM_HOME"; then + echo "error: could not claim Treehouse pool slot $WT for task $ID; refusing to launch a worker whose slot cannot later be proved to be its own; inspect window $T" >&2 + exit 1 + fi + SPAWN_SLOT_CLAIMED=1 + fi fi if [ "$RELAUNCH" -eq 0 ] && [ "$KIND" != secondmate ]; then freshen_spawn_worktree_base "$WT" || exit 1 fi -# Pre-register Claude's workspace trust for the worktree, at the first point the -# worktree is known and before any per-task state is created below. The dialog -# gates the pane before the brief is ever read, and it also gates loading the -# project settings written further down, so nothing armed below takes effect -# without it. bin/fm-claude-trust.sh owns the structural scope test and refuses -# any path that is not this project's own isolated worktree; a refusal blocks the -# spawn rather than launching a worker that would wedge on a dialog firstmate -# cannot answer. Refusing here rather than beside the arm keeps this in the same -# class as the two worktree refusals just above: no temp root, no retired -# relaunch wiring and no busy record exists yet to strand, so the refusal names -# the endpoint the same way they do and leaves nothing else behind. -if [ "$KIND" != secondmate ]; then - case "$HARNESS" in - claude*) - # Resolve a relative explicit store against the spawning process, just - # as the evidence-store owner does later. The trust helper canonicalizes - # this absolute path; passing the relative spelling would instead refuse - # a config the worker's canonical launch already supports. - CLAUDE_TRUST_CONFIG_DIR=${CLAUDE_CONFIG_DIR:-} - case "$CLAUDE_TRUST_CONFIG_DIR" in - ''|/*) ;; - *) CLAUDE_TRUST_CONFIG_DIR="$PWD/$CLAUDE_TRUST_CONFIG_DIR" ;; - esac - if ! CLAUDE_CONFIG_DIR="$CLAUDE_TRUST_CONFIG_DIR" \ - "$FM_ROOT/bin/fm-claude-trust.sh" "$WT" "$PROJ_ABS" >/dev/null; then - echo "error: could not pre-register Claude workspace trust for $WT; refusing to launch a claude worker that would wedge on the trust dialog; inspect window $T" >&2 - exit 1 - fi - ;; - esac -fi +# Pre-register Claude's workspace trust for the directory this launch starts in, +# at the first point that directory is known and before any per-task state is +# created below. The dialog gates the pane before the brief is ever read, and it +# also gates loading the project settings written further down, so nothing armed +# below takes effect without it. EVERY claude launch needs it, a secondmate's +# included: its home is just as unseen by Claude as a fresh worktree, and +# skipping the step for that kind left a standalone-clone secondmate home with +# nothing registered and a pane wedged on a dialog firstmate cannot answer. +# bin/fm-claude-trust.sh owns the structural scope test for both shapes and +# refuses anything that is neither this project's own isolated worktree nor a +# seeded secondmate home marked for this id; a refusal blocks the spawn rather +# than launching a worker that would wedge. Refusing here rather than beside the +# arm keeps this in the same class as the two worktree refusals just above: no +# temp root, no retired relaunch wiring and no busy record exists yet to strand, +# so the refusal names the endpoint the same way they do and leaves nothing else +# behind. +AGY_TRUST_PREREGISTERED=0 +case "$HARNESS" in + claude*) + if [ "$KIND" = secondmate ]; then + spawn_trust_args=(--secondmate-home "$PROJ_ABS" "$ID") + else + spawn_trust_args=("$WT" "$PROJ_ABS") + fi + CLAUDE_TRUST_CONFIG_DIR=${CLAUDE_CONFIG_DIR:-} + case "$CLAUDE_TRUST_CONFIG_DIR" in + ''|/*) ;; + *) CLAUDE_TRUST_CONFIG_DIR="$PWD/$CLAUDE_TRUST_CONFIG_DIR" ;; + esac + if ! CLAUDE_CONFIG_DIR="$CLAUDE_TRUST_CONFIG_DIR" \ + "$FM_ROOT/bin/fm-claude-trust.sh" "${spawn_trust_args[@]}" >/dev/null; then + echo "error: could not pre-register Claude workspace trust for $WT; refusing to launch a claude worker that would wedge on the trust dialog; inspect window $T" >&2 + exit 1 + fi + ;; +esac # Per-task temp root: /tmp/fm-<id>/ with Go's build temp nested at gotmp/. Go won't # create GOTMPDIR, so mkdir before it is used; fm-teardown removes the whole root. @@ -3709,6 +3909,7 @@ EOF echo "error: could not grant agy workspace trust for $WT; refusing to launch an agy crewmate that would hang on the trust modal" >&2 exit 1 fi + AGY_TRUST_PREREGISTERED=1 if [ "${FM_AGY_TRUST_ADDED:-}" = created ]; then # Arm rollback BEFORE any fallible write: the global trust entry now # exists, so from this instant an abort must roll it back. Arming after @@ -3993,6 +4194,8 @@ LAUNCH=$(fm_launch_render \ "$LAUNCH" "$MODELFLAG" "$EFFORTFLAG" "$sq_brief" "$sq_turnend" \ "$sq_piext" "$sq_piturnend" "$sq_piwatch" "$sq_opinput" "$RAW_LAUNCH" \ __PIBIN__ "$(fm_launch_shell_quote "${PI_BIN:-}")" __PITUIMODE__ "${PI_TUI_MODE:-}" \ + __CLAUDEPERMFLAG__ "$CLAUDE_PERM_FLAG" \ + __AGYBIN__ "$(fm_launch_shell_quote "${AGY_BIN:-}")" \ __CURSORBIN__ "$(fm_launch_shell_quote "${CURSOR_BIN:-}")" \ __WORKTREE__ "$(fm_launch_shell_quote "$WT")" \ __OMPBIN__ "$(fm_launch_shell_quote "${OMP_BIN:-}")" \ @@ -4199,6 +4402,18 @@ if [ "$HARNESS" = rovo ]; then exit 1 fi fi +if [ "$HARNESS" = agy ]; then + if ! agy_wait_for_working; then + if [ "$AGY_TRUST_ANSWERED" -eq 1 ]; then + agy_spawn_fail "agy did not start processing its brief after the folder-trust dialog was answered in window $T" + elif [ "$AGY_TRUST_PREREGISTERED" -eq 1 ]; then + agy_spawn_fail "agy did not start processing its brief in the pre-trusted worktree in window $T" + else + agy_spawn_fail "agy never showed its folder-trust dialog on an unregistered worktree in window $T, so the brief could not be confirmed to run there" + fi + exit 1 + fi +fi if [ "$KIND" = secondmate ] && [ "${FM_SKIP_SECONDMATE_INHERIT:-0}" != 1 ]; then if ! fm_config_reread_discard_pending "$PROJ_ABS" "$ID" "$FM_HOME"; then if fm_config_reread_quarantine_pending "$PROJ_ABS" "$ID" "$FM_HOME"; then diff --git a/bin/fm-supervision-instructions.sh b/bin/fm-supervision-instructions.sh index 94316cfb03c..d5de85133a7 100755 --- a/bin/fm-supervision-instructions.sh +++ b/bin/fm-supervision-instructions.sh @@ -13,16 +13,19 @@ DOC_DIR="$REPO_ROOT/docs/supervision-protocols" HARNESS= READ_ONLY=0 AFK=0 +AFK_MODE=away X_MODE=0 REPAIR_LINE=0 QUEUE_PENDING=0 usage() { cat <<'EOF' -Usage: fm-supervision-instructions.sh [--harness <name>] [--read-only 0|1] [--afk 0|1] [--x-mode 0|1] [--repair-line] [--queue-pending 0|1] +Usage: fm-supervision-instructions.sh [--harness <name>] [--read-only 0|1] [--afk 0|1] [--afk-mode away|quiet] [--x-mode 0|1] [--repair-line] [--queue-pending 0|1] Print the current primary harness's supervision operating instructions. With --repair-line, print one concise repair instruction for guard and hook messages. +--afk-mode only matters when --afk 1 (present); it selects the away-mode vs +quiet-mode (kunchenguid/firstmate#2356) wording, and defaults to away. EOF } @@ -50,6 +53,14 @@ while [ "$#" -gt 0 ]; do AFK=$(bool_value "$2") shift 2 ;; + --afk-mode) + [ "$#" -gt 1 ] || { echo "error: --afk-mode requires away or quiet" >&2; exit 2; } + case "$2" in + away|quiet) AFK_MODE=$2 ;; + *) AFK_MODE=away ;; + esac + shift 2 + ;; --x-mode) [ "$#" -gt 1 ] || { echo "error: --x-mode requires 0 or 1" >&2; exit 2; } X_MODE=$(bool_value "$2") @@ -125,7 +136,11 @@ repair_line() { return 0 fi if [ "$AFK" -eq 1 ]; then - printf '%s\n' 'Away mode owns watcher supervision; load /afk and ensure the daemon is running instead of starting normal supervision directly.' + if [ "$AFK_MODE" = quiet ]; then + printf '%s\n' 'Quiet mode owns watcher supervision; load /quiet and ensure the daemon is running instead of starting normal supervision directly.' + else + printf '%s\n' 'Away mode owns watcher supervision; load /afk and ensure the daemon is running instead of starting normal supervision directly.' + fi return 0 fi @@ -210,9 +225,13 @@ else printf '%s\n' '- Lock: held by this session; this session owns normal supervision unless away mode says otherwise.' fi if [ "$AFK" -eq 1 ]; then - printf '%s\n' '- Away mode: active; load /afk and keep normal harness supervision paused while the daemon owns the watcher.' + if [ "$AFK_MODE" = quiet ]; then + printf '%s\n' '- Quiet mode: active; load /quiet and keep normal harness supervision paused while the daemon owns the watcher. Ordinary captain chat does NOT exit it - only an explicit /quiet off does.' + else + printf '%s\n' '- Away mode: active; load /afk and keep normal harness supervision paused while the daemon owns the watcher.' + fi else - printf '%s\n' '- Away mode: inactive.' + printf '%s\n' '- Away/quiet mode: inactive.' fi if [ "$X_MODE" -eq 1 ]; then printf '%s%s%s\n' '- X mode: active; source ' "$x_mode_env" ' before launching any watcher process so the 30s cadence is inherited.' diff --git a/bin/fm-tasks-axi.sh b/bin/fm-tasks-axi.sh new file mode 100755 index 00000000000..b8e2844c0e5 --- /dev/null +++ b/bin/fm-tasks-axi.sh @@ -0,0 +1,127 @@ +#!/usr/bin/env bash +# fm-tasks-axi.sh - run tasks-axi against THIS home's backlog from any working directory. +# +# Usage: fm-tasks-axi.sh [<tasks-axi command> [args...]] +# fm-tasks-axi.sh --help +# +# Every routine firstmate backlog read or mutation goes through this command +# rather than a bare `tasks-axi`; `fm-tasks-axi.sh <command> --help` prints +# tasks-axi's own help. Arguments reach tasks-axi as given, apart from one +# rewrite that keeps file arguments meaning what the caller meant: a relative +# value of `--to` or any `--*-file` flag (`--body-file`, `--relation-file`, ...) +# is made absolute against the caller's working directory, because tasks-axi +# starts from the backlog root instead. `--report` stays as given: tasks-axi +# stores it verbatim as a link, which lifecycle transitions record relative to +# that same root. +# +# Why it exists: a bare `tasks-axi` resolves the tracked `.tasks.toml` paths +# against its working directory, so from the code root it forks the queue +# whenever the home lives elsewhere; docs/configuration.md ("Backlog backend") +# owns that rationale. +# +# Addressing is bin/fm-backlog-transition-lib.sh's fm_backlog_tasks_axi_addressing, +# the same resolution the lifecycle transitions use: tasks-axi runs from the +# configured data directory's parent, so that home's own `.tasks.toml` (or +# tasks-axi's built-in defaults, which keep the archive beside the backlog) +# supplies the adapter, done_keep, and the archive path; a markdown backlog is +# additionally pinned to `<data>/backlog.md` through TASKS_AXI_FILE. The +# environment carries the pin rather than a trailing --file so the no-command +# dashboard works too. A configured non-markdown adapter is addressed by that +# root alone, so an inherited TASKS_AXI_FILE is cleared for it. +# +# The data directory is FM_DATA_OVERRIDE, else $FM_HOME/data, else the code +# root's data/ (FM_HOME unset keeps the single-home layout unchanged). +# +# Refusals (exit 2, nothing run): +# - tasks-axi missing from PATH; +# - a caller-supplied --file, because this command owns the addressing and +# tasks-axi would silently let the last --file win; +# - a data directory that cannot be resolved, or whose backend configuration +# cannot be read (bin/fm-tasks-axi-lib.sh owns that diagnostic); +# - a markdown `<data>/backlog.md` that is itself a symlink, because the +# first write would replace the link with a private copy, exactly the fork +# this command exists to prevent. Lifecycle transitions refuse the same file. +# Otherwise the exit status is tasks-axi's own. +set -u + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +FM_ROOT="${FM_ROOT_OVERRIDE:-$(cd "$SCRIPT_DIR/.." && pwd)}" +FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" +DATA="${FM_DATA_OVERRIDE:-$FM_HOME/data}" +# shellcheck source=bin/fm-tasks-axi-lib.sh disable=SC1091 +. "$SCRIPT_DIR/fm-tasks-axi-lib.sh" +# shellcheck source=bin/fm-backlog-transition-lib.sh disable=SC1091 +. "$SCRIPT_DIR/fm-backlog-transition-lib.sh" + +usage() { + awk ' + NR == 1 { next } + /^#/ { sub(/^# ?/, ""); print; next } + { exit } + ' "$0" +} + +fail() { + printf 'fm-tasks-axi: %s\n' "$*" >&2 + exit 2 +} + +case "${1:-}" in + -h|--help) + usage + exit 0 + ;; +esac + +CALLER_DIR=$(pwd) + +absolute_from_caller() { # <path-value> + case "$1" in + ''|-|/*) printf '%s' "$1" ;; + *) printf '%s/%s' "$CALLER_DIR" "$1" ;; + esac +} + +ARGS=() +path_value_next=0 +for arg in "$@"; do + if [ "$path_value_next" = 1 ]; then + ARGS+=("$(absolute_from_caller "$arg")") + path_value_next=0 + continue + fi + case "$arg" in + --file|--file=*) + fail "this command always addresses this home's backlog at $DATA; drop --file, or run tasks-axi directly for another backlog" + ;; + --to|--*-file) + ARGS+=("$arg") + path_value_next=1 + ;; + --to=*|--*-file=*) + ARGS+=("${arg%%=*}=$(absolute_from_caller "${arg#*=}")") + ;; + *) + ARGS+=("$arg") + ;; + esac +done + +command -v tasks-axi >/dev/null 2>&1 || fail "tasks-axi is not on PATH; run bin/fm-bootstrap.sh for the install command" + +FM_BACKLOG_TRANSITION_ERROR= +if ! fm_backlog_tasks_axi_addressing "$DATA"; then + fail "${FM_BACKLOG_TRANSITION_ERROR:-data directory cannot be resolved: $DATA}" +fi + +if [ -n "$FM_BACKLOG_AXI_FILE" ]; then + if [ -L "$FM_BACKLOG_AXI_FILE" ]; then + fail "$FM_BACKLOG_AXI_FILE is a symlink; a tasks-axi write would replace it with a regular file and fork the backlog - make it this home's real file" + fi + export TASKS_AXI_FILE="$FM_BACKLOG_AXI_FILE" +else + unset TASKS_AXI_FILE +fi + +cd "$FM_BACKLOG_AXI_ROOT" || fail "cannot enter the backlog root $FM_BACKLOG_AXI_ROOT" +exec tasks-axi ${ARGS[@]+"${ARGS[@]}"} diff --git a/bin/fm-teardown.sh b/bin/fm-teardown.sh index cd24de3612a..d7cedc5ce71 100755 --- a/bin/fm-teardown.sh +++ b/bin/fm-teardown.sh @@ -122,8 +122,36 @@ # cleanup step, teardown verifies record exclusivity: no OTHER task record in # this home or any locally registered Firstmate home may name the same live path # in its worktree= or home=. One live path with two task records is the reuse -# collision itself, whichever record is stale. The recorded endpoint's exact -# task identity and the record's spawn incarnation are validated separately +# collision itself, whichever record is stale. +# That scan alone cannot prove THIS record is the current owner, because the task +# that took the slot next may leave no record it can reach - its own worker may +# have exited and its record been cleaned up, or it may live in a home this +# machine does not register - which is how a released-then-reassigned slot was +# returned out from under a live worker (observed 2026-09-07). So teardown also +# reads the slot's own owner claim, written by bin/fm-spawn.sh at the moment the +# slot is taken and dropped here once it is genuinely returned; bin/fm-wake-lib.sh +# owns the claim, its location, and its states. A claim naming another task is +# proof of reassignment: the slot is no longer this task's, so teardown warns, +# names the claimant, and then finishes only this task's own cleanup - endpoint, +# status, records, checks, backlog - while every step that would read or touch +# that slot is skipped: no process kill under it, no dirty or landed-work +# inspection of it, no branch or hook removal in it, no Treehouse return, and +# never the other task's claim. Skipping the inspection discards nothing of this +# task's: whatever unlanded work it had in that slot was already destroyed when +# the pool handed the slot on. Refusing instead would strand the record, because +# bin/fm-backend.sh's endpoint validation refuses an empty or missing worktree= +# unconditionally, so there is no line an operator could clear to get past it. +# A claim that cannot be read proves nothing either way and refuses; inspect or +# repair the claim file at the printed path and re-run - never remove it, since +# an absent claim proceeds and would return a slot that may be another task's. An +# absent claim - a slot taken before claims existed, or already returned - keeps +# exactly the record-scan protection it had before, because refusing it would +# strand every task in flight across that change on no evidence at all. +# Why Treehouse's own state cannot answer this for crewmate slots, and why the +# claim file sits on top of it, is owned by bin/fm-wake-lib.sh's slot-owner +# claim comment. +# The recorded endpoint's exact task identity and the record's spawn incarnation +# are validated separately # before cleanup. Its current working directory is only incidental process # state: the same worker remains the owner after changing directory, so cwd can # never veto teardown of that exact recorded endpoint. @@ -136,9 +164,10 @@ # through metadata publication, closing the publication # gap; forced secondmate teardown takes it and runs the same checks for every # descendant Treehouse slot before touching any child. -# This refusal is not relaxed by --force: --force authorizes discarding THIS -# task's unlanded work, never another task's live work. Reconcile whichever -# record is wrong and re-run. Orca is not a pool slot and proves its path through +# These refusals are not relaxed by --force: --force authorizes discarding THIS +# task's unlanded work, never another task's live work. Nothing of this task's +# own is removed by a refusal; reconcile whichever record is wrong and re-run. +# Orca is not a pool slot and proves its path through # require_orca_worktree_path_match instead. # Orca tasks use the same safety checks, then close the recorded terminal and # remove the recorded worktree through `orca worktree rm`; teardown never guesses @@ -339,23 +368,6 @@ if [ "$FORCE" = --force ] && [ "$(fm_lease_actor)" = branch ]; then fi fm_lease_guard "$ID" "teardown (fm-teardown)" -# A Treehouse slot has the managed pool's fixed <pool>/<slot>/<repo> layout. -# Require both its pool state and the same Git common directory as the recorded -# project; an ordinary linked worktree is not evidence that Treehouse owns it. -is_treehouse_pool_slot() { # <project> <worktree> - local project=$1 worktree=$2 slot pool state project_common slot_common - [ -d "$project" ] && [ -d "$worktree" ] || return 1 - slot=$(CDPATH='' cd -- "$worktree" 2>/dev/null && pwd -P) || return 1 - pool=$(dirname "$(dirname "$slot")") - state="$pool/treehouse-state.json" - [ -f "$state" ] && [ ! -L "$state" ] || return 1 - project_common=$(git -C "$project" rev-parse --path-format=absolute --git-common-dir 2>/dev/null) || return 1 - slot_common=$(git -C "$slot" rev-parse --path-format=absolute --git-common-dir 2>/dev/null) || return 1 - project_common=$(CDPATH='' cd -- "$project_common" 2>/dev/null && pwd -P) || return 1 - slot_common=$(CDPATH='' cd -- "$slot_common" 2>/dev/null && pwd -P) || return 1 - [ "$project_common" = "$slot_common" ] -} - META="$STATE/$ID.meta" TREEHOUSE_PROJECT_LOCK= TREEHOUSE_PROJECT_LOCK_HELD=0 @@ -369,7 +381,7 @@ if [ -f "$META" ] && [ ! -L "$META" ]; then TEARDOWN_LOCK_PROJECT=$(fm_meta_get "$META" project) if [ "$TEARDOWN_LOCK_KIND" != secondmate ] \ && [ "$TEARDOWN_LOCK_BACKEND" != orca ] \ - && is_treehouse_pool_slot "$TEARDOWN_LOCK_PROJECT" "$TEARDOWN_LOCK_WT"; then + && fm_treehouse_pool_slot "$TEARDOWN_LOCK_PROJECT" "$TEARDOWN_LOCK_WT"; then TREEHOUSE_SLOT_LOCK_REQUIRED=1 TREEHOUSE_PROJECT_LOCK=$(fm_treehouse_project_lock_path "$TEARDOWN_LOCK_PROJECT") || { echo "REFUSED: cannot resolve the shared Treehouse project lock for ${TEARDOWN_LOCK_PROJECT:-<missing>}; nothing was changed" >&2 @@ -1241,7 +1253,7 @@ MODE=$(grep '^mode=' "$META" | cut -d= -f2- || true) TASK_RECORDED_BRANCH=$(fm_meta_get "$META" branch) EXPECTED_TREEHOUSE_PROJECT_LOCK= if [ "$KIND" != secondmate ] && [ "$BACKEND" != orca ] \ - && is_treehouse_pool_slot "$PROJ" "$WT"; then + && fm_treehouse_pool_slot "$PROJ" "$WT"; then EXPECTED_TREEHOUSE_PROJECT_LOCK=$(fm_treehouse_project_lock_path "$PROJ") || { echo "REFUSED: cannot resolve the shared Treehouse project lock for ${PROJ:-<missing>}; nothing was changed" >&2 exit 1 @@ -1732,7 +1744,7 @@ validate_pr_poll_cleanup() { fm_task_id_path_safe "$id" || return 0 for artifact in "$state_dir/$id.check.sh" "$state_dir/$id.pr-poll" \ "$state_dir/$id.pr-poll-registration" "$state_dir/$id.pr-poll-retirement" \ - "$state_dir/$id.check-trust"; do + "$state_dir/$id.merge-authority" "$state_dir/$id.check-trust"; do [ -e "$artifact" ] || [ -L "$artifact" ] || continue has_artifact=1 done @@ -1741,11 +1753,13 @@ validate_pr_poll_cleanup() { state_device=$(fm_pr_file_device "$state_dir") || return 1 for artifact in "$state_dir/$id.check.sh" "$state_dir/$id.pr-poll" \ "$state_dir/$id.pr-poll-registration" "$state_dir/$id.pr-poll-retirement" \ - "$state_dir/$id.check-trust"; do + "$state_dir/$id.merge-authority" "$state_dir/$id.check-trust"; do [ -e "$artifact" ] || [ -L "$artifact" ] || continue if [ ! -f "$artifact" ] || [ -L "$artifact" ] \ || [ "$(fm_pr_file_device "$artifact")" != "$state_device" ] \ - || [ "$(fm_pr_file_link_count "$artifact")" != 1 ]; then + || [ "$(fm_pr_file_link_count "$artifact")" != 1 ] \ + || { [ "$artifact" = "$state_dir/$id.merge-authority" ] \ + && [ "$(fm_pr_file_mode "$artifact")" != 600 ]; }; then echo "REFUSED: unsafe task PR-check artifact; preserving task state." >&2 return 1 fi @@ -1766,7 +1780,7 @@ remove_pr_poll_artifacts() { fm_pr_poll_merge_notified_remove "$state_dir" "$id" || return 1 rm -f "$state_dir/$id.check.sh" "$state_dir/$id.pr-poll" \ "$state_dir/$id.pr-poll-registration" "$state_dir/$id.pr-poll-retirement" \ - "$state_dir/$id.check-trust" || return 1 + "$state_dir/$id.merge-authority" "$state_dir/$id.check-trust" || return 1 } # Resolve the PR number for a worktree branch via gh-axi. Echoes the number on a @@ -1977,7 +1991,7 @@ backlog_refresh_reminder() { if [ "$BACKLOG_CLOSED" = 1 ] && [ "$BACKLOG_TRANSITION" = retain ]; then printf '%s\n' "Backlog: $ID stays open in $backlog_display, still held for the captain with its deliverable recorded. Relay the question and close it only with bin/fm-captain-hold.sh answer." elif [ "$BACKLOG_CLOSED" = 1 ]; then - printf '%s\n' "Backlog: $ID is closed in $backlog_display. Run tasks-axi ready for dependency-cleared candidates, check date gates, and dispatch only work whose blockers are gone and date is due." + printf '%s\n' "Backlog: $ID is closed in $backlog_display. Run bin/fm-tasks-axi.sh ready for dependency-cleared candidates, check date gates, and dispatch only work whose blockers are gone and date is due." else printf '%s\n' "Backlog: $ID just finished ($BACKLOG_SKIP_REASON). Update $backlog_display - move $ID to Done, keep Done to the 10 most recent, then re-scan Queued and dispatch only work whose blockers are gone and date is due." fi @@ -2633,7 +2647,7 @@ require_orca_worktree_path_match_if_present() { # record with nothing live to return skips them rather than refusing. teardown_live_slot_path() { [ "$KIND" != secondmate ] || return 1 - is_treehouse_pool_slot "$PROJ" "$WT" || return 1 + fm_treehouse_pool_slot "$PROJ" "$WT" || return 1 canonical_existing_dir "$WT" } @@ -2713,6 +2727,72 @@ require_exclusive_task_worktree_slot() { require_exclusive_worktree_slot_record "$META" "$ID" "$STATE" "$slot" } +# Positive slot ownership, read from the claim the task that took the slot wrote +# into the slot itself (bin/fm-wake-lib.sh owns the claim and its states). +# +# The record scan above proves that no OTHER task record names this slot. It +# cannot prove that THIS record is not the stale one, because the task that took +# the slot next may leave no record this scan can reach: its own worker may have +# exited and its record been cleaned up, or it may belong to a home this machine +# does not register. The claim closes that gap from the other side - it names the +# task that actually took the slot, and it is written under the same project lock +# that allocates it - so a claim naming another task is proof the slot was +# reassigned after this record was written. +# +# A claim naming another task does not refuse: it means the slot is no longer +# this task's, so the record's own cleanup proceeds and every slot step is +# skipped (see the script header for why refusing would strand the record and +# why skipping discards nothing). Returns TEARDOWN_SLOT_REASSIGNED_RC for that +# state so each caller gates its slot steps on one determination; the claimant +# stays in FM_TREEHOUSE_SLOT_OWNER_ID and FM_TREEHOUSE_SLOT_OWNER_HOME. +# +# An absent claim proceeds as the slot's owner: a slot taken before claims +# existed, or already returned to the pool, carries none, and refusing those +# would strand every task in flight across the change for no evidence at all. +# Those keep exactly the record-scan protection they had before. +TEARDOWN_SLOT_REASSIGNED_RC=3 +require_owned_worktree_slot_record() { # <task-id> <worktree> + local record_id=$1 worktree=$2 marker + fm_treehouse_slot_owner_state "$worktree" "$record_id" + case "$FM_TREEHOUSE_SLOT_OWNER" in + mine|absent) return 0 ;; + other) + echo "warning: task $record_id's recorded worktree $worktree was reassigned to task $FM_TREEHOUSE_SLOT_OWNER_ID${FM_TREEHOUSE_SLOT_OWNER_HOME:+ (home $FM_TREEHOUSE_SLOT_OWNER_HOME)}, which claimed that pool slot after this record was written; that slot is no longer $record_id's, so its processes, copy, and claim are left untouched and only $record_id's own cleanup runs." >&2 + return "$TEARDOWN_SLOT_REASSIGNED_RC" + ;; + esac + marker=$(fm_treehouse_slot_owner_marker "$worktree" 2>/dev/null) || marker="beside $worktree" + echo "REFUSED: task $record_id's recorded worktree $worktree carries a slot-owner claim that cannot be read, so the slot cannot be proved to still be this task's; nothing was changed - not even with --force." >&2 + echo "Inspect or repair the claim file at $marker (task= and home= lines), then re-run teardown." >&2 + return 1 +} + +# The one ownership determination for this task's recorded slot. Every later +# step that would read or touch $WT consults teardown_owns_worktree, so a +# reassigned slot is skipped consistently rather than by each step's own guess. +TEARDOWN_SLOT_REASSIGNED=0 +TEARDOWN_SLOT_REASSIGNED_TO= +TEARDOWN_SLOT_REASSIGNED_HOME= +require_owned_task_worktree_slot() { + local slot rc=0 + slot=$(teardown_live_slot_path) || return 0 + require_owned_worktree_slot_record "$ID" "$slot" || rc=$? + case "$rc" in + 0) return 0 ;; + "$TEARDOWN_SLOT_REASSIGNED_RC") + TEARDOWN_SLOT_REASSIGNED=1 + TEARDOWN_SLOT_REASSIGNED_TO=$FM_TREEHOUSE_SLOT_OWNER_ID + TEARDOWN_SLOT_REASSIGNED_HOME=$FM_TREEHOUSE_SLOT_OWNER_HOME + return 0 + ;; + esac + return 1 +} + +teardown_owns_worktree() { + [ "$TEARDOWN_SLOT_REASSIGNED" != 1 ] +} + firstmate_home_has_treehouse_slot() { local home=$1 worktree_registered_for_project "$FM_ROOT" "$home" @@ -3180,7 +3260,7 @@ preflight_descendant_task_locks() { } preflight_descendant_treehouse_slots() { - local i state task_id meta kind backend target worktree project lock_path held + local i state task_id meta kind backend target worktree project lock_path held owner_rc for ((i=0; i < ${#DESCENDANT_TASK_IDS[@]}; i++)); do state=${DESCENDANT_TASK_STATES[$i]} task_id=${DESCENDANT_TASK_IDS[$i]} @@ -3193,7 +3273,7 @@ preflight_descendant_treehouse_slots() { if [ "$kind" = secondmate ] || [ "$backend" = orca ]; then continue fi - if ! is_treehouse_pool_slot "$project" "$worktree"; then + if ! fm_treehouse_pool_slot "$project" "$worktree"; then continue fi lock_path=$(fm_treehouse_project_lock_path "$project") || { @@ -3202,7 +3282,7 @@ preflight_descendant_treehouse_slots() { } held=0 [ "$TREEHOUSE_PROJECT_LOCK_HELD" != 1 ] || [ "$TREEHOUSE_PROJECT_LOCK" != "$lock_path" ] || held=1 - for target in "${DESCENDANT_TREEHOUSE_LOCK_PATHS[@]}"; do + for target in "${DESCENDANT_TREEHOUSE_LOCK_PATHS[@]+"${DESCENDANT_TREEHOUSE_LOCK_PATHS[@]}"}"; do [ "$target" != "$lock_path" ] || held=1 done if [ "$held" = 0 ]; then @@ -3226,11 +3306,17 @@ preflight_descendant_treehouse_slots() { if [ "$kind" = secondmate ] || [ "$backend" = orca ]; then continue fi - if ! is_treehouse_pool_slot "$project" "$worktree"; then + if ! fm_treehouse_pool_slot "$project" "$worktree"; then continue fi fm_backend_validate_task_endpoint "$meta" "$task_id" || return 1 require_exclusive_worktree_slot_record "$meta" "$task_id" "$state" "$worktree" || return 1 + owner_rc=0 + require_owned_worktree_slot_record "$task_id" "$worktree" || owner_rc=$? + case "$owner_rc" in + 0|"$TEARDOWN_SLOT_REASSIGNED_RC") ;; + *) return 1 ;; + esac done } @@ -3402,7 +3488,7 @@ preflight_firstmate_home_herdr_children() { # <home> } cleanup_firstmate_home_children() { - local home=$1 sub_state child_meta child_id child_t child_wt child_proj child_kind child_home child_backend child_orca_worktree_id child_return_rc child_busy_gen child_nativeturnend_key child_model_output + local home=$1 sub_state child_meta child_id child_t child_wt child_proj child_kind child_home child_backend child_orca_worktree_id child_return_rc child_busy_gen child_nativeturnend_key child_model_output child_owner_rc sub_state="$home/state" [ -d "$sub_state" ] || return 0 for child_meta in "$sub_state"/*.meta; do @@ -3470,22 +3556,36 @@ cleanup_firstmate_home_children() { fi fm_backend_remove_worktree "$child_backend" "$child_orca_worktree_id" || return 1 elif [ -n "$child_wt" ] && [ -d "$child_wt" ]; then - validate_child_worktree_for_removal "$child_wt" "$child_proj" >/dev/null || return 1 - rm -f "$child_wt/.claude/settings.local.json" "$child_wt/.opencode/plugins/fm-turn-end.js" \ - "$child_wt/.opencode/plugins/fm-busy-state.js" \ - "$child_wt/.fm-grok-turnend" "$child_wt/.fm-kimi-turnend" - if [ -n "$child_proj" ] && [ -d "$child_proj" ] && command -v treehouse >/dev/null 2>&1; then - if teardown_treehouse_return "$child_wt" "$child_proj" "child worktree"; then - : - else - child_return_rc=$? - if [ "$child_return_rc" -eq "$TEARDOWN_TREEHOUSE_LOCK_REFUSED" ]; then - return "$child_return_rc" + # The same ownership determination as the parent's own slot: a child + # slot reassigned to another task is not this child's to kill, reset, + # or return, so only its records are cleaned up. The preflight above + # already named the reassignment on stderr under the same lock. + child_owner_rc=0 + if fm_treehouse_pool_slot "$child_proj" "$child_wt"; then + require_owned_worktree_slot_record "$child_id" "$child_wt" 2>/dev/null || child_owner_rc=$? + fi + if [ "$child_owner_rc" -eq "$TEARDOWN_SLOT_REASSIGNED_RC" ]; then + : + elif [ "$child_owner_rc" -ne 0 ]; then + require_owned_worktree_slot_record "$child_id" "$child_wt" || return 1 + else + validate_child_worktree_for_removal "$child_wt" "$child_proj" >/dev/null || return 1 + rm -f "$child_wt/.claude/settings.local.json" "$child_wt/.opencode/plugins/fm-turn-end.js" \ + "$child_wt/.opencode/plugins/fm-busy-state.js" \ + "$child_wt/.fm-grok-turnend" "$child_wt/.fm-kimi-turnend" + if [ -n "$child_proj" ] && [ -d "$child_proj" ] && command -v treehouse >/dev/null 2>&1; then + if teardown_treehouse_return "$child_wt" "$child_proj" "child worktree"; then + fm_treehouse_slot_owner_release "$child_wt" "$child_id" + else + child_return_rc=$? + if [ "$child_return_rc" -eq "$TEARDOWN_TREEHOUSE_LOCK_REFUSED" ]; then + return "$child_return_rc" + fi + safe_rm_rf_child_worktree "$child_wt" "$child_proj" fi + else safe_rm_rf_child_worktree "$child_wt" "$child_proj" fi - else - safe_rm_rf_child_worktree "$child_wt" "$child_proj" fi fi fi @@ -3531,6 +3631,7 @@ remove_secondmate_registry_entry() { } require_exclusive_task_worktree_slot || exit 1 +require_owned_task_worktree_slot || exit 1 validate_pr_poll_cleanup "$STATE" "$ID" || exit 1 @@ -3616,7 +3717,7 @@ if [ "$FORCE" != "--force" ] \ "$SCRIPT_DIR/fm-public-followup.sh" guard-work "$PUBLIC_FOLLOWUP_WORK_HOME" "$ID" 2>/dev/null); then echo "REFUSED: task $ID still owes a public reply through the myfirstmate relay." >&2 printf '%s\n' "$PUBLIC_FOLLOWUP_BLOCKING" >&2 - echo "Deliver it with bin/fm-public-followup.sh deliver <obligation-id>, waive it with tasks-axi public-followup waive, or use --force after explicit discard approval." >&2 + echo "Deliver it with bin/fm-public-followup.sh deliver <obligation-id>, waive it with bin/fm-tasks-axi.sh public-followup waive, or use --force after explicit discard approval." >&2 exit 1 fi fi @@ -3647,7 +3748,7 @@ if [ "$BACKEND" = orca ] && [ "$KIND" != scout ] && [ "$KIND" != secondmate ] && ORCA_PATH_MATCH_VERIFIED=1 fi -if [ -d "$WT" ] && [ "$FORCE" != "--force" ]; then +if teardown_owns_worktree && [ -d "$WT" ] && [ "$FORCE" != "--force" ]; then if validate_worktree_teardown_safety; then : else @@ -3767,9 +3868,11 @@ fi # kind=secondmate: a secondmate home's own runtime lifecycle is owned by the # dedicated process-event and firstmate-home removal machinery further below, # not by task-worktree cleanup. -if [ "$KIND" != secondmate ]; then +if [ "$KIND" != secondmate ] && teardown_owns_worktree; then conclude_task_no_mistakes_run "$WT" reap_task_worktree_processes worktree "$WT" "$TASK_TMP" +elif [ "$KIND" != secondmate ]; then + reap_task_worktree_processes tasktmp "$TASK_TMP" fi # Fix 3 (see script header): sweep remote job workers abandoned by an already @@ -3838,7 +3941,7 @@ remove_owned_agy_trust "$STATE" "$ID" "task $ID" || exit 1 # Detach only this task's conventional branch before removing its worktree. # The exact branch is reaped afterwards, when the merge proof prepared above # still agrees and no worktree can need it. -if [ "$NO_VERDICT_RETAIN_WORKTREE" -eq 1 ]; then +if [ "$NO_VERDICT_RETAIN_WORKTREE" -eq 1 ] || ! teardown_owns_worktree; then : elif [ "$BACKEND" = orca ] && [ "$KIND" != secondmate ]; then if [ "$ORCA_PATH_MATCH_VERIFIED" != 1 ]; then @@ -3870,6 +3973,7 @@ elif [ -d "$WT" ] && [ "$KIND" != secondmate ]; then echo "error: treehouse return failed for worktree $WT; teardown aborted" >&2 exit 1 } + fm_treehouse_slot_owner_release "$WT" "$ID" fi reap_task_branch || true @@ -4035,7 +4139,9 @@ if [ -d "$STATE" ]; then fi if [ "$TEARDOWN_LEGACY_ACCEPTED" = 1 ]; then echo "teardown $ID complete (window $T, worktree $WT, legacy record accepted without spawn_gen: endpoint $TEARDOWN_LEGACY_ENDPOINT, incarnation $TEARDOWN_META_SPAWN_GEN)" -else +elif teardown_owns_worktree; then echo "teardown $ID complete (window $T, worktree $WT)" +else + echo "teardown $ID complete (window $T; pool slot $WT left to task $TEARDOWN_SLOT_REASSIGNED_TO${TEARDOWN_SLOT_REASSIGNED_HOME:+ (home $TEARDOWN_SLOT_REASSIGNED_HOME)}, which it was reassigned to)" fi backlog_refresh_reminder diff --git a/bin/fm-test-run.sh b/bin/fm-test-run.sh index fc0134f7098..660de20e9c2 100755 --- a/bin/fm-test-run.sh +++ b/bin/fm-test-run.sh @@ -18,6 +18,7 @@ # fm-test-run.sh --list --family <name> # fm-test-run.sh --list --lane portable-parallel-1 # fm-test-run.sh --list-scheduled --family <name> +# fm-test-run.sh --list-scheduled --lane portable-parallel-1 # fm-test-run.sh --list-families # fm-test-run.sh --list-concurrent-safe-families # fm-test-run.sh --concurrent-safe-family-jobs-max <name> @@ -35,7 +36,11 @@ # tool this host could not exercise. # --list print selected script paths (one per line) and exit 0 # --list-scheduled -# print selected paths longest-hint-first and exit 0 +# print selected paths longest-hint-first and exit 0. +# Only --lane portable-parallel-1 or portable-parallel-2 uses +# parallel hints, falling back to serial weights if missing. +# Every other selection uses serial weights alone. +# Equal weights are ordered by path under LC_ALL=C. # --base <ref> with --changed, compare against this ref (default: origin/main) # --exclude-family <name> # drop scripts whose primary family matches <name> after selection @@ -59,7 +64,7 @@ # family proofs may impose a lower cap. Individually proven # scripts share one phase; scripts admitted only by a family # proof run in a separate phase for each family. Concurrent -# phases are ordered longest-hint-first. Unproven stateful +# phases use serial weights, longest-hint-first. Unproven stateful # scripts run serially after all concurrent phases. Default is # 1 (serial) except for plain --changed and a plain list of # script paths, which use the bounded automatic scheduler. @@ -118,10 +123,20 @@ # live-capability (a live-harness guard governed by fm_live_gate, which records # unavailable tools and explicit policy skips; see tests/lib.sh), or none. # +# Every selected script runs isolated from the host's global and system Git +# configuration, including one that sources no test helper of its own; +# tests/git-config-helpers.sh owns that contract and its limits. +# # Family labels, the changed-file map, and production portable-shard composition # live in this script only (one owner). The proven-isolated candidate set remains # owned by bin/fm-test-isolation-proof.sh; portable parallel shards are a -# duration-balanced partition of that exact set (see docs/fm-test-portable-shards.md). +# duration-balanced partition of that exact set, packed from the measured hints +# in portable_parallel_weight_hints (see docs/fm-test-portable-shards.md). +# --check-coverage reports parallel_max_ms (the larger lane hint sum), +# parallel_imbalance_ms (the absolute difference between the sums), and +# parallel_unhinted (the number of members missing a parallel hint). +# These sums exclude unhinted members and are estimates, not measured job wall +# times. Missing parallel hints are reported without failing this guard. # # portable-serial stays strictly serial. Its CI shards (portable-serial-<k>of<n>) # split it across separate runners, so two of its stateful scripts still never @@ -190,10 +205,11 @@ PER_SCRIPT_TIMEOUT_SET= # with about nine minutes left for setup and runner variability. The script # bound stays unchanged; the serial job adopts upstream's 30-minute cap. # real-Herdr is 420s less its ~35s mean slot plus 480s, inside its 1200s step cap. -# portable-parallel is tighter: CI runs that lane serially, so ~545s less its -# ~45s mean slot plus 480s is ~980s, already past its 600s job cap before setup. -# That lane can therefore lose per-script attribution to a job cancellation. -# Raising the script bound would not remedy that enclosing-job limit. +# Portable-parallel CI overlaps the admitted isolated scripts, so its serial +# hint sum is no longer a job wall-time estimate. A late-starting hung script +# can still exhaust the enclosing job cap before its own bound, losing the +# artifact to job cancellation. The workflow owns worker count and job caps; +# raising the script bound would not remedy that enclosing-job limit. # # It is a guard, not a speed control: a HUNG script becomes a bounded failure # instead of an unbounded suite, which is the shape that silently outruns a @@ -303,7 +319,8 @@ family_for_basename() { fm-composer-ghost.test.sh|fm-composer-lib.test.sh|\ fm-crew-state.test.sh|fm-captain-hold-lifecycle.test.sh|fm-design-skills.test.sh|fm-model-verify.test.sh|\ fm-documentation-audiences.test.sh|fm-ensure-agents-md.test.sh|fm-grok-harness.test.sh|\ - fm-kimi-harness.test.sh|fm-muse-harness.test.sh|fm-rovo-harness.test.sh|fm-omp-harness.test.sh|fm-herdr-lab.test.sh|fm-lint.test.sh|\ + fm-harness-precedence.test.sh|\ + fm-kimi-harness.test.sh|fm-muse-harness.test.sh|fm-rovo-harness.test.sh|fm-agy-harness.test.sh|fm-omp-harness.test.sh|fm-herdr-lab.test.sh|fm-lint.test.sh|\ fm-lint-workflows.test.sh|\ fm-operational-input.test.sh|fm-pi-primary-types.test.sh|\ fm-harness-adapter-references.test.sh|\ @@ -338,7 +355,7 @@ family_for_basename() { fm-backend-herdr-focus-flash-e2e.test.sh|\ fm-backend-herdr-stale-active-tab-e2e.test.sh|\ fm-backend-herdr-agent-exit-shell-e2e.test.sh|\ - fm-herdr-session-cleanup-e2e.test.sh|\ + fm-herdr-attached-viewer-live-e2e.test.sh|fm-herdr-session-cleanup-e2e.test.sh|\ fm-backend-herdr-smoke.test.sh|fm-backend-herdr-workspace-per-home-e2e.test.sh|\ fm-control-herdr-smoke.test.sh) printf '%s\n' real-herdr-gated @@ -377,9 +394,10 @@ family_for_basename() { fm-agy-smoke.test.sh|fm-cursor-primary-live-e2e.test.sh|\ fm-grok-stop-live-e2e.test.sh|fm-cmux-claude-composer-live-e2e.test.sh|\ fm-harness-adapter-instructions-live-e2e.test.sh|\ - fm-herdr-version-floor-live-e2e.test.sh|fm-muse-signals-live-e2e.test.sh|\ fm-harness-liveness-drift-live-e2e.test.sh|\ - fm-rovo-signals-live-e2e.test.sh|\ + fm-muse-signals-live-e2e.test.sh|fm-rovo-signals-live-e2e.test.sh|fm-agy-signals-live-e2e.test.sh|\ + fm-herdr-version-floor-live-e2e.test.sh|\ + fm-herdr-pi-stale-registration-live-e2e.test.sh|\ fm-opencode-primary-live-e2e.test.sh|fm-pi-branch-live-e2e.test.sh|\ fm-pi-branch-responsiveness-live-e2e.test.sh|\ fm-pi-hung-delivery-herdr-e2e.test.sh|fm-pi-prompt-collision-live-e2e.test.sh|\ @@ -522,41 +540,87 @@ tests/fm-x-mode.test.sh EOF } -# Portable parallel shard 1: LPT balance of the proven-isolated set using the -# measured CI durations recorded in docs/fm-test-portable-shards.md. -# Execution order is longest first so wall-clock stays near the balanced sum. +# Per-script serial CI duration hints, one "<path> <ms>" per line, used to +# pack only the two portable parallel lanes. Measurement provenance and the +# refresh procedure are owned by docs/fm-test-portable-shards.md. +portable_parallel_weight_hints() { + cat <<'EOF' +tests/fm-arm-pretool-check.test.sh 30898 +tests/fm-backend-herdr.test.sh 27380 +tests/fm-brief.test.sh 23099 +tests/fm-captain-hold-lifecycle.test.sh 305369 +tests/fm-cd-pretool-check.test.sh 16964 +tests/fm-composer-ghost.test.sh 2120 +tests/fm-composer-lib.test.sh 4798 +tests/fm-crew-state.test.sh 41660 +tests/fm-ensure-agents-md.test.sh 906 +tests/fm-grok-harness.test.sh 6983 +tests/fm-herdr-lab.test.sh 9800 +tests/fm-lint.test.sh 209643 +tests/fm-pi-primary-types.test.sh 8624 +tests/fm-pr-merge.test.sh 198766 +tests/fm-review-diff.test.sh 3832 +tests/fm-send-popup-settle.test.sh 4939 +tests/fm-send-settle.test.sh 2051 +tests/fm-send-strict.test.sh 7584 +tests/fm-spawn-batch.test.sh 2505 +tests/fm-supervision-instructions.test.sh 342 +tests/fm-test-run.test.sh 157536 +tests/fm-tmux-submit-busy.test.sh 2477 +tests/fm-transition-lib.test.sh 171 +tests/fm-x-mode.test.sh 31870 +EOF +} + +# Sum the hints above for the scripts read on stdin, and report how many of +# them had no hint at all, as "<summed_ms> <unhinted_count>". +portable_parallel_lane_weight() { + awk ' + NR == FNR { if (NF) { hint[$1] = $2 } ; next } + NF { + if ($1 in hint) { total += hint[$1] } else { unhinted++ } + } + END { printf "%d %d\n", total + 0, unhinted + 0 } + ' <(portable_parallel_weight_hints) - +} + +# Portable parallel shard 1: LPT balance of the proven-isolated set over the +# hints above. Stored order agrees with this lane's --list-scheduled output. +# tests/fm-pi-primary-types.test.sh belongs to this lane because +# this is the parallel job that installs the Pi package; moving it needs that +# workflow step moved with it. list_portable_parallel_1() { cat <<'EOF' -tests/fm-captain-hold-lifecycle.test.sh -tests/fm-test-run.test.sh +tests/fm-lint.test.sh +tests/fm-pr-merge.test.sh +tests/fm-crew-state.test.sh tests/fm-x-mode.test.sh -tests/fm-brief.test.sh -tests/fm-send-strict.test.sh -tests/fm-grok-harness.test.sh -tests/fm-send-popup-settle.test.sh +tests/fm-backend-herdr.test.sh +tests/fm-cd-pretool-check.test.sh tests/fm-pi-primary-types.test.sh +tests/fm-send-popup-settle.test.sh +tests/fm-composer-lib.test.sh tests/fm-spawn-batch.test.sh tests/fm-composer-ghost.test.sh tests/fm-ensure-agents-md.test.sh -tests/fm-transition-lib.test.sh EOF } # Portable parallel shard 2: the complementary LPT half of the proven set. list_portable_parallel_2() { cat <<'EOF' -tests/fm-lint.test.sh -tests/fm-pr-merge.test.sh -tests/fm-crew-state.test.sh +tests/fm-captain-hold-lifecycle.test.sh +tests/fm-test-run.test.sh tests/fm-arm-pretool-check.test.sh -tests/fm-backend-herdr.test.sh -tests/fm-cd-pretool-check.test.sh +tests/fm-brief.test.sh tests/fm-herdr-lab.test.sh -tests/fm-composer-lib.test.sh +tests/fm-send-strict.test.sh +tests/fm-grok-harness.test.sh tests/fm-review-diff.test.sh tests/fm-tmux-submit-busy.test.sh tests/fm-send-settle.test.sh tests/fm-supervision-instructions.test.sh +tests/fm-transition-lib.test.sh EOF } @@ -655,6 +719,8 @@ tests/fm-afk-pi-herdr-return-e2e.test.sh 100 tests/fm-afk-return.test.sh 1898 tests/fm-agents-hard-rules.test.sh 422 tests/fm-agy-adapter.test.sh 25751 +tests/fm-agy-harness.test.sh 11000 +tests/fm-agy-signals-live-e2e.test.sh 23 tests/fm-agy-smoke.test.sh 136 tests/fm-agy-trust-lib.test.sh 15755 tests/fm-ask-user-authority.test.sh 223 @@ -728,6 +794,7 @@ tests/fm-guard-stale-banner.test.sh 32981 tests/fm-harness-adapter-instructions-live-e2e.test.sh 107 tests/fm-harness-adapter-references.test.sh 118 tests/fm-harness-liveness-drift-live-e2e.test.sh 856 +tests/fm-herdr-attached-viewer-live-e2e.test.sh 19000 tests/fm-herdr-session-cleanup.test.sh 6858 tests/fm-herdr-submit-confirm-live-e2e.test.sh 118 tests/fm-herdr-version-floor-live-e2e.test.sh 93 @@ -849,7 +916,7 @@ tests/fm-watch-arm.test.sh 70359 tests/fm-watch-checkpoint.test.sh 6383 tests/fm-watch-recovery-loop.test.sh 58731 tests/fm-watch-triage-waits.test.sh 292716 -tests/fm-watch-triage.test.sh 236467 +tests/fm-watch-triage.test.sh 262626 tests/fm-watcher-lock.test.sh 88554 EOF } @@ -866,6 +933,16 @@ portable_serial_unhinted() { rm -rf "$tmp" } +portable_parallel_weight_for() { + local want=$1 ms + ms=$(portable_parallel_weight_hints | awk -v want="$want" '$1 == want { print $2; exit }') + if [ -n "$ms" ]; then + printf '%s\n' "$ms" + return 0 + fi + portable_serial_weight_for "$want" +} + portable_serial_weight_for() { local want=$1 path ms while read -r path ms; do @@ -1000,6 +1077,7 @@ select_lane() { # makes `comm` emit a bogus warning and either exit 1 or return a wrong set diff. run_coverage_guard() { local tmp missing extra a b shard unhinted serial_total + local p1_ms p1_unhinted p2_ms p2_unhinted parallel_max_ms parallel_imbalance_ms local -a saved_scripts=() tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-coverage.XXXXXX") @@ -1130,9 +1208,21 @@ run_coverage_guard() { fi fi - printf 'FM_TEST_COVERAGE ok total=%s parallel=%s serial=%s serial_shards=%s serial_unhinted=%s herdr=%s\n' \ + # Keep these estimates derived from the membership and hint owners; see the + # header for the distinction between packed weights and measured job time. + read -r p1_ms p1_unhinted <<<"$(list_portable_parallel_1 | portable_parallel_lane_weight)" + read -r p2_ms p2_unhinted <<<"$(list_portable_parallel_2 | portable_parallel_lane_weight)" + parallel_max_ms=$p1_ms + [ "$p2_ms" -le "$parallel_max_ms" ] || parallel_max_ms=$p2_ms + parallel_imbalance_ms=$((p1_ms - p2_ms)) + [ "$parallel_imbalance_ms" -ge 0 ] || parallel_imbalance_ms=$((-parallel_imbalance_ms)) + + printf 'FM_TEST_COVERAGE ok total=%s parallel=%s parallel_max_ms=%s parallel_imbalance_ms=%s parallel_unhinted=%s serial=%s serial_shards=%s serial_unhinted=%s herdr=%s\n' \ "$(wc -l <"$tmp/all" | tr -d ' ')" \ "$(wc -l <"$tmp/shards_union" | tr -d ' ')" \ + "$parallel_max_ms" \ + "$parallel_imbalance_ms" \ + "$((p1_unhinted + p2_unhinted))" \ "$(wc -l <"$tmp/serial" | tr -d ' ')" \ "$PORTABLE_SERIAL_SHARDS" \ "$unhinted" \ @@ -1264,12 +1354,14 @@ select_family() { [ "$found" -eq 1 ] || die "no tests mapped to family '$want'" } -families_for_test_reference() { - local needle=$1 s +families_for_test_reference() { # <needle>... + local s needle local found=0 + local -a needles=() + for needle in "$@"; do needles+=(-e "$needle"); done while IFS= read -r s; do [ -n "$s" ] || continue - if grep -Fq "$needle" "$s"; then + if grep -Fq "${needles[@]}" "$s"; then family_for_basename "$(basename "$s")" found=1 fi @@ -1346,12 +1438,15 @@ families_for_changed_path() { # resolution in the caller; emit a marker family of __script__ printf '%s\n' "__script__:$(basename "$path")" ;; - bin/fm-test-run.sh|bin/fm-test-isolation-proof.sh) + bin/fm-test-run.sh) # Deliberately the WHOLE family, not just the two contract tests. This # runner executes every pure-contract-unit script, so a change to it is # only proven by running them: its own contract test passing says the # runner's logic is right, not that the suite it drives still runs. printf '%s\n' pure-contract-unit + # Only this script wraps each suite in run_script_bounded's fixture Git + # isolation, and only a standalone-family script proves it. + printf '%s\n' "__script__:fm-test-fixtures.test.sh" ;; # The two pointer checks share one convention (docs/one-owner.md) and split # its surface by class, so a change to either belongs with both suites. @@ -1359,7 +1454,13 @@ families_for_changed_path() { printf '%s\n' '__script__:fm-pointer-check.test.sh' printf '%s\n' '__script__:fm-documentation-audiences.test.sh' ;; - bin/backends/herdr*|bin/fm-herdr-lab.sh|tests/herdr-test-safety.sh) + bin/fm-test-isolation-proof.sh) + # Same reason as the runner above: the proof drives every + # pure-contract-unit script. It runs each candidate directly, never + # through run_script_bounded, so it cannot regress fixture Git isolation. + printf '%s\n' pure-contract-unit + ;; + bin/backends/herdr*|bin/fm-herdr-lab.sh|tests/herdr-test-safety.sh|tests/herdr-client-pair-fixture.sh) printf '%s\n' real-herdr-gated printf '%s\n' backend-dispatch printf '%s\n' pure-contract-unit @@ -1385,6 +1486,13 @@ families_for_changed_path() { printf '%s\n' backend-dispatch printf '%s\n' real-herdr-gated ;; + bin/fm-agent-process-lib.sh) + # The shared harness-process classifier feeds both the tmux and Herdr + # liveness verdicts, so a change to it is proven by both backends' suites. + printf '%s\n' backend-dispatch + printf '%s\n' real-herdr-gated + printf '%s\n' pure-contract-unit + ;; bin/fm-watch*|bin/fm-wake*|bin/fm-inactive-reconcile.sh|\ bin/fm-classify-lib.sh|bin/fm-daemon*|bin/fm-turnend-guard*|bin/fm-guard.sh) printf '%s\n' watcher-wake-lock @@ -1454,12 +1562,16 @@ families_for_changed_path() { bin/fm-stow-cascade.sh) printf '%s\n' secondmate ;; - bin/fm-session-start.sh|bin/fm-bootstrap.sh|bin/fm-fleet-sync.sh|\ + bin/fm-session-start.sh|bin/fm-fleet-sync.sh|\ bin/fm-run-attribution-legacy-transition.sh|\ bin/fm-sessionstart-nudge.sh|bin/fm-startup-network.sh|bin/fm-tangle*|bin/fm-update.sh|\ bin/fm-gate-refuse*|bin/fm-lock*) printf '%s\n' session-bootstrap ;; + bin/fm-bootstrap.sh) + printf '%s\n' session-bootstrap + printf '%s\n' "__script__:fm-brief.test.sh" + ;; bin/fm-quota-axi-lib.sh) printf '%s\n' session-bootstrap printf '%s\n' "__script__:fm-procevent-quota.test.sh" @@ -1633,6 +1745,12 @@ families_for_changed_path() { docs/configuration.md|docs/supervision-protocols/*) printf '%s\n' pure-contract-unit ;; + tests/git-config-helpers.sh) + # The reference scan is not transitive, so match the two helpers that + # source this one as well: most suites inherit it only through them. + families_for_test_reference git-config-helpers.sh lib.sh herdr-test-safety.sh \ + || printf '%s\n' "__unmapped__:$path" + ;; tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh) families_for_test_reference "$(basename "$path")" \ || printf '%s\n' "__unmapped__:$path" @@ -2152,7 +2270,14 @@ fi if [ "$LIST_ONLY" -eq 1 ] || [ "$LIST_SCHEDULED" -eq 1 ]; then if [ "$LIST_SCHEDULED" -eq 1 ]; then for s in "${SCRIPTS[@]+"${SCRIPTS[@]}"}"; do - printf '%s\t%s\n' "$(portable_serial_weight_for "$s")" "$s" + case "$MODE:$LANE" in + lane:portable-parallel-1|lane:portable-parallel-2) + printf '%s\t%s\n' "$(portable_parallel_weight_for "$s")" "$s" + ;; + *) + printf '%s\t%s\n' "$(portable_serial_weight_for "$s")" "$s" + ;; + esac done | LC_ALL=C sort -t"$(printf '\t')" -k1,1nr -k2,2 | cut -f2- else for s in "${SCRIPTS[@]+"${SCRIPTS[@]}"}"; do @@ -2472,6 +2597,9 @@ run_bound_outcome_note() { # <script> <rc> <elapsed-secs> <bound> <started-mark # would leak errexit to the caller and turn a failing test into a script exit. run_script_bounded() { # <script> <out> local script=$1 out=$2 rc began elapsed + local GIT_CONFIG_GLOBAL GIT_CONFIG_NOSYSTEM + # shellcheck source=tests/git-config-helpers.sh + . "$ROOT/tests/git-config-helpers.sh" || return local FM_TIMEOUT_PGID_FILE="$RUN_TMP/pgid.serial" local started="$RUN_TMP/started.serial" rm -f "$started" @@ -2634,6 +2762,8 @@ else unset FM_HOME FM_STATE_OVERRIDE FM_DATA_OVERRIDE FM_ROOT_OVERRIDE \ FM_PROJECTS_OVERRIDE FM_CONFIG_OVERRIDE FM_BACKEND 2>/dev/null || true cd "$ROOT" || exit 1 + # shellcheck source=tests/git-config-helpers.sh + . "$ROOT/tests/git-config-helpers.sh" || exit 1 begin_ms=$(now_ms) if [ "$PER_SCRIPT_TIMEOUT_SECS" -gt 0 ]; then # The marker records that the bound reached the script, which is what diff --git a/bin/fm-turnend-guard.sh b/bin/fm-turnend-guard.sh index 787e34c3195..295f3877c83 100755 --- a/bin/fm-turnend-guard.sh +++ b/bin/fm-turnend-guard.sh @@ -79,7 +79,12 @@ # with the repair banner, bounded to FM_CLAUDE_TURNEND_BLOCK_BUDGET # (default 3) consecutive blocks per session - safely below Claude Code's # hard 8-consecutive-block override - then allow one loud attended -# fail-open only for an already verified failure episode. +# fail-open only for an already verified failure episode. The budget +# charges each event epoch once, and it also charges every re-block +# against an epoch the auto-arm never advanced past the previous +# re-block (budget_account_current_epoch owns that rule), so an inert +# hook that leaves the ledger frozen cannot hold the guard in an +# unbounded re-block loop below that override. set -u SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -254,12 +259,31 @@ fi # The Stop-owned auto-arm fires on the same Stop event. Give it a brief bounded # window to prove it owns recovery for this event epoch before consuming one of # Claude's bounded continuations. -budget_account_current_epoch() { - local current_epoch outcome old_session old_count old_epoch tmp initialized +# +# Budget accounting, under the budget lock. Sets COUNT (the session's +# consumed continuations, including this one) and BUDGET_INITIALIZED_FAILURE. +# The ledger's epoch identity is what is charged: a new epoch charges once, +# and an epoch this same invocation already charged is never charged again, +# because the wait loop above can observe one fresh terminal epoch many times +# before the block decision. Across Stops the two callers differ: +# - observe (the allow paths in autoarm_owns_recovery): seeing an +# already-charged epoch again is free - it is the same claim, seen again. +# - block (the re-block path): a re-block against the epoch the previous +# re-block already charged is a new consumed continuation, because the +# auto-arm advanced nothing between the two Stops - it did not participate +# at all, which is exactly the absence this budget bounds. Charging only +# epoch changes let an inert hook (identity-gated, never fired, or failing +# before its generation claim) freeze the ledger and the count together, +# so the guard re-blocked without limit and the attended fail-open below +# never became reachable. +BUDGET_CHARGED_EPOCH= +budget_account_current_epoch() { # [observe|block] + local mode=${1:-observe} current_epoch outcome old_session old_count old_epoch tmp initialized charged fm_lock_try_acquire "$BUDGET_LOCK" || return 1 current_epoch=$(sed -n '1s/^epoch=\([0-9][0-9]*\) .*/\1/p' "$STATE/.claude-autoarm-epoch" 2>/dev/null || true) outcome=$(sed -n '1s/^.*outcome=\([a-z][a-z-]*\) .*$/\1/p' "$STATE/.claude-autoarm-epoch" 2>/dev/null || true) initialized=0 + charged=0 COUNT=0 if [ -f "$BUDGET_FILE" ]; then old_session=$(sed -n '1s/^session=//p' "$BUDGET_FILE" 2>/dev/null || true) @@ -271,13 +295,18 @@ budget_account_current_epoch() { if [ "$old_session" = "$SESSION_ID" ]; then COUNT=$old_count if [ -n "$current_epoch" ] && [ "$old_epoch" = "$current_epoch" ]; then - : + if [ "$mode" = block ] && [ "$BUDGET_CHARGED_EPOCH" != "$current_epoch" ]; then + COUNT=$((COUNT + 1)) + charged=1 + fi else COUNT=$((COUNT + 1)) + charged=1 fi fi fi if [ ! -f "$BUDGET_FILE" ] || [ "${old_session:-}" != "$SESSION_ID" ]; then + charged=1 case "$outcome" in failed|failed-suppressed) if [ -e "$FAILURE_NOTICE" ]; then @@ -298,6 +327,7 @@ budget_account_current_epoch() { return 1 fi rm -f "$tmp" 2>/dev/null || true + [ "$charged" -eq 0 ] || BUDGET_CHARGED_EPOCH=$current_epoch BUDGET_INITIALIZED_FAILURE=$initialized fm_lock_release "$BUDGET_LOCK" return 0 @@ -461,7 +491,7 @@ fi # The auto-arm genuinely failed to establish: consume the bounded re-block # budget before considering the verified one-time attended fail-open. -budget_account_current_epoch || block_stop +budget_account_current_epoch block || block_stop terminal_fail_open terminal_status=$? if [ "$terminal_status" -eq 0 ]; then diff --git a/bin/fm-update.sh b/bin/fm-update.sh index 621f82f7022..ce8aa279874 100755 --- a/bin/fm-update.sh +++ b/bin/fm-update.sh @@ -51,6 +51,14 @@ # A positively dead or missing endpoint has no agent to replace and is left to # the ordinary startup recovery. # +# A fast-forward that lands changes bytes under bin/ in place, which desyncs +# the trust binding of any locally armed fm-procevent-when watch whose action +# executable lives in the updated repo; left alone, the watch's next fire +# would be wrongly refused. After each home's own update (primary and every +# local secondmate), this script best-effort runs that home's own +# fm-procevent-when.sh rebind-all to republish those bindings against the new +# bytes; a failure there is swallowed rather than failing the update. +# # Usage: fm-update.sh [--help] set -eu @@ -78,8 +86,19 @@ fi reread_firstmate="no" ff_target "$FM_ROOT" "firstmate" origin no no -if [ "$FF_STATUS" = "updated" ] && [ -n "$FF_INSTR" ]; then - reread_firstmate="yes" +if [ "$FF_STATUS" = "updated" ]; then + if [ -n "$FF_INSTR" ]; then + reread_firstmate="yes" + fi + # A fast-forward changes bin/'s bytes out from under any locally armed + # fm-procevent-when watch's trust binding, with no tampering involved; left + # alone, the very next fire is refused and the watch dies silently. Refresh + # every such watch now, right after the update that broke it. FM_ROOT_OVERRIDE + # is passed explicitly rather than relying on the script's own location: this + # process's own FM_ROOT is the repo that was just updated, which is not + # always where this very script file happens to live (FM_ROOT_OVERRIDE, as + # this test suite uses to point fm-update.sh at a fixture checkout). + FM_HOME="$FM_HOME" FM_ROOT_OVERRIDE="$FM_ROOT" "$SCRIPT_DIR/fm-procevent-when.sh" rebind-all || true fi # --- secondmates ----------------------------------------------------------- @@ -137,6 +156,15 @@ claim_settled_secondmate() { # <id> # bin/fm-ff-lib.sh calls this for each local home it left AT the base with a live # endpoint - status "updated" or "current" alike. A skipped home never gets here. fm_ff_after_secondmate_settled() { # <id> <home> <window> <status> <instr> + # Same bin/-changed-out-from-under-a-watch problem as the primary home + # above, for a local secondmate's own worktree; "current" means bin/ did + # not move there this pass, so there is nothing to rebind. Run the + # secondmate's OWN copy of the script, explicitly overriding FM_ROOT to its + # own worktree rather than letting an outer FM_ROOT_OVERRIDE (this process's + # own, if the caller set one) leak into the child and misscope it. + if [ "${4:-}" = "updated" ] && [ -x "$2/bin/fm-procevent-when.sh" ]; then + FM_HOME="$2" FM_ROOT_OVERRIDE="$2" "$2/bin/fm-procevent-when.sh" rebind-all || true + fi claim_settled_secondmate "$1" } diff --git a/bin/fm-wake-lib.sh b/bin/fm-wake-lib.sh index 1ee40021360..9b41718f7c5 100755 --- a/bin/fm-wake-lib.sh +++ b/bin/fm-wake-lib.sh @@ -327,6 +327,12 @@ fm_afk_daemon_owns_supervision() { [ "$current" = "$recorded" ] } +# Compatibility entrypoint; the read-only classifier owns mode interpretation. +fm_afk_mode() { + _fm_wake_require_classify || return 1 + fm_classify_afk_mode "$@" +} + # fm_watcher_supervision_verdict <state> <watch-path> [grace] [home] [root] # Model-aware "is supervision healthy right now" verdict for the pull warning # guard (bin/fm-guard.sh), NOT the arm layer or the turn-end guard. Sets: @@ -1210,6 +1216,121 @@ fm_treehouse_project_lock_path() { # <project-dir> printf '%s/.treehouse-project-%s.lock\n' "$root/state" "$hash" } +# A Treehouse slot has the managed pool's fixed <pool>/<slot>/<repo> layout. +# Require both its pool state and the same Git common directory as the recorded +# project; an ordinary linked worktree is not evidence that Treehouse owns it. +fm_treehouse_pool_slot() { # <project-dir> <worktree> + local project=$1 worktree=$2 slot pool state project_common slot_common + [ -d "$project" ] && [ -d "$worktree" ] || return 1 + slot=$(CDPATH='' cd -- "$worktree" 2>/dev/null && pwd -P) || return 1 + pool=$(dirname "$(dirname "$slot")") + state="$pool/treehouse-state.json" + [ -f "$state" ] && [ ! -L "$state" ] || return 1 + project_common=$(git -C "$project" rev-parse --path-format=absolute --git-common-dir 2>/dev/null) || return 1 + slot_common=$(git -C "$slot" rev-parse --path-format=absolute --git-common-dir 2>/dev/null) || return 1 + project_common=$(CDPATH='' cd -- "$project_common" 2>/dev/null && pwd -P) || return 1 + slot_common=$(CDPATH='' cd -- "$slot_common" 2>/dev/null && pwd -P) || return 1 + [ "$project_common" = "$slot_common" ] +} + +# Slot-owner claim: which task a Treehouse pool slot currently belongs to. +# +# Treehouse can record ownership durably: `treehouse get --lease --lease-holder` +# reserves a slot under a label until `treehouse return --if-lease-holder` +# releases it, and Firstmate uses exactly that for secondmate homes +# (bin/fm-home-seed.sh). Crewmate spawns do not take that path: they acquire +# their slot through the interactive pane-driven `treehouse get`, whose state +# entry is a live process lease (owner_pid plus owner_started_at, and `treehouse +# status` reports in-use from the processes actually running under the path). +# That answers "is anything running here", never "which task owns this", and it +# is released by the very event that makes a task record stale - the worker +# exiting - so a slot whose lease has lapsed reads identical whether it is still +# this task's or has since been handed to another one. Firstmate therefore keeps +# its own claim on top: one file naming the task that took the slot, written by +# bin/fm-spawn.sh under the same project lock that allocates the slot and +# released by bin/fm-teardown.sh when the slot goes back to the pool. Moving +# crewmate spawns onto the durable lease is separate follow-up work. +# +# The claim lives at <pool>/<slot>/.fm-slot-owner - a sibling of the repo +# checkout rather than a file inside it - so claiming a slot can never dirty the +# copy teardown's landed-work checks inspect, and a returned slot carries no +# untracked leftover from it. +fm_treehouse_slot_owner_marker() { # <worktree> + local worktree=$1 slot + slot=$(CDPATH='' cd -- "$worktree" 2>/dev/null && pwd -P) || return 1 + printf '%s/.fm-slot-owner\n' "$(dirname "$slot")" +} + +# Claim a pool slot for a task, replacing whatever the previous holder left. +# The rename is atomic, so a reader either sees the old claim or the new one. +fm_treehouse_slot_owner_claim() { # <worktree> <task-id> <home> + local worktree=$1 id=$2 home=$3 marker tmp + [ -n "$id" ] || return 1 + marker=$(fm_treehouse_slot_owner_marker "$worktree") || return 1 + # Only a plain claim file may be replaced: renaming onto a directory would + # move the new claim inside it and leave the slot reading as unclaimable. + if { [ -e "$marker" ] || [ -L "$marker" ]; } \ + && { [ ! -f "$marker" ] || [ -L "$marker" ]; }; then + return 1 + fi + tmp="$marker.tmp.${BASHPID:-$$}" + rm -f "$tmp" || return 1 + { + printf 'task=%s\n' "$id" + printf 'home=%s\n' "$home" + } > "$tmp" 2>/dev/null || { rm -f "$tmp"; return 1; } + mv -f "$tmp" "$marker" 2>/dev/null || { rm -f "$tmp"; return 1; } +} + +# Read the claim on a pool slot and compare it with a task id. +# Sets FM_TREEHOUSE_SLOT_OWNER to one of: +# mine - the claim names this task +# other - the claim names a different task, so the slot was reassigned +# absent - no claim: the slot was taken before claims existed, or returned since +# unsafe - a claim file exists but cannot be read as a claim +# FM_TREEHOUSE_SLOT_OWNER_ID and FM_TREEHOUSE_SLOT_OWNER_HOME carry the recorded +# claimant as evidence. The home is reported, never matched: a home that moved +# must not turn a task's own slot into a refusal. +fm_treehouse_slot_owner_state() { # <worktree> <task-id> + local worktree=$1 id=$2 marker line owner_id='' owner_home='' + FM_TREEHOUSE_SLOT_OWNER=unsafe + FM_TREEHOUSE_SLOT_OWNER_ID= + FM_TREEHOUSE_SLOT_OWNER_HOME= + marker=$(fm_treehouse_slot_owner_marker "$worktree") || return 0 + if [ ! -e "$marker" ] && [ ! -L "$marker" ]; then + FM_TREEHOUSE_SLOT_OWNER=absent + return 0 + fi + [ -f "$marker" ] && [ ! -L "$marker" ] || return 0 + while IFS= read -r line || [ -n "$line" ]; do + case "$line" in + task=*) owner_id=${line#task=} ;; + home=*) owner_home=${line#home=} ;; + esac + done < "$marker" || return 0 + [ -n "$owner_id" ] || return 0 + # shellcheck disable=SC2034 # Output globals, read by the sourcing caller. + FM_TREEHOUSE_SLOT_OWNER_ID=$owner_id + # shellcheck disable=SC2034 # Output globals, read by the sourcing caller. + FM_TREEHOUSE_SLOT_OWNER_HOME=$owner_home + if [ "$owner_id" = "$id" ]; then + FM_TREEHOUSE_SLOT_OWNER=mine + else + FM_TREEHOUSE_SLOT_OWNER=other + fi +} + +# Drop a task's own claim once its slot is back in the pool. Never removes +# another task's claim, so a misdirected release cannot strip the evidence that +# protects the slot's real owner. +fm_treehouse_slot_owner_release() { # <worktree> <task-id> + local worktree=$1 id=$2 marker + fm_treehouse_slot_owner_state "$worktree" "$id" + [ "$FM_TREEHOUSE_SLOT_OWNER" = mine ] || return 0 + marker=$(fm_treehouse_slot_owner_marker "$worktree") || return 0 + rm -f "$marker" 2>/dev/null || true +} + fm_failure_episode_reset() { local state=$1 mode=${2:-acquire} lock current pid acquired=0 path lock="$state/.turnend-claude-blocks.lock" diff --git a/bin/fm-watch.sh b/bin/fm-watch.sh index 246a137474e..9a1254ec115 100755 --- a/bin/fm-watch.sh +++ b/bin/fm-watch.sh @@ -90,6 +90,22 @@ # and has not been surfaced yet; reported once per # captured generation, never again while that record # stays queued and never once it is acknowledged +# check: process-event source stranded: <keys> +# a registered process-to-event source has a claim +# reconcile will not displace and nothing collecting +# for it (bin/fm-procevent.sh reconcile queues it +# once per stranded claim generation); the queued +# payload names what clears it +# check: process-event source failed to start: <keys> +# a registered process-to-event source was launched by +# reconcile and did not prove it took the claim within +# the confirm window, so nothing is confirmed to be +# collecting for it and every cycle will relaunch it +# (bin/fm-procevent.sh reconcile queues it once per +# failure episode, and a later cycle that finds the +# source owned closes that episode); the queued +# payload names what to check. These three kinds are +# joined with `;` when more than one surfaces in a cycle # check: rejected unauthenticated state checks: <paths> # unsafe state checks were refused without execution # check: rejected unauthenticated PR poll retirement receipts: <paths> @@ -128,6 +144,10 @@ mkdir -p "$STATE" . "$SCRIPT_DIR/fm-push-transition-lib.sh" # shellcheck source=bin/fm-pr-lib.sh . "$SCRIPT_DIR/fm-pr-lib.sh" +# Only for the arm-time check on FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS below; +# the per-cycle reconcile itself runs as a separate process. +# shellcheck source=bin/fm-procevent-lib.sh +. "$SCRIPT_DIR/fm-procevent-lib.sh" # Single owner of durable merge-outcome publication, shared with # bin/fm-pr-merge.sh so self and poll origins use the same role-routed outcome. # The watcher still owns immediate delivery of its actionable poll result and @@ -139,6 +159,10 @@ mkdir -p "$STATE" # worker while adding no uncovered file. # shellcheck source=/dev/null . "$SCRIPT_DIR/fm-merge-outcome-lib.sh" +# The durable merge-authority owner is shared with bin/fm-pr-merge.sh. The +# watcher consumes only its identity-bound record after a poll observes landing. +# shellcheck source=/dev/null +. "$SCRIPT_DIR/fm-merge-authority-lib.sh" # shellcheck source=bin/fm-x-lib.sh . "$SCRIPT_DIR/fm-x-lib.sh" # shellcheck source=bin/fm-check-lib.sh @@ -1596,7 +1620,7 @@ procevent_surface_after_output() { } procevent_surface_queued() { - local key reason + local key reason captured="" stranded="" unstarted="" PROCEVENT_SURFACED= [ -s "$FM_WAKE_QUEUE" ] || return 0 fm_lock_acquire_wait "$FM_WAKE_QUEUE_LOCK" @@ -1604,12 +1628,30 @@ procevent_surface_queued() { case "$key" in procevent:*) ;; *) continue ;; esac [ -e "$(procevent_surfaced_marker "$key")" ] && continue PROCEVENT_SURFACED="$PROCEVENT_SURFACED $key" + # A stranded source or one whose launch never proved itself is the opposite + # of a captured result: nothing is collecting for it. Headlining either as + # a capture would present it as healthy, which is the shape of defect + # these wakes exist to surface. + case "$key" in + procevent:*:stranded:*) stranded="$stranded $key" ;; + procevent:*:launch-failed:*) unstarted="$unstarted $key" ;; + *) captured="$captured $key" ;; + esac done < <(fm_wake_queued_keys_locked check) if [ -z "$PROCEVENT_SURFACED" ]; then fm_lock_release "$FM_WAKE_QUEUE_LOCK" return 0 fi - reason="check: process-event result captured:$PROCEVENT_SURFACED" + reason="check:" + [ -z "$captured" ] || reason="$reason process-event result captured:$captured" + if [ -n "$stranded" ]; then + [ "$reason" = "check:" ] || reason="$reason;" + reason="$reason process-event source stranded:$stranded" + fi + if [ -n "$unstarted" ]; then + [ "$reason" = "check:" ] || reason="$reason;" + reason="$reason process-event source failed to start:$unstarted" + fi # shellcheck disable=SC2034 # Consumed by wake() in the separately linted transition owner. FM_WAKE_POST_OUTPUT_ACTION=procevent_surface_after_output wake "$reason" @@ -1967,6 +2009,24 @@ if [ "${BASH_SOURCE[0]}" != "$0" ]; then return 0 fi +# FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS is validated here, at arm time, and an +# unusable value refuses to arm. This is deliberately NOT symmetry with the +# tunables above, which this watcher only defaults and never validates. The +# reason is specific: every supervision cycle runs `fm-procevent.sh reconcile` +# with its output and exit status discarded, and reconcile refuses an unusable +# window by name before it launches anything. Under this watcher that refusal +# is invisible - every cycle would exit early, no source would ever start, and +# the whole home would sit disarmed while presenting as supervised. A watcher +# that refuses to arm is loud through an existing, independent, proven path: +# the liveness guard's WATCHER DOWN banner in firstmate's own session. The +# message shape is reconcile's own, so the operator reads one refusal in both +# places. The refusal goes to stdout because bin/fm-watch-arm.sh relays the +# child's stdout and recognises `watcher: FAILED` as the typed failure line. +if ! fm_procevent_launch_confirm_seconds >/dev/null; then + echo "watcher: FAILED - FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS must be whole seconds from $FM_PROCEVENT_LAUNCH_CONFIRM_MIN_SECONDS to $FM_PROCEVENT_LAUNCH_CONFIRM_MAX_SECONDS" + exit 1 +fi + if ! fm_lock_try_acquire "$WATCH_LOCK"; then BEAT="$STATE/.last-watcher-beat" if [ -n "${FM_LOCK_HELD_PID:-}" ]; then @@ -2062,6 +2122,13 @@ reconcile_requests_detached() { RECONCILE_REQUEST_PID=$! } +PR_POLL_CONTROL_LOCK= + +pr_poll_control_release() { + [ -z "$PR_POLL_CONTROL_LOCK" ] || fm_lock_release "$PR_POLL_CONTROL_LOCK" || return 1 + PR_POLL_CONTROL_LOCK= +} + watcher_cleanup() { # Ignore stop signals for the whole cleanup, whatever started the exit (a stop # signal, self-eviction, or an error exit). Real senders deliver stop signals in @@ -2074,6 +2141,7 @@ watcher_cleanup() { # the window it leaves open and what backstops it. trap '' HUP INT TERM local cleanup_status=0 owns_lock=0 transition=release-lock + pr_poll_control_release || cleanup_status=1 if [ "$(cat "$WATCH_LOCK/pid" 2>/dev/null || true)" = "${WATCHER_PID:-}" ]; then owns_lock=1 if [ "${WATCHER_RECOVERY_PENDING:-0}" -eq 1 ] \ @@ -2309,6 +2377,13 @@ while :; do host=$FM_PR_POLL_SNAPSHOT_HOST path=$FM_PR_POLL_SNAPSHOT_PATH number=$FM_PR_POLL_SNAPSHOT_NUMBER + PR_POLL_CONTROL_LOCK="$STATE/.control-$id.lock" + fm_lock_acquire_wait "$PR_POLL_CONTROL_LOCK" || exit 1 + if ! fm_pr_poll_snapshot_matches "$STATE" "$id" "$SCRIPT_DIR/fm-pr-poll.sh"; then + pr_poll_control_release || exit 1 + triage_log "PR poll for $id changed before its validated check; skipping the stale snapshot" + continue + fi run_check_capture "$SCRIPT_DIR/fm-pr-poll.sh" --validated \ "$provider" "$url" "$host" "$path" "$number" || exit 1 out=$FM_CHECK_RESULT @@ -2326,14 +2401,28 @@ while :; do if [ -n "$out" ]; then reason="check: $c: $out" if [ "$is_pr_poll" -eq 1 ] && [ "$out" = merged ]; then + if ! fm_merge_authority_read "$STATE" "$id" \ + "$provider" "$host" "$path" "$number"; then + triage_log "no matching persisted merge authority for $id; recording an external merge outcome" + fi + merge_authority=$FM_MERGE_AUTHORITY + merge_authority_record_identity=$FM_MERGE_AUTHORITY_RECORD_IDENTITY merge_outcome_rc=0 fm_merge_outcome_report "$FM_HOME" "$STATE" "$id" "$url" poll \ - || merge_outcome_rc=$? + "$merge_authority" || merge_outcome_rc=$? if [ "$merge_outcome_rc" -ne 0 ]; then triage_log "merge outcome for $id could not be recorded (rc=$merge_outcome_rc)" exit 1 fi + if [ -n "$merge_authority_record_identity" ] \ + && ! fm_merge_authority_remove_if_matches "$STATE" "$id" \ + "$provider" "$host" "$path" "$number" "$merge_authority" \ + "$merge_authority_record_identity"; then + triage_log "published merge outcome for $id but could not retire its authority record" + exit 1 + fi retire_merged_pr_poll "$id" + pr_poll_control_release || exit 1 touch "$STATE/.last-check" if [ "$FM_MERGE_OUTCOME_ALREADY_RECORDED" = true ]; then triage_log "absorbed duplicate merged PR poll result for $id" @@ -2341,10 +2430,12 @@ while :; do fi wake "$reason" fi + pr_poll_control_release || exit 1 fm_wake_append check "$c" "$reason" || exit 1 touch "$STATE/.last-check" wake "$reason" fi + pr_poll_control_release || exit 1 done if [ -n "$rejected_checks" ]; then reason="check: rejected unauthenticated state checks:$rejected_checks" diff --git a/bin/fm-x-link.sh b/bin/fm-x-link.sh index 13b881c0c7c..fd28ca5a11f 100755 --- a/bin/fm-x-link.sh +++ b/bin/fm-x-link.sh @@ -177,7 +177,7 @@ if [ ! -f "$META" ]; then ''|*' '*) ;; *) ROUTE_HOME_ARG="secondmate:$ROUTE_MATCHES" ;; esac - printf 'fm-x-link: bind the public promise through the promised-final path instead: tasks-axi public-followup add + bind-work, then bin/fm-public-followup.sh register <obligation-id> --relation <relation-id> --work-home %s --work-id %s --generation <n>, and put the bin/fm-public-followup.sh brief <obligation-id> command into the routed worker instructions.\n' \ + printf 'fm-x-link: bind the public promise through the promised-final path instead: bin/fm-tasks-axi.sh public-followup add + bind-work, then bin/fm-public-followup.sh register <obligation-id> --relation <relation-id> --work-home %s --work-id %s --generation <n>, and put the bin/fm-public-followup.sh brief <obligation-id> command into the routed worker instructions.\n' \ "$ROUTE_HOME_ARG" "$ID" >&2 fi exit 1 diff --git a/docs/agent-control.md b/docs/agent-control.md index 21aa5f23cf9..9c1444dd5c9 100644 --- a/docs/agent-control.md +++ b/docs/agent-control.md @@ -48,7 +48,7 @@ The clear is refused before anything is sent when the recorded backend cannot de Removing a worktree, closing an endpoint, or discarding work stays with [`bin/fm-teardown.sh`](../bin/fm-teardown.sh), which owns the landed-work test. **`resume` is not a verb.** -It is not deterministic across the verified adapters: codex, grok, and gemini resume only from a session id printed at exit, opencode continues the most recent session for the cwd, and claude, pi, pi-signed, omp, and kimi have no verified pane-resume contract. +It is not deterministic across the verified adapters: codex, grok, and gemini resume only from a session id printed at exit, opencode continues the most recent session for the cwd, and claude, pi, pi-signed, omp, kimi, and agy have no verified pane-resume contract. `relaunch` covers the same need on every adapter, because the brief on disk - not a harness-private session - is the durable instruction. ## Transactional relaunch @@ -117,11 +117,11 @@ Backend capability comes from each adapter's real surface, not from a policy cho | cmux | yes | yes | yes | yes | no | | orca | no | yes | yes | no | no | -Per-harness interrupt keys, repeat counts, composer clears, exit commands, and supported task kinds live in `bin/fm-control-lib.sh` and are exercised for every verified harness by `tests/fm-control.test.sh`. +Per-harness interrupt keys, repeat counts, composer clears, exit commands, and supported task kinds live in `bin/fm-control-lib.sh` and are exercised for every verified harness by `tests/fm-control.test.sh`, with adapters outside its lane pinning their control mechanics in their own harness suites. The empirical basis for each adapter's value is the `harness-adapters` skill's verification record for that adapter. ## Verification -- `tests/fm-control.test.sh` - the adapter contract for every verified harness, the backend capability matrix, exact-id scoping, the closed verb list, the busy, idle, dead, and idempotent lifecycle cases, and marker non-regression, all against a stubbed session provider. -- `tests/fm-control-relaunch.test.sh` - the relaunch transaction: identity preservation, including `kind=design` on the supported design runtimes, harness switching, the progress note, checkpoint refusals, and rollback after a failed launch. +- `tests/fm-control.test.sh` - the adapter contract for its verified-harness lane (adapters outside the lane pin their control mechanics in their own harness suites), the backend capability matrix, exact-id scoping, the closed verb list, the busy, idle, dead, and idempotent lifecycle cases, and marker non-regression, all against a stubbed session provider. +- `tests/fm-control-relaunch.test.sh` - the relaunch transaction: identity preservation, harness switching, the progress note, checkpoint refusals, and rollback after a failed launch. - `tests/fm-control-herdr-smoke.test.sh` - the second state-verified backend against the real herdr binary, on an isolated throwaway lab session. diff --git a/docs/architecture.md b/docs/architecture.md index 7dd66a53364..27678550052 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -218,7 +218,7 @@ Every classification returns a verdict of busy, idle, unknown, or dead together Each converted adapter reports its own turn lifecycle through a machine-readable contract the vendor already exposes, rather than through rendered footer text: Pi and pi-signed through the Firstmate-owned extension's `agent_start` and `agent_settled` confirmed by `ctx.isIdle()`, omp through its extension's `agent_start` and `agent_end` without `willContinue`, OpenCode through its plugin's semantic `session.status`, Claude through owned `UserPromptSubmit`, `Stop`, `StopFailure`, and `SessionEnd` hooks, Muse through its session log, and Cursor through its conversation transcript. Kimi behind Pi inherits Pi's lifecycle. -Codex and standalone Kimi classify unknown behind explicit probes until a semantic source is live-verified for them, and Grok keeps one clearly isolated rendered-tail fallback that can only ever classify a Grok task. +Codex and standalone Kimi classify unknown behind explicit probes until a semantic source is live-verified for them, and Grok, Rovo, and AGY each keep one clearly isolated rendered-tail fallback that can only ever classify their own task. Missing, malformed, stale, untrusted, or unverified semantic state is unknown, never idle, and unknown is never promoted to busy either. Ordinary task-state consumers act only on an exact busy verdict, so an unreadable worker surfaces for a closer look instead of being absorbed as still-working or written off as finished. @@ -357,12 +357,22 @@ The legacy `bin/fm-brief.sh --issue <number>` input remains only for a same-repo Write-back itself runs one lifecycle vocabulary through `bin/fm-work-item-milestone.sh` onto both surfaces firstmate keeps true, the work item's living status comment and the captain's project board, so the two cannot hold different opinions about a task; both are decoration that fails open and can never fail the work they describe. Every per-host forge write crosses one owner, `bin/fm-forge-lib.sh`, whose header owns the write-operation allowlist, the minimum token scope per forge, and the adapter-shaped extension path a further forge follows. The helper requires a full canonical URL and rejects malformed URLs or repo override flags before recording merge state. -A `https://github.com/<owner>/<repo>/pull/<n>` URL invokes `gh-axi pr merge <n> --repo <owner>/<repo>`, defaults to `--squash`, and preserves explicit merge-method flags, which on GitHub include `--rebase` because a GitHub rebase merge replays onto the base without rewriting the source branch. +A `https://github.com/<owner>/<repo>/pull/<n>` URL requires `gh` and `jq`, is merged only after one live read confirms the pull request is open, not a draft, mergeable, conflict-free, and every unwaived check is green at the current head, then `gh pr merge` binds that verified head with `--match-head-commit`. +A check run is green when its current run is green, because GitHub leaves a cancelled run in the rollup beside the passing re-run it triggered when the base branch advanced; `bin/fm-pr-merge.sh`'s `github_checks_not_green` owns the rule, which uses `startedAt` to clear only an older completed check run that a passing run with the same name provably replaced, while unfinished check runs and non-green status contexts stay red. +The fork allowlist always refuses `--auto` and `--admin`; `--attended-override` can admit supported branch-deletion flags for an explicit instruction, while preserving the live green check, away authority, and task hold. +An attended `--allow-red <check-name>` may appear once, waives only GitHub checks with that exact name, and is refused while the away-posture record exists. +Because away merge authority is read from that record and then acted on by the forge, the authority read and synchronous forge command share the record's cross-subsystem lock, closing the common live-owner TOCTOU. +A lock that cannot be taken refuses the merge. +While the record exists, GitHub auto-merge and any base whose rules cannot prove the absence of a merge queue are refused before submission, and GitLab auto-merge flags or scheduled state are refused while an immediate merge is forced with a final `--auto-merge=false`. +This is deliberately confused-agent-grade, as `bin/fm-lease-lib.sh` defines that grade, rather than fully atomic. +A GitHub queue-rule or PR-base change after the queue-free preflight can still enqueue a merge that lands after its away grant lapses, and killing the lock-owning shell while its forge child survives lets stale-owner recovery admit archive or replacement before that child completes. +These are accepted limitations, not oversights; durable authority, landing re-verification, and child-lock handoff are outside this boundary. +`bin/fm-afk-contract.sh` owns the lock contract, while `tests/fm-afk-contract.test.sh` and `tests/fm-pr-merge.test.sh` pin the serialization and fail-closed merge behavior. A `https://<host>/<path>/-/merge_requests/<n>` URL (see [docs/gitlab-merge-watch.md](gitlab-merge-watch.md)) invokes `glab mr merge <n> -R https://<host>/<path>`, so the instance comes from the URL, and adds no merge-method flag because the project's own merge method applies. That path merges only after one live read of the merge request confirms it is open, mergeable, conflict-free, with blocking discussions resolved and a successful pipeline at the current head, and it binds the merge to that verified head; recorded metadata is never the authority for those conditions because a rebase leaves it stale. It refuses earlier still when the project's own merge method can rebase the source branch at merge time and the source is behind the target, because that rebase lands commits whose pipeline never ran and strands the attestation this fleet merges on; only a positive reading permits, so an unreadable merge method, setting, instance version, or divergence count refuses as well, and `bin/fm-pr-merge.sh`'s header owns the GitLab version window in which that capability exists but cannot be read. After either forge command returns, the script confirms the PR or MR actually landed, and only a confirmed landing records a landed outcome; a queued or unconfirmed request records none and leaves its poll armed. -The forge read is the authority and the merge command's exit status is not: a non-zero `gh-axi pr merge` or `glab mr merge` whose follow-up read confirms the request merged is deliberately recorded as landed, reconciled out loud on both forges rather than exiting zero in silence, and returns zero, and a merge already observed as landed - by the pre-merge target read as much as by any later one - is never re-read. +The forge read is the authority and the merge command's exit status is not: a non-zero `gh pr merge` or `glab mr merge` whose follow-up read confirms the request merged is deliberately recorded as landed, reconciled out loud on both forges rather than exiting zero in silence, and returns zero, and a merge already observed as landed - by the pre-merge target read as much as by any later one - is never re-read. An outcome the read does not confirm as landed records nothing instead: on GitHub an outcome that could not be read at all refuses non-zero, while GitLab leaves the poll armed and fails the run only when the merge command failed too. That verdict is a deliberate fork divergence recorded in [docs/fork-divergence.md](fork-divergence.md); the fork's earlier refusal of a merge-queue entry was retired by captain decision on 2026-09-03, so a queued request is again reported as `verified: queued` on exit zero. Guarded merging on GitHub is limited to the repository's current default branch: a pull request targeting anything else is refused by name, naming both the target and the default, before any queue, auto-merge, or method handling, and both branches come from the same GraphQL query or, when `gh` is degraded, from gh-axi's own `api` passthrough; `HelloWorldSungin/firstmate#257` tracks extending the contract to GitLab. @@ -376,8 +386,12 @@ Every GitHub refusal states what it could not observe as plainly as what it did, A confirmed merge leaves a durable role-routed outcome instead of living only in the merging agent's memory, and [`bin/fm-merge-outcome-lib.sh`](../bin/fm-merge-outcome-lib.sh)'s header owns its destination, shape, identity, normal-case deduplication, and at-least-once recovery. The same emitter handles a merge firstmate performed and one its poll detected, while the watcher immediately delivers the emitter's local actionable poll row. After a merge succeeds, `bin/fm-pr-merge.sh` closes at most one eligible recorded work item, on whichever forge it holds a write adapter for and in the repository that item records; multiple items, a forge or host with no adapter, a credential that is absent or refused, and bookkeeping failures warn without turning a completed merge into a failed, retryable merge. +After the forge accepts firstmate's merge request, the merge path persists the resolved yolo, away-grant, or attended authority bound to the task's canonical PR identity. +A later merged poll consumes only that matching persisted value; with no match it records the landing as external rather than consulting a live away-posture record that may have been archived or replaced. +[`bin/fm-merge-authority-lib.sh`](../bin/fm-merge-authority-lib.sh)'s header owns resolution, private atomic persistence, identity-checked consumption, and retirement, while only the merge path gates on the answer. Teardown is fail-closed for ship and design worktrees: dirty worktrees refuse, and committed work must be landed before the worktree is returned. -A pool worktree is only returned after teardown passes the slot-ownership proof: a contradictory task record or supported live endpoint refuses without touching either task, and no discard authority relaxes that. +A pool worktree is only returned after teardown passes the slot-ownership proof: a contradictory task record or a supported live endpoint refuses without touching either task, and no discard authority relaxes that. +A slot's own owner claim, written by the spawn that takes it under the allocation lock and owned by [`bin/fm-wake-lib.sh`](../bin/fm-wake-lib.sh), covers a slot reassigned to a task that left no record the scan could reach: a claim naming a different task releases nothing - teardown warns, names the claimant, and finishes only the task's own cleanup - because Treehouse's own live process lease cannot answer ownership once the worker's exit releases it. Allocation and return serialize on one project lock per machine-local Firstmate tree: every home reachable through local parent links shares that lock, and a home seeded from another machine anchors its own, because a lock taken on this filesystem is neither held nor observable across that boundary. [`bin/fm-teardown.sh`](../bin/fm-teardown.sh)'s header owns the landed-work proofs, slot-ownership proof, PR-discovery fallback, pre-teardown run conclusion, and stale-lock recovery procedure; [`tests/fm-teardown-endpoint-safety.test.sh`](../tests/fm-teardown-endpoint-safety.test.sh) and [`tests/fm-secondmate-safety.test.sh`](../tests/fm-secondmate-safety.test.sh) pin the slot-collision boundary. @@ -442,7 +456,7 @@ A learning that no longer earns a startup-memory slot but is still true is captu The internal [`stow` skill](../.agents/skills/stow/SKILL.md) owns tier markers, decay, cold archival, and captain-gated offload. The same pass also persists open-work record state the session is holding - filing a thread that was never recorded and correcting one the session knows went stale - bounded to the open work that session is actually holding. It is deliberately not a reconciliation of durable records against repository or PR reality: its input is the volatile context, so it can only preserve what the session still knows, and no reconciliation that outlives a session exists today. -Task-scoped notes use `tasks-axi show <id> --full` followed by `tasks-axi update <id> --body-file <path>`, adding `--archive-body` when the prior body should remain recoverable. +Task-scoped notes use `bin/fm-tasks-axi.sh show <id> --full` followed by `bin/fm-tasks-axi.sh update <id> --body-file <path>`, adding `--archive-body` when the prior body should remain recoverable. The stow pass never writes a tracked skill; its only direct skill writes are adding to an existing user-owned local destination, creating the one bounded bootstrap destination for a home with none, and creating a validated weight or fit split, with the internal `stow` skill above owning the exact contract. A separately executed, captain-approved migration may move pinned conditional knowledge into a user-owned local skill excluded from the Firstmate clone. Invoked in a primary home, `/stow` then cascades the same sweep to every registered secondmate, enumerated through `bin/fm-stow-cascade.sh`: each home is accounted and curated against its own startup-memory allowance, a live secondmate sweeps its own session, and a slow or unreachable home is reported as an exception rather than blocking the primary. diff --git a/docs/captain-hold-lifecycle.md b/docs/captain-hold-lifecycle.md index 21a11805706..8485df4ec64 100644 --- a/docs/captain-hold-lifecycle.md +++ b/docs/captain-hold-lifecycle.md @@ -155,6 +155,7 @@ The window between a merge landing and cleanup is an accepted structural residua That local window is normally only seconds wide and requires re-holding a task whose merge has just landed. A re-hold inside the window makes cleanup retain the row rather than publish it, so the delivery is omitted until the stale hold is cleared from that row. Queued forge merges cannot be covered locally because the forge performs the merge asynchronously after the local command has returned, when no lock this code could hold would still be held. +The away-posture restriction on queued merges and its residual limits are owned by [architecture.md](architecture.md#delivery-modes-are-explicit-per-task). ## Record divergence diff --git a/docs/cd-guard.md b/docs/cd-guard.md index 2814912180c..ae4dae95eba 100644 --- a/docs/cd-guard.md +++ b/docs/cd-guard.md @@ -11,7 +11,7 @@ the watcher-arm PreToolUse seatbelt (`bin/fm-arm-pretool-check.sh`, `docs/arm-pr ## Purpose and boundary The primary firstmate shell persists its working directory across tool calls. -A stray persistent top-level `cd projects/<clone>` therefore silently relocates the shell, so the next firstmate-owned command - a backlog write, an `fm-*` lifecycle call, `tasks-axi` - runs inside a project clone instead of the home. +A stray persistent top-level `cd projects/<clone>` therefore silently relocates the shell, so the next firstmate-owned command - a backlog write, an `fm-*` lifecycle call - runs inside a project clone instead of the home. That has actually happened: a persistent top-level `cd` caused a firstmate-owned backlog write to execute inside a project clone rather than the home. The seatbelt denies exactly that command shape - a cwd change that persists to the primary shell - before it runs. diff --git a/docs/configuration.md b/docs/configuration.md index f0c0c571803..17ff7fedd98 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -267,6 +267,10 @@ A `manual` home owns its backlog file outright: the lifecycle transitions above Absent or `tasks-axi` selects the tasks-axi path. On the default markdown adapter, tasks-axi and manual edits produce the same `## In flight`, `## Queued`, and `## Done` sections. +The tracked `.tasks.toml` paths resolve against the directory tasks-axi runs in, not `FM_HOME`, so a bare `tasks-axi` run from the code root addresses the code root's `data/` whenever the home lives elsewhere. +tasks-axi writes by renaming a temp file over its target, which replaces a symlink with a regular file, so linking the code-root copy into the home forks the queue on the first such write rather than keeping the two in step. +Every routine firstmate backlog command therefore runs through [`bin/fm-tasks-axi.sh`](../bin/fm-tasks-axi.sh), which addresses this home's backlog and archive from any working directory exactly as lifecycle transitions do, and bootstrap reports a code-root `data/backlog.md` or `data/done-archive.md` that is not this home's own file as a `BACKLOG_RECONCILE: code-root ...` line even in a read-only session. + ## Runtime backend (config/backend / FM_BACKEND) For spawn-capable adapters, the runtime session-provider backend controls where task windows/endpoints are created, captured, sent to, watched, and killed. @@ -521,6 +525,17 @@ Its `remove` action excises only the marker-delimited Firstmate region and remov For Pi and pi-signed secondmate launches, `fm-spawn.sh` starts the selected executable with `-e` pointed at the secondmate home's own tracked `.pi/extensions/fm-primary-pi-watch.ts` and `.pi/extensions/fm-primary-turnend-guard.ts`, both already present from the secondmate home's git worktree. For omp secondmate launches, `fm-spawn.sh` passes no `-e` at all: omp auto-discovers the home's tracked `.omp/extensions/` with no trust gate, and naming a discovered file with `-e` as well loads it twice; every omp launch instead carries the tracked `.omp/fm-worker-overlay.yml` posture overlay through `--config`, which [`fm-spawn.sh --help`](../bin/fm-spawn.sh) owns. +## Claude permission mode (config/claude-permission-mode) + +The optional local, gitignored `config/claude-permission-mode` holds one token selecting the permission flag every Claude worker launch carries: crewmates, scouts, Claude secondmates, and control-plane relaunches alike. +The token is the file's whitespace-trimmed content. +`bypass` keeps today's launch, `claude --dangerously-skip-permissions`, and is also the default when the file is absent, so an unconfigured home launches byte-for-byte as before. +`auto` replaces that flag with `--permission-mode auto`, Claude Code's classifier-reviewed permission mode, for a captain who refuses to run workers in bypass mode; every other part of the Claude launch, including its environment prefix, inline settings, model, and effort flags, is unchanged. +Any other value, or an unreadable file, refuses every spawn from that home, whichever harness it would launch, before any endpoint, worktree, or task record exists, and names the accepted values; Firstmate never falls back to a permission posture the captain did not choose. +`bin/fm-spawn.sh` reads the file on every spawn and relaunch, so a change takes effect at the next launch without a restart. +The file is a captain-wide safety preference, so it is inherited into secondmate homes under the [`secondmate-provisioning`](../.agents/skills/secondmate-provisioning/SKILL.md) inherited-local-material contract; a secondmate's own Claude crewmates then launch on the same posture. +The [Claude adapter reference](../.agents/skills/harness-adapters/references/harness/claude.md) records the verified shape of both launches and which once-per-machine dialog each one can meet. + ## Worker launch environment (config/launch-env-allowlist) The optional local, gitignored `config/launch-env-allowlist` limits the ambient environment passed to newly launched workers, scouts, and secondmates, including relaunches. @@ -667,10 +682,11 @@ The sidecar is additive and never overrides or shadows quota evidence that `quot On session start the first mate detects what its required toolchain is missing or too old and lists each problem with either an exact install command or manual instructions. It installs automatically supported tools only after you say go; manual-only tools remain for you to install from the printed instructions. Required tools come in two parts: a universal toolchain every home needs regardless of backend, and a per-backend delta that follows the runtime backend actually resolved for this home. -The universal toolchain is node, git, gh with GitHub auth via `gh auth login`, no-mistakes v1.46.0 or newer, compatible gh-axi, compatible chrome-devtools-axi, compatible lavish-axi, compatible tasks-axi per "Backlog backend" above, and compatible quota-axi. -[`bin/fm-bootstrap.sh`](../bin/fm-bootstrap.sh) owns the axi-family floor policy and the gh-axi, chrome-devtools-axi, and lavish-axi floors, while [`bin/fm-tasks-axi-lib.sh`](../bin/fm-tasks-axi-lib.sh) and [`bin/fm-quota-axi-lib.sh`](../bin/fm-quota-axi-lib.sh) hold their own tools' floor constants. +The essential universal toolchain is node, git, jq, gh with GitHub auth via `gh auth login`, no-mistakes v1.46.0 or newer, compatible gh-axi, chrome-devtools-axi, compatible tasks-axi per "Backlog backend" above, and compatible quota-axi. +[`bin/fm-bootstrap.sh`](../bin/fm-bootstrap.sh) owns the axi-family floor policy and the gh-axi and lavish-axi floors, while [`bin/fm-tasks-axi-lib.sh`](../bin/fm-tasks-axi-lib.sh) and [`bin/fm-quota-axi-lib.sh`](../bin/fm-quota-axi-lib.sh) hold their own tools' floor constants. This section is the single owner of that universal toolchain list; backend guides' prerequisites point here and add only their backend-specific tools. -In that list, no-mistakes runs the validation pipeline, gh-axi, chrome-devtools-axi, and lavish-axi cover GitHub, browser, and rich-review operations, and tasks-axi plus quota-axi back backlog mutations and quota-aware array dispatch. +In that list, no-mistakes runs the validation pipeline, gh-axi and chrome-devtools-axi cover GitHub and browser operations, and tasks-axi plus quota-axi back backlog mutations and quota-aware array dispatch. +Lavish is a presentation-only dependency for visual decisions and reports; nonvisual work can proceed with plain text when it is unavailable. The per-backend delta is required only for the backend resolved from `FM_BACKEND`, then `config/backend`, then runtime auto-detection, then default `tmux`, so a home is never told to install a tool an inactive backend or feature would need. That delta is owned in code by `fm_backend_required_tools` in `bin/fm-backend.sh`: the resolved backend's own session-provider CLI (`tmux`, `herdr`, `zellij`, `orca`, or `cmux`), `jq` for the JSON-emitting adapters (`herdr`, `zellij`, `cmux`) whose spawn and liveness paths parse the backend's JSON output, and the `treehouse` worktree provider for every session-provider-only backend (`tmux`, `herdr`, `zellij`, `cmux`). Backend tool availability uses the adapter's own executable resolver, so bootstrap and spawn agree on supported non-`PATH` locations such as cmux's bundled CLI. @@ -679,10 +695,10 @@ Orca provides both the task worktree and terminal endpoint (see "Runtime backend A herdr, zellij, or cmux home is therefore never told `tmux` is missing, and the `treehouse` durable-lease upgrade check runs only for the backends that actually use treehouse. When `config/crew-dispatch.json` exists, bootstrap also requires `jq` for dispatch profile validation. When Relay is opted in, bootstrap also requires `curl` and `jq` before arming the relay poll shim. -`tasks-axi` and `quota-axi` are required bootstrap tools in every profile, the same class as `lavish-axi`. +`tasks-axi` and `quota-axi` are essential bootstrap tools in every profile. An absent or incompatible `tasks-axi` reports `MISSING: tasks-axi (install: npm install -g tasks-axi)`; when `config/backlog-backend` is not `manual`, a home with a configured non-markdown adapter or a markdown backlog refuses lifecycle mutation until compatible `tasks-axi` is on `PATH`, while a manual-backend home keeps its backlog hand-edited. An absent or incompatible `gh-axi` reports `MISSING: gh-axi (install: npm install -g gh-axi && gh-axi setup hooks)`. -An absent or incompatible `lavish-axi` reports `MISSING: lavish-axi (install: npm install -g lavish-axi && lavish-axi setup hooks)`. +An absent or incompatible `lavish-axi` reports `PRESENTATION_UNAVAILABLE` with its required floor, install command, and explicit text fallback; [`bootstrap-diagnostics`](../.agents/skills/bootstrap-diagnostics/SKILL.md) owns the response and compatibility check before visual use. An absent or too-old `quota-axi` reports `MISSING: quota-axi (install: npm install -g quota-axi)`; firstmate cannot resolve a profile array without a compatible binary. That floor exists because it is the first build reporting per-credential auth sources, which Firstmate uses when the candidate's authoritative catalog does not itself establish the selected authentication surface. Bootstrap also reports a `TANGLE:` line when `FM_ROOT` is on a named non-default branch; follow the printed checkout remediation rather than treating it as an installable tool problem. @@ -1018,7 +1034,8 @@ Real feedback, ended and missing sessions, any other `SERVER_ERROR`, and that sa An already-armed Lavish source keeps its registered listener command until it is retired and armed again, so re-arm a live board once to adopt this retry policy. The `when` adapter (`bin/fm-procevent-when.sh`) turns this channel into a condition->action primitive: it registers a deterministic condition and a deterministic action once, its blocking child polls the condition without waking firstmate, and a stable true fires the action at most once before one terminal outcome is durably captured and published as a wake that remains eligible for re-announcement until handled. -The (condition, action) spec is stored privately under `state/when/` and hash-bound by a trust record the same way `bin/fm-check-register.sh` binds a custom check, while the spec separately binds the resolved action executable's bytes; a mutated or unregistered spec or a changed action executable is refused before the action runs. +The (condition, action) spec is stored privately under `state/when/` and hash-bound by a trust record the same way `bin/fm-check-register.sh` binds a custom check, while the spec separately binds the resolved action executable's bytes; a mutated or unregistered spec or a changed action executable is refused before the action runs, and that binding is reloaded from disk immediately before each fire rather than trusted from when polling started. +A repo update that fast-forwards an in-repo action's bytes in place would otherwise desync every already-armed watch's trust binding with no tampering involved; `bin/fm-procevent-when.sh rebind-all` re-hashes and republishes the binding for every registered watch whose action lives under `FM_ROOT`, including one already polling, so it keeps firing across such an update instead of being refused on its next fire. Every failure path - a mutated spec or action executable, a condition error past its budget, an expired deadline, a failed action, or an earlier fire whose outcome was never captured - produces a terminal captured outcome that wakes firstmate rather than a silent retry, and a durable single-fire marker claimed before the action makes restarts and re-polls unable to fire it twice. The adapter automates only the exact deterministic subset: anything needing judgment, and anything destructive, irreversible, or security-sensitive, keeps the ordinary check-fires-then-firstmate-decides flow, and the adapter's header and `--help` own its commands, flags, and outcome document. @@ -1076,11 +1093,19 @@ A live identity-matched owner is never displaced, and release removes only the e Every stop proves ownership before its first signal: the live runner's recorded process identity must match and it must still lead its process group. Once that stop has proved ownership and sent TERM, its own escalation to KILL checks only whether the proved group still has members; it does not re-read the leader's identity or group membership, which can change or become unreadable as TERM ends the leader. This proof belongs only to that stop's own escalation and cannot authorize another caller that encounters an unproved group. -A claim counts as reclaimable only when its owner is stale and an independent process-group check finds no members; a crashed leader or reused pid whose process group still has members cannot relax ownership cleanup, so reconcile preserves the claim without signalling the ambiguous group or starting a replacement. -If the leader dies to anything other than the stop's own signal, `retire`, `reconcile`, `sweep-home`, and the guard all refuse its surviving group permanently, and the source silently stops listening. -Whether that group may ever be signalled remains an open decision; the repaired guard does not close this gap. -Reclaiming a generation that IS gone is not gated on tidying its capture-reservation records. -Those records are keyed by claim token and every replacement claims a fresh one, so a leftover that can no longer be located - a state-root identity a claim recorded before its home was re-created, for example - is stale bytes rather than an ownership hazard. +A stale claim whose process group still has members is one `reconcile` never displaces, and the two shapes it comes in recover differently. +`reconcile` preserves such a claim without signalling the ambiguous group or starting a replacement: the group check probes the runner's own process group, which contains its polling source child, so surviving members can mean that child is still attached to the session the source collects from, and a replacement would put a second destructive poller on it. +`list` reports both shapes as `orphaned`. +When the recorded pid is alive under a different identity while the group still has members, the claim boundary itself does not consult the process group, so `bin/fm-procevent.sh start <source-id>` reclaims that claim provided the dead generation's reservation records can still be tidied, and otherwise refuses with `cannot claim source`; that tidy-up is waived only for a generation proven gone, which this one is not. +That hand-run command is the recovery path, taken by someone who has checked that nothing is still polling the source. +That asymmetry between the automatic path and the deliberate one is the design rather than an inconsistency, and it is not a claim-level invariant: nothing below `reconcile` enforces it. +When the leader itself is gone and its group still has members - the leader died to anything other than the stop's own signal - `start` does not reclaim the claim either: it reports `already owned` and changes nothing, and `retire`, `reconcile`, `sweep-home`, and the guard all refuse the surviving group permanently, so the source stops listening. +Recovery there is a human verifying whether the dead runner's polling child is still attached to the source; once that process group is empty the generation reads as gone and the next `reconcile` reclaims the source on its own. +Nothing automatic signals that group, and whether it may ever be signalled remains an open decision; the repaired guard does not close this gap. +Neither shape stops listening quietly: the first `reconcile` that strands a claim generation publishes a durable `check` wake naming the source and what clears it - the `start` command for the reused pid, the check to make for the leaderless group - and later cycles stay silent for that same generation while a genuinely new stranded claim announces again. +Reclaiming a generation that IS gone is not gated on tidying anything that generation left behind: its capture-reservation records, its staging file, or the registry directory a claim recorded for them. +Every one of those is keyed by claim token and every replacement claims a fresh one, so a leftover that can no longer be located or removed - a state-root identity a claim recorded before its home was re-created, or a recorded registry directory that no longer resolves to a directory - is stale bytes rather than an ownership hazard. +Making any of them a precondition is what leaves a provably dead runner owning its source permanently, because none of those conditions clears on its own. Ordinary release and reclamation still attempt reservation cleanup and require it unless both owner staleness and whole-group absence prove the generation gone. The narrow live-owner terminal-self-retirement path also attempts cleanup but tolerates its own still-in-flight reservation, which the runner removes on the normal end-of-capture path; exact home, PID, and claim-token ownership remains mandatory before the claim is released. If identity cannot be established before the first signal, or a surviving owned group cannot be proved stopped, the operation preserves the registration and claim for safe retry rather than adding a second owner. @@ -1119,6 +1144,26 @@ Scope is the owning state root and one runner generation, never a script or proc `FM_PROCEVENT_LAUNCH_FLOOR_SECONDS` (default 1, range 1..3600) is the minimum time between consecutive launches of one registration generation's stored command, bounding the launch rate of an immediately returning source during that lease window. The generation's first launch is immediate, later launches share its monotonic pacing timestamp, a timestamp from before a reboot is treated as expired, and replacing the registration starts a fresh pacing generation. +`FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS` (default 3, range 1..600) bounds how long `reconcile` waits for the runners it just started to prove they are running: never less than the configured value, and at most one second more, because the wait is measured on a whole-second clock. +Starting a runner is detached and its errors are not visible to the caller, so `reconcile` reports a start only after the source is observed owned or its launch-pacing stamp has advanced or appeared, and reports every unconfirmed launch as `failed=` and a non-zero exit instead. +Both signals are durable evidence a runner claimed: ownership is the only evidence a runner still blocked on its source ever shows, and the stamp - written after the claim and before the source command runs, and removed only by registration replacement - covers a runner that claimed, ran and exited between two polls. +A healthy launch therefore confirms on the first poll and the window only bounds a launch that has not yet proved itself - one that died before claiming, or one merely too slow to claim inside the window; confirmation cannot tell those apart, and a launch that proves itself on a later cycle closes its failure episode without a retraction wake. +All of a cycle's launches share one window, so a home full of sources that cannot start costs the same bounded wait as one. + +Keep this window well below `FM_POLL`. +`bin/fm-watch.sh` runs `reconcile` once per supervision cycle, so a source that cannot start makes every cycle wait up to the confirm window before the rest of that cycle runs. +Raising the confirm window lengthens every supervision cycle and delays wake delivery by up to that much. + +A source that can never start is reported as `failed=` with a non-zero exit on every `reconcile`, rather than counted as `started` and retried silently as though it were healthy, so a wedged source stays visible instead of presenting as armed. +That count reaches only whoever runs the command, because `bin/fm-watch.sh` discards `reconcile`'s output and exit status, so an unconfirmed launch is also announced through the wake queue: `reconcile` publishes a durable `check` wake (`procevent:<id>:launch-failed:<registration-identity>-<episode-nonce>`) once per failure episode, and later cycles stay silent for that episode until a launch of that source confirms, after which a fresh failure announces again under a fresh key, because the watcher never re-surfaces a key it has already surfaced. +The announcement changes nothing about the launch: `reconcile` keeps relaunching the source every cycle exactly as before, and nothing is retried differently, throttled, or recovered from that signal. +The wake says only what was observed for that shape - the launch did not prove it took the claim within the window - and, if it stays that way, names the source command and adapter binary the registration names as what to check and the attached `bin/fm-procevent.sh start <source-id>` as what reproduces a refusal on stderr, where the detached launch discards it; a later cycle that finds the source owned ends the episode on its own, so a runner that was merely slow to claim needs nothing from the operator. +A source stranded on a claim nothing may automatically displace is announced the same way, once per stranded claim generation, as described above. +`bin/fm-watch.sh` surfaces both under their own headlines - `process-event source stranded` and `process-event source failed to start` - rather than as a captured result. + +A value this command cannot use is refused by name before anything is launched, the same way `FM_PROCEVENT_LAUNCH_FLOOR_SECONDS` and `FM_PROCEVENT_MAX_OUTPUT_BYTES` are refused, so a mistyped window can never present as a fleet of sources that cannot start. +`bin/fm-watch.sh` validates the same value when it arms and refuses to arm on an unusable one, naming the variable and the range: under a running watcher that refusal would otherwise repeat on every cycle into a discarded stdout and leave the whole home disarmed while presenting as supervised, whereas a watcher that will not arm is loud through the liveness guard. + `FM_PROCEVENT_MAX_OUTPUT_BYTES` (default 1048576) bounds a single captured result while the source runs; oversized output is drained but truncated with a stderr notice rather than staged or published whole or dropped. The runner proves exactly one durability boundary: output that reached the runner is stored at mode `0600` before any event referencing it is published, and a captured result with no durable handled acknowledgement remains eligible for bounded re-announcement across any number of drains and restarts, not only the crash window right after capture. @@ -1175,6 +1220,7 @@ FM_ZELLIJ_SESSION=firstmate # zellij-only: named session for normal backend ops CMUX_SOCKET_PASSWORD= # cmux-only: socket password fallback when config/cmux-socket-password is absent (docs/cmux-backend.md) FM_SESSION_START_STATUS_TAIL=5 # state/*.status lines printed per task in the session-start digest; each line is capped by bin/fm-line-cap-lib.sh FM_SESSION_START_QUEUED_LIMIT=20 # plain queued backlog rows in the session-start digest; in-flight, held, and blocked rows are never bounded and done rows are never listed +FM_BACKLOG_ROW_TIMEOUT_SECS=10 # seconds bounding each backlog row read (bin/fm-backlog-transition-lib.sh); nonpositive or invalid values fall back to 10; the first bound hit latches the sweep so later reads return immediately, each still naming its own item FM_BOOTSTRAP_DETECT_ONLY=0 # internal/read-only session-start mode: skip bootstrap's mutating sweeps and print advisory TANGLE wording FM_BOOTSTRAP_NETWORK=all # internal session-start phase split: all, skip (local steps only), or only (network steps only); see bin/fm-bootstrap.sh FM_STARTUP_NETWORK_TIMEOUT=120 # seconds bounding the deferred inactive-outcome scan plus network checks; hitting it prints an actionable NETWORK_CHECKS line @@ -1228,6 +1274,7 @@ FM_PROCEVENT_CLAIM_ROOT= # machine-wide source claim root; defaul FM_PROCEVENT_OWNER_LEASE_SECONDS=600 # how long a source runner keeps going with no activity in its owning home; 1..86400 FM_PROCEVENT_OWNER_CHECK_SECONDS=15 # a runner guard's detection interval, read twice per interval; 1..3600 FM_PROCEVENT_LAUNCH_FLOOR_SECONDS=1 # minimum interval between launches of one registration generation's source command; 1..3600 +FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS=3 # how long reconcile waits for the runners it started to prove they are running; 1..600, keep well below FM_POLL FM_WHEN_OUTPUT_TAIL_BYTES=8192 # bound on the command-output tail inside one condition->action outcome document FM_CODEX_WATCH_CHECKPOINT=180 # seconds per foreground watcher checkpoint in Codex primary supervision FM_CREW_STATE_NM_TIMEOUT=10 # seconds allowed per no-mistakes query inside fm-crew-state.sh; bin/fm-fleet-snapshot.sh derives its own value from its per-task bound for the reads it makes, so this override does not reach those, and that script's header owns the derivation @@ -1304,7 +1351,7 @@ FM_COMPOSER_CAPTURE_LINES=20 # fleet-wide bound for tail-capture composer read FM_COMPOSER_PI_MAX_LINES=8 # fleet-wide: maximum rows admitted between Pi's identity-corroborated separator pair; taller or ambiguous candidates stay unknown FM_COMPOSER_GHOST_LUMA_MAX=128 # fleet-wide: max perceived luminance (0.299R+0.587G+0.114B, 0-255) for a TRUECOLOR foreground to count as de-emphasised ghost/placeholder text and be stripped; dim/faint (SGR 2) is stripped regardless. Assumes a dark terminal theme (bin/fm-composer-lib.sh's fm_composer_strip_ghost, used by styled tmux, herdr, and Zellij reads) GROK_HOME= # optional Grok config home for firstmate's global grok turn-end hook; defaults to ~/.grok -FM_SEND_RETRIES=3 # fm-send typed-plane Enter-retry attempts after typing the line once +FM_SEND_RETRIES=3 # fm-send typed-plane Enter-retry attempts after typing the line once; agy typed targets use a longer per-harness default owned by bin/fm-send.sh FM_SEND_SLEEP=0.4 # seconds between fm-send typed-plane submit checks FM_SEND_SETTLE=1 # seconds fm-send waits after a successful typed-plane submit; 0 disables FM_PENDING_REPLY_GRACE_SECS=120 # seconds after marked-request delivery before a completed turn without a correlated parent report is eligible for its one recovery repost diff --git a/docs/documentation-audiences.json b/docs/documentation-audiences.json index dee567b6c05..3e94fe4644e 100644 --- a/docs/documentation-audiences.json +++ b/docs/documentation-audiences.json @@ -343,6 +343,10 @@ "path": ".agents/skills/project-management/SKILL.md", "audience": "agent-runtime" }, + { + "path": ".agents/skills/quiet/SKILL.md", + "audience": "agent-runtime" + }, { "path": ".agents/skills/quota-array-dispatch/SKILL.md", "audience": "agent-runtime" @@ -663,6 +667,10 @@ "path": "docs/verification/dashboard-fleet-health.md", "audience": "maintainer-verification" }, + { + "path": "docs/verification/agy.md", + "audience": "maintainer-verification" + }, { "path": "docs/verification/dispatch-auth.md", "audience": "maintainer-verification" diff --git a/docs/fm-test-isolation-proof.md b/docs/fm-test-isolation-proof.md index e8df1878041..4afb7507660 100644 --- a/docs/fm-test-isolation-proof.md +++ b/docs/fm-test-isolation-proof.md @@ -20,6 +20,21 @@ This record owns concurrent isolation evidence for the portable parallel candida | failed | 0 | | wall duration | 113278 ms | +## Portable pool with two workers + +Verified on 2026-09-14 on Linux x86_64 with Git 2.54.0, Pi 0.85.1, TypeScript 7.0.2 and Ruby 3.4.9 available on PATH. +The command was `taskset -c 0,1 bin/fm-test-isolation-proof.sh --jobs 2`, with `--json` directed to a private evidence file. +The exact completion output was: + +```text +FM_ISOLATION_SUMMARY total=24 failed=0 concurrency=2 duration_ms=434026 +``` + +Every candidate ran without a gate skip, and the proof's Git-configuration and temporary-root isolation checks passed. +After completion, the proof root was absent and no live process retained a `TMPDIR` beneath it. +The two-CPU affinity bounds this local concurrency observation; it is not a measurement of hosted-runner speed or a guarantee of CI wall-time headroom. +[The CI workflow](../.github/workflows/ci.yml) owns its worker count and deadlines, while [portable shards](fm-test-portable-shards.md) explains how to interpret lane measurements. + ## Candidate set - `tests/fm-arm-pretool-check.test.sh` @@ -119,10 +134,10 @@ Both `bin/fm-test-run.sh` and the current proof harness therefore order concurre | 1 | `FM_ISOLATION_SUMMARY total=32 failed=0 concurrency=4 duration_ms=161837` | | 2 | `FM_ISOLATION_SUMMARY total=32 failed=0 concurrency=4 duration_ms=156462` | -This family is what a change to `bin/fm-test-run.sh` itself selects, so it decides that selection's wall clock. -Before admission, 14 of its scripts fell to the serial tail and the 33-script selection measured 327.3s against a 300s budget: the concurrent group was 19 scripts totalling 273.4s while the tail alone was 215.7s, dominated by `fm-calm-pi-extension` (77.5s), `fm-vendor-auth-probe` (51.0s), and `fm-muse-harness` (39.7s). +The current runner-change selection is owned by [`bin/fm-test-run.sh`](../bin/fm-test-run.sh)'s changed-file map. +Before admission, 14 of the family's scripts fell to the serial tail and the 33-script selection measured 327.3s against a 300s budget: the concurrent group was 19 scripts totalling 273.4s while the tail alone was 215.7s, dominated by `fm-calm-pi-extension` (77.5s), `fm-vendor-auth-probe` (51.0s), and `fm-muse-harness` (39.7s). Admitting the family moves that tail into the bounded concurrent group. -Current runner-file selection was verified on 2026-08-28 with the runner and its tests bound to each measured Bash version. +The then-current runner-file selection was verified on 2026-08-28 with the runner and its tests bound to each measured Bash version. Because the runner uses `#!/usr/bin/env bash` and invokes each test with `bash` from `PATH`, the stock macOS measurement used `PATH=/bin:$PATH bin/fm-test-run.sh --changed --max-wall-ms 300000` so both resolved to `/bin/bash` 3.2.57. Two runs selected all 33 scripts, passed the five-minute result check in 153.5s and 166.8s, and reported the same two failures as `main`: `tests/fm-muse-harness.test.sh` and `tests/fm-composer-lib.test.sh`. With Bash 5.3.9 on `PATH`, three runs of `bin/fm-test-run.sh --changed --max-wall-ms 300000` selected the same 33 scripts, completed with 0 failures, and reported 163.8s, 172.0s, and 166.9s. diff --git a/docs/fm-test-portable-shards.md b/docs/fm-test-portable-shards.md index ba8207a9678..99824d0e551 100644 --- a/docs/fm-test-portable-shards.md +++ b/docs/fm-test-portable-shards.md @@ -5,6 +5,11 @@ ## Verification inputs +Balance hints come from serial measurements of the real lanes on `ubuntu-latest`. +The current workflow overlaps the admitted isolated scripts within each parallel job; [`.github/workflows/ci.yml`](../.github/workflows/ci.yml) owns its worker count and unchanged job caps. +The concurrent isolation proof in [fm-test-isolation-proof.md](fm-test-isolation-proof.md) establishes concurrency safety, not serial CI duration. +Local timings are not interchangeable with CI timings: platform and machine load can affect each script differently and change their relative weights. + The proven-isolated candidate set remains the 24-script concurrent proof recorded in [fm-test-isolation-proof.md](fm-test-isolation-proof.md). The 2026-09-14 placement refresh uses completed script measurements in [CI run 34837015817](https://github.com/HelloWorldSungin/firstmate/actions/runs/34837015817). Its first parallel lane reached the unchanged ten-minute job limit after five completed scripts, with the enlarged hold and merge suites accounting for 504 seconds together. @@ -41,15 +46,16 @@ These measurements change placement, not isolation eligibility or execution dead ## Parallel lanes -The two parallel lanes use longest-processing-time assignment from those measured durations. - -| Lane | Script count | Estimated duration | -|---|---:|---:| -| `portable-parallel-1` | 12 | 544542 ms (~9.08 min) | -| `portable-parallel-2` | 12 | 544509 ms (~9.08 min) | -| imbalance | | 33 ms | +The two parallel lanes use longest-processing-time assignment over those hints. +[`bin/fm-test-run.sh`](../bin/fm-test-run.sh) holds the duration values in `portable_parallel_weight_hints` and the ordered memberships and lane-specific prerequisite constraints beside `list_portable_parallel_1` and `list_portable_parallel_2`. +Read the derived packing estimates with that runner's `--check-coverage`; its header and `--help` own the output fields and the selection-specific `--list-scheduled` weight rules. +The largest individual hint sets a lower bound on the estimated duration of any split, regardless of how evenly the remaining work is assigned. +The CI cap and its rationale are owned by [`.github/workflows/ci.yml`](../.github/workflows/ci.yml). -`bin/fm-test-run.sh` contains the exact ordered memberships in `list_portable_parallel_1` and `list_portable_parallel_2`. +[`tests/fm-test-run.test.sh`](../tests/fm-test-run.test.sh), in `test_portable_parallel_lanes_stay_duration_balanced`, requires every parallel member to have a hint and the lane sums to differ by no more than five percent of the larger sum. +Its scheduling regressions also check stored parallel lane order and preserve serial-weight scheduling for other selections. +These checks do not detect a script outgrowing an existing hint or establish measured job headroom. +Refresh `portable_parallel_weight_hints` with the slowest completed `duration_ms` per script from several green CI runs' `fm-test-timing-portable-parallel-*` artifacts whenever the parallel set gains scripts or a member grows materially. ## Portable serial remainder @@ -83,30 +89,16 @@ That is not hypothetical: by 2026-09-01 the lane had grown from 116 to 139 scrip `bin/fm-test-run.sh --check-coverage` now reports the unmeasured share as `serial_unhinted=` and refuses past `PORTABLE_SERIAL_MAX_UNHINTED_PERCENT`, so hint drift fails the coverage guard instead of silently pushing one shard into its job cap. Refresh the hints whenever the serial lane gains scripts, rather than waiting for that bound to trip. -Shard count is sized from that total rather than left where an earlier, smaller remainder put it. -The lane grew from about 19 minutes across 69 scripts to about 58 minutes across 154, which four shards could no longer carry inside the job timeout: on the run above, `portable-serial-2of4` was cancelled at 15 minutes having finished 24 of its 32 scripts, and the hints then put a perfectly balanced quarter at 14.5 minutes, still on the tripwire rather than inside it. -The merged shared maxima and retained fork-only hints put the slowest of eight shards at about 13.69 minutes. -The current 209-script serial lane has six unhinted scripts, using the runner's conservative default, and totals 6572397 ms of assignment weight. -The 30-minute serial job cap adopted from upstream leaves setup and runner-speed margin; the fork's 480-second per-script bound remains unchanged. - -| Lane | Script count | Estimated duration | -|---|---:|---:| -| `portable-serial-1of8` | 26 | 821551 ms (~13.69 min) | -| `portable-serial-2of8` | 26 | 821546 ms (~13.69 min) | -| `portable-serial-3of8` | 26 | 821547 ms (~13.69 min) | -| `portable-serial-4of8` | 26 | 821552 ms (~13.69 min) | -| `portable-serial-5of8` | 27 | 821587 ms (~13.69 min) | -| `portable-serial-6of8` | 26 | 821546 ms (~13.69 min) | -| `portable-serial-7of8` | 26 | 821534 ms (~13.69 min) | -| `portable-serial-8of8` | 26 | 821534 ms (~13.69 min) | -| imbalance | | 53 ms | +The fork retains eight portable serial shards and its 480-second per-script bound. +`bin/fm-test-run.sh` owns the per-shard packing, so its `--check-coverage` output is the current account of lane size, shard composition, and balance rather than a copied table. +Run 34342484144 observed a shard reach about 20 minutes of passing work, so the 30-minute job cap keeps meaningful hang-tripwire margin for job setup and runner-speed spread. The watcher triage cases are split into core and wait/decision scripts with one shared fixture owner in `tests/watch-triage-helpers.sh`. All 125 original cases remain in exactly one script, with compatible new upstream progress and declared-deadline cases added beside them. Their initial passing local measurements were 160205 ms and 191775 ms after CI reached the unchanged 480-second combined-script limit while still passing cases. The refreshed CI hints are 236467 ms for core and 292716 ms for waits. -Refresh the CI-derived hints by downloading the per-shard timing artifacts from several green CI runs, replacing the `portable_serial_weight_hints` table in `bin/fm-test-run.sh` with the slowest measured `duration_ms` per `path`, and updating the table above: +Refresh the CI-derived hints by downloading the per-shard timing artifacts from several green CI runs and replacing the `portable_serial_weight_hints` table in `bin/fm-test-run.sh` with the slowest measured `duration_ms` per `path`: ```sh for run in <run-id> <run-id> <run-id>; do @@ -143,11 +135,11 @@ Portable shards, each portable serial shard, and the Herdr lane upload runner-ge | Lane | Bound | Rationale | |---|---|---| -| portable parallel 1/2 | job `timeout-minutes: 10` | The measured shard sums are about 9.08 minutes, leaving about 55 seconds for setup and runner variation. | -| portable serial 1-8 | job `timeout-minutes: 30` | The slowest estimated shard is about 13.69 minutes, leaving setup and runner-speed margin. | -| Herdr | family-run step `timeout-minutes: 20`; job `timeout-minutes: 75` backstop | The required lane is bounded independently of the per-script deadline; refresh timings from its uploaded artifacts. Previous healthy runs finished around 7 minutes, so the step bound is the hang tripwire (cleanup and timing artifacts still upload) while the job cap stays a last-resort backstop. | +| portable parallel 1/2 | See [CI workflow](../.github/workflows/ci.yml) | The workflow owns the parallel cap rationale and its evidence limits. | +| portable serial 1-8 | job `timeout-minutes: 30` | Current runners can take about 20 minutes; the 30-minute cap remains a hang tripwire while leaving margin for job setup and runner-speed spread. | +| Herdr | family-run step `timeout-minutes: 20`; job `timeout-minutes: 75` backstop | Healthy runs finished around 7 minutes before this lane gained `fm-backend-herdr-focus-flash-e2e`, which measures about 2 minutes against a real lab locally, so the step bound is still the hang tripwire (cleanup and timing artifacts still upload) while the job cap stays a last-resort backstop. Refresh this figure from the lane's uploaded timing artifact. | -Timeouts are hang tripwires rather than expected healthy durations. +Timeouts are intended as hang tripwires; a passing coverage guard does not establish a healthy job duration. `.github/workflows/ci.yml` owns the exact numbers. Inside each lane, `bin/fm-test-run.sh` applies its own default per-script bound, so a hung script usually turns red with per-script attribution before the job cap cancels the lane; its `--help` owns that bound's value and opt-out, and the rationale beside `DEFAULT_PER_SCRIPT_TIMEOUT_SECS` owns the per-lane margin arithmetic. diff --git a/docs/fork-divergence.md b/docs/fork-divergence.md index f61f03b4205..18821a49156 100644 --- a/docs/fork-divergence.md +++ b/docs/fork-divergence.md @@ -17,9 +17,12 @@ Each round brings one upstream merge through its own reviewable PR, and the conf ### Agy crew adapter The fork carries an Antigravity CLI adapter so workers can use the operator's paid Gemini subscription. -Upstream has no agy support at all, so every part of it is fork-local. +Upstream added agy detection, model validation, trust registration, and readiness checks in `kunchenguid/firstmate#4200`, with ancestry precedence refined in `kunchenguid/firstmate#3578`. +The fork adopts those compatible capabilities while retaining the following scope and native-state contracts. It is deliberately crew-only and Herdr-only because Herdr supplies the native identity, liveness, working-state, and delivery signals needed to supervise that CLI without treating screen text as authority. It is not selectable for the primary firstmate or a persistent second mate, and its raw-command bypass is rejected so the kind, backend, trust, and supervision guards cannot be skipped. +The shared settings mutation owner remains `bin/fm-agy-trust-lib.sh`, preserving created-versus-preexisting ownership, abort rollback, teardown cleanup, and locking. +Upstream's `bin/fm-agy-trust.sh` retains its linked-worktree scope checks and logical/resolved path registration while delegating writes to that owner. The adapter contract lives in the [`harness-adapters` skill's agy reference](../.agents/skills/harness-adapters/references/harness/agy.md), reachable through that skill's routing artifact, and its executable guards live in `bin/fm-spawn.sh`, `bin/fm-launch-lib.sh`, and [`tests/fm-agy-adapter.test.sh`](../tests/fm-agy-adapter.test.sh). Herdr's atomic agy prompt emits a fork-local `unverifiable` send verdict for an acceptance it can prove neither way, which `bin/fm-send.sh` keeps distinct from upstream's `pending`: both exit 3, but `unverifiable` marks the request's delivery state unknown while `pending` discards it as undelivered. Upstream's durable steering inbox (`kunchenguid/firstmate#2856`, reached 2026-08-25) moved ordinary local text off the typed submit path, so both verdicts now describe the typed plane alone - a harness-native invocation or an explicit backend target - while an ordinary steer's delivery is its durable record. @@ -31,6 +34,30 @@ Their launch templates and flag construction use the fork's existing `bin/fm-lau The dead-endpoint doorbell protection from `kunchenguid/firstmate#3823` preserves the harness-aware atomic submit path, while its shell-no-op prefix and partial-line race describe typed transport. This entry covered Cursor as well until 2026-08-18; see "Fork-local cursor crew adapter" under retired divergences. +### Attended quiet lifecycle + +Upstream introduced presentation-only quiet supervision in `kunchenguid/firstmate#4337`, but its away launcher still required a confirmed away record and its return guard refused ordinary work for any quiet flag. +Firstmate resolved this contradiction in favor of the feature's attended semantics during the pinned September 14 sync. +`bin/fm-afk-launch.sh` owns mode-aware entry through the existing single daemon, and `bin/fm-afk-return.sh` owns the read-only guard and explicit `quiet-off` exit. +Quiet grants no away authority, does not fabricate an away record, and cannot consume an actual away return or its pending catch-up gate. +Ordinary work leaves quiet active; explicit quiet exit stops the shared daemon without an away brief or archive. +The agent-facing lifecycle is owned by the [quiet skill](../.agents/skills/quiet/SKILL.md), and `tests/fm-afk-launch.test.sh` plus `tests/fm-afk-return.test.sh` exercise quiet alongside the unchanged confirmed-away entry and return gates. + +### Herdr terminal-wide agent absence + +Upstream `kunchenguid/firstmate#4191` adds a shared harness process classifier and distinguishes a stale registration from an unregistered pane. +The fork adopts those states and positive harness detection while retaining its stricter terminal-wide absence proof in `bin/backends/herdr.sh`. +A shell-only foreground cannot prove absence when a suspended, backgrounded, reparented same-terminal agent, or an agent hosting an escape shell remains. +The shared negative proof accepts the nested Treehouse shell chain only after accounting for every process on that terminal. +An unreadable process-info response now reports `unreadable` rather than `alive`; both continue to refuse recovery, while proven stale registration permits recovery but not destructive husk cleanup. +`tests/fm-backend-herdr.test.sh` exercises these distinct decisions, and `tests/fm-herdr-pi-stale-registration-live-e2e.test.sh` refreshes the actual vendor evidence in [runtime backend verification](verification/runtime-backends.md#stale-agent-registration). + +### Installed AXI compatibility floors + +The fork retains its installed-version AXI floors when upstream introduces a feature at an older compatible release. +The floor policy and current constants are owned by `bin/fm-bootstrap.sh`; this round keeps Lavish at `0.1.62` while adopting upstream's optional presentation behavior introduced against `0.1.46`. +`tests/fm-bootstrap.test.sh` exercises below-floor refusal and compatible versions, and `tests/fm-brief.test.sh` verifies that visual scout instructions follow the same effective floor. + ### Pinned ShellCheck download retry budget The fork keeps its own wall-time download retry budget in [`bin/fm-install-shellcheck.sh`](../bin/fm-install-shellcheck.sh) rather than upstream's `DOWNLOAD_ATTEMPTS` count. @@ -174,9 +201,10 @@ Serving that bound, [`bin/fm-timeout-lib.sh`](../bin/fm-timeout-lib.sh) publishe The rationale beside the constant owns why 480s and what the bound costs each CI lane, including the accepted margin on the required real-Herdr lane recorded in `HelloWorldSungin/firstmate#256`, and the script's `--help` owns the flag contract and its `0` opt-out. [`tests/fm-test-run.test.sh`](../tests/fm-test-run.test.sh) pins the default arming, the opt-out, the exit 124 versus exit 125 distinction, and the signal relay, so an upstream round that rewrites the timeout wiring cannot retire the default silently. -The upstream `kunchenguid/firstmate#3489` rebalance refreshes shared duration hints and raises the portable serial job cap to 20 minutes. -The fork retains eight shards and its fork-only timing hints. -The fork adopts the upstream 30-minute serial job cap and takes per-script maxima across both parents; the runner still owns the 480-second per-script bound and the updated margin arithmetic. +Upstream `kunchenguid/firstmate#4151` refreshes shared duration hints and raises the portable serial job cap to 30 minutes. +The fork retains eight shards and takes per-script maxima across both parents, preserving its fork-only timing hints. +The runner still owns the 480-second per-script bound and the updated margin arithmetic. +The fork overlaps already-admitted isolated scripts inside its existing portable parallel jobs; [`.github/workflows/ci.yml`](../.github/workflows/ci.yml) owns worker count and caps, with current concurrency evidence in [`docs/fm-test-isolation-proof.md`](../docs/fm-test-isolation-proof.md). Watcher triage cases are partitioned between `tests/fm-watch-triage.test.sh` and `tests/fm-watch-triage-waits.test.sh`, with shared case definitions in `tests/watch-triage-helpers.sh`, so suite growth does not weaken the per-script bound. @@ -353,7 +381,8 @@ A future round must take upstream's queue behaviour unchanged rather than re-der One consequence is left standing deliberately and is the thing to read before "fixing" it. The forwarded-argument allow-list still refuses `--auto` by name, which is a separate rule about what a caller may hand the forge and not a claim about merge queues, so upstream's restored retry guidance names flags this script would itself turn away. -That allow-list is additive rather than divergent - it constrains what a caller may pass, admitting merge-method selectors, message arguments, post-execution branch cleanup, and head-binding arguments, each recorded with why it cannot defer execution. +That allow-list is additive rather than divergent - it constrains what a caller may pass, admitting merge-method selectors, message arguments, and post-execution branch cleanup, each recorded with why it cannot defer execution. +Upstream `kunchenguid/firstmate#4199` now binds execution to the freshly verified head internally, so caller-supplied head bindings are refused to prevent replacing that proof. Reconciling the two, by retiring that allow-list entry as well or by rewording the guidance a second time, is open work left to the captain who ordered the restoration; no issue tracks it yet, and `HelloWorldSungin/firstmate#259` is the allow-list's own open issue rather than a record of this. #### ACCEPTED and standing: a failed command whose readback confirms the merge landed @@ -375,6 +404,8 @@ This fork carries the optional drift detector, bootstrap diagnostic, sync-round The feature is inert when no `upstream` git remote exists so upstream users do not acquire fork behavior merely by taking another change. The detector fetches only into a disposable repository and never changes the source repository's objects, refs, index, branch, or worktree. A sync request dispatches an isolated merge task and PR rather than merging in the primary copy or extending `/updatefirstmate` with merge behavior. +The standing Firstmate review and landing authority lives in [`sync-upstream`](../.agents/skills/sync-upstream/SKILL.md); the exact-head expected-policy exception is owned by `github_verified_upstream_sync` in [`bin/fm-pr-merge.sh`](../bin/fm-pr-merge.sh) and exercised by [`tests/fm-pr-merge.test.sh`](../tests/fm-pr-merge.test.sh). +This deliberately preserves direct-PR upstream parentage while retaining every substantive-check, publishing, task-hold, default-tip and away-authority gate. ### Repository-local validation evidence diff --git a/docs/gitlab-merge-watch.md b/docs/gitlab-merge-watch.md index 0483b0e5557..215d75c0ab9 100644 --- a/docs/gitlab-merge-watch.md +++ b/docs/gitlab-merge-watch.md @@ -239,7 +239,7 @@ $ echo $? A project that runs no pipeline at all therefore cannot merge through this path. That is the intended reading of the requirement rather than an oversight: a successful pipeline at the head is a condition, and "there is no pipeline" does not satisfy it. -Both refusals came after `pr=` was recorded and the merge poll was armed, exactly as a failing `gh-axi pr merge` does on the GitHub side, so a refusal still leaves the audit trail and the watch in place. +Both refusals came after `pr=` was recorded and the merge poll was armed, as a failed live verification or `gh pr merge` does on the GitHub side, so a refusal still leaves the audit trail and the watch in place. A recorded `pr_head=` that no longer matches the live head is reported, and the live head is what gets verified. The stale value below was written into the task record by hand, because a GitLab task never records one on its own: diff --git a/docs/herdr-backend.md b/docs/herdr-backend.md index c3eb6e87010..4ce26d056c1 100644 --- a/docs/herdr-backend.md +++ b/docs/herdr-backend.md @@ -189,6 +189,7 @@ Each entrypoint creates its own source, homes, lab, and evidence; both are selec `tests/fm-herdr-session-cleanup-e2e.test.sh` covers the restored-shell cleanup in a guarded non-default named lab. `tests/fm-backend-herdr-focus-flash-e2e.test.sh` reproduces the raw explicit-close focus steal on the installed release and proves the focus-safe emptying-close plan removes a doomed workspace with no wrong-focus interval; [`verification/runtime-backends.md`](verification/runtime-backends.md#workspace-removal-focus-safety) owns the active versioned evidence. `tests/fm-backend-herdr-stale-active-tab-e2e.test.sh` proves a persisted-focused tab still closes when no foreground client is attached. +`tests/fm-herdr-attached-viewer-live-e2e.test.sh` proves the other half against a real attached viewer, which `bin/fm-herdr-lab.sh viewer start` supplies over a pty sized before the fork; [`verification/runtime-backends.md`](verification/runtime-backends.md#attached-foreground-viewer) owns the active versioned evidence and the re-run trigger. ## Default-tab prune safety @@ -297,20 +298,16 @@ A restored same-labeled tab with a missing pane or no registered agent is a husk Create replaces only a confidently dead or no-agent husk, creates the replacement before closing the old tab, and refuses live or unknown states. This prevents closing the workspace's last tab before a replacement exists. -The generic Herdr agent-liveness probe reuses the same pane classifier, then applies one recovery-only exception. -A structurally gone pane or a pane read from a session positively reported as having no running server becomes `missing`, a restored agent-less shell becomes `dead`, a registered agent becomes `alive` unless its registration is proven stale, and every other unexpected read becomes `unreadable`. -The stopped-server exception does not widen husk detection or any close authority; those paths still refuse an unreadable pane. -Unlike tmux process-name inspection, native registration can classify Pi without guessing from a generic interpreter name. -`tests/fm-backend-herdr-agent-exit-shell-e2e.test.sh` pins the live-Pi versus leftover-shell distinction; [`verification/runtime-backends.md`](verification/runtime-backends.md#agent-lifecycle-control) owns the versioned evidence. - -A registration is not trusted on its own, because Herdr can hold one that outlives its agent. -An agent whose lifecycle reporting is hook-authoritative (Pi's `herdr:pi` extension) never deregisters on exit, Herdr skips screen detection for it, and Herdr exposes no agent-deregister verb, so `agent get` keeps reporting the last agent state forever after any exit, clean or killed. -The recovery-grade probe therefore cross-checks every registered agent against an operating-system agent-free proof: the pane's process tree must be nothing but sleeping recognized shells plus the resident treehouse worktree wrapper, with the foreground held by an idle shell that is the tree's only leaf. -That is exactly the shape a task pane is left in when its agent dies (pane shell, treehouse wrapper, worktree subshell), so a stale registration over it classifies `dead`, while a foreground, suspended, backgrounded, or shell-hosting agent breaks one of those rules and stays `alive`. -Any inconclusive read also fails the proof, so recovery stays refused unless the pane is positively agent-free. -The proof is deliberately not a screen read; a rendered prompt glyph cannot distinguish a dead shell from every agent composer. -It is a separate, classification-only predicate beside the stricter lone-idle-shell proof under [Presentation spaces](#presentation-spaces), which licenses signaling a shell pid and stays unchanged. -`bin/backends/herdr.sh`'s `fm_backend_herdr_agent_state` owns the exact contract. +A registration alone never proves an agent: hook-authoritative Pi registrations can survive both clean and killed exits. +The shared positive process classifier in `bin/fm-agent-process-lib.sh` recognizes running harnesses, while `bin/backends/herdr.sh` owns the stricter terminal-wide absence proof and its bounded settling. +That negative proof accounts for the nested Treehouse shell chain, background and suspended jobs, same-terminal processes outside the shell tree, and escape shells before labeling a registration `stale-agent`. +An unreadable process view remains `unknown`, and a working registration over a proven shell-only pane cannot authorize a busy verdict. + +The generic recovery probe maps a structurally gone pane or positively stopped server to `missing`, an unregistered shell or proven stale registration to `dead`, and a readable registered agent with live process evidence to `alive`. +Other unexpected reads remain `unreadable`. +These recovery verdicts do not widen destructive husk cleanup: `stale-agent` and unknown panes are never closed as husks. +The separate lone-idle-shell proof under [Presentation spaces](#presentation-spaces) still owns permission to signal a shell pid. +`tests/fm-backend-herdr-agent-exit-shell-e2e.test.sh` and `tests/fm-herdr-pi-stale-registration-live-e2e.test.sh` refresh the live distinction; [runtime verification](verification/runtime-backends.md#stale-agent-registration) owns measured versions and process-info compatibility evidence. The session-start sweep uses this probe. Mid-session secondmate agent-process liveness is not implemented because idle secondmates are deliberately exempt from stale-pane escalation and need a separate periodic identity signal. @@ -384,10 +381,12 @@ tests/fm-backend-herdr-launcher-workspace-e2e.test.sh tests/fm-backend-herdr-presentation-e2e.test.sh tests/fm-backend-herdr-recovery-e2e.test.sh tests/fm-backend-herdr-agent-exit-shell-e2e.test.sh +tests/fm-herdr-pi-stale-registration-live-e2e.test.sh tests/fm-backend-herdr-eventwait-smoke.test.sh tests/fm-control-herdr-smoke.test.sh tests/fm-herdr-session-cleanup.test.sh tests/fm-herdr-session-cleanup-e2e.test.sh +tests/fm-herdr-attached-viewer-live-e2e.test.sh tests/fm-afk-inject-herdr-e2e.test.sh tests/fm-afk-pi-herdr-return-e2e.test.sh tests/fm-busy-state.test.sh diff --git a/docs/scripts.md b/docs/scripts.md index f26513354fc..e2a09c93b4f 100644 --- a/docs/scripts.md +++ b/docs/scripts.md @@ -46,6 +46,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-brief.sh` | Scaffold ship or one-conversation ADR design briefs with explicit `--mode`, plus scout, secondmate-charter, and Herdr-lab briefs, with intent/spec subsections and opt-in work-item traceability | | [`fm-dod-lib.sh`](../bin/fm-dod-lib.sh) | Own ship/design/scout worker role scope, task definitions of done, and the no-mistakes `--intent` contract | | `fm-herdr-lab.sh` | Provision and guardedly operate an isolated, never-default Herdr lab session | +| `fm-herdr-lab-viewer.py` | The pty engine behind `fm-herdr-lab.sh viewer`: one real foreground Herdr client on a non-zero window grid | | `fm-install-herdr.sh` | Install CI's exact-version Herdr pin with official asset URL, SHA-256, and protocol checks | | `fm-install-treehouse.sh`| Install CI's exact-version Treehouse pin for real-Herdr E2E that needs spawn worktrees | | `fm-herdr-ci-cleanup.sh` | Snapshot and tear down only job-owned `fm-lab-*` sessions in the Herdr CI lane | @@ -73,6 +74,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-backend.sh` | Runtime-backend selection, meta helpers, selector resolution, and operation dispatch | | `fm-backend-hometag-lib.sh` | Shared per-installation home-tag derivation for zellij tab and cmux workspace titles | | `fm-composer-lib.sh` | Single fleet-wide owner of composer shapes, capability-aware screen classification, and verdicts | +| `fm-agent-process-lib.sh` | Backend-neutral harness-process name classifier shared by the tmux and herdr adapters | | `backends/tmux.sh` | Verified tmux session-provider adapter | | `backends/herdr.sh` | Herdr session-provider adapter with its own required CI lane | | `backends/zellij.sh` | Experimental zellij session-provider adapter | @@ -105,7 +107,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-watch-checkpoint.sh` | Run one bounded foreground watcher checkpoint for Codex-style supervision | | `fm-watch.sh` | Singleton-safe watcher: absorb benign wakes, detect stalled local-secondmate wake queues, and exit on actionable ones | | `fm-inactive-reconcile.sh` | Reconcile long-inactive direct crewmate terminal outcomes without forge access | -| `fm-afk-contract.sh` | Own the away-posture record: schema, mandate-clause fields and never-set scan, refusal naming the missing part, read-back, entry announcement, archive | +| `fm-afk-contract.sh` | Own the away-posture record: schema, mandate-clause fields and never-set scan, refusal naming the missing part, read-back, entry announcement, archive, and cross-subsystem authority lock | | `fm-afk-start.sh` | Run the common sourceable away-mode daemon entry in the foreground | | `fm-afk-launch.sh` | Own away-mode entry (read-back, confirm, record), exit, rollback, and any backend terminal lifecycle | | `fm-afk-return.sh` | Own deterministic return shutdown, the return brief, catch-up evidence, and the firstmate-actionable blocker gate | @@ -123,6 +125,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-lock-lib.sh` | Shared "is this git lock provably abandoned?" proof used by teardown and fleet-sync | | `fm-timeout-lib.sh` | Own bounded external command execution and the per-call share of a whole operation's budget | | `fm-config-inherit-lib.sh` | Shared primary-to-secondmate inherited local-material propagation and config-reread delivery | +| `fm-tasks-axi.sh` | Run `tasks-axi` against this home's backlog from any working directory | | `fm-tasks-axi-lib.sh` | Shared backlog-backend selector and `tasks-axi` compatibility probe | | `fm-backlog-transition-lib.sh` | Pair task-record changes with their backlog transitions and replay interrupted closes | | `fm-quota-axi-lib.sh` | Shared `quota-axi` compatibility floor and quota snapshot schema validation | @@ -169,6 +172,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-outcome-lib.sh` | Own the durable manifest, work-item, PR-status, and history wire shapes plus their atomic publication | | `fm-outcome-manifest.sh` | Write, read, and list the durable completion manifest teardown publishes before cleanup | | `fm-work-item.sh` | Maintain a task's durable forge- and host-agnostic work-item reference store | +| `fm-merge-authority-lib.sh` | Resolve merge authority at the gate, persist it against the accepted canonical PR, and identity-check its later poll consumption | | `fm-parent-channel-lib.sh` | Resolve a secondmate home's parent channel and append a captain-facing outcome line to it at most once | | `fm-promote.sh` | Promote a scout task in place to a protected ship task with an explicit delivery mode, and write the ship instructions carrying that mode's definition of done | | `fm-teardown.sh` | Fail-closed teardown: return landed ship or design worktrees, reap that one task's branch once its merge is proven, close this home's backlog item, require design decision inventory or completed scout deliverables, retire secondmate homes | diff --git a/docs/sessionstart-nudge.md b/docs/sessionstart-nudge.md index 68f9b9c4ecb..17d44c93c44 100644 --- a/docs/sessionstart-nudge.md +++ b/docs/sessionstart-nudge.md @@ -42,7 +42,8 @@ On a run-tier harness the nudge cannot also fire: `resume`, `reload`, and `fork` The run tier blocks either hook-driven session initialization or Pi's first provider preflight while the digest runs, so `bin/fm-session-start.sh` bounds itself rather than betting on an unbounded prerequisite. The digest makes no external-network call at all: every one it owes runs off the blocking path in the separately bounded deferred stage owned by `bin/fm-startup-network.sh`, so an unreachable host can no longer consume this budget. -What remains is still not individually bounded - tool version probes, the backlog listing, and the per-task endpoint reads are all local but unbounded subprocesses - so the whole digest runs as one bounded child, default 120s via `FM_SESSION_START_TIMEOUT`. +Tool version probes, the backlog listing, and the per-task endpoint reads remain local but unbounded subprocesses, so the whole digest still runs as one bounded child, default 120s via `FM_SESSION_START_TIMEOUT`. +The per-item backlog row reads inside bootstrap's reconcile and close-replay sweeps are the exception: each is bounded by `FM_BACKLOG_ROW_TIMEOUT_SECS` (default 10s) through `bin/fm-backlog-transition-lib.sh`, and the first bound hit latches the sweep so later reads return immediately while still naming their own item. The shared timeout owner falls back to a pure-Bash process-group watchdog when timeout, gtimeout, and perl are unavailable, so no supported host runs the digest unbounded. Because the child streams into the native transport as it runs, everything emitted before the bound was hit is retained for delivery; the parent then prints a `STARTUP TRUNCATED` banner naming the stage that did not finish and the stages that were therefore never emitted, and still exits 0. The registered hook timeouts sit above that budget so the harness never preempts the banner. diff --git a/docs/tmux-backend.md b/docs/tmux-backend.md index 19a7066c171..b9e7e64513a 100644 --- a/docs/tmux-backend.md +++ b/docs/tmux-backend.md @@ -48,7 +48,8 @@ Verify setup by spawning a small task and confirming its `fm-<id>` window appear A target-existence check proves only that the pane exists. The deeper tmux agent-liveness probe first verifies exact window membership, then reads process names to distinguish a running harness from a bare idle shell. -It classifies recognized harness process identities as `alive`, common shells as `dead`, an authoritatively absent window as `missing`, unreadable state as `unreadable`, and every other process as `ambiguous`. +It classifies recognized Claude, Codex, OpenCode, Pi, pi-signed, Grok, Kimi, Cursor, Muse, Rovo, and AGY process identities as `alive`, common shells as `dead`, an authoritatively absent window as `missing`, unreadable state as `unreadable`, and every other process as `ambiguous`. +The process-name vocabulary behind those verdicts is owned by `bin/fm-agent-process-lib.sh` and shared with the Herdr adapter, which proves a registered agent against the same names ([herdr-backend.md](herdr-backend.md) "Restart and liveness behavior"). Only `dead` and `missing` authorize recovery because a false dead result could launch a duplicate agent. For positive attribution, the probe combines two independent name sources rather than making either one load-bearing. @@ -61,6 +62,7 @@ The same scoping covers multi-process launchers without a special case, so the P Direct executable identities `pi`, `pi-signed`, and `Pi` remain accepted exactly, and similar or prefixed process names are not accepted through those exact Pi-family entries. Muse is likewise anchored to the exact `muse` launcher identity or the installed `muse-bin-<version>` prefix, so unrelated names such as `musescore` and `amuse` remain ambiguous. omp is anchored to the exact `omp` identity for the same reason, so `ompd` and `comp` remain ambiguous. +AGY is anchored to the exact `agy` identity for the same reason, so unrelated names containing that fragment remain ambiguous. Cursor is identified from its exact `cursor-agent` identity or versioned install tree in the foreground process path or structured argv[0]; a bare `node` or unrelated `agent` remains ambiguous. Gemini's Node bundle is identified through its script argument under the structural rules owned by [`bin/fm-gemini-lib.sh`](../bin/fm-gemini-lib.sh), rather than by a generic Node process name. diff --git a/docs/turnend-guard.md b/docs/turnend-guard.md index 4a50b2a4c20..adcbd25529f 100644 --- a/docs/turnend-guard.md +++ b/docs/turnend-guard.md @@ -15,7 +15,7 @@ Do not infer this guard's scope, loop safety, or compatibility tradeoffs for tho The turn-end guard closes the remaining gap at the primary's own turn boundary. When work, a process-event source, a registered custom check, Relay polling, or a pending wake queue needs supervision at that boundary and no identity-matched watcher has a fresh beacon, the harness integration must either block the turn end or force one bounded follow-up that uses the recovery instruction from the emitted session-start protocol. The mid-turn pull warning uses the model-aware supervision verdict described below, while the turn-end guard keeps the PID-strict watcher predicate. -Away mode is the one place the turn-end guard accepts a different supervisor: while `state/.afk` exists the away-mode daemon owns supervision, so a live identity-matched daemon with a fresh beacon satisfies that boundary in place of a watcher process holding the lock. +Away and quiet mode are the one place the turn-end guard accepts a different supervisor: while `state/.afk` exists, in either mode (`bin/fm-wake-lib.sh`'s `fm_afk_mode`), the daemon owns supervision, so a live identity-matched daemon with a fresh beacon satisfies that boundary in place of a watcher process holding the lock. The guard remains a backstop; [`watcher-continuity.md`](watcher-continuity.md) owns normal continuity. ## Guard predicates @@ -48,13 +48,13 @@ Without that proof an unheld lock alarms exactly as it did before, so an unloade Under every persistent-watcher harness a live identity-matched watcher with a fresh beacon is still required, so the pull guard keeps the same strict semantics there. Its banner names the true failing condition, either a missing live watcher process or a genuinely stale beacon with its real age, and keys the once-per-episode dedup on that condition rather than the beacon mtime. -While `state/.afk` exists the away-mode daemon (`bin/fm-supervise-daemon.sh`) owns supervision and runs the watcher one-shot: the watcher exits on every wake and the daemon starts its replacement, so a turn boundary regularly lands in a hand-off where no watcher process holds the lock and nothing is wrong. -The turn-end guard therefore accepts `fm_afk_daemon_owns_supervision` from `bin/fm-wake-lib.sh` as proof of supervision on that path: away mode must be active, and this home's `state/.supervise-daemon.lock` must name a live pid whose current process identity still matches the identity the daemon recorded for itself. +While `state/.afk` exists the daemon (`bin/fm-supervise-daemon.sh`) owns supervision and runs the watcher one-shot, in either away or quiet mode: the watcher exits on every wake and the daemon starts its replacement, so a turn boundary regularly lands in a hand-off where no watcher process holds the lock and nothing is wrong. +The turn-end guard therefore accepts `fm_afk_daemon_owns_supervision` from `bin/fm-wake-lib.sh` as proof of supervision on that path: `state/.afk` must exist (the predicate does not distinguish away from quiet mode), and this home's `state/.supervise-daemon.lock` must name a live pid whose current process identity still matches the identity the daemon recorded for itself. That is the same identity discipline the watcher lock uses, so a recycled pid, a lock left behind by a killed daemon, and a daemon that never recorded its identity all fail it. -A daemon that cannot record its own identity at startup logs a warning and keeps running, because a supervisor must not refuse to run over an unreadable `ps`; that warning is what names the cause when the guard then keeps blocking away-mode turn boundaries for the rest of that daemon's life. +A daemon that cannot record its own identity at startup logs a warning and keeps running, because a supervisor must not refuse to run over an unreadable `ps`; that warning is what names the cause when the guard then keeps blocking away/quiet-mode turn boundaries for the rest of that daemon's life. The proof covers ownership only, never freshness: the guard still requires a fresh beacon, so a daemon that stops restarting its watcher still blocks once the beacon passes grace, and a home with no daemon and no watcher blocks exactly as it did before. That beacon check uses the poll-derived grace described below rather than the flat `FM_GUARD_GRACE` default, because the daemon starts a fresh one-shot watcher only after it finishes handling the previous wake, and that handling can legitimately outrun a fixed 300-second window under load (a slow registered check, a busy supervisor pane) with the daemon perfectly healthy throughout. -With away mode off the daemon lock proves nothing and the strict watcher predicate is unchanged. +With `state/.afk` absent the daemon lock proves nothing and the strict watcher predicate is unchanged. `FM_STATE_OVERRIDE` wins over `FM_HOME/state`, and `FM_HOME` wins over repository-root `state/`. `FM_GUARD_GRACE` controls beacon freshness and defaults to 300 seconds. @@ -67,7 +67,7 @@ A fixed 300-second grace default stops correctly bounding staleness once a home' That hook and `bin/fm-watch.sh`'s own pre-acquisition staleness check (the "lock held by live pid but heartbeat is stale" refusal) both derive their default grace from the configured poll instead of a bare constant: `max(300, FM_POLL + 60)`, so the default never drops below the historical 300-second floor for the common short-poll case but grows with the poll cadence once that cadence would otherwise outrun it. `fm_poll_derived_grace` in `bin/fm-wake-lib.sh` is the single owner of that formula. The auto-arm hook additionally exports its resolved `FM_GUARD_GRACE` when it forks `bin/fm-watch-arm.sh`, so the arm wrapper and the watcher it may start judge staleness with the exact same value the hook just judged it with, whether that value came from an operator override or the poll-derived default. -`bin/fm-turnend-guard.sh`'s away-mode branch (`fm_afk_daemon_owns_supervision`, above) also derives its beacon grace from `fm_poll_derived_grace` rather than falling back to the bare 300-second default, for the same reason: the daemon's watcher-restart cadence there is not a fixed poll loop, so a flat grace misreads a daemon that is genuinely still cycling as down. +`bin/fm-turnend-guard.sh`'s daemon-ownership branch (`fm_afk_daemon_owns_supervision`, above, covering both away and quiet mode) also derives its beacon grace from `fm_poll_derived_grace` rather than falling back to the bare 300-second default, for the same reason: the daemon's watcher-restart cadence there is not a fixed poll loop, so a flat grace misreads a daemon that is genuinely still cycling as down. Every other direct `FM_GUARD_GRACE` reader (`bin/fm-guard.sh`, the strict-watcher checks in `bin/fm-turnend-guard.sh` and its harness-specific wrappers, `bin/fm-wake-lib.sh`) still falls back to the bare 300-second default unless `FM_GUARD_GRACE` is set explicitly in the environment. ## Harness integrations @@ -112,7 +112,9 @@ The first fresh exhausted-failure epoch preserves its handoff without consuming When none of those proofs appears, it re-blocks up to `FM_CLAUDE_TURNEND_BLOCK_BUDGET` times (default 3, below Claude's 8-block override). In Claude mode, positive watcher recovery clears the block budget, failure notice, and attended alarm together under the existing budget lock before either hook reports ordinary recovery. The one loud attended fail-open is available only when the auto-arm has recorded an exhausted failure, its one notice is already consumed, the block budget is exhausted, and a final check finds neither a healthy watcher nor an automatic continuation. -Each epoch identity is accounted at most once under the budget lock. +Each epoch identity is charged at most once per Stop under the budget lock, and a re-block against an epoch the auto-arm did not advance past the previous re-block is charged as well. +That second rule is what bounds an inert auto-arm: a hook kept silent by a session lock held by a live harness outside its ancestry, a hook that never fires, or a hook failing before its generation claim leaves the ledger frozen at its last outcome. +Charging only epoch changes let the count freeze with that ledger, so the guard re-blocked without limit and the attended fail-open was never reachable; `budget_account_current_epoch` in `bin/fm-turnend-guard.sh` owns the rule. Whenever both coordination locks are needed, positive auto-arm recovery and the terminal check acquire the auto-arm owner lock before the budget lock. After that alarm, the Stop auto-arm suppresses further exit-2 continuations until positive watcher recovery, so the final fail-open remains reachable. The alarm cannot repeat during that failure episode, and a later unhealthy stop blocks again. @@ -184,7 +186,7 @@ That warning uses `bin/fm-supervision-instructions.sh --repair-line`, so it alwa ## Regression coverage -`tests/fm-turnend-guard.test.sh` covers the predicate, main and secondmate primary scope, child-worktree exclusion, `FM_HOME` and `FM_STATE_OVERRIDE` precedence, the live-lock and fresh-beacon guard predicate, the cooperative `--claude` open-generation claim wait, monotonic failed-epoch progression, bounded attended fail-open, post-alarm continuation suppression, positive recovery reset, generation and legacy claim cases that must block or clear instead of allowing a blind stop, away-mode daemon ownership between watcher cycles and over a watcher lock left behind by an exited watcher, plus its dead, pid-reused, absent, stale-beacon, and away-mode-off negatives, the away-mode beacon's poll-derived grace widening for a live daemon still mid-cycle and its bound against a dead daemon, a beacon older than that wider grace, and FM_POLL's inapplicability with away mode off, Pi logical-run latching, missing-`jq` behavior, all five primary registrations, Grok native and legacy selection, typed field precedence, malformed input, and exactly-one-path safety. +`tests/fm-turnend-guard.test.sh` covers the predicate, main and secondmate primary scope, child-worktree exclusion, `FM_HOME` and `FM_STATE_OVERRIDE` precedence, the live-lock and fresh-beacon guard predicate, the cooperative `--claude` open-generation claim wait, monotonic failed-epoch progression, bounded attended fail-open, the same bound against a ledger frozen by an inert auto-arm with and without a verified failure episode, post-alarm continuation suppression, positive recovery reset, generation and legacy claim cases that must block or clear instead of allowing a blind stop, away-mode daemon ownership between watcher cycles and over a watcher lock left behind by an exited watcher, plus its dead, pid-reused, absent, stale-beacon, and away-mode-off negatives, the away-mode beacon's poll-derived grace widening for a live daemon still mid-cycle and its bound against a dead daemon, a beacon older than that wider grace, and FM_POLL's inapplicability with away mode off, Pi logical-run latching, missing-`jq` behavior, all five primary registrations, Grok native and legacy selection, typed field precedence, malformed input, and exactly-one-path safety. `tests/fm-guard-stale-banner.test.sh` covers the pull-guard predicate, including the persistent-model fresh-leftover-beacon negative control; the auto-arm model's healthy fresh-beacon-without-a-watcher case, session-and-recovery-bound long-turn rewake tolerance, independently broken tolerance signals, open-claim negative control, stale-beacon alarm, and isolation from other models; and the extension model's live-watcher path, ownership-qualified fresh hand-off, held-lock failures, independently broken ownership signals, stale-beacon alarm, queued-wake warning, and Pi and pi-signed harness routing. It also covers true-reason banner wording and reason-keyed episode dedup surviving a beacon mtime change. `tests/fm-cursor-primary.test.sh` covers the Cursor park end to end over real processes with no harness installed: each tracked Claude-shaped entrypoint standing down on a Cursor payload, both follow-up sources, the bounded repair nag and its reset, the nested loop bounds, supersession, away-mode and lock-ownership inertness, Pi-host stand-down without Cursor identity and continued parking when `PI_CODING_AGENT` leaks alongside `CURSOR_AGENT` or `CURSOR_INVOKED_AS`, child-worktree exclusion, and that the adapter never exits 2. diff --git a/docs/verification/agy.md b/docs/verification/agy.md new file mode 100644 index 00000000000..125012f4d81 --- /dev/null +++ b/docs/verification/agy.md @@ -0,0 +1,171 @@ +# Verification: the agy (Antigravity CLI) crewmate/scout adapter + +Active empirical facts for firstmate's agy adapter. +The skill tree rooted at [`.agents/skills/harness-adapters/SKILL.md`](../../.agents/skills/harness-adapters/SKILL.md) owns the operating facts through [`references/harness/agy.md`](../../.agents/skills/harness-adapters/references/harness/agy.md); this record owns how they were established and what is still unproven. + +## Subject + +| Field | Value | +|---|---| +| Version | `agy 1.2.0`; the send-confirmation timing below was re-measured on `agy 1.2.1` (2026-09-12) | +| Verified | 2026-09-10 | +| Binary | `/home/andpod/.local/bin/agy`, an ELF 64-bit Go-compiled single executable | +| Platform | Linux x64 (Arch, kernel 7.2.3) | +| Backend | Herdr, in an isolated non-`default` lab session (`fm-lab-firstmate-agy-ad-*` via `bin/fm-herdr-lab.sh`); the live `default` session was unchanged throughout | + +Every command below ran inside the disposable firstmate task worktree or the named Herdr lab session. +No captain fleet state was touched. + +## Detection: ancestry only, no marker + +``` +$ agy --version +1.2.0 +``` + +A live TUI's `/proc/<pid>/environ` carries no `AGY_*` or `ANTIGRAVITY_*` variable. +It does carry `AGENT=1` and `CLAUDECODE=1`, both inherited from the launching environment, so neither is an agy identity and neither is promoted to a marker. +Herdr's `pane process-info` for the same pane reports the foreground process as `name=agy` with `argv=["agy", ...]`, and `ps -o comm=` reports `agy`. +`bin/fm-harness.sh` therefore matches the anchored process name `agy` alone, and the spawn clears `CLAUDECODE`, `PI_CODING_AGENT`, `GROK_AGENT`, and `FM_PI_HARNESS` at the launch boundary. +`tests/fm-agy-harness.test.sh` pins the anchored match, the rejection of unrelated names containing the fragment, and that an inherited `CLAUDECODE` never outranks a real `agy` ancestor once the spawn clears it. + +## Launch: positional prompt-interactive with auto-submit + +``` +$ agy --prompt-interactive "Reply with exactly AGY_LIVE_PROBE_OK and nothing else" --model gemini-3.8-flash-low --effort low --dangerously-skip-permissions +``` + +The brief submitted itself with no extra Enter, the turn ran, and the reply rendered in the pane. +A second launch into the same directory answered a fresh prompt the same way, so the shape is repeatable, not a first-run accident. +The footer rendered `Gemini 3.8 Flash · low`, proving both flags were accepted together. + +## Trust dialog: pre-registered before launch, gated on a busy turn as the backstop + +A first launch in a fresh worktree shows this dialog: + +``` +Accessing workspace: + +/home/andpod/.treehouse/firstmate-7bab20/1/firstmate/agy-probe-tmp + +Do you trust the contents of this project? + +Antigravity CLI requires permission to read, edit, and execute files here. + +> Yes, I trust this folder + No, exit +``` +`agy --help` (1.2.0) lists no trust flag or pre-registration command, but agy honours a `trustedWorkspaces` entry written to `~/.gemini/antigravity-cli/settings.json` ahead of launch. +Verified under a throwaway `HOME` holding a copy of `~/.gemini` (the real settings file was never written): a folder appended to that array by hand launched `--prompt-interactive` straight into its turn and rendered the reply with no dialog, while an unregistered sibling folder launched the same way parked on the dialog. +agy compares the pane's logical working directory, not its resolved path: a symlinked cwd whose real path alone was registered still parked on the dialog, so `bin/fm-agy-trust.sh` records both the logical path and its resolved form when they differ. +`bin/fm-agy-trust.sh` retains that linked-worktree scope check and delegates writes to the shared `bin/fm-agy-trust-lib.sh` owner. +The fork spawn uses the library after its own worktree validation, preserves created-versus-preexisting custody for rollback and teardown, and refuses a failed trust registration before launch. +Two supervised Herdr runs in treehouse worktrees completed file-writing turns while the dialog was still unanswered at observation time (worker file and `done:` status line both verified on disk before Enter was ever sent to those panes). +Isolated runs in untrusted `/tmp` directories never reached the workspace until Enter: the turn spun through exploratory tool calls in agy's own scratch directory instead, and only the queued prompt ran after the answer. +One run left unanswered for several minutes wrote its file to agy's scratch directory instead of the workspace once finally answered. +The mechanism behind the difference was not established; path, backend, and latency were all varied across runs without isolating a single cause. +The spawn therefore does not depend on it: after pre-registration, `bin/fm-spawn.sh` runs a post-launch readiness gate (`agy_wait_for_working`) in the rovo/kimi launch-then-confirm shape as the backstop. +It polls the pane capture, answers the dialog with a single Enter the first time the `Do you trust the contents of this project?` text renders, and reports success only once `fm_busy_classify` returns a busy verdict for the pane (Herdr's native `working` status). +Because Herdr's native `working` verdict is known to coexist with an unanswered dialog, the gate is strict about order: a busy verdict counts as ready only when the worktree was pre-registered before launch or the dialog has already been seen and answered; on an unregistered path it keeps polling for the dialog instead of accepting the early busy verdict. +When the brief cannot be confirmed to run within the window (an answered dialog never turns busy, a pre-trusted pane never turns busy, or an unregistered pane never shows the dialog), the spawn fails, records `failed:` in the task status, and closes the endpoint so no orphan worker survives outside task control. +`tests/fm-agy-harness.test.sh` covers the helper's registration and scope refusals against a throwaway store, and drives a fake pane whose dialog decision reads the store the spawn just wrote: the pre-trusted launch with no dialog, a dialog that renders anyway answered exactly once, failed trust registration refusing before launch, and an unanswered or non-working launch failing and closing its endpoint. + +## Model and effort + +``` +$ agy models +Fetching available models... +gemini-3.8-flash-high Gemini 3.8 Flash (High) +gemini-3.8-flash-medium Gemini 3.8 Flash (Medium) +gemini-3.8-flash-low Gemini 3.8 Flash (Low) +... +``` + +`agy --help` documents `--effort` as `low|medium|high` and `--model` as the model for the session. +The bare `gemini-3.8-flash` id from this home's previous config is not listed; only the suffixed `-high`, `-medium`, and `-low` variants are. +`bin/fm-spawn.sh`'s `agy_model_validate` refuses a requested id a reachable `agy models` listing omits, and launches unvalidated with a stderr notice when the listing is unreachable. +The listing is a remote fetch (`Fetching available models...`), so the probe runs with stdin detached under the shared hard bound from `bin/fm-timeout-lib.sh` (15 seconds by default, `FM_AGY_MODELS_TIMEOUT`; a non-positive or non-numeric value clamps back to that default, because a non-positive bound is not a bound); a stalled fetch or a sign-in prompt is cut off and falls through to the unvalidated launch instead of blocking the spawn before any pane exists. +Print mode (`agy -p "Reply with exactly: AGY_PRINT_PROBE_OK" --model gemini-3.8-flash-low`) returned the exact reply with exit 0 in about 8 seconds, proving the credential path without a pane. + +## Busy state: the pinned status row, unknown on absence + +Mid-turn the pane rendered the status row and a spinner line at once: + +``` +⣯ Generating... +└ Tip: When reviewing a file edit, press f to see the full diff. +... +esc to cancel Gemini 3.8 Flash · low +``` + +The completed turn showed the reply, then the idle composer: + +``` +> +────────────────────────────────────────────────────────────────────────────── +? for shortcuts Gemini 3.8 Flash · low +``` + +`fm_busy_agy_tail_busy` and the delivery guard in `bin/fm-composer-lib.sh` match the `esc to cancel` token alone: the TUI pins that status row to the bottom of the pane for the whole turn, and the idle row replaces it with `? for shortcuts`. +The `Generating...` spinner word is deliberately not a signal: it is a free-floating output line, so ordinary worker output such as `Generating report...` would otherwise classify an idle worker as busy or acknowledge a submit that did not land. +No busy phase without the status row was observed live; every captured mid-turn frame carried it. +`fm_busy_classify` reports `unknown agy-regex` when the token is absent, because a long turn can scroll the marker out of the captured tail. +The signature is hardcoded with no environment override, so a stray variable can never change worker-state classification. +Herdr's own registry agreed throughout: `agent get` reported `agent_status=working` mid-turn and `idle` after, so on Herdr the native verdict carries busy with no new code. + +## Interrupt and exit + +A single `Escape` sent mid-turn through `herdr pane send-keys` cancelled it and printed this row, with the composer back at idle and no repolluted text: + +``` + ⎿ Interrupted · What should Antigravity CLI do instead? +``` + +Sending `/quit` plus Enter exited the process; the pane closed under the `exec` launch, and Herdr reported the pane gone. +`bin/fm-control-lib.sh` records `Escape` once, no clear key, no ack source, and `/quit` for agy. + +## Backend liveness: Herdr recognizes agy, tmux names it + +``` +$ herdr agent get w2:p1 --session fm-lab-firstmate-agy-ad-1599574-8823 +{"result":{"agent":{"agent":"agy","agent_status":"idle",...,"agent_session":{"agent":"agy","kind":"id","source":"herdr:antigravity_cli",...}}}} +``` + +Herdr tracks agy natively (`antigravity-cli` integration, detected as `agent=agy`). +The fork now also verifies process-level liveness and refuses an unknown startup status; [runtime verification](runtime-backends.md#antigravity-cli-agy-control-mechanics) owns the 2026-09-14 startup measurements and three complete native lifecycle trials. +The tmux adapter classifies the anchored process name `agy` as `agent` through the shared name vocabulary in `bin/fm-agent-process-lib.sh`, the muse/omp precedent for short bare-word names. +agy stays out of the session-lock name vocabulary in `bin/fm-session-lock-lib.sh`, where the other crewmate-only adapters are also absent. + +## Composer: unknown by design + +Byte-level capture of the idle pane shows a bare unstyled `>` between two full-width `─` rules, with an unstyled `? for shortcuts` cell and a dim (`SGR 2`) model cell in the status row below. +The shared classifier reads that bare `>` as `unknown` under the dead-shell rule, never `empty`. +Steering still confirms delivery: the Herdr submit core leads with the native `idle`-to-`working` transition, which agy performs, and the delivery footer regex covers the tmux path. +agy renders the busy footer late for that confirm loop - about 1.5 s after Enter for a short steer and 4-5 s for a realistic longer brief, measured live on `agy 1.2.1` (2026-09-12) against the shared budget's 3 x 0.4 s - so `bin/fm-send.sh` gives agy typed targets a longer default submit-confirm budget (20 retries, about 8 s at the default cadence); an explicit `FM_SEND_RETRIES` still wins and every other harness keeps the shared 3-retry default. +`tests/fm-send-agy-confirm.test.sh` pins the raised default and `tests/fm-agy-harness.test.sh` pins the Herdr transition path. +This is the cursor precedent, not a gap to patch in shared code. + +## Supervised task: spawn, steer, relaunch, and exit through the new path + +A trivial scout ran end to end through `bin/fm-spawn.sh --harness agy` against the same isolated lab session: `spawned agy-e2e1 harness=agy kind=scout` with a treehouse-provisioned worktree, `--model gemini-3.8-flash-low`, and `--effort low` all recorded in task metadata. +The worker wrote its worktree file and appended `done: agy e2e turn complete` to its status file, which lives outside the worktree, proving prompt processing, tool execution, outside-workspace file access, and a new completion event. +Durable steering held: a `bin/fm-send.sh` message landed in the task inbox, the worker appended the steered lines to both files, and its inbox record moved to `handled/`. +Same-copy relaunch held: `bin/fm-control.sh relaunch --note` replaced the worker in place on the identical worktree, model, and effort, the replacement verified both prior lines intact and appended `relaunched: done`. +Exit held: `bin/fm-control.sh exit` stopped the worker, the registry returned `agent_not_found`, and the pane remained a lone shell in the worktree with all work intact. +No automatic quota failover was exercised or claimed; every handoff above was an explicit supervised relaunch. + +## What is still unproven + +The unauthenticated failure mode was never observed; this host's agy runs signed in, so any auth prompt is a fail-loud credential blocker, not a handled dialog. +No slash-skill invocation form was verified, so skill invocation stays natural language. +`--continue` and `--conversation` resume were never exercised; recovery uses deterministic relaunch from the brief on disk. +No primary or secondmate behavior was built or tested, and none is claimed. + +## Refreshing this record + +Run the portable suite and the live guard after any agy upgrade, because process identity, trust-dialog text, native lifecycle state, and interrupt behavior depend on vendor-controlled surfaces: + +``` +bin/fm-test-run.sh tests/fm-agy-harness.test.sh +FM_AGY_SIGNALS_LIVE=1 bin/fm-test-run.sh tests/fm-agy-signals-live-e2e.test.sh +``` diff --git a/docs/verification/gbrain-readonly-share.md b/docs/verification/gbrain-readonly-share.md index ae84c12a0db..a650671cc84 100644 --- a/docs/verification/gbrain-readonly-share.md +++ b/docs/verification/gbrain-readonly-share.md @@ -11,6 +11,21 @@ $ gbrain version gbrain 0.42.69.0 ``` +## Read-only scope refresh on 2026-09-14 + +The complete live share regression passed against GBrain `0.46.21.0` at `649ffe5f8baf3ff7f979c77f4de3975904cfe029` on Linux x86_64. +It used two disposable brains, a loopback HTTP server, a local embedding endpoint, and an empty runtime home for child processes, with provider-key environment variables removed before registration and serving. +The runtime-home isolation covers file-backed provider credentials as well as the separate `GBRAIN_HOME` data roots. + +```sh +FM_GBRAIN_LIVE_E2E=1 bin/fm-test-run.sh tests/fm-gbrain-readonly-e2e.test.sh +``` + +The test verified world-only context packs, independent OAuth-client delta cursors, successful main-brain reads, refusal of attempted writes, and byte-for-byte preservation of the main brain. +Read-scoped `think` remained reachable but degraded without credentials and could not persist a result. +The secondmate could write its own brain and keep searching it when the main brain went offline, and generated artifacts and server logs did not contain client secrets. +The complete script passed in 18.465 seconds. + ## The installed mount model cannot express a read-only share A mount in this version carries direct transport only, so it cannot be given an OAuth credential: diff --git a/docs/verification/muse.md b/docs/verification/muse.md index 38d12653a6a..ba0d52c234b 100644 --- a/docs/verification/muse.md +++ b/docs/verification/muse.md @@ -48,7 +48,8 @@ $ grep -nE 'muse-bin|exec ' launcher.sh `ps -o comm= -p <pid>` returns the full executable path, whose basename is `muse-bin-<version>`. That is why both `bin/fm-harness.sh` and `bin/backends/tmux.sh` match the anchored prefix `muse-bin-*` rather than an exact name, and why neither can rely on an install-path component: `~/.local/bin/muse-bin-<version>` contains no `muse` path component. -The Muse launch clears `CLAUDECODE`, `PI_CODING_AGENT`, `GROK_AGENT`, `FM_PI_HARNESS`, `CURSOR_AGENT`, and `CURSOR_INVOKED_AS` before the worker starts so foreign primary markers cannot override the versioned ancestry. +The Muse launch clears `CLAUDECODE`, `PI_CODING_AGENT`, `GROK_AGENT`, `FM_PI_HARNESS`, `CURSOR_AGENT`, and `CURSOR_INVOKED_AS` before the worker starts, which is the verified launch behavior rather than what detection depends on. +[Harness detection precedence](runtime-backends.md#harness-detection-precedence) owns why a retained foreign marker cannot override the versioned ancestry. [`runtime-backends.md`](runtime-backends.md#agent-liveness-name-sources) owns the resulting tmux liveness verdict and its relationship to the portable decoy regression. diff --git a/docs/verification/process-event-sources.md b/docs/verification/process-event-sources.md index a2e52f9b5fe..279a8b19bba 100644 --- a/docs/verification/process-event-sources.md +++ b/docs/verification/process-event-sources.md @@ -86,7 +86,7 @@ Never at-least-once, no-loss, or lossless. ## What the runner does prove -Exercised by `tests/fm-procevent.test.sh` against a fake blocking source whose completion is a process event, not a timer; for the two supervision-delivery rows below, by `tests/fm-watch-triage.test.sh` driving a real `bin/fm-watch.sh` over a real capture; and for adapter-owned application, by `tests/fm-remote-reply.test.sh` driving the real remote-reply relay end to end in an isolated home: +Exercised by `tests/fm-procevent.test.sh` against a fake blocking source whose completion is a process event, not a timer; for the supervision-delivery and headline rows below, by `tests/fm-watch-triage.test.sh` driving a real `bin/fm-watch.sh` over a real capture and over queued strand and launch-failure keys, with `tests/fm-watch-arm.test.sh` covering the arm-time refusal; and for adapter-owned application, by `tests/fm-remote-reply.test.sh` driving the real remote-reply relay end to end in an isolated home: | Guarantee | How it is proven | | --- | --- | @@ -119,9 +119,13 @@ Exercised by `tests/fm-procevent.test.sh` against a fake blocking source whose c | attached owner continuity | a foreground `start` with a one-second lease remains alive beyond that lease while its caller stays attached, then captures normally when the blocking source completes | | owner-home lifetime and scope | a detached runner and its spawning descendant are observed reparented before an expired owner lease stops their whole process group and process churn; replacing the state directory at the same path cannot keep the old runner alive with a new lease because its recorded device/inode no longer matches, while an identical runner in an unchanged home whose reconcile cycle keeps its lease fresh remains alive | | launch pacing during owner-loss grace | an immediately returning source that attempts detached self-relaunches is held to the configured minimum interval between command launches and remains bounded until its expired owner lease stops the generation; replacement starts a fresh pacing generation, prunes prior pacing state, and prevents a superseded sleeping runner from recreating it | -| stale reclaim without displacement | concurrent contenders replacing one stale claim start exactly one runner, cross-home replacement removes the old generation's staging file from its recorded state directory, and a generation whose stale owner and independently empty process group prove it gone remains reclaimable when its recorded state-root identity can no longer be revalidated | -| crashed leader with a live group | `SIGKILL` on only the runner leader leaves its blocking child group alive; reconcile treats that leaderless group as ambiguous, preserves its claim without starting a replacement, and still reclaims a generation with no leader and no surviving group | -| PID-reuse safety | retirement refuses a live PID whose identity differs from the claim before signalling, and a surviving process group prevents stale-generation cleanup on both ordinary and failed reservation-removal paths | +| stale reclaim without displacement | concurrent contenders replacing one stale claim start exactly one runner, cross-home replacement removes the old generation's staging file from its recorded state directory, and a generation whose stale owner and independently empty process group prove it gone remains reclaimable when its recorded state-root identity can no longer be revalidated or its recorded registry directory no longer resolves to a directory, so `reconcile` reclaims it once, the replacement runs the source, and later cycles report nothing to do | +| confirmed launches only | `reconcile` counts a launch as `started` only after the source is observed owned or its launch-pacing stamp has moved: a registration that cannot start is reported `failed=` with a non-zero exit and its source still listed `none`, a source that claimed, ran and exited before confirmation looked is still `started`, a zero-padded confirm window reads as base 10, and an unusable `FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS` is refused by name before any runner is launched | +| launch failure announced once per episode | an unconfirmed launch queues one `check` wake keyed by source, registration identity and an episode nonce; a second failure in the same episode queues nothing, a confirmed launch queues no failure and closes the episode, a later failure opens a new episode under a fresh key, and a 64-character source id keeps that key within the watcher's marker bound | +| crashed leader with a live group | `SIGKILL` on only the runner leader leaves its blocking child group alive; reconcile treats that leaderless group as ambiguous, preserves its claim without starting or signalling anything, `start` runs nothing beside it, the strand is queued as one `check` wake keyed by source and claim token that a second cycle does not repeat, and reconcile still reclaims a generation with no leader and no surviving group | +| reused pid with a live group | a stale claim whose recorded pid is alive under a different identity while its process group still has members is listed `orphaned`, is never relaunched by `reconcile` across cycles, is announced once naming the `start` command that clears it, and `start` reclaims it while the dead generation's leftovers can be tidied and refuses with `cannot claim source`, replacing nothing, when they cannot | +| strand and failure headlines | a real `bin/fm-watch.sh` surfaces queued `stranded` and `launch-failed` keys under `process-event source stranded` and `process-event source failed to start` rather than as a captured result, joins a mixed cycle's headlines, never re-delivers a key it has already surfaced, and delivers each new failure episode; `bin/fm-watch-arm.sh` refuses to arm on an unusable confirm window, naming the variable and range, with no beacon and no running watcher | +| PID-reuse safety | retirement refuses a live PID whose identity differs from the claim before signalling, and a surviving process group keeps `reconcile` and `retire` from cleaning up the stale generation on both ordinary and failed reservation-removal paths; the reused-pid row above owns what a deliberate `start` does there | | coherent ownership reads | a claim replacement held inside the source boundary blocks `list` until one complete generation is visible | | retire-start exclusion | a queued start revalidates registration after the serialized retirement boundary and executes no child | | uncertain identity before the first signal | a live owner whose identity probe transiently fails is not signaled or released, and its registration remains for retry | @@ -139,6 +143,11 @@ Exercised by `tests/fm-procevent.test.sh` against a fake blocking source whose c | condition->action process bounds | the same suite proves action timeout terminates descendants and command-output staging remains within `FM_WHEN_OUTPUT_TAIL_BYTES` while the command runs | | silent failure handling | a nonzero exit with no output publishes nothing and leaves the source registered for retry | | inertness | a home with no registered source generates no state, starts no process, and does not need supervision | +| rebind-all refreshes an in-repo trust binding, leaves an out-of-repo one alone | after a simulated self-update rewrites an armed watch's in-repo action executable's bytes, `rebind-all` republishes exactly that watch's trust binding against the new bytes and leaves a watch whose action lives outside `FM_ROOT` untouched byte-for-byte; the rebound watch then fires cleanly against the new bytes instead of being refused, and a second `rebind-all` with nothing changed rebinds nothing (`tests/fm-procevent-when.test.sh`) | +| rebind-all matches FM_ROOT reached through a symlink | `FM_ROOT` and the resolved action executable are each canonicalized before the containment comparison, so a watch whose action is reached through a symlinked checkout path is still recognized as in-repo and rebound rather than silently skipped as out of scope (`tests/fm-procevent-when.test.sh`) | +| self-update rebinds a locally armed watch | `fm-update.sh` runs its own home's `rebind-all` best-effort immediately after a successful fast-forward of the primary repo or a local secondmate, so a watch armed against an in-repo action keeps firing across the update with no separate operator step (`tests/fm-update.test.sh`) | +| rebind-all reaches a watch already polling when the update lands | `run`'s poll loop calls `spec_load` once before entering its loop and would otherwise compare fire-time bytes against that stale in-memory hash forever; the fire-time check instead reloads the trust binding from disk immediately before firing, so a watch armed before a self-update still fires against the rebound bytes instead of being rejected as stale (`tests/fm-procevent-when.test.sh`) | +| fire-time reload serializes against rebind_one's publish | `publish_spec` renames the new spec into place and the new trust into place as two separate renames, never one atomic swap; the fire-time reload takes the same per-sid source lock `rebind_one` holds across that publish, so it can never observe the torn combination of rebound spec bytes next to a still-old trust record and instead waits for the publish to finish (`tests/fm-procevent-when.test.sh`) | | absent extension registry parity | `tests/fm-extension-binding.test.sh` drives `list` and `verify` in a fresh home while the current directory contains project files and Pi packages and an environment variable names fake package data; both commands report no bindings, create no home path, and discover nothing outside `config/extensions.d` | | complete package and binding identity | the same suite drives the public bind and verify commands through manifest duplicate/unknown/version failures, project and task-copy confinement, canonical path and symlink rejection, hard-link rejection, owner/mode checks, a non-executable entrypoint, binding mode drift, complete-tree mutation, exact executable mutation, and a missing executable; the foreign-owner fixture executes when the platform permits constructing another uid and otherwise reports that privilege limitation, while ordinary non-privileged CI does not exercise it or claim it ran | | external evidence write confinement | the same suite substitutes `state/procevent/` and `state/procevent-inbox/` with post-registration symlinks and proves an external start fails before bytes reach either outside target; it proves public lifecycle entry, environment, paths, and descriptors cannot forge capture authority; it proves live-generation claim release removes pending or consumed capture reservations only from the recorded revalidated state root, while a generation independently proved gone may leave an unreachable token-keyed reservation rather than wedging ownership; and it proves the absent-registry built-in capture path retains its legacy state-path behavior | diff --git a/docs/verification/rovo.md b/docs/verification/rovo.md index 8a588e3ad8f..2d6c722f1d6 100644 --- a/docs/verification/rovo.md +++ b/docs/verification/rovo.md @@ -184,7 +184,7 @@ $ ls "$LAB/outside/inbox/handled" ## Backend liveness: tmux verified live, herdr placement verified live with a herdr-side agent-detection gap tmux 3.6a is now installed and was exercised live in an isolated `tmux -L <private-socket>` session, so tmux pane liveness is fully verified rather than pending. -`bin/backends/tmux.sh`'s `fm_backend_tmux_classify_process_name` matches `*rovo*` alongside the other globbed harness names, so a rovo pane classifies `agent` (not `other`). +`bin/fm-agent-process-lib.sh`'s `fm_agent_process_classify_name` (then still inside `bin/backends/tmux.sh`) matches `*rovo*` alongside the other globbed harness names, so a rovo pane classifies `agent` (not `other`). The two independent name sources behaved as designed: `#{pane_current_command}` reported the truncated on-disk binary name `atlassian_cli_r` - macOS's 15-char `comm` truncation cuts `atlassian_cli_rovodev` off just before the `rovo` substring begins, the same truncation-volatility class [`runtime-backends.md`](runtime-backends.md) already documents for codex/kimi's own patch-release name drift - while the foreground ps-based `comm` correctly reported `rovo`, and `fm_backend_tmux_agent_state` correctly returned `alive` through that primary source. The two-independent-name-sources design is exactly why the truncation quirk does not break the verdict. `tmux capture-pane` correctly rendered the box composer and the `Rovo is thinking...` busy line while a real `sleep`-based bash tool call ran; `fm_busy_rovo_tail_busy` classified it busy, then idle once the tool call completed and the reply landed. The Escape/`Agent cancelled` evidence in the interrupt section above was captured in this same live tmux session. `/exit` closed the tmux window cleanly, and `fm_backend_tmux_agent_state` reported `missing` immediately afterward - a clean, unambiguous exit verdict. diff --git a/docs/verification/runtime-backends.md b/docs/verification/runtime-backends.md index 71b2af21e2f..548886b4908 100644 --- a/docs/verification/runtime-backends.md +++ b/docs/verification/runtime-backends.md @@ -130,6 +130,111 @@ all fm-subagent-pretool-check tests passed all fm-documentation-audiences tests passed all fm-pointer-check tests passed ``` +## Harness detection precedence + +Firstmate's own harness comes from two kinds of evidence, and `bin/fm-harness.sh` owns how they combine: an environment marker names its harness, and the nearest harness process in the parent chain proves who owns the process tree. +A marker alone is not proof of ownership, because it is ordinary environment state that a child inherits and a terminal multiplexer can replay into an unrelated session. +Verified on 2026-09-02 on Linux 7.1.12 with the portable regression, which builds every case from real renamed processes and no installed harness: + +```sh +bin/fm-test-run.sh tests/fm-harness-precedence.test.sh +``` + +Observed output: + +```text +ok - a markerless harness keeps its identity under an inherited foreign marker +ok - a harness that publishes a marker inside its own process tree is unchanged +ok - with ancestry silent, the marker layer and its cursor-first ordering still decide +ok - a retained cursor marker does not rename a nested claude worker +ok - an agreeing marker keeps Pi's finer identity that ancestry cannot prove +ok - an interpreter script-path match answers alone but never outranks a marker +ok - a native harness binary under an interpreter shim decides at comm strength +ok - a harness that is pid 1 of its own namespace is examined, not skipped +ok - the descent probe reaches comm strength where the top-of-session probe sees only args +ok - the descent probe reports no verdict from a sibling branch detection cannot reach +ok - a foreign args-only verdict at the deepest vantage leaves the comm-strength identity intact +ok - equal-depth descent ties prefer the comm-strength leaf regardless of spawn order +ok - session start renders the Codex protocol for a Codex primary holding a retained CLAUDECODE +FM_TEST_SUMMARY total=1 failed=0 skipped_gate=0 duration_ms=3666 +``` + +Before that boundary existed, a Codex session started from an environment that had retained `CLAUDECODE=1` reported `claude`, and session start emitted Claude's Stop-owned supervision protocol to a Codex primary. +The same live shape, reproduced with a real process named `codex` and no installed harness, now reports `codex` with the marker present and `claude` with the marker present and ancestry blinded, which is what proves the case is not vacuous. + +### A real Codex session holding a retained Claude marker + +The portable regression builds its process tree from renamed executables, so the same guarantee is proven again against the real installed Codex. +`codex sandbox` runs a command under the installed native binary with no model turn, inside a PID namespace where that binary is pid 1 and the command is pid 2. +Verified on 2026-09-01 with codex-cli 0.152.0 on Linux 7.1.10, with both Claude markers retained in the launching environment: + +```sh +CLAUDECODE=1 CLAUDE_CODE_ENTRYPOINT=cli codex sandbox bash -c \ + 'cd <checkout> && bin/fm-harness.sh; bin/fm-harness.sh ancestry; bin/fm-supervision-instructions.sh' +``` + +Against the parent commit, with `CLAUDECODE=1` and `CLAUDE_CODE_ENTRYPOINT=cli` confirmed present in the probe's own environment and the chain reading pid 2 `bash` to pid 1 `codex`: + +```text +verdict=claude +SUPERVISION OPERATING INSTRUCTIONS - primary harness: claude +Mode: Claude Stop-hook-owned supervision. +``` + +With the current boundaries in place, from the same command and the same process chain: + +```text +verdict=codex +ancestry=comm codex +SUPERVISION OPERATING INSTRUCTIONS - primary harness: codex +Mode: Codex foreground checkpoint. +``` + +Two boundaries are load-bearing here, and the marker-versus-ancestry precedence above is only the first. +The walk also used to stop as soon as the next pid was 1, on the assumption that pid 1 is always init. +That assumption inverts inside a PID namespace, where the harness is pid 1: the walk returned no ancestry at all, so the retained marker won by default even with precedence corrected. +The walk now examines that top process before stopping, which costs one `ps` call and can introduce no false positive, because a host's real pid 1 (init, systemd, launchd) matches no harness name. +The portable regression asserts both directions of that case: a host-shaped pid 1 still leaves the marker to answer, and a harness at pid 1 outranks it. + +Run on the host under Claude Code 2.1.252 with the same two markers set, the same probe reports `claude`, `comm claude`, and Claude's Stop-owned protocol, so the correction does not trade one misidentification for its inverse. + +### Real harness process names behind the walk + +The detection half of the opt-in drift guard asks the ancestry walk what it makes of each INSTALLED harness's real running process: + +```sh +FM_HARNESS_LIVENESS_DRIFT=1 bin/fm-test-run.sh tests/fm-harness-liveness-drift-live-e2e.test.sh +``` + +The guard probes the upward path between the deepest foreground descendant of the pane process and the pane process itself, and reports each distinct verdict that vantage set produces. +Refreshed on 2026-09-14 with the command above: Claude 2.1.270, Codex CLI 0.154.0, OpenCode 1.18.30, Pi 0.85.1, and Cursor 2026.09.10-fd3934a each reported alive and comm-strength ancestry identity. +The guard printed `checked 5 installed harness(es)` and its complete-suite success marker. +Pi-signed, Grok, Kimi, and Muse were explicitly absent; Gemini, Rovo, and OMP were also absent from this host. +Agy 1.2.2 is covered by its separate real lifecycle and vendor guards below. + +Observed on 2026-09-02 for the harnesses installed on that machine: + +```text +# claude 2.1.258 (Claude Code): title='claude' foreground=[claude ] +# claude 2.1.258 (Claude Code): ancestry verdicts=[comm claude] +# codex codex-cli 0.152.0: title='node' foreground=[node codex ] +# codex codex-cli 0.152.0: ancestry verdicts=[comm codex;args codex] +``` + +The verdicts are reported deepest first, so Codex's native child answers before the shim above it. +When eligible foreground descendants tie at the greatest depth, the probe prefers a leaf whose own verdict reaches comm strength; if none does, it keeps the first leaf, so process-table ordering cannot hide an equally deep native harness binary behind an args-strength interpreter. + +Codex ships as a `node` npm shim that spawns its native `codex` binary as a foreground child, which is why its two verdicts differ: the pane process is identified only from the shim's script path, and the native child is what carries the process name. +That difference is the reason the guard cannot probe the pane process alone. +The guarantee this guard holds is a strength claim, not only an identity one, because `detect_own` hands an args-strength verdict straight back to a retained foreign marker. +A pane-only probe would have observed `args codex`, passed, and gone on passing if a later release stopped spawning the native child, while real sessions silently regressed to the original bug. +Probing from below asks the question from the vantage a tool subprocess actually occupies, so the guard can require comm strength somewhere in the session and require every comm-strength vantage to name the same harness. +The vantage set stops at the upward path rather than the whole subtree, because `harness_ancestry` only ever climbs and a sibling branch is therefore a vantage firstmate's own detection can never occupy. +The reject-other-harness cross-check judges comm-strength vantages only, because an args-strength verdict is path-ambiguous by construction: a harness-spawned MCP server running as `node <home>/.claude/mcp/<server>.js` answers `args claude` purely from the `.claude` path component, and such a server is normally a child of the agent binary, so it can be the deepest descendant and sit on this path. +That narrowing changes only which vantages the cross-check judges; the comm-strength requirement itself is unchanged. +A single-process harness has no descendant that adds a distinct verdict, which is why `claude` reports one. +The portable regression pins every half without any harness installed: `tests/fm-harness-precedence.test.sh` asserts that this two-process topology decides at comm strength, that the descent probe reaches a strength the top-of-session probe cannot, that a sibling branch answering a foreign harness contributes no verdict, that a foreign args-only verdict at the deepest vantage leaves the comm-strength identity intact, and that equal-depth ties choose the comm-strength leaf regardless of process ordering. +The run did not reach `opencode`, `pi`, `pi-signed`, `grok`, `kimi`, or `muse`, which were not installed, and stopped at the same pre-existing liveness failure for `cursor` 3.18.9, whose resolved binary on that machine is the editor rather than `cursor-agent`; those adapters are unverified by this run. ## tmux @@ -188,7 +293,7 @@ The crewmate-only Muse Code 0.1.0-R708.1 adapter was verified separately on 2026 Its installed `muse-bin-0.1.0-R708.1` foreground identity classified `alive`, while `musescore`, `amuse`, `muse-binary`, and `muse-bind` remained ambiguous in the portable regression. [`muse.md`](muse.md#process-identity) owns the artifact identity and launcher evidence for that verification. -The crewmate/scout-only Rovo CLI 202609.1.2 adapter added `*rovo*` to the same glob family as `*grok*`/`*kimi*` in `fm_backend_tmux_classify_process_name`, and was relaunched live under tmux 3.6a in an isolated private socket. +The crewmate/scout-only Rovo CLI 202609.1.2 adapter added `*rovo*` to the same glob family as `*grok*`/`*kimi*` in the shared process-name classifier (now `fm_agent_process_classify_name` in `bin/fm-agent-process-lib.sh`), and was relaunched live under tmux 3.6a in an isolated private socket. `#{pane_current_command}` reported the truncated on-disk binary name `atlassian_cli_r` - macOS's 15-char `comm` truncation cuts `atlassian_cli_rovodev` off just before the `rovo` substring begins, the same truncation-volatility class codex/kimi's own patch-release name drift shows above - while the foreground ps-based `comm` correctly reported `rovo`, so `fm_backend_tmux_agent_state` returned `alive` through that primary source; the two-independent-name-sources design is exactly why the truncated title does not break the verdict. [`rovo.md`](rovo.md#backend-liveness-tmux-verified-live-herdr-placement-verified-live-with-a-herdr-side-agent-detection-gap) owns the fuller record, including the busy/interrupt/exit facts captured in that same live tmux session and the herdr agent-detection gap found when herdr placement was verified live in an isolated lab session. @@ -427,7 +532,48 @@ That warning rendered in the same shape as the trust dialog, with the selection That gate is not a production blocker, because a normal environment has already accepted it and the treatment arm above ran against the real config and saw neither dialog. This change does not address that warning and does not claim to. -`bin/fm-spawn.sh` therefore pre-registers the task worktree through `bin/fm-claude-trust.sh` before launch, and `tests/fm-claude-trust.test.sh` pins both halves of the scope contract: a fresh worktree is trusted, and an out-of-scope path is refused. +### Secondmate homes + +Verified 2026-09-11 on Claude Code 2.1.269. +A secondmate launches in its own firstmate home rather than a task worktree, and that home meets the same gate. +The control arm launched a standalone-clone secondmate home that the store had no entry for, the way `bin/fm-spawn.sh --secondmate` launches one. + +```sh +tmux -L <sock> new-session -d -s ctrl -x 180 -y 44 -c <home> \ + "CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false claude --dangerously-skip-permissions" +``` + +``` + Accessing workspace: + /private/tmp/fm-sm-trust-live-69759/fm-homes/livemate-n1 + Quick safety check: Is this a project you created or one you trust? ... + ❯ No, exit + Yes, I trust this folder +``` + +The treatment arm pre-registered that same home through the secondmate-home mode and launched it identically against the operator's real config. + +```sh +bin/fm-claude-trust.sh --secondmate-home <home> livemate-n1 +``` + +``` +trusted: /private/tmp/fm-sm-trust-live-69759/fm-homes/livemate-n1 +``` + +``` + ▐▛███▛█ Claude Code v2.1.269 +▝▜██████▀ Opus 4.8 with high effort · Claude Max + ▝▝ ▝▝ /private/tmp/fm-sm-trust-live-69759/fm-homes/livemate-n1 +... +❯ + ⏵⏵ bypass permissions on (shift+tab to cycle) · ← for agents +``` + +No dialog appeared, the composer was reached, and neither did the machine-scoped bypass warning, because this ran against the real config. +The lab home was deleted and the test entry was removed from the store and verified absent, with the same point-in-time caveat as the worktree arms above. + +`bin/fm-spawn.sh` therefore pre-registers the directory every claude launch starts in through `bin/fm-claude-trust.sh` before launch, and `tests/fm-claude-trust.test.sh` pins both halves of the scope contract for both shapes: a fresh worktree and a seeded secondmate home are trusted, and an out-of-scope path is refused. That automated spawn case runs against a fake claude, so it asserts the store entry and the launch command and nothing more; the live arms above are what establish that the entry actually suppresses the dialog. The composer-classification record below observes the same gate from the other side, where an untrusted worktree left Claude, Grok, and Muse unverified because the guard reads a first-launch trust dialog as an unreadable composer. @@ -1024,6 +1170,41 @@ Part C is the case the suite could not reach before: a doomed pane whose shell h On 0.7.5 that fallback exposed a bounded four-sample wrong-focus window and restored the anchor exactly; on 0.8.0 the same fallback exposed none, which is why default-on projection is floored at 0.8.0 rather than mitigated further below it. The suite also cross-checks its own Part A measurement against the floor classifier on whatever release it runs, so a drifted protocol-to-release mapping fails there rather than silently gating on the wrong thing. +### Attached foreground viewer + +A pseudo-terminal registers as a Herdr foreground client only when its window grid is non-zero. +`script` and a bare `pty.fork()` from a non-tty parent both start at 0x0, which is why PR #4131 could validate only the detached half of the teardown focus guard and left its four attached-client scenarios untested. +The guarded `viewer start` path fixes the pty at the proven 40-row by 120-column grid, sets that size on the master fd before the fork, and scrubs inherited `HERDR_*` variables, which makes the attached scenarios reachable from a headless runner. + +Measured on 2026-09-11 against Herdr 0.9.0 protocol 22 on macOS 26.5.2 aarch64 with Python 3.14.6: + +```sh +HERDR_LAB_HELPER=bin/fm-herdr-lab.sh \ + tests/fm-herdr-attached-viewer-live-e2e.test.sh +``` + +```text +ok - attached viewer: a pty sized before the fork registers as a real Herdr foreground client +ok - attached viewer: a live client on the target tab refuses the close and keeps the pane +ok - attached viewer: focus moving onto the target between planning and mutation still blocks the close +ok - attached viewer: a close preserves the fresh non-target focus the viewer moved to +ok - attached viewer: the projection seeded-tab prune refuses while a live client watches it +ok - attached viewer: detaching releases the refusal, so the guard tracks the client and not the pointer +``` + +The same six checks passed on 2026-09-14 against Herdr 0.8.2 on Linux x86_64 with Python 3.12.3, using the guarded helper from `kunchenguid/firstmate@b182d0f908b78d08c7ccb8dce3775bdca8c5d657`. +The isolated-session cleanup completed with the live default-session tripwire intact. + +Both halves of the recipe are load-bearing, and each was measured by removing it from the helper and re-running the guard on the same host and release. +Dropping the `TIOCSWINSZ` call and dropping the environment scrub each left startup reporting `no_foreground_client`, followed by the guard failure: + +```text +not ok - could not attach a real foreground Herdr viewer over a sized pty +``` + +Re-run this guard after every Herdr upgrade. +A release that changed the foreground-client contract, the window-grid requirement, or the nested-viewer refusal would fail here first, and the detached regressions would keep passing while saying nothing about it. + ### Presentation version floor Default-on presentation projection is floored at Herdr 0.8.0. @@ -1152,10 +1333,12 @@ Herdr is one of the two backends whose recovery-grade agent-state classifier the tests/fm-control-herdr-smoke.test.sh ``` -Observed output: +Observed output, refreshed 2026-09-10 on Herdr 0.9.0 after the stale-registration fix (the two stale-registration lines are recorded under "Stale agent registration" below): ```text ok - real herdr: exit on a pane with no registered agent is idempotent success +ok - real herdr 0.9.0: a gone session reads recoverable while a live pane and a malformed target do not +ok - real herdr: a drifted agent-free shell returns to its worktree and reuses the same endpoint ok - real herdr: interrupt refuses when herdr's own agent registry reports no agent ok - real herdr: a registration with no agent process behind it classifies dead (stale), not alive ok - real herdr: interrupt refuses a stale registration instead of keying a dead shell @@ -1163,15 +1346,15 @@ ok - real herdr: exit on a stale registration is idempotent success, so relaunch ok - real herdr: a registered agent with a live process stays alive through the cross-check ok - real herdr: interrupt delivers the harness's key and proves the agent survived it ok - real herdr: no control verb removed the endpoint or the task's local copy +ok - real herdr 0.9.0: a registration Herdr keeps after its agent exits reads stale-agent and recovers as dead +ok - real herdr: exit on a pane with a stale registration is idempotent success +ok - real herdr: a stale registration no longer blocks relaunch, and the endpoint and local copy survive ok - real herdr: an agent that does not stop fails closed instead of being reported as stopped all fm-control-herdr-smoke tests passed ``` -The registry read through `herdr pane report-agent` is the same source `fm_backend_herdr_agent_state` classifies. -In this guard, the unregistered case confirms the ordinary agent-free baseline, `report-agent` creates the stale-registration state over a plain idle shell, and a real foreground process makes the paired live case fail the agent-free proof without launching a real agent. -A separate live reading on 2026-08-15 found the other dead-pane shape the classifier must accept: pane shell, resident `treehouse` wrapper, and worktree subshell holding the foreground. -The classifier read that endpoint `dead` while four concurrently live Pi and Claude panes on the same server all read `alive`; the portable `test_agent_state_*` cases in `tests/fm-backend-herdr.test.sh` pin the chain and negative shapes. +The registry read through `herdr pane report-agent` is the same source `fm_backend_herdr_agent_state` classifies, and since 2026-09-10 that registration counts as an agent only while `pane process-info` shows a harness process behind it, so the guard backs the registration with a real process named like a harness (a symlink to `sleep`) and then stops that process, with no real harness launched. That command is the guard that refreshes this record; run it after every Herdr upgrade rather than trusting the version above. For Pi on Herdr 0.9.0, `herdr agent get` reflects whether the agent process remains live; its registration does not persist merely because the pane and parent shell do. @@ -1238,6 +1421,90 @@ ok - real herdr: a drifted agent-free shell returns to its worktree and reuses t `tests/fm-control-relaunch.test.sh` drives a tmux stub and proves that tmux retains its prior refusal without sending `cd` or any other input to the pane. The Herdr refusal when a shell accepts the command but does not move is not exercised in this change. +### Stale agent registration + +Fork verification on 2026-09-14 with Herdr 0.8.2 and Pi 0.85.1 uses the shared positive classifier and the fork's terminal-wide absence proof. +The portable companion exercises suspended, backgrounded, reparented same-terminal, escape-shell, and unreadable-process cases. +Refresh both through the isolated named-lab helper: + +```sh +HERDR_LAB_HELPER=/home/sungin/firstmate/bin/fm-herdr-lab.sh bin/fm-test-run.sh tests/fm-backend-herdr.test.sh tests/fm-herdr-pi-stale-registration-live-e2e.test.sh +``` + +```text +ok - real herdr 0.8.2 + pi 0.85.1: a running registered pi classifies alive at process level +# herdr 0.8.2 kept the pi registration (idle) after /quit under a nested shell: the stale-registration branch is exercised +ok - real herdr 0.8.2 + pi 0.85.1: the registration left behind by a quit pi reads stale-agent and recovers as dead +FM_TEST_SUMMARY total=2 failed=0 skipped_gate=0 duration_ms=22692 +``` + + +Measured 2026-09-10 on macOS aarch64 against Herdr 0.9.0 (protocol 22) and Pi 0.85.1 in an isolated `fm-lab-` session (upstream issue #4115, duplicates #3639, #3487, #2908, #3545). + +Herdr keeps a Pi registration after the Pi process has exited to a shell when a nested interactive shell sits under the pane's top shell, which is the crew shape `treehouse get` leaves behind; a plain `/quit` directly under the top shell, and a `kill -9` of Pi, both released it on this version. +Reproduced in the lab with a nested `zsh` under the pane shell, then `pi` with no prompt, then `/quit`: + +```sh +herdr pane run w1:p1 zsh --session "$LAB"; herdr pane run w1:p1 pi --session "$LAB" +herdr agent get w1:p1 --session "$LAB" | jq -c '.result.agent | {agent, agent_status}' +herdr pane process-info --pane w1:p1 --session "$LAB" | jq -c '.result.process_info | {shell_pid, fg: .foreground_process_group_id, procs: [.foreground_processes[] | {pid, name, argv0}]}' +herdr pane send-text w1:p1 '/quit' --session "$LAB"; herdr pane send-keys w1:p1 Enter --session "$LAB" +herdr agent get w1:p1 --session "$LAB" | jq -c '.result.agent | {agent, agent_status}' +herdr pane process-info --pane w1:p1 --session "$LAB" | jq -c '.result.process_info | {shell_pid, fg: .foreground_process_group_id, procs: [.foreground_processes[] | {pid, name, argv0}]}' +``` + +```text +{"agent":"pi","agent_status":"idle"} +{"shell_pid":87754,"fg":35952,"procs":[{"pid":35952,"name":"node","argv0":"pi"}]} +{"agent":"pi","agent_status":"idle"} +{"shell_pid":87754,"fg":35834,"procs":[{"pid":35834,"name":"zsh","argv0":"zsh"}]} +``` + +Before the fix `fm_backend_agent_state herdr` read that second state as `alive`, so `bin/fm-control.sh <id> relaunch` and `bin/fm-spawn.sh --relaunch` were refused for as long as the registration lived, which is hours. +The registration is still present after the wait, and Herdr's own `pane report-agent` leaves the same shape behind on any pane, which is what the lifecycle-control guard uses. + +Two vendor facts the fix rests on, both read from the outputs above and from `fm_backend_herdr_pane_process_state`'s `pane process-info` parse: + +- Pi's process presents with kernel name `node` and argv0 `pi` (its foreground group also carries Pi's child `node` helpers with argv0 such as `npm view ... version`), so a running Pi is attributed by argv[0] exactly as the tmux probe attributes it; a symlink named `claude` to `sleep` presents as name `sleep`, argv0 `claude`. +- Herdr creates the record with its own placeholder `agent_status` of `unknown` the moment it notices Pi, before Pi's extension reports `idle`; that transient reads `unknown` in the pane classifier as it always did, and only a lifecycle status is subject to the process-level proof. + +Subcommand presence below the 0.9.0 measurement, checked 2026-09-10 on macOS aarch64 against the pinned upstream release clients fetched from `https://github.com/ogulcancelik/herdr/releases/download/v<version>/herdr-macos-aarch64`: + +| Release | sha256 | +|---------|--------| +| 0.7.1 | `16f4653f0491ea1e7d2b46b5b02542f18e1b82e88daaf9e2900572e5bb634df8` | +| 0.7.3 | `b31345392d004ec1f1b2c821e1ad601019fa8385fe1e4c6931321eb58a920773` | +| 0.7.4 | `24992e1625dbdcb18354a59e299e4b263c312400b31396cdc07cd46ed57f24a7` | +| 0.7.5 | `37350546b0012555943b92eaf962665de4e264395baeb44227b8015e8ff5b0d6` | + +The command run against each client was `<client> pane --help`, which is client-side, session-independent, and opens no socket, and each printed the line: + +```text +process-info Show pane process information +``` + +This proves subcommand presence in the client only, not the server response shape, which is measured only on 0.9.0 above. + +The live guard that refreshes this record runs by default wherever Herdr and Pi are installed, spends no model token, and fails naming both versions: + +```sh +tests/fm-herdr-pi-stale-registration-live-e2e.test.sh +``` + +Observed 2026-09-10: + +```text +# pi 0.85.1 under herdr 0.9.0: registered idle, foreground [{"name":"node","argv0":"node"},{"name":"node","argv0":"node"},{"name":"node","argv0":"rpiv-ask-user-question version"},{"name":"node","argv0":"npm view gentle-engram version"},{"name":"node","argv0":"pi"}] +ok - real herdr 0.9.0 + pi 0.85.1: a running registered pi classifies alive at process level +# herdr 0.9.0 kept the pi registration (idle) after /quit under a nested shell: the stale-registration branch is exercised +ok - real herdr 0.9.0 + pi 0.85.1: the registration left behind by a quit pi reads stale-agent and recovers as dead +``` + +`tests/fm-control-herdr-smoke.test.sh` proves the same shape through the control plane with no harness launched (the two `stale` lines under "Agent lifecycle control" above): a registration over a real agent-named process reads `alive`, stopping that process makes the pane read `stale-agent` and recover as `dead` while `agent get` still reports the record, `exit` then reports `already-stopped`, and `--relaunch` reuses the same endpoint with the local copy intact. +`tests/fm-backend-herdr.test.sh` pins the logic portably with canned `process-info` bodies over real processes, driving the signals apart: the identical shell-only foreground reads `stale-agent` for a childless shell and `live` when an agent-named process is still a descendant of that shell, a `working`, `done`, or `blocked` record over a shell-only pane reads the same as `idle`, an unreadable process view reads `unknown` and refuses husk closing, a transient prompt helper beside the shell settles into `stale-agent` on the next shell-only sample while a foreground that never settles within the bound still reads `live`, and `busy_state` verifies a `working` record before reporting busy. +`tests/fm-crew-state.test.sh` pins the recovery classifier: a stale registration over a shell-only pane reports agent gone rather than alive or unreachable, and a stale `working` record never reports the pane working. +A stale-registration pane is never a husk: create, reclaim, presentation recovery, and session cleanup keep refusing it, and only recovery reuses it. + ### Away-mode transport The Pi/Herdr return and injection path was reverified on Herdr 0.7.3 and Pi 0.80.7: @@ -1249,6 +1516,7 @@ FM_AFK_PI_HERDR_E2E=1 HERDR_LAB_HELPER=bin/fm-herdr-lab.sh \ Observed guarantees: pending composer input refused injection and raised one alert; idle Pi accepted one marked escalation; the return gate refused ordinary work while a live blocker remained; resolving the blocker allowed the return flow. The dedicated Herdr daemon workspace topology is covered by `tests/fm-afk-launch.test.sh` and preserves the captain tab's pane count. +The catch-up reporting boundary is pinned by `tests/fm-afk-return.test.sh`: Bearings reports a pending catch-up while an active away window still refuses. ## Zellij @@ -1433,8 +1701,9 @@ Read from the live agent process and from a tool subprocess it spawned: | `CURSOR_CONVERSATION_ID=<uuid>` | child/tool processes | | `AGENT_TRANSCRIPTS=<projects-root>/<slug>/agent-transcripts` | child/tool processes | -Cursor does not clear an inherited `CLAUDECODE`, so ordering decides the verdict. -With both markers set, `bin/fm-harness.sh` reports `cursor`; with `CLAUDECODE` alone it still reports `claude`. +Cursor does not clear an inherited `CLAUDECODE`, so ordering decides the verdict within the marker layer. +With both markers set and ancestry silent, `bin/fm-harness.sh` reports `cursor`; with `CLAUDECODE` alone it still reports `claude`. +[Harness detection precedence](#harness-detection-precedence) owns what happens when a structural ancestor of another harness is present, which outranks either marker. ### Composer @@ -1560,31 +1829,39 @@ FM_HARNESS_LIVENESS_DRIFT=1 bin/fm-test-run.sh tests/fm-harness-liveness-drift-l ## Antigravity CLI (agy) control mechanics -Measured 2026-08-18 on agy 1.1.14, in a throwaway tmux pane on a temporary worktree whose exact path was seeded through `fm_agy_trust_add` and removed afterward through `fm_agy_trust_remove`. -agy is a crew-only, Herdr-only spawn adapter, but its interrupt and exit mechanics are properties of the CLI rather than of a session provider, so a plain pane is the correct place to read them. -This record exists because `bin/fm-control-lib.sh` refuses any lifecycle verb on a harness with no rows, which is what an agy task got before these values were measured. +Refreshed 2026-09-14 on agy 1.2.2 and Herdr 0.8.2. +The standalone vendor guard verified a real launch prompt, the busy footer, cancellation with one Escape, and process exit through `/quit`. +The former `/exit` control row is superseded by this measured command. +The production adapter remains crew-only and Herdr-only; standalone terminal checks establish CLI behavior without expanding dispatch support. +Herdr native identity and state remain the production busy source, and rendered text is not a replacement for an unknown native state. -| Fact | Measured value | -|---|---| -| Workspace trust | `--dangerously-skip-permissions` does NOT cover it: an unseeded launch renders `Do you trust the contents of this project?` with `> Yes, I trust this folder`. A seeded exact path launches straight into the session. | -| Mid-turn footer | `esc to cancel` (the idle footer reads `? for shortcuts`). | -| Interrupt | A SINGLE Escape. The turn closes with `⎿ Interrupted · What should Antigravity CLI do instead?`. | -| Interrupt clear key | None. After the interrupt the composer is a bare `>` with no restored prompt text, so nothing has to be cleared before the next send. | -| Interrupt acknowledgement source | None. The `Interrupted` line is rendered text, not a durable typed close, and agy's busy state comes from Herdr's native agent-state only. | -| Exit command | `/exit` (`/exit Exit the CLI` in the slash-command popup). One Enter selects it and the process exits. | +```sh +FM_AGY_SIGNALS_LIVE=1 bin/fm-test-run.sh tests/fm-agy-signals-live-e2e.test.sh +``` + +```text +ok - the real agy busy footer matches fm_busy_agy_tail_busy in flight +ok - the real agy worker processed its launch prompt +ok - a single Escape cancels the real agy turn +ok - /quit stops the real agy process +``` -The exact commands that produced the values above: +The portable control regression is `test_agy_has_verified_control_rows` in `tests/fm-control.test.sh`. +The full native lifecycle guard is `tests/fm-agy-smoke.test.sh`. +It waits within its existing launch window for actual generic liveness: Herdr can register agy with `agent_status=unknown` before publishing `working`. +The smoke uses a short shell wait in each real prompt so its working-to-idle observations cannot disappear inside an instantaneous one-word response. + +Three consecutive complete native trials passed on 2026-09-14, each including launch, working-to-idle settlement, confirmed atomic steering, native completion wake, trust removal, and named-lab teardown with the default-session tripwire. ```sh -tmux -L agyverify new-session -d -s v -x 200 -y 50 -c "$WT" \ - "agy --dangerously-skip-permissions --effort high --prompt-interactive '<long prompt>'" -tmux -L agyverify send-keys -t v:0.0 Escape -tmux -L agyverify capture-pane -p -t v:0.0 -tmux -L agyverify send-keys -t v:0.0 '/exit' Enter +HERDR_LAB_HELPER=/home/sungin/firstmate/bin/fm-herdr-lab.sh FM_CURSOR_AGY_LIVE_E2E=1 bin/fm-test-run.sh tests/fm-agy-smoke.test.sh ``` -The portable regression that pins the resulting rows is `test_agy_has_verified_control_rows` in `tests/fm-control.test.sh`. -The end-to-end agy lane on its real backend stays in the Herdr-gated `tests/fm-agy-smoke.test.sh`. +```text +FM_TEST_SUMMARY total=1 failed=0 skipped_gate=0 duration_ms=24444 +FM_TEST_SUMMARY total=1 failed=0 skipped_gate=0 duration_ms=20245 +FM_TEST_SUMMARY total=1 failed=0 skipped_gate=0 duration_ms=19678 +``` ## Pi supervision branch diff --git a/tests/assets/board-render-harness.mjs b/tests/assets/board-render-harness.mjs index e21a8d2dd5d..c181aed5c88 100644 --- a/tests/assets/board-render-harness.mjs +++ b/tests/assets/board-render-harness.mjs @@ -3,7 +3,9 @@ // asserted through the real template rather than by reading its source. // // Usage: node board-render-harness.mjs <built-board.html> -// Prints one JSON document: { stats:[{n,label}], charted:[{title,sub,badges,pickable}] } +// Prints one JSON document: +// { stats:[{n,label}], underway:[{title,sub,badges}], +// charted:[{title,sub,badges,pickable}], empty, more, error } import { readFileSync } from "node:fs"; const html = readFileSync(process.argv[2], "utf8"); @@ -93,18 +95,24 @@ const stats = strip.children.map((t) => ({ label: t.children.find((c) => c.className.includes("bb-stat__label"))?.textContent, })); +const rowsOf = (container) => + container.children + .filter((r) => r.className.split(/\s+/).includes("bb-row")) + .map((row) => { + const main = row.children.find((c) => c.className.includes("bb-row__main")); + return { + title: main?.children.find((c) => c.className.includes("bb-row__title"))?.textContent ?? "", + sub: main?.children.find((c) => c.className.includes("bb-row__sub"))?.textContent ?? "", + badges: badgesOf(row), + pickable: row.children.some((c) => c.className.includes("bb-pick") && !c.className.includes("spacer")), + }; + }); + +const uw = byId.get("bb-underway") || new Node("div"); +const underway = rowsOf(uw); + const ch = byId.get("bb-charted") || new Node("div"); -const charted = ch.children - .filter((r) => r.className.split(/\s+/).includes("bb-row")) - .map((row) => { - const main = row.children.find((c) => c.className.includes("bb-row__main")); - return { - title: main?.children.find((c) => c.className.includes("bb-row__title"))?.textContent ?? "", - sub: main?.children.find((c) => c.className.includes("bb-row__sub"))?.textContent ?? "", - badges: badgesOf(row), - pickable: row.children.some((c) => c.className.includes("bb-pick") && !c.className.includes("spacer")), - }; - }); +const charted = rowsOf(ch); // A fail-closed render replaces the page body instead of the board sections, so // surface it rather than reporting an empty board as a successful render. const errorText = [...byId.entries()] @@ -114,4 +122,5 @@ const errorText = [...byId.entries()] const empty = ch.children.filter((c) => c.className.includes("bb-empty")).map((c) => c.textContent); const more = ch.children.filter((c) => c.className.includes("bb-morechip")).map((c) => c.textContent); -process.stdout.write(JSON.stringify({ stats, charted, empty, more, error: errorText }) + "\n"); +process.stdout.write( + JSON.stringify({ stats, underway, charted, empty, more, error: errorText }) + "\n"); diff --git a/tests/fm-afk-contract.test.sh b/tests/fm-afk-contract.test.sh index 1ad4a6cd05b..36acf304d0c 100755 --- a/tests/fm-afk-contract.test.sh +++ b/tests/fm-afk-contract.test.sh @@ -525,6 +525,157 @@ test_inputs_are_validated() { pass "malformed inputs and foreign record versions are refused rather than guessed" } +test_merge_grants_round_trip_and_read_back() { + local home out + home=$(make_home grants-roundtrip) + out=$(contract "$home" propose --grant task-x1 --grant task-y2 --words 'merge those two when green') || fail "grant proposal failed: $out" + assert_contains "$out" 'merge when green (task ids): task-x1, task-y2' 'read-back did not list the granted ids' + [ "$(contract "$home" grants --proposal)" = "$(printf 'task-x1\ntask-y2')" ] \ + || fail "proposal grants subcommand: $(contract "$home" grants --proposal)" + contract "$home" confirm >/dev/null || fail "grant confirm failed" + [ "$(contract "$home" grants)" = "$(printf 'task-x1\ntask-y2')" ] \ + || fail "confirmed grants subcommand: $(contract "$home" grants)" + grep -q '^merge_grants:$' "$home/state/.afk-contract" || fail "confirmed record lacks merge_grants list" + grep -q ' - task-x1' "$home/state/.afk-contract" || fail "confirmed record dropped task-x1" + pass "merge grants round-trip through propose, confirm, read-back, and grants" +} + +test_merge_grants_empty_form_and_usage_errors() { + local home out rc + home=$(make_home grants-empty) + contract "$home" propose >/dev/null || fail "empty grant proposal failed" + grep -qxF 'merge_grants: -' "$home/state/.afk-contract.proposed" \ + || fail "empty grants did not write merge_grants: -" + [ -z "$(contract "$home" grants --proposal)" ] || fail "empty grants subcommand was not empty" + set +e + out=$(contract "$home" propose --grant 'bad id' 2>&1) + rc=$? + set -e + [ "$rc" -eq 2 ] || fail "invalid grant id should be usage error (rc=$rc): $out" + set +e + out=$(contract "$home" propose --grant task-x1 --grant task-x1 2>&1) + rc=$? + set -e + [ "$rc" -eq 2 ] || fail "duplicate grant id should be usage error (rc=$rc): $out" + pass "empty grants write the scalar form, and invalid or duplicate ids are usage errors" +} + +test_legacy_record_without_merge_grants_reads_empty() { + local home record + home=$(make_home grants-legacy) + contract "$home" propose >/dev/null || fail "legacy proposal failed" + contract "$home" confirm >/dev/null || fail "legacy confirm failed" + record="$home/state/.afk-contract" + awk '!/^merge_grants/' "$record" > "$home/legacy" || fail "could not strip merge_grants" + mv "$home/legacy" "$record" + contract "$home" validate >/dev/null || fail "a pre-field v1 record must still validate" + [ -z "$(contract "$home" grants)" ] || fail "a missing merge_grants field must read as an empty list" + pass "a pre-field v1 record reads as empty grants rather than skipping the field" +} + +test_malformed_merge_grants_refuse_validation() { + local home record out rc + home=$(make_home grants-malformed-scalar) + contract "$home" propose >/dev/null || fail "malformed scalar proposal failed" + contract "$home" confirm >/dev/null || fail "malformed scalar confirm failed" + record="$home/state/.afk-contract" + awk '{ print; if ($0 == "merge_grants: -") print " - task-x1" }' "$record" > "$home/malformed" + mv "$home/malformed" "$record" + set +e + out=$(contract "$home" validate 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "indented data attached to scalar merge_grants validated" + assert_contains "$out" 'invalid merge_grants field' 'attached scalar data refusal wording' + + home=$(make_home grants-malformed-duplicate) + contract "$home" propose --grant task-x1 >/dev/null || fail "duplicate field proposal failed" + contract "$home" confirm >/dev/null || fail "duplicate field confirm failed" + record="$home/state/.afk-contract" + printf 'merge_grants: -\n' >> "$record" + set +e + out=$(contract "$home" validate 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "duplicate merge_grants fields validated" + assert_contains "$out" 'invalid merge_grants field' 'duplicate field refusal wording' + pass "malformed and duplicate merge-grant fields fail record validation" +} + +test_archive_drops_live_grants() { + local home rc + home=$(make_home grants-archive) + contract "$home" propose --grant task-x1 >/dev/null || fail "archive grant proposal failed" + contract "$home" confirm >/dev/null || fail "archive grant confirm failed" + contract "$home" archive >/dev/null || fail "archive failed" + [ ! -f "$home/state/.afk-contract" ] || fail "archive left the live record" + set +e + contract "$home" grants >/dev/null 2>&1 + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "grants on the live path succeeded after archive" + pass "archive removes live grants so archived copies are not consulted" +} + +# The record-mutating commands share one lock with the subsystems that read this +# record's authority and then act on it (bin/fm-pr-merge.sh reads the grants and +# merges). While a reader holds that lock, confirm and archive must refuse and +# change nothing, so no publication, replacement, or archive can land inside the +# window between that read and the action it authorized. +test_record_changes_refuse_while_a_reader_holds_the_lock() { + local home lock holder_pid i rc out before + home=$(make_home lock-contended) + contract "$home" propose --grant task-x1 >/dev/null || fail "lock-contended: proposal failed" + contract "$home" confirm >/dev/null || fail "lock-contended: confirm failed" + before=$(cat "$home/state/.afk-contract") + lock="$home/state/.afk-contract.lock" + + FM_STATE_OVERRIDE="$home/state" bash -c ' + . "$1" + fm_lock_acquire_wait "$2" || exit 10 + printf "ready\n" > "$3" + while [ ! -e "$4" ]; do sleep 0.05; done + fm_lock_release "$2" + ' _ "$ROOT/bin/fm-wake-lib.sh" "$lock" "$home/holder.ready" "$home/release" & + holder_pid=$! + i=0 + while [ "$i" -lt 100 ] && [ ! -s "$home/holder.ready" ]; do + sleep 0.05 + i=$((i + 1)) + done + [ -s "$home/holder.ready" ] \ + || { kill "$holder_pid" 2>/dev/null || true; fail "lock-contended: the fixture never took the lock"; } + + set +e + out=$(FM_TEST_AFK_CONTRACT_LOCK_TIMEOUT=1 contract "$home" archive 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || { kill "$holder_pid" 2>/dev/null || true; fail "lock-contended: archive ran while the record was locked"; } + assert_contains "$out" 'locked by live process' "lock-contended: the archive refusal did not name the live holder" + [ -f "$home/state/.afk-contract" ] \ + || { kill "$holder_pid" 2>/dev/null || true; fail "lock-contended: the refused archive still moved the record"; } + + contract "$home" propose --grant task-other >/dev/null || fail "lock-contended: replacement proposal failed" + set +e + out=$(FM_TEST_AFK_CONTRACT_LOCK_TIMEOUT=1 contract "$home" confirm 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || { kill "$holder_pid" 2>/dev/null || true; fail "lock-contended: confirm replaced the record while it was locked"; } + assert_contains "$out" 'locked by live process' "lock-contended: the confirm refusal did not name the live holder" + [ "$(cat "$home/state/.afk-contract")" = "$before" ] \ + || { kill "$holder_pid" 2>/dev/null || true; fail "lock-contended: the refused confirm changed the standing record"; } + [ "$(contract "$home" grants)" = task-x1 ] \ + || { kill "$holder_pid" 2>/dev/null || true; fail "lock-contended: a read subcommand did not see the unchanged grants"; } + + : > "$home/release" + wait "$holder_pid" || fail "lock-contended: the fixture holder did not release cleanly" + contract "$home" confirm >/dev/null 2>&1 || fail "lock-contended: confirm failed once the lock cleared" + [ "$(contract "$home" grants)" = task-other ] \ + || fail "lock-contended: the released replacement did not take effect" + contract "$home" archive >/dev/null || fail "lock-contended: archive failed once the lock cleared" + pass "confirm and archive refuse while the record is locked, and proceed once it clears" +} + test_fields_refuse_each_missing_part_by_name test_omitted_stop_confirms_as_no_stop test_never_set_flags_without_refusing_and_never_over_matches @@ -544,3 +695,9 @@ test_validation_rejects_blank_stop_and_refused_text test_validation_rejects_damaged_words_blocks test_archive_moves_the_record_aside_and_is_idempotent test_inputs_are_validated +test_merge_grants_round_trip_and_read_back +test_merge_grants_empty_form_and_usage_errors +test_legacy_record_without_merge_grants_reads_empty +test_malformed_merge_grants_refuse_validation +test_archive_drops_live_grants +test_record_changes_refuse_while_a_reader_holds_the_lock diff --git a/tests/fm-afk-launch.test.sh b/tests/fm-afk-launch.test.sh index 1a70f571f47..177cf6b6bf9 100755 --- a/tests/fm-afk-launch.test.sh +++ b/tests/fm-afk-launch.test.sh @@ -55,6 +55,42 @@ confirm_posture() { # <home> # back, `confirm` records it and announces hold-for-return, and every daemon # path requires that confirmed record. # --------------------------------------------------------------------------- +unit_quiet_attended_lifecycle() { + local st out + st=$(mktemp -d "${TMPDIR:-/tmp}/fm-quiet.XXXXXX") + mkdir -p "$st/state" + if ! FM_HOME="$st" FM_STATE_OVERRIDE="$st/state" FM_AFK_MODE=quiet "$LAUNCH" start-native >"$st/output" 2>&1; then + fail "quiet: explicit native entry without an away record failed: $(cat "$st/output")" + fi + [ "$(head -1 "$st/state/.afk")" = quiet ] || fail "quiet: entry did not write quiet mode" + [ ! -e "$st/state/.afk-contract" ] || fail "quiet: entry fabricated an away record" + out=$(FM_HOME="$st" FM_STATE_OVERRIDE="$st/state" "$ROOT/bin/fm-afk-return.sh" guard 2>&1) \ + || fail "quiet: ordinary work guard refused attended mode: $out" + [ "$(head -1 "$st/state/.afk")" = quiet ] || fail "quiet: ordinary guard exited quiet mode" + if FM_HOME="$st" FM_STATE_OVERRIDE="$st/state" "$ROOT/bin/fm-afk-return.sh" begin >"$st/output" 2>&1; then + fail "quiet: ordinary return consumed quiet without explicit quiet-off" + fi + [ -e "$st/state/.afk" ] || fail "quiet: ordinary return removed the flag" + FM_HOME="$st" FM_STATE_OVERRIDE="$st/state" "$ROOT/bin/fm-afk-return.sh" quiet-off >"$st/output" 2>&1 \ + || fail "quiet: explicit exit failed: $(cat "$st/output")" + [ ! -e "$st/state/.afk" ] && [ ! -e "$st/state/.afk-daemon-terminal" ] \ + || fail "quiet: exit retained daemon lifecycle records" + [ ! -e "$st/state/.afk-return-catchup" ] && [ ! -d "$st/state/afk-contracts" ] \ + || fail "quiet: exit fabricated an away return" + if FM_HOME="$st" FM_STATE_OVERRIDE="$st/state" FM_AFK_MODE=away "$LAUNCH" start-native >"$st/output" 2>&1; then + fail "away: entry without a confirmed record succeeded after quiet exit" + fi + confirm_posture "$st" || fail "quiet: could not confirm away fixture" + if FM_HOME="$st" FM_STATE_OVERRIDE="$st/state" FM_AFK_MODE=quiet "$LAUNCH" start-native >"$st/output" 2>&1; then + fail "quiet: entry bypassed an active away record" + fi + if FM_HOME="$st" FM_STATE_OVERRIDE="$st/state" "$ROOT/bin/fm-afk-return.sh" quiet-off >"$st/output" 2>&1; then + fail "quiet: explicit exit consumed an active away record" + fi + [ -e "$st/state/.afk-contract" ] || fail "quiet: refused transition lost the away record" + rm -rf "$st" +} + unit_propose_confirm_records_the_posture_without_a_daemon() { local st out rc st=$(mktemp -d "${TMPDIR:-/tmp}/fm-afk-propose.XXXXXX") @@ -274,6 +310,112 @@ unit_fresh_vs_refresh() { rm -rf "$st" } +# --------------------------------------------------------------------------- +# UNIT 2a: away/quiet mode plumbing (kunchenguid/firstmate#2356). fm_afk_mode +# is the single owner of reading the mode; these pin its write side +# (fm_afk_launch_flag_write / fm_afk_flag_write) against the exact double- +# write risk a live entry hits - the launcher writes the flag, then the +# terminal-side fm-afk-start.sh entry re-writes it a second time on every +# real (non-native) entry, per UNIT 2 above. +# --------------------------------------------------------------------------- +read_mode() { # <state-dir> + bash -c '. "$1"; fm_afk_mode "$2"' _ "$ROOT/bin/fm-wake-lib.sh" "$1" +} + +unit_mode_explicit_write() { + local st out + st=$(mktemp -d "${TMPDIR:-/tmp}/fm-afk-mode-explicit.XXXXXX") + mkdir -p "$st/state" + FM_HOME="$st" FM_STATE_OVERRIDE="$st/state" FM_AFK_MODE=quiet \ + bash -c '. "$1"; fm_afk_launch_flag_write' _ "$LAUNCH" + out=$(read_mode "$st/state") + if [ "$out" = quiet ]; then + pass "mode: a fresh entry with FM_AFK_MODE=quiet writes quiet" + else + fail "mode: explicit FM_AFK_MODE=quiet fresh entry wrote '$out' instead of quiet" + fi + rm -rf "$st" +} + +unit_mode_fresh_defaults_away() { + local st out + st=$(mktemp -d "${TMPDIR:-/tmp}/fm-afk-mode-default.XXXXXX") + mkdir -p "$st/state" + FM_HOME="$st" FM_STATE_OVERRIDE="$st/state" \ + bash -c '. "$1"; fm_afk_launch_flag_write' _ "$LAUNCH" + out=$(read_mode "$st/state") + if [ "$out" = away ]; then + pass "mode: a fresh entry with FM_AFK_MODE unset defaults to away" + else + fail "mode: fresh unset-mode entry wrote '$out' instead of away" + fi + rm -rf "$st" +} + +unit_mode_refresh_preserves_quiet() { + local st sleep_pid lock out + st=$(mktemp -d "${TMPDIR:-/tmp}/fm-afk-mode-preserve.XXXXXX") + mkdir -p "$st/state" + printf 'quiet\n%s\n' "$(date '+%s')" > "$st/state/.afk" + sleep 600 & + sleep_pid=$! + lock="$st/state/.supervise-daemon.lock" + mkdir -p "$lock" + printf '%s' "$sleep_pid" > "$lock/pid" + ( . "$ROOT/bin/fm-wake-lib.sh"; fm_pid_identity "$sleep_pid" > "$lock/pid-identity" 2>/dev/null ) || true + # The exact real-entry shape: a bare direct re-write with no explicit mode, + # simulating the terminal-side fm-afk-start.sh redundant write that would + # silently clobber quiet back to away if it were not preserve-on-refresh. + FM_HOME="$st" FM_STATE_OVERRIDE="$st/state" "$START" >/dev/null 2>&1 + out=$(read_mode "$st/state") + if [ "$out" = quiet ]; then + pass "mode: a bare refresh (FM_AFK_MODE unset) of an already-running quiet daemon preserves quiet, never resets to away" + else + fail "mode: refresh incorrectly changed quiet mode to '$out'" + fi + kill "$sleep_pid" 2>/dev/null || true + wait "$sleep_pid" 2>/dev/null || true + rm -rf "$st" +} + +unit_mode_garbage_and_legacy_content_reads_away() { + local st out + st=$(mktemp -d "${TMPDIR:-/tmp}/fm-afk-mode-garbage.XXXXXX") + mkdir -p "$st/state" + + : > "$st/state/.afk" + out=$(read_mode "$st/state") + if [ "$out" = away ]; then + pass "mode: an empty (legacy pre-mode) flag reads as away" + else + fail "mode: empty flag read as '$out' instead of away" + fi + + date '+%s' > "$st/state/.afk" + out=$(read_mode "$st/state") + if [ "$out" = away ]; then + pass "mode: a bare-epoch-timestamp (legacy pre-mode) flag reads as away" + else + fail "mode: legacy timestamp flag read as '$out' instead of away" + fi + + printf 'nonsense-mode\n' > "$st/state/.afk" + out=$(read_mode "$st/state") + if [ "$out" = away ]; then + pass "mode: unrecognized content falls back to away" + else + fail "mode: unrecognized content read as '$out' instead of away" + fi + + out=$(read_mode "$st/state/missing") + if [ "$out" = away ]; then + pass "mode: a missing flag reads as away" + else + fail "mode: missing flag read as '$out' instead of away" + fi + rm -rf "$st" +} + # --------------------------------------------------------------------------- # UNIT 3: exit ordering - fm_afk_launch_stop SIGTERMs the daemon WHILE .afk is # still present (so its flush is not a no-op), and clears .afk last. @@ -1074,6 +1216,8 @@ e2e_tmux() { } unit_clear_stale +unit_quiet_attended_lifecycle + unit_propose_confirm_records_the_posture_without_a_daemon unit_pi_preserves_the_daemon_lifecycle unit_daemon_entry_requires_confirmation @@ -1081,6 +1225,10 @@ unit_failed_daemon_launch_preserves_confirmed_record unit_stop_archives_the_record_last unit_relative_paths_are_absolute_before_daemon_launch unit_fresh_vs_refresh +unit_mode_explicit_write +unit_mode_fresh_defaults_away +unit_mode_refresh_preserves_quiet +unit_mode_garbage_and_legacy_content_reads_away unit_stop_ordering unit_stop_rejects_reused_pid unit_failed_start_rolls_back_state diff --git a/tests/fm-afk-pi-herdr-return-e2e.test.sh b/tests/fm-afk-pi-herdr-return-e2e.test.sh index ccbe496860a..79b3ad476ef 100755 --- a/tests/fm-afk-pi-herdr-return-e2e.test.sh +++ b/tests/fm-afk-pi-herdr-return-e2e.test.sh @@ -269,19 +269,26 @@ set -e DAEMON_STARTED=0 [ "$RETURN_RC" -eq 3 ] || fail "return catch-up did not gate the still-live blocker (rc=$RETURN_RC): $RETURN_OUT" assert_contains "$RETURN_OUT" 'firstmate-actionable blocker: repair-task [key=synthetic-dependency]' "return gate did not assign remediation" -set +e +assert_contains "$RETURN_OUT" '=== Return brief (away ' "the return did not render the brief" +assert_contains "$RETURN_OUT" 'Supervisor health:' "the brief did not lead with supervisor health" +[ ! -f "$STATE/.afk-contract" ] || fail "the return did not archive the away-posture record" BEARINGS_OUT=$(PATH="$FAKEBIN:$ORIGINAL_PATH" HERDR_SESSION="$SESSION" FM_ROOT_OVERRIDE="$PROJECT" FM_HOME="$HOME_DIR" FM_STATE_OVERRIDE="$STATE" \ - "$ROOT/bin/fm-bearings-snapshot.sh" --json 2>&1) -BEARINGS_RC=$? -set -e -[ "$BEARINGS_RC" -eq 3 ] || fail "Bearings bypassed the return gate (rc=$BEARINGS_RC): $BEARINGS_OUT" -pass "real unmarked Pi return opens catch-up and blocks Bearings before the unresolved blocker can be deferred" + "$ROOT/bin/fm-bearings-snapshot.sh" --json 2>&1) \ + || fail "Bearings refused behind the return gate instead of reporting it: $BEARINGS_OUT" +printf '%s' "$BEARINGS_OUT" | jq -e ' + (.in_flight | any(.id == "repair-task")) + and (.gates | any(.id == "(return-catchup)" and .reason == "away-return catch-up")) + and ([.decisions_open[].id] | index("(return-catchup)") | not)' >/dev/null \ + || fail "Bearings did not surface the catch-up posture as content: $BEARINGS_OUT" +pass "real unmarked Pi return renders the brief, opens catch-up, and reports that posture through Bearings while the blocker stays Firstmate's to remediate" printf 'resolved [key=synthetic-dependency]: refreshed the synthetic token and resumed the task\n' >> "$STATE/repair-task.status" PATH="$FAKEBIN:$ORIGINAL_PATH" HERDR_SESSION="$SESSION" FM_ROOT_OVERRIDE="$PROJECT" FM_HOME="$HOME_DIR" FM_STATE_OVERRIDE="$STATE" \ "$ROOT/bin/fm-afk-return.sh" check >/dev/null || fail "remediated blocker did not clear return catch-up" PATH="$FAKEBIN:$ORIGINAL_PATH" HERDR_SESSION="$SESSION" FM_ROOT_OVERRIDE="$PROJECT" FM_HOME="$HOME_DIR" FM_STATE_OVERRIDE="$STATE" \ - "$ROOT/bin/fm-bearings-snapshot.sh" --json >/dev/null || fail "Bearings remained gated after blocker remediation" + "$ROOT/bin/fm-bearings-snapshot.sh" --json \ + | jq -e '[.gates[].id] | index("(return-catchup)") | not' >/dev/null \ + || fail "Bearings kept the catch-up posture row after the gate cleared" # A clean re-entry creates no stale delivery or alert, and an immediate return is # idempotently clear because the keyed blocker is resolved. diff --git a/tests/fm-afk-return.test.sh b/tests/fm-afk-return.test.sh index a510e70d9bf..376cf0d4044 100755 --- a/tests/fm-afk-return.test.sh +++ b/tests/fm-afk-return.test.sh @@ -4,8 +4,10 @@ # Covers the second half of the 2026-07-14 incident: an away-mode blocked event # survived in durable state, but the ordinary return request could proceed to # Bearings before Firstmate owned remediation. The shared script now stops, -# drains, preserves evidence, and refuses ordinary work until every live open -# `blocked:` event is resolved or durably reclassified. +# drains, preserves evidence, and holds ordinary WORK until every live open +# `blocked:` event is resolved or durably reclassified. Reporting is not work: +# Bearings renders behind the catch-up gate and surfaces the catch-up posture +# as content, so a returning captain still gets the picture. # The brief cases pin the away-posture redesign's return: the brief is composed # from the archived posture record, the outcome store, the held set, and the # status logs, health first, and the gate shrinks to what the away session could @@ -94,11 +96,21 @@ EOF printf 'blocked [key=%s]: firstmate can refresh the synthetic token\n' "$key" > "$dir/home/state/repair-task.status" } -test_return_gate_orders_catchup_before_bearings() { - local dir out rc gate wake_count +test_return_gate_owns_remediation_and_reports_catchup_to_bearings() { + local dir out rc gate wake_count i toon gate_header dir="$TMP_ROOT/ordering" install_runner "$dir" seed_live_blocker "$dir" herdr synthetic-dependency + { + printf '## In flight\n\n## Queued\n' + i=1 + while [ "$i" -le 20 ]; do + printf -- '- [ ] queued-%02d - Queued gate %02d (repo: sample) (kind: ship) (since 2026-06-%02d)\n' \ + "$i" "$i" "$i" + i=$((i + 1)) + done + printf '\n## Done\n' + } > "$dir/home/data/backlog.md" date +%s > "$dir/home/state/.afk" printf 'repair-task.status: blocked synthetic dependency\n' > "$dir/home/state/.subsuper-escalations" printf 'fm away-mode inject WEDGED: 4555s undelivered\n' > "$dir/home/state/.subsuper-inject-wedged" @@ -124,14 +136,40 @@ test_return_gate_orders_catchup_before_bearings() { [ -s "$dir/home/state/.fake-drain" ] || fail "blocked return acknowledged its emitted wake before handling completed" [ ! -e "$dir/home/state/.fake-drain-acks" ] || fail "blocked return crossed the post-handling acknowledgement boundary" - # The exact incident regression: Bearings is an ordinary request and must - # refuse before reading/rendering while this shared gate remains open. + # The captain is back and asking for the picture: Bearings reports the + # catch-up posture as content rather than refusing. The blocked worker still + # projects as its own Underway row, and the catch-up posture is a separate + # action-free Charted Next gate row that never becomes a Captain's Call entry. + out=$(FM_HOME="$dir/home" FM_STATE_OVERRIDE="$dir/home/state" "$ROOT/bin/fm-bearings-snapshot.sh" --json 2>&1) \ + || fail "Bearings should render behind the return catch-up gate: $out" + # The live projected state of the blocked worker follows its endpoint, which + # this fixture deliberately does not stand up; what the gate must no longer + # do is stop the fleet read, so the worker has to reach Underway at all. + printf '%s' "$out" | jq -e ' + (.in_flight | any(.id == "repair-task")) + and (.gates[0].id == "(return-catchup)" and .gates[0].filed == null) + and (.gates | length == 21) + and ([.gates[] | select(.id | startswith("queued-"))] | length == 20) + and (.gates | any(.id == "(return-catchup)" + and .owner == "(main)" + and .reason == "away-return catch-up" + and (.title | test("^1 blocker")))) + and ([.decisions_open[].id] | index("(return-catchup)") | not)' >/dev/null \ + || fail "Bearings did not reserve the catch-up posture outside bounded action-free gate rows: $out" + toon=$(FM_HOME="$dir/home" FM_STATE_OVERRIDE="$dir/home/state" "$ROOT/bin/fm-bearings-snapshot.sh" 2>&1) \ + || fail "default Bearings should render behind the return catch-up gate: $toon" + gate_header=$(printf '%s\n' "$toon" | awk '/^gates\[[0-9]+\]\{/ { print; exit }') + assert_contains "$gate_header" '{id,title,blocked_by,reason,owner,filed}' "catch-up removed filed from the TOON gate schema" + assert_contains "$toon" '2026-06-20' "catch-up removed durable gate dates from default Bearings output" + + # The guard itself still separates its two branches by exit status, so an + # active away window keeps refusing while catch-up reports. set +e - out=$(FM_HOME="$dir/home" FM_STATE_OVERRIDE="$dir/home/state" "$ROOT/bin/fm-bearings-snapshot.sh" --json 2>&1) + out=$(FM_HOME="$dir/home" FM_STATE_OVERRIDE="$dir/home/state" "$dir/bin/fm-afk-return.sh" guard 2>&1) rc=$? set -e - [ "$rc" -eq 3 ] || fail "Bearings should refuse behind the return gate (rc=$rc): $out" - assert_contains "$out" 'return catch-up is pending' "Bearings refusal did not point to the shared return owner" + [ "$rc" -eq 4 ] || fail "the catch-up branch should be distinguishable by exit status (rc=$rc): $out" + assert_contains "$out" 'return catch-up is pending' "the catch-up refusal did not point to the shared return owner" # Restart/re-entry is idempotent: no second stop, no duplicate catch-up line, # and the same unresolved blocker remains authoritative. @@ -148,6 +186,9 @@ test_return_gate_orders_catchup_before_bearings() { printf 'resolved [key=synthetic-dependency]: refreshed the synthetic token and resumed the task\n' >> "$dir/home/state/repair-task.status" out=$(run_return "$dir" check) || fail "resolved blocker did not clear return catch-up: $out" + FM_HOME="$dir/home" FM_STATE_OVERRIDE="$dir/home/state" "$ROOT/bin/fm-bearings-snapshot.sh" --json \ + | jq -e '[.gates[].id] | index("(return-catchup)") | not' >/dev/null \ + || fail "the cleared gate left the catch-up posture row in Bearings" assert_contains "$out" 'catch-up clear' "successful check did not announce that ordinary work may proceed" [ ! -e "$gate" ] || fail "successful check left the return gate behind" [ ! -e "$dir/home/state/.subsuper-escalations" ] || fail "successful check left delivered escalation state behind" @@ -162,7 +203,7 @@ test_return_gate_orders_catchup_before_bearings() { out=$(run_return "$dir" check) || fail "an already-clear repeated check should be idempotent: $out" [ ! -e "$gate" ] || fail "idempotent clear check recreated a gate" - pass "return catch-up precedes Bearings, owns live blocker remediation, preserves evidence once, and clears idempotently" + pass "return catch-up owns live blocker remediation, reports itself to Bearings as content, preserves evidence once, and clears idempotently" } test_explicit_reclassification_requires_durable_reason() { @@ -265,6 +306,21 @@ test_away_reentry_refuses_pending_return_gate() { pass "away-mode re-entry fails closed while the prior return catch-up is pending" } +test_quiet_exit_has_no_away_catchup() { + local dir out + dir="$TMP_ROOT/quiet-mode-return" + install_runner "$dir" + printf 'quiet\n%s\n' "$(date +%s)" > "$dir/home/state/.afk" + out=$(run_return "$dir" guard) || fail "quiet ordinary work guard failed: $out" + [ -e "$dir/home/state/.afk" ] || fail "quiet ordinary work guard cleared mode" + out=$(run_return "$dir" quiet-off) || fail "quiet-off failed: $out" + [ ! -e "$dir/home/state/.afk" ] || fail "quiet exit left the mode flag" + [ ! -e "$dir/home/state/.afk-return-catchup" ] || fail "quiet exit fabricated an away gate" + [ ! -d "$dir/home/state/afk-contracts" ] || fail "quiet exit fabricated an away archive" + [ "$(wc -l < "$dir/home/stop.log" | tr -d ' ')" -eq 1 ] || fail "quiet exit did not stop the daemon exactly once" + pass "explicit quiet exit stops the shared daemon without an away catch-up" +} + test_check_retries_recorded_terminal_teardown() { local dir gate out rc dir="$TMP_ROOT/terminal-teardown" @@ -508,7 +564,7 @@ test_unreadable_outcome_store_keeps_catchup_gated() { } test_failed_held_listing_keeps_catchup_gated() { - local dir out waiting rc gate + local dir out waiting rc gate guard_out guard_rc dir="$TMP_ROOT/held-list-failure" install_runner "$dir" mkdir -p "$dir/fakebin" @@ -528,6 +584,21 @@ SH set -e [ "$rc" -eq 3 ] || fail "a failed held-set read should keep catch-up gated (rc=$rc): $out" [ -f "$gate" ] || fail "a failed held-set read did not retain the return gate" + + # A gate retained for a lifecycle reason lists no blocker at all, so the + # refusal must name what actually holds it instead of promising a blocker + # list it cannot produce, and Bearings must carry that same reason. + set +e + guard_out=$(FM_HOME="$dir/home" FM_STATE_OVERRIDE="$dir/home/state" "$dir/bin/fm-afk-return.sh" guard 2>&1) + guard_rc=$? + set -e + [ "$guard_rc" -eq 4 ] || fail "a blockerless catch-up gate should use the catch-up branch (rc=$guard_rc): $guard_out" + assert_contains "$guard_out" 'no open blocker' "the blockerless refusal did not say the gate lists no blocker" + assert_contains "$guard_out" 'catch-up retained: held set unreadable' "the blockerless refusal did not name the retention reason" + assert_not_contains "$guard_out" 'every listed blocker' "the blockerless refusal still demanded an empty blocker list" + FM_HOME="$dir/home" FM_STATE_OVERRIDE="$dir/home/state" "$ROOT/bin/fm-bearings-snapshot.sh" --json \ + | jq -e '.gates | any(.id == "(return-catchup)" and (.title | startswith("catch-up retained:")))' >/dev/null \ + || fail "Bearings did not carry the blockerless catch-up retention reason" waiting=$(printf '%s\n' "$out" | awk '/^Waiting on you:/{show=1} /^Tried and failed, or could not be fixed:/{show=0} show') assert_contains "$waiting" "held listing unavailable: $dir/home/data/backlog.md: synthetic held backlog failure; catch-up stays gated" "the failed held listing was not disclosed" assert_not_contains "$waiting" '(nothing)' "an unavailable held set was also reported as empty" @@ -604,6 +675,24 @@ test_return_brief_health_leads_with_a_gap() { pass "the return brief leads with supervisor health and names every detected gap" } +test_return_brief_does_not_report_an_acked_watcher_down_marker_as_a_gap() { + local dir out + dir="$TMP_ROOT/brief-acked-marker" + install_runner "$dir" + contract_in "$dir" propose >/dev/null 2>&1 || fail "could not propose the away-posture record" + contract_in "$dir" confirm >/dev/null 2>&1 || fail "could not write the away-posture record" + # An episode that was detected and fully handled during the away window + # leaves the marker behind in an acked state (fm-wake-lib.sh + # _fm_recovery_marker_ack); that is not an open gap. + printf 'acked:downtime:fixture-generation\n' > "$dir/home/state/.watcher-down" + touch "$dir/home/state/.last-watcher-beat" + : > "$dir/home/state/.fake-drain" + out=$(run_return "$dir" begin) || fail "a clean fleet with only a handled marker should clear the gate: $out" + assert_not_contains "$out" 'GAP: watcher downtime was detected' "an acked recovery marker was reported as an open gap" + assert_contains "$out" 'no detected gap' "a fully acked window was not reported as clean" + pass "the return brief does not report an already-acked watcher-down marker as an open gap" +} + test_return_brief_without_a_record_reports_the_legacy_flag() { local dir out dir="$TMP_ROOT/brief-legacy" @@ -693,11 +782,12 @@ test_missing_final_archive_keeps_retained_contract_gated() { pass "the retained contract epoch requires its final archive on every check" } -test_return_gate_orders_catchup_before_bearings +test_return_gate_owns_remediation_and_reports_catchup_to_bearings test_explicit_reclassification_requires_durable_reason test_captain_decision_does_not_masquerade_as_firstmate_blocker test_evidence_publication_failure_preserves_wake_for_redrain test_away_reentry_refuses_pending_return_gate +test_quiet_exit_has_no_away_catchup test_check_retries_recorded_terminal_teardown test_unreadable_superseded_archive_keeps_return_gated test_missing_final_archive_keeps_retained_contract_gated @@ -710,6 +800,7 @@ test_failed_held_listing_keeps_catchup_gated test_unreadable_status_file_keeps_catchup_gated test_return_guard_refuses_while_the_record_exists test_return_brief_health_leads_with_a_gap +test_return_brief_does_not_report_an_acked_watcher_down_marker_as_a_gap test_return_brief_without_a_record_reports_the_legacy_flag printf '\nall fm-afk-return tests passed\n' diff --git a/tests/fm-agy-adapter.test.sh b/tests/fm-agy-adapter.test.sh index f8790ec160f..66c41abf811 100755 --- a/tests/fm-agy-adapter.test.sh +++ b/tests/fm-agy-adapter.test.sh @@ -495,7 +495,7 @@ set -u cmd="" for a in "$@"; do case "$a" in - status|pane|agent|prompt|get|read|capture|send-keys|send-text) cmd="$cmd $a" ;; + status|pane|agent|prompt|get|read|capture|send-keys|send-text|process-info) cmd="$cmd $a" ;; esac done dir=$(dirname "$0")/.. @@ -509,6 +509,9 @@ if [ "${1:-}" = agent ] && [ "${2:-}" = prompt ] \ fi case "$cmd" in *"status"*) printf '{"client":{"version":"0.7.5","protocol":16},"server":{"running":true}}\n' ;; + *"pane process-info"*) + jq -cn --arg name "${FM_HERDR_FAKE_AGENT:-claude}" '{result:{type:"pane_process_info",process_info:{pane_id:"w1:p1",shell_pid:4242,foreground_process_group_id:4243,foreground_processes:[{pid:4243,name:$name,argv:[$name]}]}}}' + ;; *"pane get"*) printf '{"result":{"pane":{"pane_id":"w1:p1"}}}\n' ;; *"pane send-keys"*) printf 'enter\n' >> "$dir/enter_log" diff --git a/tests/fm-agy-harness.test.sh b/tests/fm-agy-harness.test.sh new file mode 100755 index 00000000000..a2cf0916f82 --- /dev/null +++ b/tests/fm-agy-harness.test.sh @@ -0,0 +1,940 @@ +#!/usr/bin/env bash +# Behavior tests for the verified Antigravity CLI crewmate/scout adapter. +# +# The facts pinned here are the ones an agy release could silently change and +# the ones a wrong guess would make dangerous: +# 1. agy publishes no harness-identity marker of its own (a live 1.2.0 TUI +# carries no AGY_* variable; AGENT=1 there is inherited launcher state), +# so detection is ancestry alone on the anchored process name `agy`. +# 2. The anchored match must never claim unrelated commands containing the +# fragment, and a structural agy ancestor now outranks a retained or +# inherited CLAUDECODE - tests/fm-harness-precedence.test.sh owns the +# general boundary. +# 3. The launch carries the brief via --prompt-interactive with --model, +# --effort, and --dangerously-skip-permissions; a requested model a +# reachable `agy models` omits refuses loudly instead of wedging a pane, +# while a hung or unreachable listing is cut off and never blocks. +# 4. A fresh worktree would park agy on its folder-trust dialog, so the spawn +# pre-registers the worktree in agy's own trustedWorkspaces store through +# bin/fm-agy-trust.sh (scope-refused for anything but a linked worktree +# of the project) and the post-launch gate is the backstop: it answers a +# dialog that renders anyway exactly once, never counts a busy turn as +# ready on an unregistered path until the dialog has been answered (the +# Herdr native-busy-before-dialog race), and fails the spawn with endpoint +# cleanup when the brief cannot be confirmed to run in the worktree. +# 5. agy is a crewmate/scout adapter only: a secondmate launch is refused, +# and nothing is armed as busy wiring because no writer could clear it. +# 6. The busy signature is the pinned `esc to cancel` status row alone; the +# free-floating `Generating...` word must never read busy on its own. +# 7. Herdr's registry already tracks agy, and exit detection proves the +# agent at process level before trusting any registration (the shared +# post-#4115 contract in bin/backends/herdr.sh): a registered status plus +# a process view naming agy is live and refuses replacement, a registered +# status over a proven shell-only pane is the explicit stale-agent state, +# and nothing short of that shared proof flips an agy pane to agent-free. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +# bin/fm-harness.sh checks verified ENV markers before ancestry. A suite run +# from inside another harness inherits those markers, which outrank the fake +# ancestry the detection cases set up. Drop the ambient markers so the asserted +# verdict does not depend on which harness launched the suite. +unset CLAUDECODE PI_CODING_AGENT FM_PI_HARNESS GROK_AGENT CURSOR_AGENT CURSOR_INVOKED_AS \ + ATLASSIAN_AGENT_TYPE ROVODEV_CLI GEMINI_CLI AGENT FM_OMP_HARNESS + +# shellcheck source=/dev/null +. "$ROOT/bin/fm-control-lib.sh" +# shellcheck source=/dev/null +. "$ROOT/bin/fm-busy-lib.sh" +# shellcheck source=/dev/null +. "$ROOT/bin/fm-composer-lib.sh" + +HARNESS="$ROOT/bin/fm-harness.sh" +SPAWN="$ROOT/bin/fm-spawn.sh" +TRUST="$ROOT/bin/fm-agy-trust.sh" +TMP_ROOT=$(fm_test_tmproot fm-agy-harness) + +# Exercise the real spawn entrypoint with a deterministic Herdr adapter. +# The transport fake below owns the screen transitions; native backend protocol +# and atomic delivery have separate coverage in fm-agy-adapter.test.sh. +SPAWN_ROOT="$TMP_ROOT/spawn-runtime" +mkdir -p "$SPAWN_ROOT" +cp -R "$ROOT/bin" "$SPAWN_ROOT/bin" +cp "$ROOT/.tasks.toml" "$SPAWN_ROOT/" +SPAWN="$SPAWN_ROOT/bin/fm-spawn.sh" +cat > "$SPAWN_ROOT/bin/backends/herdr.sh" <<'SH' +#!/usr/bin/env bash +fm_backend_herdr_agent_prompt_capability_check() { return 0; } +fm_backend_herdr_agent_prompt_version_check() { return 0; } +fm_backend_herdr_presentation_enabled() { return 1; } +fm_backend_herdr_projection_journal_path() { printf '%s/%s.herdr-presentation\n' "$1" "$2"; } +fm_backend_herdr_workspace_label() { printf firstmate; } +fm_backend_herdr_session() { printf fixture; } +fm_backend_herdr_container_ensure() { printf 'fixture:w1\t\n'; } +fm_backend_herdr_create_task() { printf 't1 w1:p1\n'; } +fm_backend_herdr_current_path() { printf '%s\n' "$FM_FAKE_PANE_PATH"; } +fm_backend_herdr_parse_target() { printf 'fixture w1:p1\n'; } +fm_backend_herdr_capture() { tmux capture-pane -p; } +fm_backend_herdr_capture_ansi() { tmux capture-pane -p; } +fm_backend_herdr_send_literal() { tmux send-keys -t "$1" -l "$2"; } +fm_backend_herdr_send_key() { tmux send-keys -t "$1" "$2"; } +fm_backend_herdr_send_text_line() { tmux send-keys -t "$1" -l "$2"; } +fm_backend_herdr_kill() { tmux kill-window -t "$1"; } +fm_backend_herdr_busy_state() { + case "$(cat "$FM_FAKE_AGY_STATE")" in busy|racing) printf busy ;; *) printf unknown ;; esac +} +fm_backend_herdr_agent_identity_busy_state() { + printf 'agy\t'; fm_backend_herdr_busy_state +} +fm_backend_herdr_agent_state() { printf live; } +SH + +# The store is agy's own persisted settings JSON, so trust is asserted against +# the parsed trustedWorkspaces array and preservation against parsed values. +agy_trusted_paths() { # <store> + node -e 'const fs=require("node:fs");const j=fs.existsSync(process.argv[1])?JSON.parse(fs.readFileSync(process.argv[1],"utf8")):{};for(const p of (j.trustedWorkspaces||[]))console.log(p);' "$1" +} + +agy_store_value() { # <store> <key> + node -e 'const j=JSON.parse(require("node:fs").readFileSync(process.argv[1],"utf8"));console.log(JSON.stringify(j[process.argv[2]]));' "$1" "$2" +} + +assert_agy_trusted() { # <store> <path> <msg> + agy_trusted_paths "$1" | grep -Fqx "$2" || fail "$3" +} + +assert_agy_not_trusted() { # <store> <path> <msg> + agy_trusted_paths "$1" | grep -Fqx "$2" && fail "$3" + return 0 +} + +test_agy_ancestry_detects_the_native_command_name() { + local fakebin out + fakebin=$(fm_fakebin "$TMP_ROOT/anc-native") + cat > "$fakebin/ps" <<'SH' +#!/usr/bin/env bash +case "$*" in + *"comm="*) printf '%s\n' '/usr/local/bin/agy'; exit 0 ;; + *"args="*) printf '%s\n' 'agy --prompt-interactive hello'; exit 0 ;; +esac +exit 1 +SH + chmod +x "$fakebin/ps" + out=$(PATH="$fakebin:$PATH" "$HARNESS") + [ "$out" = agy ] \ + || fail "a natively-named agy command must be detected by ancestry, got '$out'" + pass "fm-harness.sh: ancestry detects a natively-named agy command" +} + +test_agy_ancestry_rejects_unrelated_mentions() { + local fakebin out + fakebin=$(fm_fakebin "$TMP_ROOT/anc-negatives") + cat > "$fakebin/ps" <<'SH' +#!/usr/bin/env bash +case "$*" in + *"comm="*) printf '%s\n' "${FAKE_PS_COMM:?}"; exit 0 ;; + *"args="*) printf '%s\n' "${FAKE_PS_ARGS:?}"; exit 0 ;; +esac +exit 1 +SH + chmod +x "$fakebin/ps" + + out=$(FAKE_PS_COMM=magyk FAKE_PS_ARGS='magyk --serve' \ + PATH="$fakebin:$PATH" "$HARNESS") + [ "$out" != agy ] \ + || fail "an unrelated magyk command must not detect agy, got '$out'" + + out=$(FAKE_PS_COMM=bash FAKE_PS_ARGS='bash -c "echo agy --help"' \ + PATH="$fakebin:$PATH" "$HARNESS") + [ "$out" != agy ] \ + || fail "a later shell argument naming agy must not detect agy, got '$out'" + pass "fm-harness.sh: ancestry rejects unrelated agy mentions" +} + +test_agy_claims_no_inherited_launcher_marker() { + local fakebin out + # AGENT=1 was observed on a live agy TUI as inherited launcher state, so it + # must never promote to an agy identity the way GEMINI_CLI does for gemini. + out=$(AGENT=1 "$HARNESS") + [ "$out" != agy ] \ + || fail "an inherited AGENT=1 must never claim the agy identity, got '$out'" + # Drive the hazard the other way: agy does not clear an inherited CLAUDECODE, + # so a structural agy ancestor must still outrank the retained marker rather + # than being renamed away from it. Pin both halves so neither can rot + # silently. + fakebin=$(fm_fakebin "$TMP_ROOT/anc-claude") + cat > "$fakebin/ps" <<'SH' +#!/usr/bin/env bash +case "$*" in + *"comm="*) printf '%s\n' agy; exit 0 ;; + *"args="*) printf '%s\n' 'agy --prompt-interactive hi'; exit 0 ;; +esac +exit 1 +SH + chmod +x "$fakebin/ps" + out=$(CLAUDECODE=1 PATH="$fakebin:$PATH" "$HARNESS") + [ "$out" = agy ] \ + || fail "a structural agy ancestor must outrank an inherited CLAUDECODE, got '$out'" + pass "fm-harness.sh: no inherited launcher marker claims the agy identity" +} + +test_agy_control_mechanics_are_the_verified_ones() { + fm_control_harness_supported agy || fail "agy must be a supported control harness" + [ "$(fm_control_harness_family agy)" = agy ] || fail "agy must map to its own family" + fm_control_harness_supports_kind agy scout || fail "agy must run scouts" + fm_control_harness_supports_kind agy ship || fail "agy must run ships" + fm_control_harness_supports_kind agy secondmate \ + && fail "agy must refuse secondmates" || true + [ "$(fm_control_interrupt_key agy)" = Escape ] || fail "agy must interrupt on Escape" + [ "$(fm_control_interrupt_repeat agy)" = 1 ] || fail "agy must interrupt on a single press" + [ -z "$(fm_control_interrupt_clear_key agy)" ] || fail "agy must need no clear key" + [ "$(fm_control_interrupt_ack_source agy)" = none ] || fail "agy must have no ack source" + [ "$(fm_control_exit_command agy)" = /quit ] || fail "agy must exit on /quit" + pass "fm-control-lib: agy mechanics are Escape once, no clear key, and /quit" +} + +test_agy_busy_tail_needs_the_pinned_status_row() { + printf 'working\nesc to cancel\n' | fm_busy_agy_tail_busy \ + || fail "the esc-to-cancel status row must read busy" + printf 'working\n Generating...\n' | fm_busy_agy_tail_busy \ + && fail "the free-floating Generating word alone must not read busy" || true + printf 'Generating report...\ndone\n? for shortcuts\n>\n' | fm_busy_agy_tail_busy \ + && fail "echoed worker output naming Generating must not read busy" || true + printf 'idle\n? for shortcuts\n>\n' | fm_busy_agy_tail_busy \ + && fail "an idle footer must not read busy" || true + printf 'Generating report...\ndone\n? for shortcuts\n>\n' | fm_busy_lines_match agy \ + && fail "the delivery guard must not acknowledge on echoed Generating output" || true + FM_BUSY_AGY_REGEX='idle' bash -c '. "$0/bin/fm-busy-lib.sh"; printf "idle\n" | fm_busy_agy_tail_busy' "$ROOT" \ + && fail "an environment override must not change the agy busy signature" || true + pass "fm-busy-lib: only the pinned esc-to-cancel row carries the agy busy verdict" +} + +test_agy_busy_signatures_are_harness_scoped() { + printf 'esc to cancel\n' | fm_busy_lines_match agy \ + || fail "harness=agy must match its own esc token" + printf 'esc to cancel\n' | fm_busy_lines_match grok \ + && fail "harness=grok must never borrow agy's esc token" || true + printf 'Ctrl+c:cancel\n' | fm_busy_lines_match agy \ + && fail "harness=agy must never borrow grok's token" || true + printf 'esc to cancel\n' | fm_busy_lines_match kimi \ + && fail "harness=kimi must never borrow agy's token" || true + printf 'esc to cancel\n' | fm_busy_lines_match spaceship \ + && fail "an unverified harness must match nothing" || true + pass "fm-composer-lib: agy delivery signatures never cross harnesses" +} + +test_agy_classify_reports_unknown_when_the_marker_scrolls_out() { + local statedir busy idle + statedir="$TMP_ROOT/classify"; mkdir -p "$statedir" + busy=$(fm_busy_classify tmux fake:win agy agy-case-1 "$statedir" 'turn running +esc to cancel Gemini 3.8 Flash · low') + [ "$busy" = "busy agy-regex" ] || fail "a busy tail must classify busy agy-regex, got '$busy'" + idle=$(fm_busy_classify tmux fake:win agy agy-case-2 "$statedir" 'reply landed +? for shortcuts Gemini 3.8 Flash · low') + [ "$idle" = "unknown agy-regex" ] || fail "a scrolled-out marker must classify unknown, got '$idle'" + pass "fm-busy-lib: agy classifies busy on its marker and unknown without it" +} + +test_agy_tmux_names_the_native_binary_an_agent() { + local got + # shellcheck source=/dev/null + . "$ROOT/bin/fm-backend.sh" + fm_backend_source tmux || fail "fm_backend_source tmux failed" + got=$(fm_agent_process_classify_name agy) + [ "$got" = agent ] || fail "tmux liveness must read the agy binary as an agent, got '$got'" + got=$(fm_agent_process_classify_name magyk) + [ "$got" = other ] || fail "tmux liveness must not read magyk as an agent, got '$got'" + got=$(fm_agent_process_classify_name bash) + [ "$got" = shell ] || fail "tmux liveness must still read bash as a shell, got '$got'" + pass "bin/fm-agent-process-lib.sh: agy is an agent, fragments are not" +} + +# Canned `pane process-info` bodies for the herdr fixtures. The shared +# exit-detection contract proves a registered agent at process level before +# trusting it (bin/backends/herdr.sh fm_backend_herdr_pane_process_state), so +# every registered-status fixture pairs its `agent get` body with a process +# view. The agy-shaped body names the foreground process exactly `agy`, which +# is the same identity surface the tmux liveness probe and the ancestry +# detector use - no real agy process is needed because the foreground branch +# answers before the descendant walk touches the process table. +agy_herdr_process_info_body() { # <shell-pid> <foreground-name> -> JSON + printf '%s\n' "{\"result\":{\"type\":\"pane_process_info\",\"process_info\":{\"pane_id\":\"w9:p1\",\"shell_pid\":$1,\"foreground_processes\":[{\"pid\":$(( $1 + 1 )),\"name\":\"$2\",\"argv\":[\"$2\",\"--prompt-interactive\"],\"argv0\":\"$2\",\"cmdline\":\"$2 --prompt-interactive\"}]}}}" +} + +agy_herdr_agent_state() { # <fixture-dir> -> verdict; logs every CLI call + local dir=$1 + : > "$dir/calls.log" + AGY_FIX_RESP="$dir/agent-get.json" AGY_FIX_PROC="$dir/process-info.json" \ + AGY_FIX_LOG="$dir/calls.log" bash -c ' + . "$0/bin/backends/herdr.sh" + fm_backend_herdr_pane_presence_state() { printf "present"; } + fm_backend_herdr_cli() { + printf "%s\n" "$*" >> "$AGY_FIX_LOG" + case "$*" in + *"agent get"*) cat "$AGY_FIX_RESP" ;; + *"pane process-info"*) cat "$AGY_FIX_PROC" ;; + *) exit 0 ;; + esac + } + fm_backend_herdr_pane_agent_state testsession w9:p1' "$ROOT" 2>&1 +} + +test_herdr_done_with_live_registry_stays_live() { + local dir out + dir="$TMP_ROOT/herdr-done"; mkdir -p "$dir" + printf '%s\n' '{"result":{"agent":{"agent":"agy","agent_status":"done","pane_id":"w9:p1"}}}' > "$dir/agent-get.json" + agy_herdr_process_info_body 424242 agy > "$dir/process-info.json" + out=$(agy_herdr_agent_state "$dir") + [ "$out" = live ] || fail "a registered done status with an agy process view must stay live, got '$out'" + grep -q "process-info" "$dir/calls.log" \ + || fail "the shared contract proves a registered agent at process level; the verdict trusted the registration alone" + out=$(AGY_FIX_RESP="$dir/agent-get.json" AGY_FIX_PROC="$dir/process-info.json" AGY_FIX_LOG="$dir/calls.log" bash -c ' + . "$0/bin/backends/herdr.sh" + fm_backend_herdr_pane_presence_state() { printf "present"; } + fm_backend_herdr_cli() { + case "$*" in + *"agent get"*) cat "$AGY_FIX_RESP" ;; + *"pane process-info"*) cat "$AGY_FIX_PROC" ;; + *) exit 0 ;; + esac + } + fm_backend_herdr_tab_is_husk testsession w9:p1 && printf husk || printf refused' "$ROOT" 2>&1) + [ "$out" = refused ] || fail "a live pane must refuse husk replacement, got '$out'" + pass "herdr exit detection: done with a live registry and an agy process view stays live and refuses replacement" +} + +test_herdr_incomplete_shell_view_cannot_prove_absence() { + local dir out shell_pid + dir="$TMP_ROOT/herdr-stale"; mkdir -p "$dir" + # The descendant walk reads the REAL process table, so the canned pane shell + # must be a process this test owns and can prove alive: a short-lived sleep. + sleep 30 & shell_pid=$! + printf '%s\n' '{"result":{"agent":{"agent":"agy","agent_status":"done","pane_id":"w9:p1"}}}' > "$dir/agent-get.json" + agy_herdr_process_info_body "$shell_pid" bash > "$dir/process-info.json" + out=$(agy_herdr_agent_state "$dir") + kill "$shell_pid" 2>/dev/null || true + [ "$out" = live ] || fail "an incomplete shell view must not prove agent absence, got '$out'" + out=$(AGY_FIX_RESP="$dir/agent-get.json" AGY_FIX_PROC="$dir/process-info.json" AGY_FIX_LOG="$dir/calls.log" bash -c ' + . "$0/bin/backends/herdr.sh" + fm_backend_herdr_pane_presence_state() { printf "present"; } + fm_backend_herdr_cli() { + case "$*" in + *"agent get"*) cat "$AGY_FIX_RESP" ;; + *"pane process-info"*) cat "$AGY_FIX_PROC" ;; + *) exit 0 ;; + esac + } + fm_backend_herdr_tab_is_husk testsession w9:p1 && printf husk || printf refused' "$ROOT" 2>&1) + [ "$out" = refused ] || fail "a stale registration must still refuse husk replacement, got '$out'" + pass "herdr exit detection: incomplete shell evidence retains conservative liveness and refuses closing" +} + +test_herdr_shell_first_with_live_registry_stays_live() { + local dir out + dir="$TMP_ROOT/herdr-idle"; mkdir -p "$dir" + printf '%s\n' '{"result":{"agent":{"agent":"agy","agent_status":"idle","pane_id":"w9:p1"}}}' > "$dir/agent-get.json" + # The pane shell is present in the process view too (shell_pid), but the + # foreground names agy: the verified harness identity outranks shell-first + # ranking, and the shared contract's process proof is satisfied. + agy_herdr_process_info_body 424242 agy > "$dir/process-info.json" + out=$(agy_herdr_agent_state "$dir") + [ "$out" = live ] || fail "a registered idle status with an agy foreground must stay live, got '$out'" + grep -q "process-info" "$dir/calls.log" \ + || fail "the shared contract proves a registered agent at process level; the verdict trusted the registration alone" + pass "herdr exit detection: a registered pane with an agy foreground stays live however its shell ranks" +} + +test_herdr_lone_unregistered_pane_is_agent_free() { + local dir out + dir="$TMP_ROOT/herdr-gone"; mkdir -p "$dir" + printf '%s\n' '{"error":{"code":"agent_not_found","message":"agent target w9:p1 not found"}}' > "$dir/agent-get.json" + out=$(agy_herdr_agent_state "$dir") + [ "$out" = no-agent ] || fail "an unregistered pane must read no-agent, got '$out'" + out=$(AGY_FIX_RESP="$dir/agent-get.json" AGY_FIX_LOG="$dir/calls.log" bash -c ' + . "$0/bin/backends/herdr.sh" + fm_backend_herdr_pane_presence_state() { printf "present"; } + fm_backend_herdr_cli() { + case "$*" in *"agent get"*) cat "$AGY_FIX_RESP" ;; *) exit 0 ;; esac + } + fm_backend_herdr_tab_is_husk testsession w9:p1 && printf husk || printf refused' "$ROOT" 2>&1) + [ "$out" = husk ] || fail "an agent-free pane must allow husk replacement, got '$out'" + pass "herdr exit detection: only a positively unregistered pane is agent-free" +} + +test_herdr_malformed_and_failed_reads_stay_unknown() { + local dir out + dir="$TMP_ROOT/herdr-malformed"; mkdir -p "$dir" + printf '%s\n' '{not json at all' > "$dir/agent-get.json" + out=$(agy_herdr_agent_state "$dir") + [ "$out" = unknown ] || fail "a malformed registry response must read unknown, got '$out'" + dir="$TMP_ROOT/herdr-failed"; mkdir -p "$dir" + printf '%s\n' '{"result":{}}' > "$dir/agent-get.json" + export AGY_FIX_FAIL=1 + out=$(AGY_FIX_RESP="$dir/agent-get.json" AGY_FIX_LOG="$dir/calls.log" bash -c ' + . "$0/bin/backends/herdr.sh" + fm_backend_herdr_pane_presence_state() { printf "present"; } + fm_backend_herdr_cli() { + printf "%s\n" "$*" >> "$AGY_FIX_LOG" + case "$*" in *"agent get"*) [ "${AGY_FIX_FAIL:-0}" = 1 ] && exit 3; cat "$AGY_FIX_RESP" ;; *) exit 0 ;; esac + } + fm_backend_herdr_pane_agent_state testsession w9:p1' "$ROOT" 2>&1) + unset AGY_FIX_FAIL + [ "$out" = unknown ] || fail "a failed registry query must read unknown, got '$out'" + pass "herdr exit detection: malformed and failed reads stay unknown" +} + +make_agy_trust_case() { # <name> -> "<case>|<proj>|<wt>|<home>" + local name=$1 case_dir proj wt home + case_dir="$TMP_ROOT/trust-$name" + proj="$case_dir/project" + wt="$case_dir/wt" + home="$case_dir/home" + mkdir -p "$home" + fm_git_worktree "$proj" "$wt" "wt-trust-$name" + printf '%s|%s|%s|%s\n' "$case_dir" "$proj" "$wt" "$home" +} + +read_agy_trust_case() { + IFS='|' read -r CASE_DIR PROJ_DIR WT_DIR HOME_DIR <<EOF +$1 +EOF +} + +run_agy_trust() { # <home> <worktree> <project> + HOME="$1" "$TRUST" "$2" "$3" 2>&1 +} + +test_agy_trust_registers_the_logical_and_resolved_worktree_paths() { + local rec store out link + rec=$(make_agy_trust_case fresh) + read_agy_trust_case "$rec" + store="$HOME_DIR/.gemini/antigravity-cli/settings.json" + mkdir -p "$(dirname "$store")" + printf '%s\n' '{"model":"Gemini 3.8 Flash (High)","allowNonWorkspaceAccess":true,"trustedWorkspaces":["/home/someone/elsewhere"]}' > "$store" + link="$CASE_DIR/wt-link" + ln -s "$WT_DIR" "$link" + out=$(run_agy_trust "$HOME_DIR" "$link" "$PROJ_DIR") || fail "a fresh linked worktree must be trusted: $out" + assert_agy_trusted "$store" "$link" "the logical (symlinked) pane path agy compares against was not registered" + assert_agy_trusted "$store" "$WT_DIR" "the resolved worktree path was not registered alongside the logical one" + assert_agy_trusted "$store" "/home/someone/elsewhere" "registration dropped an existing trustedWorkspaces entry" + [ "$(agy_store_value "$store" model)" = '"Gemini 3.8 Flash (High)"' ] \ + || fail "registration did not preserve an unrelated store key" + [ "$(agy_store_value "$store" allowNonWorkspaceAccess)" = true ] \ + || fail "registration did not preserve an unrelated boolean key" + out=$(run_agy_trust "$HOME_DIR" "$link" "$PROJ_DIR") || fail "repeat registration must succeed: $out" + [ "$(agy_trusted_paths "$store" | grep -Fcx "$WT_DIR")" -eq 1 ] \ + || fail "repeat registration duplicated the worktree entry" + pass "fm-agy-trust.sh: registers the logical and resolved worktree paths and preserves the store" +} + +test_agy_trust_creates_a_missing_store() { + local rec store out + rec=$(make_agy_trust_case nostore) + read_agy_trust_case "$rec" + store="$HOME_DIR/.gemini/antigravity-cli/settings.json" + out=$(run_agy_trust "$HOME_DIR" "$WT_DIR" "$PROJ_DIR") || fail "a missing store must be created: $out" + [ -f "$store" ] || fail "no settings store was created at $store" + assert_agy_trusted "$store" "$WT_DIR" "the worktree was not registered in the created store" + pass "fm-agy-trust.sh: creates agy's settings store when none exists" +} + +test_agy_trust_refuses_out_of_scope_paths() { + local rec store out rc plain before after + rec=$(make_agy_trust_case scope) + read_agy_trust_case "$rec" + store="$HOME_DIR/.gemini/antigravity-cli/settings.json" + mkdir -p "$(dirname "$store")" + printf '%s\n' '{"trustedWorkspaces":[]}' > "$store" + rc=0; out=$(run_agy_trust "$HOME_DIR" "$PROJ_DIR" "$PROJ_DIR") || rc=$? + [ "$rc" -ne 0 ] || fail "the primary checkout must be refused" + assert_contains "$out" "primary checkout" "primary-checkout refusal lacked its reason" + assert_agy_not_trusted "$store" "$PROJ_DIR" "a refused primary checkout was still registered" + rc=0; out=$(run_agy_trust "$HOME_DIR" "$HOME_DIR" "$PROJ_DIR") || rc=$? + [ "$rc" -ne 0 ] || fail "the home directory must be refused" + assert_agy_not_trusted "$store" "$HOME_DIR" "a refused home directory was still registered" + plain="$CASE_DIR/plain"; mkdir -p "$plain" + rc=0; out=$(run_agy_trust "$HOME_DIR" "$plain" "$PROJ_DIR") || rc=$? + [ "$rc" -ne 0 ] || fail "a plain directory must be refused" + assert_agy_not_trusted "$store" "$plain" "a refused plain directory was still registered" + rc=0; out=$(run_agy_trust "$HOME_DIR" "$WT_DIR/.git" "$PROJ_DIR") || rc=$? + [ "$rc" -ne 0 ] || fail "a path below the worktree root must be refused" + printf '%s\n' '{not json' > "$store" + before=$(cat "$store") + rc=0; out=$(run_agy_trust "$HOME_DIR" "$WT_DIR" "$PROJ_DIR") || rc=$? + [ "$rc" -ne 0 ] || fail "an unparseable store must be refused" + after=$(cat "$store") + [ "$before" = "$after" ] || fail "an unparseable store was rewritten" + pass "fm-agy-trust.sh: refuses every out-of-scope path and never rewrites a broken store" +} + +# The fake tmux renders an agy-shaped screen that advances through +# launched -> (trust dialog ->) busy as the real spawn drives it, so the launch +# command, the pre-registration, the single Enter that answers a dialog, and +# the readiness gate are exercised through their real code paths. Whether the +# dialog renders is decided the way agy decides it: the pane path is looked up +# in the trustedWorkspaces array of the store the spawn just wrote. +# FM_FAKE_AGY_IGNORE_TRUST=1 models a vendor that stopped honouring the store; +# FM_FAKE_AGY_ASSUME_TRUSTED=1 models a pane that never shows the dialog even +# though firstmate could not register the path (a busy verdict with no proof +# of where the turn runs); +# FM_FAKE_AGY_RACE=1 models Herdr's native busy verdict rendering one capture +# before the dialog paints; FM_FAKE_AGY_ANSWER=stuck models a dialog whose +# answer never turns into a busy turn. +make_agy_fakebin() { + local dir=$1 fakebin + fakebin=$(fm_fakebin "$dir") + cat > "$fakebin/tmux" <<'SH' +#!/usr/bin/env bash +set -u +printf '%s\n' "$*" >> "$FM_FAKE_TMUX_CALL_LOG" +state=$(cat "$FM_FAKE_AGY_STATE" 2>/dev/null || true) +fake_screen() { + case "$state" in + dialog) + printf 'Accessing workspace:\n\n%s\n\nDo you trust the contents of this project?\n\nAntigravity CLI requires permission to read, edit, and execute files here.\n\n> Yes, I trust this folder\n No, exit\n' "$FM_FAKE_PANE_PATH" + ;; + busy) + printf 'Generating...\n└ Tip: press f to see the full diff.\n\nesc to cancel Gemini 3.8 Flash · low\n' + ;; + racing) + printf 'esc to cancel Gemini 3.8 Flash · low\n' + printf 'dialog\n' > "$FM_FAKE_AGY_STATE" + ;; + *) + printf 'shell starting\n$ \n' + ;; + esac +} +fake_path_trusted() { + [ "${FM_FAKE_AGY_ASSUME_TRUSTED:-0}" = 1 ] && return 0 + [ "${FM_FAKE_AGY_IGNORE_TRUST:-0}" = 1 ] && return 1 + node -e 'const fs=require("node:fs");let j={};try{j=JSON.parse(fs.readFileSync(process.argv[1],"utf8"));}catch(e){process.exit(1);}process.exit(Array.isArray(j.trustedWorkspaces)&&j.trustedWorkspaces.includes(process.argv[2])?0:1);' \ + "$FM_FAKE_AGY_SETTINGS" "$FM_FAKE_PANE_PATH" +} +case "$*" in + *"#{pane_current_path}"*) printf '%s\n' "$FM_FAKE_PANE_PATH"; exit 0 ;; + *"#{cursor_y}"*) printf '1\n'; exit 0 ;; +esac +case "${1:-}" in + display-message) printf 'firstmate\n'; exit 0 ;; + list-windows) exit 0 ;; + has-session|new-session|new-window|kill-window) exit 0 ;; + send-keys) + literal= + prev= + for arg in "$@"; do + if [ "$prev" = -l ]; then literal=$arg; break; fi + prev=$arg + done + if [ -n "$literal" ]; then + case "$literal" in + *--prompt-interactive*) + printf '%s\n' "$literal" >> "$FM_FAKE_LAUNCH_LOG" + printf 'launched\n' > "$FM_FAKE_AGY_STATE" + ;; + esac + exit 0 + fi + case " $* " in + *' Enter '*) + case "$state" in + launched) + if fake_path_trusted; then + printf 'busy\n' > "$FM_FAKE_AGY_STATE" + elif [ "${FM_FAKE_AGY_RACE:-0}" = 1 ]; then + printf 'racing\n' > "$FM_FAKE_AGY_STATE" + else + printf 'dialog\n' > "$FM_FAKE_AGY_STATE" + fi + ;; + dialog) + if [ "${FM_FAKE_AGY_ANSWER:-works}" = works ]; then + printf 'busy\n' > "$FM_FAKE_AGY_STATE" + fi + ;; + esac + ;; + esac + exit 0 + ;; + capture-pane) fake_screen; exit 0 ;; +esac +exit 0 +SH + chmod +x "$fakebin/tmux" + cat > "$fakebin/agy" <<'SH' +#!/usr/bin/env bash +set -u +if [ "${1:-}" = models ]; then + if [ "${FM_FAKE_AGY_MODELS_FAIL:-0}" = 1 ]; then exit 3; fi + if [ "${FM_FAKE_AGY_MODELS_HANG:-0}" = 1 ]; then cat > /dev/null; sleep 30; exit 0; fi + printf 'gemini-3.8-flash-high\tGemini 3.8 Flash (High)\n' + printf 'gemini-3.8-flash-medium\tGemini 3.8 Flash (Medium)\n' + printf 'gemini-3.8-flash-low\tGemini 3.8 Flash (Low)\n' + exit 0 +fi +echo "fake agy must never execute" >&2 +exit 9 +SH + chmod +x "$fakebin/agy" + fm_fake_exit0 "$fakebin" treehouse gh-axi gh herdr + printf '%s\n' "$fakebin" +} + +make_agy_spawn_case() { + local name=$1 id=$2 case_dir home proj wt fakebin + case_dir="$TMP_ROOT/$name" + home="$case_dir/home" + proj="$case_dir/project" + wt="$case_dir/wt" + fakebin=$(make_agy_fakebin "$case_dir/fake") + mkdir -p "$home/data/$id" "$home/projects" "$home/state" "$home/config" + cat > "$home/data/$id/brief.md" <<'EOF' +# Task +## Captain's intent +Exercise Antigravity dispatch. + +## Firstmate spec +Verify launch and delivery behavior. +EOF + printf 'agy\n' > "$home/config/crew-harness" + mkdir -p "$home/.gemini/antigravity-cli" + printf '%s\n' '{"model":"Gemini 3.8 Flash (High)","trustedWorkspaces":["/home/someone/elsewhere"]}' \ + > "$home/.gemini/antigravity-cli/settings.json" + fm_git_worktree "$proj" "$wt" "wt-$name" + touch "$home/state/.last-watcher-beat" + : > "$case_dir/launch.log" + : > "$case_dir/tmux-calls.log" + : > "$case_dir/agy.state" + printf '%s\n' "$case_dir|$home|$proj|$wt|$fakebin" +} + +read_agy_spawn_record() { + IFS='|' read -r CASE_DIR HOME_DIR PROJ_DIR WT_DIR FAKEBIN_DIR <<EOF +$1 +EOF +} + +# The spawn drives the real bin/fm-agy-trust.sh and the fake tmux's trust +# lookup under this base PATH, and both read agy's settings store with node, +# which runners do not keep in the system bin dirs. Carry the directory the +# invoking environment resolves node from, the fm-kimi-harness shape. +NODE_BIN=$(command -v node) || fail "test needs node" +NODE_BIN_DIR=$(dirname "$NODE_BIN") +BASE_PATH=${FM_TEST_BASE_PATH:-$NODE_BIN_DIR:/usr/bin:/bin:/usr/sbin:/sbin} + +run_agy_spawn() { + local case_dir=$1 home=$2 proj=$3 wt=$4 fakebin=$5 id=$6 + shift 6 + HOME="$home" FM_ROOT_OVERRIDE='' FM_HOME="$home" \ + FM_STATE_OVERRIDE="$home/state" FM_DATA_OVERRIDE="$home/data" \ + FM_PROJECTS_OVERRIDE="$home/projects" FM_CONFIG_OVERRIDE="$home/config" \ + FM_SPAWN_NO_GUARD=1 FM_FAKE_PANE_PATH="$wt" TMUX="fake,1,0" \ + FM_FAKE_LAUNCH_LOG="$case_dir/launch.log" \ + FM_FAKE_TMUX_CALL_LOG="$case_dir/tmux-calls.log" \ + FM_FAKE_AGY_STATE="$case_dir/agy.state" \ + FM_FAKE_AGY_SETTINGS="$home/.gemini/antigravity-cli/settings.json" \ + FM_FAKE_AGY_MODELS_FAIL="${FM_FAKE_AGY_MODELS_FAIL:-0}" \ + FM_FAKE_AGY_MODELS_HANG="${FM_FAKE_AGY_MODELS_HANG:-0}" \ + FM_FAKE_AGY_IGNORE_TRUST="${FM_FAKE_AGY_IGNORE_TRUST:-0}" \ + FM_FAKE_AGY_ASSUME_TRUSTED="${FM_FAKE_AGY_ASSUME_TRUSTED:-0}" \ + FM_FAKE_AGY_RACE="${FM_FAKE_AGY_RACE:-0}" \ + FM_FAKE_AGY_ANSWER="${FM_FAKE_AGY_ANSWER:-works}" \ + FM_AGY_READY_POLLS=4 FM_AGY_POLL_INTERVAL=0 FM_AGY_MODELS_TIMEOUT=${FM_AGY_MODELS_TIMEOUT:-1} \ + PATH="$fakebin:$BASE_PATH" \ + "$SPAWN" "$id" "$proj" --backend herdr --harness agy --mode no-mistakes --yolo off "$@" 2>&1 +} + +test_agy_launch_carries_the_brief_with_model_effort_and_autonomy() { + local id rec out rc launch meta + id="agy-launch-z1-$$" + rec=$(make_agy_spawn_case launch "$id") + read_agy_spawn_record "$rec" + out=$(run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$id" \ + --model gemini-3.8-flash-low --effort low) + rc=$? + expect_code 0 "$rc" "agy spawn with a listed model should succeed: $out" + launch=$(cat "$CASE_DIR/launch.log") + assert_contains "$launch" "$FAKEBIN_DIR/agy" "agy launch did not pin the resolved absolute binary" + assert_contains "$launch" "--prompt-interactive" "agy launch did not carry the brief via --prompt-interactive" + assert_contains "$launch" "--model 'gemini-3.8-flash-low'" "agy launch did not carry the requested model" + assert_contains "$launch" "--effort 'low'" "agy launch did not carry the requested effort" + assert_contains "$launch" "--dangerously-skip-permissions" "agy launch omitted unattended autonomy" + assert_contains "$launch" "env -u CLAUDECODE" "agy launch did not clear the inherited launcher marker" + assert_not_contains "$launch" "__AGYBIN__" "agy launch left its binary placeholder unsubstituted" + assert_not_contains "$launch" "__MODELFLAG__" "agy launch left its model placeholder unsubstituted" + assert_not_contains "$launch" "__BRIEF__" "agy launch left its brief placeholder unsubstituted" + meta="$HOME_DIR/state/$id.meta" + assert_grep 'harness=agy' "$meta" "agy meta did not record its harness" + assert_grep 'model=gemini-3.8-flash-low' "$meta" "agy meta did not record its model" + assert_grep 'effort=low' "$meta" "agy meta did not record its effort" + pass "fm-spawn: agy launch carries brief, model, effort, and autonomy with cleared markers" +} + +test_agy_effort_xhigh_is_recorded_but_omitted() { + local id rec out rc launch meta + id="agy-xhigh-z2-$$" + rec=$(make_agy_spawn_case xhigh "$id") + read_agy_spawn_record "$rec" + out=$(run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$id" \ + --model gemini-3.8-flash-low --effort xhigh) + rc=$? + expect_code 0 "$rc" "agy spawn with an unsupported effort should still succeed" + launch=$(cat "$CASE_DIR/launch.log") + assert_not_contains "$launch" "--effort" "agy launch passed a known-bad effort value" + meta="$HOME_DIR/state/$id.meta" + assert_grep 'effort=xhigh' "$meta" "agy meta did not retain the unsupported effort axis" + pass "fm-spawn: agy omits xhigh from the launch but records it in task metadata" +} + +test_agy_unlisted_model_refuses_before_pane_creation() { + local id rec out rc + id="agy-badmodel-z3-$$" + rec=$(make_agy_spawn_case badmodel "$id") + read_agy_spawn_record "$rec" + rc=0 + out=$(run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$id" \ + --model gemini-3.8-flash) || rc=$? + [ "$rc" -ne 0 ] || fail "an unlisted agy model should refuse the spawn" + assert_contains "$out" "not listed by 'agy models'" "unlisted model refusal lacked its concrete reason" + [ -s "$CASE_DIR/launch.log" ] && fail "an unlisted model created a launch command" || true + pass "fm-spawn: an unlisted agy model refuses before pane creation" +} + +test_agy_unreachable_listing_launches_unvalidated() { + local id rec out rc + id="agy-nolisting-z4-$$" + rec=$(make_agy_spawn_case nolisting "$id") + read_agy_spawn_record "$rec" + rc=0 + out=$(FM_FAKE_AGY_MODELS_FAIL=1 run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" \ + "$FAKEBIN_DIR" "$id" --model gemini-3.8-flash-low) || rc=$? + expect_code 0 "$rc" "an unreachable model listing must not block the spawn" + [ -s "$CASE_DIR/launch.log" ] || fail "an unreachable listing produced no launch command" + assert_contains "$out" "listing is unreachable" "an unreachable listing launched without its notice" + pass "fm-spawn: an unreachable agy listing establishes nothing and launches" +} + +test_agy_hung_listing_is_cut_off_and_launches() { + local id rec out rc started elapsed + id="agy-hanglisting-z8-$$" + rec=$(make_agy_spawn_case hanglisting "$id") + read_agy_spawn_record "$rec" + rc=0 + started=$(date +%s) + out=$(FM_FAKE_AGY_MODELS_HANG=1 run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" \ + "$FAKEBIN_DIR" "$id" --model gemini-3.8-flash-low) || rc=$? + elapsed=$(( $(date +%s) - started )) + expect_code 0 "$rc" "a hung model listing must not block the spawn" + [ "$elapsed" -lt 20 ] || fail "the model probe was not cut off by its bound (took ${elapsed}s)" + assert_contains "$out" "did not answer within 1s" "a hung listing launched without its timeout notice" + [ -s "$CASE_DIR/launch.log" ] || fail "a hung listing produced no launch command" + assert_contains "$(cat "$CASE_DIR/launch.log")" "--model 'gemini-3.8-flash-low'" \ + "a hung listing dropped the requested model instead of launching it unvalidated" + pass "fm-spawn: a hung agy listing is cut off by the shared bound and launches unvalidated" +} + +test_agy_zero_model_timeout_is_clamped_to_the_default_bound() { + local id rec out rc started elapsed + id="agy-zerobound-z14-$$" + rec=$(make_agy_spawn_case zerobound "$id") + read_agy_spawn_record "$rec" + rc=0 + started=$(date +%s) + out=$(FM_FAKE_AGY_MODELS_HANG=1 FM_AGY_MODELS_TIMEOUT=0 \ + run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" \ + "$FAKEBIN_DIR" "$id" --model gemini-3.8-flash-low) || rc=$? + elapsed=$(( $(date +%s) - started )) + expect_code 0 "$rc" "a hung listing with a zero bound must not block the spawn" + [ "$elapsed" -lt 25 ] || fail "a zero model bound disabled the deadline (took ${elapsed}s)" + assert_contains "$out" "did not answer within 15s" \ + "a zero model bound was not clamped to the documented default" + [ -s "$CASE_DIR/launch.log" ] || fail "a zero model bound produced no launch command" + pass "fm-spawn: a zero FM_AGY_MODELS_TIMEOUT is clamped to the default bound" +} + +# Bare Enter key presses only: shell setup rides its Enter on the typed text +# (`send-keys -t <target> export X=Y Enter`), while the launch submit and the +# trust-dialog answer are lone key sends (`send-keys -t <target> Enter`). +count_enter_sends() { # <tmux-call-log> + grep -c '^send-keys -t [^ ]* Enter$' "$1" || true +} + +test_agy_fresh_worktree_is_pre_trusted_and_launches_without_a_dialog() { + local id rec out rc enters store + id="agy-trust-z9-$$" + rec=$(make_agy_spawn_case trust "$id") + read_agy_spawn_record "$rec" + store="$HOME_DIR/.gemini/antigravity-cli/settings.json" + out=$(run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$id" \ + --model gemini-3.8-flash-low) + rc=$? + expect_code 0 "$rc" "an agy spawn into a fresh worktree should succeed" + assert_contains "$out" "spawned $id harness=agy" "agy spawn did not report success" + assert_not_contains "$out" "could not pre-register" "a legitimate worktree failed trust pre-registration" + assert_agy_trusted "$store" "$WT_DIR" "the spawn did not pre-register the worktree in agy's trust store" + assert_agy_trusted "$store" "/home/someone/elsewhere" "the spawn dropped an existing trustedWorkspaces entry" + [ "$(agy_store_value "$store" model)" = '"Gemini 3.8 Flash (High)"' ] \ + || fail "the spawn did not preserve an unrelated agy setting" + [ "$(cat "$CASE_DIR/agy.state")" = busy ] \ + || fail "the spawn reported success before the pane reached a busy turn (state: $(cat "$CASE_DIR/agy.state"))" + enters=$(count_enter_sends "$CASE_DIR/tmux-calls.log") + [ "$enters" -eq 1 ] \ + || fail "a pre-trusted worktree must receive only the launch Enter, got $enters Enter sends" + assert_not_contains "$(cat "$CASE_DIR/tmux-calls.log")" "kill-window" \ + "a successful agy spawn must never tear down the endpoint it just launched" + pass "fm-spawn: agy pre-registers the worktree and launches straight into a busy turn" +} + +test_agy_dialog_despite_registration_is_answered_once() { + local id rec out rc enters + id="agy-vendor-z10-$$" + rec=$(make_agy_spawn_case vendor-dialog "$id") + read_agy_spawn_record "$rec" + out=$(FM_FAKE_AGY_IGNORE_TRUST=1 run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" \ + "$FAKEBIN_DIR" "$id" --model gemini-3.8-flash-low) + rc=$? + expect_code 0 "$rc" "an agy spawn whose dialog renders despite registration should succeed" + [ "$(cat "$CASE_DIR/agy.state")" = busy ] \ + || fail "the spawn reported success before the pane reached a busy turn (state: $(cat "$CASE_DIR/agy.state"))" + enters=$(count_enter_sends "$CASE_DIR/tmux-calls.log") + [ "$enters" -eq 2 ] \ + || fail "expected exactly one launch Enter plus one trust-dialog Enter, got $enters Enter sends" + pass "fm-spawn: agy answers a dialog that renders anyway exactly once, then confirms busy" +} + +test_agy_dialog_race_waits_for_native_working() { + local id rec out rc enters + id="agy-race-z11-$$" + rec=$(make_agy_spawn_case race "$id") + read_agy_spawn_record "$rec" + out=$(FM_FAKE_AGY_IGNORE_TRUST=1 FM_FAKE_AGY_RACE=1 run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" \ + "$FAKEBIN_DIR" "$id" --model gemini-3.8-flash-low) + rc=$? + expect_code 0 "$rc" "a delayed vendor dialog must be answered before native working is accepted: $out" + [ "$(cat "$CASE_DIR/agy.state")" = busy ] || fail "spawn completed before the dialog answer worked" + enters=$(count_enter_sends "$CASE_DIR/tmux-calls.log") + [ "$enters" -eq 2 ] || fail "delayed dialog must receive exactly one answer after launch, got $enters Enter sends" + pass "agy delayed dialog is answered before native working confirms readiness" +} + +test_agy_unregistered_path_without_a_dialog_fails_the_spawn() { + local id rec out rc store + id="agy-nodialog-z12-$$" + rec=$(make_agy_spawn_case nodialog "$id") + read_agy_spawn_record "$rec" + store="$HOME_DIR/.gemini/antigravity-cli/settings.json" + printf '%s\n' '{not json' > "$store" + rc=0 + out=$(FM_FAKE_AGY_ASSUME_TRUSTED=1 run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" \ + "$FAKEBIN_DIR" "$id" --model gemini-3.8-flash-low) || rc=$? + [ "$rc" -ne 0 ] || fail "a corrupt trust store must refuse even if a vendor would show no dialog" + assert_contains "$out" "could not grant agy workspace trust" "trust failure did not name the cause" + assert_not_contains "$out" "spawned $id" "unregistered workspace reported a successful spawn" + [ ! -s "$CASE_DIR/launch.log" ] || fail "trust failure launched an agent" + [ "$(cat "$store")" = '{not json' ] || fail "trust failure overwrote the broken store" + pass "agy trust failure refuses before launching, independently of vendor readiness" +} + +test_agy_pre_trusted_path_that_never_turns_busy_fails_the_spawn() { + local id rec out rc + id="agy-idle-z13-$$" + rec=$(make_agy_spawn_case idle "$id") + read_agy_spawn_record "$rec" + rc=0 + out=$(FM_FAKE_AGY_IGNORE_TRUST=1 FM_FAKE_AGY_ANSWER=stuck run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" \ + "$FAKEBIN_DIR" "$id" --model gemini-3.8-flash-low) || rc=$? + [ "$rc" -ne 0 ] || fail "a dialog that never turns into a busy turn must fail the spawn" + assert_contains "$out" "did not start processing its brief after the folder-trust dialog was answered" \ + "a stuck trust dialog failed without its concrete reason" + [ "$(count_enter_sends "$CASE_DIR/tmux-calls.log")" -eq 2 ] \ + || fail "the gate must answer the dialog exactly once and never hammer Enter" + assert_contains "$(cat "$CASE_DIR/tmux-calls.log")" "kill-window" \ + "a failed agy readiness gate left its launched endpoint running" + pass "fm-spawn: an agy dialog that never turns busy fails the spawn and closes the endpoint" +} + +test_agy_missing_binary_refuses_before_pane_creation() { + local id rec out rc + id="agy-missing-z5-$$" + rec=$(make_agy_spawn_case missing "$id") + read_agy_spawn_record "$rec" + rm "$FAKEBIN_DIR/agy" + rc=0 + out=$(run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$id") || rc=$? + [ "$rc" -ne 0 ] || fail "a missing agy executable should refuse the spawn" + assert_contains "$out" "agy executable not found on PATH" "missing agy diagnostic lacked its concrete reason" + [ -s "$CASE_DIR/launch.log" ] && fail "a missing agy executable created a launch command" || true + pass "fm-spawn: a missing agy executable refuses before pane creation" +} + +test_agy_secondmate_is_refused() { + local id rec out rc + id="agy-secondmate-z6-$$" + rec=$(make_agy_spawn_case secondmate-refuse "$id") + read_agy_spawn_record "$rec" + rc=0 + out=$(HOME="$HOME_DIR" FM_ROOT_OVERRIDE='' FM_HOME="$HOME_DIR" \ + FM_STATE_OVERRIDE="$HOME_DIR/state" FM_DATA_OVERRIDE="$HOME_DIR/data" \ + FM_PROJECTS_OVERRIDE="$HOME_DIR/projects" FM_CONFIG_OVERRIDE="$HOME_DIR/config" \ + FM_SPAWN_NO_GUARD=1 PATH="$FAKEBIN_DIR:$BASE_PATH" \ + "$SPAWN" "$id" --secondmate agy 2>&1) || rc=$? + [ "$rc" -ne 0 ] || fail "an agy secondmate spawn should be refused" + assert_contains "$out" "crew-only" \ + "agy secondmate refusal lacked its concrete reason" + pass "fm-spawn: agy cannot be launched as a secondmate" +} + +test_agy_spawn_arms_no_busy_wiring() { + local id rec out rc statedir + id="agy-nowiring-z7-$$" + rec=$(make_agy_spawn_case nowiring "$id") + read_agy_spawn_record "$rec" + out=$(run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$id" \ + --model gemini-3.8-flash-low) + rc=$? + expect_code 0 "$rc" "agy spawn should succeed" + statedir="$HOME_DIR/state" + [ -e "$statedir/$id.busy-gen" ] && fail "agy spawn armed a busy generation nothing could clear" || true + for sidecar in "$statedir/$id.agy-"*; do + [ -e "$sidecar" ] || continue + # Trust custody is retained for exact cleanup, independently of busy wiring. + [ "$sidecar" = "$statedir/$id.agy-trust" ] && continue + fail "agy spawn left an unexpected adapter sidecar behind: $sidecar" + done + pass "fm-spawn: agy arms no busy wiring and retains only trust custody" +} + +test_agy_ancestry_detects_the_native_command_name +test_agy_ancestry_rejects_unrelated_mentions +test_agy_claims_no_inherited_launcher_marker +test_agy_control_mechanics_are_the_verified_ones +test_agy_busy_tail_needs_the_pinned_status_row +test_agy_busy_signatures_are_harness_scoped +test_agy_classify_reports_unknown_when_the_marker_scrolls_out +test_agy_tmux_names_the_native_binary_an_agent +test_herdr_done_with_live_registry_stays_live +test_herdr_incomplete_shell_view_cannot_prove_absence +test_herdr_shell_first_with_live_registry_stays_live +test_herdr_lone_unregistered_pane_is_agent_free +test_herdr_malformed_and_failed_reads_stay_unknown +test_agy_launch_carries_the_brief_with_model_effort_and_autonomy +test_agy_effort_xhigh_is_recorded_but_omitted +test_agy_unlisted_model_refuses_before_pane_creation +test_agy_unreachable_listing_launches_unvalidated +test_agy_hung_listing_is_cut_off_and_launches +test_agy_zero_model_timeout_is_clamped_to_the_default_bound +test_agy_trust_registers_the_logical_and_resolved_worktree_paths +test_agy_trust_creates_a_missing_store +test_agy_trust_refuses_out_of_scope_paths +test_agy_fresh_worktree_is_pre_trusted_and_launches_without_a_dialog +test_agy_dialog_despite_registration_is_answered_once +test_agy_dialog_race_waits_for_native_working +test_agy_unregistered_path_without_a_dialog_fails_the_spawn +test_agy_pre_trusted_path_that_never_turns_busy_fails_the_spawn +test_agy_missing_binary_refuses_before_pane_creation +test_agy_secondmate_is_refused +test_agy_spawn_arms_no_busy_wiring +printf '\nall fm-agy-harness tests passed\n' diff --git a/tests/fm-agy-signals-live-e2e.test.sh b/tests/fm-agy-signals-live-e2e.test.sh new file mode 100755 index 00000000000..59957fe793d --- /dev/null +++ b/tests/fm-agy-signals-live-e2e.test.sh @@ -0,0 +1,193 @@ +#!/usr/bin/env bash +# Live drift guard for the Antigravity CLI adapter's vendor-controlled surface: +# process name, trust dialog, rendered busy/interrupt/exit behavior. +# Opt-in because it submits real prompts (no echo provider exists for agy). +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" +AGY_BIN=$(command -v agy 2>/dev/null || true) +REAL_TMUX=$(command -v tmux 2>/dev/null || true) +LAB= +SOCKET="fm-agy-signals-$$" +SESSION=agy-signals +TARGET="$SESSION:agy" + +cleanup() { + [ -n "$REAL_TMUX" ] && "$REAL_TMUX" -L "$SOCKET" kill-server >/dev/null 2>&1 || true + [ -z "$LAB" ] || rm -rf -- "$LAB" +} + +fail() { + printf 'not ok - %s\n' "$1" >&2 + cleanup + exit 1 +} + +pass() { + printf 'ok - %s\n' "$1" +} + +fm_live_gate opt-in FM_AGY_SIGNALS_LIVE agy tmux +[ -n "$AGY_BIN" ] || fail "agy is not installed" + +LAB=$(mktemp -d "${TMPDIR:-/tmp}/fm-agy-signals.XXXXXX") || fail "could not create the isolated agy lab" +trap cleanup EXIT +mkdir -p "$LAB/workspace" +git -C "$LAB/workspace" init -q || fail "could not initialize the isolated agy workspace" +git -C "$LAB/workspace" config user.email "guard@local" || fail "could not configure the isolated agy workspace" +git -C "$LAB/workspace" config user.name "guard" || fail "could not configure the isolated agy workspace" +git -C "$LAB/workspace" commit -q --allow-empty -m init || fail "could not seed the isolated agy workspace" +WORKSPACE=$(cd "$LAB/workspace" && pwd -P) || fail "could not resolve the isolated agy workspace" + +# The worker runs under a throwaway HOME holding a copy of ~/.gemini (the +# method recorded in docs/verification/agy.md), so its trust answer and every +# other agy write land in the lab store, never the operator's real one. +AGY_HOME="$LAB/home" +mkdir -p "$AGY_HOME" || fail "could not create the throwaway agy HOME" +[ -d "$HOME/.gemini" ] || fail "no ~/.gemini to stage for the throwaway agy HOME" +cp -R "$HOME/.gemini" "$AGY_HOME/.gemini" || fail "could not stage the throwaway agy credential copy" + +# shellcheck source=/dev/null +. "$ROOT/bin/fm-busy-lib.sh" +# shellcheck source=/dev/null +. "$ROOT/bin/fm-composer-lib.sh" + +"$REAL_TMUX" -L "$SOCKET" new-session -d -s "$SESSION" -n control -c "$WORKSPACE" \ + || fail "could not start the isolated tmux server" +"$REAL_TMUX" -L "$SOCKET" new-window -d -t "$SESSION:" -n agy -c "$WORKSPACE" \ + || fail "could not open the isolated agy window" + +capture() { + "$REAL_TMUX" -L "$SOCKET" capture-pane -p -t "$TARGET" -S -100 2>/dev/null || true +} + +# The launch prompt asks for a computed answer (12345+67890=80235) so the +# awaited token never appears in the echoed launch line itself, where a plain +# reply token would false-positive on the shell echo (including across tmux +# wrapped rows). +"$REAL_TMUX" -L "$SOCKET" send-keys -t "$TARGET" -l \ + "HOME=\"$AGY_HOME\" $AGY_BIN --prompt-interactive \"Add 12345 and 67890. Reply with exactly the sum and nothing else\" --model gemini-3.8-flash-low --effort low --dangerously-skip-permissions" \ + || fail "could not type the agy launch line" +"$REAL_TMUX" -L "$SOCKET" send-keys -t "$TARGET" Enter \ + || fail "could not submit the agy launch line" + +# A fresh workspace stops on the folder-trust dialog. Answer the preselected +# safe choice once it renders. The answer appends the workspace to +# trustedWorkspaces in the throwaway HOME's copy of the agy settings store. +screen= +for _ in $(seq 1 150); do + screen=$(capture) + case "$screen" in + *"Do you trust the contents of this project?"*|*80235*|*80,235*) break ;; + esac + sleep 0.5 +done +case "$screen" in + *"Do you trust the contents of this project?"*) + "$REAL_TMUX" -L "$SOCKET" send-keys -t "$TARGET" Enter \ + || fail "could not answer the agy trust dialog" + ;; +esac + +# The initial turn executes and its reply lands; the busy footer must render +# while it is in flight so the portable matcher has live text to prove. +# Trivial turns were observed taking one to two minutes (cold start plus model +# latency), so these windows are generous; the guard is opt-in. +busy_live= +for _ in $(seq 1 240); do + screen=$(capture) + if printf '%s' "$screen" | fm_busy_agy_tail_busy; then busy_live=1; break; fi + case "$screen" in *80235*|*80,235*) break ;; esac + sleep 1 +done +[ -n "$busy_live" ] || fail "fm_busy_agy_tail_busy never matched the real agy turn in flight" +pass "the real agy busy footer matches fm_busy_agy_tail_busy in flight" + +for _ in $(seq 1 480); do + screen=$(capture) + case "$screen" in *80235*|*80,235*) break ;; esac + sleep 0.5 +done +reply=$(capture) +case "$reply" in + *80235*|*80,235*) pass "the real agy worker processed its launch prompt" ;; + *) fail "the real agy worker never answered its launch prompt" ;; +esac +# The reply can render while the turn is still finishing: the busy footer stays +# pinned until the idle composer replaces it, so wait for the settled idle row +# before asserting what the settled pane must not match. The wait itself +# refreshes $screen: the reply-wait loop above can legitimately break on a +# frame that still carries the pinned busy footer, and asserting on that stale +# frame would fail every run whose reply lands mid-turn. +idle_settled= +for _ in $(seq 1 120); do + screen=$(capture) + case "$screen" in *"? for shortcuts"*) idle_settled=1; break ;; esac + sleep 0.5 +done +[ -n "$idle_settled" ] || fail "the agy composer never settled to its idle footer after the reply" +# Scope to the visible tail the same way the owners do: mid-turn busy rows stay +# in scrollback after the turn settles and must not count as still busy. +printf '%s' "$screen" | grep -v '^[[:space:]]*$' | tail -12 | fm_busy_lines_match agy \ + && fail "harness=agy matched its own idle footer as busy" || true +printf '%s' "$screen" | fm_busy_agy_tail_busy \ + && fail "the settled agy footer still matches the busy signature" || true + +# The dialog can outlive the turn it gated, so a still-rendered dialog must be +# dismissed before steering anything: typed text would land in it instead of +# the composer. +if case "$(capture)" in *"Do you trust the contents of this project?"*) true ;; *) false ;; esac; then + "$REAL_TMUX" -L "$SOCKET" send-keys -t "$TARGET" Enter \ + || fail "could not dismiss the residual agy trust dialog" + idle= + for _ in $(seq 1 120); do + case "$(capture)" in *"? for shortcuts"*) idle=1; break ;; esac + sleep 0.5 + done + [ -n "$idle" ] || fail "the agy composer never went idle after the trust answer" +fi + +# Interrupt a genuinely long turn: poll until busy is observed, then send +# exactly one Escape and wait only for the Interrupted row it prints; a busy +# footer that merely disappears is not cancellation and no further Escape is +# sent, so a turn that survives one Escape fails this guard. +"$REAL_TMUX" -L "$SOCKET" send-keys -t "$TARGET" -l \ + "Write a 1500-word essay on the history of glass" \ + || fail "could not type the long agy prompt" +"$REAL_TMUX" -L "$SOCKET" send-keys -t "$TARGET" Enter \ + || fail "could not submit the long agy prompt" +for _ in $(seq 1 100); do + screen=$(capture) + printf '%s' "$screen" | fm_busy_agy_tail_busy && break + sleep 0.5 +done +printf '%s' "$screen" | fm_busy_agy_tail_busy \ + || fail "the long agy turn never showed its busy footer" +"$REAL_TMUX" -L "$SOCKET" send-keys -t "$TARGET" Escape \ + || fail "could not send Escape to the real agy turn" +cancelled= +for _ in $(seq 1 120); do + screen=$(capture) + case "$screen" in *Interrupted*) cancelled=1; break ;; esac + sleep 0.5 +done +[ -n "$cancelled" ] || fail "a single Escape never cancelled the real agy turn" +pass "a single Escape cancels the real agy turn" + +"$REAL_TMUX" -L "$SOCKET" send-keys -t "$TARGET" -l "/quit" \ + || fail "could not type the agy exit command" +"$REAL_TMUX" -L "$SOCKET" send-keys -t "$TARGET" Enter \ + || fail "could not submit the agy exit command" +gone= +for _ in $(seq 1 60); do + current=$("$REAL_TMUX" -L "$SOCKET" display-message -p -t "$TARGET" '#{pane_current_command}' 2>/dev/null || true) + case "$current" in *agy*) sleep 0.5 ;; *) gone=1; break ;; esac +done +[ -n "$gone" ] || fail "/quit never stopped the real agy process" +pass "/quit stops the real agy process" + +cleanup +trap - EXIT diff --git a/tests/fm-agy-smoke.test.sh b/tests/fm-agy-smoke.test.sh index 4fe3c4f7246..ff54eac7aca 100755 --- a/tests/fm-agy-smoke.test.sh +++ b/tests/fm-agy-smoke.test.sh @@ -25,9 +25,12 @@ if [ "${FM_CURSOR_AGY_LIVE_E2E:-0}" != 1 ]; then fi ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" -LAB="$ROOT/bin/fm-herdr-lab.sh" +LAB=${HERDR_LAB_HELPER:-$ROOT/bin/fm-herdr-lab.sh} +# shellcheck source=tests/herdr-test-safety.sh +. "$ROOT/tests/herdr-test-safety.sh" +herdr_forget_inherited_pane -fail() { printf 'not ok - %s\n' "$1" >&2; cleanup; exit 1; } +fail() { printf 'not ok - %s\n' "$1" >&2; exit 1; } pass() { printf 'ok - %s\n' "$1"; } for t in herdr agy jq; do @@ -47,26 +50,40 @@ SESSION=$("$LAB" name agy-smoke) || { echo "skip: could not derive a lab session WORK= AGY_TRUST_CREATED=0 FM_AGY_TRUST_ADDED= +CLEANUP_DONE=0 +CLEANUP_STATUS=0 cleanup() { + [ "$CLEANUP_DONE" = 0 ] || return "$CLEANUP_STATUS" + CLEANUP_DONE=1 if { [ "$AGY_TRUST_CREATED" = 1 ] || [ "${FM_AGY_TRUST_ADDED:-}" = created ]; } && [ -n "$WORK" ]; then if fm_agy_trust_remove "$WORK" >/dev/null 2>&1; then AGY_TRUST_CREATED=0 FM_AGY_TRUST_ADDED= else printf 'not ok - agy workspace-trust cleanup for %s\n' "$WORK" >&2 + CLEANUP_STATUS=1 fi fi [ -z "$WORK" ] || rm -rf "$WORK" - "$LAB" teardown "$SESSION" >/dev/null 2>&1 || printf 'not ok - lab teardown for %s\n' "$SESSION" >&2 + "$LAB" teardown "$SESSION" || CLEANUP_STATUS=1 + return "$CLEANUP_STATUS" } -trap cleanup EXIT +# shellcheck disable=SC2329 # EXIT trap callback. +smoke_exit() { + local exit_status=$? + cleanup || exit_status=1 + exit "$exit_status" +} +trap smoke_exit EXIT "$LAB" provision "$SESSION" >/dev/null 2>&1 || { echo "skip: could not provision the isolated Herdr lab session"; trap - EXIT; exit 0; } export HERDR_SESSION="$SESSION" fm_backend_herdr_agent_prompt_capability_check "$SESSION" >/dev/null 2>&1 \ || { echo "skip: isolated Herdr server lacks agy atomic prompt delivery"; exit 0; } WORK=$(mktemp -d) -printf 'Reply with exactly the single word PONG and nothing else.\n' > "$WORK/brief.md" +# Keep the real turn observable across native status polling; a one-word +# cached response can finish before Herdr publishes its first working state. +printf 'Run the shell command sleep 3 once, then reply with exactly the single word PONG and nothing else.\n' > "$WORK/brief.md" LAUNCH_TARGET= @@ -101,12 +118,27 @@ EOF agent= for _ in $(seq 1 40); do agent=$("$LAB" run "$SESSION" agent get "$pane" 2>/dev/null | jq -r '.result.agent.agent // empty' 2>/dev/null) - [ -z "$agent" ] || break + # Native identity appears before its lifecycle state during real startup. + # Keep the existing launch budget, but require the production liveness + # postcondition before leaving it. Retain an observed working edge so a + # short first turn cannot disappear between readiness and settle checks. + st=$(fm_backend_agent_status herdr "$ses:$pane" 2>/dev/null) + [ "$st" = working ] && saw_working=1 + if [ "$agent" = "$harness" ] \ + && [ "$(fm_backend_herdr_agent_alive "$ses:$pane")" = alive ]; then + break + fi sleep 1 done [ "$agent" = "$harness" ] || { echo "native agent get reported '$agent', expected '$harness'" >&2; return 1; } [ "$(fm_backend_herdr_agent_alive "$ses:$pane")" = alive ] \ - || { echo "$harness pane not reported alive by the generic liveness probe" >&2; return 1; } + || { + echo "$harness pane not reported alive by the generic liveness probe for $ses:$pane" >&2 + "$LAB" run "$SESSION" pane process-info --pane "$pane" >&2 + "$LAB" run "$SESSION" agent get "$pane" >&2 + fm_backend_herdr_agent_state "$ses:$pane" >&2 + return 1 + } # Observe a working->idle/done transition: the crew picks up the brief, works, # then settles (the exact edge the native completion detector keys on). for _ in $(seq 1 60); do @@ -124,7 +156,7 @@ EOF case "$harness" in agy) token=FIRSTMATE_AGY_STEER_ACCEPTED - token_prompt="Reply with exactly FIRSTMATE_AGY_ followed immediately by STEER_ACCEPTED and nothing else." + token_prompt="Run the shell command sleep 3 once, then reply with exactly FIRSTMATE_AGY_ followed immediately by STEER_ACCEPTED and nothing else." ;; esac verdict=$(fm_backend_herdr_prompt_submit "$ses:$pane" \ @@ -140,8 +172,7 @@ EOF sleep 1 done [ "$settled" = 1 ] || { echo "$harness did not start and settle a new turn after its confirmed steer" >&2; return 1; } - readback=$(fm_backend_herdr_cli "$ses" agent read "$pane" --source recent-unwrapped --lines 120 2>/dev/null \ - | jq -r '.result.read.text // empty' 2>/dev/null) \ + readback=$(fm_backend_herdr_capture "$ses:$pane" 120) \ || { echo "$harness post-steer output could not be read" >&2; return 1; } case "$readback" in *"$token"*) : ;; @@ -211,4 +242,5 @@ if jq -e --arg p "$WORK" '(.trustedWorkspaces // []) | index($p)' "$GLOBAL_SETTI fi pass "agy workspace trust is seeded before launch and removed afterward" -echo "# all fm-agy-smoke tests passed" +cleanup || fail "agy smoke cleanup failed" +printf '\nall fm-agy-smoke tests passed\n' diff --git a/tests/fm-backend-herdr.test.sh b/tests/fm-backend-herdr.test.sh index 8697a3f2f07..ca69fc77e64 100755 --- a/tests/fm-backend-herdr.test.sh +++ b/tests/fm-backend-herdr.test.sh @@ -419,6 +419,315 @@ test_recovery_grade_read_widens_only_at_its_own_boundary() { pass "herdr recovery-grade read: a stopped server means missing there, and nowhere else" } +# --- stale agent registration over a shell-only pane (issue #4115) ----------- +# +# Herdr keeps a Pi registration (`agent get` -> agent=pi, agent_status=idle) +# after the Pi process has exited to a plain shell whenever a nested interactive +# shell sits under the pane's top shell (the `treehouse get` crew shape; +# reproduced on Herdr 0.9.0 - docs/verification/runtime-backends.md "Stale agent +# registration"). Trusting that registration alone classified the pane `live`, +# so every relaunch and recovery was refused forever. The classifier must now +# prove an agent at process level before reporting one, exactly as the tmux +# adapter does, and a registration with no live agent process is agent-free +# with an explicit reason. +# +# The fixture pairs a canned `pane process-info` body with REAL processes: +# the shell pid it names is a real process this test owns, so the descendant +# walk runs against the real operating-system process table. + +# These upstream fixtures use a sleeper as the process-tree root while the +# provider body describes a shell. Make that simulated identity explicit in +# the process table too, retaining every real descendant and its real name. +# The fork proof tests below supply complete native shell/TTY snapshots. +install_stale_shell_ps() { # <directory> <process-info-json> + local dir=$1 shell_pid + shell_pid=$(printf '%s' "$2" | jq -r '.result.process_info.shell_pid // 0' 2>/dev/null) || shell_pid=0 + cat > "$dir/ps-overlay" <<'SH' +#!/usr/bin/env bash +ps -axo pid=,ppid=,stat=,comm=,tty= | awk -v root="$(cat "$(dirname "$0")/shell-pid")" ' +{ pid[NR]=$1; parent[NR]=$2; state[NR]=$3; name[NR]=$4; tty[NR]=$5 } +END { + owned[root]=1; changed=1 + while(changed) { changed=0; for(i=1;i<=NR;i++) if(parent[i] in owned && !(pid[i] in owned)) { owned[pid[i]]=1; changed=1 } } + for(i=1;i<=NR;i++) { + if(pid[i]==root) name[i]="zsh" + if(pid[i] in owned) tty[i]="fmfixture" + print pid[i],parent[i],state[i],name[i],tty[i] + } +}' +SH + printf '%s\n' "$shell_pid" > "$dir/shell-pid" + chmod +x "$dir/ps-overlay" +} + +stale_registration_case() { # <dir-suffix> <agent_status> <process-info-body|-> [process-info-exit] + local dir="$TMP_ROOT/stale-reg-$1" resp log fb n + mkdir -p "$dir/responses"; resp="$dir/responses"; log="$dir/log"; : > "$log" + # The probe below classifies the same pane three times (pane state, the + # recovery-grade read, the husk check), and the canned fake consumes + # responses in call order, so the same three-call script is laid down for + # each pass: + for n in 0 3 6; do + # +1: pane get -> the pane structurally exists + printf '{"result":{"pane":{"pane_id":"w1:p2"}}}\n' > "$resp/$((n + 1)).out" + # +2: agent get -> a registered agent with the given status + printf '{"result":{"agent":{"agent":"pi","agent_status":"%s"}}}\n' "$2" > "$resp/$((n + 2)).out" + # +3: pane process-info -> the pane's actual process view + [ "$3" = - ] || printf '%s\n' "$3" > "$resp/$((n + 3)).out" + [ -z "${4:-}" ] || printf '%s\n' "$4" > "$resp/$((n + 3)).exit" + done + install_stale_shell_ps "$dir" "$3" + fb=$(make_herdr_fakebin "$dir") + FM_HERDR_PS_BIN="$dir/ps-overlay" FM_BACKEND_HERDR_IDLE_SHELL_PROOF_POLLS=1 \ + PATH="$fb:$PATH" FM_HERDR_LOG="$log" FM_HERDR_RESPONSES="$resp" \ + bash -c '. "$0/bin/backends/herdr.sh" + printf "%s %s " "$(fm_backend_herdr_pane_agent_state fmtest w1:p2)" "$(fm_backend_herdr_agent_state fmtest:w1:p2)" + fm_backend_herdr_tab_is_husk fmtest w1:p2 && printf husk || printf refused' "$ROOT" +} + +shell_only_process_info() { # <shell-pid> + printf '{"result":{"type":"pane_process_info","process_info":{"pane_id":"w1:p2","shell_pid":%s,"foreground_process_group_id":%s,"foreground_processes":[{"pid":%s,"name":"zsh","argv0":"zsh","argv":["-zsh"],"cmdline":"-zsh"}]}}}' "$1" "$1" "$1" +} + +test_stale_registration_over_a_shell_only_pane_is_agent_free() { + local sleep_bin shell_pid out + sleep_bin=$(command -v sleep) || fail "sleep not found" + # A real, childless process stands in for the pane's shell. + "$sleep_bin" 300 & + shell_pid=$! + out=$(stale_registration_case shell-only idle "$(shell_only_process_info "$shell_pid")") + kill "$shell_pid" 2>/dev/null || true + [ "$out" = "stale-agent dead refused" ] \ + || fail "a registered idle agent over a shell-only pane must read stale-agent, recover as dead, and still refuse husk closing; got '$out'" + pass "herdr stale registration: a shell-only pane with a lingering Pi record is agent-free with an explicit reason" +} + +test_stale_registration_ignores_status_and_reads_the_process() { + local sleep_bin shell_pid out status + sleep_bin=$(command -v sleep) || fail "sleep not found" + "$sleep_bin" 300 & + shell_pid=$! + for status in working 'done' blocked; do + out=$(stale_registration_case "shell-only-$status" "$status" "$(shell_only_process_info "$shell_pid")") + [ "$out" = "stale-agent dead refused" ] \ + || { kill "$shell_pid" 2>/dev/null; fail "a lingering '$status' record over a shell-only pane must still read stale-agent/dead, got '$out'"; } + done + kill "$shell_pid" 2>/dev/null || true + pass "herdr stale registration: no registered status can outrank a shell-only process view" +} + +test_registered_agent_with_a_live_foreground_process_stays_alive() { + local out + # The real Pi shape on Herdr 0.9.0: the kernel name is the interpreter and + # only argv0 says pi. + out=$(stale_registration_case live-pi idle \ + '{"result":{"type":"pane_process_info","process_info":{"pane_id":"w1:p2","shell_pid":4242,"foreground_process_group_id":4243,"foreground_processes":[{"pid":4243,"name":"node","argv0":"pi","argv":["pi"],"cmdline":"pi"}]}}}') + [ "$out" = "live alive refused" ] \ + || fail "a registered agent whose foreground process is Pi must stay live/alive, got '$out'" + pass "herdr stale registration: a registered agent with a live Pi foreground process still reads alive" +} + +test_registered_agent_with_a_non_shell_foreground_process_stays_alive() { + local out + # A registered agent running a foreground tool in its own process group is + # not a shell-only pane, so the registration keeps its authority. + out=$(FM_BACKEND_HERDR_IDLE_SHELL_PROOF_POLLS=1 stale_registration_case live-tool working \ + '{"result":{"type":"pane_process_info","process_info":{"pane_id":"w1:p2","shell_pid":4242,"foreground_process_group_id":4250,"foreground_processes":[{"pid":4250,"name":"git","argv0":"git","argv":["git","status"],"cmdline":"git status"}]}}}') + [ "$out" = "live alive refused" ] \ + || fail "a registered agent with a non-shell foreground process must stay live/alive, got '$out'" + pass "herdr stale registration: only a shell-only pane demotes a registration" +} + +# settle_registration_case: one pane classification over a scripted sequence +# of `pane process-info` samples, so the settle window's resampling is +# observable in the fake CLI's call log. +settle_registration_case() { # <dir-suffix> <polls> <process-info-body>... + local dir="$TMP_ROOT/settle-reg-$1" polls=$2 resp log fb n + shift 2 + mkdir -p "$dir/responses"; resp="$dir/responses"; log="$dir/log"; : > "$log" + printf '{"result":{"pane":{"pane_id":"w1:p2"}}}\n' > "$resp/1.out" + printf '{"result":{"agent":{"agent":"pi","agent_status":"idle"}}}\n' > "$resp/2.out" + n=3 + for body in "$@"; do + printf '%s\n' "$body" > "$resp/$n.out" + n=$((n + 1)) + done + install_stale_shell_ps "$dir" "$1" + fb=$(make_herdr_fakebin "$dir") + FM_HERDR_PS_BIN="$dir/ps-overlay" PATH="$fb:$PATH" FM_HERDR_LOG="$log" FM_HERDR_RESPONSES="$resp" \ + FM_BACKEND_HERDR_IDLE_SHELL_PROOF_POLLS="$polls" \ + bash -c '. "$0/bin/backends/herdr.sh" + printf "%s %s" "$(fm_backend_herdr_pane_agent_state fmtest w1:p2)" "$(grep -c "process-info" "$1")"' "$ROOT" "$log" +} + +prompt_helper_process_info() { # <shell-pid> + printf '{"result":{"type":"pane_process_info","process_info":{"pane_id":"w1:p2","shell_pid":%s,"foreground_process_group_id":%s,"foreground_processes":[{"pid":99998,"name":"starship","argv":["/usr/local/bin/starship","prompt","--continuation"]},{"pid":%s,"name":"zsh","argv0":"zsh","argv":["-zsh"],"cmdline":"-zsh"}]}}}' "$1" "$1" "$1" +} + +test_transient_prompt_helper_settles_into_stale_agent() { + local sleep_bin shell_pid out + sleep_bin=$(command -v sleep) || fail "sleep not found" + "$sleep_bin" 300 & + shell_pid=$! + # Sample 1: the shell is redrawing its prompt with starship beside it (the + # real 0.7.5 shape); sample 2: the helper is gone and the shell is alone. + out=$(settle_registration_case helper-settles 3 \ + "$(prompt_helper_process_info "$shell_pid")" "$(shell_only_process_info "$shell_pid")") + kill "$shell_pid" 2>/dev/null || true + [ "$out" = "stale-agent 2" ] \ + || fail "a transient prompt helper followed by a shell-only sample must settle into stale-agent after exactly two samples, got '$out'" + pass "herdr stale registration: a transient prompt helper settles into stale-agent instead of reading live" +} + +test_exhausted_settle_window_keeps_a_non_shell_foreground_live() { + local sleep_bin shell_pid out + sleep_bin=$(command -v sleep) || fail "sleep not found" + "$sleep_bin" 300 & + shell_pid=$! + out=$(settle_registration_case helper-persists 2 \ + "$(prompt_helper_process_info "$shell_pid")" "$(prompt_helper_process_info "$shell_pid")" \ + "$(shell_only_process_info "$shell_pid")") + kill "$shell_pid" 2>/dev/null || true + [ "$out" = "live 2" ] \ + || fail "a foreground that never settles within the bound must stay live after exactly the bounded sample count, got '$out'" + pass "herdr stale registration: an exhausted settle window still reads a non-shell foreground as live" +} + +test_registered_agent_with_an_agent_descendant_outside_the_foreground_stays_alive() { + local lab sleep_bin shell_pid out shell_verdict + sleep_bin=$(command -v sleep) || fail "sleep not found" + lab="$TMP_ROOT/stale-reg-descendant-bin"; mkdir -p "$lab" + # A symlink to a real long-running binary so the kernel records `pi` as the + # executable identity (a copied platform binary fails code signing on macOS). + ln -sf "$sleep_bin" "$lab/pi" + # A real shell whose child is that agent-named process, while the canned + # foreground view shows only the shell (a suspended or backgrounded agent). + sh -c "'$lab/pi' 300; :" & + shell_pid=$! + sleep 0.3 + out=$(stale_registration_case descendant idle "$(shell_only_process_info "$shell_pid")") + pkill -P "$shell_pid" 2>/dev/null || true + kill "$shell_pid" 2>/dev/null || true + [ "$out" = "live alive refused" ] \ + || fail "a registered agent with a live agent-named descendant must stay live/alive, got '$out'" + # The divergence itself: the identical canned foreground view reads + # stale-agent for a childless shell, so the descendant walk is what carried + # this verdict. + "$sleep_bin" 300 & + shell_pid=$! + shell_verdict=$(stale_registration_case descendant-childless idle "$(shell_only_process_info "$shell_pid")") + kill "$shell_pid" 2>/dev/null || true + [ "$shell_verdict" = "stale-agent dead refused" ] \ + || fail "the childless control must read stale-agent so the descendant case is not vacuous, got '$shell_verdict'" + pass "herdr stale registration: an agent process outside the foreground group still counts as alive" +} + +test_agent_descendant_under_a_spaced_install_path_stays_alive() { + local lab sleep_bin shell_pid out + sleep_bin=$(command -v sleep) || fail "sleep not found" + # The executable path the process table reports contains a space (the macOS + # `/Library/Application Support/...` shape), so a field-split read of the + # process table sees only a fragment of the name. + lab="$TMP_ROOT/stale-reg-spaced-bin/Application Support/Some Dir"; mkdir -p "$lab" + ln -sf "$sleep_bin" "$lab/pi" + sh -c "'$lab/pi' 300; :" & + shell_pid=$! + sleep 0.3 + out=$(stale_registration_case spaced-descendant idle "$(shell_only_process_info "$shell_pid")") + pkill -P "$shell_pid" 2>/dev/null || true + kill "$shell_pid" 2>/dev/null || true + [ "$out" = "live alive refused" ] \ + || fail "an agent-named descendant under a spaced install path must stay live/alive, got '$out'" + pass "herdr stale registration: the descendant walk reads a spaced executable path whole" +} + +test_registered_agent_with_an_unreadable_process_view_is_unknown() { + local out + out=$(stale_registration_case unreadable-exit idle 'Error: socket unavailable' 1) + [ "$out" = "unknown unreadable refused" ] \ + || fail "a failed process-info read must not demote OR trust the registration: expected unknown/unreadable, got '$out'" + out=$(stale_registration_case unreadable-empty idle -) + [ "$out" = "unknown unreadable refused" ] \ + || fail "an empty process-info read must read unknown/unreadable, got '$out'" + out=$(stale_registration_case unreadable-mismatch idle \ + '{"result":{"type":"pane_process_info","process_info":{"pane_id":"w9:p9","shell_pid":4242,"foreground_process_group_id":4242,"foreground_processes":[{"pid":4242,"name":"zsh","argv0":"zsh"}]}}}') + [ "$out" = "unknown unreadable refused" ] \ + || fail "a process view for a different pane must read unknown/unreadable, got '$out'" + out=$(stale_registration_case unreadable-no-foreground idle \ + '{"result":{"type":"pane_process_info","process_info":{"pane_id":"w1:p2","shell_pid":4242,"foreground_process_group_id":4242,"foreground_processes":[]}}}') + [ "$out" = "unknown unreadable refused" ] \ + || fail "an empty foreground list must read unknown/unreadable, got '$out'" + pass "herdr stale registration: an unreadable process view refuses instead of guessing either way" +} + +test_registered_agent_with_an_empty_foreground_over_a_real_shell_settles_via_descendant_walk() { + local sleep_bin shell_pid out + sleep_bin=$(command -v sleep) || fail "sleep not found" + # A real, childless shell process stands in for the pane's shell, and the + # foreground list is empty - the exec-to-shell handoff shape the flake fix + # targets. Unlike unreadable-no-foreground above (a synthetic pid absent + # from `ps`), this shell_pid is real, so the descendant walk can run to + # completion and prove the empty array settles to stale-agent, not + # unreadable. + "$sleep_bin" 300 & + shell_pid=$! + out=$(stale_registration_case empty-foreground idle \ + "$(printf '{"result":{"type":"pane_process_info","process_info":{"pane_id":"w1:p2","shell_pid":%s,"foreground_process_group_id":%s,"foreground_processes":[]}}}' "$shell_pid" "$shell_pid")") + kill "$shell_pid" 2>/dev/null || true + [ "$out" = "live alive refused" ] \ + || fail "an incomplete foreground list must retain conservative liveness without licensing recovery, got '$out'" + pass "herdr stale registration: an empty foreground list over a real shell is not unreadable, it settles via the descendant walk" +} + +test_projection_reclaim_rollback_refuses_a_stale_registration() { + local out + out=$(bash -c '. "$0/bin/backends/herdr.sh" + fm_backend_herdr_pane_agent_state() { printf stale-agent; } + fm_backend_herdr_projection_close_pane_focus_preserving() { printf "CLOSED %s\n" "$2" >&2; exit 99; } + fm_backend_herdr_projection_reclaim_rollback fmtest w1:p9; printf "rc=%s" "$?"' "$ROOT" 2>&1) + [ "$out" = "rc=1" ] \ + || fail "reclaim rollback must refuse (never close) a pane with a stale registration, got '$out'" + pass "herdr stale registration: presentation reclaim never closes a stale-registration pane" +} + +test_busy_state_never_reports_a_shell_only_pane_busy() { + local sleep_bin shell_pid dir resp log fb out + sleep_bin=$(command -v sleep) || fail "sleep not found" + "$sleep_bin" 300 & + shell_pid=$! + dir="$TMP_ROOT/busy-stale"; mkdir -p "$dir/responses"; resp="$dir/responses"; log="$dir/log"; : > "$log" + # 1: agent get -> a lingering working record; 2: process-info -> shell only + printf '{"result":{"agent":{"agent":"pi","agent_status":"working"}}}\n' > "$resp/1.out" + shell_only_process_info "$shell_pid" > "$resp/2.out" + install_stale_shell_ps "$dir" "$(cat "$resp/2.out")" + fb=$(make_herdr_fakebin "$dir") + out=$(FM_HERDR_PS_BIN="$dir/ps-overlay" PATH="$fb:$PATH" FM_HERDR_LOG="$log" FM_HERDR_RESPONSES="$resp" \ + bash -c '. "$0/bin/backends/herdr.sh"; fm_backend_herdr_busy_state fmtest:w1:p2' "$ROOT") + kill "$shell_pid" 2>/dev/null || true + [ "$out" = unknown ] \ + || fail "a working record over a shell-only pane must not read busy, got '$out'" + assert_contains "$(cat "$log")" $'pane\x1fprocess-info' "busy_state did not verify the working record at process level" + + # The control: the same working record with a live Pi foreground reads busy. + dir="$TMP_ROOT/busy-live"; mkdir -p "$dir/responses"; resp="$dir/responses"; log="$dir/log"; : > "$log" + printf '{"result":{"agent":{"agent":"pi","agent_status":"working"}}}\n' > "$resp/1.out" + printf '{"result":{"type":"pane_process_info","process_info":{"pane_id":"w1:p2","shell_pid":4242,"foreground_process_group_id":4243,"foreground_processes":[{"pid":4243,"name":"node","argv0":"pi","argv":["pi"],"cmdline":"pi"}]}}}\n' > "$resp/2.out" + fb=$(make_herdr_fakebin "$dir") + out=$(PATH="$fb:$PATH" FM_HERDR_LOG="$log" FM_HERDR_RESPONSES="$resp" \ + bash -c '. "$0/bin/backends/herdr.sh"; fm_backend_herdr_busy_state fmtest:w1:p2' "$ROOT") + [ "$out" = busy ] || fail "a working record with a live Pi foreground must read busy, got '$out'" + + # An idle record needs no process read: idle is never trusted as busy anyway. + dir="$TMP_ROOT/busy-idle"; mkdir -p "$dir/responses"; resp="$dir/responses"; log="$dir/log"; : > "$log" + printf '{"result":{"agent":{"agent":"pi","agent_status":"idle"}}}\n' > "$resp/1.out" + fb=$(make_herdr_fakebin "$dir") + out=$(PATH="$fb:$PATH" FM_HERDR_LOG="$log" FM_HERDR_RESPONSES="$resp" \ + bash -c '. "$0/bin/backends/herdr.sh"; fm_backend_herdr_busy_state fmtest:w1:p2' "$ROOT") + [ "$out" = idle ] || fail "an idle record should read idle without a process read, got '$out'" + assert_not_contains "$(cat "$log")" $'pane\x1fprocess-info' "busy_state ran a process read for an idle record" + pass "herdr stale registration: busy_state proves a working record at process level before reporting busy" +} + test_agent_state_bypasses_a_stale_client_shadowing_a_compatible_one() { local dir out err dir="$TMP_ROOT/client-pair-bypass"; make_herdr_client_pair "$dir" @@ -876,6 +1185,8 @@ test_create_task_refuses_duplicate_label_when_agent_live() { printf '{"result":{"pane":{"pane_id":"w1:p2"}}}\n' > "$resp/3.out" # 4: agent get -> a genuinely registered, live agent (idle, not just working) printf '{"result":{"agent":{"agent_status":"idle"}}}\n' > "$resp/4.out" + # 5: pane process-info -> a live Pi process backs that registration (#4115) + printf '%s\n' '{"result":{"type":"pane_process_info","process_info":{"pane_id":"w1:p2","shell_pid":4242,"foreground_process_group_id":4243,"foreground_processes":[{"pid":4243,"name":"node","argv0":"pi"}]}}}' > "$resp/5.out" fb=$(make_herdr_fakebin "$dir") out=$( PATH="$fb:$PATH" FM_HERDR_LOG="$log" FM_HERDR_RESPONSES="$resp" \ bash -c '. "$0/bin/backends/herdr.sh"; fm_backend_herdr_create_task fmtest:w1 fm-dup1 /tmp/proj' "$ROOT" 2>&1 ) @@ -897,6 +1208,8 @@ test_create_task_refuses_when_any_duplicate_label_is_live() { printf '{"result":{"panes":[{"pane_id":"w1:p2","tab_id":"w1:t2"},{"pane_id":"w1:p3","tab_id":"w1:t3"}]}}\n' > "$resp/5.out" printf '{"result":{"pane":{"pane_id":"w1:p3"}}}\n' > "$resp/6.out" printf '{"result":{"agent":{"agent_status":"idle"}}}\n' > "$resp/7.out" + # 8: pane process-info -> a live Pi process backs that registration (#4115) + printf '%s\n' '{"result":{"type":"pane_process_info","process_info":{"pane_id":"w1:p3","shell_pid":4242,"foreground_process_group_id":4243,"foreground_processes":[{"pid":4243,"name":"node","argv0":"pi"}]}}}' > "$resp/8.out" fb=$(make_herdr_fakebin "$dir") out=$( PATH="$fb:$PATH" FM_HERDR_LOG="$log" FM_HERDR_RESPONSES="$resp" \ bash -c '. "$0/bin/backends/herdr.sh"; fm_backend_herdr_create_task fmtest:w1 fm-mixed1 /tmp/proj' "$ROOT" 2>&1 ) @@ -2711,8 +3024,8 @@ test_agent_state_inconclusive_process_read_stays_alive() { make_agent_free_lab "$dir" fb=$(make_herdr_fakebin "$dir") out=$(run_agent_state "$fb" "$log" "$resp" "$dir/ps" 2) - [ "$out" = alive ] || fail "an inconclusive process read must keep a registered agent alive (refusing recovery), got '$out'" - pass "fm_backend_herdr_agent_state: an inconclusive process read keeps the registration alive and recovery refused" + [ "$out" = unreadable ] || fail "an inconclusive process read must refuse recovery as unreadable, got '$out'" + pass "fm_backend_herdr_agent_state: an inconclusive process read reports unreadable and refuses recovery" } test_agent_state_transient_prompt_helper_settles_to_dead() { @@ -3512,6 +3825,8 @@ test_projection_recovery_is_read_only_and_refuses_live_duplicate_risk() { printf '{"result":{"panes":[{"pane_id":"w1:p1","tab_id":"w1:t1"}]}}\n' > "$resp/2.out" printf '{"result":{"pane":{"pane_id":"w1:p1"}}}\n' > "$resp/3.out" printf '{"result":{"agent":{"agent_status":"idle"}}}\n' > "$resp/4.out" + # 5: process-info -> a live harness backs the registration (issue #4115) + printf '{"result":{"type":"pane_process_info","process_info":{"pane_id":"w1:p1","shell_pid":4242,"foreground_process_group_id":4243,"foreground_processes":[{"pid":4243,"name":"node","argv0":"pi"}]}}}\n' > "$resp/5.out" out=$(PATH="$fb:$PATH" FM_HERDR_LOG="$log" FM_HERDR_RESPONSES="$resp" \ bash -c '. "$0/bin/backends/herdr.sh"; fm_backend_herdr_projection_recovery_allows_flat fmtest "$1" task-p3' "$ROOT" "$journal" 2>&1) status=$? @@ -5290,6 +5605,18 @@ test_workspace_label_different_secondmates_get_different_labels test_cli_helper_sets_env_and_appends_trailing_session_flag test_agent_state_bypasses_a_stale_client_shadowing_a_compatible_one test_recovery_grade_read_widens_only_at_its_own_boundary +test_stale_registration_over_a_shell_only_pane_is_agent_free +test_stale_registration_ignores_status_and_reads_the_process +test_registered_agent_with_a_live_foreground_process_stays_alive +test_registered_agent_with_a_non_shell_foreground_process_stays_alive +test_transient_prompt_helper_settles_into_stale_agent +test_exhausted_settle_window_keeps_a_non_shell_foreground_live +test_registered_agent_with_an_agent_descendant_outside_the_foreground_stays_alive +test_agent_descendant_under_a_spaced_install_path_stays_alive +test_registered_agent_with_an_unreadable_process_view_is_unknown +test_registered_agent_with_an_empty_foreground_over_a_real_shell_settles_via_descendant_walk +test_projection_reclaim_rollback_refuses_a_stale_registration +test_busy_state_never_reports_a_shell_only_pane_busy test_cli_caches_the_selected_client_within_a_process test_cli_scopes_the_selected_client_to_its_session test_cli_unrelated_failure_never_triggers_reselection diff --git a/tests/fm-backlog-read-bound.test.sh b/tests/fm-backlog-read-bound.test.sh new file mode 100755 index 00000000000..ac5088f0911 --- /dev/null +++ b/tests/fm-backlog-read-bound.test.sh @@ -0,0 +1,430 @@ +#!/usr/bin/env bash +# tests/fm-backlog-read-bound.test.sh - behavior tests for the per-item bound on +# bin/fm-backlog-transition-lib.sh's backlog row read. +# +# The defect this pins: bin/fm-bootstrap.sh's reconcile and close-replay sweeps +# read the backlog backend once per item, and an unbounded read of a wedged +# backend consumed the whole FM_SESSION_START_TIMEOUT. The digest was then +# truncated before the wake queue, supervision instructions, fleet state, and +# context sections printed, leaving a whole fleet unsupervised. +# +# Both halves are proved here: +# - a deliberately hanging `tasks-axi show` cannot exceed the per-item bound, +# and the failure names the item it could not read +# - a session start against that same wedged backend still completes end to +# end, with every digest section present and a loud partial reconcile +# +# The bound must hold on its own, independent of any particular tasks-axi +# install, so the fake here simply never returns. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +BASE_PATH=${FM_TEST_BASE_PATH:-/usr/bin:/bin:/usr/sbin:/sbin} +TMP_ROOT=$(fm_test_tmproot fm-backlog-read-bound-tests) +trap fm_test_cleanup EXIT + +BOUND_SECS=2 +# Generous enough that a slow CI box never flakes, far below the unbounded hang +# (300s per read) and below the session-start budget the defect consumed. +BOUND_CEILING=30 + +# A backend whose `show` never returns. Everything the compatibility gate and the +# startup listing need still answers promptly, so the only thing under test is +# the read that hangs. +make_hanging_tasks_axi() { # <fakebin> + local fakebin=$1 + cat > "$fakebin/tasks-axi" <<'SH' +#!/usr/bin/env bash +set -u +case "${1:-}" in + --version) printf '%s\n' '0.2.5'; exit 0 ;; + update) + [ "${2:-}" = --help ] || exit 0 + printf '%s\n' 'usage: tasks-axi update <id> [flags]' ' --body-file <path>' ' --archive-body' + exit 0 + ;; + mv) + [ "${2:-}" = --help ] || exit 0 + printf '%s\n' 'usage: tasks-axi mv <id> [<id>...] --to <path-or-dir>' + exit 0 + ;; + show) + # A real backend rejects an unusable id promptly instead of wedging, which + # is what makes a dropped 124 surface as "absent" rather than as a bound. + if [ -z "${2:-}" ]; then + printf 'code: NOT_FOUND\n' >&2 + exit 1 + fi + # The wedge under test: a read that never returns. + sleep 300 + exit 0 + ;; + hold) + [ "${2:-}" = --help ] || exit 0 + printf '%s\n' 'usage: tasks-axi hold <id> [flags]' ' --kind captain' ' --until <date>' + exit 0 + ;; + add) + # Recorded, never silent: creating a row that already exists is the damage a + # timed-out read must never be spent on. + [ -z "${FM_TEST_TASKS_AXI_ADD_LOG:-}" ] || printf '%s\n' "$*" >> "$FM_TEST_TASKS_AXI_ADD_LOG" + exit 0 + ;; + list) + printf 'count: 0\n' + printf 'tasks[0]{id,state,kind,repo,title,blocked_by,hold_kind,hold_reason}:\n' + exit 0 + ;; +esac +exit 0 +SH + chmod +x "$fakebin/tasks-axi" +} + +elapsed_since() { # <start-epoch> + local now + now=$(date +%s) + printf '%s\n' "$((now - $1))" +} + +# --- half one: the per-item bound holds ------------------------------------- + +UNIT="$TMP_ROOT/unit" +UNIT_FAKEBIN=$(fm_fakebin "$UNIT") +mkdir -p "$UNIT/data" +make_hanging_tasks_axi "$UNIT_FAKEBIN" +printf '# Backlog\n' > "$UNIT/data/backlog.md" + +# Three items, so "every skipped item is still named" is actually exercised +# rather than inferred from a single skip. +PROBE_OUT="$UNIT/probe.out" +PATH="$UNIT_FAKEBIN:$BASE_PATH" FM_BACKLOG_ROW_TIMEOUT_SECS="$BOUND_SECS" \ + bash -c ' + set -u + . "$1/bin/fm-tasks-axi-lib.sh" + . "$1/bin/fm-backlog-transition-lib.sh" + for id in wedged-one wedged-two wedged-three; do + start=$(date +%s) + fm_backlog_row_probe "$2" "$id" && printf "unexpected-success\n" + printf "elapsed:%s=%s\n" "$id" "$(( $(date +%s) - start ))" + printf "error:%s=%s\n" "$id" "$FM_BACKLOG_ROW_ERROR" + done + ' _ "$ROOT" "$UNIT/data" > "$PROBE_OUT" 2>&1 + +probe_elapsed() { # <id> + sed -n "s/^elapsed:$1=//p" "$PROBE_OUT" +} + +probe_error() { # <id> + sed -n "s/^error:$1=//p" "$PROBE_OUT" +} + +grep -q '^unexpected-success$' "$PROBE_OUT" \ + && fail "a hanging tasks-axi show must not report a successful row read: $(cat "$PROBE_OUT")" + +FIRST_ELAPSED=$(probe_elapsed wedged-one) +[ -n "$FIRST_ELAPSED" ] || fail "probe produced no timing: $(cat "$PROBE_OUT")" +[ "$FIRST_ELAPSED" -lt "$BOUND_CEILING" ] \ + || fail "bounded row read took ${FIRST_ELAPSED}s, over the ${BOUND_CEILING}s ceiling: $(cat "$PROBE_OUT")" +pass "a hanging tasks-axi show returns within the per-item bound instead of running unbounded" + +FIRST_ERROR=$(probe_error wedged-one) +case "$FIRST_ERROR" in + *wedged-one*bound*) ;; + *) fail "the timed-out read must name the item and its bound, got: $FIRST_ERROR" ;; +esac +pass "a timed-out row read reports one error naming the item that timed out" + +# The latch is what keeps a home carrying a large fleet from paying N bounds and +# losing the digest anyway, so assert it strictly: a latched read must be +# FASTER than one bound, not merely under the ceiling. A ceiling-only assertion +# passes whether or not the latch works, and fm_backlog_row_show runs inside a +# command substitution whose writes die with the subshell - the exact way this +# latch can silently become inert. +for SKIPPED in wedged-two wedged-three; do + SKIPPED_ERROR=$(probe_error "$SKIPPED") + SKIPPED_ELAPSED=$(probe_elapsed "$SKIPPED") + case "$SKIPPED_ERROR" in + *"$SKIPPED"*skipped*) ;; + *) fail "every skipped item must still be named as skipped, $SKIPPED got: $SKIPPED_ERROR" ;; + esac + [ -n "$SKIPPED_ELAPSED" ] && [ "$SKIPPED_ELAPSED" -lt "$BOUND_SECS" ] \ + || fail "the latch is inert: $SKIPPED paid ${SKIPPED_ELAPSED}s against a known-wedged backend" +done +pass "after the first bound hit the sweep continues and names every remaining item without paying the bound again" + +# A padded zero is still zero, and `timeout 0` / `alarm 0` disable the deadline +# outright, so a bound that only rejects the literal 0 silently restores the +# unbounded read this whole change exists to prevent. +PADDED_OUT="$UNIT/padded.out" +PADDED_START=$(date +%s) +PATH="$UNIT_FAKEBIN:$BASE_PATH" FM_BACKLOG_ROW_TIMEOUT_SECS=00 \ + bash -c ' + set -u + . "$1/bin/fm-tasks-axi-lib.sh" + . "$1/bin/fm-backlog-transition-lib.sh" + fm_backlog_row_probe "$2" padded-zero && printf "unexpected-success\n" + printf "error=%s\n" "$FM_BACKLOG_ROW_ERROR" + ' _ "$ROOT" "$UNIT/data" > "$PADDED_OUT" 2>&1 +PADDED_ELAPSED=$(elapsed_since "$PADDED_START") + +[ "$PADDED_ELAPSED" -lt "$BOUND_CEILING" ] \ + || fail "a padded-zero bound disabled the deadline: the read ran ${PADDED_ELAPSED}s" +case "$(sed -n 's/^error=//p' "$PADDED_OUT")" in + *padded-zero*bound*) ;; + *) fail "a padded-zero bound must fall back to the default bound and report it: $(cat "$PADDED_OUT")" ;; +esac +pass "a padded-zero bound falls back to the default instead of disabling the deadline" + +# --- a bound hit is not absence --------------------------------------------- +# +# Turning a hang into a fast 124 reaches every caller that reads a non-zero row +# status as "this row does not exist". bin/fm-captain-hold.sh's hold path is the +# one where that misreading corrupts: it would create a task that already +# exists. The bound must stop the command instead. + +CAPTAIN="$TMP_ROOT/captain" +CAPTAIN_FAKEBIN=$(fm_fakebin "$CAPTAIN") +mkdir -p "$CAPTAIN/data" "$CAPTAIN/state" "$CAPTAIN/config" +make_hanging_tasks_axi "$CAPTAIN_FAKEBIN" +cp "$ROOT/.tasks.toml" "$CAPTAIN/.tasks.toml" +printf '# Backlog\n' > "$CAPTAIN/data/backlog.md" + +ADD_LOG="$CAPTAIN/add.log" +HOLD_OUT="$CAPTAIN/hold.out" +HOLD_STATUS=0 +PATH="$CAPTAIN_FAKEBIN:$BASE_PATH" FM_HOME="$CAPTAIN" \ + FM_STATE_OVERRIDE="$CAPTAIN/state" FM_DATA_OVERRIDE="$CAPTAIN/data" \ + FM_CONFIG_OVERRIDE="$CAPTAIN/config" FM_BACKLOG_ROW_TIMEOUT_SECS="$BOUND_SECS" \ + FM_TEST_TASKS_AXI_ADD_LOG="$ADD_LOG" \ + "$ROOT/bin/fm-captain-hold.sh" hold wedged-hold --title 'Wedged hold' --reason 'backend wedged' \ + > "$HOLD_OUT" 2>&1 || HOLD_STATUS=$? + +[ "$HOLD_STATUS" -ne 0 ] \ + || fail "holding a task against a wedged backend must not report success: $(cat "$HOLD_OUT")" +[ ! -s "$ADD_LOG" ] \ + || fail "a timed-out read was spent as absence: tasks-axi add ran anyway: $(cat "$ADD_LOG")" +case "$(cat "$HOLD_OUT")" in + *wedged-hold*bound*) ;; + *) fail "the refusal must name the item and the bound it hit, got: $(cat "$HOLD_OUT")" ;; +esac +pass "a bound hit stops a captain hold loudly instead of being read as a missing task" + +# The teardown gate reaches a row read through the same resolver, so the bound +# hit has to survive the command substitution that carries the resolved id. +fm_write_meta "$CAPTAIN/state/wedged-origin.meta" \ + 'window=firstmate:fm-wedged-origin' \ + 'worktree=/nonexistent/wedged-origin' \ + 'project=alpha' \ + 'harness=claude' \ + 'decisions_reviewed=1' \ + 'decision_keys=wedged-entry' + +VERIFY_OUT="$CAPTAIN/verify.out" +VERIFY_STATUS=0 +PATH="$CAPTAIN_FAKEBIN:$BASE_PATH" FM_HOME="$CAPTAIN" \ + FM_STATE_OVERRIDE="$CAPTAIN/state" FM_DATA_OVERRIDE="$CAPTAIN/data" \ + FM_CONFIG_OVERRIDE="$CAPTAIN/config" FM_BACKLOG_ROW_TIMEOUT_SECS="$BOUND_SECS" \ + "$ROOT/bin/fm-captain-hold.sh" verify wedged-origin > "$VERIFY_OUT" 2>&1 || VERIFY_STATUS=$? + +[ "$VERIFY_STATUS" -ne 0 ] \ + || fail "verify must not attest an inventory it could not read: $(cat "$VERIFY_OUT")" +case "$(cat "$VERIFY_OUT")" in + *absent*) fail "a bound hit was reported as an absent task: $(cat "$VERIFY_OUT")" ;; +esac +case "$(cat "$VERIFY_OUT")" in + *wedged-entry*bound*) ;; + *) fail "verify must name the entry it could not read and the bound it hit, got: $(cat "$VERIFY_OUT")" ;; +esac +pass "the teardown verify gate reports a bound hit by name instead of as an absent inventory entry" + +# The reconcile-requests intake reads each row with task_show in this shell and +# must stop on a bound hit by name; spending the 124 as 'refused: <id> +# (absent)' would let a wedged backend erase real rows from the reconcile +# sweep. +REQ="$TMP_ROOT/req" +REQ_FAKEBIN=$(fm_fakebin "$REQ") +mkdir -p "$REQ/data" "$REQ/state" "$REQ/config" "$REQ/state/decision-bindings" +make_hanging_tasks_axi "$REQ_FAKEBIN" +cp "$ROOT/.tasks.toml" "$REQ/.tasks.toml" +printf '# Backlog\n' > "$REQ/data/backlog.md" +printf 'schema=fm-decision-binding.v1\norigin=wedged-origin\n' \ + > "$REQ/state/decision-bindings/probe.origin" + +REQ_OUT="$REQ/req.out" +REQ_STATUS=0 +printf 'wedged-req\n' \ + | PATH="$REQ_FAKEBIN:$BASE_PATH" FM_HOME="$REQ" \ + FM_STATE_OVERRIDE="$REQ/state" FM_DATA_OVERRIDE="$REQ/data" \ + FM_CONFIG_OVERRIDE="$REQ/config" FM_BACKLOG_ROW_TIMEOUT_SECS="$BOUND_SECS" \ + "$ROOT/bin/fm-captain-hold.sh" reconcile-requests --source-id probe --source 'test capture' \ + > "$REQ_OUT" 2>&1 || REQ_STATUS=$? + +[ "$REQ_STATUS" -ne 0 ] \ + || fail "reconcile-requests must not report success against a wedged backend: $(cat "$REQ_OUT")" +case "$(cat "$REQ_OUT")" in + *absent*|*refused*) fail "the reconcile intake spent a bound hit as an absent row: $(cat "$REQ_OUT")" ;; +esac +case "$(cat "$REQ_OUT")" in + *wedged-req*bound*) ;; + *) fail "the reconcile intake must name the row and the bound it hit, got: $(cat "$REQ_OUT")" ;; +esac +pass "the reconcile-requests intake stops loudly on a bound hit instead of refusing the row as absent" + +# The migrated-prefix scan is the resolution path whose exact and legacy ids +# genuinely answer NOT_FOUND: only the prefixed migrated row wedges. A dropped +# 124 there falls through to 'no captain-held task $entry resolves to nothing' +# - the exact bound-hit-as-absence outcome the resolver's own 124 arm exists to +# prevent - so the bound must survive the prefixed scan to verify_entry_durable. +MIG="$TMP_ROOT/migrated" +MIG_FAKEBIN=$(fm_fakebin "$MIG") +mkdir -p "$MIG/data" "$MIG/state" "$MIG/config" +cat > "$MIG_FAKEBIN/tasks-axi" <<'SH' +#!/usr/bin/env bash +set -u +case "${1:-}" in + --version) printf '%s\n' '0.2.5'; exit 0 ;; + show) + [ -z "${2:-}" ] && { printf 'code: NOT_FOUND\n' >&2; exit 1; } + # Only the prefixed migrated candidates wedge; the exact and legacy ids + # answer NOT_FOUND promptly, the concrete path the prefix scan exists for. + case "$2" in $FM_TEST_PREFIXED_GLOB) sleep 300; exit 0 ;; esac + printf 'code: NOT_FOUND\n' >&2 + exit 1 + ;; + update) + [ "${2:-}" = --help ] || exit 0 + printf '%s\n' 'usage: tasks-axi update <id> [flags]' ' --body-file <path>' ' --archive-body' + exit 0 + ;; + mv) + [ "${2:-}" = --help ] || exit 0 + printf '%s\n' 'usage: tasks-axi mv <id> [<id>...] --to <path-or-dir>' + exit 0 + ;; + hold) + [ "${2:-}" = --help ] || exit 0 + printf '%s\n' 'usage: tasks-axi hold <id> [flags]' ' --kind captain' ' --until <date>' + exit 0 + ;; + list) + printf 'count: 0\n' + printf 'tasks[0]{id,state,kind,repo,title,blocked_by,hold_kind,hold_reason}:\n' + exit 0 + ;; +esac +exit 0 +SH +chmod +x "$MIG_FAKEBIN/tasks-axi" +cat > "$MIG_FAKEBIN/bd" <<'SH' +#!/usr/bin/env bash +[ "${1:-}" = list ] && { printf '[]\n'; exit 0; } +exit 1 +SH +chmod +x "$MIG_FAKEBIN/bd" +cat > "$MIG/.tasks.toml" <<'TOML' +backend = "beads" + +[beads] +prefix = "bd" +path = "graph" +binary = "bd" +TOML +printf '# Backlog\n' > "$MIG/data/backlog.md" +fm_write_meta "$MIG/state/wedged-origin.meta" \ + 'window=firstmate:fm-wedged-origin' \ + 'worktree=/nonexistent/wedged-origin' \ + 'project=alpha' \ + 'harness=claude' \ + 'decisions_reviewed=1' \ + 'decision_keys=mig-entry' + +VERIFY_MIG_OUT="$MIG/verify.out" +VERIFY_MIG_STATUS=0 +PATH="$MIG_FAKEBIN:$BASE_PATH" FM_HOME="$MIG" \ + FM_STATE_OVERRIDE="$MIG/state" FM_DATA_OVERRIDE="$MIG/data" \ + FM_CONFIG_OVERRIDE="$MIG/config" FM_BACKLOG_ROW_TIMEOUT_SECS="$BOUND_SECS" \ + FM_TEST_PREFIXED_GLOB='bd-*' \ + "$ROOT/bin/fm-captain-hold.sh" verify wedged-origin > "$VERIFY_MIG_OUT" 2>&1 || VERIFY_MIG_STATUS=$? + +[ "$VERIFY_MIG_STATUS" -ne 0 ] \ + || fail "verify must not attest an inventory whose migrated-prefix read wedged: $(cat "$VERIFY_MIG_OUT")" +case "$(cat "$VERIFY_MIG_OUT")" in + *'no captain-held task'*|*absent*) + fail "the migrated-prefix bound hit was spent as an unresolved key: $(cat "$VERIFY_MIG_OUT")" ;; +esac +case "$(cat "$VERIFY_MIG_OUT")" in + *'exceeded its read bound resolving mig-entry') ;; + *) fail "verify must name the entry it could not read and the bound it hit, got: $(cat "$VERIFY_MIG_OUT")" ;; +esac +pass "a bound hit in the migrated-prefix scan stops verify by name instead of resolving to nothing" + +# --- half two: the digest still completes end to end ------------------------ + +E2E="$TMP_ROOT/e2e" +E2E_ROOT="$E2E/root" +E2E_HOME="$E2E/home" +E2E_FAKEBIN="$E2E/fakebin" +mkdir -p "$E2E_HOME/state" "$E2E_HOME/data" "$E2E_HOME/config" "$E2E_FAKEBIN" +git init -q -b main "$E2E_ROOT" +git -C "$E2E_ROOT" commit -q --allow-empty -m init + +make_hanging_tasks_axi "$E2E_FAKEBIN" +# The reconcile sweep this half asserts on runs only under a verified fleet +# lock, and fm-lock.sh finds its holder by walking the invoking process tree +# through `ps`. A CI runner's ancestry carries no harness process, so the lock +# would be refused there and the sweep silently skipped. Pin the lock evidence +# the same way tests/fm-session-start.test.sh's make_fake_ps_harness does: +# every queried pid reports a live `claude` harness, independent of whatever +# process tree the test itself was launched from. +cat > "$E2E_FAKEBIN/ps" <<'SH' +#!/usr/bin/env bash +set -u +case "$*" in + *"comm="*) printf '%s\n' '/usr/local/bin/claude'; exit 0 ;; + *"args="*) printf '%s\n' 'claude'; exit 0 ;; + *"ppid="*) exit 1 ;; +esac +exit 1 +SH +chmod +x "$E2E_FAKEBIN/ps" +fm_fake_exit0 "$E2E_FAKEBIN" tmux node chrome-devtools-axi gh treehouse +fm_fake_version_tool "$E2E_FAKEBIN" lavish-axi FM_FAKE_LAVISH_AXI_VERSION 0.1.46 +fm_fake_version_tool "$E2E_FAKEBIN" gh-axi FM_FAKE_GH_AXI_VERSION 0.1.29 +fm_fake_version_tool "$E2E_FAKEBIN" no-mistakes FM_FAKE_NO_MISTAKES_VERSION \ + 'no-mistakes version v1.46.0 (fake) 2026-06-27T00:02:18Z' + +printf '# Backlog\n' > "$E2E_HOME/data/backlog.md" +# One owned record, so the reconcile sweep actually reads the wedged backend. +fm_write_meta "$E2E_HOME/state/wedged-task.meta" \ + 'window=firstmate:fm-wedged-task' \ + 'worktree=/nonexistent/wedged-task' \ + 'project=alpha' \ + 'harness=claude' \ + 'mode=no-mistakes' \ + 'yolo=off' + +DIGEST="$E2E/digest.out" +DIGEST_START=$(date +%s) +env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + FM_HOME="$E2E_HOME" FM_ROOT_OVERRIDE="$E2E_ROOT" PATH="$E2E_FAKEBIN:$BASE_PATH" \ + FM_BACKLOG_ROW_TIMEOUT_SECS="$BOUND_SECS" \ + "$ROOT/bin/fm-session-start.sh" > "$DIGEST" 2>&1 || true +DIGEST_ELAPSED=$(elapsed_since "$DIGEST_START") + +[ "$DIGEST_ELAPSED" -lt "$BOUND_CEILING" ] \ + || fail "session start took ${DIGEST_ELAPSED}s against a wedged backlog backend" + +for SECTION in 'WAKE QUEUE' 'SUPERVISION OPERATING INSTRUCTIONS' 'FLEET STATE' 'CONTEXT'; do + grep -q "$SECTION" "$DIGEST" \ + || fail "the digest lost its $SECTION section against a wedged backlog backend: $(cat "$DIGEST")" +done +pass "a wedged backlog backend still leaves a complete digest: wake queue, supervision instructions, fleet state, and context all print" + +grep -q '^BACKLOG_RECONCILE: wedged-task: ' "$DIGEST" \ + || fail "the wedged item must be reported by name as a partial reconcile: $(cat "$DIGEST")" +pass "an unreachable backlog backend degrades to a loud partial reconcile naming the item it could not read" + +echo "# fm-backlog-read-bound.test.sh: all assertions passed" diff --git a/tests/fm-bearings-board-render.test.sh b/tests/fm-bearings-board-render.test.sh index cf26fd31428..d32d0e9dd79 100755 --- a/tests/fm-bearings-board-render.test.sh +++ b/tests/fm-bearings-board-render.test.sh @@ -56,12 +56,14 @@ SH printf '%s\n' "$home" } -# Build the board from <charted-json> and return what the renderer produced. -render() { # <home> <charted-json> [charted_more] [charted_warning_more] - local home=$1 charted=$2 more=${3:-0} warning_more=${4:-0} data="$1/payload.json" - jq -n --argjson charted "$charted" --argjson more "$more" --argjson warning_more "$warning_more" '{ +# Build the board from <underway-json> plus <charted-json> and return what the +# renderer produced. +render_board() { # <home> <underway-json> <charted-json> [charted_more] [charted_warning_more] + local home=$1 underway=$2 charted=$3 more=${4:-0} warning_more=${5:-0} data="$1/payload.json" + jq -n --argjson underway "$underway" --argjson charted "$charted" \ + --argjson more "$more" --argjson warning_more "$warning_more" '{ schema:"fm-bearings-board.v1", home:"render-home", generated:"2026-08-26T00:00Z", - prs_live:false, captains_call:[], underway:[], landed:[], + prs_live:false, captains_call:[], underway:$underway, landed:[], charted:$charted, charted_more:$more, charted_warning_more:$warning_more}' > "$data" PATH="$home/fakebin:$PATH" FM_HOME="$home" \ FM_STATE_OVERRIDE="$home/state" FM_DATA_OVERRIDE="$home/data" \ @@ -71,6 +73,11 @@ render() { # <home> <charted-json> [charted_more] [charted_warning_more] || fail "the built board could not be rendered" } +# Build the board from <charted-json> alone and return what the renderer produced. +render() { # <home> <charted-json> [charted_more] [charted_warning_more] + render_board "$1" '[]' "$2" "${3:-0}" "${4:-0}" +} + charted_next_count() { # <render-json> printf '%s' "$1" | jq -r '.stats[] | select(.label == "charted next") | .n' } @@ -158,6 +165,73 @@ test_an_omitted_kind_keeps_the_existing_queued_rendering() { pass "an omitted kind renders exactly as queued work always did" } +test_an_underway_row_leads_with_the_task_name_and_keeps_its_run_status() { + local home out + home=$(make_home underway-name) + out=$(render_board "$home" '[ + {"id":"fm-board-name-r1","repo":"firstmate","name":"Show task names on the board", + "state":"working","kind":"ship","doing":"no-mistakes: review round 2"} + ]' '[]') + printf '%s' "$out" | jq -e ' + (.underway | length) == 1 + and (.underway[0] + | .title == "Show task names on the board" + and (.sub | test("no-mistakes: review round 2")) + and (.sub | test("ship")) and (.sub | test("firstmate")) + and [.badges[] | .text] == ["working"]) + ' >/dev/null || fail "an underway row did not lead with the task name: $out" + pass "an underway row leads with the task name and still reports its run status" +} + +test_an_underway_identifier_label_is_not_replaced_by_run_status() { + local home out + home=$(make_home underway-identifier) + out=$(render_board "$home" '[ + {"id":"mate/child-1","repo":null,"name":"mate/child-1", + "state":"working","kind":"secondmate","doing":"fixing the failing check"} + ]' '[]') + printf '%s' "$out" | jq -e ' + (.underway | length) == 1 + and (.underway[0] + | .title == "mate/child-1" + and (.sub | startswith("fixing the failing check · ")) + and (.title != "fixing the failing check")) + ' >/dev/null || fail "an identifier-labelled underway row rendered as status-only: $out" + pass "an underway identifier label is not replaced by run status" +} + +test_charted_next_reads_newest_filed_first() { + local home out + home=$(make_home charted-order) + out=$(render_board "$home" '[]' '[ + {"id":"oldest","repo":"sample","title":"Filed in June","reason":"queued","dispatchable":true,"filed":"2026-06-01"}, + {"id":"newest","repo":"sample","title":"Filed in August","reason":"queued","dispatchable":true,"filed":"2026-08-14T09:30:00Z"}, + {"id":"middle","repo":"sample","title":"Filed in July","reason":"queued","dispatchable":true,"filed":"2026-07-22"} + ]') + printf '%s' "$out" | jq -e ' + [.charted[] | .title] == ["Filed in August", "Filed in July", "Filed in June"] + ' >/dev/null || fail "charted next was not ordered newest filed first: $out" + pass "charted next renders the most recently filed work first" +} + +test_charted_rows_without_a_filed_date_follow_the_dated_rows_in_payload_order() { + local home out + home=$(make_home charted-undated) + out=$(render_board "$home" '[]' '[ + {"id":"undated-first","repo":"sample","title":"Undated one","reason":"queued","dispatchable":true}, + {"id":"dated","repo":"sample","title":"Dated","reason":"queued","dispatchable":true,"filed":"2026-07-22"}, + {"id":"undated-second","repo":"sample","title":"Undated two","reason":"queued","dispatchable":true,"filed":null} + ]') + printf '%s' "$out" | jq -e ' + [.charted[] | .title] == ["Dated", "Undated one", "Undated two"] + ' >/dev/null || fail "undated charted rows did not keep a stable trailing order: $out" + pass "charted rows with no filed date follow the dated rows in payload order" +} + +test_an_underway_row_leads_with_the_task_name_and_keeps_its_run_status +test_an_underway_identifier_label_is_not_replaced_by_run_status +test_charted_next_reads_newest_filed_first +test_charted_rows_without_a_filed_date_follow_the_dated_rows_in_payload_order test_a_warning_row_reads_as_a_repair_not_as_queued_work test_warnings_are_excluded_from_the_charted_next_count test_a_board_of_only_warnings_still_reports_nothing_queued diff --git a/tests/fm-bearings-board.test.sh b/tests/fm-bearings-board.test.sh index b59010036e9..b5254d42bfa 100644 --- a/tests/fm-bearings-board.test.sh +++ b/tests/fm-bearings-board.test.sh @@ -260,6 +260,20 @@ test_build_refuses_malformed_payloads_before_touching_the_board() { set +e; out=$(run_board "$home" build "$data" 2>&1); rc=$?; set -e [ "$rc" -ne 0 ] || fail "a fleet row without an explicit repo marker was accepted" + write_valid_payload "$data" + jq '.underway = [{"id":"sample-task","repo":"sample","state":"working", + "kind":"ship","doing":"implementing"}]' "$data" > "$data.tmp" && mv "$data.tmp" "$data" + set +e; out=$(run_board "$home" build "$data" 2>&1); rc=$?; set -e + [ "$rc" -ne 0 ] || fail "an underway row without an explicit name marker was accepted" + + for invalid_filed in "last Tuesday" "2026-13-01" "2026-08-14T99:30:00Z" "2026-02-29"; do + write_valid_payload "$data" + jq --arg filed "$invalid_filed" '.charted[0].filed = $filed' "$data" > "$data.tmp" \ + && mv "$data.tmp" "$data" + set +e; out=$(run_board "$home" build "$data" 2>&1); rc=$?; set -e + [ "$rc" -ne 0 ] || fail "an invalid filed date was accepted: $invalid_filed" + done + write_valid_payload "$data" jq '.captains_call[0].allow_freeform = "yes"' "$data" > "$data.tmp" && mv "$data.tmp" "$data" set +e; out=$(run_board "$home" build "$data" 2>&1); rc=$?; set -e diff --git a/tests/fm-bearings-snapshot.test.sh b/tests/fm-bearings-snapshot.test.sh index ac352923e79..1786759d17e 100755 --- a/tests/fm-bearings-snapshot.test.sh +++ b/tests/fm-bearings-snapshot.test.sh @@ -2463,6 +2463,139 @@ EOF pass "active children reach Underway independently of a home captain hold" } +test_nameless_legacy_summary_uses_its_durable_identifier() { + local parent remote_home fakebin json + parent=$(make_home nameless-legacy-summary) + make_remote_ledger_fleet "$parent" 1 + remote_home="$TMP_ROOT/remote-ledger-home-1" + fakebin=$(make_remote_ledger_ssh "$parent/remote-ssh") + jq ' + .active_children = [ + {id:"legacy-child",kind:"ship",state:"working",repo:null, + source:"remote-ledger",doing:"running review"}, + {id:"blank-name-child",kind:"ship",state:"working",repo:null,name:" \t ", + source:"remote-ledger",doing:"running tests"} + ] + | .counts.active_children = 2 + | .state = "active_child_work" + ' "$remote_home/state/home-summary.json" > "$remote_home/state/legacy-summary.json" + mv "$remote_home/state/legacy-summary.json" "$remote_home/state/home-summary.json" + + json=$(run_remote_ledger_bearings "$parent" "$fakebin" 1100) \ + || fail "nameless legacy summary bearings failed" + printf '%s' "$json" | jq -e ' + (.in_flight | any(.id == "ledger-1/legacy-child" + and .name == "ledger-1/legacy-child" + and .doing == "running review" + and .name != .doing)) + and (.in_flight | any(.id == "ledger-1/blank-name-child" + and .name == "ledger-1/blank-name-child" + and .doing == "running tests" + and .name != .doing)) + ' >/dev/null || fail "a blank legacy child name was not replaced by its id: $json" + pass "blank legacy summary names use their durable identifier" +} + +test_newest_filed_gates_are_selected_before_snapshot_bounds() { + local home mate fakebin json i + home=$(make_home newest-before-bounds) + : > "$home/data/secondmates.md" + printf '## In flight\n\n## Queued\n' > "$home/data/backlog.md" + i=1 + while [ "$i" -le 20 ]; do + printf -- '- [ ] old-%02d - Older gate %02d (repo: sample) (kind: ship) (since 2026-06-%02d)\n' \ + "$i" "$i" "$i" >> "$home/data/backlog.md" + i=$((i + 1)) + done + printf -- '- [ ] newest - Newest gate (repo: sample) (kind: ship) (since 2026-07-01)\n\n## Done\n' \ + >> "$home/data/backlog.md" + fakebin=$(make_fakebin "$home") + json=$(run "$home" "$fakebin" --json) + printf '%s' "$json" | jq -e ' + (.gates | length) == 20 and .gates[0].id == "newest" + and (.gates | any(.id == "old-01") | not) + ' >/dev/null || fail "the bearings gate bound dropped the newest filed row: $json" + + mate="$TMP_ROOT/newest-before-bounds-mate" + make_valid_secondmate_home bounded-mate "$mate" + : > "$home/data/backlog.md" + append_secondmate_registry "$home" bounded-mate "$mate" + cat > "$mate/data/backlog.md" <<'EOF' +## In flight + +## Queued +- [ ] mate-eligible - Eligible remote gate (repo: sample) (kind: ship) (since 2026-07-08) +- [ ] mate-call-one - Newer captain call (repo: sample) (kind: captain) (hold: choose one) (hold-kind: captain) (since 2026-07-10) +- [ ] mate-call-two - Newest captain call (repo: sample) (kind: captain) (hold: choose two) (hold-kind: captain) (since 2026-07-11) + +## Done +EOF + json=$(FM_SNAPSHOT_SECONDMATE_QUEUED=2 run "$home" "$fakebin" --json) + printf '%s' "$json" | jq -e ' + [.gates[].id] == ["mate-eligible"] + and (.decisions_open | any(.id == "bounded-mate/mate-call-one")) + and (.decisions_open | any(.id == "bounded-mate/mate-call-two")) + ' >/dev/null || fail "captain calls crowded eligible Charted work out of the bound: $json" + pass "newest filed gates are selected before snapshot bounds" +} + +# A captain scanning Underway must be able to tell WHICH task a row is, and the +# board orders Charted Next by the durable filed date, so both facts have to come +# out of fleet state rather than being invented at render time. +test_underway_and_gate_rows_carry_the_durable_name_and_filed_date() { + local home mate fakebin json + home=$(make_home durable-name-filed) + : > "$home/data/secondmates.md" + mate="$TMP_ROOT/durable-name-home" + make_valid_secondmate_home named-mate "$mate" + append_secondmate_registry "$home" named-mate "$mate" + mkdir -p "$home/projects/main-wt" + cat > "$home/data/backlog.md" <<'EOF' +## In flight +- [ ] main-ship - Rename the fleet board rows (repo: firstmate) (kind: ship) (since 2026-07-09) + +## Queued +- [ ] newer-gate - Filed later (repo: firstmate) (kind: ship) (since 2026-07-10) +- [ ] older-gate - Filed earlier (repo: firstmate) (kind: ship) (since 2026-07-01) +- [ ] undated-gate - Filed before dates were recorded (repo: firstmate) (kind: ship) + +## Done +EOF + fm_write_meta "$home/state/main-ship.meta" \ + "window=firstmate:fm-main-ship" "worktree=$home/projects/main-wt" "project=firstmate" \ + "harness=claude" "kind=ship" "mode=no-mistakes" + record_claude_state "$home/state" main-ship busy + printf 'working: no-mistakes review round 2\n' > "$home/state/main-ship.status" + + printf '## In flight\n' > "$mate/data/backlog.md" + printf -- '- [ ] mate-child - Tighten the ledger contract (repo: sample) (kind: ship) (since 2026-07-08)\n' \ + >> "$mate/data/backlog.md" + printf '\n## Queued\n\n## Done\n' >> "$mate/data/backlog.md" + mkdir -p "$mate/projects/mate-child" + fm_write_meta "$mate/state/mate-child.meta" \ + "window=firstmate:fm-mate-child" "worktree=$mate/projects/mate-child" "project=sample" \ + "harness=claude" "kind=ship" "mode=no-mistakes" + record_claude_state "$mate/state" mate-child busy + printf 'working: waiting on the pipeline\n' > "$mate/state/mate-child.status" + + fakebin=$(make_fakebin "$home") + json=$(run "$home" "$fakebin" --json) + printf '%s' "$json" | jq -e ' + (.in_flight | any(.id == "main-ship" + and .name == "Rename the fleet board rows" + and (.doing | type == "string") and (.doing | length) > 0 + and .doing != .name)) + and (.in_flight | any(.id == "named-mate/mate-child" + and .name == "Tighten the ledger contract" + and (.doing | type == "string") and (.doing | length) > 0 + and .doing != .name)) + and (.gates | any(.id == "newer-gate" and .filed == "2026-07-10")) + and (.gates | any(.id == "older-gate" and .filed == "2026-07-01")) + and (.gates | any(.id == "undated-gate" and .filed == null)) + ' >/dev/null || fail "durable Underway names or gate filed dates are missing: $json" + pass "Underway rows carry the durable task name and gates carry their filed date" +} + test_mixed_secondmate_roles_partial_state_and_captain_readiness() { local home fakebin hibit wheel sshhip ha canonical json home=$(make_home mixed-domain-regressions) @@ -3374,6 +3507,9 @@ test_main_unstructured_current_is_disclosed_with_structured_sibling test_main_orphan_counterfactual_meta_clears_inventory_warning test_working_captain_holds_keep_their_bucket_surfaces test_active_children_project_independent_of_home_captain_hold +test_nameless_legacy_summary_uses_its_durable_identifier +test_newest_filed_gates_are_selected_before_snapshot_bounds +test_underway_and_gate_rows_carry_the_durable_name_and_filed_date test_mixed_secondmate_roles_partial_state_and_captain_readiness test_main_captain_readiness_matches_secondmate_projection test_completed_scout_report_not_pending diff --git a/tests/fm-bootstrap.test.sh b/tests/fm-bootstrap.test.sh index c5d8a16d413..69d885a6f1b 100755 --- a/tests/fm-bootstrap.test.sh +++ b/tests/fm-bootstrap.test.sh @@ -5,8 +5,7 @@ # BOOTSTRAP_INFO fact, or completed bootstrap no-action fact and is silent when # all is well. firstmate consumes the exact 'MISSING: treehouse (install: ...)', # 'MISSING: tasks-axi (install: ...)', 'MISSING: quota-axi (install: ...)', -# 'MISSING: gh-axi (install: ...)', 'MISSING: chrome-devtools-axi (install: ...)', -# 'MISSING: lavish-axi (install: ...)', and +# 'MISSING: gh-axi', 'MISSING: chrome-devtools-axi', and PRESENTATION_UNAVAILABLE for Lavish. # 'BOOTSTRAP_INFO: ...' lines, so those contracts are pinned verbatim. The cases # are table-driven over the inputs that vary: whether `treehouse get --help` # advertises --lease, which (if any) tasks-axi version is on PATH, whether @@ -398,8 +397,8 @@ ROWS } test_lavish_axi_min_version() { - local label version mode case_dir fakebin out missing n - missing='MISSING: lavish-axi (install: npm install -g lavish-axi && lavish-axi setup hooks)' + local label version mode case_dir fakebin out unavailable n + unavailable='PRESENTATION_UNAVAILABLE: lavish-axi (requires >=0.1.62; install: npm install -g lavish-axi && lavish-axi setup hooks) - nonvisual work may proceed with plain-text decisions and reports; install or upgrade before using Lavish' n=0 while IFS='^' read -r label version mode; do [ -n "$label" ] || continue @@ -408,24 +407,29 @@ test_lavish_axi_min_version() { mkdir -p "$case_dir/home/config" printf '%s\n' manual > "$case_dir/home/config/backlog-backend" fakebin=$(make_fake_toolchain "$case_dir") + [ "$version" != absent ] || rm -f "$fakebin/lavish-axi" out=$(PATH="$fakebin:$BASE_PATH" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ - FM_FAKE_TREEHOUSE_LEASE_HELP=1 FM_FAKE_LAVISH_AXI_VERSION="$version" "$ROOT/bin/fm-bootstrap.sh") + FM_FAKE_TREEHOUSE_LEASE_HELP=1 FM_FAKE_LAVISH_AXI_VERSION="$version" "$ROOT/bin/fm-bootstrap.sh") \ + || fail "$label: optional presentation must not fail bootstrap" + assert_not_contains "$out" 'MISSING:' "$label: optional presentation must not block nonvisual dispatch" case "$mode" in empty) [ -z "$out" ] || fail "$label: expected silence, got: $out" ;; - missing) - [ "$out" = "$missing" ] || fail "$label: expected '$missing', got: $out" ;; + unavailable) + [ "$out" = "$unavailable" ] || fail "$label: expected '$unavailable', got: $out" ;; esac done <<'ROWS' +absent lavish-axi permits text fallback^absent^unavailable minimum lavish-axi version is accepted^0.1.62^empty newer lavish-axi patch is accepted^0.1.63^empty newer lavish-axi minor is accepted^0.2.0^empty newer lavish-axi major is accepted^1.0.0^empty -the patch just below the floor reports an upgrade^0.1.61^missing -much older lavish-axi minor reports an upgrade^0.0.9^missing -unparseable lavish-axi version reports an upgrade^lavish-axi development build^missing +the patch just below the floor permits text fallback^0.1.61^unavailable +much older lavish-axi minor permits text fallback^0.0.9^unavailable +unparseable lavish-axi version permits text fallback^lavish-axi development build^unavailable + ROWS - pass "bootstrap enforces lavish-axi minimum version" + pass "bootstrap permits nonvisual work without compatible lavish-axi and retains its presentation floor" } test_chrome_devtools_axi_min_version() { @@ -1984,6 +1988,9 @@ test_crew_dispatch_validation() { printf '%s\n' "$body" > "$case_dir/home/config/crew-dispatch.json" fakebin=$(make_fake_toolchain "$case_dir") add_real_jq "$fakebin" + # Schema acceptance uses a supported backend; the mismatch test owns refusals. + printf '%s\n' herdr > "$case_dir/home/config/backend" + fm_fake_exit0 "$fakebin" herdr out=$(PATH="$fakebin:$BASE_PATH" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ FM_FAKE_TREEHOUSE_LEASE_HELP=1 "$ROOT/bin/fm-bootstrap.sh") case "$mode" in @@ -2013,6 +2020,10 @@ pi max effort is accepted^{"rules":[{"when":"deep coding","use":{"harness":"pi", pi-signed max effort is accepted^{"rules":[{"when":"signed coding","use":{"harness":"pi-signed","model":"openai-codex/gpt-5.6-sol","effort":"max"}}]}^empty^ muse shared efforts are accepted^{"rules":[{"when":"muse low","use":{"harness":"muse","effort":"low"}},{"when":"muse medium","use":{"harness":"muse","effort":"medium"}},{"when":"muse high","use":{"harness":"muse","effort":"high"}},{"when":"muse xhigh","use":{"harness":"muse","effort":"xhigh"}},{"when":"muse max","use":{"harness":"muse","effort":"max"}}]}^empty^ unsupported muse ultra effort is flagged^{"rules":[{"when":"muse ultra","use":{"harness":"muse","effort":"ultra"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: muse:ultra +agy model profile is accepted^{"rules":[{"when":"agy work","use":{"harness":"agy","model":"gemini-3.8-flash-high"}}]}^empty^ +agy low medium high efforts are accepted^{"rules":[{"when":"agy low","use":{"harness":"agy","effort":"low"}},{"when":"agy medium","use":{"harness":"agy","effort":"medium"}},{"when":"agy high","use":{"harness":"agy","effort":"high"}}]}^empty^ +unsupported agy xhigh effort is flagged^{"rules":[{"when":"agy xhigh","use":{"harness":"agy","effort":"xhigh"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: agy:xhigh +unsupported agy max effort is flagged^{"rules":[{"when":"agy max","use":{"harness":"agy","effort":"max"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: agy:max unsupported opencode effort is flagged^{"rules":[{"when":"opencode work","use":{"harness":"opencode","model":"anthropic/claude-sonnet-4-5","effort":"high"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: opencode:high unsupported cursor effort is flagged^{"rules":[{"when":"composer work","use":{"harness":"cursor","model":"composer-2.5","effort":"high"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: cursor:high unsupported agy max effort is flagged^{"rules":[{"when":"gemini work","use":{"harness":"agy","model":"gemini-3-pro","effort":"max"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: agy:max diff --git a/tests/fm-brief.test.sh b/tests/fm-brief.test.sh index 3cfac063e8b..bee49113fae 100755 --- a/tests/fm-brief.test.sh +++ b/tests/fm-brief.test.sh @@ -294,6 +294,10 @@ brief_fingerprint() { test_no_issue_briefs_match_exact_goldens() { local home actual id + local PATH="$TMP_ROOT/golden-bin:$PATH" + mkdir -p "$TMP_ROOT/golden-bin" + printf '%s\n' '#!/usr/bin/env bash' 'echo "lavish-axi 0.1.62"' > "$TMP_ROOT/golden-bin/lavish-axi" + chmod +x "$TMP_ROOT/golden-bin/lavish-axi" home="$TMP_ROOT/no-issue-golden-home" actual="$TMP_ROOT/no-issue-golden.actual" write_registry "$home" @@ -1191,6 +1195,41 @@ test_scout_and_secondmate_load_decision_hold_policy() { pass "fm-brief.sh: investigation and visual-review completions load the shared decision policy" } +# A scout brief offers the Lavish review loop only when bootstrap confirms the +# supported lavish-axi floor at scaffold time; a missing or older build gets a +# text-report instruction instead, so a scout never drives a below-floor Lavish. +test_scout_lavish_line_follows_presentation_floor() { + local base label version expect case_dir fakebin brief n=0 + local hosting='you may host the Lavish review loop yourself' + local text_only='deliver your findings as a text report without Lavish' + base=$(fm_test_base_path_sans "${FM_TEST_BASE_PATH:-/usr/bin:/bin:/usr/sbin:/sbin}" lavish-axi) + while IFS='^' read -r label version expect; do + [ -n "$label" ] || continue + n=$((n + 1)) + case_dir="$TMP_ROOT/scout-lavish-$n" + mkdir -p "$case_dir/home/data" + fakebin=$(fm_fakebin "$case_dir") + [ "$version" = absent ] || fm_fake_version_tool "$fakebin" lavish-axi FM_FAKE_LAVISH_AXI_VERSION "$version" + PATH="$fakebin:$base" FM_HOME="$case_dir/home" \ + "$ROOT/bin/fm-brief.sh" scout-lavish alpha --scout >/dev/null \ + || fail "$label: scout scaffold failed" + brief="$case_dir/home/data/scout-lavish/brief.md" + if [ "$expect" = hosting ]; then + assert_grep "$hosting" "$brief" "$label: scout brief did not offer the Lavish review loop" + assert_no_grep "$text_only" "$brief" "$label: scout brief withheld Lavish from a compatible build" + else + assert_grep "$text_only" "$brief" "$label: scout brief did not ask for a text report" + assert_no_grep "$hosting" "$brief" "$label: scout brief offered a below-floor Lavish" + fi + done <<'ROWS' +lavish-axi at the floor^0.1.62^hosting +lavish-axi above the floor^0.2.0^hosting +lavish-axi just below the floor^0.1.61^text +absent lavish-axi^absent^text +ROWS + pass "fm-brief.sh: scout Lavish hosting follows the bootstrap lavish-axi floor" +} + # Scout and secondmate paths still scaffold well-formed briefs. test_scout_and_secondmate_scaffold() { local brief @@ -1200,8 +1239,6 @@ test_scout_and_secondmate_scaffold() { assert_present "$brief" "scout brief was not scaffolded" assert_grep "SCOUT task" "$brief" "scout brief must declare itself a scout task" assert_grep "report.md" "$brief" "scout brief must point at the report deliverable" - assert_grep "you may host the Lavish review loop yourself" "$brief" \ - "scout brief must mention the option to host a Lavish review loop" assert_grep "## Captain's intent" "$brief" "scout brief missing Captain's intent subsection" assert_grep "## Firstmate spec" "$brief" "scout brief missing Firstmate spec subsection" assert_grep "{FIRSTMATE_SPEC}" "$brief" "scout brief missing the spec placeholder" @@ -2420,4 +2457,5 @@ test_firstmate_repo_crew_persona_in_a_secondmate_home test_resolved_line_and_pr_attribution_guidance test_continue_branch_renders_setup_and_marker test_continue_branch_flag_validation +test_scout_lavish_line_follows_presentation_floor printf '\nall fm-brief tests passed\n' diff --git a/tests/fm-captain-hold-lifecycle.test.sh b/tests/fm-captain-hold-lifecycle.test.sh index 7d8f1dd8b74..6f02c3512e2 100755 --- a/tests/fm-captain-hold-lifecycle.test.sh +++ b/tests/fm-captain-hold-lifecycle.test.sh @@ -47,6 +47,19 @@ run_lavish() { # <home> <command args...> "$ROOT/bin/fm-procevent-lavish.sh" "$@" } +# The generic process-event runner, run against this suite's isolated home and +# its own claim root, so a review armed here can never contend with a real one. +run_procevent() { # <home> <command args...> + local home=$1 + shift + PATH="$home/fakebin:$PATH" REAL_TASKS_AXI="$TASKS_AXI_BIN" \ + FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$home" \ + FM_STATE_OVERRIDE="$home/state" FM_DATA_OVERRIDE="$home/data" \ + FM_CONFIG_OVERRIDE="$home/config" \ + FM_PROCEVENT_CLAIM_ROOT="$home/procevent-claims" \ + "$ROOT/bin/fm-procevent.sh" "$@" +} + run_bearings() { # <home> [extra args] local home=$1 shift @@ -90,9 +103,21 @@ configure_merged_github() { # <home> #!/usr/bin/env bash printf '%s\n' "$*" >> "$FM_TEST_GH_LOG" case "${1:-} ${2:-}" in - "pr view") printf '%s\n' 1111111111111111111111111111111111111111 ;; + "pr view") + case " $* " in + *statusCheckRollup*) + printf '%s\n' '{"state":"OPEN","isDraft":false,"mergeable":"MERGEABLE","mergeStateStatus":"CLEAN","headRefOid":"1111111111111111111111111111111111111111","baseRefName":"main","statusCheckRollup":[{"__typename":"CheckRun","name":"ci","status":"COMPLETED","conclusion":"SUCCESS"}]}' + ;; + *headRefOid*) printf '%s\n' 1111111111111111111111111111111111111111 ;; + esac + ;; + "pr merge") : > "$(dirname "$0")/../gh-merge-called"; printf 'merged:\n number: %s\n status: ok\n' "${3:-}" ;; "api graphql") - printf '%s\n' 'state=MERGED' 'merged=true' 'queued=false' 'base=main' + if [ -e "$(dirname "$0")/../gh-merge-called" ]; then + printf '%s\n' state=MERGED merged=true queued=false base=main default=main + else + printf '%s\n' state=OPEN merged=false queued=false base=main default=main + fi ;; esac SH @@ -100,11 +125,12 @@ SH #!/usr/bin/env bash printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" case "${1:-} ${2:-}" in - "pr merge") printf 'merged:\n number: %s\n status: ok\n' "${3:-}" ;; + "api "*) printf 'tip: 2222222222222222222222222222222222222222\n' ;; "pr view") printf 'pull_request:\n number: %s\n state: merged\n' "${3:-}" ;; esac SH chmod +x "$home/fakebin/gh" "$home/fakebin/gh-axi" + rm -f "$home/gh-merge-called" : > "$home/gh.log" : > "$home/gh-axi.log" } @@ -2175,6 +2201,63 @@ test_legacy_identities_keep_working() { pass "legacy identities, metadata, bindings, and the shim keep working" } +# A board answer must reach the keyed-answer intake through the RUNNER, not just +# through a hand-fed `answers` call. The captain answered ten calls on a bearings +# board, the board accepted them, and nothing collected them: the source that +# collects a board is a supervised process, and while it was not running the +# board went on presenting as armed. Everything between the captured result and +# the closed task is asserted here end to end - capture, the wake that tells +# firstmate to look, and the recorded answer - because each of those was intact +# on its own while the chain as a whole delivered nothing. +test_board_answer_reaches_the_keyed_answer_intake() { + local home sid stub out queue show + home=$(make_home board-channel) + sid=lavish-b0a4d0000000f1e2 + fm_test_track_procevent_home "$home" "$home/procevent-claims" + + run_captain "$home" hold sample-board-call --title "Choose the sample board route" \ + --reason "captain board route choice pending" --repo sample >/dev/null \ + || fail "could not register the board call" + + # One published Lavish poll response carrying the captain's structured answer, + # in the shape the adapter's own reader parses: a declared field order, an + # indented CSV row, and the versioned answer context inside its prompt. + stub="$home/board-source.sh" + cat > "$stub" <<'SH' +#!/usr/bin/env bash +cat <<'OUT' +session: + status: feedback + session_ended: false +prompts[1]{tag,text,prompt}: + "choice","Take the north route","Context data: {\"schema\":\"fm-bearings-answer.v1\",\"question\":\"sample-board-call\",\"selection\":\"north\",\"note\":\"\"}" +OUT +SH + chmod +x "$stub" + + run_procevent "$home" register lavish "$sid" -- "$stub" >/dev/null \ + || fail "could not register the board source" + run_captain "$home" bind "$sid" >/dev/null \ + || fail "could not bind the board source to the keyed-answer intake" + + out=$(run_procevent "$home" start "$sid" 2>&1) \ + || fail "the board source runner did not complete: $out" + assert_contains "$out" "$sid.1.result" "the board answer was never durably captured: $out" + assert_contains "$out" "answers-fed: $sid" \ + "the captured board answer never reached the keyed-answer intake: $out" + + queue=$(cat "$home/state/.wake-queue" 2>/dev/null || true) + assert_contains "$queue" "check: procevent lavish $sid 1" \ + "the captured board answer produced no wake: $queue" + + show=$(tasks_in "$home" show sample-board-call --full) + assert_contains "$show" "state: done" "the board answer did not close the captain call" + assert_contains "$show" "north" "the board answer lost the captain's selection" + assert_contains "$show" "the captured result $sid sequence 1" \ + "the recorded answer did not name the board result that carried it" + pass "a board answer reaches the keyed-answer intake and wakes firstmate" +} + # The intake is channel-agnostic, so chat must reach it the same way a captured # review does - for a task-id key, and for a legacy composed identity. test_chat_channel_feeds_the_same_keyed_answer_intake() { @@ -3144,14 +3227,14 @@ test_pr_merge_entrypoint_refuses_a_captain_held_task() { run_captain "$home" hold "$pr_id" --reason "captain merge approval pending" >/dev/null \ || fail "could not hold the PR entrypoint fixture" - # Without the entrypoint guard, this run reaches gh-axi and returns success - # even though the task is still held for the captain. + # Without the entrypoint guard, this run reaches gh and returns success even + # though the task is still held for the captain. set +e run_pr_merge "$home" "$pr_id" "$pr" > "$home/pr.out" 2> "$home/pr.err" rc=$? set -e [ "$rc" -ne 0 ] || fail "the PR merge entrypoint accepted a still-held task" - assert_no_grep 'pr merge 31 ' "$home/gh-axi.log" \ + assert_no_grep 'pr merge 31 ' "$home/gh.log" \ "the PR merge entrypoint reached the irreversible forge call for a held task" assert_grep "$pr_id is still held for the captain" "$home/pr.err" \ "the PR merge refusal did not name the held task" @@ -3219,7 +3302,7 @@ test_pr_merge_entrypoint_separates_an_unreadable_record_from_an_absent_one() { [ "$rc" -ne 0 ] || fail "the PR merge entrypoint accepted an unreadable captain-hold authority record" assert_grep "could not determine whether task $id is still held for the captain" "$home/missing-pr.err" \ "the PR merge refusal did not name its unreadable authority record" - assert_no_grep 'pr merge 43 ' "$home/gh-axi.log" \ + assert_no_grep 'pr merge 43 ' "$home/gh.log" \ "the PR merge entrypoint reached the forge without a readable authority record" # A home with no backlog at all records no captain calls, so nothing can be @@ -3227,7 +3310,7 @@ test_pr_merge_entrypoint_separates_an_unreadable_record_from_an_absent_one() { rm "$home/data/backlog.md" run_pr_merge "$home" "$id" "$pr" > "$home/absent-pr.out" 2> "$home/absent-pr.err" \ || fail "the PR merge entrypoint refused a home carrying no backlog" - merge_count=$(grep -c 'pr merge 43 ' "$home/gh-axi.log" || true) + merge_count=$(grep -c 'pr merge 43 ' "$home/gh.log" || true) [ "$merge_count" -eq 1 ] || fail "the absent backlog did not permit exactly one PR merge" pass "the PR merge entrypoint separates an unreadable authority record from an absent one" } @@ -3423,7 +3506,7 @@ test_merge_entrypoints_refuse_a_reused_task_incarnation() { # Without the pre-wait generation capture and locked comparison, the waiter # records and merges pull request 42 against the replacement task record. [ "$merge_rc" -ne 0 ] || fail "the PR merge accepted a replacement task incarnation" - assert_no_grep 'pr merge 42 ' "$home/gh-axi.log" \ + assert_no_grep 'pr merge 42 ' "$home/gh.log" \ "the PR merge reached the forge for a replacement task incarnation" assert_grep "changed incarnation while waiting to merge" "$home/reuse-merge.err" \ "the PR merge did not identify the replacement task incarnation" @@ -3613,7 +3696,7 @@ SH "PR cleanup was not refused by the merge's task control lock" [ "$merge_rc" -eq 0 ] || fail "the serialized PR merge failed after cleanup was refused" assert_present "$home/state/$id.meta" "the refused PR cleanup removed task metadata" - assert_grep 'pr merge 33 ' "$home/gh-axi.log" \ + assert_grep 'pr merge 33 ' "$home/gh.log" \ "the serialized PR merge did not reach the forge after cleanup was refused" local_home=$(make_home teardown-race-local-entrypoint) @@ -3813,6 +3896,7 @@ test_reconcile_closes_with_evidence_or_keeps_the_call_open test_reconcile_outcomes_retry_partial_failures_once test_unbound_source_closes_no_hold test_legacy_identities_keep_working +test_board_answer_reaches_the_keyed_answer_intake test_chat_channel_feeds_the_same_keyed_answer_intake test_origin_slug_validation_precedes_path_construction test_status_resolution_over_an_open_hold_is_signalled diff --git a/tests/fm-ci-workflow.test.sh b/tests/fm-ci-workflow.test.sh new file mode 100755 index 00000000000..fd2f7918493 --- /dev/null +++ b/tests/fm-ci-workflow.test.sh @@ -0,0 +1,159 @@ +#!/usr/bin/env bash +# Contract tests for .github/workflows/ci.yml's runner-spend safeguards. +# +# Origin: the 2026-09-12 GitHub Actions starvation incident. firstmate CI had no +# concurrency deduplication, so every superseded PR head kept its full job +# fan-out, and four jobs carried no timeout at all. These tests hold both +# safeguards: PR runs supersede within one PR while main pushes are never +# cancelled, and every CI job carries a finite hang tripwire. +# +# The workflow is parsed as YAML and its concurrency expressions are resolved +# against simulated pull_request and push contexts, so the assertions describe +# what GitHub would do, not how the file happens to be spelled. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +CI_WORKFLOW="$ROOT/.github/workflows/ci.yml" + +assert_present "$CI_WORKFLOW" ".github/workflows/ci.yml is missing" +command -v ruby >/dev/null 2>&1 \ + || fail "ruby is required to parse .github/workflows/ci.yml as YAML" + +# Resolve the workflow's concurrency contract under one simulated event and +# print "<group><TAB><cancel-in-progress>". Only the two expression constructs +# this workflow uses are resolved: an `a || b` fallback and an `==` comparison. +resolve_concurrency() { + local event=$1 pr_number=$2 run_id=$3 + ruby -ryaml -e ' +doc = YAML.load_file(ARGV[0]) +concurrency = doc.fetch("concurrency") +context = { + "github.workflow" => doc.fetch("name"), + "github.event_name" => ARGV[1], + "github.event.pull_request.number" => ARGV[2], + "github.run_id" => ARGV[3], +} + +value = lambda do |token| + token = token.strip + next token[1..-2] if token.start_with?("\x27") && token.end_with?("\x27") + raise "unresolvable context reference: #{token}" unless context.key?(token) + context.fetch(token) +end + +evaluate = lambda do |expression| + expression = expression.strip + if expression.include?("==") + left, right = expression.split("==", 2) + next value.call(left) == value.call(right) ? "true" : "false" + end + resolved = expression.split("||").map { |token| value.call(token) }.find { |v| !v.empty? } + resolved.to_s +end + +interpolate = lambda do |raw| + raw.to_s.gsub(/\$\{\{(.+?)\}\}/) { evaluate.call(Regexp.last_match(1)) } +end + +puts [interpolate.call(concurrency.fetch("group")), + interpolate.call(concurrency.fetch("cancel-in-progress"))].join("\t") +' "$CI_WORKFLOW" "$event" "$pr_number" "$run_id" +} + +job_timeout() { + ruby -ryaml -e ' +puts YAML.load_file(ARGV[0]).fetch("jobs").fetch(ARGV[1]).fetch("timeout-minutes", "none") +' "$CI_WORKFLOW" "$1" +} + +group_of() { printf '%s\n' "$1" | cut -f1; } +cancel_of() { printf '%s\n' "$1" | cut -f2; } + +test_pr_pushes_supersede_within_one_pr() { + local first second + first=$(resolve_concurrency pull_request 108 900001) || fail "could not resolve PR concurrency" + second=$(resolve_concurrency pull_request 108 900002) || fail "could not resolve PR concurrency" + [ "$(group_of "$first")" = "$(group_of "$second")" ] \ + || fail "two runs of one PR must share a concurrency group, got $(group_of "$first") and $(group_of "$second")" + [ "$(cancel_of "$first")" = true ] \ + || fail "PR runs must cancel the in-progress run, got $(cancel_of "$first")" + pass "a newer push to one PR supersedes that PR's in-flight CI" +} + +test_separate_prs_do_not_cancel_each_other() { + local one two + one=$(resolve_concurrency pull_request 108 900001) || fail "could not resolve PR concurrency" + two=$(resolve_concurrency pull_request 109 900003) || fail "could not resolve PR concurrency" + [ "$(group_of "$one")" != "$(group_of "$two")" ] \ + || fail "distinct PRs must not share a concurrency group ($(group_of "$one"))" + pass "distinct PRs get distinct concurrency groups" +} + +test_main_pushes_are_never_cancelled() { + local first second + first=$(resolve_concurrency push '' 900010) || fail "could not resolve push concurrency" + second=$(resolve_concurrency push '' 900011) || fail "could not resolve push concurrency" + [ "$(group_of "$first")" != "$(group_of "$second")" ] \ + || fail "each main push must get its own concurrency group, got $(group_of "$first") twice" + [ "$(cancel_of "$first")" = false ] \ + || fail "push runs must never cancel an in-progress run, got $(cancel_of "$first")" + pass "every main push keeps its own group and is never cancelled" +} + +test_every_job_has_a_finite_timeout() { + local reported + reported=$(ruby -ryaml -e ' +YAML.load_file(ARGV[0]).fetch("jobs").each do |name, job| + timeout = job["timeout-minutes"] + next if timeout.is_a?(Integer) && timeout > 0 + puts "#{name}: #{timeout.inspect}" +end +' "$CI_WORKFLOW") || fail "could not read job timeouts from ci.yml" + [ -z "$reported" ] || fail "these CI jobs have no finite hang tripwire:"$'\n'"$reported" + pass "every ci.yml job carries a finite timeout" +} + +# The four jobs the incident found unbounded, at the report's recommended caps. +test_previously_unbounded_jobs_keep_their_caps() { + local job expected actual + while read -r job expected; do + [ -n "$job" ] || continue + actual=$(job_timeout "$job") || fail "could not read the $job timeout" + [ "$actual" = "$expected" ] \ + || fail "$job timeout must stay $expected minutes, got $actual" + done <<'CAPS' +lint 25 +test-coverage 5 +tests-timing-aggregate 5 +invariants 5 +CAPS + pass "the incident's unbounded jobs keep their recommended caps" +} + +# Cancellation makes an undersized cap costlier: a falsely tripped job now also +# discards a run nobody replaced. These bounds were measured, not guessed. +test_measured_lanes_keep_their_existing_bounds() { + local job expected actual + while read -r job expected; do + [ -n "$job" ] || continue + actual=$(job_timeout "$job") || fail "could not read the $job timeout" + [ "$actual" = "$expected" ] \ + || fail "$job timeout must stay $expected minutes, got $actual" + done <<'CAPS' +tests-portable-parallel-1 10 +tests-portable-parallel-2 10 +tests-portable-serial 30 +tests-herdr 75 +macos-stock-bash 10 +CAPS + pass "the already-measured lane bounds are unchanged" +} + +test_pr_pushes_supersede_within_one_pr +test_separate_prs_do_not_cancel_each_other +test_main_pushes_are_never_cancelled +test_every_job_has_a_finite_timeout +test_previously_unbounded_jobs_keep_their_caps +test_measured_lanes_keep_their_existing_bounds diff --git a/tests/fm-claude-trust.test.sh b/tests/fm-claude-trust.test.sh index 94211e0e4b4..c040682516d 100755 --- a/tests/fm-claude-trust.test.sh +++ b/tests/fm-claude-trust.test.sh @@ -1,9 +1,10 @@ #!/usr/bin/env bash # Behavior tests for bin/fm-claude-trust.sh and the claude spawn that calls it. # -# Both halves of the contract are load-bearing and both are proven here: a -# legitimate fresh task worktree is trusted so a claude worker reaches its -# brief with no human, and every out-of-scope path is REFUSED rather than +# Both halves of the contract are load-bearing and both are proven here, for +# each directory a claude launch can start in: a legitimate fresh task worktree +# and a seeded secondmate home are trusted so the agent reaches its brief or +# charter with no human, and every out-of-scope path is REFUSED rather than # warned about or quietly skipped. set -u @@ -67,6 +68,39 @@ assert_store_value() { # <store> <expected-json> <msg> <key...> [ "$actual" = "$expected" ] || fail "$msg (expected $expected, got $actual)" } +# assert_all_flags <store> <path> <msg>: all three registered flags - trust, +# external-includes approved, external-includes warning-shown - are true on +# the project entry at <path>. The external-imports flags are the ones the +# running app reads only from the PROJECT-root entry, never the worktree +# entry, so this is what actually proves the dialog is suppressed. +assert_all_flags() { + local store=$1 key=$2 msg=$3 + node -e ' + const j=JSON.parse(require("node:fs").readFileSync(process.argv[1],"utf8")); + const e=(j.projects||{})[process.argv[2]]||{}; + const flags=["hasTrustDialogAccepted","hasClaudeMdExternalIncludesApproved","hasClaudeMdExternalIncludesWarningShown"]; + process.exit(flags.every((f)=>e[f]===true)?0:1); + ' "$store" "$key" || fail "$msg" +} + +# assert_trust_only_no_import_consent <store> <path> <msg>: the entry at +# <path> carries hasTrustDialogAccepted===true but NEITHER external-imports +# flag is true - the shape a registration must leave behind when the project +# entry had no prior explicit "Yes, allow" for external CLAUDE.md imports, so +# a spawn never manufactures that consent from an absent flag. +assert_trust_only_no_import_consent() { + local store=$1 key=$2 msg=$3 + node -e ' + const j=JSON.parse(require("node:fs").readFileSync(process.argv[1],"utf8")); + const e=(j.projects||{})[process.argv[2]]||{}; + const trustOk = e.hasTrustDialogAccepted === true; + const noImportConsent = + e.hasClaudeMdExternalIncludesApproved !== true && + e.hasClaudeMdExternalIncludesWarningShown !== true; + process.exit(trustOk && noImportConsent ? 0 : 1); + ' "$store" "$key" || fail "$msg" +} + # A PATH carrying the tools the scope test needs but no node, so the # missing-interpreter path is exercised without disturbing the real PATH. node_free_path() { # <case-dir> -> a bin dir holding the script's own tools but no node @@ -78,6 +112,53 @@ node_free_path() { # <case-dir> -> a bin dir holding the script's own tools but printf '%s\n' "$dir" } +# --- secondmate homes ------------------------------------------------------- + +# seed_secondmate_home <home> <id> [shape]: the on-disk shape bin/fm-home-seed.sh +# leaves behind - the identity marker, the firstmate instance files, the four +# operational directories, and a charter for the launch to carry. "clone" (the +# default) is the standalone-clone home an explicit ~/fm-homes/<id> path +# produces, a primary checkout of the firstmate repo; "worktree" is the linked +# worktree a treehouse lease produces. Both shapes are real homes, so both must +# be trusted. +seed_secondmate_home() { + local home=$1 id=$2 shape=${3:-clone} src + case "$shape" in + worktree) + src="$home.src" + fm_git_worktree "$src" "$home" "sm-$id" + ;; + *) + mkdir -p "$home" + fm_git_init_commit "$home" + ;; + esac + mkdir -p "$home/bin" "$home/data" "$home/state" "$home/config" "$home/projects" + printf '# Firstmate\n' > "$home/AGENTS.md" + printf 'charter\n' > "$home/data/charter.md" + printf '%s\n' "$id" > "$home/.fm-secondmate-home" +} + +# run_home_trust <config> <home> <id> [user-home]: invoke the secondmate-home +# mode against an isolated store. +run_home_trust() { + local config=$1 home=$2 id=$3 user_home=${4:-$1} + CLAUDE_CONFIG_DIR="$config" HOME="$user_home" "$TRUST" --secondmate-home "$home" "$id" 2>&1 +} + +# spawn_secondmate_claude <case-dir> <home> <id>: run a real --secondmate claude +# spawn against the isolated store at <case-dir>/claude-config, logging the +# launch to <case-dir>/launch.log. Echoes the spawn output. +spawn_secondmate_claude() { + local case_dir=$1 home=$2 id=$3 primary fakebin + primary="$case_dir/primary" + mkdir -p "$case_dir/claude-config" + fakebin=$(make_spawn_fakebin "$case_dir/fake" claude) + fm_test_spawn_home "$primary" claude + FM_TEST_CLAUDE_CONFIG_DIR="$case_dir/claude-config" FM_FAKE_LAUNCH_LOG="$case_dir/launch.log" \ + fm_test_run_spawn "$primary" "$home" "$fakebin" "$id" "$home" claude --secondmate +} + test_fresh_worktree_is_trusted() { local rec out rec=$(make_case fresh) @@ -92,6 +173,97 @@ test_fresh_worktree_is_trusted() { pass "fm-claude-trust.sh: a fresh task worktree is trusted" } +# The trust dialog is read only from the PROJECT-root entry, never the +# worktree entry (Claude Code's own git-root canonicalization collapses every +# linked worktree to its primary checkout for that check, with no +# ancestor-walk fallback the way the trust check has), so this proves both +# entries carry the trust flag after one registration. External-imports +# consent is a SEPARATE grant this script never manufactures: on a genuinely +# fresh project (no prior interactive answer at all) neither entry may carry +# hasClaudeMdExternalIncludesApproved or hasClaudeMdExternalIncludesWarningShown +# - see test_registration_carries_forward_existing_import_consent below for +# the case where the project already said yes. +test_fresh_worktree_also_trusts_the_project_root_without_import_consent() { + local rec out + rec=$(make_case fresh-project) + read_case "$rec" + out=$(run_trust "$CONFIG" "$WT" "$PROJ") + expect_code 0 $? "a fresh linked worktree must be trusted: $out" + assert_contains "$out" "$PROJ" "registration did not report the project root it also trusted" + assert_trust_only_no_import_consent "$CONFIG/.claude.json" "$WT" \ + "the worktree entry either lost trust or gained unearned import consent" + assert_trust_only_no_import_consent "$CONFIG/.claude.json" "$PROJ" \ + "the project-root entry either lost trust or gained unearned import consent" + pass "fm-claude-trust.sh: a fresh registration trusts the project root without manufacturing import consent" +} + +# The Greptile-flagged regression this pins: a project entry that already +# carries an explicit "Yes, allow" (hasClaudeMdExternalIncludesApproved===true) +# is exactly the standing consent this script may refresh - and refreshing it +# is what actually suppresses the external-imports dialog for the worker, +# since that check reads only the project entry (see the disassembly note at +# the top of fm-claude-trust.sh), never the worktree one. +test_registration_carries_forward_existing_import_consent() { + local rec store + rec=$(make_case import-consent-carried) + read_case "$rec" + store="$CONFIG/.claude.json" + cat > "$store" <<JSON +{"hasCompletedOnboarding":true,"projects":{"$PROJ":{"hasTrustDialogAccepted":true,"hasClaudeMdExternalIncludesApproved":true,"hasClaudeMdExternalIncludesWarningShown":true}}} +JSON + run_trust "$CONFIG" "$WT" "$PROJ" >/dev/null || fail "registration failed against a project that already approved external imports" + assert_all_flags "$store" "$WT" \ + "the worktree entry did not carry the refreshed import consent" + assert_all_flags "$store" "$PROJ" \ + "the project-root entry lost its own already-granted import consent" + pass "fm-claude-trust.sh: carries forward a project's already-granted import consent to the worktree entry" +} + +# The project-root entry is the same store the launching user's interactive +# claude sessions read and write (it is usually already present, carrying +# unrelated keys such as allowedTools or MCP config), so preservation must +# hold there exactly as it holds for the worktree entry. +test_project_root_entry_preserves_other_keys() { + local rec store + rec=$(make_case project-preserve) + read_case "$rec" + store="$CONFIG/.claude.json" + cat > "$store" <<JSON +{"hasCompletedOnboarding":true,"projects":{"$PROJ":{"hasTrustDialogAccepted":false,"allowedTools":["Read"]}}} +JSON + run_trust "$CONFIG" "$WT" "$PROJ" >/dev/null || fail "registration failed against an existing project entry" + assert_trust_only_no_import_consent "$store" "$PROJ" \ + "the project-root entry did not gain trust, or gained unearned import consent it had never been asked for" + assert_store_value "$store" '["Read"]' "the project entry's unrelated settings were lost" projects "$PROJ" allowedTools + pass "fm-claude-trust.sh: preserves unrelated keys on the project-root entry" +} + +# hasClaudeMdExternalIncludesApproved===false on the project-root entry is a +# human's explicit "No, disable" answer, recorded in the SAME store their own +# interactive sessions read. A spawn must never flip that to true on their +# behalf: doing so would grant every later interactive session in that +# checkout silent external-file inclusion the human declined. The whole +# registration refuses instead, and the store - including the worktree entry, +# which is never reached - must come back byte-for-byte unchanged. +test_project_root_entry_declined_external_imports_is_not_overridden() { + local rec store out before after + rec=$(make_case project-decline) + read_case "$rec" + store="$CONFIG/.claude.json" + cat > "$store" <<JSON +{"hasCompletedOnboarding":true,"projects":{"$PROJ":{"hasTrustDialogAccepted":true,"hasClaudeMdExternalIncludesApproved":false,"hasClaudeMdExternalIncludesWarningShown":true,"allowedTools":["Read"]}}} +JSON + before=$(cat "$store") + out=$(run_trust "$CONFIG" "$WT" "$PROJ") + expect_code 1 $? "a project that already declined external imports must be refused: $out" + assert_contains "$out" "declined external CLAUDE.md imports" \ + "the refusal did not name the declined-consent reason" + after=$(cat "$store") + [ "$before" = "$after" ] || fail "the store was modified despite the refusal" + assert_not_trusted "$store" "$WT" "the worktree entry was registered despite the refusal" + pass "fm-claude-trust.sh: refuses to override a project's declined external-imports consent" +} + test_registration_is_idempotent() { local rec out count rec=$(make_case idempotent) @@ -264,6 +436,29 @@ test_worktree_subdirectory_is_refused() { pass "fm-claude-trust.sh: refuses a subdirectory of the worktree" } +# The write target the external-imports flags depend on is only correct when +# it names the primary checkout. When <project> is itself a linked worktree +# (a secondmate home spawned from, rather than as, the primary checkout), +# writing the flags at that worktree's own path would land them at a key +# Claude Code's git-root canonicalization never reads, silently reproducing +# the bug this script exists to close - so this resolves the argument +# structurally to its primary checkout instead of refusing it. +test_project_argument_that_is_itself_a_worktree_resolves_to_the_primary_checkout() { + local rec out proj_wt + rec=$(make_case nested-project) + read_case "$rec" + proj_wt="$CASE_DIR/proj-wt" + git -C "$PROJ" worktree add --quiet -b wt-proj-wt "$proj_wt" + out=$(run_trust "$CONFIG" "$WT" "$proj_wt") + expect_code 0 $? "a project argument that is itself a linked worktree must resolve to its primary checkout: $out" + assert_contains "$out" "$PROJ" "the outcome did not name the resolved primary checkout" + assert_trust_only_no_import_consent "$CONFIG/.claude.json" "$PROJ" \ + "the resolved primary checkout either lost trust or gained unearned import consent" + assert_not_trusted "$CONFIG/.claude.json" "$proj_wt" \ + "the linked worktree argument itself was recorded as the project root" + pass "fm-claude-trust.sh: a project argument that is itself a linked worktree resolves to the primary checkout" +} + test_unrelated_store_content_is_preserved() { local rec store rec=$(make_case preserve) @@ -440,7 +635,167 @@ test_claude_spawn_pretrusts_its_worktree_and_reaches_the_brief() { pass "fm-spawn.sh: a claude spawn pre-trusts its worktree and launches with the brief" } +# A secondmate home is the second directory a claude launch starts in, and it is +# as unseen by Claude as a fresh worktree. The standalone-clone shape is the one +# that wedged in production: the trust step was skipped for every secondmate, so +# nothing was registered and the pane stopped on the dialog before it read its +# charter. +test_secondmate_standalone_clone_home_is_trusted() { + local case_dir home out + case_dir="$TMP_ROOT/sm-clone-spawn" + home="$case_dir/fm-homes/nomistakes-n1" + seed_secondmate_home "$home" nomistakes-n1 clone + out=$(spawn_secondmate_claude "$case_dir" "$home" nomistakes-n1) + expect_code 0 $? "a claude secondmate spawn into a standalone-clone home must succeed: $out" + assert_trusted "$case_dir/claude-config/.claude.json" "$home" \ + "the claude secondmate spawn did not pre-register trust for its standalone-clone home" + assert_present "$case_dir/launch.log" "the claude secondmate spawn sent no launch command" + assert_grep 'claude --dangerously-skip-permissions' "$case_dir/launch.log" \ + "the launch command was not the claude secondmate launch" + assert_grep "$home/data/charter.md" "$case_dir/launch.log" \ + "the launch command did not carry the charter the secondmate must read" + # The pane must read the SAME store the registration wrote, or the trust would + # land somewhere it never looks and the dialog would appear anyway. + assert_grep "CLAUDE_CONFIG_DIR='$case_dir/claude-config'" "$case_dir/launch.log" \ + "the launch command did not point the secondmate at the store that was trusted" + pass "fm-spawn.sh: a claude secondmate spawn pre-trusts a standalone-clone home" +} + +# The other seeded shape, a treehouse-leased linked worktree. It must be trusted +# through the same seed evidence rather than incidentally, so the registration +# does not depend on which shape the home happens to have. +test_secondmate_leased_worktree_home_is_trusted() { + local case_dir home out + case_dir="$TMP_ROOT/sm-leased-spawn" + home="$case_dir/leased/home" + mkdir -p "$case_dir/leased" + seed_secondmate_home "$home" leased-n1 worktree + out=$(spawn_secondmate_claude "$case_dir" "$home" leased-n1) + expect_code 0 $? "a claude secondmate spawn into a leased worktree home must succeed: $out" + assert_trusted "$case_dir/claude-config/.claude.json" "$home" \ + "the claude secondmate spawn did not pre-register trust for its leased worktree home" + pass "fm-spawn.sh: a claude secondmate spawn pre-trusts a leased worktree home" +} + +# The seed is the whole security boundary for home-level trust, so every path +# that is not a home seeded for THIS secondmate is refused and left untrusted. +# Each row drives one structural property apart from a genuine home. +test_secondmate_home_trust_refuses_everything_unseeded() { + local case_dir config home target out + case_dir="$TMP_ROOT/sm-refusals" + config="$case_dir/claude-config" + mkdir -p "$config" + + # A plain directory: no marker at all. + target="$case_dir/plain" + mkdir -p "$target" + out=$(run_home_trust "$config" "$target" plain-n1) + expect_code 1 $? "a plain directory must be refused: $out" + assert_contains "$out" "no .fm-secondmate-home marker" "the refusal did not name the missing marker" + assert_not_trusted "$config/.claude.json" "$target" "a plain directory was trusted" + + # A firstmate checkout that was never seeded as a secondmate home: every other + # structural signal matches and only the marker is missing. + target="$case_dir/checkout" + seed_secondmate_home "$target" checkout-n1 clone + rm -f "$target/.fm-secondmate-home" + out=$(run_home_trust "$config" "$target" checkout-n1) + expect_code 1 $? "an unseeded firstmate checkout must be refused: $out" + assert_contains "$out" "no .fm-secondmate-home marker" "the refusal did not name the missing marker" + assert_not_trusted "$config/.claude.json" "$target" "an unseeded firstmate checkout was trusted" + + # A home seeded for a DIFFERENT secondmate: one home's trust must not be + # granted while spawning another id. + target="$case_dir/other-mate" + seed_secondmate_home "$target" other-n1 clone + out=$(run_home_trust "$config" "$target" wanted-n1) + expect_code 1 $? "a home marked for another secondmate must be refused: $out" + assert_contains "$out" "other-n1" "the refusal did not name the id the home is marked for" + assert_not_trusted "$config/.claude.json" "$target" "a home marked for another secondmate was trusted" + + # A marker that is a symlink: another file's bytes must not stand in for the + # seed, even when they read as the right id. + target="$case_dir/linked-marker" + seed_secondmate_home "$target" linked-n1 clone + printf 'linked-n1\n' > "$case_dir/planted-id" + ln -sf "$case_dir/planted-id" "$target/.fm-secondmate-home" + out=$(run_home_trust "$config" "$target" linked-n1) + expect_code 1 $? "a symlinked marker must be refused: $out" + assert_contains "$out" "symlink" "the refusal did not name the symlinked marker" + assert_not_trusted "$config/.claude.json" "$target" "a home whose marker is a symlink was trusted" + + # An operational directory that escapes the home: the home's own working + # surface must stay inside it. + target="$case_dir/escaping" + seed_secondmate_home "$target" escaping-n1 clone + rm -rf "$target/projects" + mkdir -p "$case_dir/elsewhere" + ln -s "$case_dir/elsewhere" "$target/projects" + out=$(run_home_trust "$config" "$target" escaping-n1) + expect_code 1 $? "a home whose operational directory escapes it must be refused: $out" + assert_contains "$out" "outside the home" "the refusal did not name the escaping directory" + assert_not_trusted "$config/.claude.json" "$target" "a home whose projects/ escapes it was trusted" + + # The user's own home directory, seeded to prove the marker alone cannot carry + # it: HOME is refused in this mode exactly as it is for a worktree. + target="$case_dir/user-home" + seed_secondmate_home "$target" userhome-n1 clone + out=$(run_home_trust "$config" "$target" userhome-n1 "$target") + expect_code 1 $? "the user's home directory must be refused: $out" + assert_contains "$out" "home directory" "the refusal did not name the home directory" + assert_not_trusted "$config/.claude.json" "$target" "the user's home directory was trusted" + # Prove the seed really would have been accepted, so the guard above is what + # refused rather than an unrelated failure. + out=$(run_home_trust "$config" "$target" userhome-n1 "$case_dir/elsewhere-home") + expect_code 0 $? "the same seeded home must be accepted once it is not HOME: $out" + + pass "fm-claude-trust.sh: home-level trust is refused for everything but a home seeded for this secondmate" +} + +# A secondmate home is not a linked worktree, so worktree mode must keep +# refusing it rather than quietly widening to cover the new case. +test_worktree_mode_still_refuses_a_secondmate_home() { + local case_dir config home out + case_dir="$TMP_ROOT/sm-wrong-mode" + config="$case_dir/claude-config" + home="$case_dir/home" + mkdir -p "$config" + seed_secondmate_home "$home" mode-n1 clone + out=$(run_trust "$config" "$home" "$home") + expect_code 1 $? "worktree mode must still refuse a standalone-clone home: $out" + assert_contains "$out" "primary checkout" "the refusal did not name the primary checkout" + assert_not_trusted "$config/.claude.json" "$home" "worktree mode trusted a standalone-clone home" + pass "fm-claude-trust.sh: worktree mode still refuses a secondmate home" +} + +# The fail-closed half for secondmates: when the home's trust genuinely cannot be +# recorded, the spawn must refuse rather than launch a pane that would wedge on +# the dialog. This is the guard that never fired while the step was skipped. +test_secondmate_spawn_fails_closed_when_home_trust_cannot_be_recorded() { + local case_dir home out + case_dir="$TMP_ROOT/sm-failclosed" + home="$case_dir/fm-homes/failclosed-n1" + # Root owns /etc/passwd, so a store resolving to it is refused as another + # user's file. Running as root would own it and make the refusal vacuous. + if [ "$(id -u)" = 0 ]; then + pass "fm-spawn.sh: a claude secondmate spawn refuses when home trust cannot be recorded (skipped as root)" + return 0 + fi + seed_secondmate_home "$home" failclosed-n1 clone + mkdir -p "$case_dir/claude-config" + ln -s /etc/passwd "$case_dir/claude-config/.claude.json" + out=$(spawn_secondmate_claude "$case_dir" "$home" failclosed-n1) + expect_code 1 $? "a secondmate spawn whose trust registration is refused must fail: $out" + assert_contains "$out" "workspace trust" "the spawn did not report the trust refusal" + assert_absent "$case_dir/launch.log" "a secondmate was launched into a home whose trust could not be recorded" + pass "fm-spawn.sh: a claude secondmate spawn refuses when home trust cannot be recorded" +} + test_fresh_worktree_is_trusted +test_fresh_worktree_also_trusts_the_project_root_without_import_consent +test_registration_carries_forward_existing_import_consent +test_project_root_entry_preserves_other_keys +test_project_root_entry_declined_external_imports_is_not_overridden test_registration_is_idempotent test_primary_checkout_is_refused test_cdpath_cannot_defeat_the_primary_checkout_refusal @@ -452,6 +807,7 @@ test_non_git_directory_is_refused test_missing_directory_is_refused test_foreign_project_worktree_is_refused test_worktree_subdirectory_is_refused +test_project_argument_that_is_itself_a_worktree_resolves_to_the_primary_checkout test_unrelated_store_content_is_preserved test_symlinked_store_to_a_foreign_owned_target_is_refused test_symlinked_store_to_an_owned_target_is_accepted @@ -460,3 +816,8 @@ test_missing_node_is_refused test_scope_refusal_stays_fail_closed_without_node test_claude_spawn_pretrusts_its_worktree_and_reaches_the_brief test_refused_spawn_leaves_no_task_state +test_secondmate_standalone_clone_home_is_trusted +test_secondmate_leased_worktree_home_is_trusted +test_secondmate_home_trust_refuses_everything_unseeded +test_worktree_mode_still_refuses_a_secondmate_home +test_secondmate_spawn_fails_closed_when_home_trust_cannot_be_recorded diff --git a/tests/fm-control-herdr-smoke.test.sh b/tests/fm-control-herdr-smoke.test.sh index 29799fd7bf1..32bd3d25339 100755 --- a/tests/fm-control-herdr-smoke.test.sh +++ b/tests/fm-control-herdr-smoke.test.sh @@ -9,13 +9,12 @@ # an agent is running, and therefore whether a lifecycle verb may act at all, # comes from herdr's own agent registry. # -# No real agent is launched. herdr's `pane report-agent` is the same registry -# the adapter reads, so registering and not registering an agent on a plain -# shell pane exercises exactly the classification the control plane gates on. -# A registration over a plain idle shell is the STALE-registration state (the -# registry outliving its agent), which the classifier must downgrade to -# agent-free; modelling a genuinely live agent therefore additionally needs a -# real foreground process in the pane, so the idle-shell cross-check refuses. +# No real harness is launched. herdr's `pane report-agent` is the same registry +# the adapter reads, and a symlink named like a harness is the same process +# identity the adapter proves through `pane process-info`, so registering an +# agent over a real agent-named process, over a plain shell, and not at all +# exercises exactly the classification the control plane gates on - including +# the registration Herdr keeps after the agent process is gone (issue #4115). # # Always runs on a private, named, throwaway lab session, never the default # one (tests/herdr-test-safety.sh; the 2026-07-02 incident). Skips cleanly @@ -37,8 +36,11 @@ herdr_forget_inherited_pane SESSION="fm-lab-control-smoke-$$" export HERDR_SESSION="$SESSION" SCRATCH= +CLEANED=0 cleanup_all() { local status=0 + [ "$CLEANED" = 0 ] || return 0 + CLEANED=1 herdr_safe_stop_and_delete "$SESSION" || { echo "cleanup: herdr teardown failed" >&2; status=1; } if [ -n "$SCRATCH" ]; then rm -rf "$SCRATCH" 2>/dev/null || { echo "cleanup: scratch removal failed" >&2; status=1; } @@ -207,25 +209,87 @@ case "$OUT" in esac pass "real herdr: interrupt refuses when herdr's own agent registry reports no agent" -# --- a registered agent over a plain idle shell: a STALE registration ------- +# --- a registered agent WITH a live process: classification flips ------------ # -# This is the live blocker fixed by task -# fm-control-classifies-shell-as-live-agent: herdr's registry can outlive its -# agent (a hook-authoritative harness never deregisters on exit, and herdr has -# no agent-deregister verb), so a registration whose pane provably holds only -# a lone idle shell must classify agent-free, making exit idempotent success -# and recovery available. Before the fix this exact state read `alive` -# forever and every recovery verb refused. `pane report-agent` on the plain -# shell pane reproduces that state against the real registry: registered, but -# with no agent process behind it. +# A registration alone no longer proves an agent (issue #4115): the adapter +# verifies the pane's processes through the real `pane process-info` view. So +# the registered agent is backed by a real agent-named foreground process - a +# symlink to a long-running system binary named `claude`, the same construction +# tests/fm-tmux-agent-liveness.test.sh uses (a copied platform binary fails code +# signing on macOS arm64; the symlink name is what the kernel records as argv[0]). +AGENT_BIN="$SCRATCH/agentbin" +mkdir -p "$AGENT_BIN" +SLEEP_BIN=$(command -v sleep) || fail "sleep not found" +ln -s "$SLEEP_BIN" "$AGENT_BIN/claude" +printf -v AGENT_Q '%q' "$AGENT_BIN/claude" + +wait_process_state() { # <expected> <tries> + local expected=$1 tries=$2 i=0 + while [ "$i" -lt "$tries" ]; do + [ "$(fm_backend_herdr_pane_process_state "$SESSION" "$PANE_ID")" != "$expected" ] || return 0 + sleep 0.1 + i=$((i + 1)) + done + return 1 +} + +start_agent_process() { + fm_backend_herdr_send_text_line "$SESSION:$PANE_ID" "$AGENT_Q 900" \ + || fail "could not start the agent-named foreground process in the task pane" + wait_process_state agent 50 \ + || version_fail "a real agent-named foreground process reads '$(fm_backend_herdr_pane_process_state "$SESSION" "$PANE_ID")' rather than 'agent' through pane process-info" +} +start_agent_process herdr pane report-agent "$PANE_ID" --source fm-control-smoke --agent fm-control-smoke-agent \ --state idle --session "$SESSION" >/dev/null 2>&1 \ || fail "could not register an agent on the task pane" STATE=$(fm_backend_agent_state herdr "$SESSION:$PANE_ID") -[ "$STATE" = dead ] || fail "a registered agent over a provably lone idle shell is a stale registration and must classify dead, got '$STATE'" -pass "real herdr: a registration with no agent process behind it classifies dead (stale), not alive" +[ "$STATE" = alive ] || fail "herdr should classify a registered agent with a live process as alive, got '$STATE'" + +OUT=$(run_control hsmoke interrupt) || fail "interrupt against a live agent should succeed: $OUT" +case "$OUT" in + *"interrupt-delivered hsmoke harness=claude backend=herdr verified=agent-alive cancel=unconfirmed"*) : ;; + *) fail "interrupt should report the agent-alive proof on herdr, got: $OUT" ;; +esac +pass "real herdr: interrupt delivers the harness's key and proves the agent survived it" + +herdr pane get "$PANE_ID" --session "$SESSION" >/dev/null 2>&1 \ + || fail "the control plane must never remove the endpoint it was operating on" +[ -d "$WT" ] || fail "the control plane must never remove the task's local copy" +pass "real herdr: no control verb removed the endpoint or the task's local copy" + +# --- the stale registration (issue #4115): the agent process is gone, the --- +# --- record is not, and recovery must proceed anyway ------------------------ +# +# Stopping the agent-named process leaves the pane a plain shell while Herdr +# keeps the registration, which is exactly the shape a Pi crew leaves behind +# when it exits under a nested shell. Before the fix this read `alive` forever: +# exit waited out its timeout and refused, and relaunch was refused for good. +# This runs BEFORE the fail-closed exit case below, whose typed exit command +# stays buffered in the pane's tty while the stand-in ignores it and would be +# replayed into the shell the moment the stand-in died. +AGENT_PID=$(herdr pane process-info --pane "$PANE_ID" --session "$SESSION" 2>/dev/null \ + | jq -r '.result.process_info.foreground_processes[0].pid // empty') +[ -n "$AGENT_PID" ] || fail "could not read the agent-named process pid from pane process-info" +kill "$AGENT_PID" 2>/dev/null || fail "could not stop the agent-named process" +wait_process_state shell 50 \ + || version_fail "after the agent process exited the pane reads '$(fm_backend_herdr_pane_process_state "$SESSION" "$PANE_ID")' rather than 'shell' through pane process-info. Raw process-info: $(herdr pane process-info --pane "$PANE_ID" --session "$SESSION" 2>&1 | tr -d '\n')" + +# The divergence that makes this case non-vacuous: Herdr's own registry still +# reports the agent, and only the process-level view disagrees. +REGISTERED=$(herdr agent get "$PANE_ID" --session "$SESSION" 2>/dev/null | jq -r '.result.agent.agent_status // empty') +[ -n "$REGISTERED" ] \ + || version_fail "Herdr released the registration when the agent process exited, so this run cannot prove the stale-registration path; the classifier still reads dead through agent_not_found" + +PANE_STATE=$(fm_backend_herdr_pane_agent_state "$SESSION" "$PANE_ID") +[ "$PANE_STATE" = stale-agent ] \ + || version_fail "a registration over a shell-only pane reads '$PANE_STATE' rather than 'stale-agent'" +STATE=$(fm_backend_agent_state herdr "$SESSION:$PANE_ID") +[ "$STATE" = dead ] \ + || version_fail "a registration over a shell-only pane recovers as '$STATE' rather than 'dead'; every relaunch would be refused" +pass "real herdr $HERDR_VERSION: a registration Herdr keeps after its agent exits reads stale-agent and recovers as dead" if OUT=$(run_control hsmoke interrupt 2>&1); then fail "interrupt should refuse a stale registration with no agent behind it: $OUT" @@ -236,50 +300,41 @@ case "$OUT" in esac pass "real herdr: interrupt refuses a stale registration instead of keying a dead shell" -OUT=$(run_control hsmoke exit) || fail "exit against a stale registration must be idempotent success (this was the unrecoverable state): $OUT" + +OUT=$(run_control hsmoke exit) || fail "exit against a stale-registration pane should be idempotent success: $OUT" case "$OUT" in "already-stopped hsmoke"*) : ;; - *) fail "a stale registration should report already-stopped, got: $OUT" ;; + *) fail "a stale-registration pane should report already-stopped, got: $OUT" ;; esac -pass "real herdr: exit on a stale registration is idempotent success, so relaunch can proceed" +pass "real herdr: exit on a pane with a stale registration is idempotent success" -# --- a genuinely live agent: a registered agent WITH a real process --------- -# -# The paired negative: with a real long-running foreground process occupying -# the pane, the idle-shell proof fails, the registration keeps the benefit of -# the doubt, and the verbs treat the agent as alive - the cross-check must -# never let a live worker be exited or replaced. - -fm_backend_herdr_send_literal "$SESSION:$PANE_ID" "sleep 300" \ - || fail "could not type the live-process model into the task pane" -fm_backend_herdr_send_key "$SESSION:$PANE_ID" Enter \ - || fail "could not start the live-process model in the task pane" - -STATE= -for _ in 1 2 3 4 5 6 7 8 9 10; do - STATE=$(fm_backend_agent_state herdr "$SESSION:$PANE_ID") - [ "$STATE" = alive ] && break - sleep 0.3 +rm -f "$SCRATCH/codex-launched" +OUT=$(env FM_HOME="$HOME_DIR" HERDR_SESSION="$SESSION" FM_SPAWN_NO_GUARD=1 \ + "$ROOT/bin/fm-spawn.sh" hsmoke --relaunch --harness codex) \ + || fail "a stale-registration Herdr pane should be relaunched: $OUT" +for _ in $(seq 1 20); do + [ ! -e "$SCRATCH/codex-launched" ] || break + sleep 0.1 done -[ "$STATE" = alive ] || fail "a registered agent with a live foreground process must classify alive, got '$STATE'" -pass "real herdr: a registered agent with a live process stays alive through the cross-check" - -OUT=$(run_control hsmoke interrupt) || fail "interrupt against a live agent should succeed: $OUT" -case "$OUT" in - *"interrupt-delivered hsmoke harness=claude backend=herdr verified=agent-alive cancel=unconfirmed"*) : ;; - *) fail "interrupt should report the agent-alive proof on herdr, got: $OUT" ;; -esac -pass "real herdr: interrupt delivers the harness's key and proves the agent survived it" - +[ -e "$SCRATCH/codex-launched" ] || fail "the replacement harness was not launched after the stale registration" +[ "$(sed -n 's/^window=//p' "$HOME_DIR/state/hsmoke.meta" | tail -1)" = "$SESSION:$PANE_ID" ] \ + || fail "the relaunch replaced its endpoint instead of reusing it" herdr pane get "$PANE_ID" --session "$SESSION" >/dev/null 2>&1 \ - || fail "the control plane must never remove the endpoint it was operating on" -[ -d "$WT" ] || fail "the control plane must never remove the task's local copy" -pass "real herdr: no control verb removed the endpoint or the task's local copy" + || fail "the relaunch removed the endpoint it was required to reuse" +[ -d "$WT" ] || fail "the relaunch must never remove the task's local copy" +awk -F= '$1 == "harness" {$0="harness=claude"} {print}' "$HOME_DIR/state/hsmoke.meta" \ + > "$HOME_DIR/state/hsmoke.meta.tmp" +mv "$HOME_DIR/state/hsmoke.meta.tmp" "$HOME_DIR/state/hsmoke.meta" +pass "real herdr: a stale registration no longer blocks relaunch, and the endpoint and local copy survive" -# Last, because it deliberately types a harness command into a pane whose -# foreground process ignores it: the live agent cannot actually be stopped +# Last, because it deliberately types a harness command into a foreground +# process that ignores it: the registered agent cannot actually be stopped # that way, and the control plane must say so rather than report a stop it # did not achieve. +start_agent_process +herdr pane report-agent "$PANE_ID" --source fm-control-smoke --agent fm-control-smoke-agent \ + --state idle --session "$SESSION" >/dev/null 2>&1 \ + || fail "could not re-register the live agent on the task pane" if OUT=$(run_control hsmoke exit 2>&1); then fail "exit should fail closed when the agent does not stop: $OUT" fi diff --git a/tests/fm-control.test.sh b/tests/fm-control.test.sh index 74dfc0a9b73..beb5c1b243f 100755 --- a/tests/fm-control.test.sh +++ b/tests/fm-control.test.sh @@ -413,8 +413,8 @@ test_agy_has_verified_control_rows() { || fail "agy must still be a known adapter to the clear-key table, not an error" [ "$(fm_control_interrupt_ack_source agy)" = none ] \ || fail "agy exposes no durable typed close, so it must claim no acknowledgement source" - [ "$(fm_control_exit_command agy)" = '/exit' ] \ - || fail "agy exits on /exit" + [ "$(fm_control_exit_command agy)" = '/quit' ] \ + || fail "agy exits on /quit" fm_control_harness_supports_kind agy ship \ || fail "agy should be able to run a ship task" fm_control_harness_supports_kind agy secondmate \ diff --git a/tests/fm-crew-state.test.sh b/tests/fm-crew-state.test.sh index 85e40e25d3f..5da557b1f13 100755 --- a/tests/fm-crew-state.test.sh +++ b/tests/fm-crew-state.test.sh @@ -199,6 +199,18 @@ case "${1:-}" in fi printf '{"result":{"pane":{"pane_id":"%s"}}}\n' "${3:-}" exit 0 ;; + process-info) + # The process-level view a registration is verified against (#4115): + # `agent` puts a live claude in the foreground, `shell` a bare zsh whose + # pid is the test script itself (a real, long-lived process with no + # harness descendant, so the adapter's real process-table walk finds + # it), and anything else answers nothing (unreadable). + pane=""; args=("$@"); for ((i=0; i<${#args[@]}; i++)); do [ "${args[$i]}" = --pane ] && pane=${args[$((i+1))]:-}; done + case "${FM_FAKE_HERDR_PROCESS:-agent}" in + agent) printf '{"result":{"type":"pane_process_info","process_info":{"pane_id":"%s","shell_pid":%s,"foreground_process_group_id":424242,"foreground_processes":[{"pid":424242,"name":"claude","argv0":"claude"}]}}}\n' "$pane" "${FM_FAKE_HERDR_SHELL_PID:-$PPID}" ;; + shell) printf '{"result":{"type":"pane_process_info","process_info":{"pane_id":"%s","shell_pid":%s,"foreground_process_group_id":%s,"foreground_processes":[{"pid":%s,"name":"zsh","argv0":"zsh","argv":["-zsh"]}]}}}\n' "$pane" "${FM_FAKE_HERDR_SHELL_PID:-$PPID}" "${FM_FAKE_HERDR_SHELL_PID:-$PPID}" "${FM_FAKE_HERDR_SHELL_PID:-$PPID}" ;; + esac + exit 0 ;; esac ;; agent) case "${2:-}" in @@ -214,6 +226,12 @@ case "${1:-}" in esac exit 0 SH + # Match the synthetic shell process-info with a complete, isolated terminal snapshot. + cat > "$fb/herdr-ps" <<'SH' +#!/usr/bin/env bash +printf '%s 1 S zsh fmfixture\n' "$FM_FAKE_HERDR_SHELL_PID" +SH + chmod +x "$fb/herdr-ps" chmod +x "$fb/no-mistakes" "$fb/tmux" "$fb/herdr" printf '%s\n' "$fb" } @@ -232,7 +250,8 @@ make_no_timeout_toolbin() { # <dir> -> echoes toolbin path # Run the helper for one case dir. FM_FAKE_* env (run output, busy flag) are read # from the caller's environment by the fakes above. run_crew_state() { # <case-dir> <id> - PATH="$1/fakebin:$PATH" FM_STATE_OVERRIDE="$1/state" "$CREW_STATE" "$2" + PATH="$1/fakebin:$PATH" FM_STATE_OVERRIDE="$1/state" \ + FM_HERDR_PS_BIN="$1/fakebin/herdr-ps" "$CREW_STATE" "$2" } new_case() { # <name> -> echoes case dir with an empty state/ @@ -271,6 +290,8 @@ reset_fakes() { FM_FAKE_HERDR_READ_FAIL=0 FM_FAKE_HERDR_HUSK=0 FM_FAKE_HERDR_AGENT_STATUS="" + FM_FAKE_HERDR_PROCESS=agent + FM_FAKE_HERDR_SHELL_PID=$$ FM_FAKE_CI_LOGS="" FM_FAKE_NM_RC=0 FM_FAKE_RUNS_RC=0 @@ -289,6 +310,9 @@ reset_fakes() { export FM_CREW_STATE_DEGRADED_MAX_AGE FM_CREW_STATE_NM_TIMEOUT FM_CREW_STATE_RUNS_LIMIT FM_FAKE_DAEMON_DOWN=0 export FM_FAKE_DAEMON_DOWN FM_FAKE_TMUX_UNREADABLE FM_FAKE_HERDR_READ_FAIL FM_FAKE_HERDR_HUSK + export FM_FAKE_AXI_STATUS FM_FAKE_AXI_STATUS_RUN FM_FAKE_RUNS_LIST FM_FAKE_BUSY FM_FAKE_BUSY_TEXT FM_FAKE_TMUX_MISSING FM_FAKE_TMUX_UNREADABLE + export FM_FAKE_HERDR_BUSY FM_FAKE_HERDR_MISSING FM_FAKE_HERDR_READ_FAIL FM_FAKE_HERDR_HUSK FM_FAKE_HERDR_AGENT_STATUS FM_FAKE_HERDR_PROCESS FM_FAKE_HERDR_SHELL_PID FM_FAKE_CI_LOGS + export FM_FAKE_DAEMON_DOWN } # --- run-object fixtures (TOON, as `no-mistakes axi status` emits) ----------- @@ -2430,6 +2454,53 @@ test_no_run_herdr_alive_with_failed_read_stays_live() { pass "an alive endpoint whose scrollback read failed stays working" } +# Issue #4115: a registration Herdr kept after its Pi exited to a plain shell is +# not an agent. The recovery-grade read proves the process level, so the +# shell-only pane reads as positive agent-gone evidence, never as a live agent +# or as unreachable. +test_no_run_herdr_stale_registration_over_shell_reads_agent_gone() { + command -v jq >/dev/null 2>&1 || { pass "herdr stale-registration test skipped without jq"; return; } + reset_fakes + local d; d=$(new_case herdr-stale-reg) + make_repo_on_branch "$d/wt" fm/feat-herdr-stale + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-herdr-stale.meta" "window=default:w1:p2" "worktree=$d/wt" "kind=ship" \ + "backend=herdr" "harness=pi" + FM_FAKE_TMUX_MISSING=1 + FM_FAKE_HERDR_READ_FAIL=1 + FM_FAKE_HERDR_AGENT_STATUS=idle + FM_FAKE_HERDR_PROCESS=shell + local out; out=$(run_crew_state "$d" feat-herdr-stale) + assert_contains "$out" "state: unknown" "a stale registration over a shell-only pane is not a live state" + assert_contains "$out" "backend target gone" "a stale registration over a shell-only pane must read as positive agent-gone evidence" + assert_contains "$out" "agent gone, pane shell remains" "the agent-gone reason must name the remaining shell" + assert_not_contains "$out" "backend unreachable" "a readable shell-only pane is not unreachable" + pass "herdr stale registration over a shell-only pane reads agent gone, not alive" +} + +# The busy half of the same defect: a `working` record Herdr kept after the +# agent was killed mid-turn must never make a shell-only pane read as working. +test_no_run_herdr_stale_working_record_is_never_busy() { + command -v jq >/dev/null 2>&1 || { pass "herdr stale-working test skipped without jq"; return; } + reset_fakes + local d; d=$(new_case herdr-stale-working) + make_repo_on_branch "$d/wt" fm/feat-herdr-stale-working + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-herdr-stale-working.meta" "window=default:w1:p2" "worktree=$d/wt" "kind=ship" \ + "backend=herdr" "harness=pi" + FM_FAKE_TMUX_MISSING=1 + FM_FAKE_HERDR_AGENT_STATUS=working + FM_FAKE_HERDR_PROCESS=shell + local out; out=$(run_crew_state "$d" feat-herdr-stale-working) + assert_not_contains "$out" "state: working" "a stale working record over a shell-only pane must never read busy" + assert_not_contains "$out" "herdr-native" "the native busy verdict must not be trusted for a shell-only pane" + # The control: the same record with a live harness in the foreground is busy. + FM_FAKE_HERDR_PROCESS=agent + out=$(run_crew_state "$d" feat-herdr-stale-working) + assert_contains "$out" "state: working" "the same working record with a live harness process must still read working" + pass "herdr stale working record never reports a shell-only pane busy" +} + # Decision follow-up (2026-09-05 review): a husk pane (pane present, # agent_not_found) is authoritative death evidence - it keeps the gone-class # text so the stale sweep may still reclaim it, never unknown/unreachable. @@ -4512,5 +4583,7 @@ test_active_fix_round_unfetched_pipeline_head_reports_current test_unanchored_unfetched_active_row_does_not_match test_unresolved_terminal_row_is_history_not_current test_runs_list_continuation_found_when_axi_answers_other_branch +test_no_run_herdr_stale_registration_over_shell_reads_agent_gone +test_no_run_herdr_stale_working_record_is_never_busy echo "all fm-crew-state tests passed" diff --git a/tests/fm-cursor-harness.test.sh b/tests/fm-cursor-harness.test.sh index 23ecc74948b..4c59d1fdfbe 100755 --- a/tests/fm-cursor-harness.test.sh +++ b/tests/fm-cursor-harness.test.sh @@ -16,7 +16,9 @@ # 2. An unrelated `node`/`agent` pane classifies `other`, which the liveness # callers fold into `ambiguous` - NEVER `dead`. # 3. Cursor's env marker outranks an inherited CLAUDECODE, because cursor does -# not clear it and whichever marker is tested first wins. +# not clear it and whichever marker is tested first wins. That ordering +# settles the marker layer only: a nearer claude ancestor still outranks +# both (tests/fm-harness-precedence.test.sh owns that boundary). # 4. The transcript fold brackets a turn: a trailing turn_ended is idle, a # later role:user is busy, and an unresolvable binding is unknown. # 5. Cursor is a crewmate/scout adapter only and refuses a secondmate launch. @@ -156,44 +158,68 @@ test_tmux_classifies_cursor_pane_without_inferring_dead() { tree="$TMP_ROOT/tree5"; bin=$(make_cursor_tree "$tree") # shellcheck source=bin/backends/tmux.sh ( FM_BACKEND_LIB_DIR="$ROOT/bin"; . "$ROOT/bin/backends/tmux.sh" - [ "$(fm_backend_tmux_classify_process_name node "$bin/cursor-agent")" = agent ] \ + [ "$(fm_agent_process_classify_name node "$bin/cursor-agent")" = agent ] \ || fail "a cursor pane reported as node must classify agent" - [ "$(fm_backend_tmux_classify_process_name '' "$bin/cursor-agent")" = agent ] \ + [ "$(fm_agent_process_classify_name '' "$bin/cursor-agent")" = agent ] \ || fail "the argv[0]-only call must classify a cursor pane agent" # The safety half: an unrelated node is `other`, and the callers turn # `other` into `ambiguous`, never `dead`. - [ "$(fm_backend_tmux_classify_process_name node /usr/bin/node)" = other ] \ + [ "$(fm_agent_process_classify_name node /usr/bin/node)" = other ] \ || fail "an unrelated node must stay 'other', never agent" - [ "$(fm_backend_tmux_classify_process_name agent /usr/local/bin/agent)" = other ] \ + [ "$(fm_agent_process_classify_name agent /usr/local/bin/agent)" = other ] \ || fail "an unrelated agent must stay 'other', never agent" # Neighbours must not regress. - [ "$(fm_backend_tmux_classify_process_name claude '')" = agent ] || fail "claude regressed" - [ "$(fm_backend_tmux_classify_process_name zsh '')" = shell ] || fail "zsh regressed" + [ "$(fm_agent_process_classify_name claude '')" = agent ] || fail "claude regressed" + [ "$(fm_agent_process_classify_name zsh '')" = shell ] || fail "zsh regressed" ) || exit 1 pass "tmux liveness: a cursor pane is agent; an unrelated node/agent is other, never dead" } # --- 3. Detection ordering --------------------------------------------------- +# The marker ordering decides only when ancestry has nothing to say, so this +# case runs against a fake ps that reports a bash chain terminating at pid 1. +# Without it the suite would assert against whatever harness actually launched +# it, and the verdicts below would be about the runner rather than the ordering. test_cursor_marker_outranks_inherited_claudecode() { - local out + local out fakebin base_path + base_path=${FM_TEST_BASE_PATH:-/usr/bin:/bin:/usr/sbin:/sbin} + fakebin=$(fm_fakebin "$TMP_ROOT/marker-ordering") + fm_fake_blind_ancestry "$fakebin" # This is the exact hazard: cursor does NOT clear an inherited CLAUDECODE, so - # a cursor worker under a claude primary carries both markers. - out=$(CLAUDECODE=1 CURSOR_AGENT=1 "$HARNESS") + # a cursor session started by hand under a claude primary carries both markers. + out=$(PATH="$fakebin:$base_path" CLAUDECODE=1 CURSOR_AGENT=1 "$HARNESS") [ "$out" = cursor ] || fail "CLAUDECODE + CURSOR_AGENT must detect cursor, got '$out'" - out=$(CLAUDECODE=1 CURSOR_INVOKED_AS=cursor-agent "$HARNESS") + out=$(PATH="$fakebin:$base_path" CLAUDECODE=1 CURSOR_INVOKED_AS=cursor-agent "$HARNESS") [ "$out" = cursor ] || fail "CLAUDECODE + CURSOR_INVOKED_AS must detect cursor, got '$out'" # Both cursor markers stand alone, and neither steals a plain claude session. - out=$(env -u CLAUDECODE CURSOR_AGENT=1 "$HARNESS") + out=$(env -u CLAUDECODE PATH="$fakebin:$base_path" CURSOR_AGENT=1 "$HARNESS") [ "$out" = cursor ] || fail "CURSOR_AGENT alone must detect cursor, got '$out'" - out=$(env -u CURSOR_AGENT -u CURSOR_INVOKED_AS CLAUDECODE=1 "$HARNESS") + out=$(env -u CURSOR_AGENT -u CURSOR_INVOKED_AS PATH="$fakebin:$base_path" \ + CLAUDECODE=1 "$HARNESS") [ "$out" = claude ] || fail "CLAUDECODE alone must still detect claude, got '$out'" # A CURSOR_* variable that is not the invocation identity proves nothing. - out=$(env -u CURSOR_AGENT CLAUDECODE=1 CURSOR_API_ENDPOINT=https://example \ + out=$(env -u CURSOR_AGENT PATH="$fakebin:$base_path" CLAUDECODE=1 \ + CURSOR_API_ENDPOINT=https://example \ CURSOR_INVOKED_AS=something-else "$HARNESS") [ "$out" = claude ] \ || fail "an unrelated CURSOR_* setting must not claim the cursor identity, got '$out'" - pass "fm-harness.sh: cursor's marker outranks an inherited CLAUDECODE" + # The ordering is a marker-layer tiebreak, not a licence to overrule the + # process tree: with a real cursor-agent ancestor the two agree, and with a + # real claude ancestor the retained cursor marker loses. + local tree_dir + tree_dir="$TMP_ROOT/marker-ordering-trees" + mkdir -p "$tree_dir" + cp "$(command -v bash)" "$tree_dir/cursor-agent" + cp "$(command -v bash)" "$tree_dir/claude" + out=$(env -u CLAUDECODE "$tree_dir/cursor-agent" -c \ + "r=\$(CURSOR_AGENT=1 \"$HARNESS\"); printf '%s' \"\$r\"") + [ "$out" = cursor ] || fail "a real cursor-agent ancestor must detect cursor, got '$out'" + out=$("$tree_dir/claude" -c \ + "r=\$(CLAUDECODE=1 CURSOR_AGENT=1 \"$HARNESS\"); printf '%s' \"\$r\"") + [ "$out" = claude ] \ + || fail "a retained CURSOR_AGENT must not rename a real claude ancestor, got '$out'" + pass "fm-harness.sh: cursor's marker outranks an inherited CLAUDECODE when ancestry is silent" } test_harness_ancestry_rejects_cursor_named_node_script() { diff --git a/tests/fm-gbrain-readonly-e2e.test.sh b/tests/fm-gbrain-readonly-e2e.test.sh index dff218db647..5b8cb3110c0 100755 --- a/tests/fm-gbrain-readonly-e2e.test.sh +++ b/tests/fm-gbrain-readonly-e2e.test.sh @@ -38,6 +38,17 @@ curl -sf -m 5 "${EMBED_URL%/}/models" >/dev/null 2>&1 \ CLI="$ROOT/bin/fm-gbrain.sh" TMP_ROOT=$(fm_test_tmproot fm-gbrain-e2e) +# Every child gets an empty runtime home as well as an isolated GBRAIN_HOME. +# This covers file-backed provider credentials before grant-read and while serving. +RUNTIME_HOME="$TMP_ROOT/runtime-home" +mkdir -p "$RUNTIME_HOME" +runtime_env=(env) +while IFS='=' read -r _name _; do + case "$_name" in + *API_KEY*|*AUTH_TOKEN*|*API_TOKEN*|*_SECRET_KEY) runtime_env+=(-u "$_name") ;; + esac +done < <(env) +runtime_env+=("HOME=$RUNTIME_HOME") SERVE_PID="" PORT="" @@ -97,7 +108,7 @@ cp "$MAIN_HOME/config/gbrain.json" "$SM_HOME/config/gbrain.json" cp "$MAIN_HOME/config/gbrain.json" "$READER_TWO_HOME/config/gbrain.json" home_env() { # <home> <var> - FM_HOME="$1" bash "$CLI" paths --json | jq -r ".$2" + "${runtime_env[@]}" FM_HOME="$1" bash "$CLI" paths --json | jq -r ".$2" } MAIN_GBRAIN_HOME=$(home_env "$MAIN_HOME" gbrain_home) @@ -109,13 +120,13 @@ SM_PGLITE=$(home_env "$SM_HOME" pglite) mkdir -p "$MAIN_GBRAIN_HOME" "$SM_GBRAIN_HOME" init_brain() { # <gbrain-home> <pglite> - GBRAIN_HOME="$1" OLLAMA_BASE_URL="$EMBED_URL" "$GBRAIN_BIN" init --pglite \ + "${runtime_env[@]}" GBRAIN_HOME="$1" OLLAMA_BASE_URL="$EMBED_URL" "$GBRAIN_BIN" init --pglite \ --path "$2" --embedding-model "$EMBED_MODEL" --embedding-dimensions "$EMBED_DIMS" \ --non-interactive >/dev/null 2>&1 } put_page() { # <gbrain-home> <slug> <body> printf -- '---\ntype: note\n---\n\n%s\n' "$3" \ - | GBRAIN_HOME="$1" OLLAMA_BASE_URL="$EMBED_URL" "$GBRAIN_BIN" put "$2" >/dev/null 2>&1 + | "${runtime_env[@]}" GBRAIN_HOME="$1" OLLAMA_BASE_URL="$EMBED_URL" "$GBRAIN_BIN" put "$2" >/dev/null 2>&1 } init_brain "$MAIN_GBRAIN_HOME" "$MAIN_PGLITE" || { echo "skip: could not initialize a test main brain"; exit 0; } @@ -126,17 +137,17 @@ put_page "$SM_GBRAIN_HOME" sm-canary "This secondmate's own brain holds $SM_CANA || fail "could not seed the secondmate brain" WORLD_FACT='WORLD-CONTEXT-PACK-SENTINEL is visible remotely.' PRIVATE_FACT='PRIVATE-CONTEXT-PACK-SENTINEL must remain local.' -GBRAIN_HOME="$MAIN_GBRAIN_HOME" OLLAMA_BASE_URL="$EMBED_URL" "$GBRAIN_BIN" remember "$WORLD_FACT" \ +"${runtime_env[@]}" GBRAIN_HOME="$MAIN_GBRAIN_HOME" OLLAMA_BASE_URL="$EMBED_URL" "$GBRAIN_BIN" remember "$WORLD_FACT" \ --provenance 'live read-only E2E' --entity main-canary --kind commitment \ --visibility world >/dev/null 2>&1 || fail "could not seed the world-visible fact" -GBRAIN_HOME="$MAIN_GBRAIN_HOME" OLLAMA_BASE_URL="$EMBED_URL" "$GBRAIN_BIN" remember "$PRIVATE_FACT" \ +"${runtime_env[@]}" GBRAIN_HOME="$MAIN_GBRAIN_HOME" OLLAMA_BASE_URL="$EMBED_URL" "$GBRAIN_BIN" remember "$PRIVATE_FACT" \ --provenance 'live read-only E2E' --entity main-canary --kind commitment \ --visibility private >/dev/null 2>&1 || fail "could not seed the private fact" pass "two real, separately initialized brains exist, one per home" # --- grant the read-only share ---------------------------------------------- -grant_out=$(FM_HOME="$MAIN_HOME" FM_GBRAIN_BIN="$GBRAIN_BIN" \ +grant_out=$("${runtime_env[@]}" FM_HOME="$MAIN_HOME" FM_GBRAIN_BIN="$GBRAIN_BIN" \ bash "$CLI" grant-read fm-e2e-read --home "$SM_HOME" 2>&1) \ || fail "grant-read failed: $grant_out" @@ -149,7 +160,7 @@ CLIENT_ID=$(jq -r .client_id "$SM_HOME/config/gbrain-local.json") assert_not_contains "$grant_out" "$CLIENT_SECRET" "grant-read printed the credential it installed" pass "grant-read installed a read-only credential at mode 0600 without printing it" -grant_two_out=$(FM_HOME="$MAIN_HOME" FM_GBRAIN_BIN="$GBRAIN_BIN" \ +grant_two_out=$("${runtime_env[@]}" FM_HOME="$MAIN_HOME" FM_GBRAIN_BIN="$GBRAIN_BIN" \ bash "$CLI" grant-read fm-e2e-read-two --home "$READER_TWO_HOME" 2>&1) \ || fail "second grant-read failed: $grant_two_out" secret_file_two="$READER_TWO_HOME/config/gbrain-secrets/main-brain-client-secret" @@ -160,7 +171,7 @@ CLIENT_ID_TWO=$(jq -r .client_id "$READER_TWO_HOME/config/gbrain-local.json") assert_not_contains "$grant_two_out" "$CLIENT_SECRET_TWO" "the second grant printed its credential" pass "a second remote reader received a distinct read-only OAuth client" -writer_out=$(GBRAIN_HOME="$MAIN_GBRAIN_HOME" "$GBRAIN_BIN" auth register-client \ +writer_out=$("${runtime_env[@]}" GBRAIN_HOME="$MAIN_GBRAIN_HOME" "$GBRAIN_BIN" auth register-client \ fm-e2e-writer --scopes write --grant-types client_credentials 2>&1) \ || fail "writer registration failed: $writer_out" WRITER_ID=$(printf '%s\n' "$writer_out" | awk '/Client ID:/{print $NF}') @@ -169,27 +180,12 @@ WRITER_SECRET=$(printf '%s\n' "$writer_out" | awk '/Client Secret:/{print $NF}') || fail "the disposable writer registration returned no credentials" # The registration must actually be read-scoped in GBrain's own records. -scopes=$(GBRAIN_HOME="$MAIN_GBRAIN_HOME" "$GBRAIN_BIN" auth list 2>/dev/null || true) +scopes=$("${runtime_env[@]}" GBRAIN_HOME="$MAIN_GBRAIN_HOME" "$GBRAIN_BIN" auth list 2>/dev/null || true) assert_not_contains "$scopes" "gbrain_" "auth list should not expose a usable token value" # --- serve the main brain and drive real tool calls ------------------------- # -# The SERVED process is the one that would synthesize. `think` is scope:read -# since v0.42.76.0, so the guard further down really does reach it over the -# read-only share, and it would run on this process's model and credential: an -# inherited provider key would send the seeded main-brain page to a hosted -# provider from a suite that is supposed to touch nothing outside its temp root. -# So every credential-shaped variable is stripped from the served environment -# rather than the suite trusting the operator's shell to hold none, and the -# degrade is asserted at the call site instead of assumed. -serve_env=(env) -while IFS='=' read -r _name _; do - case "$_name" in - *API_KEY*|*AUTH_TOKEN*|*API_TOKEN*|*_SECRET_KEY) serve_env+=(-u "$_name") ;; - esac -done < <(env) - -"${serve_env[@]}" GBRAIN_HOME="$MAIN_GBRAIN_HOME" OLLAMA_BASE_URL="$EMBED_URL" \ +"${runtime_env[@]}" GBRAIN_HOME="$MAIN_GBRAIN_HOME" OLLAMA_BASE_URL="$EMBED_URL" \ "$GBRAIN_BIN" serve --http --port "$PORT" > "$TMP_ROOT/serve.log" 2>&1 & SERVE_PID=$! up=0 @@ -200,10 +196,10 @@ for _ in $(seq 1 90); do done [ "$up" = 1 ] || { echo "skip: the test main brain did not start serving"; exit 0; } -TOKEN=$(FM_HOME="$SM_HOME" bash "$CLI" token) \ +TOKEN=$("${runtime_env[@]}" FM_HOME="$SM_HOME" bash "$CLI" token) \ || fail "the secondmate could not obtain a read-only token" [ -n "$TOKEN" ] || fail "the secondmate received an empty token" -TOKEN_TWO=$(FM_HOME="$READER_TWO_HOME" bash "$CLI" token) \ +TOKEN_TWO=$("${runtime_env[@]}" FM_HOME="$READER_TWO_HOME" bash "$CLI" token) \ || fail "the second remote reader could not obtain a read-only token" [ -n "$TOKEN_TWO" ] || fail "the second remote reader received an empty token" WRITER_TOKEN=$(curl -sS -m 30 -X POST "http://127.0.0.1:$PORT/token" \ @@ -346,7 +342,7 @@ pass "the main brain is byte-for-byte unaffected by the refused writes" RECALL="$ROOT/bin/fm-recall.sh" recall_rc=0 -recall_out=$(FM_HOME="$SM_HOME" FM_GBRAIN_BIN="$GBRAIN_BIN" \ +recall_out=$("${runtime_env[@]}" FM_HOME="$SM_HOME" FM_GBRAIN_BIN="$GBRAIN_BIN" \ bash "$RECALL" search --json "$CANARY" 2>&1) || recall_rc=$? expect_code 0 "$recall_rc" "the wrapper should read the shared corpus: $recall_out" [ "$(printf '%s' "$recall_out" | jq -r '.sources[] | select(.source == "main") | .state')" = ok ] \ @@ -354,7 +350,7 @@ expect_code 0 "$recall_rc" "the wrapper should read the shared corpus: $recall_o assert_contains "$recall_out" '"citation": "main:main-canary"' \ "a main-brain result must arrive citable as main:<slug>" -recall_out=$(FM_HOME="$SM_HOME" FM_GBRAIN_BIN="$GBRAIN_BIN" \ +recall_out=$("${runtime_env[@]}" FM_HOME="$SM_HOME" FM_GBRAIN_BIN="$GBRAIN_BIN" \ bash "$RECALL" search --json "$SM_CANARY" 2>&1) || fail "the wrapper could not read the local corpus" assert_contains "$recall_out" '"citation": "local:sm-canary"' \ "a local result must arrive citable as local:<slug>" @@ -406,7 +402,7 @@ pass "a read-only share admits think, degrades it with no credential, cannot per put_page "$SM_GBRAIN_HOME" sm-second "The secondmate wrote this into its own brain." \ || fail "the secondmate could not write its own brain" -own=$(GBRAIN_HOME="$SM_GBRAIN_HOME" OLLAMA_BASE_URL="$EMBED_URL" \ +own=$("${runtime_env[@]}" GBRAIN_HOME="$SM_GBRAIN_HOME" OLLAMA_BASE_URL="$EMBED_URL" \ "$GBRAIN_BIN" list 2>/dev/null || true) assert_contains "$own" "sm-second" "the secondmate's own write must land in its own brain" assert_not_contains "$own" "main-canary" "the main brain's pages must not appear in the secondmate's own index" @@ -422,10 +418,10 @@ for _ in $(seq 1 30); do kill -0 "$SERVE_PID" 2>/dev/null || break; sleep 1; don SERVE_PID="" rc=0 -FM_HOME="$SM_HOME" FM_GBRAIN_TIMEOUT=3 bash "$CLI" token >/dev/null 2>&1 || rc=$? +"${runtime_env[@]}" FM_HOME="$SM_HOME" FM_GBRAIN_TIMEOUT=3 bash "$CLI" token >/dev/null 2>&1 || rc=$? [ "$rc" -ne 0 ] || fail "the main brain is stopped but a token was still issued" -local_search=$(GBRAIN_HOME="$SM_GBRAIN_HOME" OLLAMA_BASE_URL="$EMBED_URL" \ +local_search=$("${runtime_env[@]}" GBRAIN_HOME="$SM_GBRAIN_HOME" OLLAMA_BASE_URL="$EMBED_URL" \ "$GBRAIN_BIN" search "$SM_CANARY" 2>/dev/null || true) assert_contains "$local_search" "sm-canary" \ "with the main brain down, the secondmate's own search must still answer from its own index" @@ -433,7 +429,7 @@ assert_contains "$local_search" "sm-canary" \ # The same must hold through the wrapper crewmates actually use: a stopped main # brain is a degraded source, never a failed search. recall_rc=0 -recall_out=$(FM_HOME="$SM_HOME" FM_GBRAIN_BIN="$GBRAIN_BIN" FM_GBRAIN_TIMEOUT=3 \ +recall_out=$("${runtime_env[@]}" FM_HOME="$SM_HOME" FM_GBRAIN_BIN="$GBRAIN_BIN" FM_GBRAIN_TIMEOUT=3 \ bash "$RECALL" search --json --timeout 30 "$SM_CANARY" 2>&1) || recall_rc=$? expect_code 0 "$recall_rc" "a stopped main brain must not fail the wrapper's local search: $recall_out" [ "$(printf '%s' "$recall_out" | jq -r '.sources[] | select(.source == "main") | .state')" = degraded ] \ @@ -442,7 +438,7 @@ assert_contains "$recall_out" '"citation": "local:sm-canary"' \ "the home's own results must survive a stopped main brain" rc=0 -check_out=$(FM_HOME="$SM_HOME" FM_GBRAIN_TIMEOUT=3 bash "$CLI" check --json 2>/dev/null) || rc=$? +check_out=$("${runtime_env[@]}" FM_HOME="$SM_HOME" FM_GBRAIN_TIMEOUT=3 bash "$CLI" check --json 2>/dev/null) || rc=$? expect_code 0 "$rc" "a stopped main brain must not fail the reading home's check" state=$(printf '%s' "$check_out" | jq -r '.[] | select(.check == "main-brain") | .state') [ "$state" = degraded ] || fail "a stopped main brain should read as degraded, got '$state'" @@ -454,10 +450,10 @@ artifacts="$TMP_ROOT/artifacts.txt" { printf '%s\n' "$grant_out" printf '%s\n' "$grant_two_out" - FM_HOME="$SM_HOME" FM_GBRAIN_TIMEOUT=3 bash "$CLI" config - FM_HOME="$SM_HOME" FM_GBRAIN_TIMEOUT=3 bash "$CLI" config --json - FM_HOME="$SM_HOME" FM_GBRAIN_TIMEOUT=3 bash "$CLI" env - FM_HOME="$SM_HOME" FM_GBRAIN_TIMEOUT=3 bash "$CLI" check || true + "${runtime_env[@]}" FM_HOME="$SM_HOME" FM_GBRAIN_TIMEOUT=3 bash "$CLI" config + "${runtime_env[@]}" FM_HOME="$SM_HOME" FM_GBRAIN_TIMEOUT=3 bash "$CLI" config --json + "${runtime_env[@]}" FM_HOME="$SM_HOME" FM_GBRAIN_TIMEOUT=3 bash "$CLI" env + "${runtime_env[@]}" FM_HOME="$SM_HOME" FM_GBRAIN_TIMEOUT=3 bash "$CLI" check || true cat "$SM_HOME/config/gbrain.json" "$SM_HOME/config/gbrain-local.json" cat "$TMP_ROOT/serve.log" } > "$artifacts" 2>&1 diff --git a/tests/fm-gemini-harness.test.sh b/tests/fm-gemini-harness.test.sh index ad56dcc010c..b584537752f 100644 --- a/tests/fm-gemini-harness.test.sh +++ b/tests/fm-gemini-harness.test.sh @@ -31,19 +31,22 @@ HARNESS="$ROOT/bin/fm-harness.sh" TMP_ROOT=$(fm_test_tmproot fm-gemini-harness) test_gemini_marker_outranks_inherited_claudecode() { - local out + local out fakebin base_path + base_path=${FM_TEST_BASE_PATH:-/usr/bin:/bin:/usr/sbin:/sbin} + fakebin=$(fm_fakebin "$TMP_ROOT/marker-ordering") + fm_fake_blind_ancestry "$fakebin" # This is the exact hazard: gemini does not clear an inherited CLAUDECODE, so # a gemini worker under a claude primary carries both markers at once. - out=$(CLAUDECODE=1 GEMINI_CLI=1 "$HARNESS") + out=$(PATH="$fakebin:$base_path" CLAUDECODE=1 GEMINI_CLI=1 "$HARNESS") [ "$out" = gemini ] || fail "CLAUDECODE + GEMINI_CLI must detect gemini, got '$out'" # Drive the two signals apart so the case above cannot go quietly vacuous: # each marker alone must still produce its own verdict. - out=$(env -u CLAUDECODE GEMINI_CLI=1 "$HARNESS") + out=$(env -u CLAUDECODE PATH="$fakebin:$base_path" GEMINI_CLI=1 "$HARNESS") [ "$out" = gemini ] || fail "GEMINI_CLI alone must detect gemini, got '$out'" - out=$(env -u GEMINI_CLI CLAUDECODE=1 "$HARNESS") + out=$(env -u GEMINI_CLI PATH="$fakebin:$base_path" CLAUDECODE=1 "$HARNESS") [ "$out" = claude ] || fail "CLAUDECODE alone must still detect claude, got '$out'" # Cursor's marker still outranks gemini's, preserving the documented order. - out=$(CURSOR_AGENT=1 GEMINI_CLI=1 "$HARNESS") + out=$(PATH="$fakebin:$base_path" CURSOR_AGENT=1 GEMINI_CLI=1 "$HARNESS") [ "$out" = cursor ] || fail "CURSOR_AGENT must still outrank GEMINI_CLI, got '$out'" pass "fm-harness.sh: gemini's marker outranks an inherited CLAUDECODE" } diff --git a/tests/fm-gitignore-config.test.sh b/tests/fm-gitignore-config.test.sh index 0ae3568028c..21bb1de33a0 100755 --- a/tests/fm-gitignore-config.test.sh +++ b/tests/fm-gitignore-config.test.sh @@ -8,6 +8,8 @@ set -u ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" +# shellcheck source=tests/git-config-helpers.sh +. "$ROOT/tests/git-config-helpers.sh" fail() { printf 'not ok - %s\n' "$1" >&2 diff --git a/tests/fm-guard-stale-banner.test.sh b/tests/fm-guard-stale-banner.test.sh index 0ee4baac045..ca374497c91 100755 --- a/tests/fm-guard-stale-banner.test.sh +++ b/tests/fm-guard-stale-banner.test.sh @@ -169,6 +169,22 @@ test_first_stale_call_prints_full_banner() { pass "fm-guard stale banner: first stale call prints the full actionable banner" } +test_full_banner_names_quiet_mode_when_active() { + # kunchenguid/firstmate#2356: the banner's repair line must not misdirect a + # captain in quiet mode to /afk - fm-guard.sh threads the flag's declared + # mode through to fm-supervision-instructions.sh's --afk-mode. + local dir home out + dir=$(make_guard_case quiet-mode-banner) + home=$(case_home "$dir") + printf 'quiet\n%s\n' "$(date '+%s')" > "$home/state/.afk" + out=$(run_guard_case "$dir") + assert_contains "$out" "Quiet mode owns watcher supervision; load /quiet" \ + "full banner did not name /quiet for an active quiet-mode flag" + assert_not_contains "$out" "Away mode owns watcher supervision" \ + "full banner misdirected a quiet-mode captain to /afk" + pass "fm-guard stale banner: repair line is quiet-mode-aware, not hardcoded to away mode" +} + test_repeated_same_episode_prints_reminder_only() { local dir out1 out2 marker lines dir=$(make_guard_case repeated-stale) @@ -918,11 +934,15 @@ test_extension_live_watcher_is_healthy_without_ownership_evidence() { # The cases above pin the model. This one takes the end-user path instead: no # FM_SUPERVISION_MODEL at all, so bin/fm-harness.sh must route a Pi primary to the # extension model on its own. Without that routing the tolerance would never reach -# a real Pi home. The foreign markers are cleared because fm-harness.sh tests them -# ahead of Pi, and the host running this suite may carry one. +# a real Pi home. Pinning Pi takes both halves of the evidence: the foreign markers +# are cleared because the host running this suite may carry one, and the ancestry +# walk is blinded because a structural ancestor of a different harness outranks the +# Pi marker, so the harness this suite was launched from would otherwise answer. test_pi_harness_routes_itself_to_the_extension_model() { - local dir home out pid harness + local dir home out pid harness blind local -a pi_env + blind=$(fm_fakebin "$TMP_ROOT/pi-routing-blind") + fm_fake_blind_ancestry "$blind" for harness in pi pi-signed; do pi_env=(PI_CODING_AGENT=true) [ "$harness" = pi ] || pi_env+=(FM_PI_HARNESS=pi-signed) @@ -934,6 +954,7 @@ test_pi_harness_routes_itself_to_the_extension_model() { touch "$home/state/.last-watcher-beat" out=$(env -u CLAUDECODE -u CURSOR_AGENT -u CURSOR_INVOKED_AS -u GROK_AGENT -u FM_SUPERVISION_MODEL \ "${pi_env[@]}" \ + PATH="$blind:$PATH" \ FM_ROOT_OVERRIDE="$(case_root "$dir")" \ FM_HOME="$home" \ FM_GUARD_GRACE=999 \ @@ -947,6 +968,7 @@ test_pi_harness_routes_itself_to_the_extension_model() { } test_first_stale_call_prints_full_banner +test_full_banner_names_quiet_mode_when_active test_repeated_same_episode_prints_reminder_only test_pi_harness_routes_itself_to_the_extension_model test_extension_handoff_with_live_session_is_healthy diff --git a/tests/fm-harness-liveness-drift-live-e2e.test.sh b/tests/fm-harness-liveness-drift-live-e2e.test.sh index a48e0e83669..c5b5d8ee405 100755 --- a/tests/fm-harness-liveness-drift-live-e2e.test.sh +++ b/tests/fm-harness-liveness-drift-live-e2e.test.sh @@ -1,15 +1,23 @@ #!/usr/bin/env bash # tests/fm-harness-liveness-drift-live-e2e.test.sh - default-on drift guard proving # every INSTALLED harness is still classified `alive` by the tmux liveness -# probe (bin/backends/tmux.sh). +# probe (bin/backends/tmux.sh) AND still identified by the harness-detection +# ancestry walk (bin/fm-harness.sh). # -# Why this file exists: liveness classification depends on how a harness names -# its own process, which is a surface the harness vendor controls and changes -# without notice. Claude Code began reporting its version string as its process -# name and became unattributable, which silently degraded supervision. A -# regression that only a real harness release can cause needs a check that runs -# real harnesses; a stubbed agent cannot see it, and neither can a table of -# names transcribed from a previous release. +# Why this file exists: both verdicts depend on how a harness names its own +# process, which is a surface the harness vendor controls and changes without +# notice. Claude Code began reporting its version string as its process name and +# became unattributable, which silently degraded supervision. A regression that +# only a real harness release can cause needs a check that runs real harnesses; +# a stubbed agent cannot see it, and neither can a table of names transcribed +# from a previous release. +# +# Detection carries the same exposure for a second reason: a structural ancestor +# now outranks an environment marker (bin/fm-harness.sh owns that boundary), so +# a harness whose process name stops matching no longer merely loses a fast +# path - the walk keeps climbing and can reach a DIFFERENT harness that really +# is further up the tree. This guard is what catches that at the release that +# causes it. # # Each harness is launched bare, with no prompt, so this consumes no model # tokens. The launch uses whatever credentials the harness already has; an @@ -143,11 +151,87 @@ for harness in claude codex opencode pi pi-signed grok kimi cursor muse; do comms=$(fm_backend_tmux_foreground_comms "$target" | tr '\n' ' ') [ "$state" = alive ] || fail \ - "LIVENESS DRIFT: $harness $version is running but classifies '$state', not 'alive'. Supervision and lifecycle control treat this endpoint as unattributable. Observed process title '$title'; observed foreground process names [$comms]. Teach bin/backends/tmux.sh's fm_backend_tmux_classify_process_name the identity this release actually reports." + "LIVENESS DRIFT: $harness $version is running but classifies '$state', not 'alive'. Supervision and lifecycle control treat this endpoint as unattributable. Observed process title '$title'; observed foreground process names [$comms]. Teach bin/fm-agent-process-lib.sh's fm_agent_process_classify_name the identity this release actually reports." note "$harness $version: title='$title' foreground=[$comms]" pass "harness liveness: $harness $version classifies alive" + + # Detection: ask the ancestry walk what it makes of this real harness process. + # Both Pi identities share one launcher name, so ancestry can only ever prove + # the family; only the launch-boundary marker selects the signed identity. + expect_harness=$harness + [ "$harness" = pi-signed ] && expect_harness=pi + pane_pid=$("$REAL_TMUX" -L "$SOCKET" display-message -p -t "$target" '#{pane_pid}' 2>/dev/null | tr -d ' ') + [ -n "$pane_pid" ] || fail "$harness ($version): could not read the pane pid for the detection probe" + # Probe from BELOW the pane process, not the pane process alone. The shipped + # guarantee is a strength claim: detect_own hands an args-strength verdict back + # to a retained foreign marker, so a harness is only protected where the walk + # reaches it at comm strength. A harness that ships as a thin interpreter shim + # spawning its native binary as a CHILD is args strength from the pane process + # and comm strength from below that child - which is where firstmate's own + # detection actually runs, as a tool subprocess. Probing only the pane would + # therefore pass on evidence the guarantee does not rest on, and would keep + # passing if a release stopped spawning the native child at all. + # + # The vantage set is the UPWARD path from the deepest foreground descendant, not + # every descendant in the subtree, because harness_ancestry only ever climbs: a + # sibling branch is a vantage firstmate's own detection can never occupy. + # Restricting the deepest descendant to the pane tty's foreground process group + # keeps a process left running in the background out of the selection as well. + # + # The reject-other-harness cross-check below judges COMM-strength vantages only. + # An args-strength verdict is path-ambiguous by construction: harness_ancestry's + # bare-interpreter branch matches a harness name anywhere in the script path, so a + # harness-spawned MCP server running as `node <home>/.claude/mcp/<server>.js` + # answers `args claude` purely from the .claude path component, and such a server + # is normally a child of the agent binary rather than a sibling of it, so it can + # be the deepest descendant and sit ON this path. That ambiguity is the sole source + # of the false failure; a comm-strength verdict carries the real process name and + # cannot be produced that way. The comm-strength REQUIREMENT is unchanged - some + # vantage on the path must still name the expected harness at comm strength, + # because detect_own hands an args-strength verdict straight back to a retained + # foreign marker. + # The native binary can take a moment to appear, so poll for it. + pane_tty=$("$REAL_TMUX" -L "$SOCKET" display-message -p -t "$target" '#{pane_tty}' 2>/dev/null | tr -d ' ') + verdicts= + for _ in $(seq 1 150); do + fg_pids= + if [ -n "$pane_tty" ]; then + fg_pids=$(LC_ALL=C ps -t "${pane_tty#/dev/}" -o pid=,pgid=,tpgid= 2>/dev/null \ + | while read -r fg_pid fg_pgid fg_tpgid; do + [ -n "$fg_pid" ] || continue + [ "$fg_pgid" = "$fg_tpgid" ] || continue + printf '%s ' "$fg_pid" + done) + fi + # shellcheck disable=SC2086 # deliberate: the foreground pids are separate arguments + verdicts=$("$ROOT/bin/fm-harness.sh" ancestry-descent "$pane_pid" $fg_pids 2>/dev/null || true) + case "$verdicts" in *"comm $expect_harness"*) break ;; esac + sleep 0.2 + done + + drift_context="Observed process title '$title'; observed foreground process names [$comms]; observed ancestry verdicts [$(printf '%s' "$verdicts" | tr '\n' ';')]." + + [ -n "$verdicts" ] || fail \ + "DETECTION DRIFT: $harness $version is running but the ancestry walk reports nothing from the pane process or any vantage below it, so firstmate cannot identify this session at all. $drift_context Teach bin/fm-harness.sh's harness_ancestry the name this release actually reports." + + SAW_COMM=0 + while read -r strength named; do + [ -n "$strength" ] || continue + [ "$strength" = comm ] || continue + [ "$named" = "$expect_harness" ] || fail \ + "DETECTION DRIFT: $harness $version is running but a comm-strength vantage point on the upward path through its own session resolves to '$named', not '$expect_harness'. bin/fm-harness.sh lets a structural ancestor outrank an environment marker, so an unmatched process name can resolve to a DIFFERENT harness further up the tree instead of merely losing a fast path. $drift_context Teach bin/fm-harness.sh's harness_ancestry the name this release actually reports." + SAW_COMM=1 + done <<EOF +$verdicts +EOF + + [ "$SAW_COMM" = 1 ] || fail \ + "DETECTION DRIFT: $harness $version is identified only at interpreter-args strength, from no vantage point on the upward path through its session at comm strength. detect_own hands an args-strength verdict back to a retained foreign marker, so a stale CLAUDECODE would silently rename this session even though this guard sees the right identity. $drift_context Restore a process name bin/fm-harness.sh's harness_ancestry can match structurally, or teach it the name this release reports." + + note "$harness $version: ancestry verdicts=[$(printf '%s' "$verdicts" | tr '\n' ';')]" + pass "harness detection: $harness $version is identified by the ancestry walk at comm strength" CHECKED=$((CHECKED + 1)) done diff --git a/tests/fm-harness-precedence.test.sh b/tests/fm-harness-precedence.test.sh new file mode 100755 index 00000000000..0d4999984a3 --- /dev/null +++ b/tests/fm-harness-precedence.test.sh @@ -0,0 +1,761 @@ +#!/usr/bin/env bash +# Behavior tests for bin/fm-harness.sh's marker-vs-ancestry precedence boundary, +# and for the supervision protocol session start selects from it. +# +# The bug this pins: a Codex session started from an environment that had +# retained CLAUDECODE=1 detected as claude, because a verified marker outranked +# ancestry unconditionally. Session start then emitted Claude's Stop-owned +# supervision protocol to a Codex primary, which blocked every turn end. +# +# Every case drives the two evidence layers APART deliberately and asserts each +# one alone as well as the combination, so no case can pass vacuously if a layer +# silently stops working: +# marker alone - ancestry blinded by a fake ps, proving the marker is live +# and is what the old precedence would have returned. +# ancestry alone - marker cleared, proving the ancestry signal is live. +# both together - the precedence verdict this file exists to pin. +# The fake ps blinds only the ancestry walk, and the suite proves that rather +# than assuming it. fm-harness.sh reads process ancestry through ps alone; the +# one source it does not read through ps is the Cursor argv[0] probe, which on +# Linux reads /proc directly and on macOS falls back to ps. Either way it +# resolves the harmless real path of a bash-named process, which +# fm_cursor_process_matches rejects. The no-marker case below asserts `unknown` +# under the fake ps on whichever platform the run happens on, and it is exactly +# that case that fails if the blinding ever leaks a real ancestor through. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +# This suite states the markers it means to test in every case. Drop the ambient +# ones so a verdict never depends on which harness launched the suite. +unset CLAUDECODE PI_CODING_AGENT FM_PI_HARNESS GROK_AGENT CURSOR_AGENT CURSOR_INVOKED_AS + +HARNESS="$ROOT/bin/fm-harness.sh" +RENDER="$ROOT/bin/fm-supervision-instructions.sh" +TMP_ROOT=$(fm_test_tmproot fm-harness-precedence) +BASE_PATH=${FM_TEST_BASE_PATH:-/usr/bin:/bin:/usr/sbin:/sbin} + +# A real process named after a harness, asked for its verdict from a child. +# The command substitution around the probe is load-bearing: a bare `-c <cmd>` +# lets the shell exec the probe in place, which REPLACES the harness-named +# process the walk is supposed to find. +under_process() { # <named-executable> [VAR=VAL ...] + local bin=$1 + shift + env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS "$@" \ + "$bin" -c "r=\$(\"$HARNESS\"); printf '%s' \"\$r\"" +} + +# A fake ps that reports a bash ancestor terminating at pid 1, so the ancestry +# layer proves nothing and only the marker layer can answer. +blind_ancestry_bin() { # <dir> + local fakebin + fakebin=$(fm_fakebin "$1") + cat > "$fakebin/ps" <<'SH' +#!/usr/bin/env bash +case "$*" in + *'ppid='*) printf '%s\n' 1 ;; + *) printf '%s\n' bash ;; +esac +SH + chmod +x "$fakebin/ps" + printf '%s\n' "$fakebin" +} + +# A fake ps that models a PID NAMESPACE: every process reports bash with ppid 1, +# and pid 1 reports whatever FM_TEST_PID1_COMM names. This is what a harness +# looks like from inside a container or `codex sandbox`, where the harness is +# pid 1 of its own namespace rather than a child of a shell. +namespace_ancestry_bin() { # <dir> + local fakebin + fakebin=$(fm_fakebin "$1") + cat > "$fakebin/ps" <<'SH' +#!/usr/bin/env bash +pid= +prev= +for a in "$@"; do + [ "$prev" = -p ] && pid=$a + prev=$a +done +if [ "$pid" = 1 ]; then + comm=${FM_TEST_PID1_COMM:-init} + ppid=0 +else + comm=bash + ppid=1 +fi +case "$*" in + *'ppid='*) printf '%s\n' "$ppid" ;; + *) printf '%s\n' "$comm" ;; +esac +SH + chmod +x "$fakebin/ps" + printf '%s\n' "$fakebin" +} + +# Run the harness script under a fake ps, with the ambient markers dropped so +# each case states its own. +under_fake_ps() { # <fakebin> <VAR=VAL ...> -- [harness args] + local fakebin=$1 + shift + local -a assignments=() + while [ "$#" -gt 0 ] && [ "$1" != -- ]; do + assignments+=("$1") + shift + done + [ "${1:-}" = -- ] && shift + env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS "${assignments[@]}" \ + PATH="$fakebin:$BASE_PATH" "$HARNESS" "$@" +} + +with_blind_ancestry() { # <fakebin> [VAR=VAL ...] + local fakebin=$1 + shift + env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS "$@" \ + PATH="$fakebin:$BASE_PATH" "$HARNESS" +} + +named_bin() { # <dir> <name> + mkdir -p "$1" + cp "$(command -v bash)" "$1/$2" + printf '%s\n' "$1/$2" +} + +# --- 1. A foreign marker never renames a markerless harness ----------------- + +# codex, opencode, kimi, muse, and agy publish no identity marker, so before +# this boundary existed ANY retained marker renamed them outright. This is the +# reported live failure, generalized to every markerless adapter and to both +# foreign markers that can be retained. +test_markerless_ancestry_outranks_foreign_marker() { + local dir fakebin bin got name + dir="$TMP_ROOT/markerless" + fakebin=$(blind_ancestry_bin "$dir/blind") + for name in codex opencode kimi muse-bin-0.1.0 agy; do + bin=$(named_bin "$dir/$name-tree" "$name") + local expect=$name + case "$name" in muse-bin-*) expect=muse ;; esac + + got=$(under_process "$bin") + [ "$got" = "$expect" ] \ + || fail "$name ancestry alone resolved '$got', expected $expect (the ancestry signal is not live)" + + got=$(with_blind_ancestry "$fakebin" CLAUDECODE=1) + [ "$got" = claude ] \ + || fail "an inherited CLAUDECODE alone resolved '$got', expected claude (the marker signal is not live)" + + got=$(under_process "$bin" CLAUDECODE=1) + [ "$got" = "$expect" ] \ + || fail "$name ancestry with an inherited CLAUDECODE resolved '$got', expected $expect" + + got=$(under_process "$bin" CURSOR_AGENT=1) + [ "$got" = "$expect" ] \ + || fail "$name ancestry with an inherited CURSOR_AGENT resolved '$got', expected $expect" + done + pass "a markerless harness keeps its identity under an inherited foreign marker" +} + +# --- 2. A genuine harness in its own process tree still wins ---------------- + +test_genuine_marker_and_ancestry_agree() { + local dir bin got + dir="$TMP_ROOT/genuine" + + bin=$(named_bin "$dir/claude-tree" claude) + got=$(under_process "$bin" CLAUDECODE=1) + [ "$got" = claude ] || fail "a genuine claude session resolved '$got', expected claude" + + bin=$(named_bin "$dir/cursor-tree" cursor-agent) + got=$(under_process "$bin" CURSOR_AGENT=1) + [ "$got" = cursor ] || fail "a genuine cursor session resolved '$got', expected cursor" + got=$(under_process "$bin" CURSOR_INVOKED_AS=cursor-agent) + [ "$got" = cursor ] || fail "a genuine cursor session (launcher marker) resolved '$got', expected cursor" + + bin=$(named_bin "$dir/grok-tree" grok) + got=$(under_process "$bin" GROK_AGENT=1) + [ "$got" = grok ] || fail "a genuine grok session resolved '$got', expected grok" + # grok 1.0.0 hook processes carry no GROK_AGENT at all, so ancestry alone must + # still answer for them. + got=$(under_process "$bin") + [ "$got" = grok ] || fail "an unmarked grok hook process resolved '$got', expected grok" + + pass "a harness that publishes a marker inside its own process tree is unchanged" +} + +# Cursor is the case that motivated the pre-existing marker ordering: a cursor +# session started by hand under a claude primary carries BOTH markers. Ancestry +# is silent about which owns the tree there, so the ordering still decides. +test_cursor_ordering_still_decides_when_ancestry_is_silent() { + local fakebin got + fakebin=$(blind_ancestry_bin "$TMP_ROOT/cursor-ordering") + got=$(with_blind_ancestry "$fakebin" CLAUDECODE=1 CURSOR_AGENT=1) + [ "$got" = cursor ] || fail "both markers with no ancestry resolved '$got', expected cursor" + got=$(with_blind_ancestry "$fakebin" CLAUDECODE=1 CURSOR_INVOKED_AS=cursor-agent) + [ "$got" = cursor ] || fail "both markers (launcher form) with no ancestry resolved '$got', expected cursor" + got=$(with_blind_ancestry "$fakebin" CLAUDECODE=1) + [ "$got" = claude ] || fail "CLAUDECODE with no ancestry resolved '$got', expected claude" + got=$(with_blind_ancestry "$fakebin") + [ "$got" = unknown ] \ + || fail "no marker and no ancestry resolved '$got', expected unknown" + pass "with ancestry silent, the marker layer and its cursor-first ordering still decide" +} + +# The symmetric half of the same bug: cursor-agent's marker reaches a nested +# claude worker's environment, and the nearer claude ancestor must win. +test_retained_cursor_marker_does_not_rename_a_nested_claude() { + local bin got + bin=$(named_bin "$TMP_ROOT/nested-claude" claude) + got=$(under_process "$bin" CURSOR_AGENT=1 CLAUDECODE=1) + [ "$got" = claude ] \ + || fail "a claude tree carrying a retained CURSOR_AGENT resolved '$got', expected claude" + got=$(under_process "$bin" CURSOR_INVOKED_AS=cursor-agent CLAUDECODE=1) + [ "$got" = claude ] \ + || fail "a claude tree carrying a retained cursor launcher marker resolved '$got', expected claude" + pass "a retained cursor marker does not rename a nested claude worker" +} + +# --- 3. Pi keeps the marker's more specific identity ------------------------ + +# Both Pi identities share the launcher name, so ancestry can only prove the +# family. A marker that agrees on the family must keep its finer verdict rather +# than being flattened to pi by the ancestry walk. +test_pi_signed_survives_agreeing_ancestry() { + local bin got + bin=$(named_bin "$TMP_ROOT/pi-tree" pi) + got=$(under_process "$bin" PI_CODING_AGENT=true FM_PI_HARNESS=pi-signed) + [ "$got" = pi-signed ] || fail "signed Pi over pi ancestry resolved '$got', expected pi-signed" + got=$(under_process "$bin" PI_CODING_AGENT=true) + [ "$got" = pi ] || fail "plain Pi over pi ancestry resolved '$got', expected pi" + got=$(under_process "$bin") + [ "$got" = pi ] || fail "unmarked pi ancestry resolved '$got', expected pi" + + bin=$(named_bin "$TMP_ROOT/pi-signed-tree" pi-signed) + got=$(under_process "$bin" PI_CODING_AGENT=true FM_PI_HARNESS=pi-signed) + [ "$got" = pi-signed ] \ + || fail "signed Pi over shared signed-wrapper ancestry resolved '$got', expected pi-signed" + got=$(under_process "$bin") + [ "$got" = pi ] \ + || fail "unmarked signed-wrapper ancestry resolved '$got', expected pi" + pass "an agreeing marker keeps Pi's finer identity that ancestry cannot prove" +} + +# --- 4. The weakest ancestry signal does not outrank a marker --------------- + +# A bare interpreter matched only by a harness name inside the script path it +# was handed is the weakest inference in fm-harness.sh: any node process holding +# a harness-shaped path matches it. It answers when nothing else does, but it +# must not overturn a harness publishing its own identity. +test_interpreter_args_match_does_not_outrank_a_marker() { + local dir node script got + dir="$TMP_ROOT/weak-args" + node=$(named_bin "$dir" node) + script="$dir/codex-tool.sh" + cat > "$script" <<SH +r=\$("$HARNESS"); printf '%s' "\$r" +SH + + got=$(env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS "$node" "$script") + [ "$got" = codex ] \ + || fail "an unmarked interpreter holding a codex-shaped script path resolved '$got', expected codex" + + got=$(env -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS CLAUDECODE=1 "$node" "$script") + [ "$got" = claude ] \ + || fail "a published CLAUDECODE lost to a codex-shaped script path, resolving '$got'" + pass "an interpreter script-path match answers alone but never outranks a marker" +} + +# The real Codex install topology, modelled because the fix depends on it. codex +# ships as a `node` npm shim that spawns its native `codex` binary as a child and +# waits, so BOTH are in a tool subprocess's parent chain and the native name is +# the nearer one. Verified live on 2026-09-01 with codex-cli 0.152.0, whose pane +# foreground process names were [node codex]. What this case pins is that rule +# and nothing wider: a native harness binary nearer than an interpreter decides +# at comm strength, so the shim's own script path never gets to hand the verdict +# back to a retained marker. The strength assertion below is what keeps that +# non-vacuous - reaching the node shim instead would answer 'args codex'. +# A fixture cannot notice a vendor topology change; the opt-in live drift guard +# (tests/fm-harness-liveness-drift-live-e2e.test.sh) owns that. +test_native_child_of_an_interpreter_shim_decides_at_comm_strength() { + local dir node native probe entry got + dir="$TMP_ROOT/shim-topology" + node=$(named_bin "$dir" node) + native=$(named_bin "$dir/vendor" codex) + + # The probe forks so the command substitution's child is what asks, exactly as + # a tool subprocess of a real harness would. + probe="$dir/probe.sh" + cat > "$probe" <<'SH' +r=$("$FM_TEST_HARNESS" "$@") +printf '%s' "$r" +SH + # The shim SPAWNS its native binary and waits, so the node process stays alive + # above it and the walk meets the native binary first. + entry="$dir/codex-cli-entry.sh" + cat > "$entry" <<'SH' +"$FM_TEST_NATIVE" "$FM_TEST_PROBE" "$@" & +wait "$!" +SH + + # No arguments: the two cases that vary the environment or the subcommand call + # the shim entry point directly below, so this helper stays the plain no-marker + # launch. + run_shim() { + env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS \ + FM_TEST_HARNESS="$HARNESS" FM_TEST_NATIVE="$native" FM_TEST_PROBE="$probe" \ + "$node" "$entry" + } + + got=$(run_shim) + [ "$got" = codex ] \ + || fail "the shim topology without a marker resolved '$got', expected codex" + + got=$(env CLAUDECODE=1 FM_TEST_HARNESS="$HARNESS" FM_TEST_NATIVE="$native" \ + FM_TEST_PROBE="$probe" "$node" "$entry") + [ "$got" = codex ] \ + || fail "the real Codex shim topology with a retained CLAUDECODE resolved '$got', expected codex" + + got=$(env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS \ + FM_TEST_HARNESS="$HARNESS" FM_TEST_NATIVE="$native" FM_TEST_PROBE="$probe" \ + "$node" "$entry" ancestry) + [ "$got" = "comm codex" ] \ + || fail "the native child must decide at comm strength, got '$got'" + pass "a native harness binary under an interpreter shim decides at comm strength" +} + +# --- 5. A harness that is pid 1 of its own namespace ------------------------ + +# The walk used to stop as soon as the NEXT pid was 1, on the assumption that +# pid 1 is always init. Inside a PID namespace that assumption inverts: the +# harness itself is pid 1, so the one process that proves who owns the tree was +# never examined and a retained marker won by default. Verified against the real +# installed Codex, which runs as pid 1 under `codex sandbox`. +test_harness_at_namespace_pid1_is_examined() { + local fakebin got + fakebin=$(namespace_ancestry_bin "$TMP_ROOT/namespace-pid1") + + # Non-vacuity, both directions: with a host-shaped pid 1 the marker is the + # only evidence and must still answer, so the case below cannot pass by the + # ancestry layer simply matching everything. + got=$(under_fake_ps "$fakebin" FM_TEST_PID1_COMM=init CLAUDECODE=1 --) + [ "$got" = claude ] \ + || fail "a host-shaped pid 1 resolved '$got', expected claude (the marker layer is not live)" + + got=$(under_fake_ps "$fakebin" FM_TEST_PID1_COMM=codex --) + [ "$got" = codex ] \ + || fail "a Codex session at namespace pid 1 resolved '$got' with no marker, expected codex" + + got=$(under_fake_ps "$fakebin" FM_TEST_PID1_COMM=codex CLAUDECODE=1 --) + [ "$got" = codex ] \ + || fail "a Codex session at namespace pid 1 holding a retained CLAUDECODE resolved '$got', expected codex" + + got=$(under_fake_ps "$fakebin" FM_TEST_PID1_COMM=codex CLAUDECODE=1 -- ancestry) + [ "$got" = "comm codex" ] \ + || fail "the namespace pid 1 harness must decide at comm strength, got '$got'" + + pass "a harness that is pid 1 of its own namespace is examined, not skipped" +} + +# --- 6. The vantage point a probe asks from decides what strength it can see -- + +# The shipped guarantee is a strength claim, not just an identity one: detect_own +# hands an args-strength verdict back to a retained marker, so a harness is only +# protected where the walk reaches it at comm strength. Which strength is even +# REACHABLE depends on where the question is asked from. Under an interpreter +# shim the top of the session is the shim, whose own script path is args +# strength, while the native binary that carries comm strength is its CHILD. +# firstmate's own detect_own always runs from a tool subprocess below that child, +# so it sees comm; a guard that probed only the top of a real session would +# observe args, pass, and never notice a vendor release that stopped spawning the +# native child at all. `ancestry-descent` is what lets a probe ask from the same +# vantage a real session occupies, and this case pins that it reaches strictly +# further than the top-of-session probe does. +test_descent_probe_reaches_a_strength_the_top_of_session_cannot() { + local dir node native hold entry ready shim_pid got waited + dir="$TMP_ROOT/descent-vantage" + node=$(named_bin "$dir" node) + native=$(named_bin "$dir/vendor" codex) + ready="$dir/ready" + + # The native binary parks until the test releases it, so the whole topology is + # still standing while the probes run. + hold="$dir/hold.sh" + cat > "$hold" <<'SH' +touch "$FM_TEST_READY" +while [ -e "$FM_TEST_READY" ]; do sleep 0.05; done +SH + # The shim spawns its native binary and waits, so the node process stays alive + # ABOVE it exactly as the real Codex npm shim does. + entry="$dir/codex-cli-entry.sh" + cat > "$entry" <<'SH' +"$FM_TEST_NATIVE" "$FM_TEST_HOLD" & +wait "$!" +SH + + env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS \ + FM_TEST_NATIVE="$native" FM_TEST_HOLD="$hold" FM_TEST_READY="$ready" \ + "$node" "$entry" & + shim_pid=$! + + waited=0 + while [ ! -e "$ready" ] && [ "$waited" -lt 200 ]; do + sleep 0.05 + waited=$((waited + 1)) + done + [ -e "$ready" ] || { rm -f "$ready"; kill "$shim_pid" 2>/dev/null; fail "the shim fixture never reached its native child"; } + + # Non-vacuity: the top-of-session vantage really is limited to args strength + # here, which is the whole reason the descent probe has something to add. + got=$("$HARNESS" ancestry "$shim_pid") + [ "$got" = "args codex" ] \ + || fail "the shim's own vantage should see only 'args codex', got '$got'; the descent case proves nothing if the top of the session already reaches comm strength" + + got=$("$HARNESS" ancestry-descent "$shim_pid") + case "$got" in + *"comm codex"*) ;; + *) fail "the descent probe did not reach the native child at comm strength, got '$got'" ;; + esac + + # No vantage point inside the session may name a DIFFERENT harness, or a guard + # built on this probe would accept a tree it should have rejected. + while read -r strength named; do + [ -n "$strength" ] || continue + [ "$named" = codex ] \ + || fail "a vantage point inside the codex fixture reported '$strength $named'" + done <<EOF +$got +EOF + + rm -f "$ready" + wait "$shim_pid" 2>/dev/null || true + pass "the descent probe reaches comm strength where the top-of-session probe sees only args" +} + +# The other half of the vantage question: which vantages a probe must NOT ask +# from. harness_ancestry only ever climbs, so firstmate's own detection can never +# occupy a SIBLING branch of the process that runs it. A harness routinely spawns +# such branches - an MCP server started as `node <home>/.claude/mcp/<server>.js` +# matches *claude* on its script path in the bare-interpreter branch of the walk - +# and a probe that reported every descendant would answer a foreign harness from a +# process no real tool subprocess can ask from. The descent probe asks only the +# vantages on the upward path from the deepest descendant, which is exactly the set +# detection itself can reach. +test_descent_probe_ignores_a_sibling_branch_the_walk_cannot_reach() { + local dir node native worker mcp_script block hold entry ready fifo + local shim_pid mcp_pid got waited + dir="$TMP_ROOT/descent-sibling" + node=$(named_bin "$dir" node) + native=$(named_bin "$dir/vendor" codex) + worker=$(named_bin "$dir/vendor" worker) + ready="$dir/ready" + fifo="$dir/fifo" + mkdir -p "$dir/.claude/mcp" + mkfifo "$fifo" + + # Both leaves park on a fifo nothing ever writes, so they hold their position in + # the tree without spawning children of their own and the depths stay fixed. + block="$dir/block.sh" + cat > "$block" <<'SH' +read -r _ < "$FM_TEST_FIFO" +SH + # The MCP server is the sibling branch: a bare interpreter whose script path + # carries a harness name it does not belong to. + mcp_script="$dir/.claude/mcp/foo.js" + cp "$block" "$mcp_script" + + # The native binary keeps a child of its own, so the deepest descendant is + # unambiguously on the codex branch rather than tied with the sibling. + hold="$dir/hold.sh" + cat > "$hold" <<'SH' +"$FM_TEST_WORKER" "$FM_TEST_BLOCK" & +printf '%s\n' "$!" > "$FM_TEST_DIR/worker.pid" +touch "$FM_TEST_READY" +wait +SH + entry="$dir/codex-cli-entry.sh" + cat > "$entry" <<'SH' +"$FM_TEST_NATIVE" "$FM_TEST_HOLD" & +printf '%s\n' "$!" > "$FM_TEST_DIR/native.pid" +"$FM_TEST_NODE" "$FM_TEST_MCP" & +printf '%s\n' "$!" > "$FM_TEST_DIR/mcp.pid" +wait +SH + + env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS \ + FM_TEST_DIR="$dir" FM_TEST_NODE="$node" FM_TEST_NATIVE="$native" \ + FM_TEST_WORKER="$worker" FM_TEST_HOLD="$hold" FM_TEST_BLOCK="$block" \ + FM_TEST_MCP="$mcp_script" FM_TEST_READY="$ready" FM_TEST_FIFO="$fifo" \ + "$node" "$entry" & + shim_pid=$! + + waited=0 + while { [ ! -e "$ready" ] || [ ! -s "$dir/mcp.pid" ] || [ ! -s "$dir/worker.pid" ]; } \ + && [ "$waited" -lt 200 ]; do + sleep 0.05 + waited=$((waited + 1)) + done + release_sibling_fixture() { + kill "$(cat "$dir/worker.pid" 2>/dev/null)" "$(cat "$dir/mcp.pid" 2>/dev/null)" \ + "$(cat "$dir/native.pid" 2>/dev/null)" "$shim_pid" 2>/dev/null || true + wait "$shim_pid" 2>/dev/null || true + } + { [ -e "$ready" ] && [ -s "$dir/mcp.pid" ] && [ -s "$dir/worker.pid" ]; } \ + || { release_sibling_fixture; fail "the sibling fixture never reached both of its leaves"; } + mcp_pid=$(cat "$dir/mcp.pid") + + # Non-vacuity: the sibling really does answer a foreign harness when asked, so a + # probe that reported every descendant would have reported claude here. + got=$("$HARNESS" ancestry "$mcp_pid") + [ "$got" = "args claude" ] \ + || { release_sibling_fixture; fail "the sibling MCP process reported '$got', expected 'args claude'; this case proves nothing unless that branch really names a foreign harness"; } + + got=$("$HARNESS" ancestry-descent "$shim_pid") + case "$got" in + *"comm codex"*) ;; + *) release_sibling_fixture; fail "the descent probe did not reach the native child at comm strength, got '$got'" ;; + esac + while read -r strength named; do + [ -n "$strength" ] || continue + [ "$named" = codex ] \ + || { release_sibling_fixture; fail "the descent probe reported '$strength $named' from a sibling branch the ancestry walk can never climb through"; } + done <<EOF +$got +EOF + + release_sibling_fixture + pass "the descent probe reports no verdict from a sibling branch detection cannot reach" +} + +# The deeper shape the case above cannot reach, and the reason the live guard's +# reject-other-harness cross-check judges COMM-strength vantages only. A harness +# spawns its MCP servers from the AGENT BINARY, not from the npm shim, so the real +# Codex topology is shim -> native codex -> mcp server: the server inherits its +# parent's process group, passes the foreground filter, and is the deepest eligible +# descendant, which puts its own `args claude` vantage ON the descent path rather +# than off it. An args-strength verdict is path-ambiguous by construction - the +# bare-interpreter branch of the walk matches a harness name anywhere in the script +# path - so it is the comm-strength verdicts that carry a real process name and are +# the ones worth cross-checking. This case pins that the path still reaches +# `comm codex`, that every comm-strength vantage on it names codex, and that an +# `args claude` vantage really is present, which is what a cross-check applied to +# args strength would have rejected. +test_descent_probe_tolerates_an_args_only_foreign_verdict_at_the_deepest_vantage() { + local dir node native mcp_script hold entry ready fifo + local shim_pid mcp_pid got waited saw_comm + dir="$TMP_ROOT/descent-deep-mcp" + node=$(named_bin "$dir" node) + native=$(named_bin "$dir/vendor" codex) + ready="$dir/ready" + fifo="$dir/fifo" + mkdir -p "$dir/.claude/mcp" + mkfifo "$fifo" + + mcp_script="$dir/.claude/mcp/foo.js" + cat > "$mcp_script" <<'SH' +read -r _ < "$FM_TEST_FIFO" +SH + + # The native binary is what starts the MCP server, so the server sits BELOW it and + # is the deepest descendant of the whole tree. + hold="$dir/hold.sh" + cat > "$hold" <<'SH' +"$FM_TEST_NODE" "$FM_TEST_MCP" & +printf '%s\n' "$!" > "$FM_TEST_DIR/mcp.pid" +touch "$FM_TEST_READY" +wait +SH + entry="$dir/codex-cli-entry.sh" + cat > "$entry" <<'SH' +"$FM_TEST_NATIVE" "$FM_TEST_HOLD" & +printf '%s\n' "$!" > "$FM_TEST_DIR/native.pid" +wait +SH + + env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS \ + FM_TEST_DIR="$dir" FM_TEST_NODE="$node" FM_TEST_NATIVE="$native" \ + FM_TEST_HOLD="$hold" FM_TEST_MCP="$mcp_script" FM_TEST_READY="$ready" \ + FM_TEST_FIFO="$fifo" \ + "$node" "$entry" & + shim_pid=$! + + waited=0 + while { [ ! -e "$ready" ] || [ ! -s "$dir/mcp.pid" ]; } && [ "$waited" -lt 200 ]; do + sleep 0.05 + waited=$((waited + 1)) + done + release_deep_mcp_fixture() { + kill "$(cat "$dir/mcp.pid" 2>/dev/null)" "$(cat "$dir/native.pid" 2>/dev/null)" \ + "$shim_pid" 2>/dev/null || true + wait "$shim_pid" 2>/dev/null || true + } + { [ -e "$ready" ] && [ -s "$dir/mcp.pid" ]; } \ + || { release_deep_mcp_fixture; fail "the deep MCP fixture never reached its server process"; } + mcp_pid=$(cat "$dir/mcp.pid") + + got=$("$HARNESS" ancestry "$mcp_pid") + [ "$got" = "args claude" ] \ + || { release_deep_mcp_fixture; fail "the MCP server reported '$got', expected 'args claude'; this case proves nothing unless the deepest vantage really answers a foreign harness"; } + + got=$("$HARNESS" ancestry-descent "$shim_pid") + case "$got" in + *"args claude"*) ;; + *) release_deep_mcp_fixture; fail "the descent path did not include the MCP server's foreign args verdict, got '$got'; a cross-check restricted to comm strength is untested unless that vantage is on the path" ;; + esac + case "$got" in + *"comm codex"*) ;; + *) release_deep_mcp_fixture; fail "the descent probe did not reach the native binary at comm strength, got '$got'" ;; + esac + + saw_comm=0 + while read -r strength named; do + [ -n "$strength" ] || continue + [ "$strength" = comm ] || continue + [ "$named" = codex ] \ + || { release_deep_mcp_fixture; fail "a comm-strength vantage on the descent path reported '$named', expected codex"; } + saw_comm=1 + done <<EOF +$got +EOF + [ "$saw_comm" = 1 ] \ + || { release_deep_mcp_fixture; fail "no comm-strength vantage on the descent path, so the guard's strength requirement would reject this tree"; } + + release_deep_mcp_fixture + pass "a foreign args-only verdict at the deepest vantage leaves the comm-strength identity intact" +} + +# Two equally deep foreground leaves must not let ps ordering decide whether the +# chosen path reaches comm strength. The foreign MCP interpreter is spawned first +# in one pass and the native codex binary first in the other; both must resolve to +# the native leaf while the single-path shape remains intact. +test_descent_probe_prefers_comm_strength_when_deepest_leaves_tie() { + local order dir node native mcp_script block entry ready fifo + local shim_pid mcp_pid native_pid got waited + for order in mcp-first native-first; do + dir="$TMP_ROOT/descent-equal-$order" + node=$(named_bin "$dir" node) + native=$(named_bin "$dir/vendor" codex) + ready="$dir/ready" + fifo="$dir/fifo" + mkdir -p "$dir/.claude/mcp" + mkfifo "$fifo" + + block="$dir/block.sh" + cat > "$block" <<'SH' +read -r _ < "$FM_TEST_FIFO" +SH + mcp_script="$dir/.claude/mcp/foo.js" + cp "$block" "$mcp_script" + entry="$dir/codex-cli-entry.sh" + cat > "$entry" <<'SH' +if [ "$FM_TEST_ORDER" = mcp-first ]; then + "$FM_TEST_NODE" "$FM_TEST_MCP" & + printf '%s\n' "$!" > "$FM_TEST_DIR/mcp.pid" + "$FM_TEST_NATIVE" "$FM_TEST_BLOCK" & + printf '%s\n' "$!" > "$FM_TEST_DIR/native.pid" +else + "$FM_TEST_NATIVE" "$FM_TEST_BLOCK" & + printf '%s\n' "$!" > "$FM_TEST_DIR/native.pid" + "$FM_TEST_NODE" "$FM_TEST_MCP" & + printf '%s\n' "$!" > "$FM_TEST_DIR/mcp.pid" +fi +touch "$FM_TEST_READY" +wait +SH + + env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS \ + FM_TEST_ORDER="$order" FM_TEST_DIR="$dir" FM_TEST_NODE="$node" \ + FM_TEST_NATIVE="$native" FM_TEST_BLOCK="$block" FM_TEST_MCP="$mcp_script" \ + FM_TEST_READY="$ready" FM_TEST_FIFO="$fifo" "$node" "$entry" & + shim_pid=$! + + waited=0 + while { [ ! -e "$ready" ] || [ ! -s "$dir/mcp.pid" ] || [ ! -s "$dir/native.pid" ]; } \ + && [ "$waited" -lt 200 ]; do + sleep 0.05 + waited=$((waited + 1)) + done + mcp_pid=$(cat "$dir/mcp.pid" 2>/dev/null || true) + native_pid=$(cat "$dir/native.pid" 2>/dev/null || true) + release_equal_depth_fixture() { + kill "$mcp_pid" "$native_pid" "$shim_pid" 2>/dev/null || true + wait "$shim_pid" 2>/dev/null || true + } + { [ -e "$ready" ] && [ -n "$mcp_pid" ] && [ -n "$native_pid" ]; } \ + || { release_equal_depth_fixture; fail "the $order equal-depth fixture never reached both leaves"; } + + got=$("$HARNESS" ancestry "$mcp_pid") + [ "$got" = "args claude" ] \ + || { release_equal_depth_fixture; fail "the $order MCP leaf reported '$got', expected 'args claude'"; } + got=$("$HARNESS" ancestry "$native_pid") + [ "$got" = "comm codex" ] \ + || { release_equal_depth_fixture; fail "the $order native leaf reported '$got', expected 'comm codex'"; } + + got=$("$HARNESS" ancestry-descent "$shim_pid" "$mcp_pid" "$native_pid") + case "$got" in + "comm codex"*) ;; + *) release_equal_depth_fixture; fail "the $order equal-depth tie did not choose the comm-strength native leaf, got '$got'" ;; + esac + case "$got" in + *"args claude"*) release_equal_depth_fixture; fail "the $order equal-depth tie chose the foreign args-strength leaf" ;; + esac + + release_equal_depth_fixture + done + pass "equal-depth descent ties prefer the comm-strength leaf regardless of spawn order" +} + +# --- 7. Session start's supervision protocol follows the corrected verdict --- + +# The consequence the captain actually hit: the wrong verdict emitted Claude's +# Stop-owned protocol to a Codex primary, so every turn end was blocked for +# missing Claude recovery. +test_supervision_protocol_follows_corrected_verdict() { + local dir home fakebin bin got + dir="$TMP_ROOT/supervision" + home="$dir/home" + mkdir -p "$home/state" "$home/config" + bin=$(named_bin "$dir/codex-tree" codex) + fakebin=$(blind_ancestry_bin "$dir/blind") + + got=$(env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS CLAUDECODE=1 FM_HOME="$home" \ + PATH="$fakebin:$BASE_PATH" "$RENDER") + assert_contains "$got" "primary harness: claude" \ + "with ancestry blinded, the retained marker must still render claude (the case is otherwise vacuous)" + + got=$(env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS CLAUDECODE=1 FM_HOME="$home" \ + "$bin" -c "r=\$(\"$RENDER\"); printf '%s' \"\$r\"") + assert_contains "$got" "primary harness: codex" \ + "a Codex primary carrying a retained CLAUDECODE did not render the Codex protocol" + assert_contains "$got" "Mode: Codex foreground checkpoint." \ + "the rendered block is not Codex's foreground-checkpoint protocol" + assert_not_contains "$got" "Mode: Claude Stop-hook-owned supervision." \ + "the rendered block still carries Claude's Stop-owned protocol" + pass "session start renders the Codex protocol for a Codex primary holding a retained CLAUDECODE" +} + +test_markerless_ancestry_outranks_foreign_marker +test_genuine_marker_and_ancestry_agree +test_cursor_ordering_still_decides_when_ancestry_is_silent +test_retained_cursor_marker_does_not_rename_a_nested_claude +test_pi_signed_survives_agreeing_ancestry +test_interpreter_args_match_does_not_outrank_a_marker +test_native_child_of_an_interpreter_shim_decides_at_comm_strength +test_harness_at_namespace_pid1_is_examined +test_descent_probe_reaches_a_strength_the_top_of_session_cannot +test_descent_probe_ignores_a_sibling_branch_the_walk_cannot_reach +test_descent_probe_tolerates_an_args_only_foreign_verdict_at_the_deepest_vantage +test_descent_probe_prefers_comm_strength_when_deepest_leaves_tie +test_supervision_protocol_follows_corrected_verdict diff --git a/tests/fm-herdr-attached-viewer-live-e2e.test.sh b/tests/fm-herdr-attached-viewer-live-e2e.test.sh new file mode 100755 index 00000000000..061b227de50 --- /dev/null +++ b/tests/fm-herdr-attached-viewer-live-e2e.test.sh @@ -0,0 +1,259 @@ +#!/usr/bin/env bash +# Live attached-viewer regression for the Herdr teardown focus guard. +# +# PR #4131 gated the active-tab close refusal on a LIVE foreground client +# instead of the persisted `.focused` pointer, but only its two detached +# scenarios could be driven live: every pseudo-terminal the runner built +# started at a zero-sized window grid, so Herdr registered no foreground client +# and `terminal title clear` kept answering `no_foreground_client`. That was a +# harness limit, not a product one. `fm-herdr-lab.sh viewer start` now attaches +# a real Herdr TUI over a pty sized before the fork, which turns those +# untestable cases into this regression: +# +# 3. a viewer sitting on the target tab blocks the close; +# 4. a viewer that moves ONTO the target between planning and the mutation +# boundary still blocks it, because the guard re-reads focus there; +# 5. a viewer that moves OFF the target instead keeps its fresh non-target +# focus after the close, rather than being dragged back to a stale +# pre-planning pointer; +# 7. the projection's seeded-tab prune inherits the same refusal. +# +# Scenarios 4 and 5 need a focus change at one exact product boundary, so a +# PATH shim performs the real `tab focus` when the close helper issues its +# planning `pane get`. Every Herdr call, the shim's included, still routes +# through the guarded lab helper against a named non-default session. +# +# The guard submits no model prompts, so the shared live gate runs it wherever +# herdr, jq, and python3 exist. Re-run it after every Herdr upgrade: a release +# that changed the foreground-client contract would surface here first. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +LAB_HELPER=${HERDR_LAB_HELPER:-$ROOT/bin/fm-herdr-lab.sh} + +fm_live_gate default-on FM_HERDR_ATTACHED_VIEWER_LIVE_E2E herdr jq python3 + +[ -x "$LAB_HELPER" ] || { echo "skip: Herdr lab helper not executable at $LAB_HELPER"; exit 0; } + +TMP_ROOT=$(fm_test_tmproot fm-herdr-attached-viewer) +FAKEBIN=$(fm_fakebin "$TMP_ROOT") +FOCUS_SWITCH_CONTROL="$TMP_ROOT/focus-switch" +ORIGINAL_PATH=$PATH +LAB_SESSION=$("$LAB_HELPER" name fm-herdr-attached-viewer) +export LAB_HELPER LAB_SESSION ORIGINAL_PATH FOCUS_SWITCH_CONTROL + +cleanup() { + local status=$? + env PATH="$ORIGINAL_PATH" "$LAB_HELPER" viewer stop "$LAB_SESSION" >/dev/null 2>&1 || status=1 + env PATH="$ORIGINAL_PATH" "$LAB_HELPER" teardown "$LAB_SESSION" || status=1 + fm_test_cleanup + exit "$status" +} +trap cleanup EXIT +"$LAB_HELPER" provision "$LAB_SESSION" || fail "could not provision the isolated Herdr lab" + +lab() { env PATH="$ORIGINAL_PATH" "$LAB_HELPER" run "$LAB_SESSION" "$@"; } + +# The adapter under test calls `herdr` by name. This shim strips the trailing +# session flag the helper will re-append, refuses any caller-supplied one, and +# optionally performs one real focus change at the requested product boundary +# before forwarding the call. +cat > "$FAKEBIN/herdr" <<'SH' +#!/usr/bin/env bash +set -u +args=("$@") +last=$((${#args[@]} - 1)) +flag=$((last - 1)) +if [ "${#args[@]}" -ge 2 ] \ + && [ "${args[$flag]}" = --session ] \ + && [ "${args[$last]}" = "$LAB_SESSION" ]; then + unset "args[$last]" "args[$flag]" +fi +set -- "${args[@]}" +for arg in "$@"; do + case "$arg" in --session|--session=*) exit 9 ;; esac +done +if [ -f "$FOCUS_SWITCH_CONTROL/trigger" ] && [ "$*" = "$(cat "$FOCUS_SWITCH_CONTROL/trigger")" ]; then + rm -f "$FOCUS_SWITCH_CONTROL/trigger" + env PATH="$ORIGINAL_PATH" "$LAB_HELPER" run "$LAB_SESSION" \ + tab focus "$(cat "$FOCUS_SWITCH_CONTROL/tab")" >/dev/null 2>&1 + printf '%s\n' switched > "$FOCUS_SWITCH_CONTROL/done" +fi +exec env PATH="$ORIGINAL_PATH" "$LAB_HELPER" run "$LAB_SESSION" "$@" +SH +chmod +x "$FAKEBIN/herdr" + +mkdir -p "$FOCUS_SWITCH_CONTROL" + +# Arm the shim to run `tab focus <tab>` immediately before the adapter's own +# `pane get <pane>` planning read, which is the last product call before the +# close helper re-reads focus at its mutation boundary. +arm_focus_switch() { # <tab-id> <pane-id> + rm -f "$FOCUS_SWITCH_CONTROL/done" + printf '%s\n' "$1" > "$FOCUS_SWITCH_CONTROL/tab" + printf 'pane get %s\n' "$2" > "$FOCUS_SWITCH_CONTROL/trigger" +} + +assert_focus_switch_fired() { # <label> + [ -f "$FOCUS_SWITCH_CONTROL/done" ] \ + || fail "$1: the mid-close focus switch never ran, so the timing boundary was not exercised" + rm -f "$FOCUS_SWITCH_CONTROL/done" "$FOCUS_SWITCH_CONTROL/trigger" +} + +# Drive one real adapter entry point with the shim on PATH. +drive() { # <function> <argument...> + PATH="$FAKEBIN:$ORIGINAL_PATH" bash -c ' + . "$1/bin/backends/herdr.sh" + fm_backend_herdr_cli() { + local session=$1 + shift + HERDR_SESSION="$session" herdr "$@" --session "$session" + } + fn=$2 + shift 2 + "$fn" "$@" + ' _ "$ROOT" "$@" 2>&1 +} + +# These run inside command substitutions, where `fail` would exit only the +# subshell and let the script carry on with empty ids. They return non-zero +# instead, and every call site carries its own `|| fail`. +new_workspace() { # <label> -> "<workspace>\t<tab>\t<pane>" + local out + out=$(lab workspace create --cwd "$ROOT" --label "$1" --no-focus) || return 1 + printf '%s' "$out" | jq -er ' + [.result.workspace.workspace_id, .result.tab.tab_id, .result.root_pane.pane_id] | @tsv + ' +} + +new_tab() { # <workspace> <label> -> "<tab>\t<pane>" + local out + out=$(lab tab create --workspace "$1" --label "$2" --cwd "$ROOT") || return 1 + printf '%s' "$out" | jq -er '[.result.tab.tab_id, .result.root_pane.pane_id] | @tsv' +} + +focused_tab() { + lab workspace list | jq -er ' + [.result.workspaces[] | select(.focused == true)] | select(length == 1) | .[0].active_tab_id + ' +} + +pane_exists() { lab pane get "$1" >/dev/null 2>&1; } + +foreground_reason() { + lab terminal title clear | jq -er '.result.reason' +} + +# --- the attachment itself, which is what #4131 could not do ---------------- + +REASON=$(foreground_reason) || fail "could not probe the session's foreground client" +[ "$REASON" = no_foreground_client ] \ + || fail "the fresh lab already had a foreground client (reason=$REASON)" + +"$LAB_HELPER" viewer start "$LAB_SESSION" >/dev/null \ + || fail "could not attach a real foreground Herdr viewer over a sized pty" +REASON=$(foreground_reason) || fail "could not probe the session's foreground client" +[ "$REASON" = cleared ] \ + || fail "the attached pty viewer did not register as a foreground client (reason=$REASON)" +pass "attached viewer: a pty sized before the fork registers as a real Herdr foreground client" + +# --- scenario 3: a viewer on the target tab blocks the close --------------- + +FIXTURE=$(new_workspace viewer-active) || fail "could not create the scenario 3 workspace" +IFS=$'\t' read -r WS_THREE TAB_THREE_A PANE_THREE_A <<<"$FIXTURE" +# A second tab keeps the close a plain one rather than an emptying-workspace plan. +new_tab "$WS_THREE" viewer-active-b >/dev/null || fail "could not create the scenario 3 companion tab" +lab tab focus "$TAB_THREE_A" >/dev/null || fail "could not focus the scenario 3 target tab" +[ "$(focused_tab)" = "$TAB_THREE_A" ] || fail "scenario 3 did not start focused on the target tab" + +OUT=$(drive fm_backend_herdr_projection_close_pane_focus_preserving "$LAB_SESSION" "$PANE_THREE_A") +STATUS=$? +[ "$STATUS" -ne 0 ] || fail "a live viewer on the target tab did not block the close: $OUT" +assert_contains "$OUT" "target is the captain's active tab" \ + "the live-viewer refusal did not name the captain's active tab: $OUT" +pane_exists "$PANE_THREE_A" \ + || fail "the close proceeded and destroyed the tab the live viewer was watching" +pass "attached viewer: a live client on the target tab refuses the close and keeps the pane" + +# --- scenario 4: the viewer moves ONTO the target mid-close ---------------- + +FIXTURE=$(new_workspace viewer-late-on) || fail "could not create the scenario 4 workspace" +IFS=$'\t' read -r WS_FOUR TAB_FOUR_A PANE_FOUR_A <<<"$FIXTURE" +FIXTURE=$(new_tab "$WS_FOUR" viewer-late-on-b) || fail "could not create the scenario 4 companion tab" +IFS=$'\t' read -r TAB_FOUR_B _ <<<"$FIXTURE" +lab tab focus "$TAB_FOUR_B" >/dev/null || fail "could not focus away from the scenario 4 target" +[ "$(focused_tab)" = "$TAB_FOUR_B" ] || fail "scenario 4 did not start focused off the target tab" + +arm_focus_switch "$TAB_FOUR_A" "$PANE_FOUR_A" +OUT=$(drive fm_backend_herdr_projection_close_pane_focus_preserving "$LAB_SESSION" "$PANE_FOUR_A") +STATUS=$? +assert_focus_switch_fired "scenario 4" +[ "$STATUS" -ne 0 ] \ + || fail "a viewer that moved onto the target after planning did not block the close: $OUT" +assert_contains "$OUT" "target is the captain's active tab" \ + "the late-switch refusal did not come from the fresh active-tab check: $OUT" +pane_exists "$PANE_FOUR_A" \ + || fail "the close destroyed a tab the viewer had moved onto before the mutation boundary" +pass "attached viewer: focus moving onto the target between planning and mutation still blocks the close" + +# --- scenario 5: the viewer moves OFF the target mid-close ----------------- + +FIXTURE=$(new_workspace viewer-late-off) || fail "could not create the scenario 5 workspace" +IFS=$'\t' read -r WS_FIVE TAB_FIVE_A PANE_FIVE_A <<<"$FIXTURE" +FIXTURE=$(new_tab "$WS_FIVE" viewer-late-off-b) || fail "could not create the scenario 5 companion tab" +IFS=$'\t' read -r TAB_FIVE_B _ <<<"$FIXTURE" +lab tab focus "$TAB_FIVE_A" >/dev/null || fail "could not focus the scenario 5 target tab" +[ "$(focused_tab)" = "$TAB_FIVE_A" ] || fail "scenario 5 did not start focused on the target tab" + +arm_focus_switch "$TAB_FIVE_B" "$PANE_FIVE_A" +OUT=$(drive fm_backend_herdr_projection_close_pane_focus_preserving "$LAB_SESSION" "$PANE_FIVE_A") +STATUS=$? +assert_focus_switch_fired "scenario 5" +[ "$STATUS" -eq 0 ] \ + || fail "the close was refused even though the live viewer had moved off the target: $OUT" +if pane_exists "$PANE_FIVE_A"; then + fail "the close reported success but left the target pane behind" +fi +[ "$(focused_tab)" = "$TAB_FIVE_B" ] \ + || fail "the close did not preserve the viewer's fresh non-target focus (focus is $(focused_tab), expected $TAB_FIVE_B)" +pass "attached viewer: a close preserves the fresh non-target focus the viewer moved to" + +# --- scenario 7: the projection's seeded-tab prune inherits the refusal ---- + +FIXTURE=$(new_workspace viewer-seeded) || fail "could not create the scenario 7 workspace" +IFS=$'\t' read -r WS_SEVEN TAB_SEVEN_SEEDED PANE_SEVEN_SEEDED <<<"$FIXTURE" +FIXTURE=$(new_tab "$WS_SEVEN" fm-viewer-seeded-task) || fail "could not create the scenario 7 task tab" +IFS=$'\t' read -r _ PANE_SEVEN_TASK <<<"$FIXTURE" +lab tab list --workspace "$WS_SEVEN" \ + | jq -e --arg tab "$TAB_SEVEN_SEEDED" '.result.tabs[] | select(.tab_id == $tab) | .label == "1"' >/dev/null \ + || fail "the seeded tab is not the label-1 default tab the prune identifies" +lab tab focus "$TAB_SEVEN_SEEDED" >/dev/null || fail "could not focus the seeded tab" +[ "$(focused_tab)" = "$TAB_SEVEN_SEEDED" ] || fail "scenario 7 did not start focused on the seeded tab" + +OUT=$(drive fm_backend_herdr_workspace_prune_seeded_default_tab \ + "$LAB_SESSION" "$WS_SEVEN" "$TAB_SEVEN_SEEDED" focus-preserving) +STATUS=$? +[ "$STATUS" -ne 0 ] || fail "the seeded prune did not refuse the tab a live viewer was watching: $OUT" +assert_contains "$OUT" "target is the captain's active tab" \ + "the seeded prune refusal did not come from the live-viewer guard: $OUT" +pane_exists "$PANE_SEVEN_SEEDED" \ + || fail "the seeded prune closed the tab the live viewer was watching" +pane_exists "$PANE_SEVEN_TASK" || fail "the seeded prune disturbed the task pane" +pass "attached viewer: the projection seeded-tab prune refuses while a live client watches it" + +# --- detaching restores the no-client contract the detached tests rely on --- + +"$LAB_HELPER" viewer stop "$LAB_SESSION" >/dev/null \ + || fail "could not detach the lab viewer" +REASON=$(foreground_reason) || fail "could not probe the session's foreground client" +[ "$REASON" = no_foreground_client ] \ + || fail "the lab still reported a foreground client after the viewer stopped (reason=$REASON)" +OUT=$(drive fm_backend_herdr_projection_close_pane_focus_preserving "$LAB_SESSION" "$PANE_THREE_A") +STATUS=$? +[ "$STATUS" -eq 0 ] || fail "the same close was still refused after the viewer detached: $OUT" +if pane_exists "$PANE_THREE_A"; then + fail "the detached close reported success but left the pane behind" +fi +pass "attached viewer: detaching releases the refusal, so the guard tracks the client and not the pointer" diff --git a/tests/fm-herdr-lab.test.sh b/tests/fm-herdr-lab.test.sh index 116cda2495b..c1cc0ff20b9 100755 --- a/tests/fm-herdr-lab.test.sh +++ b/tests/fm-herdr-lab.test.sh @@ -64,6 +64,12 @@ case "$1 ${2:-}" in [ "${FM_FAKE_HERDR_DELETE_FAIL:-}" != 1 ] || exit 93 printf '%s\n' deleted > "$state/$session" ;; + "terminal title") + [ "${FM_FAKE_HERDR_TITLE_FAIL:-}" != 1 ] || exit 94 + reason=no_foreground_client + [ ! -f "$state/$session.foreground" ] || reason=$(cat "$state/$session.foreground") + jq -nc --arg reason "$reason" '{result:{reason:$reason,type:"client_window_title"}}' + ;; *) printf '%s\n' '{"ok":true}' ;; @@ -82,6 +88,7 @@ run_with_fake() { FM_FAKE_HERDR_SERVER_DELAY="${FM_FAKE_HERDR_SERVER_DELAY:-0}" \ FM_FAKE_HERDR_FAST_POLL="${FM_FAKE_HERDR_FAST_POLL:-}" \ FM_FAKE_HERDR_DELETE_FAIL="${FM_FAKE_HERDR_DELETE_FAIL:-}" \ + FM_FAKE_HERDR_TITLE_FAIL="${FM_FAKE_HERDR_TITLE_FAIL:-}" \ FM_HERDR_LAB_STATE_DIR="$TRIPWIRES" \ "$@" } @@ -213,6 +220,9 @@ test_timed_out_provision_cancels_late_launch() { cat > "$FAKEBIN/sleep" <<'SH' #!/usr/bin/env bash if [ "${FM_FAKE_HERDR_FAST_POLL:-}" = 1 ]; then + while [ -n "${FM_FAKE_HERDR_WAIT_MARKER:-}" ] && [ ! -f "$FM_FAKE_HERDR_WAIT_MARKER" ]; do + "$FM_FAKE_HERDR_REAL_SLEEP" 0.01 + done exit 0 fi exec "$FM_FAKE_HERDR_REAL_SLEEP" "$@" @@ -234,6 +244,260 @@ SH pass "fm-herdr-lab: timed-out provisioning cancels the launch before teardown" } + +# The pty attachment itself needs a real Herdr client, so the live guard +# tests/fm-herdr-attached-viewer-live-e2e.test.sh owns that proof. What is +# portable is who the helper will ever attach to, and who it will signal. +test_viewer_refuses_unowned_sessions() { + local name="fm-lab-viewer-guard-$$" status=0 out + : > "$FAKE_LOG" + out=$(run_with_fake fm_herdr_lab_viewer_start "$name" 2>&1) || status=$? + expect_code 1 "$status" "a session without an ownership tripwire must not be attached to" + assert_contains "$out" "does not own" \ + "the viewer refusal did not name the missing ownership record" + [ ! -s "$FAKE_LOG" ] \ + || fail "the unowned-session refusal reached Herdr instead of refusing first" + + status=0 + run_with_fake fm_herdr_lab_viewer_start default >/dev/null 2>&1 || status=$? + expect_code 1 "$status" "the default session must never be attached to" + pass "fm-herdr-lab: the viewer attaches only to a session this lab owns" +} + +start_viewer_fixture() { + local pair=$1 + ( + "$REAL_SLEEP" 20 & + printf '%s\n' "$!" > "$pair" + wait + ) & + FIXTURE_LAUNCHER_PID=$! + while [ ! -s "$pair" ]; do + "$REAL_SLEEP" 0.01 + done + FIXTURE_VIEWER_PID=$(cat "$pair") +} + +write_viewer_record() { + local record=$1 launcher_pid=$2 viewer_pid=$3 launcher_start viewer_start + launcher_start=$(fm_herdr_lab_process_start "$launcher_pid") || fail "could not identify launcher fixture process" + viewer_start=$(fm_herdr_lab_process_start "$viewer_pid") || fail "could not identify viewer fixture process" + printf 'launcher_pid=%s\nlauncher_start=%s\nviewer_pid=%s\nviewer_start=%s\n' \ + "$launcher_pid" "$launcher_start" "$viewer_pid" "$viewer_start" > "$record" +} + +test_viewer_start_cancels_an_unrecorded_launcher() { + local name="fm-lab-viewer-late-$$" out status=0 launcher_pid + local started="$TMP_ROOT/viewer-launcher-started" + run_with_fake fm_herdr_lab_provision "$name" || fail "viewer-late fixture provision failed" + cat > "$FAKEBIN/python3" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$$" > "$FM_FAKE_VIEWER_STARTED" +exec "$FM_FAKE_HERDR_REAL_SLEEP" 20 +SH + chmod +x "$FAKEBIN/python3" + out=$(FM_FAKE_HERDR_FAST_POLL=1 FM_FAKE_HERDR_WAIT_MARKER="$started" \ + FM_FAKE_VIEWER_STARTED="$started" run_with_fake fm_herdr_lab_viewer_start "$name" 2>&1) || status=$? + rm -f "$FAKEBIN/python3" + expect_code 1 "$status" "an unrecorded launcher must not outlive viewer start" + assert_present "$started" "delayed viewer launcher did not start" + launcher_pid=$(cat "$started") + kill -0 "$launcher_pid" 2>/dev/null && fail "timed-out viewer launcher remained alive" + assert_contains "$out" "did not become the foreground client" "launcher timeout was unclear" + run_with_fake fm_herdr_lab_teardown "$name" || fail "viewer-late fixture teardown failed" + pass "fm-herdr-lab: timed-out viewer startup cancels its exact launcher" +} + +test_viewer_timeout_allows_launcher_escalation() { + local launcher_pid started="$TMP_ROOT/viewer-grace-started" + local terminating="$TMP_ROOT/viewer-grace-terminating" completed="$TMP_ROOT/viewer-grace-completed" + cat > "$FAKEBIN/viewer-launcher" <<'SH' +#!/usr/bin/env bash +trap 'printf "" > "$FM_FAKE_VIEWER_TERMINATING"; "$FM_FAKE_HERDR_REAL_SLEEP" 1.2; printf "" > "$FM_FAKE_VIEWER_COMPLETED"; exit 0' TERM +printf '' > "$FM_FAKE_VIEWER_STARTED" +while :; do + "$FM_FAKE_HERDR_REAL_SLEEP" 0.1 +done +SH + chmod +x "$FAKEBIN/viewer-launcher" + FM_FAKE_HERDR_REAL_SLEEP="$REAL_SLEEP" FM_FAKE_VIEWER_STARTED="$started" \ + FM_FAKE_VIEWER_TERMINATING="$terminating" FM_FAKE_VIEWER_COMPLETED="$completed" \ + "$FAKEBIN/viewer-launcher" & + launcher_pid=$! + while [ ! -f "$started" ]; do + "$REAL_SLEEP" 0.01 + done + run_with_fake fm_herdr_lab_cancel_viewer_launcher "$launcher_pid" + assert_present "$terminating" "timed-out viewer launcher did not receive TERM" + assert_present "$completed" "viewer launcher was killed before completing child escalation" + pass "fm-herdr-lab: startup timeout allows launcher child escalation" +} + +test_viewer_start_requires_its_owned_process() { + local name="fm-lab-viewer-ownership-$$" out status=0 marker="$TMP_ROOT/viewer-launched" + run_with_fake fm_herdr_lab_provision "$name" || fail "viewer-ownership fixture provision failed" + printf '%s\n' cleared > "$FAKE_STATE/$name.foreground" + cat > "$FAKEBIN/python3" <<'SH' +#!/usr/bin/env bash +: > "$FM_FAKE_VIEWER_MARKER" +exit 0 +SH + chmod +x "$FAKEBIN/python3" + out=$(FM_FAKE_HERDR_FAST_POLL=1 FM_FAKE_VIEWER_MARKER="$marker" \ + run_with_fake fm_herdr_lab_viewer_start "$name" 2>&1) || status=$? + rm -f "$FAKEBIN/python3" + expect_code 1 "$status" "a foreign foreground client must not satisfy viewer start" + assert_present "$marker" "viewer ownership fixture did not launch" + assert_contains "$out" "did not become the foreground client" "ownership failure did not time out clearly" + assert_not_contains "$out" "viewer attached" "start claimed a foreign foreground client as its own" + printf '%s\n' no_foreground_client > "$FAKE_STATE/$name.foreground" + run_with_fake fm_herdr_lab_teardown "$name" || fail "viewer-ownership fixture teardown failed" + pass "fm-herdr-lab: viewer start requires an identity-matched owned process" +} + +test_viewer_stop_only_signals_owned_processes() { + local name="fm-lab-viewer-stop-$$" record status=0 holder_pid pair="$TMP_ROOT/viewer-stop-pair" + run_with_fake fm_herdr_lab_provision "$name" || fail "viewer-stop fixture provision failed" + record=$(run_with_fake fm_herdr_lab_viewer_record_path "$name") + + # No record: a client attached by someone else is not ours to kill. + printf '%s\n' cleared > "$FAKE_STATE/$name.foreground" + run_with_fake fm_herdr_lab_viewer_stop "$name" \ + || fail "stopping with no recorded viewer must succeed without touching a foreign client" + [ "$(cat "$FAKE_STATE/$name.foreground")" = cleared ] \ + || fail "an unrecorded foreground client was detached by the lab helper" + + # A recorded viewer is signalled until it exits and the session reports no + # foreground client again. + start_viewer_fixture "$pair" + write_viewer_record "$record" "$FIXTURE_LAUNCHER_PID" "$FIXTURE_VIEWER_PID" + status=0 + FM_FAKE_HERDR_FAST_POLL=1 run_with_fake fm_herdr_lab_viewer_stop "$name" \ + >/dev/null 2>&1 || status=$? + expect_code 1 "$status" "stop must fail while the session still reports a foreground client" + wait "$FIXTURE_LAUNCHER_PID" 2>/dev/null || true + kill -0 "$FIXTURE_VIEWER_PID" 2>/dev/null && fail "stop left the recorded viewer process running" + assert_present "$record" "a failed detach discarded the viewer record it still needs" + + sleep 20 & + holder_pid=$! + printf 'launcher_pid=%s\nlauncher_start=not-this-process\nviewer_pid=%s\nviewer_start=not-this-process\n' \ + "$holder_pid" "$holder_pid" > "$record" + printf '%s\n' no_foreground_client > "$FAKE_STATE/$name.foreground" + run_with_fake fm_herdr_lab_viewer_stop "$name" || fail "stop rejected a stale process record" + kill -0 "$holder_pid" 2>/dev/null || fail "stop signalled a PID whose recorded identity did not match" + kill "$holder_pid" 2>/dev/null || true + wait "$holder_pid" 2>/dev/null || true + + run_with_fake fm_herdr_lab_viewer_stop "$name" || fail "stop failed once the client had detached" + assert_absent "$record" "a confirmed detach left the viewer record behind" + run_with_fake fm_herdr_lab_teardown "$name" || fail "teardown after viewer stop failed" + pass "fm-herdr-lab: viewer stop signals only recorded processes and confirms the detach" +} + +test_viewer_stop_requires_the_recorded_parent() { + local name="fm-lab-viewer-parent-$$" record launcher_pid viewer_pid + run_with_fake fm_herdr_lab_provision "$name" || fail "viewer-parent fixture provision failed" + record=$(run_with_fake fm_herdr_lab_viewer_record_path "$name") + sleep 20 & + launcher_pid=$! + sleep 20 & + viewer_pid=$! + write_viewer_record "$record" "$launcher_pid" "$viewer_pid" + printf '%s\n' no_foreground_client > "$FAKE_STATE/$name.foreground" + run_with_fake fm_herdr_lab_viewer_stop "$name" || fail "parent-mismatch stop failed" + kill -0 "$launcher_pid" 2>/dev/null || fail "stop signalled a launcher without its recorded child" + kill -0 "$viewer_pid" 2>/dev/null || fail "stop signalled a viewer outside the recorded launcher" + kill "$launcher_pid" "$viewer_pid" 2>/dev/null || true + wait "$launcher_pid" 2>/dev/null || true + wait "$viewer_pid" 2>/dev/null || true + run_with_fake fm_herdr_lab_teardown "$name" || fail "viewer-parent fixture teardown failed" + pass "fm-herdr-lab: viewer ownership requires the recorded parent" +} + +test_interrupted_viewer_start_cancels_launcher() { + local name="fm-lab-viewer-interrupt-$$" command_pid launcher_pid status=0 + local started="$TMP_ROOT/viewer-interrupt-started" attached="$TMP_ROOT/viewer-interrupt-attached" + run_with_fake fm_herdr_lab_provision "$name" || fail "viewer-interrupt fixture provision failed" + cat > "$FAKEBIN/python3" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$$" > "$FM_FAKE_VIEWER_STARTED" +"$FM_FAKE_HERDR_REAL_SLEEP" 0.5 +: > "$FM_FAKE_VIEWER_ATTACHED" +printf '%s\n' cleared > "$FM_FAKE_HERDR_STATE/$FM_FAKE_VIEWER_SESSION.foreground" +exec "$FM_FAKE_HERDR_REAL_SLEEP" 20 +SH + chmod +x "$FAKEBIN/python3" + FM_FAKE_VIEWER_STARTED="$started" FM_FAKE_VIEWER_ATTACHED="$attached" \ + FM_FAKE_VIEWER_SESSION="$name" run_with_fake exec "$ROOT/bin/fm-herdr-lab.sh" \ + viewer start "$name" >/dev/null 2>&1 & + command_pid=$! + while [ ! -f "$started" ]; do + "$REAL_SLEEP" 0.01 + done + launcher_pid=$(cat "$started") + kill -TERM "$command_pid" + wait "$command_pid" || status=$? + rm -f "$FAKEBIN/python3" + [ "$status" -ne 0 ] || fail "interrupted viewer start unexpectedly succeeded" + "$REAL_SLEEP" 0.6 + kill -0 "$launcher_pid" 2>/dev/null && fail "interrupted viewer start left its launcher running" + assert_absent "$attached" "interrupted viewer start attached after its command exited" + [ ! -f "$FAKE_STATE/$name.foreground" ] || fail "interrupted viewer start left a foreground client" + run_with_fake fm_herdr_lab_teardown "$name" || fail "viewer-interrupt fixture teardown failed" + pass "fm-herdr-lab: interrupted viewer start cancels its launcher" +} + +test_teardown_refuses_while_viewer_attached() { + local name="fm-lab-viewer-teardown-$$" record status=0 pair="$TMP_ROOT/viewer-teardown-pair" + run_with_fake fm_herdr_lab_provision "$name" || fail "viewer-teardown fixture provision failed" + record=$(run_with_fake fm_herdr_lab_viewer_record_path "$name") + printf '%s\n' cleared > "$FAKE_STATE/$name.foreground" + start_viewer_fixture "$pair" + write_viewer_record "$record" "$FIXTURE_LAUNCHER_PID" "$FIXTURE_VIEWER_PID" + : > "$FAKE_LOG" + FM_FAKE_HERDR_FAST_POLL=1 run_with_fake fm_herdr_lab_teardown "$name" \ + >/dev/null 2>&1 || status=$? + expect_code 1 "$status" "teardown must refuse while an owned viewer is still attached" + [ "$(cat "$FAKE_STATE/$name")" = running ] \ + || fail "the refused teardown stopped the lab session anyway" + assert_no_grep "session delete $name" "$FAKE_LOG" \ + "the refused teardown still reached the destructive delete" + + printf '%s\n' no_foreground_client > "$FAKE_STATE/$name.foreground" + run_with_fake fm_herdr_lab_teardown "$name" || fail "teardown after the viewer detached failed" + pass "fm-herdr-lab: teardown refuses to destroy a session an attached viewer still holds" +} + +test_viewer_stop_retains_record_when_detach_is_unreadable() { + local name="fm-lab-viewer-unreadable-$$" record status=0 + run_with_fake fm_herdr_lab_provision "$name" || fail "unreadable-detach fixture provision failed" + record=$(run_with_fake fm_herdr_lab_viewer_record_path "$name") + printf 'launcher_pid=99999999\nlauncher_start=stale\nviewer_pid=99999999\nviewer_start=stale\n' > "$record" + FM_FAKE_HERDR_FAST_POLL=1 FM_FAKE_HERDR_TITLE_FAIL=1 \ + run_with_fake fm_herdr_lab_viewer_stop "$name" >/dev/null 2>&1 || status=$? + expect_code 1 "$status" "an unreadable detach result on a running session must fail closed" + assert_present "$record" "an unreadable detach result discarded the ownership record" + printf '%s\n' no_foreground_client > "$FAKE_STATE/$name.foreground" + run_with_fake fm_herdr_lab_teardown "$name" || fail "teardown after a confirmed detach failed" + pass "fm-herdr-lab: unreadable detach results fail closed on running sessions" +} + +test_viewer_launcher_refuses_unsafe_arguments() { + local launcher="$ROOT/bin/fm-herdr-lab-viewer.py" status=0 + command -v python3 >/dev/null 2>&1 || { pass "fm-herdr-lab: viewer launcher argument guard (skipped, no python3)"; return; } + python3 "$launcher" default "$TMP_ROOT/pid" >/dev/null 2>&1 || status=$? + expect_code 2 "$status" "the launcher must refuse the default session" + status=0 + python3 "$launcher" arbitrary-session "$TMP_ROOT/pid" >/dev/null 2>&1 || status=$? + expect_code 2 "$status" "the launcher must refuse a non-lab session name" + status=0 + python3 "$launcher" fm-lab-args relative-pidfile >/dev/null 2>&1 || status=$? + expect_code 2 "$status" "the launcher must refuse a relative pidfile path" + assert_absent "$TMP_ROOT/pid" "a refused launch still wrote a pid record" + pass "fm-herdr-lab: the viewer launcher refuses unsafe sessions and pidfiles" +} + test_refuses_unsafe_names test_provision_run_and_guarded_teardown test_missing_tripwire_blocks_destruction @@ -241,4 +505,14 @@ test_changed_default_trips_after_teardown test_stopped_owned_lab_can_reprovision test_failed_delete_retains_tripwire test_timed_out_provision_cancels_late_launch +test_viewer_refuses_unowned_sessions +test_viewer_start_cancels_an_unrecorded_launcher +test_viewer_timeout_allows_launcher_escalation +test_viewer_start_requires_its_owned_process +test_viewer_stop_only_signals_owned_processes +test_viewer_stop_requires_the_recorded_parent +test_interrupted_viewer_start_cancels_launcher +test_teardown_refuses_while_viewer_attached +test_viewer_stop_retains_record_when_detach_is_unreadable +test_viewer_launcher_refuses_unsafe_arguments printf '\nall fm-herdr-lab tests passed\n' diff --git a/tests/fm-herdr-pi-stale-registration-live-e2e.test.sh b/tests/fm-herdr-pi-stale-registration-live-e2e.test.sh new file mode 100755 index 00000000000..3df52274ba2 --- /dev/null +++ b/tests/fm-herdr-pi-stale-registration-live-e2e.test.sh @@ -0,0 +1,159 @@ +#!/usr/bin/env bash +# Default-on live guard for the Herdr stale-registration classifier (issue +# #4115) against the REAL Pi harness under the REAL Herdr binary. +# +# The defect: Herdr keeps a Pi registration (`agent get` -> agent=pi, +# agent_status=idle) after the Pi process has exited to a plain shell whenever +# a nested interactive shell sits under the pane's top shell - the crew shape, +# where `treehouse get` leaves a worktree shell under the pane's login shell. +# The adapter now proves an agent at process level before trusting a +# registration, and this guard measures the two vendor facts that proof rests +# on, which no fixture can prove: +# +# 1. how Pi presents in `pane process-info` (on Herdr 0.9.0 the kernel name +# is `node` and only argv0 says `pi`), so the shared process classifier +# must still attribute the running harness as `agent`; +# 2. whether this Herdr release still leaves the registration behind after +# Pi quits under a nested shell, so the stale-registration branch is +# exercised against the real record rather than a canned one. +# +# It fails naming the Herdr and Pi versions when either fact drifts. Pi is +# launched with no prompt and quit immediately, so no model token is spent and +# the shared live gate runs it by default wherever both tools are installed. +# Run it after every Herdr or Pi upgrade and before trusting a refreshed +# docs/verification/runtime-backends.md "Stale agent registration" entry. +# +# Always runs on a private, named, throwaway lab session, never the default +# one (tests/herdr-test-safety.sh; bin/fm-herdr-lab.sh owns the isolation). +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" + +fail() { printf 'not ok - %s\n' "$1" >&2; exit 1; } +pass() { printf 'ok - %s\n' "$1"; } +note() { printf '# %s\n' "$1"; } + +fm_live_gate default-on FM_HERDR_PI_STALE_REGISTRATION_LIVE_E2E herdr pi jq + +# shellcheck source=tests/herdr-test-safety.sh +. "$ROOT/tests/herdr-test-safety.sh" +herdr_forget_inherited_pane + +HERDR_VERSION=$(herdr --version 2>&1 | head -1) +HERDR_VERSION=${HERDR_VERSION#herdr } +PI_VERSION=$(pi --version 2>/dev/null | head -1 | tr -d '\r') +[ -n "$PI_VERSION" ] || PI_VERSION=unknown +version_fail() { # <message> + fail "$1 [herdr $HERDR_VERSION, pi $PI_VERSION]" +} + +SESSION="fm-lab-pi-stale-$$" +export HERDR_SESSION="$SESSION" +SCRATCH= +cleanup_all() { + local status=$? + [ -n "$SCRATCH" ] && rm -rf "$SCRATCH" + herdr_safe_stop_and_delete "$SESSION" + exit "$status" +} +trap cleanup_all EXIT +fm_herdr_lab_prepare "$SESSION" || fail "could not prepare isolated Herdr lab session" + +SCRATCH=$(mktemp -d "${TMPDIR:-/tmp}/fm-pi-stale.XXXXXX") +SCRATCH=$(cd "$SCRATCH" && pwd) +mkdir -p "$SCRATCH/cwd" + +# shellcheck source=/dev/null +. "$ROOT/bin/fm-backend.sh" +fm_backend_source herdr || fail "fm_backend_source herdr failed" + +lab() { fm_herdr_lab_cli "$SESSION" "$@"; } + +# prepare only records the tripwire; the adapter's own server-ensure starts +# the lab session's server exactly as a spawn would. +fm_backend_herdr_server_ensure "$SESSION" || fail "could not start the isolated Herdr lab server" +WS=$(lab workspace create --label fm-pi-stale --cwd "$SCRATCH/cwd" 2>&1) \ + || fail "could not create the lab workspace: $WS" +PANE_ID=$(printf '%s' "$WS" | jq -r '.result.root_pane.pane_id // empty') +[ -n "$PANE_ID" ] || fail "workspace create did not return a root pane id" +TARGET="$SESSION:$PANE_ID" + +wait_process_state() { # <expected> <tries> + local expected=$1 tries=$2 got i=0 + while [ "$i" -lt "$tries" ]; do + got=$(fm_backend_herdr_pane_process_state "$SESSION" "$PANE_ID") + [ "$got" = "$expected" ] && return 0 + sleep 0.2 + i=$((i + 1)) + done + return 1 +} + +registered_status() { + herdr agent get "$PANE_ID" --session "$SESSION" 2>/dev/null | jq -r '.result.agent.agent_status // empty' +} + +# The crew shape: a nested interactive shell under the pane's top shell, then +# the real Pi TUI with no prompt. +lab pane run "$PANE_ID" zsh >/dev/null 2>&1 || fail "could not start the nested shell in the pane" +sleep 1 +lab pane run "$PANE_ID" pi >/dev/null 2>&1 || fail "could not start pi in the pane" + +# Herdr creates the record with its own placeholder status (`unknown`, verified +# 0.9.0) the moment it notices Pi, before Pi's extension reports a lifecycle +# state; only a lifecycle state is the registration this guard is about. +STATUS= +for _ in $(seq 1 300); do + STATUS=$(registered_status) + case "$STATUS" in working|idle|done|blocked) break ;; esac + sleep 0.2 +done +case "$STATUS" in + working|idle|done|blocked) ;; + *) version_fail \ + "pi never reported a lifecycle state to Herdr in this pane (agent get read '${STATUS:-agent_not_found}' for 60s); the herdr pi integration (~/.pi/agent/extensions/herdr-agent-state.ts) is what reports it" ;; +esac + +wait_process_state agent 100 || version_fail \ + "pi is running and registered ($STATUS) but pane process-info reads '$(fm_backend_herdr_pane_process_state "$SESSION" "$PANE_ID")', not 'agent'. Observed foreground: $(herdr pane process-info --pane "$PANE_ID" --session "$SESSION" 2>/dev/null | jq -c '.result.process_info.foreground_processes'). Teach bin/fm-agent-process-lib.sh's fm_agent_process_classify the identity this release actually reports" +FOREGROUND=$(herdr pane process-info --pane "$PANE_ID" --session "$SESSION" 2>/dev/null \ + | jq -c '[.result.process_info.foreground_processes[] | {name, argv0}]') +STATE=$(fm_backend_agent_state herdr "$TARGET") +[ "$STATE" = alive ] || version_fail "a running, registered pi reads '$STATE' rather than 'alive' (registration '$(registered_status)', pane state '$(fm_backend_herdr_pane_agent_state "$SESSION" "$PANE_ID")', process state '$(fm_backend_herdr_pane_process_state "$SESSION" "$PANE_ID")', agent get: $(herdr agent get "$PANE_ID" --session "$SESSION" 2>&1 | tr -d '\n'))" +note "pi $PI_VERSION under herdr $HERDR_VERSION: registered $STATUS, foreground $FOREGROUND" +pass "real herdr $HERDR_VERSION + pi $PI_VERSION: a running registered pi classifies alive at process level" + +# Quit Pi to the nested shell. A slash command can open a completion popup that +# swallows the first Enter, so one extra Enter is allowed before judging. +lab pane send-text "$PANE_ID" '/quit' >/dev/null 2>&1 || fail "could not type /quit" +sleep 0.5 +lab pane send-keys "$PANE_ID" Enter >/dev/null 2>&1 || fail "could not submit /quit" +if ! wait_process_state shell 50; then + lab pane send-keys "$PANE_ID" Enter >/dev/null 2>&1 || true + wait_process_state shell 150 || version_fail \ + "pi did not exit to a shell within 40s of /quit; pane process-info reads '$(fm_backend_herdr_pane_process_state "$SESSION" "$PANE_ID")'" +fi + +# Let Herdr settle whatever release it is going to do, then read the record. +sleep 2 +STATUS=$(registered_status) +PANE_STATE=$(fm_backend_herdr_pane_agent_state "$SESSION" "$PANE_ID") +STATE=$(fm_backend_agent_state herdr "$TARGET") +BUSY=$(fm_backend_herdr_busy_state "$TARGET") +[ "$STATE" = dead ] || version_fail \ + "after pi quit to a shell the endpoint recovers as '$STATE' (pane state '$PANE_STATE', registration '${STATUS:-none}') rather than 'dead'; every relaunch would be refused" +[ "$BUSY" != busy ] || version_fail "a shell-only pane after pi quit reads busy (registration '${STATUS:-none}')" +if [ -n "$STATUS" ]; then + [ "$PANE_STATE" = stale-agent ] || version_fail \ + "Herdr kept the registration ($STATUS) over the shell-only pane but the classifier reads '$PANE_STATE' rather than 'stale-agent'" + note "herdr $HERDR_VERSION kept the pi registration ($STATUS) after /quit under a nested shell: the stale-registration branch is exercised" + pass "real herdr $HERDR_VERSION + pi $PI_VERSION: the registration left behind by a quit pi reads stale-agent and recovers as dead" +else + [ "$PANE_STATE" = no-agent ] || version_fail \ + "Herdr released the registration but the pane reads '$PANE_STATE' rather than 'no-agent'" + note "herdr $HERDR_VERSION released the pi registration after /quit under a nested shell; the stale-registration branch was not exercised by this release, the agent-free verdict still held through agent_not_found" + pass "real herdr $HERDR_VERSION + pi $PI_VERSION: a quit pi under a nested shell recovers as dead" +fi diff --git a/tests/fm-inactive-reconcile.test.sh b/tests/fm-inactive-reconcile.test.sh index 6a596dab169..b36f1c03858 100755 --- a/tests/fm-inactive-reconcile.test.sh +++ b/tests/fm-inactive-reconcile.test.sh @@ -233,7 +233,7 @@ test_secondmate_ledger_delivery_carries_report_and_failure() { mkdir -p "$MATE/data/scout" printf '# findings\n' > "$MATE/data/scout/report.md" write_child "$MATE" boom 'failed: build broke' - write_child "$MATE" replaced-pr $'working: old PR https://example.test/owner/repo/pull/11\ndone: replacement PR https://example.test/owner/repo/pull/22' + write_child "$MATE" replaced-pr $'working: old PR https://example.test/owner/repo/pull/11\ndone: PR https://example.test/owner/repo/pull/22' awk '$0 !~ /^pr=/' "$MATE/state/replaced-pr.meta" > "$MATE/state/replaced-pr.meta.tmp" mv "$MATE/state/replaced-pr.meta.tmp" "$MATE/state/replaced-pr.meta" FM_FAKE_CREW_STATE='unknown' run_reconcile "$MATE" @@ -244,8 +244,8 @@ test_secondmate_ledger_delivery_carries_report_and_failure() { "$MAIN/state/mate.status" || fail "scout delivery lost its report pointer: $(cat "$MAIN/state/mate.status")" grep -Fxq "failed [key=$boom_key]: child boom failed: build broke pr=https://example.test/owner/repo/pull/1 mode=no-mistakes yolo=off" \ "$MAIN/state/mate.status" || fail "failed line was not delivered under the failed verb: $(cat "$MAIN/state/mate.status")" - grep -Fxq "done [key=$replaced_key]: child replaced-pr done: replacement PR https://example.test/owner/repo/pull/22 pr=https://example.test/owner/repo/pull/22 mode=no-mistakes yolo=off" \ - "$MAIN/state/mate.status" || fail "ledger fallback did not prefer the terminal line PR: $(cat "$MAIN/state/mate.status")" + grep -Fxq "done [key=$replaced_key]: child replaced-pr done: PR https://example.test/owner/repo/pull/22 pr=https://example.test/owner/repo/pull/22 mode=no-mistakes yolo=off" \ + "$MAIN/state/mate.status" || fail "ledger fallback did not prefer the terminal ready line PR: $(cat "$MAIN/state/mate.status")" printf 'working: retrying\ndone: fixed on retry\n' >> "$MATE/state/boom.status" FM_FAKE_CREW_STATE='unknown' run_reconcile "$MATE" boom_key=$(reported_outcome_key "$MATE" boom 'done') || fail "recovered receipt key missing" @@ -256,6 +256,38 @@ test_secondmate_ledger_delivery_carries_report_and_failure() { pass "ledger delivery carries the report pointer, the failed verb, and each new terminal line" } +# A PR URL a worker only ever mentioned in prose is never claimed as the +# task's delivered PR: without a recorded PR, only a terminal line in the +# ready-signal shape carries one, and a scout never carries one at all. +test_pr_field_requires_recorded_pr_or_ready_signal_line() { + local id prose_key ready_key scout_key + make_world pr-provenance; bind_secondmate local + write_child "$MATE" prose $'working: context in https://example.test/other/repo/pull/33\ndone: cleanup finished' + write_child "$MATE" ready 'done: PR https://example.test/owner/repo/pull/44 checks green' + write_child "$MATE" lookout 'done: PR https://example.test/owner/repo/pull/55' + for id in prose ready; do + awk '$0 !~ /^pr=/' "$MATE/state/$id.meta" > "$MATE/state/$id.meta.tmp" + mv "$MATE/state/$id.meta.tmp" "$MATE/state/$id.meta" + done + awk '{ sub(/^kind=ship$/, "kind=scout"); print }' "$MATE/state/lookout.meta" \ + > "$MATE/state/lookout.meta.tmp" + mv "$MATE/state/lookout.meta.tmp" "$MATE/state/lookout.meta" + FM_FAKE_CREW_STATE='unknown' run_reconcile "$MATE" + prose_key=$(reported_outcome_key "$MATE" prose 'done') || fail "prose receipt key missing" + ready_key=$(reported_outcome_key "$MATE" ready 'done') || fail "ready receipt key missing" + scout_key=$(reported_outcome_key "$MATE" lookout 'done') || fail "scout receipt key missing" + grep -Fxq "done [key=$prose_key]: child prose done: cleanup finished mode=no-mistakes yolo=off" \ + "$MAIN/state/mate.status" \ + || fail "a PR mentioned only in prose was claimed as the delivery: $(cat "$MAIN/state/mate.status")" + grep -Fxq "done [key=$ready_key]: child ready done: PR https://example.test/owner/repo/pull/44 checks green pr=https://example.test/owner/repo/pull/44 mode=no-mistakes yolo=off" \ + "$MAIN/state/mate.status" \ + || fail "a ready-signal terminal line did not carry its PR: $(cat "$MAIN/state/mate.status")" + grep -Fxq "done [key=$scout_key]: child lookout done: PR https://example.test/owner/repo/pull/55 mode=no-mistakes yolo=off" \ + "$MAIN/state/mate.status" \ + || fail "a scout's ready-looking line carried a PR claim: $(cat "$MAIN/state/mate.status")" + pass "pr= requires the recorded PR or a ready-signal terminal line, and never a scout" +} + # If a terminal ledger line lands while the authoritative state read is in # flight, the ledger path remains the single owner on the next poll. test_terminal_line_during_state_read_yields_to_ledger_delivery() { @@ -790,6 +822,7 @@ test_main_direct_terminal_presentation_receipt test_local_secondmate_delivers_terminal_ledger_line test_busy_child_does_not_starve_later_ledger_outcomes test_secondmate_ledger_delivery_carries_report_and_failure +test_pr_field_requires_recorded_pr_or_ready_signal_line test_terminal_line_during_state_read_yields_to_ledger_delivery test_terminal_line_after_inactive_delivery_is_not_reported_twice test_progress_after_inactive_delivery_starts_a_new_event diff --git a/tests/fm-issue-writeback.test.sh b/tests/fm-issue-writeback.test.sh index 82eca83391e..416fcd80af2 100755 --- a/tests/fm-issue-writeback.test.sh +++ b/tests/fm-issue-writeback.test.sh @@ -1506,41 +1506,54 @@ test_an_unknown_milestone_is_a_usage_error() { # one command that drives two of them, and therefore the one place a second # comment would appear if the two call sites did not find each other's work. -test_the_merge_path_posts_its_own_milestones() { - local dir out rc body - dir=$(board_case mergepath) - mkdir -p "$dir/wt" "$dir/projects/widget" "$dir/data" - # The merge guard must resolve a real home before proving no delivery hold. - cp "$ROOT/.tasks.toml" "$dir/.tasks.toml" - # `gh api` is the fake GitHub; every other `gh` call fm-pr-check.sh makes - # answers as the PR-head lookup, and gh-axi records the merge. +# Model the guarded forge transaction separately from tracker API failures. +install_merge_forge() { # <case-dir> + local dir=$1 mv "$dir/fakebin/gh" "$dir/fakebin/gh-api-fake" cat > "$dir/fakebin/gh" <<'SH' #!/usr/bin/env bash -if [ "${1:-}" = api ]; then - exec "$(dirname "$0")/gh-api-fake" "$@" -fi case "${1:-} ${2:-}" in + "api graphql") + case "$*" in + *'pullRequest(number:'*) ;; + *) exec "$(dirname "$0")/gh-api-fake" "$@" ;; + esac + if [ -e "$(dirname "$0")/../merge-called" ]; then + printf '%s\n' state=MERGED merged=true queued=false base=main default=main + else + printf '%s\n' state=OPEN merged=false queued=false base=main default=main + fi + ;; "pr view") case " $* " in - *headRefOid*) printf '%s\n' deadbeefcafe ; exit 0 ;; + *statusCheckRollup*) printf '%s\n' '{"state":"OPEN","isDraft":false,"mergeable":"MERGEABLE","mergeStateStatus":"CLEAN","headRefOid":"1111111111111111111111111111111111111111","baseRefName":"main","statusCheckRollup":[{"__typename":"CheckRun","name":"ci","status":"COMPLETED","conclusion":"SUCCESS"}]}' ;; + *headRefOid*) printf '%s\n' 1111111111111111111111111111111111111111 ;; esac ;; + "pr merge") + printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" + : > "$(dirname "$0")/../merge-called" + ;; + *) exec "$(dirname "$0")/gh-api-fake" "$@" ;; esac -exit 0 SH - # Upstream's merge confirmation reads `gh-axi pr view` and records nothing - # for a merge it cannot confirm, so this stub must report the landed state - # the fixture is modelling before any bookkeeping runs. cat > "$dir/fakebin/gh-axi" <<'SH' #!/usr/bin/env bash -printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" case "${1:-} ${2:-}" in + "api "*) printf 'tip: 2222222222222222222222222222222222222222\n' ;; "pr view") printf 'pull_request:\n number: %s\n state: merged\n' "${3:-}" ;; esac -exit 0 SH chmod +x "$dir/fakebin/gh" "$dir/fakebin/gh-axi" +} + +test_the_merge_path_posts_its_own_milestones() { + local dir out rc body + dir=$(board_case mergepath) + mkdir -p "$dir/wt" "$dir/projects/widget" "$dir/data" + # The merge guard must resolve a real home before proving no delivery hold. + cp "$ROOT/.tasks.toml" "$dir/.tasks.toml" + install_merge_forge "$dir" set +e out=$(env FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$dir" FM_STATE_OVERRIDE="$dir/state" \ @@ -1568,26 +1581,7 @@ test_a_refusing_tracker_never_makes_a_completed_merge_look_retryable() { mkdir -p "$dir/wt" "$dir/projects/widget" "$dir/data" # The merge guard must resolve a real home before proving no delivery hold. cp "$ROOT/.tasks.toml" "$dir/.tasks.toml" - mv "$dir/fakebin/gh" "$dir/fakebin/gh-api-fake" - cat > "$dir/fakebin/gh" <<'SH' -#!/usr/bin/env bash -if [ "${1:-}" = api ]; then - exec "$(dirname "$0")/gh-api-fake" "$@" -fi -exit 0 -SH - # Upstream's merge confirmation reads `gh-axi pr view` and records nothing - # for a merge it cannot confirm, so this stub must report the landed state - # the fixture is modelling before any bookkeeping runs. - cat > "$dir/fakebin/gh-axi" <<'SH' -#!/usr/bin/env bash -printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" -case "${1:-} ${2:-}" in - "pr view") printf 'pull_request:\n number: %s\n state: merged\n' "${3:-}" ;; -esac -exit 0 -SH - chmod +x "$dir/fakebin/gh" "$dir/fakebin/gh-axi" + install_merge_forge "$dir" set +e out=$(env FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$dir" FM_STATE_OVERRIDE="$dir/state" \ diff --git a/tests/fm-kimi-harness.test.sh b/tests/fm-kimi-harness.test.sh index 265b9d4e604..ff679c33597 100755 --- a/tests/fm-kimi-harness.test.sh +++ b/tests/fm-kimi-harness.test.sh @@ -5,10 +5,11 @@ set -u # shellcheck source=tests/lib.sh . "$(dirname "${BASH_SOURCE[0]}")/lib.sh" -# bin/fm-harness.sh checks verified ENV markers before ancestry. A suite run -# from inside Cursor, Claude, Pi, or Grok inherits those markers, which outrank -# the fake ancestry the detection cases set up. Drop the ambient markers so the -# asserted verdict does not depend on which harness launched the suite. +# bin/fm-harness.sh answers from environment markers and process ancestry. A +# suite run from inside Cursor, Claude, Pi, or Grok inherits those markers and +# its own real ancestry, either of which can decide a case the detection cases +# meant to control. Drop the ambient markers so the asserted verdict does not +# depend on which harness launched the suite. unset CLAUDECODE PI_CODING_AGENT FM_PI_HARNESS GROK_AGENT CURSOR_AGENT CURSOR_INVOKED_AS SPAWN="$ROOT/bin/fm-spawn.sh" @@ -561,10 +562,13 @@ SH -u CURSOR_AGENT -u CURSOR_INVOKED_AS -u GEMINI_CLI \ PATH="$fakebin:$BASE_PATH" FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh") [ "$out" = kimi ] || fail "kimi ancestry detection returned '$out'" + # Kimi publishes no identity marker, so an inherited CLAUDECODE used to rename + # it outright. A structural kimi ancestor now outranks that marker; + # tests/fm-harness-precedence.test.sh owns the general boundary. out=$(env -u CURSOR_AGENT -u CURSOR_INVOKED_AS -u GEMINI_CLI \ CLAUDECODE=1 PATH="$fakebin:$BASE_PATH" FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh") - [ "$out" = claude ] || fail "verified env-marker precedence changed, got '$out'" - pass "fm-harness: markerless kimi is detected by ancestry after env-marker precedence" + [ "$out" = kimi ] || fail "an inherited CLAUDECODE renamed markerless kimi, got '$out'" + pass "fm-harness: markerless kimi keeps its ancestry identity under an inherited marker" } test_kimi_session_lock_identity() { diff --git a/tests/fm-launch-lib.test.sh b/tests/fm-launch-lib.test.sh index a2a435556dd..d2245e7c36d 100755 --- a/tests/fm-launch-lib.test.sh +++ b/tests/fm-launch-lib.test.sh @@ -62,7 +62,7 @@ test_cursor_template() { test_agy_template() { assert_eq "$(fm_launch_template agy ship)" \ - 'agy --dangerously-skip-permissions __MODELFLAG____EFFORTFLAG__--prompt-interactive "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' \ + 'env -u CLAUDECODE -u PI_CODING_AGENT -u GROK_AGENT -u FM_PI_HARNESS __AGYBIN__ --dangerously-skip-permissions __MODELFLAG____EFFORTFLAG__--prompt-interactive "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' \ "agy ship template must skip permissions, thread model+effort, and pass the brief to --prompt-interactive" assert_eq "$(fm_launch_template agy scout)" "$(fm_launch_template agy ship)" \ "agy scout template must match ship template" @@ -71,7 +71,7 @@ test_agy_template() { test_existing_templates_keep_settings_and_hooks() { assert_eq "$(fm_launch_template claude ship)" \ - 'CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --dangerously-skip-permissions --settings '\''{"feedbackDrafts":"off","attribution":{"commit":"","pr":"","sessionUrl":false}}'\'' __MODELFLAG____EFFORTFLAG__"$(__OPINPUT__ encode launch-brief < __BRIEF__)"' \ + 'CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude __CLAUDEPERMFLAG__ --settings '\''{"feedbackDrafts":"off","attribution":{"commit":"","pr":"","sessionUrl":false}}'\'' __MODELFLAG____EFFORTFLAG__"$(__OPINPUT__ encode launch-brief < __BRIEF__)"' \ "claude template drifted" assert_eq "$(fm_launch_template grok ship)" \ 'grok --always-approve __MODELFLAG____EFFORTFLAG__"$(__OPINPUT__ encode launch-brief < __BRIEF__)"' \ diff --git a/tests/fm-omp-harness.test.sh b/tests/fm-omp-harness.test.sh index 8168a98c418..0e421d25dc7 100755 --- a/tests/fm-omp-harness.test.sh +++ b/tests/fm-omp-harness.test.sh @@ -55,7 +55,7 @@ export NODE_NO_WARNINGS=1 make_named_shells() { # <dir> -> echoes <bindir> local dir=$1 name mkdir -p "$dir" - for name in omp ompd comp; do + for name in omp ompd comp claude; do ln -sf /bin/bash "$dir/$name" done printf '%s' "$dir" @@ -84,7 +84,7 @@ test_detection_anchored_name_and_marker_precedence() { # ...and is inert when it leaks into a worker with no omp ancestor. # shellcheck disable=SC2016 # the quoted body expands inside the named shell out=$(env -u PI_CODING_AGENT -u CURSOR_AGENT -u CURSOR_INVOKED_AS CLAUDECODE=1 FM_OMP_HARNESS=omp \ - bash -c '"$1"; :' _ "$HARNESS") + "$bin/claude" -c '"$1"; :' _ "$HARNESS") [ "$out" = claude ] || fail "a leaked FM_OMP_HARNESS without an omp ancestor must not relabel a claude worker, got '$out'" pass "fm-harness: omp detects by its anchored name; the marker is a precedence override that needs real omp ancestry" } @@ -97,10 +97,10 @@ test_lock_identity_and_liveness_classification() { # shellcheck source=bin/fm-backend.sh . "$ROOT/bin/fm-backend.sh" fm_backend_source tmux || fail "fm_backend_source tmux failed" - [ "$(fm_backend_tmux_classify_process_name omp)" = agent ] || fail "tmux liveness must classify omp as an agent" - [ "$(fm_backend_tmux_classify_process_name /opt/omp/bin/omp)" = agent ] || fail "tmux liveness must classify an omp path as an agent" - [ "$(fm_backend_tmux_classify_process_name ompd)" != agent ] || fail "tmux liveness must not classify ompd as an agent" - [ "$(fm_backend_tmux_classify_process_name comp)" != agent ] || fail "tmux liveness must not classify comp as an agent" + [ "$(fm_agent_process_classify_name omp)" = agent ] || fail "tmux liveness must classify omp as an agent" + [ "$(fm_agent_process_classify_name /opt/omp/bin/omp)" = agent ] || fail "tmux liveness must classify an omp path as an agent" + [ "$(fm_agent_process_classify_name ompd)" != agent ] || fail "tmux liveness must not classify ompd as an agent" + [ "$(fm_agent_process_classify_name comp)" != agent ] || fail "tmux liveness must not classify comp as an agent" pass "session lock and tmux liveness: omp is anchored, decoys stay out" } diff --git a/tests/fm-pr-check-security.test.sh b/tests/fm-pr-check-security.test.sh index aaf3ab76ac3..47fd2bb844a 100755 --- a/tests/fm-pr-check-security.test.sh +++ b/tests/fm-pr-check-security.test.sh @@ -137,19 +137,37 @@ SH printf '%s\n' "$*" >> "$FM_TEST_GH_LOG" case "${1:-} ${2:-}" in "api graphql") + if [ ! -f "$(dirname "$0")/../gh-merge-called" ]; then + printf '%s\n' state=OPEN merged=false "queued=${FM_TEST_GH_GRAPHQL_QUEUED:-false}" base=main default=main + exit 0 + fi printf '%s\n' \ - 'state=MERGED' \ - 'merged=true' \ - 'queued=false' \ + "state=${FM_TEST_GH_GRAPHQL_STATE:-MERGED}" \ + "merged=${FM_TEST_GH_GRAPHQL_MERGED:-true}" \ + "queued=${FM_TEST_GH_GRAPHQL_QUEUED:-false}" \ 'base=main' \ 'default=main' exit 0 ;; + "pr view") + case " $* " in + *statusCheckRollup*) + printf '%s\n' "{\"state\":\"OPEN\",\"isDraft\":false,\"mergeable\":\"MERGEABLE\",\"mergeStateStatus\":\"CLEAN\",\"headRefOid\":\"${FM_TEST_GH_HEAD:-0123456789abcdef0123456789abcdef01234567}\",\"baseRefName\":\"main\",\"statusCheckRollup\":[{\"__typename\":\"CheckRun\",\"name\":\"ci\",\"status\":\"COMPLETED\",\"conclusion\":\"SUCCESS\"}]}" + exit 0 + ;; + esac + ;; + "pr merge") + : > "$(dirname "$0")/../gh-merge-called" + [ -z "${FM_TEST_GH_MERGE_HOOK:-}" ] || "$FM_TEST_GH_MERGE_HOOK" + exit 0 + ;; esac case " $* " in *" headRefOid "*) printf '%s\n' "${FM_TEST_GH_HEAD:-0123456789abcdef0123456789abcdef01234567}" ;; *" state "*) [ "${FM_TEST_GH_FAIL:-0}" = 0 ] || exit 1 + [ -z "${FM_TEST_GH_STATE_STARTED:-}" ] || : > "$FM_TEST_GH_STATE_STARTED" [ "${FM_TEST_GH_SLEEP:-0}" = 0 ] || sleep "$FM_TEST_GH_SLEEP" printf '%s\n' "${FM_TEST_GH_STATE:-OPEN}" ;; @@ -174,6 +192,7 @@ case "${1:-} ${2:-}" in # not model, rather than an answer that hides the difference. case " $* " in *'{base:'*) ;; + *'{tip:'*) printf 'tip: 1111111111111111111111111111111111111111\n'; exit 0 ;; *) echo "gh-axi mock: unmodelled api query: $*" >&2 ; exit 2 ;; esac printf 'base: %s\ndef: %s\n' \ @@ -214,10 +233,14 @@ write_task_meta() { "mode=no-mistakes" } +# Extra "field=value" arguments are written before pr=, because +# fm_pr_metadata_identity_parse rejects an unrecognised line after it. write_poll_meta() { local state=$1 id=$2 url=$3 + shift 3 fm_write_meta "$state/$id.meta" \ "window=fm-$id" \ + "$@" \ "pr=$url" } @@ -536,11 +559,11 @@ test_valid_recording_and_merge_derivation() { count=$(grep -c '^pr_head=' "$dir/home/state/task-a.meta") [ "$count" -eq 1 ] || fail "duplicate pr_head metadata was appended" - : > "$dir/gh-axi.log" + : > "$dir/gh.log" run_merge_entry "$dir" task-a https://github.com/my-org/repo_name.with-dots/pull/37 -- --merge \ >/dev/null 2>/dev/null || fail "valid merge wrapper failed" - grep -qxF 'pr merge 37 --repo my-org/repo_name.with-dots --merge' "$dir/gh-axi.log" \ - || fail "merge wrapper did not preserve repository derivation and method" + grep -qxF "pr merge 37 --repo my-org/repo_name.with-dots --match-head-commit $expected --merge" "$dir/gh.log" \ + || fail "merge wrapper did not preserve repository derivation, live head, and method" # A merge this home performed leaves its own durable outcome, so the poll's # confirmation is no longer the first the captain hears of it. Acknowledge that # record before the watcher cycle below, which is what still retires the poll. @@ -637,9 +660,10 @@ SH run_watcher_bounded() { local home=$1 fakebin=$2 check_interval=${FM_TEST_CHECK_INTERVAL:-0} watch_root=${FM_TEST_WATCH_ROOT:-$ROOT} + local check_timeout=${FM_TEST_CHECK_TIMEOUT:-1} shift 2 perl -e 'my $pid=fork; die unless defined $pid; if (!$pid) { exec @ARGV } local $SIG{ALRM}=sub { kill "TERM", $pid; waitpid $pid, 0; exit 124 }; alarm 10; waitpid $pid, 0; alarm 0; exit($? >> 8)' \ - env FM_HOME="$home" FM_ROOT_OVERRIDE="$watch_root" FM_CHECK_INTERVAL="$check_interval" FM_CHECK_TIMEOUT=1 \ + env FM_HOME="$home" FM_ROOT_OVERRIDE="$watch_root" FM_CHECK_INTERVAL="$check_interval" FM_CHECK_TIMEOUT="$check_timeout" \ FM_POLL=0.02 FM_HEARTBEAT=999999 FM_SIGNAL_GRACE=0 PATH="$fakebin:$BASE_PATH" "$WATCH" "$@" } @@ -2148,12 +2172,290 @@ test_gitlab_merged_poll_retires() { pass "GitHub and GitLab exact merged results share one retirement path" } +# --- poll-path merge authority ---------------------------------------------- + +write_away_record() { # <dir> [<fm-afk-contract.sh propose args>...] + local dir=$1 + shift + FM_HOME="$dir/home" FM_STATE_OVERRIDE="$dir/home/state" \ + "$ROOT/bin/fm-afk-contract.sh" propose "$@" >/dev/null \ + || fail "could not propose an away-posture record" + FM_HOME="$dir/home" FM_STATE_OVERRIDE="$dir/home/state" \ + "$ROOT/bin/fm-afk-contract.sh" confirm >/dev/null \ + || fail "could not confirm an away-posture record" +} + +archive_away_record() { # <dir> + FM_HOME="$1/home" FM_STATE_OVERRIDE="$1/home/state" \ + "$ROOT/bin/fm-afk-contract.sh" archive >/dev/null \ + || fail "could not archive the away-posture record" +} + +# The durable queue is TSV (epoch, sequence, kind, key, payload). +merged_ledger_row() { # <state> <task-id> + awk -F'\t' -v prefix="check: merge landed: $2 " \ + 'index($5, prefix) == 1 { print $5 }' "$1/.wake-queue" +} + +run_merged_poll_cycle() { # <dir> + local dir=$1 rc=0 + add_stop_custom_check "$dir" + set +e + FM_TEST_GH_STATE=MERGED run_watcher_bounded "$dir/home" "$dir/fakebin" \ + > "$dir/watch.out" 2> "$dir/watch.err" + rc=$? + set -e + [ "$rc" -eq 0 ] || fail "merged poll watcher failed: $(cat "$dir/watch.err")" +} + +queue_merge() { # <dir> <url> + local dir=$1 url=$2 rc=0 + set +e + FM_TEST_GH_GRAPHQL_STATE=OPEN FM_TEST_GH_GRAPHQL_MERGED=false \ + FM_TEST_GH_GRAPHQL_QUEUED=true \ + run_merge_entry "$dir" task-a "$url" > "$dir/merge.out" 2> "$dir/merge.err" + rc=$? + set -e + [ "$rc" -eq 0 ] || fail "queued merge failed: $(cat "$dir/merge.err")" + assert_grep "is queued" "$dir/merge.out" "the forge did not queue the merge" + [ -f "$dir/home/state/task-a.merge-authority" ] \ + || fail "the accepted queued merge did not persist its authority" +} + +test_merged_poll_row_carries_the_merge_authority() { + local dir state url expected posture + url=https://github.com/o/r/pull/1 + + for posture in yolo grant; do + dir=$(make_case "queued-merge-authority-$posture") + state="$dir/home/state" + write_task_meta "$dir" task-a + if [ "$posture" = yolo ]; then + printf 'yolo=on\n' >> "$state/task-a.meta" + write_away_record "$dir" + expected=yolo + else + write_away_record "$dir" --grant task-a + expected=away-grant + fi + run_check_entry "$dir" task-a "$url" >/dev/null 2> "$dir/seed.err" \ + || fail "$posture: could not arm the merge poll" + queue_merge "$dir" "$url" + archive_away_record "$dir" + run_merged_poll_cycle "$dir" + [ "$(merged_ledger_row "$state" task-a)" = "check: merge landed: task-a $url $expected" ] \ + || fail "$posture: archived posture lost persisted authority: $(merged_ledger_row "$state" task-a)" + [ ! -e "$state/task-a.merge-authority" ] \ + || fail "$posture: published merge left its authority record behind" + done + + pass "queued merges retain yolo and away-grant after captain return" +} + +test_merged_poll_row_names_no_authority_when_no_record_grants_one() { + local dir state url + url=https://github.com/o/r/pull/1 + + dir=$(make_case queued-merge-authority-attended) + state="$dir/home/state" + write_task_meta "$dir" task-a + run_check_entry "$dir" task-a "$url" >/dev/null 2> "$dir/seed.err" \ + || fail "attended: could not arm the merge poll" + queue_merge "$dir" "$url" + run_merged_poll_cycle "$dir" + [ "$(merged_ledger_row "$state" task-a)" = "check: merge landed: task-a $url" ] \ + || fail "attended queued merge was tagged: $(merged_ledger_row "$state" task-a)" + + dir=$(make_case merged-poll-authority-external) + state="$dir/home/state" + write_poll_meta "$state" task-a "$url" yolo=on + write_away_record "$dir" + seed_canonical_poll "$dir" task-a "$url" + run_merged_poll_cycle "$dir" + [ "$(merged_ledger_row "$state" task-a)" = "check: merge landed: task-a $url external" ] \ + || fail "external merge was attributed from live away posture: $(merged_ledger_row "$state" task-a)" + assert_poll_absent "$state" task-a + + pass "poll distinguishes attended authorization from external landing" +} + +test_authority_persistence_refuses_rebound_metadata() { + local dir state url_a url_b rc + url_a=https://github.com/o/r/pull/1 + url_b=https://github.com/o/r/pull/2 + dir=$(make_case merge-authority-rebound-metadata) + state="$dir/home/state" + write_task_meta "$dir" task-a + run_check_entry "$dir" task-a "$url_a" >/dev/null 2> "$dir/seed.err" \ + || fail "rebind: could not arm the original poll" + cat > "$dir/rebind.sh" <<SH +#!/usr/bin/env bash +"$PR_CHECK" task-a "$url_b" >/dev/null +SH + chmod +x "$dir/rebind.sh" + set +e + FM_TEST_GH_MERGE_HOOK="$dir/rebind.sh" \ + FM_TEST_GH_GRAPHQL_STATE=OPEN FM_TEST_GH_GRAPHQL_MERGED=false \ + FM_TEST_GH_GRAPHQL_QUEUED=true \ + run_merge_entry "$dir" task-a "$url_a" > "$dir/merge.out" 2> "$dir/merge.err" + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "rebind: accepted merge persisted against rebound metadata" + grep -qxF "pr=$url_b" "$state/task-a.meta" \ + || fail "rebind: merge hook did not replace the canonical identity" + [ ! -e "$state/task-a.merge-authority" ] \ + || fail "rebind: authority was published for the wrong canonical identity" + pass "accepted merge authority refuses rebound task metadata" +} + +test_authority_persists_before_control_unlock() { + local dir state url + url=https://github.com/o/r/pull/1 + dir=$(make_case merge-authority-control-lock) + state="$dir/home/state" + write_task_meta "$dir" task-a + run_check_entry "$dir" task-a "$url" >/dev/null 2> "$dir/seed.err" \ + || fail "control lock: could not arm the merge poll" + cat > "$dir/fakebin/mv" <<'SH' +#!/usr/bin/env bash +case " $* " in + *"task-a.merge-authority "*) + [ -d "$FM_TEST_CONTROL_LOCK" ] || exit 91 + ;; +esac +exec "$FM_TEST_REAL_MV" "$@" +SH + chmod +x "$dir/fakebin/mv" + FM_TEST_CONTROL_LOCK="$state/.control-task-a.lock" FM_TEST_REAL_MV="$REAL_MV" \ + queue_merge "$dir" "$url" + pass "accepted merge authority persists under the lifecycle lock" +} + +test_teardown_cannot_race_authority_consumption() { + local dir state url watcher_pid rc i + url=https://github.com/o/r/pull/1 + dir=$(make_case merge-authority-teardown-race) + state="$dir/home/state" + fm_write_meta "$state/task-a.meta" \ + 'window=firstmate:fm-task-a' \ + 'endpoint_task_id=task-a' \ + "worktree=$dir/wt" \ + "project=$dir/project" \ + 'kind=ship' \ + 'mode=local-only' \ + 'yolo=on' + write_away_record "$dir" + run_check_entry "$dir" task-a "$url" >/dev/null 2> "$dir/seed.err" \ + || fail "teardown race: could not arm the merge poll" + queue_merge "$dir" "$url" + archive_away_record "$dir" + FM_TEST_GH_STATE_STARTED="$dir/poll-started" FM_TEST_GH_STATE=MERGED \ + FM_TEST_GH_SLEEP=0.5 FM_TEST_CHECK_TIMEOUT=3 \ + run_watcher_bounded "$dir/home" "$dir/fakebin" \ + > "$dir/watch.out" 2> "$dir/watch.err" & + watcher_pid=$! + i=0 + while [ ! -e "$dir/poll-started" ]; do + sleep 0.01 + i=$((i + 1)) + if [ "$i" -ge 500 ]; then + kill "$watcher_pid" 2>/dev/null || true + wait "$watcher_pid" 2>/dev/null || true + fail "teardown race: watcher did not begin its validated poll" + fi + done + set +e + FM_HOME="$dir/home" FM_ROOT_OVERRIDE="$ROOT" PATH="$dir/fakebin:$BASE_PATH" \ + "$TEARDOWN" task-a --force > "$dir/teardown.out" 2> "$dir/teardown.err" + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "teardown race: cleanup crossed the active poll transaction" + [ -f "$state/task-a.merge-authority" ] \ + || fail "teardown race: refused cleanup removed persisted authority" + rc=0 + wait "$watcher_pid" || rc=$? + [ "$rc" -eq 0 ] || fail "teardown race: watcher failed with $rc: $(cat "$dir/watch.err")" + [ "$(merged_ledger_row "$state" task-a)" = "check: merge landed: task-a $url yolo" ] \ + || fail "teardown race: concurrent cleanup downgraded the merge authority" + pass "teardown cannot race merged-poll authority consumption" +} + +test_authority_retirement_preserves_replacement() { + local dir state url_a url_b rc i + url_a=https://github.com/o/r/pull/1 + url_b=https://github.com/o/r/pull/2 + dir=$(make_case merge-authority-retirement-replacement) + state="$dir/home/state" + write_task_meta "$dir" task-a + run_check_entry "$dir" task-a "$url_a" >/dev/null 2> "$dir/seed.err" \ + || fail "replacement: could not arm the original poll" + queue_merge "$dir" "$url_a" + cat > "$dir/replace-authority.sh" <<SH +#!/usr/bin/env bash +"$PR_CHECK" task-a "$url_b" >/dev/null +( + FM_TEST_GH_GRAPHQL_STATE=OPEN FM_TEST_GH_GRAPHQL_MERGED=false \\ + FM_TEST_GH_GRAPHQL_QUEUED=true \\ + "$PR_MERGE" task-a "$url_b" > "$dir/replacement-merge.out" 2> "$dir/replacement-merge.err" + printf '%s\n' \$? > "$dir/replacement-merge.rc" +) & +SH + chmod +x "$dir/replace-authority.sh" + cat > "$dir/fakebin/mv" <<'SH' +#!/usr/bin/env bash +"$FM_TEST_REAL_MV" "$@" || exit $? +case " $* " in + *"task-a.pr-poll-merge-notified "*) + if [ ! -e "$FM_TEST_REPLACEMENT_RAN" ]; then + : > "$FM_TEST_REPLACEMENT_RAN" + "$FM_TEST_REPLACEMENT_SCRIPT" + fi + ;; +esac +SH + chmod +x "$dir/fakebin/mv" + add_stop_custom_check "$dir" + set +e + FM_TEST_REAL_MV="$REAL_MV" FM_TEST_REPLACEMENT_RAN="$dir/replacement-ran" \ + FM_TEST_REPLACEMENT_SCRIPT="$dir/replace-authority.sh" \ + FM_TEST_GH_STATE=MERGED run_watcher_bounded "$dir/home" "$dir/fakebin" \ + > "$dir/watch-a.out" 2> "$dir/watch-a.err" + rc=$? + set -e + [ "$rc" -eq 0 ] || fail "replacement: original poll failed: $(cat "$dir/watch-a.err")" + i=0 + while [ ! -e "$dir/replacement-merge.rc" ]; do + sleep 0.01 + i=$((i + 1)) + [ "$i" -lt 200 ] || fail "replacement: serialized replacement merge did not finish" + done + [ "$(cat "$dir/replacement-merge.rc")" -eq 0 ] \ + || fail "replacement: serialized replacement merge failed: $(cat "$dir/replacement-merge.err")" + [ -f "$state/task-a.merge-authority" ] \ + || fail "replacement: original poll retirement deleted the replacement authority" + grep -qxF "pr=$url_b" "$state/task-a.meta" \ + || fail "replacement: replacement poll was not armed" + ack_watcher_cycle "$state" || fail "replacement: could not acknowledge the original wake" + rm -f "$dir/fakebin/mv" "$state/.last-check" + run_merged_poll_cycle "$dir" + awk -F'\t' -v expected="check: merge landed: task-a $url_b" \ + '$5 == expected { found=1 } END { exit !found }' "$state/.wake-queue" \ + || fail "replacement: replacement merge lost its attended authority" + pass "poll retirement preserves a replacement authority record" +} + test_parser_matrix test_gitlab_merge_watch test_merged_poll_retires_once test_merged_poll_reregistration_after_notification_is_absorbed test_merged_poll_retries_a_failed_upward_report test_self_merge_and_poll_publish_one_outcome +test_merged_poll_row_carries_the_merge_authority +test_merged_poll_row_names_no_authority_when_no_record_grants_one +test_authority_persistence_refuses_rebound_metadata +test_authority_persists_before_control_unlock +test_teardown_cannot_race_authority_consumption +test_authority_retirement_preserves_replacement test_merged_poll_reports_upward_from_a_secondmate_home_once test_different_merged_pr_for_same_task_is_not_absorbed test_persistent_secondmate_retirement_is_poll_only diff --git a/tests/fm-pr-merge.test.sh b/tests/fm-pr-merge.test.sh index da2c5e789ca..d80acf57427 100755 --- a/tests/fm-pr-merge.test.sh +++ b/tests/fm-pr-merge.test.sh @@ -5,116 +5,8 @@ # repos with no PR CI where the usual "checks green" fm-pr-check.sh trigger # never fires. # -# Matrix: -# (a) a verified merge records pr= and pr_head= -# (b) merge is refused when gh-axi pr merge itself fails (no silent success) -# (c) extra gh-axi pr merge args are forwarded after number and --repo -# (d) merge is refused before gh-axi when task meta is missing -# (e) PR URL is parsed to number + --repo for gh-axi (defaults to --squash) -# (f) malformed PR URL fails fast without calling gh-axi -# (g) explicit merge method is not overridden by the default --squash -# (h) repo override args fail fast because the repo comes from the URL, -# including a bundled short-option cluster that carries -R -# (i) a GitLab MR URL resolves and merges through glab instead of erroring -# (j) glab is addressed by the host from the URL, never an assumed one -# (k) no merge method is imposed on GitLab, so the project's own one applies -# (l) each pre-merge condition refuses independently, and all of them report -# (m) a stale recorded pr_head= is reported and the live head is verified -# (n) an unreadable merge request state refuses rather than merging blind -# (o) glab or jq absent refuses before any state is recorded -# (p) --sha in extra GitLab args fails fast, and still forwards on GitHub -# (q) a GitLab refusal still leaves pr= recorded and the merge poll armed -# (r) GitHub success is accepted only after the PR is read back as merged -# (s) an open GitHub PR that is neither merged nor queued fails verification -# (t) a GitHub PR in the merge queue is reported as queued, not merged -# (u) a queue-required refusal names the exact compatible retry flags -# (v) a failed poll setup cannot be reported as a verified GitHub merge -# (w) a zero-exit queue-required refusal keeps merge semantics unchanged -# (x) an unreadable outcome after a successful merge call keeps the PR -# recorded and the merge poll armed -# (y) agreeing queue rules still produce exact retry flags -# (z) conflicting queue rules report ambiguous retry guidance -# (aa) gh-axi remains usable when gh is absent -# (ab) a landed merge whose fallback outcome read fails keeps its poll armed -# (ac) a successful merge in a secondmate home reports the landed PR upward -# once, on the route its parent binding names, and a repeat merge of the -# same PR does not duplicate that line -# (ad) a refused or failed merge reports nothing -# (ae) a successful merge in a main home leaves a durable wake naming the PR -# (af) a secondmate home with no usable parent binding says so loudly instead -# of merging in silence -# (ag) an accepted queued GitHub merge emits nothing and leaves its poll armed -# (ah) an accepted queued GitLab merge emits nothing and leaves its poll armed -# (ai) an uncommitted marker retry never loses the durable outcome -# (aj) distinct merged PRs for a reused task each survive queue deduplication -# (ak) pr= is already recorded when the forge call that can land the merge runs -# (al) a failed gh read falls back to the gh-axi view, which can prove a merge -# (am) a failed merge command still names an outcome read that proves a landed -# or queued pull request, without masking the forge failure -# (an) a refusal after a zero-exit merge quotes the forge's own output, marked -# apart from the wrapper's verdict and never leaked to stdout -# (ao) an outcome read that fails after a zero-exit merge still quotes the -# forge's own output, the only evidence left -# (ap) an unrecognised queue method still names the queue requirement and -# guesses no method -# (aq) unreadable branch rules are reported apart from a queue-less base -# (ar) a base branch with no queue rule says nothing about a merge queue -# (as) identical degraded evidence reaches an identical verdict whether gh is -# absent or present but broken, so the outcome never depends on which of -# the two readers happened to be missing -# (at) an open recorded issue is closed after merge and linked to the PR -# (au) an already-closed recorded issue is left alone -# (av) issue-close failure reports the merge as successful and exits zero -# (aw) a task with no recorded issue makes no issue API calls -# (ax) issue-state verification failure reports the merge as successful -# (ay) a successful close request that leaves the issue open warns -# (az) malformed or duplicate recorded issue metadata warns without API calls -# (ba) a gitea work item closes through its own host credential with the same -# linking comment, an absent credential and an empty one are reported as -# the two different facts they are with nothing sent, a verification that -# fails says why it failed rather than only that it did, and a gitea close -# failure never makes the merge look retryable -# (bb) a cached-PR-state refresh that fails hands the operator the cause the -# refresh named, bounded to one line, instead of only the symptom -# (bc) every --auto spelling is refused by name before any merge is attempted -# (bd) every path that observes a landed merge exits zero and records the -# outcome durably, over the seven routes that reach such an observation, -# each proving it reached its OWN observation point and leaving a witness -# the others cannot match, so two routes that collapse onto one read fail -# instead of both passing; and on both forges a merge whose command failed -# says so against the readback that overrides it rather than exiting zero -# in silence -# (be) a pull request whose target is not the current default branch is -# refused by name, before any queue or method handling, including one -# already merged into that branch - the contract is a precondition, so it -# is evaluated on a merged pull request too and refuses rather than -# recording a landing onto a branch guarded merging may not touch -# (bf) a target that could not be established refuses rather than permitting -# (bg) the PRE-merge half of the degraded-view seam: the degraded gh-axi -# reader reads the real target out of the api passthrough's own envelope, -# including when a broken gh is installed, and is accepted for the target -# question with no merged proof -# (bh) a landed merge this run already observed is never re-read, including one -# the PRE-merge target read is what observed -# (bi) the GitLab automatic-rebase guard refuses only inside the window where -# project-level rebase can fire, and permits outside it -# (bj) the POST-merge half of that same seam, on BOTH degraded routes: the -# degraded gh-axi view cannot answer the outcome question without a proved -# merge, and reports an outcome it could not read rather than a concrete -# not-merged verdict -# (bk) no argument position lets a refused flag reach the forge - every -# value-taking allow-list entry crossed with --auto in both orders - while -# the detached values the allow-list exists to carry still merge -# (bl) the degraded reader decodes a TOON-quoted branch name, so a ref name -# containing a comma or beginning with a dash is compared as itself rather -# than as its quotes, in both the permitting and refusing directions -# (bm) a refused target arms nothing: neither pr= nor the merge poll, which -# reads only whether the pull request merged and would otherwise record -# the non-default landing the refusal exists to keep off this ledger -# (bn) the default branch's tip is re-read immediately before the merge call, -# a tip that moved refuses naming both commits, an unmoved one still -# merges, and a tip that could not be re-read refuses rather than passing -# as unmoved +# The test_* functions below name the covered merge, refusal, live-head, +# away-authority, outcome-publication, and recovery behavior directly. set -u # shellcheck source=tests/lib.sh @@ -146,8 +38,8 @@ MOVED_DEFAULT_TIP=3434343434343434343434343434343434343434 JQ_BIN=$(command -v jq) || fail "these tests read glab's JSON with the real jq, which was not found" REAL_MV=$(command -v mv) || fail "these tests need mv to simulate a failed poll publish" -# Build a fresh sandbox for one test case: a state dir with a task meta and a -# fakebin with a gh-axi mock that records how it was invoked. Echoes the case dir. +# Build a fresh sandbox for one test case: a state dir with task metadata and a +# directory for its forge-command mocks. Echoes the case directory. make_case() { local name=$1 case_dir fakebin case_dir="$TMP_ROOT/$name" @@ -176,15 +68,77 @@ make_case() { printf '%s\n' "$case_dir" } -# gh-axi mock recording every invocation to a log file, and gh mock answering -# headRefOid for fm-pr-check.sh's pr_head lookup. Args: case_dir head_sha +# Live GitHub JSON for the pre-merge verify, plus gh-axi for the +# post-merge fallback view. Merge itself is `gh pr merge --match-head-commit`. +# Args: case_dir head_sha +write_github_live_json() { + local case_dir=$1 head=$2 + printf '%s\n' "$head" > "$case_dir/github-head" + cat > "$case_dir/github-view.json" <<JSON +{"state":"OPEN","isDraft":false,"mergeable":"MERGEABLE","mergeStateStatus":"CLEAN","headRefOid":"$head","baseRefName":"main","statusCheckRollup":[{"__typename":"CheckRun","name":"ci","status":"COMPLETED","conclusion":"SUCCESS"}]} +JSON +} + +write_github_red_json() { + local case_dir=$1 head=$2 name=$3 + printf '%s\n' "$head" > "$case_dir/github-head" + cat > "$case_dir/github-view.json" <<JSON +{"state":"OPEN","isDraft":false,"mergeable":"MERGEABLE","mergeStateStatus":"CLEAN","headRefOid":"$head","baseRefName":"main","statusCheckRollup":[{"__typename":"CheckRun","name":"$name","status":"COMPLETED","conclusion":"FAILURE"}]} +JSON +} + +# One CheckRun rollup entry the way GitHub reports it. A conclusion or timestamp +# of "-" is emitted as JSON null. Args: name status conclusion [startedAt] +# [completedAt] +check_run() { + local name=$1 status=$2 conclusion=$3 started=${4:--} completed=${5:-${4:--}} + local conclusion_json='null' started_json='null' completed_json='null' + [ "$conclusion" = - ] || conclusion_json="\"$conclusion\"" + [ "$started" = - ] || started_json="\"$started\"" + [ "$completed" = - ] || completed_json="\"$completed\"" + printf '{"__typename":"CheckRun","name":"%s","status":"%s","conclusion":%s,"startedAt":%s,"completedAt":%s}' \ + "$name" "$status" "$conclusion_json" "$started_json" "$completed_json" +} + +status_context() { + local name=$1 state=$2 + printf '{"__typename":"StatusContext","context":"%s","state":"%s"}' "$name" "$state" +} + +# Live GitHub JSON whose rollup holds the given entries verbatim, so a test can +# put several runs of one check name at the same head the way GitHub does after +# it cancels a pull request's in-flight run and re-triggers it. mergeStateStatus +# stays CLEAN because that is what GitHub reports for exactly this case. +# Args: case_dir head_sha <rollup-entry-json>... +write_github_rollup_json() { + local case_dir=$1 head=$2 entry rollup='' + shift 2 + for entry in "$@"; do + rollup="${rollup:+$rollup,}$entry" + done + printf '%s\n' "$head" > "$case_dir/github-head" + cat > "$case_dir/github-view.json" <<JSON +{"state":"OPEN","isDraft":false,"mergeable":"MERGEABLE","mergeStateStatus":"CLEAN","headRefOid":"$head","baseRefName":"main","statusCheckRollup":[$rollup]} +JSON +} + +assert_logged_gh_merge() { + local case_dir=$1 number=$2 repo=$3 head line extra= + shift 3 + head=$(cat "$case_dir/github-head") + [ "$#" -eq 0 ] || extra=" $*" + line="pr merge $number --repo $repo --match-head-commit $head$extra" + grep -qxF "$line" "$case_dir/gh.log" \ + || fail "expected gh merge line: $line"$'\n'"got: $(grep '^pr merge ' "$case_dir/gh.log" || true)" +} + add_gh_mocks() { local case_dir=$1 head=$2 + write_github_live_json "$case_dir" "$head" cat > "$case_dir/fakebin/gh-axi" <<'SH' #!/usr/bin/env bash printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" case "${1:-} ${2:-}" in - "pr merge") printf 'merged:\n number: %s\n status: ok\n' "${3:-}" ;; "pr view") [ "$#" -eq 5 ] && [ "${4:-}" = --repo ] || exit 2 printf 'pull_request:\n number: %s\n state: %s\n' "$3" "${FM_TEST_GH_MERGE_STATE:-merged}" @@ -211,56 +165,82 @@ case "${1:-} ${2:-}" in esac exit 0 SH - cat > "$case_dir/fakebin/gh" <<SH -#!/usr/bin/env bash -printf '%s\n' "\$*" >> "\$FM_TEST_GH_LOG" -case "\${1:-} \${2:-}" in - "pr view") - case " \$* " in - *headRefOid*) printf '%s\n' '$head' ; exit 0 ;; - esac - ;; - "api graphql") - cat "\$FM_TEST_GH_OUTCOME" - exit 0 - ;; - api\ *) - cat "\$FM_TEST_GH_RULES" - exit 0 - ;; -esac -exit 0 -SH - chmod +x "$case_dir/fakebin/gh-axi" "$case_dir/fakebin/gh" -} - -# gh-axi mock that fails the merge call but succeeds everything else, so a -# real merge failure is distinguishable from the recording step. -add_gh_mocks_merge_fails() { - local case_dir=$1 - cat > "$case_dir/fakebin/gh-axi" <<'SH' + cat > "$case_dir/fakebin/gh" <<'SH' #!/usr/bin/env bash -printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" +printf '%s\n' "$*" >> "$FM_TEST_GH_LOG" case "${1:-} ${2:-}" in - "pr merge") echo "error: pr merge failed" >&2 ; exit 1 ;; - api\ *) + "pr view") + if [ -e "$FM_TEST_GH_OUTCOME.cache-invalid" ]; then + case " $* " in *'state,isDraft,mergeable,mergeStateStatus,headRefOid,baseRefName,statusCheckRollup'*) ;; *statusCheckRollup*) printf 'not JSON\n'; exit 0 ;; esac + fi case " $* " in - *'{tip:'*) printf 'tip: %s\n' "${FM_TEST_GH_AXI_TIP:-$FM_TEST_DEFAULT_TIP}" ;; - *) echo "gh-axi mock: unmodelled api query: $*" >&2 ; exit 2 ;; + *statusCheckRollup*) + cat "$FM_TEST_GH_VIEW_JSON" + if [ -f "${FM_TEST_AWAY_RECORD_AFTER_VIEW:-}" ]; then + cp "$FM_TEST_AWAY_RECORD_AFTER_VIEW" "$FM_STATE_OVERRIDE/.afk-contract" + fi + exit 0 + ;; + *headRefOid*) + cat "$FM_TEST_GH_HEAD" + exit 0 + ;; esac ;; - esac - exit 0 -SH - cat > "$case_dir/fakebin/gh" <<'SH' -#!/usr/bin/env bash -printf '%s\n' "$*" >> "$FM_TEST_GH_LOG" -case "${1:-} ${2:-}" in + "pr merge") + : > "$FM_TEST_GH_OUTCOME.merge-called" + if [ -f "$FM_TEST_GH_OUTCOME.merged" ]; then + cp "$FM_TEST_GH_OUTCOME.merged" "$FM_TEST_GH_OUTCOME" + fi + if [ -n "${FM_TEST_META_AT_MERGE:-}" ] && [ -f "${FM_STATE_OVERRIDE:-}/task-x1.meta" ]; then + cat "$FM_STATE_OVERRIDE/task-x1.meta" > "$FM_TEST_META_AT_MERGE" + fi + # The forge call runs inside the merge's critical section, so a real + # away-record change attempted from here is the TOCTOU itself: whatever + # happens to it happens between the authority read and the merge. + if [ -x "${FM_TEST_AWAY_MUTATE_AT_MERGE:-}" ]; then + away_rc=0 + "$FM_TEST_AWAY_MUTATE_AT_MERGE" > "$FM_TEST_AWAY_MUTATE_OUT" 2>&1 || away_rc=$? + printf '%s\n' "$away_rc" > "$FM_TEST_AWAY_MUTATE_RC" + "$FM_TEST_ROOT/bin/fm-afk-contract.sh" grants \ + > "$FM_TEST_AWAY_GRANTS_AT_MERGE" 2>/dev/null \ + || printf 'no-live-record\n' > "$FM_TEST_AWAY_GRANTS_AT_MERGE" + fi + if [ -n "${FM_TEST_GH_MERGE_OUTPUT:-}" ]; then + printf '%s\n' "$FM_TEST_GH_MERGE_OUTPUT" + else + printf 'merged:\n number: %s\n status: ok\n' "${3:-}" + fi + merge_rc=0 + if [ -f "${FM_TEST_GH_MERGE_RC_FILE:-}" ]; then + merge_rc=$(cat "$FM_TEST_GH_MERGE_RC_FILE") + fi + exit "$merge_rc" + ;; "api graphql") - cat "$FM_TEST_GH_OUTCOME" + count_file="$FM_TEST_GH_OUTCOME.reads" + count=$(cat "$count_file" 2>/dev/null || echo 0) + count=$((count + 1)) + printf '%s\n' "$count" > "$count_file" + if [ -f "$FM_TEST_GH_OUTCOME.fail-from" ] && [ "$count" -ge "$(cat "$FM_TEST_GH_OUTCOME.fail-from")" ]; then + echo 'error: could not reach the GitHub API' >&2 + exit 1 + fi + if [ -f "${FM_TEST_GH_GRAPHQL_FAIL:-}" ]; then + echo 'error: could not reach the GitHub API' >&2 + exit 1 + fi + if [ ! -e "$FM_TEST_GH_OUTCOME.merge-called" ] && [ ! -e "$FM_TEST_GH_OUTCOME.initially-merged" ]; then + sed -e 's/^state=.*/state=OPEN/' -e 's/^merged=.*/merged=false/' "$FM_TEST_GH_OUTCOME" + else + cat "$FM_TEST_GH_OUTCOME" + fi exit 0 ;; api\ *) + if [ -f "${FM_TEST_GH_RULES_FAIL:-}" ]; then + exit 1 + fi cat "$FM_TEST_GH_RULES" exit 0 ;; @@ -282,52 +262,27 @@ SH # Args: case_dir head first_failing_read add_gh_mock_outcome_read_fails_from() { local case_dir=$1 head=$2 from=$3 - cat > "$case_dir/fakebin/gh" <<SH -#!/usr/bin/env bash -printf '%s\n' "\$*" >> "\$FM_TEST_GH_LOG" -case "\${1:-} \${2:-}" in - "pr view") - case " \$* " in - *headRefOid*) printf '%s\n' '$head' ; exit 0 ;; - esac - ;; - "api graphql") - count_file="\$FM_TEST_GH_OUTCOME.reads" - count=\$(cat "\$count_file" 2>/dev/null || echo 0) - count=\$((count + 1)) - printf '%s\n' "\$count" > "\$count_file" - if [ "\$count" -ge $from ]; then - echo 'error: could not reach the GitHub API' >&2 - exit 1 - fi - cat "\$FM_TEST_GH_OUTCOME" - exit 0 - ;; -esac -exit 0 -SH - chmod +x "$case_dir/fakebin/gh" + cp "$case_dir/fakebin/gh-axi" "$case_dir/gh-axi.saved" + add_gh_mocks "$case_dir" "$head" + mv "$case_dir/gh-axi.saved" "$case_dir/fakebin/gh-axi" + printf '%s\n' "$from" > "$case_dir/github-outcome.fail-from" } +# gh mock that fails the merge call but succeeds live verify, so a real merge +# failure is distinguishable from the recording step. +add_gh_mocks_merge_fails() { + local case_dir=$1 + local head=${2:-bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb} + add_gh_mocks "$case_dir" "$head" + printf '1\n' > "$case_dir/github-merge-rc" + printf 'error: pr merge failed\n' > "$case_dir/github-merge-output" +} + +# Flag the shared gh mock so GraphQL outcome reads fail while live verify and +# merge still succeed. Args: case_dir [head_sha ignored] add_gh_mock_outcome_read_fails() { - local case_dir=$1 head=$2 - cat > "$case_dir/fakebin/gh" <<SH -#!/usr/bin/env bash -printf '%s\n' "\$*" >> "\$FM_TEST_GH_LOG" -case "\${1:-} \${2:-}" in - "pr view") - case " \$* " in - *headRefOid*) printf '%s\n' '$head' ; exit 0 ;; - esac - ;; - "api graphql") - echo 'error: could not reach the GitHub API' >&2 - exit 1 - ;; -esac -exit 0 -SH - chmod +x "$case_dir/fakebin/gh" + local case_dir=$1 + : > "$case_dir/github-graphql-fail" } # gh-axi mock that merges but cannot answer its own view, so a case can prove @@ -365,6 +320,7 @@ add_gh_axi_mock_open_until_merged() { printf '%s\n' \ 'state=MERGED' 'merged=true' 'queued=false' 'base=main' 'default=main' \ > "$case_dir/github-outcome.merged" + printf '%s\n' "$merge_rc" > "$case_dir/github-merge-rc" cat > "$case_dir/fakebin/gh-axi" <<SH #!/usr/bin/env bash printf '%s\n' "\$*" >> "\$FM_TEST_GH_AXI_LOG" @@ -594,6 +550,7 @@ write_mr_json() { local file=$1 kv key value local state=opened detail=mergeable conflicts=false discussions=true local head=$MR_HEAD pipeline_sha=$MR_HEAD pipeline_status=success pipeline=present + local merge_when_pipeline_succeeds=false merge_after=null shift for kv in "$@"; do key=${kv%%=*} @@ -607,6 +564,8 @@ write_mr_json() { pipeline_sha) pipeline_sha=$value ;; pipeline_status) pipeline_status=$value ;; pipeline) pipeline=$value ;; + merge_when_pipeline_succeeds) merge_when_pipeline_succeeds=$value ;; + merge_after) merge_after=$value ;; *) fail "write_mr_json: unknown field '$key'" ;; esac done @@ -615,8 +574,10 @@ write_mr_json() { fi printf '{"iid":7,"state":"%s","detailed_merge_status":"%s","has_conflicts":%s,' \ "$state" "$detail" "$conflicts" > "$file" - printf '"blocking_discussions_resolved":%s,"sha":"%s","head_pipeline":%s}\n' \ + printf '"blocking_discussions_resolved":%s,"sha":"%s","head_pipeline":%s,' \ "$discussions" "$head" "$pipeline" >> "$file" + printf '"merge_when_pipeline_succeeds":%s,"merge_after":%s}\n' \ + "$merge_when_pipeline_succeeds" "$merge_after" >> "$file" } # make_gitlab_case <name> [<field>=<value> ...]: a case dir with both forge @@ -677,7 +638,19 @@ run_pr_merge() { FM_TEST_GH_OUTCOME="$case_dir/github-outcome" \ FM_TEST_GH_RULES="$case_dir/github-rules" \ FM_TEST_DEFAULT_TIP="$DEFAULT_TIP" \ + FM_TEST_GH_VIEW_JSON="$case_dir/github-view.json" \ + FM_TEST_GH_HEAD="$case_dir/github-head" \ + FM_TEST_GH_MERGE_RC_FILE="$case_dir/github-merge-rc" \ + FM_TEST_GH_MERGE_OUTPUT="$(cat "$case_dir/github-merge-output" 2>/dev/null || true)" \ + FM_TEST_GH_GRAPHQL_FAIL="$case_dir/github-graphql-fail" \ + FM_TEST_GH_RULES_FAIL="$case_dir/github-rules-fail" \ FM_TEST_META_AT_MERGE="$case_dir/meta-at-merge" \ + FM_TEST_AWAY_RECORD_AFTER_VIEW="$case_dir/away-record-after-view" \ + FM_TEST_ROOT="$ROOT" \ + FM_TEST_AWAY_MUTATE_AT_MERGE="${FM_TEST_AWAY_MUTATE_AT_MERGE:-}" \ + FM_TEST_AWAY_MUTATE_OUT="$case_dir/away-mutate-output" \ + FM_TEST_AWAY_MUTATE_RC="$case_dir/away-mutate-rc" \ + FM_TEST_AWAY_GRANTS_AT_MERGE="$case_dir/away-grants-at-merge" \ FM_TEST_REAL_MV="$REAL_MV" \ FM_TEST_GLAB_LOG="$case_dir/glab.log" \ FM_TEST_GLAB_JSON="$case_dir/mr.json" \ @@ -705,6 +678,15 @@ write_github_outcome() { # <case-dir> <state> <merged> <queued> <base> [<defaul "default=$default" > "$case_dir/github-outcome" } +write_away_record() { + local case_dir=$1 + shift + FM_HOME="$case_dir/home" FM_STATE_OVERRIDE="$case_dir/state" \ + "$ROOT/bin/fm-afk-contract.sh" propose "$@" >/dev/null + FM_HOME="$case_dir/home" FM_STATE_OVERRIDE="$case_dir/state" \ + "$ROOT/bin/fm-afk-contract.sh" confirm >/dev/null +} + test_verified_merge_records_pr_and_head() { local case_dir rc case_dir=$(make_case records-before-merge) @@ -723,8 +705,7 @@ test_verified_merge_records_pr_and_head() { "records-before-merge: pr= was not recorded" assert_grep 'pr_head=deadbeefcafefeed0000000000000000deadbeef' "$case_dir/state/task-x1.meta" \ "records-before-merge: pr_head= was not recorded" - grep -qxF 'pr merge 9 --repo example/repo --squash' "$case_dir/gh-axi.log" \ - || fail "records-before-merge: gh-axi pr merge was not invoked with number, --repo, and default --squash" + assert_logged_gh_merge "$case_dir" 9 example/repo --squash pass "fm-pr-merge records pr= and pr_head= for a verified GitHub merge" } @@ -736,21 +717,6 @@ test_pr_metadata_is_recorded_before_the_forge_call() { case_dir=$(make_case records-ahead-of-forge-call) mkdir -p "$case_dir/wt" add_gh_mocks "$case_dir" 5151515151515151515151515151515151515151 - cat > "$case_dir/fakebin/gh-axi" <<'SH' -#!/usr/bin/env bash -printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" -case "${1:-} ${2:-}" in - "pr merge") - cat "$FM_STATE_OVERRIDE/task-x1.meta" > "$FM_TEST_META_AT_MERGE" - printf 'merged:\n number: %s\n status: ok\n' "${3:-}" - ;; - "pr view") - printf 'pull_request:\n number: %s\n state: merged\n' "$3" - ;; -esac -exit 0 -SH - chmod +x "$case_dir/fakebin/gh-axi" : > "$case_dir/gh-axi.log" : > "$case_dir/meta-at-merge" @@ -761,8 +727,7 @@ SH set -e expect_code 0 "$rc" "records-ahead-of-forge-call: fm-pr-merge should succeed" - assert_grep 'pr merge 62 --repo example/repo --squash' "$case_dir/gh-axi.log" \ - "records-ahead-of-forge-call: the merge abstraction was never invoked" + assert_logged_gh_merge "$case_dir" 62 example/repo --squash assert_grep 'pr=https://github.com/example/repo/pull/62' "$case_dir/meta-at-merge" \ "records-ahead-of-forge-call: the merge ran before pr= was recorded" pass "fm-pr-merge records pr= before the forge call can land the merge" @@ -913,21 +878,8 @@ test_github_refusal_quotes_the_forge_output() { case_dir=$(make_case github-refusal-quotes-forge) mkdir -p "$case_dir/wt" add_gh_mocks "$case_dir" 6161616161616161616161616161616161616161 - cat > "$case_dir/fakebin/gh-axi" <<'SH' -#!/usr/bin/env bash -printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" -case "${1:-} ${2:-}" in - "pr merge") echo "will be added to the merge queue when all requirements are met" ;; - api\ *) - case " $* " in - *'{tip:'*) printf 'tip: %s\n' "$FM_TEST_DEFAULT_TIP" ;; - *) echo "gh-axi mock: unmodelled api query: $*" >&2 ; exit 2 ;; - esac - ;; -esac -exit 0 -SH - chmod +x "$case_dir/fakebin/gh-axi" + printf '%s\n' 'will be added to the merge queue when all requirements are met' \ + > "$case_dir/github-merge-output" write_github_outcome "$case_dir" OPEN false false main : > "$case_dir/gh-axi.log" : > "$case_dir/gh.log" @@ -979,7 +931,7 @@ test_github_auto_merge_spellings_are_refused_before_the_merge() { set +e run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/66 \ - -- "$spelling" --merge \ + --attended-override -- "$spelling" --merge \ > "$case_dir/stdout" 2> "$case_dir/stderr" rc=$? set -e @@ -1015,7 +967,7 @@ test_no_argument_position_launders_a_refused_flag() { # Every value-taking entry on the allow-list crossed with the refused flag, in # BOTH orders: the refused flag standing where a value belongs, and standing # ahead of a well-formed pair that would otherwise consume it. - for taker in --method --sha --subject --body --body-file -t -b -F; do + for taker in --method --subject --body --body-file -t -b -F; do for order in "$taker --auto" "--auto $taker value"; do number=$((number + 1)) url="https://github.com/example/repo/pull/$number" @@ -1043,7 +995,7 @@ test_no_argument_position_launders_a_refused_flag() { # widened for in the first place, and --subject=<value> is how a value that # legitimately begins with a dash is still passed: it is ONE token, so nothing # can read it as a flag standing on its own. - for order in '--sha abc123' '--subject fix' '--subject=-fix'; do + for order in '--subject fix' '--subject=-fix'; do number=$((number + 1)) url="https://github.com/example/repo/pull/$number" case_dir=$(make_case "arg-guard-admits-$number") @@ -1060,8 +1012,8 @@ test_no_argument_position_launders_a_refused_flag() { expect_code 0 "$rc" \ "arg-guard-admits: '$order' is on the allow-list and must still merge" - grep -qxF "pr merge $number --repo example/repo --squash $order" \ - "$case_dir/gh-axi.log" \ + grep -qxF "pr merge $number --repo example/repo --match-head-commit cccccccccccccccccccccccccccccccccccccccc --squash $order" \ + "$case_dir/gh.log" \ || fail "arg-guard-admits: '$order' was not forwarded to the forge unchanged" done pass "no argument position lets a refused flag reach the forge, and detached values still pass" @@ -1104,24 +1056,7 @@ test_github_unreadable_queue_rules_are_not_reported_as_no_queue() { mkdir -p "$case_dir/wt" add_gh_mocks "$case_dir" 8484848484848484848484848484848484848484 write_github_outcome "$case_dir" OPEN false false main - cat > "$case_dir/fakebin/gh" <<'SH' -#!/usr/bin/env bash -printf '%s\n' "$*" >> "$FM_TEST_GH_LOG" -case "${1:-} ${2:-}" in - "pr view") - case " $* " in - *headRefOid*) printf '%s\n' 8484848484848484848484848484848484848484 ; exit 0 ;; - esac - ;; - "api graphql") - cat "$FM_TEST_GH_OUTCOME" - exit 0 - ;; - api\ *) exit 1 ;; -esac -exit 0 -SH - chmod +x "$case_dir/fakebin/gh" + : > "$case_dir/github-rules-fail" : > "$case_dir/gh-axi.log" : > "$case_dir/gh.log" @@ -1163,6 +1098,42 @@ test_github_no_queue_rule_says_nothing_about_a_queue() { pass "fm-pr-merge says nothing about a merge queue when the base branch has no queue rule" } +test_github_unmerged_fallback_cannot_replace_queue_aware_read() { + local case_dir rc + case_dir=$(make_case github-unmerged-fallback) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 8686868686868686868686868686868686868686 + add_gh_mock_outcome_read_fails "$case_dir" + cat > "$case_dir/fakebin/gh-axi" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" +case "${1:-} ${2:-}" in + "pr view") printf 'pull_request:\n number: %s\n state: open\n' "$3" ;; +esac +exit 0 +SH + chmod +x "$case_dir/fakebin/gh-axi" + : > "$case_dir/gh-axi.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/73 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-unmerged-fallback: an unproved merge must fail" + assert_grep 'pr view 73 --repo example/repo' "$case_dir/gh-axi.log" \ + "github-unmerged-fallback: the fallback view was not consulted" + assert_grep 'the gh read failed and the gh-axi view could not prove the outcome either' \ + "$case_dir/stderr" \ + "github-unmerged-fallback: an unmerged fallback was treated as a readable outcome" + assert_no_grep 'GitHub merge outcome was not successful' "$case_dir/stderr" \ + "github-unmerged-fallback: an unmerged fallback reached detailed outcome handling" + assert_no_grep 'verified: ' "$case_dir/stdout" \ + "github-unmerged-fallback: an unproved merge was reported as verified" + pass "fm-pr-merge accepts only a proved merge from the gh-axi fallback" +} + test_github_unreadable_outcome_refusal_quotes_the_forge_output() { local case_dir rc case_dir=$(make_case github-unreadable-outcome-quotes-forge) @@ -1188,6 +1159,7 @@ SH # unreadable OUTCOME it is named for, rather than dying on an unreadable target. write_github_outcome "$case_dir" OPEN false false main add_gh_mock_outcome_read_fails_from "$case_dir" 8787878787878787878787878787878787878787 2 + printf '%s\n' 'will be added to the merge queue when all requirements are met' > "$case_dir/github-merge-output" : > "$case_dir/gh-axi.log" : > "$case_dir/gh.log" @@ -1309,18 +1281,14 @@ test_github_without_gh_still_uses_gh_axi_merge() { rc=$? set -e - expect_code 0 "$rc" "github-without-gh: gh-axi can prove a landed merge without gh" - assert_grep 'pr merge 60 --repo example/repo --squash' "$case_dir/gh-axi.log" \ - "github-without-gh: the configured merge abstraction was not invoked" - # Already merged when the run starts, so the view that answers here is the - # PRE-merge one; github|degraded-no-gh in the landed-merge invariant is what - # measures the POST-merge answer, against a fixture that reads OPEN until the - # merge runs. - assert_grep 'pr view 60 --repo example/repo' "$case_dir/gh-axi.log" \ - "github-without-gh: the gh-axi view was never consulted on a host with no gh" - assert_grep 'verified: https://github.com/example/repo/pull/60 is merged' \ - "$case_dir/stdout" "github-without-gh: the fallback did not report the proven merge" - pass "fm-pr-merge reaches and verifies the gh-axi merge path without gh" + expect_code 1 "$rc" "github-without-gh: missing gh must refuse before recording" + assert_grep 'merging a GitHub pull request requires gh on PATH' "$case_dir/stderr" \ + "github-without-gh: missing gh was not named" + assert_no_grep 'pr=' "$case_dir/state/task-x1.meta" \ + "github-without-gh: pr= was recorded without gh" + assert_absent "$case_dir/state/task-x1.check.sh" \ + "github-without-gh: a merge poll was armed without gh" + pass "fm-pr-merge refuses a GitHub merge when gh is missing, before recording" } test_github_without_gh_failed_read_keeps_bookkeeping() { @@ -1369,16 +1337,14 @@ SH rc=$? set -e - expect_code 1 "$rc" "github-without-gh-read-fails: an unreadable outcome must fail" - assert_grep 'pr merge 61 --repo example/repo --squash' "$case_dir/gh-axi.log" \ - "github-without-gh-read-fails: the merge call did not happen before the failed read" - assert_grep 'could not read the GitHub pull request outcome after the merge attempt' \ - "$case_dir/stderr" "github-without-gh-read-fails: the failed read was not reported" - assert_grep 'pr=https://github.com/example/repo/pull/61' "$case_dir/state/task-x1.meta" \ - "github-without-gh-read-fails: a landed merge lost its PR metadata" - assert_present "$case_dir/state/task-x1.check.sh" \ - "github-without-gh-read-fails: a landed merge lost its merge poll" - pass "fm-pr-merge preserves bookkeeping when gh is absent and the fallback read fails" + expect_code 1 "$rc" "github-without-gh-read-fails: missing gh must refuse before recording" + assert_grep 'merging a GitHub pull request requires gh on PATH' "$case_dir/stderr" \ + "github-without-gh-read-fails: missing gh was not named" + assert_no_grep 'pr=' "$case_dir/state/task-x1.meta" \ + "github-without-gh-read-fails: pr= was recorded without gh" + assert_absent "$case_dir/state/task-x1.check.sh" \ + "github-without-gh-read-fails: a merge poll was armed without gh" + pass "fm-pr-merge refuses a GitHub merge when gh is missing rather than merging blind" } test_github_zero_exit_queue_required_refuses_with_exact_retry() { @@ -1402,18 +1368,14 @@ test_github_zero_exit_queue_required_refuses_with_exact_retry() { "github-zero-exit-queue-required: refusal did not name the concrete observed state" assert_grep 'base branch release/2026 requires the merge queue' "$case_dir/stderr" \ "github-zero-exit-queue-required: refusal did not name the queue requirement" - assert_grep '-- --auto --rebase' "$case_dir/stderr" \ + assert_grep '--attended-override -- --auto --rebase' "$case_dir/stderr" \ "github-zero-exit-queue-required: refusal did not name the exact compatible flags" assert_grep 'api --paginate repos/example/repo/rules/branches/release%2F2026' "$case_dir/gh.log" \ "github-zero-exit-queue-required: queue rules were not read with pagination and encoded branch path" - grep -qxF 'pr merge 56 --repo example/repo --squash' "$case_dir/gh-axi.log" \ - || fail "github-zero-exit-queue-required: the attempted merge was changed unexpectedly" - # Count the MERGES rather than the whole log: the guarded path also asks gh-axi - # for the target and the base tip, and a line count silently turns "one merge" - # into "one forge call of any kind". - [ "$(count_log_lines "$case_dir/gh-axi.log" '^pr merge ')" = 1 ] \ + assert_logged_gh_merge "$case_dir" 56 example/repo --squash + [ "$(grep -c '^pr merge ' "$case_dir/gh.log")" -eq 1 ] \ || fail "github-zero-exit-queue-required: the wrapper attempted more than one merge" - assert_no_grep --auto "$case_dir/gh-axi.log" \ + assert_no_grep --auto "$case_dir/gh.log" \ "github-zero-exit-queue-required: queue flags were auto-applied to the attempted merge" assert_grep 'pr=https://github.com/example/repo/pull/56' "$case_dir/state/task-x1.meta" \ "github-zero-exit-queue-required: the attempted merge lost its PR reference" @@ -1443,7 +1405,7 @@ test_github_closed_unqueued_outcome_omits_retry_flags() { "github-closed-unqueued: refusal did not name the concrete observed state" assert_no_grep 'requires the merge queue' "$case_dir/stderr" \ "github-closed-unqueued: closed PR received unusable queue guidance" - assert_no_grep '-- --auto --merge' "$case_dir/stderr" \ + assert_no_grep '--attended-override -- --auto --merge' "$case_dir/stderr" \ "github-closed-unqueued: closed PR received retry flags" assert_grep 'pr=https://github.com/example/repo/pull/57' "$case_dir/state/task-x1.meta" \ "github-closed-unqueued: the attempted merge lost its PR reference" @@ -1504,10 +1466,9 @@ test_github_queue_required_refusal_names_retry_flags() { "github-queue-required: the original forge failure was not preserved" assert_grep 'base branch master requires the merge queue' "$case_dir/stderr" \ "github-queue-required: refusal did not name the queue requirement" - grep -F -- '-- --auto --merge' "$case_dir/stderr" >/dev/null \ + grep -F -- '--attended-override -- --auto --merge' "$case_dir/stderr" >/dev/null \ || fail "github-queue-required: refusal did not name the exact compatible flags" - grep -qxF 'pr merge 54 --repo example/repo --squash' "$case_dir/gh-axi.log" \ - || fail "github-queue-required: the wrapper silently changed the attempted merge semantics" + assert_logged_gh_merge "$case_dir" 54 example/repo --squash assert_present "$case_dir/state/task-x1.check.sh" \ "github-queue-required: the failed forge call did not leave the merge poll armed" pass "fm-pr-merge explains how to retry with the required GitHub merge queue method" @@ -1532,7 +1493,7 @@ test_github_agreeing_queue_rules_keep_retry_guidance() { expect_code 1 "$rc" "github-agreeing-queue-rules: an unproved merge must fail" assert_grep 'base branch main requires the merge queue' "$case_dir/stderr" \ "github-agreeing-queue-rules: refusal did not name the queue requirement" - assert_grep '-- --auto --rebase' "$case_dir/stderr" \ + assert_grep '--attended-override -- --auto --rebase' "$case_dir/stderr" \ "github-agreeing-queue-rules: agreeing rules omitted exact retry flags" assert_no_grep 'exact retry flags are ambiguous' "$case_dir/stderr" \ "github-agreeing-queue-rules: agreeing rules were reported as ambiguous" @@ -1560,9 +1521,9 @@ test_github_conflicting_queue_rules_report_ambiguity() { assert_grep 'base branch main has conflicting merge queue methods (MERGE, SQUASH)' \ "$case_dir/stderr" \ "github-conflicting-queue-rules: conflicting methods were not named" - assert_no_grep '-- --auto --merge' "$case_dir/stderr" \ + assert_no_grep '--attended-override -- --auto --merge' "$case_dir/stderr" \ "github-conflicting-queue-rules: an exact retry method was guessed" - assert_no_grep '-- --auto --squash' "$case_dir/stderr" \ + assert_no_grep '--attended-override -- --auto --squash' "$case_dir/stderr" \ "github-conflicting-queue-rules: an exact retry method was guessed" assert_no_grep 'SQUASH, SQUASH' "$case_dir/stderr" \ "github-conflicting-queue-rules: a repeated queue method was named twice" @@ -1576,12 +1537,25 @@ test_extra_merge_args_forwarded() { add_gh_mocks "$case_dir" 2222222222222222222222222222222222222222 : > "$case_dir/gh-axi.log" + set +e run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/15 -- --squash --delete-branch \ - > "$case_dir/stdout" 2> "$case_dir/stderr" || fail "extra-args: fm-pr-merge failed" + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "extra-args: branch deletion must be refused without --attended-override" + assert_grep 'pass --attended-override only for an explicit captain instruction' "$case_dir/stderr" \ + "extra-args: refusal did not name --attended-override" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "extra-args: gh pr merge ran despite the denylist" - grep -qxF 'pr merge 15 --repo example/repo --squash --delete-branch' "$case_dir/gh-axi.log" \ - || fail "extra-args: extra gh-axi pr merge flags were not forwarded" - pass "fm-pr-merge forwards extra flags to gh-axi pr merge after the -- separator" + case_dir=$(make_case extra-args-attended) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 2222222222222222222222222222222222222222 + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/15 \ + --attended-override -- --squash --delete-branch \ + > "$case_dir/stdout" 2> "$case_dir/stderr" || fail "extra-args-attended: attended override should merge" + assert_logged_gh_merge "$case_dir" 15 example/repo --squash --delete-branch + pass "fm-pr-merge refuses branch deletion unless --attended-override is passed" } test_missing_meta_refuses_before_merge() { @@ -1601,7 +1575,7 @@ test_missing_meta_refuses_before_merge() { expect_code 1 "$rc" "missing-meta: fm-pr-merge should refuse" assert_grep 'error: task metadata is unavailable' "$case_dir/stderr" \ "missing-meta: refusal did not explain missing meta" - [ ! -s "$case_dir/gh-axi.log" ] || fail "missing-meta: gh-axi pr merge was invoked" + [ ! -s "$case_dir/gh.log" ] || fail "missing-meta: gh pr merge was invoked" assert_absent "$case_dir/state/missing-x1.check.sh" \ "missing-meta: fm-pr-check should not arm a poll for an unknown task" pass "fm-pr-merge refuses before merging when task meta is missing" @@ -1630,7 +1604,7 @@ test_malformed_url_refuses_before_merge() { "malformed-url: malformed PR URL was recorded in meta" assert_absent "$case_dir/state/task-x1.check.sh" \ "malformed-url: malformed PR URL armed a merge poll" - assert_no_grep 'pr merge' "$case_dir/gh-axi.log" \ + assert_no_grep 'pr merge' "$case_dir/gh.log" \ "malformed-url: gh-axi pr merge was invoked for a malformed URL" pass "fm-pr-merge refuses malformed PR URLs before calling gh-axi" } @@ -1657,7 +1631,7 @@ test_rejects_unsafe_url_segments_before_recording() { "unsafe-url-segment: unsafe PR URL was recorded in meta" assert_absent "$case_dir/state/task-x1.check.sh" \ "unsafe-url-segment: unsafe PR URL armed a merge poll" - assert_no_grep 'pr merge' "$case_dir/gh-axi.log" \ + assert_no_grep 'pr merge' "$case_dir/gh.log" \ "unsafe-url-segment: gh-axi pr merge was invoked for an unsafe URL" pass "fm-pr-merge refuses unsafe PR URL segments before recording state" } @@ -1682,7 +1656,7 @@ test_repo_override_args_refuse_before_recording() { "repo-override: PR URL was recorded before rejecting repo override" assert_absent "$case_dir/state/task-x1.check.sh" \ "repo-override: repo override armed a merge poll" - assert_no_grep 'pr merge' "$case_dir/gh-axi.log" \ + assert_no_grep 'pr merge' "$case_dir/gh.log" \ "repo-override: gh-axi pr merge was invoked despite repo override" pass "fm-pr-merge refuses repo override args before recording state" } @@ -1707,7 +1681,7 @@ test_bundled_repo_override_args_refuse_before_recording() { "bundled-repo-override: PR URL was recorded before rejecting the bundled repo override" assert_absent "$case_dir/state/task-x1.check.sh" \ "bundled-repo-override: a bundled repo override armed a merge poll" - assert_no_grep 'pr merge' "$case_dir/gh-axi.log" \ + assert_no_grep 'pr merge' "$case_dir/gh.log" \ "bundled-repo-override: gh-axi pr merge was invoked despite the bundled repo override" case_dir=$(make_gitlab_case bundled-repo-override-gitlab) @@ -1735,13 +1709,23 @@ test_bundled_repo_override_args_refuse_before_recording() { add_gh_mocks "$case_dir" bcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbc : > "$case_dir/gh-axi.log" + set +e run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/8 -- -d \ - > "$case_dir/stdout" 2> "$case_dir/stderr" \ - || fail "bundled-non-repo-cluster: fm-pr-merge refused a short flag that overrides nothing" + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "bundled-non-repo-cluster: -d is branch deletion and must be refused" + assert_grep 'pass --attended-override only for an explicit captain instruction' "$case_dir/stderr" \ + "bundled-non-repo-cluster: refusal did not name --attended-override" - grep -qxF 'pr merge 8 --repo example/repo --squash -d' "$case_dir/gh-axi.log" \ - || fail "bundled-non-repo-cluster: a short flag carrying no repository override was not forwarded" - pass "fm-pr-merge refuses a bundled short-option repo override and forwards other short flags" + case_dir=$(make_case bundled-delete-attended) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" bcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbc + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/8 --attended-override -- -d \ + > "$case_dir/stdout" 2> "$case_dir/stderr" \ + || fail "bundled-delete-attended: attended override should merge" + assert_logged_gh_merge "$case_dir" 8 example/repo --squash -d + pass "fm-pr-merge refuses a bundled short-option repo override and refuses -d unless attended" } test_explicit_merge_method_not_overridden() { @@ -1754,8 +1738,7 @@ test_explicit_merge_method_not_overridden() { run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/22 -- --merge \ > "$case_dir/stdout" 2> "$case_dir/stderr" || fail "explicit-merge-method: fm-pr-merge failed" - grep -qxF 'pr merge 22 --repo example/repo --merge' "$case_dir/gh-axi.log" \ - || fail "explicit-merge-method: caller --merge was not forwarded without an extra default --squash" + assert_logged_gh_merge "$case_dir" 22 example/repo --merge pass "fm-pr-merge does not add default --squash when the caller passes an explicit merge method" } @@ -1769,8 +1752,7 @@ test_method_equals_merge_method_not_overridden() { run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/23 -- --method=merge \ > "$case_dir/stdout" 2> "$case_dir/stderr" || fail "method-equals-merge-method: fm-pr-merge failed" - grep -qxF 'pr merge 23 --repo example/repo --method=merge' "$case_dir/gh-axi.log" \ - || fail "method-equals-merge-method: caller --method=merge was not forwarded without an extra default --squash" + assert_logged_gh_merge "$case_dir" 23 example/repo --merge pass "fm-pr-merge respects --method=<value> as an explicit merge method" } @@ -1784,8 +1766,7 @@ test_parses_pr_url_for_gh_axi() { run_pr_merge "$case_dir" task-x1 https://github.com/my-org/my-repo/pull/126 \ > "$case_dir/stdout" 2> "$case_dir/stderr" || fail "url-parsing: fm-pr-merge failed" - grep -qxF 'pr merge 126 --repo my-org/my-repo --squash' "$case_dir/gh-axi.log" \ - || fail "url-parsing: gh-axi pr merge was not invoked as number + --repo + default --squash" + assert_logged_gh_merge "$case_dir" 126 my-org/my-repo --squash pass "fm-pr-merge parses a GitHub PR URL into gh-axi number and --repo arguments" } @@ -1860,8 +1841,7 @@ test_no_recorded_issue_makes_no_issue_calls() { assert_no_grep 'issue ' "$case_dir/gh-axi.log" \ "no-issue: merge path made an issue API call without recorded issue metadata" - grep -qxF 'pr merge 34 --repo example/repo --squash' "$case_dir/gh-axi.log" \ - || fail "no-issue: ordinary merge invocation changed" + assert_logged_gh_merge "$case_dir" 34 example/repo --squash pass "fm-pr-merge preserves the ordinary path when no issue is recorded" } @@ -2072,7 +2052,7 @@ test_gitea_work_item_is_closed_with_its_own_credential() { set -e expect_code 0 "$rc" "gitea-close: the merge with a gitea work item failed" - assert_grep 'pr merge 54 --repo example/repo' "$case_dir/gh-axi.log" \ + assert_grep 'pr merge 54 --repo example/repo' "$case_dir/gh.log" \ "gitea-close: the PR was never merged" [ "$(cat "$case_dir/gitea-store/issue-state" 2>/dev/null)" = closed ] \ || fail "gitea-close: the gitea issue was not closed" @@ -2108,7 +2088,7 @@ test_gitea_close_failure_keeps_merge_success_unambiguous() { set -e expect_code 0 "$rc" "gitea-down: an unreachable gitea host made a completed merge look retryable" - assert_grep 'pr merge 56 --repo example/repo' "$case_dir/gh-axi.log" \ + assert_grep 'pr merge 56 --repo example/repo' "$case_dir/gh.log" \ "gitea-down: the merge did not happen while the tracker was unreachable" assert_grep 'issue bookkeeping did not complete' "$case_dir/stderr" \ "gitea-down: the failed close was silent" @@ -2134,7 +2114,7 @@ test_gitea_verification_failure_names_its_own_reason() { set -e expect_code 0 "$rc" "gitea-403: a refused credential made a completed merge look retryable" - assert_grep 'pr merge 58 --repo example/repo' "$case_dir/gh-axi.log" \ + assert_grep 'pr merge 58 --repo example/repo' "$case_dir/gh.log" \ "gitea-403: the merge did not happen while the tracker refused the credential" assert_grep 'could not verify' "$case_dir/stderr" \ "gitea-403: the failed verification was silent" @@ -2259,6 +2239,7 @@ test_refresh_failure_warning_names_the_cause() { case_dir=$(make_case refresh-cause) mkdir -p "$case_dir/wt" add_gh_mocks "$case_dir" 4444444444444444444444444444444444444444 + : > "$case_dir/github-outcome.cache-invalid" : > "$case_dir/gh-axi.log" set +e @@ -2378,12 +2359,22 @@ test_gitlab_extra_args_forwarded() { > "$case_dir/stdout" 2> "$case_dir/stderr" rc=$? set -e + expect_code 1 "$rc" "gitlab-extra-args: source-branch deletion must be refused without --attended-override" + assert_grep 'pass --attended-override only for an explicit captain instruction' "$case_dir/stderr" \ + "gitlab-extra-args: refusal did not name --attended-override" + [ ! -s "$case_dir/glab.log" ] || fail "gitlab-extra-args: glab ran despite the denylist" - expect_code 0 "$rc" "gitlab-extra-args: merge should succeed" + case_dir=$(make_gitlab_case gitlab-extra-args-attended) + set +e + run_pr_merge "$case_dir" task-x1 "$MR_URL" --attended-override -- --remove-source-branch \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 0 "$rc" "gitlab-extra-args-attended: attended override should merge" merge_line=$(glab_merge_line "$case_dir/glab.log") [ "$merge_line" = "GITLAB_HOST=$MR_HOST mr merge 7 -R $MR_PROJECT_URL --sha $MR_HEAD --yes --remove-source-branch" ] \ - || fail "gitlab-extra-args: extra glab flags were not forwarded: '$merge_line'" - pass "fm-pr-merge forwards extra flags to glab mr merge after the -- separator" + || fail "gitlab-extra-args-attended: extra glab flags were not forwarded: '$merge_line'" + pass "fm-pr-merge refuses GitLab source-branch deletion unless --attended-override is passed" } test_gitlab_merge_failure_propagates() { @@ -2403,6 +2394,10 @@ test_gitlab_merge_failure_propagates() { pass "fm-pr-merge propagates a real glab merge failure without silently succeeding" } +# Each pre-merge condition, driven one at a time, so no condition can be +# carried by another. The refusal names that condition, no merge is attempted, +# and pr= is still recorded and the poll still armed exactly as the GitHub path +# leaves them when live verification or the gh merge fails. test_gitlab_each_condition_refuses_independently() { local case_dir rc name expected spec set -- \ @@ -2594,26 +2589,23 @@ test_gitlab_head_override_args_refuse_before_recording() { } test_github_still_forwards_sha_arg() { - local case_dir + local case_dir rc case_dir=$(make_case github-sha-arg) mkdir -p "$case_dir/wt" add_gh_mocks "$case_dir" dddddddddddddddddddddddddddddddddddddddd : > "$case_dir/gh-axi.log" - # --sha is ADMITTED to the forwarded-argument allow-list as a head-binding - # argument: it constrains the mutation to the state this run verified, cannot - # defer execution, and makes the merge strictly narrower, so a push landing - # between validation and merge makes the forge refuse rather than merge - # something nobody checked. It is forwarded because it is admitted - # categorically and recorded in the allow-list, NOT because the unbounded - # passthrough survived - that was deliberately closed, and every other - # unlisted argument is refused by name. + set +e run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/44 -- --sha abc123 \ - > "$case_dir/stdout" 2> "$case_dir/stderr" || fail "github-sha-arg: fm-pr-merge failed" - - grep -qxF 'pr merge 44 --repo example/repo --squash --sha abc123' "$case_dir/gh-axi.log" \ - || fail "github-sha-arg: an admitted head-binding argument was not forwarded" - pass "fm-pr-merge forwards --sha as an admitted head-binding argument" + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "github-sha-arg: a caller --sha must be refused on GitHub too" + assert_grep 'extra merge arguments must not override the head commit' "$case_dir/stderr" \ + "github-sha-arg: refusal did not name the head override" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-sha-arg: gh pr merge ran despite the head override" + pass "fm-pr-merge refuses a caller --sha on GitHub because the head comes from the live read" } # --- durable merge outcome --------------------------------------------------- @@ -2829,6 +2821,11 @@ test_distinct_merged_prs_keep_distinct_wakes() { rm -f "$case_dir/state/task-x1.check.sh" \ "$case_dir/state/task-x1.pr-poll" \ "$case_dir/state/task-x1.pr-poll-registration" + # Reused tasks re-bind through fm-pr-check before the next merge. Merge + # refuses a URL that is not the recorded pr=, so drop the first PR identity. + grep -vE '^(pr|pr_head)=' "$case_dir/state/task-x1.meta" \ + > "$case_dir/state/task-x1.meta.rebind" + mv "$case_dir/state/task-x1.meta.rebind" "$case_dir/state/task-x1.meta" FM_TEST_HOME="$case_dir/home" run_pr_merge "$case_dir" task-x1 "$second_url" \ >"$case_dir/stdout-2" 2>"$case_dir/stderr-2" \ || fail "distinct-merge-wakes: second merge failed" @@ -3003,8 +3000,7 @@ test_absent_backlog_still_merges() { expect_code 0 "$rc" "absent-backlog-merges: a home with no backlog must still merge" assert_no_grep 'held for the captain' "$case_dir/stderr" \ "absent-backlog-merges: an absent backlog was read as a captain hold" - grep -qxF 'pr merge 61 --repo example/repo --squash' "$case_dir/gh-axi.log" \ - || fail "absent-backlog-merges: the merge was not attempted" + assert_logged_gh_merge "$case_dir" 61 example/repo --squash pass "fm-pr-merge proceeds when the home carries no backlog at all" } @@ -3026,8 +3022,8 @@ test_unreadable_backlog_refuses_the_merge() { expect_code 1 "$rc" "unreadable-backlog-refuses: an unreadable authority record must refuse" assert_grep 'refusing to merge' "$case_dir/stderr" \ "unreadable-backlog-refuses: the refusal did not say it refused to merge" - [ ! -s "$case_dir/gh-axi.log" ] \ - || fail "unreadable-backlog-refuses: the forge was called despite an unreadable record" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "unreadable-backlog-refuses: the forge merge ran despite an unreadable record" pass "fm-pr-merge refuses when the backlog exists but cannot be read" } @@ -3050,8 +3046,8 @@ test_unreadable_backend_config_refuses_the_merge() { expect_code 1 "$rc" "unreadable-backend-config-refuses: an unreadable authority route must refuse" assert_grep 'tasks-axi backend configuration cannot be read' "$case_dir/stderr" \ "unreadable-backend-config-refuses: the unreadable authority route was not named" - [ ! -s "$case_dir/gh-axi.log" ] \ - || fail "unreadable-backend-config-refuses: the forge was called despite an unreadable authority route" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "unreadable-backend-config-refuses: the forge merge ran despite an unreadable authority route" pass "fm-pr-merge refuses when its configured backend cannot be read" } @@ -3077,8 +3073,8 @@ test_unreadable_user_backend_config_refuses_the_merge() { expect_code 1 "$rc" "unreadable-user-backend-config-refuses: an unreadable authority route must refuse" assert_grep "tasks-axi backend configuration cannot be read at $user_config" "$case_dir/stderr" \ "unreadable-user-backend-config-refuses: the unreadable authority route was not named" - [ ! -s "$case_dir/gh-axi.log" ] \ - || fail "unreadable-user-backend-config-refuses: the forge was called despite an unreadable authority route" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "unreadable-user-backend-config-refuses: the forge merge ran despite an unreadable authority route" pass "fm-pr-merge refuses when its user backend configuration cannot be read" } @@ -3104,8 +3100,8 @@ test_untraversable_user_backend_config_directory_refuses_the_merge() { expect_code 1 "$rc" "untraversable-user-backend-config-directory-refuses: an unreadable authority route must refuse" assert_grep "tasks-axi backend configuration cannot be read at $user_config" "$case_dir/stderr" \ "untraversable-user-backend-config-directory-refuses: the unreadable authority route was not named" - [ ! -s "$case_dir/gh-axi.log" ] \ - || fail "untraversable-user-backend-config-directory-refuses: the forge was called despite an unreadable authority route" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "untraversable-user-backend-config-directory-refuses: the forge merge ran despite an unreadable authority route" pass "fm-pr-merge refuses when its user backend configuration directory cannot be traversed" } @@ -3125,10 +3121,9 @@ test_absent_user_backend_config_directory_and_backlog_still_merge() { set -e expect_code 0 "$rc" "absent-user-backend-config-directory-and-backlog-merges: sound defaults and no backlog must permit merging" - [ "$(grep -c '^pr merge ' "$case_dir/gh-axi.log")" -eq 1 ] \ + [ "$(grep -c '^pr merge ' "$case_dir/gh.log")" -eq 1 ] \ || fail "absent-user-backend-config-directory-and-backlog-merges: the forge must merge exactly once" - grep -qxF 'pr merge 67 --repo example/repo --squash' "$case_dir/gh-axi.log" \ - || fail "absent-user-backend-config-directory-and-backlog-merges: the expected merge was not attempted" + assert_logged_gh_merge "$case_dir" 67 example/repo --squash pass "fm-pr-merge proceeds once when its user configuration directory and backlog are genuinely absent" } @@ -3152,11 +3147,706 @@ test_backend_override_bypasses_unreadable_user_config() { chmod 644 "$user_config" expect_code 0 "$rc" "backend-override-bypasses-unreadable-user-config: an explicit backend must bypass config" - grep -qxF 'pr merge 65 --repo example/repo --squash' "$case_dir/gh-axi.log" \ - || fail "backend-override-bypasses-unreadable-user-config: the merge was not attempted" + assert_logged_gh_merge "$case_dir" 65 example/repo --squash pass "fm-pr-merge honors a backend override over an unreadable user configuration" } +test_github_red_checks_refuse_and_allow_red_waives_named() { + local case_dir rc head + head=aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa + case_dir=$(make_case github-red-checks) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_red_json "$case_dir" "$head" lint + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/80 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "github-red: a red check must refuse" + assert_grep "check 'lint' is not green" "$case_dir/stderr" \ + "github-red: the red check was not named" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-red: gh pr merge ran on a red PR" + + case_dir=$(make_case github-allow-red) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_red_json "$case_dir" "$head" lint + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/81 \ + --allow-red lint \ + > "$case_dir/stdout" 2> "$case_dir/stderr" || fail "github-allow-red: named waiver should merge" + assert_logged_gh_merge "$case_dir" 81 example/repo --squash + pass "fm-pr-merge refuses red GitHub checks and waives only a named --allow-red check" +} + +# When the base branch advances, GitHub cancels a pull request's in-flight run +# and re-triggers it, leaving the cancelled run in the rollup beside the passing +# re-run while reporting the pull request itself CLEAN. The merge must follow the +# current run rather than the one that re-run replaced. +test_superseded_failed_check_run_no_longer_refuses() { + local case_dir head + head=cccccccccccccccccccccccccccccccccccccccc + case_dir=$(make_case github-superseded-red) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_rollup_json "$case_dir" "$head" \ + "$(check_run ci COMPLETED CANCELLED 2026-01-01T00:00:01Z)" \ + "$(check_run ci COMPLETED SUCCESS 2026-01-01T00:00:09Z)" + + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/90 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" \ + || fail "github-superseded-red: a failed run replaced by a passing re-run must merge"$'\n'"$(cat "$case_dir/stderr")" + assert_logged_gh_merge "$case_dir" 90 example/repo --squash + pass "fm-pr-merge merges when a failed check run was replaced by a passing re-run" +} + +# Legacy status contexts remain independent from check runs, even when their +# reported names match. +test_check_runs_never_supersede_status_contexts() { + local case_dir rc head + head=cdcdcdcdcdcdcdcdcdcdcdcdcdcdcdcdcdcdcdcd + case_dir=$(make_case github-cross-check-kind) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_rollup_json "$case_dir" "$head" \ + "$(status_context ci FAILURE)" \ + "$(check_run ci COMPLETED SUCCESS 2026-01-01T00:00:09Z)" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/97 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "github-cross-check-kind: a failing status context must refuse" + assert_grep "check 'ci' is not green" "$case_dir/stderr" \ + "github-cross-check-kind: the status context was not named" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-cross-check-kind: a passing check run hid a failing status context" + pass "fm-pr-merge never lets a check run supersede a legacy status context" +} + +# The inverse, and the one that matters most: a check whose current run failed is +# still red however many earlier runs of it passed. +test_current_failed_check_run_still_refuses() { + local case_dir rc head + head=dddddddddddddddddddddddddddddddddddddddd + case_dir=$(make_case github-current-red) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_rollup_json "$case_dir" "$head" \ + "$(check_run ci COMPLETED SUCCESS 2026-01-01T00:00:01Z)" \ + "$(check_run ci COMPLETED FAILURE 2026-01-01T00:00:09Z)" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/91 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "github-current-red: a currently failing check must refuse" + assert_grep "check 'ci' is not green" "$case_dir/stderr" \ + "github-current-red: the red check was not named" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-current-red: gh pr merge ran on a currently failing check" + pass "fm-pr-merge still refuses when a check's current run failed after an earlier pass" +} + +# Run generation follows startedAt rather than the order overlapping runs finish. +test_late_finishing_old_success_does_not_hide_current_failure() { + local case_dir rc head + head=dededededededededededededededededededede + case_dir=$(make_case github-old-success-finishes-last) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_rollup_json "$case_dir" "$head" \ + "$(check_run ci COMPLETED SUCCESS 2026-01-01T00:00:01Z 2026-01-01T00:00:10Z)" \ + "$(check_run ci COMPLETED FAILURE 2026-01-01T00:00:09Z 2026-01-01T00:00:09Z)" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/98 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "github-old-success-finishes-last: the later-started failure must refuse" + assert_grep "check 'ci' is not green" "$case_dir/stderr" \ + "github-old-success-finishes-last: the current failure was not named" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-old-success-finishes-last: completion order hid the current failure" + pass "fm-pr-merge uses start order when the old success finishes last" +} + +# A cancelled old run may settle after the passing re-run that superseded it. +test_late_finishing_old_cancellation_is_superseded() { + local case_dir head + head=dfdfdfdfdfdfdfdfdfdfdfdfdfdfdfdfdfdfdfdf + case_dir=$(make_case github-old-cancellation-finishes-last) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_rollup_json "$case_dir" "$head" \ + "$(check_run ci COMPLETED CANCELLED 2026-01-01T00:00:01Z 2026-01-01T00:00:10Z)" \ + "$(check_run ci COMPLETED SUCCESS 2026-01-01T00:00:09Z 2026-01-01T00:00:09Z)" + + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/99 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" \ + || fail "github-old-cancellation-finishes-last: the passing re-run must merge"$'\n'"$(cat "$case_dir/stderr")" + assert_logged_gh_merge "$case_dir" 99 example/repo --squash + pass "fm-pr-merge supersedes an old cancellation that finishes last" +} + +# A re-run that has not finished proves nothing, so it can neither be superseded +# nor supersede: the check stays red whether the run it replaces passed or failed. +test_unfinished_rerun_keeps_a_check_red() { + local case_dir rc head prior + head=eeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeee + for prior in FAILURE SUCCESS; do + case_dir=$(make_case "github-pending-rerun-$prior") + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_rollup_json "$case_dir" "$head" \ + "$(check_run ci COMPLETED "$prior" 2026-01-01T00:00:01Z)" \ + "$(check_run ci IN_PROGRESS - -)" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/92 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "github-pending-rerun-$prior: an unfinished re-run must refuse" + assert_grep "check 'ci' is not green" "$case_dir/stderr" \ + "github-pending-rerun-$prior: the pending check was not named" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-pending-rerun-$prior: gh pr merge ran with a re-run still in flight" + done + pass "fm-pr-merge keeps a check red while its re-run is still in flight" +} + +# Supersession is scoped to one check name, which is also the name --allow-red +# matches, so a newer passing check never clears a different check's failure. +test_supersession_never_crosses_check_names() { + local case_dir rc head + head=ffffffffffffffffffffffffffffffffffffffff + case_dir=$(make_case github-cross-name) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_rollup_json "$case_dir" "$head" \ + "$(check_run lint COMPLETED FAILURE 2026-01-01T00:00:01Z)" \ + "$(check_run ci COMPLETED SUCCESS 2026-01-01T00:00:09Z)" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/93 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "github-cross-name: another check passing must not clear this failure" + assert_grep "check 'lint' is not green" "$case_dir/stderr" \ + "github-cross-name: the red check was not named" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-cross-name: gh pr merge ran on a red check of a different name" + pass "fm-pr-merge never lets one check's pass clear another check's failure" +} + +# Supersession has to be proven from the forge's own start timestamps, so a run +# GitHub dated in any other way is treated as undated and clears nothing. +test_undated_runs_never_supersede() { + local case_dir rc spec label older newer + local head=0a0a0a0a0a0a0a0a0a0a0a0a0a0a0a0a0a0a0a0a + set -- \ + 'undated-failure|-|2026-01-01T00:00:09Z' \ + 'undated-pass|2026-01-01T00:00:01Z|-' \ + 'fractional-pass|2026-01-01T00:00:01Z|2026-01-01T00:00:09.500Z' \ + 'offset-pass|2026-01-01T00:00:01Z|2026-01-01T00:00:09+00:00' + for spec in "$@"; do + label=${spec%%|*} + older=${spec#*|} + older=${older%%|*} + newer=${spec##*|} + case_dir=$(make_case "github-undated-$label") + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_rollup_json "$case_dir" "$head" \ + "$(check_run ci COMPLETED FAILURE "$older")" \ + "$(check_run ci COMPLETED SUCCESS "$newer")" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/94 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "github-undated-$label: an unproven supersession must refuse" + assert_grep "check 'ci' is not green" "$case_dir/stderr" \ + "github-undated-$label: the red check was not named" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-undated-$label: gh pr merge ran on an unproven supersession" + done + pass "fm-pr-merge clears a failure only on a proven later pass of the same check" +} + +# A superseded failure changes nothing about the waiver: --allow-red still covers +# exactly the named check, still needs every other check green, and the merge is +# still bound to the verified head. +test_allow_red_still_waives_only_the_current_failure() { + local case_dir rc head + head=0b0b0b0b0b0b0b0b0b0b0b0b0b0b0b0b0b0b0b0b + case_dir=$(make_case github-superseded-allow-red-wrong-name) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_rollup_json "$case_dir" "$head" \ + "$(check_run ci COMPLETED FAILURE 2026-01-01T00:00:01Z)" \ + "$(check_run ci COMPLETED SUCCESS 2026-01-01T00:00:09Z)" \ + "$(check_run lint COMPLETED FAILURE 2026-01-01T00:00:09Z)" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/95 \ + --allow-red ci > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "superseded-allow-red-wrong-name: waiving the green check must not merge" + assert_grep "check 'lint' is not green" "$case_dir/stderr" \ + "superseded-allow-red-wrong-name: the unwaived red check was not named" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "superseded-allow-red-wrong-name: gh pr merge ran with an unwaived red check" + + case_dir=$(make_case github-superseded-allow-red-named) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_rollup_json "$case_dir" "$head" \ + "$(check_run ci COMPLETED FAILURE 2026-01-01T00:00:01Z)" \ + "$(check_run ci COMPLETED SUCCESS 2026-01-01T00:00:09Z)" \ + "$(check_run lint COMPLETED FAILURE 2026-01-01T00:00:09Z)" + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/96 \ + --allow-red lint > "$case_dir/stdout" 2> "$case_dir/stderr" \ + || fail "superseded-allow-red-named: the named waiver should merge"$'\n'"$(cat "$case_dir/stderr")" + assert_logged_gh_merge "$case_dir" 96 example/repo --squash + pass "fm-pr-merge keeps --allow-red scoped to its named check beside a superseded failure" +} + +test_allow_red_is_refused_while_away() { + local case_dir rc head + head=abababababababababababababababababababab + case_dir=$(make_case github-allow-red-away) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_red_json "$case_dir" "$head" lint + write_away_record "$case_dir" --grant task-x1 + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/82 \ + --allow-red lint \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 2 "$rc" "github-allow-red-away: --allow-red must be refused while away" + assert_grep '--allow-red is attended-only' "$case_dir/stderr" \ + "github-allow-red-away: refusal did not name attended-only" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-allow-red-away: gh pr merge ran despite away --allow-red" + + case_dir=$(make_case github-allow-red-away-after-view) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_red_json "$case_dir" "$head" lint + write_away_record "$case_dir" --grant task-x1 + mv "$case_dir/state/.afk-contract" "$case_dir/away-record-after-view" + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/82 \ + --allow-red lint \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 2 "$rc" "github-allow-red-away-after-view: late away publication must refuse --allow-red" + assert_grep '--allow-red is attended-only' "$case_dir/stderr" \ + "github-allow-red-away-after-view: late refusal did not name attended-only" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-allow-red-away-after-view: gh pr merge ran after late away publication" + pass "fm-pr-merge rechecks away presence before an attended red merge" +} + +test_allow_red_requires_one_separate_name() { + local case_dir rc head + head=afafafafafafafafafafafafafafafafafafafaf + + case_dir=$(make_case github-allow-red-equals) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_red_json "$case_dir" "$head" lint + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/87 \ + --allow-red=lint > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 2 "$rc" "github-allow-red-equals: equals form must be refused" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-allow-red-equals: gh pr merge ran for the equals alias" + + case_dir=$(make_case github-allow-red-duplicate) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_red_json "$case_dir" "$head" lint + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/88 \ + --allow-red lint --allow-red unit > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 2 "$rc" "github-allow-red-duplicate: duplicate waiver must be refused" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-allow-red-duplicate: gh pr merge ran for duplicate waivers" + pass "fm-pr-merge accepts exactly one separately named red-check waiver" +} + +test_away_grant_and_yolo_and_hold_for_return() { + local case_dir rc url head + head=acacacacacacacacacacacacacacacacacacacac + url=https://github.com/example/repo/pull/83 + + case_dir=$(make_case away-held) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_away_record "$case_dir" + set +e + run_pr_merge "$case_dir" task-x1 "$url" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "away-held: ungranted merge must refuse" + assert_grep 'task task-x1 is held for the captain return' "$case_dir/stderr" \ + "away-held: refusal did not name hold-for-return" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "away-held: gh pr merge ran without a grant" + + case_dir=$(make_case away-held-attended-override) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_away_record "$case_dir" + set +e + run_pr_merge "$case_dir" task-x1 "$url" --attended-override \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "away-held-override: --attended-override must not skip the grant" + assert_grep 'task task-x1 is held for the captain return' "$case_dir/stderr" \ + "away-held-override: override skipped the grant" + + case_dir=$(make_case away-grant) + mkdir -p "$case_dir/wt" "$case_dir/home" + add_gh_mocks "$case_dir" "$head" + write_away_record "$case_dir" --grant task-x1 + FM_TEST_HOME="$case_dir/home" run_pr_merge "$case_dir" task-x1 "$url" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" || fail "away-grant: granted green merge should succeed" + assert_logged_gh_merge "$case_dir" 83 example/repo --squash + assert_grep "merge landed: task-x1 $url away-grant" "$case_dir/state/.wake-queue" \ + "away-grant: the durable outcome did not tag away-grant" + + case_dir=$(make_case away-yolo) + mkdir -p "$case_dir/wt" "$case_dir/home" + add_gh_mocks "$case_dir" "$head" + printf '\nyolo=on\n' >> "$case_dir/state/task-x1.meta" + write_away_record "$case_dir" + FM_TEST_HOME="$case_dir/home" run_pr_merge "$case_dir" task-x1 "$url" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" || fail "away-yolo: yolo green merge should succeed" + assert_grep "merge landed: task-x1 $url yolo" "$case_dir/state/.wake-queue" \ + "away-yolo: the durable outcome did not tag yolo" + pass "away merges require yolo or a grant, and --attended-override does not skip that" +} + +test_away_posture_refuses_asynchronous_merge_paths() { + local case_dir rc url head merge_line + head=abababababababababababababababababababab + url=https://github.com/example/repo/pull/89 + + case_dir=$(make_case away-auto-refused) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_away_record "$case_dir" --grant task-x1 + set +e + run_pr_merge "$case_dir" task-x1 "$url" --attended-override -- --auto --merge \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "away-auto-refused: fork allow-list must refuse auto-merge" + assert_grep 'refusing to forward --auto' "$case_dir/stderr" \ + "away-auto-refused: refusal did not name the asynchronous flag" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "away-auto-refused: gh pr merge ran for an away auto-merge request" + + case_dir=$(make_case away-queue-refused) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + printf 'merge_method=MERGE\n' > "$case_dir/github-rules" + write_away_record "$case_dir" --grant task-x1 + set +e + run_pr_merge "$case_dir" task-x1 "$url" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 2 "$rc" "away-queue-refused: a required merge queue must refuse before submission" + assert_grep 'merge-queue state does not prove an immediate merge' "$case_dir/stderr" \ + "away-queue-refused: refusal did not explain the away restriction" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "away-queue-refused: gh received a merge that could enter its queue" + + case_dir=$(make_gitlab_case away-gitlab-auto) + write_away_record "$case_dir" --grant task-x1 + set +e + run_pr_merge "$case_dir" task-x1 "$MR_URL" --attended-override -- --auto-merge \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "away-gitlab-auto: GitLab auto-merge must refuse" + assert_grep 'refusing to forward --auto-merge' "$case_dir/stderr" \ + "away-gitlab-auto: refusal did not name auto-merge" + [ -z "$(glab_merge_line "$case_dir/glab.log")" ] \ + || fail "away-gitlab-auto: glab received an asynchronous merge" + + case_dir=$(make_gitlab_case away-gitlab-configured merge_when_pipeline_succeeds=true) + write_away_record "$case_dir" --grant task-x1 + set +e + run_pr_merge "$case_dir" task-x1 "$MR_URL" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 2 "$rc" "away-gitlab-configured: configured auto-merge must refuse" + [ -z "$(glab_merge_line "$case_dir/glab.log")" ] \ + || fail "away-gitlab-configured: glab received a configured asynchronous merge" + + case_dir=$(make_gitlab_case away-gitlab-sync) + write_away_record "$case_dir" --grant task-x1 + run_pr_merge "$case_dir" task-x1 "$MR_URL" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" \ + || fail "away-gitlab-sync: an immediate granted merge should succeed" + merge_line=$(glab_merge_line "$case_dir/glab.log") + case "$merge_line" in + *" --auto-merge=false") ;; + *) fail "away-gitlab-sync: the final glab flag did not force an immediate merge: '$merge_line'" ;; + esac + pass "away posture permits immediate merges but refuses every asynchronous path" +} + +test_away_grant_does_not_bypass_red_or_identity() { + local case_dir rc head + head=adadadadadadadadadadadadadadadadadadadad + case_dir=$(make_case away-grant-red) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_red_json "$case_dir" "$head" lint + write_away_record "$case_dir" --grant task-x1 + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/84 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "away-grant-red: a grant must not waive red checks" + assert_grep "check 'lint' is not green" "$case_dir/stderr" \ + "away-grant-red: C1 did not refuse the red check" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "away-grant-red: gh pr merge ran on a granted red PR" + + case_dir=$(make_case pr-identity-mismatch) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + printf '\npr=https://github.com/example/repo/pull/99\n' >> "$case_dir/state/task-x1.meta" + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/85 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "pr-identity: a different recorded URL must refuse" + assert_grep 'is bound to https://github.com/example/repo/pull/99' "$case_dir/stderr" \ + "pr-identity: refusal did not name the recorded URL" + pass "a grant does not bypass red checks, and a recorded pr= must match the URL" +} + +test_unreadable_away_record_refuses_merge() { + local case_dir rc + case_dir=$(make_case away-unreadable) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" aeaeaeaeaeaeaeaeaeaeaeaeaeaeaeaeaeaeaeae + printf 'not-a-contract\n' > "$case_dir/state/.afk-contract" + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/86 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "away-unreadable: an unreadable away record must refuse" + assert_grep 'away-posture record could not be read' "$case_dir/stderr" \ + "away-unreadable: refusal did not fail closed" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "away-unreadable: gh pr merge ran despite an unreadable record" + pass "an unreadable away-posture record refuses the merge instead of skipping the grant" +} + +# The race this closes: the away record is read for merge authority and the +# forge is called afterwards, so an archive (the captain's return) or a grant +# revocation landing in between would merge on authority that no longer holds. +# away_change_script writes the change the gh mock attempts from inside the +# forge call, which IS that window. Its body drives the real away-record +# commands /afk and the return use, never a file edit, and takes a one-second +# lock bound so a contended case refuses quickly instead of waiting. +away_change_script() { # <case-dir> <name>; script body on stdin + local case_dir=$1 name=$2 path + path="$case_dir/$name" + { + printf '#!/usr/bin/env bash\n' + printf 'set -eu\n' + printf 'export FM_TEST_AFK_CONTRACT_LOCK_TIMEOUT=1\n' + printf 'CONTRACT="%s/bin/fm-afk-contract.sh"\n' "$ROOT" + cat + } > "$path" + chmod +x "$path" + printf '%s\n' "$path" +} + +# Two away-record changes, each attempted from inside the merge's critical +# section: the archive a captain return performs, and the replacement that +# revokes a grant. Neither may land there, and the merge must still complete on +# the authority it read. +test_away_record_cannot_change_between_the_authority_read_and_the_merge() { + local case_dir rc mutate + case_dir=$(make_case away-archive-at-merge) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 1b1b1b1b1b1b1b1b1b1b1b1b1b1b1b1b1b1b1b1b + write_away_record "$case_dir" --grant task-x1 + mutate=$(away_change_script "$case_dir" archive-at-merge <<'SH' +"$CONTRACT" archive +SH + ) + + export FM_TEST_AWAY_MUTATE_AT_MERGE="$mutate" + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/71 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + unset FM_TEST_AWAY_MUTATE_AT_MERGE + + expect_code 0 "$rc" "away-archive-at-merge: the granted green merge should still land" + [ -s "$case_dir/away-mutate-rc" ] \ + || fail "away-archive-at-merge: the archive was never attempted inside the merge" + [ "$(cat "$case_dir/away-mutate-rc")" != 0 ] \ + || fail "away-archive-at-merge: the archive landed inside the merge's critical section" + assert_grep 'locked by live process' "$case_dir/away-mutate-output" \ + "away-archive-at-merge: the refused archive did not name the live holder" + assert_equals task-x1 "$(cat "$case_dir/away-grants-at-merge" 2>/dev/null || true)" \ + "away-archive-at-merge: the grant this merge read was not still standing at the forge call" + assert_grep "merge landed: task-x1 https://github.com/example/repo/pull/71 away-grant" \ + "$case_dir/state/.wake-queue" \ + "away-archive-at-merge: the landed merge was not recorded under the grant it read" + # The lock goes with the merge rather than leaking: the captain's return + # archives the record on its first try once the merge is done. + FM_HOME="$case_dir/home" FM_STATE_OVERRIDE="$case_dir/state" \ + "$ROOT/bin/fm-afk-contract.sh" archive >/dev/null \ + || fail "away-archive-at-merge: the record stayed locked after the merge" + + case_dir=$(make_case away-revoke-at-merge) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 2c2c2c2c2c2c2c2c2c2c2c2c2c2c2c2c2c2c2c2c + write_away_record "$case_dir" --grant task-x1 + mutate=$(away_change_script "$case_dir" revoke-at-merge <<'SH' +"$CONTRACT" propose --grant task-other +"$CONTRACT" confirm +SH + ) + export FM_TEST_AWAY_MUTATE_AT_MERGE="$mutate" + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/72 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + unset FM_TEST_AWAY_MUTATE_AT_MERGE + + expect_code 0 "$rc" "away-revoke-at-merge: the granted green merge should still land" + [ "$(cat "$case_dir/away-mutate-rc" 2>/dev/null || true)" != 0 ] \ + || fail "away-revoke-at-merge: the replacement landed inside the critical section" + assert_equals task-x1 "$(cat "$case_dir/away-grants-at-merge" 2>/dev/null || true)" \ + "away-revoke-at-merge: the grant was revoked inside the merge's critical section" + pass "no away-record archive or grant revocation lands between the authority read and the merge" +} + +# The same serialization from the other side. A revocation that wins the race +# lands BEFORE the in-lock authority read, and the merge then refuses: the lock +# decides an order, it never lets a stale grant through. +test_a_grant_revoked_before_the_merge_refuses_it() { + local case_dir rc + case_dir=$(make_case away-revoked-before-merge) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 3d3d3d3d3d3d3d3d3d3d3d3d3d3d3d3d3d3d3d3d + write_away_record "$case_dir" + mv "$case_dir/state/.afk-contract" "$case_dir/away-record-after-view" + write_away_record "$case_dir" --grant task-x1 + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/73 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "away-revoked-before-merge: a revoked grant must refuse" + assert_grep 'held for the captain return' "$case_dir/stderr" \ + "away-revoked-before-merge: refusal did not name hold-for-return" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "away-revoked-before-merge: gh pr merge ran on a revoked grant" + pass "a grant revoked before the merge's own authority read refuses the merge" +} + +# Fail closed. The lock is what makes the authority read and the merge one +# action, so a merge that cannot take it has no locked window to merge in and +# refuses - including on this attended case, where the record is absent and +# there is no grant to check at all. +test_merge_refuses_when_the_away_record_cannot_be_locked() { + local case_dir rc holder_pid i lock + case_dir=$(make_case away-lock-unavailable) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 4e4e4e4e4e4e4e4e4e4e4e4e4e4e4e4e4e4e4e4e + lock="$case_dir/state/.afk-contract.lock" + + FM_STATE_OVERRIDE="$case_dir/state" bash -c ' + . "$1" + fm_lock_acquire_wait "$2" || exit 10 + printf "ready\n" > "$3" + while [ ! -e "$4" ]; do sleep 0.05; done + fm_lock_release "$2" + ' _ "$ROOT/bin/fm-wake-lib.sh" "$lock" "$case_dir/holder.ready" "$case_dir/release-holder" & + holder_pid=$! + i=0 + while [ "$i" -lt 100 ] && [ ! -s "$case_dir/holder.ready" ]; do + sleep 0.05 + i=$((i + 1)) + done + [ -s "$case_dir/holder.ready" ] \ + || { kill "$holder_pid" 2>/dev/null || true; fail "away-lock-unavailable: the fixture never took the record lock"; } + + export FM_TEST_AFK_CONTRACT_LOCK_TIMEOUT=1 + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/74 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + unset FM_TEST_AFK_CONTRACT_LOCK_TIMEOUT + : > "$case_dir/release-holder" + wait "$holder_pid" || fail "away-lock-unavailable: the fixture holder did not release cleanly" + + expect_code 1 "$rc" "away-lock-unavailable: an unlockable away record must refuse the merge" + assert_grep 'could not be locked for the merge' "$case_dir/stderr" \ + "away-lock-unavailable: refusal did not name the lock it could not take" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "away-lock-unavailable: gh pr merge ran without the away-record lock" + pass "a merge that cannot lock the away record refuses instead of merging unlocked" +} + +test_allow_red_refused_on_gitlab() { + local case_dir rc + case_dir=$(make_gitlab_case gitlab-allow-red) + set +e + run_pr_merge "$case_dir" task-x1 "$MR_URL" --allow-red lint \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 2 "$rc" "gitlab-allow-red: --allow-red must not apply on GitLab" + assert_grep '--allow-red does not apply to GitLab' "$case_dir/stderr" \ + "gitlab-allow-red: refusal did not name GitLab" + [ ! -s "$case_dir/glab.log" ] || fail "gitlab-allow-red: glab ran despite --allow-red" + pass "fm-pr-merge refuses --allow-red on GitLab" +} + test_gitlab_head_override_args_refuse_before_recording test_github_still_forwards_sha_arg test_secondmate_merge_reports_upward_once @@ -3187,7 +3877,7 @@ test_backend_override_bypasses_unreadable_user_config # test, and the architecture document by three separate routes, and an # enumeration of the paths that get it wrong cannot keep finding them. # -# SEVEN routes, each separated from every other by ONE STATED SENTENCE naming +# SIX routes, each separated from every other by ONE STATED SENTENCE naming # WHERE it observes the landing. The count has been wrong three times under a # header rule asking for exactly that, most recently when the preflight's # carried-forward observation collapsed four of the GitHub routes onto one read: @@ -3222,9 +3912,8 @@ test_backend_override_bypasses_unreadable_user_config # after it fails # github|command-error the merge command FAILS and the read that # follows the failure confirms it landed anyway -# github|degraded-no-gh gh is absent entirely, so the gh-axi view -# answers both reads and its POST-merge answer -# is what proves the merge +# Missing gh now refuses before any observation; the required-tool tests own +# that refusal rather than treating it as a reachable landed-outcome route. # github|degraded-gh-failed gh is present and consulted, its read fails, # and the gh-axi fallback's POST-merge answer # proves the merge @@ -3248,7 +3937,6 @@ test_every_landed_observation_reaches_outcome_reporting() { github\|post-mutation \ github\|preflight-landed \ github\|command-error \ - github\|degraded-no-gh \ github\|degraded-gh-failed \ gitlab\|post-mutation \ gitlab\|command-error; do @@ -3280,6 +3968,7 @@ test_every_landed_observation_reaches_outcome_reporting() { # reporting, so this route fails outright if that observation is # discarded rather than carried. write_github_outcome "$case_dir" MERGED true false main + : > "$case_dir/github-outcome.initially-merged" add_gh_mock_outcome_read_fails_from "$case_dir" \ aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa 2 cat >"$case_dir/fakebin/gh-axi" <<'SH' @@ -3299,31 +3988,12 @@ SH # follows the failure is what observes it. add_gh_axi_mock_open_until_merged "$case_dir" 1 ;; - degraded-no-gh) - # gh absent entirely: only the gh-axi view can prove the outcome, and - # only its POST-merge answer can, because its pre-merge answer is open. - # Deleting this case's own mock is NOT enough. run_pr_merge prepends - # fakebin to the INHERITED PATH, so on any host that ships gh the real - # binary still resolves, this route quietly becomes degraded-gh-failed - # against the live forge, and the invariant's seven routes are six. The - # search path is rebuilt without gh so no gh can resolve whatever the - # host has installed. - add_gh_axi_mock_open_until_merged "$case_dir" 0 - rm -f "$case_dir/fakebin/gh" - run_path="$case_dir/path-without-gh" - mirror_path_without "$run_path" gh "$case_dir/fakebin" - ;; degraded-gh-failed) # gh present but its read fails; the gh-axi fallback proves the merge. # It logs before failing, which is how this route proves it is the one # where gh WAS consulted rather than the one where gh does not exist. add_gh_axi_mock_open_until_merged "$case_dir" 0 - cat >"$case_dir/fakebin/gh" <<'SH' -#!/usr/bin/env bash -printf '%s\n' "$*" >> "$FM_TEST_GH_LOG" -exit 1 -SH - chmod +x "$case_dir/fakebin/gh" + : > "$case_dir/github-graphql-fail" ;; esac else @@ -3365,16 +4035,6 @@ SH fi fi - # run_pr_merge PREPENDS fakebin to whatever PATH it inherits, so the search - # path this case really runs on is the one asserted here rather than the one - # built above. Deleting the case's own mock is not isolation: the host's own - # gh stays resolvable and silently turns this route into degraded-gh-failed, - # against the live forge, with the invariant's six routes quietly five. - if [ "$route" = degraded-no-gh ]; then - ! PATH="$case_dir/fakebin:${run_path:-$PATH}" command -v gh >/dev/null 2>&1 \ - || fail "landed-invariant github/degraded-no-gh: the search path this case runs on still resolves gh" - fi - set +e if [ -n "$run_path" ]; then PATH="$run_path" FM_TEST_HOME="$case_dir/home" \ @@ -3410,7 +4070,7 @@ SH # Each route must stay the route it is named for AND reach its own # observation point. Exiting zero with the outcome recorded says neither: it - # is the one thing all seven have in common. + # is the one thing all six have in common. case "$provider|$route" in github\|post-mutation) # Two reads, the preflight seeing an open pull request and the post-merge @@ -3424,10 +4084,10 @@ SH "landed-invariant github/post-mutation: the readback's landing was not reported" ;; github\|preflight-landed) - # The merge is still ATTEMPTED on this route; the preflight short-circuits - # the outcome READ, not the mutation. - assert_grep 'pr merge' "$case_dir/gh-axi.log" \ - "landed-invariant github/preflight-landed: the merge was skipped rather than attempted" + # An already-landed request needs neither another mutation nor another + # outcome read; both can introduce false failures after success. + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "landed-invariant github/preflight-landed: an already-landed request was mutated again" [ "$gh_reads" = 1 ] \ || fail "landed-invariant github/preflight-landed: a landing the preflight already observed was read back again (gh answered $gh_reads outcome reads)" [ "$axi_views" = 0 ] \ @@ -3439,12 +4099,6 @@ SH [ "$axi_views" = 0 ] \ || fail "landed-invariant github/command-error: the degraded reader answered on a route where gh works" ;; - github\|degraded-no-gh) - [ ! -s "$case_dir/gh.log" ] \ - || fail "landed-invariant github/degraded-no-gh: gh was consulted on a route that must have none" - [ "$axi_views" = 2 ] \ - || fail "landed-invariant github/degraded-no-gh: the gh-axi view did not prove the merge AFTER it landed (it answered $axi_views views)" - ;; github\|degraded-gh-failed) [ -s "$case_dir/gh.log" ] \ || fail "landed-invariant github/degraded-gh-failed: gh was never consulted, so this is the no-gh route" @@ -3482,14 +4136,14 @@ SH esac done - # THE MATRIX REFUSES TO COLLAPSE. Seven routes that reached seven different - # observation points leave seven different witnesses; two routes that ended up + # THE MATRIX REFUSES TO COLLAPSE. Six routes that reached six different + # observation points leave six different witnesses; two routes that ended up # in the same place leave the same witness twice and this fails, which is the # check the route count has been missing every time it was wrong. total=$(wc -l <"$witness_file" | tr -d '[:space:]') distinct=$(sort -u "$witness_file" | wc -l | tr -d '[:space:]') - [ "$total" = 7 ] \ - || fail "landed-invariant: $total routes ran, and the sentences above name seven" + [ "$total" = 6 ] \ + || fail "landed-invariant: $total routes ran, and the sentences above name six" [ "$total" = "$distinct" ] || { sort "$witness_file" >&2 fail "landed-invariant: only $distinct of $total routes reached a distinct observation point, so the matrix is smaller than it claims" @@ -3565,6 +4219,7 @@ test_merged_non_default_target_is_refused() { add_gh_mocks "$case_dir" bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb # Already merged, and merged into a release branch while the default is main. write_github_outcome "$case_dir" MERGED true false 'release/2026' main + : > "$case_dir/github-outcome.initially-merged" : >"$case_dir/gh-axi.log" set +e @@ -3616,9 +4271,9 @@ esac exit 0 SH chmod +x "$case_dir/fakebin/gh-axi" - rm -f "$case_dir/fakebin/gh" + : > "$case_dir/github-graphql-fail" ghless_path="$case_dir/path-without-gh" - mirror_path_without "$ghless_path" gh "$case_dir/fakebin" + ghless_path="$case_dir/fakebin:$PATH" : >"$case_dir/gh-axi.log" set +e @@ -3631,7 +4286,7 @@ SH "target-refusal-unestablished: an unestablished target must be refused" assert_grep 'target branch could not be' "$case_dir/stderr" \ "target-refusal-unestablished: the refusal did not say the target was never established" - grep -q 'pr merge' "$case_dir/gh-axi.log" \ + grep -q 'pr merge' "$case_dir/gh.log" \ && fail "target-refusal-unestablished: a merge ran without an established target" pass "fm-pr-merge refuses when the merge target cannot be established" } @@ -3642,9 +4297,9 @@ test_degraded_path_reads_the_real_target() { case_dir=$(make_case target-refusal-degraded) mkdir -p "$case_dir/wt" add_gh_mocks "$case_dir" dddddddddddddddddddddddddddddddddddddddd - rm -f "$case_dir/fakebin/gh" + : > "$case_dir/github-graphql-fail" ghless_path="$case_dir/path-without-gh" - mirror_path_without "$ghless_path" gh "$case_dir/fakebin" + ghless_path="$case_dir/fakebin:$PATH" # The shared mock emits the REAL gh-axi shape for this query: one TOON field # per line, which is what the reader is written against. The parse has been # vacuous on this path once already - an earlier reader took the last ": " in @@ -3711,11 +4366,11 @@ printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" case_dir=$(dirname "$FM_TEST_GH_AXI_LOG") case "${1:-} ${2:-}" in "pr merge") - : > "$case_dir/gh-axi-merge-attempted" + : > "$FM_TEST_GH_OUTCOME.merge-called" printf 'merged:\n number: %s\n status: ok\n' "${3:-}" ;; "pr view") - if [ -e "$case_dir/gh-axi-merge-attempted" ]; then + if [ -e "$FM_TEST_GH_OUTCOME.merge-called" ]; then printf 'pull_request:\n number: %s\n state: merged\n' "$3" else printf 'pull_request:\n number: %s\n state: open\n' "$3" @@ -3752,12 +4407,12 @@ SH "target-gh-failed-refuses: the refusal did not name the default branch" assert_no_grep 'target branch could not be read' "$case_dir/stderr" \ "target-gh-failed-refuses: a target the degraded reader supplied was called unreadable" - assert_no_grep 'pr merge' "$case_dir/gh-axi.log" \ + assert_no_grep 'pr merge' "$case_dir/gh.log" \ "target-gh-failed-refuses: a refused target still reached the forge" else expect_code 0 "$rc" \ "target-gh-failed-permits: a default target must merge even when gh's read fails" - assert_grep 'pr merge' "$case_dir/gh-axi.log" \ + assert_grep 'pr merge' "$case_dir/gh.log" \ "target-gh-failed-permits: an open pull request on the default branch never reached the merge" assert_grep "verified: $url is merged" "$case_dir/stdout" \ "target-gh-failed-permits: the landed merge was not reported" @@ -3801,11 +4456,11 @@ printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" case_dir=$(dirname "$FM_TEST_GH_AXI_LOG") case "${1:-} ${2:-}" in "pr merge") - : > "$case_dir/gh-axi-merge-attempted" + : > "$FM_TEST_GH_OUTCOME.merge-called" printf 'merged:\n number: %s\n status: ok\n' "${3:-}" ;; "pr view") - if [ -e "$case_dir/gh-axi-merge-attempted" ]; then + if [ -e "$FM_TEST_GH_OUTCOME.merge-called" ]; then printf 'pull_request:\n number: %s\n state: merged\n' "$3" else printf 'pull_request:\n number: %s\n state: open\n' "$3" @@ -3838,12 +4493,12 @@ SH "target-quoted-refuses: the refusal named a branch the encoder's quotes invented" assert_grep 'current default branch main' "$case_dir/stderr" \ "target-quoted-refuses: the default branch was not read back cleanly" - assert_no_grep 'pr merge' "$case_dir/gh-axi.log" \ + assert_no_grep 'pr merge' "$case_dir/gh.log" \ "target-quoted-refuses: a refused target still reached the forge" else expect_code 0 "$rc" \ "target-quoted-permits: a quoted name equal to the default must still merge" - assert_grep 'pr merge' "$case_dir/gh-axi.log" \ + assert_grep 'pr merge' "$case_dir/gh.log" \ "target-quoted-permits: a pull request on the default branch never reached the merge" assert_grep "verified: $url is merged" "$case_dir/stdout" \ "target-quoted-permits: the landed merge was not reported" @@ -3938,7 +4593,7 @@ SH "default-tip-moved: the refusal did not name the tip this run judged" assert_grep "to tip $MOVED_DEFAULT_TIP" "$case_dir/stderr" \ "default-tip-moved: the refusal did not name the tip the branch moved to" - assert_no_grep 'pr merge' "$case_dir/gh-axi.log" \ + assert_no_grep 'pr merge' "$case_dir/gh.log" \ "default-tip-moved: a merge ran against a base this run never judged" assert_absent "$case_dir/state/.wake-queue" \ "default-tip-moved: a refused merge was recorded as a landed outcome" @@ -3946,7 +4601,7 @@ SH steady) expect_code 0 "$rc" \ "default-tip-steady: a base that did not move must still merge" - assert_grep 'pr merge' "$case_dir/gh-axi.log" \ + assert_grep 'pr merge' "$case_dir/gh.log" \ "default-tip-steady: an unmoved base was refused, so the guard refuses always" assert_grep "verified: $url is merged" "$case_dir/stdout" \ "default-tip-steady: the landed merge was not reported" @@ -3958,7 +4613,7 @@ SH "default-tip-unreadable: a tip that could not be re-read must refuse" assert_grep 'could not be read' "$case_dir/stderr" \ "default-tip-unreadable: the refusal did not say the tip could not be read" - assert_no_grep 'pr merge' "$case_dir/gh-axi.log" \ + assert_no_grep 'pr merge' "$case_dir/gh.log" \ "default-tip-unreadable: an unreadable tip was treated as an unmoved one" ;; esac @@ -4108,18 +4763,11 @@ test_gitlab_auto_rebase_guard_refuses_and_permits # anything else must be reported as an outcome that could not be READ rather than # as a concrete not-merged verdict. # -# BOTH DEGRADED ROUTES ARE RUN, and that is the point of the loop rather than a -# second case for tidiness. The evidence is the same gh-axi view whether gh is -# ABSENT or PRESENT AND BROKEN, so a verdict that differs between them is a -# verdict about which tool happens to be installed. The seam used to be consulted -# on the gh-failed route only, and the gh-absent route answered the same evidence -# with a concrete state=open verdict instead. -# -# Without this case, dropping the merged proof on the post-merge side leaves the -# suite green and the seam collapses back into the single rule it replaced. +# The executable gh is required for head-bound mutation; its failed outcome +# read can still fall back to gh-axi, whose open state cannot prove a merge. test_degraded_view_cannot_answer_the_post_merge_question() { local case_dir rc url route run_path number=95 - for route in gh-failed no-gh; do + route=gh-failed number=$((number + 1)) url="https://github.com/example/repo/pull/$number" case_dir=$(make_case "degraded-post-merge-unanswerable-$route") @@ -4144,21 +4792,7 @@ esac exit 0 SH chmod +x "$case_dir/fakebin/gh-axi" - if [ "$route" = gh-failed ]; then - # gh answers the pre-merge target read and then fails, so the post-merge - # read falls back to a gh-axi view that reports the request still OPEN. - add_gh_mock_outcome_read_fails_from "$case_dir" ffffffffffffffffffffffffffffffffffffffff 2 - else - # gh is absent entirely, so the same gh-axi view answers both reads. - # Removing this case's own mock is not enough: run_pr_merge prepends - # fakebin to the INHERITED PATH, so on a host that ships gh the real binary - # resolves and this route quietly becomes the other one. - rm -f "$case_dir/fakebin/gh" - run_path="$case_dir/path-without-gh" - mirror_path_without "$run_path" gh "$case_dir/fakebin" - ! PATH="$case_dir/fakebin:$run_path" command -v gh >/dev/null 2>&1 \ - || fail "degraded-post-merge-unanswerable/no-gh: the search path this case runs on still resolves gh" - fi + add_gh_mock_outcome_read_fails_from "$case_dir" ffffffffffffffffffffffffffffffffffffffff 2 : >"$case_dir/gh-axi.log" set +e @@ -4182,10 +4816,135 @@ SH "degraded-post-merge-unanswerable/$route: an outcome nothing proved was stated as a not-merged verdict" assert_no_grep 'verified: ' "$case_dir/stdout" \ "degraded-post-merge-unanswerable/$route: an unproved merge was reported as verified" - done - pass "neither degraded route answers the post-merge outcome question without a proved merge" + pass "a degraded outcome read cannot establish a merge without proof" } test_degraded_view_cannot_answer_the_post_merge_question +# Real Git histories prove the exception's parentage; only forge transport is +# mocked. Every refusal must leave the forge merge command uncalled. +test_reviewed_upstream_sync_policy_exception() { + local case_dir scenario wt base target head id meta rc saved_tip url method filter + saved_tip=$DEFAULT_TIP + for scenario in attended away fix-forward absent-review stale-review wrong-target duplicate-review \ + wrong-mode wrong-task wrong-repo wrong-push wrong-upstream side-endpoint wrong-branch wrong-live-branch \ + wrong-live-repo wrong-local-head moved-base squash green-squash pending cancelled status-context lint-red lint-pending \ + extra-merge no-grant; do + case_dir=$(make_case "sync-policy-$scenario") + wt="$case_dir/wt" + id=fm-upstream-sync-2026-09-14-tail + url=https://github.com/HelloWorldSungin/firstmate/pull/96 + git init -q "$wt" + git -C "$wt" commit -qm root --allow-empty + git -C "$wt" checkout -qb upstream + git -C "$wt" commit -qm upstream --allow-empty + target=$(git -C "$wt" rev-parse HEAD) + git -C "$wt" checkout -qb "fm/$id" HEAD~1 + git -C "$wt" commit -qm fork --allow-empty + base=$(git -C "$wt" rev-parse HEAD) + git -C "$wt" merge -q --no-ff -m sync "$target" + if [ "$scenario" = extra-merge ]; then + git -C "$wt" checkout -qb extra "$base" + git -C "$wt" commit -qm extra --allow-empty + git -C "$wt" checkout -q "fm/$id" + git -C "$wt" merge -q --no-ff -m extra extra + fi + [ "$scenario" != fix-forward ] || git -C "$wt" commit -qm fix --allow-empty + head=$(git -C "$wt" rev-parse HEAD) + git -C "$wt" remote add origin https://github.com/HelloWorldSungin/firstmate.git + git -C "$wt" remote add upstream https://github.com/kunchenguid/firstmate.git + git -C "$wt" update-ref refs/remotes/upstream/main "$target" + if [ "$scenario" = side-endpoint ]; then + git -C "$wt" checkout -qb upstream-main "$target~1" + git -C "$wt" commit -qm mainline --allow-empty + git -C "$wt" merge -q --no-ff -m side-target "$target" + git -C "$wt" update-ref refs/remotes/upstream/main HEAD + git -C "$wt" checkout -q "fm/$id" + fi + DEFAULT_TIP=$base + add_gh_mocks "$case_dir" "$head" + write_github_red_json "$case_dir" "$head" 'PR must be raised via no-mistakes' + jq --arg branch "fm/$id" '. + {headRefName:$branch,headRepository:{nameWithOwner:"HelloWorldSungin/firstmate"}}' \ + "$case_dir/github-view.json" > "$case_dir/view.tmp" + mv "$case_dir/view.tmp" "$case_dir/github-view.json" + [ "$scenario" != wrong-task ] || id=ordinary-task + meta="$case_dir/state/$id.meta" + mv "$case_dir/state/task-x1.meta" "$meta" + sed 's/^mode=.*/mode=direct-PR/' "$meta" > "$case_dir/meta.tmp" + mv "$case_dir/meta.tmp" "$meta" + printf 'upstream_sync_review=%s:%s:%s\n' "$base" "$target" "$head" >> "$meta" + method=--merge + case "$scenario" in + away|no-grant) write_away_record "$case_dir" --grant "$id" + [ "$scenario" != no-grant ] || write_away_record "$case_dir" ;; + absent-review) sed '/^upstream_sync_review=/d' "$meta" > "$case_dir/meta.tmp"; mv "$case_dir/meta.tmp" "$meta" ;; + stale-review|wrong-target) + sed '/^upstream_sync_review=/d' "$meta" > "$case_dir/meta.tmp"; mv "$case_dir/meta.tmp" "$meta" + if [ "$scenario" = stale-review ]; then + printf 'upstream_sync_review=%s:%s:%s\n' "$base" "$target" "$base" >> "$meta" + else + printf 'upstream_sync_review=%s:%s:%s\n' "$base" "$base" "$head" >> "$meta" + fi ;; + duplicate-review) printf 'upstream_sync_review=%s:%s:%s\n' "$base" "$target" "$head" >> "$meta" ;; + wrong-mode) printf 'mode=no-mistakes\n' >> "$meta" ;; + wrong-repo) url=https://github.com/example/repo/pull/96 ;; + wrong-push) git -C "$wt" remote set-url --push origin https://github.com/kunchenguid/firstmate.git ;; + wrong-upstream) git -C "$wt" remote set-url upstream https://github.com/example/repo.git ;; + wrong-branch) git -C "$wt" checkout -qb unrelated ;; + wrong-local-head) git -C "$wt" commit -qm unreviewed --allow-empty ;; + moved-base) DEFAULT_TIP=$target ;; + squash|green-squash) method=--squash ;; + esac + case "$scenario" in + green-squash) filter='.statusCheckRollup[0].conclusion="SUCCESS"' ;; + pending) filter='.statusCheckRollup[0].status="IN_PROGRESS"' ;; + cancelled) filter='.statusCheckRollup[0].conclusion="CANCELLED"' ;; + status-context) filter='.statusCheckRollup += [{__typename:"StatusContext",context:"PR must be raised via no-mistakes",state:"PENDING"}]' ;; + lint-red) filter='.statusCheckRollup += [{__typename:"CheckRun",name:"lint",status:"COMPLETED",conclusion:"FAILURE"}]' ;; + lint-pending) filter='.statusCheckRollup += [{__typename:"CheckRun",name:"lint",status:"QUEUED",conclusion:null}]' ;; + wrong-live-branch) filter='.headRefName="unrelated"' ;; + wrong-live-repo) filter='.headRepository.nameWithOwner="example/repo"' ;; + *) filter='.' ;; + esac + jq "$filter" "$case_dir/github-view.json" > "$case_dir/view.tmp" + mv "$case_dir/view.tmp" "$case_dir/github-view.json" + rc=0 + run_pr_merge "$case_dir" "$id" "$url" -- "$method" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" || rc=$? + case "$scenario" in + attended|away|fix-forward) + expect_code 0 "$rc" "sync-policy-$scenario: verified round should merge: $(cat "$case_dir/stderr")" + assert_logged_gh_merge "$case_dir" 96 HelloWorldSungin/firstmate --merge ;; + *) + [ "$rc" -ne 0 ] || fail "sync-policy-$scenario: invalid proof or check was accepted" + assert_no_grep 'pr merge' "$case_dir/gh.log" "sync-policy-$scenario: forge merge ran" ;; + esac + done + DEFAULT_TIP=$saved_tip + pass "fm-pr-merge limits the sync exception to reviewed graphs and completed policy failures" +} + +test_reviewed_upstream_sync_policy_exception + +test_github_red_checks_refuse_and_allow_red_waives_named +test_superseded_failed_check_run_no_longer_refuses +test_check_runs_never_supersede_status_contexts +test_current_failed_check_run_still_refuses +test_late_finishing_old_success_does_not_hide_current_failure +test_late_finishing_old_cancellation_is_superseded +test_unfinished_rerun_keeps_a_check_red +test_supersession_never_crosses_check_names +test_undated_runs_never_supersede +test_allow_red_still_waives_only_the_current_failure +test_allow_red_is_refused_while_away +test_allow_red_requires_one_separate_name +test_away_grant_and_yolo_and_hold_for_return +test_away_posture_refuses_asynchronous_merge_paths +test_away_grant_does_not_bypass_red_or_identity +test_unreadable_away_record_refuses_merge +test_away_record_cannot_change_between_the_authority_read_and_the_merge +test_a_grant_revoked_before_the_merge_refuses_it +test_merge_refuses_when_the_away_record_cannot_be_locked +test_allow_red_refused_on_gitlab + printf '\nall fm-pr-merge tests passed\n' diff --git a/tests/fm-procevent-when.test.sh b/tests/fm-procevent-when.test.sh index 396c48df8d9..36ddc955f78 100755 --- a/tests/fm-procevent-when.test.sh +++ b/tests/fm-procevent-when.test.sh @@ -17,6 +17,8 @@ set -u ROOT=$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd) TMP_ROOT=$(fm_test_tmproot fm-procevent-when-tests) export FM_PROCEVENT_CLAIM_ROOT="$TMP_ROOT/claims" +# shellcheck source=bin/fm-pr-lib.sh +. "$ROOT/bin/fm-pr-lib.sh" pe() { FM_HOME="$1" "$ROOT/bin/fm-procevent.sh" "${@:2}"; } when() { FM_HOME="$1" "$ROOT/bin/fm-procevent-when.sh" "${@:2}"; } @@ -390,4 +392,258 @@ assert_absent "$ACTION_TAMPER_LOG" "the mutated action was not executed" assert_absent "$H/state/when/when-action-tamper.fired" "no fire was claimed for mutated action bytes" pass "mutated action bytes are refused before claiming the fire" +# --- rebind-all refreshes a watch's action hash after a self-update ---------- +# A self-update fast-forwards bin/ in place, changing an in-repo action +# executable's bytes with no tampering involved. Without rebind-all the next +# fire is refused as "does not match the registered trust binding" (see the +# mutated-action-bytes case above); rebind-all exists to follow that update +# and republish a trust binding that matches the new bytes, but only for an +# action living under the simulated repo root, never for one outside it. +H="$TMP_ROOT/h-rebind"; new_home "$H" +REPO_ROOT="$TMP_ROOT/rebind-repo" +mkdir -p "$REPO_ROOT/bin" +IN_REPO_ACT="$REPO_ROOT/bin/act.sh" +cat > "$IN_REPO_ACT" <<'SH' +#!/usr/bin/env bash +log=$1 +echo v1 >> "$log" +SH +chmod +x "$IN_REPO_ACT" +OUT_OF_REPO_ACT="$TMP_ROOT/rebind-outside-act.sh" +cat > "$OUT_OF_REPO_ACT" <<'SH' +#!/usr/bin/env bash +log=$1 +echo v1 >> "$log" +SH +chmod +x "$OUT_OF_REPO_ACT" +when_ro() { FM_HOME="$1" FM_ROOT_OVERRIDE="$REPO_ROOT" "$ROOT/bin/fm-procevent-when.sh" "${@:2}"; } + +when_ro "$H" arm rebind-in-repo --interval 0.1 --stable 1 \ + --condition true --action "$IN_REPO_ACT" "$TMP_ROOT/rebind-in-repo.log" >/dev/null +when_ro "$H" arm rebind-out-of-repo --interval 0.1 --stable 1 \ + --condition true --action "$OUT_OF_REPO_ACT" "$TMP_ROOT/rebind-out-of-repo.log" >/dev/null + +SPEC_IN="$H/state/when/when-rebind-in-repo.spec" +TRUST_IN="$H/state/when/when-rebind-in-repo.trust" +TRUST_OUT="$H/state/when/when-rebind-out-of-repo.trust" + +OUT=$(when_ro "$H" rebind-all) || fail "rebind-all failed with nothing to rebind: $OUT" +assert_contains "$OUT" "0 rebound, 2 unchanged or out of scope, 0 failed" \ + "rebind-all should be a no-op before any action bytes change" + +# Simulate the self-update: rewrite both action scripts' bytes in place. +OLD_IN_REPO_SHA=$(fm_pr_sha256 "$IN_REPO_ACT") +cat > "$IN_REPO_ACT" <<'SH' +#!/usr/bin/env bash +log=$1 +echo v2 >> "$log" +echo "action ran v2 against $log" +SH +chmod +x "$IN_REPO_ACT" +NEW_HASH=$(fm_pr_sha256 "$IN_REPO_ACT") +[ "$OLD_IN_REPO_SHA" != "$NEW_HASH" ] || fail "test fixture error: mutation did not change the in-repo action's hash" +printf "#!/usr/bin/env bash\necho v2 >> \"\$1\"\n" > "$OUT_OF_REPO_ACT" +chmod +x "$OUT_OF_REPO_ACT" + +OLD_TRUST_OUT=$(cat "$TRUST_OUT") +OUT=$(when_ro "$H" rebind-all) || fail "rebind-all reported a failure: $OUT" +assert_contains "$OUT" "rebound: when-rebind-in-repo" "the in-repo watch was rebound" +assert_contains "$OUT" "1 rebound, 1 unchanged or out of scope, 0 failed" \ + "exactly the in-repo watch should rebind; the out-of-repo one stays out of scope" +[ "$(cat "$TRUST_OUT")" = "$OLD_TRUST_OUT" ] \ + || fail "rebind-all must never touch a watch whose action lives outside FM_ROOT" + +grep -qx "action_sha256=$NEW_HASH" "$SPEC_IN" \ + || fail "rebind-all did not record the action's current bytes in the spec" +SPEC_HASH=$(fm_pr_sha256 "$SPEC_IN") +TRUST_WANT=$(sed -n '2p' "$TRUST_IN") +[ "$SPEC_HASH" = "$TRUST_WANT" ] \ + || fail "the republished spec must still match its own trust binding" + +# The watch actually works again: a fresh run fires cleanly against the new +# bytes instead of being rejected. +pe "$H" reconcile >/dev/null +wait_for_result "$H" "when-rebind-in-repo" || fail "the rebound watch captured no outcome" +RESULT=$(first_result "$H" "when-rebind-in-repo") +assert_grep 'status: fired' "$RESULT" "the rebound watch fires instead of being rejected" +assert_grep 'action ran v2 against' "$RESULT" "the fired action ran the new bytes, not a stale copy" +pass "rebind-all refreshes an in-repo watch's trust binding after a self-update and leaves an out-of-repo one alone" + +# --- rebind-all matches an action reached through a symlinked FM_ROOT ------- +H="$TMP_ROOT/h-rebind-symlink"; new_home "$H" +REPO_REAL="$TMP_ROOT/rebind-symlink-real" +mkdir -p "$REPO_REAL/bin" +REPO_LINK="$TMP_ROOT/rebind-symlink-link" +ln -s "$REPO_REAL" "$REPO_LINK" +SYMLINK_ACT="$REPO_LINK/bin/act.sh" +cat > "$REPO_REAL/bin/act.sh" <<'SH' +#!/usr/bin/env bash +log=$1 +echo v1 >> "$log" +SH +chmod +x "$REPO_REAL/bin/act.sh" +when_symlink_ro() { FM_HOME="$1" FM_ROOT_OVERRIDE="$REPO_LINK" "$ROOT/bin/fm-procevent-when.sh" "${@:2}"; } + +when_symlink_ro "$H" arm rebind-symlink --interval 0.1 --stable 1 \ + --condition true --action "$SYMLINK_ACT" "$TMP_ROOT/rebind-symlink.log" >/dev/null + +cat > "$REPO_REAL/bin/act.sh" <<'SH' +#!/usr/bin/env bash +log=$1 +echo v2 >> "$log" +SH +chmod +x "$REPO_REAL/bin/act.sh" + +OUT=$(when_symlink_ro "$H" rebind-all) || fail "rebind-all reported a failure through a symlinked FM_ROOT: $OUT" +assert_contains "$OUT" "rebound: when-rebind-symlink" \ + "rebind-all must rebind an action reached through a symlinked FM_ROOT, not report it out of scope" +pass "rebind-all matches FM_ROOT through a symlinked checkout path" + +# --- rebind-all reaches a watch whose poller is already running ------------- +# The self-update race the fire-time revalidation targets: `run` calls +# spec_load once before entering its poll loop and caches the action hash in +# memory for the rest of its life. If the self-update (and its rebind-all) +# land while that poll loop is still running, only rewriting the on-disk spec +# and trust is not enough - the fire-time check must re-read the binding from +# disk, or the still-running poller compares against its stale in-memory hash +# and rejects a perfectly legitimate post-update fire. +H="$TMP_ROOT/h-live-rebind"; new_home "$H" +REPO_ROOT="$TMP_ROOT/live-rebind-repo" +mkdir -p "$REPO_ROOT/bin" +LIVE_ACT="$REPO_ROOT/bin/act.sh" +cat > "$LIVE_ACT" <<'SH' +#!/usr/bin/env bash +echo v1 >> "$1" +echo "v1 ran against $1" +SH +chmod +x "$LIVE_ACT" +LIVE_TRIGGER="$TMP_ROOT/live-rebind-trigger" +LIVE_COUNTER="$TMP_ROOT/live-rebind-count" +LIVE_LOG="$TMP_ROOT/live-rebind.log" +when_live_ro() { FM_HOME="$1" FM_ROOT_OVERRIDE="$REPO_ROOT" "$ROOT/bin/fm-procevent-when.sh" "${@:2}"; } + +when_live_ro "$H" arm live-rebind --interval 0.1 --stable 1 \ + --condition "$COND" "$LIVE_TRIGGER" "$LIVE_COUNTER" \ + --action "$LIVE_ACT" "$LIVE_LOG" >/dev/null + +# Start the poller now, before the simulated self-update, so its one-time +# spec_load caches the pre-update (v1) action hash in memory. +pe "$H" reconcile >/dev/null +wait_for_file "$LIVE_COUNTER" || fail "the live-rebind poller never evaluated its condition" + +# Simulate the self-update while that poller is still running: rewrite the +# action's bytes in place, then rebind-all republishes the on-disk trust +# binding to match. The already-running poller's in-memory hash is untouched. +cat > "$LIVE_ACT" <<'SH' +#!/usr/bin/env bash +echo v2 >> "$1" +echo "v2 ran against $1" +SH +chmod +x "$LIVE_ACT" +OUT=$(when_live_ro "$H" rebind-all) || fail "rebind-all reported a failure during a live poll: $OUT" +assert_contains "$OUT" "rebound: when-live-rebind" "the live watch's trust binding was rebound on disk" + +# Let the condition go true; the still-running poller must pick up the fresh +# binding at fire time instead of comparing against its stale cached hash. +: > "$LIVE_TRIGGER" +wait_for_result "$H" when-live-rebind || fail "the live poller never captured an outcome after rebind-all" +RESULT=$(first_result "$H" when-live-rebind) +assert_grep 'status: fired' "$RESULT" \ + "a watch whose poller was already running when rebind-all ran must still fire, not be rejected as stale" +assert_grep 'v2 ran against' "$RESULT" "the fired action ran the post-update bytes, not the ones cached at poll start" +pass "rebind-all reaches a watch whose run process was already polling when the self-update landed" + +# --- the fire-time reload never observes rebind_one's publish mid-rename ---- +# publish_spec is not an atomic swap: it renames the new spec into place, then +# separately renames the new trust into place. A `run` process reloading the +# binding at fire time must serialize against that window instead of reading +# a spec already rebound to v2 next to a trust record still bound to v1 - the +# exact torn combination that would otherwise report the rebind itself as a +# trust violation. This test builds that torn state under a held per-sid lock +# (the same lock rebind_one takes) so the reload's timing is deterministic, +# not a race that only sometimes reproduces. +H="$TMP_ROOT/h-torn-race"; new_home "$H" +TORN_ACT="$TMP_ROOT/torn-act.sh" +cat > "$TORN_ACT" <<'SH' +#!/usr/bin/env bash +echo v1 >> "$1" +echo "v1 ran against $1" +SH +chmod +x "$TORN_ACT" +TORN_TRIGGER="$TMP_ROOT/torn-race-trigger" +TORN_COUNTER="$TMP_ROOT/torn-race-count" +TORN_LOG="$TMP_ROOT/torn-race.log" +when "$H" arm torn-race --interval 0.05 --stable 1 \ + --condition "$COND" "$TORN_TRIGGER" "$TORN_COUNTER" \ + --action "$TORN_ACT" "$TORN_LOG" >/dev/null +SID=$(when "$H" source-id torn-race) +SPEC_TORN="$H/state/when/$SID.spec" +TRUST_TORN="$H/state/when/$SID.trust" + +# Start the poller now, with the condition still false, so reconcile's own +# brief use of this same per-sid lock (to claim and launch the source) is +# already done and released well before the holder below ever takes it. +pe "$H" reconcile >/dev/null +wait_for_file "$TORN_COUNTER" || fail "the torn-race poller never evaluated its condition" + +# Simulate the self-update, then build the rebound (v2) spec+trust pair ahead +# of time exactly as publish_spec would (same fields, only action_sha256 +# differs), so the background holder below only performs the two renames. +cat > "$TORN_ACT" <<'SH' +#!/usr/bin/env bash +echo v2 >> "$1" +echo "v2 ran against $1" +SH +chmod +x "$TORN_ACT" +NEW_HASH=$(fm_pr_sha256 "$TORN_ACT") +NEW_SPEC="$TMP_ROOT/torn-race-new.spec" +sed "s/^action_sha256=.*/action_sha256=$NEW_HASH/" "$SPEC_TORN" > "$NEW_SPEC" +NEW_SPEC_HASH=$(fm_pr_sha256 "$NEW_SPEC") +NEW_TRUST="$TMP_ROOT/torn-race-new.trust" +printf 'fm-when-trust-v1\n%s\n' "$NEW_SPEC_HASH" > "$NEW_TRUST" +chmod 0600 "$NEW_SPEC" "$NEW_TRUST" + +TORN_READY="$TMP_ROOT/torn-ready" +TORN_RELEASE="$TMP_ROOT/torn-release" +rm -f "$TORN_READY" "$TORN_RELEASE" +parent=$$ +FM_HOME="$TMP_ROOT/torn-race-lock-helper-home" bash -c ' + . "$1/bin/fm-pr-lib.sh" + . "$1/bin/fm-wake-lib.sh" + . "$1/bin/fm-procevent-lib.sh" + fm_procevent_source_lock_acquire "$2" || exit 1 + trap "fm_procevent_source_lock_release \"$2\"" EXIT + mv -f -- "$3" "$5" + printf "ready\n" > "$6" + while [ ! -e "$7" ]; do + kill -0 "$8" 2>/dev/null || exit 0 + sleep 0.02 + done + mv -f -- "$4" "$9" +' _ "$ROOT" "$SID" "$NEW_SPEC" "$NEW_TRUST" "$SPEC_TORN" "$TORN_READY" "$TORN_RELEASE" "$parent" "$TRUST_TORN" & +HOLDER_PID=$! + +wait_for_file "$TORN_READY" || fail "the torn-write holder never installed the rebound spec" +grep -qx "action_sha256=$NEW_HASH" "$SPEC_TORN" \ + || fail "test fixture error: the torn window did not actually install the rebound spec" +[ "$(sed -n '2p' "$TRUST_TORN")" != "$NEW_SPEC_HASH" ] \ + || fail "test fixture error: the trust file was rebound before the torn window began" + +# The still-running poller now sees its condition go true and reaches the +# fire-time reload while the torn state above is live and the lock is held. +: > "$TORN_TRIGGER" +sleep 0.3 +if first_result "$H" "$SID" >/dev/null 2>&1; then + fail "the reload must block on the source lock instead of reading the torn spec/trust pair" +fi + +: > "$TORN_RELEASE" +wait "$HOLDER_PID" 2>/dev/null || true +wait_for_result "$H" "$SID" || fail "the watch never captured an outcome after the torn window closed" +RESULT=$(first_result "$H" "$SID") +assert_grep 'status: fired' "$RESULT" \ + "the reload must wait past the torn spec/trust window, not reject a legitimate rebind mid-publish" +assert_grep 'v2 ran against' "$RESULT" "the fired action ran the rebound (v2) bytes, not a rejection from a torn read" +pass "the fire-time reload never observes rebind_one's spec/trust publish mid-rename" + printf 'all fm-procevent-when tests passed\n' diff --git a/tests/fm-procevent.test.sh b/tests/fm-procevent.test.sh index 436b72449b7..c75f4c54ac5 100755 --- a/tests/fm-procevent.test.sh +++ b/tests/fm-procevent.test.sh @@ -53,6 +53,43 @@ pe_register() { # <home> <adapter> <source-id> -- <argv>... new_home() { mkdir -p "$1/state"; } wake_payloads() { awk -F '\t' '{print $5}' "$1/state/.wake-queue" 2>/dev/null; } +# The wake queue is a durable tab-separated record firstmate consumes: +# <epoch> <sequence> <kind> <key> <payload>. These read the rows reconcile +# publishes for a source it stranded, keyed by that source and its claim +# generation. +stranded_wake_keys() { # <home> <source-id> + [ -e "$1/state/.wake-queue" ] || return 0 + awk -F '\t' -v id="$2" \ + '$3 == "check" && index($4, "procevent:" id ":stranded:") == 1 { print $4 }' \ + "$1/state/.wake-queue" +} +stranded_wake_count() { # <home> <source-id> + stranded_wake_keys "$1" "$2" | grep -c . || true +} +stranded_wake_payloads() { # <home> <source-id> + [ -e "$1/state/.wake-queue" ] || return 0 + awk -F '\t' -v id="$2" \ + '$3 == "check" && index($4, "procevent:" id ":stranded:") == 1 { print $5 }' \ + "$1/state/.wake-queue" +} +# The same rows for a launch reconcile could not confirm, keyed by that source +# and the registration identity the launch ran under. +launch_failed_wake_keys() { # <home> <source-id> + [ -e "$1/state/.wake-queue" ] || return 0 + awk -F '\t' -v id="$2" \ + '$3 == "check" && index($4, "procevent:" id ":launch-failed:") == 1 { print $4 }' \ + "$1/state/.wake-queue" +} +launch_failed_wake_count() { # <home> <source-id> + launch_failed_wake_keys "$1" "$2" | grep -c . || true +} +launch_failed_wake_payloads() { # <home> <source-id> + [ -e "$1/state/.wake-queue" ] || return 0 + awk -F '\t' -v id="$2" \ + '$3 == "check" && index($4, "procevent:" id ":launch-failed:") == 1 { print $5 }' \ + "$1/state/.wake-queue" +} + first_result() { # <home> <source-id>: print the first captured result, if any local g for g in "$1/state/procevent-inbox/$2".*.result; do @@ -1138,6 +1175,42 @@ assert_contains "$orphan_out" "started=0" \ [ "$(wc -l < "$ORPHAN_LOG" | tr -d ' ')" = 1 ] \ || fail "reconcile started a source beside an ambiguous leaderless group" assert_absent "$ORPHAN_OVERLAP" "no replacement source starts while the leaderless group remains" +# This is the ordinary crash shape, and it is refused permanently: `orphaned` +# in a listing and `uncertain=1` in output the supervision cycle discards +# reach nobody, so the strand has to announce itself durably, exactly once, +# under a key the watcher can tell apart from a captured result. +orphan_token=$(sed -n '3p' "$FM_PROCEVENT_CLAIM_ROOT/orphan-src.claim") +[ -n "$orphan_token" ] || fail "could not read the leaderless claim's token" +[ "$(stranded_wake_count "$HG" orphan-src)" = 1 ] \ + || fail "reconcile stranded a leaderless source without announcing it: $orphan_out" +[ "$(stranded_wake_keys "$HG" orphan-src)" = "procevent:orphan-src:stranded:$orphan_token" ] \ + || fail "the stranded wake is not keyed by source and claim generation: $(stranded_wake_keys "$HG" orphan-src)" +orphan_wake=$(stranded_wake_payloads "$HG" orphan-src) +assert_contains "$orphan_wake" "orphan-src" \ + "the leaderless stranded wake does not name the source it is about: $orphan_wake" +assert_contains "$orphan_wake" "polling" \ + "the leaderless stranded wake does not say what a human should check: $orphan_wake" +# `start` reports this claim as owned and reclaims nothing, so a wake that +# named it as the clearing command would send someone to a no-op. +case "$orphan_wake" in + *"start orphan-src"*) fail "the leaderless stranded wake names start as clearing it: $orphan_wake" ;; +esac +orphan_start=$(pe "$HG" start orphan-src 2>&1) +assert_contains "$orphan_start" "already owned" \ + "start displaced a leaderless group's claim: $orphan_start" +[ "$(wc -l < "$ORPHAN_LOG" | tr -d ' ')" = 1 ] \ + || fail "start ran the source beside an ambiguous leaderless group" +orphan_again=$(pe "$HG" reconcile) +assert_contains "$orphan_again" "started=0" \ + "the second cycle replaced an ambiguous leaderless generation: $orphan_again" +assert_contains "$orphan_again" "uncertain=1" \ + "the second cycle stopped reporting the claim it could not settle: $orphan_again" +[ "$(stranded_wake_count "$HG" orphan-src)" = 1 ] \ + || fail "reconcile re-announced the same leaderless strand: $orphan_again" +[ "$(wc -l < "$ORPHAN_LOG" | tr -d ' ')" = 1 ] \ + || fail "the second cycle started a source beside an ambiguous leaderless group" +kill -0 -"$orphan_leader" 2>/dev/null \ + || fail "announcing the strand signalled the leaderless process group" kill -KILL -"$orphan_leader" 2>/dev/null || true for _ in $(seq 1 50); do kill -0 -"$orphan_leader" 2>/dev/null || break; sleep 0.1; done kill -0 -"$orphan_leader" 2>/dev/null && fail "could not clean up the leaderless fixture group" @@ -1277,8 +1350,66 @@ sr4_out=$(pe "$HSR4" reconcile) sleep 0.5 [ "$(wc -l < "$SR4_LOG" | tr -d ' ')" = 1 ] \ || fail "reconcile started a replacement beside a reused pid's live group: $sr4_out" +# Ownership cannot move here by design, so a replacement could only die on the +# claim it cannot take - once per reconcile cycle, forever. +assert_contains "$sr4_out" "started=0" \ + "reconcile reported a start into a claim nothing can take: $sr4_out" +assert_contains "$sr4_out" "uncertain=1" \ + "reconcile did not report the claim it could not settle: $sr4_out" [ "$(sed -n '2p' "$sr4_claim")" = "$sr4_leader" ] \ || fail "reconcile replaced the reused-pid generation's claim" +# Nothing can take this source, so reporting it as unowned reads like an idle +# source waiting to be started - the reassuring answer this surface gave while a +# review board collected nothing. +sr4_owner=$(pe "$HSR4" list | awk '$1 == "reused-group-src" { print $3 }') +[ "$sr4_owner" = orphaned ] \ + || fail "a source no caller can claim is listed as '$sr4_owner'" +# `orphaned` in a listing and `uncertain=1` in output the supervision cycle +# discards reach nobody. The strand has to announce itself durably, exactly +# once, and say which command clears it. +[ "$(stranded_wake_count "$HSR4" reused-group-src)" = 1 ] \ + || fail "reconcile stranded a source without announcing it: $sr4_out" +# The key carries the source and its claim generation in a shape the watcher +# can tell apart from a captured result, so the strand is never headlined as one. +[ "$(stranded_wake_keys "$HSR4" reused-group-src)" = "procevent:reused-group-src:stranded:$(sed -n '3p' "$sr4_claim")" ] \ + || fail "the stranded wake is not keyed by source and claim generation: $(stranded_wake_keys "$HSR4" reused-group-src)" +sr4_wake=$(stranded_wake_payloads "$HSR4" reused-group-src) +assert_contains "$sr4_wake" "reused-group-src" \ + "the stranded wake does not name the source it is about: $sr4_wake" +assert_contains "$sr4_wake" "bin/fm-procevent.sh start reused-group-src" \ + "the stranded wake does not name the command that clears it: $sr4_wake" +# A wake nobody can silence is as unusable as one nobody gets: the same stranded +# generation must not re-announce on every supervision cycle. +sr4_again=$(pe "$HSR4" reconcile) +assert_contains "$sr4_again" "uncertain=1" \ + "the second cycle stopped reporting the claim it could not settle: $sr4_again" +[ "$(stranded_wake_count "$HSR4" reused-group-src)" = 1 ] \ + || fail "reconcile re-announced the same stranded generation: $sr4_again" +[ "$(wc -l < "$SR4_LOG" | tr -d ' ')" = 1 ] \ + || fail "the second cycle started a replacement beside a reused pid's live group: $sr4_again" +# The wake names `start` as the recovery, so run it against the state it will +# actually meet. The earlier end-to-end demonstration of that command used an +# UNDRIFTED fixture and therefore proved only the easy case; on this one the +# state root has drifted, so the dead generation's reservation records cannot +# be tidied, and the claim path waives that tidy-up only for a generation +# proven gone - which a surviving group is not. `start` must refuse here, keep +# the claim, and start no second source beside the live group, and the wake +# must have said so rather than promising a reclaim. +assert_contains "$sr4_wake" "cannot claim source" \ + "the stranded wake promises an unconditional reclaim: $sr4_wake" +set +e +sr4_start=$(pe "$HSR4" start reused-group-src 2>&1) +sr4_start_rc=$? +set -e +[ "$sr4_start_rc" -ne 0 ] \ + || fail "start reported success against a claim it could not tidy: $sr4_start" +assert_contains "$sr4_start" "cannot claim source" \ + "start did not refuse by name on the drifted reused-pid fixture: $sr4_start" +[ "$(sed -n '2p' "$sr4_claim")" = "$sr4_leader" ] \ + || fail "a refused start replaced the reused-pid generation's claim" +sleep 0.3 +[ "$(wc -l < "$SR4_LOG" | tr -d ' ')" = 1 ] \ + || fail "a refused start ran a second source beside a reused pid's live group: $(cat "$SR4_LOG")" set +e sr4_retire=$(pe "$HSR4" retire reused-group-src 2>&1) sr4_rc=$? @@ -1299,6 +1430,347 @@ kill -0 -"$sr4_leader" 2>/dev/null \ && fail "retirement left the restored reused-group fixture running" pass "a reused pid never makes its surviving process group reclaimable" +# --- the easy case the wake promises: an undrifted reused-pid claim ---------- +# Same strand, no state-root drift: the dead generation's reservation records +# can be tidied, so the attached `start` the wake names takes the claim and +# runs the source. The claim path does not consult the process group; that is +# the documented asymmetry between reconcile and a deliberate start. +HSR5="$TMP_ROOT/hsr5"; new_home "$HSR5" +SR5_TRIGGER="$TMP_ROOT/reused-plain-trigger" +SR5_LOG="$TMP_ROOT/reused-plain-executions" +pe_register "$HSR5" lavish reused-plain-src -- "$RACE_BLOCKER" "$SR5_LOG" "$SR5_TRIGGER" >/dev/null +pe "$HSR5" reconcile >/dev/null +wait_for "$FM_PROCEVENT_CLAIM_ROOT/reused-plain-src.claim" \ + || fail "undrifted reused-pid fixture never claimed its source" +wait_for "$SR5_LOG" || fail "undrifted reused-pid fixture source never started" +sr5_claim="$FM_PROCEVENT_CLAIM_ROOT/reused-plain-src.claim" +sr5_leader=$(sed -n '2p' "$sr5_claim") +awk 'NR == 4 { print "different-live-process-identity"; next } { print }' \ + "$sr5_claim" > "$sr5_claim.tmp" && mv "$sr5_claim.tmp" "$sr5_claim" +chmod 0600 "$sr5_claim" +sr5_out=$(pe "$HSR5" reconcile) +assert_contains "$sr5_out" "uncertain=1" \ + "reconcile did not strand the undrifted reused-pid claim: $sr5_out" +[ "$(stranded_wake_count "$HSR5" reused-plain-src)" = 1 ] \ + || fail "the undrifted strand was not announced: $sr5_out" +pe "$HSR5" start reused-plain-src > "$TMP_ROOT/reused-plain-start.out" 2>&1 & +sr5_start_pid=$! +wait_for_lines "$SR5_LOG" 2 \ + || fail "start did not reclaim the undrifted reused-pid claim: $(cat "$TMP_ROOT/reused-plain-start.out")" +[ "$(sed -n '2p' "$sr5_claim")" != "$sr5_leader" ] \ + || fail "start ran the source without taking the claim from the dead generation" +: > "$SR5_TRIGGER" +wait "$sr5_start_pid" \ + || fail "start failed after reclaiming the undrifted claim: $(cat "$TMP_ROOT/reused-plain-start.out")" +assert_contains "$(cat "$TMP_ROOT/reused-plain-start.out")" "captured:" \ + "the reclaiming start did not capture the source's result" +for _ in $(seq 1 50); do kill -0 -"$sr5_leader" 2>/dev/null || break; sleep 0.1; done +pe "$HSR5" retire reused-plain-src >/dev/null 2>&1 || true +pass "start reclaims a reused-pid claim whose leftovers can still be tidied" + +# --- a launch that cannot confirm is announced once per failure episode ------ +# `bin/fm-watch.sh` discards reconcile's `failed=` count and exit status, so a +# runner that dies before claiming - for any cause, not only the claim wedge - +# would be relaunched and reported failed every cycle with nobody told: armed +# in appearance, a dead drop in fact. The episode is keyed by the registration +# identity the launch ran under and ends when a launch of that source confirms, +# so the registration below is damaged and repaired IN PLACE to keep that +# identity fixed across the whole sequence. The wake changes nothing about the +# launch: every failing cycle below still relaunches and still reports failed. +HEP="$TMP_ROOT/hep"; new_home "$HEP" +EP_SOURCE_CMD="$TMP_ROOT/episode-source.sh" +cat > "$EP_SOURCE_CMD" <<'SH' +#!/usr/bin/env bash +printf 'episode result\n' +SH +chmod +x "$EP_SOURCE_CMD" +pe_register "$HEP" lavish episode-src -- "$EP_SOURCE_CMD" >/dev/null +EP_SOURCE="$HEP/state/procevent/episode-src.source" +cp "$EP_SOURCE" "$TMP_ROOT/episode-good.source" +awk '/^argv:$/ { print; exit } { print }' "$EP_SOURCE" > "$TMP_ROOT/episode-bad.source" \ + || fail "could not prepare the damaged episode registration" +ep_damage() { cat "$TMP_ROOT/episode-bad.source" > "$EP_SOURCE"; } +ep_repair() { cat "$TMP_ROOT/episode-good.source" > "$EP_SOURCE"; } +ep_reconcile() { # <expected-fragment> <expected-exit-nonzero:0|1> <msg>; sets ep_out + local rc=0 + ep_out=$(FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS=2 pe "$HEP" reconcile) || rc=$? + assert_contains "$ep_out" "$1" "$3: $ep_out" + if [ "$2" -eq 1 ]; then + [ "$rc" -ne 0 ] || fail "$3 (reconcile exited 0): $ep_out" + else + [ "$rc" -eq 0 ] || fail "$3 (reconcile exited $rc): $ep_out" + fi +} +ep_damage +ep_reconcile "failed=1" 1 "a launch that never proved its claim was not reported failed" +[ "$(launch_failed_wake_count "$HEP" episode-src)" = 1 ] \ + || fail "a launch that could not confirm was not announced: $ep_out" +ep_key=$(launch_failed_wake_keys "$HEP" episode-src) +# <registration identity>-<per-episode nonce>: the watcher remembers every key +# it has surfaced for good, so the identity alone would announce only the first +# episode of a registration (tests/fm-watch-triage.test.sh proves delivery). +[[ "$ep_key" =~ ^(procevent:episode-src:launch-failed:[0-9]+-[0-9]+)-[0-9]+$ ]] \ + || fail "the launch-failed wake is not keyed by source, registration identity and episode: $ep_key" +ep_episode_prefix=${BASH_REMATCH[1]} +ep_wake=$(launch_failed_wake_payloads "$HEP" episode-src) +assert_contains "$ep_wake" "episode-src" \ + "the launch-failed wake does not name the source it is about: $ep_wake" +# The payload may state only what confirmation observed: no claim proved +# inside the window. It cannot know whether the runner died or was slow, so it +# must not assert a cause, must not present `start` as the fix, and must say +# that a later cycle finding the source owned closes the episode by itself. +assert_contains "$ep_wake" "did not prove it took the source's claim within FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS" \ + "the launch-failed wake does not state what confirmation observed: $ep_wake" +assert_contains "$ep_wake" "attached bin/fm-procevent.sh start episode-src to reproduce a refusal" \ + "the launch-failed wake does not say start reproduces rather than fixes: $ep_wake" +assert_contains "$ep_wake" "adapter binary" \ + "the launch-failed wake does not name what to check: $ep_wake" +assert_contains "$ep_wake" "finds the source owned ends this episode on its own" \ + "the launch-failed wake does not say a slow runner closes its own episode: $ep_wake" +case "$ep_wake" in + *"never claimed"*|*"exited without"*|*"runner died"*) + fail "the launch-failed wake asserts a cause confirmation cannot observe: $ep_wake" ;; +esac +ep_reconcile "failed=1" 1 "the second cycle stopped relaunching a source that cannot start" +[ "$(launch_failed_wake_count "$HEP" episode-src)" = 1 ] \ + || fail "the same failure episode was announced twice: $ep_out" +ep_repair +ep_reconcile "started=1" 0 "a repaired source did not confirm" +assert_contains "$ep_out" "failed=0" "a repaired source was still reported failed: $ep_out" +[ "$(launch_failed_wake_count "$HEP" episode-src)" = 1 ] \ + || fail "a confirmed launch produced a launch-failed wake: $ep_out" +for _ in $(seq 1 100); do + [ -e "$FM_PROCEVENT_CLAIM_ROOT/episode-src.claim" ] || break + sleep 0.1 +done +[ ! -e "$FM_PROCEVENT_CLAIM_ROOT/episode-src.claim" ] \ + || fail "the confirmed episode runner never released its claim" +ep_damage +ep_reconcile "failed=1" 1 "a source that failed again after recovering was not reported failed" +[ "$(launch_failed_wake_count "$HEP" episode-src)" = 2 ] \ + || fail "a new failure episode after a confirmed launch was not announced: $ep_out" +# The earlier version of this assertion locked in ONE key for both episodes, +# which is exactly the collision that left every episode after the first +# unsurfaced: both keys must carry the same registration identity and still +# differ, or the watcher's seen marker for episode one suppresses episode two. +ep_key_again=$(launch_failed_wake_keys "$HEP" episode-src | sed -n '2p') +[ "$ep_key_again" != "$ep_key" ] \ + || fail "a new failure episode reused the first episode's queue key: $ep_key_again" +case "$ep_key_again" in + "$ep_episode_prefix"-*) ;; + *) fail "the second episode ran under a different registration identity: $ep_key_again (first: $ep_key)" ;; +esac +ep_repair +pe "$HEP" retire episode-src >/dev/null 2>&1 || true +pass "a launch that cannot confirm is announced once per failure episode" + +# --- the launch-failed key fits the watcher's seen marker at the id limit ---- +# bin/fm-watch.sh names the marker for a surfaced key `.seen-procevent-<hex>`, +# 16 + 2 * keylen bytes against NAME_MAX 255, so a key longer than 119 chars +# cannot be marked and its wake would re-surface every cycle. The longest id +# the validator accepts is 64 chars; the executed key for such an id must fit. +HLK="$TMP_ROOT/hlk"; new_home "$HLK" +LK_ID=$(printf 'k%.0s' $(seq 1 64)) +[ "${#LK_ID}" -eq 64 ] || fail "fixture invalid: long source id is ${#LK_ID} chars" +pe_register "$HLK" lavish "$LK_ID" -- "$EP_SOURCE_CMD" >/dev/null +LK_SOURCE="$HLK/state/procevent/$LK_ID.source" +if ! { awk '/^argv:$/ { print; exit } { print }' "$LK_SOURCE" > "$LK_SOURCE.tmp" \ + && cat "$LK_SOURCE.tmp" > "$LK_SOURCE" && rm -f -- "$LK_SOURCE.tmp"; }; then + fail "could not damage the long-id registration" +fi +lk_out=$(FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS=2 pe "$HLK" reconcile) || true +assert_contains "$lk_out" "failed=1" "the long-id launch was not reported failed: $lk_out" +lk_key=$(launch_failed_wake_keys "$HLK" "$LK_ID") +[ -n "$lk_key" ] || fail "the long-id launch failure was not announced: $lk_out" +[ "${#lk_key}" -le 119 ] \ + || fail "a 64-char source id yields a ${#lk_key}-char launch-failed key, which the watcher cannot mark: $lk_key" +pe "$HLK" retire "$LK_ID" >/dev/null 2>&1 || true +pass "a 64-char source id keeps the launch-failed key within the watcher's marker bound" + +# --- reconcile reports only launches it actually confirmed ------------------- +# The reported incident. A review board the captain had answered sat collecting +# nothing while `reconcile` reported a start on every run: `detach_runner` is +# fire-and-forget with the child's stderr discarded, so a runner that died +# before it could claim was counted exactly like one that is listening. A +# surface that presents as armed while being a dead drop is worse than one that +# visibly fails, because the answers look recorded. +# +# The damaged registration below makes the runner die BEFORE it claims, which +# is what keeps this deterministic: a runner that claims and then dies would +# race the confirmation either way, and the next reconcile cycle is what covers +# that case. +HUF="$TMP_ROOT/huf"; new_home "$HUF" +UF_TRIGGER="$TMP_ROOT/unstartable-trigger" +pe_register "$HUF" lavish unstartable-src -- "$BLOCKER" "$UF_TRIGGER" "unstartable" >/dev/null +UF_SOURCE="$HUF/state/procevent/unstartable-src.source" +if ! awk '/^argv:$/ { print; exit } { print }' "$UF_SOURCE" > "$UF_SOURCE.tmp"; then + fail "could not damage the unstartable registration" +fi +mv "$UF_SOURCE.tmp" "$UF_SOURCE" || fail "could not damage the unstartable registration" +chmod 0600 "$UF_SOURCE" +uf_rc=0 +uf_out=$(FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS=2 pe "$HUF" reconcile) || uf_rc=$? +assert_contains "$uf_out" "started=0" \ + "reconcile counted a runner that never started as a start: $uf_out" +assert_contains "$uf_out" "failed=1" \ + "reconcile did not report the launch it could not confirm: $uf_out" +[ "$uf_rc" -ne 0 ] || fail "reconcile reported success while a source could not start: $uf_out" +uf_owner=$(pe "$HUF" list | awk '$1 == "unstartable-src" { print $3 }') +[ "$uf_owner" = none ] || fail "the unstartable source reports an owner: $uf_owner" +pe "$HUF" retire unstartable-src >/dev/null 2>&1 || true +pass "reconcile reports a launch it could not confirm instead of counting it as a start" + +# --- a launch that finished before the first poll is still confirmed --------- +# Confirmation has to read evidence a finished runner leaves behind. A runner +# removes its own runner record on the way out, so a source that claims, runs +# and exits before confirmation looks at it once returns every transient signal +# to exactly what it was before the launch - and a good run gets reported as a +# failure, on every cycle, for a source that is working perfectly. +# +# The second registration is what makes that deterministic rather than a race: +# reconcile launches the fast source first, then blocks acquiring the held +# lock of the second source, and the holder is released only once the fast +# runner has captured its result and let go of both its claim and its runner +# record. Confirmation therefore starts strictly after the fast runner is gone. +HFC="$TMP_ROOT/hfc"; new_home "$HFC" +FC_FAST="$TMP_ROOT/fast-source.sh" +cat > "$FC_FAST" <<'SH' +#!/usr/bin/env bash +printf 'fast payload\n' +SH +chmod +x "$FC_FAST" +FC_TRIGGER="$TMP_ROOT/fast-hold-trigger" +pe_register "$HFC" lavish aa-fast-src -- "$FC_FAST" >/dev/null +pe_register "$HFC" lavish zz-hold-src -- "$BLOCKER" "$FC_TRIGGER" "held" >/dev/null +FC_READY="$TMP_ROOT/fast-hold-ready"; FC_RELEASE="$TMP_ROOT/fast-hold-release" +hold_source_lock zz-hold-src "$FC_READY" "$FC_RELEASE" +wait_for "$FC_READY" || fail "the fast-source fixture could not hold a source lock" +( + for _ in $(seq 1 600); do + if first_result "$HFC" aa-fast-src >/dev/null 2>&1 \ + && [ ! -e "$HFC/state/procevent/aa-fast-src.runner" ] \ + && [ ! -e "$FM_PROCEVENT_CLAIM_ROOT/aa-fast-src.claim" ]; then + break + fi + sleep 0.05 + done + : > "$FC_RELEASE" +) & +FC_RELEASER=$! +fc_rc=0 +fc_out=$(FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS=2 pe "$HFC" reconcile) || fc_rc=$? +wait "$FC_RELEASER" 2>/dev/null || true +wait "$HOLDER_PID" 2>/dev/null || true +first_result "$HFC" aa-fast-src >/dev/null \ + || fail "fixture invalid: the fast source never produced a result: $fc_out" +assert_contains "$fc_out" "started=2" \ + "reconcile did not report both launches as started: $fc_out" +assert_contains "$fc_out" "failed=0" \ + "reconcile reported a launch that ran to completion as a failure: $fc_out" +[ "$fc_rc" -eq 0 ] || fail "reconcile exited non-zero with every launch confirmed: $fc_out" +: > "$FC_TRIGGER" +pe "$HFC" retire aa-fast-src >/dev/null 2>&1 || true +pe "$HFC" retire zz-hold-src >/dev/null 2>&1 || true +pass "a launch that finished before confirmation looked is still reported as started" + +# --- a zero-padded confirm window is read as base 10 ------------------------- +# The window's validator reads base 10, so `08` is a value it accepts. Read as +# octal in arithmetic it is not a number at all, which under `set -u` takes the +# confirmation down with it and turns every launch of the cycle - including a +# perfectly healthy one - into a reported failure and a non-zero exit. +HZP="$TMP_ROOT/hzp"; new_home "$HZP" +ZP_TRIGGER="$TMP_ROOT/zeropad-trigger" +pe_register "$HZP" lavish zeropad-src -- "$BLOCKER" "$ZP_TRIGGER" "zeropad" >/dev/null +zp_rc=0 +zp_out=$(FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS=08 pe "$HZP" reconcile 2>/dev/null) || zp_rc=$? +assert_contains "$zp_out" "started=1" \ + "a zero-padded confirm window lost the launch reconcile started: $zp_out" +assert_contains "$zp_out" "failed=0" \ + "a zero-padded confirm window reported a healthy launch as failed: $zp_out" +[ "$zp_rc" -eq 0 ] || fail "a zero-padded confirm window made reconcile exit non-zero: $zp_out" +: > "$ZP_TRIGGER" +pe "$HZP" retire zeropad-src >/dev/null 2>&1 || true +pass "a zero-padded launch confirm window is honored as base 10" + +# --- an unusable confirm window is refused by name -------------------------- +# A window this command cannot use makes every launch unconfirmable. Reported +# from inside the confirmation it comes out as a fleet of healthy runners that +# all "could not start", blaming the sources instead of the typo. Every other +# tunable on this path - the launch floor, the output bound - refuses a bad +# value by name before anything runs, and so does this one. +HIW="$TMP_ROOT/hiw"; new_home "$HIW" +IW_TRIGGER="$TMP_ROOT/invalid-window-trigger" +pe_register "$HIW" lavish invalid-window-src -- "$BLOCKER" "$IW_TRIGGER" "window" >/dev/null +for iw_value in 5s 0 700; do + iw_rc=0 + iw_out=$(FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS="$iw_value" pe "$HIW" reconcile 2>&1) || iw_rc=$? + [ "$iw_rc" -ne 0 ] \ + || fail "reconcile accepted the unusable confirm window '$iw_value': $iw_out" + assert_contains "$iw_out" "FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS" \ + "the unusable confirm window '$iw_value' was not named by what refused it: $iw_out" + case "$iw_out" in + *failed=*) fail "the unusable confirm window '$iw_value' was blamed on the sources: $iw_out" ;; + esac + [ ! -e "$FM_PROCEVENT_CLAIM_ROOT/invalid-window-src.claim" ] \ + || fail "reconcile launched a runner before refusing the confirm window '$iw_value'" +done +# The refusal costs the source nothing: it still arms on the next run with a +# usable value. +iw_ok=$(pe "$HIW" reconcile) +assert_contains "$iw_ok" "started=1" \ + "the source did not arm once its confirm window was usable: $iw_ok" +assert_contains "$iw_ok" "failed=0" \ + "the source was reported as failed once its confirm window was usable: $iw_ok" +: > "$IW_TRIGGER" +pe "$HIW" retire invalid-window-src >/dev/null 2>&1 || true +pass "an unusable launch confirm window is refused by name instead of blamed on the sources" + +# --- a dead generation's untidyable leftovers never wedge ownership ---------- +# The same wedge as the state-root case above, reached through the sibling +# cleanups in the stale-claim branch rather than the capture reservation. Every +# one of them tidies leftovers keyed by the DEAD generation's claim token, so +# none can collide with the replacement, yet a failure in any of them used to +# refuse the claim outright - permanently, because the condition never clears on +# its own. Here the recorded registry directory no longer resolves to a +# directory at all, which is what a claim recorded before its home was replaced +# looks like. +HUW="$TMP_ROOT/huw"; new_home "$HUW" +UW_TRIGGER="$TMP_ROOT/untidyable-trigger" +UW_LOG="$TMP_ROOT/untidyable-executions" +pe_register "$HUW" lavish untidyable-src -- "$RACE_BLOCKER" "$UW_LOG" "$UW_TRIGGER" >/dev/null +UW_REG_FILE="$TMP_ROOT/untidyable-recorded-registry" +: > "$UW_REG_FILE" +uw_identity=$(bash -c '. "$1/bin/fm-pr-lib.sh"; fm_pr_file_identity "$2"' _ \ + "$ROOT" "$HUW/state/procevent/untidyable-src.source") \ + || fail "could not read the untidyable fixture registration identity" +UW_CLAIM="$FM_PROCEVENT_CLAIM_ROOT/untidyable-src.claim" +{ + printf '%s\n%s\nuntidyable-token\nuntidyable-identity\n' "$HUW" 999999 + printf '%s\n%s\nactive\n' "$UW_REG_FILE" "$uw_identity" + printf '%s\n%s\n%s\n%s\n%s\n' "$HUW/state" \ + "$(bash -c '. "$1/bin/fm-pr-lib.sh"; fm_pr_file_device "$2"' _ "$ROOT" "$HUW/state")" \ + "$(bash -c '. "$1/bin/fm-pr-lib.sh"; fm_pr_file_inode "$2"' _ "$ROOT" "$HUW/state")" \ + "$(id -u)" 755 +} > "$UW_CLAIM" +chmod 0600 "$UW_CLAIM" +kill -0 999999 2>/dev/null && fail "fixture invalid: the untidyable claim names a live pid" +kill -0 -999999 2>/dev/null && fail "fixture invalid: the untidyable claim's process group is alive" +uw_rc=0 +uw_out=$(pe "$HUW" reconcile) || uw_rc=$? +[ "$uw_rc" -eq 0 ] || fail "reconcile could not repair a provably dead generation: $uw_out" +# Reporting a start is not the same fact as listening, so prove the listening +# half first: before this fix reconcile reported exactly this start on every run +# while the dead generation kept the claim and nothing ever attached. +wait_for "$UW_LOG" || fail "reconcile reported a start but no replacement source ever ran: $uw_out" +uw_new=$(sed -n '2p' "$UW_CLAIM") +[ "$uw_new" != 999999 ] || fail "the dead generation kept owning the source: $uw_out" +kill -0 "$uw_new" 2>/dev/null || fail "the replacement runner did not take ownership: $uw_out" +assert_contains "$uw_out" "started=1" "reconcile did not report the replacement it started: $uw_out" +assert_contains "$uw_out" "failed=0" "reconcile could not confirm the replacement: $uw_out" +: > "$UW_TRIGGER" +pe "$HUW" retire untidyable-src >/dev/null +pass "a dead generation whose leftovers cannot be tidied never keeps owning its source" + HJ="$TMP_ROOT/hj"; new_home "$HJ" TORN_TRIGGER="$TMP_ROOT/torn-trigger" pe_register "$HJ" lavish torn-src -- "$BLOCKER" "$TORN_TRIGGER" "torn" >/dev/null @@ -1792,6 +2264,15 @@ assert_present "$HMISS/state/procevent/$missing_id.source" \ MISSING_CAPTURED=$(first_result "$HMISS" "$missing_id" || true) assert_contains "$("$ROOT/bin/fm-procevent-lavish.sh" classify "$MISSING_CAPTURED")" artifact-missing \ "the recurring result from a deleted artifact classifies as artifact-missing" +# Capture precedes runner exit and claim release. This case tests retiring the +# registered id after completed polls, not racing the runner's identity proof +# against its exit; live-runner retirement has separate blocking-source cases. +for _ in $(seq 1 100); do + [ ! -e "$FM_PROCEVENT_CLAIM_ROOT/$missing_id.claim" ] && break + sleep 0.1 +done +assert_absent "$FM_PROCEVENT_CLAIM_ROOT/$missing_id.claim" \ + "the completed missing-artifact runner releases its claim before retirement" # Retire by source id stops the recurring wakes (the wake text carries the id). # Retire returns only after the runner it owns is stopped and its claim # released, so the count taken after it is a settled baseline. diff --git a/tests/fm-public-followup.test.sh b/tests/fm-public-followup.test.sh index e2950fa7abb..6fb605b1052 100755 --- a/tests/fm-public-followup.test.sh +++ b/tests/fm-public-followup.test.sh @@ -292,6 +292,55 @@ expect_failure() { fi } +# --- 0. the suite's seeding never reaches a real backlog ------------------------ + +# An operator shell exports TASKS_AXI_FILE at its live home's backlog, and +# tasks-axi resolves that env ahead of the fixture's .tasks.toml. This suite +# seeds obligations with bare `tasks-axi` from the fixture home, so before +# tests/lib.sh cleared the ambient overrides, every seed landed in the operator's +# real backlog while the consumer read the empty fixture. Re-enter the suite's +# seeding exactly as a test process begins (source tests/lib.sh, then seed from +# the fixture) under a decoy "live" backlog and a backend tasks-axi refuses: the +# decoy must stay byte-identical and the fixture must hold the obligation. +test_ambient_tasks_axi_env_never_reaches_a_real_backlog() { + local home decoy_dir decoy fixture_state + home=$(make_home ambient-env) + decoy_dir="$TMP_ROOT/ambient-live/data" + decoy="$decoy_dir/backlog.md" + mkdir -p "$decoy_dir" + printf '## In flight\n\n## Queued\n\n## Done\n' > "$decoy" + cp "$decoy" "$decoy.expected" + jq -n '{request_id:"req-amb", platform:"discord", + context_binding:{version:"ctx1", value:"ctx1_req-amb"}, + public_safe_summary:"seeded under an ambient tasks-axi override", + received_at:"2026-07-30T10:00:00Z", + followup_expires_at:"2026-08-06T10:00:00Z", + reservation_expires_at:"2026-08-06T10:00:00Z"}' > "$home/request.json" + jq -n '{type:"pr-merged", project:"firstmate", + required_deliverables:["pr_url"], completion_policy:"all-required"}' \ + > "$home/expected.json" + + TASKS_AXI_FILE="$decoy" TASKS_AXI_BACKEND=no-such-backend bash -c ' + set -u + . "$1/tests/lib.sh" + [ -z "${TASKS_AXI_FILE+x}" ] || { echo "TASKS_AXI_FILE survived tests/lib.sh"; exit 1; } + [ -z "${TASKS_AXI_BACKEND+x}" ] || { echo "TASKS_AXI_BACKEND survived tests/lib.sh"; exit 1; } + cd "$2" && tasks-axi public-followup add pf-ambient \ + --request-context-file "$2/request.json" --purpose promised-final \ + --expected-final-file "$2/expected.json" --expires-at 2026-10-01T00:00:00Z >/dev/null + ' _ "$ROOT" "$home" \ + || fail "seeding under an ambient tasks-axi override did not reach the fixture backlog" + + cmp -s "$decoy" "$decoy.expected" \ + || fail "the suite's seeding wrote the ambient TASKS_AXI_FILE backlog instead of the fixture" + [ "$(find "$decoy_dir" -mindepth 1 -maxdepth 1 -exec basename {} \; | LC_ALL=C sort | tr '\n' ' ')" = "backlog.md backlog.md.expected " ] \ + || fail "the suite's seeding left an artifact beside the ambient TASKS_AXI_FILE backlog: $(find "$decoy_dir" -mindepth 1 -maxdepth 1 -exec basename {} \; | LC_ALL=C sort | tr '\n' ' ')" + fixture_state=$(task_state "$home" pf-ambient) + [ "$fixture_state" != absent ] \ + || fail "the fixture backlog does not hold the obligation seeded under the ambient override" + pass "the suite's seeding never reaches a backlog named by ambient TASKS_AXI_FILE/BACKEND" +} + # --- 0. bounded, single-line, character-safe outcome text ----------------------- # The outcome sentence becomes a public reply, so bounding it must not mangle @@ -3118,6 +3167,7 @@ if [ -n "${FM_TEST_ONLY:-}" ]; then exit 0 fi +test_ambient_tasks_axi_env_never_reaches_a_real_backlog test_outcome_text_is_bounded_without_corrupting_characters test_restart_e2e_delivers_exactly_once test_duplicate_event_and_replay_are_noops diff --git a/tests/fm-remote-secondmate-trace-context.test.sh b/tests/fm-remote-secondmate-trace-context.test.sh index 8efa24291cd..2be7e5c7116 100755 --- a/tests/fm-remote-secondmate-trace-context.test.sh +++ b/tests/fm-remote-secondmate-trace-context.test.sh @@ -96,7 +96,11 @@ git -C "$REMOTE_ROOT" init -q -b main git -C "$REMOTE_ROOT" config user.email test@example.com git -C "$REMOTE_ROOT" config user.name Test git -C "$REMOTE_ROOT" add . -git -C "$REMOTE_ROOT" commit -qm 'remote fixture root' +# Complete fixture maintenance before the remote route clones this repository. +# A detached repack can remove loose objects while a local clone copies them. +# The gc setting also covers Git versions predating maintenance.autoDetach. +git -C "$REMOTE_ROOT" -c maintenance.autoDetach=false -c gc.autoDetach=false \ + commit -qm 'remote fixture root' cat > "$FAKEBIN/fake-ssh" <<'SH' #!/usr/bin/env bash @@ -307,4 +311,4 @@ try_flag 'requires a non-empty value' \ --secondmate --traceparent= pass "delivery: a parent-supplied carrier is accepted only for a secondmate launch and only as a strict W3C value" -echo "ALL TESTS PASSED" +printf '\nall fm-remote-secondmate-trace-context tests passed\n' diff --git a/tests/fm-rovo-harness.test.sh b/tests/fm-rovo-harness.test.sh index c625c9e043e..0ce9acd7738 100644 --- a/tests/fm-rovo-harness.test.sh +++ b/tests/fm-rovo-harness.test.sh @@ -5,7 +5,9 @@ set -u # shellcheck source=tests/lib.sh . "$(dirname "${BASH_SOURCE[0]}")/lib.sh" -# bin/fm-harness.sh checks verified ENV markers before ancestry. A suite run +# bin/fm-harness.sh checks verified ENV markers before ancestry, but that +# ordering settles the marker layer only: a structural (comm-strength) +# ancestor of a different harness still outranks either marker. A suite run # from inside Cursor, Claude, Pi, or Grok inherits those markers, which outrank # the fake ancestry the detection cases set up. Drop the ambient markers so the # asserted verdict does not depend on which harness launched the suite. @@ -386,8 +388,15 @@ SH PATH="$fakebin:$BASE_PATH" FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh") [ "$out" = rovo ] || fail "rovo's ROVODEV_CLI marker did not outrank an inherited CLAUDECODE, got '$out'" + # CLAUDECODE alone, with no rovo marker, is a marker-layer question, not an + # ancestry one: blind the walk so the rovo-resolving fake ps above (needed + # for the markerless-ancestry and marker+ancestry cases) cannot also decide + # this assertion, matching the sibling-file pattern. + local blind_fakebin + blind_fakebin=$(fm_fakebin "$dir/blind-ancestry") + fm_fake_blind_ancestry "$blind_fakebin" out=$(env -u CURSOR_AGENT -u CURSOR_INVOKED_AS \ - CLAUDECODE=1 PATH="$fakebin:$BASE_PATH" FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh") + CLAUDECODE=1 PATH="$blind_fakebin:$BASE_PATH" FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh") [ "$out" = claude ] || fail "verified env-marker precedence changed, got '$out'" pass "fm-harness: rovo's markers outrank an inherited CLAUDECODE, and markerless ancestry still resolves rovo" } diff --git a/tests/fm-secondmate-harness.test.sh b/tests/fm-secondmate-harness.test.sh index dac5c6cbe65..527aacc7f7b 100755 --- a/tests/fm-secondmate-harness.test.sh +++ b/tests/fm-secondmate-harness.test.sh @@ -51,12 +51,11 @@ set -u . "$ROOT/bin/fm-config-inherit-lib.sh" # The harness-detection cases below fake `ps` so process ancestry is fully -# controlled, but bin/fm-harness.sh checks verified ENV markers before ancestry. -# A suite run from inside one of those harnesses inherits its marker, and the -# highest-precedence one wins over everything these cases set up: with an -# ambient CLAUDECODE=1, the pi-signed ancestry case resolves "claude". Drop the -# ambient markers so what this suite asserts does not depend on which harness it -# was launched from; every case states the marker it means to test. +# controlled, but bin/fm-harness.sh also reads verified ENV markers. A suite run +# from inside one of those harnesses inherits its marker, and it wins over +# everything these cases set up wherever ancestry is silent. Drop the ambient +# markers so what this suite asserts does not depend on which harness it was +# launched from; every case states the marker it means to test. unset CLAUDECODE PI_CODING_AGENT FM_PI_HARNESS GROK_AGENT CURSOR_AGENT CURSOR_INVOKED_AS BASE_PATH=${FM_TEST_BASE_PATH:-/usr/bin:/bin:/usr/sbin:/sbin} @@ -64,11 +63,25 @@ fm_git_identity fmtest fmtest@example.com TMP_ROOT=$(fm_test_tmproot fm-secondmate-harness) export FM_BACKEND=tmux +# Every claude launch pre-registers workspace trust for the directory it starts +# in, and for a secondmate that directory is the home (bin/fm-claude-trust.sh). +# Several cases here resolve claude, so every spawn below pins a throwaway HOME +# with an empty CLAUDE_CONFIG_DIR and puts node on the spawn's PATH; without the +# first, this suite would write the developer's real ~/.claude.json. +# Dropping the ambient markers is only half the isolation: a structural ancestor +# outranks a marker, so a case that PINS detect_own with CLAUDECODE=1 also has to +# blind the ancestry walk, or the harness this suite was launched from answers +# instead of the pin. BLIND_BIN goes AFTER a case's own fakebin in PATH, so a +# fixture that deliberately supplies its own ps or a harness-named ancestor keeps +# it (tests/fm-harness-precedence.test.sh owns the precedence boundary itself). +BLIND_BIN=$(fm_fakebin "$TMP_ROOT/blind-ancestry") +fm_fake_blind_ancestry "$BLIND_BIN" + # =========================================================================== # A) fm-harness.sh secondmate resolution + fallback (deterministic detect_own) # =========================================================================== -# detect_own is pinned to claude via CLAUDECODE=1 so the "fall through to own" -# cases are reproducible. Each row sets crew-harness / secondmate-harness in a +# detect_own is pinned to claude via CLAUDECODE=1 over a blinded ancestry walk so +# the "fall through to own" cases are reproducible on any host harness. Each row sets crew-harness / secondmate-harness in a # fresh config dir (a literal '-' means leave the file absent) and asserts BOTH # the secondmate resolution AND that crew resolution is unchanged (backward-compat). # <label>^<crew-harness>^<secondmate-harness>^<expect-secondmate>^<expect-crew> @@ -83,8 +96,8 @@ test_harness_resolution() { mkdir -p "$cfg" [ "$crew" = "-" ] || printf '%s\n' "$crew" > "$cfg/crew-harness" [ "$sm" = "-" ] || printf '%s\n' "$sm" > "$cfg/secondmate-harness" - got_sm=$(CLAUDECODE=1 FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh" secondmate) - got_crew=$(CLAUDECODE=1 FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh" crew) + got_sm=$(PATH="$BLIND_BIN:$BASE_PATH" CLAUDECODE=1 FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh" secondmate) + got_crew=$(PATH="$BLIND_BIN:$BASE_PATH" CLAUDECODE=1 FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh" crew) [ "$got_sm" = "$exp_sm" ] || fail "$label: secondmate resolved '$got_sm', expected '$exp_sm'" [ "$got_crew" = "$exp_crew" ] || fail "$label: crew resolved '$got_crew', expected '$exp_crew'" done <<'ROWS' @@ -140,9 +153,9 @@ test_secondmate_model_effort_tokens() { cfg="$case_dir/config" mkdir -p "$cfg" [ "$line" = ABSENT ] || printf '%b\n' "$line" > "$cfg/secondmate-harness" - got_h=$(CLAUDECODE=1 FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh" secondmate) - got_m=$(CLAUDECODE=1 FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh" secondmate-model) - got_e=$(CLAUDECODE=1 FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh" secondmate-effort) + got_h=$(PATH="$BLIND_BIN:$BASE_PATH" CLAUDECODE=1 FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh" secondmate) + got_m=$(PATH="$BLIND_BIN:$BASE_PATH" CLAUDECODE=1 FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh" secondmate-model) + got_e=$(PATH="$BLIND_BIN:$BASE_PATH" CLAUDECODE=1 FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh" secondmate-effort) [ "$got_h" = "$exp_harness" ] || fail "$label: harness resolved '$got_h', expected '$exp_harness'" [ "$got_m" = "$exp_model" ] || fail "$label: model resolved '$got_m', expected '$exp_model'" [ "$got_e" = "$exp_effort" ] || fail "$label: effort resolved '$got_e', expected '$exp_effort'" @@ -426,6 +439,10 @@ make_noop_tmux() { exit 0 SH chmod +x "$fakebin/tmux" + # BASE_PATH deliberately omits the developer's node, which the trust + # registration below needs, so link the real one in rather than presenting a + # node-less spawn host no real fleet member looks like. + ln -sf "$(command -v node)" "$fakebin/node" printf '%s\n' "$fakebin" } @@ -442,8 +459,8 @@ make_seeded_home() { # spawn_secondmate <world> <id> <home> [explicit-harness] # Runs fm-spawn.sh in secondmate mode. FM_ROOT is the real repo (so fm-harness.sh -# resolves), the primary config dir is <world>/home/config, and CLAUDECODE pins -# detect_own. stderr is discarded (the local-HEAD ff sync harmlessly skips a +# resolves), the primary config dir is <world>/home/config, and CLAUDECODE over a +# blinded ancestry walk pins detect_own. stderr is discarded (the local-HEAD ff sync harmlessly skips a # non-worktree home). Inspect <world>/home/state/<id>.meta and <home>/config after. spawn_secondmate() { local world=$1 id=$2 home=$3 harness=${4:-} fakebin @@ -454,8 +471,8 @@ spawn_secondmate() { local spawn_args=("$id" "$home") [ -n "$harness" ] && spawn_args+=("$harness") spawn_args+=(--secondmate) - PATH="$fakebin:$BASE_PATH" TMUX='' CLAUDECODE=1 \ - FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$world/home" \ + PATH="$fakebin:$BLIND_BIN:$BASE_PATH" TMUX='' CLAUDECODE=1 \ + FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$world/home" HOME="$world/home/user-home" CLAUDE_CONFIG_DIR='' \ FM_STATE_OVERRIDE="$world/home/state" FM_DATA_OVERRIDE="$world/home/data" \ FM_PROJECTS_OVERRIDE="$world/home/projects" FM_CONFIG_OVERRIDE="$world/home/config" \ FM_SPAWN_NO_GUARD=1 \ @@ -568,7 +585,7 @@ test_spawn_unverified_secondmate_harness_refused() { err="$w/spawn.err" rc=0 PATH="$fakebin:$BASE_PATH" TMUX='' CLAUDECODE=1 \ - FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$w/home" \ + FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$w/home" HOME="$w/home/user-home" CLAUDE_CONFIG_DIR='' \ FM_STATE_OVERRIDE="$w/home/state" FM_DATA_OVERRIDE="$w/home/data" \ FM_PROJECTS_OVERRIDE="$w/home/projects" FM_CONFIG_OVERRIDE="$w/home/config" \ FM_SPAWN_NO_GUARD=1 \ @@ -595,7 +612,7 @@ test_spawn_cursor_secondmate_launches_with_its_primary_contract() { : > "$launchlog" rc=0 PATH="$fakebin:$BASE_PATH" TMUX='' CLAUDECODE=1 \ - FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$w/home" \ + FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$w/home" HOME="$w/home/user-home" CLAUDE_CONFIG_DIR='' \ FM_STATE_OVERRIDE="$w/home/state" FM_DATA_OVERRIDE="$w/home/data" \ FM_PROJECTS_OVERRIDE="$w/home/projects" FM_CONFIG_OVERRIDE="$w/home/config" \ FM_SPAWN_NO_GUARD=1 FM_FAKE_LAUNCH_LOG="$launchlog" FM_FAKE_PANE_PATH="$sm" \ @@ -661,6 +678,10 @@ exit 0 SH chmod +x "$fakebin/tmux" fm_fake_exit0 "$fakebin" pi + # BASE_PATH deliberately omits the developer's node, which the trust + # registration below needs, so link the real one in rather than presenting a + # node-less spawn host no real fleet member looks like. + ln -sf "$(command -v node)" "$fakebin/node" printf '%s\n' "$fakebin" } @@ -673,8 +694,8 @@ spawn_secondmate_capture() { mkdir -p "$world/home/state" "$world/home/data" fakebin=$(make_launch_capturing_tmux "$world/tmux-$id") : > "$launchlog" - PATH="$fakebin:$BASE_PATH" TMUX='' CLAUDECODE=1 \ - FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$world/home" \ + PATH="$fakebin:$BLIND_BIN:$BASE_PATH" TMUX='' CLAUDECODE=1 \ + FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$world/home" HOME="$world/home/user-home" CLAUDE_CONFIG_DIR='' \ FM_STATE_OVERRIDE="$world/home/state" FM_DATA_OVERRIDE="$world/home/data" \ FM_PROJECTS_OVERRIDE="$world/home/projects" FM_CONFIG_OVERRIDE="$world/home/config" \ FM_SPAWN_NO_GUARD=1 FM_FAKE_LAUNCH_LOG="$launchlog" \ @@ -1007,6 +1028,7 @@ new_world() { [ "$dispatch_ignore" = no ] || printf 'config/crew-dispatch.json\n' printf 'config/crew-harness\nconfig/secondmate-harness\nconfig/backlog-backend\n' printf 'config/backend\nconfig/herdr-presentation-spaces\nconfig/startup-memory-budget\n' + printf 'config/claude-permission-mode\n' } > "$w/main/.gitignore" printf 'v1\n' > "$w/main/AGENTS.md" printf 'r1\n' > "$w/main/README.md" @@ -1392,6 +1414,52 @@ test_bootstrap_sweep_materializes_and_inherits_memory_default() { } # config/backend: present and absent primary state converges exactly. +# config/claude-permission-mode=auto reaches a Claude SECONDMATE launch too: the +# same template swap as a crewmate, with model/effort untouched. +test_spawn_secondmate_claude_permission_mode_auto() { + local w sm meta launchlog launch out status + w="$TMP_ROOT/spawn-claude-permmode" + sm="$w/sm" + launchlog="$w/launch.log" + mkdir -p "$w/home/config" + printf 'claude opus\n' > "$w/home/config/secondmate-harness" + printf 'auto\n' > "$w/home/config/claude-permission-mode" + make_seeded_home "$sm" sm + + out=$(spawn_secondmate_capture "$w" sm "$sm" "$launchlog" 2>&1); status=$? + expect_code 0 "$status" "claude secondmate spawn under claude-permission-mode=auto should succeed" + + meta="$w/home/state/sm.meta" + [ "$(meta_field "$meta" harness)" = claude ] || fail "permmode: meta harness not claude" + launch=$(cat "$launchlog") + assert_contains "$launch" "claude --permission-mode auto --settings '{\"feedbackDrafts\":\"off\",\"attribution\":{\"commit\":\"\",\"pr\":\"\",\"sessionUrl\":false}}' --model 'opus'" \ + "permmode: secondmate launch did not swap the permission flag while keeping --model" + assert_not_contains "$launch" "--dangerously-skip-permissions" "permmode: secondmate launch must not request bypass mode" + pass "C2b spawn: config/claude-permission-mode=auto reaches a Claude secondmate launch" +} + +# The file is a captain-wide safety preference, so it inherits like +# config/backend: present values converge exactly and primary absence mirrors. +test_claude_permission_mode_inheritance_present_and_absent() { + local w head out err status + w=$(new_world permmode-inherit) + head=$(git -C "$w/main" rev-parse HEAD) + add_sm_worktree "$w" sm "$head" + + printf 'auto\n' > "$w/home/config/claude-permission-mode" + err="$w/permmode-inherit.err" + out=$(run_config_push "$w" 2>"$err"); status=$? + expect_code 0 "$status" "claude-permission-mode present push should succeed" + assert_contains "$out" "claude-permission-mode: pushed" "present value should report pushed" + [ "$(cat "$w/sm/config/claude-permission-mode")" = auto ] || fail "claude-permission-mode present value not pushed" + + rm -f "$w/home/config/claude-permission-mode" + out=$(run_config_push "$w" 2>"$err"); status=$? + expect_code 0 "$status" "claude-permission-mode absence push should succeed" + [ -e "$w/sm/config/claude-permission-mode" ] && fail "claude-permission-mode not removed on primary absence" + pass "B12c claude-permission-mode inheritance: present values and primary absence converge exactly" +} + test_backend_inheritance_present_and_absent() { local w head out err status instruction w=$(new_world backend-inherit) @@ -2526,7 +2594,7 @@ SH chmod +x "$fakebin/rm" launchlog="$w/spawn-quarantine.launch.log" out=$(PATH="$fakebin:$BASE_PATH" TMUX='' CLAUDECODE=1 \ - FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$w/home" \ + FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$w/home" HOME="$w/home/user-home" CLAUDE_CONFIG_DIR='' \ FM_STATE_OVERRIDE="$w/home/state" FM_DATA_OVERRIDE="$w/home/data" \ FM_PROJECTS_OVERRIDE="$w/home/projects" FM_CONFIG_OVERRIDE="$w/home/config" \ FM_SPAWN_NO_GUARD=1 FM_FAKE_LAUNCH_LOG="$launchlog" \ @@ -2589,6 +2657,8 @@ test_bootstrap_sweep_propagates_when_tracked_current test_bootstrap_sweep_defers_dispatch_on_stale_unignored_home test_bootstrap_sweep_materializes_and_inherits_memory_default test_backend_inheritance_present_and_absent +test_spawn_secondmate_claude_permission_mode_auto +test_claude_permission_mode_inheritance_present_and_absent test_presentation_inheritance_default_on_and_opt_out test_bootstrap_sweep_surfaces_config_propagation_failure test_bootstrap_rereads_after_partial_propagation diff --git a/tests/fm-secondmate-liveness.test.sh b/tests/fm-secondmate-liveness.test.sh index 11b7b2222fa..aa374b0f01c 100755 --- a/tests/fm-secondmate-liveness.test.sh +++ b/tests/fm-secondmate-liveness.test.sh @@ -162,20 +162,18 @@ SH test_herdr_agent_state_preserves_husk_classifier() { local pane_state expected out - # A live registration additionally runs the stale-registration cross-check - # (the idle-shell process proof); stub it inconclusive here so this table - # keeps pinning the raw husk mapping. The cross-check's own verdicts are - # pinned in tests/fm-backend-herdr.test.sh and in the next case below. + # The pane classifier owns process-level proof. This table isolates its + # recovery mapping from any real host session named sess. for row in 'dead missing' 'no-agent dead' 'live alive' 'unknown unreadable'; do pane_state=${row%% *} expected=${row#* } - out=$(FM_TEST_PANE_STATE="$pane_state" bash -c '. "$0/bin/backends/herdr.sh"; fm_backend_herdr_pane_agent_state() { printf "%s" "$FM_TEST_PANE_STATE"; }; fm_backend_herdr_pane_agent_free_proof() { return 1; }; fm_backend_herdr_agent_state "sess:p1"' "$ROOT") + out=$(FM_TEST_PANE_STATE="$pane_state" bash -c '. "$0/bin/backends/herdr.sh"; fm_backend_herdr_pane_agent_state() { printf "%s" "$FM_TEST_PANE_STATE"; }; fm_backend_herdr_server_running_state() { printf "unknown"; }; fm_backend_herdr_agent_state "sess:p1"' "$ROOT") [ "$out" = "$expected" ] || fail "Herdr pane state $pane_state should map to $expected, got '$out'" done # A registration whose pane provably holds only a lone idle shell is stale: # the agent process is gone, so the recovery-grade verdict is dead. - out=$(bash -c '. "$0/bin/backends/herdr.sh"; fm_backend_herdr_pane_agent_state() { printf "live"; }; fm_backend_herdr_pane_agent_free_proof() { return 0; }; fm_backend_herdr_agent_state "sess:p1"' "$ROOT") + out=$(bash -c '. "$0/bin/backends/herdr.sh"; fm_backend_herdr_pane_agent_state() { printf "stale-agent"; }; fm_backend_herdr_agent_state "sess:p1"' "$ROOT") [ "$out" = dead ] || fail "a live registration over a proven lone idle shell should classify dead, got '$out'" out=$(bash -c '. "$0/bin/backends/herdr.sh"; fm_backend_herdr_agent_state "no-colon-target"' "$ROOT") diff --git a/tests/fm-send-agy-confirm.test.sh b/tests/fm-send-agy-confirm.test.sh new file mode 100755 index 00000000000..1a5a5240509 --- /dev/null +++ b/tests/fm-send-agy-confirm.test.sh @@ -0,0 +1,165 @@ +#!/usr/bin/env bash +# fm-send typed-plane submit-confirm budget for agy targets. +# +# A typed send to an explicit tmux agy endpoint is acknowledged only by the +# submit core's idle-to-busy transition poll: agy's bare `>` composer verdict +# is `unknown` (dead-shell rule), so the poll watching the pane's verified +# `esc to cancel` busy footer is the only proof a landed Enter can get. agy +# renders that footer ~1.5s after Enter for a short steer and ~4s for a +# multi-line brief (live-measured on agy 1.2.1), while the shared default +# confirm budget is 3 retries x 0.4s - so fm-send used to exit 1 "not +# submitted" for a message that landed and ran, inviting a duplicate resend. +# fm-send now gives agy typed targets a longer default budget (20 retries, +# ~8s at the default cadence); an explicit FM_SEND_RETRIES still wins and every +# other harness keeps the +# shared 3-retry default. These tests pin that behavior hermetically (stubbed +# tmux + sleep, no real agent): the fake tmux renders the busy footer only +# from the BUSY_AT-th plain pane capture, so the number of logged 0.4s waits +# stands in for wall-clock latency and each case is deterministic: +# 1. agy target, busy footer at the 5th poll (short-steer latency): the send +# exits 0 (confirmed idle-to-busy), and the sleep log shows the poll +# reaching that read. +# 2. agy target, busy footer only at the 15th poll (the live-measured long +# brief latency): the default budget still reaches it and exits 0. +# 3. agy target with an explicit FM_SEND_RETRIES=3: the operator knob wins +# and the send keeps the loud exit-1 verdict=unknown refusal. +# 4. agy target whose busy footer never renders: no confirmation is +# fabricated - exit 1 verdict=unknown. +# 5. claude target, same late-busy pane: the shared 3-retry default is +# untouched, so the send still exits 1 verdict=unknown. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +SEND="$ROOT/bin/fm-send.sh" + +TMP_ROOT=$(fm_test_tmproot fm-send-agy-confirm) + +# A fake tmux that models agy's late busy render, plus a fake sleep that +# records every requested duration (one per line) into FM_SLEEP_LOG instead of +# sleeping. The styled capture (-e) always shows agy's idle bare-`>` composer +# (verdict `unknown`); the plain capture - the one fm_pane_busy_state polls - +# shows the idle screen until its BUSY_AT-th call and the verified `esc to +# cancel` busy row from then on. The BUSY_AT threshold is read from the +# per-case dir so cases are independent. +make_stubs() { # <dir> <busy-at> -> echoes fakebin dir + local dir=$1 busy_at=$2 fb="$1/fakebin" + mkdir -p "$fb" + cat > "$fb/tmux" <<SH +#!/usr/bin/env bash +set -u +cnt_file="$dir/plain.count" +case "\${1:-}" in + send-keys) exit 0 ;; + display-message) + for a in "\$@"; do + case "\$a" in + *cursor_y*) printf '0\n'; exit 0 ;; + *pane_tty*) printf '\n'; exit 0 ;; + esac + done + printf 'fakepane\n'; exit 0 ;; + capture-pane) + styled=0 + for a in "\$@"; do [ "\$a" = -e ] && styled=1; done + if [ "\$styled" = 1 ]; then + printf '> \n? for shortcuts\n' + exit 0 + fi + n=\$(( \$(cat "\$cnt_file" 2>/dev/null || echo 0) + 1 )) + printf '%s' "\$n" > "\$cnt_file" + if [ "\$n" -ge $busy_at ]; then + printf '> \n? for shortcuts\n ⏺ 5s · esc to cancel · gemini-3.8-flash-low\n' + else + printf '> \n? for shortcuts\n' + fi + exit 0 ;; + list-windows) printf 'win\n'; exit 0 ;; +esac +exit 0 +SH + chmod +x "$fb/tmux" + cat > "$fb/sleep" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "${1:-}" >> "$FM_SLEEP_LOG" +exit 0 +SH + chmod +x "$fb/sleep" + printf '%s\n' "$fb" +} + +# run_send <harness> <busy-at> <extra-env-assignments...>: build a fresh home +# whose recorded task targets sess:win on the tmux backend with <harness> meta, +# then run the real fm-send typed plane against the stubs. FM_SEND_SETTLE=0 +# strips the post-submit pause so the sleep log holds only the popup settle +# plus the 0.4 submit waits, keeping the poll arithmetic visible. FM_ROOT_OVERRIDE +# points at the case dir so fm-guard's tangle check stays silent. Emits +# "rc <exit>" and leaves the send's stderr in $dir/err and the sleep log in +# $dir/sleep.log for the caller to assert on. +run_send() { # <harness> <busy-at> [env=val ...] + local harness=$1 busy_at=$2 dir fb log + shift 2 + dir="$TMP_ROOT/case-$RANDOM-$RANDOM"; mkdir -p "$dir/state" + fb=$(make_stubs "$dir" "$busy_at") + log="$dir/sleep.log"; : > "$log" + fm_write_meta "$dir/state/agyw.meta" "window=sess:win" "harness=$harness" + ( + export FM_GATE_REFUSE_BYPASS=1 FM_SEND_SETTLE=0 + export PATH="$fb:$PATH" FM_ROOT_OVERRIDE="$dir" FM_HOME="$dir" FM_SLEEP_LOG="$log" + for a in "$@"; do eval "export $a"; done + "$SEND" sess:win 'Append steer1 line to notes.md' 2>"$dir/err" + printf 'rc %s\n' "$?" + ) +} + +# agy, default budget, busy footer renders at the 5th poll (the 6th plain +# capture): the raised default (20 retries) must reach that read and exit 0. +# Under the old shared default (3 retries) this exact shape exited 1 +# "verdict=unknown" - the regression this suite pins. +out=$(run_send agy 6) +expect_code 0 "$(printf '%s' "$out" | sed -n 's/^rc //p')" \ + "agy typed send with late busy footer confirms idle-to-busy and exits 0" +grep -q 'not submitted' "$TMP_ROOT"/*/err 2>/dev/null && \ + fail "agy typed send: refusal text present despite confirmed submit" +pass "agy typed send: no not-submitted refusal on confirmed idle-to-busy" +case_dir=$(printf '%s\n' "$TMP_ROOT"/case-* | head -1) +settles=$(grep -cv '^0\.4$' "$case_dir/sleep.log" || true) +waits=$(grep -c '^0\.4$' "$case_dir/sleep.log" || true) +[ "$settles" = 1 ] || fail "agy typed send: expected exactly 1 non-wait sleep (popup settle), got $settles" +[ "$waits" = 5 ] || fail "agy typed send: expected the poll to reach the 5th busy read (5 x 0.4s: Enter wait + 4 poll waits), got $waits" +pass "agy typed send: sleep log shows the confirm poll running to the late busy render" + +# agy with an explicit FM_SEND_RETRIES=3: the operator knob wins over the agy +# default, the budget expires before the late footer, and the loud refusal +# boundary is preserved. +out=$(run_send agy 6 'FM_SEND_RETRIES=3') +expect_code 1 "$(printf '%s' "$out" | sed -n 's/^rc //p')" \ + "agy typed send honors an explicit FM_SEND_RETRIES=3" +grep -q 'verdict=unknown' "$TMP_ROOT"/*/err || fail "agy typed send FM_SEND_RETRIES=3: expected verdict=unknown refusal" +pass "agy typed send: explicit FM_SEND_RETRIES=3 keeps the exit-1 verdict=unknown refusal" + +# agy whose busy footer never renders: the raised budget must time out into +# the same loud refusal, never fabricate a confirmation. +out=$(run_send agy 999) +expect_code 1 "$(printf '%s' "$out" | sed -n 's/^rc //p')" \ + "agy typed send with no busy footer refuses exit 1" +grep -q 'verdict=unknown' "$TMP_ROOT"/*/err || fail "agy typed send never-busy: expected verdict=unknown refusal" +pass "agy typed send: never-rendering busy footer still refuses with verdict=unknown" + +# agy, busy footer renders only at the 15th poll: the live long-brief case. +# The default budget must still reach that read and exit 0; under the shared +# 3-retry default this shape refused for a message that landed. +out=$(run_send agy 16) +expect_code 0 "$(printf '%s' "$out" | sed -n 's/^rc //p')" \ + "agy typed send with long-brief late busy footer confirms and exits 0" +pass "agy typed send: long-brief render (15th poll) still confirms idle-to-busy" + +# claude on the identical late-busy pane: the shared 3-retry default is +# untouched, so the same latency still refuses - the raised budget is +# agy-scoped, not a global slowdown. +out=$(run_send claude 6) +expect_code 1 "$(printf '%s' "$out" | sed -n 's/^rc //p')" \ + "claude typed send keeps the shared 3-retry default" +grep -q 'verdict=unknown' "$TMP_ROOT"/*/err || fail "claude typed send: expected verdict=unknown refusal" +pass "claude typed send: late busy footer still refuses (agy budget is agy-scoped)" diff --git a/tests/fm-session-lock-ancestry.test.sh b/tests/fm-session-lock-ancestry.test.sh index 0c86e17a378..613fe085f07 100755 --- a/tests/fm-session-lock-ancestry.test.sh +++ b/tests/fm-session-lock-ancestry.test.sh @@ -87,6 +87,52 @@ SH pass "session-lock: a version-named Claude Code session is identified from its install path and argv[0]" } +# A harness that is pid 1 of its own PID namespace - a container, or the +# `codex sandbox` this shape was verified in - used to be invisible: the walk +# stopped as soon as the NEXT pid was 1, so the one process that identifies the +# session was never examined and the session could not recognize its own lock. +test_harness_at_namespace_pid1_is_examined() { + local dir fakebin got + dir="$TMP_ROOT/namespace-pid1" + fakebin=$(fm_fakebin "$dir") + mkdir -p "$dir/state" + cat > "$fakebin/ps" <<'SH' +#!/usr/bin/env bash +set -u +field= pid= +while [ "$#" -gt 0 ]; do + case "$1" in + -o) field=$2; shift 2 ;; + -p) pid=$2; shift 2 ;; + *) shift ;; + esac +done +case "$pid:$field" in + 1:comm=) printf '%s\n' "${FM_TEST_PID1_COMM:-claude}" ;; + 1:args=) printf '%s\n' "${FM_TEST_PID1_COMM:-claude}" ;; + 1:ppid=) printf '%s\n' 0 ;; + *:comm=) printf '%s\n' bash ;; + *:args=) printf '%s\n' 'bash /repo/bin/fm-watch.sh' ;; + *:ppid=) printf '%s\n' 1 ;; +esac +SH + chmod +x "$fakebin/ps" + printf '1\n' > "$dir/state/.lock" + + # Non-vacuity: with a host-shaped pid 1 the same table must find nothing, so + # this case cannot pass by the walk matching everything it reaches. + if FM_TEST_PID1_COMM=systemd lib_eval "$fakebin" 'fm_harness_ancestry_pid' >/dev/null 2>&1; then + fail "a host-shaped pid 1 was read as a harness process" + fi + + got=$(lib_eval "$fakebin" 'fm_harness_ancestry_pid') \ + || fail "the harness at namespace pid 1 was not found in the ancestry at all" + [ "$got" = 1 ] || fail "ancestry resolved '$got', expected the namespace harness pid 1" + lib_eval "$fakebin" "fm_session_lock_owned_by_self '$dir/state'" \ + || fail "the session holding the lock at namespace pid 1 did not recognize itself as the owner" + pass "session-lock: a harness that is pid 1 of its own namespace is examined, not skipped" +} + test_ordinary_paths_are_never_harness_processes() { local dir fakebin shape dir="$TMP_ROOT/ordinary-paths" @@ -359,6 +405,7 @@ test_e2e_daemon_parented_version_named_session_keeps_its_lock() { } test_version_named_session_is_identified_on_both_platforms +test_harness_at_namespace_pid1_is_examined test_ordinary_paths_are_never_harness_processes test_harness_beyond_a_gap_never_owns_the_lock test_competing_version_named_session_is_seen_as_live diff --git a/tests/fm-session-start.test.sh b/tests/fm-session-start.test.sh index 87e8316dce7..72b39080091 100755 --- a/tests/fm-session-start.test.sh +++ b/tests/fm-session-start.test.sh @@ -217,10 +217,15 @@ make_fake_ps_claude() { make_fake_ps_harness() { local fakebin=$1 harness=$2 - cat > "$fakebin/ps" <<'SH' + cat > "$fakebin/ps" <<SH #!/usr/bin/env bash set -u -harness=${FM_FAKE_HARNESS:-claude} +# The ancestry this stub reports defaults to the harness the fixture was built +# for, so a case that builds a pi (or codex) fixture gets pi (or codex) ancestry +# without having to repeat it per run; FM_FAKE_HARNESS still overrides it. +harness=\${FM_FAKE_HARNESS:-$harness} +SH + cat >> "$fakebin/ps" <<'SH' pid= previous= for argument in "$@"; do @@ -1763,7 +1768,7 @@ EOF assert_not_contains "$out" "DONE-ROW-LINE" "tasks-axi compact digest listed a done row at startup" assert_contains "$out" "--- compact-startup ---" "in-flight meta identity disappeared from startup recovery digest" assert_contains "$out" "worktree=$home/projects/firstmate" "in-flight recovery worktree identity disappeared from startup digest" - assert_contains "$out" "Full task bodies remain available on demand: tasks-axi show <id> --full" \ + assert_contains "$out" "Full task bodies remain available on demand: bin/fm-tasks-axi.sh show <id> --full" \ "compact digest omitted the full-body lookup pointer" assert_contains "$out" "ready_public_followups: 0 delivery-ready obligations" \ "the composed listing dropped a real signal from the dispatchable set" @@ -1807,7 +1812,7 @@ EOF assert_not_contains "$out" "ready-4,queued" "the queued bound did not actually bound the ready listing" assert_contains "$out" "(shown 3 of 7 ready queued item(s))" \ "the bounded queued listing did not report what it showed" - assert_contains "$out" "(4 more queued - tasks-axi ready --file $home/data/backlog.md)" \ + assert_contains "$out" "(4 more queued - bin/fm-tasks-axi.sh ready)" \ "the bounded queued listing did not disclose an exact remainder and how to see it" # The bound is for dispatchable work only: held and blocked rows stay whole. @@ -2434,6 +2439,48 @@ EOF pass "next step delegates watcher ownership to the AFK daemon" } +test_next_step_quiet_mode_delegates_to_daemon() { + local rec root home fakebin out + rec=$(new_world next-step-quiet) + IFS='|' read -r root home fakebin <<EOF +$rec +EOF + make_fake_toolchain "$fakebin" + make_fake_ps_claude "$fakebin" + printf 'quiet\n%s\n' "$(date '+%s')" > "$home/state/.afk" + + out=$(run_session_start "$home" "$root" "$fakebin:$BASE_PATH") + + assert_contains "$out" "quiet-mode supervision is active" "AFK digest did not report quiet mode for a quiet-content flag" + assert_contains "$out" "only an explicit /quiet off exits it" "AFK digest lost the explicit-only exit rule" + assert_contains "$out" "Quiet mode is active" "next step did not switch to quiet-mode guidance" + assert_contains "$out" "load /quiet" "next step did not name the /quiet skill" + assert_contains "$out" "- Quiet mode: active" "supervision block did not include active quiet state" + assert_not_contains "$out" "Away mode is active" "quiet-mode flag was misreported as away mode" + assert_not_contains "$out" " bin/fm-watch-arm.sh" "quiet next step still told the agent to arm the watcher directly" + + pass "next step delegates watcher ownership to the daemon in quiet mode, distinctly from away mode" +} + +test_next_step_afk_legacy_empty_flag_defaults_away() { + local rec root home fakebin out + rec=$(new_world next-step-afk-legacy) + IFS='|' read -r root home fakebin <<EOF +$rec +EOF + make_fake_toolchain "$fakebin" + make_fake_ps_claude "$fakebin" + : > "$home/state/.afk" + + out=$(run_session_start "$home" "$root" "$fakebin:$BASE_PATH") + + assert_contains "$out" "away-mode supervision is active" "a legacy empty .afk flag was not read as away mode" + assert_contains "$out" "Away mode is active" "a legacy empty .afk flag did not drive away-mode next-step guidance" + assert_not_contains "$out" "Quiet mode" "a legacy empty .afk flag leaked quiet-mode text" + + pass "a legacy empty .afk flag (written before mode existed) still reads as away mode" +} + test_supervision_block_exactly_one_and_pi_diagnostic() { local rec root home fakebin out block_count wake_line sup_line context_line rec=$(new_world pi-supervision-block) @@ -2668,6 +2715,8 @@ test_backlog_compact_tasks_axi_unavailable_uses_manual_fallback test_fleet_digest_empty_fleet test_next_step_sources_x_mode_cadence test_next_step_afk_delegates_to_daemon +test_next_step_quiet_mode_delegates_to_daemon +test_next_step_afk_legacy_empty_flag_defaults_away test_supervision_block_exactly_one_and_pi_diagnostic test_pi_signed_primary_uses_pi_extensions_without_identity_normalization test_pi_diagnostic_rejects_stale_loaded_marker diff --git a/tests/fm-sessionstart-nudge.test.sh b/tests/fm-sessionstart-nudge.test.sh index ae9f9a2fcef..6ea27b2845c 100755 --- a/tests/fm-sessionstart-nudge.test.sh +++ b/tests/fm-sessionstart-nudge.test.sh @@ -131,6 +131,40 @@ test_owned_lock_is_silent() { pass "fm-sessionstart-nudge: a lock holder in process ancestry is already run" } +# A firstmate running inside a PID namespace - a container, or `codex sandbox` - +# holds its home lock from a harness that IS pid 1, so this hook must recognize +# that owner. The old walk rejected a lock pid of 1 outright and stopped before +# comparing pid 1, so the hook nudged a session that had already run. +# A fake ps cannot reach this path: `kill -0` is a shell builtin gating the lock +# pid, and on a host `kill -0 1` fails for an unprivileged user, so the case +# needs a real namespace where pid 1 is this user's own process. +test_namespace_pid1_lock_holder_is_silent() { + local root="$TMP_ROOT/namespace-pid1" out status=0 + if ! command -v unshare >/dev/null 2>&1 \ + || ! unshare -rpf --mount-proc true >/dev/null 2>&1; then + printf '# skip: unprivileged PID namespaces are unavailable here, so the namespace pid 1 lock owner is unverified\n' + return 0 + fi + make_primary "$root" + + # Non-vacuity: inside the same namespace, with no lock at all, the hook must + # still produce its nudge, so silence below means the owner was recognized. + out=$(unshare -rpf --mount-proc bash -c \ + "FM_GATE_REFUSE_BYPASS=0 FM_ROOT_OVERRIDE='$root' FM_HOME='$root' '$NUDGE'; exit \$?") || status=$? + expect_code 0 "$status" "namespace nudge without a lock" + [ "$out" = "$NUDGE_LINE" ] \ + || fail "the namespace fixture did not nudge without a lock, so its silence proves nothing: $out" + + printf '1\n' > "$root/state/.lock" + status=0 + out=$(unshare -rpf --mount-proc bash -c \ + "FM_GATE_REFUSE_BYPASS=0 FM_ROOT_OVERRIDE='$root' FM_HOME='$root' '$NUDGE'; exit \$?") || status=$? + expect_code 0 "$status" "namespace pid 1 lock nudge" + [ -z "$out" ] \ + || fail "a lock held by the harness at namespace pid 1 was not recognized, got: $out" + pass "fm-sessionstart-nudge: a lock holder that is pid 1 of its own namespace is already run" +} + test_opencode_plugin_delivers_exact_nudge_once() { local root="$TMP_ROOT/opencode-primary" out status=0 make_primary "$root" @@ -1025,6 +1059,7 @@ test_unmarked_linked_worktree_is_silent test_linked_secondmate_primary_nudges test_missing_state_is_silent test_owned_lock_is_silent +test_namespace_pid1_lock_holder_is_silent test_opencode_plugin_delivers_exact_nudge_once test_run_startup_runs_the_full_digest test_run_clear_and_compact_reemit diff --git a/tests/fm-spawn-dispatch-profile.test.sh b/tests/fm-spawn-dispatch-profile.test.sh index 29541f6c513..146f67db65b 100755 --- a/tests/fm-spawn-dispatch-profile.test.sh +++ b/tests/fm-spawn-dispatch-profile.test.sh @@ -1569,6 +1569,99 @@ SH pass "fm-spawn: actual ship/scout launch commands deliver the worker role contract" } +# config/claude-permission-mode (bin/fm-spawn.sh header): absent and `bypass` +# must both produce today's launch byte-for-byte, `auto` swaps only the +# permission flag, and any other token refuses before endpoint or metadata. +claude_expected_launch() { # <home> <id> <permission-flag> + local home=$1 id=$2 flag=$3 + printf '%s' "env -u CURSOR_AGENT -u CURSOR_INVOKED_AS -u GEMINI_CLI CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude $flag --settings '{\"feedbackDrafts\":\"off\",\"attribution\":{\"commit\":\"\",\"pr\":\"\",\"sessionUrl\":false}}' \"\$('${ROOT}/bin/fm-operational-input.sh' encode launch-brief < '$home/data/$id/launch-brief.md')\"" +} + +test_claude_permission_mode_bypass_matches_absent_launch() { + local rec id out status launch expected + id=permmode-bypass-z19 + rec=$(make_spawn_case permmode-bypass claude "$id") + read_case_record "$rec" + printf 'bypass\n' > "$HOME_DIR/config/claude-permission-mode" + + out=$(run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR") + status=$? + expect_code 0 "$status" "claude spawn with claude-permission-mode=bypass should succeed" + launch=$(cat "$LAUNCH_LOG") + expected=$(claude_expected_launch "$HOME_DIR" "$id" --dangerously-skip-permissions) + [ "$launch" = "$expected" ] || fail "explicit bypass did not reproduce the absent-file launch"$'\n'"expected: $expected"$'\n'"actual: $launch" + pass "config/claude-permission-mode=bypass launches exactly as an absent file does" +} + +test_claude_permission_mode_auto_swaps_only_the_permission_flag() { + local rec id out status launch expected + id=permmode-auto-z20 + rec=$(make_spawn_case permmode-auto claude "$id") + read_case_record "$rec" + # Surrounding whitespace is trimmed, so an editor's trailing newline or indent is fine. + printf ' auto\n' > "$HOME_DIR/config/claude-permission-mode" + + out=$(run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR") + status=$? + expect_code 0 "$status" "claude spawn with claude-permission-mode=auto should succeed" + assert_contains "$out" "spawned $id harness=claude" "auto spawn did not report claude" + launch=$(cat "$LAUNCH_LOG") + expected=$(claude_expected_launch "$HOME_DIR" "$id" '--permission-mode auto') + [ "$launch" = "$expected" ] || fail "auto changed more than the permission flag"$'\n'"expected: $expected"$'\n'"actual: $launch" + assert_not_contains "$launch" "--dangerously-skip-permissions" "auto launch must not request bypass mode" + pass "config/claude-permission-mode=auto replaces --dangerously-skip-permissions with --permission-mode auto" +} + +test_claude_permission_mode_auto_reaches_scout_launch() { + local rec id out status launch + id=permmode-scout-z21 + rec=$(make_spawn_case permmode-scout claude "$id") + read_case_record "$rec" + printf 'auto\n' > "$HOME_DIR/config/claude-permission-mode" + + out=$(run_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --scout) + status=$? + expect_code 0 "$status" "claude scout spawn with claude-permission-mode=auto should succeed" + launch=$(cat "$LAUNCH_LOG") + assert_contains "$launch" "claude --permission-mode auto --settings" "scout launch did not carry --permission-mode auto" + assert_not_contains "$launch" "--dangerously-skip-permissions" "scout launch must not request bypass mode" + pass "config/claude-permission-mode=auto reaches scout launches too" +} + +test_claude_permission_mode_invalid_refuses_before_endpoint_or_metadata() { + local rec id out status + id=permmode-invalid-z22 + rec=$(make_spawn_case permmode-invalid claude "$id") + read_case_record "$rec" + printf 'yolo\n' > "$HOME_DIR/config/claude-permission-mode" + + out=$(run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR") + status=$? + expect_code 1 "$status" "an unrecognized claude-permission-mode token must refuse the spawn" + assert_contains "$out" "config/claude-permission-mode holds 'yolo'" "refusal must name the file and the offending token" + assert_contains "$out" "bypass" "refusal must list bypass as an accepted value" + assert_contains "$out" "--permission-mode auto" "refusal must list auto as an accepted value" + [ ! -s "$LAUNCH_LOG" ] || fail "an invalid permission mode must launch nothing (got: $(cat "$LAUNCH_LOG"))" + assert_absent "$HOME_DIR/state/$id.meta" "refusal must happen before meta is written" + pass "an unrecognized config/claude-permission-mode token refuses before any endpoint or metadata" +} + +test_non_claude_harness_ignores_claude_permission_mode() { + local rec id out status launch + id=permmode-codex-z23 + rec=$(make_spawn_case permmode-codex codex "$id") + read_case_record "$rec" + printf 'auto\n' > "$HOME_DIR/config/claude-permission-mode" + + out=$(run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness codex) + status=$? + expect_code 0 "$status" "codex spawn under claude-permission-mode=auto should succeed" + launch=$(cat "$LAUNCH_LOG") + assert_contains "$launch" "codex " "codex launch did not run codex" + assert_not_contains "$launch" "--permission-mode" "the claude permission flag must not leak into a codex launch" + pass "config/claude-permission-mode changes claude launches only" +} + test_worker_launch_delivers_role_scope test_no_profile_keeps_claude_profile_defaults test_non_cursor_launch_clears_inherited_cursor_markers @@ -1605,6 +1698,12 @@ test_pi_signed_persistent_secondmate_uses_pi_extensions_and_identity test_batch_forwards_shared_profile_flags test_claude_forwards_firstmate_config_dir_when_set test_claude_default_uses_home_config_and_records_evidence_store +test_claude_omits_config_dir_prefix_when_unset +test_claude_permission_mode_bypass_matches_absent_launch +test_claude_permission_mode_auto_swaps_only_the_permission_flag +test_claude_permission_mode_auto_reaches_scout_launch +test_claude_permission_mode_invalid_refuses_before_endpoint_or_metadata +test_non_claude_harness_ignores_claude_permission_mode test_non_claude_harness_ignores_config_dir test_claude_crewmate_launch_carries_the_attribution_policy test_claude_secondmate_launch_carries_the_attribution_policy diff --git a/tests/fm-spawn-pool-base-freshen.test.sh b/tests/fm-spawn-pool-base-freshen.test.sh index aeb5a193149..2d39f3607ca 100755 --- a/tests/fm-spawn-pool-base-freshen.test.sh +++ b/tests/fm-spawn-pool-base-freshen.test.sh @@ -676,7 +676,75 @@ test_stale_pin_beside_other_dirt_reports_one_verdict() { pass "a stale pin beside other dirt yields the conservative refusal alone, with no stale-pin line" } +# Re-lay a case's pooled worktree as a managed Treehouse slot: <pool>/<slot>/<repo> +# with the pool's state file beside the slot, which is the shape fm-spawn claims +# for its task. Rewrites POOL_DIR to the relocated checkout. +lay_out_as_pool_slot() { + local slot_root="$CASE_DIR/slots" + mkdir -p "$slot_root/1" + git -C "$PROJECT_DIR" worktree move "$POOL_DIR" "$slot_root/1/project" + printf '{"worktrees":[{"name":"1","path":"%s"}]}\n' "$slot_root/1/project" \ + > "$slot_root/treehouse-state.json" + POOL_DIR="$slot_root/1/project" + SLOT_CLAIM="$slot_root/1/.fm-slot-owner" +} + +# The spawn side of the slot-owner claim that bin/fm-teardown.sh later reads: +# a launched task's claim names it, a slot that cannot be claimed refuses before +# anything is published, and an abort while the allocation lock is still held +# leaves no claim naming a task with no record. +test_pool_slot_claim_follows_the_spawn_outcome() { + local rec id out status before + + id='pool-slot-claim-r1' + rec=$(make_case slot-claim "$id") + read_case_record "$rec" + lay_out_as_pool_slot + out=$(run_spawn "$id" --scout) + status=$? + expect_code 0 "$status" "spawn from a Treehouse slot should launch"$'\n'"$out" + assert_grep "worktree=$POOL_DIR" "$HOME_DIR/state/$id.meta" \ + "spawn did not publish the relocated slot as its worktree" + [ -f "$SLOT_CLAIM" ] || fail "spawn left its Treehouse slot unclaimed: $out" + grep -Fxq -- "task=$id" "$SLOT_CLAIM" \ + || fail "the slot claim does not name the spawned task: $(cat "$SLOT_CLAIM")" + grep -Fxq -- "home=$HOME_DIR" "$SLOT_CLAIM" \ + || fail "the slot claim does not name the spawning home: $(cat "$SLOT_CLAIM")" + + id='pool-slot-unclaimable-r1' + rec=$(make_case slot-unclaimable "$id") + read_case_record "$rec" + lay_out_as_pool_slot + mkdir -p "$SLOT_CLAIM" + before=$(git -C "$POOL_DIR" rev-parse HEAD) + out=$(run_spawn "$id" --scout) + status=$? + [ "$status" -ne 0 ] || fail "spawn launched a worker on a slot it could not claim" + assert_contains "$out" "could not claim Treehouse pool slot" \ + "spawn did not name the unclaimable slot as the reason" + [ -d "$SLOT_CLAIM" ] || fail "spawn replaced the directory blocking its slot claim" + [ ! -e "$HOME_DIR/state/$id.meta" ] || fail "spawn published a record for an unclaimable slot" + [ "$(git -C "$POOL_DIR" rev-parse HEAD)" = "$before" ] \ + || fail "spawn moved the slot's HEAD after failing to claim it" + + id='pool-slot-claim-aborted-r1' + rec=$(make_originless_case slot-claim-aborted "$id") + read_case_record "$rec" + lay_out_as_pool_slot + git -C "$POOL_DIR" config remote.origin.fetch '+refs/heads/*:refs/remotes/origin/*' + out=$(run_spawn "$id" --mode no-mistakes --yolo off) + status=$? + [ "$status" -ne 0 ] || fail "spawn succeeded despite an unusable origin on the slot" + assert_contains "$out" "could not fetch origin" \ + "the aborted spawn did not refuse on its unusable origin" + [ ! -e "$HOME_DIR/state/$id.meta" ] || fail "the aborted spawn published task metadata" + [ ! -e "$SLOT_CLAIM" ] && [ ! -L "$SLOT_CLAIM" ] \ + || fail "the aborted spawn left a slot claim naming a task with no record: $(cat "$SLOT_CLAIM")" + pass "a Treehouse slot claim names the launched task, refuses when unclaimable, and is dropped by a locked abort" +} + test_remote_seeded_home_spawns_from_treehouse_pool +test_pool_slot_claim_follows_the_spawn_outcome test_linked_spawning_home_rejects_primary_before_refresh test_stale_pool_base_refreshes_before_branching test_non_main_default_branch_refreshes_before_branching diff --git a/tests/fm-supervision-instructions.test.sh b/tests/fm-supervision-instructions.test.sh index 9d996cf37c5..a6953660cd6 100755 --- a/tests/fm-supervision-instructions.test.sh +++ b/tests/fm-supervision-instructions.test.sh @@ -42,6 +42,31 @@ test_conditional_stanzas() { pass "renderer includes read-only, afk, and effective x-mode current-state stanzas" } +test_quiet_mode_stanzas() { + local home config out + home="$TMP_ROOT/quiet-home" + config="$TMP_ROOT/quiet-config" + mkdir -p "$home/state" "$config" + out=$(FM_HOME="$home" FM_CONFIG_OVERRIDE="$config" "$RENDER" --harness codex --afk 1 --afk-mode quiet) + assert_contains "$out" "- Quiet mode: active" "quiet stanza missing" + assert_contains "$out" "load /quiet" "quiet stanza did not name the /quiet skill" + assert_contains "$out" "Ordinary captain chat does NOT exit it" "quiet stanza lost the explicit-only exit rule" + assert_not_contains "$out" "- Away mode: active" "quiet mode incorrectly rendered as away mode" + out=$(FM_HOME="$home" "$RENDER" --harness codex --afk 1 --afk-mode quiet --repair-line) + assert_contains "$out" "Quiet mode owns watcher supervision; load /quiet" "quiet repair line did not name /quiet" + + out=$(FM_HOME="$home" "$RENDER" --harness codex --afk 1) + assert_contains "$out" "- Away mode: active" "omitting --afk-mode did not default to away (regression)" + assert_not_contains "$out" "Quiet mode" "omitting --afk-mode leaked quiet-mode text" + + out=$(FM_HOME="$home" "$RENDER" --harness codex --afk 1 --afk-mode not-a-real-mode) + assert_contains "$out" "- Away mode: active" "unrecognized --afk-mode value did not fall back to away" + + out=$(FM_HOME="$home" "$RENDER" --harness codex --afk 0) + assert_contains "$out" "- Away/quiet mode: inactive" "inactive stanza missing" + pass "renderer's away/quiet stanzas are mode-aware, default to away, and fall back safely on garbage input" +} + test_repair_lines() { local home out home="$TMP_ROOT/repair-home" @@ -196,6 +221,7 @@ test_pi_snippet_uses_effective_extension_path() { test_selected_harness_block_only test_unknown_fallback test_conditional_stanzas +test_quiet_mode_stanzas test_repair_lines test_cross_harness_ordinary_continuation_and_repair_matrix test_pi_signed_preserves_identity_with_pi_supervision_protocol diff --git a/tests/fm-tasks-axi.test.sh b/tests/fm-tasks-axi.test.sh new file mode 100755 index 00000000000..5ceeabea836 --- /dev/null +++ b/tests/fm-tasks-axi.test.sh @@ -0,0 +1,230 @@ +#!/usr/bin/env bash +# Behavior tests for bin/fm-tasks-axi.sh home addressing and bootstrap's +# shadow-backlog check, over the split layout where the operational home lives +# outside the code root that carries the tracked .tasks.toml. +# +# The fork these guard against: .tasks.toml names data/backlog.md relative to +# the caller's working directory, and tasks-axi writes by renaming a temp file +# over its target, so a bare tasks-axi run from the code root turns a code-root +# symlink into the home's backlog into a private regular copy. The suite proves +# that every write through bin/fm-tasks-axi.sh lands in $FM_HOME/data from the +# code root (including archiving and relative --body-file arguments), +# that the command refuses addressing it cannot keep correct, and that bootstrap +# reports any code-root copy that is not this home's own file while staying +# silent for a link into the home, an absent copy, and the single-home layout. +set -u + +# shellcheck source=tests/lib.sh disable=SC1091 +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +WRAPPER="$ROOT/bin/fm-tasks-axi.sh" +BOOTSTRAP="$ROOT/bin/fm-bootstrap.sh" +TMP_ROOT=$(fm_test_tmproot fm-tasks-axi) +BASE_PATH=${FM_TEST_BASE_PATH:-/usr/bin:/bin:/usr/sbin:/sbin} + +# The developer shell may pin any of these; each case states its own layout. +unset TASKS_AXI_FILE TASKS_AXI_BACKEND FM_HOME FM_ROOT_OVERRIDE \ + FM_DATA_OVERRIDE FM_STATE_OVERRIDE FM_CONFIG_OVERRIDE FM_PROJECTS_OVERRIDE + +HAVE_TASKS_AXI=0 +command -v tasks-axi >/dev/null 2>&1 && HAVE_TASKS_AXI=1 + +empty_backlog() { # <path> + printf '## In flight\n\n## Queued\n\n## Done\n' > "$1" +} + +# A code root carrying the tracked .tasks.toml and an operational home beside +# it, with the code-root backlog linked into the home the way an operator +# would try to keep the two in sync. +make_split() { # <name>; prints the case directory + local dir="$TMP_ROOT/$1" + mkdir -p "$dir/code/data" "$dir/home/data" "$dir/home/state" "$dir/home/config" + cp "$ROOT/.tasks.toml" "$dir/code/.tasks.toml" + empty_backlog "$dir/home/data/backlog.md" + ln -s "$dir/home/data/backlog.md" "$dir/code/data/backlog.md" + printf '%s\n' "$dir" +} + +# Run the wrapper from the code root, as firstmate does. +wrapper_from_code() { # <case-dir> <tasks-axi args...> + local dir=$1 + shift + (cd "$dir/code" && FM_HOME="$dir/home" FM_ROOT_OVERRIDE="$dir/code" "$WRAPPER" "$@") +} + +# Only the shadow-backlog lines matter here; the rest of a detect-only local +# bootstrap pass reports this host's toolchain, which is not under test, so it +# runs on the bare base PATH where every tool probe is a fast miss. +bootstrap_backlog_lines() { # <code-root> [<home>] + local code=$1 home=${2:-} + if [ -n "$home" ]; then + PATH="$BASE_PATH" FM_HOME="$home" FM_ROOT_OVERRIDE="$code" FM_BOOTSTRAP_DETECT_ONLY=1 \ + FM_BOOTSTRAP_NETWORK=skip "$BOOTSTRAP" 2>&1 | grep '^BACKLOG_RECONCILE: code-root' || true + else + PATH="$BASE_PATH" FM_ROOT_OVERRIDE="$code" FM_BOOTSTRAP_DETECT_ONLY=1 \ + FM_BOOTSTRAP_NETWORK=skip "$BOOTSTRAP" 2>&1 | grep '^BACKLOG_RECONCILE: code-root' || true + fi +} + +test_guard_reports_regular_code_root_backlog() { + local dir out + dir=$(make_split guard-regular) + out=$(bootstrap_backlog_lines "$dir/code" "$dir/home") + assert_equals "" "$out" "a code-root link into this home must stay silent" + + rm "$dir/code/data/backlog.md" + out=$(bootstrap_backlog_lines "$dir/code" "$dir/home") + assert_equals "" "$out" "an absent code-root backlog must stay silent" + + printf '## In flight\n\n## Queued\n\n- [ ] stray: written from the code root\n\n## Done\n' \ + > "$dir/code/data/backlog.md" + out=$(bootstrap_backlog_lines "$dir/code" "$dir/home") + assert_contains "$out" "BACKLOG_RECONCILE: code-root $dir/code/data/backlog.md is not this home's $dir/home/data/backlog.md" \ + "a regular code-root backlog beside a separate home was not reported" + assert_not_contains "$out" "done-archive.md" "an absent code-root archive was reported" + pass "bootstrap reports a regular code-root backlog and stays silent for a link into the home or no copy" +} + +test_guard_reports_foreign_link_and_archive() { + local dir out + dir=$(make_split guard-foreign) + empty_backlog "$dir/elsewhere.md" + rm "$dir/code/data/backlog.md" + ln -s "$dir/elsewhere.md" "$dir/code/data/backlog.md" + printf '## Done\n' > "$dir/code/data/done-archive.md" + out=$(bootstrap_backlog_lines "$dir/code" "$dir/home") + assert_contains "$out" "code-root $dir/code/data/backlog.md is not this home's" \ + "a code-root backlog linked outside this home was not reported" + assert_contains "$out" "code-root $dir/code/data/done-archive.md is not this home's $dir/home/data/done-archive.md" \ + "a regular code-root archive beside a separate home was not reported" + pass "bootstrap reports a code-root backlog linked elsewhere and a forked archive" +} + +test_guard_silent_for_single_home() { + local dir out + dir="$TMP_ROOT/single-guard" + mkdir -p "$dir/data" + cp "$ROOT/.tasks.toml" "$dir/.tasks.toml" + empty_backlog "$dir/data/backlog.md" + printf '## Done\n' > "$dir/data/done-archive.md" + out=$(bootstrap_backlog_lines "$dir") + assert_equals "" "$out" "the single-home layout's own backlog was reported as a fork" + out=$(bootstrap_backlog_lines "$dir" "$dir") + assert_equals "" "$out" "FM_HOME naming the code root was reported as a fork" + pass "bootstrap stays silent when the code root is the home" +} + +# The end-to-end fork: a bare tasks-axi write from the code root. Whatever the +# installed tasks-axi does to the link, bootstrap must agree with the result: +# a replaced link is reported, a written-through link is not. +test_bare_tasks_axi_fork_is_detected() { + local dir out + dir=$(make_split bare-fork) + (cd "$dir/code" && tasks-axi add bare-1 "written from the code root" >/dev/null 2>&1) \ + || fail "bare tasks-axi add failed in the code root" + out=$(bootstrap_backlog_lines "$dir/code" "$dir/home") + if [ -L "$dir/code/data/backlog.md" ]; then + assert_grep "bare-1" "$dir/home/data/backlog.md" "a written-through link lost the row" + assert_equals "" "$out" "a written-through link was reported as a fork" + pass "bare tasks-axi wrote through the code-root link and bootstrap stayed silent" + else + assert_no_grep "bare-1" "$dir/home/data/backlog.md" "the replaced link still reached the home" + assert_contains "$out" "code-root $dir/code/data/backlog.md is not this home's" \ + "bootstrap missed the fork a bare tasks-axi write left behind" + pass "bare tasks-axi replaced the code-root link and bootstrap reported the fork" + fi +} + +test_wrapper_writes_through_to_home() { + local dir i + dir=$(make_split wrapper-home) + for i in 1 2; do + wrapper_from_code "$dir" add "ship-$i" "ship $i" >/dev/null || fail "add ship-$i failed" + wrapper_from_code "$dir" start "ship-$i" >/dev/null || fail "start ship-$i failed" + wrapper_from_code "$dir" "done" "ship-$i" >/dev/null || fail "done ship-$i failed" + done + wrapper_from_code "$dir" add call-1 "captain call" >/dev/null || fail "add call-1 failed" + wrapper_from_code "$dir" hold call-1 --reason "awaiting the captain" --kind captain >/dev/null \ + || fail "hold call-1 failed" + printf 'RELATIVE-BODY-MARKER\n' > "$dir/code/body.md" + wrapper_from_code "$dir" update call-1 --body-file body.md >/dev/null \ + || fail "update with a caller-relative --body-file failed" + wrapper_from_code "$dir" prune --keep 1 >/dev/null || fail "prune failed" + + [ -L "$dir/code/data/backlog.md" ] || fail "a wrapper write replaced the code-root link" + [ "$dir/code/data/backlog.md" -ef "$dir/home/data/backlog.md" ] \ + || fail "the code-root link no longer names the home's backlog" + assert_grep "call-1" "$dir/home/data/backlog.md" "the held row did not land in the home" + assert_grep "RELATIVE-BODY-MARKER" "$dir/home/data/backlog.md" \ + "a caller-relative --body-file was not read from the caller's directory" + assert_present "$dir/home/data/done-archive.md" "archiving did not reach the home" + assert_grep "ship-1" "$dir/home/data/done-archive.md" "the oldest closed row was not archived in the home" + assert_absent "$dir/code/data/done-archive.md" "archiving wrote a code-root archive" + assert_equals "" "$(bootstrap_backlog_lines "$dir/code" "$dir/home")" \ + "bootstrap reported a fork after only wrapper writes" + pass "fm-tasks-axi.sh writes, holds, archives, and reads relative body files through to the home from the code root" +} + +test_wrapper_overrides_ambient_file() { + local dir + dir=$(make_split wrapper-ambient) + empty_backlog "$dir/decoy.md" + (cd "$dir/code" && TASKS_AXI_FILE="$dir/decoy.md" FM_HOME="$dir/home" "$WRAPPER" add amb-1 "ambient" >/dev/null) \ + || fail "add under an ambient TASKS_AXI_FILE failed" + assert_grep "amb-1" "$dir/home/data/backlog.md" "an ambient TASKS_AXI_FILE diverted the write from the home" + assert_no_grep "amb-1" "$dir/decoy.md" "an ambient TASKS_AXI_FILE received the write" + wrapper_from_code "$dir" >/dev/null || fail "the no-command dashboard failed" + pass "fm-tasks-axi.sh pins the home's backlog over an ambient TASKS_AXI_FILE and serves the dashboard" +} + +test_wrapper_refusals() { + local dir out rc before + dir=$(make_split wrapper-refuse) + before=$(cat "$dir/home/data/backlog.md") + out=$(wrapper_from_code "$dir" add r-1 "explicit" --file "$dir/home/data/backlog.md" 2>&1) + rc=$? + expect_code 2 "$rc" "--file" + assert_contains "$out" "drop --file" "--file refusal did not explain itself" + out=$(wrapper_from_code "$dir" list --file="$dir/home/data/backlog.md" 2>&1) + rc=$? + expect_code 2 "$rc" "--file=" + + mv "$dir/home/data/backlog.md" "$dir/home/real-backlog.md" + ln -s "$dir/home/real-backlog.md" "$dir/home/data/backlog.md" + out=$(wrapper_from_code "$dir" add r-2 "through a link" 2>&1) + rc=$? + expect_code 2 "$rc" "symlinked home backlog" + assert_contains "$out" "is a symlink" "the symlinked home backlog refusal did not name the link" + [ -L "$dir/home/data/backlog.md" ] || fail "a refused call still replaced the home link" + assert_equals "$before" "$(cat "$dir/home/real-backlog.md")" "a refused call changed the backlog" + + out=$(cd "$dir/code" && FM_HOME="$dir/missing-home" "$WRAPPER" list 2>&1) + rc=$? + expect_code 2 "$rc" "missing data directory" + pass "fm-tasks-axi.sh refuses caller --file, a symlinked home backlog, and an unresolvable home" +} + +test_wrapper_single_home() { + local dir + dir="$TMP_ROOT/single-wrapper" + mkdir -p "$dir/data" + cp "$ROOT/.tasks.toml" "$dir/.tasks.toml" + empty_backlog "$dir/data/backlog.md" + (cd "$dir" && FM_ROOT_OVERRIDE="$dir" "$WRAPPER" add solo-1 "single home" >/dev/null) \ + || fail "add in the single-home layout failed" + assert_grep "solo-1" "$dir/data/backlog.md" "the single-home layout lost its own backlog write" + pass "fm-tasks-axi.sh keeps the single-home layout addressing its own code-root backlog" +} + +test_guard_reports_regular_code_root_backlog +test_guard_reports_foreign_link_and_archive +test_guard_silent_for_single_home +if [ "$HAVE_TASKS_AXI" = 1 ]; then + test_bare_tasks_axi_fork_is_detected + test_wrapper_writes_through_to_home + test_wrapper_overrides_ambient_file + test_wrapper_refusals + test_wrapper_single_home +else + echo "skip: tasks-axi not found; home-addressing cases not run" +fi diff --git a/tests/fm-teardown-endpoint-safety.test.sh b/tests/fm-teardown-endpoint-safety.test.sh index 279d585601c..d88e36dc6f1 100755 --- a/tests/fm-teardown-endpoint-safety.test.sh +++ b/tests/fm-teardown-endpoint-safety.test.sh @@ -48,6 +48,11 @@ mark_case_as_treehouse_pool() { # <case> : > "$dir/worktree/sentinel" } +claim_pool_slot() { # <case> <task-id> [home] + local dir=$1 id=$2 home=${3:-$1/home} + printf 'task=%s\nhome=%s\n' "$id" "$home" > "$dir/pool/1/.fm-slot-owner" +} + run_case() { # <case> <id> local dir=$1 id=$2 FM_HOME="$dir/home" FM_ROOT_OVERRIDE="$ROOT" \ @@ -825,6 +830,148 @@ test_remote_layout_homes_serialize_on_one_project_lock() { pass "Treehouse project locking still serializes two homes across the remote-seeded boundary" } +# The slot-reuse sequence with only ONE discoverable record: the finished task's +# worker exited, its slot was granted to another task, and that task leaves no +# record this home can enumerate. Nothing in the record scan contradicts the +# stale worktree= line, so the slot's own owner claim is the only evidence that +# it was reassigned. The slot is no longer this task's, so teardown finishes the +# task's own cleanup and leaves the slot - its worker, its copy, its claim - +# exactly as it found it. +assert_reassigned_slot_left_alone() { # <case> <id> <other> <description> + local dir=$1 id=$2 other=$3 description=$4 + assert_absent "$dir/home/state/$id.meta" "$description: the stale task's own record was not removed" + assert_present "$dir/pool/1/.fm-slot-owner" "$description: another task's slot claim was removed" + assert_contains "$(cat "$dir/pool/1/.fm-slot-owner")" "task=$other" \ + "$description: another task's slot claim was rewritten" + assert_present "$dir/pool/1/project/.git" "$description: the reassigned slot's checkout was removed" + ! grep -Fq "treehouse <return>" "$dir/runtime.log" \ + || fail "$description: the reassigned slot was returned to the pool: $(cat "$dir/runtime.log")" + assert_contains "$(cat "$dir/stderr")" "$other" \ + "$description: the warning should name the task the slot was reassigned to" + assert_contains "$(cat "$dir/stderr")" "reassigned" \ + "$description: the warning should name the reassignment as the cause" +} + +test_reassigned_pool_slot_finishes_own_cleanup_without_touching_the_slot() { + local dir id=stale-task other=reassigned-task worker rc + + # Dirty slot, --force, and a live worker inside it: --force authorizes + # discarding this task's unlanded work, which is already gone with the slot, + # never the other task's live work. + dir=$(make_case slot-reassigned) + mark_case_as_treehouse_pool "$dir" + fm_write_meta "$dir/home/state/$id.meta" \ + "window=firstmate:fm-$id" "endpoint_task_id=$id" \ + "worktree=$dir/worktree" "project=$dir/project" "kind=scout" + claim_pool_slot "$dir" "$other" "$dir/other-home" + # Staged in this shell, not a command substitution: a background child of a + # $(...) subshell does not outlive it, and the point of this worker is to be + # alive in the slot while teardown runs. + ( cd "$dir/worktree" && exec sleep 30 ) & + worker=$! + + set +e + run_case "$dir" "$id" > "$dir/stdout" 2> "$dir/stderr" + rc=$? + set -e + + [ "$rc" -eq 0 ] || fail "teardown of a task whose slot was reassigned failed: $(cat "$dir/stderr")" + kill -0 "$worker" 2>/dev/null || fail "teardown killed the worker holding the reassigned pool slot" + assert_present "$dir/worktree/sentinel" "teardown reset a pool slot another task had claimed" + assert_reassigned_slot_left_alone "$dir" "$id" "$other" "dirty reassigned slot with --force" + assert_contains "$(cat "$dir/stderr")" "$dir/other-home" \ + "the warning should name the claimant's home" + kill "$worker" 2>/dev/null || true + wait "$worker" 2>/dev/null || true + + # The same reassignment on a CLEAN slot: a landed ship task torn down without + # --force, which is the shape of the real incident. A clean, fully landed copy + # passes every unlanded-work check, so only the ownership determination can + # keep this slot out of the pool; a guard keyed off dirtiness would return it + # and destroy the live task's copy. + dir=$(make_case slot-reassigned-clean) + mark_case_as_treehouse_pool "$dir" + rm -f "$dir/worktree/sentinel" + [ -z "$(git -C "$dir/worktree" status --porcelain)" ] \ + || fail "clean-slot fixture is not clean: $(git -C "$dir/worktree" status --porcelain)" + fm_write_meta "$dir/home/state/$id.meta" \ + "window=firstmate:fm-$id" "endpoint_task_id=$id" \ + "worktree=$dir/worktree" "project=$dir/project" "kind=ship" + claim_pool_slot "$dir" "$other" "$dir/other-home" + ( cd "$dir/worktree" && exec sleep 30 ) & + worker=$! + + set +e + FM_HOME="$dir/home" FM_ROOT_OVERRIDE="$ROOT" \ + FM_RUNTIME_LOG="$dir/runtime.log" PATH="$dir/fakebin:$PATH" \ + "$TEARDOWN" "$id" > "$dir/stdout" 2> "$dir/stderr" + rc=$? + set -e + [ "$rc" -eq 0 ] || fail "teardown of a clean ship task whose slot was reassigned failed: $(cat "$dir/stderr")" + kill -0 "$worker" 2>/dev/null || fail "teardown killed the worker holding the clean reassigned pool slot" + assert_reassigned_slot_left_alone "$dir" "$id" "$other" "clean reassigned slot without --force" + kill "$worker" 2>/dev/null || true + wait "$worker" 2>/dev/null || true + + # A claim that exists but cannot be read as a claim proves nothing either way, + # so it refuses rather than guessing the slot is still this task's. + dir=$(make_case slot-claim-unreadable) + mark_case_as_treehouse_pool "$dir" + fm_write_meta "$dir/home/state/$id.meta" \ + "window=firstmate:fm-$id" "endpoint_task_id=$id" \ + "worktree=$dir/worktree" "project=$dir/project" "kind=scout" + printf 'not-a-claim\n' > "$dir/pool/1/.fm-slot-owner" + + set +e + run_case "$dir" "$id" > "$dir/stdout" 2> "$dir/stderr" + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "teardown returned a pool slot whose claim could not be read" + assert_present "$dir/worktree/sentinel" "teardown reset a pool slot whose claim could not be read" + assert_present "$dir/pool/1/.fm-slot-owner" "teardown removed an unreadable slot claim" + assert_present "$dir/home/state/$id.meta" "teardown removed the task record on an unreadable claim" + [ ! -s "$dir/runtime.log" ] \ + || fail "teardown reached the runtime on an unreadable slot claim: $(cat "$dir/runtime.log")" + assert_contains "$(cat "$dir/stderr")" "$dir/pool/1/.fm-slot-owner" \ + "unreadable-claim refusal should name the claim file to inspect" + + pass "fm-teardown: a pool slot claimed by another task is left alone while the task's own cleanup finishes" +} + +# The two states that must never become a false refusal: the task's own claim, +# and no claim at all (a slot taken before claims existed, or already returned). +test_own_and_absent_slot_claims_still_tear_down() { + local dir id=owned-task + + dir=$(make_case slot-claim-own) + mark_case_as_treehouse_pool "$dir" + fm_write_meta "$dir/home/state/$id.meta" \ + "window=firstmate:fm-$id" "endpoint_task_id=$id" \ + "worktree=$dir/worktree" "project=$dir/project" "kind=scout" + claim_pool_slot "$dir" "$id" + + run_case "$dir" "$id" > "$dir/stdout" 2> "$dir/stderr" \ + || fail "teardown of a task holding its own slot claim failed: $(cat "$dir/stderr")" + assert_absent "$dir/home/state/$id.meta" "own-claim teardown left the task record" + assert_absent "$dir/pool/1/.fm-slot-owner" "own-claim teardown left its spent slot claim behind" + grep -Fq "treehouse <return>" "$dir/runtime.log" \ + || fail "own-claim teardown did not return its own pool slot: $(cat "$dir/runtime.log")" + + dir=$(make_case slot-claim-absent) + mark_case_as_treehouse_pool "$dir" + fm_write_meta "$dir/home/state/$id.meta" \ + "window=firstmate:fm-$id" "endpoint_task_id=$id" \ + "worktree=$dir/worktree" "project=$dir/project" "kind=scout" + + run_case "$dir" "$id" > "$dir/stdout" 2> "$dir/stderr" \ + || fail "teardown of an unclaimed slot failed: $(cat "$dir/stderr")" + assert_absent "$dir/home/state/$id.meta" "unclaimed-slot teardown left the task record" + grep -Fq "treehouse <return>" "$dir/runtime.log" \ + || fail "unclaimed-slot teardown did not return its pool slot: $(cat "$dir/runtime.log")" + + pass "fm-teardown: a task's own slot claim, and an unclaimed slot, both still tear down" +} + test_invalid_endpoint_records_refuse_before_mutation test_control_lock_contention_refuses_before_mutation test_non_pool_teardown_ignores_task_set_lock @@ -837,6 +984,8 @@ test_bare_relative_origin_shares_project_lock_with_clone test_reused_pool_slot_refuses_before_touching_the_other_task test_cross_home_pool_slot_collision_refuses test_sole_slot_record_still_tears_down +test_reassigned_pool_slot_finishes_own_cleanup_without_touching_the_slot +test_own_and_absent_slot_claims_still_tear_down test_recorded_endpoint_that_changed_directory_still_tears_down test_project_lock_anchors_at_the_local_root_across_home_layouts test_remote_seeded_home_returns_its_uncontested_slot diff --git a/tests/fm-teardown.test.sh b/tests/fm-teardown.test.sh index 05abc0e51e0..271d1abff1a 100755 --- a/tests/fm-teardown.test.sh +++ b/tests/fm-teardown.test.sh @@ -1870,7 +1870,7 @@ test_teardown_closes_the_backlog_item_itself() { "closed backlog item did not record the task's PR" assert_absent "$case_dir/state/task-x1.backlog-close" \ "a landed close left its pending-close record behind" - printf '%s\n' "$out" | grep -F 'tasks-axi ready' >/dev/null \ + printf '%s\n' "$out" | grep -F 'bin/fm-tasks-axi.sh ready' >/dev/null \ || fail "teardown dropped the dependency-cleared follow-up: $out" printf '%s\n' "$out" | grep -F 'check date gates' >/dev/null \ || fail "teardown did not preserve date-gate check: $out" diff --git a/tests/fm-test-fixtures.test.sh b/tests/fm-test-fixtures.test.sh index 9176771dd5d..5d9c67c8b1a 100755 --- a/tests/fm-test-fixtures.test.sh +++ b/tests/fm-test-fixtures.test.sh @@ -6,6 +6,12 @@ # filesystem effects - never on helper source text. Migrated spawn suites cover # fm_test_run_spawn through the real fm-spawn.sh; this file pins the shared # primitives and stubs those suites use. +# +# It is also the fixture Git-config isolation regression, with host signing +# armed on a scratch config file: it drives every entry point that must reach +# tests/git-config-helpers.sh - the shared helpers, bin/fm-test-run.sh's +# per-suite wrapper, and the standalone scripts runnable without a live vendor. +# That helper's header owns the contract and the layers it leaves in force. set -u # shellcheck source=tests/fixtures.sh @@ -13,6 +19,140 @@ set -u TMP_ROOT=$(fm_test_tmproot fm-test-fixtures) +test_git_config_isolation() ( + local dir="$TMP_ROOT/git-config" helper jobs timeout fakebin rc + mkdir -p "$dir/runner/bin" "$dir/runner/tests" + git init -q "$dir/caller" + git -C "$dir/caller" config commit.gpgsign false + cd "$dir/caller" || exit 1 + cp "$ROOT/bin/fm-test-run.sh" "$ROOT/bin/fm-timeout-lib.sh" "$dir/runner/bin/" + cp "$ROOT/tests/git-config-helpers.sh" "$dir/runner/tests/" + fakebin=$(fm_fakebin "$dir/standalone") + fm_fake_exit0 "$fakebin" pi + cat > "$fakebin/tmux" <<'SH' +#!/usr/bin/env bash +set -eu +while [ "$#" -gt 0 ]; do + if [ "$1" = -c ]; then + git -C "$2" log -1 --format=%s > "${FM_TEST_STANDALONE_COMMIT:?}" + exit 1 + fi + shift +done +SH + chmod +x "$fakebin/tmux" + cat > "$dir/runner/tests/fm-test-run.test.sh" <<'SH' +#!/usr/bin/env bash +set -eu +repo=$(mktemp -d "${TMPDIR:-/tmp}/fm-git-runner.XXXXXX") +trap 'rm -rf "$repo"' EXIT +git init -q "$repo" +git -C "$repo" config user.name 'Runner Fixture' +git -C "$repo" config user.email runner@example.invalid +git -C "$repo" commit -q --allow-empty -m initial +[ "$(git -C "$repo" log -1 --format='%s:%an:%ae')" = 'initial:Runner Fixture:runner@example.invalid' ] +[ "$(git -C "$repo" config --get fixture.input)" = preserved ] +[ "$(GIT_CONFIG_GLOBAL="$FM_TEST_GIT_CONFIG" git config --global --get commit.gpgsign)" = true ] +SH + chmod +x "$dir/runner/tests/fm-test-run.test.sh" + export GIT_CONFIG_GLOBAL="$dir/global" GIT_CONFIG_SYSTEM="$dir/system" + export GIT_CONFIG_NOSYSTEM=0 + unset GIT_CONFIG_COUNT GIT_CONFIG_PARAMETERS + + # A failing signer exposes inherited config without requiring GPG or keys. + arm_host_signing() { # <scope>: only this layer carries the failing signer + : > "$dir/global" + : > "$dir/system" + git config --file "$dir/$1" commit.gpgsign true + git config --file "$dir/$1" gpg.format openpgp + git config --file "$dir/$1" gpg.program /usr/bin/false + cp "$dir/$1" "$dir/expected" + } + + assert_helper_isolates() { # <helper> <scope> + bash -eus -- "$ROOT/tests/$1.sh" "$dir/$2-$1" "$dir/$2" <<'SH' || exit 1 +. "$1" +fm_git_init_commit "$2" +[ "$(git -C "$2" log -1 --format=%s)" = initial ] || fail "fixture has no initial commit" +fm_git_identity +# Child Git processes and direct commits inherit the same isolation. +bash -eu -c 'git -C "$1" commit -q --allow-empty -m child' _ "$2" +# Repository-local config and explicit command inputs remain authoritative. +git -C "$2" config commit.gpgsign true +git -C "$2" config gpg.program /usr/bin/false +if git -C "$2" commit -q --allow-empty -m signed > "$2/signing.log" 2>&1; then + fail "repository-local signing config was ignored" +fi +assert_grep 'gpg failed to sign' "$2/signing.log" "local signing was not attempted" +git -C "$2" -c commit.gpgsign=false commit -q --allow-empty -m explicit +GIT_CONFIG_COUNT=1 GIT_CONFIG_KEY_0=commit.gpgsign GIT_CONFIG_VALUE_0=false \ + git -C "$2" commit -q --allow-empty -m environment +# A config test can deliberately supply its own global file after sourcing. +[ "$(GIT_CONFIG_GLOBAL="$3" git config --global --get commit.gpgsign)" = true ] || fail "explicit global config was ignored" +SH + } + + assert_host_config_still_governs() { # <scope> + # Sourcing in test subprocesses cannot change the caller or its config files. + [ "$(git config --"$1" --get commit.gpgsign)" = true ] || fail "caller lost signing preference" + cmp -s "$dir/$1" "$dir/expected" || fail "host config file was changed" + git init -q "$dir/$1-outside" + if git -C "$dir/$1-outside" -c user.name=test -c user.email=test@example.invalid \ + commit -q --allow-empty -m outside > "$dir/outside.log" 2>&1; then + fail "commit outside fixtures bypassed signing" + fi + assert_grep 'gpg failed to sign' "$dir/outside.log" "outside commit did not attempt signing" + } + + # Every fixture entry point, once. Each only has to reach the shared helper; + # which layers that helper neutralizes is the helper's own property, settled + # by the system-layer case below. + arm_host_signing global + for helper in lib fixtures secondmate-helpers wake-helpers; do + assert_helper_isolates "$helper" global + done + bash -eus -- "$ROOT/tests/herdr-test-safety.sh" "$dir/global-herdr" <<'SH' || exit 1 +. "$1" +git init -q "$2" +git -C "$2" -c user.name=test -c user.email=test@example.invalid \ + commit -q --allow-empty -m initial +[ "$(git -C "$2" log -1 --format=%s)" = initial ] +SH + for jobs in 1 2; do + for timeout in 0 30; do + GIT_CONFIG_COUNT=1 GIT_CONFIG_KEY_0=fixture.input GIT_CONFIG_VALUE_0=preserved \ + FM_TEST_GIT_CONFIG="$dir/global" \ + "$dir/runner/bin/fm-test-run.sh" --jobs "$jobs" --per-script-timeout-secs "$timeout" \ + tests/fm-test-run.test.sh > "$dir/runner.log" 2>&1 \ + || fail "runner inherited global config (jobs=$jobs, timeout=$timeout): $(cat "$dir/runner.log")" + assert_grep 'FM_TEST_SUMMARY total=1 failed=0 skipped_gate=0' "$dir/runner.log" \ + "runner did not execute the Git fixture" + done + done + rc=0 + FM_SESSIONSTART_INSTRUCTION_REFRESH_LIVE_E2E=1 FM_SESSIONSTART_INSTRUCTION_REFRESH_REF=HEAD \ + FM_SESSIONSTART_INSTRUCTION_REFRESH_EXPECT=updated \ + FM_TEST_STANDALONE_COMMIT="$dir/global-standalone-commit" PATH="$fakebin:$PATH" \ + bash "$ROOT/tests/fm-sessionstart-instruction-refresh-live-e2e.test.sh" \ + > "$dir/standalone.log" 2>&1 || rc=$? + [ "$rc" = 1 ] || fail "standalone fixture did not stop at the tmux launch" + assert_grep 'could not start isolated Pi session' "$dir/standalone.log" \ + "standalone fixture failed before the tmux launch: $(cat "$dir/standalone.log")" + [ "$(cat "$dir/global-standalone-commit")" = 'test: initial instruction contract' ] \ + || fail "standalone fixture did not create its initial commit" + bash "$ROOT/tests/fm-gitignore-config.test.sh" > "$dir/gitignore.log" 2>&1 \ + || fail "standalone gitignore fixture inherited global config: $(cat "$dir/gitignore.log")" + assert_host_config_still_governs global + + # The system layer is the shared helper's other half: one entry point settles + # it, and the caller still signing proves the layer was genuinely armed. + arm_host_signing system + assert_helper_isolates lib system + assert_host_config_still_governs system + + pass "runner and shared helpers isolate host Git config and preserve explicit config and outside commits" +) + test_touch_epoch_preserves_repeated_dst_hour() { local TZ=Europe/Paris epoch path actual export TZ @@ -139,6 +279,7 @@ test_spawn_home_layout() { pass "spawn-home layout writes harness pin, beat, and brief" } +test_git_config_isolation || fail "Git fixture config isolation" test_touch_epoch_preserves_repeated_dst_hour test_no_mistakes_version_constant test_no_mistakes_init_doctor_markers diff --git a/tests/fm-test-run.test.sh b/tests/fm-test-run.test.sh index f78a931a918..3bc3c20cd76 100755 --- a/tests/fm-test-run.test.sh +++ b/tests/fm-test-run.test.sh @@ -26,6 +26,29 @@ skip_without() { # <interpreter> <case-title> "$2" "$1" } +# Every synthetic runner installation uses this owner, including constructors +# that deliberately omit or replace the optional timeout helper. Exercise the +# public runner immediately so missing runtime dependencies fail at construction. +install_runner_fixture() { # <repository> + local repo=$1 output probe=tests/fixture-install.test.sh + mkdir -p "$repo/bin" "$repo/tests" + cp "$RUNNER" "$repo/bin/fm-test-run.sh" + cp "$ROOT/tests/git-config-helpers.sh" "$repo/tests/" + chmod +x "$repo/bin/fm-test-run.sh" + cat > "$repo/$probe" <<'SH' +#!/usr/bin/env bash +[ "${GIT_CONFIG_GLOBAL:-}" = /dev/null ] || exit 1 +[ "${GIT_CONFIG_NOSYSTEM:-}" = 1 ] || exit 1 +printf 'ok - fixture runner isolates inherited Git configuration\n' +SH + chmod +x "$repo/$probe" + output=$(cd "$repo" && FM_TASK_ID='' GIT_CONFIG_GLOBAL=/fixture-invalid-config \ + GIT_CONFIG_NOSYSTEM=0 bin/fm-test-run.sh --per-script-timeout-secs 0 "$probe" 2>&1) \ + || fail "synthetic runner cannot execute its dependency probe in $repo: $output" + assert_contains "$output" 'FM_TEST_SUMMARY total=1 failed=0' "synthetic runner dependency probe did not complete" + rm "$repo/$probe" +} + test_list_all_exact_suite_coverage() { local listed expected missing extra f listed=$("$RUNNER" --list --all | LC_ALL=C sort) @@ -104,7 +127,8 @@ test_changed_file_selection_is_conservative() { init_changed_fixture_repo() { local repo=$1 script mkdir -p "$repo/bin" "$repo/tests" - cp "$RUNNER" "$repo/bin/fm-test-run.sh" + install_runner_fixture "$repo" + chmod +x "$repo/bin/fm-test-run.sh" for script in \ fm-brief.test.sh \ @@ -112,6 +136,7 @@ init_changed_fixture_repo() { fm-documentation-audiences.test.sh \ fm-test-isolation-proof.test.sh \ fm-test-run.test.sh \ + fm-test-fixtures.test.sh \ fm-cd-pretool-check.test.sh \ fm-daemon.test.sh \ fm-harness-adapter-instructions-live-e2e.test.sh \ @@ -143,6 +168,7 @@ init_changed_fixture_repo() { : >"$repo/bin/fm-supervisor-target-lib.sh" : >"$repo/bin/fm-control-lib.sh" cp "$ROOT/bin/fm-timeout-lib.sh" "$repo/bin/fm-timeout-lib.sh" + : >"$repo/bin/fm-procevent-quota.sh" : >"$repo/bin/fm-quota-axi-lib.sh" : >"$repo/bin/fm-quota-choose.sh" @@ -190,8 +216,9 @@ init_primary_and_linked_worktree() { git -C "$repo" worktree add --quiet -b linked-probe "$linked" for tree in "$repo" "$linked"; do mkdir -p "$tree/bin" "$tree/tests" - cp "$RUNNER" "$tree/bin/fm-test-run.sh" + install_runner_fixture "$tree" cp "$ROOT/bin/fm-timeout-lib.sh" "$tree/bin/fm-timeout-lib.sh" + chmod +x "$tree/bin/fm-test-run.sh" cat >"$tree/tests/probe.test.sh" <<PROBE #!/usr/bin/env bash @@ -270,6 +297,12 @@ test_changed_runner_surfaces_select_their_family() { *tests/fm-ask-user-authority.test.sh*) ;; *) fail "runner change did not select its pure-contract-unit family: $listed" ;; esac + # The suite that proves the runner's per-suite fixture Git isolation lives in + # the standalone family, which pure-contract-unit never reaches. + case "$listed" in + *tests/fm-test-fixtures.test.sh*) ;; + *) fail "runner change did not select its fixture-isolation regression: $listed" ;; + esac git -C "$repo" add bin/fm-test-run.sh git -C "$repo" -c user.name=test -c user.email=test@example.invalid commit -qm runner-change @@ -317,6 +350,14 @@ test_changed_dependency_selection_and_unmapped_failure() { git -C "$repo" add tests/lib.sh git -C "$repo" -c user.name=test -c user.email=test@example.invalid commit -qm helper-change + printf '\n' >>"$repo/tests/git-config-helpers.sh" + listed=$(cd "$repo" && bin/fm-test-run.sh --list --changed --base HEAD) + assert_contains "$listed" "tests/fm-pr-merge.test.sh" "git-config helper selects lib.sh dependents" + assert_contains "$listed" "tests/fm-secondmate-safety.test.sh" "git-config helper selects secondmate dependents" + assert_contains "$listed" "tests/fm-bearings-snapshot.test.sh" "git-config helper selects snapshot dependents" + git -C "$repo" add tests/git-config-helpers.sh + git -C "$repo" -c user.name=test -c user.email=test@example.invalid commit -qm git-config-helper-change + printf '\n' >>"$repo/tests/fm-backend-herdr-eventwait.test.py" listed=$(cd "$repo" && bin/fm-test-run.sh --list --changed --base HEAD) assert_contains "$listed" "tests/fm-backend-herdr-smoke.test.sh" "eventwait test selects Herdr coverage" @@ -512,7 +553,8 @@ PY timeout_repo="$tmp/timeout-repo" timeout_script=tests/fm-calm-pi-extension.test.sh mkdir -p "$timeout_repo/bin" "$timeout_repo/tests" - cp "$RUNNER" "$timeout_repo/bin/fm-test-run.sh" + install_runner_fixture "$timeout_repo" + cat >"$timeout_repo/bin/fm-timeout-lib.sh" <<'SH' fm_run_timed() { [ "$1" -eq 900 ] || return 99 @@ -662,8 +704,10 @@ test_family_proofs_run_in_separate_concurrent_phases() { tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run-family-phases.XXXXXX") repo="$tmp/repo" mkdir -p "$repo/bin" "$repo/tests" - cp "$RUNNER" "$repo/bin/fm-test-run.sh" + install_runner_fixture "$repo" + cp "$ROOT/bin/fm-timeout-lib.sh" "$repo/bin/fm-timeout-lib.sh" + chmod +x "$repo/bin/fm-test-run.sh" for script in \ fm-calm-pi-extension.test.sh fm-vendor-auth-probe.test.sh \ @@ -987,6 +1031,63 @@ test_exclude_family() { pass "exclude-family drops the named primary family after selection" } +test_list_scheduled_proven_isolated_uses_serial_weights() { + local tmp + tmp=$(fm_test_tmproot fm-test-run-proven-schedule) + "$RUNNER" --list --proven-isolated | LC_ALL=C sort >"$tmp/expected" + "$RUNNER" --list-scheduled --proven-isolated >"$tmp/actual" \ + || fail "--list-scheduled --proven-isolated failed" + cmp -s "$tmp/expected" "$tmp/actual" \ + || fail "proven-isolated scheduling must break serial-default ties by path" + pass "proven-isolated scheduling ignores parallel hints" +} + +test_list_scheduled_non_lane_selections_use_serial_weights() { + local tmp repo script selection + local -a scripts=( + tests/fm-operational-input.test.sh + tests/fm-lint.test.sh + tests/fm-muse-harness.test.sh + tests/fm-captain-hold-lifecycle.test.sh + tests/fm-kimi-harness.test.sh + tests/fm-brief.test.sh + ) + tmp=$(fm_test_tmproot fm-test-run-non-lane-schedule) + repo="$tmp/repo" + mkdir -p "$repo/bin" "$repo/tests" + install_runner_fixture "$repo" + for script in "${scripts[@]}"; do + printf '#!/usr/bin/env bash\nexit 0\n' >"$repo/$script" + chmod +x "$repo/$script" + done + git -C "$repo" init -q + git -C "$repo" add . + git -C "$repo" -c user.name=test -c user.email=test@example.invalid commit -qm baseline + for script in "${scripts[@]}"; do + printf '\n' >>"$repo/$script" + done + printf '%s\n' \ + tests/fm-muse-harness.test.sh \ + tests/fm-brief.test.sh \ + tests/fm-captain-hold-lifecycle.test.sh \ + tests/fm-lint.test.sh \ + tests/fm-kimi-harness.test.sh \ + tests/fm-operational-input.test.sh >"$tmp/expected" + for selection in family all changed scripts; do + case "$selection" in + family) set -- --family pure-contract-unit ;; + all) set -- --all ;; + changed) set -- --changed --base HEAD ;; + scripts) set -- "${scripts[@]}" ;; + esac + "$repo/bin/fm-test-run.sh" --list-scheduled "$@" >"$tmp/actual" \ + || fail "--list-scheduled $selection failed" + cmp -s "$tmp/expected" "$tmp/actual" \ + || fail "$selection scheduling must use serial hints and path-ordered default ties" + done + pass "family, all, changed, and script selections ignore parallel hints" +} + test_portable_shard_union_and_coverage_guard() { local s1 s2 proven serial herdr optin all_count union_count overlap out first s1=$("$RUNNER" --list --lane portable-parallel-1) @@ -1021,13 +1122,38 @@ test_portable_shard_union_and_coverage_guard() { # No duplicates across the four partitions. [ "$(printf '%s\n' "$s1" "$s2" "$serial" "$herdr" | LC_ALL=C sort | uniq -d | wc -l | tr -d ' ')" = "0" ] \ || fail "lanes must not duplicate scripts" - # LPT order: first script of shard 1 is the longest proven script. - first=$(printf '%s\n' "$s1" | head -n 1) - [ "$first" = "tests/fm-captain-hold-lifecycle.test.sh" ] \ - || fail "shard 1 must start with the longest proven script, got $first" + # LPT execution order, asserted against the runner's own measured schedule + # rather than against a script name: naming the current longest script here is + # what let the recorded lane duration go stale unnoticed in the first place. + for lane in portable-parallel-1 portable-parallel-2; do + [ "$("$RUNNER" --list --lane "$lane")" = "$("$RUNNER" --list-scheduled --lane "$lane")" ] \ + || fail "$lane membership must be stored longest-measured-first" + done pass "portable shard union, disjointness, and coverage guard hold" } +# The two parallel lanes are only "duration-balanced" while every member has a +# measured hint and the packing over those hints stays even. Both halves went +# unchecked until one lane grew past its CI job cap and was cancelled on every +# run, so assert them through the guard's own reported numbers. +test_portable_parallel_lanes_stay_duration_balanced() { + local out max imbalance unhinted + out=$("$RUNNER" --check-coverage) + unhinted=$(printf '%s\n' "$out" | sed -n 's/.*parallel_unhinted=\([0-9]*\).*/\1/p') + max=$(printf '%s\n' "$out" | sed -n 's/.*parallel_max_ms=\([0-9]*\).*/\1/p') + imbalance=$(printf '%s\n' "$out" | sed -n 's/.*parallel_imbalance_ms=\([0-9]*\).*/\1/p') + [ -n "$unhinted" ] && [ -n "$max" ] && [ -n "$imbalance" ] \ + || fail "coverage guard must report parallel_unhinted, parallel_max_ms, parallel_imbalance_ms: $out" + [ "$unhinted" = "0" ] \ + || fail "$unhinted proven-isolated scripts have no measured parallel hint, so the lanes are packed on a guess" + [ "$max" -gt 0 ] || fail "parallel_max_ms must be a positive packed duration, got $max" + # 5% of the worst lane: wide enough that one script's growth does not trip it, + # narrow enough that a lopsided partition cannot call itself balanced. + [ "$((imbalance * 20))" -le "$max" ] \ + || fail "parallel lanes differ by ${imbalance}ms against a ${max}ms worst lane, more than 5%" + pass "portable parallel lanes are fully hinted and packed within 5% of each other" +} + test_portable_serial_shards_partition_the_serial_lane() { local lanes count serial shard listed union dups shard_lane total cap lanes=$("$RUNNER" --list-lanes) @@ -1240,8 +1366,10 @@ test_unmapped_new_test_never_inherits_family_concurrency() { tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run-unmapped.XXXXXX") repo="$tmp/repo" mkdir -p "$repo/bin" "$repo/tests" - cp "$RUNNER" "$repo/bin/fm-test-run.sh" + install_runner_fixture "$repo" cp "$ROOT/bin/fm-timeout-lib.sh" "$repo/bin/fm-timeout-lib.sh" + + chmod +x "$repo/bin/fm-test-run.sh" # Two members of the proven residual family, plus a test basename the family # map has never seen - the shape of any test added tomorrow. @@ -1318,8 +1446,10 @@ test_per_script_timeout_bounds_a_hang() { runner="$repo/bin/fm-test-run.sh" hang=tests/fm-hang-fixture.test.sh mkdir -p "$repo/bin" "$repo/tests" - cp "$RUNNER" "$runner" + install_runner_fixture "$repo" + cp "$ROOT/bin/fm-timeout-lib.sh" "$repo/bin/fm-timeout-lib.sh" + grandchild_pid="$tmp/grandchild.pid" cat >"$repo/$hang" <<'SH' #!/usr/bin/env bash @@ -1386,8 +1516,10 @@ test_max_wall_ms_is_a_result_not_advice() { runner="$repo/bin/fm-test-run.sh" fast=tests/fm-budget-fixture.test.sh mkdir -p "$repo/bin" "$repo/tests" - cp "$RUNNER" "$runner" + install_runner_fixture "$repo" cp "$ROOT/bin/fm-timeout-lib.sh" "$repo/bin/fm-timeout-lib.sh" + + cat >"$repo/$fast" <<'SH' #!/usr/bin/env bash sleep 1 @@ -1452,10 +1584,12 @@ test_jobs_parallel_scheduler_and_failure_propagation() { c=tests/fm-lint.test.sh d=tests/fm-supervision-instructions.test.sh mkdir -p "$repo/bin" "$repo/tests" "$evidence" "$fake_bin" - cp "$RUNNER" "$runner" + install_runner_fixture "$repo" # The copied runner's per-script timeout default (see bin/fm-test-run.sh) needs # the timeout helper next to it; the runner refuses the run otherwise. cp "$ROOT/bin/fm-timeout-lib.sh" "$repo/bin/fm-timeout-lib.sh" + + cat >"$fake_bin/stat" <<'SH' #!/usr/bin/env bash if [ "$1" = "-c" ] && [ "$2" = "%a" ]; then @@ -1614,8 +1748,9 @@ test_per_script_timeout_bounds_a_hung_script_under_jobs() { # A real proven-isolated name, so --jobs accepts it; the body is a fixture. proven=tests/fm-brief.test.sh mkdir -p "$repo/bin" "$repo/tests" - cp "$RUNNER" "$runner" + install_runner_fixture "$repo" cp "$ROOT/bin/fm-timeout-lib.sh" "$repo/bin/fm-timeout-lib.sh" + cat >"$repo/$proven" <<'SH' #!/usr/bin/env bash echo "ok - hung parallel fixture started" @@ -1650,7 +1785,7 @@ test_per_script_timeout_zero_is_a_real_opt_out() { runner="$repo/bin/fm-test-run.sh" fixture=tests/slow.test.sh mkdir -p "$repo/bin" "$repo/tests" - cp "$RUNNER" "$runner" + install_runner_fixture "$repo" chmod +x "$runner" cat >"$repo/$fixture" <<'SH' #!/usr/bin/env bash @@ -1695,6 +1830,7 @@ SH # With the helper available the opt-out still runs unbounded rather than # quietly falling back to the default. cp "$ROOT/bin/fm-timeout-lib.sh" "$repo/bin/fm-timeout-lib.sh" + rc=0 ( cd "$repo" && "$runner" --per-script-timeout-secs 0 "$fixture" ) \ >"$tmp/out2" 2>"$tmp/err2" || rc=$? @@ -1715,7 +1851,7 @@ SH changed_repo="$tmp/changed" changed_script=tests/fm-calm-pi-extension.test.sh mkdir -p "$changed_repo/bin" "$changed_repo/tests" - cp "$RUNNER" "$changed_repo/bin/fm-test-run.sh" + install_runner_fixture "$changed_repo" cat >"$changed_repo/bin/fm-timeout-lib.sh" <<'SH' fm_run_timed() { return 124 @@ -1779,8 +1915,9 @@ test_signal_relay_reports_the_signal_that_stopped_the_sweep() { runner="$repo/bin/fm-test-run.sh" fixture=tests/slow.test.sh mkdir -p "$repo/bin" "$repo/tests" - cp "$RUNNER" "$runner" + install_runner_fixture "$repo" cp "$ROOT/bin/fm-timeout-lib.sh" "$repo/bin/fm-timeout-lib.sh" + chmod +x "$runner" # The fixture distinguishes the two signals the way several suites in this # repo do, so the relayed signal is observable and not just the runner's own @@ -1869,7 +2006,7 @@ test_per_script_bound_that_could_not_be_armed_reports_the_script_as_not_run() { # A real proven-isolated name, so --jobs accepts it; the body is a fixture. proven=tests/fm-brief.test.sh mkdir -p "$repo/bin" "$repo/tests" - cp "$RUNNER" "$runner" + install_runner_fixture "$repo" chmod +x "$runner" # The bound owner's documented "no bounded runner could start" outcome, in # the one file the runner sources to get it. @@ -1922,6 +2059,7 @@ SH # With a real bound owner, a script that exits 125 itself must keep its own # status and its own output, and must not be described as never run. cp "$ROOT/bin/fm-timeout-lib.sh" "$repo/bin/fm-timeout-lib.sh" + cat >"$repo/$fixture" <<'SH' #!/usr/bin/env bash echo "ok - self-inflicted 125 fixture ran" @@ -1962,8 +2100,9 @@ test_a_script_that_exits_124_itself_is_not_reported_as_a_bound_kill() { # A real proven-isolated name, so --jobs accepts it; the body is a fixture. proven=tests/fm-brief.test.sh mkdir -p "$repo/bin" "$repo/tests" - cp "$RUNNER" "$runner" + install_runner_fixture "$repo" cp "$ROOT/bin/fm-timeout-lib.sh" "$repo/bin/fm-timeout-lib.sh" + chmod +x "$runner" cat >"$repo/$fixture" <<'SH' #!/usr/bin/env bash @@ -2024,7 +2163,7 @@ test_per_script_timeout_default_arms_for_standard_modes() { runner="$repo/bin/fm-test-run.sh" fixture=tests/fm-brief.test.sh mkdir -p "$repo/bin" "$repo/tests" - cp "$RUNNER" "$runner" + install_runner_fixture "$repo" chmod +x "$runner" cat >"$repo/$fixture" <<'SH' #!/usr/bin/env bash @@ -2051,6 +2190,7 @@ SH # And with the helper present the armed default leaves healthy work alone. cp "$ROOT/bin/fm-timeout-lib.sh" "$repo/bin/fm-timeout-lib.sh" + rc=0 out="$tmp/green.out" ( cd "$repo" && "$runner" "$fixture" ) >"$out" 2>"$tmp/green.err" || rc=$? @@ -2159,7 +2299,10 @@ test_a_run_that_ran_records_no_skip_reason test_live_guards_expect_a_capability_skip_class test_fail_on_gate_skip_token test_exclude_family +test_list_scheduled_proven_isolated_uses_serial_weights +test_list_scheduled_non_lane_selections_use_serial_weights test_portable_shard_union_and_coverage_guard +test_portable_parallel_lanes_stay_duration_balanced test_portable_serial_shards_partition_the_serial_lane test_portable_serial_hint_coverage_is_reported_and_bounded test_portable_serial_shard_lane_refusals diff --git a/tests/fm-tmux-agent-liveness.test.sh b/tests/fm-tmux-agent-liveness.test.sh index d8c60708274..f87c0a41d66 100755 --- a/tests/fm-tmux-agent-liveness.test.sh +++ b/tests/fm-tmux-agent-liveness.test.sh @@ -153,7 +153,7 @@ wait_for_state() { # <target> <expected> [tries] title_classifies_agent() { # <target> local name name=$(fm_backend_tmux_current_command "$1" 2>/dev/null) - [ "$(fm_backend_tmux_classify_process_name "$name")" = agent ] + [ "$(fm_agent_process_classify_name "$name")" = agent ] } # Does the foreground-process-group identity, including argv[0], name one? @@ -161,13 +161,13 @@ comms_classify_agent() { # <target> local name while IFS= read -r name; do [ -n "$name" ] || continue - [ "$(fm_backend_tmux_classify_process_name "$name")" = agent ] && return 0 + [ "$(fm_agent_process_classify_name "$name")" = agent ] && return 0 done <<EOF $(fm_backend_tmux_foreground_comms "$1") EOF while IFS= read -r name; do [ -n "$name" ] || continue - [ "$(fm_backend_tmux_classify_process_name '' "$name")" = agent ] && return 0 + [ "$(fm_agent_process_classify_name '' "$name")" = agent ] && return 0 done <<EOF $(fm_backend_tmux_foreground_argv0s "$1") EOF diff --git a/tests/fm-trace-context-spawn.test.sh b/tests/fm-trace-context-spawn.test.sh index d38fbf787d1..b9a61736246 100755 --- a/tests/fm-trace-context-spawn.test.sh +++ b/tests/fm-trace-context-spawn.test.sh @@ -214,8 +214,12 @@ run_two_level() { smlog="$base/sm-launch.log" smfake=$(make_spawn_fakebin "$base/sm-fake") : > "$smlog" + # A claude secondmate spawn pre-registers workspace trust for the HOME it + # launches into (bin/fm-claude-trust.sh), so this runs against a throwaway + # HOME; without it this suite would write the developer's real ~/.claude.json. + mkdir -p "$base/user-home" env FM_TRACE_CONTEXT="$penv" \ - FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$prim" \ + FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$prim" HOME="$base/user-home" CLAUDE_CONFIG_DIR='' \ FM_STATE_OVERRIDE="$prim/state" FM_DATA_OVERRIDE="$prim/data" \ FM_PROJECTS_OVERRIDE="$prim/projects" FM_CONFIG_OVERRIDE="$prim/config" \ FM_SPAWN_NO_GUARD=1 CLAUDECODE=1 TMUX="fake,1,0" \ @@ -387,8 +391,12 @@ test_duplicate_secondmate_spawn_does_not_converge_trace_context() { printf 'charter\n' > "$sm/data/charter.md" fake=$(make_spawn_fakebin "$base/fake") + # A claude secondmate spawn pre-registers workspace trust for the HOME it + # launches into (bin/fm-claude-trust.sh), so this runs against a throwaway + # HOME; without it this suite would write the developer's real ~/.claude.json. + mkdir -p "$base/user-home" out=$(env -u FM_TRACE_CONTEXT \ - FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$prim" \ + FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$prim" HOME="$base/user-home" CLAUDE_CONFIG_DIR='' \ FM_STATE_OVERRIDE="$prim/state" FM_DATA_OVERRIDE="$prim/data" \ FM_PROJECTS_OVERRIDE="$prim/projects" FM_CONFIG_OVERRIDE="$prim/config" \ FM_SPAWN_NO_GUARD=1 CLAUDECODE=1 TMUX="fake,1,0" \ diff --git a/tests/fm-turnend-guard.test.sh b/tests/fm-turnend-guard.test.sh index c577bb6dc3f..911c0a6e0e4 100755 --- a/tests/fm-turnend-guard.test.sh +++ b/tests/fm-turnend-guard.test.sh @@ -22,6 +22,14 @@ fm_git_identity fmtest fmtest@example.invalid REQUIRED_REASON='watcher supervision needs Stop-owned automatic recovery; inspect the hook registration and startup status before ending the turn' AWAY_REQUIRED_REASON='Away mode owns watcher supervision' +# REQUIRED_REASON is the CLAUDE repair line, so the cases asserting it need +# detect_own to answer claude. CLAUDECODE=1 alone no longer pins that - a +# structural ancestor of a different harness outranks a marker - so those +# invocations also blind the ancestry walk. Only per-pid comm/args/ppid queries +# are answered here; watcher liveness still reaches the real ps. +BLIND_BIN=$(fm_fakebin "$TMP_ROOT/blind-ancestry") +fm_fake_blind_ancestry "$BLIND_BIN" + # --- PREDICATE: bin/fm-supervision-lib.sh ----------------------------------- test_predicate_healthy_no_inflight() { @@ -286,7 +294,7 @@ make_secondmate_linked_home_dir() { run_hook() { local dir=$1 stop_active=$2 home home=$(cd "$dir" && pwd) - printf '{"stop_hook_active":%s}' "$stop_active" | CLAUDECODE=1 FM_HOME="$home" bash "$dir/bin/fm-turnend-guard.sh" 2>&1 + printf '{"stop_hook_active":%s}' "$stop_active" | PATH="$BLIND_BIN:$PATH" CLAUDECODE=1 FM_HOME="$home" bash "$dir/bin/fm-turnend-guard.sh" 2>&1 } nonexistent_pid() { @@ -464,7 +472,7 @@ test_hook_blocks_from_fm_home_state() { home="$TMP_ROOT/hook-fm-home-op" mkdir -p "$home/state" : > "$home/state/task1.meta" - out=$(printf '{"stop_hook_active":false}' | CLAUDECODE=1 FM_HOME="$home" bash "$dir/bin/fm-turnend-guard.sh" 2>&1); status=$? + out=$(printf '{"stop_hook_active":false}' | PATH="$BLIND_BIN:$PATH" CLAUDECODE=1 FM_HOME="$home" bash "$dir/bin/fm-turnend-guard.sh" 2>&1); status=$? expect_code 2 "$status" "hook must inspect the active FM_HOME state dir" assert_contains "$out" "$REQUIRED_REASON" "block reason must contain the exact required instruction" pass "fm-turnend-guard: blocks from active FM_HOME state, not only repo-root state" @@ -541,7 +549,7 @@ test_hook_uses_state_override() { state="$TMP_ROOT/hook-state-override-active" mkdir -p "$home/state" "$state" : > "$state/task1.meta" - out=$(printf '{"stop_hook_active":false}' | CLAUDECODE=1 FM_HOME="$home" FM_STATE_OVERRIDE="$state" bash "$dir/bin/fm-turnend-guard.sh" 2>&1); status=$? + out=$(printf '{"stop_hook_active":false}' | PATH="$BLIND_BIN:$PATH" CLAUDECODE=1 FM_HOME="$home" FM_STATE_OVERRIDE="$state" bash "$dir/bin/fm-turnend-guard.sh" 2>&1); status=$? expect_code 2 "$status" "hook must let FM_STATE_OVERRIDE win over FM_HOME/state" assert_contains "$out" "$REQUIRED_REASON" "block reason must contain the exact required instruction" pass "fm-turnend-guard: uses FM_STATE_OVERRIDE ahead of FM_HOME/state" @@ -1268,6 +1276,17 @@ run_integrated_autoarm() { ' 2>&1 } +# The same real hook, fired from a harness-named process that does NOT write +# state/.lock: whoever already holds that lock decides whether this firing is +# the owning session's or a competing one. +run_integrated_autoarm_unowned() { + local dir=$1 home + home=$(cd "$dir" && pwd) + # shellcheck disable=SC2016 # the fake harness expands FM_HOME inside its child shell. + printf '{"session_id":"sess-claude-mode","stop_hook_active":false}\n' \ + | FM_HOME="$home" "$dir/fake-claude" -c '"$FM_HOME/bin/fm-claude-stop-autoarm.sh"' 2>&1 +} + write_integrated_failed_arm() { local dir=$1 cat > "$dir/bin/fm-watch-arm.sh" <<'SH' @@ -1662,6 +1681,115 @@ test_hook_claude_mode_integrated_monotonic_fail_open() { pass "fm-turnend-guard --claude: integrated fresh failures reach one bounded fail-open, stop continuation, and reset on recovery" } +# The auto-arm's ledger epoch advances only when the hook reaches its +# generation claim. A live harness-named process outside the hook's ancestry +# holding state/.lock keeps the hook inert by its identity contract, so the +# ledger stays at the exhausted-failure epoch the hook wrote before it went +# quiet. The block budget used to advance only on an epoch change, so this +# shape re-blocked without limit and the attended fail-open never fired: the +# budget must count consecutive re-blocks against an unchanged epoch instead. +hold_session_lock_from_foreign_harness() { # sets FOREIGN_LOCK_HOLDER + local dir=$1 + # `bash -c` execs a single command in place, which would rename the process + # to sleep; the trailing no-op keeps the harness-named shell as the holder. + # Started in this shell, not a command substitution, so the caller can reap + # it and no inherited pipe keeps a substitution waiting on the sleeper. + "$dir/fake-claude" -c 'sleep 60; true' >/dev/null 2>&1 & + FOREIGN_LOCK_HOLDER=$! + printf '%s\n' "$FOREIGN_LOCK_HOLDER" > "$dir/state/.lock" +} + +test_hook_claude_mode_frozen_epoch_reaches_bounded_fail_open() { + local dir out status guard_out guard_status holder i pid identity count epoch_line + dir=$(make_primary_dir "$TMP_ROOT/hook-claude-frozen-epoch") + : > "$dir/state/task1.meta" + install_integrated_autoarm "$dir" + write_integrated_failed_arm "$dir" + + out=$(run_integrated_autoarm "$dir"); status=$? + expect_code 2 "$status" "the exhausted auto-arm cycle must emit its one failure notice before going quiet" + guard_out=$(FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=100 run_hook_claude "$dir" true); guard_status=$? + expect_code 0 "$guard_status" "the first failed epoch must own its Stop handoff" + epoch_line=$(sed -n '1p' "$dir/state/.claude-autoarm-epoch") + + hold_session_lock_from_foreign_harness "$dir" + holder=$FOREIGN_LOCK_HOLDER + for i in 1 2 3 4; do + out=$(run_integrated_autoarm_unowned "$dir"); status=$? + expect_code 0 "$status" "an auto-arm outside the lock owner's ancestry must stay inert at stop $i" + [ -z "$out" ] || fail "inert auto-arm produced output at stop $i: $out" + [ "$(sed -n '1p' "$dir/state/.claude-autoarm-epoch")" = "$epoch_line" ] \ + || fail "the ledger epoch advanced at stop $i, so this case no longer drives a frozen epoch" + guard_out=$(FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=100 run_hook_claude "$dir" true); guard_status=$? + if [ "$i" -lt 4 ]; then + expect_code 2 "$guard_status" "frozen-epoch stop $i must still re-block within the budget" + assert_contains "$guard_out" "TURN WOULD END BLIND" "frozen-epoch re-block $i lost the blind-turn banner" + assert_not_contains "$guard_out" 'FIRSTMATE SUPERVISION IS GENUINELY DOWN' "fail-open fired before the frozen-epoch budget was spent" + assert_absent "$dir/state/.claude-autoarm-failure-alarmed" "frozen-epoch re-block $i consumed the attended alarm early" + else + expect_code 0 "$guard_status" "the frozen-epoch progression must reach the attended fail-open" + assert_contains "$guard_out" 'FIRSTMATE SUPERVISION IS GENUINELY DOWN' "the frozen-epoch fail-open alarm is missing" + assert_present "$dir/state/.claude-autoarm-failure-alarmed" "the frozen-epoch fail-open did not consume its episode alarm" + fi + done + + guard_out=$(FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=100 run_hook_claude "$dir" true); guard_status=$? + expect_code 2 "$guard_status" "a later unhealthy stop after the frozen-epoch alarm must remain attended" + assert_not_contains "$guard_out" 'FIRSTMATE SUPERVISION IS GENUINELY DOWN' "the attended alarm repeated against the frozen epoch" + + # The other direction: the bound must not outlive the failure. A verified + # healthy watcher still lets the stop through and clears the whole episode. + sleep 60 & + pid=$! + identity=$(watcher_identity "$dir" "$pid") || { + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + kill "$holder" 2>/dev/null || true + wait "$holder" 2>/dev/null || true + fail "could not identify the frozen-epoch recovery watcher" + } + record_watcher_lock "$dir" "$pid" "$identity" + touch "$dir/state/.last-watcher-beat" + guard_out=$(run_hook_claude "$dir" true); guard_status=$? + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + kill "$holder" 2>/dev/null || true + wait "$holder" 2>/dev/null || true + rm -rf "$dir/state/.watch.lock" + expect_code 0 "$guard_status" "a healthy watcher must still allow the stop after a frozen-epoch alarm" + [ -z "$guard_out" ] || fail "healthy allow after the frozen-epoch alarm produced output: $guard_out" + assert_absent "$dir/state/.turnend-claude-blocks" "positive recovery left the frozen-epoch block budget" + assert_absent "$dir/state/.claude-autoarm-failure-notified" "positive recovery left the failure notice" + assert_absent "$dir/state/.claude-autoarm-failure-alarmed" "positive recovery left the attended alarm" + guard_out=$(FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=100 run_hook_claude "$dir" true); guard_status=$? + expect_code 2 "$guard_status" "a later unhealthy stop must re-block from a fresh budget" + count=$(sed -n '2s/^count=//p' "$dir/state/.turnend-claude-blocks") + [ "$count" = 1 ] || fail "the post-recovery episode must restart its budget at 1, got $count" + pass "fm-turnend-guard --claude: an inert auto-arm's frozen epoch reaches one bounded fail-open and resets on recovery" +} + +# The same frozen ledger without a verified failure episode: the budget must +# still provably run out, and the verified-failure gate - not a stuck counter - +# is what keeps the stop blocking after that. +test_hook_claude_mode_frozen_epoch_without_verified_failure_spends_budget_and_keeps_blocking() { + local dir out status i count + dir=$(make_primary_dir "$TMP_ROOT/hook-claude-frozen-unverified") + : > "$dir/state/task1.meta" + printf 'epoch=7 owner_pid=999 outcome=clean updated_at=1\n' > "$dir/state/.claude-autoarm-epoch" + touch -t 202001010000 "$dir/state/.claude-autoarm-epoch" + for i in 1 2 3 4 5; do + out=$(FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=100 run_hook_claude "$dir" false); status=$? + expect_code 2 "$status" "frozen unverified stop $i must keep blocking" + assert_not_contains "$out" 'systemMessage' "an unverified frozen epoch must never fail open" + [ "$(sed -n '1p' "$dir/state/.claude-autoarm-epoch")" = 'epoch=7 owner_pid=999 outcome=clean updated_at=1' ] \ + || fail "the guard rewrote the frozen ledger at stop $i" + done + count=$(sed -n '2s/^count=//p' "$dir/state/.turnend-claude-blocks") + [ "$count" -gt 3 ] 2>/dev/null || fail "the block budget must run out against a frozen epoch, but the recorded count is ${count:-absent}" + assert_absent "$dir/state/.claude-autoarm-failure-alarmed" "an unverified frozen epoch recorded an attended alarm" + pass "fm-turnend-guard --claude: a frozen unverified epoch spends the budget yet still blocks" +} + test_hook_claude_mode_recovery_contention_is_not_ordinary_allow() { local dir pid identity holder out status dir=$(make_primary_dir "$TMP_ROOT/hook-claude-recovery-contention") @@ -1751,6 +1879,8 @@ test_hook_claude_mode_budget_without_verified_failure_keeps_blocking() { out=$(FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=100 run_hook_claude "$dir" false); status=$? expect_code 2 "$status" "--claude block $i must exit 2 within the budget" done + count=$(sed -n '2s/^count=//p' "$dir/state/.turnend-claude-blocks") + [ "$count" -gt 3 ] 2>/dev/null || fail "four consecutive blocks must spend the budget, but the recorded count is ${count:-absent}" assert_not_contains "$out" 'systemMessage' "budget exhaustion without verified auto-arm failure must not fail open" assert_absent "$dir/state/.claude-autoarm-failure-alarmed" "unverified budget exhaustion recorded an attended alarm" pass "fm-turnend-guard --claude: budget exhaustion alone cannot permit a blind stop" @@ -2205,6 +2335,8 @@ test_hook_claude_mode_blocks_on_stuck_generation_claim test_hook_claude_mode_terminal_fail_open_clears_abandoned_claim test_hook_claude_mode_preserves_fresh_failed_progression test_hook_claude_mode_integrated_monotonic_fail_open +test_hook_claude_mode_frozen_epoch_reaches_bounded_fail_open +test_hook_claude_mode_frozen_epoch_without_verified_failure_spends_budget_and_keeps_blocking test_hook_claude_mode_recovery_contention_is_not_ordinary_allow test_hook_claude_mode_concurrent_recovery_resets_are_idempotent test_hook_claude_mode_stale_rewake_epoch_blocks diff --git a/tests/fm-update.test.sh b/tests/fm-update.test.sh index f39d078025a..f45c3d595e6 100755 --- a/tests/fm-update.test.sh +++ b/tests/fm-update.test.sh @@ -472,6 +472,42 @@ test_unsafe_secondmate_home_skipped_before_git_update() { pass "T11 unsafe secondmate home is not fast-forwarded" } +# --- T12: a self-update rebinds a locally armed watch on the primary -------- +# A self-update fast-forwards bin/ in place, changing bytes an armed +# fm-procevent-when watch's trust binding was hashed against with no +# tampering involved; without a rebind the very next fire would be refused. +test_primary_update_rebinds_local_watch() { + local w before_hash after_hash out spec + w=$(new_world t12) + mkdir -p "$w/seed/bin" + printf "#!/usr/bin/env bash\necho v1 >> \"\$1\"\n" > "$w/seed/bin/watched-action.sh" + chmod +x "$w/seed/bin/watched-action.sh" + git -C "$w/seed" add -A + git -C "$w/seed" commit -qm add-watched-action + git -C "$w/seed" push -q origin main + git -C "$w/main" pull -q origin main + + FM_ROOT_OVERRIDE="$w/main" FM_HOME="$w/home" "$ROOT/bin/fm-procevent-when.sh" \ + arm rebind-primary --interval 60 --stable 1 \ + --condition true --action "$w/main/bin/watched-action.sh" "$w/rebind.log" >/dev/null + spec="$w/home/state/when/when-rebind-primary.spec" + before_hash=$(grep '^action_sha256=' "$spec") + + printf "#!/usr/bin/env bash\necho v2 >> \"\$1\"\n" > "$w/seed/bin/watched-action.sh" + git -C "$w/seed" add -A + git -C "$w/seed" commit -qm bump-watched-action + git -C "$w/seed" push -q origin main + + out=$(run_update "$w") + + assert_contains "$out" "firstmate: updated " "the primary still advanced" + assert_contains "$out" "rebound: when-rebind-primary" "the primary self-update rebound its own locally armed watch" + after_hash=$(grep '^action_sha256=' "$spec") + [ "$before_hash" != "$after_hash" ] \ + || fail "the watch's trust binding was not refreshed to match the updated action bytes" + pass "T12 a self-update rebinds a locally armed watch on the primary" +} + test_updates_main_and_secondmate test_reread_gate_is_instruction_only test_bin_only_advance_restarts @@ -486,5 +522,6 @@ test_registry_backstop_dedup_and_self_exclusion test_firstmate_wrong_branch_skipped test_firstmate_detached_head_skipped test_unsafe_secondmate_home_skipped_before_git_update +test_primary_update_rebinds_local_watch echo "# all fm-update tests passed" diff --git a/tests/fm-watch-arm.test.sh b/tests/fm-watch-arm.test.sh index 940b16cb2de..bd3d02a4505 100755 --- a/tests/fm-watch-arm.test.sh +++ b/tests/fm-watch-arm.test.sh @@ -799,8 +799,51 @@ test_downtime_marker_does_not_follow_symlink() { pass "watch-arm: downtime marker publication does not follow symlinks" } +# The watcher validates FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS when it arms and +# refuses to arm on an unusable value. Under a running watcher that value would +# make every per-cycle reconcile refuse by name into a discarded stdout, so no +# source would ever start and the home would sit disarmed while presenting as +# supervised; refusing to arm is loud through the liveness guard instead. This +# drives the real arm entry and asserts the arm STOPPED - non-zero exit, no +# started line, no lock holder, no beacon - and that its refusal names the +# variable, so a validator that merely returned false somewhere would not pass. +test_arm_refuses_an_unusable_launch_confirm_window() { + local dir home state fakebin armout status lock_pid + dir=$(make_case confirm-window-refusal) + home="$dir/home" + state="$dir/state" + fakebin="$dir/fakebin" + armout="$dir/arm.out" + mkdir -p "$home/data" + + PATH="$fakebin:$PATH" FM_HOME="$home" FM_STATE_OVERRIDE="$state" \ + FM_POLL=1 FM_SIGNAL_GRACE=0 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + FM_ARM_CONFIRM_TIMEOUT=5 FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS=5s \ + "$WATCH_ARM" > "$armout" 2>&1 & + ARM_PID=$! + wait_for_exit "$ARM_PID" 200 + status=$? + [ "$status" -ne 124 ] || fail "arm with an unusable confirm window never stopped: $(cat "$armout")" + [ "$status" -ne 0 ] || fail "arm reported success with an unusable confirm window: $(cat "$armout")" + grep -q '^watcher: FAILED' "$armout" \ + || fail "arm did not report the typed failure line: $(cat "$armout")" + grep -qF 'FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS' "$armout" \ + || fail "the refusal did not name the variable: $(cat "$armout")" + grep -qF "must be whole seconds from 1 to 600" "$armout" \ + || fail "the refusal did not name the accepted range: $(cat "$armout")" + ! grep -q '^watcher: started' "$armout" \ + || fail "arm reported a started watcher despite the refusal: $(cat "$armout")" + [ ! -e "$state/.last-watcher-beat" ] \ + || fail "a refused watcher still published a liveness beacon" + lock_pid=$(cat "$state/.watch.lock/pid" 2>/dev/null || true) + [ -z "$lock_pid" ] || ! kill -0 "$lock_pid" 2>/dev/null \ + || fail "a refused watcher is still running as pid $lock_pid" + pass "watch-arm: an unusable launch confirm window refuses to arm by name" +} + test_attached_arm_reports_the_delivered_wake test_attached_arm_reports_the_delivered_wake_after_drain +test_arm_refuses_an_unusable_launch_confirm_window test_attached_arm_still_fails_on_a_wake_it_did_not_deliver test_rearm_resurfaces_durable_queue_and_remote_open_decision test_marker_publish_failure_retains_recovery_evidence diff --git a/tests/fm-watch-triage.test.sh b/tests/fm-watch-triage.test.sh index 8994c60503d..52b1fce61a8 100755 --- a/tests/fm-watch-triage.test.sh +++ b/tests/fm-watch-triage.test.sh @@ -110,6 +110,8 @@ test_live_captain_held_first_sight_silenced_by_away_record() { test_backlog_hold_never_rechecked_while_away_record_exists() { local dir out capture wakes + command -v tasks-axi >/dev/null 2>&1 \ + || { echo "skip: tasks-axi not found (away-record backlog hold)"; return 0; } dir=$(make_hold_home away-record-backlog-hold 'done: PR https://example.test/pr/9 checks green' hold) \ || fail "could not build the backlog-hold fixture" out="$dir/watch.out"; capture="$dir/pane.txt" @@ -314,6 +316,8 @@ test_triage_log_size_cap_accepts_spaced_wc_counts test_procevent_captured_result_surfaces_proactively test_procevent_unacknowledged_result_redrains_until_handled test_procevent_marker_keys_are_injective +test_procevent_headlines_classify_queue_keys +test_procevent_launch_failed_episodes_are_each_delivered test_procevent_surface_serializes_with_drain test_procevent_surface_crash_boundaries test_procevent_marker_failure_exits_and_replays diff --git a/tests/fm-watcher-lock.test.sh b/tests/fm-watcher-lock.test.sh index 2389ff90d18..3f2b4508e24 100755 --- a/tests/fm-watcher-lock.test.sh +++ b/tests/fm-watcher-lock.test.sh @@ -133,19 +133,25 @@ test_guard_warnings() { # warning follows it, and the guidance is repair-after-drain (never the # old conflicting "restart NOW first"). # (2) a fresh watcher and an empty queue: total silence. - local dir state err first banner_line queue_line pid identity + local dir state err first banner_line queue_line pid identity blind dir=$(make_case guard) + # The repair line the cases below assert is the CLAUDE one, so detect_own has to + # answer claude. A marker alone no longer pins that: a structural ancestor of a + # different harness outranks it, so the harness this suite was launched from + # would otherwise choose the wording. Blind the ancestry walk as well; every + # other ps query (watcher liveness below) still reaches the real ps. + blind=$(fm_fakebin "$dir/blind") + fm_fake_blind_ancestry "$blind" state="$dir/state" err="$dir/guard.err" # (1) watcher down (no beacon) + two in-flight tasks + a queued wake. # FM_ROOT_OVERRIDE points the worktree-tangle check at a non-git dir so it stays # inert here; this case is about the watcher-down banner, not the tangle guard. - # Pin Claude so the host test runner's harness ancestry cannot change this fixture. printf 'project=x\n' > "$state/task.meta" printf 'project=y\n' > "$state/task2.meta" append_wake "$state" heartbeat heartbeat heartbeat || fail "guard heartbeat append failed" - CLAUDECODE=1 PI_CODING_AGENT='' GROK_AGENT='' FM_ROOT_OVERRIDE="$dir" FM_STATE_OVERRIDE="$state" FM_GUARD_GRACE=1 "$ROOT/bin/fm-guard.sh" 2> "$err" >/dev/null || fail "guard failed" + PATH="$blind:$PATH" CLAUDECODE=1 PI_CODING_AGENT='' GROK_AGENT='' FM_ROOT_OVERRIDE="$dir" FM_STATE_OVERRIDE="$state" FM_GUARD_GRACE=1 "$ROOT/bin/fm-guard.sh" 2> "$err" >/dev/null || fail "guard failed" first=$(grep -v '^[[:space:]]*$' "$err" | head -1) case "$first" in '●'*) ;; @@ -171,7 +177,7 @@ test_guard_warnings() { mkdir -p "$dir/config" printf 'project=x\n' > "$state/task.meta" : > "$dir/config/x-mode.env" - CLAUDECODE=1 PI_CODING_AGENT='' GROK_AGENT='' FM_ROOT_OVERRIDE="$dir" FM_STATE_OVERRIDE="$state" FM_GUARD_GRACE=1 "$ROOT/bin/fm-guard.sh" 2> "$err" >/dev/null || fail "guard failed" + PATH="$blind:$PATH" CLAUDECODE=1 PI_CODING_AGENT='' GROK_AGENT='' FM_ROOT_OVERRIDE="$dir" FM_STATE_OVERRIDE="$state" FM_GUARD_GRACE=1 "$ROOT/bin/fm-guard.sh" 2> "$err" >/dev/null || fail "guard failed" grep -F "source '$dir/config/x-mode.env' first" "$err" >/dev/null || fail "guard repair line did not source the X-mode cadence config" # (2) live watcher plus fresh beacon, empty queue -> silence. diff --git a/tests/fm-x-mode.test.sh b/tests/fm-x-mode.test.sh index 9f3eeb85a05..37e4048cf6a 100755 --- a/tests/fm-x-mode.test.sh +++ b/tests/fm-x-mode.test.sh @@ -949,8 +949,14 @@ test_reply_text_file_and_stdin() { } test_bootstrap_opt_out_cleanup() { - local home out + local home out blind home="$TMP_ROOT/boot-optout"; mkdir -p "$home" + # The remediation wording asserted below is the CLAUDE one, so detect_own has to + # answer claude. A marker alone no longer pins that - a structural ancestor of a + # different harness outranks it - so blind the ancestry walk too, or the harness + # this suite was launched from picks the wording. + blind=$(fm_fakebin "$TMP_ROOT/boot-optout-blind") + fm_fake_blind_ancestry "$blind" # Opt in, artifacts appear. printf 'FMX_PAIRING_TOKEN=tok-out\n' > "$home/.env" FM_HOME="$home" "$ROOT/bin/fm-bootstrap.sh" >/dev/null 2>&1 @@ -958,7 +964,7 @@ test_bootstrap_opt_out_cleanup() { assert_present "$home/config/x-mode.env" "opt-in must create the cadence config" # Opt out: empty the token, re-run bootstrap -> artifacts removed + one off line. printf 'FMX_PAIRING_TOKEN=\n' > "$home/.env" - out=$(CLAUDECODE=1 FM_HOME="$home" "$ROOT/bin/fm-bootstrap.sh" 2>/dev/null) + out=$(PATH="$blind:$PATH" CLAUDECODE=1 FM_HOME="$home" "$ROOT/bin/fm-bootstrap.sh" 2>/dev/null) assert_contains "$out" "FMX: X mode off" "opt-out must announce X mode off when it removed artifacts" assert_contains "$out" "watcher supervision needs Stop-owned automatic recovery" "opt-out remediation must use neutral automatic-recovery guidance" assert_not_contains "$out" "is broken" "opt-out remediation claimed an unverified mechanism failure" diff --git a/tests/git-config-helpers.sh b/tests/git-config-helpers.sh new file mode 100644 index 00000000000..8fde3e7f85e --- /dev/null +++ b/tests/git-config-helpers.sh @@ -0,0 +1,26 @@ +#!/usr/bin/env bash +# tests/git-config-helpers.sh - fixture Git isolation from the host's global and +# system configuration. +# +# Source this before a fixture's first Git operation: +# # shellcheck source=tests/git-config-helpers.sh +# . "$(dirname "${BASH_SOURCE[0]}")/git-config-helpers.sh" +# +# Fixture Git processes must not inherit host signing, hooks, or other global and +# system preferences: with commit.gpgsign=true set globally and no secret key for +# the fixture identities, every fixture commit fails before its assertion. +# Isolation is limited to those two layers on purpose, so repository-local config, +# inline `git -c`, GIT_CONFIG_COUNT, and a GIT_CONFIG_GLOBAL the caller supplies +# after sourcing all stay authoritative - the Git-config suites assert on them. +# The export reaches only the sourcing shell and its children, so the developer's +# own config files are never written and real project commits made outside the +# fixtures keep their configuration and signing. +# +# tests/lib.sh and tests/herdr-test-safety.sh source this for every suite that +# uses them, bin/fm-test-run.sh sources it per suite in run_script_bounded, and a +# suite reaching none of those sources it directly so a hand-run invocation is +# isolated too. tests/fm-test-fixtures.test.sh is the regression - it drives the +# shared helpers, the runner, and the standalone entry points that run without a +# live vendor - and the changed-file map selects it for a change to this file. + +export GIT_CONFIG_GLOBAL=/dev/null GIT_CONFIG_NOSYSTEM=1 diff --git a/tests/herdr-client-pair-fixture.sh b/tests/herdr-client-pair-fixture.sh index 52a7fde89c1..8d2410aade8 100644 --- a/tests/herdr-client-pair-fixture.sh +++ b/tests/herdr-client-pair-fixture.sh @@ -52,6 +52,10 @@ SH fi ;; "agent get") printf '{"id":"cli:agent:get","result":{"agent":{"agent":"claude","agent_status":"idle","pane_id":"wCY:p2"},"type":"agent_info"}}\n' ;; + "pane process-info") + # The registration above is only trusted once a live harness process backs + # it (issue #4115), so the compatible client also serves the process view. + printf '{"id":"cli:pane:process_info","result":{"process_info":{"pane_id":"wCY:p2","shell_pid":4242,"foreground_process_group_id":4243,"foreground_processes":[{"pid":4243,"name":"claude","argv0":"claude","argv":["claude"],"cmdline":"claude"}]},"type":"pane_process_info"}}\n' ;; *) : ;; esac exit 0 diff --git a/tests/herdr-test-safety.sh b/tests/herdr-test-safety.sh index 59a2bb46cc2..810b4a518fe 100644 --- a/tests/herdr-test-safety.sh +++ b/tests/herdr-test-safety.sh @@ -4,6 +4,9 @@ # fleet-state tripwire contract is bin/fm-herdr-lab.sh. set -u +# shellcheck source=tests/git-config-helpers.sh +. "$(dirname "${BASH_SOURCE[0]}")/git-config-helpers.sh" + # Herdr backend tests drive the real fm-spawn/fm-teardown but do not source # tests/lib.sh, so exempt them from the gate-lifecycle refusal here too (see # tests/lib.sh and bin/fm-gate-refuse-lib.sh for why firstmate's own suite, diff --git a/tests/lib.sh b/tests/lib.sh index 9db0195c8e4..73a4fd390aa 100644 --- a/tests/lib.sh +++ b/tests/lib.sh @@ -33,6 +33,11 @@ FM_TEST_LIB_SOURCED=1 # suite's fixtures were written against. umask 022 +# Fixture Git isolation for every suite that reaches this library; the helper's +# header owns the invariant and the layers it deliberately leaves in force. +# shellcheck source=tests/git-config-helpers.sh +. "$(dirname "${BASH_SOURCE[0]}")/git-config-helpers.sh" + # Exempt firstmate's own test suite from the gate-lifecycle refusal # (bin/fm-gate-refuse-lib.sh). The no-mistakes gate runs this suite FROM a gate # worktree - the exact environment that guard refuses - so without this every @@ -77,6 +82,16 @@ export FM_DASHBOARD_EVENTS_CONFIG="$FM_TEST_ISOLATION_ROOT/dashboard-events.json export FM_DASHBOARD_EVENT_DB="$FM_TEST_ISOLATION_ROOT/events.db" export FM_DASHBOARD_AUTH_FILE="$FM_TEST_ISOLATION_ROOT/dashboard-auth.json" +# Clear the tasks-axi env overrides. An operator shell exports TASKS_AXI_FILE +# (and may export TASKS_AXI_BACKEND) at its real home's backlog, and tasks-axi +# resolves that env AHEAD of the .tasks.toml a fixture copies, so a suite that +# seeds a temp home with bare `tasks-axi` would silently write the operator's +# live backlog instead - tests/fm-public-followup.test.sh did exactly that. Every +# fixture addresses its own data/backlog.md through its copied .tasks.toml, an +# explicit --file, or bin/fm-tasks-axi.sh; a case that verifies the wrapper +# against an ambient override sets TASKS_AXI_FILE itself. +unset TASKS_AXI_FILE TASKS_AXI_BACKEND + # Resolve the repo root from this library's own location. Consumed by sourcing # test files, not by this library, so it reads as "unused" here. # shellcheck disable=SC2034 @@ -584,6 +599,34 @@ SH chmod +x "$fakebin/fm-crash-inject" } +# fm_fake_blind_ancestry <fakebin> +# Blind the parent-chain walks: a query of the FIELD-FIRST per-pid form those walks +# use - `ps -o comm=|args=|ppid= -p <pid>`, the shape in bin/fm-harness.sh, +# bin/fm-session-lock-lib.sh, bin/fm-sessionstart-nudge.sh and bin/fm-backend.sh's +# cmux ancestor detection - reports a bash ancestor terminating at pid 1, so ancestry +# proves nothing and the marker a case sets is the only evidence left. A case that pins +# its harness with a marker (CLAUDECODE=1 and friends) needs this, because a structural +# ancestor of a DIFFERENT harness outranks a marker - without it, the harness the SUITE +# was launched from decides the verdict. +# Every other ps query reaches the real ps untouched, and the pid-first form is +# deliberately among them: bin/fm-tmux-lib.sh and bin/backends/tmux.sh read pane and +# cursor identity with `ps -p <pid> -o args=`, so intercepting that shape too would make +# a pane assertion under a PATH-wide blind read `bash` and reject every cursor pane. +fm_fake_blind_ancestry() { + local fakebin=$1 real_ps + real_ps=$(command -v ps) || return 1 + cat > "$fakebin/ps" <<SH +#!/usr/bin/env bash +case "\$*" in + '-o comm= -p '*) printf '%s\n' bash ;; + '-o args= -p '*) printf '%s\n' bash ;; + '-o ppid= -p '*) printf '%s\n' 1 ;; + *) exec "$real_ps" "\$@" ;; +esac +SH + chmod +x "$fakebin/ps" +} + # fm_fake_version_tool <fakebin> <tool> <override-env-var> <default-version> # The stub answers `--version` with <override-env-var> when that variable is set # and non-empty, and with <default-version> otherwise; every other invocation diff --git a/tests/remote-herdr-fixture.sh b/tests/remote-herdr-fixture.sh index e849e8d6e61..cf40a57a03f 100644 --- a/tests/remote-herdr-fixture.sh +++ b/tests/remote-herdr-fixture.sh @@ -41,13 +41,14 @@ SH printf '%s\n' "$*" >> "$LOG" jq_state() { jq "$@" "$STATE"; } save() { tmp="$STATE.tmp.$$"; cat > "$tmp" && mv "$tmp" "$STATE"; } -ws=""; label=""; cwd="" +ws=""; label=""; cwd=""; pane="" args=("$@") for ((i=0; i<${#args[@]}; i++)); do case "${args[$i]}" in --workspace) ws=${args[$((i+1))]:-} ;; --label) label=${args[$((i+1))]:-} ;; --cwd) cwd=${args[$((i+1))]:-} ;; + --pane) pane=${args[$((i+1))]:-} ;; esac done case "${1:-} ${2:-}" in @@ -97,7 +98,9 @@ case "${1:-} ${2:-}" in [ ! -f "$SEND_FAIL" ] || exit 1 jq_state --arg p "${3:-}" '.typed[$p] = true | .working[$p] = true' | save ;; "pane read") printf '\n' ;; - "pane process-info") printf '{"result":{"process":{"name":"codex"}}}\n' ;; + "pane process-info") + printf '{"result":{"type":"pane_process_info","process_info":{"pane_id":"%s","shell_pid":%s,"foreground_process_group_id":%s,"foreground_processes":[{"pid":%s,"name":"codex","argv0":"codex","argv":["codex"],"cmdline":"codex"}]}}}\n' \ + "$pane" "$$" "$$" "$$" ;; "agent get") pane=${3:-} if [ "$(jq_state -r --arg p "$pane" '.working[$p] // false')" = true ]; then diff --git a/tests/watch-triage-helpers.sh b/tests/watch-triage-helpers.sh index 86e6cff1eec..788ba275556 100644 --- a/tests/watch-triage-helpers.sh +++ b/tests/watch-triage-helpers.sh @@ -5164,6 +5164,19 @@ seed_captured_procevent_result() { # <dir> sleep 0.1 i=$((i + 1)) done + # The runner publishes that wake BEFORE it releases its claim and exits, so a + # retire that lands in that gap reads the exiting runner's ownership as + # uncertain and refuses with "cannot confirm runner identity" - the pipeline + # saw exactly that under load. Wait, bounded, for the release the publish + # promises, so retire meets a source nothing owns instead of racing the + # runner's last milliseconds. The bound keeps a runner that never releases a + # real failure at retire rather than a hang here. + i=0 + while [ "$i" -lt 100 ]; do + [ -e "$dir/claims/delivery-src.claim" ] || break + sleep 0.1 + i=$((i + 1)) + done pe_case "$dir" retire delivery-src >/dev/null || return 1 [ -s "$dir/state/.wake-queue" ] } @@ -5267,6 +5280,108 @@ test_procevent_marker_keys_are_injective() { pass "complete process-event queue keys map to distinct seen markers" } +# The reason line is the headline firstmate reads before the payload. Every +# procevent:* key used to surface as "process-event result captured", which +# presents a source that is collecting NOTHING as a healthy capture - the exact +# shape of the incident these wakes exist to expose. These assertions read the +# reason the watcher actually printed, so a typo in either classifying glob +# fails here instead of silently falling back to the healthy-looking headline. +surface_once() { # <dir> <out> [limit-ticks]: run one watcher to its wake, return its status + local dir=$1 out=$2 limit=${3:-100} pid + procevent_watch_bg "$dir" "$out" + pid=$! + wait_for_exit "$pid" "$limit" +} + +test_procevent_headlines_classify_queue_keys() { + local dir state out + dir=$(make_case procevent-headline-captured); state="$dir/state"; out="$dir/watch.out" + append_wake "$state" check "procevent:cap-src:1" "check: procevent lavish cap-src 1" + surface_once "$dir" "$out" || fail "a captured-result key was not surfaced: $(cat "$out")" + grep -F "check: process-event result captured: procevent:cap-src:1" "$out" >/dev/null \ + || fail "a captured result did not surface under its own headline: $(cat "$out")" + ! grep -F "source stranded" "$out" >/dev/null \ + || fail "a captured result was headlined as a strand: $(cat "$out")" + ! grep -F "failed to start" "$out" >/dev/null \ + || fail "a captured result was headlined as a failed start: $(cat "$out")" + FM_STATE_OVERRIDE="$state" "$DRAIN" >/dev/null 2>&1 || fail "captured headline fixture drain failed" + + dir=$(make_case procevent-headline-stranded); state="$dir/state"; out="$dir/watch.out" + append_wake "$state" check "procevent:str-src:stranded:tok-1" "check: process-event source str-src is registered but nothing can arm it" + surface_once "$dir" "$out" || fail "a stranded key was not surfaced: $(cat "$out")" + grep -F "check: process-event source stranded: procevent:str-src:stranded:tok-1" "$out" >/dev/null \ + || fail "a stranded source did not surface under its own headline: $(cat "$out")" + ! grep -F "result captured" "$out" >/dev/null \ + || fail "a stranded source was headlined as a captured result: $(cat "$out")" + FM_STATE_OVERRIDE="$state" "$DRAIN" >/dev/null 2>&1 || fail "stranded headline fixture drain failed" + + dir=$(make_case procevent-headline-joined); state="$dir/state"; out="$dir/watch.out" + append_wake "$state" check "procevent:cap2-src:1" "check: procevent lavish cap2-src 1" + append_wake "$state" check "procevent:str2-src:stranded:tok-2" "check: process-event source str2-src is registered but nothing can arm it" + surface_once "$dir" "$out" || fail "a mixed cycle was not surfaced: $(cat "$out")" + grep -F "check: process-event result captured: procevent:cap2-src:1; process-event source stranded: procevent:str2-src:stranded:tok-2" "$out" >/dev/null \ + || fail "a cycle with a capture and a strand did not carry both headlines joined: $(cat "$out")" + FM_STATE_OVERRIDE="$state" "$DRAIN" >/dev/null 2>&1 || fail "joined headline fixture drain failed" + pass "process-event queue keys surface under their own headlines" +} + +# Delivery, not queue rows, is what proves a launch-failure episode reaches +# firstmate. The watcher remembers every procevent key it has surfaced for +# good, so reconcile keys each episode with a fresh suffix beyond the +# registration identity: this test would fail if a second episode reused the +# first one's key, because the watcher would keep polling and never wake. +test_procevent_launch_failed_episodes_are_each_delivered() { + local dir state out status + dir=$(make_case procevent-launch-failed-episodes); state="$dir/state"; out="$dir/watch.out" + append_wake "$state" check "procevent:lf-src:launch-failed:1-2-100-7" \ + "check: process-event source lf-src is registered but its launch did not prove it took the claim" + surface_once "$dir" "$out" || fail "a launch-failed key was not surfaced: $(cat "$out")" + grep -F "check: process-event source failed to start: procevent:lf-src:launch-failed:1-2-100-7" "$out" >/dev/null \ + || fail "a failed launch did not surface under its own headline: $(cat "$out")" + ! grep -F "result captured" "$out" >/dev/null \ + || fail "a failed launch was headlined as a captured result: $(cat "$out")" + ack_stopped_cycle "$state" >/dev/null || fail "launch-failed fixture could not be handled and acknowledged" + + # The same key again is what a registration-identity-only key would produce + # for the next episode: already surfaced, so the process-event surface never + # delivers it under its headline again. A fresh watcher still recovers the + # unacknowledged queue row through the generic `check: rearm-resurface` + # path (the contract test_procevent_unacknowledged_result_redrains_until_handled + # proves), so what this asserts is the headline, not silence. + append_wake "$state" check "procevent:lf-src:launch-failed:1-2-100-7" \ + "check: process-event source lf-src is registered but its launch did not prove it took the claim" + : > "$out" + status=0 + surface_once "$dir" "$out" 30 || status=$? + case "$status" in + 124) ;; + 0) + # The one wake this tolerates is the recovery path named above, by its + # exact reason line. A wake for any other reason would mean either that + # the ordinary surface delivered the repeated key after all, or that + # something unrelated fired inside the window - and both are failures of + # exactly what this test guards, so neither may pass as "recovery". + grep -F 'check: rearm-resurface' "$out" >/dev/null \ + || fail "an already-surfaced launch-failed key woke the watcher, and the reason was not the one tolerated recovery path (expected the exact line 'check: rearm-resurface'; if that path was reworded, update this expectation, do not restore the strict silence check): $(cat "$out")" + ;; + *) fail "the watcher failed on an already-surfaced launch-failed key (status $status): $(cat "$out")" ;; + esac + ! grep -F "failed to start: procevent:lf-src:launch-failed:1-2-100-7" "$out" >/dev/null \ + || fail "an already-surfaced launch-failed key was delivered again under its headline: $(cat "$out")" + ack_stopped_cycle "$state" >/dev/null || fail "repeated-key fixture could not be handled and acknowledged" + + # A later episode of the same registration carries the same identity under a + # fresh suffix, and that one must be delivered. + append_wake "$state" check "procevent:lf-src:launch-failed:1-2-160-9" \ + "check: process-event source lf-src is registered but its launch did not prove it took the claim" + : > "$out" + surface_once "$dir" "$out" || fail "a second launch-failure episode was not surfaced: $(cat "$out")" + grep -F "check: process-event source failed to start: procevent:lf-src:launch-failed:1-2-160-9" "$out" >/dev/null \ + || fail "a second launch-failure episode did not surface under its own headline: $(cat "$out")" + ack_stopped_cycle "$state" >/dev/null || fail "second episode fixture could not be handled and acknowledged" + pass "every launch-failure episode is delivered under the failed-to-start headline" +} + install_marker_mv_fault() { # <dir> local dir=$1 REAL_MV=$(command -v mv)