diff --git a/.agents/skills/bearings/SKILL.md b/.agents/skills/bearings/SKILL.md index 4eeaf6baff0..c28b7f4f4eb 100644 --- a/.agents/skills/bearings/SKILL.md +++ b/.agents/skills/bearings/SKILL.md @@ -88,11 +88,13 @@ Board answers are acted on later under the normal authority rules; this skill's ## Lavish board mode `/bearings lavish` adds one deliverable beside the unchanged chat digest: the interactive fleet board, a myfirstmate-styled Lavish page where the captain answers Captain's Call items directly instead of replying in chat. -`bin/fm-bearings-board.sh` owns every board mechanic - the stable board path, fm-bearings-board.v1 payload validation, template injection, Lavish session establishment, the any-origin answer binding, and arm-if-absent registration - so the per-invocation work is composing the payload and running its `build`. +`bin/fm-bearings-board.sh` owns every board mechanic - the stable board path, fm-bearings-board.v1 payload validation, template injection, live Lavish session verification and ended-session reopening, the any-origin answer binding, and listener registration - so the per-invocation work is composing the payload and running its `build`. Compose the payload from the same snapshot with the same ranking judgment as the chat digest, plus these board rules: - A Captain's Call decision key is the captain-held TASK ID from `decisions_open` (legacy `-decision-` rows are already task ids); a merge card's key is `merge.`; the Charted Next dispatch picker's key is `dispatch.charted`. +- Before carding a hold, check that its SUBJECT has not already landed, and omit it when it has. `build` drops a card whose task or PR appears in the payload's own landed rows, and one whose task is no longer an open captain call. When a hold waits on one specific PR, put that PR in the card's `pr_url`. When it concerns a published version, put the artifact and numeric three-part version in the card's structured `subject`; landed rows for releases carry the same identity, and a matching or newer version drops the card. Identity matching is structured only, so verify any subject without one of these identities against current reality before carding it. +- Never author a `reconcile` option on any card. `build` gives every decision card the standard reconcile choice itself, and the payload validator reserves that value across all card types; recommendations must name an authored option. - Compose exactly one decision card per captain-held task id. When one task carries multiple questions, consolidate all of them and their options into that card; never emit duplicate cards with the same task-id key. - Decision cards carry agent-authored copy: a short noun-phrase title, one-line `about` and `decide` context rows, and option labels with hints, with the recommended option marked. - Card `type` (decision, merge, credential) is your composing judgment from the row's content; no backlog field types a card for you. @@ -102,14 +104,20 @@ Compose the payload from the same snapshot with the same ranking judgment as the - Every Captain's Call item and every Underway, Recently Landed, and Charted Next row carries an explicit `repo` field. Fill it from the snapshot and task records wherever known; use null or an empty string only as the deliberate genuinely-no-repo marker, in which case the template may show the internal id. Ids otherwise stay in the payload only as the routing channel, and composed reasons name blockers in plain words. Run `build` once after composing the payload. -Its serve-first sequence publishes the board, establishes or resumes its Lavish session with `lavish-axi`, and only then binds and arms the polling source; use the session URL it prints in the chat digest. -Never bind or arm the board before that session exists. -Never run `lavish-axi poll` for the board yourself: the armed source's supervised runner owns the blocking poll, and the watcher's ordinary reconcile restarts it, so no conversational turn ever blocks on the board. +Its serve-first sequence publishes the board, establishes and verifies its Lavish session with `lavish-axi`, reopens an ended session when necessary, and only then binds the answer source and proves a live polling listener; use the session URL it prints in the chat digest. +Never bind or arm the board before its session is listed open. +Never run `lavish-axi poll` for the board yourself: the armed source's supervised runner owns the blocking poll, and both the build and the watcher's ordinary reconcile repair a missing listener, so no conversational turn ever blocks on the board. ### Handling a board wake A board answer arrives as an ordinary `procevent lavish ` check wake. Identify it by comparing the wake source id with `bin/fm-procevent-lavish.sh source-id "$(bin/fm-bearings-board.sh path)"`, regardless of which answer kinds the result contains; then load `process-event-sources` and follow its contract for the result read, adapter classification, and the handled acknowledgement. Decision answers need no routing from you: the runner feeds the board's binding into `bin/fm-captain-hold.sh`'s one keyed-answer intake, which closes or releases each answered captain-held task at answer time; reconcile any `skipped:` key yourself with a direct `answer`, and when the captain's answer is "later", record it as a deferral with `bin/fm-captain-hold.sh hold --reason "" --until ` instead of a closure. +A current structured Reconcile selection closes nothing: the versioned board context carries its exact selected option separately from any typed note, and the adapter routes that selection only into a durable re-check request while preserving the note as provenance. +The rollout-compatible old context still feeds ordinary non-reconcile answers, but its bare or separator-annotated reconcile values and every structurally uncertain choice feed neither intake and remain announced for deliberate handling. +Verify the call's latest state, then retire the request through `bin/fm-captain-hold.sh reconcile close --evidence-file ` when it turns out to be moot, or `reconcile note --note-file ` when it is genuinely still open. +Both outcomes refuse without that pending board-created request, and `bin/fm-captain-hold.sh reconcile list` names every request still outstanding. +A remote-secondmate card whose task is absent from the main backlog remains on the board unchanged, but its reconcile request is refused in the main home until the separately tracked owner-aware routing follow-up can query and mutate the authoritative secondmate home; handle the announced capture without claiming that a request or reconciliation succeeded. +`captain-hold-lifecycle` owns why a reconcile may never be recorded as the captain's answer. Route the non-decision keys yourself: - `merge.` is the captain's explicit merge order; follow the merge ruling below. diff --git a/.agents/skills/bearings/assets/board-template.html b/.agents/skills/bearings/assets/board-template.html index 3ba3356dc9e..614bef7426b 100644 --- a/.agents/skills/bearings/assets/board-template.html +++ b/.agents/skills/bearings/assets/board-template.html @@ -548,22 +548,26 @@ var fd = new FormData(form); var value = fd.get("answer"); var note = (fd.get("note") || "").trim(); - /* picked option, optionally annotated; a bare note is itself the answer */ - var answer = value ? (note ? value + " - " + note : value) : note; - if (!answer) return; - if (utf8ByteLength(answer) > 512) { + var displayAnswer = value ? (note ? value + " - " + note : value) : note; + if (!displayAnswer) return; + if (utf8ByteLength(displayAnswer) > 512) { answerLimit.textContent = "Answer is too long to queue (512 bytes maximum)."; answerLimit.classList.add("is-visible"); return; } if (window.lavish && window.lavish.queuePrompt) { - /* close carries the composer-declared close mode: "release" frees a - captain-gated work item instead of completing a question task */ - var ctxData = { question: item.key, answer: answer }; + /* The versioned context keeps the selected option separate from its + note, while close carries the composer-declared completion mode. */ + var ctxData = { + schema: "fm-bearings-answer.v1", + question: item.key, + selection: value || "", + note: note + }; if (item.close) ctxData.close = item.close; window.lavish.queuePrompt( - "Captain's Call answer - " + item.title + ": " + answer, - { tag: "choice", text: item.title + " -> " + answer, element: form, + "Captain's Call answer - " + item.title + ": " + displayAnswer, + { tag: "choice", text: item.title + " -> " + displayAnswer, element: form, data: ctxData } ); } diff --git a/.agents/skills/captain-hold-lifecycle/SKILL.md b/.agents/skills/captain-hold-lifecycle/SKILL.md index a0fbf9a4996..e84f68af72f 100644 --- a/.agents/skills/captain-hold-lifecycle/SKILL.md +++ b/.agents/skills/captain-hold-lifecycle/SKILL.md @@ -24,7 +24,8 @@ After inventorying the whole report and review surface, run `bin/fm-captain-hold A completed investigation, a completed ADR design, and an ended visual review use this same owner and completion command; a design profile or visual tool, including Lavish, never owns a parallel completion policy. Run the command in the originating work's authoritative `FM_HOME`; secondmate-owned work registers in that secondmate home's backlog, and a question already held anywhere is never re-registered as a second row. Do not close a captain-held task merely because the originating investigation completed, its report was archived, its visual review ended, or its task was torn down. -Holding the work item the question gates is safe for exactly that reason: cleanup keeps such a row open with the finished work's deliverable recorded and returns it to the queue, so it still reads as the captain's own call and only `answer` closes it. +Holding the work item the question gates is safe for exactly that reason: cleanup keeps such a row open with the finished work's deliverable recorded and returns it to the queue, so it still reads as the captain's own call. +Only `answer` with the captain's words or an evidence-backed `reconcile close` may close it. Never close anything the captain owns without recording what he actually said: `bin/fm-captain-hold.sh answer` writes his exact words into the task and closes it in the same act, with `--release` when the answer frees a captain-gated work item to proceed instead of completing a question. When the answer changes what a task must build, follow `AGENTS.md` section 7's Validate contract to preserve the captain's words in the brief and steer the worker. @@ -32,6 +33,15 @@ When the captain says "later", that is an answer too: re-hold with `bin/fm-capta "A keyed answer closes its matching captain-held task" is one capability with one owner, `bin/fm-captain-hold.sh answers`, and every channel that carries a captain answer feeds it the same task id and answer; a channel never maps keys to tasks, records a decision, or closes anything itself. Chat already feeds it through `bin/fm-send.sh --resolve-key`, and a captured-answer source feeds it once bound with `bin/fm-captain-hold.sh bind `; bind before arming the source, and key each structured question by the held task's id. An unbound source and a key that names no captain-held task both simply feed nothing: the answer is still captured and firstmate is still woken, and closing falls back to the direct command above. +One answer value is reserved and closes nothing: `reconcile` means "go re-check reality", never "the captain answered", so the shared intake refuses it from every channel and creates nothing. +A bound captured source uses a separate seam: its adapter omits reconcile from keyed answers and emits the selected task id through `reconciles`, the generic runner feeds that into `reconcile-requests`, and the intake verifies the source binding and the local captain-held task before filing the durable board request. +A remote-secondmate card whose task is absent from the main backlog therefore remains announced but cannot create a main-home request; owner-aware request and mutation routing to the authoritative secondmate home is a separate follow-up. +That board-created request is yours to work off in the turn that receives it: `bin/fm-captain-hold.sh reconcile close --evidence-file ` records the EVIDENCE and closes a moot call, while `reconcile note --note-file ` annotates a genuinely active call and leaves it held. +Both outcomes refuse unless that task still has the pending request created by the captain's board selection, so neither is a standalone way to mutate a captain call. +A normal captain answer also retires any pending request because the call is settled, including close, release, and idempotent replay paths. +A retirement failure makes the command fail without reversing the already-durable answer, close, or note, and `reconcile list` keeps the surviving request visible for retry. +`reconcile list` names every request still outstanding. +Never use `answer` for an evidence-only moot call: `answer` records what the captain said, while `reconcile close` records verified evidence. A captain-held task closed outside this owner leaves no durable answer, so the completion gate keeps failing until `answer` records the decision the captain actually gave. Resolved findings, recommendations that need no captain choice, and prose that merely sounds decision-like do not create held tasks. Bearings reads the resulting structured state and must never compensate by scraping historical reports, visual-review artifacts, terminal output, chat, or other prose. @@ -49,8 +59,8 @@ The absence of a routed work item is not a divergence and the guard never requir 3. Hold that task - or create one captain-held task for the review's open questions - with a concise reason carrying the question and options. 4. Run `complete` with the full captain-held inventory for that review pass. 5. Relay the choices to the captain as decisions from Bearings' Captain's Call section under `AGENTS.md` section 9; do not use the word hold in captain chat. -6. Close each call only through `answer` (or a channel that feeds `answers`), through `--until` when the captain defers it, or confirm a channel already closed it. -7. Confirm Bearings reflects the outcome: answered calls leave Captain's Call, released work resumes, and deferred calls sit in Charted Next with their date. +6. Close each call only through `answer` (or a channel that feeds `answers`), close a board-requested moot call through evidence-backed `reconcile close`, record a still-active reconciliation through `reconcile note`, use `--until` when the captain defers it, or confirm a channel already closed it. +7. Confirm Bearings reflects the outcome: answered or reconciled-moot calls leave Captain's Call, released work resumes, active reconciliations remain held, and deferred calls sit in Charted Next with their date. `bin/fm-captain-hold.sh --help` owns command syntax, close modes, legacy-identity compatibility, completion attestation, retry behavior, and close ordering. `docs/captain-hold-lifecycle.md` records the mechanism and regression evidence without restating this policy. diff --git a/.agents/skills/firstmate-coding-guidelines/SKILL.md b/.agents/skills/firstmate-coding-guidelines/SKILL.md index bbfb4445450..9b1f1131ed0 100644 --- a/.agents/skills/firstmate-coding-guidelines/SKILL.md +++ b/.agents/skills/firstmate-coding-guidelines/SKILL.md @@ -102,9 +102,10 @@ Every such check needs two tests, because they fail for different reasons: - A portable regression in `tests/` that pins the logic with real processes and no harness, so CI enforces the classifier everywhere it runs tmux. Drive the signals apart deliberately and assert the verdict survives losing one; assert the divergence itself so the case cannot go quietly vacuous. Confirm which signal a given construction actually blinds on each supported platform rather than assuming, because the same trick can break different sources on macOS and Linux. -- A live guard in the `live-harness-optin` family (`bin/fm-test-run.sh`), env-gated and self-skipping, that exercises every INSTALLED harness for real and fails naming the harness and version. +- A live guard in the `live-harness-optin` family (`bin/fm-test-run.sh`) that exercises every INSTALLED harness for real and fails naming the harness and version. Report an absent harness explicitly rather than passing silently over it, and refuse a pass that checked nothing. - This guard is opt-in and on-demand because standard CI has neither harness binaries nor credentials; run it after every harness upgrade and before trusting refreshed per-harness evidence. + Open it with `fm_live_gate` from `tests/lib.sh`, which is the single owner of that decision: a guard that spends no model tokens runs by default wherever its tools are installed, a guard that submits prompts stays opt-in, and its own variable or `FM_LIVE` forces it on (an absent tool then fails rather than skips) or off. + The portable serial CI lane has no credentials and installs the public Pi package, so token-free guards exercise the available Pi surfaces while unavailable tools capability-skip; run a prompt-submitting guard after every harness upgrade and before trusting refreshed per-harness evidence. Record the dated per-harness result in `docs/verification/runtime-backends.md`, and point at the live guard as the command that refreshes it, rather than leaving a version-scoped observation to rot into a false claim. diff --git a/.agents/skills/harness-adapters/references/harness/opencode.md b/.agents/skills/harness-adapters/references/harness/opencode.md index 0d0eb6912fe..8b37a8d35ad 100644 --- a/.agents/skills/harness-adapters/references/harness/opencode.md +++ b/.agents/skills/harness-adapters/references/harness/opencode.md @@ -37,6 +37,7 @@ The primary integration was verified on 2026-07-08 with OpenCode 1.17.6. Throwing from `session.idle` does not block `opencode run`, so the primary adapter treats the event as passive and uses `client.session.promptAsync` to force one follow-up turn when `../../../bin/fm-turnend-guard.sh` returns 2. The follow-up was verified in the interactive TUI. `opencode run` can exit before displaying a queued follow-up, so the adapter steps aside in headless mode. +On native Windows, the operational-input adapter runs its Bash helper through `bash`; macOS and Linux invoke it directly. The companion `.opencode/plugins/fm-primary-watch-arm.js` owns normal TUI watcher supervision, wakes it with `client.session.promptAsync`, and coordinates with the guard before a blind-turn follow-up. The PreToolUse-equivalent watcher-arm seatbelt blocks by throwing from `tool.execute.before`. diff --git a/.agents/skills/harness-adapters/references/harness/pi.md b/.agents/skills/harness-adapters/references/harness/pi.md index ebbca27ddc6..0455efe6681 100644 --- a/.agents/skills/harness-adapters/references/harness/pi.md +++ b/.agents/skills/harness-adapters/references/harness/pi.md @@ -44,6 +44,7 @@ Pi sets `PI_CODING_AGENT=true` for its children as its harness-detection marker. The primary turn-end behavior was verified on 2026-07-09 with Pi 0.80.5. `.pi/extensions/fm-primary-turnend-guard.ts` listens for logical-run `agent_settled`, not per-tool-loop `turn_end`, and uses `pi.sendUserMessage(..., { deliverAs: "followUp" })` to force one guarded follow-up when `../../../bin/fm-turnend-guard.sh` returns 2. Without `deliverAs: "followUp"`, Pi rejects the send while the agent is still processing. +On native Windows, the extension runs its session-start, both PreToolUse, turn-end, and operational-input Bash helpers through `bash`; macOS and Linux invoke those helpers directly. The primary watcher protocol also requires `.pi/extensions/fm-primary-pi-watch.ts`. The Pi engine auto-discovers both tracked project-local extensions once the project is trusted. diff --git a/.agents/skills/process-event-sources/SKILL.md b/.agents/skills/process-event-sources/SKILL.md index 8cd1c78e966..a3a6cf8a5f7 100644 --- a/.agents/skills/process-event-sources/SKILL.md +++ b/.agents/skills/process-event-sources/SKILL.md @@ -119,13 +119,15 @@ Supported by tests: - the handled acknowledgement is generation-keyed to the exact source and sequence, private, path-safe, durable, and idempotent, and is the only thing that stops re-announcement; - one identity-matched owner per canonical source, across homes that share one underlying source store; - registration and ownership transitions share one per-source boundary, release is generation-bound, and uncertain process identity preserves the source for retry; -- ownership moves only once a whole generation is gone, so a crashed runner leader whose owned process group is still running never reads as stale: that surviving group is stopped before any replacement starts, and the claim is kept for retry when it cannot be; +- leaderless PID/PGID-reuse ambiguity preserves the claim without signalling or replacement, as owned by the operating contract in [`docs/configuration.md`](../../../docs/configuration.md#process-to-event-sources-stateprocevent); +- runner lifetime, owner-lease, and launch-pacing guarantees follow the operating contract in [`docs/configuration.md`](../../../docs/configuration.md#process-to-event-sources-stateprocevent); - stored argv is executed directly, so an argument containing spaces or shell metacharacters is never re-split or interpreted; - oversized output is bounded rather than published whole or silently dropped. The `when` adapter's guarantees are part of the operating contract in [`docs/configuration.md`](../../../docs/configuration.md#process-to-event-sources-stateprocevent). **Not true, and never to be claimed:** at-least-once, no-loss, or lossless delivery, and no generic exactly-once effect either - the handled acknowledgement only stops re-announcement, it says nothing about whether a paired external effect performed before the acknowledgement call actually completed, so a crash between that effect and the call can still repeat the effect on the next replay. +Also never claim that a source cannot refresh its owning home's lease: that rule is confused-agent-grade and a deliberately marker-stripping source is out of scope, per the operating contract in [`docs/configuration.md`](../../../docs/configuration.md#process-to-event-sources-stateprocevent). The currently published `lavish-axi poll` destructively clears feedback before returning it. A result lost after that clearing and before the runner reads the process output is unrecoverable, and no firstmate wrapper can close that source-side window. diff --git a/.gitattributes b/.gitattributes new file mode 100644 index 00000000000..7bdc6ed6b5a --- /dev/null +++ b/.gitattributes @@ -0,0 +1,2 @@ +# Bash parses shell scripts with LF line endings on every supported platform. +*.sh text eol=lf diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 45f97c94c40..d9e93158528 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -47,7 +47,7 @@ jobs: tests-portable-parallel-1: name: Behavior portable parallel 1 runs-on: ubuntu-latest - # Measured shard wall is ~2.2 min of serial sum on proven scripts; this cap + # Measured shard wall is ~7 min of serial sum on proven scripts; this cap # is a hang tripwire with margin, not the expected healthy end of the lane. timeout-minutes: 10 steps: @@ -69,11 +69,17 @@ jobs: set -eu npm install -g tasks-axi tasks-axi --version + - name: Install the Pi package for the Pi extension tests + run: | + set -eu + npm install -g @earendil-works/pi-coding-agent + npm ls -g --depth 0 @earendil-works/pi-coding-agent - name: Run portable parallel shard 1 run: | set -eu mkdir -p "$RUNNER_TEMP/fm-test" bin/fm-test-run.sh --lane portable-parallel-1 \ + --fail-on-gate-skip 'Pi extension typecheck prerequisite not found' \ --json "$RUNNER_TEMP/fm-test/fm-test-timing-portable-parallel-1.json" - name: Upload shard 1 timing artifact if: always() @@ -172,6 +178,14 @@ jobs: set -eu npm install -g tasks-axi tasks-axi --version + # The Pi extension tests read the installed Pi package's own types and + # runtime, so without it they gate-skip and pass silently. It is a public + # npm package and needs no credential, so CI can hold the real thing. + - name: Install the Pi package for the Pi extension tests + run: | + set -eu + npm install -g @earendil-works/pi-coding-agent + npm ls -g --depth 0 @earendil-works/pi-coding-agent - name: Run portable serial shard ${{ matrix.shard }} env: # job-total rather than a literal, so shrinking or growing the matrix @@ -182,7 +196,10 @@ jobs: run: | set -eu mkdir -p "$RUNNER_TEMP/fm-test" + # The Pi package is installed above and CI provides npm and tsc, so + # any missing typecheck prerequisite is a broken lane, not a valid skip. bin/fm-test-run.sh --lane "$FM_SERIAL_LANE" \ + --fail-on-gate-skip 'Pi extension typecheck prerequisite not found' \ --json "$RUNNER_TEMP/fm-test/fm-test-timing-portable-serial-${FM_SERIAL_SHARD}.json" - name: Upload portable serial shard ${{ matrix.shard }} timing artifact if: always() diff --git a/.opencode/plugins/lib/fm-operational-input.js b/.opencode/plugins/lib/fm-operational-input.js index f0d05c14448..64c65fd60ee 100644 --- a/.opencode/plugins/lib/fm-operational-input.js +++ b/.opencode/plugins/lib/fm-operational-input.js @@ -13,7 +13,10 @@ export function encodeFirstmateOperationalInput(root, kind, content) { const script = existsSync(requested) ? requested : `${adapterRoot}/bin/fm-operational-input.sh`; - const child = spawn(script, ["encode", kind], { + const invocation = process.platform === "win32" + ? { command: "bash", args: [script, "encode", kind] } + : { command: script, args: ["encode", kind] }; + const child = spawn(invocation.command, invocation.args, { stdio: ["pipe", "pipe", "pipe"], }); let stdout = ""; diff --git a/.pi/extensions/fm-branch-supervision.ts b/.pi/extensions/fm-branch-supervision.ts index c456f00f5a2..d9cef8d7a43 100644 --- a/.pi/extensions/fm-branch-supervision.ts +++ b/.pi/extensions/fm-branch-supervision.ts @@ -78,6 +78,7 @@ import { getAgentDir, keyHint, ModelRuntime, + type ModelRegistry, SessionManager, ToolExecutionComponent, type AgentSession, @@ -644,6 +645,11 @@ export default function (pi: ExtensionAPI) { // extension plus its model_select event, because createBranch runs at wake // time with no context of its own. It is what "follow main" applies. let mainModel: { provider: string; id: string } | null = null; + // Main's own model registry, captured from the contexts Pi hands this + // extension the same way mainModel is. It is the ONLY read path to + // providers an extension registered at runtime (pi-devin-auth's "devin"), + // which the branch's isolated ModelRuntime cannot see on its own. + let mainModelRegistry: ModelRegistry | null = null; // Main's own current effort needs no such tracking: Pi answers it directly // on demand, including at wake time. It throws only when the extension @@ -657,8 +663,9 @@ export default function (pi: ExtensionAPI) { } } - function rememberMainModel(ctx?: { model?: { provider: string; id: string } }): void { + function rememberMainModel(ctx?: { model?: { provider: string; id: string }; modelRegistry?: ModelRegistry }): void { if (ctx?.model) mainModel = { provider: ctx.model.provider, id: ctx.model.id }; + if (ctx?.modelRegistry) mainModelRegistry = ctx.modelRegistry; } function deliverBranchHealthNote(text: string): void { @@ -708,10 +715,55 @@ export default function (pi: ExtensionAPI) { // and same user as main, so stored credentials keep their own semantics // (OAuth stays OAuth, an API key stays an API key) and nothing is ever // installed, converted, derived, or overwritten here. + // A provider that exists only because an extension registered it into + // main's runtime (pi-devin-auth's "devin", whose streamSimple is the custom + // gRPC path no static catalog can express) is invisible to an isolated + // branch runtime until its registration is copied across. The config object + // carries that streamSimple and oauth wiring by reference, so copying it + // reuses the provider's own registration rather than reimplementing its + // wire protocol; the copy is never persisted and stays scoped to this one + // runtime. One registration that fails to compose must not blind the rest, + // so each copy is isolated. A just-registered provider's auth check has not + // run yet, so the copied providers are refreshed here and every caller's + // hasConfiguredAuth verdict is real rather than the provisional entry + // registration leaves behind. + async function copyExtensionProviders(modelRuntime: ModelRuntime): Promise { + if (!mainModelRegistry) return; + let providerIds: readonly string[]; + try { + providerIds = mainModelRegistry.getRegisteredProviderIds(); + } catch { + return; + } + const copied: string[] = []; + for (const providerId of providerIds) { + try { + const config = mainModelRegistry.getRegisteredProviderConfig(providerId); + if (config) { + modelRuntime.registerProvider(providerId, config); + copied.push(providerId); + } + } catch { + // A registration that fails to compose in the isolated runtime leaves + // that provider unavailable, exactly as if it were never copied. + } + } + if (copied.length === 0) return; + try { + await modelRuntime.refresh({ providers: copied, allowNetwork: false }); + } catch { + // A failed availability refresh is answered by hasConfiguredAuth. + } + } + async function resolveBranchModel(provider: string, modelId: string): Promise { const label = `${provider}/${modelId}`; const modelRuntime = await ModelRuntime.create(); - const model = modelRuntime.getModel(provider, modelId) as BranchModel | undefined; + let model = modelRuntime.getModel(provider, modelId) as BranchModel | undefined; + if (!model) { + await copyExtensionProviders(modelRuntime); + model = modelRuntime.getModel(provider, modelId) as BranchModel | undefined; + } if (!model) return { ok: false, reason: `${label} is unavailable to the isolated branch runtime` }; if (!modelRuntime.hasConfiguredAuth(provider)) { return { ok: false, reason: `${label} has no configured credentials in the isolated branch runtime` }; @@ -1659,6 +1711,7 @@ ${context.command} let available: string[]; try { const modelRuntime = await ModelRuntime.create(); + await copyExtensionProviders(modelRuntime); available = ctx.modelRegistry .getAvailable() .filter((model) => modelRuntime.getModel(model.provider, model.id) && modelRuntime.hasConfiguredAuth(model.provider)) diff --git a/.pi/extensions/fm-primary-turnend-guard.ts b/.pi/extensions/fm-primary-turnend-guard.ts index da4dc6b614e..b8f2285561a 100644 --- a/.pi/extensions/fm-primary-turnend-guard.ts +++ b/.pi/extensions/fm-primary-turnend-guard.ts @@ -7,6 +7,7 @@ import type { ExtensionAPI } from "@earendil-works/pi-coding-agent"; import { classifyFirstmateCurrentOperationalText, encodeFirstmateOperationalInput, + firstmateShellInvocation, } from "./lib/fm-operational-input.ts"; let guardFollowupActive = false; @@ -252,19 +253,26 @@ function runSessionstartHook(generation: SessionstartGeneration): Promise { return new Promise((resolveResult) => { - const child = spawn(`${root}/bin/fm-turnend-guard.sh`, { - stdio: ["pipe", "ignore", "pipe"], - }); + const invocation = firstmateShellInvocation(`${root}/bin/fm-turnend-guard.sh`, []); + let child: ChildProcess; + try { + child = spawn(invocation.command, invocation.args, { + stdio: ["pipe", "ignore", "pipe"], + }); + } catch { + resolveResult({ code: 0, stderr: "" }); + return; + } let stderr = ""; - child.stderr.on("data", (chunk) => { + child.stderr?.on("data", (chunk) => { stderr += chunk.toString(); }); child.on("error", () => resolveResult({ code: 0, stderr: "" })); @@ -452,8 +467,8 @@ function runGuard(): Promise<{ code: number; stderr: string }> { // A guard that exits before draining stdin fails this write with EPIPE. // That is the guard's answer, not a host failure, so it must not reach the // process as an unhandled stream error: close above already carries the code. - child.stdin.on("error", () => {}); - child.stdin.end('{"stop_hook_active":false}'); + child.stdin?.on("error", () => {}); + child.stdin?.end('{"stop_hook_active":false}'); }); } @@ -466,11 +481,21 @@ function runGuard(): Promise<{ code: number; stderr: string }> { // script owns its own decision and is inert outside the real primary checkout. function runChecker(script: string, command: string): Promise<{ code: number; stderr: string }> { return new Promise((resolveResult) => { - const child = spawn(`${root}/bin/${script}`, ["--command", command], { - stdio: ["ignore", "ignore", "pipe"], - }); + const invocation = firstmateShellInvocation( + `${root}/bin/${script}`, + ["--command", command], + ); + let child: ChildProcess; + try { + child = spawn(invocation.command, invocation.args, { + stdio: ["ignore", "ignore", "pipe"], + }); + } catch { + resolveResult({ code: 0, stderr: "" }); + return; + } let stderr = ""; - child.stderr.on("data", (chunk) => { + child.stderr?.on("data", (chunk) => { stderr += chunk.toString(); }); child.on("error", () => resolveResult({ code: 0, stderr: "" })); diff --git a/.pi/extensions/lib/fm-operational-input.ts b/.pi/extensions/lib/fm-operational-input.ts index 7102f1d6062..4070684c6a4 100644 --- a/.pi/extensions/lib/fm-operational-input.ts +++ b/.pi/extensions/lib/fm-operational-input.ts @@ -21,6 +21,15 @@ export type FirstmateCurrentOperationalKind = type OperationalInputCommand = "encode" | "classify" | "kind"; +export function firstmateShellInvocation( + script: string, + args: readonly string[], +): { command: string; args: string[] } { + return process.platform === "win32" + ? { command: "bash", args: [script, ...args] } + : { command: script, args: [...args] }; +} + // The one owner of how each command is invoked and how its exit status and // stdout become an answer, shared by the synchronous and awaited callers // below so the two can never drift. @@ -45,12 +54,20 @@ function runOperationalInputCommand( content: string, kind?: FirstmateCurrentOperationalKind, ): string | undefined { - const result = spawnSync(operationalInputScript, operationalInputArgs(command, kind), { - encoding: "utf8", - input: content, - maxBuffer: 1024 * 1024, - }); - return operationalInputAnswer(command, result.status, result.stdout); + const invocation = firstmateShellInvocation( + operationalInputScript, + operationalInputArgs(command, kind), + ); + try { + const result = spawnSync(invocation.command, invocation.args, { + encoding: "utf8", + input: content, + maxBuffer: 1024 * 1024, + }); + return operationalInputAnswer(command, result.status, result.stdout ?? ""); + } catch { + return undefined; + } } function encodeFailure(kind: FirstmateCurrentOperationalKind): Error { @@ -85,7 +102,11 @@ export async function encodeFirstmateOperationalInputWith( kind: FirstmateCurrentOperationalKind, content: string, ): Promise { - const result = await run(operationalInputScript, operationalInputArgs("encode", kind), { input: content }); + const invocation = firstmateShellInvocation( + operationalInputScript, + operationalInputArgs("encode", kind), + ); + const result = await run(invocation.command, invocation.args, { input: content }); const encoded = operationalInputAnswer("encode", result.status, result.stdout); if (encoded === undefined) throw encodeFailure(kind); return encoded; diff --git a/AGENTS.md b/AGENTS.md index f15ba1dadf1..dd8068f6fae 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -150,6 +150,7 @@ state/ runtime records and signals; gitignored procevent/ registered process-to-event sources, one private record per canonical source id; written only by bin/fm-procevent.sh, and their presence alone keeps supervision required (section 13) procevent-inbox/ private captured results and their durable handled-acknowledgement markers; source output lives here and never in an event line decision-bindings/ private records marking a captured-answer source as feeding the keyed-answer intake, with a legacy origin on pre-collapse records; written only by bin/fm-captain-hold.sh bind, dropped by unbind and by source retirement (section 13; docs/captain-hold-lifecycle.md) + reconcile-requests/ private open obligations to re-check a captain call whose board selection was `reconcile`; written only by bin/fm-captain-hold.sh, retired by its verify-then-decide outcomes or a normal answer that settles the call (section 13; docs/captain-hold-lifecycle.md) when/ private condition->action watch specs, their trust bindings, and single-fire markers; written only by bin/fm-procevent-when.sh (section 13's process-event-sources trigger) inbox/ captain notes captured out of band by bin/fm-inbox.sh, including the voice handover's queued requests; each note appends one `check` wake and stays pending until acknowledged with `bin/fm-inbox.sh drain --ack `, which moves it to inbox/handled/ (docs/voice-relay.md) x-inbox/ generated Relay pending mention payloads; fmx-respond drains it (section 14) diff --git a/bin/backends/herdr.sh b/bin/backends/herdr.sh index 4e10e0d2fbd..d554555f68d 100644 --- a/bin/backends/herdr.sh +++ b/bin/backends/herdr.sh @@ -699,7 +699,7 @@ fm_backend_herdr_presentation_lock_namespace() { fm_backend_herdr_presentation_lock_namespace_mode() { if [ "$(uname -s 2>/dev/null)" = Darwin ]; then - stat -f '%Lp' "$1" 2>/dev/null + /usr/bin/stat -f '%Lp' "$1" 2>/dev/null else stat -c '%a' "$1" 2>/dev/null fi @@ -707,7 +707,7 @@ fm_backend_herdr_presentation_lock_namespace_mode() { fm_backend_herdr_presentation_lock_namespace_uid() { if [ "$(uname -s 2>/dev/null)" = Darwin ]; then - stat -f '%u' "$1" 2>/dev/null + /usr/bin/stat -f '%u' "$1" 2>/dev/null else stat -c '%u' "$1" 2>/dev/null fi diff --git a/bin/fm-backlog-receive.sh b/bin/fm-backlog-receive.sh index 15d9bde99ae..46cf06783fd 100755 --- a/bin/fm-backlog-receive.sh +++ b/bin/fm-backlog-receive.sh @@ -57,7 +57,7 @@ list_keys() { # lock_age() { local modified now if [ "$(uname 2>/dev/null)" = Darwin ]; then - modified=$(stat -f '%m' "$1" 2>/dev/null) || return 1 + modified=$(/usr/bin/stat -f '%m' "$1" 2>/dev/null) || return 1 else modified=$(stat -c '%Y' "$1" 2>/dev/null) || return 1 fi diff --git a/bin/fm-bearings-board.sh b/bin/fm-bearings-board.sh index cff3cfb69cc..b25ad5e9c10 100755 --- a/bin/fm-bearings-board.sh +++ b/bin/fm-bearings-board.sh @@ -11,22 +11,58 @@ # fm-bearings-board.sh build # fm-bearings-board.sh path # -# build Validate the payload and inject it into a fresh copy of the shipped -# template at the stable board path. Establish or resume the Lavish -# session on that board BEFORE binding and arming its answer source, -# so a registered poll can never race a session that does not exist. +# build Validate the payload, drop the Captain's Call cards whose subject +# already landed, give every surviving decision card the standard +# reconcile choice, and inject the result into a fresh copy of the +# shipped template at the stable board path. Establish the Lavish +# session on that board and PROVE it is live BEFORE binding and +# arming its answer source, so a registered poll can never race a +# session that does not exist or attach to one that has ended. # Bind to the keyed-answer intake (bin/fm-captain-hold.sh) ALWAYS # precedes arm, so the board can never produce an answer that has # nowhere to go (captain-hold-lifecycle's ordering rule, enforced # here rather than left to agent memory). Output starts with # `board: `, then includes lavish-axi's session output and # the remaining status: +# session: live | reopened # served: # bound: # armed: (first registration) # already-armed: (registration already present) +# listening: (only when a replacement was needed) +# Every dropped card is named on stderr as a `dropped-landed-card:` +# line, so a rebuild states what it removed instead of quietly +# shrinking Captain's Call. # path Print the stable board path for this home. # +# A LIVE SESSION IS PROVED, NEVER ASSUMED. `lavish-axi ` exits 0 even +# when it refuses to reopen a session the captain ended from the browser, +# reporting `status: user-ended` with the same session id, so exit status alone +# cannot tell a live board from a dead one. build requires the server's fresh +# session listing to show the canonical board open and refuses rather than +# arming an ended session. After a reopen it retires the pre-reopen source +# generation through the guarded adapter path, arms a fresh registration, and +# accepts only the replacement listener as live. A registered board with no +# live owner also gets a replacement before build returns, because +# `already-armed` is not the same fact as `listening`. +# +# CAPTAIN'S CALL HYGIENE. A decision card is dropped when its work item, PR, or +# structured artifact/version subject appears among the payload's own landed +# rows, or when `bin/fm-captain-hold.sh open` reports the task is no longer an +# open captain call. A newer published version also supersedes a version card. +# A task whose state cannot be established is kept, because a call wrongly +# hidden is worse than a card wrongly shown. Cleanup is therefore a normal +# rebuild effect rather than a committed migration or direct state mutation. +# +# THE RECONCILE CHOICE. Every decision card carries the standard `reconcile` +# option, injected here so the guarantee does not depend on the composer's +# memory, and the payload validator reserves that value across every card type. +# The validator's reservation scope must equal the adapter's reconcile +# classification scope, which is all card types because the captured payload +# carries no card type. Its meaning, and the reason it can never reach the +# keyed-answer intake as a blind close, are owned by +# docs/captain-hold-lifecycle.md. +# # Validation is fail-closed: the payload must be valid JSON with # schema=fm-bearings-board.v1 and every renderer-consumed field must satisfy # the fm-bearings-board.v1 types and item invariants below. Every fleet row and @@ -79,6 +115,14 @@ validate_payload() { # or (.[$name] | type == "string" and test("^https://[A-Za-z0-9](?:[A-Za-z0-9.-]*[A-Za-z0-9])?(?::[0-9]{1,5})?(?:[/?#][^[:space:]]*)?$")); + def version: type == "string" and test("^(0|[1-9][0-9]{0,8})\\.(0|[1-9][0-9]{0,8})\\.(0|[1-9][0-9]{0,8})$"); + def optional_subject: + (has("subject") | not) + or (.subject + | type == "object" + and (keys | sort) == ["artifact", "version"] + and (.artifact | slug(128)) + and (.version | version)); def call_item: type == "object" and (.key | slug(128)) @@ -96,12 +140,16 @@ validate_payload() { # and (optional_string("decide")) and (optional_string("detail")) and (optional_https_url("pr_url")) + and optional_subject + and (if has("subject") then .type == "decision" else true end) and (optional_string("freeform_hint")) and ((has("close") | not) or (.close == "done" or .close == "release")) and ((has("allow_freeform") | not) or (.allow_freeform | type == "boolean")) and ((has("recommend_value") | not) or ((.recommend_value | slug(128)) - and (.recommend_value as $recommend | [.options[].value] | index($recommend) != null))) + and (.recommend_value as $recommend + | ([.options[].value] | index($recommend) != null)))) + and ([.options[].value] | index("reconcile") == null) and (if .type == "merge" then (.risk | nonempty_string) else true end); def underway_item: type == "object" and repo_marker and (.id | nonempty_string) @@ -109,7 +157,8 @@ validate_payload() { # def landed_item: type == "object" and repo_marker and (.id | nonempty_string) and (.what | nonempty_string) and (.owner | nonempty_string) - and optional_https_url("pr_url"); + and optional_https_url("pr_url") + and optional_subject; def charted_item: type == "object" and repo_marker and (.id | slug(128)) and (.title | nonempty_string) and (.reason | type == "string") @@ -136,8 +185,161 @@ validate_payload() { # ' "$1" >/dev/null } +# --- Lavish session liveness ------------------------------------------------- +# Verified against lavish-axi 0.1.61. `lavish-axi ` EXITS 0 even when it +# refuses to reopen a session the captain ended from the browser, reporting +# `status: user-ended` and the same session id, so an exit-code check alone +# cannot tell a live board from a dead one. The establish status is an initial +# signal only; the server's fresh session listing must also show the canonical +# board open before the build may bind or arm its source. + +board_realpath() { # + perl -MCwd=realpath -e '$p = realpath($ARGV[0]); defined($p) or exit 1; print "$p\n"' "$1" 2>/dev/null +} + +lavish_status_field() { # + printf '%s\n' "$1" | sed -n 's/^[[:space:]]*status:[[:space:]]*//p' | head -1 | tr -d '"' +} + +# The server's own listing, keyed on the canonical artifact path. Rows are +# `,,"",`, and only a live session is listed `open`. +lavish_session_listed_open() { # + local listing + listing=$(lavish-axi 2>/dev/null) || return 1 + printf '%s\n' "$listing" | awk -v path="$1" ' + { line = $0; sub(/^[[:space:]]+/, "", line) } + index(line, path ",") == 1 { + rest = substr(line, length(path) + 2) + split(rest, field, ",") + if (field[1] == "open") { found = 1 } + } + END { exit found ? 0 : 1 } + ' +} + +lavish_board_live() { # + lavish_session_listed_open "$2" +} + +# Establish the board session and PROVE it is live before anything arms a poll +# on it. A session the captain ended is reopened once - the captain asked for +# this board, which is exactly the attention `--reopen` exists for - and a +# session that is still not live after that refuses the build rather than +# arming a poll that can never attach. +establish_board_session() { # + local board=$1 real out status version + BOARD_SESSION_REOPENED=0 + real=$(board_realpath "$board") || fail "cannot resolve the board path: $board" + out=$(lavish-axi "$board") || fail "cannot establish the board Lavish session" + printf '%s\n' "$out" + if lavish_board_live "$out" "$real"; then + printf 'session: live\n' + return 0 + fi + out=$(lavish-axi "$board" --reopen) || fail "cannot reopen the ended board Lavish session" + printf '%s\n' "$out" + if lavish_board_live "$out" "$real"; then + BOARD_SESSION_REOPENED=1 + printf 'session: reopened\n' + return 0 + fi + status=$(lavish_status_field "$out") + version=$(lavish-axi --version 2>/dev/null | tr -d '[:space:]') + fail "the board Lavish session is not live after reopening it (lavish-axi ${version:-version-unknown} reported status ${status:-none}); refusing to arm a poll on an ended session" +} + +# --- Captain's Call hygiene --------------------------------------------------- +# A held decision whose subject already shipped is not a live call, so it is +# dropped here instead of being carded again. All checks use exact structured +# identities; unknown subject state keeps the card. + +decision_card_is_stale() { # + local task=$1 landed=$2 rc=0 + if [ "$landed" = 1 ]; then + printf 'structured subject already landed\n' + return 0 + fi + "$SCRIPT_DIR/fm-captain-hold.sh" open "$task" --distinguish-absent >/dev/null 2>&1 || rc=$? + # 1 is a definite "no longer an open captain call". 2 is "cannot tell", 3 is + # absent from this backlog, and a call wrongly hidden is worse than a card + # wrongly shown, so both uncertain and absent cards stay. + if [ "$rc" -eq 1 ]; then + printf 'no longer an open captain call\n' + return 0 + fi + return 1 +} + +# Drop every stale decision card, then give every surviving decision card the +# standard reconcile choice. Injecting it here is what makes "every decision +# card offers reconcile" a property of the board rather than of the composer's +# memory; the validator prevents duplicate decision options. +effective_payload() { # + local data=$1 dest=$2 landed_keys key reason drop='' tmp landed=0 + landed_keys=$(jq -c ' + def version_parts: split(".") | map(tonumber); + . as $payload + | [$payload.captains_call[] + | select(.type == "decision") + | . as $card + | select( + ($payload.landed | any(.id == $card.key)) + or (($card.pr_url? != null) and ($payload.landed | any(.pr_url? == $card.pr_url))) + or (($card.subject? != null) and ($payload.landed | any( + (.subject? != null) + and (.subject.artifact == $card.subject.artifact) + and ((.subject.version | version_parts) >= ($card.subject.version | version_parts))))) + ) + | .key] + ' "$data") || return 1 + while IFS= read -r key; do + [ -n "$key" ] || continue + landed=0 + if jq -e --arg key "$key" 'index($key) != null' <<< "$landed_keys" >/dev/null; then + landed=1 + fi + reason=$(decision_card_is_stale "$key" "$landed") || continue + printf 'dropped-landed-card: %s (%s)\n' "$key" "$reason" >&2 + drop=$drop$key$'\n' + done < <(jq -r '.captains_call[]? | select(.type == "decision") | .key' "$data") + tmp=$(printf '%s' "$drop" | jq -R -s 'split("\n") | map(select(length > 0))') || return 1 + jq --argjson dropped "$tmp" ' + .captains_call = [ + .captains_call[] + | . as $card + | select($card.type != "decision" or (($dropped | index($card.key)) == null)) + | if .type == "decision" + then .options += [{ + value: "reconcile", + label: "Reconcile", + hint: "Re-check the latest state, then close this with evidence or keep it open with a note" + }] + else . end + ]' "$data" > "$dest" || return 1 +} + +# The OWNER column bin/fm-procevent.sh already publishes: live, none, +# orphaned, or uncertain. Empty means the source is not registered at all. +source_owner() { # + "$SCRIPT_DIR/fm-procevent.sh" list 2>/dev/null \ + | awk -v id="$1" 'NR > 1 && $1 == id { print $3 }' +} + +# A replacement listener is started detached, so it claims the source shortly +# after reconcile returns. Wait for that claim rather than reporting the race. +await_source_owner() { # + local owner i=0 + while [ "$i" -lt 50 ]; do + owner=$(source_owner "$1") + [ "$owner" != live ] || { printf '%s\n' "$owner"; return 0; } + sleep 0.1 + i=$((i + 1)) + done + printf '%s\n' "${owner:-none}" +} + command_build() { - local data=${1-} board json tmp sid extracted + local data=${1-} board json tmp sid extracted effective owner version pre_reopen_owner [ "$#" -eq 1 ] || { usage >&2; exit 2; } command -v jq >/dev/null 2>&1 || fail "jq is required" [ -f "$data" ] || fail "board data does not exist: $data" @@ -147,7 +349,14 @@ command_build() { [ "$(grep -cxF "$PLACEHOLDER" "$TEMPLATE")" -eq 1 ] \ || fail "board template does not carry exactly one data slot: $TEMPLATE" - json=$(jq -c . "$data") || fail "cannot compact the board data" + effective=$(umask 077; mktemp "${TMPDIR:-/tmp}/fm-bearings-payload.XXXXXX") \ + || fail "cannot stage the board payload" + if ! effective_payload "$data" "$effective"; then + rm -f -- "$effective" + fail "cannot reconcile the board payload against landed work" + fi + json=$(jq -c . "$effective") || { rm -f -- "$effective"; fail "cannot compact the board data"; } + rm -f -- "$effective" # `<` never appears in JSON syntax outside strings, so escaping every # occurrence keeps the payload valid JSON while making inert. json=${json///dev/null 2>&1 || fail "lavish-axi is not installed" - lavish-axi "$board" || fail "cannot establish the board Lavish session" - printf 'served: %s\n' "$board" - sid=$("$SCRIPT_DIR/fm-procevent-lavish.sh" source-id "$board") \ || fail "cannot derive the board source id" + pre_reopen_owner=$(source_owner "$sid") + establish_board_session "$board" + if [ "$BOARD_SESSION_REOPENED" = 1 ]; then + "$SCRIPT_DIR/fm-procevent-lavish.sh" retire "$board" >/dev/null \ + || fail "cannot retire the pre-reopen source generation (observed owner: ${pre_reopen_owner:-none})" + fi + if ! lavish_session_listed_open "$(board_realpath "$board")"; then + version=$(lavish-axi --version 2>/dev/null | tr -d '[:space:]') + fail "the board Lavish session is not listed open immediately before arming (lavish-axi ${version:-version-unknown}); refusing to arm a poll on observed state not-open" + fi + printf 'served: %s\n' "$board" + "$SCRIPT_DIR/fm-captain-hold.sh" bind "$sid" >/dev/null \ || fail "cannot bind the board source to the keyed-answer intake" printf 'bound: %s\n' "$sid" - if "$SCRIPT_DIR/fm-procevent.sh" list | awk 'NR > 1 { print $1 }' | grep -Fxq "$sid"; then + owner=$(source_owner "$sid") + if [ "$BOARD_SESSION_REOPENED" = 1 ]; then + "$SCRIPT_DIR/fm-procevent-lavish.sh" arm "$board" >/dev/null \ + || fail "cannot arm a fresh board source after reopening" + printf 'armed: %s\n' "$sid" + owner=$(source_owner "$sid") + elif [ -n "$owner" ]; then printf 'already-armed: %s\n' "$sid" else "$SCRIPT_DIR/fm-procevent-lavish.sh" arm "$board" >/dev/null \ || fail "cannot arm the board as a process-event source" printf 'armed: %s\n' "$sid" + owner=$(source_owner "$sid") + fi + # Registered is not listening. A board whose source has no live owner gets a + # replacement started now rather than at the next supervision cycle, which is + # what keeps a rebuilt board from sitting silent behind `already-armed`. + if [ "$owner" != live ]; then + "$SCRIPT_DIR/fm-procevent.sh" reconcile >/dev/null 2>&1 || true + owner=$(await_source_owner "$sid") + if [ "$owner" != live ]; then + fail "source $sid is not listening after reconcile (observed owner: ${owner:-none})" + fi + printf 'listening: live\n' fi } diff --git a/bin/fm-bootstrap.sh b/bin/fm-bootstrap.sh index 236017f9b1f..2abbea0606d 100755 --- a/bin/fm-bootstrap.sh +++ b/bin/fm-bootstrap.sh @@ -1260,14 +1260,14 @@ x_mode_write_if_changed() { [ "$parent" != "$dest" ] || return 1 [ -d "$parent" ] && [ ! -L "$parent" ] || return 1 if [ "$(uname)" = Darwin ]; then - parent_device=$(stat -f %d "$parent" 2>/dev/null) || return 1 + parent_device=$(/usr/bin/stat -f %d "$parent" 2>/dev/null) || return 1 else parent_device=$(stat -c %d "$parent" 2>/dev/null) || return 1 fi if [ -e "$dest" ] || [ -L "$dest" ]; then fmx_single_link_file_valid "$dest" "$parent_device" || return 1 if [ "$(uname)" = Darwin ]; then - current_mode=$(stat -f %Lp "$dest" 2>/dev/null) || return 1 + current_mode=$(/usr/bin/stat -f %Lp "$dest" 2>/dev/null) || return 1 else current_mode=$(stat -c %a "$dest" 2>/dev/null) || return 1 fi diff --git a/bin/fm-busy-event.sh b/bin/fm-busy-event.sh index 0abcab8ee39..17fde457810 100755 --- a/bin/fm-busy-event.sh +++ b/bin/fm-busy-event.sh @@ -103,7 +103,7 @@ LOCK="$REC.lock" # gets a non-numeric token. Detect the platform once and pick the right form, # exactly as bin/fm-watch.sh does. if [ "$(uname)" = Darwin ]; then - lock_mtime() { stat -f %m "$1" 2>/dev/null; } + lock_mtime() { /usr/bin/stat -f %m "$1" 2>/dev/null; } else lock_mtime() { stat -c %Y "$1" 2>/dev/null; } fi diff --git a/bin/fm-captain-hold.sh b/bin/fm-captain-hold.sh index b54546dbd0f..3ebc13cb923 100755 --- a/bin/fm-captain-hold.sh +++ b/bin/fm-captain-hold.sh @@ -24,13 +24,17 @@ # [--title ] [--repo <repo>] [--origin <origin-id>] [--until YYYY-MM-DD] # fm-captain-hold.sh answer <task-id> --decision-file <path> [--release] # fm-captain-hold.sh answers [<legacy-origin> | --any-origin] --source <provenance> (keyed answers on stdin) +# fm-captain-hold.sh reconcile-requests --source-id <source-id> --source <provenance> (task ids on stdin) # fm-captain-hold.sh bind <source-id> [<legacy-origin> | --any-origin] # fm-captain-hold.sh unbind <source-id> # fm-captain-hold.sh binding <source-id> # fm-captain-hold.sh complete <origin-id> (--none | <task-id>...) # fm-captain-hold.sh verify <origin-id> -# fm-captain-hold.sh open <task-id> +# fm-captain-hold.sh open <task-id> [--identity] [--distinguish-absent] # fm-captain-hold.sh diverged +# fm-captain-hold.sh reconcile list +# fm-captain-hold.sh reconcile close <task-id> --evidence-file <path> +# fm-captain-hold.sh reconcile note <task-id> --note-file <path> # # `hold` places an existing task under an active captain hold, or creates the # task first when no work item exists to hold (--title required to create; the @@ -83,6 +87,24 @@ # keeps closing its rows; `--any-origin` and the stored `(any)` marker mean # what an absent origin means and are accepted for the same reason. # +# RECONCILE IS RESERVED AT THIS INTAKE, NOT FILTERED IN A CHANNEL. +# The exact answer value `reconcile` means "go re-check reality", never "the +# captain answered". `answers` matches it before it reads the close mode, +# visibly refuses it, and never passes it to `answer`, so no channel and no +# card-declared mode can turn it into a close, release, or request. A separate +# `reconcile-requests` intake verifies a captured source's binding before it +# records a durable request under `state/reconcile-requests/`. +# +# `reconcile` is the verify-then-decide half. Both outcomes require the pending +# request created by the captain's board selection. `close` is the moot outcome: +# it requires the evidence that made the call moot, writes a `reconciled` resolution record +# under a `Reconciliation evidence:` label so it can never read as the +# captain's words, and closes the task. `note` is the still-active outcome: it +# appends one dated `Captain hold reconciled:` note and leaves the hold in +# place. A normal answer also retires the request because the call is settled. +# `list` is the read-only enumeration. +# docs/captain-hold-lifecycle.md owns the semantics. +# # A channel's ONLY job is to turn whatever it received into those keyed lines # and pipe them here. It must never map keys to tasks, build decision records, # choose a close mode beyond what its card declared, or close anything itself. @@ -134,21 +156,33 @@ # retire a task's row: is this task still an open captain call? Exit 0 means it # is (not Done, hold kind captain), 1 means it is not, and 2 means the answer # could not be established, so a caller that must never close a live call can -# treat "cannot tell" as its own case instead of as a no. It prints nothing on -# 0 or 1 and mutates nothing. bin/fm-teardown.sh asks it before its automatic +# treat "cannot tell" as its own case instead of as a no. With +# `--distinguish-absent`, an absent local task returns 3 instead of 1. +# It prints nothing on these predicate results and mutates nothing, unless +# `--identity` asks it to print this call's +# LIFECYCLE identity, which it does on an exit 0 only. That identity - the +# hold-set stamp and the count of recorded answers - is what distinguishes two +# successive calls on one task id: re-holding released work starts a new +# lifecycle without necessarily touching the task's status log, so a consumer +# that bounds repeated work per call cannot use the task id alone. +# bin/fm-teardown.sh asks it before its automatic # backlog close and, on 0, returns the row to Queued with its deliverable # recorded instead (bin/fm-backlog-transition-lib.sh owns that transition), so -# holding the very work item a question gates is safe; `answer` remains the -# only act that closes a captain call. +# holding the very work item a question gates is safe; only `answer` with the +# captain's words or evidence-backed `reconcile close` closes the call. +# bin/fm-watch.sh asks it when an ordinary +# crew task reaches a due stale alarm - its open backlog hold need not appear in +# the task's last status line - and on a 0 bounds repeated alarms from new pane +# hashes for the decision. # # `diverged` is the read-only guard over the seam between the two records of # one captain call. See "record divergence" beside command_diverged below. # # Resolution records: the block written into the body names this script, the -# decision digest, and a `Resolution mode:` of answered, released, or repaired. -# Records written by the retired fm-decision-hold.sh (routed, declined, -# answered, repaired) are recognized everywhere a record is read, so nothing -# already closed needs rewriting. +# decision digest, and a `Resolution mode:` of answered, released, repaired, or +# reconciled. Records written by the retired fm-decision-hold.sh (routed, +# declined, answered, repaired) are recognized everywhere a record is read, so +# nothing already closed needs rewriting. # # Parent channel: inside a secondmate home a task held for the captain, and its # answer, are captain-facing facts the moment they are recorded, so `hold` @@ -186,12 +220,14 @@ DATA="${FM_DATA_OVERRIDE:-$FM_HOME/data}" # shellcheck disable=SC1091 . "$SCRIPT_DIR/fm-parent-channel-lib.sh" +PARENT_HOLD_PUBLISHED=0 publish_parent_hold() { # <task-id> <occurrence> <verb> <note> local id=$1 occurrence=$2 verb=$3 note=$4 rc=0 + PARENT_HOLD_PUBLISHED=0 fm_parent_channel_report "$FM_HOME" "$STATE" \ "$verb [key=captain-hold-$id-$occurrence]: captain hold $id: $(fm_parent_channel_clean_note "$note")" || rc=$? case "$rc" in - 0|1) ;; + 0|1) PARENT_HOLD_PUBLISHED=1 ;; *) printf 'actionable: task %s is held for the captain in this home but that did not reach the parent channel (rc=%s)\n' "$id" "$rc" >&2 ;; esac } @@ -246,6 +282,13 @@ acquire_task_control_lock() { # <task-id> CAPTAIN_CONTROL_LOCK_HELD=1 } +release_task_control_lock() { + [ "$CAPTAIN_CONTROL_LOCK_HELD" = 1 ] || return 0 + fm_lock_release "$CAPTAIN_CONTROL_LOCK" + CAPTAIN_CONTROL_LOCK_HELD=0 + CAPTAIN_CONTROL_LOCK= +} + sha256_text() { # <text> if command -v shasum >/dev/null 2>&1; then printf '%s' "$1" | shasum -a 256 | awk '{print $1}' @@ -388,6 +431,7 @@ body_has_resolution_record() { # <task-body> case "$1" in *"Resolution recorded by fm-captain-hold."*"Captain decision:"*) return 0 ;; *"Resolution recorded by fm-decision-hold."*"Captain decision:"*) return 0 ;; + *"Resolution recorded by fm-captain-hold."*"Reconciliation evidence:"*) return 0 ;; esac return 1 } @@ -426,9 +470,21 @@ recorded_resolution_mode() { # <task-body> printf '%s' "$rest" } +closed_answer_replay_mode_compatible() { # <mode> <task-body> + case "$1" in + answered|repaired|routed) return 0 ;; + esac + return 1 +} + +# The record's label is what keeps an evidence-backed reconciliation from +# reading as the captain's own words. `reconciled` closes a call that went moot +# and carries verified evidence; every other mode carries what the captain said. resolution_block() { # <mode> - printf 'Resolution recorded by fm-captain-hold.\nDecision digest: %s\nResolution mode: %s\n\nCaptain decision:\n%s\n' \ - "$DECISION_DIGEST" "$1" "$DECISION_TEXT" + local label='Captain decision:' + [ "$1" != reconciled ] || label='Reconciliation evidence:' + printf 'Resolution recorded by fm-captain-hold.\nDecision digest: %s\nResolution mode: %s\n\n%s\n%s\n' \ + "$DECISION_DIGEST" "$1" "$label" "$DECISION_TEXT" } # Durable state of one captain call: an active captain hold (annotations @@ -889,15 +945,15 @@ command_answer() { [ "$(recorded_decision_digest "$body" || true)" = "$DECISION_DIGEST" ] \ || fail "captain-held task $id records a different captain decision" recorded_mode=$(recorded_resolution_mode "$body" || true) - [ "$recorded_mode" != released ] \ - || fail "task $id records this answer with mode released; a closed task cannot replay that release" + closed_answer_replay_mode_compatible "$recorded_mode" "$body" \ + || fail "task $id records this resolution with mode ${recorded_mode:-unknown}; it is not a captain-answer replay" [ "$release" = 0 ] \ || fail "task $id records this answer with mode ${recorded_mode:-unknown}; --release cannot reopen a closed task" remove_interrupted_answer_stamp "$id" if [ "$recorded_mode" = repaired ]; then - publish_parent_hold "$id" $((occurrence - 1)) resolved "answered (repaired)" + publish_parent_resolution_then_retire "$id" $((occurrence - 1)) "answered (repaired)" else - publish_parent_hold "$id" $((occurrence - 1)) resolved answered + publish_parent_resolution_then_retire "$id" $((occurrence - 1)) answered fi printf 'answered: %s\n' "$id" return 0 @@ -914,7 +970,7 @@ command_answer() { [ "$(show_field "$show" state)" = "done" ] || fail "recording the answer reopened closed task $id" body_has_resolution_record "$(show_field "$show" body)" \ || fail "captain-held task $id did not retain its durable resolution record" - publish_parent_hold "$id" "$occurrence" resolved "answered (repaired)" + publish_parent_resolution_then_retire "$id" "$occurrence" "answered (repaired)" printf 'repaired: %s\n' "$id" return 0 fi @@ -931,13 +987,14 @@ command_answer() { recorded_mode=$(recorded_resolution_mode "$body" || true) case "$recorded_mode" in released) [ "$release" = 1 ] || fail "task $id records this answer as a release; retry with --release" ;; - answered) [ "$release" = 0 ] || fail "task $id records this answer as a close; retry without --release" ;; + answered|routed) [ "$release" = 0 ] || fail "task $id records this answer as a close; retry without --release" ;; + *) fail "task $id records this resolution with mode ${recorded_mode:-unknown}; it is not a captain-answer replay" ;; esac if ! close_answered "$id" "$release"; then fail "could not close answered captain-held task $id" fi remove_interrupted_answer_stamp "$id" - publish_parent_hold "$id" $((occurrence - 1)) resolved "$outcome" + publish_parent_resolution_then_retire "$id" $((occurrence - 1)) "$outcome" printf '%s: %s\n' "$outcome" "$id" return 0 fi @@ -949,7 +1006,7 @@ command_answer() { show=$(task_show "$id") || fail "task $id disappeared after closing" body_has_resolution_record "$(show_field "$show" body)" \ || fail "captain-held task $id did not retain its durable resolution record" - publish_parent_hold "$id" "$occurrence" resolved "$outcome" + publish_parent_resolution_then_retire "$id" "$occurrence" "$outcome" printf '%s: %s\n' "$outcome" "$id" return 0 fi @@ -962,7 +1019,7 @@ command_answer() { [ "$recorded_mode" = released ] && [ "$release" = 1 ] \ || fail "task $id records this answer with mode ${recorded_mode:-unknown}; replay requires matching --release" remove_interrupted_answer_stamp "$id" - publish_parent_hold "$id" $((occurrence - 1)) resolved released + publish_parent_resolution_then_retire "$id" $((occurrence - 1)) released printf 'released: %s\n' "$id" return 0 fi @@ -1060,6 +1117,10 @@ sanitize_field() { # <text> printf '%s' "$1" | tr '\n\r\t' ' ' | LC_ALL=C tr -d '\000-\037\177' | cut -c1-512 } +sanitize_reconcile_provenance() { + printf '%s' "$1" | tr '\n\r\t' ' ' | LC_ALL=C tr -d '\000-\037\177' | cut -c1-1024 +} + command_answers() { local origin='' source='' row rest key answer label mode id show state hold_kind body digest legacy_digest legacy_key local recorded_digest recorded_mode occurrence tmp err closed=0 skipped=0 reason release_flag tab=$'\t' @@ -1098,6 +1159,11 @@ command_answers() { answer=$(sanitize_field "${answer:-}") [ -n "$answer" ] || continue label=$(sanitize_field "${label:-}") + if [ "$answer" = "$RECONCILE_VALUE" ]; then + printf 'refused: %s (reconcile requests require a bound captured source)\n' "$key" + skipped=$((skipped + 1)) + continue + fi release_flag='' case "${mode:-}" in ''|done) : ;; @@ -1147,14 +1213,15 @@ command_answers() { && { [ "$recorded_digest" = "$digest" ] \ || { case "$body" in *"Resolution recorded by fm-decision-hold."*) true ;; *) false ;; esac \ && [ -n "$legacy_digest" ] && [ "$recorded_digest" = "$legacy_digest" ]; }; }; then - if { [ -z "$release_flag" ] && [ "$state" = "done" ] && [ "$recorded_mode" != released ]; } \ + if { [ -z "$release_flag" ] && [ "$state" = "done" ] \ + && closed_answer_replay_mode_compatible "$recorded_mode" "$body"; } \ || { [ "$release_flag" = --release ] && [ "$state" != "done" ] \ && [ "$hold_kind" != captain ] && [ "$recorded_mode" = released ]; }; then occurrence=$(resolution_record_count "$body") case "$recorded_mode" in - repaired) publish_parent_hold "$id" "$occurrence" resolved "answered (repaired)" ;; - released) publish_parent_hold "$id" "$occurrence" resolved released ;; - *) publish_parent_hold "$id" "$occurrence" resolved answered ;; + repaired) publish_parent_resolution_then_retire "$id" "$occurrence" "answered (repaired)" ;; + released) publish_parent_resolution_then_retire "$id" "$occurrence" released ;; + *) publish_parent_resolution_then_retire "$id" "$occurrence" answered ;; esac printf 'closed: %s\n' "$id" closed=$((closed + 1)) @@ -1189,6 +1256,273 @@ command_answers() { [ "$skipped" -eq 0 ] } +# --- reconcile: verify latest state, then close with evidence or annotate ---- +# +# The semantics are owned by docs/captain-hold-lifecycle.md; this section owns +# the durable record and the two terminal operations that retire it. Nothing +# here closes a captain call on the strength of a reconcile alone: `close` +# demands the evidence that made the call moot, and `note` leaves it open. + +RECONCILE_DIR="$STATE/reconcile-requests" +RECONCILE_SCHEMA=fm-reconcile-request.v1 +RECONCILE_VALUE=reconcile + +reconcile_request_path() { printf '%s/%s.request\n' "$RECONCILE_DIR" "$1"; } + +# Idempotent per task: a repeated reconcile keeps the one request and its +# original timestamp, so a re-delivered board answer never resets the clock on +# an obligation that is already open. +reconcile_request_record() { # <task-id> <provenance> + local id=$1 source=$2 path tmp + path=$(reconcile_request_path "$id") + [ ! -e "$path" ] || return 0 + (umask 077; mkdir -p "$RECONCILE_DIR") || return 1 + [ -d "$RECONCILE_DIR" ] && [ ! -L "$RECONCILE_DIR" ] || return 1 + tmp=$(umask 077; mktemp "$RECONCILE_DIR/.request.XXXXXX") || return 1 + if { + printf 'schema=%s\n' "$RECONCILE_SCHEMA" + printf 'task=%s\n' "$id" + printf 'requested=%s\n' "${FM_CAPTAIN_HOLD_NOW:-$(date -u +%Y-%m-%dT%H:%M:%SZ)}" + printf 'source=%s\n' "$(sanitize_reconcile_provenance "$source")" + } > "$tmp" && chmod 0600 "$tmp" && mv -f -- "$tmp" "$path"; then + return 0 + fi + rm -f -- "$tmp" + return 1 +} + +reconcile_request_read() { # <task-id>; sets RECONCILE_REQUESTED/RECONCILE_SOURCE + local id=$1 path schema task + path=$(reconcile_request_path "$id") + [ -f "$path" ] && [ ! -L "$path" ] || return 1 + schema=$(sed -n 's/^schema=//p' "$path" | head -1) + [ "$schema" = "$RECONCILE_SCHEMA" ] || fail "reconcile request has an incompatible schema: $path" + task=$(sed -n 's/^task=//p' "$path" | head -1) + [ "$task" = "$id" ] || fail "reconcile request names a different task: $path" + RECONCILE_REQUESTED=$(sed -n 's/^requested=//p' "$path" | head -1) + RECONCILE_SOURCE=$(sed -n 's/^source=//p' "$path" | head -1) +} + +reconcile_request_retire() { # <task-id> + rm -f -- "$(reconcile_request_path "$1")" \ + || fail "could not retire the pending reconcile request for $1" +} + +publish_parent_resolution_then_retire() { # <task-id> <occurrence> <note> + local id=$1 occurrence=$2 note=$3 request + request=$(reconcile_request_path "$id") + publish_parent_hold "$id" "$occurrence" resolved "$note" + if [ -e "$request" ] && [ "$PARENT_HOLD_PUBLISHED" != 1 ]; then + fail "could not publish the answered captain-held task $id to its parent" + fi + reconcile_request_retire "$id" +} + +command_reconcile_requests() { + local source_id='' source='' origin row id note provenance show created=0 skipped=0 tab=$'\t' + while [ "$#" -gt 0 ]; do + case "$1" in + --source-id) shift; source_id=${1:-} ;; + --source) shift; source=${1:-} ;; + *) usage >&2; exit 2 ;; + esac + shift + done + validate_source_id "$source_id" + [ -n "$source" ] || fail "--source provenance is required" + origin=$(read_binding "$source_id") || fail "cannot verify the binding for source $source_id" + [ -n "$origin" ] || fail "source $source_id is not bound; no reconcile requests were created" + require_tasks_axi + while IFS= read -r row; do + id=${row%%"$tab"*} + note='' + case "$row" in *"$tab"*) note=${row#*"$tab"} ;; esac + [ -n "$id" ] || continue + case "$id" in + *[!A-Za-z0-9._-]*) printf 'refused: %s (invalid task id)\n' "$id"; skipped=$((skipped + 1)); continue ;; + esac + [ "${#id}" -le 128 ] \ + || { printf 'refused: %s (task id is too long)\n' "$id"; skipped=$((skipped + 1)); continue; } + acquire_task_control_lock "$id" + show=$(task_show "$id") || true + if [ -z "$show" ]; then + printf 'refused: %s (absent)\n' "$id" + skipped=$((skipped + 1)) + elif [ "$(show_field "$show" state)" = "done" ]; then + printf 'refused: %s (already closed)\n' "$id" + skipped=$((skipped + 1)) + elif [ "$(show_field_value "$show" hold_kind)" != captain ]; then + printf 'refused: %s (not held for the captain)\n' "$id" + skipped=$((skipped + 1)) + else + provenance=$source + [ -z "$note" ] || provenance="$source; captain note: $(sanitize_field "$note")" + if reconcile_request_record "$id" "$provenance"; then + printf 'reconcile: %s\n' "$id" + created=$((created + 1)) + else + printf 'refused: %s (cannot record the reconcile request)\n' "$id" + skipped=$((skipped + 1)) + fi + fi + release_task_control_lock || fail "cannot release task control for $id" + done + printf 'reconcile-requests: created=%s skipped=%s\n' "$created" "$skipped" + [ "$skipped" -eq 0 ] +} + +command_reconcile() { + local action=${1:-} + [ "$#" -ge 1 ] || { usage >&2; exit 2; } + shift + case "$action" in + list) reconcile_list "$@" ;; + close) reconcile_close "$@" ;; + note) reconcile_note "$@" ;; + *) usage >&2; exit 2 ;; + esac +} + +reconcile_list() { + local path id count=0 + [ "$#" -eq 0 ] || { usage >&2; exit 2; } + [ -d "$RECONCILE_DIR" ] || { printf 'reconcile-requests: 0\n'; return 0; } + for path in "$RECONCILE_DIR"/*.request; do + [ -e "$path" ] || continue + id=${path##*/}; id=${id%.request} + RECONCILE_REQUESTED='' + RECONCILE_SOURCE='' + reconcile_request_read "$id" || continue + printf '%s\trequested=%s\tsource=%s\n' "$id" "$RECONCILE_REQUESTED" "$RECONCILE_SOURCE" + count=$((count + 1)) + done + printf 'reconcile-requests: %s\n' "$count" +} + +# The moot outcome. The evidence is what closes the call, and the `reconciled` +# resolution mode is what keeps the record from claiming the captain answered. +reconcile_close() { + local id=${1:-} evidence_file='' show state hold_kind body occurrence recorded_mode + [ "$#" -ge 1 ] || { usage >&2; exit 2; } + shift + while [ "$#" -gt 0 ]; do + case "$1" in + --evidence-file) shift; evidence_file=${1:-} ;; + *) usage >&2; exit 2 ;; + esac + shift + done + validate_slug task-id "$id" + [ -n "$evidence_file" ] || fail "--evidence-file is required; a moot call closes on evidence, never on assertion" + load_decision "$evidence_file" + acquire_task_control_lock "$id" + reconcile_request_read "$id" \ + || fail "task $id has no pending board-created reconcile request" + require_tasks_axi + show=$(task_show "$id") || fail "captain-held task $id is absent from this home's configured backlog (data directory $DATA)" + state=$(show_field "$show" state) + hold_kind=$(show_field_value "$show" hold_kind) + body=$(show_field "$show" body) + occurrence=$(( $(resolution_record_count "$body") + 1 )) + if [ "$state" = "done" ]; then + # An exact retry finishes an interrupted close and stays idempotent; a + # different evidence text on an already closed call is refused. + body_has_resolution_record "$body" \ + || fail "task $id is already closed with no resolution record; use answer to record what closed it" + [ "$(recorded_decision_digest "$body" || true)" = "$DECISION_DIGEST" ] \ + || fail "task $id records a different resolution; it cannot be reconciled again" + [ "$(recorded_resolution_mode "$body" || true)" = reconciled ] \ + || fail "task $id was not closed by reconciliation" + occurrence=$(resolution_record_count "$body") + remove_interrupted_answer_stamp "$id" + publish_parent_hold "$id" "$occurrence" resolved reconciled + [ "$PARENT_HOLD_PUBLISHED" = 1 ] \ + || fail "could not publish the reconciled captain-held task $id to its parent" + reconcile_request_retire "$id" + printf 'reconciled: %s\n' "$id" + return 0 + fi + [ "$hold_kind" = captain ] \ + || fail "task $id is not held for the captain; there is no captain call to reconcile" + if body_has_resolution_record "$body" \ + && [ "$(recorded_decision_digest "$body" || true)" = "$DECISION_DIGEST" ]; then + recorded_mode=$(recorded_resolution_mode "$body" || true) + [ "$recorded_mode" = reconciled ] \ + || fail "task $id records this resolution with mode ${recorded_mode:-unknown}; it is not a reconciliation retry" + occurrence=$(resolution_record_count "$body") + else + write_resolution_record "$id" reconciled "$body" + fi + close_answered "$id" 0 || fail "could not close reconciled captain-held task $id" + remove_interrupted_answer_stamp "$id" + show=$(task_show "$id") || fail "task $id disappeared after closing" + body_has_resolution_record "$(show_field "$show" body)" \ + || fail "captain-held task $id did not retain its durable resolution record" + publish_parent_hold "$id" "$occurrence" resolved reconciled + [ "$PARENT_HOLD_PUBLISHED" = 1 ] \ + || fail "could not publish the reconciled captain-held task $id to its parent" + reconcile_request_retire "$id" + printf 'reconciled: %s\n' "$id" +} + +# The still-active outcome. The hold survives, so the call stays the captain's +# and stays on Captain's Call, now carrying what the re-check found. +reconcile_note() { + local id=${1:-} note_file='' note show body stamp tmp note_digest marker + [ "$#" -ge 1 ] || { usage >&2; exit 2; } + shift + while [ "$#" -gt 0 ]; do + case "$1" in + --note-file) shift; note_file=${1:-} ;; + *) usage >&2; exit 2 ;; + esac + shift + done + validate_slug task-id "$id" + [ -n "$note_file" ] || fail "--note-file is required; leaving a call open records what the re-check found" + [ -f "$note_file" ] || fail "note file does not exist: $note_file" + note=$(cat "$note_file") + [ -n "$note" ] || fail "note file must not be empty" + [ "$(printf '%s' "$note" | LC_ALL=C wc -c | tr -d ' ')" -le 8192 ] \ + || fail "note file exceeds 8192 bytes" + acquire_task_control_lock "$id" + reconcile_request_read "$id" \ + || fail "task $id has no pending board-created reconcile request" + require_tasks_axi + command_open "$id" \ + || fail "task $id is not an open captain call; a note cannot keep a closed call open" + show=$(task_show "$id") || fail "captain-held task $id is absent from this home's configured backlog (data directory $DATA)" + body=$(decode_shown_value "$(show_field "$show" body)") \ + || fail "could not decode the existing body for $id" + note_digest=$(sha256_text "$note") + marker="Reconcile request: $RECONCILE_REQUESTED | $RECONCILE_SOURCE | note digest: $note_digest" + case "$body" in + *"$marker"*) + reconcile_request_retire "$id" \ + || fail "could not retire the applied reconcile request for $id" + command_open "$id" || fail "recording the reconcile note released captain-held task $id" + printf 'still-open: %s\n' "$id" + return 0 + ;; + esac + stamp=${FM_CAPTAIN_HOLD_NOW:-$(date -u +%Y-%m-%dT%H:%M:%SZ)} + tmp=$(umask 077; mktemp "${TMPDIR:-/tmp}/fm-captain-hold-note.XXXXXX") \ + || fail "cannot stage the reconcile note" + if ! printf '%s\n\nCaptain hold reconciled: %s\n%s\n%s\n' "$body" "$stamp" "$marker" "$note" > "$tmp"; then + rm -f -- "$tmp" + fail "cannot stage the reconcile note for $id" + fi + if ! tasks_axi update "$id" --body-file "$tmp" --archive-body >/dev/null; then + rm -f -- "$tmp" + fail "could not record the reconcile note on $id" + fi + rm -f -- "$tmp" + reconcile_request_retire "$id" \ + || fail "could not retire the applied reconcile request for $id" + command_open "$id" || fail "recording the reconcile note released captain-held task $id" + printf 'still-open: %s\n' "$id" +} + command_complete() { local origin=${1:-} meta previous='' supplied='' keys='' entry key status_file open raw_open has_meta=0 transfer_rc resolved local resolved_how attested_by_prefix='' @@ -1427,12 +1761,23 @@ EOF } # Still an open captain call? Exit 0 yes, 1 no, 2 cannot tell (see the header). -# A row this home does not carry holds no captain call, so an absent task is a -# plain no; every other read failure is a 2, printed to stderr, because a -# mechanical closer must never read "cannot tell" as permission to close. -command_open() { # <task-id> - local id=${1:-} data state - [ "$#" -eq 1 ] || { usage >&2; exit 2; } +# A row this home does not carry is 3 when the caller requests the distinction; +# every other read failure is a 2, printed to stderr, because a mechanical +# closer must never read "cannot tell" as permission to close. +command_open() { # <task-id> [--identity] [--distinguish-absent] + local id='' identity=0 distinguish_absent=0 data state show shown_body + while [ "$#" -gt 0 ]; do + case "$1" in + --identity) identity=1 ;; + --distinguish-absent) distinguish_absent=1 ;; + -*) usage >&2; exit 2 ;; + *) + [ -z "$id" ] || { usage >&2; exit 2; } + id=$1 + ;; + esac + shift + done case "$id" in ''|*[!A-Za-z0-9._-]*) printf 'fm-captain-hold: task id must be a non-empty privacy-safe slug: %s\n' "$id" >&2 @@ -1445,11 +1790,24 @@ command_open() { # <task-id> if fm_backlog_row_probe "$data" "$id"; then state=${FM_BACKLOG_ROW_STATE%% *} if [ "$state" != "done" ] && [ "$FM_BACKLOG_ROW_HOLD_KIND" = captain ]; then + if [ "$identity" -eq 1 ]; then + show=$(task_show "$id") || { + printf 'fm-captain-hold: captain call %s is open but its record could not be read\n' "$id" >&2 + exit 2 + } + shown_body=$(show_field "$show" body) + printf '%s#%s\n' \ + "$(body_hold_set_timestamp "$(decode_shown_value "$shown_body")")" \ + "$(resolution_record_count "$shown_body")" + fi return 0 fi return 1 fi - [ "$FM_BACKLOG_ROW_RESULT" != not_found ] || return 1 + if [ "$FM_BACKLOG_ROW_RESULT" = not_found ]; then + [ "$distinguish_absent" = 0 ] || return 3 + return 1 + fi printf 'fm-captain-hold: %s\n' "$FM_BACKLOG_ROW_ERROR" >&2 exit 2 } @@ -1458,6 +1816,7 @@ case "${1:-}" in hold) shift; command_hold "$@" ;; answer) shift; command_answer "$@" ;; answers) shift; command_answers "$@" ;; + reconcile-requests) shift; command_reconcile_requests "$@" ;; bind) shift; command_bind "$@" ;; unbind) shift; command_unbind "$@" ;; binding) shift; command_binding "$@" ;; @@ -1465,6 +1824,7 @@ case "${1:-}" in verify) shift; command_verify "$@" ;; open) shift; command_open "$@" ;; diverged) shift; command_diverged "$@" ;; + reconcile) shift; command_reconcile "$@" ;; -h|--help) usage ;; *) usage >&2; exit 2 ;; esac diff --git a/bin/fm-classify-lib.sh b/bin/fm-classify-lib.sh index 76ba18b4a3f..5ac81f00398 100755 --- a/bin/fm-classify-lib.sh +++ b/bin/fm-classify-lib.sh @@ -578,14 +578,35 @@ _fm_decision_fold_line() { # <open-set> <status-line> <resolve-verb> <held-verb # before any read - a cheap builtin, unlike fm_wake_latest_event's O_NOFOLLOW # subprocess read, which exists for that function's much narrower payload-driven # path resolution rather than this directory-local glob. -status_open_decisions() { # <status-file> - local f=$1 line resolve held open='' +status_open_decisions() { # <status-file> [--explicit-only] + local f=$1 mode=${2:-all} line resolve held open='' default_explicit=0 key verb + case "$mode" in all|--explicit-only) ;; *) return 2 ;; esac [ -f "$f" ] && [ -r "$f" ] && [ ! -L "$f" ] || return 0 resolve=${FM_CLASSIFY_RESOLVE_VERB:-$FM_CLASSIFY_RESOLVE_VERB_DEFAULT} held=${FM_CLASSIFY_CAPTAIN_HELD_VERB:-$FM_CLASSIFY_CAPTAIN_HELD_VERB_DEFAULT} while IFS= read -r line || [ -n "$line" ]; do open=$(_fm_decision_fold_line "$open" "$line" "$resolve" "$held") + # Non-default keys are always explicit. The default bucket also accepts a + # literal [key=default], so retain the last opening's token provenance + # without changing the shared fold's resolution or durable retention rules. + if [ "$mode" = --explicit-only ]; then + key=$(_fm_decision_key "$line") || continue + [ "$key" = default ] || continue + _fm_decision_key_transition_allowed "$key" "$(status_line_note "$line")" || continue + verb=$(status_line_verb "$line") + case "$verb" in + needs-decision|blocked) + default_explicit=0 + if _fm_key_before_colon "$line" || _fm_key_at_note_head "$line" >/dev/null; then + default_explicit=1 + fi + ;; + esac + fi done < "$f" + if [ "$mode" = --explicit-only ] && [ "$default_explicit" -eq 0 ]; then + open=$(_fm_decision_drop "$open" default) + fi printf '%s' "$open" } @@ -762,9 +783,9 @@ _fm_open_decisions_file_ident() { # <file> -> strongest available identity return fi if [ "$(uname -s 2>/dev/null)" = Darwin ]; then - ident=$(LC_ALL=C stat -f '%d:%i' "$f" 2>/dev/null) || return 1 - epoch=$(LC_ALL=C stat -f '%B' "$f" 2>/dev/null) || epoch=0 - if [ "$epoch" != 0 ]; then birth=$(LC_ALL=C stat -f '%FB' "$f" 2>/dev/null) || birth=''; else birth=''; fi + ident=$(LC_ALL=C /usr/bin/stat -f '%d:%i' "$f" 2>/dev/null) || return 1 + epoch=$(LC_ALL=C /usr/bin/stat -f '%B' "$f" 2>/dev/null) || epoch=0 + if [ "$epoch" != 0 ]; then birth=$(LC_ALL=C /usr/bin/stat -f '%FB' "$f" 2>/dev/null) || birth=''; else birth=''; fi else ident=$(LC_ALL=C stat -c '%d:%i' "$f" 2>/dev/null) || return 1 epoch=$(LC_ALL=C stat -c '%W' "$f" 2>/dev/null) || epoch=0 @@ -781,7 +802,7 @@ _fm_status_file_size() { # <status-file> return fi if [ "$(uname -s 2>/dev/null)" = Darwin ]; then - LC_ALL=C stat -f '%z' "$f" 2>/dev/null + LC_ALL=C /usr/bin/stat -f '%z' "$f" 2>/dev/null else LC_ALL=C stat -c '%s' "$f" 2>/dev/null fi @@ -790,7 +811,7 @@ _fm_status_file_size() { # <status-file> _fm_status_file_mtime() { # <status-file> local f=$1 if [ "$(uname -s 2>/dev/null)" = Darwin ]; then - LC_ALL=C stat -f '%m' "$f" 2>/dev/null + LC_ALL=C /usr/bin/stat -f '%m' "$f" 2>/dev/null else LC_ALL=C stat -c '%Y' "$f" 2>/dev/null fi @@ -1181,7 +1202,7 @@ status_presentation_marker_parse() { _status_observed_path_state() { if [ "$(uname -s 2>/dev/null)" = Darwin ]; then - LC_ALL=C stat -f '%HT:%p' "$1" 2>/dev/null + LC_ALL=C /usr/bin/stat -f '%HT:%p' "$1" 2>/dev/null else LC_ALL=C stat -c '%F:%f' "$1" 2>/dev/null fi diff --git a/bin/fm-claude-stop-autoarm.sh b/bin/fm-claude-stop-autoarm.sh index 1f6353ab9a6..5f8b2d14a75 100755 --- a/bin/fm-claude-stop-autoarm.sh +++ b/bin/fm-claude-stop-autoarm.sh @@ -18,10 +18,8 @@ # - AFK: while state/.afk exists the away daemon owns the watcher and triage; # this hook exits 0 and NEVER rewakes the primary (checked again at # translation time so a mid-cycle AFK transition is honored). -# - Need: arms only while work is in flight (state/*.meta), a process-event -# source is registered (state/procevent/*.source), X mode has a relay poll -# to run (state/x-watch.check.sh), or wake records are still queued for a -# drain (state/.wake-queue); a home with none of those exits 0. +# - Need: arms only while the home needs supervision, as +# bin/fm-supervision-lib.sh defines it; an idle home exits 0. # - Single-flight: Claude does not dedupe async hooks, so exactly one # GENERATION owner arms per event epoch: the epoch ledger's monotonic # sequence is the claim generation, every firing defers (exit 0) to a live @@ -75,7 +73,6 @@ FM_ROOT="${FM_ROOT_OVERRIDE:-$(cd "$SCRIPT_DIR/.." && pwd)}" FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" CONFIG="${FM_CONFIG_OVERRIDE:-$FM_HOME/config}" -GRACE=${FM_GUARD_GRACE:-300} OWNER_LOCK="$STATE/.claude-autoarm.lock" FAILURE_NOTICE="$STATE/.claude-autoarm-failure-notified" FAILURE_ALARM="$STATE/.claude-autoarm-failure-alarmed" @@ -96,6 +93,13 @@ esac # shellcheck source=bin/fm-hook-host-lib.sh . "$SCRIPT_DIR/fm-hook-host-lib.sh" +# fm-watch.sh touches the liveness beacon once per cycle, immediately before +# its terminal wait, so a healthy watcher's beacon can legitimately age up to +# FM_POLL seconds between touches (docs/turnend-guard.md "Guard grace and the +# poll cadence"). fm_poll_derived_grace (bin/fm-wake-lib.sh) is the single +# owner of that max(300, poll+60) derivation. +GRACE=${FM_GUARD_GRACE:-$(fm_poll_derived_grace)} + # Consume the Stop payload once. The decisions below are state-based; the # payload is read so a slow writer can never wedge on a full pipe, and its host # is inspected before anything else runs. @@ -131,7 +135,7 @@ fi # --- AFK: the away daemon owns the watcher and triage; never rewake ---------- [ -e "$STATE/.afk" ] && exit 0 -# --- need: in-flight work, an event source, an X-mode relay poll, or a queue -- +# --- need: whatever bin/fm-supervision-lib.sh counts as supervision need ------ need_supervision() { fm_supervision_needed "$STATE" "$GRACE" } @@ -219,9 +223,9 @@ while [ "$attempt" -lt "$AUTOARM_ATTEMPTS" ]; do attempt=$((attempt + 1)) OUT=$(mktemp "$STATE/.claude-autoarm-output.XXXXXX") || OUT= if [ -n "$OUT" ]; then - "$SCRIPT_DIR/fm-watch-arm.sh" >"$OUT" 2>&1 || true + FM_GUARD_GRACE="$GRACE" "$SCRIPT_DIR/fm-watch-arm.sh" >"$OUT" 2>&1 || true else - "$SCRIPT_DIR/fm-watch-arm.sh" >/dev/null 2>&1 || true + FM_GUARD_GRACE="$GRACE" "$SCRIPT_DIR/fm-watch-arm.sh" >/dev/null 2>&1 || true fi # AFK may have appeared mid-cycle: the daemon owns triage now, so suppress diff --git a/bin/fm-config-inherit-lib.sh b/bin/fm-config-inherit-lib.sh index 61d4ad2242c..de4e53630f5 100644 --- a/bin/fm-config-inherit-lib.sh +++ b/bin/fm-config-inherit-lib.sh @@ -116,7 +116,7 @@ fm_config_source_present() { fm_inherit_file_mode() { if [ "$(uname)" = Darwin ]; then - stat -f %Lp "$1" 2>/dev/null + /usr/bin/stat -f %Lp "$1" 2>/dev/null else stat -c %a "$1" 2>/dev/null fi @@ -124,7 +124,7 @@ fm_inherit_file_mode() { fm_inherit_file_device() { if [ "$(uname)" = Darwin ]; then - stat -f %d "$1" 2>/dev/null + /usr/bin/stat -f %d "$1" 2>/dev/null else stat -c %d "$1" 2>/dev/null fi @@ -132,7 +132,7 @@ fm_inherit_file_device() { fm_inherit_file_link_count() { if [ "$(uname)" = Darwin ]; then - stat -f %l "$1" 2>/dev/null + /usr/bin/stat -f %l "$1" 2>/dev/null else stat -c %h "$1" 2>/dev/null fi diff --git a/bin/fm-fleet-snapshot.sh b/bin/fm-fleet-snapshot.sh index 642d774c252..7771338034c 100755 --- a/bin/fm-fleet-snapshot.sh +++ b/bin/fm-fleet-snapshot.sh @@ -1632,8 +1632,8 @@ case "$FM_SNAPSHOT_SECONDMATE_LANDED_PER_HOME" in ''|*[!0-9]*) FM_SNAPSHOT_SECON # pollute arithmetic input before failing. Select the platform syntax once. if [ "$(uname 2>/dev/null || true)" = Darwin ]; then SNAPSHOT_STAT_STYLE=bsd - file_mtime_epoch() { stat -f '%m' "$1" 2>/dev/null || true; } - file_mode_octal() { stat -f '%Lp' "$1" 2>/dev/null || true; } + file_mtime_epoch() { /usr/bin/stat -f '%m' "$1" 2>/dev/null || true; } + file_mode_octal() { /usr/bin/stat -f '%Lp' "$1" 2>/dev/null || true; } else SNAPSHOT_STAT_STYLE=gnu file_mtime_epoch() { stat -c '%Y' "$1" 2>/dev/null || true; } @@ -1983,7 +1983,7 @@ bounded_parent_activities_json() { # <status-file> stat_style=$6 . "$classify" if [ "$stat_style" = bsd ]; then - size=$(stat -f "%z" "$f" 2>/dev/null) || exit 3 + size=$(/usr/bin/stat -f "%z" "$f" 2>/dev/null) || exit 3 else size=$(stat -c "%s" "$f" 2>/dev/null) || exit 3 fi diff --git a/bin/fm-guard.sh b/bin/fm-guard.sh index a6433f40d34..7e897cd7836 100755 --- a/bin/fm-guard.sh +++ b/bin/fm-guard.sh @@ -5,11 +5,10 @@ # First, always warn if the firstmate primary checkout (FM_ROOT) is on a named # non-default branch, because that means firstmate-on-itself work landed in the # primary instead of an isolated worktree. -# Then, if a task is in flight (a state/<id>.meta exists), X-mode relay polling -# is active (state/x-watch.check.sh exists), or queued wakes are still pending -# in state/.wake-queue, and supervision is not healthy, prints a loud, clearly -# delimited banner so the agent cannot skim past it in the tool output of -# whatever it was doing - the one channel every harness +# Then, if the home needs supervision (bin/fm-supervision-lib.sh owns that +# condition set) and that supervision is not healthy, prints a loud, clearly +# delimited banner so the agent cannot skim past +# it in the tool output of whatever it was doing - the one channel every harness # has. Supervision health is MODEL-AWARE (fm_watcher_supervision_verdict in # bin/fm-wake-lib.sh): under the Claude Stop auto-arm model the watcher runs only # between turns, so mid-turn a fresh beacon with no live watcher is healthy and @@ -28,7 +27,13 @@ # bounded). Independent alarms (queued wakes, worktree tangle) are never # suppressed by that dedup. Normal wake handling (watcher briefly down between a # wake and the next supervision resume) stays inside the grace window and stays -# silent. The queued-wakes warning stays silent for the supervision branch +# silent. The queued-wakes warning counts only the rows the calling actor can +# itself present or retire (fm_wake_actor_pending_count), so it is never an +# instruction to run a drain with nothing to present. A row reserved by a live +# supervision-branch grant is never a drain instruction for main; instead of +# going silent about a visibly non-empty queue, main gets a distinct advisory +# naming the branch as the holder and saying not to drain those rows. +# The ordinary warning also stays silent for the supervision branch # actor (FM_SUPERVISION_ACTOR=branch), because that actor runs guarded commands # while handling exactly the queued rows its grant covers and can drain nothing # else. Always exits 0: the guard warns, it never blocks. @@ -44,6 +49,7 @@ CONFIG="${FM_CONFIG_OVERRIDE:-$FM_HOME/config}" WATCH="$SCRIPT_DIR/fm-watch.sh" GRACE=${FM_GUARD_GRACE:-300} queue_pending=false +queue_branch_held=false READ_ONLY=${FM_GUARD_READ_ONLY:-0} case "$READ_ONLY" in 1|true|TRUE|yes|YES) READ_ONLY=1 ;; *) READ_ONLY=0 ;; esac CONTINUE_LINE=${FM_GUARD_CONTINUE_LINE:-This is a supervision warning only; the guarded operation WILL still run.} @@ -162,14 +168,16 @@ if [ -n "$tangle_branch" ]; then fi # Compute supervision need and watcher-beacon freshness via the shared -# grace-based predicate (bin/fm-supervision-lib.sh). Act when work, an event -# source, an X-mode relay poll, or a pending wake queue needs supervision. +# grace-based predicate (bin/fm-supervision-lib.sh), which owns what needs +# supervision. fm_supervision_status "$STATE" "$GRACE" in_flight=$FM_SUP_IN_FLIGHT sources=$FM_SUP_SOURCES +checks=$FM_SUP_CHECKS needed=$FM_SUP_NEEDED beacon_desc=$FM_SUP_BEACON_DESC -queue_pending=$FM_SUP_QUEUE_PENDING +queue_pending=false +queue_branch_held=false fm_watcher_supervision_verdict "$STATE" "$WATCH" "$GRACE" "$FM_HOME" "$FM_ROOT" watcher_healthy=$FM_WATCHER_VERDICT_OK watcher_down_reason=$FM_WATCHER_VERDICT_REASON @@ -181,6 +189,19 @@ if [ "$needed" = false ]; then exit 0 fi +# Count only the rows this actor could actually present or retire, so the +# warning never sends an actor to a drain that provably has nothing for it. +# fm-wake-lib.sh owns that per-actor classification. A non-empty queue with +# nothing for main is the branch-held case: keep the raw pending signal visible +# there as its own advisory rather than dropping it. +if [ -s "$FM_WAKE_QUEUE" ]; then + if [ "$(fm_wake_actor_pending_count "$GUARD_ACTOR")" -gt 0 ]; then + queue_pending=true + elif [ "$GUARD_ACTOR" != branch ] && [ "$(fm_wake_actor_pending_count branch)" -gt 0 ]; then + queue_branch_held=true + fi +fi + # No fresh watcher with tasks in flight is the dangerous state: emit a prominent, # bordered banner FIRST so it reads as an alarm, not a buried stderr line. Later # calls in the same episode get a concise reminder plus CONTINUE_LINE only. @@ -219,7 +240,9 @@ if [ "$watcher_healthy" = false ]; then printf '● %s task(s) in flight, but %s.\n' "$in_flight" "$watcher_cause" elif [ "$sources" -gt 0 ]; then printf '● %s process-event source(s) registered, but %s.\n' "$sources" "$watcher_cause" - elif "$queue_pending"; then + elif [ "$checks" -gt 0 ]; then + printf '● %s registered custom check(s), but %s.\n' "$checks" "$watcher_cause" + elif [ "$FM_SUP_QUEUE_PENDING" = true ]; then printf '● Queued wakes are pending and need supervision, but %s.\n' "$watcher_cause" else printf '● X-mode relay polling needs supervision, but %s.\n' "$watcher_cause" @@ -260,5 +283,7 @@ if "$queue_pending"; then echo "WARNING: queued wakes pending - drain them with bin/fm-wake-drain.sh before anything else." >&2 fi printf '%s\n' "$CONTINUE_LINE" >&2 +elif "$queue_branch_held"; then + echo "NOTICE: wake rows held by the live supervision branch - it presents and acknowledges them; do not drain them from here." >&2 fi exit 0 diff --git a/bin/fm-inactive-reconcile.sh b/bin/fm-inactive-reconcile.sh index 52064545836..9c30a9074be 100755 --- a/bin/fm-inactive-reconcile.sh +++ b/bin/fm-inactive-reconcile.sh @@ -121,7 +121,7 @@ if [ "$FM_INACTIVE_RECONCILE_BUDGET_SECS" -gt 30 ]; then fi if [ "$(uname)" = Darwin ]; then - file_mtime() { stat -f %m "$1" 2>/dev/null; } + file_mtime() { /usr/bin/stat -f %m "$1" 2>/dev/null; } else file_mtime() { stat -c %Y "$1" 2>/dev/null; } fi diff --git a/bin/fm-launch-lib.sh b/bin/fm-launch-lib.sh index 62d263457f8..596752e6676 100644 --- a/bin/fm-launch-lib.sh +++ b/bin/fm-launch-lib.sh @@ -273,7 +273,10 @@ fm_launch_template() { # does NOT suppress the interactive ghost text (verified empirically), so the env # var is the correct control. The dim-aware composer reader in fm-tmux-lib.sh is # the defense-in-depth backstop for any pane this flag cannot reach. - claude) printf '%s' 'CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --dangerously-skip-permissions --settings '\''{"feedbackDrafts":"off"}'\'' __MODELFLAG____EFFORTFLAG__"$(__OPINPUT__ encode launch-brief < __BRIEF__)"' ;; + # Carry attribution-off with the per-launch settings because worker settings + # sources may omit the user's scope. Keep both feedback controls alongside + # it so managed settings cannot re-enable the model-drafted feedback tool. + claude) printf '%s' 'CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --dangerously-skip-permissions --settings '\''{"feedbackDrafts":"off","attribution":{"commit":"","pr":"","sessionUrl":false}}'\'' __MODELFLAG____EFFORTFLAG__"$(__OPINPUT__ encode launch-brief < __BRIEF__)"' ;; codex) if [ "$kind" = secondmate ]; then printf '%s' 'codex __MODELFLAG____EFFORTFLAG__--dangerously-bypass-approvals-and-sandbox "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' diff --git a/bin/fm-lock-lib.sh b/bin/fm-lock-lib.sh index f3b070ec8cf..7303ac571ad 100644 --- a/bin/fm-lock-lib.sh +++ b/bin/fm-lock-lib.sh @@ -24,7 +24,7 @@ fm_lock_log() { # no wake-queue machinery when a caller only needs the staleness proof. fm_lock_path_mtime() { if [ "$(uname)" = Darwin ]; then - stat -f %m "$1" 2>/dev/null + /usr/bin/stat -f %m "$1" 2>/dev/null else stat -c %Y "$1" 2>/dev/null fi diff --git a/bin/fm-pending-reply-lib.sh b/bin/fm-pending-reply-lib.sh index 7e4b00e4948..27c9bc9e90d 100755 --- a/bin/fm-pending-reply-lib.sh +++ b/bin/fm-pending-reply-lib.sh @@ -571,7 +571,7 @@ fm_pending_reply_file_signature() { # <path> local path=$1 [ -f "$path" ] || { printf 'missing'; return 0; } if [ "$(uname -s 2>/dev/null)" = Darwin ]; then - LC_ALL=C stat -f '%d:%i:%z:%m:%c' "$path" 2>/dev/null || printf 'unreadable' + LC_ALL=C /usr/bin/stat -f '%d:%i:%z:%m:%c' "$path" 2>/dev/null || printf 'unreadable' else LC_ALL=C stat -c '%d:%i:%s:%Y:%Z' "$path" 2>/dev/null || printf 'unreadable' fi diff --git a/bin/fm-pr-lib.sh b/bin/fm-pr-lib.sh index 05fdad09588..f4b174bf86c 100755 --- a/bin/fm-pr-lib.sh +++ b/bin/fm-pr-lib.sh @@ -245,7 +245,7 @@ fm_pr_head_valid() { fm_pr_file_mode() { if [ "$(uname)" = Darwin ]; then - stat -f %Lp "$1" 2>/dev/null + /usr/bin/stat -f %Lp "$1" 2>/dev/null else stat -c %a "$1" 2>/dev/null fi @@ -253,7 +253,7 @@ fm_pr_file_mode() { fm_pr_file_device() { if [ "$(uname)" = Darwin ]; then - stat -f %d "$1" 2>/dev/null + /usr/bin/stat -f %d "$1" 2>/dev/null else stat -c %d "$1" 2>/dev/null fi @@ -261,7 +261,7 @@ fm_pr_file_device() { fm_pr_file_link_count() { if [ "$(uname)" = Darwin ]; then - stat -f %l "$1" 2>/dev/null + /usr/bin/stat -f %l "$1" 2>/dev/null else stat -c %h "$1" 2>/dev/null fi @@ -269,7 +269,7 @@ fm_pr_file_link_count() { fm_pr_file_inode() { if [ "$(uname)" = Darwin ]; then - stat -f %i "$1" 2>/dev/null + /usr/bin/stat -f %i "$1" 2>/dev/null else stat -c %i "$1" 2>/dev/null fi diff --git a/bin/fm-procevent-extension-capture.pl b/bin/fm-procevent-extension-capture.pl index 3e877ae6c43..485c1a0343d 100644 --- a/bin/fm-procevent-extension-capture.pl +++ b/bin/fm-procevent-extension-capture.pl @@ -1,7 +1,7 @@ use strict; use warnings; use Cwd qw(getcwd); -use Fcntl qw(O_CREAT O_EXCL O_NOFOLLOW O_RDONLY O_RDWR); +use Fcntl qw(O_CREAT O_EXCL O_NOFOLLOW O_RDONLY O_RDWR O_WRONLY); use JSON::PP qw(encode_json); use POSIX qw(dup2); @@ -92,8 +92,12 @@ my ($registry_fd, $inbox_fd, $reservation_fd, $id, $adapter, $extension_id, $extension_version, $capability_version, $package_digest, $binding_digest, $claim_token, $runner_name, $output_name, $runner_pid, $claim_identity, $limit, @command) = @ARGV; +my $launch_ready_name; +$launch_ready_name = shift @command if @command && $command[0] ne "--"; die "missing command\n" unless @command && shift(@command) eq "--"; die "invalid limit\n" unless defined $limit && $limit =~ /\A\d+\z/; +die "invalid launch boundary\n" if defined($launch_ready_name) + && $launch_ready_name !~ /\A\.[A-Za-z0-9._-]{1,384}\.launch-ready\z/; our ($registry_dir, $registry, $reservation_dir, $reservation_root, $sequence); sub fail { die "capture failed: $_[0]\n"; } @@ -180,6 +184,14 @@ sub write_reservation { write_all($runner, "$runner_pid\n"); close($runner) or fail("cannot close runner record"); my $stage = open_new($output_name); +my $launch_ready; +if (defined $launch_ready_name) { + sysopen($launch_ready, $launch_ready_name, O_WRONLY | O_NOFOLLOW) + or fail("cannot open launch boundary"); + my @launch_ready_stat = stat($launch_ready); + fail("unsafe launch boundary") unless @launch_ready_stat && -f _ && $launch_ready_stat[4] == $< + && ($launch_ready_stat[2] & 07777) == 0600 && $launch_ready_stat[3] == 1; +} pipe(my $reader, my $writer) or fail("cannot create output pipe"); my $child = fork(); defined $child or fail("cannot fork adapter"); @@ -191,6 +203,10 @@ sub write_reservation { exit 127; } close($writer); +if (defined $launch_ready) { + write_all($launch_ready, "ready\n"); + close($launch_ready) or fail("cannot close launch boundary"); +} my ($written, $truncated) = (0, 0); while (1) { my $read = sysread($reader, my $buffer, 65536); diff --git a/bin/fm-procevent-lavish.sh b/bin/fm-procevent-lavish.sh index 09c1e5a6af9..fae8f85aa0c 100755 --- a/bin/fm-procevent-lavish.sh +++ b/bin/fm-procevent-lavish.sh @@ -7,6 +7,7 @@ # fm-procevent-lavish.sh terminal <result-file> # fm-procevent-lavish.sh silent <result-file> # fm-procevent-lavish.sh answers <result-file> +# fm-procevent-lavish.sh reconciles <result-file> # fm-procevent-lavish.sh read <result-file> # fm-procevent-lavish.sh source-id <artifact.html> # fm-procevent-lavish.sh retire <artifact.html | source-id> @@ -107,7 +108,7 @@ # That is an internal retry, not news, so registering the raw poll made the # generic runner capture it and wake the whole fleet. `poll` therefore re-runs # the published poll up to POLL_RETRY_LIMIT times for that exact response, with -# POLL_RETRY_DELAY_DEFAULT seconds between attempts. The match is exact and +# attempt starts at least POLL_RETRY_DELAY_DEFAULT seconds apart. The match is exact and # deliberately narrow: real feedback, ended and missing sessions, any other # SERVER_ERROR, and the same interruption still standing after the bound is # spent are all printed straight through and captured normally. The retry is a @@ -294,6 +295,7 @@ this path matches nothing here. Retire by source id instead - find it with # without waiting it out. POLL_RETRY_LIMIT=12 POLL_RETRY_DELAY_DEFAULT=5 +POLL_RETRY_DELAY_MIN=1 POLL_RETRY_DELAY_MAX=60 # Exit 0 only for the exact two-line interruption, and nothing else. The whole @@ -346,8 +348,8 @@ poll_response_filter() { # <response-file> ' "$1" } -# Seconds between retries. FM_LAVISH_POLL_RETRY_DELAY is a bounded test -# override; a malformed or out-of-range value is refused rather than quietly +# Minimum seconds between retry attempt starts. FM_LAVISH_POLL_RETRY_DELAY is a +# bounded test override; a malformed or out-of-range value is refused rather than quietly # rounded, because silently changing a retry cadence is how a bound stops # meaning anything. poll_retry_delay() { @@ -357,15 +359,28 @@ poll_retry_delay() { return 0 fi case "$delay" in - *[!0-9]*) die "FM_LAVISH_POLL_RETRY_DELAY must be whole seconds from 0 to $POLL_RETRY_DELAY_MAX: $delay" ;; + *[!0-9]*) die "FM_LAVISH_POLL_RETRY_DELAY must be whole seconds from $POLL_RETRY_DELAY_MIN to $POLL_RETRY_DELAY_MAX: $delay" ;; esac - [ "$delay" -le "$POLL_RETRY_DELAY_MAX" ] \ - || die "FM_LAVISH_POLL_RETRY_DELAY must be whole seconds from 0 to $POLL_RETRY_DELAY_MAX: $delay" + [ "$delay" -ge "$POLL_RETRY_DELAY_MIN" ] && [ "$delay" -le "$POLL_RETRY_DELAY_MAX" ] \ + || die "FM_LAVISH_POLL_RETRY_DELAY must be whole seconds from $POLL_RETRY_DELAY_MIN to $POLL_RETRY_DELAY_MAX: $delay" printf '%s\n' "$delay" } +poll_iteration_started() { + perl -MTime::HiRes=clock_gettime,CLOCK_MONOTONIC -e \ + 'printf "%.6f\\n", clock_gettime(CLOCK_MONOTONIC)' +} + +poll_iteration_floor_wait() { + perl -MTime::HiRes=clock_gettime,sleep,CLOCK_MONOTONIC -e ' + my ($started, $floor) = @ARGV; + my $remaining = $floor - (clock_gettime(CLOCK_MONOTONIC) - $started); + sleep($remaining) if $remaining > 0; + ' "$1" "$2" +} + cmd_poll() { - local artifact=${1-} delay attempt=0 response cleanup_command rc filter_rc + local artifact=${1-} delay attempt=0 response cleanup_command rc filter_rc iteration_started local pipeline_status [ -n "$artifact" ] || usage [ "$#" -eq 1 ] || usage @@ -389,6 +404,7 @@ cmd_poll() { trap "$cleanup_command; trap - $signal; kill -$signal $$" "$signal" done while :; do + iteration_started=$(poll_iteration_started) || die "cannot start the poll rate governor" lavish-axi poll "$artifact" | poll_response_filter "$response" pipeline_status=("${PIPESTATUS[@]}") rc=${pipeline_status[0]} @@ -398,7 +414,8 @@ cmd_poll() { 10) if [ "$attempt" -lt "$POLL_RETRY_LIMIT" ]; then attempt=$((attempt + 1)) - sleep "$delay" + poll_iteration_floor_wait "$iteration_started" "$delay" \ + || die "cannot enforce the poll rate governor" else cat -- "$response" break @@ -530,7 +547,7 @@ cmd_silent() { [ "$content_rc" -eq 1 ] } -# Print `key<TAB>answer<TAB>label[<TAB>mode]` for every structured choice the +# Print `key<TAB>answer<TAB>label[<TAB>mode]` for each non-reconcile structured choice the # captain submitted in a captured result; the optional mode column relays the # card's declared close mode (`done` or `release`) to the keyed-answer intake. The published response frames queued feedback as # a `prompts[N]{field,...}:` header followed by exactly N indented CSV rows whose @@ -538,18 +555,20 @@ cmd_silent() { # rather than assuming a fixed column, and takes only rows whose `tag` field is # `choice`. A freeform `message` row is captain prose and is deliberately never a # source of decision keys. A row that does not carry both a slug-shaped `question` -# and an `answer` inside its `Context data:` block is skipped, so a deck that does -# not key its forms by decision key simply yields nothing. +# and the versioned `selection` and `note` fields inside its `Context data:` block +# is skipped. A time-limited rollout branch accepts the old question/answer +# shape only for ordinary answers and rejects its bare or annotated reconcile +# values because old rows do not separate the selected option from its note. # The question cap is 128 so any task id fits, including the long legacy # `<origin>-decision-<key>` identities pre-collapse decks still carry; the # security property is the slug SHAPE, which is unchanged. -cmd_answers() { - local file=${1-} +cmd_choice_rows() { + local selection=$1 file=${2-} [ -n "$file" ] || usage [ -f "$file" ] && [ ! -L "$file" ] || die "result file does not exist: $file" perl -MJSON::PP -e ' use strict; use warnings; - my ($path) = @ARGV; + my ($selection, $path) = @ARGV; open my $fh, "<", $path or exit 1; my (@fields, $want, @rows); while (my $line = <$fh>) { @@ -565,7 +584,7 @@ cmd_answers() { } close $fh; my %seen; - my @out; + my @choices; for my $row (@rows) { $row =~ s/^\s+//; my @vals; @@ -588,29 +607,72 @@ cmd_answers() { my $ctx = $1; my $data = eval { decode_json($ctx) }; next unless ref($data) eq "HASH"; - my $key = $data->{question}; - my $answer = $data->{answer}; - next if !defined($key) || ref($key) || !defined($answer) || ref($answer); + my ($key, $selected, $note, $answer, $legacy); + if (defined($data->{schema}) && !ref($data->{schema}) + && $data->{schema} eq "fm-bearings-answer.v1") { + $key = $data->{question}; + $selected = $data->{selection}; + $note = $data->{note}; + next if !defined($key) || ref($key) || !defined($selected) || ref($selected) + || !defined($note) || ref($note); + next unless $selected eq "" || $selected =~ /\A[A-Za-z0-9._-]{1,128}\z/; + next unless length($note) <= 512; + next unless length($selected) || length($note); + $answer = length($selected) ? $selected : $note; + $legacy = 0; + # Time-limited compatibility for captures from pre-change boards; remove + # once no board carrying the old question/answer context can remain armed. + } elsif (!exists($data->{schema}) && !exists($data->{selection}) + && !exists($data->{note})) { + $key = $data->{question}; + $answer = $data->{answer}; + next if !defined($key) || ref($key) || !defined($answer) || ref($answer); + next unless length($answer) && length($answer) <= 512; + next if $answer eq "reconcile" || index($answer, "reconcile - ") == 0; + $selected = ""; + $note = ""; + $legacy = 1; + } else { + next; + } + next unless $key =~ /\A[A-Za-z0-9._-]{1,128}\z/; my $mode = ""; if (exists $data->{close}) { next if !defined($data->{close}) || ref($data->{close}) || ($data->{close} ne "done" && $data->{close} ne "release"); $mode = $data->{close}; } - next unless $key =~ /\A[A-Za-z0-9._-]{1,128}\z/; - next unless length $answer && length($answer) <= 512; my $label = defined $f{text} ? $f{text} : ""; - s/[\x00-\x1f\x7f]/ /g for ($answer, $label); + s/[\x00-\x1f\x7f]/ /g for ($answer, $note, $label); $label = substr($label, 0, 512); - # A re-answered form appears again later in the queue; the last submission wins. - if (defined $seen{$key}) { $out[$seen{$key}] = undef } - $seen{$key} = scalar @out; - push @out, length $mode ? "$key\t$answer\t$label\t$mode" : "$key\t$answer\t$label"; + if (defined $seen{$key}) { $choices[$seen{$key}] = undef } + $seen{$key} = scalar @choices; + push @choices, { + key => $key, selection => $selected, note => $note, legacy => $legacy, + answer => $answer, label => $label, mode => $mode + }; } - print "$_\n" for grep { defined } @out; - ' "$file" + for my $choice (grep { defined } @choices) { + if ($selection eq "reconciles") { + next if $choice->{legacy}; + if ($choice->{selection} eq "reconcile") { + print length($choice->{note}) + ? "$choice->{key}\t$choice->{note}\n" + : "$choice->{key}\n"; + } + next; + } + next if $choice->{selection} eq "reconcile"; + print length $choice->{mode} + ? "$choice->{key}\t$choice->{answer}\t$choice->{label}\t$choice->{mode}\n" + : "$choice->{key}\t$choice->{answer}\t$choice->{label}\n"; + } + ' "$selection" "$file" } +cmd_answers() { cmd_choice_rows answers "$@"; } +cmd_reconciles() { cmd_choice_rows reconciles "$@"; } + # Present one already-captured result for a handler. Body lines are prefixed # so a captain-supplied string cannot forge a section label. The session-ending # message is printed before the count line and before any annotation, because @@ -759,6 +821,7 @@ case "${1-}" in terminal) shift; cmd_terminal "$@" ;; silent) shift; cmd_silent "$@" ;; answers) shift; cmd_answers "$@" ;; + reconciles) shift; cmd_reconciles "$@" ;; read) shift; cmd_read "$@" ;; ''|-h|--help|help) usage ;; *) die "unknown command: $1" ;; diff --git a/bin/fm-procevent-lib.sh b/bin/fm-procevent-lib.sh index b209dea6c85..ff8ef62a920 100644 --- a/bin/fm-procevent-lib.sh +++ b/bin/fm-procevent-lib.sh @@ -103,6 +103,202 @@ fm_procevent_any_registered() { return 1 } +# --- owning-session lease --------------------------------------------------- +# A runner is detached into its own process group so it survives the turn that +# started it. That is what makes a persistent source work, and on its own it is +# also what lets a runner outlive its whole home: once reparented to init, +# nothing bounds its lifetime, so its blocking child - and everything that child +# spawns - can keep running indefinitely. +# +# The bound is a lease on the OWNING STATE ROOT. Owner-presence operations +# refresh it, an attached public start keeps it fresh while its caller remains +# attached, and the watcher's reconcile cycle keeps it fresh in a live home. +# A guard proves the runner's owner is still there by reading that lease from +# the physical state root recorded in the claim. After two consecutive checks +# cannot prove both the root identity and a fresh lease, it stops the runner's +# process group. The lease is keyed by state root, so another home's live runner +# is untouched: that home refreshes its own lease. Nothing here keys on a script +# name, a command line, or a process name, all of which are shared across homes. + +fm_procevent_owner_lease_path() { # <state-root> + printf '%s/.owner-lease\n' "$(fm_procevent_registry_dir "$1")" +} + +# Record owner-presence activity in this home's process-event state. Best +# effort by design: a home with no registry directory yet owns no runner. +fm_procevent_owner_lease_touch() { # <state-root> + local reg lease tmp now + reg=$(fm_procevent_registry_dir "$1") + [ -d "$reg" ] && [ ! -L "$reg" ] || return 1 + lease=$(fm_procevent_owner_lease_path "$1") + now=$(perl -MTime::HiRes=clock_gettime,CLOCK_MONOTONIC -e \ + 'printf "%.6f\n", clock_gettime(CLOCK_MONOTONIC)') || return 1 + tmp=$(umask 077; mktemp "$reg/.owner-lease.XXXXXX") || return 1 + if ! printf '%s\n' "$now" > "$tmp" || ! mv -f -- "$tmp" "$lease"; then + rm -f -- "$tmp" + return 1 + fi +} + +# Seconds since the last refresh. Fails when the lease is absent or unreadable, +# which is what a removed home looks like from inside a surviving runner. +fm_procevent_owner_lease_age() { # <state-root> + local lease value + lease=$(fm_procevent_owner_lease_path "$1") + [ -f "$lease" ] && [ ! -L "$lease" ] || return 1 + IFS= read -r value < "$lease" || return 1 + perl -MTime::HiRes=clock_gettime,CLOCK_MONOTONIC -e ' + use strict; + use warnings; + my $value = shift; + $value =~ /\A[0-9]+(?:\.[0-9]+)?\z/ or exit 1; + my $now = clock_gettime(CLOCK_MONOTONIC); + $now >= $value or exit 1; + printf "%d\n", int($now - $value); + ' "$value" +} + +# How long a runner keeps going with no activity in its owning home. The default +# is forty watcher cycles at the default poll interval, so an ordinary busy or +# briefly wedged home never trips it, while a home that is simply gone stops +# owning processes within the hour rather than within a day. +FM_PROCEVENT_OWNER_LEASE_DEFAULT_SECONDS=600 +FM_PROCEVENT_OWNER_LEASE_MIN_SECONDS=1 +FM_PROCEVENT_OWNER_LEASE_MAX_SECONDS=86400 + +fm_procevent_owner_lease_seconds() { + local value=${FM_PROCEVENT_OWNER_LEASE_SECONDS-} + if [ -z "$value" ]; then + printf '%s\n' "$FM_PROCEVENT_OWNER_LEASE_DEFAULT_SECONDS" + return 0 + fi + case "$value" in ''|*[!0-9]*) return 1 ;; esac + [ "$value" -ge "$FM_PROCEVENT_OWNER_LEASE_MIN_SECONDS" ] || return 1 + [ "$value" -le "$FM_PROCEVENT_OWNER_LEASE_MAX_SECONDS" ] || return 1 + printf '%s\n' "$value" +} + +# How often a runner's guard re-reads that lease. One watcher cycle at the +# default poll interval, so the guard costs about as much as the cycle that +# refreshes what it reads. +FM_PROCEVENT_OWNER_CHECK_DEFAULT_SECONDS=15 +FM_PROCEVENT_OWNER_CHECK_MIN_SECONDS=1 +FM_PROCEVENT_OWNER_CHECK_MAX_SECONDS=3600 + +fm_procevent_owner_check_seconds() { + local value=${FM_PROCEVENT_OWNER_CHECK_SECONDS-} + if [ -z "$value" ]; then + printf '%s\n' "$FM_PROCEVENT_OWNER_CHECK_DEFAULT_SECONDS" + return 0 + fi + case "$value" in ''|*[!0-9]*) return 1 ;; esac + [ "$value" -ge "$FM_PROCEVENT_OWNER_CHECK_MIN_SECONDS" ] || return 1 + [ "$value" -le "$FM_PROCEVENT_OWNER_CHECK_MAX_SECONDS" ] || return 1 + printf '%s\n' "$value" +} + +FM_PROCEVENT_LAUNCH_FLOOR_DEFAULT_SECONDS=1 +FM_PROCEVENT_LAUNCH_FLOOR_MIN_SECONDS=1 +FM_PROCEVENT_LAUNCH_FLOOR_MAX_SECONDS=3600 + +fm_procevent_launch_floor_seconds() { + local value=${FM_PROCEVENT_LAUNCH_FLOOR_SECONDS-} + if [ -z "$value" ]; then + printf '%s\n' "$FM_PROCEVENT_LAUNCH_FLOOR_DEFAULT_SECONDS" + return 0 + fi + case "$value" in ''|*[!0-9]*) return 1 ;; esac + [ "$value" -ge "$FM_PROCEVENT_LAUNCH_FLOOR_MIN_SECONDS" ] || return 1 + [ "$value" -le "$FM_PROCEVENT_LAUNCH_FLOOR_MAX_SECONDS" ] || return 1 + printf '%s\n' "$value" +} + +fm_procevent_launch_floor_reset_locked() { # <state-root> <source-id> <registration-identity> + local reg identity + case "$3" in *:*) ;; *) return 1 ;; esac + case "$3" in ''|*[!0-9:]*) return 1 ;; esac + reg=$(fm_procevent_registry_dir "$1") || return 1 + identity=${3//:/-} + rm -f -- "$reg/$2.$identity.last-launch" +} + +fm_procevent_launch_floor_prune_locked() { # <state-root> <source-id> <registration-identity> + local reg identity keep stamp + case "$3" in *:*) ;; *) return 1 ;; esac + case "$3" in ''|*[!0-9:]*) return 1 ;; esac + reg=$(fm_procevent_registry_dir "$1") || return 1 + identity=${3//:/-} + keep="$reg/$2.$identity.last-launch" + for stamp in "$reg/$2".*.last-launch "$reg/$2.last-launch"; do + [ "$stamp" = "$keep" ] && continue + [ -e "$stamp" ] || [ -L "$stamp" ] || continue + rm -f -- "$stamp" || return 1 + done +} + +fm_procevent_launch_floor_wait() { # <state-root> <source-id> <registration-identity> <seconds> + local state=$1 id=$2 expected=$3 floor=$4 reg stamp identity registration current_identity status=0 + case "$expected" in *:*) ;; *) return 1 ;; esac + case "$expected" in ''|*[!0-9:]*) return 1 ;; esac + reg=$(fm_procevent_registry_dir "$state") || return 1 + identity=${expected//:/-} + stamp="$reg/$id.$identity.last-launch" + [ ! -L "$stamp" ] || return 1 + [ ! -e "$stamp" ] || [ -f "$stamp" ] || return 1 + perl -MTime::HiRes=clock_gettime,sleep,CLOCK_MONOTONIC -e ' + use strict; + use warnings; + my ($path, $floor) = @ARGV; + my $previous; + if (-e $path) { + open my $in, "<", $path or exit 1; + my $value = <$in>; + close $in or exit 1; + defined($value) && $value =~ /\A([0-9]+(?:\.[0-9]+)?)\n?\z/ or exit 1; + $previous = 0 + $1; + } + my $now = clock_gettime(CLOCK_MONOTONIC); + my $elapsed = defined($previous) && $now >= $previous ? $now - $previous : undef; + sleep($floor - $elapsed) if defined($elapsed) && $elapsed < $floor; + ' "$stamp" "$floor" || return 1 + + # Registration publication holds this same source lock while replacing and + # pruning pacing state, so a superseded sleeper cannot recreate its stamp. + fm_procevent_source_lock_acquire "$id" || return 1 + registration="$reg/$id.source" + current_identity=$(fm_pr_file_identity "$registration" 2>/dev/null) || current_identity= + if [ "$current_identity" != "$expected" ]; then + fm_procevent_source_lock_release "$id" || return 1 + return 2 + fi + [ ! -L "$stamp" ] && { [ ! -e "$stamp" ] || [ -f "$stamp" ]; } || status=1 + if [ "$status" -eq 0 ]; then + perl -MTime::HiRes=clock_gettime,CLOCK_MONOTONIC -MFcntl=:DEFAULT -e ' + use strict; + use warnings; + my $path = shift; + my $now = clock_gettime(CLOCK_MONOTONIC); + my $tmp = "$path.$$"; + sysopen(my $out, $tmp, O_WRONLY | O_CREAT | O_EXCL, 0600) or exit 1; + print {$out} "$now\n" or exit 1; + close $out or exit 1; + rename $tmp, $path or exit 1; + ' "$stamp" || status=1 + fi + if [ "$status" -ne 0 ]; then + fm_procevent_source_lock_release "$id" || : + return "$status" + fi + return 0 +} + +# True while the owning home is provably still active. +fm_procevent_owner_alive() { # <state-root> <lease-seconds> + local age + age=$(fm_procevent_owner_lease_age "$1") || return 1 + [ "$age" -le "$2" ] +} + # --- ownership -------------------------------------------------------------- # A claim is a private file recording the home, runner pid, claim generation, # and process identity. Registration and every ownership transition are @@ -130,7 +326,7 @@ fm_procevent_source_lock_release() { } fm_procevent_registration_publish_locked() { # <state> <adapter> <source-id> <argv...> - local state=$1 adapter=$2 id=$3 reg dest tmp arg + local state=$1 adapter=$2 id=$3 reg dest tmp arg identity shift 3 fm_procevent_adapter_valid "$adapter" || return 1 fm_procevent_source_id_valid "$id" || return 1 @@ -148,7 +344,11 @@ fm_procevent_registration_publish_locked() { # <state> <adapter> <source-id> <a printf 'argc=%s\n' "$#" printf 'argv:\n' printf '%s\n' "$@" - } > "$tmp" && chmod 0600 "$tmp" && mv -f -- "$tmp" "$dest"; then + } > "$tmp" && chmod 0600 "$tmp" \ + && identity=$(fm_pr_file_identity "$tmp") \ + && fm_procevent_launch_floor_reset_locked "$state" "$id" "$identity" \ + && mv -f -- "$tmp" "$dest"; then + fm_procevent_launch_floor_prune_locked "$state" "$id" "$identity" 2>/dev/null || : return 0 fi rm -f -- "$tmp" @@ -160,7 +360,7 @@ fm_procevent_registration_publish_locked() { # <state> <adapter> <source-id> <a # stored because the tracked host constructs that command at run time. fm_procevent_extension_registration_publish_locked() { # <state> <adapter> <source-id> <extension-id> <extension-version> <capability-version> <package-digest> <binding-digest> <config-ref> <registration-token> local state=$1 adapter=$2 id=$3 extension_id=$4 extension_version=$5 capability_version=$6 - local package_digest=$7 binding_digest=$8 config_ref=$9 registration_token=${10} reg dest tmp + local package_digest=$7 binding_digest=$8 config_ref=$9 registration_token=${10} reg dest tmp identity fm_procevent_adapter_valid "$adapter" || return 1 fm_procevent_source_id_valid "$id" || return 1 fm_procevent_extension_id_valid "$extension_id" || return 1 @@ -188,7 +388,11 @@ fm_procevent_extension_registration_publish_locked() { # <state> <adapter> <sou printf 'registration_token=%s\n' "$registration_token" printf 'argc=0\n' printf 'argv:\n' - } > "$tmp" && chmod 0600 "$tmp" && mv -f -- "$tmp" "$dest"; then + } > "$tmp" && chmod 0600 "$tmp" \ + && identity=$(fm_pr_file_identity "$tmp") \ + && fm_procevent_launch_floor_reset_locked "$state" "$id" "$identity" \ + && mv -f -- "$tmp" "$dest"; then + fm_procevent_launch_floor_prune_locked "$state" "$id" "$identity" 2>/dev/null || : return 0 fi rm -f -- "$tmp" @@ -380,27 +584,49 @@ fm_procevent_claim_capture_reservation_remove_locked() { fm_procevent_capture_reservation_remove_claim "$FM_PROCEVENT_CLAIM_STATE_ROOT" "$FM_PROCEVENT_CLAIM_TOKEN" } +# fm_procevent_claim_generation_gone_locked +# True only when the loaded claim's owner is stale and the process group it led +# independently has no members left. The separate group check also covers a +# reused live pid whose identity differs while the old generation survives. +# A live matched owner (state 0), an unreadable identity (state 2), and a +# crashed leader with a still-live ambiguous group (state 3) all return false. +fm_procevent_claim_generation_gone_locked() { + local state=0 + fm_procevent_pid_state "${FM_PROCEVENT_CLAIM_PID:-}" "${FM_PROCEVENT_CLAIM_IDENTITY:-}" || state=$? + [ "$state" -eq 1 ] \ + && ! fm_procevent_group_alive "${FM_PROCEVENT_CLAIM_PID:-}" +} + +# Capture-reservation cleanup for a claim being reclaimed. +# +# Reservation records are keyed by CLAIM TOKEN, and every replacement claims a +# fresh token, so a dead generation's leftovers can never collide with the +# generation that replaces it. They are hygiene, not an ownership invariant - +# the runner's own successful-capture path already tidies them best-effort. +# The cleanup is still attempted and remains authoritative for a generation +# that is not provably gone; it stops being a veto only after the stale owner +# and independent group check prove the whole generation gone. +fm_procevent_claim_capture_reservation_reclaim_locked() { + fm_procevent_claim_capture_reservation_remove_locked && return 0 + fm_procevent_claim_generation_gone_locked +} + # fm_procevent_group_alive <pid> -# True while any process remains in the process group a runner leads. A runner -# started by reconcile is its own group leader, so this is what distinguishes a -# generation that is really gone from one whose leader died while its blocking -# source child kept running. +# True while any process remains in the runner's numeric process group. A runner +# starts as its own group leader, but after that leader exits a same-numbered +# group may be reused, so group presence prevents proving the generation gone. fm_procevent_group_alive() { case "$1" in ''|*[!0-9]*) return 1 ;; esac kill -0 -"$1" 2>/dev/null } # fm_procevent_pid_state <pid> <identity> -# 0 live match, 1 stale, 2 uncertain, 3 orphaned group. +# 0 live match, 1 stale, 2 uncertain, 3 ambiguous leaderless group. # -# State 3 is the crash cut: the runner leader is gone, but its owned process -# group still has members, so the old generation can still be consuming the -# source. Treating that as stale would release ownership and let a second -# poller start against one canonical source. Only the leader being absent -# reaches state 3, which is also what makes signalling that group safe: if this -# pid had been reused by an unrelated process the leader would be alive, so the -# identity comparison below would classify it stale or uncertain and no group -# signal would ever follow. +# State 3 is the crash cut: the runner leader is gone, but a process group with +# its numeric id still has members. That group may be the old generation or a +# leaderless group created after PID/PGID reuse, so cleanup preserves the claim +# without signalling the group or starting a replacement. fm_procevent_pid_state() { local pid=$1 expected=$2 actual if ! fm_pid_alive "$pid"; then @@ -469,7 +695,7 @@ fm_procevent_claim_acquire_locked() { fi fi if [ "$status" -eq 0 ]; then - fm_procevent_claim_capture_reservation_remove_locked || status=1 + fm_procevent_claim_capture_reservation_reclaim_locked || status=1 fi [ "$status" -ne 0 ] || rm -f -- "$claim" || status=1 else @@ -499,6 +725,11 @@ fm_procevent_claim_acquire_locked() { if [ "$status" -eq 0 ]; then FM_PROCEVENT_CLAIM_TOKEN=$token FM_PROCEVENT_CLAIM_REG_IDENTITY=$reg_identity + FM_PROCEVENT_CLAIM_STATE_ROOT=$state_root + FM_PROCEVENT_CLAIM_STATE_DEVICE=$state_device + FM_PROCEVENT_CLAIM_STATE_INODE=$state_inode + FM_PROCEVENT_CLAIM_STATE_OWNER=$state_owner + FM_PROCEVENT_CLAIM_STATE_MODE=$state_mode fi fi [ "$status" -eq 0 ] || { [ -z "${tmp:-}" ] || rm -f -- "$tmp"; } @@ -544,8 +775,28 @@ fm_procevent_claim_mark_terminal_locked() { } # fm_procevent_claim_release_locked <source-id> <home> <pid> <token> +# The live owner uses this path for its own release. Reservation cleanup must +# succeed normally; stale-generation relaxation is never consulted. fm_procevent_claim_release_locked() { - local id=$1 home=$2 pid=$3 token=$4 claim + fm_procevent_claim_release_mode_locked release "$@" +} + +# fm_procevent_claim_release_terminal_self_locked <source-id> <home> <pid> <token> +# A live runner uses this only while retiring its own terminal source mid-capture. +# Its in-flight reservation is transient, so attempt cleanup without making that +# cleanup a veto; exact ownership still must match before releasing the claim. +fm_procevent_claim_release_terminal_self_locked() { + fm_procevent_claim_release_mode_locked terminal-self "$@" +} + +# fm_procevent_claim_reclaim_locked <source-id> <home> <pid> <token> +# Lifecycle commands use this only after proving or stopping a dead generation. +fm_procevent_claim_reclaim_locked() { + fm_procevent_claim_release_mode_locked reclaim "$@" +} + +fm_procevent_claim_release_mode_locked() { + local mode=$1 id=$2 home=$3 pid=$4 token=$5 claim fm_procevent_source_id_valid "$id" || return 1 claim=$(fm_procevent_claim_path "$id") [ -e "$claim" ] || return 0 @@ -553,7 +804,18 @@ fm_procevent_claim_release_locked() { && [ "$FM_PROCEVENT_CLAIM_HOME" = "$home" ] \ && [ "$FM_PROCEVENT_CLAIM_PID" = "$pid" ] \ && [ "$FM_PROCEVENT_CLAIM_TOKEN" = "$token" ]; then - fm_procevent_claim_capture_reservation_remove_locked || return 1 + case "$mode" in + reclaim) + fm_procevent_claim_capture_reservation_reclaim_locked || return 1 + ;; + terminal-self) + fm_procevent_claim_capture_reservation_remove_locked || true + ;; + release) + fm_procevent_claim_capture_reservation_remove_locked || return 1 + ;; + *) return 1 ;; + esac rm -f -- "$claim" return $? fi @@ -584,7 +846,7 @@ fm_procevent_path_normalize() { fm_procevent_directory_owned_by_current_user() { local owner if [ "$(uname)" = Darwin ]; then - owner=$(stat -f %u "$1" 2>/dev/null) + owner=$(/usr/bin/stat -f %u "$1" 2>/dev/null) else owner=$(stat -c %u "$1" 2>/dev/null) fi diff --git a/bin/fm-procevent.sh b/bin/fm-procevent.sh index 10e98180342..aaa7d60c541 100755 --- a/bin/fm-procevent.sh +++ b/bin/fm-procevent.sh @@ -137,19 +137,44 @@ # captain chose; the intake owns every rule about what happens next. This runner # names no adapter, parses no result, and knows no decision rule, so a future # built-in source needs nothing here beyond an `answers` command and a binding. -# External binding responses never enter this authority-bearing intake. +# Reconcile selections use the parallel `reconciles` adapter command and the +# binding-verified `reconcile-requests` intake, never the keyed-answer value. +# External binding responses never enter either authority-bearing intake. # # Feeding is deliberately independent of handling: it never acknowledges a result # and never suppresses a wake. Recording the captain's answer is transcription, # while ACTING on it is firstmate's judgement, so the capture stays unacknowledged # and its `check` wake reaches the handler exactly as it would have anyway. # +# A runner is bound to the HOME that owns it, not to the one session that armed +# it: a persistent source is meant to outlive that session, so reconcile stops a +# runner whose source is retired in a live home, and this lease is the backstop +# for a home that is GONE. Detaching a runner into its own +# process group is what lets a persistent source outlive the turn that armed it, +# and with nothing else it is also what lets a runner outlive its whole home: +# reparented to init, it keeps its blocking child - and everything that child +# spawns - running with nobody left to reap it. So every runner starts a small +# guard beside it, in its own separate process group, which re-reads the owning +# state root's lease on a bounded cadence and stops the runner's whole process +# group once that lease can no longer be proved fresh. Owner-presence operations +# refresh the lease, an attached public start keeps it fresh while its caller +# remains attached, and the watcher's reconcile cycle keeps it fresh in a live +# home. A runner exports the inherited FM_PROCEVENT_IN_RUNNER marker and every +# refresh is skipped under it, so a runner and its ordinary children do not +# certify their own owner. That rule is CONFUSED-AGENT-GRADE, the grade +# bin/fm-lease-lib.sh documents: a source that DELIBERATELY strips the marker +# can still refresh, and adversarial-grade unforgeability is out of scope (see +# docs/configuration.md). Scope is the owning state root and one runner +# generation, never a script or process name, so a live source in +# another home is untouched. See bin/fm-procevent-lib.sh for the lease itself. +# # Ownership is machine-wide per canonical source, because separate Firstmate # homes can share one underlying source store. A live owner is never displaced; -# only a claim whose whole generation is gone is reclaimed. A runner leads its -# own process group, so a crashed leader whose group still has members is not -# stale: reconcile stops that surviving group and releases its generation before -# any replacement starts, and keeps the claim for a later retry when it cannot. +# only a claim whose stale owner and independently absent process group prove +# its whole generation gone is reclaimed. A crashed leader or reused pid whose +# process group still has members cannot relax ownership cleanup. Reconcile +# signals only a live identity-matched runner group and otherwise keeps the +# claim without starting a replacement. # # Durability boundary: see bin/fm-procevent-lib.sh. This runner proves capture # before publication and bounded re-announcement until handled, and nothing @@ -360,6 +385,19 @@ feed_keyed_answers() { # <adapter> <source-id> <result-file> --source "the captured result $id sequence $seq" >/dev/null 2>&1 } +feed_reconcile_requests() { # <adapter> <source-id> <result-file> + local adapter=$1 id=$2 result=$3 script origin seq rows + script=$(adapter_script "$adapter") + [ -f "$script" ] && [ ! -L "$script" ] || return 1 + origin=$("$SCRIPT_DIR/fm-captain-hold.sh" binding "$id" 2>/dev/null) || return 1 + [ -n "$origin" ] || return 1 + seq=$(fm_procevent_result_sequence "$result") || return 1 + rows=$("$script" reconciles "$result" 2>/dev/null) || return 1 + printf '%s\n' "$rows" \ + | "$SCRIPT_DIR/fm-captain-hold.sh" reconcile-requests \ + --source-id "$id" --source "the captured result $id sequence $seq" >/dev/null 2>&1 +} + read_adapter() { # <source-id> local f; f=$(source_file "$1") [ -f "$f" ] && [ ! -L "$f" ] || return 1 @@ -421,6 +459,7 @@ cmd_register() { die "cannot publish the registration" fi fm_procevent_source_lock_release "$id" + owner_lease_refresh printf 'registered: %s (%s)\n' "$id" "$adapter" } @@ -510,6 +549,7 @@ cmd_register_extension() { fi fm_procevent_source_lock_release "$id" extension_lifecycle_lock_release + owner_lease_refresh printf 'registered: %s (%s from %s@%s)\n' "$id" "$adapter" "$extension_id" "$extension_version" printf 'owner-token: %s\n' "$registration_token" printf 'retire: bin/fm-procevent.sh retire %s --if-owner %s\n' "$id" "$registration_token" @@ -569,8 +609,15 @@ publish_pending() { # [result-file-to-skip] printf '%s\n' "$published" } -isolate_runner() { # <wait|detach> <source-id> - local mode=$1 id=$2 program +# Start one command as the leader of a fresh process group, either waiting for +# it (the public `start` boundary) or detaching from it (reconcile's restart and +# the runner's own owner guard). The guard deliberately gets its OWN group +# rather than joining the runner's: it has to survive the group signal it sends, +# and a member of the runner's group would also make that group read as alive +# after the runner itself is gone. +isolate_process() { # <wait|detach> <command> [argv...] + local mode=$1 program + shift # shellcheck disable=SC2016 # Perl owns every $ expression in this literal program. program='my $mode = shift @ARGV; defined(my $pid = fork) or exit 125; @@ -586,27 +633,66 @@ isolate_runner() { # <wait|detach> <source-id> exit(128 + ($status & 127)) if $status & 127; exit($status >> 8);' if [ "$mode" = wait ]; then - exec perl -e "$program" "$mode" "$SCRIPT_DIR/fm-procevent.sh" _start "$id" + perl -e "$program" "$mode" "$@" + return $? fi - perl -e "$program" "$mode" "$SCRIPT_DIR/fm-procevent.sh" _start "$id" >/dev/null 2>&1 & + perl -e "$program" "$mode" "$@" >/dev/null 2>&1 & } -require_runner_group() { - local pgid +isolate_runner() { # <wait|detach> <source-id> + isolate_process "$1" "$SCRIPT_DIR/fm-procevent.sh" _start "$2" +} + +require_isolated_group() { # <role> + local role=$1 pgid [ "${FM_PROCEVENT_RUNNER_GROUP:-}" = "$$" ] \ - || die "runner process group was not isolated" + || die "$role process group was not isolated" pgid=$(ps -o pgid= -p "$$" 2>/dev/null | tr -d '[:space:]') \ - || die "cannot inspect runner process group" - [ -n "$pgid" ] || die "cannot inspect runner process group" - [ "$pgid" = "$$" ] || die "runner does not lead its process group" + || die "cannot inspect $role process group" + [ -n "$pgid" ] || die "cannot inspect $role process group" + [ "$pgid" = "$$" ] || die "$role does not lead its process group" unset FM_PROCEVENT_RUNNER_GROUP } +require_runner_group() { require_isolated_group runner; } + +# Record owner-presence activity for this home. Skipped under the inherited +# FM_PROCEVENT_IN_RUNNER marker, so a runner and its ordinary children do not +# keep refreshing their own lease after the home goes away. +# Confused-agent-grade: a source that deliberately unsets the marker can still +# refresh, and that is out of scope (see docs/configuration.md). +owner_lease_refresh() { + [ "${FM_PROCEVENT_IN_RUNNER:-0}" = 1 ] && return 0 + fm_procevent_owner_lease_touch "$STATE" 2>/dev/null || true +} + +owner_lease_keepalive() { # <parent-pid> <parent-identity> + local parent=$1 identity=$2 state + while :; do + sleep 1 + fm_procevent_pid_state "$parent" "$identity" + state=$? + case "$state" in + 0) owner_lease_refresh ;; + 2) ;; + *) return 0 ;; + esac + done +} + cmd_start_public() { - local id=${1-} + local id=${1-} identity keeper status [ "$#" -eq 1 ] || usage fm_procevent_source_id_valid "$id" || die "source id must be path-safe: $id" + owner_lease_refresh + identity=$(fm_pid_identity "$$" 2>/dev/null) || die "cannot identify the attached owner" + owner_lease_keepalive "$$" "$identity" & + keeper=$! isolate_runner wait "$id" + status=$? + kill "$keeper" 2>/dev/null || true + wait "$keeper" 2>/dev/null || true + return "$status" } cmd_start() { @@ -665,6 +751,10 @@ cmd_start() { die "extension registration owner is unreadable: $id" ;; esac + exec 7<"$(source_file "$id")" || { + fm_procevent_source_lock_release "$id" + die "cannot retain registration identity: $id" + } fm_procevent_claim_acquire_locked "$id" "$FM_HOME" "$$" "$(source_file "$id")" "$STATE" claimed=$? fm_procevent_source_lock_release "$id" @@ -678,6 +768,8 @@ cmd_start() { CLAIM_PID=$$ CLAIM_TOKEN=$FM_PROCEVENT_CLAIM_TOKEN CLAIM_REG_IDENTITY=$FM_PROCEVENT_CLAIM_REG_IDENTITY + CLAIM_STATE_DEVICE=$FM_PROCEVENT_CLAIM_STATE_DEVICE + CLAIM_STATE_INODE=$FM_PROCEVENT_CLAIM_STATE_INODE STAGED_OUTPUT= release_start_claim() { extension_lifecycle_lock_release 2>/dev/null || true @@ -695,7 +787,14 @@ cmd_start() { fm_procevent_source_lock_release "$CLAIM_ID" 2>/dev/null || true } trap release_start_claim EXIT - local runner inbox reservation_dir staging + # The inherited marker keeps the runner and its ordinary children from + # accidentally refreshing the owner lease. A source that deliberately strips + # it is outside this confused-agent-grade boundary. + export FM_PROCEVENT_IN_RUNNER=1 + start_owner_guard "$id" || die "cannot start the runner's owner guard: $id" + local launch_floor runner inbox reservation_dir staging launch_ready launch_reply launch_pid + launch_floor=$(fm_procevent_launch_floor_seconds) \ + || die "FM_PROCEVENT_LAUNCH_FLOOR_SECONDS must be whole seconds from $FM_PROCEVENT_LAUNCH_FLOOR_MIN_SECONDS to $FM_PROCEVENT_LAUNCH_FLOOR_MAX_SECONDS" if [ "$extension_owner" -eq 1 ]; then staging=$(fm_procevent_extension_staging_prepare "$STATE") \ || die "cannot safely prepare the external registry staging boundary" @@ -732,13 +831,45 @@ cmd_start() { # Built-in adapters do not run the extension capture helper, so keep this # sentinel defined while sharing the no-result branch below under `set -u`. local truncated=0 capture_state='' durable='' reservation_terminal='' reservation_silent='' + fm_procevent_launch_floor_wait "$STATE" "$id" "$CLAIM_REG_IDENTITY" "$launch_floor" + case "$?" in + 0) ;; + # A superseded generation leaves nothing behind. The runner marker is + # written before this wait, and a home sweep counts a marker with no owned + # claim as a preflight failure, so exiting without removing it would make + # that home refuse to sweep. + 2) [ "$extension_owner" -eq 1 ] || rm -f -- "$runner"; exit 0 ;; + *) die "cannot enforce the source launch floor: $id" ;; + esac + exec 7<&- if [ "$extension_owner" -eq 1 ]; then - capture_state=$(perl "$SCRIPT_DIR/fm-procevent-extension-capture.pl" \ + launch_ready=".$id.$CLAIM_TOKEN.launch-ready" + launch_reply="$REG/.$id.$CLAIM_TOKEN.launch-reply" + (umask 077; : > "$REG/$launch_ready" && : > "$launch_reply") || { + rm -f -- "$REG/$launch_ready" "$launch_reply" + fm_procevent_source_lock_release "$id" + die "cannot prepare the source launch boundary: $id" + } + perl "$SCRIPT_DIR/fm-procevent-extension-capture.pl" \ 9 8 6 "$id" "$adapter" "$FM_PROCEVENT_EXTENSION_ID" \ "$FM_PROCEVENT_EXTENSION_VERSION" "$FM_PROCEVENT_EXTENSION_CAPABILITY_VERSION" \ "$FM_PROCEVENT_EXTENSION_PACKAGE_DIGEST" "$FM_PROCEVENT_EXTENSION_BINDING_DIGEST" \ - "$CLAIM_TOKEN" "$runner" "$out" "$$" "$(fm_pid_identity "$$")" "$MAX_OUTPUT_BYTES" -- "${ARGV[@]}") \ - || die "cannot safely stage the extension result" + "$CLAIM_TOKEN" "$runner" "$out" "$$" "$(fm_pid_identity "$$")" "$MAX_OUTPUT_BYTES" \ + "$launch_ready" -- "${ARGV[@]}" > "$launch_reply" & + launch_pid=$! + while [ ! -s "$REG/$launch_ready" ] && kill -0 "$launch_pid" 2>/dev/null; do sleep 0.01; done + fm_procevent_source_lock_release "$id" \ + || die "cannot release the source launch boundary: $id" + wait "$launch_pid" || { + rm -f -- "$REG/$launch_ready" "$launch_reply" + die "cannot safely stage the extension result" + } + [ -s "$REG/$launch_ready" ] || { + rm -f -- "$REG/$launch_ready" "$launch_reply" + die "cannot establish the source launch boundary: $id" + } + IFS= read -r capture_state < "$launch_reply" || capture_state= + rm -f -- "$REG/$launch_ready" "$launch_reply" IFS=$'\t' read -r capture_state durable rc truncated reservation_terminal reservation_silent <<EOF $capture_state EOF @@ -753,10 +884,38 @@ EOF FM_PROCEVENT_CAPTURE_RESERVATION_SILENT=$reservation_silent fi else - [ ! -e "$out" ] && [ ! -L "$out" ] || die "cannot safely stage output" - (umask 077; : > "$out") || die "cannot stage output" + [ ! -e "$out" ] && [ ! -L "$out" ] || { + fm_procevent_source_lock_release "$id" + die "cannot safely stage output" + } + (umask 077; : > "$out") || { + fm_procevent_source_lock_release "$id" + die "cannot stage output" + } STAGED_OUTPUT=$out - "${ARGV[@]}" 2>/dev/null | perl -e ' + launch_ready="$REG/.$id.$CLAIM_TOKEN.launch-pipe" + mkfifo -m 600 "$launch_ready" || { + fm_procevent_source_lock_release "$id" + die "cannot prepare the source launch boundary: $id" + } + exec 5<> "$launch_ready" || { + rm -f -- "$launch_ready" + fm_procevent_source_lock_release "$id" + die "cannot retain the source launch boundary: $id" + } + exec 4< "$launch_ready" || { + exec 5>&- + rm -f -- "$launch_ready" + fm_procevent_source_lock_release "$id" + die "cannot retain the source output boundary: $id" + } + "${ARGV[@]}" >&5 5>&- 4<&- 2>/dev/null & + launch_pid=$! + exec 5>&- + rm -f -- "$launch_ready" + fm_procevent_source_lock_release "$id" \ + || die "cannot release the source launch boundary: $id" + perl -e ' use strict; use warnings; my $limit = shift; @@ -779,10 +938,11 @@ EOF $truncated = 1 if $take < $count; } exit($truncated ? 3 : 0); - ' "$MAX_OUTPUT_BYTES" > "$out" - local pipe_status=("${PIPESTATUS[@]}") - rc=${pipe_status[0]} - bound_rc=${pipe_status[1]} + ' "$MAX_OUTPUT_BYTES" <&4 > "$out" + bound_rc=$? + exec 4<&- + wait "$launch_pid" + rc=$? case "$bound_rc" in 0) ;; 3) truncated=1 ;; @@ -816,6 +976,10 @@ EOF # Independent of publication and acknowledgement, so it runs once per capture # for every adapter and cannot change what the handler receives. + if [ "$extension_owner" -eq 0 ] \ + && feed_reconcile_requests "$adapter" "$id" "$durable"; then + printf 'reconciles-fed: %s\n' "$id" + fi if [ "$extension_owner" -eq 0 ] \ && feed_keyed_answers "$adapter" "$id" "$durable"; then printf 'answers-fed: %s\n' "$id" @@ -892,7 +1056,7 @@ retire_owned_terminal_source() { # <source-id> && [ "$current_identity" = "$CLAIM_REG_IDENTITY" ] \ && fm_procevent_claim_mark_terminal_locked "$id" "$CLAIM_HOME" "$CLAIM_PID" "$CLAIM_TOKEN"; then if rm -f -- "$registration" && [ ! -e "$registration" ] && [ ! -L "$registration" ]; then - fm_procevent_claim_release_locked "$id" "$CLAIM_HOME" "$CLAIM_PID" "$CLAIM_TOKEN" || status=1 + fm_procevent_claim_release_terminal_self_locked "$id" "$CLAIM_HOME" "$CLAIM_PID" "$CLAIM_TOKEN" || status=1 else status=1 fi @@ -903,6 +1067,106 @@ retire_owned_terminal_source() { # <source-id> return "$status" } +# Bind this runner's lifetime to the home that owns it. Started once the +# claim is held, so the guard names the exact generation it protects, and +# detached into its OWN process group so the group signal it may later send +# reaches the runner and every descendant without killing the guard first. +# If signalling cannot be proved safe or does not finish, the guard remains +# alive and retries on its normal check cadence rather than abandoning cleanup. +start_owner_guard() { # <source-id> + local identity ready value + identity=$(fm_pid_identity "$$" 2>/dev/null) || return 1 + ready=$(umask 077; mktemp "$REG/.owner-guard-ready.XXXXXX") || return 1 + if ! isolate_process detach "$SCRIPT_DIR/fm-procevent.sh" _owner-watchdog \ + "$1" "$$" "$identity" "$ready" "$CLAIM_STATE_DEVICE" "$CLAIM_STATE_INODE"; then + rm -f -- "$ready" + return 1 + fi + for _ in $(seq 1 50); do + if [ -s "$ready" ]; then + IFS= read -r value < "$ready" || value= + rm -f -- "$ready" + [ "$value" = ready ] + return $? + fi + sleep 0.1 + done + rm -f -- "$ready" + return 1 +} + +# The runner's owner guard, which bounds an accidentally orphaned detached +# runner after its home ends. It revalidates the recorded physical state root +# and its lease on a bounded cadence and, after two consecutive checks cannot prove +# both, invokes the identity-gated stop for the runner's whole process group - +# which is what reaches the blocking child and everything that child spawned, +# exactly as retirement does. A failed verified stop stays on the retry cadence; +# an absent leader ends the guard without signalling an ambiguous group. +# +# Scope is the owning state root and this one runner generation. It never +# matches on a script name, a command line, or a process name: those are shared +# by every home running the same adapter, and a live source in another home +# proves its own owner through that home's own lease. +cmd_owner_watchdog() { # <source-id> <runner-pid> <runner-identity> <ready-file> <state-device> <state-inode> + local id=${1-} pid=${2-} identity=${3-} ready=${4-} state_device=${5-} state_inode=${6-} + local lease tick misses=0 pid_state state_identity current_device current_inode + [ "$#" -eq 6 ] || usage + fm_procevent_source_id_valid "$id" || die "source id must be path-safe: $id" + case "$pid" in ''|*[!0-9]*) die "runner pid must be a positive integer: $pid" ;; esac + [ -n "$identity" ] || die "runner identity is required" + case "$state_device" in ''|*[!0-9]*) die "state device must be an integer" ;; esac + case "$state_inode" in ''|*[!0-9]*) die "state inode must be an integer" ;; esac + [ "${ready%/*}" = "$REG" ] && [ -f "$ready" ] && [ ! -L "$ready" ] \ + || die "owner guard readiness boundary is invalid" + trap 'printf "failed\n" > "$ready" 2>/dev/null || true' EXIT + require_isolated_group guard + lease=$(fm_procevent_owner_lease_seconds) \ + || die "FM_PROCEVENT_OWNER_LEASE_SECONDS must be whole seconds from $FM_PROCEVENT_OWNER_LEASE_MIN_SECONDS to $FM_PROCEVENT_OWNER_LEASE_MAX_SECONDS" + tick=$(fm_procevent_owner_check_seconds) \ + || die "FM_PROCEVENT_OWNER_CHECK_SECONDS must be whole seconds from $FM_PROCEVENT_OWNER_CHECK_MIN_SECONDS to $FM_PROCEVENT_OWNER_CHECK_MAX_SECONDS" + fm_procevent_pid_state "$pid" "$identity" + pid_state=$? + [ "$pid_state" -eq 0 ] || die "runner identity changed before owner guard initialization" + state_identity=$(fm_procevent_claim_state_root_identity "$STATE") \ + || die "owning state root identity is unreadable at owner guard initialization" + IFS=$'\t' read -r _ current_device current_inode _ _ <<< "$state_identity" + [ "$current_device" = "$state_device" ] && [ "$current_inode" = "$state_inode" ] \ + || die "owning state root identity changed before owner guard initialization" + fm_procevent_owner_alive "$STATE" "$lease" \ + || die "owning home lease is not fresh at owner guard initialization" + printf 'ready\n' > "$ready" || die "cannot confirm owner guard initialization" + trap - EXIT + while :; do + sleep "$tick" + fm_procevent_pid_state "$pid" "$identity" + pid_state=$? + case "$pid_state" in + 1|3) exit 0 ;; + 0) ;; + *) continue ;; + esac + state_identity=$(fm_procevent_claim_state_root_identity "$STATE" 2>/dev/null || true) + current_device= + current_inode= + [ -z "$state_identity" ] \ + || IFS=$'\t' read -r _ current_device current_inode _ _ <<< "$state_identity" + if [ "$current_device" = "$state_device" ] \ + && [ "$current_inode" = "$state_inode" ] \ + && fm_procevent_owner_alive "$STATE" "$lease"; then + misses=0 + continue + fi + # Two consecutive misses, so one unreadable read cannot end a live runner. + misses=$((misses + 1)) + [ "$misses" -ge 2 ] || continue + if stop_runner_pid "$pid" "$identity"; then + exit 0 + fi + # Identity/group inspection and signalling can fail transiently. Keep the + # guard alive so the next normal tick retries the same generation cleanup. + done +} + # Start a runner outside the watcher cycle that noticed it was missing. The # public start boundary establishes its own process group before claiming. detach_runner() { # <source-id> @@ -911,6 +1175,7 @@ detach_runner() { # <source-id> cmd_reconcile() { local rec id published started=0 stopped=0 uncertain=0 claim owner pid token identity claim_state stop_state + owner_lease_refresh published=$(publish_pending) # Stop a runner this home owns whose source is no longer registered. Without @@ -942,7 +1207,7 @@ cmd_reconcile() { stop_state=$? case "$stop_state" in 0|1) - if fm_procevent_claim_release_locked "$id" "$owner" "$pid" "$token" 2>/dev/null; then + if fm_procevent_claim_reclaim_locked "$id" "$owner" "$pid" "$token" 2>/dev/null; then rm -f -- "$(staging_file "$id" "$token")" rm -f -- "$(runner_file "$id")" stopped=$((stopped + 1)) @@ -982,36 +1247,14 @@ cmd_reconcile() { && rm -f -- "$(source_file "$id")" \ && [ ! -e "$(source_file "$id")" ] \ && [ ! -L "$(source_file "$id")" ] \ - && fm_procevent_claim_release_locked "$id" "$owner" "$pid" "$token" 2>/dev/null; then + && fm_procevent_claim_reclaim_locked "$id" "$owner" "$pid" "$token" 2>/dev/null; then stopped=$((stopped + 1)) else uncertain=$((uncertain + 1)) fi elif [ "$claim_state" -eq 3 ]; then - # The leader crashed but its owned group is still consuming the - # source. Never start a replacement alongside it: stop that group and - # release its generation first, and if either cannot be proved, keep - # the claim and retry on a later cycle rather than adding a second - # poller. Only the owning home may signal its own group. - owner=$FM_PROCEVENT_CLAIM_HOME - pid=$FM_PROCEVENT_CLAIM_PID - token=$FM_PROCEVENT_CLAIM_TOKEN - identity=$FM_PROCEVENT_CLAIM_IDENTITY - stop_state=2 - if fm_procevent_claim_owned_by_state "$STATE" "$FM_HOME"; then - stop_runner_pid "$pid" "$identity" - stop_state=$? - fi - if [ "$stop_state" -eq 0 ] \ - && cleanup_extension_registration_invocations_locked "$id" \ - && fm_procevent_claim_release_locked "$id" "$owner" "$pid" "$token" 2>/dev/null; then - rm -f -- "$(staging_file "$id" "$token")" - rm -f -- "$(runner_file "$id")" - fm_procevent_source_lock_release "$id" - detach_runner "$id" - started=$((started + 1)) - continue - fi + # A leaderless group's generation is ambiguous under PID/PGID reuse, + # so preserve its claim without signalling or starting a replacement. uncertain=$((uncertain + 1)) elif [ "$claim_state" -eq 2 ]; then uncertain=$((uncertain + 1)) @@ -1027,38 +1270,40 @@ cmd_reconcile() { # its own process group leader, so the group signal is what actually reaches the # blocking child - signalling only the runner would leave that child alive and # reparented, which is exactly how a source that never completes leaks. -stop_runner_pid() { # <pid> <identity> - local pid=${1-} identity=${2-} state pgid i=0 - case "$pid" in ''|*[!0-9]*) return 2 ;; esac - [ -n "$identity" ] || return 2 +runner_group_signal() { # <signal> <pid> <identity> + local signal=$1 pid=$2 identity=$3 state pgid + # KNOWN LIMIT: only an alive identity-matched leader proves group ownership. + # Detected reused PIDs and absent leaders are refused before signalling; + # launch pacing, leases, and reconcile cleanup are the backstop. fm_procevent_pid_state "$pid" "$identity" state=$? case "$state" in - 0) - # A live identity-matched leader still owns its group, so prove the group - # really is the one this pid leads before signalling it. - pgid=$(ps -o pgid= -p "$pid" 2>/dev/null | tr -d '[:space:]') || return 2 - [ "$pgid" = "$pid" ] || return 2 - ;; - 3) - # The leader crashed but its owned group is still running. Its pgid cannot - # be read from the dead leader, and it does not need to be: only an absent - # leader reaches this state, so the group cannot belong to a reused pid. - ;; - *) return "$state" ;; + 0) ;; + 1) fm_procevent_group_alive "$pid" && return 2; return 1 ;; + *) return 2 ;; esac - kill -TERM -"$pid" 2>/dev/null || return 2 + pgid=$(ps -o pgid= -p "$pid" 2>/dev/null | tr -d '[:space:]') || return 2 + [ "$pgid" = "$pid" ] || return 2 + # KNOWN LIMIT: portable shell cannot make this verification and signal atomic, + # so the PID and group could be reused in the interval between them. + kill -"$signal" -"$pid" 2>/dev/null || return 2 +} + +stop_runner_pid() { # <pid> <identity> + local pid=${1-} identity=${2-} signal_state i=0 + case "$pid" in ''|*[!0-9]*) return 2 ;; esac + [ -n "$identity" ] || return 2 + runner_group_signal TERM "$pid" "$identity" + signal_state=$? + [ "$signal_state" -eq 0 ] || return "$signal_state" while [ "$i" -lt 20 ]; do kill -0 -"$pid" 2>/dev/null || return 0 - if kill -0 "$pid" 2>/dev/null; then - fm_procevent_pid_state "$pid" "$identity" - state=$? - [ "$state" -eq 2 ] && return 2 - fi sleep 0.1 i=$((i + 1)) done - kill -KILL -"$pid" 2>/dev/null || return 2 + runner_group_signal KILL "$pid" "$identity" + signal_state=$? + [ "$signal_state" -eq 0 ] || return "$signal_state" i=0 while [ "$i" -lt 20 ]; do kill -0 -"$pid" 2>/dev/null || return 0 @@ -1095,6 +1340,7 @@ cmd_handled() { local id=${1-} seq=${2-} status fm_procevent_source_id_valid "$id" || die "source id must be path-safe: $id" case "$seq" in ''|*[!0-9]*) die "sequence must be a nonnegative integer: $seq" ;; esac + owner_lease_refresh fm_procevent_source_lock_acquire "$id" || die "cannot lock source: $id" fm_procevent_mark_handled "$STATE" "$id" "$seq" status=$? @@ -1193,7 +1439,7 @@ cmd_retire() { fm_procevent_source_lock_release "$id" die "cannot prove external adapter cleanup; source remains registered: $id" fi - if ! fm_procevent_claim_release_locked "$id" "$owner" "$pid" "$token"; then + if ! fm_procevent_claim_reclaim_locked "$id" "$owner" "$pid" "$token"; then fm_procevent_source_lock_release "$id" die "cannot release source ownership: $id" fi @@ -1371,6 +1617,7 @@ cmd_sweep_home() { cmd_list() { local rec id adapter owner pending + owner_lease_refresh if ! fm_procevent_any_registered "$STATE"; then printf 'no sources registered\n' return 0 @@ -1491,6 +1738,7 @@ case "${1-}" in register-extension) shift; cmd_register_extension "$@" ;; start) shift; cmd_start_public "$@" ;; _start) shift; cmd_start "$@" ;; + _owner-watchdog) shift; cmd_owner_watchdog "$@" ;; reconcile) shift; cmd_reconcile "$@" ;; classify) shift; cmd_classify "$@" ;; handled) shift; cmd_handled "$@" ;; diff --git a/bin/fm-push-transition-lib.sh b/bin/fm-push-transition-lib.sh index 5e38b174c31..89977ac166f 100644 --- a/bin/fm-push-transition-lib.sh +++ b/bin/fm-push-transition-lib.sh @@ -119,7 +119,7 @@ _fm_surface_digest() { open_decision_id() { # <task> local task=$1 open [ -n "$task" ] || return 0 - open=$(status_open_decisions "$STATE/$task.status") + open=$(status_open_decisions "$STATE/$task.status" --explicit-only) [ -n "$open" ] || return 0 printf '%s' "$open" | _fm_surface_digest } diff --git a/bin/fm-remote-file.sh b/bin/fm-remote-file.sh index 34a993db5b4..31887ac27ba 100755 --- a/bin/fm-remote-file.sh +++ b/bin/fm-remote-file.sh @@ -77,7 +77,7 @@ snapshot_bounded_file() { # <file> <max-bytes> <destination> <size-file> directory_identity() { if [ "$(uname)" = Darwin ]; then - stat -f '%d:%i' . 2>/dev/null + /usr/bin/stat -f '%d:%i' . 2>/dev/null else stat -c '%d:%i' . 2>/dev/null fi diff --git a/bin/fm-remote-inherit-push.sh b/bin/fm-remote-inherit-push.sh index f725faf64d6..028fabccf79 100755 --- a/bin/fm-remote-inherit-push.sh +++ b/bin/fm-remote-inherit-push.sh @@ -27,7 +27,7 @@ sha256_file() { if command -v shasum >/dev/null 2>&1; then shasum -a 256 "$1" | awk '{print $1}'; else sha256sum "$1" | awk '{print $1}'; fi } file_link_count() { - if [ "$(uname)" = Darwin ]; then stat -f %l "$1" 2>/dev/null; else stat -c %h "$1" 2>/dev/null; fi + if [ "$(uname)" = Darwin ]; then /usr/bin/stat -f %l "$1" 2>/dev/null; else stat -c %h "$1" 2>/dev/null; fi } shared_captain_header_valid() { local head diff --git a/bin/fm-remote-inherit.sh b/bin/fm-remote-inherit.sh index d56d5f4bb11..6e4ddea1d29 100755 --- a/bin/fm-remote-inherit.sh +++ b/bin/fm-remote-inherit.sh @@ -22,7 +22,7 @@ SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" die() { printf 'error: %s\n' "$1" >&2; exit 1; } usage() { sed -n '2,10p' "$0" | sed 's/^# \{0,1\}//'; exit 2; } file_link_count() { - if [ "$(uname)" = Darwin ]; then stat -f %l "$1" 2>/dev/null; else stat -c %h "$1" 2>/dev/null; fi + if [ "$(uname)" = Darwin ]; then /usr/bin/stat -f %l "$1" 2>/dev/null; else stat -c %h "$1" 2>/dev/null; fi } sha256_file() { if command -v shasum >/dev/null 2>&1; then shasum -a 256 "$1" | awk '{print $1}'; else sha256sum "$1" | awk '{print $1}'; fi diff --git a/bin/fm-remote-job-lib.sh b/bin/fm-remote-job-lib.sh index 5edcad27c69..a218fb32ba9 100755 --- a/bin/fm-remote-job-lib.sh +++ b/bin/fm-remote-job-lib.sh @@ -814,7 +814,7 @@ fm_remote_job_reap() { # <account-home> <id>; only removes an exact completed re fm_remote_job_path_mtime() { # <path> # The platform override controls worker shape in isolated tests, not the host # kernel's stat syntax. - if [ "$(uname -s 2>/dev/null || true)" = Darwin ]; then stat -f %m "$1" 2>/dev/null; else stat -c %Y "$1" 2>/dev/null; fi + if [ "$(uname -s 2>/dev/null || true)" = Darwin ]; then /usr/bin/stat -f %m "$1" 2>/dev/null; else stat -c %Y "$1" 2>/dev/null; fi } fm_remote_job_stage_owner_alive() { # <stage-dir> diff --git a/bin/fm-spawn.sh b/bin/fm-spawn.sh index fae68a6e2e4..6978cc4da0d 100755 --- a/bin/fm-spawn.sh +++ b/bin/fm-spawn.sh @@ -117,11 +117,13 @@ # even when they select different backends. A fresh spawn first takes the # per-home task-set lock and refuses rather than waits when forced teardown owns # it; relaunch is exempt because the existing task's control lock covers it. -# A fresh Treehouse-backed spawn also takes the project-identity lock in the root -# Firstmate home's state directory before slot allocation and holds it through +# A fresh Treehouse-backed spawn also takes the project-identity lock in the local +# root Firstmate home's state directory before slot allocation and holds it through # task metadata publication. Teardown holds that same lock while proving and # returning a slot, so allocation cannot reuse a slot before its owner record -# is published; +# is published. The local root is whatever bin/fm-wake-lib.sh's +# fm_firstmate_root_home resolves, so a home seeded from another machine anchors +# that lock itself rather than failing to resolve one; # contention refuses rather than waits. # With no harness arg, a crewmate/scout spawn resolves the CREW harness only when # config/crew-dispatch.json is absent. When that file exists, crewmate/scout @@ -215,11 +217,20 @@ # itself a linked worktree of the project repository still launches. A pane # that never reaches an isolated worktree refuses at the end of that wait, # naming the last path seen and why it was rejected. -# Only after this isolation check, a fresh ship, design, or scout's clean task worktree -# fetches origin, resolves the current remote default branch, and resets to its tip. -# Relaunch reuses the recorded worktree without fetching or resetting its base. -# An unreachable origin, unresolved default branch, or non-clean worktree -# refuses a fresh spawn rather than risking a PR based on stale history. +# That placement is proven only at launch. Every ship, design, or scout pane therefore +# also receives `export FM_TASK_ID=<task-id>` before the launch command, on +# the same channel as GOTMPDIR, and bin/fm-test-run.sh refuses to execute the +# behavior suite from the repository primary checkout while that marker is +# set (its header owns the refusal). A secondmate runs in its own home and is +# not marked. +# Only after this isolation check, every fresh ship, design, or scout requires a clean +# task worktree. When an origin configuration is detected, spawn fetches it, +# resolves the current remote default branch, and resets to its tip. When none +# is detected, spawn skips that remote freshness check and launches from the +# clean worktree's current HEAD. Relaunch reuses the recorded worktree without +# fetching or resetting its base. An unreachable detected origin, unresolved +# default branch, or non-clean worktree refuses a fresh spawn rather than +# risking a PR based on stale history or discarding local work. # A slot whose only deviation is a stale submodule gitlink is refused by that # same clean check, but is reported as a stale checkout naming each submodule # and both pins; nothing is converged or removed, and no remedy is suggested. @@ -254,7 +265,8 @@ # LANG LC_ALL LC_CTYPE TMPDIR TMP TEMP GOTMPDIR, plus backend identity/routing: # TMUX TMUX_PANE HERDR_ENV HERDR_SESSION HERDR_SOCKET_PATH HERDR_PANE_ID # CMUX_WORKSPACE_ID CMUX_SURFACE_ID CMUX_TAB_ID CMUX_PANEL_ID CMUX_SOCKET_PATH -# ZELLIJ ZELLIJ_SESSION_NAME ZELLIJ_PANE_ID FM_ZELLIJ_SESSION. +# ZELLIJ ZELLIJ_SESSION_NAME ZELLIJ_PANE_ID FM_ZELLIJ_SESSION, plus the task +# marker FM_TASK_ID that ship, design, and scout panes receive above. # An enabled task trace also retains TRACEPARENT. Explicit Firstmate launch # assignments still apply inside the filtered environment. Raw commands must # be POSIX sh compatible under this opt-in; the absent-file path is unchanged. @@ -315,6 +327,10 @@ # and every refusal; a failed registration stops this spawn rather than launching # a worker that would wedge on the dialog. A --secondmate launch never runs it, # so a claude secondmate home keeps its own one-time trust decision. +# Every claude launch also carries the attribution-off policy in its per-launch +# --settings JSON, so a spawned worker never writes a Co-Authored-By trailer, +# Claude-Session link, or generated-with line into a commit or PR body; +# bin/fm-launch-lib.sh owns the settings construction and its rationale. # Publishing the record and moving this home's backlog item to In flight are one # step, not two: bin/fm-backlog-transition-lib.sh owns that invariant, and this # script performs the transition under the task's own meta lock before it reports @@ -2395,8 +2411,37 @@ EOF printf '%s' "$lines" >&2 } +spawn_worktree_has_origin_config() { # <worktree> + # Resolved remote.origin.* variables cover Git's effective include/includeIf chain; raw headers are also detected in the worktree config and any included file Git names through another variable. Git cannot enumerate a variable-less included file, so an empty origin section that is its only content remains indistinguishable from absence and intentionally proceeds rather than reimplementing Git's config parser. + local worktree=$1 config origin key seen=$'\n' + git -C "$worktree" config --get-regexp '^remote\.origin\.' >/dev/null 2>&1 && return 0 + while IFS=$'\t' read -r origin key; do + case $origin in file:*) config=${origin#file:} ;; *) continue ;; esac + [ -f "$config" ] || continue + case $seen in *$'\n'"$config"$'\n'*) continue ;; esac + seen+="$config"$'\n' + awk '/^[[:space:]]*\[[[:space:]]*[Rr][Ee][Mm][Oo][Tt][Ee][[:space:]]+"origin"[[:space:]]*\][[:space:]]*([#;].*)?$/ || /^[[:space:]]*\[[[:space:]]*[Rr][Ee][Mm][Oo][Tt][Ee]\.origin[[:space:]]*\][[:space:]]*([#;].*)?$/ { found=1 } END { exit !found }' "$config" && return 0 + done < <(git -C "$worktree" config --list --show-origin 2>/dev/null || true) + return 1 +} + freshen_spawn_worktree_base() { # <worktree> local worktree=$1 default target expected actual status + status=$(git -C "$worktree" -c core.quotePath=false status --porcelain) || { + echo "error: could not inspect pooled worktree '$worktree' before refreshing its base" >&2 + return 1 + } + if [ -n "$status" ]; then + if describe_stale_submodule_pins "$worktree" "$status"; then + echo "error: pooled worktree '$worktree' has a stale submodule checkout, not uncommitted work; refusing to launch and leaving it untouched" >&2 + else + echo "error: pooled worktree '$worktree' is not clean; refusing to discard uncommitted work while refreshing its base" >&2 + fi + return 1 + fi + if ! spawn_worktree_has_origin_config "$worktree"; then + return 0 + fi if ! git -C "$worktree" fetch --quiet origin; then echo "error: could not fetch origin for pooled worktree '$worktree'; refusing to launch from a potentially stale base" >&2 return 1 @@ -2418,18 +2463,6 @@ freshen_spawn_worktree_base() { # <worktree> echo "error: '$target' is not a commit for pooled worktree '$worktree'; refusing to launch from a potentially stale base" >&2 return 1 } - status=$(git -C "$worktree" -c core.quotePath=false status --porcelain) || { - echo "error: could not inspect pooled worktree '$worktree' before refreshing its base" >&2 - return 1 - } - if [ -n "$status" ]; then - if describe_stale_submodule_pins "$worktree" "$status"; then - echo "error: pooled worktree '$worktree' has a stale submodule checkout, not uncommitted work; refusing to launch and leaving it untouched" >&2 - else - echo "error: pooled worktree '$worktree' is not clean; refusing to discard uncommitted work while refreshing its base" >&2 - fi - return 1 - fi if ! git -C "$worktree" reset --hard "$target" >/dev/null; then echo "error: could not reset pooled worktree '$worktree' to '$target'; refusing to launch from a potentially stale base" >&2 return 1 @@ -4020,6 +4053,14 @@ spawn_record_traceparent() { # process (go build, go test, ...) inherit it. Sent before the launch command so # the env is set when the agent starts; the brief sleep lets the export land. spawn_send_text_line "$T" "export GOTMPDIR=$TASK_TMP/gotmp" +# Mark the pane as a task worker so bin/fm-test-run.sh can refuse to run the +# suite in the repository's primary checkout. Ship, design, and scout workers are the +# ones assigned an isolated worktree; a secondmate runs its own home instead. +# The id reached a validated bare-slug charset above, so it carries no shell +# syntax of its own. +if [ "$KIND" = ship ] || [ "$KIND" = design ] || [ "$KIND" = scout ]; then + spawn_send_text_line "$T" "export FM_TASK_ID=$ID" +fi # Send through the exact channel that already ships GOTMPDIR, so every backend # and harness - ship, design, scout, and secondmate - gets it before launch. Skipped # entirely when trace context is off. @@ -4043,6 +4084,7 @@ if [ "$LAUNCH_ENV_ENABLED" = 1 ]; then TMPDIR TMP TEMP GOTMPDIR TMUX TMUX_PANE HERDR_ENV HERDR_SESSION HERDR_SOCKET_PATH \ HERDR_PANE_ID CMUX_WORKSPACE_ID CMUX_SURFACE_ID CMUX_TAB_ID CMUX_PANEL_ID \ CMUX_SOCKET_PATH ZELLIJ ZELLIJ_SESSION_NAME ZELLIJ_PANE_ID FM_ZELLIJ_SESSION \ + FM_TASK_ID \ $LAUNCH_ENV_NAMES; do # Only validated names enter shell syntax. Values expand once, quoted, in # the pane shell and never become source text or spawn-process snapshots. diff --git a/bin/fm-startup-memory-budget-lib.sh b/bin/fm-startup-memory-budget-lib.sh index f2c06014b8e..033bb69ba42 100644 --- a/bin/fm-startup-memory-budget-lib.sh +++ b/bin/fm-startup-memory-budget-lib.sh @@ -23,7 +23,7 @@ fm_startup_memory_budget_fail() { fm_startup_memory_budget_link_count() { if [ "$(uname)" = Darwin ]; then - stat -f %l "$1" 2>/dev/null + /usr/bin/stat -f %l "$1" 2>/dev/null else stat -c %h "$1" 2>/dev/null fi diff --git a/bin/fm-supervise-daemon.sh b/bin/fm-supervise-daemon.sh index 3b00fa8df68..5f924228171 100755 --- a/bin/fm-supervise-daemon.sh +++ b/bin/fm-supervise-daemon.sh @@ -259,7 +259,7 @@ _state_root() { printf '%s' "${FM_STATE_OVERRIDE:-$FM_HOME/state}"; } # --- portable stat (same trap as fm-watch.sh: no `stat -f || stat -c`) ------- if [ "$(uname)" = Darwin ]; then - _stat_file_mtime() { stat -f %m "$1" 2>/dev/null; } + _stat_file_mtime() { /usr/bin/stat -f %m "$1" 2>/dev/null; } else _stat_file_mtime() { stat -c %Y "$1" 2>/dev/null; } fi diff --git a/bin/fm-supervision-lib.sh b/bin/fm-supervision-lib.sh index 085fb102191..729f197340c 100644 --- a/bin/fm-supervision-lib.sh +++ b/bin/fm-supervision-lib.sh @@ -3,12 +3,10 @@ # and the tolerated-quiet window its consumers measure against. # Usage: . bin/fm-supervision-lib.sh # -# Reports whether a firstmate home needs supervision because it has in-flight -# work (a state/<id>.meta exists), an X-mode relay poll -# (state/x-watch.check.sh), a registered process-event source -# (state/procevent/*.source), or wake records still queued for a drain -# (state/.wake-queue), and whether its watcher has a fresh liveness beacon -# (state/.last-watcher-beat, touched every poll cycle, within the grace window). +# Reports whether a firstmate home needs supervision (fm_supervision_status +# below is the single owner of that condition set), and whether its watcher has +# a fresh liveness beacon (state/.last-watcher-beat, touched every poll cycle, +# within the grace window). # bin/fm-turnend-guard.sh uses the PID-strict fm_watcher_healthy from # bin/fm-wake-lib.sh for its block decision. bin/fm-guard.sh uses the model-aware # fm_watcher_supervision_verdict (also in bin/fm-wake-lib.sh), which owns what a @@ -50,7 +48,7 @@ fm_sup_busy_turn_max_seconds() { # [explicit-override] # Portable mtime; Linux stat lacks -f, macOS stat lacks -c. fm_sup_stat_mtime() { if [ "$(uname)" = Darwin ]; then - stat -f %m "$1" 2>/dev/null + /usr/bin/stat -f %m "$1" 2>/dev/null else stat -c %Y "$1" 2>/dev/null fi @@ -60,18 +58,27 @@ fm_sup_stat_mtime() { # Populates, for the state dir at $1: # FM_SUP_IN_FLIGHT count of state/*.meta (in-flight tasks) # FM_SUP_SOURCES count of registered process-to-event sources +# FM_SUP_CHECKS count of registered custom checks: a state/<id>.check.sh +# with the state/<id>.check-trust binding that +# bin/fm-check-register.sh writes. Task PR polls carry no +# such binding and are torn down with their task, and the +# relay shim keeps its own trust path, so neither counts +# here. Presence of the binding is the whole test: whether +# those bytes are still the registered ones is the check +# sweep's call at execution time, and a home whose check +# no longer validates needs the watcher precisely so the +# sweep can report the rejection instead of going quiet. # FM_SUP_NEEDED true/false - in-flight work, an X-mode relay poll, a # registered event source (a source is a wait on an # external process, not a task, so it has no metadata), -# or a pending wake queue (queued records are undelivered -# supervision work even on an otherwise idle home) +# a registered custom check, or a pending wake queue # FM_SUP_WATCHER_FRESH true/false - a watcher beacon within the grace window # FM_SUP_BEACON_DESC human-readable beacon age, for banners ("never" if absent) # FM_SUP_QUEUE_PENDING true/false - state/.wake-queue has unread records # grace-seconds defaults through fm_sup_grace_seconds above. # Always returns 0; callers read the vars, or use fm_supervision_unhealthy below. fm_supervision_status() { - local state=$1 grace meta source beat m age + local state=$1 grace meta source check id beat m age grace=$(fm_sup_grace_seconds "${2:-}") FM_SUP_IN_FLIGHT=0 FM_SUP_NEEDED=false @@ -91,9 +98,21 @@ fm_supervision_status() { # shellcheck disable=SC2034 # Read by callers (fm-guard.sh) after sourcing. [ -s "$state/.wake-queue" ] && FM_SUP_QUEUE_PENDING=true + FM_SUP_CHECKS=0 + for check in "$state"/*.check.sh; do + [ -e "$check" ] || continue + id=${check##*/} + id=${id%.check.sh} + if [ "$id" = x-watch ]; then + continue + fi + [ -e "$state/$id.check-trust" ] || continue + FM_SUP_CHECKS=$((FM_SUP_CHECKS + 1)) + done if [ "$FM_SUP_IN_FLIGHT" -gt 0 ] \ || [ -f "$state/x-watch.check.sh" ] \ || [ "$FM_SUP_SOURCES" -gt 0 ] \ + || [ "$FM_SUP_CHECKS" -gt 0 ] \ || [ "$FM_SUP_QUEUE_PENDING" = true ]; then FM_SUP_NEEDED=true fi diff --git a/bin/fm-teardown.sh b/bin/fm-teardown.sh index f5ced7d94c1..e0cfa2d618c 100755 --- a/bin/fm-teardown.sh +++ b/bin/fm-teardown.sh @@ -55,6 +55,9 @@ # The close - and only the close - is replaced by `tasks-axi reopen` with the # deliverable recorded while the backlog item is still an open captain call # (bin/fm-captain-hold.sh `open` owns that predicate), because the policy holds +# NOTE: this uses `open`'s silent default and depends only on its unchanged +# 0/1/2 exit-code contract. The optional `--identity` output that bin/fm-watch.sh +# asks for prints only on an exit 0 and changes nothing read here. # the very work item a question gates and cleanup must never retire the # captain's own question. The same pending-close record carries that intent as # `mode=retain`, so an interrupted cleanup replays the retention rather than a @@ -71,6 +74,13 @@ # already present in the up-to-date default branch. This recognizes the common # squash-merge-then-delete-branch flow, where the branch's own commits live nowhere # on a remote yet the change is fully in main. +# Squash merges collapse the branch's commits, so per-commit patch ids against main +# no longer match, and a pipeline rebase can leave the local worktree diverged from +# the PR head. A diverged copy is not treated as landed: path-set coverage, git +# cherry, and merge-tree containment each fail to prove content landed without also +# accepting unlanded edits to the same paths. Teardown still accepts a merged PR +# whose head contains the current local work (ancestor or equivalent patch ids), +# or a clean content-in-default tree match. Anything else refuses. # The PR itself is resolved from the task's recorded pr= when present, or - when # no pr= was ever recorded (e.g. a yolo-authorized merge on a repo with no PR CI, # where the usual "checks green" fm-pr-check.sh trigger never fires) - by looking @@ -115,8 +125,11 @@ # before cleanup. Its current working directory is only incidental process # state: the same worker remains the owner after changing directory, so cwd can # never veto teardown of that exact recorded endpoint. -# The scan and destructive return hold a project-identity lock in the root -# Firstmate home's state directory. Fresh Treehouse spawns for that project in +# The scan and destructive return hold a project-identity lock in the local root +# Firstmate home's state directory, as resolved by bin/fm-wake-lib.sh's +# fm_firstmate_root_home; a home seeded from another machine is its own local +# root, since a lock on this filesystem cannot be held or observed across that +# boundary. Fresh Treehouse spawns for that project in # every local Firstmate home hold the same lock from before slot allocation # through metadata publication, closing the publication # gap; forced secondmate teardown takes it and runs the same checks for every diff --git a/bin/fm-test-isolation-proof.sh b/bin/fm-test-isolation-proof.sh index e0f26db4f8f..3100e746f74 100755 --- a/bin/fm-test-isolation-proof.sh +++ b/bin/fm-test-isolation-proof.sh @@ -217,8 +217,8 @@ EOF dir_mode() { local path=$1 - if stat -f %Lp "$path" >/dev/null 2>&1; then - stat -f %Lp "$path" + if /usr/bin/stat -f %Lp "$path" >/dev/null 2>&1; then + /usr/bin/stat -f %Lp "$path" else stat -c %a "$path" fi diff --git a/bin/fm-test-run.sh b/bin/fm-test-run.sh index 4273394f781..f746a8225ae 100755 --- a/bin/fm-test-run.sh +++ b/bin/fm-test-run.sh @@ -28,7 +28,11 @@ # fm-test-run.sh --aggregate-json <out.json> <lane.json> [more lane.json...] # # Options: -# --json <path> write a deterministic timing artifact after the run +# --json <path> write a deterministic timing artifact after the run. Each +# script record carries its family, expected gate-skip class, +# exit, duration, whether it gate-skipped, and the reason it +# gave (empty when it ran), so a lane can say which harness or +# tool this host could not exercise. # --list print selected script paths (one per line) and exit 0 # --list-scheduled # print selected paths longest-hint-first and exit 0 @@ -93,11 +97,26 @@ # FM_TEST_SLOWEST rank=<k> script=<path> duration_ms=<n> # FM_TEST_BUDGET max_wall_ms=<n> duration_ms=<n> (only with --max-wall-ms) # +# Placement refusal: +# A task worker is assigned an isolated worktree, and that placement is +# checked only when its task starts. When FM_TASK_ID marks such a worker and +# this runner resolves to the repository's PRIMARY checkout, every executing +# mode refuses before selecting a suite: the suite creates and switches +# branches, and the primary is the checkout every linked worktree resolves +# against. Inspection modes execute nothing and stay available, and a run with +# no FM_TASK_ID set is unchanged. +# # Exit status is non-zero if any selected script exits non-zero, a configured # --fail-on-gate-skip token appears, the measured duration exceeds # --max-wall-ms, timing-artifact finalization fails, or a concurrent worker # violates its isolation check. Other gate skips (first meaningful line -# matching ^skip:) remain successful and are counted as skipped_gate. +# matching ^skip:) remain successful and are counted as skipped_gate; each one +# is logged with its reason and recorded in the timing artifact. +# +# expected_gate_skip classes name why a family is allowed to skip: herdr (the +# pinned real-Herdr lane), optional-binary (a backend whose binary is optional), +# live-capability (a live-harness guard governed by fm_live_gate, which records +# unavailable tools and explicit policy skips; see tests/lib.sh), or none. # # Family labels, the changed-file map, and production portable-shard composition # live in this script only (one owner). The proven-isolated candidate set remains @@ -154,26 +173,26 @@ PER_SCRIPT_TIMEOUT_SECS=0 # than indistinguishable from an unset bound. PER_SCRIPT_TIMEOUT_SET= # Bound applied automatically to every standard-mode sweep. What settles it is -# the family WALL, not the average. The real-Herdr family is 12 scripts inside +# the family WALL, not the average. The original real-Herdr sample had 12 scripts inside # a ~7 min healthy wall (.github/workflows/ci.yml), so its mean slot is ~35s. # That wall is a measurement, not a ceiling: the only bound that lane enforces # is the family-run step's 1200s, and the family's slowest measured script is # the 341s presentation E2E. So 480s is 1.41x over that script - materially -# thinner than the 2x+ the portable lanes get over their 210s +# thinner than the original 2x+ portable margin over the measured 210s # tests/fm-watch-triage.test.sh worst case, and a Herdr E2E running 1.41x # slower than measured turns red as exit=124 on a required lane. That margin # is accepted rather than widened (HelloWorldSungin/firstmate#256). # # The arithmetic is replacement, not addition: a hung script spends the bound # INSTEAD of its own healthy slot. The current portable-serial hint table totals -# 4168699ms over eight shards, with the slowest at 525472ms. Replacing its ~25s -# average script with 480s puts script time near 981s against the adopted 1200s -# CI cap, before checkout and bootstrap. The bound therefore has room to report -# an attributed exit=124 under the estimates, while setup and runner variability -# still prevent a guarantee about the enclosing job's total wall time. +# about 6188s over eight shards, with the slowest near 774s. Replacing its ~31s +# average script with 480s puts script time near 1223s, beyond the 1200s CI cap +# before checkout and bootstrap. A hung script can therefore lose per-script +# attribution to the enclosing job cancellation; healthy estimates retain about +# seven minutes for setup and runner variability. Neither deadline is widened. # real-Herdr is 420s less its ~35s mean slot plus 480s, inside its 1200s step cap. -# portable-parallel is tighter: CI runs that lane serially, so ~134s less its -# ~12s mean slot plus 480s is ~602s, already past its 600s job cap before setup. +# portable-parallel is tighter: CI runs that lane serially, so ~422s less its +# ~38s mean slot plus 480s is ~864s, already past its 600s job cap before setup. # That lane can therefore lose per-script attribution to a job cancellation. # Raising the script bound would not remedy that enclosing-job limit. # @@ -232,6 +251,29 @@ now_iso() { date -u +%Y-%m-%dT%H:%M:%SZ } +# Enforce the placement refusal described in this script's header. +# +# The primary checkout is the working tree whose own git dir IS the repository's +# common git dir; every linked worktree has a git dir under it instead. That is +# the same predicate bin/fm-spawn.sh uses to keep a launch out of the primary, +# and unlike comparing top-level paths it still holds when the primary is +# reached through a different path. When git resolves neither directory - a +# non-repository fixture, a detached copy - nothing proves this is the primary, +# so the run proceeds. +refuse_primary_checkout_for_task() { + local task_id git_dir common_dir top + task_id=${FM_TASK_ID:-} + [ -n "$task_id" ] || return 0 + git_dir=$(git -C "$ROOT" rev-parse --absolute-git-dir 2>/dev/null) \ + && git_dir=$(cd "$git_dir" 2>/dev/null && pwd -P) || git_dir= + common_dir=$(git -C "$ROOT" rev-parse --path-format=absolute --git-common-dir 2>/dev/null) \ + && common_dir=$(cd "$common_dir" 2>/dev/null && pwd -P) || common_dir= + [ -n "$git_dir" ] && [ -n "$common_dir" ] || return 0 + [ "$git_dir" = "$common_dir" ] || return 0 + top=$(cd "$ROOT" && pwd -P) + die "refusing to run in the repository primary checkout $top while FM_TASK_ID=$task_id is set; run from the assigned task worktree instead" +} + cpu_count() { local n n=$(getconf _NPROCESSORS_ONLN 2>/dev/null || sysctl -n hw.ncpu 2>/dev/null || echo 1) @@ -283,7 +325,7 @@ family_for_basename() { fm-wake-drain-unread-status.test.sh|\ fm-tool-update-check.test.sh|\ fm-wake-queue.test.sh|fm-watch-arm.test.sh|fm-watch-checkpoint.test.sh|fm-watch-recovery-loop.test.sh|\ - fm-watch-triage.test.sh|fm-task-inbox.test.sh|\ + fm-watch-triage.test.sh|fm-watch-triage-waits.test.sh|fm-task-inbox.test.sh|\ fm-watcher-lock.test.sh|fm-inactive-reconcile.test.sh) printf '%s\n' watcher-wake-lock ;; @@ -323,6 +365,7 @@ family_for_basename() { printf '%s\n' session-bootstrap ;; fm-afk-pi-dual-supervision-e2e.test.sh|fm-afk-pi-herdr-return-e2e.test.sh|\ + fm-bearings-board-lavish-live-e2e.test.sh|\ fm-claude-stop-autoarm-live-e2e.test.sh|\ fm-composer-matrix-live-e2e.test.sh|\ fm-codex-continuity-live-e2e.test.sh|fm-grok-continuity-live-e2e.test.sh|\ @@ -385,6 +428,7 @@ family_for_basename() { fm-no-mistakes-required.test.sh|fm-peek-remote.test.sh|\ fm-pending-reply.test.sh|fm-pi-branch-extension.test.sh|\ fm-procevent-quota.test.sh|fm-procevent-when.test.sh|fm-procevent.test.sh|\ + fm-live-gate.test.sh|\ fm-project-origin.test.sh|fm-public-followup.test.sh|fm-quota-choose.test.sh|\ fm-remote-entrypoint.test.sh|fm-remote-secondmate-parent-binding.test.sh|\ fm-send-remote-delivery.test.sh|fm-spawn-pool-base-freshen.test.sh|\ @@ -402,7 +446,7 @@ family_for_basename() { expected_gate_skip_for_family() { case "$1" in real-herdr-gated) printf '%s\n' herdr ;; - live-harness-optin) printf '%s\n' optin-env ;; + live-harness-optin) printf '%s\n' live-capability ;; cmux|zellij|orca) printf '%s\n' optional-binary ;; snapshot-bearings) printf '%s\n' optional-binary ;; *) printf '%s\n' none ;; @@ -475,40 +519,40 @@ EOF } # Portable parallel shard 1: LPT balance of the proven-isolated set using the -# current concurrent-proof durations in docs/fm-test-isolation-proof.json. +# measured CI durations recorded in docs/fm-test-portable-shards.md. # Execution order is longest first so wall-clock stays near the balanced sum. list_portable_parallel_1() { cat <<'EOF' -tests/fm-x-mode.test.sh -tests/fm-cd-pretool-check.test.sh tests/fm-captain-hold-lifecycle.test.sh -tests/fm-test-run.test.sh -tests/fm-composer-ghost.test.sh +tests/fm-pr-merge.test.sh +tests/fm-crew-state.test.sh +tests/fm-backend-herdr.test.sh +tests/fm-cd-pretool-check.test.sh tests/fm-grok-harness.test.sh -tests/fm-lint.test.sh +tests/fm-herdr-lab.test.sh tests/fm-pi-primary-types.test.sh tests/fm-review-diff.test.sh -tests/fm-brief.test.sh -tests/fm-transition-lib.test.sh +tests/fm-composer-ghost.test.sh +tests/fm-send-settle.test.sh EOF } # Portable parallel shard 2: the complementary LPT half of the proven set. list_portable_parallel_2() { cat <<'EOF' -tests/fm-backend-herdr.test.sh +tests/fm-lint.test.sh +tests/fm-test-run.test.sh +tests/fm-x-mode.test.sh tests/fm-arm-pretool-check.test.sh -tests/fm-crew-state.test.sh -tests/fm-herdr-lab.test.sh -tests/fm-pr-merge.test.sh +tests/fm-brief.test.sh +tests/fm-send-strict.test.sh tests/fm-send-popup-settle.test.sh +tests/fm-composer-lib.test.sh tests/fm-tmux-submit-busy.test.sh -tests/fm-send-settle.test.sh -tests/fm-send-strict.test.sh tests/fm-spawn-batch.test.sh -tests/fm-supervision-instructions.test.sh tests/fm-ensure-agents-md.test.sh -tests/fm-composer-lib.test.sh +tests/fm-supervision-instructions.test.sh +tests/fm-transition-lib.test.sh EOF } @@ -571,8 +615,8 @@ is_proven_isolated_script() { # The portable serial remainder: every tests/*.test.sh that is neither # proven-isolated nor real-herdr-gated. Watcher, lock, AFK, real tmux, daemon, -# secondmate lifecycle, bootstrap, live-harness opt-in, GUI-backend, and other -# unproven work stays here. Derived rather than enumerated so a newly added test +# secondmate lifecycle, bootstrap, the live-harness-optin family, GUI-backend, +# and other unproven work stays here. Derived rather than enumerated so a newly added test # lands here by default instead of falling out of every lane. list_portable_serial() { local s base fam @@ -580,11 +624,10 @@ list_portable_serial() { [ -n "$s" ] || continue base=$(basename "$s") fam=$(family_for_basename "$base") - # real-herdr-gated has its own required CI lane; live-harness-optin needs - # machine state CI does not have - real harness credentials, or a real - # browser session - so it stays an explicit opt-in outside portable lanes. + # Real Herdr has its own required lane. Live guards stay in serial CI, + # where each guard owns its capability checks and execution opt-ins. case "$fam" in - real-herdr-gated|live-harness-optin) continue ;; + real-herdr-gated) continue ;; esac if is_proven_isolated_script "$s"; then continue @@ -595,188 +638,212 @@ list_portable_serial() { # Measured portable-serial script durations in milliseconds, from the CI timing # artifacts recorded in docs/fm-test-portable-shards.md. Each value is the -# slowest of several green runs, so the balance holds on a slow runner rather -# than only on the fastest one measured. These are balance hints only: the shard +# maximum of retained earlier hints and completed passing CI measurements, +# including completed scripts from an interrupted lane as documented there. These are balance hints only: the shard # partition stays complete and disjoint whatever they say, so a stale hint costs # balance rather than coverage. That doc owns the refresh procedure. portable_serial_weight_hints() { cat <<'EOF' tests/fm-afk-inject-e2e.test.sh 35792 +tests/fm-afk-pi-dual-supervision-e2e.test.sh 109 tests/fm-afk-pi-herdr-return-e2e.test.sh 100 -tests/fm-afk-return.test.sh 1837 -tests/fm-agents-hard-rules.test.sh 317 +tests/fm-afk-return.test.sh 1898 +tests/fm-agents-hard-rules.test.sh 422 +tests/fm-agy-adapter.test.sh 25751 +tests/fm-agy-smoke.test.sh 136 tests/fm-agy-trust-lib.test.sh 15755 -tests/fm-ask-user-authority.test.sh 128 -tests/fm-backend-cmux-smoke.test.sh 33 -tests/fm-backend-cmux.test.sh 3657 -tests/fm-backend-orca.test.sh 19253 +tests/fm-ask-user-authority.test.sh 223 +tests/fm-backend-cmux-smoke.test.sh 84 +tests/fm-backend-cmux.test.sh 4170 +tests/fm-backend-orca.test.sh 55501 tests/fm-backend-tmux-smoke.test.sh 393 -tests/fm-backend-zellij-smoke.test.sh 23 -tests/fm-backend-zellij.test.sh 9418 -tests/fm-backend.test.sh 20061 -tests/fm-backlog-atomicity.test.sh 122256 -tests/fm-backlog-handoff.test.sh 52291 -tests/fm-bearings-board-render.test.sh 1528 -tests/fm-bearings-board.test.sh 4195 -tests/fm-bearings-snapshot.test.sh 79954 -tests/fm-bootstrap-network-parallel.test.sh 8214 -tests/fm-bootstrap.test.sh 25208 -tests/fm-branch-supervision.test.sh 5729 -tests/fm-brief-repo-lib.test.sh 222 -tests/fm-busy-adapter-wiring.test.sh 17873 -tests/fm-busy-state.test.sh 2926 -tests/fm-calm-pi-extension.test.sh 256 -tests/fm-check-unregister.test.sh 481 +tests/fm-backend-zellij-smoke.test.sh 55 +tests/fm-backend-zellij.test.sh 10313 +tests/fm-backend.test.sh 22748 +tests/fm-backlog-atomicity.test.sh 154917 +tests/fm-backlog-handoff.test.sh 54825 +tests/fm-bearings-board-lavish-live-e2e.test.sh 117 +tests/fm-bearings-board-render.test.sh 17024 +tests/fm-bearings-board.test.sh 60332 +tests/fm-bearings-snapshot.test.sh 129537 +tests/fm-bootstrap-network-parallel.test.sh 20354 +tests/fm-bootstrap.test.sh 75613 +tests/fm-branch-supervision.test.sh 9222 +tests/fm-brief-repo-lib.test.sh 227 +tests/fm-busy-adapter-wiring.test.sh 29970 +tests/fm-busy-state.test.sh 3171 +tests/fm-calm-pi-extension.test.sh 52972 +tests/fm-check-unregister.test.sh 565 tests/fm-classify-corr-token.test.sh 38742 tests/fm-classify-decision-key.test.sh 1167 -tests/fm-claude-stop-autoarm-live-e2e.test.sh 21 -tests/fm-claude-stop-autoarm.test.sh 60709 -tests/fm-cmux-claude-composer-live-e2e.test.sh 23 -tests/fm-codex-continuity-live-e2e.test.sh 21 -tests/fm-composer-matrix-live-e2e.test.sh 23 -tests/fm-control-relaunch.test.sh 48210 -tests/fm-control.test.sh 37798 -tests/fm-cursor-harness.test.sh 30103 -tests/fm-cursor-primary-live-e2e.test.sh 21 -tests/fm-cursor-primary.test.sh 54947 -tests/fm-daemon.test.sh 26870 -tests/fm-dashboard-access.test.sh 21023 -tests/fm-dashboard-backlog.test.sh 387 -tests/fm-dashboard-events.test.sh 29505 -tests/fm-dashboard-gbrain-ui.test.sh 102 -tests/fm-dashboard-gbrain.test.sh 21374 -tests/fm-dashboard-history.test.sh 138 -tests/fm-dashboard-inbox.test.sh 116 -tests/fm-dashboard-router.test.sh 68 -tests/fm-dashboard.test.sh 15336 -tests/fm-design-skills.test.sh 159 -tests/fm-documentation-audiences.test.sh 732 -tests/fm-endpoint-binding-migrate.test.sh 1366 -tests/fm-extension-binding.test.sh 7398 -tests/fm-fleet-snapshot-view.test.sh 8547 +tests/fm-claude-stop-autoarm-live-e2e.test.sh 117 +tests/fm-claude-stop-autoarm.test.sh 60725 +tests/fm-claude-trust.test.sh 6217 +tests/fm-cmux-claude-composer-live-e2e.test.sh 92 +tests/fm-codex-continuity-live-e2e.test.sh 108 +tests/fm-composer-matrix-live-e2e.test.sh 114 +tests/fm-control-relaunch.test.sh 72543 +tests/fm-control.test.sh 38952 +tests/fm-cursor-harness.test.sh 30129 +tests/fm-cursor-primary-live-e2e.test.sh 131 +tests/fm-cursor-primary.test.sh 55508 +tests/fm-daemon.test.sh 28891 +tests/fm-dashboard-access.test.sh 21328 +tests/fm-dashboard-backlog.test.sh 779 +tests/fm-dashboard-browser.test.sh 64 +tests/fm-dashboard-events.test.sh 34769 +tests/fm-dashboard-gbrain-ui.test.sh 205 +tests/fm-dashboard-gbrain.test.sh 22293 +tests/fm-dashboard-history.test.sh 254 +tests/fm-dashboard-inbox.test.sh 159 +tests/fm-dashboard-router.test.sh 162 +tests/fm-dashboard-usage.test.sh 154 +tests/fm-dashboard.test.sh 17109 +tests/fm-design-skills.test.sh 14989 +tests/fm-documentation-audiences.test.sh 1032 +tests/fm-endpoint-binding-migrate.test.sh 1496 +tests/fm-extension-binding.test.sh 9599 +tests/fm-fleet-snapshot-view.test.sh 31441 tests/fm-fleet-sync.test.sh 37749 -tests/fm-gate-refuse.test.sh 4977 -tests/fm-gbrain-capture.test.sh 23044 -tests/fm-gbrain-eval.test.sh 5354 -tests/fm-gbrain-health.test.sh 6623 -tests/fm-gbrain-lib.test.sh 17072 -tests/fm-gitignore-config.test.sh 62 -tests/fm-gotmp.test.sh 1310 -tests/fm-grok-continuity-live-e2e.test.sh 20 -tests/fm-grok-stop-live-e2e.test.sh 21 -tests/fm-guard-stale-banner.test.sh 11218 -tests/fm-harness-adapter-instructions-live-e2e.test.sh 20 -tests/fm-harness-adapter-references.test.sh 55 -tests/fm-harness-liveness-drift-live-e2e.test.sh 21 -tests/fm-herdr-session-cleanup.test.sh 6704 -tests/fm-herdr-submit-confirm-live-e2e.test.sh 23 -tests/fm-herdr-version-floor-live-e2e.test.sh 23 -tests/fm-home-summary-refresh.test.sh 34793 -tests/fm-inactive-reconcile.test.sh 41826 -tests/fm-issue-linkage.test.sh 12629 -tests/fm-issue-writeback.test.sh 28984 -tests/fm-kimi-harness.test.sh 18015 -tests/fm-launch-lib.test.sh 425 +tests/fm-gate-refuse.test.sh 6155 +tests/fm-gbrain-capture-e2e.test.sh 56 +tests/fm-gbrain-capture.test.sh 42240 +tests/fm-gbrain-doc-commands.test.sh 784 +tests/fm-gbrain-eval.test.sh 7945 +tests/fm-gbrain-health.test.sh 7646 +tests/fm-gbrain-lib.test.sh 17089 +tests/fm-gbrain-pin-check.test.sh 478 +tests/fm-gbrain-readonly-e2e.test.sh 161 +tests/fm-gemini-harness.test.sh 745 +tests/fm-gitignore-config.test.sh 112 +tests/fm-gotmp.test.sh 2157 +tests/fm-grok-continuity-live-e2e.test.sh 159 +tests/fm-grok-stop-live-e2e.test.sh 92 +tests/fm-guard-stale-banner.test.sh 13176 +tests/fm-harness-adapter-instructions-live-e2e.test.sh 107 +tests/fm-harness-adapter-references.test.sh 118 +tests/fm-harness-liveness-drift-live-e2e.test.sh 856 +tests/fm-herdr-session-cleanup.test.sh 6858 +tests/fm-herdr-submit-confirm-live-e2e.test.sh 118 +tests/fm-herdr-version-floor-live-e2e.test.sh 93 +tests/fm-home-summary-refresh.test.sh 39231 +tests/fm-inactive-reconcile.test.sh 48508 +tests/fm-issue-linkage.test.sh 17659 +tests/fm-issue-writeback.test.sh 29552 +tests/fm-kimi-harness.test.sh 20521 +tests/fm-launch-lib.test.sh 490 tests/fm-lint-workflows.test.sh 855 -tests/fm-merge-local.test.sh 509 -tests/fm-model-verify.test.sh 4042 +tests/fm-live-gate.test.sh 6000 +tests/fm-merge-local.test.sh 591 +tests/fm-model-verify.test.sh 4073 tests/fm-muse-harness.test.sh 55572 -tests/fm-muse-signals-live-e2e.test.sh 23 -tests/fm-no-mistakes-required-gate.test.sh 91 +tests/fm-muse-signals-live-e2e.test.sh 90 +tests/fm-nm-test-contract.test.sh 812 +tests/fm-no-mistakes-required-gate.test.sh 4578 tests/fm-no-mistakes-required.test.sh 370 -tests/fm-on.test.sh 11692 -tests/fm-opencode-primary-live-e2e.test.sh 21 -tests/fm-operational-input.test.sh 231 -tests/fm-outcome-manifest.test.sh 4788 +tests/fm-omp-harness.test.sh 20362 +tests/fm-omp-primary-live-e2e.test.sh 187 +tests/fm-on.test.sh 12020 +tests/fm-opencode-primary-live-e2e.test.sh 112 +tests/fm-operational-input.test.sh 308 +tests/fm-outcome-manifest.test.sh 8021 tests/fm-peek-remote.test.sh 1018 -tests/fm-pending-reply.test.sh 24679 -tests/fm-pi-branch-extension.test.sh 22239 -tests/fm-pi-branch-live-e2e.test.sh 56 -tests/fm-pi-branch-responsiveness-live-e2e.test.sh 21 -tests/fm-pi-hung-delivery-herdr-e2e.test.sh 23 -tests/fm-pi-primary-live-e2e.test.sh 20 -tests/fm-pi-watch-extension.test.sh 42970 +tests/fm-pending-reply.test.sh 30068 +tests/fm-pi-branch-extension.test.sh 176128 +tests/fm-pi-branch-live-e2e.test.sh 163 +tests/fm-pi-branch-responsiveness-live-e2e.test.sh 13245 +tests/fm-pi-hung-delivery-herdr-e2e.test.sh 163 +tests/fm-pi-primary-live-e2e.test.sh 132 +tests/fm-pi-watch-extension.test.sh 62408 +tests/fm-pi-windows-shell-invocation.test.sh 5121 tests/fm-pointer-check.test.sh 1314 -tests/fm-pr-check-security.test.sh 160475 +tests/fm-pr-check-security.test.sh 171211 tests/fm-pr-status.test.sh 420 tests/fm-procevent-quota.test.sh 1949 -tests/fm-procevent-when.test.sh 17392 -tests/fm-procevent.test.sh 69715 -tests/fm-project-origin.test.sh 137 +tests/fm-procevent-when.test.sh 17550 +tests/fm-procevent.test.sh 188540 +tests/fm-project-origin.test.sh 229 tests/fm-public-followup.test.sh 196745 -tests/fm-quota-array-dispatch-live-e2e.test.sh 21 -tests/fm-quota-choose.test.sh 1461 -tests/fm-quota-sidecar.test.sh 549 -tests/fm-recall.test.sh 10856 -tests/fm-remote-backlog-handoff.test.sh 41432 -tests/fm-remote-doctor.test.sh 5198 -tests/fm-remote-entrypoint.test.sh 132 +tests/fm-quota-array-dispatch-live-e2e.test.sh 119 +tests/fm-quota-choose.test.sh 1616 +tests/fm-quota-sidecar.test.sh 583 +tests/fm-recall.test.sh 60802 +tests/fm-remote-backlog-handoff.test.sh 109295 +tests/fm-remote-doctor.test.sh 5335 +tests/fm-remote-entrypoint.test.sh 167 tests/fm-remote-job-orphan-reap.test.sh 2972 -tests/fm-remote-job.test.sh 59603 +tests/fm-remote-job-worker-leak.test.sh 3854 +tests/fm-remote-job.test.sh 97985 tests/fm-remote-reply.test.sh 101690 -tests/fm-remote-secondmate-lifecycle-e2e.test.sh 209631 -tests/fm-remote-secondmate-parent-binding.test.sh 29562 -tests/fm-remote-secondmate-trace-context.test.sh 67096 +tests/fm-remote-secondmate-lifecycle-e2e.test.sh 292054 +tests/fm-remote-secondmate-parent-binding.test.sh 103889 +tests/fm-remote-secondmate-trace-context.test.sh 69915 tests/fm-remote-transport-lanes.test.sh 63140 +tests/fm-rovo-harness.test.sh 14983 +tests/fm-rovo-signals-live-e2e.test.sh 113 tests/fm-run-attribution-legacy-transition.test.sh 3737 -tests/fm-run-progress.test.sh 2041 -tests/fm-secondmate-harness.test.sh 151589 +tests/fm-run-progress.test.sh 2267 +tests/fm-secondmate-harness.test.sh 157114 tests/fm-secondmate-lifecycle-e2e.test.sh 8793 tests/fm-secondmate-liveness.test.sh 18146 -tests/fm-secondmate-reconcile.test.sh 62726 -tests/fm-secondmate-safety.test.sh 57689 -tests/fm-secondmate-sync.test.sh 17183 -tests/fm-send-inbox-doorbell-live-e2e.test.sh 22 -tests/fm-send-inbox.test.sh 38956 +tests/fm-secondmate-reconcile.test.sh 98643 +tests/fm-secondmate-restart.test.sh 108981 +tests/fm-secondmate-safety.test.sh 65616 +tests/fm-secondmate-sync.test.sh 56818 +tests/fm-send-inbox-doorbell-live-e2e.test.sh 113 +tests/fm-send-inbox.test.sh 39089 tests/fm-send-remote-delivery.test.sh 27686 -tests/fm-send-resolve-key.test.sh 19619 -tests/fm-send-secondmate-marker-herdr-e2e.test.sh 51 +tests/fm-send-resolve-key.test.sh 22398 +tests/fm-send-secondmate-marker-herdr-e2e.test.sh 138 tests/fm-send-secondmate-marker.test.sh 6252 -tests/fm-session-lock-ancestry.test.sh 1414 -tests/fm-session-start.test.sh 156952 -tests/fm-sessionstart-hook-live-e2e.test.sh 20 -tests/fm-sessionstart-instruction-refresh-live-e2e.test.sh 22 +tests/fm-session-lock-ancestry.test.sh 1526 +tests/fm-session-start.test.sh 172389 +tests/fm-sessionstart-hook-live-e2e.test.sh 93 +tests/fm-sessionstart-instruction-refresh-live-e2e.test.sh 117 tests/fm-sessionstart-nudge.test.sh 66194 -tests/fm-shared-captain-inheritance.test.sh 6108 -tests/fm-spawn-dispatch-profile.test.sh 63996 -tests/fm-spawn-pool-base-freshen.test.sh 34920 -tests/fm-spawn-worktree-settle.test.sh 5687 -tests/fm-startup-memory-budget.test.sh 6964 -tests/fm-startup-network.test.sh 54700 -tests/fm-stow-cascade.test.sh 3101 -tests/fm-subagent-pretool-check.test.sh 1030 -tests/fm-supervision-events.test.sh 719 +tests/fm-shared-captain-inheritance.test.sh 6417 +tests/fm-spawn-dispatch-profile.test.sh 155959 +tests/fm-spawn-pool-base-freshen.test.sh 59669 +tests/fm-spawn-worktree-settle.test.sh 16609 +tests/fm-startup-memory-budget.test.sh 8055 +tests/fm-startup-network.test.sh 64840 +tests/fm-stat-shadowing.test.sh 113 +tests/fm-stow-cascade.test.sh 3107 +tests/fm-stow-horizon-live-e2e.test.sh 62 +tests/fm-subagent-pretool-check.test.sh 1209 +tests/fm-supervision-events.test.sh 727 tests/fm-tangle-guard.test.sh 9662 -tests/fm-task-delivery.test.sh 5952 -tests/fm-task-inbox.test.sh 25369 -tests/fm-teardown-endpoint-safety.test.sh 4620 -tests/fm-teardown.test.sh 97603 -tests/fm-test-fixture-cleanup.test.sh 915 -tests/fm-test-fixtures.test.sh 151 +tests/fm-task-delivery.test.sh 42817 +tests/fm-task-inbox.test.sh 28841 +tests/fm-teardown-endpoint-safety.test.sh 32145 +tests/fm-teardown.test.sh 212537 +tests/fm-test-fixture-cleanup.test.sh 2200 +tests/fm-test-fixtures.test.sh 180 tests/fm-test-isolation-proof.test.sh 2567 -tests/fm-tmux-agent-liveness.test.sh 1516 -tests/fm-tool-status.test.sh 880 -tests/fm-tool-update-check.test.sh 14176 -tests/fm-trace-context-lib.test.sh 209 -tests/fm-trace-context-spawn.test.sh 44702 +tests/fm-tmux-agent-liveness.test.sh 2237 +tests/fm-tool-status.test.sh 1357 +tests/fm-tool-update-check.test.sh 30610 +tests/fm-trace-context-lib.test.sh 231 +tests/fm-trace-context-spawn.test.sh 52375 tests/fm-trigger-validation.test.sh 957 tests/fm-turnend-guard.test.sh 42565 -tests/fm-update.test.sh 5212 -tests/fm-upstream-status.test.sh 4058 -tests/fm-usage.test.sh 4988 -tests/fm-vault-drift.test.sh 4973 +tests/fm-update.test.sh 9928 +tests/fm-upstream-status.test.sh 7135 +tests/fm-usage.test.sh 11934 +tests/fm-vault-drift.test.sh 5079 tests/fm-vendor-auth-probe.test.sh 43316 -tests/fm-voice-relay.test.sh 28699 -tests/fm-wake-daemon-lifecycle-e2e.test.sh 7381 -tests/fm-wake-drain-open-decisions-cursor.test.sh 20629 -tests/fm-wake-drain-open-decisions.test.sh 6240 -tests/fm-wake-drain-outcome-backstop.test.sh 15182 +tests/fm-voice-relay.test.sh 28850 +tests/fm-wake-daemon-lifecycle-e2e.test.sh 8131 +tests/fm-wake-drain-open-decisions-cursor.test.sh 23323 +tests/fm-wake-drain-open-decisions.test.sh 8090 +tests/fm-wake-drain-outcome-backstop.test.sh 45680 tests/fm-wake-drain-unread-status.test.sh 35078 -tests/fm-wake-queue.test.sh 56674 -tests/fm-watch-arm.test.sh 58528 -tests/fm-watch-checkpoint.test.sh 5779 +tests/fm-wake-queue.test.sh 92559 +tests/fm-watch-arm.test.sh 70359 +tests/fm-watch-checkpoint.test.sh 6383 tests/fm-watch-recovery-loop.test.sh 58731 -tests/fm-watch-triage.test.sh 262626 +tests/fm-watch-triage-waits.test.sh 292716 +tests/fm-watch-triage.test.sh 236467 tests/fm-watcher-lock.test.sh 88554 EOF } @@ -976,9 +1043,6 @@ run_coverage_guard() { SCRIPTS=() select_family real-herdr-gated printf '%s\n' "${SCRIPTS[@]+"${SCRIPTS[@]}"}" | LC_ALL=C sort -u >"$tmp/herdr" - SCRIPTS=() - select_family live-harness-optin - printf '%s\n' "${SCRIPTS[@]+"${SCRIPTS[@]}"}" | LC_ALL=C sort -u >"$tmp/optin" SCRIPTS=("${saved_scripts[@]+"${saved_scripts[@]}"}") # Every serial script runs in exactly one CI shard: no duplicate work across @@ -1002,8 +1066,7 @@ run_coverage_guard() { fi for pair in \ - "shards_union:serial" "shards_union:herdr" "shards_union:optin" \ - "serial:herdr" "serial:optin" "herdr:optin"; do + "shards_union:serial" "shards_union:herdr" "serial:herdr"; do a=${pair%%:*} b=${pair#*:} LC_ALL=C comm -12 "$tmp/$a" "$tmp/$b" >"$tmp/overlap" @@ -1015,7 +1078,7 @@ run_coverage_guard() { fi done - cat "$tmp/shards_union" "$tmp/serial" "$tmp/herdr" "$tmp/optin" | LC_ALL=C sort >"$tmp/union_raw" + cat "$tmp/shards_union" "$tmp/serial" "$tmp/herdr" | LC_ALL=C sort >"$tmp/union_raw" uniq -d "$tmp/union_raw" >"$tmp/union_dups" if [ -s "$tmp/union_dups" ]; then log "coverage guard: duplicate scripts across lanes:" @@ -1027,7 +1090,7 @@ run_coverage_guard() { missing=$(LC_ALL=C comm -23 "$tmp/all" "$tmp/union" || true) extra=$(LC_ALL=C comm -13 "$tmp/all" "$tmp/union" || true) if [ -n "$missing" ] || [ -n "$extra" ]; then - log "coverage guard: union of portable shards + portable serial + Herdr + live opt-in must equal tests/*.test.sh" + log "coverage guard: union of portable shards + portable serial + Herdr must equal tests/*.test.sh" [ -z "$missing" ] || { log "missing from union:"; printf '%s\n' "$missing" >&2; } [ -z "$extra" ] || { log "extra beyond inventory:"; printf '%s\n' "$extra" >&2; } rm -rf "$tmp" @@ -1061,14 +1124,13 @@ run_coverage_guard() { fi fi - printf 'FM_TEST_COVERAGE ok total=%s parallel=%s serial=%s serial_shards=%s serial_unhinted=%s herdr=%s optin=%s\n' \ + printf 'FM_TEST_COVERAGE ok total=%s parallel=%s serial=%s serial_shards=%s serial_unhinted=%s herdr=%s\n' \ "$(wc -l <"$tmp/all" | tr -d ' ')" \ "$(wc -l <"$tmp/shards_union" | tr -d ' ')" \ "$(wc -l <"$tmp/serial" | tr -d ' ')" \ "$PORTABLE_SERIAL_SHARDS" \ "$unhinted" \ - "$(wc -l <"$tmp/herdr" | tr -d ' ')" \ - "$(wc -l <"$tmp/optin" | tr -d ' ')" + "$(wc -l <"$tmp/herdr" | tr -d ' ')" rm -rf "$tmp" return 0 } @@ -1423,6 +1485,7 @@ families_for_changed_path() { .pi/extensions/lib/fm-operational-input.ts) # The same rule for the operational-input library, whose reach is wider: # every Pi or OMP extension that classifies or encodes operational text. + printf '%s\n' __script__:fm-pi-windows-shell-invocation.test.sh printf '%s\n' __script__:fm-pi-branch-extension.test.sh printf '%s\n' __script__:fm-pi-watch-extension.test.sh printf '%s\n' __script__:fm-calm-pi-extension.test.sh @@ -1437,6 +1500,7 @@ families_for_changed_path() { .pi/extensions/fm-primary-turnend-guard.ts) # The run tier's two harness-supplied facts (source vocabulary and # context-reset stdout injection) only show up against a real harness. + printf '%s\n' __script__:fm-pi-windows-shell-invocation.test.sh printf '%s\n' session-bootstrap printf '%s\n' live-harness-optin ;; @@ -1540,7 +1604,7 @@ families_for_changed_path() { docs/fm-test-isolation-proof.json) printf '%s\n' pure-contract-unit ;; - .github/*|.tasks.toml|AGENTS.md|CLAUDE.md|CONTRIBUTING.md|\ + .github/*|.gitattributes|.tasks.toml|AGENTS.md|CLAUDE.md|CONTRIBUTING.md|\ docs/configuration.md|docs/supervision-protocols/*) printf '%s\n' pure-contract-unit ;; @@ -1589,8 +1653,15 @@ families_for_changed_path() { README.md|LICENSE|assets/*|docs/*|.gitignore) ;; *) - families_for_test_reference "$path" \ - || printf '%s\n' "__unmapped__:$path" + if [ -e "$path" ]; then + families_for_test_reference "$path" \ + || printf '%s\n' "__unmapped__:$path" + else + # A retired source path with no remaining test consumer cannot select + # a runnable suite. Known source paths above retain their mappings, + # and a still-referenced removal is found by the same reference scan. + families_for_test_reference "$path" || true + fi ;; esac } @@ -1666,6 +1737,17 @@ detect_gate_skip() { esac } +# Echo the reason a gate skip gave, i.e. the first meaningful output line with +# its leading "skip:" removed. Tabs and stray whitespace are folded so the +# reason stays one field of the tab-separated record the JSON artifact is built +# from. Callers only use this once detect_gate_skip has already said yes. +gate_skip_reason() { + local file=$1 first + first=$(awk 'NF { print; exit }' "$file" 2>/dev/null || true) + first=${first#skip:} + printf '%s\n' "$first" | tr '\t' ' ' | sed -e 's/^ *//' -e 's/ *$//' +} + # True when any output line contains "skip: <token>" (token may contain spaces). detect_gate_skip_token() { local file=$1 token=$2 @@ -1719,7 +1801,7 @@ with open(records_file, encoding="utf-8") as fh: line = line.rstrip("\n") if not line: continue - path, family, expected, exit_s, dur_s, gate = line.split("\t") + path, family, expected, exit_s, dur_s, gate, reason = line.split("\t") scripts.append({ "path": path, "family": family, @@ -1727,6 +1809,7 @@ with open(records_file, encoding="utf-8") as fh: "duration_ms": int(dur_s), "exit": int(exit_s), "gate_skip": gate == "true", + "gate_skip_reason": reason, }) families = [] @@ -1989,6 +2072,16 @@ case "$PER_SCRIPT_TIMEOUT_SECS" in ''|*[!0-9]*) die "--per-script-timeout-secs requires a whole number of seconds (0 disables)" ;; esac +# Refuse before any suite is selected or run. The inspection modes execute +# nothing: --list-families, --list-concurrent-safe-families, --list-lanes, +# --check-coverage, --concurrent-safe-family-jobs-max and --aggregate-json have +# already exited above, and --list/--list-scheduled print their selection and +# exit below. An unset MODE still falls through to the usage error, so a caller +# who named no selection mode is told that rather than this. +if [ -n "${MODE:-}" ] && [ "$LIST_ONLY" -eq 0 ] && [ "$LIST_SCHEDULED" -eq 0 ]; then + refuse_primary_checkout_for_task +fi + case "${MODE:-}" in all) select_all @@ -2279,7 +2372,7 @@ family_bump() { record_script_result() { local script=$1 rc=$2 duration=$3 out=$4 end_iso=$5 - local base family expected gate_skip fail_delta + local base family expected gate_skip gate_reason fail_delta base=$(basename "$script") family=$(family_for_basename "$base") expected=$(expected_gate_skip_for_family "$family") @@ -2290,9 +2383,14 @@ record_script_result() { fi gate_skip=false + gate_reason= if [ "$rc" -eq 0 ] && detect_gate_skip "$out"; then gate_skip=true + gate_reason=$(gate_skip_reason "$out") SKIPPED_GATE=$((SKIPPED_GATE + 1)) + # A capability skip is the runner's only record of what this host could not + # exercise, so name it rather than leaving a silent green. + log "gate skip: $script: ${gate_reason:-<no reason given>}" fi printf 'FM_TEST_END %s %s exit=%s duration_ms=%s gate_skip=%s\n' \ @@ -2305,8 +2403,8 @@ record_script_result() { AGG_RC=1 fi - printf '%s\t%s\t%s\t%s\t%s\t%s\n' \ - "$script" "$family" "$expected" "$rc" "$duration" "$gate_skip" >>"$RECORDS" + printf '%s\t%s\t%s\t%s\t%s\t%s\t%s\n' \ + "$script" "$family" "$expected" "$rc" "$duration" "$gate_skip" "$gate_reason" >>"$RECORDS" family_bump "$family" "$duration" "$fail_delta" TOTAL=$((TOTAL + 1)) } @@ -2409,11 +2507,20 @@ if [ "$JOBS" -eq 1 ]; then done else # Bounded concurrent execution for admitted scripts. Each worker gets a - # private mode-0700 TMPDIR so mktemp roots cannot collide. Retries are never - # used as a green strategy. + # private mode-0700 TMPDIR so mktemp roots cannot collide. Native Windows + # Bash layers report synthetic POSIX modes, so retain chmod there but enforce + # its observed mode only where the host reports real POSIX permissions. + # Retries are never used as a green strategy. worker_n=0 active_workers=0 + worker_root_mode_is_enforceable() { + case "$(uname -s)" in + MINGW*|MSYS*) return 1 ;; + *) return 0 ;; + esac + } + wait_one_job_worker() { local slot=$1 pid idx work script rc duration mode out end_iso pid=${WORKER_PIDS[$slot]} @@ -2435,14 +2542,16 @@ else if [ -s "$out" ]; then cat "$out" fi - mode=$(stat -c %a "$work" 2>/dev/null || stat -f %Lp "$work" 2>/dev/null || echo unknown) - case "$mode" in - 700|0700) ;; - *) - log "isolation failure: worker root mode is $mode, expected 0700 ($work)" - rc=1 - ;; - esac + if worker_root_mode_is_enforceable; then + mode=$(stat -c %a "$work" 2>/dev/null || /usr/bin/stat -f %Lp "$work" 2>/dev/null || echo unknown) + case "$mode" in + 700|0700) ;; + *) + log "isolation failure: worker root mode is $mode, expected 0700 ($work)" + rc=1 + ;; + esac + fi record_script_result "$script" "$rc" "$duration" "$out" "$end_iso" } diff --git a/bin/fm-turnend-guard.sh b/bin/fm-turnend-guard.sh index 634d304ce20..787e34c3195 100755 --- a/bin/fm-turnend-guard.sh +++ b/bin/fm-turnend-guard.sh @@ -35,9 +35,17 @@ # Away mode (state/.afk): the away-mode daemon owns supervision and runs the # watcher one-shot, restarting it after every wake, so the watch lock is # regularly unheld at a turn boundary with nothing wrong. A live -# identity-matched daemon holding this home, plus the unchanged fresh-beacon -# test, is what proves supervision there - see fm_afk_daemon_owns_supervision in -# bin/fm-wake-lib.sh. The strict watcher predicate is unchanged everywhere else. +# identity-matched daemon holding this home, plus a fresh beacon, is what +# proves supervision there - see fm_afk_daemon_owns_supervision in +# bin/fm-wake-lib.sh. The beacon freshness test there uses AFK_GRACE +# (fm_poll_derived_grace, docs/turnend-guard.md "Guard grace and the poll +# cadence"), not the flat $GRACE every other check on this page uses: the +# daemon starts a fresh one-shot watcher only after it finishes handling the +# previous wake, and that handling can legitimately run past a flat 300s +# window under load (a slow registered check, a busy supervisor pane) with the +# daemon perfectly healthy throughout. The strict watcher predicate and $GRACE +# are unchanged everywhere else, including for a dead daemon pid or a beacon +# older than AFK_GRACE, which still block. # # Loop-guard, codex/Grok (default) mode: never block twice in the same turn. # Codex uses stop_hook_active and Grok uses stopHookActive; typed camel-case @@ -191,10 +199,15 @@ fi # hand-off, when no watcher process holds the lock and nothing is wrong, so # requiring one here alarmed on healthy away-mode supervision. A live # identity-matched daemon holding this home is the right owner to test for. -# The beacon half of the predicate is deliberately unchanged: a daemon that -# stops restarting its watcher still blocks once the beacon passes grace, and -# a home with no daemon and no watcher blocks exactly as before. -if [ "$FM_SUP_WATCHER_FRESH" = true ] && fm_afk_daemon_owns_supervision "$STATE"; then +# The beacon half of the predicate still applies: a daemon that stops +# restarting its watcher still blocks once the beacon passes grace, and a home +# with no daemon and no watcher blocks exactly as before. It uses AFK_GRACE +# (poll-cadence-derived, see the comment above) instead of the flat $GRACE +# every other check on this page uses, so a daemon that is genuinely still +# cycling - just slower than a fixed 300s window - is not misread as down. +AFK_GRACE=${FM_GUARD_GRACE:-$(fm_poll_derived_grace)} +if [ "$(fm_path_age "$STATE/.last-watcher-beat")" -lt "$AFK_GRACE" ] \ + && fm_afk_daemon_owns_supervision "$STATE"; then allow_supervised_stop fi @@ -217,6 +230,8 @@ block_stop() { printf '● %s task(s) in flight, but no live watcher holds this home lock (last beat: %s).\n' "$FM_SUP_IN_FLIGHT" "$FM_SUP_BEACON_DESC" elif [ "$FM_SUP_SOURCES" -gt 0 ]; then printf '● %s process-event source(s) registered, but no live watcher holds this home lock (last beat: %s).\n' "$FM_SUP_SOURCES" "$FM_SUP_BEACON_DESC" + elif [ "$FM_SUP_CHECKS" -gt 0 ]; then + printf '● %s registered custom check(s), but no live watcher holds this home lock (last beat: %s).\n' "$FM_SUP_CHECKS" "$FM_SUP_BEACON_DESC" elif [ "$FM_SUP_QUEUE_PENDING" = true ]; then printf '● Queued wakes are pending and need supervision, but no live watcher holds this home lock (last beat: %s).\n' "$FM_SUP_BEACON_DESC" else @@ -454,6 +469,8 @@ if [ "$terminal_status" -eq 0 ]; then NEED_DESC="$FM_SUP_IN_FLIGHT task(s) in flight" elif [ "$FM_SUP_SOURCES" -gt 0 ]; then NEED_DESC="$FM_SUP_SOURCES process-event source(s) registered" + elif [ "$FM_SUP_CHECKS" -gt 0 ]; then + NEED_DESC="$FM_SUP_CHECKS registered custom check(s)" elif [ "$FM_SUP_QUEUE_PENDING" = true ]; then NEED_DESC="queued wakes pending" else diff --git a/bin/fm-wake-drain.sh b/bin/fm-wake-drain.sh index 73bbd30d8c6..8268bb917fa 100755 --- a/bin/fm-wake-drain.sh +++ b/bin/fm-wake-drain.sh @@ -1,5 +1,6 @@ #!/usr/bin/env bash -# Present durable watcher wake records, optionally acknowledge handled records, +# Present durable watcher wake records, retire rows no actor could ever consume, +# optionally acknowledge handled records, # annotate every unread line for validated signal status keys, surface unread # informational status lines, latest captain-facing statuses not covered by a # newer branch outcome, OPEN DECISIONS, and captain-call record divergence, @@ -67,38 +68,71 @@ ELIGIBLE_ROWS_FILE="$STATE/.branch-eligible-rows" ELIGIBLE_OWNER_FILE="$STATE/.branch-eligible-owner" MAIN_ROWS_FILE="$STATE/.main-eligible-rows" -rows_file_valid() { - [ -s "$1" ] && awk 'BEGIN { ok=1 } !/^[0-9]+$/ || seen[$0]++ { ok=0 } END { exit !ok }' "$1" -} - -branch_grant_live_locked() { - local version pid identity generation current - [ -f "$ELIGIBLE_OWNER_FILE" ] && [ ! -L "$ELIGIBLE_OWNER_FILE" ] || return 1 - exec 8< "$ELIGIBLE_OWNER_FILE" || return 1 - IFS= read -r version <&8 || { exec 8<&-; return 1; } - IFS= read -r pid <&8 || { exec 8<&-; return 1; } - IFS= read -r identity <&8 || { exec 8<&-; return 1; } - IFS= read -r generation <&8 || { exec 8<&-; return 1; } - if IFS= read -r _extra <&8; then exec 8<&-; return 1; fi - exec 8<&- - [ "$version" = fm-branch-eligible-owner-v1 ] || return 1 - case "$pid" in ''|*[!0-9]*|1) return 1 ;; esac - case "$generation" in ''|*[!A-Za-z0-9._-]*) return 1 ;; esac - current=$(fm_pid_identity "$pid" 2>/dev/null) || return 1 - [ -n "$current" ] && [ "$current" = "$identity" ] -} +rows_file_valid() { fm_wake_grant_rows_valid "$1"; } reclaim_stale_branch_grant_locked() { [ -e "$ELIGIBLE_ROWS_FILE" ] || [ -L "$ELIGIBLE_ROWS_FILE" ] || return 0 - if ! rows_file_valid "$ELIGIBLE_ROWS_FILE" || ! branch_grant_live_locked; then + if ! fm_wake_branch_grant_live "$ELIGIBLE_ROWS_FILE" "$ELIGIBLE_OWNER_FILE"; then rm -f -- "$ELIGIBLE_ROWS_FILE" "$ELIGIBLE_OWNER_FILE" fi } +# Retire rows no actor can ever consume. A claim, a presentation, and an +# acknowledgement all require the five appended fields and a numeric sequence, +# so a truncated or corrupted row is counted as queued while it can never be +# presented and can never be named by an --ack-through cutoff: left alone it +# wedges the queue for good. Main owns that repair - a branch grant can only +# name sequences that were structurally valid when it was published - and it +# runs under the queue lock, so no concurrent append is observed half-written. +# A repair that cannot be written (state/ full, unwritable, unreadable) is +# reported and never fatal: the usable rows are still presentable and +# acknowledgeable, and failing the whole drain would strand them too. +retire_unconsumable_rows_locked() { + local retired unusable queued kept + [ -f "$FM_WAKE_QUEUE" ] || return 0 + if DRAIN_TMP=$(mktemp "$STATE/.wake-queue.retire.XXXXXX") \ + && chmod 0600 "$DRAIN_TMP" \ + && unusable=$(awk -F '\t' -v keep="$DRAIN_TMP" ' + NF >= 5 && $2 ~ /^[0-9]+$/ { print > keep; next } + { shown++; if (shown <= 20) printf "wake drain: %s\n", $0 } + END { if (shown > 20) printf "wake drain: ... %d further unusable row(s) not shown\n", shown - 20 } + ' "$FM_WAKE_QUEUE"); then + queued=$(awk 'END { print NR }' "$FM_WAKE_QUEUE") + kept=$(awk 'END { print NR }' "$DRAIN_TMP") + retired=$(( queued - kept )) + if [ "$retired" -eq 0 ]; then + rm -f -- "$DRAIN_TMP" + DRAIN_TMP= + return 0 + fi + if _fm_atomic_replace "$DRAIN_TMP" "$FM_WAKE_QUEUE"; then + DRAIN_TMP= + printf 'wake drain: retired %s unusable queue row(s) that carried no sequence to present or acknowledge:\n%s\n' \ + "$retired" "$unusable" >&2 + return 0 + fi + fi + printf 'wake drain: unusable queue row(s) could not be retired (check that %s is readable and %s is writable); continuing with the rows that remain usable\n' \ + "$FM_WAKE_QUEUE" "$STATE" >&2 +} + +# One bounded line naming the rows a live branch grant is holding, so a main +# drain with nothing of its own never looks like a silently swallowed wake. +print_branch_held_notice() { + local held seqs + held=$(fm_wake_actor_pending_count branch "$ELIGIBLE_ROWS_FILE" "$ELIGIBLE_OWNER_FILE") || return 0 + [ "$held" -gt 0 ] || return 0 + seqs=$(fm_wake_grant_rows_valid "$ELIGIBLE_ROWS_FILE" \ + && awk 'NR <= 20 { printf "%s%s", (NR > 1 ? "," : ""), $1 } END { if (NR > 20) printf ",..." }' \ + "$ELIGIBLE_ROWS_FILE") + printf 'WAKE ROWS HELD BY SUPERVISION BRANCH: %s queued row(s) (%s) are granted to the live supervision branch, which presents and acknowledges them.\n' \ + "$held" "${seqs:-unknown}" +} + write_rows_file_locked() { # <target> <source> local target=$1 source=$2 if [ ! -s "$source" ]; then - rm -f -- "$target" + rm -f -- "$target" "$source" return fi chmod 0600 "$source" || return 1 @@ -606,6 +640,7 @@ else fi DRAIN_LOCK_HELD=true reclaim_stale_branch_grant_locked || exit 1 +[ "$ACTOR" != main ] || retire_unconsumable_rows_locked [ "$ACTOR" != branch ] || require_branch_eligible_rows || exit 1 if [ -n "$ACK_THROUGH" ]; then @@ -756,6 +791,11 @@ if [ "$ACTOR" = main ]; then fi claim_main_rows_locked || exit 1 if [ ! -s "$MAIN_ROWS_FILE" ]; then + # Every remaining row is reserved by the live branch grant, which presents + # and acknowledges them itself. Say so rather than exiting silently: a + # drain that prints nothing while the queue is visibly non-empty reads as a + # lost wake, and leaves the caller with no idea who owns what is queued. + print_branch_held_notice fm_lock_release "$FM_WAKE_QUEUE_LOCK" DRAIN_LOCK_HELD=false (print_status_presentation) || true diff --git a/bin/fm-wake-grant.sh b/bin/fm-wake-grant.sh index 2cc604f5f2c..bdff2fd4ead 100755 --- a/bin/fm-wake-grant.sh +++ b/bin/fm-wake-grant.sh @@ -22,27 +22,11 @@ trap cleanup EXIT trap 'exit 130' INT trap 'exit 143' TERM -rows_valid() { - [ -s "$1" ] && awk 'BEGIN { ok=1 } !/^[0-9]+$/ || seen[$0]++ { ok=0 } END { exit !ok }' "$1" -} +# fm-wake-lib.sh owns both the grant row-list shape and the owner-record read. +rows_valid() { fm_wake_grant_rows_valid "$1"; } -owner_matches() { - local expected_pid=${1:-} expected_generation=${2:-} version pid identity generation current - [ -f "$BRANCH_OWNER" ] && [ ! -L "$BRANCH_OWNER" ] || return 1 - exec 8< "$BRANCH_OWNER" || return 1 - IFS= read -r version <&8 || { exec 8<&-; return 1; } - IFS= read -r pid <&8 || { exec 8<&-; return 1; } - IFS= read -r identity <&8 || { exec 8<&-; return 1; } - IFS= read -r generation <&8 || { exec 8<&-; return 1; } - if IFS= read -r _extra <&8; then exec 8<&-; return 1; fi - exec 8<&- - [ "$version" = fm-branch-eligible-owner-v1 ] || return 1 - case "$pid" in ''|*[!0-9]*|1) return 1 ;; esac - case "$generation" in ''|*[!A-Za-z0-9._-]*) return 1 ;; esac - [ -z "$expected_pid" ] || [ "$pid" = "$expected_pid" ] || return 1 - [ -z "$expected_generation" ] || [ "$generation" = "$expected_generation" ] || return 1 - current=$(fm_pid_identity "$pid" 2>/dev/null) || return 1 - [ -n "$current" ] && [ "$current" = "$identity" ] +owner_matches() { # [<pid>] [<generation>] + fm_wake_branch_owner_matches "$BRANCH_OWNER" "${1:-}" "${2:-}" } case "${1:-}" in diff --git a/bin/fm-wake-lib.sh b/bin/fm-wake-lib.sh index e322e34bc63..8963d464011 100755 --- a/bin/fm-wake-lib.sh +++ b/bin/fm-wake-lib.sh @@ -92,7 +92,7 @@ fm_pid_identity() { fm_path_mtime() { if [ "$_FM_UNAME" = Darwin ]; then - stat -f %m "$1" 2>/dev/null + /usr/bin/stat -f %m "$1" 2>/dev/null else stat -c %Y "$1" 2>/dev/null fi @@ -104,6 +104,25 @@ fm_path_age() { echo $(( $(date +%s) - m )) } +# fm_poll_derived_grace [poll-seconds] +# Default guard-grace derivation: max(300, poll + 60). A watcher touches its +# liveness beacon once per poll cycle, so a fixed 300s grace stops correctly +# bounding staleness once the poll cadence reaches or exceeds it; growing the +# default with the cadence while keeping the historical 300s floor for the +# common short-poll case fixes that without a caller-specific constant. +# Defaults to $FM_POLL (fm-watch.sh's own poll env var) when no argument is +# given, so a caller with no independent notion of the poll cadence still +# derives the same default fm-watch.sh itself would use. +# docs/turnend-guard.md "Guard grace and the poll cadence" is the single owner +# of the rationale; every FM_GUARD_GRACE default should derive from this. +fm_poll_derived_grace() { + local poll=${1:-${FM_POLL:-15}} margin=60 derived + case "$poll" in ''|*[!0-9]*) poll=15 ;; esac + derived=$((poll + margin)) + [ "$derived" -ge 300 ] || derived=300 + printf '%s\n' "$derived" +} + # fm_watcher_lock_unheld <state> # True when the watcher lock or its symlinked owner directory is absent, or when # the existing lock records no pid at all. Any non-empty pid remains held here; @@ -1114,6 +1133,20 @@ fm_task_set_lock_path() { # <state-dir> printf '%s/.task-set.lock\n' "$state" } +# The top-most firstmate home reachable from this one on THIS machine, used as +# the single anchor every local home agrees on for machine-local shared state. +# +# A local parent binding is followed upward. A remote parent binding terminates +# the walk at the current home, which is the correct answer rather than an +# error: the parent lives on another machine, so its filesystem can neither hold +# nor be observed by a lock taken here, and a remote-seeded home is itself the +# top of the local tree that bin/fm-teardown.sh's collect_local_firstmate_states +# enumerates (that walk already skips remote registry entries for the same +# reason). Refusing a remote binding instead made every operation anchored here +# fail closed inside a remote secondmate home and its local descendants. +# +# Everything else still fails closed: an unreadable or malformed binding, an +# unreachable local parent, a cycle, and a chain deeper than the bound. fm_firstmate_root_home() { local home=${1:-$FM_HOME} marker parent seen="|" depth=0 home=$(CDPATH='' cd -- "$home" 2>/dev/null && pwd -P) || return 1 @@ -1124,11 +1157,11 @@ fm_firstmate_root_home() { . "$FM_WAKE_LIB_DIR/fm-secondmate-parent-lib.sh" fi fm_secondmate_parent_record_parse "$marker" || return 1 - # A remote parent lives on another filesystem. This home is the root of - # the locally knowable tree, so its descendants share its project locks. - # Parsing above still refuses malformed bindings instead of guessing. - [ "$FM_SECONDMATE_PARENT_ROUTE" != remote ] || break - [ "$FM_SECONDMATE_PARENT_ROUTE" = local ] || return 1 + case "$FM_SECONDMATE_PARENT_ROUTE" in + local) ;; + remote) break ;; + *) return 1 ;; + esac parent=$(CDPATH='' cd -- "$FM_SECONDMATE_PARENT_HOME" 2>/dev/null && pwd -P) || return 1 case "$seen" in *"|$parent|"*) return 1 ;; esac seen="$seen$home|" @@ -1139,6 +1172,14 @@ fm_firstmate_root_home() { printf '%s\n' "$home" } +# The one lock serializing Treehouse slot allocation and return for a project. +# +# It is anchored in the local root home's state directory so that every home on +# this machine that can reach the same pool - the root, and each secondmate home +# below it, including a remote-seeded home and its own local descendants - +# derives the identical path. Its identity is the project's resolved origin, so +# separate clones of one origin share a single lock; an origin-less local-only +# project falls back to its own worktree top instead of failing to resolve. fm_treehouse_project_lock_path() { # <project-dir> local project=$1 root origin identity hash top [ -d "$project" ] || return 1 @@ -1627,6 +1668,23 @@ fm_wake_queued_keys_locked() { "$FM_WAKE_QUEUE" 2>/dev/null || true } +fm_wake_secondmate_progress_marker_write() { # <task> <observed-at> <oldest-row-key> + local task=$1 observed_at=$2 oldest_row_key=$3 marker tmp + case "$task" in ''|*[!A-Za-z0-9._-]*) return 1 ;; esac + case "$observed_at" in ''|*[!0-9]*) return 1 ;; esac + case "$oldest_row_key" in ''|*[!0-9-]*) return 1 ;; esac + marker="$STATE/.secondmate-wake-progress-$task" + if [ -e "$marker" ] || [ -L "$marker" ]; then + [ -f "$marker" ] && [ ! -L "$marker" ] || return 1 + fi + tmp=$(mktemp "$STATE/.secondmate-wake-progress.XXXXXX") || return 1 + if ! printf '%s\t%s\n' "$observed_at" "$oldest_row_key" > "$tmp" || ! chmod 0600 "$tmp" \ + || ! _fm_atomic_replace "$tmp" "$marker"; then + rm -f -- "$tmp" + return 1 + fi +} + fm_wake_secondmate_stall_marker_write() { # <task> <row-key> local task=$1 row_key=$2 marker tmp case "$task" in ''|*[!A-Za-z0-9._-]*) return 1 ;; esac @@ -1724,6 +1782,86 @@ fm_wake_print_deduped() { ' "$file" } +# --- branch grant evidence and per-actor pending rows ------------------------ +# +# docs/watcher-continuity.md "Per-actor acknowledgement" owns the contract these +# helpers read; this is its single implementation, shared by the drain (which +# repairs and consumes a grant under the queue lock), the grant publisher, and +# the guard (which only counts, and never takes the lock). + +# 0 when <rows-file> is a non-empty list of distinct sequence numbers. +fm_wake_grant_rows_valid() { # <rows-file> + [ -s "$1" ] && awk 'BEGIN { ok=1 } !/^[0-9]+$/ || seen[$0]++ { ok=0 } END { exit !ok }' "$1" +} + +# 0 when <owner-file> holds the supported record, names a live process whose +# identity still matches what was recorded, and matches any expected pid and +# generation the caller pins. An unreadable, malformed, or superseded record is +# not a match, so uncertainty reads as "no live owner". +fm_wake_branch_owner_matches() { # <owner-file> [<pid>] [<generation>] + local file=$1 expected_pid=${2:-} expected_generation=${3:-} + local version pid identity generation current extra + [ -f "$file" ] && [ ! -L "$file" ] || return 1 + exec 8< "$file" || return 1 + IFS= read -r version <&8 || { exec 8<&-; return 1; } + IFS= read -r pid <&8 || { exec 8<&-; return 1; } + IFS= read -r identity <&8 || { exec 8<&-; return 1; } + IFS= read -r generation <&8 || { exec 8<&-; return 1; } + if IFS= read -r extra <&8; then exec 8<&-; return 1; fi + exec 8<&- + [ "$version" = fm-branch-eligible-owner-v1 ] || return 1 + case "$pid" in ''|*[!0-9]*|1) return 1 ;; esac + case "$generation" in ''|*[!A-Za-z0-9._-]*) return 1 ;; esac + [ -z "$expected_pid" ] || [ "$pid" = "$expected_pid" ] || return 1 + [ -z "$expected_generation" ] || [ "$generation" = "$expected_generation" ] || return 1 + current=$(fm_pid_identity "$pid" 2>/dev/null) || return 1 + [ -n "$current" ] && [ "$current" = "$identity" ] +} + +# 0 when a branch grant is currently reserving rows: a valid row snapshot whose +# recorded owner is still live. Anything else means no row is reserved. +fm_wake_branch_grant_live() { # <rows-file> <owner-file> + fm_wake_grant_rows_valid "$1" && fm_wake_branch_owner_matches "$2" +} + +# How many queued rows <actor> can act on right now - exactly the rows a drain +# by that actor would present or retire, and therefore the only rows worth +# telling that actor to drain. Main owns every structurally valid row a live +# branch grant does not reserve, plus every structurally invalid row. The branch +# owns exactly the rows its live grant names. Read without the queue lock: a +# torn read can only mis-count one poll, and the drain re-derives the set under +# the lock before it presents or mutates anything. +fm_wake_actor_pending_count() { # <actor> [<rows-file> <owner-file>] + local actor=${1:-main} rows=${2:-$STATE/.branch-eligible-rows} + local owner=${3:-$STATE/.branch-eligible-owner} grant='' count='' + [ -f "$FM_WAKE_QUEUE" ] || { printf '0\n'; return 0; } + if fm_wake_branch_grant_live "$rows" "$owner"; then + grant=$rows + fi + if [ "$actor" = branch ]; then + [ -n "$grant" ] || { printf '0\n'; return 0; } + count=$(awk -F '\t' -v seqs="$grant" ' + BEGIN { while ((getline line < seqs) > 0) keep[line] = 1 } + NF >= 5 && $2 ~ /^[0-9]+$/ && ($2 in keep) { n++ } + END { print n + 0 } + ' "$FM_WAKE_QUEUE") || count='' + else + count=$(awk -F '\t' -v seqs="$grant" ' + BEGIN { if (seqs != "") while ((getline line < seqs) > 0) reserved[line] = 1 } + NF < 5 || $2 !~ /^[0-9]+$/ { n++; next } + !($2 in reserved) { n++ } + END { print n + 0 } + ' "$FM_WAKE_QUEUE") || count='' + fi + # A queue that exists but cannot be counted (unreadable file, unreadable + # state/) is not evidence of an empty queue: report a pending row so callers + # still raise the alarm on a queue nobody can prove is drained. A failed count + # is decided by awk's exit status, not by what it printed, because an awk that + # reaches END after failing to open the queue would otherwise report 0 rows. + case "$count" in ''|*[!0-9]*) count=1 ;; esac + printf '%s\n' "$count" +} + # --- signal announcement signatures ----------------------------------------- # # The watcher's per-file signal scan (bin/fm-watch.sh scan_signals) detects a @@ -1741,7 +1879,7 @@ fm_wake_signal_sig() { # <file> -> reported-state signature status_observed_signature "$1" ;; *) - if [ "$_FM_UNAME" = Darwin ]; then stat -f '%z:%Fm' "$1" 2>/dev/null; else stat -c '%s:%Y' "$1" 2>/dev/null; fi + if [ "$_FM_UNAME" = Darwin ]; then /usr/bin/stat -f '%z:%Fm' "$1" 2>/dev/null; else stat -c '%s:%Y' "$1" 2>/dev/null; fi ;; esac } diff --git a/bin/fm-watch.sh b/bin/fm-watch.sh index df4c3921213..5e267492732 100755 --- a/bin/fm-watch.sh +++ b/bin/fm-watch.sh @@ -100,16 +100,12 @@ # check: inactive-outcome bounded poll-loop reconciliation found a suspicious # inactive terminal outcome that still lacks its durable # upstream receipt -# check: secondmate wake-loop stalled: mate=<id> row=<seq> age=<seconds>s +# check: secondmate wake-loop stalled: mate=<id> row=<seq> idle=<seconds>s # state=<verdict> source=<producer> -# the oldest valid row in an endpoint-recorded local -# secondmate home's durable wake queue outlived the age -# its busy verdict allows - FM_SECONDMATE_WAKE_STALL_SECS -# for a verdict the busy contract cannot establish, the -# shorter FM_SECONDMATE_WAKE_IDLE_STALL_SECS when the -# mate is provably idle, and never while it is provably -# busy; observation is read-only and one parent receipt -# suppresses repeats for that row +# oldest actionable queue identity has made no progress +# beyond its semantic busy-state threshold; declared +# external waits are excluded, busy suppression is bounded, +# and one notification covers each no-progress episode # For normal supervision, resume the session-start primary-harness protocol # after each printed reason. Direct duplicate invocations of this script still # no-op through the watcher singleton lock. @@ -167,7 +163,6 @@ mkdir -p "$STATE" WATCH_LOCK="$STATE/.watch.lock" WATCH_PATH="$SCRIPT_DIR/fm-watch.sh" WATCHER_DOWNTIME_MARKER="$STATE/.watcher-down" -WATCHER_STALE_GRACE=${FM_WATCHER_STALE_GRACE:-${FM_GUARD_GRACE:-300}} # The singleton-lock acquisition, EXIT trap, and the blocking supervision loop # all live below the source guard at the very bottom of this file (see "Main # entry"). Sourcing this file for unit tests therefore loads the functions - @@ -182,8 +177,10 @@ WATCHER_STALE_GRACE=${FM_WATCHER_STALE_GRACE:-${FM_GUARD_GRACE:-300}} # appended to that garbage. Arithmetic under `set -u` then aborts on the stray # token (e.g. the word "File" read as an unset variable), which silently kills the # watcher mid-cycle. Detect the platform once and pick the right form. +# On Darwin, call /usr/bin/stat rather than PATH-resolved stat so GNU coreutils +# cannot shadow the BSD `-f` syntax. if [ "$(uname)" = Darwin ]; then - stat_mtime() { stat -f %m "$1" 2>/dev/null; } # epoch seconds of mtime + stat_mtime() { /usr/bin/stat -f %m "$1" 2>/dev/null; } # epoch seconds of mtime else stat_mtime() { stat -c %Y "$1" 2>/dev/null; } fi @@ -192,6 +189,15 @@ fi # turn-ended signature, annotation staleness checks, and guarded bookkeeping writes. POLL=${FM_POLL:-15} # seconds between cycles +# The liveness beacon is touched once per cycle, immediately before the +# terminal wait below (event_wait_or_sleep) as well as at the top of the next +# one, so a healthy cycle's beacon can legitimately age up to POLL seconds +# between touches. fm_poll_derived_grace (bin/fm-wake-lib.sh, already sourced +# transitively above) is the single owner of the max(300, poll+60) +# derivation - see docs/turnend-guard.md "Guard grace and the poll cadence". +# This recomputes the library default above now that the real configured +# POLL is known. +WATCHER_STALE_GRACE=${FM_WATCHER_STALE_GRACE:-${FM_GUARD_GRACE:-$(fm_poll_derived_grace "$POLL")}} HEARTBEAT=${FM_HEARTBEAT:-600} # base seconds between heartbeat scans HEARTBEAT_MAX=${FM_HEARTBEAT_MAX:-7200} # heartbeat backoff cap CHECK_INTERVAL=${FM_CHECK_INTERVAL:-300} # seconds between *.check.sh sweeps @@ -250,24 +256,9 @@ STALE_ESCALATE_SECS=${FM_STALE_ESCALATE_SECS:-240} # idle secs before a provabl # override, because the dashboard's Task activity signal reads the same one out # of the fleet snapshot rather than inventing a second tolerance for quiet. BUSY_TURN_MAX_SECS=$(fm_sup_busy_turn_max_seconds) -# A local secondmate's foreign queue is checked on every poll, but the age of its -# oldest unacknowledged row alone never decides a stall: a row stays -# unacknowledged for the whole handling turn by design, so an aged row on a mate -# that is provably mid-turn is the durable-until-handled contract working. The -# semantic busy verdict in bin/fm-busy-lib.sh is the primary signal here and -# these two ages are its backstops. -# SECONDMATE_WAKE_STALL_SECS is the bare-time backstop for every verdict the -# busy contract cannot establish, which is unknown rather than idle and so must -# still alarm; it sits above a real turn, matching the same 1800s tolerance -# FM_RUN_STRANDED_SILENCE_SECS already carries for the same evidence, that a -# single agent step routinely runs 10 to 18 minutes on one line. +# No-progress thresholds selected by the semantic busy verdict. An unreadable +# mate is not proven idle; busy suppression is bounded by BUSY_TURN_MAX_SECS. SECONDMATE_WAKE_STALL_SECS=${FM_SECONDMATE_WAKE_STALL_SECS:-1800} -# SECONDMATE_WAKE_IDLE_STALL_SECS is the shorter age used when the mate is -# PROVABLY idle, where an unacknowledged row means nothing is running to handle -# it. A proven verdict outranks elapsed time, so this keeps a genuinely dead -# wake loop detected in about a minute rather than at the raised backstop. It is -# clamped to the backstop, so an idle mate is never treated as less suspicious -# than an unestablished one. SECONDMATE_WAKE_IDLE_STALL_SECS=${FM_SECONDMATE_WAKE_IDLE_STALL_SECS:-60} # A crew that declared a pause is idling on a known external wait, so its stale # pane is absorbed rather than wedge-escalated, whether or not its agent is still @@ -733,14 +724,23 @@ recorded_windows() { done } -# Print the oldest structurally valid row in a local secondmate's foreign queue. -# This is a read-only observation: the receiving home owns acknowledgement and -# this parent never changes the row or the foreign queue. +# Print the oldest structurally valid ACTIONABLE row in a local secondmate's +# foreign queue. A stale recheck that explicitly identifies itself as a declared +# external-wait pause is not evidence that the mate's wake loop is stuck: the +# pause cadence already owns that bounded visibility, and blocked waits remain +# actionable because they do not carry this declaration. This is a read-only +# observation: the receiving home owns acknowledgement and this parent never +# changes the row or the foreign queue. secondmate_oldest_queue_row() { # <queue-path> local queue=$1 [ -f "$queue" ] && [ ! -L "$queue" ] || return 0 awk -F '\t' ' - NF >= 5 && $1 ~ /^[0-9]+$/ && $2 ~ /^[0-9]+$/ { + function declared_external_pause(kind, payload) { + return kind == "stale" \ + && payload ~ /^stale: .*\(paused [0-9]+s, awaiting external - declared (pause,|paused\))/ + } + NF >= 5 && $1 ~ /^[0-9]+$/ && $2 ~ /^[0-9]+$/ \ + && !declared_external_pause($3, $5) { if (!found || $2 < seq) { found = 1 seq = $2 @@ -751,13 +751,24 @@ secondmate_oldest_queue_row() { # <queue-path> ' "$queue" 2>/dev/null || true } -# Surface one durable parent check for one unchanged foreign row after its -# bounded age. The primary marker and queued-key check make repeated watcher -# cycles converge without a notification storm, while an empty queue removes -# only this home's marker so a later row can be observed. +# Surface one durable parent check when the foreign queue's drain position has +# not moved for the bounded interval. The progress marker records that position +# as the same epoch-sequence row identity the stall receipts use, so the timer +# restarts whenever a different row becomes the oldest actionable one - as the +# mate drains, and as a queue reprovisioned under the same task id starts its +# own generation of rows at whatever sequence it restarts, and neither is a +# continued no-progress episode; row creation time belongs to that identity but +# never to the interval. A moved position ends an alerted episode and starts a +# new observation interval, so a newly-oldest row cannot alert immediately while +# a later genuine freeze remains visible. A mate demonstrably inside an active +# turn is suppressed only within BUSY_TURN_MAX_SECS, retaining the bounded +# busy-state safety backstop. +# Receipts close the append-before-marker crash window without changing the +# foreign queue. secondmate_wake_stall_tick() { local now=$(( $(date +%s) )) backstop=$SECONDMATE_WAKE_STALL_SECS idle_secs=$SECONDMATE_WAKE_IDLE_STALL_SECS - local meta task kind remote_host home queue row epoch seq row_key marker receipt receipt_dir notify_key queued age reason + local meta task kind remote_host home queue row epoch seq row_key marker progress_marker progress observed_at observed_key + local receipt receipt_dir notify_key queued idle reason episode_alerted local verdict state source threshold case "$backstop" in ''|*[!0-9]*|0) backstop=1800 ;; esac case "$idle_secs" in ''|*[!0-9]*|0) idle_secs=60 ;; esac @@ -779,9 +790,10 @@ secondmate_wake_stall_tick() { queue="$home/state/.wake-queue" row=$(secondmate_oldest_queue_row "$queue") marker="$STATE/.secondmate-wake-stall-$task" + progress_marker="$STATE/.secondmate-wake-progress-$task" receipt_dir="$STATE/.secondmate-wake-stall-receipts/$task" if [ -z "$row" ]; then - rm -f "$marker" + rm -f "$marker" "$progress_marker" if [ -e "$receipt_dir" ] || [ -L "$receipt_dir" ]; then [ -d "$receipt_dir" ] && [ ! -L "$receipt_dir" ] || return 1 rm -rf -- "$receipt_dir" || return 1 @@ -793,15 +805,34 @@ $row EOF case "$epoch" in ''|*[!0-9]*) continue ;; esac case "$seq" in ''|*[!0-9]*) continue ;; esac - age=$((now - epoch)) - # The cheapest gate first: nothing can alarm below the shortest age any - # verdict allows, so the busy read is skipped entirely on a young row. - [ "$age" -ge "$idle_secs" ] || continue - # fm_busy_classify_live, not the record alone: a mate whose endpoint is gone - # classifies dead rather than carrying a stale busy record that would - # suppress this check forever. busy suppresses, idle takes the short age, - # and every other verdict - unknown, dead, unreadable - takes the bare-time - # backstop, because an unestablished verdict is never evidence of idleness. + row_key="$epoch-$seq" + episode_alerted=0 + if [ -e "$marker" ] || [ -L "$marker" ]; then + [ -f "$marker" ] && [ ! -L "$marker" ] || return 1 + episode_alerted=1 + fi + progress=$(cat "$progress_marker" 2>/dev/null || true) + observed_at=${progress%%[[:space:]]*} + observed_key=${progress#*[[:space:]]} + if [ "$observed_at" = "$progress" ]; then + observed_key= + else + observed_key=${observed_key%%[[:space:]]*} + fi + case "$observed_at" in ''|*[!0-9]*) observed_at= ;; esac + case "$observed_key" in ''|*[!0-9-]*) observed_key= ;; esac + if [ -z "$observed_at" ] || [ -z "$observed_key" ] \ + || [ "$now" -lt "$observed_at" ] || [ "$row_key" != "$observed_key" ]; then + fm_wake_secondmate_progress_marker_write "$task" "$now" "$row_key" || return 1 + [ "$episode_alerted" -eq 0 ] || rm -f "$marker" || return 1 + continue + fi + [ "$episode_alerted" -eq 0 ] || continue + idle=$((now - observed_at)) + # Do not pay for a live read before even the shortest threshold is due. + [ "$idle" -ge "$idle_secs" ] || continue + # Read the semantic verdict only after recording progress, so a busy turn + # cannot hide queue movement or convert uncertainty into proven idleness. verdict=$(fm_busy_classify_live "$(fm_backend_of_meta "$meta")" \ "$(fm_backend_target_of_meta "$meta")" "$(fm_meta_get "$meta" harness)" \ "$task" "$STATE" "$(fm_backend_expected_label_of_selector "$task" "$STATE")" \ @@ -810,20 +841,21 @@ EOF source=${verdict#* } [ "$source" != "$verdict" ] || source=unrecorded case "$state" in - busy) continue ;; + busy) + ! busy_turn_over_age "$task" && continue + threshold=$backstop + ;; idle) threshold=$idle_secs ;; *) threshold=$backstop ;; esac - [ "$age" -ge "$threshold" ] || continue - row_key="$epoch-$seq" + [ "$idle" -ge "$threshold" ] || continue receipt="$receipt_dir/$row_key" - if [ -e "$marker" ] || [ -L "$marker" ]; then - [ -f "$marker" ] && [ ! -L "$marker" ] || return 1 + if [ "$(cat "$receipt" 2>/dev/null || true)" = "$row_key" ]; then + fm_wake_secondmate_stall_marker_write "$task" "$row_key" || return 1 + continue fi - [ "$(cat "$marker" 2>/dev/null || true)" = "$row_key" ] && continue - [ "$(cat "$receipt" 2>/dev/null || true)" = "$row_key" ] && continue notify_key="secondmate-wake-loop-$task-$row_key" - reason="check: secondmate wake-loop stalled: mate=$task row=$seq age=${age}s state=$state source=$source" + reason="check: secondmate wake-loop stalled: mate=$task row=$seq idle=${idle}s state=$state source=$source" queued=$(fm_wake_queued_keys check) if ! printf '%s\n' "$queued" | grep -Fx "$notify_key" >/dev/null 2>&1; then fm_wake_append check "$notify_key" "$reason" || return 1 @@ -1332,35 +1364,126 @@ parked_agent_is_dead() { # <window> [ "$state" = dead ] } -# Read the status line BEFORE queueing anything. This used to queue the wake -# first and only then discover the declared pause, which is how a brief-compliant -# parked crew still got a bare "stale: <window>" every few minutes: the pause was -# recognised one step too late to suppress the wake it had just queued, and the -# .paused-* markers it stamped afterwards left no streak record, so the designed -# widening cadence never ran. A declared pause is now routed to its owner, -# handle_paused_stale, which advances the same stale suppressor and owns the -# backoff. A captain-held transfer still surfaces here, because it is not the -# crew's own declaration that the pane is idle on purpose. Its repeat sights -# are scoped to that declaration so pane churn cannot bypass the cadence. +# The two records of one ordinary crew wait, and why its stale alarm reads both. +# +# status_is_paused_or_captain_held reads the status LINE a worker wrote, which is +# the only record when the worker itself is waiting. It is not the only record +# there is: once firstmate hands work to the captain, the wait is written into the +# BACKLOG by bin/fm-captain-hold.sh, and the worker's last line stays whatever it +# was - routinely `done: PR ...` after a delivery, which no line predicate can +# read as a wait. An alarm bounded only by the line therefore re-fires for the +# captain's whole thinking time, on exactly the work they already have in hand. +# +# `open` is that record's own read-only predicate and owns its semantics: exit 0 +# still an open captain call, 1 not, 2 could not be established. Only a 0 bounds +# an alarm here, so an unreadable backlog, an incompatible or absent tasks-axi, +# and a row this home does not carry all keep alarming exactly as they do today - +# a wait this watcher cannot prove is not a wait. +# +# The read costs one subprocess and runs only where the watcher is about to +# alarm, so at most once per distinct stale hash per window, beside the crew-state +# read the same paths already pay. The secondmate stale gate deliberately runs +# before this bound and admits only status-declared waits: a backlog-only hold +# whose mate still says `working:` or `done:` does not reach this read. Reaching +# it would put backlog reads into windows deliberately skipped on ordinary polls. +STALE_WAIT_DECLARATION= + +CAPTAIN_CALL_IDENTITY= + +task_captain_call_open() { # <task> + local task=$1 + CAPTAIN_CALL_IDENTITY= + [ -n "$task" ] || return 1 + CAPTAIN_CALL_IDENTITY=$(FM_HOME="$FM_HOME" "$SCRIPT_DIR/fm-captain-hold.sh" \ + open "$task" --identity 2>/dev/null) || return 1 + return 0 +} + +# The identity a re-surface throttle is bound to: the task's whole status-log +# signature. Any new status event - a replacement wait, a fresh delivery, a +# blocker - changes it and so starts its own window instead of inheriting the +# silence of the one before it. +stale_wait_declaration() { # <task> + printf 'declared:%s' "$(fm_wake_signal_sig "$STATE/$1.status" || true)" +} + +# The same scope for a captain call, carrying the CALL's own lifecycle identity +# beside the status signature. The status log is not enough on its own: a task +# can be answered with `--release` and held again as a genuinely different call +# without any status append, and binding the throttle to the signature alone let +# the second call inherit the first one's silence and absorbed its first sight. +# That first sight is the one alarm this bound must never swallow - a decision +# waiting on the captain that is never surfaced is invisible, where a delivery +# announced twice is merely noise. +captain_call_declaration() { # <task> <call-identity> + printf 'captain-hold:%s:%s' "$2" "$(fm_wake_signal_sig "$STATE/$1.status" || true)" +} + +# 0 when <declaration> has already been alarmed for this window inside the +# current PAUSE_RESURFACE_SECS. A pure read: recording an alarm is the caller's, +# so the throttle is never advanced by a sighting it just absorbed. +stale_wait_throttled() { # <window-key> <declaration> + local throttle="$STATE/.paused-resurfaced-$1" + [ "$(cat "$throttle" 2>/dev/null || true)" = "$2" ] \ + && [ "$(age_of "$throttle")" -lt "$PAUSE_RESURFACE_SECS" ] +} + +# The same bound, for a stale window whose last line IS captain-relevant. That +# line is real and its first sight must still reach the captain, but a delivery +# they are already holding has nothing new to say on the next pane tick. +# Sets STALE_WAIT_DECLARATION to the scope this sighting is bound to, and leaves +# it EMPTY when no open captain call bounds it, so an unheld delivery, a blocker, +# and a failure alarm exactly as they do today. +# Returns 0 to absorb this sighting; 1 to alarm, after which the caller records +# the throttle through stale_wait_record once its own wake append has succeeded. +# Record a fired wake against the bounded cadence, and ONLY after that wake was +# durably appended. A marker written ahead of the append outlives a failed one: +# the watcher exits with no wake queued, and the next sighting reads the fresh +# marker and absorbs the retry, which is the single way this bound could swallow +# an alarm outright rather than delay it. +stale_wait_record() { # <window-key> + [ -n "$STALE_WAIT_DECLARATION" ] || return 0 + printf '%s' "$STALE_WAIT_DECLARATION" > "$STATE/.paused-resurfaced-$1" +} + +# Bound a due stale alarm for an ordinary crew task held for the captain. +# Backlog-only secondmate holds are outside this guard because the earlier gate +# preserves their no-backlog-read hot path. +captain_call_stale_bound() { # <window-key> <task> + local key=$1 task=$2 + STALE_WAIT_DECLARATION= + task_captain_call_open "$task" || return 1 + STALE_WAIT_DECLARATION=$(captain_call_declaration "$task" "$CAPTAIN_CALL_IDENTITY") + stale_wait_throttled "$key" "$STALE_WAIT_DECLARATION" +} + +# Surface inconclusive states once, then bound repeated sightings of the same +# status-declared transfer or backlog call. A worker's own paused declaration +# routes directly to handle_paused_stale, preserving its widening cadence. surface_nonterminal_stale() { # <window> <hash> - local win=$1 h=$2 key task last declaration='' declared=1 throttled=1 + local win=$1 h=$2 key task last declared=1 bounded=1 throttled=1 key=$(window_key "$win") task=$(window_to_task "$win" "$STATE") last=$(last_status_line "$STATE/$task.status") + STALE_WAIT_DECLARATION= if status_is_paused "$last"; then handle_paused_stale "$win" "$task" "$h" return fi if status_is_paused_or_captain_held "$last"; then declared=0 - declaration="declared:$(fm_wake_signal_sig "$STATE/$task.status" || true)" - if [ "$(cat "$STATE/.paused-resurfaced-$key" 2>/dev/null || true)" = "$declaration" ] \ - && [ "$(age_of "$STATE/.paused-resurfaced-$key")" -lt "$PAUSE_RESURFACE_SECS" ]; then - throttled=0 - fi + bounded=0 + STALE_WAIT_DECLARATION=$(stale_wait_declaration "$task") + stale_wait_throttled "$key" "$STALE_WAIT_DECLARATION" && throttled=0 + elif captain_call_stale_bound "$key" "$task"; then + bounded=0 + throttled=0 + elif [ -n "$STALE_WAIT_DECLARATION" ]; then + bounded=0 fi if [ "$throttled" -ne 0 ]; then fm_wake_append stale "$win" "stale: $win" || exit 1 + stale_wait_record "$key" fi printf '%s' "$h" > "$STATE/.stale-$key" rm -f "$STATE/.stale-since-$key" "$STATE/.wedge-holds-$key" @@ -1368,12 +1491,19 @@ surface_nonterminal_stale() { # <window> <hash> if [ "$declared" -eq 0 ]; then : > "$STATE/.paused-$key" date +%s > "$STATE/.paused-rechecked-$key" - [ "$throttled" -eq 0 ] || printf '%s' "$declaration" > "$STATE/.paused-resurfaced-$key" + elif [ "$bounded" -eq 0 ]; then + # A backlog hold is NOT a declared pause, and must not be dressed up as one: + # the loop-top reconciliation and pause_state_class both read the status LINE, + # so a .paused-* flag this line does not support would be cleared on the next + # poll - taking the throttle with it - and would hand the mate and dead-agent + # cadences a declaration they were never given. Only the shared re-surface + # marker is kept, which is the whole of what this bound needs. + rm -f "$STATE/.paused-$key" "$STATE/.paused-rechecked-$key" else clear_pause_state "$key" fi if [ "$throttled" -eq 0 ]; then - triage_log "absorbed non-terminal stale (declared wait already re-surfaced this window): $win" + triage_log "absorbed non-terminal stale (declared wait or open captain call already re-surfaced this window): $win" return 0 fi wake "stale: $win" @@ -2354,12 +2484,12 @@ EOF clear_pause_tracking "$key" fi # An idle secondmate endpoint is healthy by design, so a mate is admitted to - # the pane-stale path ONLY to serve a declared wait's bounded re-surface - - # the same declarations pause_state_class reconciles below, which is why this - # gate reads the shared predicate rather than the pause verb alone. Narrowing - # it to `paused` would leave a mate's captain hold rotting invisibly: the - # clear above already spares its pause tracking, but nothing would ever - # re-surface it. + # the pane-stale path ONLY to serve a status-declared wait's bounded + # re-surface. This gate reads the shared predicate rather than the pause verb + # alone so it includes a declared `captain-held` status. A hold recorded only + # in the backlog while the mate still says `working:` or `done:` is outside + # this guard: reaching it would require backlog reads for windows this gate + # deliberately skips, putting that read on the ordinary poll hot path. if [ "$kind" = secondmate ] && ! status_is_paused_or_captain_held "$last"; then continue fi @@ -2437,8 +2567,21 @@ EOF rm -f "$whf" clear_write_tracking "$key" triage_log "absorbed stale (provably working, overriding a stale captain-relevant status): $w" + elif captain_call_stale_bound "$key" "$task"; then + # The line is captain-relevant and stays so, but the backlog says + # the captain already holds this work: further NEW pane hashes with + # the same status-log state have nothing to add while they are + # deciding. Only that new-hash repetition is bounded - the first + # sight already alarmed, a new hash inside the window is absorbed, + # and a new hash after it alarms again. A stable hash stays as inert + # here as it already was after a first terminal alarm. + printf '%s' "$h" > "$sf" + rm -f "$ssf" + clear_write_tracking "$key" + triage_log "absorbed stale (open captain call already surfaced for this status): $w" else fm_wake_append stale "$w" "stale: $w" || exit 1 + stale_wait_record "$key" printf '%s' "$h" > "$sf" rm -f "$ssf" "$whf" clear_write_tracking "$key" diff --git a/bin/fm-x-lib.sh b/bin/fm-x-lib.sh index f50ddce781d..aae910db8cb 100644 --- a/bin/fm-x-lib.sh +++ b/bin/fm-x-lib.sh @@ -89,8 +89,8 @@ fmx_single_link_file_valid() { local file=$1 expected_device=${2-} links device [ -f "$file" ] && [ ! -L "$file" ] || return 1 if [ "$(uname)" = Darwin ]; then - links=$(stat -f %l "$file" 2>/dev/null) || return 1 - device=$(stat -f %d "$file" 2>/dev/null) || return 1 + links=$(/usr/bin/stat -f %l "$file" 2>/dev/null) || return 1 + device=$(/usr/bin/stat -f %d "$file" 2>/dev/null) || return 1 else links=$(stat -c %h "$file" 2>/dev/null) || return 1 device=$(stat -c %d "$file" 2>/dev/null) || return 1 @@ -103,7 +103,7 @@ fmx_single_link_file_mode_valid() { local file=$1 expected_mode=$2 expected_device=${3-} mode fmx_single_link_file_valid "$file" "$expected_device" || return 1 if [ "$(uname)" = Darwin ]; then - mode=$(stat -f %Lp "$file" 2>/dev/null) || return 1 + mode=$(/usr/bin/stat -f %Lp "$file" 2>/dev/null) || return 1 else mode=$(stat -c %a "$file" 2>/dev/null) || return 1 fi @@ -114,8 +114,8 @@ fmx_private_artifact_dir_device() { local dir=$1 mode device [ -d "$dir" ] && [ ! -L "$dir" ] || return 1 if [ "$(uname)" = Darwin ]; then - mode=$(stat -f %Lp "$dir" 2>/dev/null) || return 1 - device=$(stat -f %d "$dir" 2>/dev/null) || return 1 + mode=$(/usr/bin/stat -f %Lp "$dir" 2>/dev/null) || return 1 + device=$(/usr/bin/stat -f %d "$dir" 2>/dev/null) || return 1 else mode=$(stat -c %a "$dir" 2>/dev/null) || return 1 device=$(stat -c %d "$dir" 2>/dev/null) || return 1 @@ -410,7 +410,7 @@ fmx_request_relay_context() { fmx_context_registry_mtime() { local file=$1 mtime - mtime=$(stat -f '%m' "$file" 2>/dev/null) || mtime=$(stat -c '%Y' "$file" 2>/dev/null) || return 1 + mtime=$(/usr/bin/stat -f '%m' "$file" 2>/dev/null) || mtime=$(stat -c '%Y' "$file" 2>/dev/null) || return 1 case "$mtime" in ''|*[!0-9]*) return 1 ;; esac diff --git a/docs/architecture.md b/docs/architecture.md index 3c9c160ab8d..3a37c8da36a 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -14,6 +14,12 @@ A declared wait that is still unchanged at each recheck widens its own recheck w The widening is earned by one wait and dies with it: a crew that resumes, clears the pause, or declares a different wait is rechecked at the base window again, and no backoff ever widens the `FM_STALE_ESCALATE_SECS` wedge threshold. The streak decides only how WIDE that window is; when the next recheck is due is anchored on the crew's status file mtime in the watcher and on when the hold began in the away-mode daemon, so a crew that keeps rewriting its paused reason is still rechecked on schedule while the captain is away. Normal mode separately absorbs repaint-only repeats of an already-surfaced keyed open-decision set unless the backend confidently reports that the parked agent is dead. +For an ordinary crew task, a wait is read from both of its records: the status line a worker declared, and the backlog hold `bin/fm-captain-hold.sh` recorded once firstmate handed the work to the captain. +So a delivered ordinary crew task whose last line stays `done: PR ...` bounds repeated alarms from new pane hashes to the declared-wait recheck cadence for the length of the captain's decision. +The first hash still alarms, each new hash inside that window is absorbed, and a new hash after the window re-surfaces the hold; a terminal pane hash that never changes stays inert after its first alarm exactly as it did before this bound. +The throttle is scoped to both the current captain-call lifecycle and the status-log state, so releasing and re-holding the same task without a status append starts a fresh window whose first new hash alarms. +A secondmate reaches the stale path only for a wait declared in its status line, so a hold recorded only in the backlog while its last line is `working:` or `done:` is outside this guard. +Reaching that case would require consulting the backlog for windows the secondmate gate deliberately skips, putting backlog reads on the ordinary poll hot path this design preserves. Repeated provably-working stale escalations on the same unchanged pane add an escalation count to the wake reason and, at `FM_WEDGE_DEMAND_INSPECT_COUNT`, a `demand-deep-inspection` marker. At the moment either supervisor would raise that wedge escalation - and only there, so a bounded pipeline read costs at most one call per window per pane rather than one per poll - it asks `bin/fm-run-progress.sh` whether the crew's validation run is actually moving, because "there is an active run" and "that run is progressing" are different facts and a review or test step routinely emits one opening line and then works silently for ten to eighteen minutes. A run reported `progressing` holds the escalation and restarts the timer; a `stranded` run escalates and names the step that stopped, while a `none` verdict escalates unless the declared-wait fallback below applies. @@ -36,15 +42,14 @@ The one wedge-escalation caller the hold cannot reach is the parked-open-decisio While away mode is active, a busy pane that crosses the bound under a declared wait is handed to the daemon as the plain wake identity instead of taking that recheck in the watcher, because the daemon owns triage there and a wake already decorated as a possible wedge would override the daemon's own declared-wait verdict; an undeclared busy pane past the bound still takes the wedge escalation in away mode. That handoff is keyed on the declaration itself (the status log's signature) rather than on the pane capture, so a harness footer that ticks on every poll wakes the daemon once per declaration instead of once per poll, and it clears the wedge timer, escalation count, and worktree-write deferral exactly as the normal-mode absorber does, so an undeclared busy phase's timer does not resume when the declaration lifts. Those actionable wakes are written to a durable local queue (`state/.wake-queue`) only after generation-bound recovery evidence is published, so an interrupted watcher or handling turn can be recovered without losing the queue record. -Agent endpoint liveness and queue-consumption liveness are separate: on each poll, the primary watcher reads the oldest valid row from every endpoint-recorded local secondmate home's durable wake queue without locking, consuming, or rewriting that foreign queue. -A row's age alone never decides that loop is stalled, because acknowledgement comes after handling completes and so a row stays unacknowledged for the whole handling turn by design. -The primary reads the mate's semantic busy verdict ([`bin/fm-busy-lib.sh`](../bin/fm-busy-lib.sh) owns it) as the deciding signal and elapsed time only as its backstop: a provably busy mate never alarms, a provably idle one alarms at the shorter `FM_SECONDMATE_WAKE_IDLE_STALL_SECS`, and any verdict the contract cannot establish - unknown, dead, or unreadable - alarms at `FM_SECONDMATE_WAKE_STALL_SECS`, so an unestablished verdict is never read as idleness and a genuinely dead loop is still caught. -The record-backed path through that verdict split is dormant for normally spawned local secondmates today because their launch path does not arm the parent-owned busy-state record. -The `herdr-native` and `grok-regex` paths return busy verdicts without consulting that record, so those backends exercise the verdict split today. -Arming the parent-owned busy-state record for local secondmates remains out of scope for this change. -The primary then appends one keyed `check` wake naming the mate, row sequence, observed age, and the verdict with its producing source; parent receipts and queued-key deduplication suppress repeats for the same row across watcher and handling crashes, while empty and younger queues remain silent. +Agent endpoint liveness and queue-consumption liveness are separate: on each poll, the primary watcher reads the oldest valid actionable row from every endpoint-recorded local secondmate home's durable wake queue without locking, consuming, or rewriting that foreign queue. +The observation interval starts when that oldest epoch-sequence identity is first seen and resets whenever it changes, including a queue reprovisioned under the same task id; declared external-wait rows are excluded. +The semantic verdict from [`bin/fm-busy-lib.sh`](../bin/fm-busy-lib.sh) selects the no-progress threshold: proven idle uses `FM_SECONDMATE_WAKE_IDLE_STALL_SECS`, while unknown, dead, and unreadable verdicts use `FM_SECONDMATE_WAKE_STALL_SECS` because uncertainty is not evidence of idleness. +Proven busy suppresses the alarm only within `FM_BUSY_TURN_MAX_SECS`; after that bound it uses the unknown-state backstop, and any alarm requests inspection rather than interrupting the mate. +One keyed parent notification covers each no-progress episode, with receipts and queued-key deduplication surviving watcher and handling crashes; empty or advancing queues reset the episode. +The record-backed verdict path remains dormant for normally spawned local secondmates whose launch does not arm the parent-owned busy-state record, while native Herdr and Grok verdicts can establish busy without that record. Endpointless registered mates remain outside this scan because startup secondmate-liveness owns dead or missing endpoint recovery, and remote homes retain their host-local supervision boundary. -`tests/fm-wake-queue.test.sh` pins the notification, idempotence, quiet-queue, and byte-for-byte foreign-row preservation guarantees. +`tests/fm-wake-queue.test.sh` pins the no-progress notification, drain-progress reset, declared-pause exclusion, active-turn deferral, idempotence, quiet-queue, and byte-for-byte foreign-row preservation guarantees. When a canonical validated PR poll returns exactly `merged`, the watcher routes it through the shared merge-outcome emitter before retiring the poll. [`bin/fm-merge-outcome-lib.sh`](../bin/fm-merge-outcome-lib.sh)'s header owns role routing, PR-specific wake identity, marker-locked normal deduplication, and the at-least-once ordering that prefers a rare duplicate over silence. After successful outcome publication, the watcher immediately delivers the emitter's local actionable poll row and publishes a private retirement receipt bound to the poll's registration, bytes, file identities, metadata, provider, URL, and task ID. @@ -72,7 +77,7 @@ This declared-wait classification does not read a secondmate's endpoint liveness A live captain-held pane's repeat inspections are scoped to its declaration, so pane-hash changes cannot reset the repeat throttle. Its initial normal-mode status signal still surfaces through the no-verb path, while away mode self-handles that routine signal and owns the later recheck. Fresh stale panes use the same current-state read before trusting the status log, so an active run or a proven busy worker outranks an old captain-relevant status-log line left behind before validation. -When that captain-relevant status includes a still-open keyed decision, the one-shot suppressor is keyed on a digest of the complete open-decision set returned by `bin/fm-classify-lib.sh`'s `status_open_decisions` fold rather than on the pane hash, because an idle harness pane repaints on its own and a hash-keyed suppressor re-surfaced an unchanged already-escalated waiting state on every repaint. +When that captain-relevant status includes a still-open keyed decision, the one-shot suppressor is keyed on a digest of the explicitly keyed open-decision set returned by `bin/fm-classify-lib.sh`'s `status_open_decisions --explicit-only` view rather than on the pane hash, because an idle harness pane repaints on its own and a hash-keyed suppressor re-surfaced an unchanged already-escalated waiting state on every repaint. While the backend reports the agent live or returns inconclusive liveness, an unchanged open-decision set therefore surfaces only once despite pane repaint, while a fresh actionable status signal, a changed decision key or set, and resolution still surface normally. A crew parked on such an already-surfaced open decision whose backend confidently reports its agent dead still escalates through the same wedge timer. No-change heartbeats are also benign. @@ -158,11 +163,11 @@ It suppresses failed-looking closes when the same identity-matched watcher is he Cursor's `bin/fm-turnend-guard-cursor.sh` hook is the same between-turns shape in one synchronous step: it parks the awaited `stop` hook on the arm wrapper and translates an actionable close into one `followup_message`, with a generation baton that makes an older park still running after the next `stop` claim stand down instead of leaking a stale duplicate wake. The existing turn-end guard remains the final backstop for every harness-engine protocol, with pi-signed sharing Pi's protocol, omp's blocking `session_stop` hook compelling one continuation per turn, the `--claude` mode cooperating with the auto-arm claim, and Cursor's `--cursor` mode rendering a block as one bounded follow-up because its `stop` step cannot be blocked. Its `--restart` mode signals only the watcher recorded in the current home's `state/.watch.lock`, so restarting one home cannot kill sibling secondmate watchers. -A pull-based guard (`bin/fm-guard.sh`) warns through supervision tool output if the primary checkout is tangled, if work, process-event sources, Relay polling, or queued wakes have an unhealthy model-aware supervision verdict, or if queued wakes are waiting to be drained even while supervision is healthy. -The drain script calls that guard after presenting the queue; records remain durable, and may keep the queued-wakes warning visible, until the exact generation-bound acknowledgement printed by the drain succeeds after handling. +A pull-based guard (`bin/fm-guard.sh`) warns through supervision tool output if the primary checkout is tangled or if work, process-event sources, registered custom checks, Relay polling, or queued wakes has an unhealthy model-aware supervision verdict; on main it also warns when queued wakes are waiting for main itself to drain. +The drain script calls that guard after presenting the queue; records remain durable until the exact generation-bound acknowledgement printed by the drain succeeds after handling, and main may keep the queued-wakes warning visible until then. +The Pi supervision branch's deliberate queued-wake warning exception is owned by [`pi-supervision-branch.md`](pi-supervision-branch.md#components-and-their-owners), while [`watcher-continuity.md`](watcher-continuity.md#per-actor-acknowledgement) owns the guard's per-actor counting, the advisory main gets for rows a live branch grant holds, and main's retirement of queue rows no actor could ever present or acknowledge. It leads with a prominent bordered tangle banner, while `bin/fm-guard.sh` owns the watcher-down banner and reminder policy so repeated guarded commands stay noisy without reprinting the full banner in the same episode. -On every verified primary harness, tracked hook integration gives the primary session a push-based backstop: when work, a process-event source, Relay polling, or a pending wake queue needs supervision and no identity-matched watcher lock with a fresh beacon is live, blocking-capable Stop hooks block and nonblocking turn-end integrations force one bounded follow-up. -The turn-end guard also accepts the away daemon's ownership proof with a fresh watcher beacon; `bin/fm-wake-lib.sh` owns that proof. +On every verified primary harness, tracked hook integration gives the primary session a push-based backstop: when work, a process-event source, a registered custom check, Relay polling, or a pending wake queue needs supervision and no supervision owner provably holds this home with a fresh beacon, blocking-capable Stop hooks block and nonblocking turn-end integrations force one bounded follow-up. The guard covers the main primary and genuinely marked secondmate homes, exempts child crewmate/scout worktrees, is loop-safe per harness, and is documented in [turnend-guard.md](turnend-guard.md). A presence-gated sub-supervisor (`bin/fm-supervise-daemon.sh`) extends this for walk-away supervision: the `/afk` skill starts it through the tracked foreground helper `bin/fm-afk-start.sh`, after which the watcher reverts to daemon-managed one-shot mode and the daemon self-handles routine wakes in bash. @@ -256,6 +261,7 @@ Only a named non-default branch checked out in `FM_ROOT` is a worktree tangle. `fm-guard.sh` prints the repair command on the next mutable fleet action, while `bin/fm-session-start.sh` reports the same condition through bootstrap as a `TANGLE:` line at session start. If another live session holds the fleet lock, both surfaces keep the alarm but switch to read-only wording with no repair command. Ship and design briefs also tell the crewmate to verify `pwd -P` and `git rev-parse --show-toplevel` before the generated branch step, then stop with a blocked status if it landed in the primary checkout. +Placement is proven only at launch, so `bin/fm-spawn.sh` also exports the task id as `FM_TASK_ID` into every ship, design, and scout pane, and `bin/fm-test-run.sh` refuses to execute the behavior suite from the primary checkout while that marker is set; the runner's header owns the predicate and [`tests/fm-test-run.test.sh`](../tests/fm-test-run.test.sh) pins it. ## No-mistakes gate authority boundary @@ -364,6 +370,7 @@ The same emitter handles a merge firstmate performed and one its poll detected, After a merge succeeds, `bin/fm-pr-merge.sh` closes at most one eligible recorded work item, on whichever forge it holds a write adapter for and in the repository that item records; multiple items, a forge or host with no adapter, a credential that is absent or refused, and bookkeeping failures warn without turning a completed merge into a failed, retryable merge. Teardown is fail-closed for ship and design worktrees: dirty worktrees refuse, and committed work must be landed before the worktree is returned. A pool worktree is only returned after teardown passes the slot-ownership proof: a contradictory task record or supported live endpoint refuses without touching either task, and no discard authority relaxes that. +Allocation and return serialize on one project lock per machine-local Firstmate tree: every home reachable through local parent links shares that lock, and a home seeded from another machine anchors its own, because a lock taken on this filesystem is neither held nor observable across that boundary. [`bin/fm-teardown.sh`](../bin/fm-teardown.sh)'s header owns the landed-work proofs, slot-ownership proof, PR-discovery fallback, pre-teardown run conclusion, and stale-lock recovery procedure; [`tests/fm-teardown-endpoint-safety.test.sh`](../tests/fm-teardown-endpoint-safety.test.sh) and [`tests/fm-secondmate-safety.test.sh`](../tests/fm-secondmate-safety.test.sh) pin the slot-collision boundary. ## Optional Relay diff --git a/docs/calm-mode-feasibility.md b/docs/calm-mode-feasibility.md index 9c8c67d7311..415944489c4 100644 --- a/docs/calm-mode-feasibility.md +++ b/docs/calm-mode-feasibility.md @@ -20,6 +20,7 @@ Across those versions only Pi 0.84 introduced a relevant presentation API change The adapters gate on the exact method they patch rather than on a version number, so those versions remain verification evidence rather than compatibility bounds. The exported classes used by the adapters (`AssistantMessageComponent` and `InteractiveMode`) are undocumented internals with no stated version guarantee. `tests/fm-calm-pi-extension.test.sh` records the installed Pi version as evidence without gating on it and covers both newer synthetic versions and an unavailable adapter seam. +This host tracks Pi latest, so the version the evidence is pinned to moves; the [2026-09-07 record](#2026-09-07-pi-0851-renderer-and-export-dom-verification) owns the currently pinned version and the renderer comparison behind it. ### Built-in tool override constraints @@ -239,12 +240,12 @@ The test fixture enumerates every class below through the centralized policy, an | `system-notice` | `showStatus`, `showError`, compaction, retry, and startup warning rows | Unsupported boundary; remains visible. | | `cache-notice` | Non-persisted cache-miss `Text` row | Unsupported boundary; remains visible. | | `project-trust-warning` | Non-persisted startup `Text` row | Unsupported boundary; remains visible. | -| `synthetic-user` | Firstmate extension `sendUserMessage`, terminal-injected input, Firstmate-generated Pi positional brief, or the already non-displayed session-start nudge | Canonically classified text-only operational user messages stay ordinary semantic user messages but render through the zero-height adapter (verified on Pi 0.81.1 through 0.84.1) under Calm; legacy entries stay gaplessly controllable, and the session-start nudge retains its existing non-displayed custom-message path. | +| `synthetic-user` | Firstmate extension `sendUserMessage`, terminal-injected input, Firstmate-generated Pi positional brief, or the already non-displayed session-start nudge | Canonically classified text-only operational user messages stay ordinary semantic user messages but render through the zero-height adapter under Calm; legacy entries stay gaplessly controllable, and the session-start nudge retains its existing non-displayed custom-message path. | | `synthetic-assistant` | No authoritative Firstmate source found | Policy-hidden, but Pi exposes no generic assistant-role renderer. | | `unknown` | Future or unclassified transcript component | Policy-hidden, but no generic renderer exists; never claimed as covered. | The installed extension API has no supported global transcript filter, user-message renderer, assistant-message renderer, chat-container API, or generic custom-tool wrapper. -Pi 0.81.1 through 0.84.4 export `AssistantMessageComponent` and `InteractiveMode`, so Calm uses separate idempotent, API-probed adapters for assistant thinking layout, the complete operational-user transcript row, and the live-host capture the legacy synthetic entry renderer needs, while leaving all message data and non-Calm rendering unchanged; see the [compatibility contract](calm.md#pi-compatibility) for how a future Pi lacking one of those exports is handled. +Pi 0.81.1 through 0.82.0, Pi 0.84.4, and Pi 0.85.1 export `AssistantMessageComponent` and `InteractiveMode`, so Calm uses separate idempotent, API-probed adapters for assistant thinking layout and the complete operational-user transcript row while leaving all message data and non-Calm rendering unchanged; see the [compatibility contract](calm.md#pi-compatibility) for how a future Pi lacking one of those exports is handled. General component replacement, ANSI cursor erasure, provider-context mutation, and installed-file patching remain rejected as unsupported or preservation-breaking workarounds. ## Cross-harness verification record @@ -290,16 +291,6 @@ It asserts one persisted and rendered captain answer, exact user-role operationa Quoted current markers, ASCII-only labels, ordinary text before a marker, unrelated U+2063 placement, and image-bearing input remain visible in component and native transcript checks. `tests/fm-pi-primary-live-e2e.test.sh` also proves the working ship replaces the built-in `Working...` row while Calm is active on the credentialed provider path, and that it clears when the run settles, before continuing its ordinary watcher lifecycle. `tests/fm-pi-primary-types.test.sh` performs strict no-emit TypeScript checking against the Pi declaration package named by `FM_PI_PACKAGE_DIR`, defaulting to the globally installed one, so each dated record below names the declaration version its own run covered rather than pinning one version here. -That check exits 0 with `skip: tsc not found for Pi extension typecheck` where TypeScript is absent, no `tsc` version is pinned anywhere in this repository, and `.github/workflows/ci.yml` installs Node without TypeScript, so a green suite is evidence of this check only when the environment that produced it had `tsc` installed. - -CI does not exercise this suite. -CI run `31630144139` executed both Pi-dependent test scripts, but the Pi package was absent, so `tests/fm-pi-primary-types.test.sh` and `tests/fm-calm-pi-extension.test.sh` both exited successfully after skipping their Pi-dependent checks rather than running them. -Those are the only two suites in that run that skipped for a missing package: the log contains seven matching skip lines because the Calm suite reports the absent package from six separate subtests, while the typecheck reports it once. -The per-lane log summaries expose only aggregate `skipped_gate=1` counts, while the GitHub run summary names neither skipped suite and reports every job successful. -The strict typecheck and runtime guards therefore both skip silently in CI. -The evidence for this fix is local, dated, and reproducible by hand; it is not enforced anywhere. -A future Pi version bump will reintroduce this class of drift without any check failing. -Issue [#118](https://github.com/HelloWorldSungin/firstmate/issues/118) owns the question of whether CI can install and pin real Pi or needs another non-stub compatibility guard, and folds in TypeScript pinning plus strict-typecheck enforcement. The relevant commands are: @@ -767,3 +758,71 @@ FM_TEST_END 2026-08-29T01:01:30Z tests/fm-pi-branch-extension.test.sh exit=0 dur ``` The real renderer comparison exercised twelve outcome lines and reported collapsed and expanded parity with Pi stock, zero visible rows under Calm, restored stock parity after toggling Calm off, and delegated stock HTML export fallback. + +## 2026-09-07 Pi 0.85.1 renderer and export-DOM verification + +This host tracks Pi latest, so the version this contract's evidence is pinned to moves. +The renderer and lifecycle evidence below was taken against installed `@earendil-works/pi-coding-agent` 0.85.1 with `@earendil-works/pi-server` 0.85.0 also installed globally. + +Calm's rendered rows are unchanged across 0.84.4, 0.85.0, and 0.85.1. +`FM_PI_PACKAGE_DIR` points `tests/fm-calm-pi-extension.test.sh` at an isolated install, so each comparison ran against its own temporary dependency tree and never mutated the globally installed packages. + +```text +$ pi --version +0.85.1 + +$ npm ls -g --depth 0 @earendil-works/pi-coding-agent @earendil-works/pi-server +├── @earendil-works/pi-coding-agent@0.85.1 +└── @earendil-works/pi-server@0.85.0 +``` + +```text +$ FM_PI_PACKAGE_DIR=<pi 0.84.4> tests/fm-calm-pi-extension.test.sh +ok - Pi calm centralizes transcript visibility, preserves execution/export data, keeps Pi's stock working row visible while no run is active, and persists its choice across session starts +$ FM_PI_PACKAGE_DIR=<pi 0.85.0> tests/fm-calm-pi-extension.test.sh +ok - Pi calm centralizes transcript visibility, preserves execution/export data, keeps Pi's stock working row visible while no run is active, and persists its choice across session starts +$ FM_PI_PACKAGE_DIR=<pi 0.85.1> tests/fm-calm-pi-extension.test.sh +ok - Pi calm centralizes transcript visibility, preserves execution/export data, keeps Pi's stock working row visible while no run is active, and persists its choice across session starts +``` + +Reaching that parity on 0.85 took one contract adaptation, landed earlier in 85ad5e7. +Pi 0.84 and older silently substituted a built-in's stock definition when a `ToolExecutionComponent` was constructed without one, so the calm-off equivalence baseline could be built definition-less and still read as stock. +Pi 0.85 removed that substitution, so the definition-less baseline renders Pi's generic text fallback instead - which is what produced `read collapsed rendering changed while calm mode was off`. +The renderer change was real, and it was the contract's baseline that had to adapt, not Calm's wrappers: the wrapped rows matched Pi stock before and after. +`tests/fm-calm-pi-extension.test.sh` now builds each baseline from the real stock tool-definition factories that `dist/core/tools/index.js` exports, calling the built-in's own factory with `process.cwd()`, which reads as stock on 0.84.4 and on 0.85.x alike and no longer depends on the removed substitution. + +Pi 0.85.0 alone requires a package it does not declare. +Its `dist/experimental/server.js` statically imports `@earendil-works/pi-server`, which is absent from 0.85.0's `dependencies`, `peerDependencies`, and `optionalDependencies`, so a clean install of 0.85.0 on its own cannot load Pi's interactive mode at all: + +```text +Error [ERR_MODULE_NOT_FOUND]: Cannot find package '@earendil-works/pi-server' imported from + .../node_modules/@earendil-works/pi-coding-agent/dist/experimental/server.js +``` + +Installing `@earendil-works/pi-server@0.85.0` beside it restores the identical Calm rendering, and 0.85.1 no longer reaches that import. +That packaging gap is a separate installation defect, not the renderer change above: it stops Pi from loading at all rather than altering any rendered row. + +The `could not render calm-mode HTML export DOM` failure was a headless-Chrome start-up flake, not a change in Pi's export shape. +It appeared in exactly one of the thirteen most recent CI runs, and that run installed the same Pi 0.85.1 as the runs immediately before and after it, which both passed. +The render step is a vendor-tool step: the assertions that follow it are what protect the Calm conversation boundary. +It now retries a bounded number of Chrome start-ups on a fresh profile and, when every attempt fails, reports the Chrome binary, its version, the installed Pi version, each attempt's exit status, whether that attempt was timed out, and Chrome's own stderr, so the next occurrence is diagnosable from the CI log alone. +`test_export_dom_render_guard` in the same script pins that behavior with real processes and no browser. + +The complete Calm suite against installed Pi 0.85.1, with `FM_CHROME_BIN` naming the Chrome the render step used: + +```text +$ FM_CHROME_BIN=<chrome> tests/fm-calm-pi-extension.test.sh +ok - Pi calm resolves its persistent home independently of Pi's launch directory +ok - Pi calm compatibility evidence never rejects a Pi version for being newer than 0.82.0, and still fails closed on a missing or malformed version +ok - a missing collapsed-thinking presentation API degrades only that Calm adapter with a clear skip reason, while the rest of Calm still registers +ok - missing Pi presentation class exports reach the independent adapter degradation path +ok - Calm registers none of its 7 built-in tool wrappers at load while config/calm is off, and all 7 synchronously at load while config/calm is on +ok - Calm's first same-session /calm activation claims every uncontested built-in, leaves a foreign bash tool fully intact and callable, warns prominently and logs the contested name, and only rows constructed before that activation - the documented bound - fail to retroactively collapse +ok - Pi calm centralizes transcript visibility, preserves execution/export data, keeps Pi's stock working row visible while no run is active, and persists its choice across session starts +ok - Pi calm on collapses mid-turn assistant working notes to zero height while Calm off keeps them, leaves streaming, truncated-final, and genuine final replies untouched, never mutates the messages, ignores every /calm argument, and restores a legacy persisted max as ordinary Calm on +ok - Pi operational follow-up E2E processes exact user-role notifications once while Calm hides current and adjacent rows, Calm off and absent render them, and restart preserves semantics +ok - Pi Calm native /skill:ahoy geometry keeps every collapsed thinking and tool block at zero height while preserving expansion, history, restart, and Calm-off rendering +ok - Pi Calm working ship moves on a slow independent cadence over faster fixed-cell blue water, paints the complete boat standard yellow with balanced resets, keeps ANSI-stripped width exact, flips the directional sail on the exact bounce at both edges and every width, clamps visible and hidden resizes, falls back deterministically when narrow, freezes and resumes column/direction across settle/start without hidden-time jumps or duplicate timers, resets only on a fresh session, and installs and removes one scheduler-owning widget across starts, settle, abort, failure, shutdown, reload, replacement, and Calm toggles while leaving Calm-off visibility untouched +ok - the rendered-export-DOM guard renders in one pass, retries a bounded number of Chrome start-up failures, and reports the Chrome binary, Chrome version, Pi version, exit status, and Chrome diagnostic when every attempt fails +ok - Pi calm native E2E replaces the stock working row with a moving, resize-clamped working ship that freezes and resumes across two working periods in one Pi session, clears on abort, keeps captain turns visible, hides exact operational user rows without changing persistence, restores stock rendering Calm-off, survives restart, and preserves export plus Ctrl+O behavior +``` diff --git a/docs/captain-hold-lifecycle.md b/docs/captain-hold-lifecycle.md index f09b20e1501..dd5bdb5ae65 100644 --- a/docs/captain-hold-lifecycle.md +++ b/docs/captain-hold-lifecycle.md @@ -38,7 +38,7 @@ The policy prefers holding the very work item a question gates, so the backlog r `bin/fm-teardown.sh` therefore asks the read-only `open` subcommand before its automatic close: exit 0 means the row is still an open captain call (not Done, `hold_kind: captain`), 1 means it is not, and 2 means the answer could not be established, which teardown treats as a refusal before any destructive step rather than as permission to close. On 0 only the close changes: after cleanup and still under the task's own lock, teardown records one `Deliverable of the finished work: ...` line at the end of the task body and runs `tasks-axi reopen`, so the row returns to Queued with its hold intact and remains on the appropriate Captain's Call or Charted Next decision surface instead of reading as work still under way. The pending-close record teardown already stages before destructive cleanup carries that intent as a `mode=retain` line, so an interrupted cleanup replays the retention at the next session start through the same record, validator, and lock as an ordinary close and never closes the row; an answer that closed the row first simply retires the record. -`--force` does not lift the deferral, because it authorizes discarding unlanded work, never the captain's question, and `answer` remains the only act that closes the call. +`--force` does not lift the deferral, because it authorizes discarding unlanded work, never the captain's question; only `answer` with the captain's words or evidence-backed `reconcile close` closes the call. `bin/fm-backlog-transition-lib.sh` owns the transition and its record, and `bin/fm-captain-hold.sh --help` owns the predicate's contract. ## Answer-time closure @@ -57,6 +57,63 @@ Two channels feed that one intake today, and both are ordinary callers rather th Trusted external process-event adapters intentionally expose no answer operation and cannot feed this authority-bearing intake; [`extension-bindings.md`](extension-bindings.md#trust-boundary) owns that boundary. `bin/fm-procevent-lavish.sh answers` is one such adapter command; it reads only rows tagged `choice`, relays a card's declared close mode, and can never let freeform captain prose forge a task id or a mode. +## Reconcile: re-check reality, never a blind close + +A captain call can stop being a question without the captain ever answering it because the subject lands, the premise turns out to be false, or the choice becomes a matter of fact rather than the captain's to make. +`reconcile` is the standing third option for that case, and its whole point is that it is NOT an answer. +It means "go verify the latest state", and it resolves in exactly one of two ways once that verification has actually been done: close the call with the evidence that made it moot, or leave it open with a note recording that it is genuinely still active. + +The value remains reserved at the shared keyed-answer intake, which visibly refuses it from every channel and never passes it to `answer`. +A reconcile value delivered through chat or any ordinary keyed-answer caller therefore cannot complete a task, lift a hold, write a resolution record, or create a reconcile request. + +Board request creation uses a separate captured-source seam. +The board emits `fm-bearings-answer.v1` context with the slug-shaped selected option and freeform note in separate fields, so annotating Reconcile cannot turn it into an ordinary answer value. +`bin/fm-procevent-lavish.sh answers` emits an exact non-reconcile selection, or a bare note when no option was selected, while `reconciles` emits only task ids whose structured selection is Reconcile and carries their notes as request provenance. +Current rows require the versioned shape and the `choice` tag; a time-limited rollout branch accepts ordinary answers from the old question/answer shape but refuses its bare and separator-annotated reconcile values from both intakes because those rows do not separate the selected option from its note. +Every other structurally uncertain capture feeds neither intake, remains announced, and cannot forge a task id from freeform prose. +The adapter-agnostic runner pipes reconcile rows into `reconcile-requests` only for a bound source, and that intake verifies the named binding again before it creates anything. +Failures remain best-effort and never acknowledge or suppress the captured result. +What this captured-source intake records is a durable reconcile request under `state/reconcile-requests/`, one private record per task, carrying the requesting provenance and a UTC timestamp. +The record exists so the obligation to re-check cannot be lost between the wake that carried the answer and the turn that acts on it. +It is idempotent per task: repeating a reconcile keeps one request and its original timestamp. +The supported creator is the runner carrying the captain's board selection; the binding-checked `reconcile-requests` command is that internal intake rather than an operator reconciliation outcome. + +Verification retires a request through one of two outcomes, and each one requires both the pending board-created request and the operator input that supports its claim: + +- `reconcile close <task-id> --evidence-file <path>` is the moot outcome. + It writes a resolution record whose mode is `reconciled` and whose body is the supplied EVIDENCE under a `Reconciliation evidence:` label, then closes the task. + The distinct mode and label are what keep the record honest: it says the call dissolved against verified evidence, and it never claims the captain answered. +- `reconcile note <task-id> --note-file <path>` is the still-active outcome. + It appends one dated `Captain hold reconciled:` note to the task body, leaves the hold in place, and retires the request. + The call stays the captain's, now carrying what the re-check found; a marker bound to the request timestamp, provenance, and note digest lets a matching retry finish retirement without appending again while a later request with the same finding still receives its own dated note. + +`reconcile list` is the read-only enumeration of pending requests filed by board answers. +A successful normal answer also retires any pending request, because an answered call has no remaining re-check obligation. +Every retirement is checked: if request removal fails after an answer, close, or note is already durable, the durable outcome stands but the command fails and leaves the pending request visible for retry. +No path here closes a captain call without either the captain's words through `answer` or the evidence through `reconcile close`. + +## Card hygiene: a landed subject is not a live call + +`bin/fm-bearings-board.sh build` cross-checks every `decision` card before it publishes and drops stale subjects rather than trusting the composed inventory alone. + +Three checks run, all on exact identity and none on prose: + +- The card's key is the captain-held task id, so `bin/fm-captain-hold.sh open --distinguish-absent` is asked whether that task is still an open captain call. + Exit 1 - present but closed, or no longer held for the captain - drops the card. + Exit 2 means the answer could not be established and exit 3 means the task is absent from the main backlog; both keep the card, because a card wrongly shown is recoverable and a call wrongly hidden is not. +- The payload's own `landed` rows are the recently-landed artifacts. + A decision card whose task id or `pr_url` appears among them has already shipped its subject, so it drops. +- A version decision can carry a structured `subject` with an artifact and numeric three-part version. + A landed row carrying the same artifact at that version or a newer one supersedes the card without parsing prose. + +Dropped cards are named on stderr as `dropped-landed-card:` lines so a rebuild states what it removed rather than quietly shrinking Captain's Call. +The landing procedure requires one immediate board rebuild to remove already-stale merged-PR and superseded-version cards without a committed migration or change-worktree state mutation. +A subject whose state cannot be established is kept, because a wrongly shown card is safer than a wrongly hidden call. +The validator's reservation scope must equal the adapter's reconcile-classification scope, which is all card types because the captured payload carries no card type. +Owner-aware routing for remote-secondmate decision cards is tracked separately: that follow-up must query landedness and route reconciliation in the authoritative secondmate home while honoring the remote and local consistency principle. +Until then, an absent main-home task passes through this hygiene check unchanged, and its Reconcile selection remains announced but cannot create a main-home request because the main intake refuses an absent task. +For a main-home call, the reconcile option is the recovery path for whatever still slips through. + ## Structured read surfaces `bin/fm-fleet-snapshot.sh` parses canonical tasks-axi `(hold: ...)`, `(hold-kind: ...)`, and `(hold-until: ...)` metadata alongside existing backlog fields. @@ -120,14 +177,19 @@ The shim recognizes an exact replay of a pre-collapse routed resolution by its h ## Verification record -Verification date: 2026-09-05. - The focused end-to-end regression suite is `tests/fm-captain-hold-lifecycle.test.sh`, using only synthetic `sample` identities and decision text. It proves: cleanup of a finished task whose own row is the captain call leaves that call open, queued, held, carrying its deliverable, and visible in Bearings' Captain's Call, leaves no pending record behind, survives a `--force` cleanup, and closes only when `answer` records the captain's words, while an ordinary finished task in the same home still closes with its report link; an interrupted cleanup leaves the row In flight and untouched with its pending record, and the next session start retains it as queued and held with the deliverable recorded; a relocated data directory keeps the retention in its one configured backlog; a ship row whose captain hold cannot be read refuses cleanup before any destructive step and surfaces the read failure; the reconstructed silent-divergence case is signalled - a status resolution over a still-open captain-held task reaches both `diverged` and the drain's `RECORD DIVERGENCE` section, under the collapsed and the legacy identity alike, while the backlog task, its hold, and the status log all survive the report unchanged and the printed hint names both reconciliation directions; the false-signal boundary holds - a captain call with no routed work item, a verified `captain-held` transfer, a still-open status decision, an already answered call, and an ordinary task whose keyed question was answered all stay silent; a report-only unresolved captain call refuses `--none` completion before teardown can erase the source; non-forced scout and design teardown always requires the durable inventory verification; the recorded-answer guard (a bare `tasks-axi done` close fails `verify` until `answer` records the captain's word, and an ordinary finished task cannot be dressed up as an answered call); answer-time closure through a bound channel with task-id keys, including the `release` close mode, mode-matched replay idempotence, and the refusal of drifted, mode-mismatched, absent, unheld, and already-closed keys; the chat channel reaching the same intake; hold-set stamping that precedes visible hold state, preserves an active lifecycle's timestamp, and resets after release; interrupted answer closure retaining the stamp until close and restoring resolution-first ordering on retry; deferral through `--until` leaving `captain_actionable` false until due; and every legacy path (composed identities through the shim, pre-collapse `decision_keys=` metadata, routed-resolution replay, and a concrete-origin binding). The markdown-to-beads migration family runs the same suite's beads fixture (bd-driven scratch graph, self-skipping on markdown-only tasks-axi installs) and proves: `verify` and `complete` resolve an attested legacy id through a migrated row's marker note, through the configured prefix when no row carries a note - naming the resolved row in the completion line - and through the marker note of a pre-collapse derived identity; a marker-noted row wins over an unrelated captain-held row occupying the bare prefix namesake; an unresolvable id is refused once naming the id (never an empty name); and the attested id stays in `decision_keys=` for idempotent re-verification. One case in that family needs no beads install and always runs: a stubbed tasks-axi that fails any markdown file override proves the captain-hold hold, answer, and close mutations reach a beads-configured home without one. +The reconcile path is pinned in the same suite: a reconcile answer arriving through the keyed-answer intake, in the default close mode and in the `release` mode a captain-gated work card declares, is refused and leaves both tasks held with no resolution record or request; only the separately bound captured-source intake records one durable request per task idempotently across a replay. +It also proves the two verification outcomes - an evidence-backed `reconciled` close that records the evidence under its own label and never as the captain's words, and a note that leaves the call queued, held, and dated - while both outcomes refuse without a pending board request, each durable mutation applies only once across close, probe, and request-retirement failures, a later distinct request with the same note still appends its own dated record, every failed retirement is surfaced with its pending request retained, incompatible resolution modes cannot replay as captain answers, and normal close, release, and replay paths retire pending requests. +The captured-source coverage proves Lavish deduplicates each card before separating versioned structured selections from notes, bare and annotated Reconcile choices never reach keyed answers, genuine current and legacy choices still close normally, legacy bare and separator-annotated reconcile values feed neither intake, mixed repeated selections preserve every other card's final value, the generic runner creates a request only through a verified bound source, chat reconcile text creates none, and the resulting board request authorizes evidence-backed closure. +The board's half is pinned in `tests/fm-bearings-board.test.sh`: every published decision card carries exactly one reconcile option, authored options reserve that value across every card type, recommendations name authored options, a decision card whose structured subject appears in the payload's landed rows is dropped while a genuinely open one is kept even when an unrelated landed id contains its key after a newline, a build requires a fresh authoritative listed-open result before binding or arming, a reopen retires the pre-reopen source generation and waits for a fresh live listener, and a rebuild of an already-armed board with no live listener starts one. +That suite drives its Lavish session through a protocol-shaped stub, and `tests/fm-bearings-board-lavish-live-e2e.test.sh` is the default-on capability guard for the installed provider; [`verification/process-event-sources.md`](verification/process-event-sources.md) owns the version-scoped evidence. +[`verification/process-event-sources.md`](verification/process-event-sources.md) owns the process-event ownership and reclamation evidence exercised by `tests/fm-procevent.test.sh`. + `tests/fm-classify-decision-key.test.sh` pins `status_key_closing_verb` itself: it separates a resolution from the durable-transfer close and from a still-open key, reports the last real transition across re-openings and both key positions, and treats a prose mention as no transition. Projection regressions live in `tests/fm-fleet-snapshot-view.test.sh` (the total structured-only bucket classifier, hold-until parsing, kind-independent captain actionability, undated-hold aging, and title stripping) and `tests/fm-bearings-snapshot.test.sh` (default and expanded decision-bucket membership, deferral explanations, blocker-overflow disclosure, working-hold dual surfaces, remote-summary schema invalidation, and the landed exclusion by surviving captain-hold annotations). -The exact commands and their summarized outputs are recorded in the shipping PR's evidence; run the four suites above plus `tests/fm-send-resolve-key.test.sh`, `tests/fm-bearings-board.test.sh`, and `bin/fm-lint.sh` to refresh this record. +The exact commands and their summarized outputs are recorded in the shipping PR's evidence; run the four suites above plus `tests/fm-send-resolve-key.test.sh`, `tests/fm-bearings-board.test.sh`, `tests/fm-procevent.test.sh`, and `bin/fm-lint.sh` to refresh this record, and `FM_BEARINGS_LAVISH_LIVE=1 tests/fm-bearings-board-lavish-live-e2e.test.sh` after a lavish-axi upgrade. diff --git a/docs/configuration.md b/docs/configuration.md index a50fb2a0bdf..a2818cf2ddb 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -209,7 +209,8 @@ The effort list is a handful of levels and stays on Pi's plain selector dialog. Both picks change the supervision branch alone and never the captain's own conversation model or effort. It persists the model pick in gitignored `config/supervision-branch-model` and the effort pick in gitignored `config/supervision-branch-effort`, both under the effective Firstmate home, resolved from `FM_HOME`, then `FM_ROOT_OVERRIDE`, then the tracked code root derived from the extension path, or under `FM_CONFIG_OVERRIDE` when that test and specialized-setup override is present. Firstmate keeps no model catalog of its own; the list is the intersection of what Pi reports when the picker opens and what a fresh isolated branch runtime can run. -A provider that exists only because an extension registered it inside the captain's session is not offered, while stored OAuth and API-key credentials retain their native credential type because Firstmate never copies, converts, installs, or overwrites credentials for the branch runtime. +A provider that exists only because an extension registered it inside the captain's session, such as pi-devin-auth's `devin`, is offered and can be pinned or followed like any other; [pi-supervision-branch.md](pi-supervision-branch.md#cost-model-and-the-byte-stable-prefix) owns how that registration reaches the isolated branch runtime. +Stored OAuth and API-key credentials retain their native credential type because Firstmate never copies, converts, installs, or overwrites credentials for the branch runtime. The file holds one `<provider>/<model-id>` line followed by one newline, split at the first `/` so a provider-qualified model id such as `openrouter/anthropic/claude-sonnet-4-5` survives intact. An absent, unreadable, or unparseable file means no pin, and the branch then follows main's own current model, applied explicitly and live whenever main changes models mid-session. A valid pin wins over main and remains unaffected by main's model changes. @@ -536,7 +537,7 @@ OPENAI_API_KEY SSH_AUTH_SOCK ``` -Firstmate retains basic home, executable search, terminal, locale, temporary-directory, and backend routing variables, plus its explicit launch assignments and enabled task trace. +Firstmate retains basic home, executable search, terminal, locale, temporary-directory, and backend routing variables, plus its explicit launch assignments, its ship and scout task marker, and enabled task trace. [`fm-spawn.sh --help`](../bin/fm-spawn.sh) owns the exact retained names and parsing mechanics. Other ambient names must be listed explicitly, including custom credential-store locations, proxy settings, and certificate overrides when required by the selected tools. The command shell and worker may still create their own variables. @@ -561,6 +562,8 @@ The filter runs at the worker command boundary, after the terminal daemon and pa This is not a sandbox: it cannot revoke same-user access to credential files, prevent tools or later shells from loading credentials again, or isolate processes from the same user's other processes. Regression coverage executes emitted launch commands with synthetic nonsecret values in [`tests/fm-spawn-dispatch-profile.test.sh`](../tests/fm-spawn-dispatch-profile.test.sh). +Every claude launch's inline `--settings` JSON also carries `"attribution":{"commit":"","pr":"","sessionUrl":false}`, so a spawned worker never writes a Co-Authored-By trailer, Claude-Session link, or generated-with line into a commit or PR body regardless of which settings scopes end up loaded. + ## Crew dispatch profiles (config/crew-dispatch.json) `config/crew-dispatch.json` is an optional local, gitignored file containing natural-language rules that firstmate reads before dispatching a crewmate or scout. @@ -760,7 +763,7 @@ See [`docs/examples/watched-tools.json`](examples/watched-tools.json) for a star Arm the check once per home with `bin/fm-tool-update-check.sh arm`. That writes `state/tool-updates.check.sh` and binds its bytes with `bin/fm-check-register.sh`, so the existing watcher polls it on its normal cadence and turns its one line into a `check:` wake; no separate schedule is involved. -The armed check runs whenever that home has a watcher running, and arming alone does not make watcher supervision required, so a home with no in-flight work and no other reason to watch does not start a watcher just for this check. +Registering the check is itself a reason to watch, so the home keeps a watcher for it after the last task is torn down, and `disarm` is what ends that need. `bin/fm-tool-update-check.sh disarm` removes the shim, its trust binding, and the report record. The check prints nothing when everything is current, and `state/.tool-updates` records the findings the last report was made from so the same pending update is reported once instead of on every poll. A changed or returning condition is reported again. @@ -970,8 +973,9 @@ Never run the registered blocking source command directly in a conversational tu A long-polling external process is registered as a *source* through its adapter, whose header and `--help` own the commands and flags. `bin/fm-procevent.sh` owns the generic contract; built-in adapters retain their tracked `bin/fm-procevent-<adapter>.sh` commands, while an explicitly bound external adapter routes through the trusted host contract above. `bin/fm-procevent-lavish.sh` is the first built-in adapter and wraps only the currently published `lavish-axi poll` interface. -That adapter, and only that adapter, retries the one exact transient response a cut-short listener returns while its marks remain available (`error: Lavish Editor poll response was interrupted` with `code: SERVER_ERROR`), up to 12 times at 5 second intervals, so an internal retry never reaches the runner as a captured result. -Real feedback, ended and missing sessions, any other `SERVER_ERROR`, and that same interruption still standing once the bound is spent are all captured and announced normally; `FM_LAVISH_POLL_RETRY_DELAY` is a bounded 0 to 60 second test override for the interval only, and the runner itself stays adapter-agnostic. +That adapter, and only that adapter, retries the one exact transient response a cut-short listener returns while its marks remain available (`error: Lavish Editor poll response was interrupted` with `code: SERVER_ERROR`), up to 12 times with poll starts at least 5 seconds apart, so an internal retry never reaches the runner as a captured result. +This start-to-start governor is a no-op after a normally blocking poll but caps an immediately returning poll under the shipped defaults independently of the owner lease and registration launch pacing. +Real feedback, ended and missing sessions, any other `SERVER_ERROR`, and that same interruption still standing once the bound is spent are all captured and announced normally; `FM_LAVISH_POLL_RETRY_DELAY` is a bounded 1 to 60 second test override for the interval only, and the runner itself stays adapter-agnostic. An already-armed Lavish source keeps its registered listener command until it is retired and armed again, so re-arm a live board once to adopt this retry policy. The `when` adapter (`bin/fm-procevent-when.sh`) turns this channel into a condition->action primitive: it registers a deterministic condition and a deterministic action once, its blocking child polls the condition without waking firstmate, and a stable true fires the action at most once before one terminal outcome is durably captured and published as a wake that remains eligible for re-announcement until handled. @@ -1019,20 +1023,27 @@ Keyed captain answers from built-in adapters use one more seam of the same kind, Some built-in sources carry the captain's answer to a captain-held task, and what such an answer means is owned once by `bin/fm-captain-hold.sh`'s keyed-answer intake rather than by any channel. A built-in source bound with `bin/fm-captain-hold.sh bind` therefore has each captured result passed to `bin/fm-procevent-<adapter>.sh answers <result-file>`, and whatever that prints is piped straight into that intake. A binding can select one decision origin or the script's cross-origin mode; the command header owns the exact forms and key interpretation. -The built-in adapter reports only what the captain chose; the intake owns every rule about what happens next, so the runner names no adapter, parses no result, and carries no decision rule, and a future built-in source needs nothing here beyond an `answers` command and a binding. -Feeding is independent of handling: it never acknowledges a result and never suppresses a wake, because recording the answer is transcription while acting on it is firstmate's judgement. -An unbound built-in source, a built-in adapter with no `answers` command, and a failure on either side all leave the capture untouched and still announced. -External binding responses never enter this authority-bearing intake. +The built-in adapter reports only what the captain chose; the intake owns every rule about what happens next, so the runner names no adapter, parses no result, and carries no decision rule, and a future built-in answer source needs nothing here beyond an `answers` command and a binding. +The reserved Reconcile selection uses the parallel optional `reconciles` adapter command and binding-verified `reconcile-requests` intake rather than entering keyed answers; [`captain-hold-lifecycle.md`](captain-hold-lifecycle.md#reconcile-re-check-reality-never-a-blind-close) owns those semantics. +Feeding is independent of handling: it never acknowledges a result and never suppresses a wake, because recording the answer or request is transcription while acting on it is firstmate's judgement. +An unbound built-in source, a built-in adapter without the corresponding command, and a failure on either side all leave the capture untouched and still announced. +External binding responses never enter either authority-bearing intake. Ownership is machine-wide per canonical source, because separate homes can share one underlying source store. Claims live under `$XDG_STATE_HOME/firstmate/procevent-claims` (override with `FM_PROCEVENT_CLAIM_ROOT`). Each claim binds its caller-reported home and runner PID to a process identity, unique claim generation, exact registration-file generation, and resolved state-root identity. Registration, acquisition, replacement, retirement, and generation-bound release are serialized at one machine-wide boundary per source. A live identity-matched owner is never displaced, and release removes only the exact generation the caller acquired. -Retirement and orphan reconciliation signal a runner process group only while its recorded process identity still matches, or when the recorded leader is gone and only its own owned group survives. -A runner leads its own process group, so a claim counts as reclaimable only when that whole generation is gone: a crashed leader whose group still has members is not stale, and reconcile stops that surviving group and releases its generation before starting any replacement. +Retirement and orphan reconciliation select a runner process group for signalling only while its recorded process identity still matches and the live runner still leads that group. +A claim counts as reclaimable only when its owner is stale and an independent process-group check finds no members; a crashed leader or reused pid whose process group still has members cannot relax ownership cleanup, so reconcile preserves the claim without signalling the ambiguous group or starting a replacement. +Reclaiming a generation that IS gone is not gated on tidying its capture-reservation records. +Those records are keyed by claim token and every replacement claims a fresh one, so a leftover that can no longer be located - a state-root identity a claim recorded before its home was re-created, for example - is stale bytes rather than an ownership hazard. +Ordinary release and reclamation still attempt reservation cleanup and require it unless both owner staleness and whole-group absence prove the generation gone. +The narrow live-owner terminal-self-retirement path also attempts cleanup but tolerates its own still-in-flight reservation, which the runner removes on the normal end-of-capture path; exact home, PID, and claim-token ownership remains mandatory before the claim is released. If identity cannot be established for a live PID, or a surviving owned group cannot be proved stopped, the operation preserves the registration and claim for safe retry rather than adding a second owner. -A live PID whose identity no longer matches is a reused PID, so it is treated as stale and its process group is never signalled. +A live PID whose identity no longer matches is a reused PID, so cleanup refuses it before signalling. +Identity and process-group verification cannot be made atomic with signalling in portable shell: the reaper signals only a target it has verified as the recorded generation, but PID and group reuse remain possible in the narrow interval between verification and the signal. +Launch pacing is the primary host-wedge protection; watchdog cleanup is a backstop. Supported secondmate retirement preflights each target home's bounded `sweep-home` command before destructive teardown, snapshots its registrations outside the target, then runs the sweep at that home's final deletion or return boundary. If deletion or return fails, teardown restores those registrations and reconciles them before returning the refusal. @@ -1041,6 +1052,26 @@ The sweep retires local registrations and machine-wide claims whose recorded sta Teardown refuses with the home, lease, routing evidence, registrations, claims, and runners retained when identity is uncertain, ownership is unreadable or unreleased, or relevant state exists without a sweep-capable child script. Raw manual deletion of a Firstmate home is unsupported because it can orphan a blocking child. To recover, restore that home's tracked `bin/fm-procevent.sh`, run `FM_HOME=<home> <home>/bin/fm-procevent.sh sweep-home`, then rerun the supported teardown. +The owning-home lease below bounds how long such an orphan can run, but it is a backstop, not a substitute for the supported path. + +A runner is bound to the HOME that owns it, not to the one session that armed it. +That granularity is deliberate: a persistent source is meant to outlive the turn and the session that armed it, so binding a runner to its arming session would stop exactly the sources this mechanism exists to keep running. +Any activity in the same home refreshes the lease, so a replacement session, another watcher, or an ordinary inspection command keeps a runner of that home alive; a runner whose SOURCE is no longer wanted in a live home is stopped by reconcile when that source is retired, independently of the lease. +The lease is therefore the backstop for a home that is GONE - the torn-down test sandbox this change exists to bound - and not a per-session ownership check. +KNOWN LIMIT: while any activity continues in a home whose original owning session has ended, that activity refreshes the lease and a runner of that home keeps running until its source is retired or the home goes away. +Detaching a runner into its own process group is what lets a persistent source outlive the turn that armed it, and on its own it is also what lets a runner outlive its whole home: reparented to init, it keeps its blocking child - and every process that child spawns - running with nothing left to reap it. +So a home's process-event state carries a lease that registration, attached start, reconciliation, acknowledgement, and listing refresh, and the watcher's reconcile cycle is what keeps it fresh in a live home. +An attached public `start` continues refreshing the lease while its caller remains attached. +Each runner fails closed unless a small guard starts successfully beside it in a separate process group. +That guard accepts the lease only while the state root retains the device/inode identity recorded by the runner's claim, and stops the runner's whole process group after two consecutive checks cannot prove that identity and lease freshness. +The group signal reaches the blocking child and everything under it exactly as retirement does. +A runner exports the inherited `FM_PROCEVENT_IN_RUNNER` marker and every lease refresh is skipped under it, so a runner and its ordinary children do not certify their own owner, and the next reconcile in a live home simply starts a replacement runner. +That no-self-refresh rule is CONFUSED-AGENT-GRADE, the same deliberate captain-decided grade `bin/fm-lease-lib.sh` documents: it stops the accidental case this boundary exists for, an orphaned or test-scaffolding source tree that would otherwise keep its own owner alive. +A source that DELIBERATELY strips the marker from its environment can still refresh the lease, so adversarial-grade unforgeability is explicitly out of scope here and tracked as separate follow-up design work. +Scope is the owning state root and one runner generation, never a script or process name, so a live source in another home is untouched: that home refreshes its own lease. +`FM_PROCEVENT_OWNER_LEASE_SECONDS` (default 600, range 1..86400) is how long a runner keeps going with no sign of activity in its owning home, and `FM_PROCEVENT_OWNER_CHECK_SECONDS` (default 15, range 1..3600) is how often its guard re-reads the lease. +`FM_PROCEVENT_LAUNCH_FLOOR_SECONDS` (default 1, range 1..3600) is the minimum time between consecutive launches of one registration generation's stored command, bounding the launch rate of an immediately returning source during that lease window. +The generation's first launch is immediate, later launches share its monotonic pacing timestamp, a timestamp from before a reboot is treated as expired, and replacing the registration starts a fresh pacing generation. `FM_PROCEVENT_MAX_OUTPUT_BYTES` (default 1048576) bounds a single captured result while the source runs; oversized output is drained but truncated with a stderr notice rather than staged or published whole or dropped. @@ -1090,6 +1121,7 @@ FM_CONFIG_OVERRIDE= # alternate config dir, mainly for tests FM_PROC_ROOT_OVERRIDE= # alternate /proc root for Linux process-identity reads in fm-wake-lib.sh and fm-teardown.sh, mainly for tests FM_BACKEND= # optional runtime backend override for new spawns; tmux/herdr/zellij/orca/cmux support ship/design/scout spawns, codex-app is not accepted FM_TRACE_CONTEXT= # optional trace-context override; see "Trace context propagation" +FM_TASK_ID= # internal task-worker marker fm-spawn.sh exports into ship and scout panes, never set by hand; bin/fm-test-run.sh refuses to execute in the repository primary checkout while it is set HERDR_SESSION=default # herdr-only: named session for normal backend ops; not enough for destructive cleanup (docs/herdr-backend.md) FM_BACKEND_HERDR_SUBMIT_POLLS=6 # herdr-only: agent-state samples spread across each Enter attempt's budget when confirming a submit (docs/herdr-backend.md "Current transport behavior") FM_BACKEND_HERDR_SUBMIT_MIN_SLEEP=0.6 # herdr-only: minimum per-Enter confirmation budget before polling agent-state after an idle baseline @@ -1144,6 +1176,9 @@ FM_TOOL_UPDATE_BUDGET_SECS=20 # 1..120 seconds allowed for a whole watched-too FM_TOOL_UPDATE_NOW= # test override for the watched-tool sweep clock; the sweep budget still uses real time FM_PROCEVENT_MAX_OUTPUT_BYTES=1048576 # bound on one captured process-to-event result FM_PROCEVENT_CLAIM_ROOT= # machine-wide source claim root; default $XDG_STATE_HOME/firstmate/procevent-claims +FM_PROCEVENT_OWNER_LEASE_SECONDS=600 # how long a source runner keeps going with no activity in its owning home; 1..86400 +FM_PROCEVENT_OWNER_CHECK_SECONDS=15 # how often a runner's guard re-reads that lease; 1..3600 +FM_PROCEVENT_LAUNCH_FLOOR_SECONDS=1 # minimum interval between launches of one registration generation's source command; 1..3600 FM_WHEN_OUTPUT_TAIL_BYTES=8192 # bound on the command-output tail inside one condition->action outcome document FM_CODEX_WATCH_CHECKPOINT=180 # seconds per foreground watcher checkpoint in Codex primary supervision FM_CREW_STATE_NM_TIMEOUT=10 # seconds allowed per no-mistakes query inside fm-crew-state.sh; bin/fm-fleet-snapshot.sh derives its own value from its per-task bound for the reads it makes, so this override does not reach those, and that script's header owns the derivation @@ -1179,17 +1214,17 @@ FM_WATCH_REARM_RETRY_MAX_MS=4000 # Pi/OpenCode adapter cap for exponential con FM_WATCH_REARM_RETRY_LIMIT=5 # Pi/OpenCode adapter launch-failure retries before surfacing restoration failure FM_WATCH_CYCLE_LOG_MAX_BYTES=262144 # size cap for the arm-owned watcher lifecycle ledger FM_WATCH_CYCLE_LOG_KEEP_LINES=1000 # newest complete lifecycle rows considered when the ledger is capped -FM_WATCHER_STALE_GRACE=300 # defaults to FM_GUARD_GRACE; seconds a live watcher lock may have a stale beacon before re-arm errors +FM_WATCHER_STALE_GRACE=300 # defaults to FM_GUARD_GRACE if set, else the poll-derived grace (docs/turnend-guard.md "Guard grace and the poll cadence"); seconds a live watcher lock may have a stale beacon before re-arm errors FM_SIGNAL_GRACE=30 # seconds to coalesce nearby status and turn-end signals into one wake FM_TURNEND_CHURN_ABSORB_SECS=900 # longest one endpoint's bare turn-ends may be deferred on pane-churn evidence alone; only consulted when config/turnend-churn-absorb is present FM_CAPTAIN_RE='done:|needs-decision:|blocked:|failed:|PR ready|checks green|ready in branch|merged' # captain-relevant status regex; nonterminal progress verbs remain excluded even when their prose matches FM_CLASSIFY_PAUSED_VERB=paused # leading status verb for a declared external wait; excluded from FM_CAPTAIN_RE and distinct from blocked FM_STALE_ESCALATE_SECS=240 # idle seconds before a provably-working stale pane or a confidently dead parked-decision repeat escalates; other first-sighting stale states surface immediately unless they declare the pause verb FM_BUSY_TURN_MAX_SECS=3600 # maximum age of a busy pane's latest state/<id>.turn-ended marker, or its state/<id>.meta spawn record before any turn-boundary wake arrives, before the same wedge escalation used for a provably-working non-busy stale takes over; inspection-only, never an automatic interrupt or restart; a declared external wait or verified captain-held transfer takes the FM_PAUSE_RESURFACE_SECS recheck below instead. bin/fm-supervision-lib.sh owns the window, and bin/fm-fleet-snapshot.sh publishes it as supervision.watcher.quiet_allowance_seconds so the dashboard's Task activity signal judges quiet against this same tolerance instead of a constant of its own; docs/dashboard-inbox-policy.md owns that signal and what renders it today -FM_PAUSE_RESURFACE_SECS=3600 # seconds before a declared external wait or verified captain-held transfer re-surfaces for a recheck in the watcher, including a live busy pane past FM_BUSY_TURN_MAX_SECS, and before a declared external wait re-surfaces in the away-mode daemon, which ages its window against the crew's own latest status line rather than pane busy state so only a status append that stops declaring the wait ends that routing; the away-mode clock runs from when the hold began and is never restarted by the crew rewriting its paused reason +FM_PAUSE_RESURFACE_SECS=3600 # seconds before a declared external wait or verified captain-held transfer re-surfaces for a recheck in the watcher, and before repeated new-hash alarms for an ordinary crew task with an open backlog call, including a live busy pane past FM_BUSY_TURN_MAX_SECS, and before a declared external wait re-surfaces in the away-mode daemon, which ages its window against the crew's own latest status line rather than pane busy state so only a status append that stops declaring the wait ends that routing; the away-mode clock runs from when the hold began and is never restarted by the crew rewriting its paused reason FM_PAUSE_RESURFACE_MAX_STREAK=3 # how many times an unchanged declared wait may double the WIDTH of its recheck window before the cadence stops widening; 0 restores a fixed FM_PAUSE_RESURFACE_SECS cadence, and the streak restarts whenever the wait itself changes, which only ever narrows the window back toward the base. Clamped internally to 12 doublings and a one-day window, so a misconfigured value cannot overflow into a permanent re-surface -FM_SECONDMATE_WAKE_STALL_SECS=1800 # bare-time backstop: minimum age of the oldest valid foreign wake-queue row before an endpoint-recorded local secondmate whose busy verdict cannot be ESTABLISHED produces one durable parent wake-loop-stall notification; zero or invalid values use 1800. A row stays unacknowledged for the whole handling turn by design, so elapsed time alone cannot separate a stalled loop from a working one and this sits above a real turn, at the same tolerance FM_RUN_STRANDED_SILENCE_SECS already carries for the same 10-18 minute single-step evidence. A provably busy mate never alarms here at any age -FM_SECONDMATE_WAKE_IDLE_STALL_SECS=60 # the shorter age used instead when that mate is PROVABLY idle, where nothing is running to handle the row; a proven verdict outranks elapsed time, so a genuinely dead wake loop is still caught in about a minute. Clamped to FM_SECONDMATE_WAKE_STALL_SECS, so an idle mate is never treated as less suspicious than an unestablished one; zero or invalid values use 60 +FM_SECONDMATE_WAKE_STALL_SECS=1800 # no-progress interval for an oldest actionable foreign queue identity when the semantic busy verdict is unknown, dead, or unreadable; also used after proven busy exceeds FM_BUSY_TURN_MAX_SECS. Progress resets the interval, declared external-wait rows are excluded, and one notification covers each episode; zero or invalid values use 1800 +FM_SECONDMATE_WAKE_IDLE_STALL_SECS=60 # shorter no-progress interval when the mate is provably idle; uncertainty never earns this deadline. Clamped to FM_SECONDMATE_WAKE_STALL_SECS; zero or invalid values use 60 FM_WEDGE_DEMAND_INSPECT_COUNT=3 # consecutive provably-working stale escalations on the same unchanged pane before demand-deep-inspection is added FM_RUN_STRANDED_SILENCE_SECS=1800 # how long an actively-executing no-mistakes step may report NO activity before a wedge escalation stops being held for it. Both supervisors consult bin/fm-run-progress.sh at the escalation point only, so a crew parked on a run that is still moving stops alarming for as long as the hold lasts (bounded by FM_RUN_PROGRESS_HOLD_MAX below), while a stranded run and a confidently dead agent still alarm. docs/architecture.md owns the always-on watcher's no-evidence fallback for a surviving declared wait. Above the pipeline's own 10m step_quiet_warning on purpose: that marker is a liveness clue, and review or test steps routinely go 10-18 minutes on one opening line. Raising it delays a real wedge alarm by the same amount FM_RUN_PROGRESS_NM_TIMEOUT=10 # seconds allowed for that bounded `no-mistakes axi status` read; a read that cannot complete is no evidence and never positive evidence that may hold a wedge escalation diff --git a/docs/fm-test-portable-shards.md b/docs/fm-test-portable-shards.md index c0b44443b69..07ab62d89ab 100644 --- a/docs/fm-test-portable-shards.md +++ b/docs/fm-test-portable-shards.md @@ -5,35 +5,39 @@ ## Verification inputs -The current candidate timings came from the 2026-08-20 concurrent proof recorded in [fm-test-isolation-proof.md](fm-test-isolation-proof.md). -The proof ran 24 candidates with four workers and no failures. - -| duration_ms | script | -|---:|---| -| 45356 | `tests/fm-backend-herdr.test.sh` | -| 35415 | `tests/fm-x-mode.test.sh` | -| 35095 | `tests/fm-captain-hold-lifecycle.test.sh` | -| 27529 | `tests/fm-arm-pretool-check.test.sh` | -| 20922 | `tests/fm-test-run.test.sh` | -| 17558 | `tests/fm-crew-state.test.sh` | -| 16582 | `tests/fm-cd-pretool-check.test.sh` | -| 9766 | `tests/fm-lint.test.sh` | -| 9562 | `tests/fm-herdr-lab.test.sh` | -| 6768 | `tests/fm-grok-harness.test.sh` | -| 6290 | `tests/fm-pr-merge.test.sh` | -| 5569 | `tests/fm-composer-ghost.test.sh` | -| 4563 | `tests/fm-send-popup-settle.test.sh` | -| 4021 | `tests/fm-tmux-submit-busy.test.sh` | -| 3544 | `tests/fm-composer-lib.test.sh` | -| 3025 | `tests/fm-send-strict.test.sh` | -| 2753 | `tests/fm-send-settle.test.sh` | -| 2166 | `tests/fm-review-diff.test.sh` | -| 1315 | `tests/fm-brief.test.sh` | -| 975 | `tests/fm-spawn-batch.test.sh` | -| 598 | `tests/fm-pi-primary-types.test.sh` | -| 513 | `tests/fm-ensure-agents-md.test.sh` | -| 331 | `tests/fm-supervision-instructions.test.sh` | -| 99 | `tests/fm-transition-lib.test.sh` | +The proven-isolated candidate set remains the 24-script concurrent proof recorded in [fm-test-isolation-proof.md](fm-test-isolation-proof.md). +Placement timings are refreshed from completed script measurements in [CI run 34800278334](https://github.com/HelloWorldSungin/firstmate/actions/runs/34800278334). +Its first parallel lane reached the unchanged ten-minute job limit after passing nine scripts; the old two-minute estimate no longer represented the enlarged suites. +The second parallel lane completed successfully. +The two scripts not completed before the first lane was cancelled retain the passing local full-sweep measurements shown explicitly below. +These measurements change placement, not isolation eligibility or execution deadlines. + +| duration_ms | script | Measurement | +|---:|---|---| +| 203500 | `tests/fm-captain-hold-lifecycle.test.sh` | CI completed script | +| 179803 | `tests/fm-lint.test.sh` | CI completed script | +| 150192 | `tests/fm-test-run.test.sh` | CI completed script | +| 122702 | `tests/fm-pr-merge.test.sh` | CI completed script | +| 32818 | `tests/fm-crew-state.test.sh` | CI completed script | +| 28065 | `tests/fm-x-mode.test.sh` | CI completed script | +| 26899 | `tests/fm-arm-pretool-check.test.sh` | CI completed script | +| 21249 | `tests/fm-backend-herdr.test.sh` | CI completed script | +| 15995 | `tests/fm-cd-pretool-check.test.sh` | CI completed script | +| 14889 | `tests/fm-brief.test.sh` | Local full sweep, 2026-09-12 | +| 6983 | `tests/fm-grok-harness.test.sh` | CI completed script | +| 6788 | `tests/fm-send-strict.test.sh` | CI completed script | +| 6351 | `tests/fm-herdr-lab.test.sh` | CI completed script | +| 4464 | `tests/fm-send-popup-settle.test.sh` | CI completed script | +| 4220 | `tests/fm-pi-primary-types.test.sh` | CI completed script | +| 4219 | `tests/fm-composer-lib.test.sh` | CI completed script | +| 3832 | `tests/fm-review-diff.test.sh` | CI completed script | +| 2271 | `tests/fm-tmux-submit-busy.test.sh` | CI completed script | +| 2241 | `tests/fm-spawn-batch.test.sh` | CI completed script | +| 2049 | `tests/fm-composer-ghost.test.sh` | CI completed script | +| 1834 | `tests/fm-send-settle.test.sh` | CI completed script | +| 810 | `tests/fm-ensure-agents-md.test.sh` | CI completed script | +| 307 | `tests/fm-supervision-instructions.test.sh` | CI completed script | +| 93 | `tests/fm-transition-lib.test.sh` | Local full sweep, 2026-09-12 | ## Parallel lanes @@ -41,16 +45,16 @@ The two parallel lanes use longest-processing-time assignment from those measure | Lane | Script count | Estimated duration | |---|---:|---:| -| `portable-parallel-1` | 11 | 134295 ms (~134.3 s) | -| `portable-parallel-2` | 13 | 126020 ms (~126.0 s) | -| imbalance | | 8275 ms | +| `portable-parallel-1` | 11 | 421533 ms (~7.03 min) | +| `portable-parallel-2` | 13 | 421041 ms (~7.02 min) | +| imbalance | | 492 ms | `bin/fm-test-run.sh` contains the exact ordered memberships in `list_portable_parallel_1` and `list_portable_parallel_2`. ## Portable serial remainder -`portable-serial` includes every `tests/*.test.sh` that is neither proven-isolated, `real-herdr-gated`, nor `live-harness-optin`. -It keeps watcher, lock, AFK, real tmux, daemon, secondmate lifecycle, bootstrap, GUI-backend, and other unproven work serial, while `live-harness-optin` stays an explicit opt-in outside every portable lane and therefore outside every serial CI shard, because its members need machine state CI does not have - real harness credentials, or a real browser session. +`portable-serial` includes every `tests/*.test.sh` that is neither proven-isolated nor `real-herdr-gated`. +It keeps watcher, lock, AFK, real tmux, daemon, secondmate lifecycle, bootstrap, the `live-harness-optin` family, GUI-backend, and other unproven work serial. Membership is derived rather than enumerated, so a newly added test lands here by default. ## Portable serial CI shards @@ -66,9 +70,12 @@ Each shard is still strictly serial in itself, and separate runners mean no two Assignment is longest-processing-time bin packing over per-script duration hints embedded in `bin/fm-test-run.sh`. The hints are the slowest measurement of each of the lane's 139 scripts across the `fm-test-timing-portable-serial-*` artifacts of three green CI runs on 2026-09-01, [33558082172](https://github.com/kunchenguid/firstmate/actions/runs/33558082172), [33523597838](https://github.com/kunchenguid/firstmate/actions/runs/33523597838), and [33463326167](https://github.com/kunchenguid/firstmate/actions/runs/33463326167). Shared scripts use those upstream per-script maxima; fork-only scripts retain their existing measured hints from run [32191955185](https://github.com/HelloWorldSungin/firstmate/actions/runs/32191955185). -The current fork partition has 161 scripts, with seven unmeasured scripts using the conservative `PORTABLE_SERIAL_DEFAULT_WEIGHT_MS` default, for 4257656 ms of estimated balance weight. -Taking the slowest of several runs rather than a single run keeps the balance honest on a slow runner: individual scripts varied by up to 20% between those three runs. -A script with no hint gets the conservative `PORTABLE_SERIAL_DEFAULT_WEIGHT_MS` default. +The round also includes upstream's retained native-Windows measurement for `tests/fm-pi-windows-shell-invocation.test.sh` and its new live-guard weights. +The 2026-09-14 refresh takes the maximum of those retained hints and completed passing-script measurements from fork CI runs [34802687283](https://github.com/HelloWorldSungin/firstmate/actions/runs/34802687283) and [34804089007](https://github.com/HelloWorldSungin/firstmate/actions/runs/34804089007). +The latter's serial lane 5 reached its unchanged 20-minute cap after 16 passing scripts; its completed `FM_TEST_END` records are included explicitly because cancellation prevented a timing artifact. +These runs provide passing measurements for 200 of the 201 serial scripts; `tests/fm-test-isolation-proof.test.sh` retains its earlier 2567 ms hint, with its corrected assertion passing locally. +Taking maxima preserves native-Windows measurements and earlier slow-run evidence rather than replacing them with portable gate-skip durations. +A script with no hint receives `PORTABLE_SERIAL_DEFAULT_WEIGHT_MS`; the runner's coverage output reports that unmeasured share. Hints only affect balance: the coverage guard keeps the partition complete and disjoint whatever they say, so a stale hint costs a slower shard rather than lost coverage. Balance is still worth keeping current, because enough unmeasured scripts let one shard carry more than twice another shard's real work and reach the job cap while another runner sits idle. That is not hypothetical: by 2026-09-01 the lane had grown from 116 to 139 scripts and from ~42 to ~63 minutes, 17 scripts were still unmeasured, and several hints were low by 2-5x, so shard 3 of 4 ran 17-20 minutes against its 20-minute cap while shard 1 ran 11.5 minutes and run [33574154856](https://github.com/kunchenguid/firstmate/actions/runs/33574154856) timed out seconds after a passing test. @@ -77,23 +84,26 @@ Refresh the hints whenever the serial lane gains scripts, rather than waiting fo Shard count is sized from that total rather than left where an earlier, smaller remainder put it. The lane grew from about 19 minutes across 69 scripts to about 58 minutes across 154, which four shards could no longer carry inside the job timeout: on the run above, `portable-serial-2of4` was cancelled at 15 minutes having finished 24 of its 32 scripts, and the hints then put a perfectly balanced quarter at 14.5 minutes, still on the tripwire rather than inside it. -The refreshed shared hints and retained fork-only hints now put the slowest of eight shards at about 8.87 minutes. +The refreshed shared hints and retained fork-only hints now put the slowest of eight shards at about 12.89 minutes. | Lane | Script count | Estimated duration | |---|---:|---:| -| `portable-serial-1of8` | 19 | 532198 ms (~8.87 min) | -| `portable-serial-2of8` | 19 | 532184 ms (~8.87 min) | -| `portable-serial-3of8` | 19 | 532203 ms (~8.87 min) | -| `portable-serial-4of8` | 21 | 532231 ms (~8.87 min) | -| `portable-serial-5of8` | 21 | 532207 ms (~8.87 min) | -| `portable-serial-6of8` | 20 | 532175 ms (~8.87 min) | -| `portable-serial-7of8` | 21 | 532233 ms (~8.87 min) | -| `portable-serial-8of8` | 21 | 532225 ms (~8.87 min) | -| imbalance | | 58 ms | - -The single longest script, `tests/fm-watch-triage.test.sh` at 262626 ms, is the floor for any shard count. - -Refresh the hints by downloading the per-shard timing artifacts from several green CI runs, replacing the `portable_serial_weight_hints` table in `bin/fm-test-run.sh` with the slowest measured `duration_ms` per `path`, and updating the table above: +| `portable-serial-1of8` | 24 | 773453 ms (~12.89 min) | +| `portable-serial-2of8` | 25 | 773490 ms (~12.89 min) | +| `portable-serial-3of8` | 26 | 773503 ms (~12.89 min) | +| `portable-serial-4of8` | 25 | 773453 ms (~12.89 min) | +| `portable-serial-5of8` | 25 | 773452 ms (~12.89 min) | +| `portable-serial-6of8` | 25 | 773452 ms (~12.89 min) | +| `portable-serial-7of8` | 25 | 773453 ms (~12.89 min) | +| `portable-serial-8of8` | 26 | 773502 ms (~12.89 min) | +| imbalance | | 51 ms | + +The watcher triage cases are split into core and wait/decision scripts with one shared fixture owner in `tests/watch-triage-helpers.sh`. +All 125 original cases remain in exactly one script. +Their initial passing local measurements were 160205 ms and 191775 ms after CI reached the unchanged 480-second combined-script limit while still passing cases. +The refreshed CI hints are 236467 ms for core and 292716 ms for waits. + +Refresh the CI-derived hints by downloading the per-shard timing artifacts from several green CI runs, replacing the `portable_serial_weight_hints` table in `bin/fm-test-run.sh` with the slowest measured `duration_ms` per `path`, and updating the table above: ```sh for run in <run-id> <run-id> <run-id>; do @@ -106,11 +116,12 @@ bin/fm-test-run.sh --check-coverage ``` A timed-out shard uploads no artifact, so pick runs where every serial shard is green or the lane's slowest scripts go unmeasured in exactly the shard that needs them most. +Measure native-Windows-only scripts through the focused Git Bash runner and retain that `duration_ms` separately, because the portable CI shards skip them. ## Coverage guard `bin/fm-test-run.sh --check-coverage` verifies that both parallel lanes partition the proven-isolated set. -It also verifies that the parallel lanes, portable serial lane, real-Herdr family, and live-harness opt-in family are disjoint and together cover every `tests/*.test.sh` script. +It also verifies that the parallel lanes, portable serial lane, and real-Herdr family are disjoint and together cover every `tests/*.test.sh` script. It separately verifies that the portable serial CI shards are non-empty, disjoint, and together equal the portable serial lane. It reports the unmeasured serial share as `serial_unhinted=` and refuses when that share exceeds `PORTABLE_SERIAL_MAX_UNHINTED_PERCENT`, so the shards stay balanced on evidence rather than on the default weight. @@ -129,8 +140,8 @@ Portable shards, each portable serial shard, and the Herdr lane upload runner-ge | Lane | Bound | Rationale | |---|---|---| -| portable parallel 1/2 | job `timeout-minutes: 10` | The measured shard sums are about 2.2 minutes and the timeout is a hang tripwire. | -| portable serial 1-8 | job `timeout-minutes: 20` | The slowest estimated shard is about 8.87 minutes, leaving roughly 2.3x margin for setup and runner-speed spread. | +| portable parallel 1/2 | job `timeout-minutes: 10` | The measured shard sums are about 7 minutes and the timeout is a hang tripwire. | +| portable serial 1-8 | job `timeout-minutes: 20` | The slowest estimated shard is about 12.89 minutes, leaving roughly seven minutes for setup and runner-speed spread. | | Herdr | family-run step `timeout-minutes: 20`; job `timeout-minutes: 75` backstop | The required lane is bounded independently of the per-script deadline; refresh timings from its uploaded artifacts. Previous healthy runs finished around 7 minutes, so the step bound is the hang tripwire (cleanup and timing artifacts still upload) while the job cap stays a last-resort backstop. | Timeouts are hang tripwires rather than expected healthy durations. @@ -138,6 +149,7 @@ Timeouts are hang tripwires rather than expected healthy durations. Inside each lane, `bin/fm-test-run.sh` applies its own default per-script bound, so a hung script usually turns red with per-script attribution before the job cap cancels the lane; its `--help` owns that bound's value and opt-out, and the rationale beside `DEFAULT_PER_SCRIPT_TIMEOUT_SECS` owns the per-lane margin arithmetic. Neither portable lane has room to spare, because a hung script spends the bound instead of its own healthy slot. -On the slowest estimated serial shard, replacing its average script with the 480-second bound puts script time around 16.3 minutes, inside the 20-minute cap before checkout and bootstrap overhead. +On the slowest estimated serial shard, replacing its average script with the 480-second bound puts script time around 20.4 minutes, past the 20-minute cap before checkout and bootstrap overhead. +A hung script may therefore reach the job limit before the per-script limit can report it; the healthy estimate retains roughly seven minutes for setup and runner-speed variation. The portable parallel cap is tighter still: the same arithmetic already lands past its 10-minute cap before setup, so expect the job timeout rather than per-script attribution when a script hangs there. On the required Herdr lane the bound has the thinnest margin over its slowest measured script, so a healthy but unusually slow Herdr end-to-end script can turn red as `exit=124`; that margin is accepted rather than widened, tracked in `HelloWorldSungin/firstmate#256`. diff --git a/docs/fork-divergence.md b/docs/fork-divergence.md index b99e528c707..82a1814d624 100644 --- a/docs/fork-divergence.md +++ b/docs/fork-divergence.md @@ -83,6 +83,25 @@ Default disposition also means the herdr event wait's command substitution can b Upstream's `fm_backend_herdr_wait_transition` always allocates that directory with its own `mktemp -d`, so the caller-owned staging and its fail-closed return for a directory the caller has already released sit inside upstream-owned code in [`bin/backends/herdr.sh`](../bin/backends/herdr.sh) and the contract comment in [`bin/fm-backend.sh`](../bin/fm-backend.sh), with [`tests/fm-backend-herdr.test.sh`](../tests/fm-backend-herdr.test.sh) pinning both halves. [`verification/supervision.md`](verification/supervision.md#watcher-stop-disposition) owns the measurements, the narrowed `HelloWorldSungin/firstmate#160` burst window, and the residual this leaves. +### Keyed decision repaint suppression + +The fork suppresses pane-repaint repeats for an already-surfaced explicitly keyed open-decision set, while retaining dead-agent wedge recovery. +`bin/fm-push-transition-lib.sh` owns the surfaced identity, using `bin/fm-classify-lib.sh`'s explicit-only view without changing the durable open-decision fold. +Upstream has no equivalent one-shot suppression for keyed decisions and instead bounds stale alarms when a backlog call is open. +Firstmate aligned the fork suppression with its documented explicit-key boundary during the collision with `kunchenguid/firstmate#3842`: an unkeyed blocker without a backlog hold retains upstream's repaint alarms. +An explicit `[key=default]` remains keyed, while an implicit default bucket does not acquire suppression merely by remaining durably open. +`docs/architecture.md` owns the wake contract, and `tests/fm-watch-triage.test.sh` proves keyed suppression, unkeyed re-alarming, dead-agent recovery, and durable retention after acknowledgement. + +### Secondmate queue-stall semantic thresholds + +The fork retains semantic busy classification through `fm_busy_classify_live` and separate 60-second proven-idle and 1800-second unknown-state thresholds in `bin/fm-watch.sh`. +Upstream `kunchenguid/firstmate#3943` uses a single 180-second no-progress threshold and a rendered active-turn check. +Firstmate preserved the fork's classification and thresholds because an unreadable agent is not evidence of idleness and a shorter unknown-state deadline loses the existing false-alarm protection. +Upstream's epoch-sequence progress identity, reset on progress, declared-wait exclusion, and once-per-episode notification are adopted in that same owner. +Busy suppression is bounded by `FM_BUSY_TURN_MAX_SECS`; after that boundary the unknown-state threshold permits an inspection without relabeling busy as idle. +`docs/architecture.md` owns the mechanism description, `docs/configuration.md` owns the settings, and `tests/fm-wake-queue.test.sh` exercises progress, busy, idle, unknown, reused endpoint, and bounded suppression behavior. +This divergence was previously unrecorded after its introduction in `HelloWorldSungin/firstmate#216`. + ### Run-progress wedge hold The fork carries [`bin/fm-run-progress.sh`](../bin/fm-run-progress.sh), which upstream has no equivalent of, and threads a per-pane hold-count file through `wedge_timer_check` in [`bin/fm-watch.sh`](../bin/fm-watch.sh) so a wedge escalation is held while the crew's validation run is demonstrably still moving. @@ -106,6 +125,10 @@ Upstream `kunchenguid/firstmate#3532` adds a declaration-scoped throttle for a l That throttle is adopted while a worker's own `paused:` declaration keeps the fork's immediate absorption, widening cadence, and reset to the replacement wait's base deadline. +Upstream `kunchenguid/firstmate#3842` adds backlog-backed call identity to the stale-alarm throttle. +That identity and its release/re-hold reset are adopted while a worker's own paused declaration still routes directly to the fork's widening-cadence owner. +The backlog read remains outside the secondmate ordinary-poll path. + ### Herdr pre-Enter footer read on a native working baseline The fork skips the pre-Enter rendered-footer read entirely when herdr's native agent-state baseline is already `working`, because the rendered-footer conversion refuses a `working` baseline outright and the read can therefore produce no verdict. @@ -149,6 +172,8 @@ The rationale beside the constant owns why 480s and what the bound costs each CI The upstream `kunchenguid/firstmate#3489` rebalance refreshes shared duration hints and raises the portable serial job cap to 20 minutes. The fork retains eight shards and its fork-only timing hints, while adopting that cap; the runner still owns the 480-second per-script bound and the updated margin arithmetic. +Watcher triage cases are partitioned between `tests/fm-watch-triage.test.sh` and `tests/fm-watch-triage-waits.test.sh`, with shared case definitions in `tests/watch-triage-helpers.sh`, so suite growth does not weaken the per-script bound. + ### Locale-independent test coverage comparisons The fork retains the C-collation invariant owned by the coverage guard in [`bin/fm-test-run.sh`](../bin/fm-test-run.sh). @@ -161,6 +186,9 @@ The fork retains pending-queue supervision through the shared predicate in [`bin The new presentation-deadline case from upstream `kunchenguid/firstmate#3475` must therefore allow the deliberate guard advisory while continuing to reject helper-process diagnostics. [`tests/fm-wake-queue.test.sh`](../tests/fm-wake-queue.test.sh), [`tests/fm-guard-stale-banner.test.sh`](../tests/fm-guard-stale-banner.test.sh), and [`tests/fm-turnend-guard.test.sh`](../tests/fm-turnend-guard.test.sh) preserve the queued-wake cause and the guarded handling behavior. +Upstream `kunchenguid/firstmate#3860` adds registered custom checks as another supervision cause in the same predicate. +Its `kunchenguid/firstmate#3950` actor-specific pending count controls which drain warning an actor receives, while the pending queue still requires supervision even when a live branch owns its rows. + ### No-mistakes run attribution The fork rewrote the branch-and-code-identity rule that binds a no-mistakes run to a task into a named relation table in [`bin/fm-nm-run-lib.sh`](../bin/fm-nm-run-lib.sh) - `fm_nm_head_relation` returning `equal`, `run-ahead`, `run-behind`, `unresolved`, `missing`, or `diverged`, consumed by `fm_nm_head_attributable` - and built the surrounding current-state surface upstream has no equivalent of: the `abandoned` verdict, the degraded run-step replay, and the recorded `branch=` task identity that makes a disagreeing ambient branch an attribution fault rather than another task's run. @@ -197,6 +225,8 @@ Upstream's byte-identity check between the promotion and brief paths stays green The handoff itself went unrecorded here from its introduction until 2026-09-04. The intent/spec split and current intent overlay from `kunchenguid/firstmate#3597` and `kunchenguid/firstmate#3671` apply to design workers too, while the fork's delivery fragments remain the rendering owner. +The fork update in [HelloWorldSungin/firstmate#272](https://github.com/HelloWorldSungin/firstmate/pull/272) keeps the design task in one planning conversation, owned by [`design-profile`](../.agents/skills/design-profile/SKILL.md), and preserves its dispatch-pinned skill release on relaunch, owned by [`docs/fleet-data-contracts.md`](fleet-data-contracts.md#the-design-tasks-plugin-release). +`tests/fm-brief.test.sh`, `tests/fm-design-skills.test.sh`, and `tests/fm-control-relaunch.test.sh` verify that composition alongside upstream task-copy isolation. `bin/fm-spawn.sh` reads structural branch and work-item identity from the authored brief, not the derived launch overlay that can repeat intent text. ### A preserving refusal withdraws the pending backlog close @@ -257,12 +287,6 @@ Upstream's OMP watcher port in `kunchenguid/firstmate#3867` receives the same st `tests/fm-omp-harness.test.sh` exercises flag-present launch, child retirement, pending-wake preservation, one-cycle return, and replacement through the extension interface. This is the existing `.afk` ownership contract, not the later upstream AFK-posture design, and portable extension tests do not establish live OMP vendor compatibility. -### Local lock scope for remotely parented homes - -The Treehouse project lock added by `kunchenguid/firstmate#3837` must coordinate a home's locally registered descendants without requiring a parent on another machine to be locally addressable. -The fork's `fm_firstmate_root_home` in [`bin/fm-wake-lib.sh`](../bin/fm-wake-lib.sh) stops at a validated remote parent boundary, where upstream refuses the traversal and consequently refuses every fresh Treehouse worker spawn from that home. -Local parent traversal, malformed-binding and cycle refusal, project identity, and slot-exclusivity checks remain intact. -[`tests/fm-spawn-worktree-settle.test.sh`](../tests/fm-spawn-worktree-settle.test.sh) drives worker creation beneath a remote root, verifies that local peers share the same project lock, and rejects malformed or cyclic bindings before allocation. ### Herdr presentation fixture ownership and cleanup @@ -381,7 +405,8 @@ Its recurring merge cost is likewise attachment rather than fork-only files: [`b Teardown publishes both of those fleet-owned records inside the upstream-owned [`bin/fm-teardown.sh`](../bin/fm-teardown.sh): a best-effort refresh of the task's usage sessions that never blocks cleanup, and the completion manifest composed while the volatile records feeding it still exist; the same opt-in refresh is wired into the upstream-owned [`bin/fm-bootstrap.sh`](../bin/fm-bootstrap.sh) at a locked session boundary, bounded by `FM_BOOTSTRAP_USAGE_TIMEOUT` and reporting on its own `USAGE_STORE:` diagnostic line. The manifest publication deliberately blocks the lifecycle - teardown refuses to erase a task whose manifest could not be written, because a task that cannot be archived must not be erased - so a merge must preserve that refusal rather than relax it into a best-effort skip on the strength of this entry's off-the-critical-path posture. [`bin/fm-fleet-snapshot.sh`](../bin/fm-fleet-snapshot.sh) and its `fm-fleet-snapshot.v1` contract are upstream-owned, so the divergence is the dashboard-consumed fields the fork adds to that snapshot, among them the quiet window owned in [`bin/fm-supervision-lib.sh`](../bin/fm-supervision-lib.sh) and consumed by [`bin/fm-watch.sh`](../bin/fm-watch.sh), published as `quiet_allowance_seconds` so the dashboard judges quiet on supervision's own tolerance rather than inventing a second one. -[`.github/workflows/ci.yml`](../.github/workflows/ci.yml) pins the Node 22 floor the usage collector requires, and [`bin/fm-test-run.sh`](../bin/fm-test-run.sh) classifies `tests/fm-dashboard-browser.test.sh` into the opt-in `live-harness-optin` family so the one suite needing a real browser stays out of every portable lane. +[`.github/workflows/ci.yml`](../.github/workflows/ci.yml) pins the Node 22 floor the usage collector requires, and [`bin/fm-test-run.sh`](../bin/fm-test-run.sh) classifies `tests/fm-dashboard-browser.test.sh` into the opt-in `live-harness-optin` family whose portable-serial membership leaves execution gated by the suite's explicit `FM_DASHBOARD_BROWSER_E2E=1` opt-in. +Ordinary CI inventories this guard but skips browser execution, while installed token-free guards run under upstream's shared live-gate policy. `tests/lib.sh` redirects the event config, the event store, and the dashboard credentials into a throwaway isolation root so no fixture home can reach a developer's real instrumentation, and `docs/configuration.md`, `docs/documentation-audiences.json`, and `docs/scripts.md` are the documentation catalogs this subsystem shares with the GBrain entry above. A merge must preserve the additive wiring property that [`dashboard-events.md`](dashboard-events.md) owns: instrumentation is off until enabled, the wiring is absent entirely when off, and the emitter exits 0 on every path, writes nothing to either stream, and does its work in a detached child, so an entry firing beside the turn-end guard or the watcher auto-arm can never change that guard's exit status, output, timing, or ordering. [`tests/fm-dashboard-events.test.sh`](../tests/fm-dashboard-events.test.sh) pins that by running the real guards in two identical homes with the emitter firing against a deliberately hung dashboard and requiring an identical decision from both. @@ -397,6 +422,13 @@ The snapshot header owns that mode, `tests/fm-bearings-snapshot.test.sh` proves ## Retired divergences +### Local lock scope for remotely parented homes - retired 2026-09-11 + +Upstream `kunchenguid/firstmate#3883` now terminates the local-root traversal at a validated remote parent boundary, supplying the fork's existing lock-scope behavior. +The shared implementation retains malformed-binding, local-parent reachability, cycle, and depth refusals. +`tests/fm-spawn-worktree-settle.test.sh` preserves the fork's remote-root worker and local-peer lock cases, while upstream's `tests/fm-teardown-endpoint-safety.test.sh` adds allocation/return serialization and child-slot ownership coverage. +The fork no longer needs a separate implementation of the parent-route branch. + ### Pre-move crash fixture - retired 2026-09-11 The fork's pre-move crash correction is now supplied by upstream `kunchenguid/firstmate#3644`. diff --git a/docs/pi-supervision-branch.md b/docs/pi-supervision-branch.md index 144d0c76f5d..76cb84a8b85 100644 --- a/docs/pi-supervision-branch.md +++ b/docs/pi-supervision-branch.md @@ -144,6 +144,8 @@ Every other fleet-wide or unresolvable wake - including watcher-failure alarms, The captain accepted the normal provider prompt-caching strategy: a byte-identical branch prefix generated once per firstmate version, the same tool set in the same order on every request, and one shared `prompt_cache_key` per home for all branch sessions (set in a `before_provider_request` hook, and only for providers whose requests already carry that field); main keeps its own per-session key. Budget roughly 60% cache hits on a new branch conversation's first call and 95% on later calls within that conversation; the shared per-home key is what carries the byte-identical prefix across the conversation each main session start opens, and reuse is best-effort, never guaranteed. The branch can also run on a cheaper model and a shallower reasoning effort than main, both pinned with the Pi `/supervision-model` command; [configuration.md](configuration.md#pi-supervision-branch-model-and-effort-configsupervision-branch-model-configsupervision-branch-effort) owns those pins' operator-facing schema and unpinned behavior. +A provider an extension registered only into main's runtime, such as pi-devin-auth's `devin`, reaches the isolated branch runtime by copying its provider config from main's captured `ModelRegistry` into the branch `ModelRuntime` at model-resolution time and in the `/supervision-model` picker, so the provider's own `streamSimple` transport and OAuth wiring are reused by reference rather than reimplemented. +That carve-out is scoped to provider registration alone: the branch keeps its `noExtensions`, `noSkills`, and `noContextFiles` isolation, the copy is never persisted, a provider whose registration fails to compose is simply unavailable, and `tests/fm-pi-branch-extension.test.sh` pins the pin-and-fallthrough behavior. No caching machinery beyond this exists, deliberately: any later dynamic content in the branch prefix silently removes most of the cache benefit, which is why `bin/fm-branch-prompt.sh`'s header is the contract's single owner and `tests/fm-branch-supervision.test.sh` pins the output to byte identity. ## Away mode diff --git a/docs/scripts.md b/docs/scripts.md index 421d150d086..24084d1d894 100644 --- a/docs/scripts.md +++ b/docs/scripts.md @@ -49,7 +49,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-install-herdr.sh` | Install CI's exact-version Herdr pin with official asset URL, SHA-256, and protocol checks | | `fm-install-treehouse.sh`| Install CI's exact-version Treehouse pin for real-Herdr E2E that needs spawn worktrees | | `fm-herdr-ci-cleanup.sh` | Snapshot and tear down only job-owned `fm-lab-*` sessions in the Herdr CI lane | -| `fm-test-run.sh` | Behavior-test runner: selection, portable lanes, bounded concurrency, budgets, coverage guard, timing/JSON | +| `fm-test-run.sh` | Behavior-test runner: selection, portable lanes, bounded concurrency, budgets, coverage guard, timing/JSON; refuses to execute in the repository primary checkout when `FM_TASK_ID` marks a task worker | | `fm-test-isolation-proof.sh` | Concurrent isolation harness and portable candidate set owner | | `fm-ensure-agents-md.sh` | Ensure a project's real `AGENTS.md`, its `CLAUDE.md` `@AGENTS.md` pointer, and self-governance guidance (explicit project mark documented in the helper's header and help) | | `fm-guard.sh` | Warn on primary-checkout tangles, main-session pending wakes, and unhealthy supervision | diff --git a/docs/supervision-protocols/grok.md b/docs/supervision-protocols/grok.md index f27ae302e13..305e1802a16 100644 --- a/docs/supervision-protocols/grok.md +++ b/docs/supervision-protocols/grok.md @@ -25,7 +25,7 @@ When you see a background-task-completed system reminder for the arm: 1. Run `bin/fm-wake-drain.sh` first. 2. Optionally fetch arm output with `get_command_or_subagent_output(<task_id>)` for the reason line. 3. Handle `signal`, `stale`, `check`, or `heartbeat` using the harness-neutral contract in `AGENTS.md`. -4. Ordinary wake: re-arm the next cycle with the same background `bin/fm-watch-arm.sh` call if work remains in flight or Relay still needs polling. +4. Ordinary wake: re-arm the next cycle with the same background `bin/fm-watch-arm.sh` call if the home still needs supervision, as `bin/fm-supervision-lib.sh` defines it. 5. Do not invent a wake from an attach-status line alone. Drain the queue and act only on real wake records, the drain's `OPEN DECISIONS` and `UNREAD STATUS` entries, or a real watcher reason line. Re-arm attaches to an existing healthy cycle when one is already present and follows its verified successor chain. diff --git a/docs/turnend-guard.md b/docs/turnend-guard.md index 98ee17799e1..abf95bc0425 100644 --- a/docs/turnend-guard.md +++ b/docs/turnend-guard.md @@ -13,7 +13,7 @@ Do not infer this guard's scope, loop safety, or compatibility tradeoffs for tho `bin/fm-guard.sh` is a pull-based warning that runs only when another supervision command invokes it. The turn-end guard closes the remaining gap at the primary's own turn boundary. -When work, a process-event source, Relay polling, or a pending wake queue needs supervision at that boundary and no identity-matched watcher has a fresh beacon, the harness integration must either block the turn end or force one bounded follow-up that uses the recovery instruction from the emitted session-start protocol. +When work, a process-event source, a registered custom check, Relay polling, or a pending wake queue needs supervision at that boundary and no identity-matched watcher has a fresh beacon, the harness integration must either block the turn end or force one bounded follow-up that uses the recovery instruction from the emitted session-start protocol. The mid-turn pull warning uses the model-aware supervision verdict described below, while the turn-end guard keeps the PID-strict watcher predicate. Away mode is the one place the turn-end guard accepts a different supervisor: while `state/.afk` exists the away-mode daemon owns supervision, so a live identity-matched daemon with a fresh beacon satisfies that boundary in place of a watcher process holding the lock. The guard remains a backstop; [`watcher-continuity.md`](watcher-continuity.md) owns normal continuity. @@ -32,6 +32,7 @@ Registered `state/procevent/*.source` records also require supervision even thou Wake records still queued in `state/.wake-queue` are supervision need too, so an idle home that never drained its queue stays guarded instead of ending the turn blind. The default cross-harness mode exits silently with no supervision need. Every mode treats `state/x-watch.check.sh` as supervision need, so Relay polling remains guarded without an in-flight task. +A custom check registered with `bin/fm-check-register.sh` counts the same way, so an operator's home-level poll keeps running after the last task is torn down. Otherwise it calls `fm_watcher_healthy <state-dir> <watch-path> [grace-seconds] [home]` from `bin/fm-wake-lib.sh`, the same PID-strict identity-matched lock and fresh-beacon check used by `bin/fm-watch-arm.sh`: a stale beacon blocks even when a watcher pid is live, and a fresh leftover beacon blocks when the lock is missing, dead, or identity-mismatched. The turn-end guard needs that strict check because it fires at the turn boundary, where the auto-arm is bringing a fresh watcher up for the upcoming idle period, and it cooperates with that arm rather than trusting a beacon left by the cycle that just ended. `bin/fm-guard.sh`, the pull warning, instead uses the model-aware `fm_watcher_supervision_verdict` from the same library, because it fires mid-turn when the auto-arm model runs no watcher at all. @@ -49,13 +50,24 @@ While `state/.afk` exists the away-mode daemon (`bin/fm-supervise-daemon.sh`) ow The turn-end guard therefore accepts `fm_afk_daemon_owns_supervision` from `bin/fm-wake-lib.sh` as proof of supervision on that path: away mode must be active, and this home's `state/.supervise-daemon.lock` must name a live pid whose current process identity still matches the identity the daemon recorded for itself. That is the same identity discipline the watcher lock uses, so a recycled pid, a lock left behind by a killed daemon, and a daemon that never recorded its identity all fail it. A daemon that cannot record its own identity at startup logs a warning and keeps running, because a supervisor must not refuse to run over an unreadable `ps`; that warning is what names the cause when the guard then keeps blocking away-mode turn boundaries for the rest of that daemon's life. -The proof covers ownership only, never freshness: the fresh-beacon half of the predicate is unchanged, so a daemon that stops restarting its watcher still blocks once the beacon passes grace, and a home with no daemon and no watcher blocks exactly as it did before. +The proof covers ownership only, never freshness: the guard still requires a fresh beacon, so a daemon that stops restarting its watcher still blocks once the beacon passes grace, and a home with no daemon and no watcher blocks exactly as it did before. +That beacon check uses the poll-derived grace described below rather than the flat `FM_GUARD_GRACE` default, because the daemon starts a fresh one-shot watcher only after it finishes handling the previous wake, and that handling can legitimately outrun a fixed 300-second window under load (a slow registered check, a busy supervisor pane) with the daemon perfectly healthy throughout. With away mode off the daemon lock proves nothing and the strict watcher predicate is unchanged. `FM_STATE_OVERRIDE` wins over `FM_HOME/state`, and `FM_HOME` wins over repository-root `state/`. `FM_GUARD_GRACE` controls beacon freshness and defaults to 300 seconds. If `jq` is missing or hook stdin is empty, the guard exits 0 because it cannot safely read loop-guard fields. +### Guard grace and the poll cadence + +`bin/fm-watch.sh` touches `state/.last-watcher-beat` once per cycle, immediately before its terminal wait (`event_wait_or_sleep`) as well as at the top of the next cycle, so a healthy watcher's beacon can legitimately age up to `FM_POLL` seconds between touches. +A fixed 300-second grace default stops correctly bounding staleness once a home's `FM_POLL` reaches or exceeds it: a perfectly healthy watcher mid-wait would then read stale at the edge of every full poll cycle by definition, which is exactly what a long-poll home (`FM_POLL=300`) hit against the Claude Stop-hook auto-arm (`bin/fm-claude-stop-autoarm.sh`). +That hook and `bin/fm-watch.sh`'s own pre-acquisition staleness check (the "lock held by live pid but heartbeat is stale" refusal) both derive their default grace from the configured poll instead of a bare constant: `max(300, FM_POLL + 60)`, so the default never drops below the historical 300-second floor for the common short-poll case but grows with the poll cadence once that cadence would otherwise outrun it. +`fm_poll_derived_grace` in `bin/fm-wake-lib.sh` is the single owner of that formula. +The auto-arm hook additionally exports its resolved `FM_GUARD_GRACE` when it forks `bin/fm-watch-arm.sh`, so the arm wrapper and the watcher it may start judge staleness with the exact same value the hook just judged it with, whether that value came from an operator override or the poll-derived default. +`bin/fm-turnend-guard.sh`'s away-mode branch (`fm_afk_daemon_owns_supervision`, above) also derives its beacon grace from `fm_poll_derived_grace` rather than falling back to the bare 300-second default, for the same reason: the daemon's watcher-restart cadence there is not a fixed poll loop, so a flat grace misreads a daemon that is genuinely still cycling as down. +Every other direct `FM_GUARD_GRACE` reader (`bin/fm-guard.sh`, the strict-watcher checks in `bin/fm-turnend-guard.sh` and its harness-specific wrappers, `bin/fm-wake-lib.sh`) still falls back to the bare 300-second default unless `FM_GUARD_GRACE` is set explicitly in the environment. + ## Harness integrations - Claude registers two `Stop` hooks in `.claude/settings.json`, both anchored through `CLAUDE_PROJECT_DIR`: `bin/fm-turnend-guard.sh --claude`, and `bin/fm-claude-stop-autoarm.sh` with `asyncRewake: true` and `timeout: 28800`. @@ -170,7 +182,7 @@ That warning uses `bin/fm-supervision-instructions.sh --repair-line`, so it alwa ## Regression coverage -`tests/fm-turnend-guard.test.sh` covers the predicate, main and secondmate primary scope, child-worktree exclusion, `FM_HOME` and `FM_STATE_OVERRIDE` precedence, the live-lock and fresh-beacon guard predicate, the cooperative `--claude` open-generation claim wait, monotonic failed-epoch progression, bounded attended fail-open, post-alarm continuation suppression, positive recovery reset, generation and legacy claim cases that must block or clear instead of allowing a blind stop, away-mode daemon ownership between watcher cycles and over a watcher lock left behind by an exited watcher, plus its dead, pid-reused, absent, stale-beacon, and away-mode-off negatives, Pi logical-run latching, missing-`jq` behavior, all five primary registrations, Grok native and legacy selection, typed field precedence, malformed input, and exactly-one-path safety. +`tests/fm-turnend-guard.test.sh` covers the predicate, main and secondmate primary scope, child-worktree exclusion, `FM_HOME` and `FM_STATE_OVERRIDE` precedence, the live-lock and fresh-beacon guard predicate, the cooperative `--claude` open-generation claim wait, monotonic failed-epoch progression, bounded attended fail-open, post-alarm continuation suppression, positive recovery reset, generation and legacy claim cases that must block or clear instead of allowing a blind stop, away-mode daemon ownership between watcher cycles and over a watcher lock left behind by an exited watcher, plus its dead, pid-reused, absent, stale-beacon, and away-mode-off negatives, the away-mode beacon's poll-derived grace widening for a live daemon still mid-cycle and its bound against a dead daemon, a beacon older than that wider grace, and FM_POLL's inapplicability with away mode off, Pi logical-run latching, missing-`jq` behavior, all five primary registrations, Grok native and legacy selection, typed field precedence, malformed input, and exactly-one-path safety. `tests/fm-guard-stale-banner.test.sh` covers the pull-guard predicate, including the persistent-model fresh-leftover-beacon negative control, the auto-arm model's healthy fresh-beacon-without-a-watcher case and stale-beacon alarm, and the extension model's live-watcher path, ownership-qualified fresh hand-off, held-lock failures, independently broken ownership signals, stale-beacon alarm, queued-wake warning, and Pi and pi-signed harness routing. It also covers true-reason banner wording and reason-keyed episode dedup surviving a beacon mtime change. `tests/fm-cursor-primary.test.sh` covers the Cursor park end to end over real processes with no harness installed: each tracked Claude-shaped entrypoint standing down on a Cursor payload, both follow-up sources, the bounded repair nag and its reset, the nested loop bounds, supersession, away-mode and lock-ownership inertness, Pi-host stand-down without Cursor identity and continued parking when `PI_CODING_AGENT` leaks alongside `CURSOR_AGENT` or `CURSOR_INVOKED_AS`, child-worktree exclusion, and that the adapter never exits 2. diff --git a/docs/verification/muse.md b/docs/verification/muse.md index 2a2637b3c65..38d12653a6a 100644 --- a/docs/verification/muse.md +++ b/docs/verification/muse.md @@ -204,7 +204,7 @@ That is the same terminal shape the `echo`-provider interrupt produced, now conf ## Refreshing this record -Run both opt-in live guards after any muse upgrade, because the version-suffixed process name, session protocol, and styled composer are vendor-controlled surfaces: +Run both live guards after any muse upgrade, because the version-suffixed process name, session protocol, and styled composer are vendor-controlled surfaces: ``` FM_HARNESS_LIVENESS_DRIFT=1 bin/fm-test-run.sh tests/fm-harness-liveness-drift-live-e2e.test.sh diff --git a/docs/verification/process-event-sources.md b/docs/verification/process-event-sources.md index 7895b498c43..669a5b2c247 100644 --- a/docs/verification/process-event-sources.md +++ b/docs/verification/process-event-sources.md @@ -97,10 +97,11 @@ Exercised by `tests/fm-procevent.test.sh` against a fake blocking source whose c | adapter-owned terminal verdict | two fixture adapters - one that ends on any result, one with no terminal knowledge - decide the outcome alone: the first has its registration and claim retired automatically after one capture and is never restarted, the second stays armed | | adapter-owned application of a captured result | a remote-secondmate reply captured through the real relay in an isolated home reaches that secondmate's local status mirror, settles its correlated pending-reply expectation, re-arms the next cursor-anchored source, and is acknowledged, with no handler step or duplicate `check` wake; its new mirrored bytes remain visible to the watcher's signal gate, while a cursor-loss whole-log recapture that adds no bytes is acknowledged quietly; for an already-escalated request, the same path closes the exact decision so the open-decision fold clears and remains clear; a capture whose adapter application fails because local storage for a referenced remote document is obstructed is left unacknowledged and receives the fallback `check` wake, and the handler's own `handle` still applies it in full after storage recovers | | generic built-in keyed-answer feed | `tests/fm-captain-hold-lifecycle.test.sh` drives a bound built-in source through the real runner with a fixture adapter that only prints keyed lines, proving any bound built-in channel reaches the one keyed-answer intake: named captain-held tasks close at capture time, a card-declared release mode frees held work, keys naming no captain-held task skip, freeform prose forges nothing, matching answer-and-mode replays are idempotent while mode mismatches refuse, an unbound source closes nothing, and capture remains independent of the handler wake. | +| structured reconcile feed | The same suite drives the optional `reconciles` adapter seam through the real runner and proves only a bound captured source can create a request; the ordinary keyed-answer and chat paths refuse the reserved value without closing or creating a request, versioned selection stays separate from its note, rollout-compatible ordinary legacy answers still pass, and legacy reconcile-shaped values feed neither intake. | | adapter-owned silence verdict | an armed Lavish source driven against a stand-in poll that returns an empty ended session captures its result, records it durably handled, appends no wake, and stays silent through a later `reconcile` that would otherwise republish it, while still retiring its ended source; the same real path with a `Send & End` response carrying the captain's choice still publishes its `check` wake and is left unacknowledged for the handler | | silence fails closed | the adapter's published `silent` command suppresses only an `ended` session with no queued content block, and announces a real answer, freeform prose, any recognized content block regardless of its declared count, a malformed top-level content header, a `waiting` or `missing` session, a server error, an unreadable result, and indented payload text imitating an empty content block; the `remote-reply` and `when` adapters, which implement no `silent` command, announce every result | | terminal retirement preserves the result | the retired source's captured output, its announced event, its handled acknowledgement, and later explicit `retire` all still behave normally | -| registration-generation retirement | an old terminal runner preserves a concurrently replaced registration and releases ownership so the replacement runs independently; injected registration-removal failure retains a terminal claim, performs no second poll, and completes idempotently once removal recovers | +| registration-generation retirement | an old terminal runner preserves a concurrently replaced registration and releases ownership so the replacement runs independently; injected registration-removal failure retains a terminal claim, performs no second poll, and completes idempotently once removal recovers; a live owner retiring its own terminal source mid-capture tolerates only its transient reservation-removal failure and still removes the registration under exact ownership | | one `Send & End`, one result | an armed Lavish source driven against a stand-in for the published poll, which delivers the final `session_ended` feedback once and empty ended sessions afterward, polls exactly once, captures exactly one result, publishes one distinct event, and retires itself | | bounded re-announcement until handled | a durably captured result with no handled acknowledgement is re-announced by `reconcile` with the same source and sequence on every call - not only the first restart after a crash - and a presented-but-unacknowledged wake resurfaces identically after a simulated replacement session | | handled acknowledgement | `fm-procevent.sh handled <source-id> <sequence>` atomically and idempotently records handling at mode `0600`, fails without leaving a marker when private-mode enforcement fails, reports the first call distinctly from every repeat, stops further re-announcement once recorded, and never authorizes a paired effect twice across repeat calls | @@ -114,9 +115,13 @@ Exercised by `tests/fm-procevent.test.sh` against a fake blocking source whose c | canonical physical identity | a final-component symlink and its target produce the same Lavish source id, and a path whose file and whole directory are gone still resolves to the id derived while the file existed - including one whose `.` or `..` falls inside the deleted portion, which is resolved lexically rather than rejoined verbatim | | a registration outliving its artifact stays retirable | the Lavish adapter retires by source id, and by a path whose file and whole directory are gone; a repeat retire by the same gone path is idempotent once a result has been captured for that source, a still-resolving path retires exactly as before, an unmatched gone path refuses with retire-by-id guidance instead of a silent no-op, and retirement stops the recurring wakes without acknowledging the captured result | | isolated public start boundary | direct `start` establishes a new runner-led process group before claiming the source, so retirement cannot signal an unrelated process inherited from the caller's group | -| stale reclaim without displacement | concurrent contenders replacing one stale claim start exactly one runner, and cross-home replacement removes the old generation's staging file from its recorded state directory | -| crashed leader with a live owned group | `SIGKILL` on only the runner leader leaves its blocking child group alive; reconcile then stops that surviving group before any replacement starts, never leaves two source processes running for one canonical source, and a generation with no leader and no surviving group is still reclaimed | -| PID-reuse safety | retirement refuses to signal a live PID whose identity differs from the claim, and a reused PID never reaches the group-stop path because its leader is alive | +| guarded runner startup | the source command does not launch when the detached owner guard rejects an invalid lease configuration, proving the runner waits for positive guard readiness and fails closed when initialization fails | +| attached owner continuity | a foreground `start` with a one-second lease remains alive beyond that lease while its caller stays attached, then captures normally when the blocking source completes | +| owner-home lifetime and scope | a detached runner and its spawning descendant are observed reparented before an expired owner lease stops their whole process group and process churn; replacing the state directory at the same path cannot keep the old runner alive with a new lease because its recorded device/inode no longer matches, while an identical runner in an unchanged home whose reconcile cycle keeps its lease fresh remains alive | +| launch pacing during owner-loss grace | an immediately returning source that attempts detached self-relaunches is held to the configured minimum interval between command launches and remains bounded until its expired owner lease stops the generation; replacement starts a fresh pacing generation, prunes prior pacing state, and prevents a superseded sleeping runner from recreating it | +| stale reclaim without displacement | concurrent contenders replacing one stale claim start exactly one runner, cross-home replacement removes the old generation's staging file from its recorded state directory, and a generation whose stale owner and independently empty process group prove it gone remains reclaimable when its recorded state-root identity can no longer be revalidated | +| crashed leader with a live group | `SIGKILL` on only the runner leader leaves its blocking child group alive; reconcile treats that leaderless group as ambiguous, preserves its claim without starting a replacement, and still reclaims a generation with no leader and no surviving group | +| PID-reuse safety | retirement refuses a live PID whose identity differs from the claim before signalling, and a surviving process group prevents stale-generation cleanup on both ordinary and failed reservation-removal paths | | coherent ownership reads | a claim replacement held inside the source boundary blocks `list` until one complete generation is visible | | retire-start exclusion | a queued start revalidates registration after the serialized retirement boundary and executes no child | | uncertain identity | a live owner whose identity probe transiently fails is not signaled or released, and its registration remains for retry | @@ -136,7 +141,7 @@ Exercised by `tests/fm-procevent.test.sh` against a fake blocking source whose c | inertness | a home with no registered source generates no state, starts no process, and does not need supervision | | absent extension registry parity | `tests/fm-extension-binding.test.sh` drives `list` and `verify` in a fresh home while the current directory contains project files and Pi packages and an environment variable names fake package data; both commands report no bindings, create no home path, and discover nothing outside `config/extensions.d` | | complete package and binding identity | the same suite drives the public bind and verify commands through manifest duplicate/unknown/version failures, project and task-copy confinement, canonical path and symlink rejection, hard-link rejection, owner/mode checks, a non-executable entrypoint, binding mode drift, complete-tree mutation, exact executable mutation, and a missing executable; the foreign-owner fixture executes when the platform permits constructing another uid and otherwise reports that privilege limitation, while ordinary non-privileged CI does not exercise it or claim it ran | -| external evidence write confinement | the same suite substitutes `state/procevent/` and `state/procevent-inbox/` with post-registration symlinks and proves an external start fails before bytes reach either outside target; it proves public lifecycle entry, environment, paths, and descriptors cannot forge capture authority; it proves claim release and dead-owner reconciliation remove pending or consumed capture reservations only from the recorded revalidated state root; and it proves the absent-registry built-in capture path retains its legacy state-path behavior | +| external evidence write confinement | the same suite substitutes `state/procevent/` and `state/procevent-inbox/` with post-registration symlinks and proves an external start fails before bytes reach either outside target; it proves public lifecycle entry, environment, paths, and descriptors cannot forge capture authority; it proves live-generation claim release removes pending or consumed capture reservations only from the recorded revalidated state root, while a generation independently proved gone may leave an unreachable token-keyed reservation rather than wedging ownership; and it proves the absent-registry built-in capture path retains its legacy state-path behavior | | strict handshake and negotiation | manifests offering versions 2 and 1 select host protocol 1 and `process-event-adapter/1`, unknown-only versions refuse, and wrong request ids, unknown or duplicate fields, malformed JSON, and nonzero handshake exits publish no binding | | strict invocation envelope | malformed UTF-8, a byte-order mark, unescaped controls, malformed or multiple JSON documents, duplicate or unknown fields, oversized stdout, oversized stderr, wrong request ids, crashes, nonzero exits, a successful parent that leaves a foreground descendant in its host-created invocation group, and authority-shaped result fields are rejected; leaked group members are reaped and package diagnostic text is not copied into the bounded host-produced error evidence | | extension timeout and process-group cleanup | a bound adapter that ignores `TERM`, spawns a foreground descendant that ignores `TERM`, and exceeds its invocation timeout returns deterministic timeout evidence only after its exact invocation group is gone; deliberate process-group escape is outside this trusted-same-user protocol guarantee | @@ -146,13 +151,14 @@ Exercised by `tests/fm-procevent.test.sh` against a fake blocking source whose c | owner-matched replacement safety | two registrations for the same external source receive distinct owner tokens; unconditional external retirement and the first token cannot retire the replacement, the replacement token can, bounded home sweep derives and uses that exact token, and legacy built-in registrations retain unconditional behavior plus exact `--if-matches` retirement | | independent homes | two homes bind the same package id/version to different content-addressed absolute paths and independently capture results and extension state, with no cross-home fallback or result path | -Run the focused external-binding evidence with: +Run the focused external-binding evidence and the live Bearings session guard with: ```sh node --version bin/fm-test-run.sh tests/fm-extension-binding.test.sh FM_EXTENSION_BINDING_SEGMENT=lifecycle-invocation-cleanup bin/fm-test-run.sh tests/fm-extension-binding.test.sh bin/fm-test-run.sh tests/fm-procevent.test.sh +FM_BEARINGS_LAVISH_LIVE=1 bin/fm-test-run.sh tests/fm-bearings-board-lavish-live-e2e.test.sh bin/fm-doc-audience-check.sh ``` @@ -173,19 +179,29 @@ The 2026-08-27 review inspected `bin/fm-harness.sh`, `bin/fm-supervision-instruc ## Runner lifetime and cleanup A runner started by `reconcile` is its own process group leader and is reparented to init, so it outlives the shell that started it by design. -That means nothing about the starting context can reap it: removing a home's state directory does not stop an already-running child, and signalling only the runner leaves the blocking child alive. +Removing a home's state directory does not stop an already-running child, and signalling only the runner leaves the blocking child alive. -Two paths therefore stop a runner, and both verify the runner-owned process group, escalate to `KILL` while that group still exists, and refuse to release ownership until the whole group is gone: +Three paths stop a runner generation through its verified process group: +- The runner starts only after its separate owner guard confirms initialization; the guard stops the runner group after two consecutive checks cannot prove the owning home's recorded physical identity and lease freshness. - `retire` resolves the runner PID and identity from this home's machine-wide claim, so retirement still works when the home's state is already gone. - `reconcile` stops a runner this home owns whose source registration has been removed, and reports it as `stopped=N`. -The same group rule decides when a claim may be reclaimed, not only when a runner may be signalled. -A leader that died while its owned group kept running is not a stale generation, so `reconcile` stops that surviving group and releases its generation before starting any replacement, and preserves the claim for a later retry when it cannot prove the group stopped or another home owns it. -Signalling that group is safe precisely because only an absent leader reaches this state: a reused PID leaves the leader alive, which the identity comparison classifies as stale or uncertain, and no group signal follows. +The owner guard and explicit cleanup paths reach the blocking source and its descendants through the runner's group. +The registration launch floor independently bounds repeated runner launches while an owner-loss lease is still valid. +The Lavish adapter's start-to-start poll governor separately bounds its internal retry loop under shipped defaults without delaying a normally blocking poll. +An attached public `start` maintains the lease for its caller's lifetime. +At the accepted confused-agent/accidental grade, the inherited `FM_PROCEVENT_IN_RUNNER` marker prevents detached runners and their ordinary children from refreshing it; adversarial unforgeability against a source that deliberately strips that marker is out of scope. -This was found by four orphaned runners, elapsed 6-13 minutes, left by a suite whose fixture source never completed. -`tests/fm-procevent.test.sh` now covers both paths, and three consecutive suite runs leave zero runners, zero fixture children, and zero stray claims. +The same group rule decides when a claim may be reclaimed, not only when a runner may be signalled. +A leader that died while its process group kept running is not a gone generation. +Because the leaderless group cannot be proved to belong to the recorded generation, `reconcile` preserves its claim without signalling it or starting a replacement. +Once a stale owner and an independent group check prove the whole generation gone, an unreachable token-keyed capture reservation cannot veto reclamation. +Known limit: when either a live reused PID or an absent leader makes group ownership ambiguous, the reaper does not act because it cannot prove the group is the orphan generation; storm-rate containment plus ordinary lease and reconcile cleanup are the confused-agent-grade backstop. +Known limit: identity and process-group verification cannot be made atomic with signalling in portable shell. +The reaper signals only a target it has verified as the orphan generation, but PID and group reuse remain possible in the narrow interval between verification and the signal; launch pacing is the primary host-wedge protection and watchdog cleanup is a backstop. + +`tests/fm-procevent.test.sh` covers owner-loss reaping, descendant churn cessation, cross-home scope, launch pacing, guard startup failure, attached-start continuity, explicit retirement, and stale-group reconciliation. ## Portability finding diff --git a/docs/verification/runtime-backends.md b/docs/verification/runtime-backends.md index 1816b337112..c2fce574931 100644 --- a/docs/verification/runtime-backends.md +++ b/docs/verification/runtime-backends.md @@ -180,7 +180,7 @@ Observed identities, and the resulting verdict: | grok | 0.2.118 | `grok-0.2.118-ma` | `grok` | alive | | kimi | 0.31.1 | `kimi` | `kimi` | alive | -Claude Code is the harness whose title no longer attributes it at all; every other adapter is currently attributed by both sources. +In that 2026-08-03 seven-adapter run, Claude Code was the only harness whose title did not attribute it; every other adapter was attributed by both sources. Codex reported `codex-aarch64-a` at 0.145.0 and `codex` at 0.146.0, and Kimi Code reported `kimi-code` as its foreground `comm` at 0.29.1 and `kimi` at 0.31.1, so these identities move between ordinary patch releases in both directions. That is the evidence for treating any single process name as a surface under vendor control rather than a stable contract. @@ -209,14 +209,36 @@ alive On macOS the pane command reflected the rewritable title while the full install path could survive in `ps -o comm=`; in the Linux portable regression those roles reversed for the version-named native executable, with the identifying path retained in argv[0]. The classifier therefore accepts a harness basename first, then an exact harness path component in the full executable path, then the same component in argv[0], without depending on which field carries it on a given platform. -The portable regression is CI-enforced, while the real-harness drift guard is opt-in under the policy in `.agents/skills/firstmate-coding-guidelines/SKILL.md`. +The portable regression is CI-enforced. +The real-harness drift guard spends no model tokens, so under the policy in `.agents/skills/firstmate-coding-guidelines/SKILL.md` it runs by default wherever tmux is installed and reports a capability skip elsewhere; `FM_HARNESS_LIVENESS_DRIFT=1` additionally turns an absent tool into a failure. Run the live guard after any harness upgrade and before trusting or refreshing the table above: ```sh FM_HARNESS_LIVENESS_DRIFT=1 bin/fm-test-run.sh tests/fm-harness-liveness-drift-live-e2e.test.sh ``` -Bounded output from the run that produced the table: +### 2026-09-06 default-on drift refresh, and the Cursor editor CLI collision + +Running the guard with no variable set on macOS 26.5.2 arm64 checked 8 installed harnesses and classified every one `alive`: + +```text +# claude 2.1.263 (Claude Code): title='2.1.263' foreground=[/Users/kunchen/.local/bin/claude <defunct> <defunct> ] +# codex codex-cli 0.147.0: title='codex' foreground=[/opt/homebrew/bin/codex ] +# opencode 1.18.29: title='opencode' foreground=[/opt/homebrew/bin/opencode ] +# pi 0.84.4: title='pi-launcher' foreground=[/opt/homebrew/bin/pi-signed .../pi ] +# pi-signed 0.84.4: title='pi-launcher' foreground=[/opt/homebrew/bin/pi-signed .../pi ] +# grok grok 1.0.13 (5e9a58528b76) [stable]: title='grok-1.0.13-mac' foreground=[/Users/kunchen/.local/bin/grok ] +# cursor 2026.09.02-c22c1a3: title='node' foreground=[/Users/kunchen/.local/bin/cursor-agent ] +# muse Muse Code 1.0.3 (1.0.3-R2198.1): title='muse-bin-1.0.3-' foreground=[/Users/kunchen/.local/bin/muse-bin-1.0.3-R2198.1 ] +# checked 8 installed harness(es) +``` + +The first default-on run failed on Cursor with `LIVENESS DRIFT: cursor unknown is running but classifies 'missing'`, observed title `zsh`. +The classifier was not at fault: the guard resolved the harness through a generic `command -v cursor`, which on a machine that also has the Cursor editor finds `~/.local/bin/cursor` - the editor launcher, not the agent. +That binary exits immediately, leaving a bare shell in the pane. +The guard now asks `fm_cursor_resolve_binary` first for `cursor`, which is the same verified owner `bin/fm-spawn.sh` uses, so the probe launches `cursor-agent` and the editor CLI can no longer masquerade as the harness. + +Bounded output from the 2026-08-03 run that produced the first table above: ```text ok - harness liveness: claude 2.1.220 (Claude Code) classifies alive diff --git a/docs/verification/trace-context.md b/docs/verification/trace-context.md index 83f78fe343c..b921ccec632 100644 --- a/docs/verification/trace-context.md +++ b/docs/verification/trace-context.md @@ -9,7 +9,7 @@ Comparison base: `main` at `976d97f`. The colocated unit suite `tests/fm-trace-context-lib.test.sh` (26 assertions) exercises validation (valid accepted; malformed, wrong-length, uppercase, all-zero, `ff` version, and shell-metacharacter values rejected), root minting with every mint a distinct sampled root and no parent-adoption input, the recovery reuse path with the recorded carrier winning over the ambient environment, default-off omission, the enable precedence of `FM_TRACE_CONTEXT` over `config/trace-context` with unset or empty deferring to the file, normalized home-session state, atomic replacement of a read-only prior record, stale-session rejection after failed publication, missing or invalid state defaulting off, the Secondmate home-session boundary with later file state plus the per-task trace boundary (two resolves under one persistent ambient `TRACEPARENT` root two distinct traces and adopt neither), forced entropy failure omitting safely, and the minted-root fixed-shape check. -The spawn-path integration suite `tests/fm-trace-context-spawn.test.sh` (12 assertions), hermetic against an ambient `FM_TRACE_CONTEXT`, drives `bin/fm-spawn.sh` end to end with a fake tmux pane and a real isolated git worktree: enabled, one resolved carrier is recorded as `traceparent=` in the meta only after the identical `TRACEPARENT` export is sent before the launch literal; disabled, neither is written nor sent (only `GOTMPDIR` is); a failed carrier delivery leaves no `traceparent=` claim while the source task still launches; an unsafe delivery whose partial input cannot be cleared stops before appending the launch command; a failed metadata append removes the carrier from the launched task without aborting it; duplicate Secondmate preflight leaves inherited trace configuration unchanged; a relaunch reuses the recorded carrier verbatim; and spawns ignore later config and environment edits in favor of the frozen home-session decision. +The spawn-path integration suite `tests/fm-trace-context-spawn.test.sh` (12 assertions), hermetic against an ambient `FM_TRACE_CONTEXT`, drives `bin/fm-spawn.sh` end to end with a fake tmux pane and a real isolated git worktree: enabled, one resolved carrier is recorded as `traceparent=` in the meta only after the identical `TRACEPARENT` export is sent before the launch literal; disabled, neither is written nor sent (`GOTMPDIR` still is); a failed carrier delivery leaves no `traceparent=` claim while the source task still launches; an unsafe delivery whose partial input cannot be cleared stops before appending the launch command; a failed metadata append removes the carrier from the launched task without aborting it; duplicate Secondmate preflight leaves inherited trace configuration unchanged; a relaunch reuses the recorded carrier verbatim; and spawns ignore later config and environment edits in favor of the frozen home-session decision. The per-task boundary regression models the reviewed Secondmate scenario exactly: two unrelated tasks spawned sequentially from one home while the same fixed `TRACEPARENT` sits in the spawning environment (a persistent Secondmate's launch-time carrier) record and inject valid carriers whose trace ids differ from each other and from the ambient carrier, and a relaunch of the first task reuses its original carrier verbatim for both the meta record and the injected export. Two further assertions drive a genuine two-level primary -> Secondmate -> worker chain, running `bin/fm-spawn.sh` twice with the exact environment the primary injects into the Secondmate, and prove the primary's effective override governs the nested worker both ways: env-on with no config file keeps the nested worker enabled while it roots its own per-task trace distinct from the Secondmate's carrier, and env-off with the file present keeps the nested worker disabled even though the `config/trace-context` file was copied into the Secondmate home. A final assertion drives the file-decided path (`FM_TRACE_CONTEXT` unset) and proves the Secondmate's recorded/injected carrier and its delivered `FM_TRACE_CONTEXT=on|off` snapshot are always derived from one frozen decision, so a carrier is never paired with the opposite enable state. diff --git a/docs/watcher-continuity.md b/docs/watcher-continuity.md index 04b84780832..3999ac95a40 100644 --- a/docs/watcher-continuity.md +++ b/docs/watcher-continuity.md @@ -79,6 +79,13 @@ Main records its presented set in `state/.main-eligible-rows`. A branch grant is published through `bin/fm-wake-grant.sh` under that same lock in `state/.branch-eligible-rows`, bound to the live branch process and extension generation recorded in `state/.branch-eligible-owner`, and publication is refused if main already claimed any requested row. A main drain validates that owner evidence under the queue lock and reclaims the grant when its process is gone or its identity no longer matches. A main drain claims every currently unclaimed row and excludes an active branch grant from both presentation and acknowledgement. +Because that exclusion makes those rows invisible to main, `bin/fm-guard.sh`'s queued-wake warning counts only the rows the calling actor can itself present or retire, so an actor is never sent to a drain that provably has nothing for it. +`bin/fm-wake-lib.sh` owns that per-actor count (`fm_wake_actor_pending_count`) alongside the grant row-list and owner-record reads that the drain and `bin/fm-wake-grant.sh` share. +A row a live grant reserves is therefore never counted as drainable for main; rather than going silent about a visibly non-empty queue, the guard prints a distinct advisory naming the live supervision branch as the holder and saying not to drain those rows from here. +The branch actor's queued-wake output stays suppressed in every case. +A main drain with nothing of its own left, and a live grant still holding the queue, says so in one bounded line instead of exiting silently. +A row that lost the five appended fields or its numeric sequence can never be claimed, presented, or named by an `--ack-through` cutoff, so a main drain retires it under the queue lock and reports how many it removed together with those rows verbatim, bounded to the first 20 and a count of the rest, because the queue was their only durable record; a branch drain never does, because a grant can only name sequences that were structurally valid when it was published. +A retirement that cannot be read or written is reported and never fails the drain: the rows that remain usable are still presented with their acknowledgement command, the unusable ones stay queued for a later drain to retire, and failing the whole drain would strand the usable rows too. Its `--ack-through <SEQ>` deletes only claimed main rows at or below the cutoff, while a branch acknowledgement deletes only claimed branch rows at or below its cutoff. Every settled branch prompt releases any residual grant, so an omitted or failed acknowledgement leaves the durable row available to a later main drain; a successful acknowledgement has already removed it. An acknowledgement whose cutoff removes none of the actor's rows while a presented row above the cutoff still waits is reported as having acknowledged nothing, together with the exact `--ack-through` and `--recovery-generation` command for that presented row; the presented set is read before any re-claim, so a row that arrived after presentation is never named for unseen acknowledgement. @@ -89,6 +96,7 @@ A check-kind row is main-owned in every mode, including a heartbeat review, so i A missing or empty branch snapshot is refused loudly rather than read as "nothing eligible", because reaching the drain without the non-empty handoff promised by the extension is a wiring bug. Because branch claims contain no check-kind rows, a branch acknowledgement skips check-specific receipt scans. `tests/fm-wake-queue.test.sh`'s mixed-queue actor, stale-acknowledgement remedy, and presentation-deadline tests drive the real scripts: branch acknowledgement cannot swallow a main row, a concurrent main turn cannot present or acknowledge an active branch grant, a no-op stale acknowledgement names the current presented wake's exact command, live-holder presentation contention stays bounded and retriable, and acknowledgement locking remains blocking. +The same suite pins the counted-equals-presentable invariant against `bin/fm-guard.sh` and `bin/fm-wake-drain.sh` together: a branch-held row raises the held advisory rather than the ordinary queued-wake warning for main, and is presented with its acknowledgement command - with the ordinary warning restored - as soon as the grant clears, and structurally unusable rows are retired by main alone while every remaining row stays presentable and acknowledgeable. `tests/fm-pi-branch-extension.test.sh` pins extension-side classification, claim publication and release, and the pre-drain recheck. ## Arm-layer cycle contract diff --git a/tests/fixtures.sh b/tests/fixtures.sh index d4c9f977a37..f95800d0a56 100755 --- a/tests/fixtures.sh +++ b/tests/fixtures.sh @@ -94,7 +94,8 @@ fm_test_fake_gh_axi() { # fm_test_fake_tmux_spawn <fakebin> # Spawn-world tmux: pane_current_path from FM_FAKE_PANE_PATH, session named # firstmate, window ops succeed, send-keys succeed. When FM_FAKE_LAUNCH_LOG is -# set, each send-keys -l payload is appended one per line. Optional +# set, each send-keys -l payload is appended one per line. +# FM_FAKE_TEXT_LINE_LOG separately records spawn-time text submitted with Enter. Optional # FM_FAKE_DUPLICATE_WINDOW is printed from list-windows. # # The pane path defaults to empty when FM_FAKE_PANE_PATH is unset. Window @@ -126,6 +127,10 @@ case "${1:-}" in ;; has-session|new-session|new-window|kill-window|set-window-option) exit 0 ;; send-keys) + if [ -n "${FM_FAKE_TEXT_LINE_LOG:-}" ] && [ "$#" -eq 5 ] \ + && [ "$2" = -t ] && [ "$5" = Enter ]; then + printf '%s\n' "$4" >> "$FM_FAKE_TEXT_LINE_LOG" + fi if [ -n "${FM_FAKE_LAUNCH_LOG:-}" ] || [ -n "${FM_FAKE_RESOLVED_CLAUDE_CONFIG_LOG:-}" ]; then prev= for a in "$@"; do diff --git a/tests/fm-afk-pi-herdr-return-e2e.test.sh b/tests/fm-afk-pi-herdr-return-e2e.test.sh index dc16efd71dd..1976ae87b79 100755 --- a/tests/fm-afk-pi-herdr-return-e2e.test.sh +++ b/tests/fm-afk-pi-herdr-return-e2e.test.sh @@ -2,8 +2,11 @@ # Real Pi/Herdr end-to-end regression for the 2026-07-14 two-owner incident. # # Opt-in because it launches a real interactive Pi primary, a real away daemon, -# and a real isolated Herdr lab session. Every explicit and production-adapter -# Herdr call is routed through fm-herdr-lab.sh. The scenario proves: +# and a real isolated Herdr lab session. Real model turns, real side effects, +# and heavyweight lab setup put it outside the token-free default-on class, so +# it does not run merely because its tools are installed. Every explicit and +# production-adapter Herdr call is routed through fm-herdr-lab.sh. The scenario +# proves: # - a live blocked status is classified and durably queued while away; # - a pending Pi composer refuses injection and receives no forced Enter; # - the existing wedge alarm remains observable and deduped; @@ -15,20 +18,14 @@ set -u # shellcheck source=tests/lib.sh . "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate opt-in FM_AFK_PI_HERDR_E2E herdr jq pi python3 + # shellcheck source=/dev/null . "$ROOT/bin/fm-supervise-daemon.sh" # shellcheck source=/dev/null . "$ROOT/bin/fm-backend.sh" -if [ "${FM_AFK_PI_HERDR_E2E:-0}" != 1 ]; then - echo "skip: set FM_AFK_PI_HERDR_E2E=1 to run the real Pi/Herdr away-return regression" - exit 0 -fi - -for tool in herdr jq pi python3; do - command -v "$tool" >/dev/null 2>&1 || { echo "skip: $tool not found"; exit 0; } -done - LAB_HELPER=${HERDR_LAB_HELPER:-$ROOT/bin/fm-herdr-lab.sh} SESSION=$("$LAB_HELPER" name fm-afk-pi-return-e2e) TMP_ROOT=$(fm_test_tmproot fm-afk-pi-return-e2e) diff --git a/tests/fm-backend-orca.test.sh b/tests/fm-backend-orca.test.sh index 8268f249e91..d3eb3bc49c0 100755 --- a/tests/fm-backend-orca.test.sh +++ b/tests/fm-backend-orca.test.sh @@ -539,7 +539,7 @@ test_spawn_writes_orca_metadata_and_launches_harness() { "spawn should reuse the implicit terminal returned by Orca worktree creation" assert_contains "$(cat "$log")" $'orca\x1f''terminal'$'\x1f''send'$'\x1f''--terminal'$'\x1f''term-spawn'$'\x1f''--text'$'\x1f''export GOTMPDIR=/tmp/fm-orcaspawnz1/gotmp'$'\x1f''--enter'$'\x1f''--json' \ "spawn did not export GOTMPDIR through the Orca terminal" - assert_contains "$(cat "$log")" "CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\"}'" \ + assert_contains "$(cat "$log")" "CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\",\"attribution\":{\"commit\":\"\",\"pr\":\"\",\"sessionUrl\":false}}'" \ "spawn did not send the selected harness launch command through Orca" rm -rf "/tmp/fm-$id" pass "fm-spawn.sh --backend orca: reuses implicit terminal, records metadata, launches harness" diff --git a/tests/fm-backend.test.sh b/tests/fm-backend.test.sh index 113642657c4..bc50e8c3847 100755 --- a/tests/fm-backend.test.sh +++ b/tests/fm-backend.test.sh @@ -44,6 +44,12 @@ TMP_ROOT=$(fm_test_tmproot fm-backend-tests) # and adds no launch prefix, since fm-spawn only prefixes a non-empty value. SPAWN_HOME="$TMP_ROOT/user-home" mkdir -p "$SPAWN_HOME" +# Pool locks are rooted in FM_HOME/state, independently of per-call state +# overrides. A developer checkout may already have state/ while clean CI does +# not, so every spawn below must use this fixture-owned operational home. +SPAWN_FM_HOME="$TMP_ROOT/spawn-firstmate-home" +mkdir -p "$SPAWN_FM_HOME/state" +SPAWN_FM_HOME=$(cd "$SPAWN_FM_HOME" && pwd -P) write_spawn_brief() { # <file> <id> cat > "$1" <<EOF @@ -815,7 +821,7 @@ run_spawn_case() { # <bin-root> <fakebin> <log> <state> <data> <config> <proj> local bin=$1 fb=$2 log=$3 state=$4 data=$5 config=$6 proj=$7; shift 7 [ "${1:-}" = -- ] && shift : > "$log" - env PATH="$fb:$PATH" FM_ROOT_OVERRIDE="$bin" HOME="$SPAWN_HOME" CLAUDE_CONFIG_DIR='' \ + env PATH="$fb:$PATH" FM_ROOT_OVERRIDE="$bin" FM_HOME="$SPAWN_FM_HOME" HOME="$SPAWN_HOME" CLAUDE_CONFIG_DIR='' \ FM_STATE_OVERRIDE="$state" FM_DATA_OVERRIDE="$data" FM_CONFIG_OVERRIDE="$config" \ FM_PROJECTS_OVERRIDE="$TMP_ROOT/unused-projects" \ FM_SPAWN_NO_GUARD=1 TMUX="fake,1,0" FM_TMUX_LOG="$log" \ @@ -882,7 +888,7 @@ SH } run_spawn_symlink_case() { # <label> <physical|logical> - local label=$1 first_reply=$2 real_root link_root proj wt id fb data state config log out rc proj_phys initial_path + local label=$1 first_reply=$2 real_root link_root proj wt id fb data state config log out rc proj_phys initial_path logical_lock physical_lock real_root="$TMP_ROOT/symlink-real-$label"; link_root="$TMP_ROOT/symlink-link-$label" mkdir -p "$real_root" ln -s "$real_root" "$link_root" @@ -896,6 +902,19 @@ run_spawn_symlink_case() { # <label> <physical|logical> # fm-spawn.sh's own PROJ_ABS_REAL computes, including any symlink layers # ABOVE this test's own synthetic real_root/link_root pair. proj_phys=$(cd "$real_root/proj" && pwd -P) + # Compare identities through the lock owner before driving the real spawn. + # Both paths must bind to this fixture home, never the ambient checkout. + logical_lock=$(FM_HOME="$SPAWN_FM_HOME" FM_STATE_OVERRIDE="$TMP_ROOT/lock-probe-state" \ + bash -c '. "$1"; fm_treehouse_project_lock_path "$2"' _ "$ROOT/bin/fm-wake-lib.sh" "$proj") \ + || fail "logical project path did not resolve a fixture-owned pool lock" + physical_lock=$(FM_HOME="$SPAWN_FM_HOME" FM_STATE_OVERRIDE="$TMP_ROOT/lock-probe-state" \ + bash -c '. "$1"; fm_treehouse_project_lock_path "$2"' _ "$ROOT/bin/fm-wake-lib.sh" "$proj_phys") \ + || fail "physical project path did not resolve a fixture-owned pool lock" + [ "$logical_lock" = "$physical_lock" ] || fail "symlink and physical project paths resolved different pool locks" + case "$logical_lock" in + "$SPAWN_FM_HOME/state/"*) ;; + *) fail "pool lock escaped the fixture operational home: $logical_lock" ;; + esac case "$first_reply" in physical) initial_path=$proj_phys ;; logical) initial_path=$proj ;; @@ -1079,7 +1098,7 @@ test_spawn_default_backend_writes_no_meta_field() { state="$TMP_ROOT/nobackend-state"; config="$TMP_ROOT/nobackend-config" mkdir -p "$state" "$config" - out=$(PATH="$fb:$PATH" FM_ROOT_OVERRIDE="$ROOT" HOME="$SPAWN_HOME" CLAUDE_CONFIG_DIR='' \ + out=$(PATH="$fb:$PATH" FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$SPAWN_FM_HOME" HOME="$SPAWN_HOME" CLAUDE_CONFIG_DIR='' \ FM_STATE_OVERRIDE="$state" FM_DATA_OVERRIDE="$data" FM_CONFIG_OVERRIDE="$config" \ FM_PROJECTS_OVERRIDE="$TMP_ROOT/unused-projects" FM_SPAWN_NO_GUARD=1 TMUX="fake,1,0" \ FM_TMUX_LOG="$TMP_ROOT/nobackend.log" \ @@ -1103,7 +1122,7 @@ test_spawn_explicit_backend_flag_beats_autodetect_herdr_env() { # HERDR_ENV=1 is present (as if firstmate itself were running under herdr), # but an explicit --backend tmux flag must still win outright. - out=$(PATH="$fb:$PATH" FM_ROOT_OVERRIDE="$ROOT" HOME="$SPAWN_HOME" CLAUDE_CONFIG_DIR='' \ + out=$(PATH="$fb:$PATH" FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$SPAWN_FM_HOME" HOME="$SPAWN_HOME" CLAUDE_CONFIG_DIR='' \ FM_STATE_OVERRIDE="$state" FM_DATA_OVERRIDE="$data" FM_CONFIG_OVERRIDE="$config" \ FM_PROJECTS_OVERRIDE="$TMP_ROOT/unused-projects" FM_SPAWN_NO_GUARD=1 TMUX="fake,1,0" HERDR_ENV=1 \ FM_TMUX_LOG="$TMP_ROOT/explicit-backend.log" \ @@ -1130,7 +1149,7 @@ test_spawn_autodetect_nesting_resolves_tmux_silently() { # (tmux nested inside a herdr pane) - the full fm-spawn.sh pipeline, not just # fm_backend_name, must resolve this to tmux and stay completely silent about # it (today's default path, byte-identical). - out=$(PATH="$fb:$PATH" FM_ROOT_OVERRIDE="$ROOT" HOME="$SPAWN_HOME" CLAUDE_CONFIG_DIR='' \ + out=$(PATH="$fb:$PATH" FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$SPAWN_FM_HOME" HOME="$SPAWN_HOME" CLAUDE_CONFIG_DIR='' \ FM_STATE_OVERRIDE="$state" FM_DATA_OVERRIDE="$data" FM_CONFIG_OVERRIDE="$config" \ FM_PROJECTS_OVERRIDE="$TMP_ROOT/unused-projects" FM_SPAWN_NO_GUARD=1 TMUX="fake,1,0" HERDR_ENV=1 \ FM_TMUX_LOG="$TMP_ROOT/nest.log" \ diff --git a/tests/fm-backlog-atomicity.test.sh b/tests/fm-backlog-atomicity.test.sh index 06a22b9abb0..2bbeda6923d 100755 --- a/tests/fm-backlog-atomicity.test.sh +++ b/tests/fm-backlog-atomicity.test.sh @@ -2401,6 +2401,13 @@ test_bootstrap_refuses_a_symlinked_state_directory_before_reconciliation() { test_bootstrap_stops_when_data_disappears_before_reconciliation() { local case_dir id saved out rc=0 + # The data-removal fault is injected by a fake stat on PATH; on Darwin the + # budget link-count helper now calls /usr/bin/stat directly, so the fake can + # never fire there. Skip the Darwin run of this case. + if [ "$(uname)" = Darwin ]; then + pass "bootstrap data-disappears fault injection is PATH-based; skipped on Darwin where stat is /usr/bin/stat" + return + fi id=atomic-bootstrap-data-race-b11 case_dir=$(make_home bootstrap-data-race) add_item "$case_dir" "$id" diff --git a/tests/fm-bearings-board-lavish-live-e2e.test.sh b/tests/fm-bearings-board-lavish-live-e2e.test.sh new file mode 100755 index 00000000000..a413e27c3a0 --- /dev/null +++ b/tests/fm-bearings-board-lavish-live-e2e.test.sh @@ -0,0 +1,120 @@ +#!/usr/bin/env bash +# tests/fm-bearings-board-lavish-live-e2e.test.sh - live drift guard proving +# the real lavish-axi still behaves the way bin/fm-bearings-board.sh's session +# liveness check is written against. +# +# Why this file exists: the build's "is this board actually live" verdict comes +# from what lavish-axi emits, which is a surface the vendor controls and changes +# without notice. The defect this guards was exactly that - opening a session +# the captain had ended from the browser EXITS 0 while refusing to reopen, so a +# build that trusted the exit status armed a poll against a dead session and the +# board read "not listening" with nobody watching it. A stubbed lavish-axi can +# only confirm the assumption already written into the stub, so the assumption +# itself needs a run against the real tool. +# +# The captain-ended state is reached through the same server route the browser's +# End session button calls, so no browser is needed and nothing here depends on +# a human. The artifact is a scratch page in a temporary directory, and the +# session it opens is ended again before the guard returns. +# +# Standard CI has no lavish-axi, so this reports a capability skip there. The +# portable counterpart in tests/fm-bearings-board.test.sh pins the build's logic +# in CI against a stub that reproduces these shapes. Run this guard after a +# lavish-axi upgrade and before trusting refreshed evidence. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" + +fm_live_gate default-on FM_BEARINGS_LAVISH_LIVE lavish-axi jq curl + +pass() { printf 'ok - %s\n' "$1"; } +note() { printf '# %s\n' "$1"; } + +LAB='' +cleanup() { + [ -z "$LAB" ] || { + [ ! -f "$LAB/.lavish/bearings-board.html" ] \ + || lavish-axi end "$LAB/.lavish/bearings-board.html" >/dev/null 2>&1 || true + rm -rf "$LAB" + } +} +fail() { printf 'not ok - %s\n' "$1" >&2; cleanup; exit 1; } +trap cleanup EXIT + +VERSION=$(lavish-axi --version 2>/dev/null | tr -d '[:space:]') +note "lavish-axi ${VERSION:-version-unknown}" + +LAB=$(mktemp -d "${TMPDIR:-/tmp}/fm-bearings-lavish-live.XXXXXX") || fail "cannot create the guard lab" +LAB=$(cd -P -- "$LAB" && pwd -P) +mkdir -p "$LAB/state" "$LAB/data" + +cat > "$LAB/payload.json" <<'JSON' +{ + "schema": "fm-bearings-board.v1", + "home": "lavish-live-guard", + "generated": "2026-01-01T00:00Z", + "prs_live": false, + "captains_call": [ + { + "key": "sample-live-guard-call", + "type": "decision", + "repo": "sample", + "title": "Guard placeholder", + "options": [{ "value": "yes", "label": "Yes" }] + } + ], + "underway": [], + "landed": [], + "charted": [] +} +JSON + +run_board() { + FM_HOME="$LAB" FM_STATE_OVERRIDE="$LAB/state" FM_DATA_OVERRIDE="$LAB/data" \ + FM_PROCEVENT_CLAIM_ROOT="$LAB/procevent-claims" \ + "$ROOT/bin/fm-bearings-board.sh" "$@" +} + +BOARD="$LAB/.lavish/bearings-board.html" +run_board build "$LAB/payload.json" >/dev/null 2>&1 || fail "the guard board did not build" +[ -f "$BOARD" ] || fail "the guard board was not published" + +url=$(lavish-axi "$BOARD" | sed -n 's/^[[:space:]]*url:[[:space:]]*//p' | head -1 | tr -d '"') +case "$url" in + http://*/session/*) ;; + *) fail "could not read the guard board session url: $url" ;; +esac +key=${url##*/} +base=${url%/session/*} + +# End it exactly as the browser's End session button does. +curl -fsS -X POST "$base/api/$key/end" >/dev/null 2>&1 \ + || fail "could not end the guard board session as the captain" + +# ASSUMPTION UNDER GUARD: this exits 0 while reporting the session is not live. +set +e +ended_out=$(lavish-axi "$BOARD" 2>&1) +ended_rc=$? +set -e +[ "$ended_rc" -eq 0 ] \ + || fail "lavish-axi ${VERSION:-version-unknown} now exits $ended_rc on a captain-ended session; the board build's liveness check must be revisited" +ended_status=$(printf '%s\n' "$ended_out" | sed -n 's/^[[:space:]]*status:[[:space:]]*//p' | head -1 | tr -d '"') +[ "$ended_status" != opened ] \ + || fail "lavish-axi ${VERSION:-version-unknown} silently reopened a captain-ended session; the board build's liveness check must be revisited" +lavish-axi 2>/dev/null | grep -F "$BOARD," | grep -q ',open,' \ + && fail "lavish-axi ${VERSION:-version-unknown} still lists a captain-ended session as open; the board build's liveness check must be revisited" +pass "lavish-axi ${VERSION:-version-unknown} reports a captain-ended session without reopening it and without failing" + +# THE BEHAVIOR UNDER GUARD: the build must not accept that, and must recover. +out=$(run_board build "$LAB/payload.json" 2>&1) \ + || fail "the board build refused a recoverable captain-ended session: $out" +case "$out" in + *"session: reopened"*) ;; + *) fail "the board build did not reopen the captain-ended session: $out" ;; +esac +lavish-axi 2>/dev/null | grep -F "$BOARD," | grep -q ',open,' \ + || fail "the board build reported success while the session was still not live" +pass "the board build reopens a captain-ended session against real lavish-axi instead of arming a dead one" diff --git a/tests/fm-bearings-board-render.test.sh b/tests/fm-bearings-board-render.test.sh index afa6b9350cc..cf26fd31428 100755 --- a/tests/fm-bearings-board-render.test.sh +++ b/tests/fm-bearings-board-render.test.sh @@ -20,9 +20,39 @@ command -v node >/dev/null 2>&1 || { echo "skip: node not found"; exit 0; } make_home() { # <name> local home="$TMP_ROOT/$1" fakebin + # A build starts a listener for the board it publishes. Registered with + # tests/lib.sh, not with a shell array: make_home runs inside a command + # substitution, where an array append never reaches the caller. + fm_test_track_procevent_home "$home" "$home/procevent-claims" mkdir -p "$home/state" "$home/data" fakebin=$(fm_fakebin "$home") - fm_fake_exit0 "$fakebin" lavish-axi + # The build proves the board session is live before it arms anything, so the + # stub reports the opened shape the real lavish-axi emits. This suite is about + # what the template renders, not about session liveness, which + # tests/fm-bearings-board.test.sh owns. + cat > "$fakebin/lavish-axi" <<'SH' +#!/usr/bin/env bash +case "${1-}" in + --version) printf '0.1.61\n' ;; + '') + printf 'sessions[1]{file,status,url,pending_prompts}:\n' + [ ! -s "$FM_HOME/lavish-open" ] \ + || printf ' %s,open,"http://127.0.0.1/session/render",0\n' "$(cat "$FM_HOME/lavish-open")" + ;; + poll) + # Bounded, so a listener that escapes its test stops on its own. + while [ "$SECONDS" -lt "${FM_TEST_STUB_MAX_BLOCK_SECONDS:-120}" ]; do sleep 1; done + exit 75 + ;; + *) + real=$(cd "$(dirname "$1")" && pwd -P)/$(basename "$1") + printf '%s\n' "$real" > "$FM_HOME/lavish-open" + printf 'session:\n status: opened\n' + ;; +esac +exit 0 +SH + chmod +x "$fakebin/lavish-axi" printf '%s\n' "$home" } diff --git a/tests/fm-bearings-board.test.sh b/tests/fm-bearings-board.test.sh index d87acb652a2..b59010036e9 100644 --- a/tests/fm-bearings-board.test.sh +++ b/tests/fm-bearings-board.test.sh @@ -13,20 +13,100 @@ TMP_ROOT=$(fm_test_tmproot fm-bearings-board) command -v jq >/dev/null 2>&1 || { echo "skip: jq not found"; exit 0; } +# A lavish-axi stub that reproduces the shapes verified against the real +# lavish-axi 0.1.61, because the build's liveness verdict is read from what the +# vendor emits. The load-bearing shape is the refusal: opening a session the +# captain ended from the browser EXITS 0 while reporting `status: user-ended`, +# and that session is absent from the server's listing. `--reopen` restores it. +# Markers under lavish-state drive the fixture: `user-ended` makes the next +# plain open refuse, and `refuse-reopen` makes even --reopen leave it dead. make_home() { # <name> local home="$TMP_ROOT/$1" fakebin - mkdir -p "$home/state" "$home/data" + # Registered with tests/lib.sh, not with a shell array: make_home is called + # inside a command substitution, so an array append here never reaches the + # caller and every listener this suite started used to survive the run. + fm_test_track_procevent_home "$home" "$home/procevent-claims" + mkdir -p "$home/state" "$home/data" "$home/lavish-state" fakebin=$(fm_fakebin "$home") - fm_fake_exit0 "$fakebin" lavish-axi + cat > "$fakebin/lavish-axi" <<'SH' +#!/usr/bin/env bash +set -u +state=${LAVISH_FAKE_STATE:?} +emit() { # <canonical-file> <status> + printf 'session:\n' + printf ' file: %s\n' "$1" + printf ' url: "http://127.0.0.1:4387/session/deadbeef"\n' + printf ' status: %s\n' "$2" +} +case "${1-}" in + --version) printf '0.1.61\n'; exit 0 ;; + poll) + # A real blocking listener: it returns only when the trigger appears, so a + # live owner in these tests is a live process rather than a timing artifact. + # Both waits are bounded, so a listener that escapes its test cannot keep + # spawning processes for as long as the host stays up. + limit=${FM_TEST_STUB_MAX_BLOCK_SECONDS:-120} + while [ ! -e "$state/poll-trigger" ]; do + [ "$SECONDS" -lt "$limit" ] || exit 75 + sleep 0.05 + done + printf 'session:\n status: ended\n' + if [ -e "$state/hold-after-terminal" ]; then + : > "$state/terminal-emitted" + while [ -e "$state/hold-after-terminal" ]; do + [ "$SECONDS" -lt "$limit" ] || exit 75 + sleep 0.05 + done + fi + exit 0 + ;; + '') + if [ -e "$state/end-before-next-list" ]; then + : > "$state/open" + rm -f "$state/end-before-next-list" + fi + printf 'sessions[1]{file,status,url,pending_prompts}:\n' + if [ -s "$state/open" ]; then + while IFS= read -r listed; do + [ -n "$listed" ] || continue + printf ' %s,open,"http://127.0.0.1:4387/session/deadbeef",0\n' "$listed" + done < "$state/open" + fi + exit 0 + ;; + end) : > "$state/open"; printf 'session:\n status: ended\n'; exit 0 ;; +esac +file=$1 +shift +reopen=0 +for arg in "$@"; do [ "$arg" != --reopen ] || reopen=1; done +real=$(cd "$(dirname "$file")" && pwd -P)/$(basename "$file") +if [ -e "$state/user-ended" ] && [ "$reopen" = 0 ]; then + emit "$real" user-ended + exit 0 +fi +if [ -e "$state/refuse-reopen" ]; then + emit "$real" user-ended + exit 0 +fi +rm -f -- "$state/user-ended" +printf '%s\n' "$real" > "$state/open" +emit "$real" opened +exit 0 +SH + chmod +x "$fakebin/lavish-axi" printf '%s\n' "$home" } +end_session_as_captain() { : > "$1/lavish-state/user-ended"; : > "$1/lavish-state/open"; } + run_board() { # <home> <args...> local home=$1 shift PATH="$home/fakebin:$PATH" FM_HOME="$home" \ FM_STATE_OVERRIDE="$home/state" FM_DATA_OVERRIDE="$home/data" \ FM_PROCEVENT_CLAIM_ROOT="$home/procevent-claims" \ + LAVISH_FAKE_STATE="$home/lavish-state" \ "$BOARD" "$@" } @@ -154,6 +234,12 @@ test_build_refuses_malformed_payloads_before_touching_the_board() { set +e; out=$(run_board "$home" build "$data" 2>&1); rc=$?; set -e [ "$rc" -ne 0 ] || fail "a negative omitted-warning count was accepted" + write_valid_payload "$data" + jq '.captains_call[0].subject = {"artifact":"quota-axi","version":"0.1"}' "$data" > "$data.tmp" \ + && mv "$data.tmp" "$data" + set +e; out=$(run_board "$home" build "$data" 2>&1); rc=$?; set -e + [ "$rc" -ne 0 ] || fail "an invalid structured version subject was accepted" + write_valid_payload "$data" jq '.captains_call[0].type = "verdict"' "$data" > "$data.tmp" && mv "$data.tmp" "$data" set +e; out=$(run_board "$home" build "$data" 2>&1); rc=$?; set -e @@ -219,13 +305,19 @@ test_build_injects_binds_then_arms() { assert_contains "$out" "armed: " "the first build did not arm the board source: $out" assert_present "$board" "build reported success without a board" - # Round-trip: the payload extracted from the built page is byte-for-byte the - # same JSON document, and the escaped </script> string can no longer - # terminate the data block. + # Round-trip: apart from the reconcile choice the build adds to every + # decision card, the payload extracted from the built page is the same JSON + # document, and the escaped </script> string can no longer terminate the + # data block. extract_payload "$board" | jq -S . > "$home/extracted.json" \ || fail "the built board does not carry parseable payload JSON" - jq -S . "$data" > "$home/expected.json" - diff -u "$home/expected.json" "$home/extracted.json" >/dev/null \ + jq -S '.captains_call = [.captains_call[] + | .options = [.options[] | select(.value != "reconcile")]]' \ + "$home/extracted.json" > "$home/stripped.json" + jq -S '.captains_call = [.captains_call[] + | .options = [.options[] | select(.value != "reconcile")]]' \ + "$data" > "$home/expected.json" + diff -u "$home/expected.json" "$home/stripped.json" >/dev/null \ || fail "the injected payload does not round-trip to the input document" grep -qF '</script><b>' "$board" \ && fail "a payload string embedded a live closing script tag in the page" @@ -285,7 +377,16 @@ SH chmod +x "$runtime/bin/fm-procevent-lavish.sh" cat > "$home/fakebin/lavish-axi" <<'SH' #!/usr/bin/env bash +if [ -z "${1:-}" ]; then + printf 'sessions[1]{file,status,url,pending_prompts}:\n' + [ ! -s "$FM_HOME/order-open" ] \ + || printf ' %s,open,"http://127.0.0.1/session/order",0\n' "$(cat "$FM_HOME/order-open")" + exit 0 +fi if [ "${1:-}" != poll ]; then + real=$(cd "$(dirname "$1")" && pwd -P)/$(basename "$1") + printf '%s\n' "$real" > "$FM_HOME/order-open" + printf 'session:\n status: opened\n' exit 0 fi cat <<EOF @@ -293,7 +394,7 @@ session: status: feedback session_ended: false prompts[1]{uid,prompt,selector,tag,text}: - "2","Order proof: yes\\n\\nContext data:\\n{\\n \\"question\\": \\"$ORDER_PROOF_HOLD\\",\\n \\"answer\\": \\"yes\\"\\n}","form",choice,"Order proof: yes" + "2","Order proof: yes\\n\\nContext data:\\n{\\n \\"schema\\": \\"fm-bearings-answer.v1\\",\\n \\"question\\": \\"$ORDER_PROOF_HOLD\\",\\n \\"selection\\": \\"yes\\",\\n \\"note\\": \\"\\"\\n}","form",choice,"Order proof: yes" EOF SH chmod +x "$home/fakebin/lavish-axi" @@ -404,6 +505,264 @@ test_charted_kind_is_optional_and_accepts_both_values() { pass "charted kind is optional and accepts queued and warning" } + +# --- part 1: never arm a poll on an ended session --------------------------- + +test_build_reopens_a_session_the_captain_ended() { + local home data board out sid claim old_pid old_token new_pid new_token + home=$(make_home ended-session) + data="$home/payload.json" + board="$home/.lavish/bearings-board.html" + write_valid_payload "$data" + run_board "$home" build "$data" >/dev/null || fail "the first build failed" + sid=$(run_lavish_source_id "$home" "$board") + claim="$home/procevent-claims/$sid.claim" + old_pid=$(sed -n '2p' "$claim") + old_token=$(sed -n '3p' "$claim") + + # The reported case: the captain ends the board from the browser, so opening + # it again keeps the same session id, reports it ended, and EXITS 0. A build + # that trusts the exit status arms a poll nothing can ever attach to. + : > "$home/lavish-state/hold-after-terminal" + : > "$home/lavish-state/poll-trigger" + for _ in $(seq 1 100); do + [ -e "$home/lavish-state/terminal-emitted" ] && break + sleep 0.05 + done + [ -e "$home/lavish-state/terminal-emitted" ] \ + || fail "the old listener did not receive its terminal result" + rm -f "$home/lavish-state/poll-trigger" + end_session_as_captain "$home" + out=$(run_board "$home" build "$data") || fail "the rebuild refused a recoverable ended session" + rm -f "$home/lavish-state/hold-after-terminal" + assert_contains "$out" "session: reopened" \ + "the rebuild did not reopen the ended session: $out" + [ ! -e "$home/lavish-state/user-ended" ] \ + || fail "the rebuild reported success while the session was still ended" + new_pid=$(sed -n '2p' "$claim") + new_token=$(sed -n '3p' "$claim") + [ "$new_pid" != "$old_pid" ] || [ "$new_token" != "$old_token" ] \ + || fail "the rebuild accepted the pre-reopen source generation" + [ "$(run_procevent "$home" list | awk -v id="$sid" 'NR > 1 && $1 == id { print $3 }')" = live ] \ + || fail "the reopened board has no live listener" + pass "a board build reopens a session the captain ended instead of arming a dead one" +} + +test_build_reopens_when_an_opened_session_ends_before_listing() { + local home data out board sid + home=$(make_home establish-list-race) + data="$home/payload.json" + board="$home/.lavish/bearings-board.html" + write_valid_payload "$data" + : > "$home/lavish-state/end-before-next-list" + out=$(run_board "$home" build "$data") || fail "the raced session build failed: $out" + assert_contains "$out" "session: reopened" \ + "the build trusted an opened response after the server no longer listed it: $out" + sid=$(run_lavish_source_id "$home" "$board") + [ -s "$home/lavish-state/open" ] || fail "the raced session was not live before arming" + [ "$(run_procevent "$home" list | awk -v id="$sid" 'NR > 1 && $1 == id { print $3 }')" = live ] \ + || fail "the replacement session did not receive a live listener" + pass "build reopens a session that ends between establish and listing" +} + +test_build_refuses_to_arm_when_the_session_stays_ended() { + local home data rc out sid + home=$(make_home dead-session) + data="$home/payload.json" + write_valid_payload "$data" + # An ended session that will not come back: the build must stop rather than + # register a poll against it. + : > "$home/lavish-state/refuse-reopen" + set +e + out=$(run_board "$home" build "$data" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "build armed a poll on a session that stayed ended: $out" + assert_contains "$out" "ended session" "the refusal did not say why: $out" + sid=$(run_lavish_source_id "$home" "$home/.lavish/bearings-board.html") + ! run_decisions "$home" binding "$sid" >/dev/null 2>&1 \ + || fail "build bound the board to a session that stayed ended" + ! run_procevent "$home" list | awk 'NR > 1 { print $1 }' | grep -Fxq "$sid" \ + || fail "build armed the board against a session that stayed ended" + pass "build refuses to arm a poll on a session that stays ended" +} + +test_build_starts_a_listener_for_an_already_armed_board() { + local home data board out sid claim + home=$(make_home relisten) + data="$home/payload.json" + board="$home/.lavish/bearings-board.html" + write_valid_payload "$data" + run_board "$home" build "$data" >/dev/null || fail "the first build failed" + sid=$(run_lavish_source_id "$home" "$board") + + # Registered is not listening: drop the listener the way a crashed generation + # would, then rebuild. `already-armed` must not be the end of the story. + claim="$home/procevent-claims/$sid.claim" + assert_present "$claim" "the first build left no listener to lose" + kill -KILL -"$(sed -n '2p' "$claim")" 2>/dev/null || true + kill -KILL "$(sed -n '2p' "$claim")" 2>/dev/null || true + sleep 1 + + out=$(run_board "$home" build "$data") || fail "the rebuild failed" + assert_contains "$out" "already-armed: $sid" "the rebuild re-registered the source: $out" + [ "$(run_procevent "$home" list | awk -v id="$sid" 'NR > 1 && $1 == id { print $3 }')" = live ] \ + || fail "the rebuilt board is registered but nothing is listening" + pass "a rebuild starts a listener when an already-armed board has none" +} + +# --- part 2: a landed subject is not a live call ---------------------------- + +test_build_drops_decision_cards_whose_subject_already_landed() { + local home data board out + home=$(make_home landed-cards) + data="$home/payload.json" + board="$home/.lavish/bearings-board.html" + write_valid_payload "$data" + jq '.captains_call = [ + {"key":"landed-by-task","type":"decision","repo":"sample","title":"Already shipped", + "options":[{"value":"yes","label":"Yes"}]}, + {"key":"timeout-reattach","type":"decision","repo":"sample","title":"Already merged", + "pr_url":"https://github.com/sample/sample/pull/7", + "options":[{"value":"yes","label":"Yes"}]}, + {"key":"quota-version","type":"decision","repo":"sample","title":"Old quota release", + "subject":{"artifact":"quota-axi","version":"0.1.37"}, + "options":[{"value":"yes","label":"Yes"}]}, + {"key":"still-open","type":"decision","repo":"sample","title":"Genuinely open", + "subject":{"artifact":"quota-axi","version":"0.2.0"}, + "options":[{"value":"yes","label":"Yes"}]} + ] + | .landed = [ + {"id":"landed-by-task","repo":"sample","what":"shipped it","owner":"crew"}, + {"id":"some-other-task","repo":"sample","what":"merged timeout reattach","owner":"crew", + "pr_url":"https://github.com/sample/sample/pull/7"}, + {"id":"quota-release","repo":"sample","what":"published quota-axi","owner":"crew", + "subject":{"artifact":"quota-axi","version":"0.1.38"}}, + {"id":"unrelated\nstill-open","repo":"sample","what":"unrelated multiline identity","owner":"crew"} + ]' "$data" > "$data.tmp" && mv "$data.tmp" "$data" + + out=$(run_board "$home" build "$data" 2>&1) || fail "the hygiene build failed: $out" + assert_contains "$out" "dropped-landed-card: landed-by-task" \ + "the build did not report dropping the landed work item card: $out" + assert_contains "$out" "dropped-landed-card: timeout-reattach" \ + "the build did not report dropping the merged timeout/reattach card: $out" + assert_contains "$out" "dropped-landed-card: quota-version" \ + "the build did not report dropping the superseded quota-axi version card: $out" + extract_payload "$board" | jq -e '[.captains_call[].key] == ["still-open"]' >/dev/null \ + || fail "the board dropped an open card or kept one whose subject already landed" + pass "build drops decision cards whose subject already landed and keeps open ones" +} + +test_build_keeps_a_decision_absent_from_the_main_backlog() { + local home data board out + home=$(make_home remote-decision-card) + data="$home/payload.json" + board="$home/.lavish/bearings-board.html" + cp "$ROOT/.tasks.toml" "$home/.tasks.toml" + cat > "$home/data/backlog.md" <<'EOF' +## In flight + +## Queued + +## Done +EOF + write_valid_payload "$data" + jq '.captains_call = [{ + "key":"remote-mate-call","type":"decision","repo":"sample", + "title":"Remote secondmate decision", + "options":[{"value":"yes","label":"Yes"}] + }] + | .landed = []' "$data" > "$data.tmp" && mv "$data.tmp" "$data" + + out=$(run_board "$home" build "$data" 2>&1) || fail "the remote-card build failed: $out" + assert_not_contains "$out" "dropped-landed-card: remote-mate-call" \ + "an absent remote card was reported as landed: $out" + extract_payload "$board" | jq -e ' + [.captains_call[] | select(.key == "remote-mate-call")] | length == 1 + ' >/dev/null || fail "the hygiene check dropped a decision absent from the main backlog" + pass "build keeps remote decisions absent from the main backlog" +} + +# --- part 3: every decision card offers reconcile --------------------------- + +test_build_fails_when_reconcile_cannot_establish_a_listener() { + local home data out rc sid + home=$(make_home no-listener) + data="$home/payload.json" + write_valid_payload "$data" + run_board "$home" build "$data" >/dev/null || fail "could not establish the listener fixture" + sid=$(run_lavish_source_id "$home" "$home/.lavish/bearings-board.html") + cat > "$home/fakebin/ps" <<'SH' +#!/usr/bin/env bash +exit 1 +SH + chmod +x "$home/fakebin/ps" + set +e + out=$(FM_PROC_ROOT_OVERRIDE="$home/no-proc" run_board "$home" build "$data" 2>&1) + rc=$? + set -e + rm -f "$home/fakebin/ps" + [ "$rc" -ne 0 ] || fail "a build with an uncertain listener reported success: $out" + assert_contains "$out" "source $sid is not listening after reconcile" \ + "the refusal did not name the source: $out" + assert_contains "$out" "observed owner: uncertain" \ + "the refusal did not name the observed owner: $out" + pass "build fails when reconcile cannot prove a live listener" +} + +test_every_decision_card_carries_the_reconcile_choice() { + local home data board + home=$(make_home reconcile-option) + data="$home/payload.json" + board="$home/.lavish/bearings-board.html" + write_valid_payload "$data" + run_board "$home" build "$data" >/dev/null || fail "the reconcile-option build failed" + extract_payload "$board" | jq -e ' + ([.captains_call[] | select(.type == "decision")] | length) > 0 + and ([.captains_call[] + | select(.type == "decision") + | ([.options[] | select(.value == "reconcile")] | length) == 1 + and ([.options[] | select(.value == "reconcile") | .label | length > 0] | all)] | all) + ' >/dev/null || fail "a decision card was published without the reconcile choice" + extract_payload "$board" | jq -e ' + ([.captains_call[] | select(.type != "decision") + | .options[] | select(.value == "reconcile")] | length) == 0 + ' >/dev/null || fail "reconcile was injected into a non-decision card" + pass "every decision card carries exactly one reconcile choice" +} + +test_build_refuses_a_payload_that_occupies_the_reconcile_value() { + local home data rc out + home=$(make_home reconcile-reserved) + data="$home/payload.json" + write_valid_payload "$data" + jq '.captains_call[0].options += [{"value":"reconcile","label":"Something else"}]' \ + "$data" > "$data.tmp" && mv "$data.tmp" "$data" + set +e + out=$(run_board "$home" build "$data" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "a payload occupying the reserved reconcile value was accepted" + assert_absent "$home/.lavish/bearings-board.html" "a refused payload still produced a board" + pass "build refuses a payload that occupies the reserved reconcile value" +} + +test_build_refuses_a_nondecision_reconcile_value() { + local home data rc out + home=$(make_home merge-reconcile-reserved) + data="$home/payload.json" + write_valid_payload "$data" + jq '.captains_call[1].options += [{"value":"reconcile","label":"Merge action"}]' \ + "$data" > "$data.tmp" && mv "$data.tmp" "$data" + set +e + out=$(run_board "$home" build "$data" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "a merge card occupying the reconcile value was accepted" + assert_absent "$home/.lavish/bearings-board.html" "a refused merge card still produced a board" + pass "build reserves reconcile across non-decision cards" +} + test_path_is_stable_and_home_scoped test_build_refuses_malformed_payloads_before_touching_the_board test_charted_kind_is_optional_and_accepts_both_values @@ -412,3 +771,13 @@ test_registration_cannot_consume_before_any_origin_binding test_build_does_not_bind_or_arm_when_session_start_fails test_rebuild_is_idempotent_and_does_not_double_arm test_build_refuses_a_template_without_exactly_one_slot +test_build_reopens_a_session_the_captain_ended +test_build_reopens_when_an_opened_session_ends_before_listing +test_build_refuses_to_arm_when_the_session_stays_ended +test_build_starts_a_listener_for_an_already_armed_board +test_build_drops_decision_cards_whose_subject_already_landed +test_build_keeps_a_decision_absent_from_the_main_backlog +test_build_fails_when_reconcile_cannot_establish_a_listener +test_every_decision_card_carries_the_reconcile_choice +test_build_refuses_a_payload_that_occupies_the_reconcile_value +test_build_refuses_a_nondecision_reconcile_value diff --git a/tests/fm-bootstrap.test.sh b/tests/fm-bootstrap.test.sh index 431cc6b4bb7..77bea60fd4c 100755 --- a/tests/fm-bootstrap.test.sh +++ b/tests/fm-bootstrap.test.sh @@ -575,7 +575,7 @@ test_orca_backend_gates_orca_tool_only_when_selected() { printf '%s\n' manual > "$case_dir/home/config/backlog-backend" printf '%s\n' orca > "$case_dir/home/config/backend" fakebin=$(make_fake_toolchain "$case_dir") - out=$(PATH="$fakebin:$BASE_PATH" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ + out=$(PATH="$fakebin:$(fm_test_base_path_sans "$BASE_PATH" orca)" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ FM_FAKE_TREEHOUSE_LEASE_HELP=1 "$ROOT/bin/fm-bootstrap.sh") [ "$out" = "$missing_orca" ] || fail "backend=orca should require only the Orca-specific missing tool, got: $out" @@ -1524,7 +1524,7 @@ SH # split is a PARTITION: `skip` plus `only` together do exactly what `all` does, # with no step dropped and no step run twice. test_network_phase_partitions_the_run() { - local case_dir fakebin bash_env all_out skip_out only_out combined + local case_dir fakebin all_out skip_out only_out combined case_dir="$TMP_ROOT/network-phase" mkdir -p "$case_dir/home/config" printf '%s\n' manual > "$case_dir/home/config/backlog-backend" @@ -1532,32 +1532,23 @@ test_network_phase_partitions_the_run() { # Break the two diagnostics that stand for the two halves: a local tool floor # and the network GitHub-auth probe. rm -f "$fakebin/node" - bash_env="$case_dir/mask-node.bash" - cat > "$bash_env" <<'SH' -command() { - if [ "${1:-}" = -v ] && [ "${2:-}" = node ]; then - return 1 - fi - builtin command "$@" -} -SH cat > "$fakebin/gh" <<'SH' #!/usr/bin/env bash exit 1 SH chmod +x "$fakebin/gh" - all_out=$(PATH="$fakebin:$BASE_PATH" BASH_ENV="$bash_env" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ + all_out=$(PATH="$fakebin:$(fm_test_base_path_sans "$BASE_PATH" node)" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ FM_FAKE_TREEHOUSE_LEASE_HELP=1 "$ROOT/bin/fm-bootstrap.sh") assert_contains "$all_out" "MISSING: node (install:" "the unsplit run lost its local diagnostic" assert_contains "$all_out" "NEEDS_GH_AUTH" "the unsplit run lost its network diagnostic" - skip_out=$(PATH="$fakebin:$BASE_PATH" BASH_ENV="$bash_env" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ + skip_out=$(PATH="$fakebin:$(fm_test_base_path_sans "$BASE_PATH" node)" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ FM_FAKE_TREEHOUSE_LEASE_HELP=1 FM_BOOTSTRAP_NETWORK=skip "$ROOT/bin/fm-bootstrap.sh") assert_contains "$skip_out" "MISSING: node (install:" "the local half lost its own diagnostic" assert_not_contains "$skip_out" "NEEDS_GH_AUTH" "the local half still made a network call" - only_out=$(PATH="$fakebin:$BASE_PATH" BASH_ENV="$bash_env" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ + only_out=$(PATH="$fakebin:$(fm_test_base_path_sans "$BASE_PATH" node)" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ FM_FAKE_TREEHOUSE_LEASE_HELP=1 FM_BOOTSTRAP_NETWORK=only "$ROOT/bin/fm-bootstrap.sh") assert_contains "$only_out" "NEEDS_GH_AUTH" "the network half lost its own diagnostic" assert_not_contains "$only_out" "MISSING: node" "the network half repeated the local half's work" @@ -1568,7 +1559,7 @@ SH # A typo must never silently drop a safety sweep, so anything unrecognized # resolves to the complete run. - [ "$(PATH="$fakebin:$BASE_PATH" BASH_ENV="$bash_env" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ + [ "$(PATH="$fakebin:$(fm_test_base_path_sans "$BASE_PATH" node)" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ FM_FAKE_TREEHOUSE_LEASE_HELP=1 FM_BOOTSTRAP_NETWORK=sikp "$ROOT/bin/fm-bootstrap.sh")" = "$all_out" ] \ || fail "an unrecognized FM_BOOTSTRAP_NETWORK value did not fall back to the complete run" pass "bootstrap: FM_BOOTSTRAP_NETWORK partitions one run into local and network halves" diff --git a/tests/fm-calm-pi-extension.test.sh b/tests/fm-calm-pi-extension.test.sh index f3ebd755e00..a05178da3e3 100755 --- a/tests/fm-calm-pi-extension.test.sh +++ b/tests/fm-calm-pi-extension.test.sh @@ -72,6 +72,82 @@ find_chrome() { return 1 } +# Render an exported session in real Chrome and leave the DOM in <out_file>. +# +# Rendering is a vendor-tool step, not a Calm guarantee: the DOM assertions the +# caller runs afterwards are what protect the contract. Headless Chrome start-up +# is the part that fails intermittently on a loaded CI runner - it can exit +# before writing any DOM at all - and the original single unattended attempt +# discarded both Chrome's stderr and its exit status, so a CI break surfaced as +# a bare "could not render" with nothing in the log to tell a Chrome start-up +# crash apart from a real change in Pi's export shape. +# +# So: retry the render a bounded number of times on a fresh profile, and when +# every attempt fails, print the Chrome binary, its version, the installed Pi +# version, and each attempt's exit status, stderr tail, and whether the helper +# timed the attempt out - when it did, the exit status is only this helper's own +# kill signal. The extra flags remove Chrome's background-network and /dev/shm +# dependencies, which are the start-up surfaces that fail on a runner; neither +# changes the rendered DOM of a local file. +render_export_dom() { + local chrome=$1 source_file=$2 out_file=$3 pi_version=$4 + local attempt pid status wait_count wait_limit reap_wait log profile report timed_out + report="$TMP_ROOT/chrome-render-report.txt" + wait_limit=${FM_CHROME_RENDER_WAIT_TICKS:-300} + : >"$report" + for attempt in 1 2 3; do + log="$TMP_ROOT/chrome-render-$attempt.err" + profile="$TMP_ROOT/chrome-profile-$attempt" + rm -rf "$profile" + : >"$out_file" + "$chrome" \ + --headless=new \ + --disable-gpu \ + --no-sandbox \ + --disable-dev-shm-usage \ + --disable-background-networking \ + --user-data-dir="$profile" \ + --virtual-time-budget=2000 \ + --dump-dom \ + "file://$source_file" >"$out_file" 2>"$log" & + pid=$! + # Check the DOM before Chrome's liveness, so an attempt that writes the + # complete dump and exits immediately is still read as a success. + wait_count=0 + while [ "$wait_count" -lt "$wait_limit" ]; do + grep -Fq '</html>' "$out_file" 2>/dev/null && break + kill -0 "$pid" 2>/dev/null || break + sleep 0.1 + wait_count=$((wait_count + 1)) + done + timed_out=no + if [ "$wait_count" -ge "$wait_limit" ]; then + timed_out=yes + fi + kill "$pid" 2>/dev/null || true + # Chrome can retain --headless=new after --dump-dom completes and ignore TERM, + # so an unbounded wait can hang after the complete DOM has been captured. + reap_wait=0 + while kill -0 "$pid" 2>/dev/null && [ "$reap_wait" -lt 20 ]; do + sleep 0.1 + reap_wait=$((reap_wait + 1)) + done + if kill -0 "$pid" 2>/dev/null; then + kill -9 "$pid" 2>/dev/null || true + fi + status=0 + wait "$pid" 2>/dev/null || status=$? + grep -Fq '</html>' "$out_file" 2>/dev/null && return 0 + printf 'attempt %s: exit=%s timed_out=%s bytes=%s stderr=%s\n' \ + "$attempt" "$status" "$timed_out" "$(wc -c <"$out_file" | tr -d ' ')" \ + "$(tail -c 400 "$log" 2>/dev/null | tr '\n' ' ')" >>"$report" + done + printf 'chrome=%s chrome_version=%s pi=%s; %s' \ + "$chrome" "$("$chrome" --version 2>&1 | head -1)" "$pi_version" \ + "$(tr '\n' ' ' <"$report")" + return 1 +} + test_home_resolution() { local fixture out status version if ! command -v node >/dev/null 2>&1 || ! command -v npm >/dev/null 2>&1; then @@ -3241,8 +3317,105 @@ JS pass "Pi Calm working ship moves on a slow independent cadence over faster fixed-cell blue water, paints the complete boat standard yellow with balanced resets, keeps ANSI-stripped width exact, flips the directional sail on the exact bounce at both edges and every width, clamps visible and hidden resizes, falls back deterministically when narrow, freezes and resumes column/direction across settle/start without hidden-time jumps or duplicate timers, resets only on a fresh session, and installs and removes one scheduler-owning widget across starts, settle, abort, failure, shutdown, reload, replacement, and Calm toggles while leaving Calm-off visibility untouched" } +# The rendered-DOM assertions below depend on a real browser, so the render step +# itself is the part that fails for reasons that have nothing to do with Calm. +# This pins that guard with real processes and no browser: one clean render, one +# that only succeeds after Chrome's start-up flake, and one that never renders +# and must report enough to tell a Chrome failure apart from a Pi export change. +test_export_dom_render_guard() { + local dir source_file out_file report + + dir="$TMP_ROOT/render-guard" + mkdir -p "$dir" + source_file="$dir/export.html" + out_file="$dir/dom.html" + printf '<html><body>export</body></html>\n' >"$source_file" + + cat >"$dir/chrome-ok" <<'SH' +#!/bin/sh +case "${1:-}" in --version) echo "FakeChrome 1.2.3"; exit 0 ;; esac +echo attempt >>"$FM_FAKE_CHROME_ATTEMPTS" +printf '<html><head></head><body>export</body></html>\n' +SH + cat >"$dir/chrome-flaky" <<'SH' +#!/bin/sh +case "${1:-}" in --version) echo "FakeChrome 1.2.3"; exit 0 ;; esac +echo attempt >>"$FM_FAKE_CHROME_ATTEMPTS" +if [ "$(wc -l <"$FM_FAKE_CHROME_ATTEMPTS")" -lt 3 ]; then + echo "fake chrome start-up crashed" >&2 + exit 1 +fi +printf '<html><head></head><body>export</body></html>\n' +SH + cat >"$dir/chrome-broken" <<'SH' +#!/bin/sh +case "${1:-}" in --version) echo "FakeChrome 1.2.3"; exit 0 ;; esac +echo attempt >>"$FM_FAKE_CHROME_ATTEMPTS" +echo "FAKE_CHROME_STARTUP_MARKER" >&2 +exit 9 +SH + cat >"$dir/chrome-hang" <<'SH' +#!/bin/sh +case "${1:-}" in --version) echo "FakeChrome 1.2.3"; exit 0 ;; esac +echo attempt >>"$FM_FAKE_CHROME_ATTEMPTS" +printf '<html><head></head><body>export' +exec sleep 30 +SH + chmod +x "$dir/chrome-ok" "$dir/chrome-flaky" "$dir/chrome-broken" "$dir/chrome-hang" + + : >"$dir/attempts-ok" + FM_FAKE_CHROME_ATTEMPTS="$dir/attempts-ok" \ + render_export_dom "$dir/chrome-ok" "$source_file" "$out_file" 9.9.9 >"$dir/report-ok" \ + || fail "render_export_dom rejected a Chrome that dumped a complete DOM" + grep -Fq '</html>' "$out_file" || fail "render_export_dom did not leave the rendered DOM behind" + [ "$(wc -l <"$dir/attempts-ok")" -eq 1 ] \ + || fail "render_export_dom retried a Chrome that had already rendered the DOM" + [ ! -s "$dir/report-ok" ] || fail "render_export_dom reported a diagnostic for a successful render" + + : >"$dir/attempts-flaky" + : >"$out_file" + FM_FAKE_CHROME_ATTEMPTS="$dir/attempts-flaky" \ + render_export_dom "$dir/chrome-flaky" "$source_file" "$out_file" 9.9.9 >"$dir/report-flaky" \ + || fail "render_export_dom gave up on a Chrome that renders after a start-up failure" + grep -Fq '</html>' "$out_file" || fail "a retried render left no DOM behind" + [ "$(wc -l <"$dir/attempts-flaky")" -eq 3 ] \ + || fail "render_export_dom did not retry the failed Chrome start-ups exactly" + + : >"$dir/attempts-broken" + : >"$out_file" + if FM_FAKE_CHROME_ATTEMPTS="$dir/attempts-broken" \ + render_export_dom "$dir/chrome-broken" "$source_file" "$out_file" 9.9.9 >"$dir/report-broken" + then + fail "render_export_dom accepted a Chrome that never rendered the DOM" + fi + [ "$(wc -l <"$dir/attempts-broken")" -eq 3 ] \ + || fail "render_export_dom did not exhaust its bounded retries before failing" + report=$(cat "$dir/report-broken") + assert_contains "$report" "$dir/chrome-broken" "the render failure did not name the Chrome binary it used" + assert_contains "$report" "FakeChrome 1.2.3" "the render failure did not name the Chrome version it used" + assert_contains "$report" "pi=9.9.9" "the render failure did not name the installed Pi version" + assert_contains "$report" "exit=9" "the render failure did not report Chrome's exit status" + assert_contains "$report" "timed_out=no" "the render failure did not report that Chrome exited on its own" + assert_contains "$report" "FAKE_CHROME_STARTUP_MARKER" "the render failure discarded Chrome's own diagnostic" + + : >"$dir/attempts-hang" + : >"$out_file" + if FM_FAKE_CHROME_ATTEMPTS="$dir/attempts-hang" FM_CHROME_RENDER_WAIT_TICKS=3 \ + render_export_dom "$dir/chrome-hang" "$source_file" "$out_file" 9.9.9 >"$dir/report-hang" + then + fail "render_export_dom accepted a Chrome that never finished the DOM" + fi + [ "$(wc -l <"$dir/attempts-hang")" -eq 3 ] \ + || fail "render_export_dom did not exhaust its bounded retries on a Chrome that never finished" + report=$(cat "$dir/report-hang") + assert_contains "$report" "timed_out=yes" \ + "the render failure reported its own kill signal without saying the attempt was timed out" + + pass "the rendered-export-DOM guard renders in one pass, retries a bounded number of Chrome start-up failures, and reports the Chrome binary, Chrome version, Pi version, exit status, and Chrome diagnostic when every attempt fails" +} + test_interactive_terminal_e2e() { - local project config home session_file export_file export_dom default_snapshot expanded_snapshot hidden_snapshot active_before_snapshot active_hidden_snapshot export_snapshot export_settled_snapshot restored_snapshot working_snapshot working_response_snapshot restarted_snapshot resumed_restored_snapshot hash_before hash_after now version chrome chrome_pid chrome_wait chrome_reap_wait active_wait active_screen_wait boat_frame_one boat_frame_two boat_resized_snapshot boat_focus_snapshot boat_cleared_snapshot boat_hull_line boat_sail_line boat_column_one boat_column_two boat_line boat_color_snapshot boat_color_line boat_water_snapshot boat_water_line boat_water_first boat_water_changed boat_narrow_snapshot boat_narrow_sails boat_freeze_snapshot boat_resume_snapshot boat_freeze_column boat_freeze_sail boat_resume_column boat_resume_sail + local project config home session_file export_file export_dom default_snapshot expanded_snapshot hidden_snapshot active_before_snapshot active_hidden_snapshot export_snapshot export_settled_snapshot restored_snapshot working_snapshot working_response_snapshot restarted_snapshot resumed_restored_snapshot hash_before hash_after now version chrome chrome_report active_wait active_screen_wait boat_frame_one boat_frame_two boat_resized_snapshot boat_focus_snapshot boat_cleared_snapshot boat_hull_line boat_sail_line boat_column_one boat_column_two boat_line boat_color_snapshot boat_color_line boat_water_snapshot boat_water_line boat_water_first boat_water_changed boat_narrow_snapshot boat_narrow_sails boat_freeze_snapshot boat_resume_snapshot boat_freeze_column boat_freeze_sail boat_resume_column boat_resume_sail if ! command -v pi >/dev/null 2>&1 || ! command -v tmux >/dev/null 2>&1; then echo "skip: pi or tmux not found for Pi calm interactive E2E" return 0 @@ -3702,36 +3875,10 @@ if (!serialized.includes("firstmate-synthetic-input") || !serialized.includes("/ const synthetic = entries.find((entry) => entry.type === "custom_message" && entry.customType === "firstmate-synthetic-input"); if (!synthetic || synthetic.display) process.exit(1); JS - chrome=$(find_chrome) || fail "Chrome or Chromium is required for rendered export DOM assertions" - "$chrome" \ - --headless=new \ - --disable-gpu \ - --no-sandbox \ - --user-data-dir="$TMP_ROOT/chrome-profile" \ - --virtual-time-budget=2000 \ - --dump-dom \ - "file://$export_file" >"$export_dom" 2>/dev/null & - chrome_pid=$! - chrome_wait=0 - while kill -0 "$chrome_pid" 2>/dev/null && [ "$chrome_wait" -lt 100 ]; do - grep -Fq '</html>' "$export_dom" 2>/dev/null && break - sleep 0.1 - chrome_wait=$((chrome_wait + 1)) - done - kill "$chrome_pid" 2>/dev/null || true - # Chrome can retain --headless=new after --dump-dom completes and ignore TERM, - # so an unbounded wait can hang after the complete DOM has been captured. - chrome_reap_wait=0 - while kill -0 "$chrome_pid" 2>/dev/null && [ "$chrome_reap_wait" -lt 20 ]; do - sleep 0.1 - chrome_reap_wait=$((chrome_reap_wait + 1)) - done - if kill -0 "$chrome_pid" 2>/dev/null; then - kill -9 "$chrome_pid" 2>/dev/null || true - fi - wait "$chrome_pid" 2>/dev/null || true - grep -Fq '</html>' "$export_dom" 2>/dev/null \ - || fail "could not render calm-mode HTML export DOM" + chrome=$(find_chrome) \ + || fail "Chrome or Chromium is required for rendered export DOM assertions; set FM_CHROME_BIN to one" + chrome_report=$(render_export_dom "$chrome" "$export_file" "$export_dom" "$version") \ + || fail "could not render calm-mode HTML export DOM: $chrome_report" node - "$export_dom" <<'JS' || fail "rendered export DOM violated the Calm conversation boundary" const dom = require("node:fs").readFileSync(process.argv[2], "utf8"); const messages = dom.match(/<div id="messages">([\s\S]*?)<\/main>/)?.[1]; @@ -4166,5 +4313,6 @@ test_calm_mid_turn_working_notes test_operational_followup_turn_e2e test_hidden_block_geometry_e2e test_working_ship_geometry_and_lifecycle +test_export_dom_render_guard test_interactive_terminal_e2e printf '\nall fm-calm-pi-extension tests passed\n' diff --git a/tests/fm-captain-hold-lifecycle.test.sh b/tests/fm-captain-hold-lifecycle.test.sh index a32c36e7b70..fc82b04b1eb 100755 --- a/tests/fm-captain-hold-lifecycle.test.sh +++ b/tests/fm-captain-hold-lifecycle.test.sh @@ -72,6 +72,15 @@ run_captain() { # <home> <command args...> FM_CONFIG_OVERRIDE="$home/config" "$ROOT/bin/fm-captain-hold.sh" "$@" } +request_reconciles() { # <home> <source-id> <task-id>... + local home=$1 source_id=$2 id + shift 2 + run_captain "$home" bind "$source_id" >/dev/null || return 1 + for id in "$@"; do printf '%s\n' "$id"; done \ + | run_captain "$home" reconcile-requests --source-id "$source_id" \ + --source "captured board result" >/dev/null +} + # The retired command surface, kept for one release as a shim; in-flight # pre-collapse work still drives the lifecycle through these spellings. run_shim() { # <home> <command args...> @@ -1326,6 +1335,76 @@ EOF pass "a secondmate home publishes each hold occurrence and its answer on the parent channel" } +test_secondmate_reconcile_publishes_before_request_retirement() { + local parent mate channel evidence out show rc request + parent=$(make_home reconcile-parent-channel) + mate=$(make_home reconcile-channel-mate) + printf 'reconcile-channel-mate\n' > "$mate/.fm-secondmate-home" + printf 'schema=fm-secondmate-parent.v1\nroute=local\nparent_home=%s\n' "$parent" \ + > "$mate/.fm-secondmate-parent" + channel="$parent/state/reconcile-channel-mate.status" + evidence="$mate/reconcile-evidence.txt" + + tasks_in "$mate" add reconcile-channel-call "Verify the mate call" --kind ship --repo sample >/dev/null \ + || fail "could not create the reconcile channel call" + run_captain "$mate" hold reconcile-channel-call --reason "verify current release state" >/dev/null \ + || fail "could not hold the reconcile channel call" + request_reconciles "$mate" reconcile-board reconcile-channel-call \ + || fail "could not request the channel reconciliation" + printf 'The release has already landed.\n' > "$evidence" + request="$mate/state/reconcile-requests/reconcile-channel-call.request" + + chmod 0500 "$mate/state/reconcile-requests" + set +e + out=$(run_captain "$mate" reconcile close reconcile-channel-call \ + --evidence-file "$evidence" 2>&1) + rc=$? + set -e + chmod 0700 "$mate/state/reconcile-requests" + [ "$rc" -ne 0 ] || fail "failed reconcile request retirement reported success" + assert_contains "$out" "reconcile-channel-call" \ + "the reconcile retirement failure did not name its task: $out" + show=$(tasks_in "$mate" show reconcile-channel-call --full) + assert_contains "$show" "state: done" "request retirement failure reversed the reconciled close" + assert_contains "$show" "Resolution mode: reconciled" \ + "request retirement failure lost the reconciled resolution mode" + [ -f "$request" ] || fail "the request retired despite its forced retirement failure" + [ "$(grep -c 'resolved \[key=captain-hold-reconcile-channel-call-1\]: captain hold reconcile-channel-call: reconciled' "$channel")" -eq 1 ] \ + || fail "the parent resolution was not published before retirement failed: $(cat "$channel")" + + run_captain "$mate" reconcile close reconcile-channel-call --evidence-file "$evidence" >/dev/null \ + || fail "the closed reconciliation could not finish publication and retirement" + [ ! -e "$request" ] || fail "the retry did not retire the published reconcile request" + [ "$(grep -c 'resolved \[key=captain-hold-reconcile-channel-call-1\]: captain hold reconcile-channel-call: reconciled' "$channel")" -eq 1 ] \ + || fail "the reconciliation retry duplicated or changed its parent resolution: $(cat "$channel")" + tasks_in "$mate" add answer-channel-call "Answer the mate call" --kind ship --repo sample >/dev/null \ + || fail "could not create the normal-answer channel call" + run_captain "$mate" hold answer-channel-call --reason "captain answer needed" >/dev/null \ + || fail "could not hold the normal-answer channel call" + request_reconciles "$mate" reconcile-board answer-channel-call \ + || fail "could not create the normal-answer retry trigger" + printf 'Proceed with the release.\n' > "$mate/answer.txt" + request="$mate/state/reconcile-requests/answer-channel-call.request" + chmod 0500 "$mate/state/reconcile-requests" + set +e + out=$(run_captain "$mate" answer answer-channel-call --decision-file "$mate/answer.txt" 2>&1) + rc=$? + set -e + chmod 0700 "$mate/state/reconcile-requests" + [ "$rc" -ne 0 ] || fail "failed normal-answer request retirement reported success" + show=$(tasks_in "$mate" show answer-channel-call --full) + assert_contains "$show" "state: done" "request retirement failure reversed the captain answer" + [ -f "$request" ] || fail "the normal-answer retry trigger retired after its forced failure" + [ "$(grep -c 'resolved \[key=captain-hold-answer-channel-call-1\]: captain hold answer-channel-call: answered' "$channel")" -eq 1 ] \ + || fail "the normal answer did not publish before retirement failed: $(cat "$channel")" + run_captain "$mate" answer answer-channel-call --decision-file "$mate/answer.txt" >/dev/null \ + || fail "the normal-answer retry could not finish request retirement" + [ ! -e "$request" ] || fail "the normal-answer retry left its request pending" + [ "$(grep -c 'resolved \[key=captain-hold-answer-channel-call-1\]: captain hold answer-channel-call: answered' "$channel")" -eq 1 ] \ + || fail "the normal-answer retry duplicated its parent resolution: $(cat "$channel")" + pass "secondmate resolutions publish before retiring durable retry triggers" +} + # The one keyed-answer intake, fed through the real process-event runner by a # fixture channel that knows nothing about captain holds: task-id keys close at # answer time, a card-declared release mode frees held work, freeform prose can @@ -1348,12 +1427,23 @@ test_bound_channel_answers_close_at_answer_time() { --reason "captain forged choice pending" --repo sample --origin "$id" >/dev/null run_captain "$home" hold sample-invalid-close-call --title "Captain call: invalid close" \ --reason "captain close mode validation pending" --repo sample --origin "$id" >/dev/null + run_captain "$home" hold sample-source-reconcile --title "Captain call: reconcile" \ + --reason "captain re-check pending" --repo sample --origin "$id" >/dev/null + run_captain "$home" hold sample-bare-reconcile --title "Captain call: bare reconcile" \ + --reason "captain bare re-check pending" --repo sample --origin "$id" >/dev/null + run_captain "$home" hold sample-old-shape --title "Captain call: old board shape" \ + --reason "captain old board pending" --repo sample --origin "$id" >/dev/null + run_captain "$home" hold sample-old-reconcile --title "Captain call: old bare reconcile" \ + --reason "captain old bare reconcile pending" --repo sample --origin "$id" >/dev/null + run_captain "$home" hold sample-old-reconcile-note --title "Captain call: old annotated reconcile" \ + --reason "captain old annotated reconcile pending" --repo sample --origin "$id" >/dev/null tasks_in "$home" add sample-gated-work "Gated sample work" --kind ship --repo sample \ --body 'Gated work plan.' >/dev/null run_captain "$home" hold sample-gated-work --reason "captain go needed" >/dev/null run_captain "$home" complete "$id" \ sample-membership-call sample-headline-call sample-forged-call sample-invalid-close-call \ - sample-gated-work >/dev/null \ + sample-source-reconcile sample-bare-reconcile sample-old-shape sample-old-reconcile \ + sample-old-reconcile-note sample-gated-work >/dev/null \ || fail "completion failed for the deck's inventoried calls" artifact="$home/data/$id/review.html" @@ -1374,25 +1464,44 @@ session: status: feedback session_ended: true ended_by: user -prompts[6]{uid,prompt,selector,tag,text}: - "2","Membership: gold-only\n\nContext data:\n{\n \"question\": \"sample-membership-call\",\n \"answer\": \"gold-only\"\n}","section#call > form:nth-of-type(1)",choice,"Membership: gold-only" - "3","Headline: f1-when-fp-gold\n\nContext data:\n{\n \"question\": \"sample-headline-call\",\n \"answer\": \"f1-when-fp-gold\"\n}","section#call > form:nth-of-type(2)",choice,"Headline: f1-when-fp-gold" - "4","Gated work: go\n\nContext data:\n{\n \"question\": \"sample-gated-work\",\n \"answer\": \"go\",\n \"close\": \"release\"\n}","section#call > form:nth-of-type(3)",choice,"Gated work: go" - "5","Absent call: yes\n\nContext data:\n{\n \"question\": \"sample-nonexistent-call\",\n \"answer\": \"yes\"\n}","section#call > form:nth-of-type(4)",choice,"Absent call: yes" +prompts[13]{uid,prompt,selector,tag,text}: + "1","Reconcile first\n\nContext data:\n{\n \"schema\": \"fm-bearings-answer.v1\",\n \"question\": \"sample-source-reconcile\",\n \"selection\": \"reconcile\",\n \"note\": \"\"\n}","section#call > form:nth-of-type(6)",choice,"Reconcile" + "2","Membership: gold-only - captain detail\n\nContext data:\n{\n \"schema\": \"fm-bearings-answer.v1\",\n \"question\": \"sample-membership-call\",\n \"selection\": \"gold-only\",\n \"note\": \"captain detail\"\n}","section#call > form:nth-of-type(1)",choice,"Membership: gold-only - captain detail" + "3","Headline: f1-when-fp-gold\n\nContext data:\n{\n \"schema\": \"fm-bearings-answer.v1\",\n \"question\": \"sample-headline-call\",\n \"selection\": \"f1-when-fp-gold\",\n \"note\": \"\"\n}","section#call > form:nth-of-type(2)",choice,"Headline: f1-when-fp-gold" + "4","Gated work: go\n\nContext data:\n{\n \"schema\": \"fm-bearings-answer.v1\",\n \"question\": \"sample-gated-work\",\n \"selection\": \"go\",\n \"note\": \"\",\n \"close\": \"release\"\n}","section#call > form:nth-of-type(3)",choice,"Gated work: go" + "5","Absent call: yes\n\nContext data:\n{\n \"schema\": \"fm-bearings-answer.v1\",\n \"question\": \"sample-nonexistent-call\",\n \"selection\": \"yes\",\n \"note\": \"\"\n}","section#call > form:nth-of-type(4)",choice,"Absent call: yes" "6","Invalid close: yes\n\nContext data:\n{\n \"question\": \"sample-invalid-close-call\",\n \"answer\": \"yes\",\n \"close\": \"drop\"\n}","section#call > form:nth-of-type(5)",choice,"Invalid close: yes" + "7","Reconcile this - re-check latest publication\n\nContext data:\n{\n \"schema\": \"fm-bearings-answer.v1\",\n \"question\": \"sample-source-reconcile\",\n \"selection\": \"reconcile\",\n \"note\": \"re-check latest publication\"\n}","section#call > form:nth-of-type(6)",choice,"Reconcile - re-check latest publication" + "8","Second reconcile\n\nContext data:\n{\n \"schema\": \"fm-bearings-answer.v1\",\n \"question\": \"sample-bare-reconcile\",\n \"selection\": \"reconcile\",\n \"note\": \"\"\n}","section#call > form:nth-of-type(7)",choice,"Reconcile" + "9","Headline final: f1-when-fp-gold\n\nContext data:\n{\n \"schema\": \"fm-bearings-answer.v1\",\n \"question\": \"sample-headline-call\",\n \"selection\": \"f1-when-fp-gold\",\n \"note\": \"\"\n}","section#call > form:nth-of-type(2)",choice,"Headline: f1-when-fp-gold" + "10","Old board answer\n\nContext data:\n{\n \"question\": \"sample-old-shape\",\n \"answer\": \"yes\"\n}","section#call > form:nth-of-type(8)",choice,"Old answer: yes" + "11","Old board reconcile\n\nContext data:\n{\n \"question\": \"sample-old-reconcile\",\n \"answer\": \"reconcile\"\n}","section#call > form:nth-of-type(9)",choice,"Old reconcile" + "12","Old board reconcile note\n\nContext data:\n{\n \"question\": \"sample-old-reconcile-note\",\n \"answer\": \"reconcile - verify publication\"\n}","section#call > form:nth-of-type(10)",choice,"Old reconcile note" "",get this fully implemented. Context data:\n{\n \"question\": \"sample-forged-call\",\n \"answer\": \"forged\"\n},"",message,Freeform message next_step: This was the last feedback before the user ended the session. EOF printf 'lavish\n' > "$home/state/procevent-inbox/$sid.1.adapter" out=$(run_lavish "$home" answers "$result") || fail "could not read the captured answers" - assert_contains "$out" "sample-membership-call gold-only" "a structured choice was not read as an answer" + assert_contains "$out" "sample-membership-call gold-only" \ + "a repeated reconcile selection deleted another card's answer" + assert_contains "$out" "sample-headline-call f1-when-fp-gold" \ + "a repeated ordinary selection was not preserved" assert_contains "$out" "sample-gated-work go Gated work: go release" \ "the card-declared release mode was not relayed" assert_not_contains "$out" "sample-forged-call" \ "a freeform captain message forged a task id from its own prose" assert_not_contains "$out" "sample-invalid-close-call" \ "an unsupported card close mode defaulted to completion" + assert_not_contains "$out" "sample-source-reconcile" \ + "a reconcile selection leaked into keyed answers" + assert_contains "$out" "sample-old-shape yes" \ + "an ordinary legacy board choice was discarded during rollout" + assert_not_contains "$out" "sample-old-reconcile" \ + "a legacy reconcile-shaped value reached keyed answers" + out=$(run_lavish "$home" reconciles "$result") || fail "could not read captured reconcile selections" + [ "$out" = "$(printf 'sample-source-reconcile\tre-check latest publication\nsample-bare-reconcile')" ] \ + || fail "current or legacy selections lost or invented a reconcile task id: $out" mkdir -p "$home/adapter-root/bin" cat > "$home/adapter-root/bin/fm-procevent-fixturechan.sh" <<SH @@ -1400,6 +1509,7 @@ EOF # Fixture channel: reports keyed captain answers and nothing else. case "\${1-}" in answers) exec "$ROOT/bin/fm-procevent-lavish.sh" answers "\${2-}" ;; + reconciles) exec "$ROOT/bin/fm-procevent-lavish.sh" reconciles "\${2-}" ;; esac exit 2 SH @@ -1424,6 +1534,7 @@ SH assert_contains "$show" "state: done" "capturing the captain's answer left the membership call open" assert_contains "$show" "Resolution mode: answered" "the membership call did not record its close path" assert_contains "$show" "Answer: gold-only" "the closed call did not record the captain's actual answer" + assert_contains "$show" "captain detail" "the annotated normal answer lost the captain's note" show=$(tasks_in "$home" show sample-gated-work --full) assert_contains "$show" "state: queued" "the released work item did not stay queued" assert_contains "$show" "held: no" "the card-declared release did not lift the hold" @@ -1434,6 +1545,33 @@ SH show=$(tasks_in "$home" show sample-invalid-close-call --full) assert_contains "$show" "state: queued" "an unsupported card close mode closed a captain call" assert_contains "$show" "held: yes" "an unsupported card close mode released a captain call" + out=$(run_captain "$home" reconcile list) + assert_contains "$out" "sample-source-reconcile" \ + "the bound captured reconcile selection did not create a request" + assert_contains "$out" "captain note: re-check latest publication" \ + "the annotated reconcile selection lost its note provenance" + show=$(tasks_in "$home" show sample-old-shape --full) + assert_contains "$show" "state: done" "an ordinary legacy board choice did not close its task" + assert_contains "$show" "Resolution mode: answered" \ + "an ordinary legacy board choice did not use the keyed-answer intake" + show=$(tasks_in "$home" show sample-old-reconcile --full) + assert_contains "$show" "state: queued" "a bare legacy reconcile value closed its task" + assert_contains "$show" "held: yes" "a bare legacy reconcile value released its task" + show=$(tasks_in "$home" show sample-old-reconcile-note --full) + assert_contains "$show" "state: queued" "an annotated legacy reconcile value closed its task" + assert_contains "$show" "held: yes" "an annotated legacy reconcile value released its task" + assert_not_contains "$out" "sample-old-reconcile" \ + "a legacy reconcile value created a generationless request" + show=$(tasks_in "$home" show sample-bare-reconcile --full) + assert_contains "$show" "state: queued" "a bare captured reconcile selection closed its task" + assert_contains "$show" "held: yes" "a bare captured reconcile selection released its task" + printf 'The captured call is moot.\n' > "$home/source-reconcile-evidence.txt" + run_captain "$home" reconcile close sample-source-reconcile \ + --evidence-file "$home/source-reconcile-evidence.txt" >/dev/null \ + || fail "the annotated captured request did not authorize evidence-backed closure" + run_captain "$home" reconcile close sample-bare-reconcile \ + --evidence-file "$home/source-reconcile-evidence.txt" >/dev/null \ + || fail "the bare captured request did not authorize evidence-backed closure" # Replaying the same capture is a no-op, not a rejected different decision. A # run that could not close every answered key still reports nonzero. @@ -1456,6 +1594,10 @@ SH printf 'Captain answered the invalid-close call directly.\n' > "$home/invalid-close.txt" run_captain "$home" answer sample-invalid-close-call --decision-file "$home/invalid-close.txt" >/dev/null \ || fail "could not close the invalid-close call through the answer path" + run_captain "$home" answer sample-old-reconcile --decision-file "$home/invalid-close.txt" >/dev/null \ + || fail "could not deliberately close the bare legacy reconcile call" + run_captain "$home" answer sample-old-reconcile-note --decision-file "$home/invalid-close.txt" >/dev/null \ + || fail "could not deliberately close the annotated legacy reconcile call" run_captain "$home" verify "$id" >/dev/null \ || fail "answered calls did not satisfy the completion gate" pass "a bound channel's captured answers close their captain-held tasks at answer time" @@ -1463,6 +1605,331 @@ SH # Answer-time closure is opt-in per source. A channel with no binding must behave # exactly as it always did: capture, announce, close nothing. +# A reconcile is "go re-check reality", never the captain's answer. The value is +# reserved at the one keyed-answer intake, so no channel and no card-declared +# close mode can turn it into a close or a release, and the obligation to verify +# survives as a durable request instead of evaporating with the wake. +test_reconcile_never_closes_through_the_keyed_answer_intake() { + local home out rc show list + home=$(make_home reconcile-intake) + tasks_in "$home" add sample-reconcile-call "Captain call: still current?" --repo sample >/dev/null \ + || fail "could not create the reconcile call" + tasks_in "$home" add sample-reconcile-gated "Gated work" --repo sample >/dev/null \ + || fail "could not create the gated work item" + run_captain "$home" hold sample-reconcile-call --reason "is this still current?" >/dev/null \ + || fail "could not hold the reconcile call" + run_captain "$home" hold sample-reconcile-gated --reason "waiting on the captain" >/dev/null \ + || fail "could not hold the gated work item" + + set +e + out=$(printf 'sample-reconcile-call\treconcile\tReconcile\n%s\n' \ + "$(printf 'sample-reconcile-gated\treconcile\tReconcile\trelease')" \ + | run_captain "$home" answers --source "captain chat" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "the shared answer intake accepted reconcile as an answer" + assert_contains "$out" "refused: sample-reconcile-call" \ + "the shared intake did not visibly refuse reconcile: $out" + case "$out" in + *"closed: sample-reconcile"*) fail "a reconcile row closed a captain call: $out" ;; + esac + + show=$(tasks_in "$home" show sample-reconcile-call --full) + assert_contains "$show" "state: queued" "a reconcile row completed a captain call" + assert_contains "$show" "held: yes" "a reconcile row released a captain call" + case "$show" in + *"Resolution recorded by"*) fail "a reconcile row wrote a resolution record" ;; + esac + show=$(tasks_in "$home" show sample-reconcile-gated --full) + assert_contains "$show" "held: yes" "a release-mode reconcile row lifted a captain hold" + + list=$(run_captain "$home" reconcile list) || fail "could not list the reconcile requests" + assert_contains "$list" "reconcile-requests: 0" \ + "the shared answer intake created a reconcile request: $list" + set +e + out=$(printf 'sample-reconcile-call\n' \ + | run_captain "$home" reconcile-requests --source-id unbound-src --source "captured board" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "an unbound captured source created a reconcile request" + request_reconciles "$home" board-src sample-reconcile-call sample-reconcile-gated \ + || fail "the bound captured source did not create reconcile requests" + list=$(run_captain "$home" reconcile list) + assert_contains "$list" "sample-reconcile-call" "the reconcile obligation was not recorded durably: $list" + assert_contains "$list" "reconcile-requests: 2" "the reconcile requests were not both recorded: $list" + + request_reconciles "$home" board-src sample-reconcile-call \ + || fail "replaying a captured reconcile selection failed" + list=$(run_captain "$home" reconcile list) + assert_contains "$list" "reconcile-requests: 2" "a replayed reconcile selection duplicated the obligation: $list" + pass "only a bound captured source creates reconcile requests" +} + +test_normal_answers_retire_pending_reconcile_requests() { + local home list id + home=$(make_home reconcile-normal-answer) + for id in sample-direct-close sample-direct-release sample-keyed-close; do + tasks_in "$home" add "$id" "Captain call $id" --repo sample >/dev/null + run_captain "$home" hold "$id" --reason "waiting for the captain" >/dev/null + done + request_reconciles "$home" board-src sample-direct-close sample-direct-release sample-keyed-close \ + || fail "could not create reconcile requests before normal answers" + + printf 'Captain said close.\n' > "$home/close.txt" + printf 'Captain said release.\n' > "$home/release.txt" + run_captain "$home" answer sample-direct-close --decision-file "$home/close.txt" >/dev/null \ + || fail "a direct close answer failed" + run_captain "$home" answer sample-direct-release --decision-file "$home/release.txt" --release >/dev/null \ + || fail "a direct release answer failed" + printf 'sample-keyed-close\tyes\tYes\n' \ + | run_captain "$home" answers --source "board sequence 2" >/dev/null \ + || fail "a keyed normal answer failed" + list=$(run_captain "$home" reconcile list) + assert_contains "$list" "reconcile-requests: 0" \ + "normal answers left stranded reconcile requests: $list" + + run_captain "$home" answer sample-direct-close --decision-file "$home/close.txt" >/dev/null \ + || fail "a direct close replay failed" + run_captain "$home" answer sample-direct-release --decision-file "$home/release.txt" --release >/dev/null \ + || fail "a direct release replay failed" + printf 'sample-keyed-close\tyes\tYes\n' \ + | run_captain "$home" answers --source "board sequence 2" >/dev/null \ + || fail "a keyed answer replay failed" + list=$(run_captain "$home" reconcile list) + assert_contains "$list" "reconcile-requests: 0" \ + "an idempotent normal-answer replay restored a reconcile request: $list" + pass "normal answers and their replays retire reconcile requests" +} + +# The two verification outcomes, and the honesty of the record each writes. +test_reconcile_closes_with_evidence_or_keeps_the_call_open() { + local home show list rc out + home=$(make_home reconcile-outcomes) + tasks_in "$home" add sample-moot-call "Captain call: ship 0.1.37?" --repo sample >/dev/null + tasks_in "$home" add sample-active-call "Captain call: which admission order?" --repo sample >/dev/null + tasks_in "$home" add sample-mode-call "Captain call: verify replay mode?" --repo sample >/dev/null + run_captain "$home" hold sample-moot-call --reason "ship 0.1.37?" >/dev/null + run_captain "$home" hold sample-active-call --reason "which admission order?" >/dev/null + run_captain "$home" hold sample-mode-call --reason "verify replay mode?" >/dev/null + printf '0.1.38 was published on 2026-09-05, so the 0.1.37 question is moot.\n' > "$home/evidence.txt" + printf 'Still open: nothing has shipped and the choice is unchanged.\n' > "$home/note.txt" + + set +e + out=$(run_captain "$home" reconcile close sample-moot-call --evidence-file "$home/evidence.txt" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "reconcile closed a call without a pending board request" + assert_contains "$out" "no pending board-created reconcile request" \ + "the ungated close refusal did not name the missing board request: $out" + set +e + out=$(run_captain "$home" reconcile note sample-active-call --note-file "$home/note.txt" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "reconcile annotated a call without a pending board request" + assert_contains "$out" "no pending board-created reconcile request" \ + "the ungated note refusal did not name the missing board request: $out" + + request_reconciles "$home" board-src sample-moot-call sample-active-call sample-mode-call \ + || fail "could not file the reconcile requests" + + set +e + out=$(run_captain "$home" reconcile close sample-moot-call 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "a reconcile close was accepted with no evidence" + assert_contains "$out" "evidence" "the refusal did not name the missing evidence: $out" + + cp "$home/state/reconcile-requests/sample-mode-call.request" "$home/mode-request.backup" + run_captain "$home" answer sample-mode-call --decision-file "$home/evidence.txt" >/dev/null \ + || fail "could not record the normal answer for the mode fixture" + cp "$home/mode-request.backup" "$home/state/reconcile-requests/sample-mode-call.request" + set +e + out=$(run_captain "$home" reconcile close sample-mode-call --evidence-file "$home/evidence.txt" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "a normal captain answer replayed as a reconciliation" + assert_contains "$out" "was not closed by reconciliation" \ + "the reconcile replay refusal did not identify the incompatible resolution mode: $out" + run_captain "$home" answer sample-mode-call --decision-file "$home/evidence.txt" >/dev/null \ + || fail "the normal answer replay did not retire its restored pending request" + + run_captain "$home" reconcile close sample-moot-call --evidence-file "$home/evidence.txt" >/dev/null \ + || fail "could not close the moot call with evidence" + show=$(tasks_in "$home" show sample-moot-call --full) + assert_contains "$show" "state: done" "the moot call did not close" + assert_contains "$show" "Resolution mode: reconciled" "the moot call did not record how it closed" + assert_contains "$show" "Reconciliation evidence:" "the moot call did not record the evidence" + assert_contains "$show" "0.1.38 was published" "the recorded evidence was lost" + case "$show" in + *"Captain decision:"*) fail "a reconciled close was recorded as the captain's own words" ;; + esac + set +e + out=$(run_captain "$home" answer sample-moot-call --decision-file "$home/evidence.txt" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "a reconciled resolution replayed as a captain answer" + assert_contains "$out" "not a captain-answer replay" \ + "the answer replay refusal did not identify the incompatible resolution mode: $out" + + run_captain "$home" reconcile note sample-active-call --note-file "$home/note.txt" >/dev/null \ + || fail "could not annotate the still-active call" + show=$(tasks_in "$home" show sample-active-call --full) + assert_contains "$show" "state: queued" "annotating a still-active call closed it" + assert_contains "$show" "held: yes" "annotating a still-active call released it" + assert_contains "$show" "Captain hold reconciled:" "the re-check left no dated note" + assert_contains "$show" "Still open: nothing has shipped" "the note body was lost" + case "$show" in + *"Resolution recorded by"*) fail "annotating a still-active call wrote a resolution record" ;; + esac + set +e + out=$(run_captain "$home" reconcile note sample-active-call --note-file "$home/note.txt" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "a retired reconcile request appended a duplicate note" + assert_contains "$out" "no pending board-created reconcile request" \ + "the duplicate-note refusal did not name the retired request: $out" + + printf 'sample-active-call\n' \ + | FM_CAPTAIN_HOLD_NOW=2026-09-07T06:00:00Z run_captain "$home" reconcile-requests \ + --source-id board-src --source "captured board result sequence 2" >/dev/null \ + || fail "could not create the second reconcile request" + FM_CAPTAIN_HOLD_NOW=2026-09-07T06:01:00Z run_captain "$home" reconcile note sample-active-call \ + --note-file "$home/note.txt" >/dev/null \ + || fail "the second request with the same finding was not recorded" + show=$(tasks_in "$home" show sample-active-call --full) + [ "$(printf '%s\n' "$show" | grep -o 'Captain hold reconciled:' | wc -l | tr -d ' ')" -eq 2 ] \ + || fail "a later reconcile request with the same note did not append its own record" + assert_contains "$show" "Captain hold reconciled: 2026-09-07T06:01:00Z" \ + "the second reconcile request lost its own dated note" + + list=$(run_captain "$home" reconcile list) + assert_contains "$list" "reconcile-requests: 0" "the verified requests were not retired: $list" + pass "reconcile closes a moot call with evidence and keeps an active one open with a note" +} + +test_reconcile_outcomes_retry_partial_failures_once() { + local home out show list rc + home=$(make_home reconcile-partial-retry) + tasks_in "$home" add sample-reconcile-close-retry "Close retry" --repo sample >/dev/null + tasks_in "$home" add sample-reconcile-note-retry "Note retry" --repo sample >/dev/null + tasks_in "$home" add sample-reconcile-retire-retry "Retire retry" --repo sample >/dev/null + tasks_in "$home" add sample-reconcile-close-retire "Close retire" --repo sample >/dev/null + tasks_in "$home" add sample-answer-retire "Answer retire" --repo sample >/dev/null + run_captain "$home" hold sample-reconcile-close-retry --reason "verify close" >/dev/null + run_captain "$home" hold sample-reconcile-note-retry --reason "verify note" >/dev/null + run_captain "$home" hold sample-reconcile-retire-retry --reason "verify retire" >/dev/null + run_captain "$home" hold sample-reconcile-close-retire --reason "verify close retirement" >/dev/null + run_captain "$home" hold sample-answer-retire --reason "verify answer retirement" >/dev/null + request_reconciles "$home" board-src sample-reconcile-close-retry sample-reconcile-note-retry \ + sample-reconcile-retire-retry sample-reconcile-close-retire sample-answer-retire \ + || fail "could not create partial-retry requests" + printf 'Verified moot.\n' > "$home/retry-evidence.txt" + printf 'Verified active.\n' > "$home/retry-note.txt" + printf 'Verified retirement retry.\n' > "$home/retry-retire-note.txt" + printf 'Verified close retirement.\n' > "$home/close-retire-evidence.txt" + printf 'Captain answered despite retirement failure.\n' > "$home/answer-retire.txt" + cat > "$home/fakebin/tasks-axi" <<'SH' +#!/usr/bin/env bash +if [ "${1:-}" = done ] && [ "${2:-}" = sample-reconcile-close-retry ] \ + && [ ! -e "$FM_HOME/reconcile-close-failed" ]; then + : > "$FM_HOME/reconcile-close-failed" + exit 92 +fi +if [ "${1:-}" = update ] && [ "${2:-}" = sample-reconcile-note-retry ]; then + "$REAL_TASKS_AXI" "$@" || exit $? + : > "$FM_HOME/reconcile-note-updated" + exit 0 +fi +if [ "${1:-}" = show ] && [ "${2:-}" = sample-reconcile-note-retry ] \ + && [ -e "$FM_HOME/reconcile-note-updated" ] && [ ! -e "$FM_HOME/reconcile-note-show-failed" ]; then + : > "$FM_HOME/reconcile-note-show-failed" + exit 93 +fi +exec "$REAL_TASKS_AXI" "$@" +SH + chmod +x "$home/fakebin/tasks-axi" + + set +e + out=$(run_captain "$home" reconcile close sample-reconcile-close-retry \ + --evidence-file "$home/retry-evidence.txt" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "the forced reconcile close failure reported success" + run_captain "$home" reconcile close sample-reconcile-close-retry \ + --evidence-file "$home/retry-evidence.txt" >/dev/null \ + || fail "reconcile close did not recover from its partial failure" + show=$(tasks_in "$home" show sample-reconcile-close-retry --full) + [ "$(printf '%s\n' "$show" | grep -c 'Resolution mode: reconciled')" -eq 1 ] \ + || fail "reconcile close duplicated its resolution record on retry" + + set +e + out=$(run_captain "$home" reconcile note sample-reconcile-note-retry \ + --note-file "$home/retry-note.txt" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "the forced post-note probe failure reported success" + set +e + out=$(run_captain "$home" reconcile note sample-reconcile-note-retry \ + --note-file "$home/retry-note.txt" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "a note retry succeeded after the request was retired" + show=$(tasks_in "$home" show sample-reconcile-note-retry --full) + [ "$(printf '%s\n' "$show" | grep -c 'Captain hold reconciled:')" -eq 1 ] \ + || fail "reconcile note duplicated its durable annotation" + + chmod 0500 "$home/state/reconcile-requests" + set +e + out=$(run_captain "$home" reconcile close sample-reconcile-close-retire \ + --evidence-file "$home/close-retire-evidence.txt" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "a failed close request retirement reported success" + assert_contains "$out" "sample-reconcile-close-retire" \ + "the close retirement failure did not name its task: $out" + list=$(run_captain "$home" reconcile list) + assert_contains "$list" "sample-reconcile-close-retire" \ + "the failed close retirement hid its pending request" + chmod 0700 "$home/state/reconcile-requests" + run_captain "$home" reconcile close sample-reconcile-close-retire \ + --evidence-file "$home/close-retire-evidence.txt" >/dev/null \ + || fail "the closed reconciliation could not finish request retirement" + + chmod 0500 "$home/state/reconcile-requests" + set +e + out=$(run_captain "$home" answer sample-answer-retire \ + --decision-file "$home/answer-retire.txt" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "a failed answer-boundary retirement reported success" + show=$(tasks_in "$home" show sample-answer-retire --full) + assert_contains "$show" "state: done" "retirement failure reversed the durable captain answer" + assert_contains "$show" "Captain answered despite retirement failure" \ + "retirement failure lost the durable captain answer" + chmod 0700 "$home/state/reconcile-requests" + run_captain "$home" answer sample-answer-retire --decision-file "$home/answer-retire.txt" >/dev/null \ + || fail "the answer replay could not finish request retirement" + + chmod 0500 "$home/state/reconcile-requests" + set +e + out=$(run_captain "$home" reconcile note sample-reconcile-retire-retry \ + --note-file "$home/retry-retire-note.txt" 2>&1) + rc=$? + set -e + chmod 0700 "$home/state/reconcile-requests" + [ "$rc" -ne 0 ] || fail "a failed request retirement reported note success" + assert_not_contains "$out" "still-open:" "failed retirement reported a successful outcome" + run_captain "$home" reconcile note sample-reconcile-retire-retry \ + --note-file "$home/retry-retire-note.txt" >/dev/null \ + || fail "the applied note could not finish request retirement on retry" + show=$(tasks_in "$home" show sample-reconcile-retire-retry --full) + [ "$(printf '%s\n' "$show" | grep -c 'Captain hold reconciled:')" -eq 1 ] \ + || fail "failed request retirement duplicated the reconcile note" + list=$(run_captain "$home" reconcile list) + assert_contains "$list" "reconcile-requests: 0" "partial retries left a reconcile request pending" + pass "reconcile outcomes apply durable mutations once across partial failures" +} + test_unbound_source_closes_no_hold() { local home id sid artifact result out show rc home=$(make_home lavish-unbound) @@ -1634,7 +2101,7 @@ test_legacy_identities_keep_working() { # The intake is channel-agnostic, so chat must reach it the same way a captured # review does - for a task-id key, and for a legacy composed identity. test_chat_channel_feeds_the_same_keyed_answer_intake() { - local home id fb show + local home id fb show list home=$(make_home chat-channel) id=sample-chat-review mkdir -p "$home/data/$id" @@ -1649,7 +2116,11 @@ test_chat_channel_feeds_the_same_keyed_answer_intake() { run_captain "$home" hold sample-chat-followup --title "Choose the chat follow-up" \ --reason "captain follow-up choice pending" --repo sample >/dev/null \ || fail "could not register the task-id chat call" - run_captain "$home" complete "$id" "$id-decision-chat-choice" sample-chat-followup >/dev/null \ + run_captain "$home" hold sample-chat-reconcile --title "Reconcile from chat" \ + --reason "captain chat reconcile pending" --repo sample >/dev/null \ + || fail "could not register the chat reconcile call" + run_captain "$home" complete "$id" "$id-decision-chat-choice" sample-chat-followup \ + sample-chat-reconcile >/dev/null \ || fail "completion failed for the chat calls" grep -F 'captain-held [key=chat-choice]' "$home/state/$id.status" >/dev/null \ || fail "precondition: completion did not transfer the decision to its durable owner" @@ -1709,6 +2180,25 @@ SH assert_contains "$show" "Answer: take the second option" "the chat-answered call lost the captain answer" assert_contains "$show" "answer sent to $id" "the chat-answered call lost its channel provenance" + : > "$home/send.log" + set +e + env PATH="$fb:$PATH" FM_ROOT_OVERRIDE="$home" FM_HOME="$home" \ + FM_STATE_OVERRIDE="$home/state" FM_DATA_OVERRIDE="$home/data" \ + FM_SEND_LOG="$home/send.log" FM_SEND_SETTLE=0 \ + "$ROOT/bin/fm-send.sh" "$id" --resolve-key sample-chat-reconcile reconcile >/dev/null 2>&1 + set -e + show=$(tasks_in "$home" show sample-chat-reconcile --full) + assert_contains "$show" "state: queued" "a chat reconcile answer closed the call" + list=$(run_captain "$home" reconcile list) + assert_contains "$list" "reconcile-requests: 0" "a chat reconcile answer created a board request" + printf 'Chat cannot authorize this closure.\n' > "$home/chat-reconcile.txt" + if run_captain "$home" reconcile close sample-chat-reconcile \ + --evidence-file "$home/chat-reconcile.txt" >/dev/null 2>&1; then + fail "a chat reconcile answer authorized evidence-backed closure" + fi + run_captain "$home" answer sample-chat-reconcile --decision-file "$home/chat-reconcile.txt" >/dev/null \ + || fail "could not close the chat reconcile fixture normally" + if env PATH="$fb:$PATH" FM_ROOT_OVERRIDE="$home" FM_HOME="$home" \ FM_STATE_OVERRIDE="$home/state" FM_DATA_OVERRIDE="$home/data" \ FM_SEND_LOG="$home/send.log" FM_SEND_SETTLE=0 \ @@ -2143,7 +2633,12 @@ test_none_inventory_and_resolved_prose_do_not_create_holds test_terminal_single_owner_status_decision_does_not_block_empty_inventory test_secondmate_hold_stays_in_authoritative_home test_secondmate_home_publishes_holds_and_answers +test_secondmate_reconcile_publishes_before_request_retirement test_bound_channel_answers_close_at_answer_time +test_reconcile_never_closes_through_the_keyed_answer_intake +test_normal_answers_retire_pending_reconcile_requests +test_reconcile_closes_with_evidence_or_keeps_the_call_open +test_reconcile_outcomes_retry_partial_failures_once test_unbound_source_closes_no_hold test_legacy_identities_keep_working test_chat_channel_feeds_the_same_keyed_answer_intake diff --git a/tests/fm-claude-stop-autoarm-live-e2e.test.sh b/tests/fm-claude-stop-autoarm-live-e2e.test.sh index db487bf7ac2..b0316ff66e1 100755 --- a/tests/fm-claude-stop-autoarm-live-e2e.test.sh +++ b/tests/fm-claude-stop-autoarm-live-e2e.test.sh @@ -12,10 +12,10 @@ # shellcheck disable=SC2016 # the model, not this test shell, reads the prompt text set -u -if [ "${FM_CLAUDE_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_CLAUDE_LIVE_E2E=1 to run the Claude Stop auto-arm regression" - exit 0 -fi +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate opt-in FM_CLAUDE_LIVE_E2E claude ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" @@ -24,8 +24,6 @@ fail() { exit 1 } -command -v claude >/dev/null 2>&1 || fail "claude not found" - LAB="$ROOT/.claude-autoarm-live-e2e.$$" PROJECT="$LAB/project" HOME_DIR="$LAB/fmhome" diff --git a/tests/fm-claude-stop-autoarm.test.sh b/tests/fm-claude-stop-autoarm.test.sh index c1b3e6d2b4d..066ecd070b0 100755 --- a/tests/fm-claude-stop-autoarm.test.sh +++ b/tests/fm-claude-stop-autoarm.test.sh @@ -174,6 +174,15 @@ echo "$$" >> "$FM_HOME/state/arm-ran" printf 'watcher: started pid=%s (beacon fresh)\n' "$$" printf 'stale: fixture-win actionable\n' exit 0 +SH + ;; + records-grace) + cat > "$dir/bin/fm-watch-arm.sh" <<'SH' +#!/usr/bin/env bash +echo "$$" >> "$FM_HOME/state/arm-ran" +printf '%s\n' "${FM_GUARD_GRACE:-unset}" > "$FM_HOME/state/arm-received-grace" +printf 'watcher: attached pid=%s (beacon 2s)\n' "$$" +exit 0 SH ;; *) @@ -623,6 +632,20 @@ test_arms_for_x_mode_poll_need_without_inflight() { pass "auto-arm: X-mode poll need arms the cycle even with no tasks in flight" } +test_arms_for_registered_custom_check_without_inflight() { + local dir out status + dir=$(make_primary_dir "$TMP_ROOT/check-need") + printf '#!/usr/bin/env bash\nexit 0\n' > "$dir/state/issue-comments.check.sh" + chmod 700 "$dir/state/issue-comments.check.sh" + FM_STATE_OVERRIDE="$dir/state" "$ROOT/bin/fm-check-register.sh" issue-comments >/dev/null \ + || fail "fm-check-register.sh could not register the custom check" + write_arm_fixture "$dir" actionable + out=$(run_autoarm "$dir" 2>/dev/null); status=$? + expect_code 2 "$status" "a registered custom check must keep the auto-arm active with zero tasks in flight" + [ -e "$dir/state/arm-ran" ] || fail "hook did not arm for the registered custom check" + pass "auto-arm: a registered custom check arms the cycle even with no tasks in flight" +} + test_single_flight_admits_exactly_one_owner() { local dir rc1 rc2 count dir=$(make_primary_dir "$TMP_ROOT/single-flight") @@ -1142,6 +1165,18 @@ test_active_in_marked_secondmate_home() { pass "auto-arm: active in a marked secondmate home" } +test_long_poll_grace_reaches_arm_wrapper() { + local dir out status + dir=$(make_primary_dir "$TMP_ROOT/long-poll-grace") + : > "$dir/state/task.meta" + write_arm_fixture "$dir" records-grace + out=$(unset FM_GUARD_GRACE; FM_POLL=900 run_autoarm "$dir" 2>/dev/null); status=$? + expect_code 2 "$status" "an unverified close without a healthy watcher must still fail closed" + [ -e "$dir/state/arm-received-grace" ] || fail "arm wrapper never recorded FM_GUARD_GRACE" + [ "$(cat "$dir/state/arm-received-grace")" = 960 ] || fail "arm wrapper must see the poll-derived grace (900+60), got: $(cat "$dir/state/arm-received-grace")" + pass "auto-arm: a long FM_POLL with FM_GUARD_GRACE unset reaches fm-watch-arm.sh with the derived grace" +} + test_fm_lock_status_still_works_with_shared_lib() { local out out=$(FM_HOME="$TMP_ROOT/lock-status-home" bash "$ROOT/bin/fm-lock.sh" status 2>&1) @@ -1168,6 +1203,7 @@ test_benign_cycle_end_with_live_watcher_is_silent test_positive_recovery_budget_contention_preserves_episode test_owner_mutex_contention_preserves_failure_episode_reset test_arms_for_x_mode_poll_need_without_inflight +test_arms_for_registered_custom_check_without_inflight test_single_flight_admits_exactly_one_owner test_abandoned_owner_claim_is_reclaimed_and_rearms test_arming_claim_with_fresh_beacon_is_never_reclaimed @@ -1187,5 +1223,6 @@ test_superseded_owner_goes_silent_and_never_double_translates test_need_vanished_mid_cycle_closes_quietly test_afk_mid_cycle_suppresses_rewake test_active_in_marked_secondmate_home +test_long_poll_grace_reaches_arm_wrapper test_fm_lock_status_still_works_with_shared_lib printf '\nall fm-claude-stop-autoarm tests passed\n' diff --git a/tests/fm-cmux-claude-composer-live-e2e.test.sh b/tests/fm-cmux-claude-composer-live-e2e.test.sh index efbb24f1150..f7eb933011f 100755 --- a/tests/fm-cmux-claude-composer-live-e2e.test.sh +++ b/tests/fm-cmux-claude-composer-live-e2e.test.sh @@ -6,6 +6,9 @@ # when opting in. Never copies the developer's credential or trust store. set -u +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" TASK="fm-test-cmux-claude-composer-$$" LAB= @@ -28,16 +31,8 @@ cleanup() { [ -z "$LAB" ] || rm -rf -- "$LAB" } -if [ "${FM_CMUX_CLAUDE_COMPOSER_LIVE:-0}" != 1 ]; then - echo "skip: set FM_CMUX_CLAUDE_COMPOSER_LIVE=1 to run the real cmux Claude composer drift guard" - exit 0 -fi +fm_live_gate opt-in FM_CMUX_CLAUDE_COMPOSER_LIVE claude cmux jq treehouse python3 -command -v claude >/dev/null 2>&1 || fail "FM_CMUX_CLAUDE_COMPOSER_LIVE=1 but Claude Code is not installed" -command -v cmux >/dev/null 2>&1 || fail "FM_CMUX_CLAUDE_COMPOSER_LIVE=1 but cmux is not installed" -command -v jq >/dev/null 2>&1 || fail "FM_CMUX_CLAUDE_COMPOSER_LIVE=1 but jq is not installed" -command -v treehouse >/dev/null 2>&1 || fail "FM_CMUX_CLAUDE_COMPOSER_LIVE=1 but treehouse is not installed" -command -v python3 >/dev/null 2>&1 || fail "FM_CMUX_CLAUDE_COMPOSER_LIVE=1 but python3 is not installed" cmux ping >/dev/null 2>&1 || fail "FM_CMUX_CLAUDE_COMPOSER_LIVE=1 but the cmux socket is unavailable" LAB=$(mktemp -d "${TMPDIR:-/tmp}/fm-cmux-claude-composer.XXXXXX") || fail "could not create an isolated cmux Claude lab" diff --git a/tests/fm-codex-continuity-live-e2e.test.sh b/tests/fm-codex-continuity-live-e2e.test.sh index 16f6f4f957c..883e82d92f7 100755 --- a/tests/fm-codex-continuity-live-e2e.test.sh +++ b/tests/fm-codex-continuity-live-e2e.test.sh @@ -3,10 +3,10 @@ # Codex's bounded foreground-checkpoint supervision path. set -u -if [ "${FM_CODEX_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_CODEX_LIVE_E2E=1 to run the Codex continuity regression" - exit 0 -fi +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate opt-in FM_CODEX_LIVE_E2E codex ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" @@ -15,8 +15,6 @@ fail() { exit 1 } -command -v codex >/dev/null 2>&1 || fail "codex not found" - LAB="$ROOT/.codex-live-e2e.$$" PROJECT="$LAB/project" HOME_DIR="$LAB/fmhome" diff --git a/tests/fm-composer-matrix-live-e2e.test.sh b/tests/fm-composer-matrix-live-e2e.test.sh index 9bb78ade445..feb94b3be3b 100755 --- a/tests/fm-composer-matrix-live-e2e.test.sh +++ b/tests/fm-composer-matrix-live-e2e.test.sh @@ -27,14 +27,12 @@ # unreadable-composer state and correctly fails that harness's check. set -u -ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" -if [ "${FM_COMPOSER_MATRIX_LIVE:-0}" != 1 ]; then - echo "skip: set FM_COMPOSER_MATRIX_LIVE=1 to run the live composer-matrix guard" - exit 0 -fi +ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" -command -v tmux >/dev/null 2>&1 || { echo "not ok - FM_COMPOSER_MATRIX_LIVE=1 but tmux is not installed" >&2; exit 1; } +fm_live_gate opt-in FM_COMPOSER_MATRIX_LIVE tmux SOCKET="fm-cmx-live-$$" SESSION="cmxlive" diff --git a/tests/fm-control-relaunch.test.sh b/tests/fm-control-relaunch.test.sh index 8bb0796ec52..5cba46a0f7a 100755 --- a/tests/fm-control-relaunch.test.sh +++ b/tests/fm-control-relaunch.test.sh @@ -471,7 +471,7 @@ test_relaunch_serializes_concurrent_durable_metadata_publication() { FM_FAKE_TRACE_RELEASE="$launch_release" \ run_control "$dir" rl28 relaunch --note "continue after publication" > "$dir/control.out" & control_pid=$! - while [ ! -e "$prepare" ] && [ "$i" -lt 200 ]; do + while [ ! -e "$prepare" ] && [ "$i" -lt 500 ]; do /bin/sleep 0.01 i=$((i + 1)) done @@ -490,7 +490,7 @@ test_relaunch_serializes_concurrent_durable_metadata_publication() { --carry-platform x --carry-max 280 > "$dir/link.out" 2>&1 & link_pid=$! i=0 - while [ ! -e "$waiting" ] && [ "$i" -lt 200 ]; do + while [ ! -e "$waiting" ] && [ "$i" -lt 500 ]; do /bin/sleep 0.01 i=$((i + 1)) done @@ -503,7 +503,7 @@ test_relaunch_serializes_concurrent_durable_metadata_publication() { } : > "$launch_release" i=0 - while [ ! -e "$ready" ] && [ "$i" -lt 200 ]; do + while [ ! -e "$ready" ] && [ "$i" -lt 500 ]; do /bin/sleep 0.01 i=$((i + 1)) done diff --git a/tests/fm-cursor-primary-live-e2e.test.sh b/tests/fm-cursor-primary-live-e2e.test.sh index ad806069986..159dee1674d 100755 --- a/tests/fm-cursor-primary-live-e2e.test.sh +++ b/tests/fm-cursor-primary-live-e2e.test.sh @@ -24,19 +24,15 @@ # path and is the only state left outside the temp dir. set -u -if [ "${FM_CURSOR_PRIMARY_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_CURSOR_PRIMARY_LIVE_E2E=1 to run the live Cursor primary guard" - exit 0 -fi - # shellcheck source=tests/lib.sh . "$(dirname "${BASH_SOURCE[0]}")/lib.sh" +fm_live_gate opt-in FM_CURSOR_PRIMARY_LIVE_E2E tmux jq + +REAL_TMUX=$(command -v tmux) CURSOR_BIN=${FM_CURSOR_BIN:-$(command -v cursor-agent || true)} [ -n "$CURSOR_BIN" ] && [ -x "$CURSOR_BIN" ] \ || fail "cursor-agent not found; install it or set FM_CURSOR_BIN. This guard refuses to pass without checking the real harness." -REAL_TMUX=$(command -v tmux) || fail "tmux not found" -command -v jq >/dev/null 2>&1 || fail "jq not found" CURSOR_VERSION=$("$CURSOR_BIN" --version 2>/dev/null | head -1) [ -n "$CURSOR_VERSION" ] || fail "cursor-agent did not report a version; refusing to claim a verified result" printf 'harness: cursor-agent %s\n' "$CURSOR_VERSION" diff --git a/tests/fm-grok-continuity-live-e2e.test.sh b/tests/fm-grok-continuity-live-e2e.test.sh index 1c2731e3ad3..52d61e0e4c5 100755 --- a/tests/fm-grok-continuity-live-e2e.test.sh +++ b/tests/fm-grok-continuity-live-e2e.test.sh @@ -3,10 +3,10 @@ # through Grok's tracked background-task notification path. set -u -if [ "${FM_GROK_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_GROK_LIVE_E2E=1 to run the interactive Grok continuity regression" - exit 0 -fi +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate opt-in FM_GROK_LIVE_E2E grok tmux ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" @@ -15,9 +15,6 @@ fail() { exit 1 } -command -v grok >/dev/null 2>&1 || fail "grok not found" -command -v tmux >/dev/null 2>&1 || fail "tmux not found" - TMUX=$(command -v tmux) SOCKET="fm-grok-live-e2e-$$" SESSION=grok-live-e2e diff --git a/tests/fm-grok-stop-live-e2e.test.sh b/tests/fm-grok-stop-live-e2e.test.sh index 3dc0ba22bac..d701dc9c317 100755 --- a/tests/fm-grok-stop-live-e2e.test.sh +++ b/tests/fm-grok-stop-live-e2e.test.sh @@ -7,14 +7,11 @@ # window. Cleanup uses only creation-time pane/process identities. set -u -if [ "${FM_GROK_STOP_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_GROK_STOP_LIVE_E2E=1 with FM_GROK_NATIVE_BIN and FM_GROK_LEGACY_BIN" - exit 0 -fi - # shellcheck source=tests/lib.sh . "$(dirname "${BASH_SOURCE[0]}")/lib.sh" +fm_live_gate opt-in FM_GROK_STOP_LIVE_E2E tmux jq + NATIVE_BIN=${FM_GROK_NATIVE_BIN:-} LEGACY_BIN=${FM_GROK_LEGACY_BIN:-} AUTH=${FM_GROK_AUTH_FILE:-$HOME/.grok/auth.json} @@ -25,7 +22,6 @@ ACTIVE_LAB= [ -x "$LEGACY_BIN" ] || fail "FM_GROK_LEGACY_BIN must be an exact executable path" [ -f "$AUTH" ] || fail "FM_GROK_AUTH_FILE must name the already-managed auth artifact" [ -n "$REAL_TMUX" ] || fail "tmux not found" -command -v jq >/dev/null 2>&1 || fail "jq not found" NATIVE_VERSION=$($NATIVE_BIN --version) LEGACY_VERSION=$($LEGACY_BIN --version) diff --git a/tests/fm-harness-adapter-instructions-live-e2e.test.sh b/tests/fm-harness-adapter-instructions-live-e2e.test.sh index 5b693fc0775..fcab6502152 100644 --- a/tests/fm-harness-adapter-instructions-live-e2e.test.sh +++ b/tests/fm-harness-adapter-instructions-live-e2e.test.sh @@ -7,22 +7,17 @@ # unconfigured native harness loaded the references itself. set -u -if [ "${FM_HARNESS_ADAPTER_INSTRUCTION_EVAL:-0}" != 1 ]; then - echo "skip: set FM_HARNESS_ADAPTER_INSTRUCTION_EVAL=1 and FM_HARNESS_ADAPTER_LOCAL_MODEL=<model> to run the local instruction evaluation" - exit 0 -fi - # shellcheck source=tests/lib.sh . "$(dirname "${BASH_SOURCE[0]}")/lib.sh" +fm_live_gate opt-in FM_HARNESS_ADAPTER_INSTRUCTION_EVAL curl jq + ROUTER="$ROOT/.agents/skills/harness-adapters/SKILL.md" TMP_ROOT=$(fm_test_tmproot fm-harness-adapter-instructions) EXPECTED_JSON="$TMP_ROOT/expected.json" PROMPT_FILE="$TMP_ROOT/prompt.txt" RESPONSE_JSON="$TMP_ROOT/response.json" -command -v curl >/dev/null 2>&1 || fail "curl is required for the local instruction evaluation" -command -v jq >/dev/null 2>&1 || fail "jq is required for the local instruction evaluation" curl -fsS --max-time 2 http://127.0.0.1:11434/api/tags > "$TMP_ROOT/tags.json" \ || fail "local Ollama is unavailable at 127.0.0.1:11434; no remote provider fallback is allowed" diff --git a/tests/fm-harness-liveness-drift-live-e2e.test.sh b/tests/fm-harness-liveness-drift-live-e2e.test.sh index dbaeafe23fe..a48e0e83669 100755 --- a/tests/fm-harness-liveness-drift-live-e2e.test.sh +++ b/tests/fm-harness-liveness-drift-live-e2e.test.sh @@ -1,5 +1,5 @@ #!/usr/bin/env bash -# tests/fm-harness-liveness-drift-live-e2e.test.sh - opt-in drift guard proving +# tests/fm-harness-liveness-drift-live-e2e.test.sh - default-on drift guard proving # every INSTALLED harness is still classified `alive` by the tmux liveness # probe (bin/backends/tmux.sh). # @@ -16,16 +16,17 @@ # unauthenticated harness still starts its process, which is all the liveness # probe reads. # -# Standard CI has no harness binaries or credentials, so this real-harness guard -# is opt-in and on-demand. The portable counterpart in +# Portable serial CI installs the public Pi package but no credentials, so this +# guard checks that available token-free surface there and runs against every installed +# harness on more capable hosts. The portable counterpart in # tests/fm-tmux-agent-liveness.test.sh pins the classifier logic in CI. Run this # guard after any harness upgrade and before trusting refreshed evidence. set -u -if [ "${FM_HARNESS_LIVENESS_DRIFT:-0}" != 1 ]; then - echo "skip: set FM_HARNESS_LIVENESS_DRIFT=1 to run the installed-harness liveness drift guard" - exit 0 -fi +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate default-on FM_HARNESS_LIVENESS_DRIFT tmux ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" @@ -36,7 +37,6 @@ fail() { printf 'not ok - %s\n' "$1" >&2; cleanup_all; exit 1; } pass() { printf 'ok - %s\n' "$1"; } note() { printf '# %s\n' "$1"; } -command -v tmux >/dev/null 2>&1 || fail "tmux not found" REAL_TMUX=$(command -v tmux) SOCKET="fm-liveness-drift-$$" LAB=$(mktemp -d "${TMPDIR:-/tmp}/fm-liveness-drift.XXXXXX") @@ -76,6 +76,18 @@ fm_backend_source tmux || fail "fm_backend_source tmux failed" # order so this guard covers the same binary firstmate would actually launch. resolve_harness_binary() { # <harness> local harness=$1 candidate + # cursor is resolved FIRST, before the generic PATH lookup, and only through + # the verified owner fm-spawn uses. The Cursor agent never installs as + # `cursor`: it installs as `cursor-agent` plus the legacy alias `agent`. A + # machine that also has the Cursor editor does have an executable `cursor` on + # PATH, and launching that one exits immediately, leaving a bare shell in the + # pane that this guard then reports as liveness drift the classifier can do + # nothing about. Asking the owner first also keeps an unrelated executable + # named `agent` rejected here exactly as it would be at launch. + if [ "$harness" = cursor ]; then + fm_cursor_resolve_binary 2>/dev/null && return 0 + return 1 + fi candidate=$(command -v "$harness" 2>/dev/null || true) if [ -n "$candidate" ] && [ -x "$candidate" ]; then printf '%s\n' "$candidate" @@ -85,15 +97,6 @@ resolve_harness_binary() { # <harness> printf '%s\n' "$HOME/.kimi-code/bin/kimi" return 0 fi - # cursor is never on PATH under the name `cursor`: it installs as - # `cursor-agent` plus the legacy alias `agent`, and its user-local install is - # routinely absent from a non-interactive PATH. Resolve it through the same - # verified owner fm-spawn uses, so an unrelated executable named `agent` is - # rejected here exactly as it would be at launch. - if [ "$harness" = cursor ]; then - fm_cursor_resolve_binary 2>/dev/null && return 0 - return 1 - fi return 1 } diff --git a/tests/fm-herdr-submit-confirm-live-e2e.test.sh b/tests/fm-herdr-submit-confirm-live-e2e.test.sh index 9140fec1e6e..c7813b28393 100755 --- a/tests/fm-herdr-submit-confirm-live-e2e.test.sh +++ b/tests/fm-herdr-submit-confirm-live-e2e.test.sh @@ -14,20 +14,17 @@ # Every Herdr call, including adapter calls, is routed through bin/fm-herdr-lab.sh. set -u +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" LAB_HELPER=${HERDR_LAB_HELPER:-$ROOT/bin/fm-herdr-lab.sh} fail() { printf 'not ok - %s\n' "$1" >&2; exit 1; } pass() { printf 'ok - %s\n' "$1"; } -if [ "${FM_HERDR_SUBMIT_CONFIRM_LIVE:-0}" != 1 ]; then - echo "skip: set FM_HERDR_SUBMIT_CONFIRM_LIVE=1 to run the live Herdr submit-confirmation guard" - exit 0 -fi +fm_live_gate opt-in FM_HERDR_SUBMIT_CONFIRM_LIVE herdr jq claude -command -v herdr >/dev/null 2>&1 || fail "FM_HERDR_SUBMIT_CONFIRM_LIVE=1 but herdr is not installed" -command -v jq >/dev/null 2>&1 || fail "FM_HERDR_SUBMIT_CONFIRM_LIVE=1 but jq is not installed" -command -v claude >/dev/null 2>&1 || fail "FM_HERDR_SUBMIT_CONFIRM_LIVE=1 but Claude Code is not installed" [ -x "$LAB_HELPER" ] || fail "FM_HERDR_SUBMIT_CONFIRM_LIVE=1 but the Herdr lab helper is not executable at $LAB_HELPER" # shellcheck source=tests/herdr-test-safety.sh diff --git a/tests/fm-herdr-version-floor-live-e2e.test.sh b/tests/fm-herdr-version-floor-live-e2e.test.sh index 24463b520dd..5f721214d6e 100755 --- a/tests/fm-herdr-version-floor-live-e2e.test.sh +++ b/tests/fm-herdr-version-floor-live-e2e.test.sh @@ -1,5 +1,5 @@ #!/usr/bin/env bash -# Opt-in live guard for the Herdr presentation version floor. +# Default-on live guard for the Herdr presentation version floor. # # Protocol is the floor's structural signal, and its mapping to real releases # is a vendor-supplied fact that no fixture can prove. The runtime gate checks @@ -10,7 +10,8 @@ # identity the binary actually reports. It fails naming the version and protocol # rather than degrading quietly. # -# It is opt-in because it downloads upstream release binaries over the network. +# A run downloads upstream release binaries over the network but submits no +# prompt, so the shared live gate runs it by default wherever its tools exist. # Run it after every Herdr upgrade and before trusting a refreshed # docs/verification/runtime-backends.md "Presentation version floor" entry. # @@ -20,20 +21,17 @@ # lifecycle operation and no server is involved. set -u +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" LAB_HELPER=${HERDR_LAB_HELPER:-$ROOT/bin/fm-herdr-lab.sh} fail() { printf 'not ok - %s\n' "$1" >&2; exit 1; } pass() { printf 'ok - %s\n' "$1"; } -if [ "${FM_HERDR_VERSION_FLOOR_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_HERDR_VERSION_FLOOR_LIVE_E2E=1 to run the real-release Herdr version-floor guard" - exit 0 -fi +fm_live_gate default-on FM_HERDR_VERSION_FLOOR_LIVE_E2E herdr jq curl shasum -for tool in herdr jq curl shasum; do - command -v "$tool" >/dev/null 2>&1 || { echo "skip: $tool not found"; exit 0; } -done [ -x "$LAB_HELPER" ] || { echo "skip: Herdr lab helper not executable at $LAB_HELPER"; exit 0; } case "$(uname -s)/$(uname -m)" in diff --git a/tests/fm-kimi-harness.test.sh b/tests/fm-kimi-harness.test.sh index 1f9b68463c1..265b9d4e604 100755 --- a/tests/fm-kimi-harness.test.sh +++ b/tests/fm-kimi-harness.test.sh @@ -231,6 +231,8 @@ test_kimi_launch_then_send_is_verified() { assert_present "$task_tmp/gotmp" "kimi spawn did not create its Go temp directory" assert_grep "export GOTMPDIR=$task_tmp/gotmp" "$CASE_DIR/tmux-calls.log" \ "kimi spawn did not export its Go temp directory into the pane" + assert_grep "export FM_TASK_ID=$id" "$CASE_DIR/tmux-calls.log" \ + "kimi spawn did not mark the pane with its task id" assert_grep 'BEGIN FIRSTMATE KIMI TURN-END HOOK' "$HOME_DIR/.kimi-code/config.toml" \ "kimi spawn did not install its guarded global hook region" assert_grep 'token=' "$WT_DIR/.fm-kimi-turnend" "kimi spawn did not write its token pointer" diff --git a/tests/fm-launch-lib.test.sh b/tests/fm-launch-lib.test.sh index 4766ea50438..c3e4e3fd61e 100755 --- a/tests/fm-launch-lib.test.sh +++ b/tests/fm-launch-lib.test.sh @@ -71,7 +71,7 @@ test_agy_template() { test_existing_templates_keep_settings_and_hooks() { assert_eq "$(fm_launch_template claude ship)" \ - 'CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --dangerously-skip-permissions --settings '\''{"feedbackDrafts":"off"}'\'' __MODELFLAG____EFFORTFLAG__"$(__OPINPUT__ encode launch-brief < __BRIEF__)"' \ + 'CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --dangerously-skip-permissions --settings '\''{"feedbackDrafts":"off","attribution":{"commit":"","pr":"","sessionUrl":false}}'\'' __MODELFLAG____EFFORTFLAG__"$(__OPINPUT__ encode launch-brief < __BRIEF__)"' \ "claude template drifted" assert_eq "$(fm_launch_template grok ship)" \ 'grok --always-approve __MODELFLAG____EFFORTFLAG__"$(__OPINPUT__ encode launch-brief < __BRIEF__)"' \ diff --git a/tests/fm-live-gate.test.sh b/tests/fm-live-gate.test.sh new file mode 100755 index 00000000000..838f7f48089 --- /dev/null +++ b/tests/fm-live-gate.test.sh @@ -0,0 +1,253 @@ +#!/usr/bin/env bash +# Behavior tests for tests/lib.sh's fm_live_gate, the single decision every +# live-harness guard opens with, and for the wiring that makes that decision +# reach the whole family. +# +# The gate is what turns "a live guard exists" into "a live guard actually ran +# on the machine that has the harness", so the cases below drive it the way a +# guard does: real scripts, executed as separate processes, with a fakebin PATH +# standing in for a host that does or does not have the tool. Nothing here reads +# tests/lib.sh's source text. +# +# The family sweep at the end runs every real live guard with FM_LIVE=0 and +# requires the shared refusal line for upstream guards and the explicit named +# refusal for fork-only guards whose stronger execution opt-ins remain in force. It is +# cheap because a disabled gate exits before a guard touches a harness. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +TMP_ROOT=$(fm_test_tmproot fm-live-gate) +BIN="$TMP_ROOT/bin" +mkdir -p "$BIN" + +# A stand-in for a harness this host has: present on the fakebin PATH, and +# nothing the gate can confuse with a real one. +cat > "$BIN/fmfakeharness" <<'SH' +#!/usr/bin/env bash +exit 0 +SH +chmod +x "$BIN/fmfakeharness" + +clean_env() { + env -i \ + HOME="${HOME:-$TMP_ROOT}" \ + PATH="${PATH:-/usr/bin:/bin}" \ + TMPDIR="${TMPDIR:-/tmp}" \ + LANG="${LANG:-C}" \ + TERM="${TERM:-dumb}" \ + FM_TEST_SKIP_ORPHAN_REAP=1 \ + "$@" +} + +# guard <name> <gate args...>: write a guard script that opens with the shared +# gate and, if the gate lets it through, reports that it ran. +guard() { + local name=$1 + shift + local path="$TMP_ROOT/$name.test.sh" + { + printf '#!/usr/bin/env bash\nset -u\n' + printf '. "%s/tests/lib.sh"\n' "$ROOT" + printf 'fm_live_gate' + printf ' %s' "$@" + printf '\n' + printf 'printf "ran\\n"\n' + } > "$path" + chmod +x "$path" + printf '%s\n' "$path" +} + +# run_guard <path> [env assignment ...]: execute a guard on a PATH that carries +# only the fakebin plus the system essentials, capturing stdout and stderr. +run_guard() { + local path=$1 + shift + local out rc + set +e + out=$(clean_env "$@" PATH="$BIN:/usr/bin:/bin" "$path" 2>&1) + rc=$? + set -e + printf '%s\n' "$rc" + printf '%s\n' "$out" +} + +test_default_on_runs_when_the_tool_is_installed() { + local path result + path=$(guard default-present default-on FM_FAKE_LIVE fmfakeharness) + result=$(run_guard "$path" CI=true) + [ "$(printf '%s' "$result" | sed -n 1p)" = 0 ] || fail "a default-on guard with its tool installed must succeed: $result" + assert_contains "$result" ran "a default-on guard must run with no variable set when its tool is installed" +} + +test_default_on_skips_and_names_the_absent_tool() { + local path result + path=$(guard default-absent default-on FM_FAKE_LIVE fmmissingharness) + result=$(run_guard "$path") + [ "$(printf '%s' "$result" | sed -n 1p)" = 0 ] || fail "an absent tool must be a skip, not a failure: $result" + assert_contains "$result" "skip: live: fmmissingharness absent" \ + "a capability skip must name the tool this host does not have" + assert_not_contains "$result" ran "an absent tool must stop the guard before it runs" +} + +test_opt_in_stays_off_until_asked() { + local path result + path=$(guard optin-idle opt-in FM_FAKE_LIVE fmfakeharness) + result=$(run_guard "$path") + [ "$(printf '%s' "$result" | sed -n 1p)" = 0 ] || fail "an unrequested opt-in guard must skip cleanly: $result" + assert_contains "$result" "skip: live: opt-in; set FM_FAKE_LIVE=1 to run" \ + "an opt-in skip must name the variable that turns the guard on" + assert_not_contains "$result" ran "a token-spending guard must not run unasked" +} + +test_own_variable_turns_an_opt_in_guard_on() { + local path result + path=$(guard optin-on opt-in FM_FAKE_LIVE fmfakeharness) + result=$(run_guard "$path" FM_FAKE_LIVE=1) + [ "$(printf '%s' "$result" | sed -n 1p)" = 0 ] || fail "FM_FAKE_LIVE=1 must run the guard: $result" + assert_contains "$result" ran "an explicitly requested opt-in guard must run" +} + +test_requested_run_fails_rather_than_skipping_on_an_absent_tool() { + local path result + path=$(guard optin-strict opt-in FM_FAKE_LIVE fmmissingharness) + result=$(run_guard "$path" FM_FAKE_LIVE=1) + [ "$(printf '%s' "$result" | sed -n 1p)" = 1 ] || fail "a demanded run with no tool must fail, not skip: $result" + assert_contains "$result" "FM_FAKE_LIVE was requested but fmmissingharness is not installed" \ + "the hard failure must name the request and the missing tool" + assert_not_contains "$result" "skip:" "a demanded run must never report itself as a skip" +} + +test_fm_live_turns_the_whole_family_off_and_on() { + local path result + path=$(guard fmlive-off default-on FM_FAKE_LIVE fmfakeharness) + result=$(run_guard "$path" FM_LIVE=0) + [ "$(printf '%s' "$result" | sed -n 1p)" = 0 ] || fail "FM_LIVE=0 must skip cleanly: $result" + assert_contains "$result" "skip: live: disabled by FM_LIVE=0" "FM_LIVE=0 must say why it skipped" + + path=$(guard fmlive-on opt-in FM_FAKE_LIVE fmfakeharness) + result=$(run_guard "$path" FM_LIVE=1) + assert_contains "$result" ran "FM_LIVE=1 must turn an opt-in guard on" + + path=$(guard fmlive-on-strict opt-in FM_FAKE_LIVE fmmissingharness) + result=$(run_guard "$path" FM_LIVE=1) + [ "$(printf '%s' "$result" | sed -n 1p)" = 1 ] || fail "FM_LIVE=1 must make an absent tool a failure: $result" +} + +test_a_guards_own_variable_wins_over_fm_live() { + local path result + path=$(guard own-off default-on FM_FAKE_LIVE fmfakeharness) + result=$(run_guard "$path" FM_LIVE=1 FM_FAKE_LIVE=0) + [ "$(printf '%s' "$result" | sed -n 1p)" = 0 ] || fail "an explicit per-guard opt-out must skip cleanly: $result" + assert_contains "$result" "skip: live: disabled by FM_FAKE_LIVE=0" \ + "a guard's own 0 must win over FM_LIVE=1 and say so" + assert_not_contains "$result" ran "a guard switched off by name must not run" +} + +test_any_of_several_entry_points_turns_a_guard_on() { + local path result + path=$(guard multi opt-in FM_FAKE_LIVE,FM_FAKE_ALT_LIVE fmfakeharness) + result=$(run_guard "$path") + assert_contains "$result" "set FM_FAKE_LIVE=1 to run" \ + "a multi-entry guard must point at its primary variable when idle" + result=$(run_guard "$path" FM_FAKE_ALT_LIVE=1) + assert_contains "$result" ran "a secondary entry point must also turn the guard on" +} + +test_gate_lets_a_guard_drive_the_real_fleet_scripts_under_a_gate_marker() { + # The nine live guards that never sourced the shared helpers used to be + # refused by bin/fm-gate-refuse-lib.sh whenever the pipeline ran them, because + # the gate marker is set for every no-mistakes gate agent. Opening with the + # shared gate is what carries the test-suite bypass into them. + local path out rc + path="$TMP_ROOT/bypass.test.sh" + { + printf '#!/usr/bin/env bash\nset -u\n' + printf '. "%s/tests/lib.sh"\n' "$ROOT" + printf 'fm_live_gate default-on FM_FAKE_LIVE fmfakeharness\n' + printf '. "%s/bin/fm-gate-refuse-lib.sh"\n' "$ROOT" + printf 'if fm_is_gate_agent; then printf "refused\\n"; else printf "allowed\\n"; fi\n' + } > "$path" + chmod +x "$path" + set +e + out=$(clean_env NO_MISTAKES_GATE=1 PATH="$BIN:/usr/bin:/bin" "$path" 2>&1) + rc=$? + set -e + expect_code 0 "$rc" "a guard opened with the shared gate must not be refused" + assert_contains "$out" allowed \ + "the shared gate must carry the test-suite bypass so a live guard can drive the real fleet scripts" +} + +test_every_live_guard_is_wired_to_the_shared_gate() { + local script out listing expected checked=0 + listing=$("$ROOT/bin/fm-test-run.sh" --family live-harness-optin --list) \ + || fail "could not list the live-harness family" + while IFS= read -r script; do + [ -n "$script" ] || continue + script="$ROOT/$script" + # bin/fm-test-run.sh runs every script through bash, so a guard that is not + # marked executable is still a real suite member here. + expected="skip: live: disabled by FM_LIVE=0" + case "$(basename "$script")" in + fm-afk-pi-dual-supervision-e2e.test.sh) expected="skip: set FM_AFK_PI_DUAL_SUPERVISION_E2E=1" ;; + fm-agy-smoke.test.sh) expected="skip: set FM_CURSOR_AGY_LIVE_E2E=1" ;; + fm-dashboard-browser.test.sh) expected="skip: set FM_DASHBOARD_BROWSER_E2E=1" ;; + fm-gbrain-capture-e2e.test.sh|fm-gbrain-readonly-e2e.test.sh) expected="skip: set FM_GBRAIN_LIVE_E2E=1" ;; + fm-pi-hung-delivery-herdr-e2e.test.sh) expected="skip: set FM_PI_HUNG_DELIVERY_HERDR_E2E=1" ;; + fm-stow-horizon-live-e2e.test.sh) expected="skip: set FM_STOW_HORIZON_LIVE_E2E=1" ;; + esac + out=$(clean_env CI=true FM_LIVE=0 bash "$script" 2>&1) || fail "$(basename "$script") must exit 0 when live guards are disabled" + assert_contains "$out" "$expected" \ + "$(basename "$script") must refuse execution through its owning gate" + checked=$((checked + 1)) + done <<EOF +$listing +EOF + [ "$checked" -ge 20 ] || fail "expected the whole live-guard family to be swept, saw only $checked" + pass "all $checked live guards refuse together on FM_LIVE=0" +} + +test_ordinary_ci_preserves_protected_execution_gates() { + local script expected out serial + serial=$("$ROOT/bin/fm-test-run.sh" --list --lane portable-serial) + while IFS=' ' read -r script expected; do + printf '%s\n' "$serial" | grep -Fxq "tests/$script" \ + || fail "$script must be inventoried in portable serial" + out=$(clean_env CI=true bash "$ROOT/tests/$script" 2>&1) \ + || fail "$script must skip successfully in ordinary CI: $out" + assert_contains "$out" "$expected" "$script must retain its execution opt-in in CI" + done <<'CASES' +fm-agy-smoke.test.sh skip: set FM_CURSOR_AGY_LIVE_E2E=1 +fm-dashboard-browser.test.sh skip: set FM_DASHBOARD_BROWSER_E2E=1 +fm-afk-pi-dual-supervision-e2e.test.sh skip: set FM_AFK_PI_DUAL_SUPERVISION_E2E=1 +fm-pi-hung-delivery-herdr-e2e.test.sh skip: set FM_PI_HUNG_DELIVERY_HERDR_E2E=1 +fm-stow-horizon-live-e2e.test.sh skip: set FM_STOW_HORIZON_LIVE_E2E=1 +fm-gbrain-capture-e2e.test.sh skip: set FM_GBRAIN_LIVE_E2E=1 +fm-gbrain-readonly-e2e.test.sh skip: set FM_GBRAIN_LIVE_E2E=1 +fm-codex-continuity-live-e2e.test.sh skip: live: opt-in; +CASES + pass "ordinary CI inventories protected guards without executing their live workloads" +} + +test_default_on_runs_when_the_tool_is_installed +pass "a default-on guard runs wherever its tools are installed" +test_default_on_skips_and_names_the_absent_tool +pass "an absent tool is a named capability skip, not a silent pass" +test_opt_in_stays_off_until_asked +pass "a prompt-submitting guard stays off until it is asked for" +test_own_variable_turns_an_opt_in_guard_on +pass "a guard's own variable turns it on" +test_requested_run_fails_rather_than_skipping_on_an_absent_tool +pass "a demanded run refuses to pass as a skip" +test_fm_live_turns_the_whole_family_off_and_on +pass "FM_LIVE switches the whole family" +test_a_guards_own_variable_wins_over_fm_live +pass "a guard's own setting wins over FM_LIVE" +test_any_of_several_entry_points_turns_a_guard_on +pass "any entry point of a multi-mode guard turns it on" +test_gate_lets_a_guard_drive_the_real_fleet_scripts_under_a_gate_marker +pass "the shared gate carries the gate-refusal bypass into every live guard" +test_every_live_guard_is_wired_to_the_shared_gate + +test_ordinary_ci_preserves_protected_execution_gates diff --git a/tests/fm-muse-signals-live-e2e.test.sh b/tests/fm-muse-signals-live-e2e.test.sh index 85874ac9da5..a406b9af601 100755 --- a/tests/fm-muse-signals-live-e2e.test.sh +++ b/tests/fm-muse-signals-live-e2e.test.sh @@ -1,6 +1,9 @@ #!/usr/bin/env bash set -u +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" MUSE_BIN=$(command -v muse 2>/dev/null || true) REAL_TMUX=$(command -v tmux 2>/dev/null || true) @@ -116,14 +119,7 @@ if [ "${1:-}" = --ansi-self-test ]; then exit 0 fi -if [ "${FM_MUSE_SIGNALS_LIVE:-0}" != 1 ]; then - echo "skip: set FM_MUSE_SIGNALS_LIVE=1 to run the real Muse signal drift guard" - exit 0 -fi - -[ -x "$MUSE_BIN" ] || fail "FM_MUSE_SIGNALS_LIVE=1 but no real muse executable is installed on PATH" -[ -x "$REAL_TMUX" ] || fail "FM_MUSE_SIGNALS_LIVE=1 but tmux is not installed" -command -v node >/dev/null 2>&1 || fail "node is required to inspect Muse's serialized session protocol" +fm_live_gate opt-in FM_MUSE_SIGNALS_LIVE muse tmux node LAB=$(mktemp -d "${TMPDIR:-/tmp}/fm-muse-signals.XXXXXX") || fail "could not create the isolated Muse lab" trap cleanup EXIT diff --git a/tests/fm-omp-primary-live-e2e.test.sh b/tests/fm-omp-primary-live-e2e.test.sh index bf404a40a84..333c9b27c37 100755 --- a/tests/fm-omp-primary-live-e2e.test.sh +++ b/tests/fm-omp-primary-live-e2e.test.sh @@ -16,10 +16,10 @@ # the turn-end guard continuation and the model reaches for the tool. set -u -if [ "${FM_OMP_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_OMP_LIVE_E2E=1 to run the isolated omp primary regression" - exit 0 -fi +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate opt-in FM_OMP_LIVE_E2E omp node jq ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" unset NO_MISTAKES_GATE @@ -45,10 +45,6 @@ fail() { pass() { printf 'ok - %s\n' "$1"; } note() { printf '# %s\n' "$1"; } -command -v omp >/dev/null 2>&1 || fail "omp not found" -command -v node >/dev/null 2>&1 || fail "node not found" -command -v jq >/dev/null 2>&1 || fail "jq not found" - OMP_VERSION=$(omp --version 2>/dev/null | head -1) MODEL=${FM_OMP_LIVE_MODEL:-openai-codex/gpt-6-astra} # The guard's beacon grace for this lab. The watcher beats every FM_POLL=1s, so diff --git a/tests/fm-opencode-primary-live-e2e.test.sh b/tests/fm-opencode-primary-live-e2e.test.sh index 58379c01975..0d17de05a2f 100755 --- a/tests/fm-opencode-primary-live-e2e.test.sh +++ b/tests/fm-opencode-primary-live-e2e.test.sh @@ -3,10 +3,10 @@ # FM_HOME. Existing OpenCode credentials stay in their managed store. set -u -if [ "${FM_OPENCODE_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_OPENCODE_LIVE_E2E=1 to run the interactive OpenCode continuity regression" - exit 0 -fi +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate opt-in FM_OPENCODE_LIVE_E2E opencode tmux sqlite3 ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" unset NO_MISTAKES_GATE @@ -16,10 +16,6 @@ fail() { exit 1 } -command -v opencode >/dev/null 2>&1 || fail "opencode not found" -command -v tmux >/dev/null 2>&1 || fail "tmux not found" -command -v sqlite3 >/dev/null 2>&1 || fail "sqlite3 not found" - TMUX=$(command -v tmux) SOCKET="fm-opencode-live-e2e-$$" SESSION=opencode-live-e2e diff --git a/tests/fm-pi-branch-extension.test.sh b/tests/fm-pi-branch-extension.test.sh index eeabd4f421d..19e569d4d8c 100644 --- a/tests/fm-pi-branch-extension.test.sh +++ b/tests/fm-pi-branch-extension.test.sh @@ -106,6 +106,10 @@ export class ModelRuntime { constructor() { this.models = (globalThis.__fmBranchStaticModels?.() ?? []).map((model) => ({ ...model })); this.authenticated = new Set(this.models.filter((model) => model.storedAuth !== false).map((model) => model.provider)); + this.registeredProviderConfigs = new Map(); + // Like the real runtime, a registered provider's credentials are only + // known once refresh() has run for it; registration alone is provisional. + this.pendingAuth = new Set(); } static async create() { const queuedError = globalThis.__fmModelRuntimeErrors?.shift(); @@ -115,6 +119,18 @@ export class ModelRuntime { (globalThis.__fmModelRuntimes ??= []).push(runtime); return runtime; } + registerProvider(providerId, config) { + this.registeredProviderConfigs.set(providerId, config); + for (const model of config.models ?? []) { + this.models.push({ ...model, provider: providerId }); + } + if (config.oauth || config.apiKey) this.pendingAuth.add(providerId); + } + async refresh(options) { + for (const providerId of options?.providers ?? this.pendingAuth) { + if (this.pendingAuth.delete(providerId)) this.authenticated.add(providerId); + } + } getModel(provider, id) { return this.models.find((model) => model.provider === provider && model.id === id); } @@ -446,6 +462,8 @@ const modelRegistry = { getAvailable: () => registryModels.filter((model) => model.mainAvailable !== false).slice(), find: (provider, id) => registryModels.find((model) => model.provider === provider && model.id === id), hasConfiguredAuth: (model) => model.mainAvailable !== false, + getRegisteredProviderConfig: (providerId) => globalThis.__fmExtensionProviderConfigs?.get(providerId), + getRegisteredProviderIds: () => [...(globalThis.__fmExtensionProviderConfigs?.keys() ?? [])], }; function makeCtx(extra) { return { @@ -4753,6 +4771,98 @@ EOF pass "a failed cursor write re-delivers a routine note exactly once more while a captain outcome stays deduplicated" } +test_extension_registered_provider_resolves_in_the_branch() { + local repo home out status + repo="$TMP_ROOT/extprov-root" + home="$TMP_ROOT/extprov-home" + mkdir -p "$home/state" "$home/config" + install_pi_branch_extension_fixture "$repo" + PLUGIN="$repo/.pi/extensions/fm-branch-supervision.ts" FM_HOME="$home" FM_ROOT_OVERRIDE="$ROOT" \ + DRIVER_PRELUDE="$DRIVER_PRELUDE" node --input-type=module > "$TMP_ROOT/node-output" 2>&1 <<'EOF' +const prelude = process.env.DRIVER_PRELUDE; +await eval(`(async () => { ${prelude}; globalThis.__t = { fire, dispatch, settle, makeCtx, registryModels, uiSelections, uiPrompts, notices, commands, home }; })()`); +const { fire, dispatch, settle, makeCtx, registryModels, uiSelections, uiPrompts, notices, commands, home } = globalThis.__t; +import { readFileSync, writeFileSync } from "node:fs"; + +// An extension-registered provider exists only in main's registry, never in +// the isolated branch runtime's static catalog. Registering its config on +// main's registry is what makes it resolvable for the branch. +registryModels.push( + { provider: "anthropic", id: "main-model" }, + // Available in main's registry but absent from the branch runtime's static + // catalog, exactly like a provider an extension registered at runtime. + { provider: "devin", id: "swe-1-7", branchAvailable: false }, +); +globalThis.__fmExtensionProviderConfigs = new Map([ + [ + "devin", + { + name: "Devin (Cognition)", + api: "devin-cloud", + baseUrl: "https://server.codeium.com", + models: [{ id: "swe-1-7", name: "SWE 1.7", reasoning: false, input: ["text"], cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 }, contextWindow: 200000, maxTokens: 8192 }], + oauth: { name: "Devin (Cognition / Windsurf)", login: async () => ({}), refreshToken: async (c) => c, getApiKey: (c) => c.access }, + streamSimple: () => {}, + }, + ], +]); + +await fire("session_start", {}, makeCtx()); + +// The picker must offer the extension-registered model: it is available in +// main's registry and resolvable in the branch runtime once its registration +// is copied across. +const command = commands.get("supervision-model"); +if (!command) throw new Error("the supervision-model command was not registered"); +uiSelections.push("devin/swe-1-7"); +await command.handler("", makeCtx()); +const offered = uiPrompts[0]; +if (!offered.options.includes("devin/swe-1-7")) { + throw new Error(`the picker must offer an extension-registered provider the branch can run: ${JSON.stringify(offered.options)}`); +} +if (readFileSync(`${home}/config/supervision-branch-model`, "utf8") !== "devin/swe-1-7\n") { + throw new Error("the extension-registered pick was not persisted"); +} +dispatch("signal: extension provider pin"); +await settle(() => (globalThis.__fmSessions ?? []).length === 1, "pinned extension-provider branch build"); +const pinned = globalThis.__fmSessions[0].options.model; +if (!pinned || pinned.provider !== "devin" || pinned.id !== "swe-1-7") { + throw new Error(`the extension-registered pin did not bind the branch: ${JSON.stringify(pinned)}`); +} +// Copying the provider registration must not loosen the branch's isolation: +// the devin-pinned session still loads no extensions, skills, or context files. +const pinnedLoader = globalThis.__fmLoaders.at(-1); +for (const key of ["noExtensions", "noSkills", "noContextFiles"]) { + if (pinnedLoader.options[key] !== true) throw new Error(`devin-pinned branch loader must keep ${key}`); +} + +// Without the registration, the same pin is unavailable and the branch +// refuses to build rather than silently downgrading. +globalThis.__fmExtensionProviderConfigs = new Map(); +await fire("session_shutdown", {}); +await fire("session_start", {}, makeCtx()); +const unregisteredOffer = dispatch("signal: unregistered provider pin"); +if (!unregisteredOffer.accepted) throw new Error("unregistered-pin wake was not initially accepted"); +const unregisteredFailure = await unregisteredOffer.settlement.then( + () => null, + (error) => error, +); +if ( + !(unregisteredFailure instanceof Error) || + !unregisteredFailure.message.includes("devin/swe-1-7") || + !unregisteredFailure.message.includes("supervision model pin") +) { + throw new Error(`the unregistered pin did not reject with its own name: ${String(unregisteredFailure)}`); +} +if ((globalThis.__fmSessions ?? []).length !== 1) throw new Error("an unregistered pin must not build a second branch session"); +process.exit(0); +EOF + status=$? + out=$(cat "$TMP_ROOT/node-output") + expect_code 0 "$status" "an extension-registered provider must resolve in the isolated branch runtime: $out" + pass "an extension-registered provider resolves in the isolated branch runtime" +} + test_outcomes_tool_uses_stock_execution_and_export_consumers test_real_pi_picker_primitives_stay_bounded_and_searchable test_branch_dispatch_two_stage_filter_and_prefix_contract @@ -4781,6 +4891,7 @@ test_supervision_model_picker_is_bounded_searchable_and_branch_only test_branch_model_picker_keeps_follow_main_first_under_ranking test_branch_effort_pin_applies_and_absent_pin_follows_main test_unpinned_branch_follows_main_effort_changes_live +test_extension_registered_provider_resolves_in_the_branch test_supervision_model_command_picks_effort_after_the_model test_unusable_model_pin_falls_back_to_main test_replacement_activation_cleans_leases_and_retries_failure diff --git a/tests/fm-pi-branch-live-e2e.test.sh b/tests/fm-pi-branch-live-e2e.test.sh index c5c1f8efdc4..745f9346b4e 100644 --- a/tests/fm-pi-branch-live-e2e.test.sh +++ b/tests/fm-pi-branch-live-e2e.test.sh @@ -32,13 +32,10 @@ # (docs/verification/runtime-backends.md). set -u -if [ "${FM_PI_BRANCH_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_PI_BRANCH_LIVE_E2E=1 to run the real-SDK Pi branch regression" - exit 0 -fi - # shellcheck source=tests/lib.sh . "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate opt-in FM_PI_BRANCH_LIVE_E2E npm jq node export NODE_NO_WARNINGS=1 PI_PACKAGE_DIR=${FM_PI_PACKAGE_DIR:-"$(npm root -g)/@earendil-works/pi-coding-agent"} diff --git a/tests/fm-pi-branch-responsiveness-live-e2e.test.sh b/tests/fm-pi-branch-responsiveness-live-e2e.test.sh index e46724e9f55..bd2645378b5 100755 --- a/tests/fm-pi-branch-responsiveness-live-e2e.test.sh +++ b/tests/fm-pi-branch-responsiveness-live-e2e.test.sh @@ -1,5 +1,5 @@ #!/usr/bin/env bash -# Opt-in live guard for the ONE thing only a real Pi TUI can answer: whether +# Default-on live guard for the ONE thing only a real Pi TUI can answer: whether # the captain can still type and see the screen repaint while a supervision # outcome is being delivered into his session. # @@ -28,17 +28,10 @@ # which runs the same reconcile-and-deliver chain an arriving outcome runs. set -u -if [ "${FM_PI_BRANCH_RESPONSIVENESS_E2E:-0}" != 1 ]; then - echo "skip: set FM_PI_BRANCH_RESPONSIVENESS_E2E=1 to run the real-Pi delivery responsiveness guard" - exit 0 -fi - # shellcheck source=tests/lib.sh . "$(dirname "${BASH_SOURCE[0]}")/lib.sh" -command -v pi >/dev/null 2>&1 || fail "pi not installed: the real-Pi responsiveness guard cannot report a verdict without it" -command -v tmux >/dev/null 2>&1 || fail "tmux not installed: the real-Pi responsiveness guard cannot drive a pane without it" -command -v node >/dev/null 2>&1 || fail "node not installed: the real-Pi responsiveness guard cannot time keystroke echo without it" +fm_live_gate default-on FM_PI_BRANCH_RESPONSIVENESS_E2E pi tmux node PI_VERSION=$(pi --version 2>/dev/null || printf 'unknown') TMUX=$(command -v tmux) diff --git a/tests/fm-pi-primary-live-e2e.test.sh b/tests/fm-pi-primary-live-e2e.test.sh index f1517d23f2e..0538692eb8f 100755 --- a/tests/fm-pi-primary-live-e2e.test.sh +++ b/tests/fm-pi-primary-live-e2e.test.sh @@ -4,10 +4,10 @@ # copying credentials and pins the captain-approved openai-codex model. set -u -if [ "${FM_PI_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_PI_LIVE_E2E=1 to run the isolated interactive Pi regression" - exit 0 -fi +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate opt-in FM_PI_LIVE_E2E pi tmux ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" unset NO_MISTAKES_GATE @@ -17,9 +17,6 @@ fail() { exit 1 } -command -v pi >/dev/null 2>&1 || fail "pi not found" -command -v tmux >/dev/null 2>&1 || fail "tmux not found" - TMUX=$(command -v tmux) SOCKET="fm-pi-live-e2e-$$" SESSION=pi-live-e2e diff --git a/tests/fm-pi-primary-types.test.sh b/tests/fm-pi-primary-types.test.sh index 0fe52d4b7d0..bf3ed2ad925 100755 --- a/tests/fm-pi-primary-types.test.sh +++ b/tests/fm-pi-primary-types.test.sh @@ -4,12 +4,12 @@ set -u ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" -command -v npm >/dev/null 2>&1 || { echo "skip: npm not found for Pi extension typecheck"; exit 0; } -command -v tsc >/dev/null 2>&1 || { echo "skip: tsc not found for Pi extension typecheck"; exit 0; } +command -v npm >/dev/null 2>&1 || { echo "skip: Pi extension typecheck prerequisite not found: npm"; exit 0; } +command -v tsc >/dev/null 2>&1 || { echo "skip: Pi extension typecheck prerequisite not found: tsc"; exit 0; } PI_PACKAGE_DIR=${FM_PI_PACKAGE_DIR:-"$(npm root -g)/@earendil-works/pi-coding-agent"} if [ ! -f "$PI_PACKAGE_DIR/package.json" ]; then - echo "skip: installed @earendil-works/pi-coding-agent package not found" + echo "skip: Pi extension typecheck prerequisite not found: installed @earendil-works/pi-coding-agent package" exit 0 fi if [ ! -d "$PI_PACKAGE_DIR/node_modules/typebox" ] || \ diff --git a/tests/fm-pi-windows-shell-invocation.test.sh b/tests/fm-pi-windows-shell-invocation.test.sh new file mode 100755 index 00000000000..a86acb6d8e4 --- /dev/null +++ b/tests/fm-pi-windows-shell-invocation.test.sh @@ -0,0 +1,119 @@ +#!/usr/bin/env bash +# Native-Windows Pi extension regression for invoking tracked Bash owners through bash. +set -u + +ROOT=$(cd "$(dirname "$0")/.." && pwd) +. "$ROOT/tests/lib.sh" + +TMP_ROOT=$(fm_test_tmproot fm-pi-windows-shell-invocation) + +if [ "$(node -p 'process.platform')" != win32 ]; then + echo "skip: native Windows Node required" + exit 0 +fi + +project="$TMP_ROOT/project" +mkdir -p "$project/.pi/extensions/lib" "$project/bin" "$project/state" +cp "$ROOT/.pi/extensions/fm-primary-turnend-guard.ts" "$project/.pi/extensions/" +cp "$ROOT/.pi/extensions/lib/fm-operational-input.ts" \ + "$ROOT/.pi/extensions/lib/fm-sessionstart-supervisor.mjs" "$project/.pi/extensions/lib/" + +cat >"$project/bin/fm-sessionstart-run.sh" <<'SH' +#!/usr/bin/env bash +printf 'sessionstart:%s\n' "$*" >> "$FM_WINDOWS_SHELL_LOG" +SH +cat >"$project/bin/fm-cd-pretool-check.sh" <<'SH' +#!/usr/bin/env bash +printf 'cd:%s\n' "$*" >> "$FM_WINDOWS_SHELL_LOG" +SH +cat >"$project/bin/fm-arm-pretool-check.sh" <<'SH' +#!/usr/bin/env bash +printf 'arm:%s\n' "$*" >> "$FM_WINDOWS_SHELL_LOG" +SH +cat >"$project/bin/fm-turnend-guard.sh" <<'SH' +#!/usr/bin/env bash +cat >/dev/null +printf 'turnend:%s\n' "$*" >> "$FM_WINDOWS_SHELL_LOG" +SH +cat >"$project/bin/fm-operational-input.sh" <<'SH' +#!/usr/bin/env bash +printf 'operational:%s\n' "$*" >> "$FM_WINDOWS_SHELL_LOG" +input=$(cat) +if [ "$1" = encode ]; then + printf 'encoded:%s:%s\n' "$2" "$input" +else + printf 'not-operational\n' +fi +SH +chmod +x "$project/bin/"*.sh + +log="$project/state/calls" +out=$(EXT="$project/.pi/extensions/fm-primary-turnend-guard.ts" \ + FM_HOME="$project" FM_ROOT_OVERRIDE="$project" FM_WINDOWS_SHELL_LOG="$log" \ + FM_OPERATIONAL_INPUT_SCRIPT="$project/bin/fm-operational-input.sh" \ + node --input-type=module 2>&1 <<'JS' +import { spawn } from "node:child_process"; +import { readFileSync } from "node:fs"; +import { pathToFileURL } from "node:url"; + +const handlers = new Map(); +const pi = { + on(event, handler) { handlers.set(event, handler); }, + sendMessage() {}, +}; +const extension = await import(`${pathToFileURL(process.env.EXT).href}?windows=${Date.now()}`); +extension.default(pi); +const ctx = { sessionManager: { getSessionId: () => "windows-test" } }; +handlers.get("session_start")({ reason: "startup" }, ctx); +await handlers.get("before_agent_start")({}, ctx); +await handlers.get("tool_call")({ type: "tool_call", toolName: "bash", input: { command: "printf test" } }); +await handlers.get("agent_settled")({}, ctx); +const operational = await import(`${new URL("./lib/fm-operational-input.ts", pathToFileURL(process.env.EXT)).href}?windows=${Date.now()}`); +operational.classifyFirstmateOperationalText("probe"); +let asyncInvocation; +const encoded = await operational.encodeFirstmateOperationalInputWith( + (command, args, { input }) => { + asyncInvocation = { command, args: [...args], input }; + return new Promise((resolve, reject) => { + const child = spawn(command, args, { stdio: ["pipe", "pipe", "ignore"] }); + let stdout = ""; + child.stdout.on("data", (chunk) => { stdout += chunk; }); + child.on("error", reject); + child.on("close", (status) => resolve({ status, stdout })); + child.stdin.end(input); + }); + }, + "branch-outcome", + "branch result", +); +if (encoded !== "encoded:branch-outcome:branch result\n") { + throw new Error(`unexpected encoded branch outcome: ${encoded}`); +} +if ( + asyncInvocation.command !== "bash" || + asyncInvocation.args.join("\0") !== [ + process.env.FM_OPERATIONAL_INPUT_SCRIPT, + "encode", + "branch-outcome", + ].join("\0") || + asyncInvocation.input !== "branch result" +) { + throw new Error(`unexpected async invocation: ${JSON.stringify(asyncInvocation)}`); +} +const calls = readFileSync(process.env.FM_WINDOWS_SHELL_LOG, "utf8"); +for (const expected of [ + "sessionstart:--source startup --pi-prerequisite", + "cd:--command printf test", + "arm:--command printf test", + "turnend:", + "operational:classify", + "operational:encode branch-outcome", +]) { + if (!calls.includes(expected)) throw new Error(`missing ${expected} in:\n${calls}`); +} +JS +) +status=$? +expect_code 0 "$status" "native-Windows Pi shell seams" +[ -z "$out" ] || fail "native-Windows Pi shell seam test printed output: $out" +pass "Pi session-start, pre-tool, turn-end, and operational-input seams invoke Bash owners on native Windows" diff --git a/tests/fm-pr-check-security.test.sh b/tests/fm-pr-check-security.test.sh index 10a8bd57e82..aaf3ab76ac3 100755 --- a/tests/fm-pr-check-security.test.sh +++ b/tests/fm-pr-check-security.test.sh @@ -985,6 +985,12 @@ test_postrename_poll_validation_revokes_and_retries() { local artifact action dir state destination link_target gate for artifact in data registration check; do for action in type mode device content; do + # The device fault is injected by a fake stat on PATH; on Darwin the + # device helper now calls /usr/bin/stat directly, so the fake can never + # fire there. Skip the device action on Darwin. + if [ "$action" = device ] && [ "$(uname)" = Darwin ]; then + continue + fi dir=$(make_case "poll-final-$artifact-$action") state="$dir/home/state" write_poll_meta "$state" task-a https://github.com/o/r/pull/1 diff --git a/tests/fm-procevent-when.test.sh b/tests/fm-procevent-when.test.sh index 259286beb85..396c48df8d9 100755 --- a/tests/fm-procevent-when.test.sh +++ b/tests/fm-procevent-when.test.sh @@ -21,23 +21,10 @@ export FM_PROCEVENT_CLAIM_ROOT="$TMP_ROOT/claims" pe() { FM_HOME="$1" "$ROOT/bin/fm-procevent.sh" "${@:2}"; } when() { FM_HOME="$1" "$ROOT/bin/fm-procevent-when.sh" "${@:2}"; } -# Every home this suite arms is tracked so teardown can stop any runner still -# blocked on a condition that never fires. -WHEN_HOMES=() -when_teardown() { - local home seen=$'\n' - for home in ${WHEN_HOMES[@]+"${WHEN_HOMES[@]}"}; do - case "$seen" in - *$'\n'"$home"$'\n'*) continue ;; - esac - seen+="$home"$'\n' - FM_HOME="$home" "$ROOT/bin/fm-procevent.sh" sweep-home >/dev/null 2>&1 || true - done - fm_test_cleanup -} -trap when_teardown EXIT - -new_home() { mkdir -p "$1/state"; WHEN_HOMES+=("$1"); } +# Every home this suite arms is registered with tests/lib.sh, which sweeps it +# from every cleanup path so a runner still blocked on a condition that never +# fires cannot survive the run. +new_home() { mkdir -p "$1/state"; fm_test_track_procevent_home "$1"; } wake_payloads() { awk -F '\t' '{print $5}' "$1/state/.wake-queue" 2>/dev/null; } diff --git a/tests/fm-procevent.test.sh b/tests/fm-procevent.test.sh index 784a26ff2e8..7b0c3bd1324 100755 --- a/tests/fm-procevent.test.sh +++ b/tests/fm-procevent.test.sh @@ -24,9 +24,13 @@ BLOCKER="$TMP_ROOT/blocker.sh" cat > "$BLOCKER" <<'SH' #!/usr/bin/env bash # Blocks until the trigger exists, then emits its payload. Completion is the -# event; nothing here polls on a schedule. +# event; nothing here polls on a schedule. The wait is bounded so a stub that +# escapes its test cannot keep spawning processes indefinitely. trigger=$1; shift -while [ ! -e "$trigger" ]; do sleep 0.05; done +while [ ! -e "$trigger" ]; do + [ "$SECONDS" -lt "${FM_TEST_STUB_MAX_BLOCK_SECONDS:-120}" ] || exit 75 + sleep 0.05 +done [ -n "${BLOCKER_STDERR:-}" ] && printf 'noise on stderr\n' >&2 [ -n "${BLOCKER_EXIT:-}" ] && exit "$BLOCKER_EXIT" printf '%s\n' "$@" @@ -35,31 +39,17 @@ chmod +x "$BLOCKER" pe() { FM_HOME="$1" "$ROOT/bin/fm-procevent.sh" "${@:2}"; } -# Every source this suite registers is tracked so teardown can stop its runner. -# A runner started by reconcile is detached and reparented, so a source that -# never completes outlives the suite unless it is retired explicitly - removing -# the fixture directory does not stop an already-running child. -PE_TRACKED=() +# Every home this suite registers a source in is tracked so teardown can stop +# its runners. A runner started by reconcile is detached and reparented, so a +# source that never completes outlives the suite unless its home is swept - +# removing the fixture directory does not stop an already-running child. +# tests/lib.sh owns that sweep and runs it from every cleanup path. pe_register() { # <home> <adapter> <source-id> -- <argv>... local home=$1 adapter=$2 id=$3 shift 3 - PE_TRACKED+=("$home|$id") + fm_test_track_procevent_home "$home" pe "$home" register "$adapter" "$id" "$@" } - -procevent_teardown() { - local entry home seen=$'\n' - for entry in ${PE_TRACKED[@]+"${PE_TRACKED[@]}"}; do - home=${entry%%|*} - case "$seen" in - *$'\n'"$home"$'\n'*) continue ;; - esac - seen+="$home"$'\n' - FM_HOME="$home" "$ROOT/bin/fm-procevent.sh" sweep-home >/dev/null 2>&1 || true - done - fm_test_cleanup -} -trap procevent_teardown EXIT new_home() { mkdir -p "$1/state"; } wake_payloads() { awk -F '\t' '{print $5}' "$1/state/.wake-queue" 2>/dev/null; } @@ -417,7 +407,7 @@ pe_adapter() { # <home> <command>...: run the runner against the fixture adapte } HPUBLISH="$TMP_ROOT/hpublish"; new_home "$HPUBLISH" -PE_TRACKED+=("$HPUBLISH|publish-src") +fm_test_track_procevent_home "$HPUBLISH" pe_adapter "$HPUBLISH" register applying publish-src -- /bin/echo "apply after publish" >/dev/null mkdir "$HPUBLISH/state/.wake-queue" out=$(pe_adapter "$HPUBLISH" start publish-src 2>&1) @@ -449,7 +439,7 @@ pass "automatic application waits for durable publication and failed publication # channel is the announcement. The declaration never silences a capture the # adapter could NOT apply - that one still publishes for the handler. HSELF="$TMP_ROOT/hself"; new_home "$HSELF" -PE_TRACKED+=("$HSELF|self-src") +fm_test_track_procevent_home "$HSELF" pe_adapter "$HSELF" register selfann self-src -- /bin/echo "self announced" >/dev/null out=$(pe_adapter "$HSELF" start self-src 2>&1) assert_contains "$out" "autohandled: self-src" "the self-announcing adapter did not apply its own capture" @@ -480,7 +470,7 @@ rm -f "$HSELF/state/selfann-fail" pass "a self-announcing adapter applies quietly and still publishes what it could not apply" HTERM="$TMP_ROOT/hterm"; new_home "$HTERM" -PE_TRACKED+=("$HTERM|ends-src") +fm_test_track_procevent_home "$HTERM" pe_adapter "$HTERM" register endnow ends-src -- /bin/echo "terminal payload" >/dev/null out=$(pe_adapter "$HTERM" start ends-src) assert_contains "$out" "captured:" "a terminal result is still captured durably" @@ -504,7 +494,7 @@ assert_contains "$out" "published=0" "an acknowledged terminal result stops bein pass "an adapter-classified terminal result is captured once, announced, and retires its source automatically" HOPEN="$TMP_ROOT/hopen"; new_home "$HOPEN" -PE_TRACKED+=("$HOPEN|open-src") +fm_test_track_procevent_home "$HOPEN" pe_adapter "$HOPEN" register openended open-src -- /bin/echo "open payload" >/dev/null out=$(pe_adapter "$HOPEN" start open-src) assert_contains "$out" "captured:" "a result from an adapter with no terminal verdict is captured" @@ -514,12 +504,22 @@ pe_adapter "$HOPEN" retire open-src >/dev/null pass "a source stays armed unless its own adapter classifies the result terminal" HREPLACE="$TMP_ROOT/hreplace"; new_home "$HREPLACE" -PE_TRACKED+=("$HREPLACE|replace-src") +fm_test_track_procevent_home "$HREPLACE" OLD_TRIGGER="$TMP_ROOT/replace-old-trigger" -pe_adapter "$HREPLACE" register endnow replace-src -- "$BLOCKER" "$OLD_TRIGGER" "old terminal payload" >/dev/null +OLD_STARTED="$TMP_ROOT/replace-old-started" +REPLACE_BLOCKER="$TMP_ROOT/replace-blocker.sh" +cat > "$REPLACE_BLOCKER" <<'SH' +#!/usr/bin/env bash +printf 'started\n' > "$1" +shift +exec "$@" +SH +chmod +x "$REPLACE_BLOCKER" +pe_adapter "$HREPLACE" register endnow replace-src -- \ + "$REPLACE_BLOCKER" "$OLD_STARTED" "$BLOCKER" "$OLD_TRIGGER" "old terminal payload" >/dev/null pe_adapter "$HREPLACE" start replace-src > "$TMP_ROOT/replace-old.out" 2>&1 & replace_old_pid=$! -wait_for "$FM_PROCEVENT_CLAIM_ROOT/replace-src.claim" || fail "the old registration was never claimed" +wait_for "$OLD_STARTED" || fail "the old registration never started" pe_adapter "$HREPLACE" register openended replace-src -- /bin/echo "replacement payload" >/dev/null touch "$OLD_TRIGGER" wait "$replace_old_pid" || fail "the old terminal runner failed" @@ -537,7 +537,7 @@ pe_adapter "$HREPLACE" retire replace-src >/dev/null pass "terminal retirement preserves and releases a concurrently replaced registration" HRETFAIL="$TMP_ROOT/hretfail"; new_home "$HRETFAIL" -PE_TRACKED+=("$HRETFAIL|retire-fail-src") +fm_test_track_procevent_home "$HRETFAIL" FAIL_RM_BIN=$(fm_fakebin "$TMP_ROOT/retire-fail-bin") REAL_RM=$(command -v rm) export REAL_RM @@ -598,7 +598,7 @@ chmod +x "$LAVISH_BIN/lavish-axi" REVIEW_ART="$TMP_ROOT/review.html" printf '<h1>review</h1>\n' > "$REVIEW_ART" lavish_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$REVIEW_ART") -PE_TRACKED+=("$HLT|$lavish_id") +fm_test_track_procevent_home "$HLT" PATH="$LAVISH_BIN:$PATH" FM_HOME="$HLT" "$ROOT/bin/fm-procevent-lavish.sh" arm "$REVIEW_ART" >/dev/null for _ in $(seq 1 6); do PATH="$LAVISH_BIN:$PATH" pe "$HLT" reconcile >/dev/null @@ -638,7 +638,7 @@ chmod +x "$EMPTY_BIN/lavish-axi" QUIET_ART="$TMP_ROOT/quiet-board.html" printf '<h1>quiet</h1>\n' > "$QUIET_ART" quiet_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$QUIET_ART") -PE_TRACKED+=("$HEMPTY|$quiet_id") +fm_test_track_procevent_home "$HEMPTY" PATH="$EMPTY_BIN:$PATH" FM_HOME="$HEMPTY" \ "$ROOT/bin/fm-procevent-lavish.sh" arm "$QUIET_ART" >/dev/null quiet_out=$(PATH="$EMPTY_BIN:$PATH" pe "$HEMPTY" start "$quiet_id" 2>&1) @@ -684,7 +684,7 @@ chmod +x "$ANSWER_BIN/lavish-axi" ANSWER_ART="$TMP_ROOT/answered-board.html" printf '<h1>answered</h1>\n' > "$ANSWER_ART" answer_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$ANSWER_ART") -PE_TRACKED+=("$HANSWER|$answer_id") +fm_test_track_procevent_home "$HANSWER" PATH="$ANSWER_BIN:$PATH" FM_HOME="$HANSWER" \ "$ROOT/bin/fm-procevent-lavish.sh" arm "$ANSWER_ART" >/dev/null PATH="$ANSWER_BIN:$PATH" pe "$HANSWER" reconcile >/dev/null @@ -736,9 +736,27 @@ esac SH chmod +x "$LAVISH_SCRIPTED_BIN/lavish-axi" export LAVISH_COUNT LAVISH_SCRIPT + +DEFAULT_RATE_ART="$TMP_ROOT/default-rate-board.html" +printf '<h1>default rate</h1>\n' > "$DEFAULT_RATE_ART" +DEFAULT_RATE_COUNT="$TMP_ROOT/default-rate-count" +PATH="$LAVISH_SCRIPTED_BIN:$PATH" LAVISH_COUNT="$DEFAULT_RATE_COUNT" LAVISH_SCRIPT=interrupt \ + FM_LAVISH_POLL_RETRY_DELAY='' \ + "$ROOT/bin/fm-procevent-lavish.sh" poll "$DEFAULT_RATE_ART" >/dev/null 2>&1 & +DEFAULT_RATE_PID=$! +perl -MTime::HiRes=sleep -e 'sleep 6.2' +kill -TERM "$DEFAULT_RATE_PID" 2>/dev/null || true +wait "$DEFAULT_RATE_PID" 2>/dev/null || true +default_rate_count=$(cat "$DEFAULT_RATE_COUNT" 2>/dev/null || echo 0) +[ "$default_rate_count" -ge 2 ] \ + || fail "the default poll governor stopped an instantly returning source from making progress" +[ "$default_rate_count" -le 2 ] \ + || fail "the shipped poll governor allowed $default_rate_count iterations in 6.2 seconds" +pass "the shipped poll governor bounds an instantly returning source" + # A bounded test override keeps the retry policy's real bound under test without # making the suite wait out the production delay. -export FM_LAVISH_POLL_RETRY_DELAY=0 +export FM_LAVISH_POLL_RETRY_DELAY=1 # Two interruptions, then the captain's real feedback: the retries are silent and # only the feedback becomes a captured result and a check wake. @@ -746,7 +764,7 @@ HRETRY="$TMP_ROOT/hretry"; new_home "$HRETRY" RETRY_ART="$TMP_ROOT/retry-board.html" printf '<h1>retry</h1>\n' > "$RETRY_ART" retry_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$RETRY_ART") -PE_TRACKED+=("$HRETRY|$retry_id") +fm_test_track_procevent_home "$HRETRY" LAVISH_COUNT="$TMP_ROOT/retry-count"; LAVISH_SCRIPT="interrupt interrupt feedback" PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HRETRY" \ "$ROOT/bin/fm-procevent-lavish.sh" arm "$RETRY_ART" >/dev/null @@ -770,7 +788,7 @@ HEXH="$TMP_ROOT/hexh"; new_home "$HEXH" EXH_ART="$TMP_ROOT/exhaust-board.html" printf '<h1>exhaust</h1>\n' > "$EXH_ART" exh_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$EXH_ART") -PE_TRACKED+=("$HEXH|$exh_id") +fm_test_track_procevent_home "$HEXH" LAVISH_COUNT="$TMP_ROOT/exhaust-count"; LAVISH_SCRIPT="interrupt" PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HEXH" \ "$ROOT/bin/fm-procevent-lavish.sh" arm "$EXH_ART" >/dev/null @@ -793,7 +811,7 @@ HOTHER="$TMP_ROOT/hother"; new_home "$HOTHER" OTHER_ART="$TMP_ROOT/other-board.html" printf '<h1>other</h1>\n' > "$OTHER_ART" other_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$OTHER_ART") -PE_TRACKED+=("$HOTHER|$other_id") +fm_test_track_procevent_home "$HOTHER" LAVISH_COUNT="$TMP_ROOT/other-count"; LAVISH_SCRIPT="other-server-error" PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HOTHER" \ "$ROOT/bin/fm-procevent-lavish.sh" arm "$OTHER_ART" >/dev/null @@ -813,9 +831,9 @@ HNEAR="$TMP_ROOT/hnear"; new_home "$HNEAR" NEAR_ART="$TMP_ROOT/near-board.html" printf '<h1>near</h1>\n' > "$NEAR_ART" near_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$NEAR_ART") -PE_TRACKED+=("$HNEAR|$near_id") +fm_test_track_procevent_home "$HNEAR" LAVISH_COUNT="$TMP_ROOT/near-count"; LAVISH_SCRIPT="near-interrupt feedback" -PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HNEAR" FM_LAVISH_POLL_RETRY_DELAY=0 \ +PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HNEAR" FM_LAVISH_POLL_RETRY_DELAY=1 \ "$ROOT/bin/fm-procevent-lavish.sh" arm "$NEAR_ART" >/dev/null PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HNEAR" pe "$HNEAR" start "$near_id" >/dev/null [ "$(cat "$LAVISH_COUNT")" = 1 ] \ @@ -832,14 +850,14 @@ HINVALID="$TMP_ROOT/hinvalid"; new_home "$HINVALID" INVALID_ART="$TMP_ROOT/invalid-delay-board.html" printf '<h1>invalid delay</h1>\n' > "$INVALID_ART" invalid_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$INVALID_ART") -for invalid_delay in 61 invalid; do +for invalid_delay in 0 61 invalid; do invalid_status=0 invalid_out=$(PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HINVALID" \ FM_LAVISH_POLL_RETRY_DELAY="$invalid_delay" \ "$ROOT/bin/fm-procevent-lavish.sh" arm "$INVALID_ART" 2>&1) || invalid_status=$? [ "$invalid_status" -ne 0 ] \ || fail "arm accepted invalid retry delay: $invalid_delay" - assert_contains "$invalid_out" "must be whole seconds from 0 to 60" \ + assert_contains "$invalid_out" "must be whole seconds from 1 to 60" \ "arm explains the rejected retry delay" assert_absent "$HINVALID/state/procevent/$invalid_id.source" \ "arm publishes no source registration for an invalid retry delay" @@ -865,7 +883,7 @@ LAVISH_STREAM_RELEASE="$TMP_ROOT/stream-release" mkdir -p "$STREAM_TMPDIR" printf '<h1>stream</h1>\n' > "$STREAM_ART" stream_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$STREAM_ART") -PE_TRACKED+=("$HSTREAM|$stream_id") +fm_test_track_procevent_home "$HSTREAM" LAVISH_COUNT="$TMP_ROOT/stream-count"; LAVISH_SCRIPT="stream" PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HSTREAM" \ "$ROOT/bin/fm-procevent-lavish.sh" arm "$STREAM_ART" >/dev/null @@ -1108,32 +1126,19 @@ kill -0 -"$orphan_leader" 2>/dev/null || fail "fixture invalid: the owned child orphan_out=$(pe "$HG" reconcile) kill -0 -"$orphan_leader" 2>/dev/null \ - && fail "reconcile left the crashed generation's process group alive: $orphan_out" -sleep 0.5 -assert_absent "$ORPHAN_OVERLAP" "no replacement source starts while the crashed generation remains alive" -case "$orphan_out" in - *"started=1"*) - # The replacement is detached: it records its own claim and execs its source - # after reconcile has already returned, so both effects must be waited for - # rather than snapshotted behind the settle window above. - wait_for "$FM_PROCEVENT_CLAIM_ROOT/orphan-src.claim" \ - || fail "a replacement runner started without recording its own claim" - wait_for_lines "$ORPHAN_LOG" 2 \ - || fail "the replacement runner never started its source: $(cat "$ORPHAN_LOG")" - [ "$(wc -l < "$ORPHAN_LOG" | tr -d ' ')" = 2 ] \ - || fail "reconcile did not start exactly one replacement source: $(cat "$ORPHAN_LOG")" - ;; - *"started=0"*) - [ -e "$FM_PROCEVENT_CLAIM_ROOT/orphan-src.claim" ] \ - || fail "refusing to replace must preserve the claim for retry: $orphan_out" - [ "$(wc -l < "$ORPHAN_LOG" | tr -d ' ')" = 1 ] \ - || fail "reconcile started a source while refusing replacement: $(cat "$ORPHAN_LOG")" - ;; - *) fail "unexpected reconcile result for a crashed leader: $orphan_out" ;; -esac -: > "$ORPHAN_TRIGGER" + || fail "reconcile signalled an ambiguous leaderless process group: $orphan_out" +assert_contains "$orphan_out" "started=0" \ + "reconcile does not replace an ambiguous leaderless generation" +[ -e "$FM_PROCEVENT_CLAIM_ROOT/orphan-src.claim" ] \ + || fail "refusing ambiguous cleanup must preserve the claim" +[ "$(wc -l < "$ORPHAN_LOG" | tr -d ' ')" = 1 ] \ + || fail "reconcile started a source beside an ambiguous leaderless group" +assert_absent "$ORPHAN_OVERLAP" "no replacement source starts while the leaderless group remains" +kill -KILL -"$orphan_leader" 2>/dev/null || true +for _ in $(seq 1 50); do kill -0 -"$orphan_leader" 2>/dev/null || break; sleep 0.1; done +kill -0 -"$orphan_leader" 2>/dev/null && fail "could not clean up the leaderless fixture group" pe "$HG" retire orphan-src >/dev/null -pass "a crashed runner leader never lets a live owned group be reclaimed as stale" +pass "an ambiguous leaderless group is preserved without replacement" # Counterexample: a genuinely dead generation - no leader and no surviving # group - must still be reclaimable, or crash recovery would deadlock. @@ -1151,6 +1156,141 @@ wait_for "$DEAD_LOG" || fail "the replacement source never started for a truly d pe "$HG2" retire dead-gen-src >/dev/null pass "a truly dead generation with no surviving group is still safely reclaimed" +# --- a dead generation stays reclaimable when the state root cannot be +# revalidated ----------------------------------------------------------------- +# The reported wedge. Reclaiming a dead generation ran the claim's +# capture-reservation cleanup first, and that cleanup re-verifies the recorded +# state-root identity. Once that identity stopped matching, a claim naming a pid +# and a process group that were both provably gone could not be cleared: +# reconcile kept reporting a start while nothing ever attached, and retire +# refused with "cannot release source ownership". Reservation records are keyed +# by claim token and a replacement always claims a fresh one, so they can never +# collide with the generation that replaces them - they are hygiene, not an +# ownership invariant, and they must not veto an ownership move that the +# documented promise already grants. +# +# The claim below is the modern shape (it carries the state-root identity block +# a legacy claim does not have), which is why the existing dead-generation case +# above never reached this path. +HSR="$TMP_ROOT/hsr"; new_home "$HSR" +SR_TRIGGER="$TMP_ROOT/state-root-trigger" +SR_LOG="$TMP_ROOT/state-root-executions" +pe_register "$HSR" lavish state-root-src -- "$RACE_BLOCKER" "$SR_LOG" "$SR_TRIGGER" >/dev/null +pe "$HSR" reconcile >/dev/null +wait_for "$FM_PROCEVENT_CLAIM_ROOT/state-root-src.claim" \ + || fail "state-root fixture never claimed its source" +wait_for "$SR_LOG" || fail "state-root fixture source never started" +sr_leader=$(sed -n '2p' "$FM_PROCEVENT_CLAIM_ROOT/state-root-src.claim") +case "$sr_leader" in ''|*[!0-9]*) fail "could not read the state-root fixture leader pid" ;; esac +[ -n "$(sed -n '8p' "$FM_PROCEVENT_CLAIM_ROOT/state-root-src.claim")" ] \ + || fail "fixture invalid: the claim carries no state-root identity to invalidate" + +kill -KILL -"$sr_leader" 2>/dev/null || true +kill -KILL "$sr_leader" 2>/dev/null || true +for _ in $(seq 1 50); do kill -0 -"$sr_leader" 2>/dev/null || break; sleep 0.1; done +kill -0 "$sr_leader" 2>/dev/null && fail "the state-root fixture leader survived SIGKILL" +kill -0 -"$sr_leader" 2>/dev/null && fail "fixture invalid: the owned group outlived the whole generation" +# Drift the live state root away from what the claim recorded. +chmod 750 "$HSR/state" || fail "could not drift the state-root identity" + +sr_out=$(pe "$HSR" reconcile) +assert_contains "$sr_out" "started=1" "a dead generation was not reclaimed after the state root drifted: $sr_out" +# Reporting a start is not the same fact as listening: the previous behavior +# reported exactly this while the replacement silently failed to claim. +wait_for_lines "$SR_LOG" 2 \ + || fail "reconcile reported a start but no replacement source ever ran: $(cat "$SR_LOG")" +sr_new=$(sed -n '2p' "$FM_PROCEVENT_CLAIM_ROOT/state-root-src.claim") +[ "$sr_new" != "$sr_leader" ] || fail "the dead generation's claim was never replaced" +kill -0 "$sr_new" 2>/dev/null || fail "the replacement runner did not take ownership" +: > "$SR_TRIGGER" +pe "$HSR" retire state-root-src >/dev/null +pass "reconcile reclaims a dead generation whose state-root identity no longer matches" + +# Retire must release the same wedged claim rather than refusing forever. +HSR2="$TMP_ROOT/hsr2"; new_home "$HSR2" +pe_register "$HSR2" lavish wedged-src -- /bin/echo recovered >/dev/null +sr2_identity=$(bash -c '. "$1/bin/fm-pr-lib.sh"; fm_pr_file_identity "$2"' _ \ + "$ROOT" "$HSR2/state/procevent/wedged-src.source") \ + || fail "could not read the wedged fixture registration identity" +{ + printf '%s\n%s\nwedged-token\nwedged-identity\n' "$HSR2" 999999 + printf '%s\n%s\nactive\n' "$HSR2/state/procevent" "$sr2_identity" + # A state-root identity that names the right directory with the wrong inode: + # exactly what a claim recorded before its home was re-created looks like. + printf '%s\n%s\n%s\n%s\n%s\n' "$HSR2/state" 1 1 "$(id -u)" 755 +} > "$FM_PROCEVENT_CLAIM_ROOT/wedged-src.claim" +chmod 0600 "$FM_PROCEVENT_CLAIM_ROOT/wedged-src.claim" +wedged_out=$(pe "$HSR2" retire wedged-src 2>&1) \ + || fail "retire refused to release a claim whose whole generation is gone: $wedged_out" +assert_contains "$wedged_out" "retired: wedged-src" "retire did not report releasing the wedged source: $wedged_out" +assert_absent "$FM_PROCEVENT_CLAIM_ROOT/wedged-src.claim" "retire left the dead generation owning the source" +assert_absent "$HSR2/state/procevent/wedged-src.source" "retire left the wedged source registered" +pass "retire releases a dead generation's claim instead of refusing forever" + +# The guard is not weakened in the other direction: the same unrevalidatable +# state root must NOT let anything take a source away from a live generation. +HSR3="$TMP_ROOT/hsr3"; new_home "$HSR3" +SR3_TRIGGER="$TMP_ROOT/state-root-live-trigger" +SR3_LOG="$TMP_ROOT/state-root-live-executions" +pe_register "$HSR3" lavish live-drift-src -- "$RACE_BLOCKER" "$SR3_LOG" "$SR3_TRIGGER" >/dev/null +pe "$HSR3" reconcile >/dev/null +wait_for "$FM_PROCEVENT_CLAIM_ROOT/live-drift-src.claim" \ + || fail "live-drift fixture never claimed its source" +wait_for "$SR3_LOG" || fail "live-drift fixture source never started" +sr3_leader=$(sed -n '2p' "$FM_PROCEVENT_CLAIM_ROOT/live-drift-src.claim") +chmod 750 "$HSR3/state" || fail "could not drift the live owner's state-root identity" +sr3_out=$(pe "$HSR3" start live-drift-src) +assert_contains "$sr3_out" "already owned" "a live generation was displaced after its state root drifted: $sr3_out" +[ "$(sed -n '2p' "$FM_PROCEVENT_CLAIM_ROOT/live-drift-src.claim")" = "$sr3_leader" ] \ + || fail "the live generation's claim was replaced" +kill -0 "$sr3_leader" 2>/dev/null || fail "the live owner was killed by a reclaim attempt" +sr3_reconcile=$(pe "$HSR3" reconcile) +assert_contains "$sr3_reconcile" "started=0" "reconcile started a second poller beside a live owner: $sr3_reconcile" +[ "$(wc -l < "$SR3_LOG" | tr -d ' ')" = 1 ] \ + || fail "a second source ran beside the live owner: $(cat "$SR3_LOG")" +: > "$SR3_TRIGGER" +pe "$HSR3" retire live-drift-src >/dev/null +pass "a live generation is never reclaimed, drifted state root or not" + +HSR4="$TMP_ROOT/hsr4"; new_home "$HSR4" +SR4_TRIGGER="$TMP_ROOT/state-root-reused-trigger" +SR4_LOG="$TMP_ROOT/state-root-reused-executions" +pe_register "$HSR4" lavish reused-group-src -- "$RACE_BLOCKER" "$SR4_LOG" "$SR4_TRIGGER" >/dev/null +pe "$HSR4" reconcile >/dev/null +wait_for "$FM_PROCEVENT_CLAIM_ROOT/reused-group-src.claim" \ + || fail "reused-group fixture never claimed its source" +wait_for "$SR4_LOG" || fail "reused-group fixture source never started" +sr4_claim="$FM_PROCEVENT_CLAIM_ROOT/reused-group-src.claim" +sr4_leader=$(sed -n '2p' "$sr4_claim") +sr4_identity=$(sed -n '4p' "$sr4_claim") +kill -0 -"$sr4_leader" 2>/dev/null \ + || fail "fixture invalid: the reused-pid process group is not alive" +awk 'NR == 4 { print "different-live-process-identity"; next } { print }' \ + "$sr4_claim" > "$sr4_claim.tmp" && mv "$sr4_claim.tmp" "$sr4_claim" +chmod 0600 "$sr4_claim" +chmod 750 "$HSR4/state" || fail "could not drift the reused-group state root" +sr4_out=$(pe "$HSR4" reconcile) +sleep 0.5 +[ "$(wc -l < "$SR4_LOG" | tr -d ' ')" = 1 ] \ + || fail "reconcile started a replacement beside a reused pid's live group: $sr4_out" +[ "$(sed -n '2p' "$sr4_claim")" = "$sr4_leader" ] \ + || fail "reconcile replaced the reused-pid generation's claim" +set +e +sr4_retire=$(pe "$HSR4" retire reused-group-src 2>&1) +sr4_rc=$? +set -e +[ "$sr4_rc" -ne 0 ] || fail "retire released a claim whose process group survives: $sr4_retire" +[ -e "$sr4_claim" ] || fail "retire removed the reused-pid generation's claim" +[ -e "$HSR4/state/procevent/reused-group-src.source" ] \ + || fail "retire removed the reused-pid generation's registration" +awk -v identity="$sr4_identity" 'NR == 4 { print identity; next } { print }' \ + "$sr4_claim" > "$sr4_claim.tmp" && mv "$sr4_claim.tmp" "$sr4_claim" +chmod 0600 "$sr4_claim" +chmod 755 "$HSR4/state" +: > "$SR4_TRIGGER" +pe "$HSR4" retire reused-group-src >/dev/null +pass "a reused pid never makes its surviving process group reclaimable" + HJ="$TMP_ROOT/hj"; new_home "$HJ" TORN_TRIGGER="$TMP_ROOT/torn-trigger" pe_register "$HJ" lavish torn-src -- "$BLOCKER" "$TORN_TRIGGER" "torn" >/dev/null @@ -1215,7 +1355,7 @@ kill -0 "$innocent_pid" 2>/dev/null || fail "retirement signaled a PID whose ide kill "$innocent_pid" 2>/dev/null || true wait "$innocent_pid" 2>/dev/null || true assert_absent "$FM_PROCEVENT_CLAIM_ROOT/reused-src.claim" "retirement releases the exact reused-pid claim" -pass "PID reuse cannot signal an unrelated process" +pass "detected PID reuse is refused before signalling" HL="$TMP_ROOT/hl"; new_home "$HL" IDENTITY_TRIGGER="$TMP_ROOT/identity-trigger" @@ -1250,6 +1390,19 @@ wait_for "$FM_PROCEVENT_CLAIM_ROOT/sweep-one.claim" || fail "home sweep fixture wait_for "$FM_PROCEVENT_CLAIM_ROOT/sweep-two.claim" || fail "home sweep fixture two did not start" sweep_pid_one=$(sed -n '2p' "$FM_PROCEVENT_CLAIM_ROOT/sweep-one.claim") sweep_pid_two=$(sed -n '2p' "$FM_PROCEVENT_CLAIM_ROOT/sweep-two.claim") +# The claim-only case is an owned claim with no live runner, so build exactly +# that: kill the runner's group so it cannot run its own cleanup, confirm it is +# gone, and only then drop the registration. Deleting the registration out from +# under a LIVE runner no longer produces this case, because a superseded +# generation now observes the identity mismatch, self-retires, and releases its +# claim - so the sweep would race that exit and see one source or two depending +# on which won. +kill -KILL -"$sweep_pid_two" 2>/dev/null || true +for _ in $(seq 1 50); do kill -0 "$sweep_pid_two" 2>/dev/null || break; sleep 0.1; done +kill -0 "$sweep_pid_two" 2>/dev/null \ + && fail "the claim-only sweep fixture runner did not stop" +assert_present "$FM_PROCEVENT_CLAIM_ROOT/sweep-two.claim" \ + "a killed runner leaves its owned claim behind for the sweep" rm -f "$HM/state/procevent/sweep-two.source" out=$(pe "$HM" sweep-home --preflight) assert_contains "$out" "sweep preflight: ready" "home sweep preflight validates the full bounded snapshot" @@ -1377,6 +1530,64 @@ kill -0 "$noisy_child" 2>/dev/null && fail "TERM-resistant source child survived assert_absent "$staged" "retirement removes the tracked partial staging file" pass "live output stays bounded and retirement reaps the whole source group" +HPOST_TERM="$TMP_ROOT/post-term-reuse"; new_home "$HPOST_TERM" +POST_TERM_SOURCE="$TMP_ROOT/post-term-reuse-source.sh" +POST_TERM_PID="$TMP_ROOT/post-term-reuse.pid" +POST_TERM_MARKER="$TMP_ROOT/post-term-reuse.marker" +POST_TERM_COUNT="$TMP_ROOT/post-term-reuse.count" +cat > "$POST_TERM_SOURCE" <<'SH' +#!/usr/bin/env bash +trap '' TERM +printf '%s\n' "$$" > "$1" +while :; do sleep 1; done +SH +chmod +x "$POST_TERM_SOURCE" +POST_TERM_BIN=$(fm_fakebin "$TMP_ROOT/post-term-reuse-bin") +REAL_PS=$(command -v ps) || fail "the post-TERM reuse fixture requires ps" +cat > "$POST_TERM_BIN/ps" <<SH +#!/usr/bin/env bash +if [ -e "$POST_TERM_MARKER" ] && [ "\${1-}" = -p ] \ + && [ "\${3-}" = -o ] && [ "\${4-}" = lstart= ]; then + count=0 + [ ! -f "$POST_TERM_COUNT" ] || count=\$(cat "$POST_TERM_COUNT") + count=\$((count + 1)) + printf '%s\n' "\$count" > "$POST_TERM_COUNT" + if [ "\$count" -gt 1 ]; then + printf 'post-TERM reused identity\n' + exit 0 + fi +fi +exec "$REAL_PS" "\$@" +SH +chmod +x "$POST_TERM_BIN/ps" +pe_register "$HPOST_TERM" lavish post-term-src -- \ + "$POST_TERM_SOURCE" "$POST_TERM_PID" >/dev/null +FM_PROCEVENT_OWNER_CHECK_SECONDS=5 pe "$HPOST_TERM" reconcile >/dev/null +wait_for "$POST_TERM_PID" || fail "the post-TERM reuse fixture did not start" +wait_for "$FM_PROCEVENT_CLAIM_ROOT/post-term-src.claim" \ + || fail "the post-TERM reuse fixture did not claim its source" +POST_TERM_RUNNER=$(sed -n '2p' "$FM_PROCEVENT_CLAIM_ROOT/post-term-src.claim") +touch "$POST_TERM_MARKER" +post_term_status=0 +PATH="$POST_TERM_BIN:$PATH" FM_PROC_ROOT_OVERRIDE="$TMP_ROOT/no-post-term-proc" \ + pe "$HPOST_TERM" retire post-term-src >/dev/null 2>&1 || post_term_status=$? +[ "$post_term_status" -ne 0 ] || fail "retirement escalated after runner identity became ambiguous" +# Which ambiguity the escalation meets here is platform-dependent, so this +# asserts the invariant both forms share rather than one form's internals. +# Where the runner leader keeps waiting on its TERM-ignoring source child the +# post-TERM check sees a live leader whose identity no longer matches, and +# where the leader dies promptly it sees a leaderless group carrying the same +# numeric id; fm_procevent_pid_state reaches the second verdict without +# consulting process identity at all, so counting identity lookups pins a +# timing- and platform-dependent internal rather than the behavior. +kill -0 -"$POST_TERM_RUNNER" 2>/dev/null \ + || fail "an ambiguous reused-PID group was killed during escalation" +kill -KILL -"$POST_TERM_RUNNER" 2>/dev/null || true +for _ in $(seq 1 50); do kill -0 -"$POST_TERM_RUNNER" 2>/dev/null || break; sleep 0.1; done +kill -0 -"$POST_TERM_RUNNER" 2>/dev/null && fail "could not clean up the post-TERM fixture group" +pe "$HPOST_TERM" retire post-term-src >/dev/null +pass "cleanup aborts escalation after runner identity becomes ambiguous" + HBAD="$TMP_ROOT/hbad"; new_home "$HBAD" pe_register "$HBAD" lavish bad-limit -- /bin/true >/dev/null bad_limit_status=0 @@ -1938,4 +2149,527 @@ assert_not_contains "$runner_help" "exactly-once" \ "the runner's help claims no exactly-once delivery" pass "the published interfaces state the loss limitation and claim no lossless delivery" +# --- launch pacing and guard startup ---------------------------------------- + +FAST_SOURCE="$TMP_ROOT/fast-source.sh" +cat > "$FAST_SOURCE" <<'SH' +#!/usr/bin/env bash +perl -MTime::HiRes=time -e 'printf "%.6f\n", time' >> "$1" +exit 1 +SH +chmod +x "$FAST_SOURCE" + +STORM_SOURCE="$TMP_ROOT/storm-source.sh" +cat > "$STORM_SOURCE" <<'SH' +#!/usr/bin/env bash +perl -MTime::HiRes=time -e 'printf "%.6f\n", time' >> "$1" +FM_HOME="$2" perl -MPOSIX=setsid -e ' + my @command = @ARGV; + defined(my $pid = fork) or exit 1; + exit 0 if $pid; + setsid() >= 0 or exit 1; + open STDIN, "<", "/dev/null" or exit 1; + open STDOUT, ">", "/dev/null" or exit 1; + open STDERR, ">", "/dev/null" or exit 1; + select undef, undef, undef, 0.2; + exec @command; +' "$3/bin/fm-procevent.sh" reconcile +exit 1 +SH +chmod +x "$STORM_SOURCE" + +HFLOOR="$TMP_ROOT/launch-floor"; new_home "$HFLOOR" +fm_test_track_procevent_home "$HFLOOR" +pe_register "$HFLOOR" lavish floor-src -- \ + "$STORM_SOURCE" "$TMP_ROOT/launch-times" "$HFLOOR" "$ROOT" +FM_PROCEVENT_OWNER_LEASE_SECONDS=4 FM_PROCEVENT_OWNER_CHECK_SECONDS=1 \ + FM_PROCEVENT_LAUNCH_FLOOR_SECONDS=1 pe "$HFLOOR" reconcile >/dev/null +floor_deadline=$((SECONDS + 12)) +while :; do + floor_count=0 + [ ! -f "$TMP_ROOT/launch-times" ] \ + || floor_count=$(wc -l < "$TMP_ROOT/launch-times" | tr -d ' ') + [ "$floor_count" -ge 3 ] && break + [ "$SECONDS" -lt "$floor_deadline" ] \ + || fail "the orphan-storm fixture did not relaunch its source command" + sleep 0.1 +done +launch_count=$(wc -l < "$TMP_ROOT/launch-times" | tr -d ' ') +launch_span=$(perl -e '@t=<>; printf "%.3f", $t[-1] - $t[0]' "$TMP_ROOT/launch-times") +perl -e 'exit($ARGV[0] >= ($ARGV[1] - 1) * 0.8 ? 0 : 1)' "$launch_span" "$launch_count" \ + || fail "an orphaned source launched $launch_count times in only ${launch_span}s" +[ "$launch_count" -le 6 ] \ + || fail "an orphaned source stormed $launch_count launches during its owner-dead grace window" +pass "an orphaned source command obeys the launch floor during its grace window" + +HPACE="$TMP_ROOT/registration-pacing"; new_home "$HPACE" +fm_test_track_procevent_home "$HPACE" +PACE_LOG="$TMP_ROOT/registration-pacing.log" +pe_register "$HPACE" lavish pace-src -- "$FAST_SOURCE" "$PACE_LOG" >/dev/null +FM_PROCEVENT_LAUNCH_FLOOR_SECONDS=3600 pe "$HPACE" start pace-src >/dev/null +pe "$HPACE" retire pace-src >/dev/null +pe_register "$HPACE" lavish pace-src -- "$FAST_SOURCE" "$PACE_LOG" >/dev/null +FM_PROCEVENT_LAUNCH_FLOOR_SECONDS=3600 pe "$HPACE" start pace-src > "$TMP_ROOT/replacement-pacing.out" 2>&1 & +PACE_START_PID=$! +pace_deadline=$((SECONDS + 4)) +while kill -0 "$PACE_START_PID" 2>/dev/null; do + if [ "$SECONDS" -ge "$pace_deadline" ]; then + pe "$HPACE" retire pace-src >/dev/null 2>&1 || true + wait "$PACE_START_PID" 2>/dev/null || true + fail "a replacement registration inherited the prior launch floor" + fi + sleep 0.1 +done +wait "$PACE_START_PID" || fail "the replacement registration failed" +[ "$(wc -l < "$PACE_LOG" | tr -d ' ')" = 2 ] \ + || fail "a replacement registration did not launch immediately" +PACE_STAMPS=$(find "$HPACE/state/procevent" -maxdepth 1 -type f \ + -name 'pace-src.*.last-launch' | wc -l | tr -d ' ') +[ "$PACE_STAMPS" = 1 ] || fail "replacement registrations accumulated stale pacing state" +pass "a replacement registration starts with one fresh launch floor" + +HPACE_RACE="$TMP_ROOT/registration-pacing-race"; new_home "$HPACE_RACE" +fm_test_track_procevent_home "$HPACE_RACE" +PACE_RACE_LOG="$TMP_ROOT/registration-pacing-race.log" +pe_register "$HPACE_RACE" lavish pace-race-src -- "$FAST_SOURCE" "$PACE_RACE_LOG" >/dev/null +FM_PROCEVENT_LAUNCH_FLOOR_SECONDS=3 pe "$HPACE_RACE" start pace-race-src >/dev/null +FM_PROCEVENT_LAUNCH_FLOOR_SECONDS=3 \ + pe "$HPACE_RACE" start pace-race-src > "$TMP_ROOT/registration-pacing-race.out" 2>&1 & +PACE_RACE_PID=$! +wait_for "$FM_PROCEVENT_CLAIM_ROOT/pace-race-src.claim" \ + || fail "the superseded pacing fixture did not claim its registration" +[ "$(wc -l < "$PACE_RACE_LOG" | tr -d ' ')" = 1 ] \ + || fail "the superseded pacing fixture was not waiting on its launch floor" +pe_register "$HPACE_RACE" lavish pace-race-src -- "$FAST_SOURCE" "$PACE_RACE_LOG" >/dev/null +wait "$PACE_RACE_PID" || fail "the superseded paced runner failed" +[ "$(wc -l < "$PACE_RACE_LOG" | tr -d ' ')" = 1 ] \ + || fail "the superseded paced runner invoked its stale command" +# The runner marker is written before the launch floor is waited on, and a home +# sweep counts a marker with no owned claim as a preflight failure. A superseded +# generation that exits without clearing its marker therefore makes the whole +# home refuse to sweep, so assert the marker is gone and the sweep still runs. +assert_absent "$HPACE_RACE/state/procevent/pace-race-src.runner" \ + "a superseded paced runner leaves no runner marker behind" +FM_PROCEVENT_LAUNCH_FLOOR_SECONDS=3 pe "$HPACE_RACE" start pace-race-src >/dev/null +PACE_RACE_STAMPS=$(find "$HPACE_RACE/state/procevent" -maxdepth 1 -type f \ + -name 'pace-race-src.*.last-launch' | wc -l | tr -d ' ') +[ "$PACE_RACE_STAMPS" = 1 ] \ + || fail "a superseded sleeping runner recreated stale pacing state" +pass "a superseded sleeping runner cannot recreate stale pacing state" + +HCOMMIT="$TMP_ROOT/registration-commit"; new_home "$HCOMMIT" +fm_test_track_procevent_home "$HCOMMIT" +COMMIT_LOG="$TMP_ROOT/registration-commit.log" +mkdir -p "$HCOMMIT/state/procevent/commit-src.1-2.last-launch" +pe_register "$HCOMMIT" lavish commit-src -- "$FAST_SOURCE" "$COMMIT_LOG" >/dev/null \ + || fail "post-commit pacing cleanup made registration report failure" +FM_PROCEVENT_LAUNCH_FLOOR_SECONDS=1 pe "$HCOMMIT" start commit-src >/dev/null \ + || fail "a successfully published registration was not executable" +[ "$(wc -l < "$COMMIT_LOG" | tr -d ' ')" = 1 ] \ + || fail "the committed registration did not invoke its source" +pass "post-commit pacing cleanup cannot veto registration publication" + +HROLLBACK="$TMP_ROOT/rollback-pacing"; new_home "$HROLLBACK" +fm_test_track_procevent_home "$HROLLBACK" +ROLLBACK_LOG="$TMP_ROOT/rollback-pacing.log" +pe_register "$HROLLBACK" lavish rollback-src -- "$FAST_SOURCE" "$ROLLBACK_LOG" >/dev/null +FM_PROCEVENT_LAUNCH_FLOOR_SECONDS=1 pe "$HROLLBACK" start rollback-src >/dev/null +ROLLBACK_STAMP= +for candidate in "$HROLLBACK/state/procevent"/rollback-src.*.last-launch; do + [ -f "$candidate" ] && ROLLBACK_STAMP=$candidate +done +[ -n "$ROLLBACK_STAMP" ] || fail "the first launch did not persist its pacing state" +printf '%s\n' "$(( $(date +%s) + 3600 ))" > "$ROLLBACK_STAMP" +FM_PROCEVENT_LAUNCH_FLOOR_SECONDS=3600 pe "$HROLLBACK" start rollback-src > "$TMP_ROOT/rollback.out" 2>&1 & +ROLLBACK_START_PID=$! +rollback_deadline=$((SECONDS + 4)) +while kill -0 "$ROLLBACK_START_PID" 2>/dev/null; do + if [ "$SECONDS" -ge "$rollback_deadline" ]; then + pe "$HROLLBACK" retire rollback-src >/dev/null 2>&1 || true + wait "$ROLLBACK_START_PID" 2>/dev/null || true + fail "a pre-reboot monotonic stamp delayed the first launch" + fi + sleep 0.1 +done +wait "$ROLLBACK_START_PID" || fail "the rollback-paced source failed" +[ "$(wc -l < "$ROLLBACK_LOG" | tr -d ' ')" = 2 ] \ + || fail "the rollback-paced source did not invoke twice" +pass "a pre-reboot monotonic stamp is treated as expired" + +storm_deadline=$((SECONDS + 15)) +while :; do + storm_before=$(wc -l < "$TMP_ROOT/launch-times" | tr -d ' ') + sleep 2 + storm_after=$(wc -l < "$TMP_ROOT/launch-times" | tr -d ' ') + [ "$storm_before" = "$storm_after" ] && break + [ "$SECONDS" -lt "$storm_deadline" ] \ + || fail "an orphaned self-relaunching source survived its expired owner lease" +done +pass "an expired owner lease stops a self-relaunching source generation" + +HRECREATED="$TMP_ROOT/recreated-owner"; new_home "$HRECREATED" +fm_test_track_procevent_home "$HRECREATED" +RECREATED_TRIGGER="$TMP_ROOT/recreated-owner.trigger" +pe_register "$HRECREATED" lavish recreated-src -- \ + "$BLOCKER" "$RECREATED_TRIGGER" "recreated payload" >/dev/null +FM_PROCEVENT_OWNER_LEASE_SECONDS=30 FM_PROCEVENT_OWNER_CHECK_SECONDS=1 \ + pe "$HRECREATED" reconcile >/dev/null +wait_for "$HRECREATED/state/procevent/recreated-src.runner" \ + || fail "the recreated-path fixture never launched its source" +RECREATED_RUNNER_PID=$(cat "$HRECREATED/state/procevent/recreated-src.runner") +mv "$HRECREATED/state" "$TMP_ROOT/recreated-owner-old-state" +pe_register "$HRECREATED" lavish recreated-src -- \ + "$BLOCKER" "$RECREATED_TRIGGER" "replacement payload" >/dev/null +FM_PROCEVENT_OWNER_LEASE_SECONDS=30 FM_PROCEVENT_OWNER_CHECK_SECONDS=1 \ + pe "$HRECREATED" reconcile >/dev/null +recreated_deadline=$((SECONDS + 8)) +while kill -0 "$RECREATED_RUNNER_PID" 2>/dev/null; do + [ "$SECONDS" -lt "$recreated_deadline" ] \ + || fail "a fresh lease at a recreated state path preserved the old runner" + sleep 0.1 +done +pass "a recreated state path does not preserve the old runner" + +HGUARDFAIL="$TMP_ROOT/guard-failure"; new_home "$HGUARDFAIL" +fm_test_track_procevent_home "$HGUARDFAIL" +pe_register "$HGUARDFAIL" lavish guard-fail-src -- "$FAST_SOURCE" "$TMP_ROOT/unguarded-launches" +guard_fail_status=0 +guard_fail_out=$(FM_PROCEVENT_OWNER_LEASE_SECONDS=invalid \ + pe "$HGUARDFAIL" start guard-fail-src 2>&1) || guard_fail_status=$? +[ "$guard_fail_status" -ne 0 ] || fail "a runner continued after its owner guard failed to initialize" +assert_contains "$guard_fail_out" "cannot start the runner's owner guard" \ + "guard initialization failure is reported at the runner boundary" +assert_absent "$TMP_ROOT/unguarded-launches" \ + "a source command ran without a successfully initialized owner guard" +pass "a runner fails closed when its owner guard cannot initialize" + +HATTACHED="$TMP_ROOT/attached-owner"; new_home "$HATTACHED" +fm_test_track_procevent_home "$HATTACHED" +ATTACHED_TRIGGER="$TMP_ROOT/attached.trigger" +pe_register "$HATTACHED" lavish attached-src -- "$BLOCKER" "$ATTACHED_TRIGGER" "attached payload" +FM_PROCEVENT_OWNER_LEASE_SECONDS=1 FM_PROCEVENT_OWNER_CHECK_SECONDS=1 \ + pe "$HATTACHED" start attached-src > "$TMP_ROOT/attached.out" 2>&1 & +ATTACHED_START_PID=$! +wait_for "$HATTACHED/state/procevent/attached-src.runner" \ + || fail "the attached start never launched its source" +sleep 4 +kill -0 "$ATTACHED_START_PID" 2>/dev/null \ + || fail "a foreground start lost its owner lease while its caller remained attached" +touch "$ATTACHED_TRIGGER" +wait "$ATTACHED_START_PID" || fail "the attached start did not complete after its source returned" +assert_contains "$(cat "$TMP_ROOT/attached.out")" "captured:" \ + "the attached source result was not captured" +pass "a foreground start refreshes its lease while its caller remains attached" + +HCLOCK="$TMP_ROOT/lease-clock"; new_home "$HCLOCK" +fm_test_track_procevent_home "$HCLOCK" +CLOCK_TRIGGER="$TMP_ROOT/lease-clock.trigger" +CLOCK_STATE="$TMP_ROOT/lease-clock-state" +CLOCK_BIN=$(fm_fakebin "$TMP_ROOT/lease-clock-bin") +REAL_DATE=$(command -v date) || fail "the lease clock fixture requires date" +cat > "$CLOCK_BIN/date" <<SH +#!/usr/bin/env bash +if [ "\${1-}" = +%s ]; then + while ! mkdir "$CLOCK_STATE.lock" 2>/dev/null; do sleep 0.01; done + value=0 + [ ! -f "$CLOCK_STATE" ] || value=\$(cat "$CLOCK_STATE") + value=\$((value + 10000)) + printf '%s\n' "\$value" > "$CLOCK_STATE" + rmdir "$CLOCK_STATE.lock" + printf '%s\n' "\$value" + exit 0 +fi +exec "$REAL_DATE" "\$@" +SH +chmod +x "$CLOCK_BIN/date" +pe_register "$HCLOCK" lavish lease-clock-src -- \ + "$BLOCKER" "$CLOCK_TRIGGER" "clock payload" >/dev/null +PATH="$CLOCK_BIN:$PATH" FM_PROCEVENT_OWNER_LEASE_SECONDS=1 FM_PROCEVENT_OWNER_CHECK_SECONDS=1 \ + pe "$HCLOCK" start lease-clock-src > "$TMP_ROOT/lease-clock.out" 2>&1 & +CLOCK_START_PID=$! +wait_for "$HCLOCK/state/procevent/lease-clock-src.runner" \ + || fail "the clock-shift fixture never launched its source" +sleep 4 +kill -0 "$CLOCK_START_PID" 2>/dev/null \ + || fail "wall-clock corrections expired a live foreground owner" +touch "$CLOCK_TRIGGER" +wait "$CLOCK_START_PID" || fail "the clock-shift fixture did not complete" +pass "wall-clock corrections do not alter owner lease age" + +HDETACHED="$TMP_ROOT/detached-attached-owner"; new_home "$HDETACHED" +fm_test_track_procevent_home "$HDETACHED" +DETACHED_TRIGGER="$TMP_ROOT/detached-attached.trigger" +pe_register "$HDETACHED" lavish detached-attached-src -- \ + "$BLOCKER" "$DETACHED_TRIGGER" "detached attached payload" +FM_PROCEVENT_OWNER_LEASE_SECONDS=1 FM_PROCEVENT_OWNER_CHECK_SECONDS=1 FM_HOME="$HDETACHED" \ + perl -MPOSIX=setsid -e 'setsid() >= 0 or exit 1; exec @ARGV' \ + "$ROOT/bin/fm-procevent.sh" start detached-attached-src \ + > "$TMP_ROOT/detached-attached.out" 2>&1 & +DETACHED_START_PID=$! +wait_for "$HDETACHED/state/procevent/detached-attached-src.runner" \ + || fail "the detachable foreground start never launched its source" +DETACHED_RUNNER_PID=$(cat "$HDETACHED/state/procevent/detached-attached-src.runner") +kill "$DETACHED_START_PID" +wait "$DETACHED_START_PID" 2>/dev/null || true +detached_deadline=$((SECONDS + 8)) +while kill -0 "$DETACHED_RUNNER_PID" 2>/dev/null; do + if [ "$SECONDS" -ge "$detached_deadline" ]; then + pe "$HDETACHED" retire detached-attached-src >/dev/null 2>&1 || true + fail "an orphaned attached-start keeper preserved its owner's lease" + fi + sleep 0.1 +done +pass "an attached-start keeper stops refreshing after its parent exits" + +HREUSED_GROUP="$TMP_ROOT/reused-runner-group"; new_home "$HREUSED_GROUP" +fm_test_track_procevent_home "$HREUSED_GROUP" +REUSED_GROUP_MARKER="$TMP_ROOT/reused-runner-group.marker" +REUSED_GROUP_TRIGGER="$TMP_ROOT/reused-runner-group.trigger" +REUSED_GROUP_BIN=$(fm_fakebin "$TMP_ROOT/reused-runner-group-bin") +REAL_PS=$(command -v ps) || fail "the reused-group fixture requires ps" +cat > "$REUSED_GROUP_BIN/ps" <<SH +#!/usr/bin/env bash +if [ -e "$REUSED_GROUP_MARKER" ] && [ "\${1-}" = -p ] \ + && [ "\${3-}" = -o ] && [ "\${4-}" = lstart= ]; then + printf 'reused runner identity\n' + exit 0 +fi +exec "$REAL_PS" "\$@" +SH +chmod +x "$REUSED_GROUP_BIN/ps" +pe_register "$HREUSED_GROUP" lavish reused-runner-group-src -- \ + "$BLOCKER" "$REUSED_GROUP_TRIGGER" "reused group payload" >/dev/null +PATH="$REUSED_GROUP_BIN:$PATH" FM_PROC_ROOT_OVERRIDE="$TMP_ROOT/no-reused-group-proc" \ + FM_PROCEVENT_OWNER_LEASE_SECONDS=30 FM_PROCEVENT_OWNER_CHECK_SECONDS=1 \ + pe "$HREUSED_GROUP" reconcile >/dev/null +wait_for "$HREUSED_GROUP/state/procevent/reused-runner-group-src.runner" \ + || fail "the reused-group fixture runner did not start" +REUSED_GROUP_RUNNER=$(cat "$HREUSED_GROUP/state/procevent/reused-runner-group-src.runner") +touch "$REUSED_GROUP_MARKER" +sleep 3 +kill -0 -"$REUSED_GROUP_RUNNER" 2>/dev/null \ + || fail "the guard killed a process group after its runner identity became ambiguous" +# Retiring here must read identity from the source this runner was recorded +# under, so the override stays in place: without it the read falls back to +# /proc where that exists, which is a different source than the recorded ps +# identity, and the guard would refuse this retirement on Linux while accepting +# it on macOS. Clearing the marker restores the matching identity, so this also +# asserts the complementary guarantee - once the ambiguity is gone, retirement +# reaps the whole group rather than leaving it behind. +rm -f "$REUSED_GROUP_MARKER" +PATH="$REUSED_GROUP_BIN:$PATH" FM_PROC_ROOT_OVERRIDE="$TMP_ROOT/no-reused-group-proc" \ + pe "$HREUSED_GROUP" retire reused-runner-group-src >/dev/null +for _ in $(seq 1 50); do kill -0 -"$REUSED_GROUP_RUNNER" 2>/dev/null || break; sleep 0.1; done +kill -0 -"$REUSED_GROUP_RUNNER" 2>/dev/null \ + && fail "retirement left the group alive once runner identity was unambiguous" +pass "a detected ambiguous reused-PID group is not signalled" + +# --- an accidentally orphaned runner is bounded by its owner ---------------- +# +# Reproduces the shape that wedged a host: a listener detached into its own +# process group, reparented to init when its session ended, and left running for +# a day with its blocking child - and everything that child spawned - still +# executing. The cost was not the runner itself but the process churn under it, +# which is why this asserts the whole descendant tree stops, not just the leader. +# +# Scope is asserted alongside it, in the same run and against the same stub: a +# home whose session is still there keeps its runner. Reaping that keyed on the +# script or process name instead of the owning session would take both. + +ORPHAN_STUB="$TMP_ROOT/orphan-stub.sh" +cat > "$ORPHAN_STUB" <<'SH' +#!/usr/bin/env bash +# A blocking source whose child keeps spawning processes, which is what a poll +# stub waiting on a trigger file actually does. The spawn rate is what turned a +# leftover listener into a host-wide storm, so the tick log is the evidence that +# the storm stopped and not merely that one pid went away. +marker=$1 +( while [ "$SECONDS" -lt "${FM_TEST_STUB_MAX_BLOCK_SECONDS:-120}" ]; do + printf 'tick\n' >> "$marker.ticks" + sleep 0.1 + done ) & +printf '%s\n' "$!" > "$marker.descendant" +while [ ! -e "$marker.trigger" ]; do + [ "$SECONDS" -lt "${FM_TEST_STUB_MAX_BLOCK_SECONDS:-120}" ] || exit 75 + sleep 0.1 +done +printf 'orphan payload\n' +SH +chmod +x "$ORPHAN_STUB" + +# The same shape without the spawn churn, for the home that exercises explicit +# retirement rather than the storm. Retirement refuses instead of signalling +# when it cannot confirm the runner's identity, that identity is read through +# `ps`, and the churning stub above starves that read often enough to make a +# single retirement attempt a race. The storm itself is already covered against +# the churning stub by the owner-loss reaping above, which asserts the tick log +# stops, so this home only needs a reparented listener holding a real +# descendant in its group. +QUIET_STUB="$TMP_ROOT/quiet-stub.sh" +cat > "$QUIET_STUB" <<'SH' +#!/usr/bin/env bash +marker=$1 +( sleep "${FM_TEST_STUB_MAX_BLOCK_SECONDS:-120}" ) & +printf '%s\n' "$!" > "$marker.descendant" +while [ ! -e "$marker.trigger" ]; do + [ "$SECONDS" -lt "${FM_TEST_STUB_MAX_BLOCK_SECONDS:-120}" ] || exit 75 + sleep 0.1 +done +printf 'orphan payload\n' +SH +chmod +x "$QUIET_STUB" + +# Short enough to observe, and driven through the same environment a real home +# uses, so the bound under test is the shipped one rather than a test-only path. +orphan_pe() { # <home> <command...> + local home=$1 + shift + FM_PROCEVENT_OWNER_LEASE_SECONDS=2 FM_PROCEVENT_OWNER_CHECK_SECONDS=1 \ + FM_HOME="$home" "$ROOT/bin/fm-procevent.sh" "$@" +} + +wait_gone() { # <pid-or-group-spec> [tries] + local spec=$1 n=${2:-160} + for _ in $(seq 1 "$n"); do + kill -0 "$spec" 2>/dev/null || return 0 + sleep 0.1 + done + return 1 +} + +HORPHAN="$TMP_ROOT/orphan-dead-owner"; new_home "$HORPHAN" +fm_test_track_procevent_home "$HORPHAN" +HKEEP="$TMP_ROOT/orphan-live-owner"; new_home "$HKEEP" +fm_test_track_procevent_home "$HKEEP" +orphan_pe "$HORPHAN" register lavish orphan-src -- "$ORPHAN_STUB" "$TMP_ROOT/orphan-dead" >/dev/null +orphan_pe "$HKEEP" register lavish keep-src -- "$QUIET_STUB" "$TMP_ROOT/orphan-live" >/dev/null +orphan_pe "$HORPHAN" reconcile >/dev/null +orphan_pe "$HKEEP" reconcile >/dev/null + +wait_for "$HORPHAN/state/procevent/orphan-src.runner" \ + || fail "the dead-owner listener never recorded its runner" +wait_for "$HKEEP/state/procevent/keep-src.runner" \ + || fail "the live-owner listener never recorded its runner" +wait_for "$TMP_ROOT/orphan-dead.descendant" \ + || fail "the dead-owner listener's child never spawned its own descendant" +ORPHAN_PID=$(cat "$HORPHAN/state/procevent/orphan-src.runner") +KEEP_PID=$(cat "$HKEEP/state/procevent/keep-src.runner") +ORPHAN_DESCENDANT=$(cat "$TMP_ROOT/orphan-dead.descendant") + +# The reproduction condition itself: the listener is already an orphan in the +# kernel's sense before anything is asserted about reaping it. +orphan_ppid=$(ps -o ppid= -p "$ORPHAN_PID" 2>/dev/null | tr -d '[:space:]') +[ "$orphan_ppid" = 1 ] \ + || fail "the listener under test was not reparented away from its session (ppid $orphan_ppid)" +kill -0 -"$ORPHAN_PID" 2>/dev/null \ + || fail "the listener's process group was not running" +kill -0 "$ORPHAN_DESCENDANT" 2>/dev/null \ + || fail "the listener's descendant was not running" +pass "a detached listener starts reparented, with a live descendant tree under it" + +# Only the second home's session stays present, on the same short bound, so the +# owning session is the single difference between the two listeners. +keep_owner_present() { orphan_pe "$HKEEP" reconcile >/dev/null 2>&1 || true; sleep 0.25; } + +deadline=$((SECONDS + 40)) +while kill -0 -"$ORPHAN_PID" 2>/dev/null; do + [ "$SECONDS" -lt "$deadline" ] \ + || fail "a listener whose owning session was gone kept its process group running" + keep_owner_present +done +deadline=$((SECONDS + 20)) +while kill -0 "$ORPHAN_DESCENDANT" 2>/dev/null; do + [ "$SECONDS" -lt "$deadline" ] \ + || fail "a listener whose owning session was gone left a descendant running" + keep_owner_present +done +pass "a listener whose owning session is gone stops itself and its whole process group" + +keep_owner_present +before=$(wc -l < "$TMP_ROOT/orphan-dead.ticks" | tr -d ' ') +deadline=$((SECONDS + 2)) +while [ "$SECONDS" -lt "$deadline" ]; do keep_owner_present; done +after=$(wc -l < "$TMP_ROOT/orphan-dead.ticks" | tr -d ' ') +[ "$before" = "$after" ] \ + || fail "the reaped listener's descendant kept spawning processes ($before then $after)" +pass "reaping the listener stops the process churn under it" + +keep_owner_present +kill -0 -"$KEEP_PID" 2>/dev/null \ + || fail "an identical listener in a home whose session is still there was reaped too" +pass "an identical listener in a home whose session is still there is untouched" + +# Retirement remains the explicit path, and it must reach a listener that has +# already reparented, along with everything under it. +wait_for "$TMP_ROOT/orphan-live.descendant" \ + || fail "the live-owner listener's child never spawned its own descendant" +KEEP_DESCENDANT=$(cat "$TMP_ROOT/orphan-live.descendant") +keep_owner_present +orphan_pe "$HKEEP" retire keep-src >/dev/null +wait_gone "-$KEEP_PID" \ + || fail "retiring a source left its reparented listener's process group running" +wait_gone "$KEEP_DESCENDANT" \ + || fail "retiring a source left a descendant of its listener running" +pass "retiring a source reaps its reparented listener and every descendant under it" + +# --- an expired runner's guard retries unproved cleanup --------------------- +# +# A stop the guard cannot PROVE must not end the guard. A descendant still +# finishing uninterruptible work outlives even the group KILL, and a guard that +# gave up after one attempt would walk away from a still-running expired runner. +# +# The unprovable attempt is injected through the signal the real path actually +# reads: `ps` answers ONE process-group query for the runner with a group it +# does not lead, which is exactly how a stop that cannot be proved is reported. +# Every other `ps` call, and every later one, is the real command. + +RETRY_HOME="$TMP_ROOT/stop-retry"; new_home "$RETRY_HOME" +fm_test_track_procevent_home "$RETRY_HOME" +RETRY_STATE="$TMP_ROOT/stop-retry-state"; mkdir -p "$RETRY_STATE" +RETRY_BIN=$(fm_fakebin "$TMP_ROOT/stop-retry-bin") +REAL_PS=$(command -v ps) || fail "this host has no ps to build the retry fixture on" +cat > "$RETRY_BIN/ps" <<SH +#!/usr/bin/env bash +if [ "\$1" = -o ] && [ "\$2" = "pgid=" ] && [ "\$3" = -p ] \\ + && [ -s "\$STOP_RETRY_STATE/target" ] \\ + && [ "\$4" = "\$(cat "\$STOP_RETRY_STATE/target")" ] \\ + && [ ! -e "\$STOP_RETRY_STATE/spent" ]; then + : > "\$STOP_RETRY_STATE/spent" + printf ' 999999\n' + exit 0 +fi +exec "$REAL_PS" "\$@" +SH +chmod +x "$RETRY_BIN/ps" + +retry_pe() { # <command...> + PATH="$RETRY_BIN:$PATH" STOP_RETRY_STATE="$RETRY_STATE" \ + FM_PROCEVENT_OWNER_LEASE_SECONDS=2 FM_PROCEVENT_OWNER_CHECK_SECONDS=1 \ + FM_HOME="$RETRY_HOME" "$ROOT/bin/fm-procevent.sh" "$@" +} + +retry_pe register lavish retry-src -- "$ORPHAN_STUB" "$TMP_ROOT/stop-retry-marker" >/dev/null +retry_pe reconcile >/dev/null +wait_for "$RETRY_HOME/state/procevent/retry-src.runner" \ + || fail "the retry listener never recorded its runner" +RETRY_PID=$(cat "$RETRY_HOME/state/procevent/retry-src.runner") +# Armed only now: the runner already proved its own process group at startup, +# and arming earlier would fail that assertion instead of the stop under test. +printf '%s\n' "$RETRY_PID" > "$RETRY_STATE/target" +wait_for "$TMP_ROOT/stop-retry-marker.descendant" \ + || fail "the retry listener's child never spawned its own descendant" +RETRY_DESCENDANT=$(cat "$TMP_ROOT/stop-retry-marker.descendant") + +deadline=$((SECONDS + 60)) +while kill -0 -"$RETRY_PID" 2>/dev/null; do + [ "$SECONDS" -lt "$deadline" ] \ + || fail "the guard gave up on an expired runner after a stop it could not prove" + sleep 0.5 +done +[ -e "$RETRY_STATE/spent" ] \ + || fail "the unprovable stop attempt this test injects never happened" +wait_gone "$RETRY_DESCENDANT" \ + || fail "the guard stopped retrying before the expired runner's descendant was reaped" +pass "a stop the guard cannot prove is retried until the expired runner is reaped" + printf '\nall procevent tests passed\n' diff --git a/tests/fm-quota-array-dispatch-live-e2e.test.sh b/tests/fm-quota-array-dispatch-live-e2e.test.sh index c6d30242aff..88a42c67764 100755 --- a/tests/fm-quota-array-dispatch-live-e2e.test.sh +++ b/tests/fm-quota-array-dispatch-live-e2e.test.sh @@ -8,10 +8,10 @@ # fall back without the call log catching it. set -u -if [ "${FM_QUOTA_ARRAY_DISPATCH_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_QUOTA_ARRAY_DISPATCH_LIVE_E2E=1 to run the credentialed Pi dispatch-selection regression" - exit 0 -fi +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate opt-in FM_QUOTA_ARRAY_DISPATCH_LIVE_E2E pi python3 ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" OWNER="$ROOT/.agents/skills/quota-array-dispatch/SKILL.md" @@ -21,8 +21,6 @@ fail() { exit 1 } -command -v pi >/dev/null 2>&1 || fail "pi not found" -command -v python3 >/dev/null 2>&1 || fail "python3 not found" [ -f "$OWNER" ] || fail "quota-array-dispatch skill not found" LAB=$(mktemp -d "${TMPDIR:-/tmp}/fm-quota-array-dispatch-live.XXXXXX") diff --git a/tests/fm-remote-job-orphan-reap.test.sh b/tests/fm-remote-job-orphan-reap.test.sh index a914e57ff03..e6f9e21898f 100755 --- a/tests/fm-remote-job-orphan-reap.test.sh +++ b/tests/fm-remote-job-orphan-reap.test.sh @@ -61,6 +61,30 @@ wait_child() { # <pid> <seconds> return 1 } +# True when <pid>'s parent is a reaper for orphaned processes: init itself, or +# a subreaper systemd registers one hop below init (PR_SET_CHILD_SUBREAPER, +# e.g. `systemd --user`) - a live host's per-user manager adopts orphans there +# instead of letting them reach real init, and that is just as orphaned for +# this fixture's purpose. +is_orphaned() { # <pid> + local parent + parent=$(ppid_of "$1") + case "$parent" in ''|*[!0-9]*) return 1 ;; esac + [ "$parent" = 1 ] && return 0 + [ "$(ppid_of "$parent")" = 1 ] +} + +# Wait up to <seconds> for <pid> to be reparented to an orphan reaper (see +# is_orphaned) after its launching shell exits; 0 when it does. +wait_orphaned() { # <pid> <seconds> + local pid=$1 deadline=$(( $(date +%s) + $2 )) + while [ "$(date +%s)" -lt "$deadline" ]; do + is_orphaned "$pid" && return 0 + sleep 0.1 + done + return 1 +} + # --- a real worker fixture, launched exactly the way fm-on's Linux start does - # build_remote_root <dir>: a minimal but genuine Firstmate code root carrying diff --git a/tests/fm-rovo-signals-live-e2e.test.sh b/tests/fm-rovo-signals-live-e2e.test.sh index 9cef364fb74..3d9896849e8 100644 --- a/tests/fm-rovo-signals-live-e2e.test.sh +++ b/tests/fm-rovo-signals-live-e2e.test.sh @@ -18,6 +18,9 @@ # places, so it is what this guard drives. set -u +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" ROVO_BIN=$(command -v rovo 2>/dev/null || true) [ -x "${ROVO_BIN:-}" ] || ROVO_BIN="$HOME/.local/bin/rovo" @@ -31,13 +34,9 @@ pass() { printf 'ok - %s\n' "$1" } -if [ "${FM_ROVO_SIGNALS_LIVE:-0}" != 1 ]; then - echo "skip: set FM_ROVO_SIGNALS_LIVE=1 to run the real Rovo signal drift guard" - exit 0 -fi +fm_live_gate opt-in FM_ROVO_SIGNALS_LIVE python3 [ -x "$ROVO_BIN" ] || fail "FM_ROVO_SIGNALS_LIVE=1 but no real rovo executable is installed" -command -v python3 >/dev/null 2>&1 || fail "python3 is required to drive rovo through a PTY" VERSION_OUT=$("$ROVO_BIN" --version 2>&1) || fail "rovo --version failed: $VERSION_OUT" echo "BOOTSTRAP_INFO: live rovo version: $VERSION_OUT" diff --git a/tests/fm-secondmate-harness.test.sh b/tests/fm-secondmate-harness.test.sh index b72f994e6ef..dac5c6cbe65 100755 --- a/tests/fm-secondmate-harness.test.sh +++ b/tests/fm-secondmate-harness.test.sh @@ -769,7 +769,7 @@ test_spawn_secondmate_harness_model_token() { [ "$(meta_field "$meta" model)" = opus ] || fail "model-token: meta model not opus (got '$(meta_field "$meta" model)')" [ "$(meta_field "$meta" effort)" = default ] || fail "model-token: meta effort not default (got '$(meta_field "$meta" effort)')" launch=$(cat "$launchlog") - assert_contains "$launch" "claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\"}' --model 'opus'" \ + assert_contains "$launch" "claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\",\"attribution\":{\"commit\":\"\",\"pr\":\"\",\"sessionUrl\":false}}' --model 'opus'" \ "model-token: launch did not carry --model opus" assert_not_contains "$launch" "--effort" "model-token: launch must not carry an --effort flag" pass "C3 spawn: config/secondmate-harness's model token threads --model into the launch and meta" @@ -791,7 +791,7 @@ test_spawn_secondmate_harness_model_and_effort_tokens() { [ "$(meta_field "$meta" model)" = opus ] || fail "model-effort-tokens: meta model not opus" [ "$(meta_field "$meta" effort)" = high ] || fail "model-effort-tokens: meta effort not high (got '$(meta_field "$meta" effort)')" launch=$(cat "$launchlog") - assert_contains "$launch" "claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\"}' --model 'opus' --effort 'high'" \ + assert_contains "$launch" "claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\",\"attribution\":{\"commit\":\"\",\"pr\":\"\",\"sessionUrl\":false}}' --model 'opus' --effort 'high'" \ "model-effort-tokens: launch did not carry both --model opus and --effort high" pass "C4 spawn: config/secondmate-harness's model+effort tokens thread into the launch and meta" } diff --git a/tests/fm-send-inbox-doorbell-live-e2e.test.sh b/tests/fm-send-inbox-doorbell-live-e2e.test.sh index e6a5d696fe7..36de34f56b4 100644 --- a/tests/fm-send-inbox-doorbell-live-e2e.test.sh +++ b/tests/fm-send-inbox-doorbell-live-e2e.test.sh @@ -28,14 +28,13 @@ # unready state and correctly fails that harness's check. set -u +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" -if [ "${FM_SEND_INBOX_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_SEND_INBOX_LIVE_E2E=1 to run the live steering-inbox doorbell guard" - exit 0 -fi +fm_live_gate opt-in FM_SEND_INBOX_LIVE_E2E tmux -command -v tmux >/dev/null 2>&1 || { echo "not ok - FM_SEND_INBOX_LIVE_E2E=1 but tmux is not installed" >&2; exit 1; } unset NO_MISTAKES_GATE SOCKET="fm-inbox-live-$$" diff --git a/tests/fm-send-secondmate-marker-herdr-e2e.test.sh b/tests/fm-send-secondmate-marker-herdr-e2e.test.sh index 3ac165e4b53..404fbb49454 100755 --- a/tests/fm-send-secondmate-marker-herdr-e2e.test.sh +++ b/tests/fm-send-secondmate-marker-herdr-e2e.test.sh @@ -22,14 +22,7 @@ set -u # shellcheck source=/dev/null . "$ROOT/bin/fm-backend.sh" -if [ "${FM_SEND_MARKER_HERDR_E2E:-0}" != 1 ]; then - echo "skip: set FM_SEND_MARKER_HERDR_E2E=1 to run the real Pi/Herdr secondmate-marker regression" - exit 0 -fi - -for tool in git herdr jq pi; do - command -v "$tool" >/dev/null 2>&1 || { echo "skip: $tool not found"; exit 0; } -done +fm_live_gate opt-in FM_SEND_MARKER_HERDR_E2E git herdr jq pi LAB_HELPER=${HERDR_LAB_HELPER:-$ROOT/bin/fm-herdr-lab.sh} SESSION=$("$LAB_HELPER" name fm-send-secondmate-marker-v7) diff --git a/tests/fm-session-start.test.sh b/tests/fm-session-start.test.sh index 06fe3355d70..0f3e273b163 100755 --- a/tests/fm-session-start.test.sh +++ b/tests/fm-session-start.test.sh @@ -110,25 +110,6 @@ SH printf '%s\n' manual > "${fakebin%/*}/home-placeholder" 2>/dev/null || true } -# hide_node <fakebin> <home>: make node genuinely undetectable so bootstrap emits -# its real "MISSING: node" diagnostic. Deleting the fake binary alone is not -# enough on hosts that ship node in the base PATH, so this also masks the lookup -# the same way the Herdr cases mask tmux. Echoes the mask path for BASH_ENV. -hide_node() { - local fakebin=$1 home=$2 mask - rm -f "$fakebin/node" - mask="$home/mask-node.bash" - cat > "$mask" <<'SH' -command() { - if [ "${1:-}" = -v ] && [ "${2:-}" = node ]; then - return 1 - fi - builtin command "$@" -} -SH - printf '%s\n' "$mask" -} - # make_fake_tasks_axi_compact <fakebin>: a tasks-axi boundary that answers the # four group filters the startup listing composes (in-flight, held, blocked # queued, and the dispatchable ready set) and REFUSES anything the recovery @@ -980,7 +961,7 @@ SH # still leads, live fleet identity now outranks curated memory, and the # read-once contract arrives before the payload it governs. test_output_ordering_diagnostics_lead() { - local rec root home fakebin mask out lock_line boot_line wake_line read_once_line + local rec root home fakebin out lock_line boot_line wake_line read_once_line local context_line fleet_line next_line inventory_line missing_line rec=$(new_world ordering) IFS='|' read -r root home fakebin <<EOF @@ -993,14 +974,13 @@ EOF # rather than a common tool like node that a host may install into $BASE_PATH # (/usr/bin) and thereby mask the forced-missing condition. rm -f "$fakebin/no-mistakes" - # Also mask node, so a host that ships it in the base PATH cannot change which - # other diagnostics the digest emits around the asserted line. - mask=$(hide_node "$fakebin" "$home") + # The filtered base PATH below also excludes host node installations. + rm -f "$fakebin/node" printf 'window=fm-sess:w1\nkind=ship\n' > "$home/state/task-a.meta" printf 'Captain memory that may be truncated away safely.\n' > "$home/data/captain.md" - out=$(BASH_ENV="$mask" run_session_start "$home" "$root" "$fakebin:$BASE_PATH") + out=$(run_session_start "$home" "$root" "$fakebin:$(fm_test_base_path_sans "$BASE_PATH" node)") lock_line=$(printf '%s\n' "$out" | grep -n '^LOCK$' | head -1 | cut -d: -f1) boot_line=$(printf '%s\n' "$out" | grep -n '^BOOTSTRAP$' | head -1 | cut -d: -f1) @@ -1392,7 +1372,7 @@ EOF # --- composition: real scripts run, not reimplemented ------------------------ test_composition_invokes_real_scripts() { - local rec root home fakebin mask out + local rec root home fakebin out rec=$(new_world composition) IFS='|' read -r root home fakebin <<EOF $rec @@ -1402,12 +1382,12 @@ EOF # Force a MISSING via a firstmate-specific tool never on a standard system PATH # (node can leak from $BASE_PATH; see test_output_ordering_diagnostics_lead). rm -f "$fakebin/no-mistakes" - mask=$(hide_node "$fakebin" "$home") + rm -f "$fakebin/node" printf 'needs-decision: pick a library\n' > "$home/state/task-z.status" append_wake "$home/state" signal task-z.status "needs-decision: pick a library" - out=$(BASH_ENV="$mask" run_session_start "$home" "$root" "$fakebin:$BASE_PATH") + out=$(run_session_start "$home" "$root" "$fakebin:$(fm_test_base_path_sans "$BASE_PATH" node)") # fm-lock.sh's own exact success text. assert_contains "$out" "lock acquired: harness pid" "fm-lock.sh's real output did not appear (composition, not reimplementation)" diff --git a/tests/fm-sessionstart-hook-live-e2e.test.sh b/tests/fm-sessionstart-hook-live-e2e.test.sh index ba38197a0d2..acc295666d6 100755 --- a/tests/fm-sessionstart-hook-live-e2e.test.sh +++ b/tests/fm-sessionstart-hook-live-e2e.test.sh @@ -40,11 +40,10 @@ # FM_PI_SESSIONSTART_RACE_LIVE_E2E=1 tests/fm-sessionstart-hook-live-e2e.test.sh set -u -if [ "${FM_SESSIONSTART_HOOK_LIVE_E2E:-0}" != 1 ] && \ - [ "${FM_PI_SESSIONSTART_RACE_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_SESSIONSTART_HOOK_LIVE_E2E=1 for the cross-harness guard or FM_PI_SESSIONSTART_RACE_LIVE_E2E=1 for the offline Pi /new race regression" - exit 0 -fi +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate opt-in FM_SESSIONSTART_HOOK_LIVE_E2E,FM_PI_SESSIONSTART_RACE_LIVE_E2E tmux ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" unset NO_MISTAKES_GATE @@ -56,8 +55,6 @@ fail() { pass() { printf 'ok - %s\n' "$1"; } note() { printf '# %s\n' "$1"; } -command -v tmux >/dev/null 2>&1 || fail "tmux not found; the context-reset checks drive real interactive harnesses" - # Outside the repo on purpose: each lab is its own git repo, and nesting one # inside the checkout would show up as an embedded repository in a working tree # a maintainer may be committing from while this guard runs. diff --git a/tests/fm-sessionstart-instruction-refresh-live-e2e.test.sh b/tests/fm-sessionstart-instruction-refresh-live-e2e.test.sh index 0ab68bc2cec..42819acc1bb 100755 --- a/tests/fm-sessionstart-instruction-refresh-live-e2e.test.sh +++ b/tests/fm-sessionstart-instruction-refresh-live-e2e.test.sh @@ -24,10 +24,10 @@ # This costs real Pi model turns and requires its normal authenticated profile. set -u -if [ "${FM_SESSIONSTART_INSTRUCTION_REFRESH_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_SESSIONSTART_INSTRUCTION_REFRESH_LIVE_E2E=1 to run the isolated real-Pi instruction-refresh regression" - exit 0 -fi +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate opt-in FM_SESSIONSTART_INSTRUCTION_REFRESH_LIVE_E2E pi tmux git ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" TMUX_SOCKET="fm-sessionstart-instruction-refresh-$$" @@ -107,10 +107,6 @@ cleanup() { } trap cleanup EXIT INT TERM -command -v pi >/dev/null 2>&1 || fail "pi not found" -command -v tmux >/dev/null 2>&1 || fail "tmux not found" -command -v git >/dev/null 2>&1 || fail "git not found" - mkdir -p "$LAB" git clone --quiet --no-hardlinks "$ROOT" "$PROJECT" || fail "could not create isolated Firstmate checkout" git -C "$PROJECT" checkout -q -B main "$TEST_COMMIT" \ diff --git a/tests/fm-spawn-dispatch-profile.test.sh b/tests/fm-spawn-dispatch-profile.test.sh index 05c25b4718c..9126ca0dc49 100755 --- a/tests/fm-spawn-dispatch-profile.test.sh +++ b/tests/fm-spawn-dispatch-profile.test.sh @@ -107,11 +107,12 @@ run_spawn() { local home=$1 wt=$2 fakebin=$3 launchlog=$4 shift 4 : > "$launchlog" + : > "$launchlog.text-lines" # An explicitly set CLAUDE_CONFIG_DIR is forwarded onto Claude launches, so # pin it empty by default instead of leaking the invoking shell's value. # A test opts in to the set case via FM_TEST_CLAUDE_CONFIG_DIR. CLAUDE_CONFIG_DIR="${FM_TEST_CLAUDE_CONFIG_DIR:-}" \ - FM_FAKE_LAUNCH_LOG="$launchlog" FM_FAKE_PI_VERSION="${FM_TEST_PI_VERSION:-0.84.0}" \ + FM_FAKE_LAUNCH_LOG="$launchlog" FM_FAKE_TEXT_LINE_LOG="$launchlog.text-lines" FM_FAKE_PI_VERSION="${FM_TEST_PI_VERSION:-0.84.0}" \ FM_FAKE_DAEMON_CLAUDE_CONFIG_DIR="${FM_FAKE_DAEMON_CLAUDE_CONFIG_DIR:-}" \ FM_FAKE_RESOLVED_CLAUDE_CONFIG_LOG="${FM_FAKE_RESOLVED_CLAUDE_CONFIG_LOG:-}" \ FM_FAKE_CURSOR_MODELS="${FM_TEST_CURSOR_MODELS:-}" \ @@ -157,7 +158,7 @@ test_no_profile_keeps_claude_profile_defaults() { "spawn did not copy the brief's exact task branch into metadata" launch=$(cat "$LAUNCH_LOG") - expected="env -u CURSOR_AGENT -u CURSOR_INVOKED_AS -u GEMINI_CLI CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\"}' \"\$('${ROOT}/bin/fm-operational-input.sh' encode launch-brief < '$HOME_DIR/data/$id/launch-brief.md')\"" + expected="env -u CURSOR_AGENT -u CURSOR_INVOKED_AS -u GEMINI_CLI CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\",\"attribution\":{\"commit\":\"\",\"pr\":\"\",\"sessionUrl\":false}}' \"\$('${ROOT}/bin/fm-operational-input.sh' encode launch-brief < '$HOME_DIR/data/$id/launch-brief.md')\"" [ "$launch" = "$expected" ] || fail "no-profile claude launch did not use the canonical launch kind"$'\n'"expected: $expected"$'\n'"actual: $launch" pass "no --model/--effort records defaults and types the claude launch instructions" } @@ -421,7 +422,7 @@ test_claude_threads_model_and_effort() { expect_code 0 "$status" "claude spawn with profile flags should succeed" assert_meta_profile "$HOME_DIR/state/$id.meta" claude sonnet high launch=$(cat "$LAUNCH_LOG") - assert_contains "$launch" "claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\"}' --model 'sonnet' --effort 'high'" \ + assert_contains "$launch" "claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\",\"attribution\":{\"commit\":\"\",\"pr\":\"\",\"sessionUrl\":false}}' --model 'sonnet' --effort 'high'" \ "claude launch did not thread model and effort flags" assert_not_contains "$launch" "--tui-mode" "non-Pi launches must not receive Pi's TUI mode override" pass "claude receives --model and --effort profile flags" @@ -993,6 +994,8 @@ EOF assert_contains "$out" "spawned $id harness=$harness kind=design" \ "design spawn did not retain kind=design on $harness" schema_launch=$(cat "$LAUNCH_LOG") + assert_contains "$(cat "$LAUNCH_LOG.text-lines")" "export FM_TASK_ID=$id" \ + "design worker did not receive the primary-checkout isolation marker" schema_meta="$HOME_DIR/state/$id.meta" pass "design matrix $harness 1/5: dispatch schema selects $harness/$model/$effort" if [ "$harness" = codex ]; then @@ -1050,6 +1053,49 @@ EOF done } +# The captain's attribution policy lives in the `user` settings scope, which a +# spawned worker's settings sources are not guaranteed to load. Every claude +# launch must therefore carry the policy itself, or a spawned worker writes +# Co-Authored-By and Claude-Session trailers into commits and PR bodies. +assert_attribution_policy() { # <launch-command> <what> + local launch=$1 what=$2 + assert_contains "$launch" '"attribution":' "$what launch carries no attribution policy" + assert_contains "$launch" '"commit":""' "$what launch does not silence the commit trailer" + assert_contains "$launch" '"pr":""' "$what launch does not silence the PR-body attribution" + assert_contains "$launch" '"sessionUrl":false' "$what launch does not silence the session URL" +} + +test_claude_crewmate_launch_carries_the_attribution_policy() { + local rec id out status launch + id="profile-claude-attribution-z22" + rec=$(make_spawn_case profile-claude-attribution claude "$id") + read_case_record "$rec" + + out=$(run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR") + status=$? + expect_code 0 "$status" "claude crewmate spawn should succeed"$'\n'"$out" + launch=$(cat "$LAUNCH_LOG") + assert_attribution_policy "$launch" "claude crewmate" + pass "a claude crewmate launch carries the attribution-off policy in its own settings" +} + +test_claude_secondmate_launch_carries_the_attribution_policy() { + local rec id sm out status launch + id="profile-secondmate-attribution-z23" + rec=$(make_spawn_case profile-secondmate-attribution claude "$id") + read_case_record "$rec" + sm="$CASE_DIR/secondmate-home" + make_seeded_secondmate_home "$sm" "$id" + + out=$(FM_TEST_CLAUDE_CONFIG_DIR="$CASE_DIR/claude-work" \ + run_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$sm" --secondmate) + status=$? + expect_code 0 "$status" "secondmate claude spawn should succeed"$'\n'"$out" + launch=$(cat "$LAUNCH_LOG") + assert_attribution_policy "$launch" "claude secondmate" + pass "a claude secondmate launch carries the attribution-off policy too" +} + test_active_dispatch_profile_does_not_block_secondmate_launch() { local rec id sm out status id='profile-secondmate-z16' @@ -1485,6 +1531,8 @@ test_batch_forwards_shared_profile_flags test_claude_forwards_firstmate_config_dir_when_set test_claude_default_uses_home_config_and_records_evidence_store test_non_claude_harness_ignores_config_dir +test_claude_crewmate_launch_carries_the_attribution_policy +test_claude_secondmate_launch_carries_the_attribution_policy test_design_profile_resolves_on_claude_codex_and_pi test_active_dispatch_profile_does_not_block_secondmate_launch test_spawn_copies_continued_task_branch_from_brief_flag diff --git a/tests/fm-spawn-pool-base-freshen.test.sh b/tests/fm-spawn-pool-base-freshen.test.sh index df80f90a276..aeb5a193149 100755 --- a/tests/fm-spawn-pool-base-freshen.test.sh +++ b/tests/fm-spawn-pool-base-freshen.test.sh @@ -4,8 +4,8 @@ # A treehouse pool can return a clean detached worktree whose origin/main was # advanced after the worktree was allocated. # These tests drive the real spawn path with a fake terminal, then prove it -# starts the worker from the fetched origin/main tip or stops when origin is -# unreachable. +# starts the worker from the fetched origin tip, launches a clean origin-less +# pool as-is, or stops when a configured origin is unusable. set -u # shellcheck source=tests/fixtures.sh @@ -59,6 +59,41 @@ run_spawn() { "$id" "$PROJECT_DIR" "$@" } +test_remote_seeded_home_spawns_from_treehouse_pool() { + local rec id out status lock + id='pool-remote-seeded-r13' + rec=$(make_case remote-seeded "$id") + read_case_record "$rec" + cat > "$HOME_DIR/.fm-secondmate-parent" <<'REC' +schema=fm-secondmate-parent.v1 +route=remote +parent_host=parent-machine +REC + + out=$(run_spawn "$id" --scout) + status=$? + expect_code 0 "$status" \ + "a remote-seeded secondmate home should allocate and launch from its Treehouse pool"$'\n'"$out" + assert_contains "$out" "spawned $id" \ + "the remote-seeded spawn did not report success" + assert_grep "worktree=$POOL_DIR" "$HOME_DIR/state/$id.meta" \ + "the remote-seeded spawn did not publish its allocated pool worktree" + lock=$(FM_HOME="$HOME_DIR" bash -c '. "$1"; fm_treehouse_project_lock_path "$2"' _ \ + "$ROOT/bin/fm-wake-lib.sh" "$PROJECT_DIR") \ + || fail "the launched remote-seeded home could not resolve its Treehouse project lock" + case "$lock" in + "$HOME_DIR/state/"*) ;; + *) fail "the remote-seeded spawn anchored its lock outside its local root: $lock" ;; + esac + if [ "${FM_TEST_EVIDENCE:-0}" = 1 ]; then + printf '# remote-seeded Treehouse spawn command\n' + printf '$ FM_HOME=%s bin/fm-spawn.sh %s %s --scout\n%s\nexit=%s\n' \ + "$HOME_DIR" "$id" "$PROJECT_DIR" "$out" "$status" + printf 'published worktree=%s\nresolved project lock=%s\n' "$POOL_DIR" "$lock" + fi + pass "a remote-seeded secondmate home allocates and launches from its Treehouse pool" +} + test_linked_spawning_home_rejects_primary_before_refresh() { local rec id out status returned primary spawning before_reflog for returned in primary primary-alias spawning scout; do @@ -180,6 +215,154 @@ test_non_main_default_branch_refreshes_before_branching() { pass "a stale pooled worktree resolves and refreshes a non-main default branch" } +make_originless_case() { # <name> <id> + local name=$1 id=$2 case_dir home project pool fakebin initial + case_dir="$TMP_ROOT/$name" + home="$case_dir/home" + project="$case_dir/project" + pool="$case_dir/pool" + fakebin=$(make_spawn_fakebin "$case_dir/fake") + + mkdir -p "$home/data/$id" "$home/projects" "$home/state" "$home/config" + printf 'codex\n' > "$home/config/crew-harness" + fm_test_spawn_brief "$home" "$id" + touch "$home/state/.last-watcher-beat" + + git init --quiet -b main "$project" + printf 'base\n' > "$project/README.md" + git -C "$project" add README.md + git -C "$project" -c user.name='Firstmate Tests' -c user.email='tests@example.invalid' commit -qm initial + initial=$(git -C "$project" rev-parse HEAD) + git -C "$project" worktree add --quiet --detach "$pool" "$initial" + + printf '%s\n' "$case_dir|$home|$project|$pool|$fakebin|$initial|main" +} + +test_originless_pool_launches_without_a_freshness_fetch() { + local rec id out status before + id='pool-originless-r6' + rec=$(make_originless_case originless "$id") + read_case_record "$rec" + ! git -C "$POOL_DIR" remote get-url origin >/dev/null 2>&1 \ + || fail "fixture unexpectedly configured an origin remote" + before=$(git -C "$POOL_DIR" rev-parse HEAD) + + out=$(run_spawn "$id" --mode no-mistakes --yolo off) + status=$? + expect_code 0 "$status" "spawn should launch a local-only pooled worktree with no origin"$'\n'"$out" + assert_contains "$out" "spawned $id" "spawn did not report success for the origin-less pool" + assert_not_contains "$out" "could not fetch origin" \ + "spawn attempted a freshness fetch against a nonexistent origin" + [ ! -e "$POOL_DIR/.git/FETCH_HEAD" ] || fail "spawn fetched against a pooled worktree with no origin" + [ "$(git -C "$POOL_DIR" rev-parse HEAD)" = "$before" ] \ + || fail "spawn moved HEAD on an origin-less pooled worktree that had nothing to refresh against" + if [ "${FM_TEST_EVIDENCE:-0}" = 1 ]; then + printf '# observed origin-less launch: %s\n' "$(printf '%s\n' "$out" | tail -n 1)" + fi + pass "an origin-less pooled worktree launches as-is, skipping the freshness gate" +} + +test_originless_dirty_pool_refuses_without_discarding_work() { + local rec id out status before + id='pool-originless-dirty-r1' + rec=$(make_originless_case originless-dirty "$id") + read_case_record "$rec" + before=$(git -C "$POOL_DIR" rev-parse HEAD) + printf 'keep this local work\n' > "$POOL_DIR/uncommitted.txt" + + out=$(run_spawn "$id" --mode no-mistakes --yolo off) + status=$? + [ "$status" -ne 0 ] || fail "spawn succeeded despite a dirty origin-less pooled worktree" + assert_contains "$out" "is not clean" \ + "spawn did not clearly refuse a dirty origin-less pooled worktree" + [ "$(git -C "$POOL_DIR" rev-parse HEAD)" = "$before" ] \ + || fail "spawn moved HEAD while refusing a dirty origin-less pooled worktree" + assert_grep 'keep this local work' "$POOL_DIR/uncommitted.txt" \ + "spawn discarded local work from an origin-less pool" + pass "a dirty origin-less pooled worktree is refused without discarding its local work" +} + +test_origin_config_without_url_refuses_pool() { + local rec id out status before + id='pool-origin-without-url-r1' + rec=$(make_originless_case origin-without-url "$id") + read_case_record "$rec" + git -C "$POOL_DIR" config remote.origin.fetch '+refs/heads/*:refs/remotes/origin/*' + before=$(git -C "$POOL_DIR" rev-parse HEAD) + + out=$(run_spawn "$id" --mode no-mistakes --yolo off) + status=$? + [ "$status" -ne 0 ] || fail "spawn succeeded despite an origin configuration with no URL" + assert_contains "$out" "could not fetch origin" \ + "spawn did not refuse an origin configuration with no URL as unusable" + [ "$(git -C "$POOL_DIR" rev-parse HEAD)" = "$before" ] \ + || fail "spawn moved HEAD after finding an unusable origin configuration" + [ ! -e "$HOME_DIR/state/$id.meta" ] || fail "refused spawn published task metadata" + pass "an origin configuration without a URL refuses the pooled worktree" +} + +test_empty_origin_config_section_refuses_pool() { + local rec id out status before config + id='pool-empty-origin-section-r1' + rec=$(make_originless_case empty-origin-section "$id") + read_case_record "$rec" + config=$(git -C "$POOL_DIR" rev-parse --path-format=absolute --git-path config) + printf '\n[remote "origin"]\n' >> "$config" + before=$(git -C "$POOL_DIR" rev-parse HEAD) + + out=$(run_spawn "$id" --mode no-mistakes --yolo off) + status=$? + [ "$status" -ne 0 ] || fail "spawn succeeded despite an empty origin configuration section" + assert_contains "$out" "could not fetch origin" \ + "spawn did not refuse an empty origin configuration section as unusable" + [ "$(git -C "$POOL_DIR" rev-parse HEAD)" = "$before" ] \ + || fail "spawn moved HEAD after finding an empty origin configuration section" + [ ! -e "$HOME_DIR/state/$id.meta" ] || fail "refused spawn published task metadata" + pass "an empty origin configuration section refuses the pooled worktree" +} + +test_empty_only_included_origin_config_section_launches_pool() { + local rec id out status before config included + id='pool-empty-only-included-origin-section-r1' + rec=$(make_originless_case empty-only-included-origin-section "$id") + read_case_record "$rec" + config=$(git -C "$POOL_DIR" rev-parse --path-format=absolute --git-path config) + included=$(dirname "$config")/empty-origin.inc + printf '[remote "origin"]\n' > "$included" + git -C "$POOL_DIR" config include.path "$(basename "$included")" + before=$(git -C "$POOL_DIR" rev-parse HEAD) + + out=$(run_spawn "$id" --mode no-mistakes --yolo off) + status=$? + expect_code 0 "$status" "spawn should proceed when an included empty origin section is not enumerable"$'\n'"$out" + assert_contains "$out" "spawned $id" "spawn did not report success for the undetectable included section" + assert_not_contains "$out" "could not fetch origin" \ + "spawn treated an undetectable included empty section as a configured origin" + [ "$(git -C "$POOL_DIR" rev-parse HEAD)" = "$before" ] \ + || fail "spawn moved HEAD despite treating the included empty section as origin-less" + pass "an empty-only included origin section documents the accepted detection boundary" +} + +test_inactive_conditional_origin_include_launches_pool() { + local rec id out status before config included + id='pool-inactive-origin-include-r1' + rec=$(make_originless_case inactive-origin-include "$id") + read_case_record "$rec" + config=$(git -C "$POOL_DIR" rev-parse --path-format=absolute --git-path config) + included=$(dirname "$config")/inactive-origin.inc + printf '[fm-test]\n\tmarker = true\n[remote "origin"]\n' > "$included" + git -C "$POOL_DIR" config 'includeIf.gitdir:/never/matches/this/worktree/.path' "$included" + before=$(git -C "$POOL_DIR" rev-parse HEAD) + + out=$(run_spawn "$id" --mode no-mistakes --yolo off) + status=$? + expect_code 0 "$status" "spawn should ignore an inactive conditional origin include"$'\n'"$out" + assert_contains "$out" "spawned $id" "spawn did not report success with an inactive origin include" + [ "$(git -C "$POOL_DIR" rev-parse HEAD)" = "$before" ] \ + || fail "spawn moved HEAD despite having no effective origin" + pass "an inactive conditional origin include leaves the pooled worktree origin-less" +} + test_unreachable_origin_refuses_stale_pool_base() { local rec id out status before after id='pool-unreachable-origin-r2' @@ -353,6 +536,7 @@ test_stale_submodule_pin_explains_itself() { rec=$(make_submodule_case stale-pin "$id") read_submodule_case "$rec" strand_submodule_pin_via_spawn 'pool-stale-pin-seed-r7' + git -C "$POOL_DIR" remote remove origin before=$(git -C "$POOL_DIR" rev-parse HEAD) before_sub=$(git -C "$POOL_DIR/ui" rev-parse HEAD) @@ -378,7 +562,7 @@ test_stale_submodule_pin_explains_itself() { if [ "${FM_TEST_EVIDENCE:-0}" = 1 ]; then printf '# observed stale-pin refusal: %s\n' "$(printf '%s\n' "$out" | grep 'submodule' | head -n 1)" fi - pass "two consecutive spawns across a moved submodule pin end in a refusal naming both pins and no remedy" + pass "an origin-less pool with a stale submodule pin refuses while naming both pins and no remedy" } test_unpushed_submodule_commit_is_still_uncommitted_work() { @@ -492,6 +676,7 @@ test_stale_pin_beside_other_dirt_reports_one_verdict() { pass "a stale pin beside other dirt yields the conservative refusal alone, with no stale-pin line" } +test_remote_seeded_home_spawns_from_treehouse_pool test_linked_spawning_home_rejects_primary_before_refresh test_stale_pool_base_refreshes_before_branching test_non_main_default_branch_refreshes_before_branching @@ -499,6 +684,12 @@ test_direct_pr_and_scout_refresh_before_launch test_dirty_pool_refuses_without_discarding_work test_unresolved_remote_default_refuses_pool test_unreachable_origin_refuses_stale_pool_base +test_originless_pool_launches_without_a_freshness_fetch +test_originless_dirty_pool_refuses_without_discarding_work +test_origin_config_without_url_refuses_pool +test_empty_origin_config_section_refuses_pool +test_empty_only_included_origin_config_section_launches_pool +test_inactive_conditional_origin_include_launches_pool test_stale_submodule_pin_explains_itself test_unpushed_submodule_commit_is_still_uncommitted_work test_work_inside_submodule_is_still_uncommitted_work diff --git a/tests/fm-stat-shadowing.test.sh b/tests/fm-stat-shadowing.test.sh new file mode 100644 index 00000000000..ec8ef4f72b7 --- /dev/null +++ b/tests/fm-stat-shadowing.test.sh @@ -0,0 +1,139 @@ +#!/usr/bin/env bash +# tests/fm-stat-shadowing.test.sh - verify Darwin BSD-stat helpers ignore a GNU +# stat earlier on PATH. +# +# On Darwin, GNU coreutils can put a GNU stat earlier on PATH than /usr/bin/stat. +# A bare `stat -f <fmt>` then reaches GNU stat, where `-f` means filesystem stat +# rather than BSD-format output and can leak a filesystem dump into callers. +# Runtime Darwin BSD-format calls use /usr/bin/stat so their syntax stays tied +# to the system BSD implementation. +# +# This test proves the invariant by installing a fake GNU-like stat that shadows +# /usr/bin/stat and asserting the helpers still return correct values. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +# Darwin-only: the shadowing assertions require a real BSD /usr/bin/stat to +# shadow; on Linux the `stat -f` semantics differ and the helpers take the +# `stat -c` branch instead, so there is nothing meaningful to assert. Skip +# visibly (after lib.sh so `pass`/`fail` exist) rather than silently. +if [ "$(uname)" != Darwin ]; then + pass "Darwin-only test: shadowing assertions require BSD /usr/bin/stat; skipping on $(uname -s)" + exit 0 +fi + +TMP_ROOT=$(mktemp -d "${TMPDIR:-/tmp}/fm-stat-shadowing.XXXXXX") || exit 1 +trap 'rm -rf "$TMP_ROOT"' EXIT + +# --- fake GNU stat that mimics ~/.local/bin/stat shadowing /usr/bin/stat ------- + +# A shadowed GNU stat does not interpret `-f <fmt>` as BSD-format output. +# We fail the tokens used in our helpers so callers cannot accidentally use the +# shadowed stat instead of /usr/bin/stat. +FAKE_STAT="$TMP_ROOT/fakebin/stat" +mkdir -p "$(dirname "$FAKE_STAT")" + +cat > "$FAKE_STAT" <<'FAKESTAT' +#!/usr/bin/env bash +# Mimics GNU coreutils stat when it shadows BSD /usr/bin/stat. +# On Darwin: BSD stat uses %m (mtime), %z (size), etc. +# GNU stat -f treats its argument as a filesystem-path option, not a format. +# For the format tokens our code uses, this fake exits non-zero before a helper +# could consume shadowed output. +opt1=${1:-} opt2=${2:-} +if [ "$opt1" = "-f" ]; then + case "$opt2" in + %m|%l|%z|%d|%Lp|%i|%u|%B|%FB|%HT:%p|%d:%i|%d:%i:%z:%m:%c) + printf 'File: "%s"\n' "${3:-}" >&2 + exit 1 + ;; + *) + printf 'GNU stat: unknown -f format: %s\n' "$opt2" >&2 + exit 1 + ;; + esac +fi +printf 'GNU stat: unknown invocation: %s\n' "$*" >&2 +exit 1 +FAKESTAT +chmod +x "$FAKE_STAT" + +# Prepend fakebin so the fake GNU stat shadows /usr/bin/stat +ORIGINAL_PATH="$PATH" +export PATH="$TMP_ROOT/fakebin:$ORIGINAL_PATH" + +# Verify the shadowing is active: a bare `stat -f %m /` must fail (not use BSD) +if stat -f %m / >/dev/null 2>&1; then + # The fake stat didn't catch this, so something is wrong with the PATH setup + PATH="$ORIGINAL_PATH" + fail "shadowing sanity check: bare stat -f %m / should fail under GNU-shadow but did not" +fi + +# Also verify /usr/bin/stat still works when called directly +REAL_MTIME=$(/usr/bin/stat -f %m "$0" 2>/dev/null) || true +[ -n "$REAL_MTIME" ] && [ "$REAL_MTIME" -ge 0 ] || { + PATH="$ORIGINAL_PATH" + fail "/usr/bin/stat -f %m sanity check failed — /usr/bin/stat is not working" +} +pass "shadowing: fake GNU stat shadows /usr/bin/stat in PATH" + +# --- test the fixed helpers under shadowing ---------------------------------- + +. "$ROOT/bin/fm-supervision-lib.sh" +. "$ROOT/bin/fm-startup-memory-budget-lib.sh" + +TESTFILE="$TMP_ROOT/testfile" +printf 'hello world\n' > "$TESTFILE" + +# 1. fm_sup_stat_mtime from bin/fm-supervision-lib.sh +RESULT_MTIME=$(fm_sup_stat_mtime "$TESTFILE") || true +EXPECTED_MTIME=$(/usr/bin/stat -f %m "$TESTFILE" 2>/dev/null) +if [ -z "$RESULT_MTIME" ] || [ "$RESULT_MTIME" != "$EXPECTED_MTIME" ]; then + PATH="$ORIGINAL_PATH" + fail "fm_sup_stat_mtime: expected $EXPECTED_MTIME, got '$RESULT_MTIME'" +fi +pass "fm_sup_stat_mtime returns correct epoch mtime under GNU stat shadowing" + +# 2. fm_startup_memory_budget_link_count from bin/fm-startup-memory-budget-lib.sh +RESULT_LINKS=$(fm_startup_memory_budget_link_count "$TESTFILE") || true +EXPECTED_LINKS=$(/usr/bin/stat -f %l "$TESTFILE" 2>/dev/null) +if [ -z "$RESULT_LINKS" ] || [ "$RESULT_LINKS" != "$EXPECTED_LINKS" ]; then + PATH="$ORIGINAL_PATH" + fail "fm_startup_memory_budget_link_count: expected $EXPECTED_LINKS, got '$RESULT_LINKS'" +fi +pass "fm_startup_memory_budget_link_count returns correct link count under GNU stat shadowing" + +# 3. _fm_status_file_size from bin/fm-classify-lib.sh +# We source it and call the internal function directly. +RESULT_SIZE=$(LC_ALL=C /usr/bin/stat -f '%z' "$TESTFILE" 2>/dev/null) || true +# The fixed code uses /usr/bin/stat so it should produce the same value as direct call +if [ -z "$RESULT_SIZE" ]; then + PATH="$ORIGINAL_PATH" + fail "_fm_status_file_size: could not get size via /usr/bin/stat" +fi +# Verify the helper itself works by checking that calling it with /usr/bin/stat prefix matches +. "$ROOT/bin/fm-classify-lib.sh" +HELPER_SIZE=$(_fm_status_file_size "$TESTFILE") || true +if [ -z "$HELPER_SIZE" ] || [ "$HELPER_SIZE" != "$RESULT_SIZE" ]; then + PATH="$ORIGINAL_PATH" + fail "_fm_status_file_size: expected $RESULT_SIZE, got '$HELPER_SIZE'" +fi +pass "_fm_status_file_size returns correct byte size under GNU stat shadowing" + +# 4. stat_mtime from bin/fm-watch.sh +# fm-watch.sh runs a top-level `mkdir -p` on its state dir when sourced; pin it +# to the temp root via FM_STATE_OVERRIDE so no artifact escapes into the repo's +# git-ignored state/ directory. +export FM_STATE_OVERRIDE="$TMP_ROOT/state" +. "$ROOT/bin/fm-watch.sh" +RESULT_WATCH_MTIME=$(stat_mtime "$TESTFILE") || true +if [ -z "$RESULT_WATCH_MTIME" ] || [ "$RESULT_WATCH_MTIME" != "$EXPECTED_MTIME" ]; then + PATH="$ORIGINAL_PATH" + fail "stat_mtime (fm-watch.sh): expected $EXPECTED_MTIME, got '$RESULT_WATCH_MTIME'" +fi +pass "stat_mtime from fm-watch.sh returns correct epoch mtime under GNU stat shadowing" + +# Restore PATH before exit +PATH="$ORIGINAL_PATH" diff --git a/tests/fm-teardown-endpoint-safety.test.sh b/tests/fm-teardown-endpoint-safety.test.sh index b927cf5057e..279d585601c 100755 --- a/tests/fm-teardown-endpoint-safety.test.sh +++ b/tests/fm-teardown-endpoint-safety.test.sh @@ -598,6 +598,233 @@ SH pass "fm-teardown: an exact recorded endpoint still tears down after changing cwd outside its worktree" } +# --- Treehouse project-lock anchoring across home layouts -------------------- +# +# The lock is anchored at the local root home, so every home on this machine +# that can reach the same pool must derive the identical file. A remote parent +# binding terminates that walk at the home holding it: its parent is on another +# machine and can neither hold nor observe a lock taken here. + +write_local_parent_record() { # <home> <parent-home> + cat > "$1/.fm-secondmate-parent" <<REC +schema=fm-secondmate-parent.v1 +route=local +parent_home=$2 +REC +} + +write_remote_parent_record() { # <home> + cat > "$1/.fm-secondmate-parent" <<'REC' +schema=fm-secondmate-parent.v1 +route=remote +parent_host=machine-a +REC +} + +make_home() { # <path> + mkdir -p "$1/state" "$1/data" "$1/config" "$1/projects" +} + +resolve_project_lock() { # <home> <project> + FM_HOME="$1" bash -c '. "$1"; fm_treehouse_project_lock_path "$2"' _ \ + "$ROOT/bin/fm-wake-lib.sh" "$2" +} + +test_project_lock_anchors_at_the_local_root_across_home_layouts() { + local dir main_home main_project local_mate remote_mate remote_child + local main_lock mate_lock remote_lock child_lock orphan_lock rc + dir=$(make_case project-lock-anchoring) + git -C "$dir/project" -c user.name=test -c user.email=test@example.invalid \ + commit --allow-empty -qm anchor-fixture + + # Main-home layout: a root home and a local secondmate beneath it. + main_home="$dir/home" + main_project="$main_home/projects/project" + make_home "$main_home" + git clone -q "$dir/project" "$main_project" + local_mate="$dir/local-mate" + make_home "$local_mate" + write_local_parent_record "$local_mate" "$main_home" + git clone -q "$dir/project" "$local_mate/projects/project" + + # Remote layout: a home seeded from another machine, plus its own local child. + remote_mate="$dir/remote-mate" + make_home "$remote_mate" + write_remote_parent_record "$remote_mate" + git clone -q "$dir/project" "$remote_mate/projects/project" + remote_child="$dir/remote-mate-child" + make_home "$remote_child" + write_local_parent_record "$remote_child" "$remote_mate" + git clone -q "$dir/project" "$remote_child/projects/project" + + main_lock=$(resolve_project_lock "$main_home" "$main_project") \ + || fail "the root home could not resolve its project lock" + mate_lock=$(resolve_project_lock "$local_mate" "$local_mate/projects/project") \ + || fail "a local secondmate home could not resolve its project lock" + remote_lock=$(resolve_project_lock "$remote_mate" "$remote_mate/projects/project") \ + || fail "a remote-seeded secondmate home could not resolve its project lock" + child_lock=$(resolve_project_lock "$remote_child" "$remote_child/projects/project") \ + || fail "a local child of a remote-seeded home could not resolve its project lock" + + [ "$main_lock" = "$mate_lock" ] \ + || fail "the root home and its local secondmate derived different project locks" + [ "$remote_lock" = "$child_lock" ] \ + || fail "a remote-seeded home and its local child derived different project locks" + case "$remote_lock" in + "$remote_mate/state/"*) ;; + *) fail "a remote-seeded home anchored its project lock outside its own state: $remote_lock" ;; + esac + + # An origin-less local-only project still resolves, keyed on its worktree top. + git init -q "$remote_mate/projects/local-only" + orphan_lock=$(resolve_project_lock "$remote_mate" "$remote_mate/projects/local-only") \ + || fail "an origin-less local-only project could not resolve its lock in a remote-seeded home" + [ "$orphan_lock" != "$remote_lock" ] \ + || fail "an origin-less project shared the lock identity of an unrelated origin" + + # Everything other than a remote route still fails closed. + printf 'schema=fm-secondmate-parent.v1\nroute=sideways\n' \ + > "$remote_child/.fm-secondmate-parent" + set +e + resolve_project_lock "$remote_child" "$remote_child/projects/project" >/dev/null 2>&1 + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "an unsupported parent route resolved a project lock instead of refusing" + + pass "Treehouse project locking anchors at the local root for main-home, local-secondmate, and remote-seeded layouts" +} + +test_remote_seeded_home_returns_its_uncontested_slot() { + local dir id=remote-task rc + dir=$(make_case remote-home-teardown) + mark_case_as_treehouse_pool "$dir" + write_remote_parent_record "$dir/home" + fm_write_meta "$dir/home/state/$id.meta" \ + "window=firstmate:fm-$id" "endpoint_task_id=$id" \ + "worktree=$dir/worktree" "project=$dir/project" "kind=scout" + + set +e + run_case "$dir" "$id" > "$dir/stdout" 2> "$dir/stderr" + rc=$? + set -e + [ "$rc" -eq 0 ] \ + || fail "teardown in a remote-seeded home refused its own uncontested slot: $(cat "$dir/stderr")" + assert_absent "$dir/home/state/$id.meta" "remote-seeded teardown left the task record" + grep -Fq "treehouse <return>" "$dir/runtime.log" \ + || fail "remote-seeded teardown did not return its own pool slot: $(cat "$dir/runtime.log")" + if [ "${FM_TEST_EVIDENCE:-0}" = 1 ]; then + printf '# remote-seeded Treehouse teardown command\n' + printf '$ FM_HOME=%s bin/fm-teardown.sh %s --force\n' "$dir/home" "$id" + printf 'stdout:\n'; cat "$dir/stdout" + printf 'stderr:\n'; cat "$dir/stderr" + printf 'exit=%s\nruntime calls:\n' "$rc"; cat "$dir/runtime.log" + printf 'task metadata=%s\nslot sentinel=%s\n' \ + "$([ -e "$dir/home/state/$id.meta" ] && printf present || printf removed)" \ + "$([ -e "$dir/worktree/sentinel" ] && printf present || printf removed)" + fi + + pass "fm-teardown: a remote-seeded secondmate home returns its own uncontested pool slot" +} + +test_remote_seeded_home_still_refuses_a_slot_its_child_holds() { + local dir id=remote-stale other=child-task child_home child_project rc + dir=$(make_case remote-home-collision) + mark_case_as_treehouse_pool "$dir" + write_remote_parent_record "$dir/home" + printf 'fixture\n' > "$dir/project/tracked" + git -C "$dir/project" add tracked + git -C "$dir/project" -c user.name=test -c user.email=test@example.invalid commit -qm fixture + child_home="$dir/child-home" + child_project="$child_home/projects/project" + make_home "$child_home" + write_local_parent_record "$child_home" "$dir/home" + git clone -q "$dir/project" "$child_project" + printf '%s\n' "- mate - fixture (home: $child_home; scope: test; projects: project; added 2026-01-01)" \ + > "$dir/home/data/secondmates.md" + fm_write_meta "$dir/home/state/$id.meta" \ + "window=firstmate:fm-$id" "endpoint_task_id=$id" \ + "worktree=$dir/worktree" "project=$dir/project" "kind=scout" + fm_write_meta "$child_home/state/$other.meta" \ + "window=firstmate:fm-$other" "endpoint_task_id=$other" \ + "worktree=$dir/worktree" "project=$child_project" "kind=scout" + + set +e + run_case "$dir" "$id" > "$dir/stdout" 2> "$dir/stderr" + rc=$? + set -e + [ "$rc" -ne 0 ] \ + || fail "a remote-seeded home returned a pool slot its own local child still holds" + assert_present "$dir/home/state/$id.meta" "remote-layout collision removed stale metadata" + assert_present "$child_home/state/$other.meta" "remote-layout collision removed live metadata" + assert_present "$dir/worktree/sentinel" "remote-layout collision reset the shared slot" + assert_contains "$(cat "$dir/stderr")" "$other" \ + "remote-layout refusal should name the task holding the slot" + if [ "${FM_TEST_EVIDENCE:-0}" = 1 ]; then + printf '# remote-seeded cross-home collision command\n' + printf '$ FM_HOME=%s bin/fm-teardown.sh %s --force\n' "$dir/home" "$id" + printf 'stderr:\n'; cat "$dir/stderr" + printf 'exit=%s\nruntime calls=%s\n' "$rc" \ + "$([ -s "$dir/runtime.log" ] && cat "$dir/runtime.log" || printf none)" + printf 'remote metadata=%s\nchild metadata=%s\nslot sentinel=%s\n' \ + "$([ -e "$dir/home/state/$id.meta" ] && printf preserved || printf removed)" \ + "$([ -e "$child_home/state/$other.meta" ] && printf preserved || printf removed)" \ + "$([ -e "$dir/worktree/sentinel" ] && printf preserved || printf removed)" + fi + + pass "fm-teardown: slot ownership across a remote-seeded home and its local child still refuses" +} + +test_remote_layout_homes_serialize_on_one_project_lock() { + local dir id=remote-serialize child_home child_project lock holder rc waited=0 + dir=$(make_case remote-lock-exclusion) + mark_case_as_treehouse_pool "$dir" + write_remote_parent_record "$dir/home" + printf 'fixture\n' > "$dir/project/tracked" + git -C "$dir/project" add tracked + git -C "$dir/project" -c user.name=test -c user.email=test@example.invalid commit -qm fixture + child_home="$dir/child-home" + child_project="$child_home/projects/project" + make_home "$child_home" + write_local_parent_record "$child_home" "$dir/home" + git clone -q "$dir/project" "$child_project" + fm_write_meta "$dir/home/state/$id.meta" \ + "window=firstmate:fm-$id" "endpoint_task_id=$id" \ + "worktree=$dir/worktree" "project=$dir/project" "kind=scout" + + # The local child takes the lock its own home derives and stays alive holding + # it, standing in for a slot allocation running in that home right now. + lock=$(resolve_project_lock "$child_home" "$child_project") \ + || fail "the local child could not resolve the shared project lock" + FM_HOME="$child_home" bash -c \ + '. "$1"; fm_lock_try_acquire "$2" || exit 1; : > "$3"; exec sleep 30' _ \ + "$ROOT/bin/fm-wake-lib.sh" "$lock" "$dir/lock-held" & + holder=$! + while [ ! -e "$dir/lock-held" ] && [ "$waited" -lt 100 ]; do + kill -0 "$holder" 2>/dev/null || break + sleep 0.1 + waited=$((waited + 1)) + done + [ -e "$dir/lock-held" ] || fail "the local child never took the shared project lock" + + set +e + run_case "$dir" "$id" > "$dir/stdout" 2> "$dir/stderr" + rc=$? + set -e + kill "$holder" 2>/dev/null || true + wait "$holder" 2>/dev/null || true + + [ "$rc" -ne 0 ] \ + || fail "a remote-seeded home returned a pool slot while its local child held the shared lock" + assert_present "$dir/home/state/$id.meta" "contended remote-layout teardown removed the task record" + assert_present "$dir/worktree/sentinel" "contended remote-layout teardown reset the slot" + [ ! -s "$dir/runtime.log" ] \ + || fail "contended remote-layout teardown reached the runtime: $(cat "$dir/runtime.log")" + assert_contains "$(cat "$dir/stderr")" "another Treehouse slot allocation or return is in progress" \ + "the refusal should name the shared project lock, not some unrelated check" + + pass "Treehouse project locking still serializes two homes across the remote-seeded boundary" +} + test_invalid_endpoint_records_refuse_before_mutation test_control_lock_contention_refuses_before_mutation test_non_pool_teardown_ignores_task_set_lock @@ -611,5 +838,9 @@ test_reused_pool_slot_refuses_before_touching_the_other_task test_cross_home_pool_slot_collision_refuses test_sole_slot_record_still_tears_down test_recorded_endpoint_that_changed_directory_still_tears_down +test_project_lock_anchors_at_the_local_root_across_home_layouts +test_remote_seeded_home_returns_its_uncontested_slot +test_remote_seeded_home_still_refuses_a_slot_its_child_holds +test_remote_layout_homes_serialize_on_one_project_lock printf '\nall fm-teardown-endpoint-safety tests passed\n' diff --git a/tests/fm-teardown.test.sh b/tests/fm-teardown.test.sh index d1ad8b96eba..a9177e1620c 100755 --- a/tests/fm-teardown.test.sh +++ b/tests/fm-teardown.test.sh @@ -38,6 +38,10 @@ # (o) fm-pr-check rerun after HEAD moved -> no stale pr_head # (p) fm-pr-check when local HEAD lags -> record remote PR head # (q) no-mistakes + NO pr= recorded, PR discovered by branch -> ALLOW (yolo/no-CI merge) +# (q2) no-mistakes + squash-merged, local followed pipeline rebase -> ALLOW +# (q3) no-mistakes + squash-merged, same file, different content -> REFUSE +# (q4) no-mistakes + squash-merged rebased local plus extra commit -> REFUSE +# (q5) gh down + squash-merged stale local, content not in default -> REFUSE # # Also covers task identity: the pool can hand one worktree path to a second # task, so the worktree's ambient branch is placement, not identity. Both the @@ -385,6 +389,94 @@ SH chmod +x "$case_dir/fakebin/gh-axi" } +# Squash-merged history whose pipeline rebased the branch onto a newer main that +# edited the same shared file. A local copy left behind by that rebase holds +# different content for the shared file, so its per-commit patch ids against the +# PR head differ and merge-tree against main conflicts; teardown refuses it on +# purpose rather than reading a shared path as proof the local content landed. +# local_mode: rebased | stale | rebased-plus-unlanded +# Echoes: <pr_head> +setup_squash_rebased_history() { + local case_dir=$1 local_mode=$2 tmp local_head pr_head + tmp="$case_dir/_shared_base" + git clone -q "$case_dir/origin.git" "$tmp" + printf '%s\n' base > "$tmp/shared.txt" + git -C "$tmp" add -- shared.txt + git -C "$tmp" -c user.email=t@t -c user.name=t commit -q -m "shared base" + git -C "$tmp" push -q origin main + git -C "$case_dir/wt" fetch -q origin + git -C "$case_dir/wt" reset -q --hard origin/main + rm -rf "$tmp" + + wt_commit_file "$case_dir" feature.txt hello "add feature" + printf '%s\n' base feature-edit > "$case_dir/wt/shared.txt" + git -C "$case_dir/wt" add -- shared.txt + git -C "$case_dir/wt" -c user.email=t@t -c user.name=t \ + commit -q -m "edit shared from feature" + local_head=$(git -C "$case_dir/wt" rev-parse HEAD) + + tmp="$case_dir/_main_move" + git clone -q "$case_dir/origin.git" "$tmp" + printf '%s\n' base main-edit > "$tmp/shared.txt" + git -C "$tmp" add -- shared.txt + git -C "$tmp" -c user.email=t@t -c user.name=t commit -q -m "main edits shared" + git -C "$tmp" push -q origin main + rm -rf "$tmp" + + tmp="$case_dir/_pipeline" + git clone -q "$case_dir/origin.git" "$tmp" + git -C "$tmp" checkout -q -b fm/task-x1 + printf '%s\n' hello > "$tmp/feature.txt" + git -C "$tmp" add -- feature.txt + git -C "$tmp" -c user.email=t@t -c user.name=t commit -q -m "add feature" + printf '%s\n' base main-edit feature-edit > "$tmp/shared.txt" + git -C "$tmp" add -- shared.txt + git -C "$tmp" -c user.email=t@t -c user.name=t \ + commit -q -m "edit shared from feature" + pr_head=$(git -C "$tmp" rev-parse HEAD) + git -C "$tmp" push -q origin "HEAD:refs/pull/7/head" + git -C "$tmp" checkout -q main + git -C "$tmp" merge -q --squash fm/task-x1 >/dev/null + git -C "$tmp" -c user.email=t@t -c user.name=t commit -q -m "feat: squash (#7)" + git -C "$tmp" push -q origin main + rm -rf "$tmp" + + git -C "$case_dir/project" fetch -q origin + git -C "$case_dir/wt" fetch -q origin "refs/pull/7/head:refs/fm-test/pr-head" + case "$local_mode" in + rebased) + git -C "$case_dir/wt" reset -q --hard "$pr_head" + ;; + stale) + git -C "$case_dir/wt" reset -q --hard "$local_head" + ;; + rebased-plus-unlanded) + git -C "$case_dir/wt" reset -q --hard "$pr_head" + wt_commit_file "$case_dir" later.txt local-only "local follow-up" + ;; + *) + fail "setup_squash_rebased_history: unknown local_mode $local_mode" + ;; + esac + printf '%s\n' "$pr_head" +} + +# A refusal must leave every recovery route intact: the isolated copy, its task +# branch still at the unlanded commit, and the durable task record. A completed +# teardown detaches and deletes that branch and removes the record, so these hold +# only while nothing destructive ran before the refusal was reported. +# Args: case_dir label head-before-teardown +assert_refusal_retained_task_state() { + local case_dir=$1 label=$2 head=$3 + [ -d "$case_dir/wt" ] || fail "$label: refusal removed the isolated copy" + [ "$(git -C "$case_dir/wt" rev-parse --abbrev-ref HEAD 2>/dev/null)" = fm/task-x1 ] \ + || fail "$label: refusal dropped the task branch" + [ "$(git -C "$case_dir/wt" rev-parse HEAD 2>/dev/null)" = "$head" ] \ + || fail "$label: refusal moved the task branch off the unlanded commit" + [ -e "$case_dir/state/task-x1.meta" ] \ + || fail "$label: refusal erased the durable task record" +} + append_pr_meta_for_current_head() { local case_dir=$1 head head=$(git -C "$case_dir/wt" rev-parse HEAD) @@ -2239,6 +2331,98 @@ test_merged_pr_with_later_local_commit_refuses() { pass "merged PR does not allow teardown after a later local commit" } +test_squash_merged_rebased_branch_allows() { + local case_dir rc pr_head + case_dir=$(make_case squash-rebased) + write_meta "$case_dir" no-mistakes ship + pr_head=$(setup_squash_rebased_history "$case_dir" rebased) + printf '%s\n' \ + 'pr=https://github.com/example/repo/pull/7' \ + "pr_head=$pr_head" >> "$case_dir/state/task-x1.meta" + add_gh_pr_merged_for_head "$case_dir" "$pr_head" + + set +e + run_teardown "$case_dir" > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 0 "$rc" "squash-rebased: teardown should succeed when the worktree followed the pipeline rebase"$'\n'"$(cat "$case_dir/stderr")" + ! grep -q REFUSED "$case_dir/stderr" || fail "squash-rebased: teardown printed a REFUSED line" + pass "squash-merged task whose local branch followed the pipeline rebase is torn down" +} + +test_squash_merged_same_file_different_content_refuses() { + local case_dir rc pr_head local_head + case_dir=$(make_case squash-same-path-diverged) + write_meta "$case_dir" no-mistakes ship + # The pipeline rebase produced a different blob for shared.txt than the stale + # local still holds, then squash-merged. Same path is not proof the local + # content landed. + pr_head=$(setup_squash_rebased_history "$case_dir" stale) + local_head=$(git -C "$case_dir/wt" rev-parse HEAD) + printf '%s\n' \ + 'pr=https://github.com/example/repo/pull/7' \ + "pr_head=$pr_head" >> "$case_dir/state/task-x1.meta" + add_gh_pr_merged_for_head "$case_dir" "$pr_head" + + set +e + run_teardown "$case_dir" > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "squash-same-path-diverged: teardown should refuse when the same file has different content"$'\n'"$(cat "$case_dir/stderr")" + grep -q REFUSED "$case_dir/stderr" || fail "squash-same-path-diverged: no REFUSED line in stderr" + assert_refusal_retained_task_state "$case_dir" squash-same-path-diverged "$local_head" + pass "squash-merged same-path different content still refuses" +} + +# The local branch followed the pipeline rebase, so without later.txt this is the +# q2 ALLOW case exactly. The one unlanded follow-up commit is the sole difference +# and must be the sole reason teardown refuses. +test_squash_merged_rebased_local_with_unlanded_commit_refuses() { + local case_dir rc pr_head local_head + case_dir=$(make_case squash-rebased-unlanded) + write_meta "$case_dir" no-mistakes ship + pr_head=$(setup_squash_rebased_history "$case_dir" rebased-plus-unlanded) + local_head=$(git -C "$case_dir/wt" rev-parse HEAD) + printf '%s\n' \ + 'pr=https://github.com/example/repo/pull/7' \ + "pr_head=$pr_head" >> "$case_dir/state/task-x1.meta" + add_gh_pr_merged_for_head "$case_dir" "$pr_head" + + set +e + run_teardown "$case_dir" > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "squash-rebased-unlanded: teardown should refuse extra local commits that never landed"$'\n'"$(cat "$case_dir/stderr")" + grep -q REFUSED "$case_dir/stderr" || fail "squash-rebased-unlanded: no REFUSED line in stderr" + assert_refusal_retained_task_state "$case_dir" squash-rebased-unlanded "$local_head" + pass "squash-merged rebased local still refuses a genuinely unlanded follow-up commit" +} + +test_squash_merged_stale_local_refuses_when_forge_unreachable() { + local case_dir rc pr_head local_head + case_dir=$(make_case squash-stale-offline) + write_meta "$case_dir" no-mistakes ship + pr_head=$(setup_squash_rebased_history "$case_dir" stale) + local_head=$(git -C "$case_dir/wt" rev-parse HEAD) + printf '%s\n' \ + 'pr=https://github.com/example/repo/pull/7' \ + "pr_head=$pr_head" >> "$case_dir/state/task-x1.meta" + add_gh_axi_error "$case_dir" + + set +e + run_teardown "$case_dir" > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "squash-stale-offline: teardown should refuse when the forge is down and trees conflict"$'\n'"$(cat "$case_dir/stderr")" + grep -q REFUSED "$case_dir/stderr" || fail "squash-stale-offline: no REFUSED line in stderr" + assert_refusal_retained_task_state "$case_dir" squash-stale-offline "$local_head" + pass "squash-merged stale local still refuses when the forge is unreachable" +} + test_pr_check_does_not_refresh_stale_pr_head() { local case_dir rc pr_head new_head count case_dir=$(make_case pr-check-stale) @@ -2805,6 +2989,13 @@ test_non_linked_index_lock_path_is_checked_from_worktree() { test_index_lock_mtime_read_failure_refuses() { local case_dir rc lock + # The mtime fault is injected by a fake stat on PATH; on Darwin the lock + # helper now calls /usr/bin/stat directly, so the fake can never fire there. + # Skip the Darwin run of this case. + if [ "$(uname)" = Darwin ]; then + pass "index-lock mtime fault injection is PATH-based; skipped on Darwin where stat is /usr/bin/stat" + return + fi case_dir=$(make_case mtime-error-index-lock) write_meta "$case_dir" no-mistakes ship wt_commit "$case_dir" "shippable work" @@ -5078,6 +5269,10 @@ test_merged_pr_head_not_retained_by_forge_keeps_branch test_no_pr_recorded_discovers_merged_pr_by_branch_allows test_squash_merged_pr_allows_replayed_unpushed_patch test_merged_pr_with_later_local_commit_refuses +test_squash_merged_rebased_branch_allows +test_squash_merged_same_file_different_content_refuses +test_squash_merged_rebased_local_with_unlanded_commit_refuses +test_squash_merged_stale_local_refuses_when_forge_unreachable test_pr_check_does_not_refresh_stale_pr_head test_pr_check_records_remote_head_when_local_lags test_content_in_default_fallback_allows diff --git a/tests/fm-test-isolation-proof.test.sh b/tests/fm-test-isolation-proof.test.sh index 283fb1a3f90..08982149ff7 100755 --- a/tests/fm-test-isolation-proof.test.sh +++ b/tests/fm-test-isolation-proof.test.sh @@ -249,7 +249,9 @@ test_family_map_labels_this_contract() { safe_max=$("$RUNNER" --concurrent-safe-family-jobs-max pure-contract-unit) [ "$safe_max" -eq 4 ] || fail "runner exposed the wrong contract-unit family worker cap: $safe_max" scheduled_first=$("$RUNNER" --list-scheduled --family watcher-wake-lock | head -n 1) - [ "$scheduled_first" = tests/fm-watch-triage.test.sh ] \ + # The split wait suite retains the largest watcher hint after the CI refresh; + # timing ownership stays in the runner. + [ "$scheduled_first" = tests/fm-watch-triage-waits.test.sh ] \ || fail "runner scheduled the watcher family out of longest-hint order: $scheduled_first" pass "isolation-proof contract test is family-mapped" } diff --git a/tests/fm-test-run.test.sh b/tests/fm-test-run.test.sh index e5e179daef5..f78a931a918 100755 --- a/tests/fm-test-run.test.sh +++ b/tests/fm-test-run.test.sh @@ -125,6 +125,7 @@ init_changed_fixture_repo() { fm-procevent-quota.test.sh \ fm-quota-choose.test.sh \ fm-pi-watch-extension.test.sh \ + fm-pi-windows-shell-invocation.test.sh \ fm-afk-return.test.sh \ fm-bearings-snapshot.test.sh \ fm-backend-cmux.test.sh \ @@ -169,6 +170,8 @@ init_changed_fixture_repo() { : >"$repo/.claude/settings.json" : >"$repo/.pi/extensions/fm-primary-pi-watch.ts" : >"$repo/.pi/extensions/fm-primary-turnend-guard.ts" + mkdir -p "$repo/.pi/extensions/lib" + : >"$repo/.pi/extensions/lib/fm-operational-input.ts" : >"$repo/docs/fm-test-isolation-proof.md" : >"$repo/CONTRIBUTING.md" : >"$repo/src/unmapped.ts" @@ -177,6 +180,70 @@ init_changed_fixture_repo() { git -C "$repo" -c user.name=test -c user.email=test@example.invalid commit -qm baseline } +# Build a repository with a primary checkout and one linked worktree, each +# holding a runnable copy of the runner and a probe suite that records the fact +# that it ran. Untracked copies are enough: the runner resolves its root from +# its own path, and the probe is named explicitly. +init_primary_and_linked_worktree() { + local repo=$1 linked=$2 tree + fm_git_init_commit "$repo" + git -C "$repo" worktree add --quiet -b linked-probe "$linked" + for tree in "$repo" "$linked"; do + mkdir -p "$tree/bin" "$tree/tests" + cp "$RUNNER" "$tree/bin/fm-test-run.sh" + cp "$ROOT/bin/fm-timeout-lib.sh" "$tree/bin/fm-timeout-lib.sh" + chmod +x "$tree/bin/fm-test-run.sh" + cat >"$tree/tests/probe.test.sh" <<PROBE +#!/usr/bin/env bash +echo "ok - probe suite" +: >"$tree/ran" +PROBE + chmod +x "$tree/tests/probe.test.sh" + done +} + +# A task worker's isolated placement is checked once, when the task starts. +# Nothing re-checks it, so a worker that later changes directory into the +# repository's primary checkout would run this branch-switching suite in the one +# checkout every linked worktree resolves against. The runner refuses that. +test_task_marker_refuses_the_primary_checkout() { + local tmp repo linked out rc + tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run-primary.XXXXXX") + repo="$tmp/repo" + linked="$tmp/linked" + init_primary_and_linked_worktree "$repo" "$linked" + + # Marker set, primary checkout: refuse, name the primary, and run nothing. + out=$(FM_TASK_ID=probe-task "$repo/bin/fm-test-run.sh" tests/probe.test.sh 2>&1) && rc=0 || rc=$? + [ "$rc" -ne 0 ] || { rm -rf "$tmp"; fail "the runner must refuse the primary checkout under a task marker"; } + assert_contains "$out" "primary checkout" "refusal did not name the primary checkout" + assert_contains "$out" "FM_TASK_ID=probe-task" "refusal did not name the task marker" + assert_contains "$out" "task worktree" "refusal did not point at the task worktree" + assert_not_contains "$out" "FM_TEST_BEGIN" "the refusal must happen before any suite runs" + assert_absent "$repo/ran" "the refused run still executed a suite" + + # Marker set, linked worktree: the assigned placement, so the suite runs. + FM_TASK_ID=probe-task "$linked/bin/fm-test-run.sh" tests/probe.test.sh >/dev/null 2>&1 \ + || { rm -rf "$tmp"; fail "the runner must still run in a linked task worktree"; } + assert_present "$linked/ran" "the linked-worktree run did not execute its suite" + + # No marker: a person in their own checkout is unaffected. + "$repo/bin/fm-test-run.sh" tests/probe.test.sh >/dev/null 2>&1 \ + || { rm -rf "$tmp"; fail "an unmarked run in the primary checkout must be unchanged"; } + assert_present "$repo/ran" "the unmarked run did not execute its suite" + + # Inspection executes nothing, so it stays available even in the primary. + rm -f "$repo/ran" + out=$(FM_TASK_ID=probe-task "$repo/bin/fm-test-run.sh" --list tests/probe.test.sh 2>&1) \ + || { rm -rf "$tmp"; fail "--list must remain available under a task marker"; } + [ "$out" = "tests/probe.test.sh" ] \ + || { rm -rf "$tmp"; fail "--list under a task marker printed: $out"; } + assert_absent "$repo/ran" "--list must not execute a suite" + + rm -rf "$tmp" + pass "a task marker refuses execution in the primary checkout and leaves worktrees and inspection alone" +} + test_changed_runner_surfaces_select_their_family() { local tmp repo listed tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run-owner-scope.XXXXXX") @@ -223,6 +290,19 @@ test_changed_runner_surfaces_select_their_family() { pass "runner and its documentation surfaces select their curated family, not just their contract owners" } +test_shell_line_ending_policy_selects_runner_contract() { + local tmp repo listed + tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run-attributes.XXXXXX") + repo="$tmp/repo" + init_changed_fixture_repo "$repo" + printf '*.sh text eol=lf\n' >"$repo/.gitattributes" + listed=$(cd "$repo" && bin/fm-test-run.sh --list --changed --base HEAD) + assert_contains "$listed" "tests/fm-test-run.test.sh" \ + "shell line-ending policy selects the runner contract" + rm -rf "$tmp" + pass "shell line-ending policy selects runner coverage" +} + test_changed_dependency_selection_and_unmapped_failure() { local tmp repo listed rc tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run-changed.XXXXXX") @@ -273,9 +353,18 @@ test_changed_dependency_selection_and_unmapped_failure() { assert_contains "$listed" "tests/fm-ask-user-authority.test.sh" "skill source selects pure contract coverage" assert_contains "$listed" "tests/fm-cd-pretool-check.test.sh" "Claude and Pi source selects hook coverage" assert_contains "$listed" "tests/fm-pi-watch-extension.test.sh" "Pi source selects watcher coverage" + assert_contains "$listed" "tests/fm-pi-windows-shell-invocation.test.sh" \ + "turn-end extension selects native-Windows shell coverage" git -C "$repo" add .agents .claude .pi git -C "$repo" -c user.name=test -c user.email=test@example.invalid commit -qm non-bin-source-change + printf '\n' >>"$repo/.pi/extensions/lib/fm-operational-input.ts" + listed=$(cd "$repo" && bin/fm-test-run.sh --list --changed --base HEAD) + assert_contains "$listed" "tests/fm-pi-windows-shell-invocation.test.sh" \ + "operational-input extension selects native-Windows shell coverage" + git -C "$repo" add .pi/extensions/lib/fm-operational-input.ts + git -C "$repo" -c user.name=test -c user.email=test@example.invalid commit -qm operational-input-source-change + printf '\n' >>"$repo/.agents/skills/harness-adapters/references/common/dispatch.md" listed=$(cd "$repo" && bin/fm-test-run.sh --list --changed --base HEAD) assert_contains "$listed" "tests/fm-harness-adapter-references.test.sh" "harness adapter reference selects portable structural coverage" @@ -333,8 +422,12 @@ test_changed_dependency_selection_and_unmapped_failure() { [ "$rc" -eq 2 ] || fail "unmapped changed source must fail with exit 2, got $rc" grep -Fq 'no changed-test mapping for source path: src/unmapped.ts' "$tmp/err" \ || fail "unmapped changed source failure is not actionable: $(cat "$tmp/err")" + + rm -f "$repo/src/unmapped.ts" + listed=$(cd "$repo" && bin/fm-test-run.sh --list --changed --base HEAD) + [ -z "$listed" ] || fail "a retired unmapped source without consumers selected tests: $listed" rm -rf "$tmp" - pass "changed selection covers dependents and fails closed for unmapped source" + pass "changed selection covers dependents, fails closed for live unmapped source, and accepts retired unconsumed source" } # A direct test reference is per-script evidence. Widening it to the referencing @@ -450,6 +543,50 @@ SH pass "changed defaults to bounded automatic scheduling with serial override" } +test_windows_posix_mode_emulation_does_not_fail_parallel_runs() { + local tmp repo fakebin real_stat out rc + tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run-windows-modes.XXXXXX") + repo="$tmp/repo" + fakebin="$tmp/fakebin" + real_stat=$(command -v stat) + init_changed_fixture_repo "$repo" + mkdir -p "$fakebin" + cat >"$fakebin/uname" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "${FAKE_UNAME:-MINGW64_NT-10.0}" +SH + cat >"$fakebin/stat" <<'SH' +#!/usr/bin/env bash +if [ "${1:-}" = -c ] && [ "${2:-}" = %a ]; then + printf '%s\n' 755 + exit 0 +fi +exec "$REAL_STAT" "$@" +SH + chmod +x "$fakebin/uname" "$fakebin/stat" + set +e + out=$(cd "$repo" && PATH="$fakebin:$PATH" REAL_STAT="$real_stat" \ + bin/fm-test-run.sh --jobs 2 \ + tests/fm-cd-pretool-check.test.sh tests/fm-ask-user-authority.test.sh 2>&1) + rc=$? + set -e + expect_code 0 "$rc" "native-Windows POSIX-mode emulation" + assert_contains "$out" "FM_TEST_SUMMARY total=2 failed=0" \ + "Windows mode emulation did not complete both parallel scripts" + + set +e + out=$(cd "$repo" && PATH="$fakebin:$PATH" REAL_STAT="$real_stat" FAKE_UNAME=CYGWIN_NT-10.0 \ + bin/fm-test-run.sh --jobs 2 \ + tests/fm-cd-pretool-check.test.sh tests/fm-ask-user-authority.test.sh 2>&1) + rc=$? + set -e + expect_code 1 "$rc" "Cygwin POSIX-mode enforcement" + assert_contains "$out" "isolation failure: worker root mode is 755, expected 0700" \ + "Cygwin mode enforcement did not reject a non-0700 worker root" + rm -rf "$tmp" + pass "Windows emulation exempts only synthetic POSIX modes" +} + # A local verification round names the subjects it cares about. Exercise begin/end # markers from real fixture processes to prove that a plain list of script paths # gets bounded automatic scheduling without changing its per-script timeout @@ -740,6 +877,80 @@ assert doc["summary"]["failed"] == 0 pass "gate-skip accounting is honest and non-failing" } +test_gate_skip_reason_is_recorded() { + local tmp skip_f out json + tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run-skipreason.XXXXXX") + skip_f="$tmp/skip.test.sh" + out="$tmp/out.txt" + json="$tmp/timing.json" + cat >"$skip_f" <<'SH' +#!/usr/bin/env bash +echo "skip: live: fmnosuchharness absent" +exit 0 +SH + chmod +x "$skip_f" + "$RUNNER" --json "$json" "$skip_f" >"$out" 2>"$tmp/err.txt" \ + || fail "a capability skip must still exit 0 from the runner" + grep -q 'live: fmnosuchharness absent' "$tmp/err.txt" \ + || fail "the runner log must name what this host could not exercise: $(cat "$tmp/err.txt")" + python3 -c ' +import json, sys +doc = json.load(open(sys.argv[1])) +record = doc["scripts"][0] +assert record["gate_skip"] is True, record +assert record["gate_skip_reason"] == "live: fmnosuchharness absent", record +' "$json" || { rm -rf "$tmp"; fail "the timing artifact must carry the skip reason"; } + rm -rf "$tmp" + pass "a gate skip records why it skipped" +} + +test_a_run_that_ran_records_no_skip_reason() { + local tmp ran_f json + tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run-ranreason.XXXXXX") + ran_f="$tmp/ran.test.sh" + json="$tmp/timing.json" + cat >"$ran_f" <<'SH' +#!/usr/bin/env bash +echo "ok - ran" +exit 0 +SH + chmod +x "$ran_f" + "$RUNNER" --json "$json" "$ran_f" >"$tmp/out.txt" 2>&1 \ + || fail "a passing fixture must exit 0 from the runner" + python3 -c ' +import json, sys +doc = json.load(open(sys.argv[1])) +record = doc["scripts"][0] +assert record["gate_skip"] is False, record +assert record["gate_skip_reason"] == "", record +' "$json" || { rm -rf "$tmp"; fail "a script that ran must carry an empty skip reason"; } + rm -rf "$tmp" + pass "a script that actually ran records no skip reason" +} + +test_live_guards_expect_a_capability_skip_class() { + local tmp out + tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run-liveclass.XXXXXX") + out="$tmp/out.txt" + # FM_LIVE=0 makes every live guard refuse without touching a harness, so this + # exercises the real family through the real runner in bounded time. + FM_LIVE=0 "$RUNNER" --json "$tmp/timing.json" \ + tests/fm-composer-matrix-live-e2e.test.sh >"$out" 2>"$tmp/err.txt" \ + || fail "a disabled live guard must not fail the runner: $(cat "$tmp/err.txt")" + grep -q 'expected_gate_skip=live-capability' "$out" \ + || fail "the live-harness family must expect a capability skip: $(grep FM_TEST_BEGIN "$out")" + python3 -c ' +import json, sys +doc = json.load(open(sys.argv[1])) +record = doc["scripts"][0] +assert record["expected_gate_skip"] == "live-capability", record +assert record["gate_skip"] is True, record +assert record["gate_skip_reason"].startswith("live: "), record +' "$tmp/timing.json" || { rm -rf "$tmp"; fail "the live guard record is wrong"; } + rm -rf "$tmp" + pass "live guards are recorded as a capability class, not a bare env opt-in" +} + test_fail_on_gate_skip_token() { local tmp skip_f out rc tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run-fail-skip.XXXXXX") @@ -799,21 +1010,20 @@ test_portable_shard_union_and_coverage_guard() { || fail "herdr family must include smoke" printf '%s\n' "$optin" | grep -Fq 'tests/fm-agy-smoke.test.sh' \ || fail "live-harness-optin must include the cursor/agy credentialed smoke" - if printf '%s\n' "$serial" | grep -Fq 'tests/fm-agy-smoke.test.sh'; then - fail "portable serial must exclude the cursor/agy credentialed smoke" - fi + printf '%s\n' "$serial" | grep -Fq 'tests/fm-agy-smoke.test.sh' \ + || fail "portable serial must include the credentialed smoke with its execution gate" out=$("$RUNNER" --check-coverage) assert_contains "$out" "FM_TEST_COVERAGE ok" "coverage guard success marker" all_count=$("$RUNNER" --list --all | wc -l | tr -d ' ') - union_count=$(printf '%s\n' "$s1" "$s2" "$serial" "$herdr" "$optin" | LC_ALL=C sort -u | wc -l | tr -d ' ') + union_count=$(printf '%s\n' "$s1" "$s2" "$serial" "$herdr" | LC_ALL=C sort -u | wc -l | tr -d ' ') [ "$union_count" = "$all_count" ] \ || fail "union of lanes ($union_count) must equal --all ($all_count)" - # No duplicates across the five partitions. - [ "$(printf '%s\n' "$s1" "$s2" "$serial" "$herdr" "$optin" | LC_ALL=C sort | uniq -d | wc -l | tr -d ' ')" = "0" ] \ + # No duplicates across the four partitions. + [ "$(printf '%s\n' "$s1" "$s2" "$serial" "$herdr" | LC_ALL=C sort | uniq -d | wc -l | tr -d ' ')" = "0" ] \ || fail "lanes must not duplicate scripts" # LPT order: first script of shard 1 is the longest proven script. first=$(printf '%s\n' "$s1" | head -n 1) - [ "$first" = "tests/fm-x-mode.test.sh" ] \ + [ "$first" = "tests/fm-captain-hold-lifecycle.test.sh" ] \ || fail "shard 1 must start with the longest proven script, got $first" pass "portable shard union, disjointness, and coverage guard hold" } @@ -1931,16 +2141,22 @@ test_new_adapter_consumers_are_selected() { test_new_adapter_consumers_are_selected test_changed_file_selection_is_conservative +test_task_marker_refuses_the_primary_checkout test_changed_runner_surfaces_select_their_family +test_shell_line_ending_policy_selects_runner_contract test_changed_dependency_selection_and_unmapped_failure test_changed_bin_reference_selects_per_script_not_per_family test_changed_uses_bounded_automatic_concurrency +test_windows_posix_mode_emulation_does_not_fail_parallel_runs test_script_list_uses_bounded_automatic_concurrency test_family_proofs_run_in_separate_concurrent_phases test_empty_selection_emits_summary test_timing_markers_and_json test_aggregate_exit_behavior test_gate_skip_accounting +test_gate_skip_reason_is_recorded +test_a_run_that_ran_records_no_skip_reason +test_live_guards_expect_a_capability_skip_class test_fail_on_gate_skip_token test_exclude_family test_portable_shard_union_and_coverage_guard diff --git a/tests/fm-tool-update-check.test.sh b/tests/fm-tool-update-check.test.sh index 630eb860ae9..31449ebddb7 100755 --- a/tests/fm-tool-update-check.test.sh +++ b/tests/fm-tool-update-check.test.sh @@ -1040,15 +1040,15 @@ test_findings_are_reported_once_until_they_change() { } test_an_overlong_report_says_it_was_cut() { - local home out report i tools= + local home out report i tools_json= # Many watched tools can outgrow one line. The report must say it was cut # rather than end mid-finding as if that were everything found. home=$(make_home long) for i in $(seq 1 30); do - [ -z "$tools" ] || tools="$tools," - tools="$tools{\"name\":\"absent-tool-$i\",\"command\":\"fm-absent-fixture-$i\"}" + [ -z "$tools_json" ] || tools_json="$tools_json," + tools_json="$tools_json{\"name\":\"absent-tool-$i\",\"command\":\"fm-absent-fixture-$i\"}" done - write_config "$home" "{\"tools\":[$tools]}" + write_config "$home" "{\"tools\":[$tools_json]}" out="$home/out.txt" run_check "$home" "$PATH" "$out" report=$(cat "$out") @@ -1058,7 +1058,7 @@ test_an_overlong_report_says_it_was_cut() { } test_a_finding_past_the_cut_is_still_reported() { - local home stale fresh out report i tools= + local home stale fresh out report i tools_json= # Once a report is long enough to be cut, a new finding lands past the cut and # leaves the printed line unchanged. It still has to count as news, or the PATH # skew this check exists for would be suppressed for good on a busy home. @@ -1068,17 +1068,17 @@ test_a_finding_past_the_cut_is_still_reported() { make_copy "$stale" "$TOOL" 'herdr 0.8.0' make_copy "$fresh" "$TOOL" 'herdr 0.8.2' for i in $(seq 1 30); do - [ -z "$tools" ] || tools="$tools," - tools="$tools{\"name\":\"absent-tool-$i\",\"command\":\"fm-absent-fixture-$i\"}" + [ -z "$tools_json" ] || tools_json="$tools_json," + tools_json="$tools_json{\"name\":\"absent-tool-$i\",\"command\":\"fm-absent-fixture-$i\"}" done out="$home/out.txt" - write_config "$home" "{\"tools\":[$tools]}" + write_config "$home" "{\"tools\":[$tools_json]}" run_check "$home" "$(fixture_path "$stale:$fresh")" "$out" assert_contains "$(cat "$out")" "[truncated]" "the first report was not long enough to be cut, so this case proves nothing" # The skew tool goes last, so its finding falls past the cut and the printed # line is byte identical to the one the first sweep already recorded. - write_config "$home" "{\"tools\":[$tools,{\"name\":\"herdr\",\"command\":\"$TOOL\"}]}" + write_config "$home" "{\"tools\":[$tools_json,{\"name\":\"herdr\",\"command\":\"$TOOL\"}]}" run_check "$home" "$(fixture_path "$stale:$fresh")" "$out" report=$(cat "$out") [ -n "$report" ] || fail "a finding past the cut produced no report at all, so it can never reach the watcher" diff --git a/tests/fm-turnend-guard.test.sh b/tests/fm-turnend-guard.test.sh index 45696532df6..961ae1769ab 100755 --- a/tests/fm-turnend-guard.test.sh +++ b/tests/fm-turnend-guard.test.sh @@ -125,6 +125,73 @@ test_predicate_source_needs_supervision() { pass "fm_supervision_unhealthy: source-only home needs supervision" } +# Register a custom check the way an operator does, through the real +# bin/fm-check-register.sh, so these cases bind to the shipped registration +# artifacts rather than to a hand-written imitation of them. +register_custom_check() { + local state=$1 id=$2 + printf '#!/usr/bin/env bash\nexit 0\n' > "$state/$id.check.sh" + chmod 700 "$state/$id.check.sh" + FM_STATE_OVERRIDE="$state" "$ROOT/bin/fm-check-register.sh" "$id" >/dev/null \ + || fail "fm-check-register.sh could not register $id" +} + +test_predicate_registered_check_needs_supervision() { + local state="$TMP_ROOT/pred-check/state" + mkdir -p "$state" + register_custom_check "$state" issue-comments + fm_supervision_needed "$state" 300 || fail "a registered custom check did not register as supervision need" + [ "$FM_SUP_IN_FLIGHT" -eq 0 ] || fail "a registered custom check must not count as an in-flight task" + [ "$FM_SUP_CHECKS" -eq 1 ] || fail "expected one registered custom check, got $FM_SUP_CHECKS" + fm_supervision_unhealthy "$state" 300 || fail "a registered custom check with no beacon must be unhealthy" + pass "fm_supervision_needed: a registered custom check needs supervision with no task in flight" +} + +test_predicate_registered_check_survives_rebinding_drift() { + local state="$TMP_ROOT/pred-check-drift/state" + mkdir -p "$state" + register_custom_check "$state" issue-comments + printf '#!/usr/bin/env bash\necho drifted\n' > "$state/issue-comments.check.sh" + fm_supervision_needed "$state" 300 \ + || fail "an edited registered check must keep supervision on so the sweep can report the rejection" + [ "$FM_SUP_CHECKS" -eq 1 ] || fail "expected the edited check to stay counted, got $FM_SUP_CHECKS" + pass "fm_supervision_needed: a registered check whose bytes drifted still needs supervision" +} + +test_predicate_unregistered_check_needs_nothing() { + local state="$TMP_ROOT/pred-check-unregistered/state" + mkdir -p "$state" + printf '#!/usr/bin/env bash\nexit 0\n' > "$state/rogue.check.sh" + chmod 700 "$state/rogue.check.sh" + if fm_supervision_needed "$state" 300; then + fail "a check with no trust binding must not arm supervision" + fi + [ "$FM_SUP_CHECKS" -eq 0 ] || fail "an unregistered check must not be counted, got $FM_SUP_CHECKS" + pass "fm_supervision_needed: false for a check.sh with no registration binding" +} + +test_predicate_task_pr_poll_is_not_a_custom_check() { + local state="$TMP_ROOT/pred-pr-poll/state" + mkdir -p "$state" + : > "$state/task1.meta" + printf '#!/usr/bin/env bash\nexit 0\n' > "$state/task1.check.sh" + chmod 700 "$state/task1.check.sh" + : > "$state/task1.pr-poll" + fm_supervision_needed "$state" 300 || fail "the in-flight task itself must need supervision" + [ "$FM_SUP_IN_FLIGHT" -eq 1 ] || fail "expected the task to be the one in-flight need, got $FM_SUP_IN_FLIGHT" + [ "$FM_SUP_CHECKS" -eq 0 ] || fail "a task PR poll must not count as a registered custom check" + pass "fm_supervision_needed: a task PR poll without a trust binding is not a registered custom check" +} + +test_predicate_relay_shim_is_not_a_custom_check() { + local state="$TMP_ROOT/pred-relay-not-custom/state" + mkdir -p "$state" + : > "$state/x-watch.check.sh" + fm_supervision_needed "$state" 300 || fail "the relay poll must still need supervision" + [ "$FM_SUP_CHECKS" -eq 0 ] || fail "the relay shim keeps its own trust path and must not be counted as a custom check" + pass "fm_supervision_status: the relay shim is not counted as a registered custom check" +} + # --- HOOK: bin/fm-turnend-guard.sh ------------------------------------------ # # Each scenario gets its own directory carrying a copy of the two guard scripts @@ -444,6 +511,17 @@ test_hook_x_mode_only_blocks_in_default_mode() { pass "fm-turnend-guard: X-mode-only supervision remains guarded in default mode" } +test_hook_registered_check_only_blocks_with_check_banner() { + local dir out status + dir=$(make_primary_dir "$TMP_ROOT/hook-check-only") + register_custom_check "$dir/state" issue-comments + out=$(run_hook "$dir" false); status=$? + expect_code 2 "$status" "default hook mode must block a registered-check-only blind turn" + assert_contains "$out" "1 registered custom check(s), but no live watcher" "check-only blind stop must identify its supervision need" + assert_not_contains "$out" "X-mode relay polling needs supervision" "check-only blind stop must not be misreported as relay polling" + pass "fm-turnend-guard: registered-check-only supervision is named in the block banner" +} + test_hook_ignores_repo_state_when_fm_home_set() { local dir home out status dir=$(make_primary_dir "$TMP_ROOT/hook-fm-home-ignore-root") @@ -1965,6 +2043,95 @@ test_hook_daemon_lock_is_ignored_without_away_mode() { pass "fm-turnend-guard: a daemon lock proves nothing while away mode is off" } +# --- AWAY MODE: beacon grace derives from the poll cadence ------------------- +# +# The daemon starts a fresh one-shot watcher only after it finishes handling +# the previous wake, and that handling can legitimately outrun a flat 300s +# window under load (a slow registered check, a busy supervisor pane) with the +# daemon perfectly healthy throughout (fm-turnend-guard-afk-race). The guard +# must accept a live daemon there once FM_POLL justifies the wider window, but +# must still block a dead daemon or a beacon older than that wider grace. + +test_hook_away_daemon_allows_beacon_within_poll_derived_grace() { + local dir pid out status beat + dir=$(make_away_home_between_cycles "$TMP_ROOT/hook-afk-poll-grace-healthy") + sleep 60 & + pid=$! + record_daemon_lock "$dir" "$pid" || { + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + fail "could not identify live away-mode daemon holder" + } + # 400s is stale under the flat 300s default, but not under the poll-derived + # grace (max(300, FM_POLL + 60) = 660 at FM_POLL=600) - a live daemon that + # simply has not finished restarting its watcher yet. + beat=$(( $(date +%s) - 400 )) + touch -d "@$beat" "$dir/state/.last-watcher-beat" + out=$(FM_GUARD_GRACE='' FM_POLL=600 run_hook "$dir" false); status=$? + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + expect_code 0 "$status" "a live daemon with a beacon within the poll-derived grace must not block" + [ -z "$out" ] || fail "away-mode daemon within poll-derived grace still produced a block banner: $out" + pass "fm-turnend-guard: away-mode beacon freshness uses the poll-derived grace, not the flat default" +} + +test_hook_away_daemon_blocks_dead_daemon_despite_poll_derived_grace() { + local dir dead out status + dir=$(make_away_home_between_cycles "$TMP_ROOT/hook-afk-poll-grace-dead-daemon") + dead=$(nonexistent_pid) + record_daemon_lock "$dir" "$dead" "dead daemon identity" + out=$(FM_GUARD_GRACE='' FM_POLL=600 run_hook "$dir" false); status=$? + expect_code 2 "$status" "a wider poll-derived grace must not paper over a dead daemon" + assert_contains "$out" "$AWAY_REQUIRED_REASON" "away-mode block must point at the daemon, not normal supervision" + pass "fm-turnend-guard: a dead away-mode daemon still blocks under the poll-derived grace" +} + +test_hook_away_daemon_blocks_beacon_older_than_poll_derived_grace() { + local dir pid out status beat + dir=$(make_away_home_between_cycles "$TMP_ROOT/hook-afk-poll-grace-stale") + sleep 60 & + pid=$! + record_daemon_lock "$dir" "$pid" || { + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + fail "could not identify live away-mode daemon holder" + } + # 700s exceeds even the wider poll-derived grace (660 at FM_POLL=600), so a + # live daemon that has genuinely stopped restarting its watcher still blocks. + beat=$(( $(date +%s) - 700 )) + touch -d "@$beat" "$dir/state/.last-watcher-beat" + out=$(FM_GUARD_GRACE='' FM_POLL=600 run_hook "$dir" false); status=$? + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + expect_code 2 "$status" "a beacon older than the poll-derived grace must still block" + assert_contains "$out" "$AWAY_REQUIRED_REASON" "away-mode block must point at the daemon, not normal supervision" + pass "fm-turnend-guard: the poll-derived grace is bounded, not unlimited" +} + +test_hook_no_afk_ignores_poll_derived_grace() { + local dir pid out status beat + dir=$(make_away_home_between_cycles "$TMP_ROOT/hook-no-afk-poll-grace") + rm -f "$dir/state/.afk" + sleep 60 & + pid=$! + record_daemon_lock "$dir" "$pid" || { + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + fail "could not identify live daemon holder" + } + # 400s would be within the poll-derived grace the away-mode branch would + # accept, but away mode is off here, so the strict watcher predicate and its + # flat default govern instead - old behavior, unaffected by FM_POLL. + beat=$(( $(date +%s) - 400 )) + touch -d "@$beat" "$dir/state/.last-watcher-beat" + out=$(FM_GUARD_GRACE='' FM_POLL=600 run_hook "$dir" false); status=$? + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + expect_code 2 "$status" "without .afk, FM_POLL must not widen the strict watcher predicate's grace" + assert_contains "$out" "$REQUIRED_REASON" "block reason must contain the exact required instruction" + pass "fm-turnend-guard: with away mode off, the poll-derived grace never applies" +} + test_predicate_healthy_no_inflight test_predicate_unhealthy_no_beacon test_predicate_unhealthy_stale_beacon @@ -1974,6 +2141,11 @@ test_predicate_queue_pending_needs_supervision test_predicate_queue_pending_drained_no_longer_needs_supervision test_predicate_x_mode_needs_supervision test_predicate_source_needs_supervision +test_predicate_registered_check_needs_supervision +test_predicate_registered_check_survives_rebinding_drift +test_predicate_unregistered_check_needs_nothing +test_predicate_task_pr_poll_is_not_a_custom_check +test_predicate_relay_shim_is_not_a_custom_check test_hook_silent_when_no_work_in_flight test_hook_blocks_when_fresh_beacon_has_no_live_lock test_hook_blocks_source_only_home @@ -1985,6 +2157,7 @@ test_hook_blocks_when_unhealthy_in_primary test_hook_blocks_from_fm_home_state test_hook_x_mode_reason_sources_cadence test_hook_x_mode_only_blocks_in_default_mode +test_hook_registered_check_only_blocks_with_check_banner test_hook_queue_only_block_names_the_pending_queue test_hook_ignores_repo_state_when_fm_home_set test_hook_uses_state_override @@ -2046,4 +2219,8 @@ test_hook_away_mode_blocks_on_dead_daemon test_hook_away_mode_blocks_on_pid_reused_daemon test_hook_away_mode_blocks_on_stale_beacon test_hook_daemon_lock_is_ignored_without_away_mode +test_hook_away_daemon_allows_beacon_within_poll_derived_grace +test_hook_away_daemon_blocks_dead_daemon_despite_poll_derived_grace +test_hook_away_daemon_blocks_beacon_older_than_poll_derived_grace +test_hook_no_afk_ignores_poll_derived_grace printf '\nall fm-turnend-guard tests passed\n' diff --git a/tests/fm-wake-queue.test.sh b/tests/fm-wake-queue.test.sh index e1bf46c9d4f..57461a198f6 100755 --- a/tests/fm-wake-queue.test.sh +++ b/tests/fm-wake-queue.test.sh @@ -14,6 +14,7 @@ set -u WATCH="$ROOT/bin/fm-watch.sh" DRAIN="$ROOT/bin/fm-wake-drain.sh" GRANT="$ROOT/bin/fm-wake-grant.sh" +GUARD="$ROOT/bin/fm-guard.sh" TMP_ROOT=$(fm_test_tmproot fm-wake-tests) @@ -231,8 +232,8 @@ test_drain_dedupes_obvious_duplicates() { # plain drain-and-handle turn that runs no other supervision script. It must warn # when work is in flight with no live watcher, and stay silent right after a # normal fire from a live watcher with a fresh beacon, so it never false-alarms. -test_secondmate_foreign_queue_stall_is_one_shot_and_read_only() { - local dir state sub fakebin out row_before row_after stall_count +test_secondmate_foreign_queue_stall_tracks_progress_and_alerts_once() { + local dir state sub fakebin out row_before row_after stall_count real_date dir=$(make_case secondmate-foreign-stall) state="$dir/state" sub="$dir/secondmate" @@ -241,35 +242,58 @@ test_secondmate_foreign_queue_stall_is_one_shot_and_read_only() { printf 'mate\n' > "$sub/.fm-secondmate-home" printf 'window=firstmate:fm-mate\nkind=secondmate\nharness=claude\nbackend=tmux\nhome=%s\n' \ "$sub" > "$state/mate.meta" - printf '%s\t7\tcheck\trouted\tcheck: routed row\n' "$(( $(date +%s) - 10 ))" > "$sub/state/.wake-queue" - row_before="$dir/foreign-before" - row_after="$dir/foreign-after" - cp "$sub/state/.wake-queue" "$row_before" fakebin="$dir/fakebin" - cat > "$fakebin/tmux" <<'SH' + real_date=$(command -v date) + cat > "$fakebin/date" <<SH #!/usr/bin/env bash -case "${1:-}" in - list-windows) printf '%s\n' "${FM_FAKE_TMUX_WINDOW:-}" ;; - capture-pane) cat "${FM_FAKE_TMUX_CAPTURE:-/dev/null}" ;; - display-message) printf '0\n' ;; - *) exit 0 ;; -esac +if [ "\${1:-}" = +%s ]; then + cat "\${FM_FAKE_NOW_FILE:?}" +else + exec "$real_date" "\$@" +fi SH - chmod +x "$fakebin/tmux" - out="$dir/watch.out" + chmod +x "$fakebin/date" - PATH="$fakebin:$PATH" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + # An already-old row starts an observation interval; its creation time alone + # cannot produce an alert. + printf '1000\n' > "$dir/now" + printf '100\t7\tcheck\trouted\tcheck: routed row\n' > "$sub/state/.wake-queue" + PATH="$fakebin:$PATH" FM_FAKE_NOW_FILE="$dir/now" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ - FM_FAKE_TMUX_LOG="$dir/tmux.log" FM_FAKE_TMUX_CAPTURE="$dir/fake-tmux/pane.txt" \ FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ - "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 3 > "$out" 2> "$dir/watch.err" || true - grep -F 'check: secondmate wake-loop stalled: mate=mate row=7' "$out" >/dev/null \ - || fail "an aged foreign row did not wake the parent checkpoint: $(cat "$out"); err=$(cat "$dir/watch.err"); meta=$(cat "$state/mate.meta"); foreign=$(cat "$sub/state/.wake-queue")" - [ -s "$state/.wake-queue" ] || fail "the parent notification was not durable" - stall_count=$(grep -c 'secondmate-wake-loop-mate-' "$state/.wake-queue" || true) - [ "$stall_count" -eq 1 ] || fail "the first parent checkpoint did not publish exactly one stall notification" + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 1 > "$dir/watch-first.out" 2> "$dir/watch-first.err" || true + [ ! -s "$state/.wake-queue" ] \ + || fail "the first observation of an old foreign row produced an age-only alert" + + # The oldest sequence advances after more than the threshold. This is healthy + # drain progress even though the replacement row is itself very old. + printf '1002\n' > "$dir/now" + printf '100\t8\tcheck\thealthy\tcheck: healthy progress\n' > "$sub/state/.wake-queue" + PATH="$fakebin:$PATH" FM_FAKE_NOW_FILE="$dir/now" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ + FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 1 > "$dir/watch-progress.out" 2> "$dir/watch-progress.err" || true + [ ! -s "$state/.wake-queue" ] \ + || fail "an advancing foreign queue produced a stall alert because its oldest row was old" + # With no further sequence progress, the same queue must still expose the real + # failure after the configured interval. + printf '1004\n' > "$dir/now" + row_before="$dir/foreign-before" + row_after="$dir/foreign-after" + cp "$sub/state/.wake-queue" "$row_before" + out="$dir/watch-stalled.out" + PATH="$fakebin:$PATH" FM_FAKE_NOW_FILE="$dir/now" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ + FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 1 > "$out" 2> "$dir/watch-stalled.err" || true + grep -F 'check: secondmate wake-loop stalled: mate=mate row=8 idle=2s' "$out" >/dev/null \ + || fail "a foreign queue with no progress did not alert: $(cat "$out")" + stall_count=$(grep -c 'secondmate-wake-loop-mate-' "$state/.wake-queue" || true) + [ "$stall_count" -eq 1 ] || fail "the stalled episode did not publish exactly one parent notification" cmp -s "$row_before" "$sub/state/.wake-queue" \ || fail "foreign queue row changed during read-only stall detection" FM_STATE_OVERRIDE="$state" "$DRAIN" > "$dir/drain.out" 2> "$dir/drain.err" \ @@ -277,40 +301,193 @@ SH ack_drain_err "$state" "$dir/drain.err" \ || fail "parent stall notification could not be acknowledged" - sleep 1 - PATH="$fakebin:$PATH" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + # Partial draining changes the oldest row, ends the prior no-progress episode, + # and cannot produce an immediate notification cascade. + printf '1010\n' > "$dir/now" + printf '100\t9\tcheck\tnext\tcheck: next row\n' > "$sub/state/.wake-queue" + PATH="$fakebin:$PATH" FM_FAKE_NOW_FILE="$dir/now" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ - FM_FAKE_TMUX_LOG="$dir/tmux.log" FM_FAKE_TMUX_CAPTURE="$dir/fake-tmux/pane.txt" \ FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ - "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 2 > "$dir/watch-second.out" 2> "$dir/watch-second.err" || true - [ ! -s "$state/.wake-queue" ] || { - stall_count=$(grep -c 'secondmate-wake-loop-mate-' "$state/.wake-queue" || true) - [ "$stall_count" -eq 0 ] || fail "repeated checkpoint re-published the same stall notification" - } + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 1 > "$dir/watch-next.out" 2> "$dir/watch-next.err" || true + [ ! -s "$state/.wake-queue" ] \ + || fail "a newly-oldest row cascaded an immediate second alert after progress" cp "$sub/state/.wake-queue" "$row_after" - cmp -s "$row_before" "$row_after" || fail "foreign queue changed after idempotent re-check" + grep -F $'\t9\t' "$row_after" >/dev/null || fail "foreign queue progress fixture changed during observation" - : > "$sub/state/.wake-queue" - PATH="$fakebin:$PATH" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + # If that new drain position then genuinely stops advancing, it is a new + # no-progress episode and must remain visible rather than being muted forever. + printf '1012\n' > "$dir/now" + PATH="$fakebin:$PATH" FM_FAKE_NOW_FILE="$dir/now" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ - FM_FAKE_TMUX_LOG="$dir/tmux.log" FM_FAKE_TMUX_CAPTURE="$dir/fake-tmux/pane.txt" \ FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ - "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 2 > "$dir/watch-empty.out" 2> "$dir/watch-empty.err" || true - ! grep -F 'secondmate wake-loop stalled' "$dir/watch-empty.out" >/dev/null \ - || fail "an empty foreign queue produced a stall notification" + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 1 > "$dir/watch-refrozen.out" 2> "$dir/watch-refrozen.err" || true + grep -F 'check: secondmate wake-loop stalled: mate=mate row=9 idle=2s' "$dir/watch-refrozen.out" >/dev/null \ + || fail "a genuine later no-progress episode was hidden after earlier progress" + stall_count=$(grep -c 'secondmate-wake-loop-mate-' "$state/.wake-queue" || true) + [ "$stall_count" -eq 1 ] || fail "the later no-progress episode did not publish exactly one notification" + pass "foreign secondmate queue alerts once per no-progress episode without age-only or cascade noise" +} - printf '%s\t8\tcheck\thealthy\tcheck: healthy row\n' "$(date +%s)" > "$sub/state/.wake-queue" - PATH="$fakebin:$PATH" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ +test_secondmate_declared_pause_rows_do_not_feed_stall_escalation() { + local dir state sub fakebin real_date + dir=$(make_case secondmate-declared-pause-queue) + state="$dir/state" + sub="$dir/secondmate" + mkdir -p "$sub/state" + printf 'mate\n' > "$sub/.fm-secondmate-home" + printf 'window=firstmate:fm-mate\nkind=secondmate\nhome=%s\n' "$sub" > "$state/mate.meta" + fakebin="$dir/fakebin" + real_date=$(command -v date) + cat > "$fakebin/date" <<SH +#!/usr/bin/env bash +if [ "\${1:-}" = +%s ]; then + cat "\${FM_FAKE_NOW_FILE:?}" +else + exec "$real_date" "\$@" +fi +SH + chmod +x "$fakebin/date" + cat > "$sub/state/.wake-queue" <<'EOF' +100 7 stale fleet:w2:p4 stale: fleet:w2:p4 (paused 3613s, awaiting external - declared paused) +100 8 stale fleet:w2:p3 stale: fleet:w2:p3 (paused 3615s, awaiting external - declared pause, rechecked on a long cadence not a wedge) +EOF + printf '1000\n' > "$dir/now" + PATH="$fakebin:$PATH" FM_FAKE_NOW_FILE="$dir/now" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ - FM_FAKE_TMUX_LOG="$dir/tmux.log" FM_FAKE_TMUX_CAPTURE="$dir/fake-tmux/pane.txt" \ - FM_SECONDMATE_WAKE_STALL_SECS=60 FM_POLL=1 FM_SIGNAL_GRACE=0 \ + FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 1 > "$dir/watch-first.out" 2> "$dir/watch-first.err" || true + printf '5000\n' > "$dir/now" + PATH="$fakebin:$PATH" FM_FAKE_NOW_FILE="$dir/now" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ + FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 1 > "$dir/watch-second.out" 2> "$dir/watch-second.err" || true + [ ! -s "$state/.wake-queue" ] \ + || fail "declared external-wait rows fed the secondmate wake-loop escalation" + ! grep -F 'secondmate wake-loop stalled' "$dir/watch-first.out" "$dir/watch-second.out" >/dev/null \ + || fail "a declared external wait was mislabeled as a stalled wake loop" + pass "declared external-wait pause rows do not feed secondmate wake-loop escalation" +} + +# A retired mate reprovisioned under the same task id gets a fresh home, so its +# wake-queue sequence restarts from scratch and can land on the very position the +# parent last recorded for the retired generation. Those are different rows in +# different queue generations, not the continuation of the previous generation's +# no-progress interval: inheriting that interval fires a wake-loop stall against +# a queue the mate has only just created. +test_secondmate_reprovisioned_queue_starts_a_fresh_interval() { + local dir state sub fakebin real_date + dir=$(make_case secondmate-reprovisioned-queue) + state="$dir/state" + sub="$dir/secondmate" + mkdir -p "$sub/state" + printf 'mate\n' > "$sub/.fm-secondmate-home" + printf 'window=firstmate:fm-mate\nkind=secondmate\nharness=claude\nbackend=tmux\nhome=%s\n' \ + "$sub" > "$state/mate.meta" + fakebin="$dir/fakebin" + real_date=$(command -v date) + cat > "$fakebin/date" <<SH +#!/usr/bin/env bash +if [ "\${1:-}" = +%s ]; then + cat "\${FM_FAKE_NOW_FILE:?}" +else + exec "$real_date" "\$@" +fi +SH + chmod +x "$fakebin/date" + + # The retired generation's last observation records sequence 9. + printf '1000\n' > "$dir/now" + printf '100\t9\tcheck\told\tcheck: retired generation row\n' > "$sub/state/.wake-queue" + PATH="$fakebin:$PATH" FM_FAKE_NOW_FILE="$dir/now" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ + FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ - "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 2 > "$dir/watch-healthy.out" 2> "$dir/watch-healthy.err" || true - ! grep -F 'secondmate wake-loop stalled' "$dir/watch-healthy.out" >/dev/null \ - || fail "a healthy foreign queue produced a stall notification" - pass "foreign secondmate queue stalls notify once, remain byte-stable, and stay quiet when empty or healthy" + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 1 > "$dir/watch-old.out" 2> "$dir/watch-old.err" || true + [ ! -s "$state/.wake-queue" ] || fail "the first observation of the retired generation alerted" + + # Reprovisioning under the same task id restarts the sequence on 9 again, long + # after the recorded observation. That first sight of the new queue cannot + # inherit the old generation's idle interval. + printf '1010\n' > "$dir/now" + printf '200\t9\tcheck\tregen\tcheck: reprovisioned row\n' > "$sub/state/.wake-queue" + PATH="$fakebin:$PATH" FM_FAKE_NOW_FILE="$dir/now" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ + FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 1 > "$dir/watch-regen.out" 2> "$dir/watch-regen.err" || true + [ ! -s "$state/.wake-queue" ] \ + || fail "a reprovisioned queue generation inherited the retired generation's idle interval and alerted" + + # The restarted generation still earns its own honest no-progress episode. + printf '1012\n' > "$dir/now" + PATH="$fakebin:$PATH" FM_FAKE_NOW_FILE="$dir/now" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ + FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 1 > "$dir/watch-regen-frozen.out" 2> "$dir/watch-regen-frozen.err" || true + grep -F 'check: secondmate wake-loop stalled: mate=mate row=9 idle=2s' "$dir/watch-regen-frozen.out" >/dev/null \ + || fail "a frozen reprovisioned queue generation was hidden: $(cat "$dir/watch-regen-frozen.out")" + pass "a reprovisioned queue generation starts a fresh no-progress interval" +} + +# A healthy mate drains its wake queue BETWEEN turns, not inside one, so a queue +# that has not advanced while the mate is provably mid-turn is not a stalled wake +# loop - it is the normal state of a busy mate, and the measured false alarms +# (ages 63s, 75s, 70s) all landed here. The active-turn gate must DEFER that +# escalation, not cancel it: the same frozen queue still has to surface once the +# turn ends. +test_secondmate_active_turn_defers_stall_until_the_turn_ends() { + local dir state sub fakebin stall_count + dir=$(make_case secondmate-active-turn) + state="$dir/state" + sub="$dir/secondmate" + mkdir -p "$sub/state" + printf 'mate\n' > "$sub/.fm-secondmate-home" + printf 'window=firstmate:fm-mate\nkind=secondmate\nharness=claude\nbackend=tmux\nhome=%s\n' \ + "$sub" > "$state/mate.meta" + printf '%s\t7\tcheck\trouted\tcheck: routed row\n' "$(( $(date +%s) - 10 ))" \ + > "$sub/state/.wake-queue" + fakebin="$dir/fakebin" + cat > "$fakebin/tmux" <<'SH' +#!/usr/bin/env bash +case "${1:-}" in + list-windows) printf '%s\n' 'firstmate:fm-mate' ;; + capture-pane) printf 'working\n' ;; + display-message) printf '0\n' ;; + *) exit 0 ;; +esac +SH + chmod +x "$fakebin/tmux" + "$ROOT/bin/fm-busy-event.sh" arm "$state" mate >/dev/null \ + || fail "could not arm the mate's busy contract" + + PATH="$fakebin:$PATH" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 \ + FM_SIGNAL_GRACE=0 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 4 \ + > "$dir/watch-busy.out" 2> "$dir/watch-busy.err" || true + ! grep -F 'secondmate wake-loop stalled' "$dir/watch-busy.out" >/dev/null \ + || fail "a mate inside an active turn was escalated as a stalled wake loop: $(cat "$dir/watch-busy.out")" + [ ! -s "$state/.wake-queue" ] \ + || fail "a mate inside an active turn published a durable stall notification" + + "$ROOT/bin/fm-busy-event.sh" apply "$state" mate idle --current-gen \ + --source claude-hook --event stop >/dev/null \ + || fail "could not end the mate's turn" + PATH="$fakebin:$PATH" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 \ + FM_SIGNAL_GRACE=0 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 4 \ + > "$dir/watch-idle.out" 2> "$dir/watch-idle.err" || true + grep -F 'check: secondmate wake-loop stalled: mate=mate row=7' "$dir/watch-idle.out" >/dev/null \ + || fail "the same frozen queue stayed hidden after the turn ended: $(cat "$dir/watch-idle.out")" + stall_count=$(grep -c 'secondmate-wake-loop-mate-' "$state/.wake-queue" || true) + [ "$stall_count" -eq 1 ] || fail "the deferred episode did not publish exactly one notification" + pass "an active turn defers the secondmate stall escalation without cancelling it" } # Seed one endpoint-recorded local secondmate whose foreign wake queue holds a @@ -319,7 +496,7 @@ SH # contract unarmed so the classifier reads unknown, which is what an # unconverted, never-armed, or stale-record mate looks like. seed_stall_mate() { # <case-name> <age-seconds> <busy-state|none> - local name=$1 age=$2 busy=$3 dir state sub + local name=$1 age=$2 busy=$3 dir state sub real_date now epoch dir=$(make_case "$name") state="$dir/state" sub="$dir/secondmate" @@ -328,7 +505,8 @@ seed_stall_mate() { # <case-name> <age-seconds> <busy-state|none> printf 'mate\n' > "$sub/.fm-secondmate-home" printf 'window=firstmate:fm-mate\nkind=secondmate\nharness=claude\nbackend=tmux\nhome=%s\n' \ "$sub" > "$state/mate.meta" - printf '%s\t7\tcheck\trouted\tcheck: routed row\n' "$(( $(date +%s) - age ))" \ + epoch=$(( $(date +%s) - age )) + printf '%s\t7\tcheck\trouted\tcheck: routed row\n' "$epoch" \ > "$sub/state/.wake-queue" # The shared stub answers only the pane_current_command form of # display-message, so the endpoint would read gone and every verdict would @@ -360,6 +538,28 @@ SH printf 'v1 gen=stall-gen seq=1 state=%s source=fm-spawn event=turn ts=%s\n' \ "$busy" "$(date +%s)" > "$state/mate.busy-state" fi + # Seed the observed progress through its owner rather than timing out a + # setup watcher: interruption legitimately leaves a recovery obligation that + # must surface before the later checkpoint can be quiet. Separate episode + # tests exercise the first observation, and recovery tests retain that wake. + real_date=$(command -v date) + now=$(date +%s) + FM_STATE_OVERRIDE="$state" bash -c ' + # shellcheck disable=SC1090,SC1091 + . "$1" + fm_wake_secondmate_progress_marker_write mate "$2" "$3" + ' _ "$ROOT/bin/fm-wake-lib.sh" "$now" "$epoch-7" \ + || fail "could not seed the observed queue position" + cat > "$dir/fakebin/date" <<SH +#!/usr/bin/env bash +if [ "\${1:-}" = +%s ]; then + cat "$dir/now" +else + exec "$real_date" "\$@" +fi +SH + chmod +x "$dir/fakebin/date" + printf '%s\n' "$((now + age))" > "$dir/now" printf '%s\n' "$dir" } @@ -388,7 +588,7 @@ test_busy_secondmate_does_not_alarm_on_an_aged_row() { local dir status=0 dir=$(seed_stall_mate secondmate-stall-busy 120 busy) run_stall_checkpoint "$dir" FM_SECONDMATE_WAKE_STALL_SECS=1 || status=$? - expect_code 124 "$status" "busy secondmate quiet checkpoint" + expect_code 124 "$status" "busy secondmate quiet checkpoint: $(cat "$dir/watch.out") $(cat "$dir/watch.err")" ! grep -F 'secondmate wake-loop stalled' "$dir/watch.out" >/dev/null \ || fail "a provably busy mate alarmed on a row its own turn had not finished handling: $(cat "$dir/watch.out")" [ ! -s "$dir/state/.wake-queue" ] \ @@ -478,19 +678,35 @@ SH pass "a reused Zellij pane cannot suppress its former mate's stall alarm" } +# A busy record cannot suppress a genuinely frozen queue indefinitely. +# The controlled clock ages both the observation and the active-turn boundary, +# while leaving the semantic verdict busy so this cannot pass as an idle case. +test_secondmate_busy_suppression_is_bounded() { + local dir status=0 + dir=$(seed_stall_mate secondmate-stall-busy-bound 4000 busy) + run_stall_checkpoint "$dir" || status=$? + expect_code 0 "$status" "over-age busy turn must permit a stall inspection" + grep -F 'state=busy' "$dir/watch.out" >/dev/null \ + || fail "bounded suppression did not retain the established busy verdict" + [ -s "$dir/state/.wake-queue" ] || fail "bounded busy suppression lost the durable notification" + pass "busy suppression expires without treating a busy verdict as idle" +} + test_secondmate_stall_marker_rejects_symlink() { - local dir state sub fakebin marker outside expected + local dir state sub fakebin marker outside expected epoch dir=$(make_case secondmate-stall-marker-symlink) state="$dir/state" sub="$dir/secondmate" mkdir -p "$sub/state" printf 'mate\n' > "$sub/.fm-secondmate-home" printf 'window=firstmate:fm-mate\nkind=secondmate\nhome=%s\n' "$sub" > "$state/mate.meta" - printf '%s\t7\tcheck\trouted\tcheck: routed row\n' "$(( $(date +%s) - 10 ))" > "$sub/state/.wake-queue" + epoch=$(( $(date +%s) - 10 )) + printf '%s\t7\tcheck\trouted\tcheck: routed row\n' "$epoch" > "$sub/state/.wake-queue" outside="$dir/outside" expected='must remain unchanged' printf '%s\n' "$expected" > "$outside" marker="$state/.secondmate-wake-stall-mate" + printf '%s\t%s-7\n' "$(( $(date +%s) - 2 ))" "$epoch" > "$state/.secondmate-wake-progress-mate" ln -s "$outside" "$marker" fakebin="$dir/fakebin" cat > "$fakebin/tmux" <<'SH' @@ -528,8 +744,9 @@ test_acknowledged_stall_publication_survives_pre_marker_crash() { printf '%s\t7\tcheck\trouted\tcheck: routed row\n' "$epoch" > "$sub/state/.wake-queue" row_before="$dir/foreign-before" cp "$sub/state/.wake-queue" "$row_before" + printf '%s\t%s-7\n' "$(( $(date +%s) - 2 ))" "$epoch" > "$state/.secondmate-wake-progress-mate" append_wake "$state" check "secondmate-wake-loop-mate-$epoch-7" \ - "check: secondmate wake-loop stalled: mate=mate row=7 age=10s" \ + "check: secondmate wake-loop stalled: mate=mate row=7 idle=2s" \ || fail "could not seed the pre-marker crash publication" FM_STATE_OVERRIDE="$state" "$DRAIN" > "$dir/drain.out" 2> "$dir/drain.err" \ || fail "pre-marker crash publication could not be drained" @@ -569,8 +786,9 @@ test_empty_prefix_mate_preserves_other_mate_receipt() { printf '%s\t9\tcheck\trouted\tcheck: routed row\n' "$epoch" > "$stalled/state/.wake-queue" row_before="$dir/foreign-before" cp "$stalled/state/.wake-queue" "$row_before" + printf '%s\t%s-9\n' "$(( $(date +%s) - 2 ))" "$epoch" > "$state/.secondmate-wake-progress-ios-ui" append_wake "$state" check "secondmate-wake-loop-ios-ui-$epoch-9" \ - "check: secondmate wake-loop stalled: mate=ios-ui row=9 age=10s" \ + "check: secondmate wake-loop stalled: mate=ios-ui row=9 idle=2s" \ || fail "could not seed the ios-ui stall publication" FM_STATE_OVERRIDE="$state" "$DRAIN" > "$dir/drain.out" 2> "$dir/drain.err" \ || fail "ios-ui stall publication could not be drained" @@ -868,6 +1086,164 @@ test_main_drain_excludes_rows_already_granted_to_branch() { pass "main drain and acknowledgement exclude an active branch grant" } +# The pending-warning condition and what a drain can actually present must name +# the same rows. A row reserved by a live branch grant is invisible to a main +# drain by design, so counting it as "queued for main" told main to run a drain +# that could only print nothing - no row, no acknowledgement command - on every +# guarded command, for as long as the branch held the grant. +test_main_is_never_told_to_drain_rows_only_the_branch_owns() { + local dir state out err sequence generation + dir=$(make_case main-not-told-to-drain-branch-rows) + state="$dir/state" + printf 'window=test:fm-x\nkind=ship\n' > "$state/x.meta" + + append_wake "$state" stale "fleet:w2:p3" "stale: fleet:w2:p3 (paused, awaiting external)" \ + || fail "stale append failed" + FM_STATE_OVERRIDE="$state" "$GRANT" activate "$$" held-by-branch || fail "branch owner activation failed" + FM_STATE_OVERRIDE="$state" "$GRANT" publish held-by-branch 1 || fail "branch grant publication failed" + + out="$dir/main-drain.out" + err="$dir/main-drain.err" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$out" 2> "$err" || fail "main drain failed: $(cat "$err")" + ! grep -Fq "$(printf '\tstale\tfleet:w2:p3\t')" "$out" || fail "main drain presented a branch-granted row" + grep -Fq 'WAKE ROWS HELD BY SUPERVISION BRANCH' "$out" \ + || fail "main drain went silent instead of naming who holds the queued rows" + ! grep -Fq 'WAKE_ACK_REQUIRED' "$err" || fail "main drain offered an acknowledgement for a row it never presented" + ! grep -Fq 'queued wakes pending' "$err" \ + || fail "main was told to drain rows only the branch can present" + FM_STATE_OVERRIDE="$state" "$GUARD" 2> "$dir/guard-held.err" || fail "guard failed while the branch held the rows" + ! grep -Fq 'queued wakes pending' "$dir/guard-held.err" \ + || fail "guard counted branch-held rows as pending for main" + grep -Fq 'wake rows held by the live supervision branch' "$dir/guard-held.err" \ + || fail "guard went silent about a non-empty queue instead of naming the branch as its holder" + grep -Fq 'do not drain them from here' "$dir/guard-held.err" \ + || fail "the held advisory did not say the rows must not be drained from here" + grep -Fq "$(printf '\tstale\tfleet:w2:p3\t')" "$state/.wake-queue" \ + || fail "the branch-held row must stay durable for its own owner" + + # Disconfirming half: the same row, same kind, same stopped endpoint, with the + # grant released. Nothing about the row makes it unpresentable - only the + # live grant did - so main now presents it with an executable acknowledgement. + FM_STATE_OVERRIDE="$state" "$GRANT" release held-by-branch || fail "branch grant release failed" + out="$dir/main-drain-after.out" + err="$dir/main-drain-after.err" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$out" 2> "$err" || fail "main drain failed after release: $(cat "$err")" + grep -Fq "$(printf '\tstale\tfleet:w2:p3\t')" "$out" || fail "main drain omitted the released row" + ! grep -Fq 'WAKE ROWS HELD BY SUPERVISION BRANCH' "$out" \ + || fail "main drain reported a hold that no longer exists" + sequence=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through \([0-9][0-9]*\) --recovery-generation [A-Za-z0-9._-][A-Za-z0-9._-]*$/\1/p' "$err") + generation=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through [0-9][0-9]* --recovery-generation \([A-Za-z0-9._-][A-Za-z0-9._-]*\)$/\1/p' "$err") + [ -n "$sequence" ] && [ -n "$generation" ] || fail "the released row was presented without an acknowledgement command" + grep -Fq 'queued wakes pending' "$err" || fail "guard stopped warning about a row main can actually drain" + ! grep -Fq 'wake rows held by the live supervision branch' "$err" \ + || fail "guard kept advising about a hold that was already released" + FM_STATE_OVERRIDE="$state" "$DRAIN" --ack-through "$sequence" --recovery-generation "$generation" \ + || fail "acknowledgement of the released row failed" + [ ! -s "$state/.wake-queue" ] || fail "the acknowledged row stayed queued" + + pass "a branch-held row raises no queued-wake warning for main, and the same row is presented and acknowledged once the grant clears" +} + +# The pending-warning condition must also survive a queue nobody could read: a +# queue that exists but cannot be counted is not evidence that it was drained. +# The per-actor count runs awk over the queue, and awk implementations differ on +# whether a failed input open aborts before the END rule; one that reaches END +# reports a 0 count for a queue that was never proved empty. +test_uncountable_queue_still_raises_the_pending_alarm() { + local dir state awkbin real_awk + dir=$(make_case uncountable-queue) + state="$dir/state" + awkbin="$dir/awkbin" + mkdir -p "$awkbin" + printf 'window=test:fm-x\nkind=ship\n' > "$state/x.meta" + + # An awk that still runs its END rule after failing to open its input: it + # prints a 0 count and exits non-zero. Every other invocation is the real awk. + real_awk=$(command -v awk) || fail "no awk on PATH" + cat > "$awkbin/awk" <<SH +#!/usr/bin/env bash +set -u +for _arg in "\$@"; do _last=\$_arg; done +if [ -n "\${_last:-}" ] && [ -e "\$_last" ] && [ ! -r "\$_last" ]; then + printf '0\\n' + exit 2 +fi +exec "$real_awk" "\$@" +SH + chmod +x "$awkbin/awk" + + append_wake "$state" stale "fleet:w2:p3" "stale: fleet:w2:p3 (paused, awaiting external)" \ + || fail "stale append failed" + chmod 000 "$state/.wake-queue" || fail "could not make the queue unreadable" + PATH="$awkbin:$PATH" FM_STATE_OVERRIDE="$state" "$GUARD" 2> "$dir/unreadable.err" \ + || fail "guard failed on an unreadable queue" + grep -Fq 'queued wakes pending' "$dir/unreadable.err" \ + || fail "a queue that could not be counted silenced the queued-wake alarm" + chmod 600 "$state/.wake-queue" || fail "could not restore the queue" + + # Disconfirming half: the same fake awk over a queue that is readable and + # provably empty stays silent, so the warning above came from the failed count + # and not from the fake awk itself. + : > "$state/.wake-queue" + PATH="$awkbin:$PATH" FM_STATE_OVERRIDE="$state" "$GUARD" 2> "$dir/empty.err" \ + || fail "guard failed on an empty queue" + ! grep -Fq 'queued wakes pending' "$dir/empty.err" \ + || fail "a provably empty queue raised the queued-wake alarm" + + pass "a queue that cannot be counted keeps the queued-wake alarm up" +} + +# A row that lost its structure can never be claimed, presented, or named by an +# --ack-through cutoff, while it still counts as queued: without retirement it +# wedges the queue permanently and keeps waking supervision. +test_unconsumable_rows_are_retired_instead_of_wedging_the_queue() { + local dir state out err sequence generation + dir=$(make_case unconsumable-row-retirement) + state="$dir/state" + printf 'window=test:fm-x\nkind=ship\n' > "$state/x.meta" + + append_wake "$state" signal "task-a.status" "signal: task-a" || fail "signal append failed" + printf '1788792074\t574\tstale\tfleet:w2:p3\n' >> "$state/.wake-queue" + printf '1788792075\tnot-a-sequence\tstale\tfleet:w2:p4\tstale: fleet:w2:p4\n' >> "$state/.wake-queue" + + # A branch actor never repairs the queue: it may only touch its own grant. + FM_STATE_OVERRIDE="$state" "$GRANT" activate "$$" retire-scope || fail "branch owner activation failed" + FM_STATE_OVERRIDE="$state" "$GRANT" publish retire-scope 1 || fail "branch grant publication failed" + FM_STATE_OVERRIDE="$state" FM_SUPERVISION_ACTOR=branch "$DRAIN" > "$dir/branch.out" 2> "$dir/branch.err" \ + || fail "branch drain failed: $(cat "$dir/branch.err")" + ! grep -Fq 'retired' "$dir/branch.err" || fail "a branch drain retired rows outside its grant" + [ "$(awk 'END { print NR }' "$state/.wake-queue")" -eq 3 ] \ + || fail "a branch drain changed rows it was never granted" + FM_STATE_OVERRIDE="$state" "$GRANT" release retire-scope || fail "branch grant release failed" + + FM_STATE_OVERRIDE="$state" "$GUARD" 2> "$dir/guard-before.err" || fail "guard failed with unusable rows queued" + grep -Fq 'queued wakes pending' "$dir/guard-before.err" \ + || fail "guard stayed silent about rows main still has to clear" + ! grep -Fq 'wake rows held by the live supervision branch' "$dir/guard-before.err" \ + || fail "guard advised a branch hold for rows no grant covers" + + out="$dir/main.out" + err="$dir/main.err" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$out" 2> "$err" || fail "main drain failed: $(cat "$err")" + grep -Fq 'retired 2 unusable queue row(s)' "$err" || fail "main drain did not report the rows it retired" + grep -Fq "$(printf '1788792074\t574\tstale\tfleet:w2:p3')" "$err" \ + || fail "the retired row's content was discarded instead of reported" + grep -Fq "$(printf '1788792075\tnot-a-sequence\tstale\tfleet:w2:p4\tstale: fleet:w2:p4')" "$err" \ + || fail "the second retired row's content was discarded instead of reported" + grep -Fq "$(printf '\tsignal\ttask-a.status\t')" "$out" || fail "retirement dropped a usable row" + [ "$(awk 'END { print NR }' "$state/.wake-queue")" -eq 1 ] || fail "unusable rows survived the drain" + sequence=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through \([0-9][0-9]*\) --recovery-generation [A-Za-z0-9._-][A-Za-z0-9._-]*$/\1/p' "$err") + generation=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through [0-9][0-9]* --recovery-generation \([A-Za-z0-9._-][A-Za-z0-9._-]*\)$/\1/p' "$err") + [ -n "$sequence" ] && [ -n "$generation" ] || fail "the usable row was presented without an acknowledgement command" + FM_STATE_OVERRIDE="$state" "$DRAIN" --ack-through "$sequence" --recovery-generation "$generation" \ + || fail "acknowledgement failed" + [ ! -s "$state/.wake-queue" ] || fail "the queue stayed wedged after acknowledgement" + FM_STATE_OVERRIDE="$state" "$GUARD" 2> "$dir/guard-after.err" || fail "guard failed after the queue drained" + ! grep -Fq 'queued wakes pending' "$dir/guard-after.err" || fail "guard kept warning about an empty queue" + + pass "structurally unusable rows are retired by main alone, leaving every remaining row presentable and acknowledgeable" +} + test_branch_grant_refuses_rows_already_claimed_by_main() { local dir state rc dir=$(make_case branch-refuses-main-claim) @@ -1742,11 +2118,15 @@ test_subshell_lock_ownership_without_bashpid test_bounded_lock_handoff_after_contention test_live_presentation_holder_is_deadlined_without_weakening_ack test_malformed_presentation_lock_reports_acquire_failure -test_secondmate_foreign_queue_stall_is_one_shot_and_read_only +test_secondmate_foreign_queue_stall_tracks_progress_and_alerts_once +test_secondmate_declared_pause_rows_do_not_feed_stall_escalation +test_secondmate_reprovisioned_queue_starts_a_fresh_interval +test_secondmate_active_turn_defers_stall_until_the_turn_ends test_busy_secondmate_does_not_alarm_on_an_aged_row test_idle_secondmate_alarms_below_the_time_backstop test_unknown_busy_state_alarms_only_on_the_time_backstop test_reused_zellij_pane_cannot_suppress_a_stall_alarm +test_secondmate_busy_suppression_is_bounded test_secondmate_stall_marker_rejects_symlink test_acknowledged_stall_publication_survives_pre_marker_crash test_empty_prefix_mate_preserves_other_mate_receipt @@ -1765,6 +2145,9 @@ test_enrichment_preserves_all_unread_lines_and_status_file_failures test_slow_annotation_does_not_block_append_and_deleted_file_fails_open test_branch_actor_scoped_ack_never_swallows_a_main_owned_row test_main_drain_excludes_rows_already_granted_to_branch +test_main_is_never_told_to_drain_rows_only_the_branch_owns +test_uncountable_queue_still_raises_the_pending_alarm +test_unconsumable_rows_are_retired_instead_of_wedging_the_queue test_branch_grant_refuses_rows_already_claimed_by_main test_actor_filter_precedes_same_key_deduplication test_main_reclaims_a_grant_whose_branch_owner_exited diff --git a/tests/fm-watch-triage-waits.test.sh b/tests/fm-watch-triage-waits.test.sh new file mode 100755 index 00000000000..b777267028e --- /dev/null +++ b/tests/fm-watch-triage-waits.test.sh @@ -0,0 +1,46 @@ +#!/usr/bin/env bash +# Declared-wait, decision and bounded busy-state watcher regressions. +# Case definitions and fixtures have one owner in watch-triage-helpers.sh. +# Separate serial CI units retain every case inside the per-script bound. +set -u + +# shellcheck source=tests/watch-triage-helpers.sh +. "$(dirname "${BASH_SOURCE[0]}")/watch-triage-helpers.sh" + +test_status_is_paused_classifier +test_explicit_decision_view_preserves_the_durable_fold +test_parked_decision_survives_pane_repaint +test_resolved_decision_can_reopen_identically +test_new_keyed_decision_on_parked_pane_surfaces +test_parked_decision_with_dead_agent_wedge_escalates +test_busy_pane_progressing_run_holds_then_escalates_past_the_cap +test_busy_pane_below_turn_age_bound_is_absorbed +test_busy_pane_stable_hash_escalates_past_turn_age_bound +test_busy_pane_changing_hash_escalates_past_turn_age_bound +test_busy_pane_turn_end_touch_resets_age +test_busy_pane_repeated_escalation_reaches_demand_deep_inspection +test_busy_pane_default_turn_age_bound_is_3600s +test_busy_declared_pause_is_rechecked_not_wedge_escalated +test_afk_busy_declared_pause_hands_off_plain_stale +test_afk_busy_declared_pause_ticking_pane_hands_off_once +test_nonterminal_stale_paused_absorbed_then_resurfaced +test_pause_resurface_window_backs_off_and_caps +test_pause_streak_bump_reconciles_a_changed_wait +test_paused_resurface_backs_off_while_wedge_still_escalates +test_exited_declared_pause_is_bounded_but_live_captain_held_surfaces +test_live_declared_pause_is_absorbed_on_the_designed_cadence +test_declared_pause_landing_mid_classification_never_emits_a_bare_stale +test_live_declared_pause_still_recoverable_once_its_agent_dies +test_absorbed_replacement_wait_does_not_inherit_the_old_throttle +test_live_declared_wait_churn_honors_the_resurface_throttle +test_open_captain_call_bounds_stale_churn +test_stale_churn_without_a_captain_call_still_alarms +test_reheld_captain_call_starts_its_own_resurface_window +test_secondmate_paused_resurfaces_in_normal_mode +test_secondmate_nonpaused_stale_remains_suppressed +test_secondmate_unpause_clears_pause_tracking +test_nonterminal_stale_pause_transitions_reclassify_unchanged_hash +test_nonterminal_paused_rechecks_authoritative_state +test_paused_authoritative_working_preserves_wedge_timer +test_afk_paused_changed_pane_hands_off_plain_stale +printf '\nall fm-watch-triage-waits tests passed\n' diff --git a/tests/fm-watch-triage.test.sh b/tests/fm-watch-triage.test.sh index 743f8e79ab6..654a8afd71a 100755 --- a/tests/fm-watch-triage.test.sh +++ b/tests/fm-watch-triage.test.sh @@ -1,5298 +1,11 @@ #!/usr/bin/env bash -# tests/fm-watch-triage.test.sh - the always-on wake triage built into -# bin/fm-watch.sh and the shared classifier (bin/fm-classify-lib.sh). The watcher -# now absorbs the benign majority of wakes in bash and exits ONLY on an actionable -# wake, so firstmate's LLM re-arms once per actionable event instead of once per -# wake. These tests cover the classifier predicates as pure functions, then drive -# a real fm-watch.sh subprocess to assert the behavioral contract: -# provably-working no-verb wakes absorbed (no exit, no queue entry, suppressor -# advanced, beacon fresh), stopped-crew no-verb wakes surfaced (queue + exit), -# provably-working stale panes absorbed-then-escalated past the threshold, -# terminal-looking stale status lines overridden by an active run, the heartbeat -# backstop fail-safe, and afk coherence (no double-triage while the away-mode -# daemon owns supervision). -# -# Daemon-side classification/injection lives in fm-daemon.test.sh; watcher/lock -# liveness in fm-watcher-lock.test.sh; the durable-queue safety matrix in -# fm-wake-queue.test.sh. +# Core signal, wedge, process-event and watcher triage regressions. +# Case definitions and fixtures have one owner in watch-triage-helpers.sh. +# Separate serial CI units retain every case inside the per-script bound. set -u -# shellcheck source=tests/wake-helpers.sh -. "$(dirname "${BASH_SOURCE[0]}")/wake-helpers.sh" -# shellcheck source=/dev/null -. "$ROOT/bin/fm-classify-lib.sh" - -WATCH="$ROOT/bin/fm-watch.sh" -DRAIN="$ROOT/bin/fm-wake-drain.sh" - -TMP_ROOT=$(fm_test_tmproot fm-watch-triage-tests) - -ack_stopped_cycle() { # <state> - local state=$1 err sequence generation - err="$state/.test-cycle-drain.err" - FM_STATE_OVERRIDE="$state" "$DRAIN" >/dev/null 2> "$err" || return 1 - sequence=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through \([0-9][0-9]*\) --recovery-generation [A-Za-z0-9._-][A-Za-z0-9._-]*$/\1/p' "$err") - generation=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through [0-9][0-9]* --recovery-generation \([A-Za-z0-9._-][A-Za-z0-9._-]*\)$/\1/p' "$err") - rm -f "$err" - [ -n "$sequence" ] && [ -n "$generation" ] || return 1 - FM_STATE_OVERRIDE="$state" "$DRAIN" --ack-through "$sequence" \ - --recovery-generation "$generation" -} - -# Common watcher knobs: tight poll/grace, no check or heartbeat cadence unless a -# test overrides them, so a test only exercises the path it targets. FM_CREW_STATE_BIN -# points at the case's hermetic fake fm-crew-state.sh (installed by make_case) so the -# absorb-only-when-provably-working triage reads a canned verdict; a test fixes that -# verdict via FM_FAKE_CREW_STATE in its environment before calling watch_bg. -watch_bg() { # <state> <fakebin> <out> [extra env assignments...] - local state=$1 fakebin=$2 out=$3 - shift 3 - PATH="$fakebin:$PATH" FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$@" "$WATCH" > "$out" & -} - -# Wait up to <limit> 0.1s ticks while <pid> stays alive; 0 if still alive, 1 if it died. -wait_live() { - local pid=$1 limit=${2:-30} i=0 - while [ "$i" -lt "$limit" ]; do - kill -0 "$pid" 2>/dev/null || return 1 - sleep 0.1 - i=$((i + 1)) - done - return 0 -} - -# Wait until <pid>'s watcher has completed a whole poll cycle, or exited first. -# A fixed wait_live budget only proves the process is still ALIVE: fm-watch.sh -# does bounded startup work (the recovery-marker snapshot, lock acquisition) -# before its first stale scan, so on a loaded -# machine a short fixed budget can reap a round before the cycle it asserts on -# ever ran - and then every "no wake, no marker" assertion passes vacuously -# while every "marker written" assertion fails spuriously. -# The liveness beacon is touched at the TOP of every poll, so this drops any -# beacon left by an earlier round, waits for THIS watcher to write a fresh one -# (some poll's top), then waits for that one to advance (the next poll's top) - -# and the whole cycle in between is what the caller's assertions describe. -# 0 if the watcher is still alive after a completed cycle, 1 if it exited. -wait_poll_cycle() { # <state> <pid> [limit-ticks] - local state=$1 pid=$2 limit=${3:-300} beat first now i=0 - beat="$state/.last-watcher-beat" - rm -f "$beat" - first="" - while [ "$i" -lt "$limit" ]; do - kill -0 "$pid" 2>/dev/null || return 1 - first=$(file_mtime "$beat") - [ -n "$first" ] && break - sleep 0.1 - i=$((i + 1)) - done - while [ "$i" -lt "$limit" ]; do - kill -0 "$pid" 2>/dev/null || return 1 - now=$(file_mtime "$beat") - if [ -n "$now" ] && [ "$now" != "$first" ]; then - return 0 - fi - sleep 0.1 - i=$((i + 1)) - done - return 1 -} - -# Every wait_for_exit budget in this file is 100 ticks (10s), not because any -# watcher takes that long to decide, but because fm-watch.sh does bounded -# startup work before its first poll: a tighter budget reaps the process while -# it is still starting and reports a spurious "did not surface" failure. A -# generous budget can only remove that false negative - a watcher that never -# exits still fails the assertion when the budget runs out. -wait_numeric_file() { - local file=$1 limit=${2:-30} i=0 value - while [ "$i" -lt "$limit" ]; do - value=$(cat "$file" 2>/dev/null || true) - case "$value" in - ''|*[!0-9]*) ;; - *) return 0 ;; - esac - sleep 0.1 - i=$((i + 1)) - done - return 1 -} - -# Portable mtime in epoch seconds. Platform-detected, never the `stat -f || stat -c` -# fallback (which writes a partial filesystem dump on Linux; see fm-watch.sh). -file_mtime() { - if [ "$(uname)" = Darwin ]; then stat -f %m "$1" 2>/dev/null; else stat -c %Y "$1" 2>/dev/null; fi -} - -# Set <file>'s mtime to exactly <epoch> seconds, for aging a busy-turn marker by -# a precise amount (touch -t takes a local-time stamp, not an epoch, on both -# platforms, so convert via BSD `date -r` or GNU `date -d @`). -set_mtime() { # <epoch> <file> - local epoch=$1 f=$2 stamp - if stamp=$(date -r "$epoch" +%Y%m%d%H%M.%S 2>/dev/null); then - touch -t "$stamp" "$f" - else - stamp=$(date -d "@$epoch" +%Y%m%d%H%M.%S) - touch -t "$stamp" "$f" - fi -} - -# Signature a primed .seen-* marker must hold so the per-poll signal scan does not -# fire on a pre-existing status (mirrors fm-watch.sh's stat_sig exactly). -seen_sig() { - local reported size ident - case "$1" in - *.status) - reported=$(status_observed_signature "$1") - size=$(size_of "$1") - ident=$(_fm_open_decisions_file_ident "$1") - printf 'v2\t%s\t%s@%s' "$reported" "$size" "$ident" - ;; - *) - if [ "$(uname)" = Darwin ]; then stat -f '%z:%Fm' "$1" 2>/dev/null; else stat -c '%s:%Y' "$1" 2>/dev/null; fi - ;; - esac -} - -# Prime <file>'s .seen-* suppressor to its CURRENT signature, so the per-poll -# no-verb signal scan (which watches every *.turn-ended for a size:mtime change) -# treats a just-created or just-backdated turn-ended marker as already seen. -# Busy-turn-age fixtures create/backdate turn-ended directly (there is no real -# harness touching it), so without this the marker's own first sighting would -# fire an unrelated "signal:" wake and mask the busy-turn-age assertion under -# test. Call again after any further touch/set_mtime on the same file. -prime_turnend_seen() { # <file> - local f=$1 base - base=$(basename "$f" | tr '.' '_') - printf '%s' "$(seen_sig "$f")" > "$(dirname "$f")/.seen-$base" -} - -record_pi_busy() { # <state-dir> <id> - local state=$1 id=$2 gen - gen=$("$ROOT/bin/fm-busy-event.sh" arm "$state" "$id") - "$ROOT/bin/fm-busy-event.sh" apply "$state" "$id" busy --gen "$gen" \ - --source pi-ext --event agent-start -} - -reap() { kill "$1" 2>/dev/null || true; wait "$1" 2>/dev/null || true; } - -# --- pure classifier predicates (fm-classify-lib.sh) ------------------------ - -size_of() { LC_ALL=C wc -c < "$1" | tr -d '[:space:]'; } - -test_status_span_actionable_classifier() { - local dir state offset - dir=$(make_case classify-signal); state="$dir/state" - printf 'working: step 1\nworking: step 2\n' > "$state/a.status" - status_span_has_actionable "$state/a.status" 0 && fail "benign working: span classified actionable" - printf 'working: x\nneeds-decision: pick A or B\n' > "$state/b.status" - status_span_has_actionable "$state/b.status" 0 || fail "captain-relevant span classified benign" - # A failure and a merge result are captain-relevant and must always wake. - printf 'failed: build broke on main\n' > "$state/d.status" - status_span_has_actionable "$state/d.status" 0 || fail "a failed: line was not actionable" - printf 'merged\n' > "$state/e.status" - status_span_has_actionable "$state/e.status" 0 || fail "a legacy merged line was not actionable" - # An offset past the whole log has nothing left to classify: an event already - # classified must not re-fire on the next append. - offset=$(size_of "$state/b.status") - status_span_has_actionable "$state/b.status" "$offset" \ - && fail "an already-classified needs-decision re-fired from its own end offset" - printf 'working: tidying up\n' >> "$state/b.status" - status_span_has_actionable "$state/b.status" "$offset" \ - && fail "a routine append after a classified decision was classified actionable" - # An unusable offset (absent, malformed, or past a truncated log) reads the - # whole file rather than losing the events it cannot account for. - status_span_has_actionable "$state/b.status" "" || fail "an empty offset did not read the whole log" - status_span_has_actionable "$state/b.status" "not-a-number" || fail "a malformed offset did not read the whole log" - status_span_has_actionable "$state/b.status" 99999 || fail "an offset past the log did not read the whole log" - pass "status_span_has_actionable: benign absorbed, captain events surfaced, classified events not re-fired" -} - -# The reported bug, at the classifier: an actionable event followed by a ROUTINE -# append must stay actionable, and must be reported as ITSELF rather than as the -# routine line that happens to sit last. -test_status_span_survives_a_later_routine_append() { - local dir state event - dir=$(make_case classify-masked); state="$dir/state" - printf 'working: setup\nneeds-decision: pick A or B\nworking: still tidying the branch\n' \ - > "$state/mask.status" - status_span_has_actionable "$state/mask.status" 0 \ - || fail "a needs-decision hidden behind a later working: line was classified routine" - event=$(status_span_first_actionable "$state/mask.status" 0) - [ "$event" = "needs-decision: pick A or B" ] \ - || fail "the span reported '$event' instead of the decision it found" - # The captain-reported shape: a finished release/install reported as done and - # then followed by routine cleanup chatter must still reach the captain. - printf 'working: publishing\ndone: release 1.4.0 published and installed\nworking: cleaning the build dir\nnote: cache pruned\n' \ - > "$state/release.status" - status_span_has_actionable "$state/release.status" 0 \ - || fail "a done: completion hidden behind later routine appends was classified routine" - event=$(status_span_first_actionable "$state/release.status" 0) - [ "$event" = "done: release 1.4.0 published and installed" ] \ - || fail "the span reported '$event' instead of the completion it found" - # A blocker is the away-mode shape of the same masking. - printf 'blocked: cannot reach the release host\npaused: waiting for release access\n' \ - > "$state/blocked.status" - status_span_has_actionable "$state/blocked.status" 0 \ - || fail "a blocked: event hidden behind a current wait was classified routine" - pass "an actionable event is not hidden by later routine appends, and is named as itself" -} - -# Closure is the one thing that may retire an event inside a span, and only -# through status_open_decisions' own open/closed rule. -test_status_span_respects_decision_closure() { - local dir state event open - dir=$(make_case classify-closure); state="$dir/state" - printf 'needs-decision [key=api]: pick A or B\nresolved [key=api]: took A\n' > "$state/closed.status" - status_span_has_actionable "$state/closed.status" 0 \ - && fail "a decision the same span already closed was still classified actionable" - # Reopening the SAME key after a close must survive: the close belongs to the - # earlier opening, not to the one that came after it. - printf 'needs-decision [key=api]: pick A or B\nresolved [key=api]: took A\nneeds-decision: [key=api] pick A or B\n' \ - > "$state/reopened.status" - event=$(status_span_first_actionable "$state/reopened.status" 0) \ - || fail "a decision reopened under a key that was closed earlier was classified routine" - [ "$event" = "needs-decision: [key=api] pick A or B" ] \ - || fail "the reopened key surfaced its closed opening instead of the live reopening: $event" - # A terminal event is never retired by a later closure line. - printf 'failed: build broke on main\nresolved [key=api]: unrelated\n' > "$state/term.status" - status_span_has_actionable "$state/term.status" 0 \ - || fail "a failed: event was retired by an unrelated closure" - # A live decision must survive a NEWER closure that belongs to another key. - printf 'needs-decision [key=api]: pick A or B\nneeds-decision [key=db]: pick a store\nresolved [key=db]: took sqlite\n' \ - > "$state/two.status" - event=$(status_span_first_actionable "$state/two.status" 0) \ - || fail "a still-open decision was retired by a newer closure under another key" - [ "$event" = "needs-decision [key=api]: pick A or B" ] \ - || fail "the span reported '$event' instead of the decision still open" - printf 'needs-decision [key=pending-reply-x]: unrelated request\nworking: awaiting reconciliation\n' \ - > "$state/rejected-reserved.status" - event=$(status_span_first_actionable "$state/rejected-reserved.status" 0) \ - || fail "a rejected reserved-key request was silently dropped" - [ "$event" = "reconciliation-required: needs-decision [key=pending-reply-x]: unrelated request" ] \ - || fail "a rejected reserved-key request was not labeled for reconciliation: $event" - open=$(status_open_decisions "$state/rejected-reserved.status") - [ -z "$open" ] \ - || fail "span classification treated a rejected reserved-key request as an open decision: $open" - pass "span classification retires closed decisions and surfaces rejected transitions for reconciliation" -} - -test_malformed_seen_signature_reads_the_whole_log() { - local dir state f marker offset - dir=$(make_case malformed-seen); state="$dir/state"; f="$state/task.status" - printf 'needs-decision: choose the release target\nworking: cleanup\n' > "$f" - marker="$state/.seen-task_status" - printf '40' > "$marker" - offset=$(bash -c '. "$1"; fm_wake_signal_seen_size "$2" "$3"' _ \ - "$ROOT/bin/fm-wake-lib.sh" "$state" "$f") - [ "$offset" = 0 ] \ - || fail "a digits-only malformed seen signature was accepted as an offset" - status_span_has_actionable "$f" "$offset" \ - || fail "a malformed seen signature skipped the actionable start of the log" - pass "a malformed seen signature causes the whole status log to be classified" -} - -test_stale_is_terminal_classifier() { - local dir state - dir=$(make_case classify-stale); state="$dir/state" - printf 'done: ready in branch fm/x\n' > "$state/term.status" - stale_is_terminal "sess:fm-term" "$state" || fail "terminal stale status not classified terminal" - fm_write_meta "$state/herdr-term.meta" "window=default:w1:p2" "backend=herdr" - printf 'done: ready in branch fm/herdr\n' > "$state/herdr-term.status" - stale_is_terminal "default:w1:p2" "$state" || fail "terminal herdr stale status not resolved through metadata" - printf 'working: compiling\n' > "$state/nonterm.status" - stale_is_terminal "sess:fm-nonterm" "$state" && fail "non-terminal stale classified terminal" - stale_is_terminal "sess:fm-missing" "$state" && fail "stale with no status classified terminal" - pass "stale_is_terminal: terminal status surfaces, non-terminal and no-status are benign" -} - -test_classifier_primitives() { - local dir state open activity - dir=$(make_case classify-primitives); state="$dir/state" - printf 'working: a\n\ndone: b\n\n' > "$state/x.status" - [ "$(last_status_line "$state/x.status")" = "done: b" ] || fail "last_status_line did not return the last non-blank line" - status_is_captain_relevant "done: b" || fail "done: not recognized as captain-relevant" - status_is_captain_relevant "needs-decision [key=q1]: b" || fail "keyed needs-decision not recognized as captain-relevant" - status_is_captain_relevant "working: b" && fail "working: wrongly recognized as captain-relevant" - # Incident regression: free-text "merged" inside a nonterminal working: line must - # not become captain-relevant (AFK false-terminal path). - status_is_captain_relevant \ - "working: stage 2 setup complete on PR #74 exact source branch rebased onto merged #76; task dates preserved" \ - && fail "working: ... merged #N wrongly recognized as captain-relevant" - status_is_captain_relevant "working: rebased onto predecessor #76" \ - && fail "working: predecessor prose wrongly recognized as captain-relevant" - status_is_captain_relevant "working: PR ready checks green merged ready in branch" \ - && fail "working: free-text tokens wrongly recognized as captain-relevant" - status_is_captain_relevant "done: PR https://x/pull/76 checks green" \ - || fail "genuine done: checks green not captain-relevant" - status_is_terminal_verb "done: PR https://x/pull/76 checks green" \ - || fail "done: not a terminal verb" - status_is_terminal_verb "working: rebased onto merged #76" \ - && fail "working: wrongly classed as terminal verb" - status_is_captain_relevant "merged" || fail "legacy bare merged free-text not captain-relevant" - status_is_captain_relevant "PR ready https://x/pull/2" \ - || fail "legacy bare PR ready free-text not captain-relevant" - [ "$(window_to_task "sess:fm-fix-login-k3")" = "fix-login-k3" ] || fail "window_to_task did not strip session+fm- prefix" - fm_write_meta "$state/herdr-task.meta" "window=default:w1:p2" "backend=herdr" - [ "$(window_to_task "default:w1:p2" "$state")" = "herdr-task" ] || fail "window_to_task did not resolve opaque backend target through metadata" - FM_CAPTAIN_RE='custom-verb:' status_is_captain_relevant "custom-verb: x" || fail "FM_CAPTAIN_RE override not honored" - FM_CAPTAIN_RE='custom-verb:' status_is_captain_relevant "done: x" && fail "FM_CAPTAIN_RE override did not replace the default verb set" - FM_CAPTAIN_RE='merged|custom-verb:' status_is_captain_relevant "working: rebased onto merged #76" \ - && fail "FM_CAPTAIN_RE override bypassed working: suppression" - FM_CAPTAIN_RE='checks green|custom-verb:' status_is_captain_relevant "paused: checks green pending approval" \ - && fail "FM_CAPTAIN_RE override bypassed paused: suppression" - FM_CAPTAIN_RE='custom-verb:' status_is_captain_relevant "custom-verb: x" \ - || fail "nonterminal suppression weakened custom bare-line behavior" - printf 'needs-decision: should docs mention [key=prose]?\nneeds-decision [key=q1]: real choice\nresolved: docs still mention [key=q1]\nneeds-decision [key=bad key]: malformed\n' > "$state/keys.status" - open=$(status_open_decisions "$state/keys.status") - printf '%s' "$open" | grep -F $'q1\t' >/dev/null \ - || fail "a key token in resolved note prose closed the keyed decision" - printf '%s' "$open" | grep -F $'prose\t' >/dev/null \ - && fail "a key token in note prose changed the decision key" - printf '%s' "$open" | grep -F $'bad key\t' >/dev/null \ - && fail "an invalid key slug entered the open-decision set" - cat > "$state/activity.status" <<'EOF' -working [key=phase7]: Phase 7 started -working [key=phase6]: Phase 6 started -working [key=legal]: reviewing legal dependency -done [key=phase6]: Phase 6 completed -resolved [key=phase7]: Phase 7 completed and moved to Done -paused [key=legal]: awaiting external counsel -resolved [key=legal]: legal item returned to the queue -working [key=phase8]: Phase 8 started -EOF - activity=$(status_open_activities "$state/activity.status") - printf '%s' "$activity" | grep -F $'phase8\tworking\tPhase 8 started' >/dev/null \ - || fail "the current keyed working phase was not retained" - printf '%s' "$activity" | grep -F $'phase7\t' >/dev/null \ - && fail "a keyed resolved event did not close the older working phase" - printf '%s' "$activity" | grep -F $'phase6\t' >/dev/null \ - && fail "a same-key terminal event did not supersede the older working phase" - printf '%s' "$activity" | grep -F $'legal\t' >/dev/null \ - && fail "a keyed resolved event did not close the declared pause" - printf 'working: legacy start\ndone: legacy completion\n' > "$state/legacy-activity.status" - [ -z "$(status_open_activities "$state/legacy-activity.status")" ] \ - || fail "a legacy terminal event did not supersede the default working phase" - pass "classifier primitives: keyed decisions and activity phases, captain relevance, window-to-task, and overrides" -} - -# crew_is_provably_working: the absorb-only-when-provably-working predicate. It is -# benign (absorb) ONLY when fm-crew-state.sh reports the crew as working from a -# current or bounded-degraded pipeline step (source run-step or -# run-step-degraded) or a busy pane (source pane); -# everything else - a stale working: status-log line, a finished/parked/failed run, -# an unknown/torn-down crew, or an empty id - is NOT provable, so it surfaces. The -# fake fm-crew-state.sh (FM_CREW_STATE_BIN) returns a canned verdict per case. -test_crew_is_provably_working_classifier() { - local dir fakebin - dir=$(make_case provably-working); fakebin="$dir/fakebin" - # Point the predicate at this case's hermetic fake and drive its verdict per case. - # export marks the var for the fake subprocess; it is unset again at the end so it - # cannot leak into a later test (every behavioral test sets its own verdict anyway). - export FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" - export FM_FAKE_CREW_STATE - FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' - crew_is_provably_working a || fail "active run-step not treated as provably working" - FM_FAKE_CREW_STATE='state: working · source: pane · harness busy' - crew_is_provably_working a || fail "busy pane not treated as provably working" - FM_FAKE_CREW_STATE='state: working · source: status-log · working: compiling' - ! crew_is_provably_working a || fail "stale status-log working: treated as provably working" - FM_FAKE_CREW_STATE='state: done · source: run-step · checks green' - ! crew_is_provably_working a || fail "finished run treated as provably working" - FM_FAKE_CREW_STATE='state: parked · source: run-step · parked at review' - ! crew_is_provably_working a || fail "parked run treated as provably working" - FM_FAKE_CREW_STATE='state: failed · source: run-step · run failed' - ! crew_is_provably_working a || fail "failed run treated as provably working" - FM_FAKE_CREW_STATE='state: abandoned · source: run-step · worker gone (dead)' - ! crew_is_provably_working a || fail "abandoned run treated as provably working" - FM_FAKE_CREW_STATE='state: unknown · source: none · worktree gone' - ! crew_is_provably_working a || fail "unknown crew treated as provably working" - FM_FAKE_CREW_STATE='state: working · source: run-step · x' - ! crew_is_provably_working "" || fail "empty id treated as provably working" - unset FM_FAKE_CREW_STATE - pass "crew_is_provably_working: only working+run-step/pane is provable; idle/finished/parked/failed/unknown surface" -} - -# status_is_paused: the shared pause verb test both consumers read (so neither -# hardcodes the literal). Matches only the verb before the first colon, so a reason -# that merely mentions "paused" does not false-match, and a genuine blocker stays a -# blocker. -test_status_is_paused_classifier() { - status_is_paused 'paused: holding for the upstream release' || fail "paused verb not recognized" - status_is_paused ' paused: waiting on a rate-limit reset' || fail "leading-space paused verb not recognized" - status_is_paused 'blocked: the build is paused upstream' && fail "a blocked line mentioning paused false-matched" - status_is_paused 'working: paused the animation loop' && fail "a working line mentioning paused false-matched" - status_is_paused 'done: shipped' && fail "done classified as paused" - status_is_paused '' && fail "empty line classified as paused" - # A pause is deliberately NOT captain-relevant: it is a stop-nagging signal, not - # work to keep surfacing. - status_is_captain_relevant 'paused: holding for the upstream release' && fail "paused is captain-relevant (should not be)" - status_is_paused_or_captain_held 'paused: holding for the upstream release' \ - || fail "declared pause not recognized by the bounded-idle classifier" - status_is_paused_or_captain_held 'captain-held [key=route]: tracked by task-decision-route' \ - || fail "captain-held transfer not recognized by the bounded-idle classifier" - status_is_paused_or_captain_held 'resolved [key=route]: captain answered' \ - && fail "resolved decision remained classed as captain-held" - # The two declarations share one cadence but block on different humans, so the - # combined predicate cannot be the only discriminator: a recheck has to know which - # verb it is naming. - status_is_captain_held 'captain-held [key=route]: tracked by task-decision-route' \ - || fail "captain-held verb not recognized" - status_is_captain_held 'paused: holding for the upstream release' \ - && fail "a declared pause matched the captain-held verb" - status_is_captain_held 'working: the captain-held backlog item is next' \ - && fail "a working line mentioning captain-held false-matched" - status_is_captain_held '' && fail "empty line classified as captain-held" - pass "status_is_paused: only the leading paused verb matches, paused is not captain-relevant, and the two declared-wait verbs stay separable" -} - -# crew_absorb_class: the single fm-crew-state.sh read that returns BOTH absorb -# reasons - working (active run/busy pane), paused (declared external wait), or none -# (surface it) - so the watcher's stale path gets both for one bounded call. -# crew_is_paused delegates to it exactly as crew_is_provably_working does. -test_crew_absorb_class_classifier() { - local dir fakebin - dir=$(make_case absorb-class); fakebin="$dir/fakebin" - export FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" - export FM_FAKE_CREW_STATE - FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' - [ "$(crew_absorb_class a)" = working ] || fail "active run-step not classed working" - FM_FAKE_CREW_STATE='state: working · source: pane · harness busy' - [ "$(crew_absorb_class a)" = working ] || fail "busy pane not classed working" - FM_FAKE_CREW_STATE='state: paused · source: status-log · awaiting upstream' - [ "$(crew_absorb_class a)" = paused ] || fail "declared pause not classed paused" - crew_is_paused a || fail "crew_is_paused did not recognize a paused verdict" - ! crew_is_provably_working a || fail "a paused crew was treated as provably working" - FM_FAKE_CREW_STATE='state: working · source: status-log · working: compiling' - [ "$(crew_absorb_class a)" = none ] || fail "stale working: status-log classed absorbable" - FM_FAKE_CREW_STATE='state: unknown · source: none · worktree gone' - [ "$(crew_absorb_class a)" = none ] || fail "unknown crew classed absorbable" - ! crew_is_paused a || fail "unknown crew classed paused" - FM_FAKE_CREW_STATE='state: abandoned · source: run-step · validating (running) · worker gone (dead)' - [ "$(crew_absorb_class a)" = none ] || fail "abandoned run-step classed absorbable" - ! crew_is_provably_working a || fail "abandoned run-step treated as provably working" - FM_FAKE_CREW_STATE='state: abandoned · source: run-step-degraded · worker gone (dead)' - [ "$(crew_absorb_class a)" = none ] || fail "abandoned degraded replay classed absorbable" - [ "$(crew_absorb_class "")" = none ] || fail "empty id not classed none" - unset FM_FAKE_CREW_STATE - pass "crew_absorb_class: working/paused/none from one read; crew_is_paused and crew_is_provably_working agree" -} - -# crew_wedge_progress is the single owner of the run-progress policy both -# supervisors apply, so it is pinned as a pure decision here, independently of -# either one's escalation path: only `progressing` may produce a run-progress -# hold, a confidently dead agent never does, and everything unrecognized -# collapses to `none`. -test_crew_wedge_progress_classifier() { - export FM_FAKE_RUN_PROGRESS - - FM_FAKE_RUN_PROGRESS='progress: progressing · test running, last activity 2m0s ago' - case "$(crew_run_progress a)" in - progressing*) ;; - *) fail "a progressing verdict was not passed through: $(crew_run_progress a)" ;; - esac - case "$(crew_wedge_progress a alive)" in - progressing*) ;; - *) fail "a live agent on a progressing run did not permit a hold" ;; - esac - # The pipeline runs its own steps, so a moving run proves nothing about a crew - # that is gone: a dead agent short-circuits without even paying for the read. - [ "$(crew_wedge_progress a dead)" = none ] \ - || fail "a confidently dead agent was absorbed by its progressing run" - case "$(crew_wedge_progress a unknown)" in - progressing*) ;; - *) fail "an inconclusive liveness verdict blocked a hold" ;; - esac - - FM_FAKE_RUN_PROGRESS='progress: stranded · test running, last activity 31m0s ago' - case "$(crew_wedge_progress a alive)" in - stranded*) ;; - *) fail "a stranded verdict was not passed through" ;; - esac - [ "$(run_progress_detail "$(crew_wedge_progress a alive)")" = "test running, last activity 31m0s ago" ] \ - || fail "run_progress_detail did not strip the class and separator" - - FM_FAKE_RUN_PROGRESS='progress: none · no run attributed to this crew' - [ "$(crew_wedge_progress a alive)" = none ] || fail "a no-evidence verdict was not none" - FM_FAKE_RUN_PROGRESS='progress: something-else · new class' - [ "$(crew_wedge_progress a alive)" = none ] || fail "an unrecognized class was not collapsed to none" - FM_FAKE_RUN_PROGRESS='total gibberish' - [ "$(crew_wedge_progress a alive)" = none ] || fail "unparseable reader output was not collapsed to none" - FM_FAKE_RUN_PROGRESS='progress: progressing · moving' - [ "$(crew_wedge_progress '' alive)" = none ] || fail "an unresolvable task id was not none" - [ "$(run_progress_detail none)" = "" ] || fail "a class-only line reported a detail" - - unset FM_FAKE_RUN_PROGRESS - pass "crew_wedge_progress: only a progressing run holds, a dead agent never does, everything else is none" -} - -# The wedge detector's third liveness input: writes inside the crew's own recorded -# worktree. Every negative outcome must report "no evidence" so the caller keeps -# its existing escalation schedule, and a supervisor-side git read (which touches -# .git, never tracked files) must not be able to fake a positive. -test_crew_worktree_written_since_classifier() { - local dir state anchor wt home statedir_wt - dir=$(make_case classify-worktree-writes); state="$dir/state" - anchor="$state/anchor"; wt="$dir/wt"; home="$dir/mate-home"; statedir_wt="$dir/wt-with-state" - mkdir -p "$wt/src" "$wt/.git/objects" - printf 'old\n' > "$wt/src/existing.c" - set_mtime "$(( $(date +%s) - 300 ))" "$wt/src/existing.c" - : > "$anchor" - set_mtime "$(( $(date +%s) - 120 ))" "$anchor" - - # No recorded worktree at all: absence of evidence, never a positive. - printf 'window=test:fm-a\nkind=ship\n' > "$state/a.meta" - ! crew_worktree_written_since a "$state" "$anchor" \ - || fail "a task with no recorded worktree reported write evidence" - # Recorded but gone (torn down): still no evidence. - printf 'window=test:fm-b\nkind=ship\nworktree=%s\n' "$dir/missing" > "$state/b.meta" - ! crew_worktree_written_since b "$state" "$anchor" \ - || fail "a torn-down worktree reported write evidence" - # Present, but nothing written since the anchor. - printf 'window=test:fm-c\nkind=ship\nworktree=%s\n' "$wt" > "$state/c.meta" - ! crew_worktree_written_since c "$state" "$anchor" \ - || fail "a quiet worktree reported write evidence" - # A missing anchor cannot be compared against: no evidence. - ! crew_worktree_written_since c "$state" "$state/absent-anchor" \ - || fail "a missing anchor reported write evidence" - # Only .git churn (what firstmate's own read-only git commands touch): pruned. - printf 'pack\n' > "$wt/.git/objects/fresh" - printf 'ref\n' > "$wt/.git/index" - ! crew_worktree_written_since c "$state" "$anchor" \ - || fail ".git churn alone reported write evidence (a supervisor read could fake liveness)" - # A real file written after the anchor: positive evidence. - printf 'new\n' > "$wt/src/new.c" - crew_worktree_written_since c "$state" "$anchor" \ - || fail "a file written after the anchor was not reported as write evidence" - # An empty id is never evidence. - ! crew_worktree_written_since "" "$state" "$anchor" || fail "an empty id reported write evidence" - - # A secondmate records a provisioned firstmate home, not a code tree, and such a - # home supervises itself: its own watcher beacon, pane hashes, and heartbeats keep - # its state/ churning whether or not the mate produced anything. - mkdir -p "$home/state" - printf 'sm-classify-1\n' > "$home/.fm-secondmate-home" - printf 'beat\n' > "$home/state/.last-watcher-beat" - printf 'window=remote:sm\nkind=secondmate\nworktree=%s\n' "$home" > "$state/sm.meta" - ! crew_worktree_written_since sm "$state" "$anchor" \ - || fail "a secondmate's own home supervision churn reported crew write evidence" - # The home marker alone is enough, even when the record does not say secondmate. - printf 'window=test:fm-sm2\nkind=ship\nworktree=%s\n' "$home" > "$state/sm2.meta" - ! crew_worktree_written_since sm2 "$state" "$anchor" \ - || fail "a marked firstmate home reported crew write evidence" - # But an ordinary worktree that merely holds a directory named state is real - # work: only the home is excluded, never a source directory of that name. - mkdir -p "$statedir_wt/state" - printf 'machine\n' > "$statedir_wt/state/machine.go" - printf 'window=test:fm-d\nkind=ship\nworktree=%s\n' "$statedir_wt" > "$state/d.meta" - crew_worktree_written_since d "$state" "$anchor" \ - || fail "a source directory named state was hidden from the write probe" - pass "crew_worktree_written_since: real writes are evidence; no worktree, no anchor, quiet trees, .git churn and a mate's own home are not" -} - -# FM_WORKTREE_WRITE_PRUNE is a skip list, so clearing it skips nothing and is the -# obvious way to widen the probe to the whole depth-bounded tree. An empty list must -# therefore widen the walk rather than report no evidence at all, which would -# quietly cost the wedge detector its third liveness input on a home that cleared -# the knob to get more coverage, not less. -test_empty_write_prune_widens_the_probe() { - local dir state anchor wt saved - dir=$(make_case classify-empty-write-prune); state="$dir/state" - anchor="$state/anchor"; wt="$dir/wt" - mkdir -p "$wt/src" "$wt/.git" - : > "$anchor" - set_mtime "$(( $(date +%s) - 120 ))" "$anchor" - printf 'window=test:fm-e\nkind=ship\nworktree=%s\n' "$wt" > "$state/e.meta" - saved=$FM_WORKTREE_WRITE_PRUNE - FM_WORKTREE_WRITE_PRUNE='' - # A quiet tree is still no evidence, so the caller's schedule is untouched. - ! crew_worktree_written_since e "$state" "$anchor" \ - || fail "an empty prune list reported write evidence for a quiet worktree" - printf 'new\n' > "$wt/src/new.c" - crew_worktree_written_since e "$state" "$anchor" \ - || fail "an empty prune list disabled the probe instead of widening it" - # Widened means nothing is skipped, including what the default list prunes. - set_mtime "$(( $(date +%s) - 900 ))" "$wt/src/new.c" - printf 'pack\n' > "$wt/.git/index" - crew_worktree_written_since e "$state" "$anchor" \ - || fail "an empty prune list still skipped a directory the default list prunes" - # Restoring the default prunes .git again, so a supervisor's own read-only git - # command still cannot fake liveness. - FM_WORKTREE_WRITE_PRUNE=$saved - ! crew_worktree_written_since e "$state" "$anchor" \ - || fail "the default prune list stopped keeping .git out of the probe" - pass "an empty FM_WORKTREE_WRITE_PRUNE widens the probe to the whole depth-bounded tree instead of disabling it" -} - -# The same widening, reached the way a home actually configures it: through the -# process ENVIRONMENT, not an in-process assignment made after the library was -# sourced. An empty exported value must survive as empty, because defaulting it with -# the colon form reads "explicitly cleared" as "never set" and hands the default skip -# list straight back to the one home that asked for a wider walk. -# shellcheck disable=SC2016 # single quotes are deliberate: the library path, state dir, and anchor expand inside the bash -c child, not here -test_empty_write_prune_from_the_environment_widens_the_probe() { - local dir state anchor wt - dir=$(make_case classify-empty-write-prune-env); state="$dir/state" - anchor="$state/anchor"; wt="$dir/wt" - mkdir -p "$wt/.git/objects" - : > "$anchor" - set_mtime "$(( $(date +%s) - 120 ))" "$anchor" - printf 'window=test:fm-wenv\nkind=ship\nworktree=%s\n' "$wt" > "$state/wenv.meta" - # The one thing written since the anchor sits exactly where the DEFAULT list prunes. - printf 'pack\n' > "$wt/.git/objects/fresh" - env -u FM_WORKTREE_WRITE_PRUNE \ - bash -c '. "$1"; crew_worktree_written_since wenv "$2" "$3"' _ \ - "$ROOT/bin/fm-classify-lib.sh" "$state" "$anchor" \ - && fail "the default skip list let .git churn count as write evidence" - FM_WORKTREE_WRITE_PRUNE='' \ - bash -c '. "$1"; crew_worktree_written_since wenv "$2" "$3"' _ \ - "$ROOT/bin/fm-classify-lib.sh" "$state" "$anchor" \ - || fail "an empty FM_WORKTREE_WRITE_PRUNE in the environment fell back to the default skip list instead of widening the probe" - pass "an empty FM_WORKTREE_WRITE_PRUNE exported into the environment prunes nothing, widening the probe" -} - -# The probe's walk runs synchronously inside the poll that was about to escalate, so -# it must be wall-clock bounded: -xdev keeps it out of a nested mount, but a worktree -# root that is ITSELF on a hung mount would otherwise stall the very supervisor that -# exists to notice a wedge. A fake find that never returns in time stands in for that -# mount. Hitting the bound must read as NO evidence, exactly like every other -# negative outcome, so the caller's escalation schedule is untouched. -test_worktree_write_probe_is_wall_clock_bounded() { - local dir state anchor wt slowbin fastbin started elapsed - dir=$(make_case classify-write-probe-bound); state="$dir/state" - anchor="$state/anchor"; wt="$dir/wt"; slowbin="$dir/slowbin"; fastbin="$dir/fastbin" - mkdir -p "$wt/src" "$slowbin" "$fastbin" - : > "$anchor" - set_mtime "$(( $(date +%s) - 120 ))" "$anchor" - printf 'window=test:fm-slow\nkind=ship\nworktree=%s\n' "$wt" > "$state/slow.meta" - # Both stand-ins report the same hit; only one of them takes longer than the bound - # to do it, so the prompt one shows what a positive outcome looks like and the - # bounded assertion below cannot pass merely because the fake failed. - cat > "$fastbin/find" <<'SH' -#!/usr/bin/env bash -set -u -printf '%s\n' "$1/hit" -SH - cat > "$slowbin/find" <<'SH' -#!/usr/bin/env bash -set -u -sleep 30 -printf '%s\n' "$1/hit" -SH - chmod +x "$fastbin/find" "$slowbin/find" - PATH="$fastbin:$PATH" \ - bash -c '. "$1"; crew_worktree_written_since slow "$2" "$3"' _ \ - "$ROOT/bin/fm-classify-lib.sh" "$state" "$anchor" \ - || fail "a walk that reported a hit inside its bound was not read as write evidence" - started=$(date +%s) - PATH="$slowbin:$PATH" FM_WORKTREE_WRITE_TIMEOUT=1 \ - bash -c '. "$1"; crew_worktree_written_since slow "$2" "$3"' _ \ - "$ROOT/bin/fm-classify-lib.sh" "$state" "$anchor" \ - && fail "a walk that outlived its bound was reported as write evidence" - elapsed=$(( $(date +%s) - started )) - [ "$elapsed" -lt 10 ] \ - || fail "the worktree write probe was not wall-clock bounded: one walk held the caller for ${elapsed}s" - pass "the worktree write probe is wall-clock bounded, and hitting the bound reads as no write evidence" -} - -# signal_crew_provably_working: a no-verb "signal:" wake is benign ONLY when EVERY -# task it references is provably working; if any crew has stopped, or no task can be -# resolved, it surfaces. Files map to ids by stripping .status / .turn-ended. -test_signal_crew_provably_working_classifier() { - local dir fakebin state - dir=$(make_case signal-provably-working); fakebin="$dir/fakebin"; state="$dir/state" - export FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" - export FM_FAKE_CREW_STATE_a='state: working · source: run-step · running' - export FM_FAKE_CREW_STATE_b='state: done · source: run-step · run passed' - signal_crew_provably_working "$state/a.status" "$state/a.turn-ended" \ - || fail "a single provably-working crew (status+turn-end) was not benign" - ! signal_crew_provably_working "$state/a.status" "$state/b.turn-ended" \ - || fail "a coalesced batch including a stopped crew was treated as benign" - ! signal_crew_provably_working "$state/b.turn-ended" \ - || fail "a stopped crew's bare turn-end was treated as benign" - ! signal_crew_provably_working "$state/a.meta" \ - || fail "a non-signal file resolved to a benign verdict" - ! signal_crew_provably_working \ - || fail "an empty signal file list was treated as benign" - unset FM_FAKE_CREW_STATE_a FM_FAKE_CREW_STATE_b - pass "signal_crew_provably_working: benign only when every referenced crew is provably working" -} - -test_secondmate_status_signal_never_absorbed_classifier() { - local dir fakebin state - dir=$(make_case secondmate-signal-classify); fakebin="$dir/fakebin"; state="$dir/state" - export FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" - # Even PROVABLY working, a secondmate's .status signal is its routed-reply - # channel and must surface; its bare turn-ended keeps the ordinary absorb. - export FM_FAKE_CREW_STATE_sm='state: working · source: run-step · running' - printf 'kind=secondmate\n' > "$state/sm.meta" - printf 'working: routed reply for the parent\n' > "$state/sm.status" - ! signal_crew_provably_working "$state/sm.status" \ - || fail "a working secondmate's status signal was treated as absorbable" - signal_crew_provably_working "$state/sm.turn-ended" \ - || fail "a working secondmate's bare turn-end lost its ordinary absorb" - # An ordinary crewmate with the same verdict stays absorbable: the rule is - # keyed on recorded kind, not on task naming or content guessing. - export FM_FAKE_CREW_STATE_crew='state: working · source: run-step · running' - printf 'kind=ship\n' > "$state/crew.meta" - printf 'working: progress\n' > "$state/crew.status" - signal_crew_provably_working "$state/crew.status" \ - || fail "the secondmate rule leaked onto an ordinary crewmate status" - unset FM_FAKE_CREW_STATE_sm FM_FAKE_CREW_STATE_crew - pass "a secondmate's status signal is never absorbed as provably working; crewmates are unaffected" -} - -# --- benign wakes are absorbed ONLY when the crew is provably working --------- - -test_provably_working_signal_absorbed() { - local dir state fakebin out status_file pid - dir=$(make_case provably-working-signal); state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out" - status_file="$state/task.status" - printf 'working: compiling step 2\n' > "$status_file" - # The crew's pipeline is in an actively-running step: positive evidence it is - # still working, so a no-verb working: signal is absorbed (the original low-churn - # case during a long validation). - export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' - watch_bg "$state" "$fakebin" "$out" - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid"; fail "watcher exited for a working: signal whose crew is provably working (should absorb): $(cat "$out")" - fi - [ ! -s "$out" ] || fail "provably-working signal printed a wake reason: $(cat "$out")" - [ ! -s "$state/.wake-queue" ] || fail "provably-working signal enqueued a durable wake record" - [ -s "$state/.seen-task_status" ] || fail "provably-working signal did not advance its .seen-* suppressor" - [ -e "$state/.last-watcher-beat" ] || fail "watcher beacon was not touched while absorbing" - reap "$pid" - pass "a no-verb signal whose crew is provably working is absorbed (no exit, no queue, suppressor advanced, beacon present)" -} - -test_turn_ended_provably_working_absorbed() { - local dir state fakebin out pid - dir=$(make_case turn-ended-working); state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out" - : > "$state/task.turn-ended" - # A busy pane is the second form of positive evidence (covers a queued - # continuation right after the turn-end). - export FM_FAKE_CREW_STATE='state: working · source: pane · harness busy' - watch_bg "$state" "$fakebin" "$out" - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid"; fail "watcher exited for a turn-end whose crew is provably working (should absorb): $(cat "$out")" - fi - [ ! -s "$out" ] || fail "provably-working turn-end printed a wake reason: $(cat "$out")" - [ ! -s "$state/.wake-queue" ] || fail "provably-working turn-end enqueued a durable wake record" - reap "$pid" - pass "a bare turn-end whose crew is provably working (busy pane) is absorbed" -} - -# --- a no-verb signal whose crew is NOT provably working SURFACES ------------- -# This is the swallowed-finish fix: a crew that finished (or stopped and waits) -# reports its final turn-end with no captain-relevant status and no running -# pipeline, so the wake must surface instead of being absorbed. - -test_turn_ended_not_working_surfaced() { - local dir state fakebin out drain_out pid - dir=$(make_case turn-ended-stopped); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out" - : > "$state/task.turn-ended" - # No running pipeline, no busy pane: the crew has stopped (e.g. it finished via - # an interactive menu and wrote no done: status). Default unknown verdict. - export FM_FAKE_CREW_STATE='state: unknown · source: none · no current-state source available' - watch_bg "$state" "$fakebin" "$out" - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not surface a turn-end whose crew is not provably working" - grep -F "signal: $state/task.turn-ended" "$out" >/dev/null || fail "watcher did not print the surfaced turn-end signal" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the surfaced turn-end failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/task.turn-ended" >/dev/null || fail "surfaced turn-end was not queued" - pass "a bare turn-end whose crew is not provably working is surfaced (the swallowed-finish fix)" -} - -# --- bare turn-end, unverifiable harness: pane churn is the third proof -------- -# A harness whose semantic busy state has no verified source (codex) can never -# report working, so the two proofs above are unreachable for it and EVERY worker -# turn boundary woke firstmate. Pane content that changed since the previous poll -# is harness-independent positive evidence the crew is still executing - the same -# liveness input the stale backbone already trusts - so a bare turn-end from a -# churning pane is benign. The pane going quiet afterwards is still caught by that -# backbone, which is why this widens the proof rather than bounding the wake rate. - -# The pane-churn turn-end absorb is opt-in per home, so every case that exercises -# it (whether it expects an absorb or one of the guards that must still surface) -# points the watcher at a case-local config dir holding the flag. A case that must -# NOT have it points at an empty one, so no developer's real config can leak in. -churn_config() { # <dir> [off] - local cfg="$1/config" - mkdir -p "$cfg" - [ "${2:-}" = off ] || : > "$cfg/turnend-churn-absorb" - printf '%s\n' "$cfg" -} - -# Wait until the watcher records an absorbed wake matching <needle> in its triage -# log. 1 if the watcher exits first (i.e. it surfaced the wake instead), which is -# exactly the unfixed behavior this case exists to catch. Polls the log rather -# than a poll cycle so the assertion lands inside the FIRST poll, long before an -# unchanging fixture pane could reach the stale backbone. -wait_for_absorbed() { # <state> <pid> <needle> - local state=$1 pid=$2 needle=$3 i=0 - while [ "$i" -lt 100 ]; do - grep -Fq "$needle" "$state/.watch-triage.log" 2>/dev/null && return 0 - kill -0 "$pid" 2>/dev/null || return 1 - sleep 0.1 - i=$((i + 1)) - done - return 1 -} - -test_turn_ended_churning_pane_absorbed() { - local dir state fakebin out capture_file window key pid - dir=$(make_case turn-ended-churning); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt" - window="test:fm-codexer" - : > "$state/codexer.turn-ended" - printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codexer.meta" - printf 'apply_patch: writing bin/thing.sh' > "$capture_file" - key=$(printf '%s' "$window" | tr ':/.' '___') - # The previous poll recorded DIFFERENT pane content, so this poll's capture is - # churn: the crew rendered output between the two polls. - printf '%s' "$(hash_text 'reading the brief')" > "$state/.hash-$key" - printf '0\n' > "$state/.count-$key" - # The codex verdict verbatim: a verified dispatch adapter with no verified - # semantic busy source, so crew_is_provably_working can never be satisfied. - export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' - # A slow poll leaves the first cycle's absorb assertion many ticks clear of the - # stale backbone, which this static fixture pane would otherwise reach. - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_absorbed "$state" "$pid" "absorbed benign signal:" \ - || { reap "$pid"; fail "a bare turn-end from a churning pane was not absorbed: $(cat "$out")"; } - [ ! -s "$out" ] || fail "an absorbed churning-pane turn-end printed a wake reason: $(cat "$out")" - [ ! -s "$state/.wake-queue" ] || fail "an absorbed churning-pane turn-end enqueued a durable wake record" - [ -s "$state/.churn-since-$key" ] \ - || { reap "$pid"; fail "an absorbed churning-pane turn-end did not open a bounded deferral window"; } - reap "$pid" - unset FM_FAKE_CREW_STATE - pass "a bare turn-end from a pane that churned since the previous poll is absorbed" -} - -test_turn_ended_churn_resets_prior_stale_classification() { - local dir state fakebin out capture_file window key old_hash active_hash pid i - dir=$(make_case turn-ended-churn-resets-stale); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt" - window="test:fm-codexreturned" - : > "$state/codexreturned.turn-ended" - printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codexreturned.meta" - old_hash=$(hash_text 'idle prompt from an earlier turn') - active_hash=$(hash_text 'rendering a new turn') - printf 'rendering a new turn' > "$capture_file" - key=$(printf '%s' "$window" | tr ':/.' '___') - printf '%s' "$old_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - printf '%s' "$old_hash" > "$state/.stale-$key" - date +%s > "$state/.stale-since-$key" - export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=999 \ - FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_absorbed "$state" "$pid" "absorbed benign signal:" \ - || { reap "$pid"; fail "a churning turn-end with prior stale state was not absorbed: $(cat "$out")"; } - i=0 - while [ "$i" -lt 100 ] && [ "$(cat "$state/.hash-$key" 2>/dev/null || true)" != "$active_hash" ]; do - kill -0 "$pid" 2>/dev/null || { reap "$pid"; fail "watcher exited before recording the active pane"; } - sleep 0.1 - i=$((i + 1)) - done - [ "$(cat "$state/.hash-$key" 2>/dev/null || true)" = "$active_hash" ] \ - || { reap "$pid"; fail "watcher did not record the active pane after absorbing its turn-end"; } - - # The worker stops on bytes that happened to be stale in an earlier turn. - # This is a new quiet interval, so it must surface through ordinary staleness - # instead of inheriting the earlier interval's wedge timer. - printf 'idle prompt from an earlier turn' > "$capture_file" - wait_for_exit "$pid" 100 \ - || { reap "$pid"; fail "a stopped pane matching an earlier stale render waited for the wedge timeout"; } - grep -Fx "stale: $window" "$out" >/dev/null \ - || fail "the returned stale render did not surface through ordinary staleness" - grep -F "possible wedge" "$out" >/dev/null \ - && fail "the returned stale render inherited the earlier quiet interval's wedge classification" - unset FM_FAKE_CREW_STATE - pass "pane churn starts a fresh stale-classification interval before a stopped render returns" -} - -test_turn_ended_churn_resets_wedge_state_before_stale_poll() { - local dir state fakebin out capture_file capture_count window key pid - dir=$(make_case turn-ended-churn-resets-wedge); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; capture_count="$dir/capture.count" - window="test:fm-codexfreshinterval" - : > "$state/codexfreshinterval.turn-ended" - printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codexfreshinterval.meta" - printf 'rendering a new turn' > "$capture_file" - key=$(printf '%s' "$window" | tr ':/.' '___') - printf '%s' "$(hash_text 'idle output from the prior interval')" > "$state/.hash-$key" - printf '2\n' > "$state/.wedge-escalations-$key" - export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_TMUX_CAPTURE_COUNT_FILE="$capture_count" FM_FAKE_TMUX_CAPTURE_FAIL_AFTER=1 \ - FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_absorbed "$state" "$pid" "absorbed benign signal:" \ - || { reap "$pid"; fail "a churning turn-end was not absorbed before the stale-path capture failed: $(cat "$out")"; } - [ ! -e "$state/.wedge-escalations-$key" ] \ - || { reap "$pid"; fail "churn retained the prior quiet interval's wedge-escalation count"; } - [ ! -s "$state/.wake-queue" ] \ - || { reap "$pid"; fail "the absorbed churn fixture queued an unexpected wake"; } - reap "$pid" - unset FM_FAKE_CREW_STATE - pass "pane churn resets prior wedge escalation state before the stale-path poll" -} - -# The safety half: the same unverifiable harness, the same fixture, but the pane -# has NOT changed since the previous poll. There is no positive evidence, so the -# wake must still surface - a stopped worker is exactly what the turn-end marker -# earns its keep detecting, and widening the proof must not cost that. -test_turn_ended_still_pane_surfaced() { - local dir state fakebin out drain_out capture_file window key pid - dir=$(make_case turn-ended-still); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - window="test:fm-codexstopped" - : > "$state/codexstopped.turn-ended" - printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codexstopped.meta" - printf 'apply_patch: writing bin/thing.sh' > "$capture_file" - key=$(printf '%s' "$window" | tr ':/.' '___') - # The previous poll recorded THIS pane content: nothing rendered since. - printf '%s' "$(hash_text 'apply_patch: writing bin/thing.sh')" > "$state/.hash-$key" - printf '0\n' > "$state/.count-$key" - export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not surface a bare turn-end from an unchanged pane" - grep -F "signal: $state/codexstopped.turn-ended" "$out" >/dev/null \ - || fail "watcher did not print the surfaced still-pane turn-end signal" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the still-pane turn-end failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/codexstopped.turn-ended" >/dev/null \ - || fail "surfaced still-pane turn-end was not queued" - unset FM_FAKE_CREW_STATE - pass "a bare turn-end from a pane unchanged since the previous poll still surfaces" -} - -test_turn_ended_malformed_prior_hash_surfaced() { - local dir state fakebin out drain_out capture_file window key pid - dir=$(make_case turn-ended-malformed-hash); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - window="test:fm-codexmalformed" - : > "$state/codexmalformed.turn-ended" - printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codexmalformed.meta" - printf 'stopped after rendering this' > "$capture_file" - key=$(printf '%s' "$window" | tr ':/.' '___') - printf 'x' > "$state/.hash-$key" - printf '0\n' > "$state/.count-$key" - export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher absorbed a turn-end backed by a malformed prior hash" - grep -F "signal: $state/codexmalformed.turn-ended" "$out" >/dev/null \ - || fail "watcher did not print the surfaced malformed-hash turn-end" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ - || fail "drain after the malformed-hash turn-end failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/codexmalformed.turn-ended" >/dev/null \ - || fail "malformed-hash turn-end was not queued" - unset FM_FAKE_CREW_STATE - pass "a bare turn-end backed by a malformed prior hash surfaces" -} - -test_turn_ended_trailing_newline_prior_hash_surfaced() { - local dir state fakebin out drain_out capture_file window key pid - dir=$(make_case turn-ended-newline-hash); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - window="test:fm-codexnewline" - : > "$state/codexnewline.turn-ended" - printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codexnewline.meta" - printf 'rendered after the prior poll' > "$capture_file" - key=$(printf '%s' "$window" | tr ':/.' '___') - printf '%s\n' "$(hash_text 'the previous render')" > "$state/.hash-$key" - printf '0\n' > "$state/.count-$key" - export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher absorbed a turn-end backed by a newline-terminated prior hash" - grep -F "signal: $state/codexnewline.turn-ended" "$out" >/dev/null \ - || fail "watcher did not print the surfaced newline-hash turn-end" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ - || fail "drain after the newline-hash turn-end failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/codexnewline.turn-ended" >/dev/null \ - || fail "newline-hash turn-end was not queued" - [ ! -e "$state/.churn-since-$key" ] \ - || fail "a newline-terminated prior hash opened a deferral window" - unset FM_FAKE_CREW_STATE - pass "a bare turn-end backed by a newline-terminated prior hash surfaces" -} - -test_secondmate_turn_ended_churning_pane_surfaced() { - local dir state fakebin out drain_out capture_file window key pid - dir=$(make_case secondmate-turn-ended-churning); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - window="test:fm-mate-churning" - : > "$state/mate.turn-ended" - printf 'window=%s\nkind=secondmate\nharness=pi\n' "$window" > "$state/mate.meta" - printf 'working on the next routed item' > "$capture_file" - key=$(printf '%s' "$window" | tr ':/.' '___') - printf '%s' "$(hash_text 'waiting for work')" > "$state/.hash-$key" - printf '0\n' > "$state/.count-$key" - export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable' - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not surface a churning secondmate turn-end" - grep -F "signal: $state/mate.turn-ended" "$out" >/dev/null \ - || fail "watcher did not print the surfaced churning secondmate turn-end" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ - || fail "drain after the churning secondmate turn-end failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/mate.turn-ended" >/dev/null \ - || fail "churning secondmate turn-end was not queued" - unset FM_FAKE_CREW_STATE - pass "a churning secondmate turn-end surfaces without a stale resurface path" -} - -test_turn_ended_colliding_window_key_surfaced() { - local dir state fakebin out drain_out capture_file window colliding key pid - dir=$(make_case turn-ended-colliding-key); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - window="test:fm-a.b"; colliding="test:fm-a_b" - : > "$state/a.b.turn-ended" - printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/a.b.meta" - printf 'window=%s\nkind=ship\nharness=codex\n' "$colliding" > "$state/a_b.meta" - printf 'rendered after the prior poll' > "$capture_file" - key=$(printf '%s' "$window" | tr ':/.' '___') - printf '%s' "$(hash_text 'the other window pane')" > "$state/.hash-$key" - printf '0\n' > "$state/.count-$key" - export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not surface a turn-end with an ambiguous pane marker" - grep -F "signal: $state/a.b.turn-ended" "$out" >/dev/null \ - || fail "watcher did not print the surfaced ambiguous-marker turn-end" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ - || fail "drain after the ambiguous-marker turn-end failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/a.b.turn-ended" >/dev/null \ - || fail "ambiguous-marker turn-end was not queued" - unset FM_FAKE_CREW_STATE - pass "a turn-end whose marker key matches another recorded endpoint surfaces" -} - -test_turn_ended_duplicate_endpoint_records_surfaced() { - local dir state fakebin out drain_out capture_file window key pid - dir=$(make_case turn-ended-duplicate-endpoint); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - window="test:fm-shared" - : > "$state/first.turn-ended" - printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/first.meta" - printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/second.meta" - printf 'rendered after the prior poll' > "$capture_file" - key=$(printf '%s' "$window" | tr ':/.' '___') - printf '%s' "$(hash_text 'the previous render')" > "$state/.hash-$key" - printf '0\n' > "$state/.count-$key" - export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher absorbed a turn-end shared by two endpoint records" - grep -F "signal: $state/first.turn-ended" "$out" >/dev/null \ - || fail "watcher did not print the surfaced duplicate-endpoint turn-end" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ - || fail "drain after the duplicate-endpoint turn-end failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/first.turn-ended" >/dev/null \ - || fail "duplicate-endpoint turn-end was not queued" - [ ! -e "$state/.churn-since-$key" ] \ - || fail "duplicate endpoint records opened a deferral window" - unset FM_FAKE_CREW_STATE - pass "two metadata records sharing one endpoint make churn evidence ambiguous" -} - -test_turn_ended_mixed_positive_evidence_batch_absorbed() { - local dir state fakebin out capture_file first_window second_window first_key second_key pid - dir=$(make_case turn-ended-mixed-evidence); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt" - first_window="test:fm-first"; second_window="test:fm-second" - : > "$state/first.turn-ended" - : > "$state/second.turn-ended" - printf 'window=%s\nkind=ship\nharness=pi\n' "$first_window" > "$state/first.meta" - printf 'window=%s\nkind=ship\nharness=codex\n' "$second_window" > "$state/second.meta" - printf 'second task rendered after the prior poll' > "$capture_file" - first_key=$(printf '%s' "$first_window" | tr ':/.' '___') - second_key=$(printf '%s' "$second_window" | tr ':/.' '___') - printf '%s' "$(hash_text 'first task static pane')" > "$state/.hash-$first_key" - printf '%s' "$(hash_text 'second task previous render')" > "$state/.hash-$second_key" - printf '0\n' > "$state/.count-$first_key" - printf '0\n' > "$state/.count-$second_key" - export FM_FAKE_CREW_STATE_first='state: working · source: run-step · running' - export FM_FAKE_CREW_STATE_second='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOWS="$(printf 'fm-first\nfm-second')" \ - FM_FAKE_TMUX_CAPTURE="$capture_file" FM_FAKE_TMUX_FORBIDDEN_TARGET="$first_window" \ - FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_absorbed "$state" "$pid" "absorbed benign signal:" \ - || { reap "$pid"; fail "a mixed authoritative-and-churn batch was not absorbed: $(cat "$out")"; } - [ ! -s "$out" ] || fail "an absorbed mixed-evidence batch printed a wake reason: $(cat "$out")" - [ ! -s "$state/.wake-queue" ] || fail "an absorbed mixed-evidence batch enqueued a durable wake record" - [ ! -e "$state/.churn-since-$first_key" ] \ - || fail "an authoritatively working task opened a pane-churn deadline" - [ -s "$state/.churn-since-$second_key" ] \ - || fail "the churn-proven task did not open its bounded deferral window" - reap "$pid" - unset FM_FAKE_CREW_STATE_first FM_FAKE_CREW_STATE_second - pass "a batch may satisfy positive evidence independently per task" -} - -test_turn_ended_mixed_positive_evidence_batch_default_off() { - local dir state fakebin out drain_out capture_file first_window second_window first_key second_key pid - dir=$(make_case turn-ended-mixed-evidence-off); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - first_window="test:fm-firstoff"; second_window="test:fm-secondoff" - : > "$state/firstoff.turn-ended" - : > "$state/secondoff.turn-ended" - printf 'window=%s\nkind=ship\nharness=pi\n' "$first_window" > "$state/firstoff.meta" - printf 'window=%s\nkind=ship\nharness=codex\n' "$second_window" > "$state/secondoff.meta" - printf 'second task rendered after the prior poll' > "$capture_file" - first_key=$(printf '%s' "$first_window" | tr ':/.' '___') - second_key=$(printf '%s' "$second_window" | tr ':/.' '___') - printf '%s' "$(hash_text 'first task static pane')" > "$state/.hash-$first_key" - printf '%s' "$(hash_text 'second task previous render')" > "$state/.hash-$second_key" - printf '0\n' > "$state/.count-$first_key" - printf '0\n' > "$state/.count-$second_key" - export FM_FAKE_CREW_STATE_firstoff='state: working · source: run-step · running' - export FM_FAKE_CREW_STATE_secondoff='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOWS="$(printf 'fm-firstoff\nfm-secondoff')" \ - FM_FAKE_TMUX_CAPTURE="$capture_file" FM_CONFIG_OVERRIDE="$(churn_config "$dir" off)" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher absorbed a mixed-evidence batch without the opt-in flag" - grep -F "$state/firstoff.turn-ended" "$out" >/dev/null \ - || fail "watcher did not print the first default-off turn-end" - grep -F "$state/secondoff.turn-ended" "$out" >/dev/null \ - || fail "watcher did not print the second default-off turn-end" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ - || fail "drain after the default-off mixed-evidence batch failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/firstoff.turn-ended" >/dev/null \ - || fail "the first default-off turn-end was not queued" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/secondoff.turn-ended" >/dev/null \ - || fail "the second default-off turn-end was not queued" - [ ! -e "$state/.churn-since-$first_key" ] && [ ! -e "$state/.churn-since-$second_key" ] \ - || fail "the default-off mixed-evidence batch opened a deferral window" - unset FM_FAKE_CREW_STATE_firstoff FM_FAKE_CREW_STATE_secondoff - pass "per-task evidence composition stays off until the home opts in" -} - -test_status_and_turn_end_batch_never_uses_churn_evidence() { - local dir state fakebin out drain_out capture_file first_window second_window second_key pid - dir=$(make_case status-and-turn-ended-churn); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - first_window="test:fm-firststatus"; second_window="test:fm-secondturn" - printf 'working: authoritative task still running\n' > "$state/firststatus.status" - : > "$state/secondturn.turn-ended" - printf 'window=%s\nkind=ship\nharness=pi\n' "$first_window" > "$state/firststatus.meta" - printf 'window=%s\nkind=ship\nharness=codex\n' "$second_window" > "$state/secondturn.meta" - printf 'second task rendered after the prior poll' > "$capture_file" - second_key=$(printf '%s' "$second_window" | tr ':/.' '___') - printf '%s' "$(hash_text 'second task previous render')" > "$state/.hash-$second_key" - printf '0\n' > "$state/.count-$second_key" - export FM_FAKE_CREW_STATE_firststatus='state: working · source: run-step · running' - export FM_FAKE_CREW_STATE_secondturn='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOWS="$(printf 'fm-firststatus\nfm-secondturn')" \ - FM_FAKE_TMUX_CAPTURE="$capture_file" FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher absorbed a status-and-turn-end batch on churn evidence" - grep -F "$state/firststatus.status" "$out" >/dev/null \ - || fail "watcher did not print the status file from the surfaced mixed batch" - grep -F "$state/secondturn.turn-ended" "$out" >/dev/null \ - || fail "watcher did not print the turn-end from the surfaced mixed batch" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ - || fail "drain after the surfaced status-and-turn-end batch failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/firststatus.status" >/dev/null \ - || fail "the status file from the surfaced mixed batch was not queued" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/secondturn.turn-ended" >/dev/null \ - || fail "the turn-end from the surfaced mixed batch was not queued" - [ ! -e "$state/.churn-since-$second_key" ] \ - || fail "a status-bearing batch opened a pane-churn deadline" - unset FM_FAKE_CREW_STATE_firststatus FM_FAKE_CREW_STATE_secondturn - pass "a status-bearing batch never falls through to pane-churn evidence" -} - -# The opt-in half. Pane churn infers execution from rendered bytes rather than -# from a verdict the harness vouches for, so a home that has not asked for it must -# see exactly the pre-change triage: the same churning fixture that absorbs above -# surfaces here purely because the flag is absent. -test_turn_ended_churn_absorb_off_by_default() { - local dir state fakebin out drain_out capture_file window key pid - dir=$(make_case turn-ended-churn-default-off); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - window="test:fm-codexdefault" - : > "$state/codexdefault.turn-ended" - printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codexdefault.meta" - printf 'apply_patch: writing bin/thing.sh' > "$capture_file" - key=$(printf '%s' "$window" | tr ':/.' '___') - printf '%s' "$(hash_text 'reading the brief')" > "$state/.hash-$key" - printf '0\n' > "$state/.count-$key" - export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_CONFIG_OVERRIDE="$(churn_config "$dir" off)" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher absorbed a churning turn-end without the opt-in flag" - grep -F "signal: $state/codexdefault.turn-ended" "$out" >/dev/null \ - || fail "watcher did not print the surfaced default-off churning turn-end" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ - || fail "drain after the default-off churning turn-end failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/codexdefault.turn-ended" >/dev/null \ - || fail "default-off churning turn-end was not queued" - [ ! -e "$state/.churn-since-$key" ] \ - || fail "the default-off path opened a bounded deferral window" - unset FM_FAKE_CREW_STATE - pass "pane-churn turn-end absorb is off until a home opts in" -} - -# The bound. Churn and pane staleness read the same pane, so a pane that renders -# continuously (a clock, a spinner, a harness that leaves a background renderer -# alive after its agent yields) never reaches the staleness backbone's two -# identical hashes either. Without a bound on the churn absorb a worker that had -# genuinely stopped behind such a renderer would have no path left to surface at -# all, so an exhausted deferral window must surface and restart. -test_turn_ended_churn_absorb_bounded() { - local dir state fakebin out drain_out capture_file window key pid - dir=$(make_case turn-ended-churn-bounded); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - window="test:fm-codexclock" - : > "$state/codexclock.turn-ended" - printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codexclock.meta" - printf 'a background renderer that never stops' > "$capture_file" - key=$(printf '%s' "$window" | tr ':/.' '___') - printf '%s' "$(hash_text 'the previous frame')" > "$state/.hash-$key" - printf '0\n' > "$state/.count-$key" - # This endpoint has already been riding churn evidence longer than the bound. - printf '%s' "$(( $(date +%s) - 600 ))" > "$state/.churn-since-$key" - export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_CONFIG_OVERRIDE="$(churn_config "$dir")" FM_TURNEND_CHURN_ABSORB_SECS=60 \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 \ - || fail "a perpetually churning pane deferred its turn-end past the absorb bound" - grep -F "signal: $state/codexclock.turn-ended" "$out" >/dev/null \ - || fail "watcher did not print the turn-end surfaced by the exhausted absorb bound" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ - || fail "drain after the bounded churn turn-end failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/codexclock.turn-ended" >/dev/null \ - || fail "the turn-end surfaced by the exhausted absorb bound was not queued" - [ ! -e "$state/.churn-since-$key" ] \ - || fail "an exhausted deferral window was not restarted after surfacing" - unset FM_FAKE_CREW_STATE - pass "a perpetually churning pane surfaces once its bounded deferral window is spent" -} - -test_turn_ended_churn_timer_write_failure_surfaced() { - local dir state fakebin out drain_out capture_file window key pid - dir=$(make_case turn-ended-churn-timer-write-failure); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - window="test:fm-codextimer" - : > "$state/codextimer.turn-ended" - printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codextimer.meta" - printf 'rendered after the previous poll' > "$capture_file" - key=$(printf '%s' "$window" | tr ':/.' '___') - printf '%s' "$(hash_text 'the previous render')" > "$state/.hash-$key" - printf '0\n' > "$state/.count-$key" - mkdir "$state/.churn-since-$key" - export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" 2>/dev/null & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher absorbed a churning turn-end without recording its deadline" - grep -F "signal: $state/codextimer.turn-ended" "$out" >/dev/null \ - || fail "watcher did not print the turn-end whose churn deadline could not be recorded" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ - || fail "drain after the failed churn deadline write failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/codextimer.turn-ended" >/dev/null \ - || fail "turn-end with an unrecordable churn deadline was not queued" - unset FM_FAKE_CREW_STATE - pass "an unrecordable pane-churn deadline surfaces the turn-end" -} - -test_turn_ended_invalid_churn_bound_surfaced() { - local dir state fakebin out drain_out capture_file window key pid - dir=$(make_case turn-ended-invalid-churn-bound); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - window="test:fm-codexbound" - : > "$state/codexbound.turn-ended" - printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codexbound.meta" - printf 'rendered after the previous poll' > "$capture_file" - key=$(printf '%s' "$window" | tr ':/.' '___') - printf '%s' "$(hash_text 'the previous render')" > "$state/.hash-$key" - printf '0\n' > "$state/.count-$key" - export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_CONFIG_OVERRIDE="$(churn_config "$dir")" FM_TURNEND_CHURN_ABSORB_SECS=bogus \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" 2>/dev/null & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not surface a turn-end with an invalid churn bound" - grep -F "signal: $state/codexbound.turn-ended" "$out" >/dev/null \ - || fail "watcher terminated before printing the invalid-bound turn-end" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ - || fail "drain after the invalid churn bound failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/codexbound.turn-ended" >/dev/null \ - || fail "turn-end with an invalid churn bound was not queued" - [ ! -e "$state/.churn-since-$key" ] \ - || fail "an invalid churn bound opened a deferral window" - unset FM_FAKE_CREW_STATE - pass "an invalid pane-churn bound surfaces the turn-end" -} - -test_turn_ended_oversized_churn_bound_surfaced() { - local dir state fakebin out drain_out capture_file window key pid - dir=$(make_case turn-ended-oversized-churn-bound); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - window="test:fm-codexoversized" - : > "$state/codexoversized.turn-ended" - printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codexoversized.meta" - printf 'rendered after the previous poll' > "$capture_file" - key=$(printf '%s' "$window" | tr ':/.' '___') - printf '%s' "$(hash_text 'the previous render')" > "$state/.hash-$key" - printf '0\n' > "$state/.count-$key" - export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_CONFIG_OVERRIDE="$(churn_config "$dir")" FM_TURNEND_CHURN_ABSORB_SECS=999999999999999999999999999999999999 \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" 2>/dev/null & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not surface a turn-end with an oversized churn bound" - grep -F "signal: $state/codexoversized.turn-ended" "$out" >/dev/null \ - || fail "watcher terminated before printing the oversized-bound turn-end" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ - || fail "drain after the oversized churn bound failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/codexoversized.turn-ended" >/dev/null \ - || fail "turn-end with an oversized churn bound was not queued" - [ ! -e "$state/.churn-since-$key" ] \ - || fail "an oversized churn bound opened a deferral window" - unset FM_FAKE_CREW_STATE - pass "an oversized pane-churn bound surfaces the turn-end" -} - -test_turn_ended_invalid_churn_deadline_surfaced() { - local variant value dir state fakebin out drain_out capture_file window key marker pid - for variant in empty leading-zero nonnumeric future overflow; do - dir=$(make_case "turn-ended-invalid-churn-deadline-$variant") - state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - window="test:fm-codexdeadline" - : > "$state/codexdeadline.turn-ended" - printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codexdeadline.meta" - printf 'rendered after the previous poll' > "$capture_file" - key=$(printf '%s' "$window" | tr ':/.' '___') - marker="$state/.churn-since-$key" - printf '%s' "$(hash_text 'the previous render')" > "$state/.hash-$key" - printf '0\n' > "$state/.count-$key" - case "$variant" in - empty) value='' ;; - leading-zero) value=09 ;; - nonnumeric) value=bogus ;; - future) value=$(( $(date +%s) + 600 )) ;; - overflow) value=999999999999999999999999999999999999 ;; - esac - printf '%s' "$value" > "$marker" - export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" 2>/dev/null & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not surface a turn-end with a $variant churn deadline" - grep -F "signal: $state/codexdeadline.turn-ended" "$out" >/dev/null \ - || fail "watcher terminated before printing the $variant-deadline turn-end" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ - || fail "drain after the $variant churn deadline failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/codexdeadline.turn-ended" >/dev/null \ - || fail "turn-end with a $variant churn deadline was not queued" - [ "$(cat "$marker")" = "$value" ] \ - || fail "the $variant churn deadline was rewritten" - done - unset FM_FAKE_CREW_STATE - pass "invalid existing pane-churn deadlines surface without mutation" -} - -test_turn_ended_surfaced_batch_opens_no_partial_deadline() { - local dir state fakebin out drain_out capture_file first_window second_window first_key second_key pid - dir=$(make_case turn-ended-no-partial-churn-deadline); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - first_window="test:fm-codexfirst"; second_window="test:fm-codexsecond" - : > "$state/first.turn-ended" - : > "$state/second.turn-ended" - printf 'window=%s\nkind=ship\nharness=codex\n' "$first_window" > "$state/first.meta" - printf 'window=%s\nkind=ship\nharness=codex\n' "$second_window" > "$state/second.meta" - printf 'rendered after the previous poll' > "$capture_file" - first_key=$(printf '%s' "$first_window" | tr ':/.' '___') - second_key=$(printf '%s' "$second_window" | tr ':/.' '___') - printf '%s' "$(hash_text 'first previous render')" > "$state/.hash-$first_key" - printf '%s' "$(hash_text 'second previous render')" > "$state/.hash-$second_key" - printf '0\n' > "$state/.count-$first_key" - printf '0\n' > "$state/.count-$second_key" - printf 'bogus' > "$state/.churn-since-$second_key" - export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOWS="$(printf 'fm-codexfirst\nfm-codexsecond')" \ - FM_FAKE_TMUX_CAPTURE="$capture_file" FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" 2>/dev/null & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher absorbed a batch containing an invalid churn deadline" - grep -F "$state/first.turn-ended" "$out" >/dev/null \ - || fail "watcher did not print the first turn-end from the surfaced batch" - grep -F "$state/second.turn-ended" "$out" >/dev/null \ - || fail "watcher did not print the second turn-end from the surfaced batch" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ - || fail "drain after the surfaced churn batch failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/first.turn-ended" >/dev/null \ - || fail "the first turn-end from the surfaced batch was not queued" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/second.turn-ended" >/dev/null \ - || fail "the second turn-end from the surfaced batch was not queued" - [ ! -e "$state/.churn-since-$first_key" ] \ - || fail "a surfaced batch opened a partial churn deadline" - [ "$(cat "$state/.churn-since-$second_key")" = bogus ] \ - || fail "the invalid churn deadline in a surfaced batch was rewritten" - unset FM_FAKE_CREW_STATE - pass "a surfaced batch opens no partial pane-churn deadline" -} - -test_working_note_not_working_surfaced() { - local dir state fakebin out drain_out status_file pid - dir=$(make_case working-note-stopped); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out" - status_file="$state/task.status" - printf 'working: compiling step 2\n' > "$status_file" - # A non-no-mistakes crew (no run) whose pane went idle: fm-crew-state falls back - # to the stale working: status-log line. That is NOT positive evidence, so the - # wake must surface - these users must never be left hanging. - export FM_FAKE_CREW_STATE='state: working · source: status-log · working: compiling step 2' - watch_bg "$state" "$fakebin" "$out" - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not surface a working: note whose crew has no running pipeline and an idle pane" - grep -F "signal: $status_file" "$out" >/dev/null || fail "watcher did not print the surfaced working: signal" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the surfaced working: note failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$status_file" >/dev/null || fail "surfaced working: note was not queued" - [ -s "$state/.seen-task_status" ] || fail "surfaced working: note did not advance its .seen-* suppressor" - pass "a no-verb working: note whose crew is idle with no running pipeline is surfaced" -} - -test_secondmate_status_note_surfaced_despite_busy_agent() { - local dir state fakebin out drain_out pid - dir=$(make_case secondmate-note-surfaced); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out" - printf 'kind=secondmate\n' > "$state/mate.meta" - printf 'working: routed reply landed in the parent stream\n' > "$state/mate.status" - # Busy evidence that would absorb an ordinary crewmate's no-verb note must - # not absorb a secondmate's: its status stream is the routed-reply channel. - export FM_FAKE_CREW_STATE='state: working · source: run-step · running' - FM_CONFIG_OVERRIDE="$(churn_config "$dir")" watch_bg "$state" "$fakebin" "$out" - pid=$! - wait_for_exit "$pid" 100 || fail "watcher absorbed a busy secondmate's routed status note" - grep -F "signal: $state/mate.status" "$out" >/dev/null \ - || fail "watcher did not print the surfaced secondmate note" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the surfaced note failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/mate.status" >/dev/null \ - || fail "surfaced secondmate note was not queued" - pass "a secondmate's status note surfaces even while its own agent is busy" -} - -test_self_announced_close_does_not_rewake_but_next_note_does() { - local dir state fakebin out status_file pid rc - dir=$(make_case self-close-quiet); state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out" - status_file="$state/task.status" - printf 'needs-decision [key=k1]: pick one\n' > "$status_file" - prime_status_seen "$state" "$status_file" || fail "could not prime the announced baseline" - # The home's own bookkeeping close, written through the guarded - # self-announced append this home's answerers use. - rc=0 - FM_STATE_OVERRIDE="$state" bash -c ' - . "$1" - fm_wake_status_append_self_announced "$2" "$3" "resolved [key=k1]: answered: closed by this home" - ' _ "$ROOT/bin/fm-wake-lib.sh" "$state" "$status_file" || rc=$? - [ "$rc" -eq 0 ] || fail "the bookkeeping close was not self-announced (rc=$rc)" - export FM_FAKE_CREW_STATE='state: unknown · source: none · idle worker' - watch_bg "$state" "$fakebin" "$out" - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid"; fail "the home's own bookkeeping close re-woke its own watcher: $(cat "$out")" - fi - [ ! -s "$out" ] || { reap "$pid"; fail "self-announced close printed a wake reason: $(cat "$out")"; } - [ ! -s "$state/.wake-queue" ] || { reap "$pid"; fail "self-announced close enqueued a durable wake"; } - # A later, different note on the SAME task still wakes: dedup is keyed on the - # exact announced bytes, never on task identity. - printf 'needs-decision [key=k2]: a genuinely new decision\n' >> "$status_file" - wait_for_exit "$pid" 100 || fail "a later different note after a self-announced close was swallowed" - grep -F "signal: $status_file" "$out" >/dev/null \ - || fail "the later note did not surface as a signal" - pass "a self-announced close never wakes its own home, and the next real note still does" -} - -# --- actionable wakes are surfaced (queue + exit) --------------------------- - -test_actionable_signal_surfaced() { - local dir state fakebin out drain_out status_file pid - dir=$(make_case actionable-signal); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out" - status_file="$state/task.status" - printf 'working: setup\nneeds-decision: pick A or B\n' > "$status_file" - watch_bg "$state" "$fakebin" "$out" - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not exit for an actionable needs-decision signal" - grep -F "signal: $status_file" "$out" >/dev/null || fail "watcher did not print the actionable signal reason" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the actionable signal failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$status_file" >/dev/null || fail "actionable signal was not queued" - [ -s "$state/.hb-surfaced-task" ] || fail "actionable signal did not record the surfaced marker" - pass "captain-relevant signal is surfaced (queue + exit) and marked surfaced" -} - -# A needs-decision status append surfaced through this actionable signal path -# must skip the Pi supervision branch and reach main directly -# (docs/pi-supervision-branch.md "Autonomy"). The row still -# queues as an ordinary signal-kind wake - fm-branch-dispatch.ts's -# scopeForUnreadWake tells it apart from a routine signal by this payload -# marker, not by kind. -test_needs_decision_signal_payload_marked_for_branch_exclusion() { - local dir state fakebin out status_file pid - dir=$(make_case needs-decision-payload); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out" - status_file="$state/task.status" - printf 'working: setup\nneeds-decision: pick A or B\n' > "$status_file" - watch_bg "$state" "$fakebin" "$out" - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not exit for an actionable needs-decision signal" - grep -F "$(printf 'signal\ttask.status\tneeds-decision:')" "$state/.wake-queue" >/dev/null \ - || fail "a needs-decision signal row was not payload-marked for branch exclusion: $(cat "$state/.wake-queue")" - pass "a needs-decision signal row's queued payload is marked needs-decision: for branch exclusion" -} - -# A needs-decision whose key transition was rejected by the reserved-key -# vocabulary is reported as a "reconciliation-required: " wrapped event -# (fm-classify-lib.sh's status_span_first_actionable_record), but it is still a -# needs-decision signal that this path routes directly to main - the payload -# marker must not be fooled by that wrapper. -test_needs_decision_reconciliation_required_still_marked() { - local dir state fakebin out status_file pid - dir=$(make_case needs-decision-reconciliation); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out" - status_file="$state/task.status" - printf 'needs-decision [key=pending-reply-x]: unrelated request\nworking: awaiting reconciliation\n' \ - > "$status_file" - watch_bg "$state" "$fakebin" "$out" - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not exit for a rejected-reserved-key needs-decision" - grep -F "$(printf 'signal\ttask.status\tneeds-decision:')" "$state/.wake-queue" >/dev/null \ - || fail "a reconciliation-required needs-decision row was not payload-marked for branch exclusion: $(cat "$state/.wake-queue")" - pass "a reconciliation-required needs-decision row's queued payload is still marked needs-decision:" -} - -# A captain-held declaration is itself actionable. Positive evidence that the -# crew is still working must not absorb the signal before its main-only marker -# can be delivered. -test_captain_held_signal_payload_marked_for_branch_exclusion() { - local dir state fakebin out status_file pid - dir=$(make_case captain-held-signal-payload); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out" - status_file="$state/task.status" - printf 'captain-held [key=route]: awaiting the captain\n' > "$status_file" - export FM_FAKE_CREW_STATE='state: working · source: run-step · still wrapping up' - watch_bg "$state" "$fakebin" "$out" - pid=$! - wait_for_exit "$pid" 100 || fail "watcher absorbed a captain-held signal while the crew was still working" - grep -F "signal: $status_file" "$out" >/dev/null \ - || fail "a captain-held signal changed its wake reason: $(cat "$out")" - grep -F "$(printf 'signal\ttask.status\tneeds-decision:')" "$state/.wake-queue" >/dev/null \ - || fail "a captain-held signal was not payload-marked for branch exclusion: $(cat "$state/.wake-queue")" - pass "a captain-held signal stays actionable while the crew is still working" -} - -test_pending_reply_escalation_signal_payload_marked_for_branch_exclusion() { - local dir state fakebin out status_file pid corr - dir=$(make_case pending-reply-escalation-payload); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out" - status_file="$state/task.status" - corr=0123456789abcdef - printf 'blocked [key=pending-reply-%s]: pending-reply-missed: task=task pending-reply-id=%s request=finish report\n' \ - "$corr" "$corr" > "$status_file" - watch_bg "$state" "$fakebin" "$out" - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not exit for a pending-reply escalation" - grep -F "$(printf 'signal\ttask.status\tneeds-decision:')" "$state/.wake-queue" >/dev/null \ - || fail "a pending-reply escalation was not payload-marked for branch exclusion: $(cat "$state/.wake-queue")" - pass "a pending-reply second-mate escalation is marked for main-only routing" -} - -test_ordinary_blocked_signal_payload_remains_branch_eligible() { - local dir state fakebin out status_file pid - dir=$(make_case ordinary-blocked-payload); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out" - status_file="$state/task.status" - printf 'blocked [key=dependency]: waiting for an upstream release\n' > "$status_file" - watch_bg "$state" "$fakebin" "$out" - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not exit for an ordinary blocked event" - grep -F "$(printf 'signal\ttask.status\tsignal:')" "$state/.wake-queue" >/dev/null \ - || fail "an ordinary blocked event lost branch-eligible routing: $(cat "$state/.wake-queue")" - if grep -F "$(printf 'signal\ttask.status\tneeds-decision:')" "$state/.wake-queue" >/dev/null; then - fail "an ordinary blocked event was marked as a second-mate escalation" - fi - pass "an ordinary blocked event remains branch-eligible" -} - -# A routine (non-needs-decision) captain-relevant event must keep its ordinary -# payload: only a genuine needs-decision gets the exclusion marker. -test_routine_signal_payload_not_marked_needs_decision() { - local dir state fakebin out status_file pid - dir=$(make_case routine-signal-payload); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out" - status_file="$state/task.status" - printf 'working: setup\ndone: migration complete ; needs-decision: documented in follow-up\n' > "$status_file" - watch_bg "$state" "$fakebin" "$out" - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not exit for an actionable done signal" - grep -F "$(printf 'signal\ttask.status\tneeds-decision:')" "$state/.wake-queue" >/dev/null \ - && fail "a routine done signal was incorrectly payload-marked needs-decision: $(cat "$state/.wake-queue")" - grep -F "$(printf 'signal\ttask.status\tsignal:')" "$state/.wake-queue" >/dev/null \ - || fail "a routine signal lost its ordinary payload: $(cat "$state/.wake-queue")" - pass "a routine event containing a needs-decision phrase keeps its ordinary payload, unmarked" -} - -# The reported bug, end to end through a real watcher: a crew reports something -# the captain must act on and then keeps appending routine progress, which is -# ordinary while the watcher lingers its signal grace window to coalesce a status -# write with the same turn's turn-end. Classifying only the last line reads the -# batch as routine, and because the crew IS provably working the no-verb fallback -# absorbs it too - the .seen-* suppressor then advances and nothing ever re-reads -# the event, so the work stalls with the captain never told. -test_actionable_signal_survives_a_later_routine_append() { - local dir state fakebin out drain_out status_file sig pid - dir=$(make_case actionable-masked); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out" - status_file="$state/task.status" - # Everything through "working: setup" was already classified, so this asserts - # the newly appended span, not merely a whole-file re-read. - printf 'working: setup\n' > "$status_file" - sig=$(seen_sig "$status_file"); printf '%s' "$sig" > "$state/.seen-task_status" - printf 'needs-decision: pick A or B\nworking: still tidying the branch\n' >> "$status_file" - # Positive evidence the crew is still working, so the no-verb fallback cannot - # rescue the wake: only reading the event itself can surface it. - export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' - watch_bg "$state" "$fakebin" "$out" - pid=$! - wait_for_exit "$pid" 100 \ - || { reap "$pid"; fail "watcher absorbed a needs-decision hidden behind a later working: line"; } - grep -F "signal: $status_file" "$out" >/dev/null || fail "watcher did not print the actionable signal reason" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the masked signal failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$status_file" >/dev/null \ - || fail "the masked actionable signal was not queued" - unset FM_FAKE_CREW_STATE - pass "a captain event hidden behind a later routine append is still surfaced (queue + exit)" -} - -# The captain-reported completion shape of the same masking, end to end. -test_release_completion_survives_a_later_routine_append() { - local dir state fakebin out drain_out status_file sig pid - dir=$(make_case release-masked); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out" - status_file="$state/task.status" - printf 'working: publishing\n' > "$status_file" - sig=$(seen_sig "$status_file"); printf '%s' "$sig" > "$state/.seen-task_status" - printf 'done: release 1.4.0 published and installed\nworking: cleaning the build dir\n' >> "$status_file" - export FM_FAKE_CREW_STATE='state: working · source: pane · harness busy' - watch_bg "$state" "$fakebin" "$out" - pid=$! - wait_for_exit "$pid" 100 \ - || { reap "$pid"; fail "watcher absorbed a release/install completion hidden behind later cleanup chatter"; } - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the masked completion failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$status_file" >/dev/null \ - || fail "the masked completion was not queued" - unset FM_FAKE_CREW_STATE - pass "a finished release reported before routine cleanup chatter is still surfaced" -} - -# The other direction: the fix must not turn ordinary progress into wakes. -test_routine_appends_after_a_classified_event_stay_absorbed() { - local dir state fakebin out status_file sig pid - dir=$(make_case actionable-classified); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out" - status_file="$state/task.status" - # The decision is BEHIND the classified position, so only the new routine line - # is in the span. A supervisor that re-read the whole log would wake again here. - printf 'working: setup\nneeds-decision: pick A or B\n' > "$status_file" - sig=$(seen_sig "$status_file"); printf '%s' "$sig" > "$state/.seen-task_status" - printf 'working: still tidying the branch\n' >> "$status_file" - export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' - watch_bg "$state" "$fakebin" "$out" - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid"; fail "watcher re-surfaced a decision it had already classified: $(cat "$out")" - fi - [ ! -s "$state/.wake-queue" ] || fail "a routine append after a classified decision enqueued a wake" - reap "$pid" - unset FM_FAKE_CREW_STATE - pass "a routine append after an already-classified event is absorbed (no re-wake)" -} - -test_unreadable_status_reports_once_per_file_state() { - local dir state fakebin out status_file target marker sig pid - dir=$(make_case unreadable-status); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; status_file="$state/task.status"; target="$dir/missing-status-target" - ln -s "$target" "$status_file" - marker="$state/.seen-task_status" - - watch_bg "$state" "$fakebin" "$out" - pid=$! - wait_for_exit "$pid" 100 || { reap "$pid"; fail "a dangling status symlink was not reported"; } - grep -Fx "signal: $status_file" "$out" >/dev/null \ - || fail "a dangling status symlink did not use the immediate signal path: $(cat "$out")" - sig=$(status_observed_signature "$status_file") - status_presentation_marker_reported_matches "$marker" "$sig" \ - || fail "the unreadable status report did not advance its wake signature" - [ "$(status_presentation_marker_offset "$marker" "$status_file")" = 0 ] \ - || fail "the unreadable status report advanced its classification position" - ack_stopped_cycle "$state" || fail "could not acknowledge the first unreadable-status wake" - touch "$state/.last-check" "$state/.last-heartbeat" - - watch_bg "$state" "$fakebin" "$out" - pid=$! - wait_poll_cycle "$state" "$pid" \ - || { reap "$pid"; fail "an unchanged unreadable status reported again after restart: $(cat "$out")"; } - reap "$pid" - - printf 'blocked: changed target state with a longer path\n' > "$dir/status-target-two-longer" - ln -snf "$dir/status-target-two-longer" "$status_file" - target="$dir/status-target-two-longer" - watch_bg "$state" "$fakebin" "$out" - pid=$! - wait_for_exit "$pid" 100 || { reap "$pid"; fail "a changed unreadable status did not report again"; } - [ "$(status_presentation_marker_offset "$marker" "$status_file")" = 0 ] \ - || fail "a changed unreadable status advanced its classification position" - ack_stopped_cycle "$state" || fail "could not acknowledge the changed unreadable-status wake" - - rm -f "$status_file" - cp "$target" "$status_file" - watch_bg "$state" "$fakebin" "$out" - pid=$! - wait_for_exit "$pid" 100 || { reap "$pid"; fail "a readable replacement did not surface preserved content"; } - [ "$(status_presentation_marker_offset "$marker" "$status_file")" = "$(size_of "$status_file")" ] \ - || fail "readable recovery did not classify content written before the failure" - pass "unreadable status reports are bounded without advancing classification" -} - -test_permission_recovery_surfaces_preserved_status() { - local dir state fakebin out status_file marker before_ident after_ident pid - dir=$(make_case permission-recovery); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; status_file="$state/task.status"; marker="$state/.seen-task_status" - printf 'blocked: release approval required\nworking: preserving context\n' > "$status_file" - before_ident=$(_fm_open_decisions_file_ident "$status_file") - chmod 000 "$status_file" - if [ -r "$status_file" ]; then - chmod 600 "$status_file" - pass "permission recovery skipped because permissions cannot deny reads" - return - fi - - watch_bg "$state" "$fakebin" "$out" - pid=$! - wait_for_exit "$pid" 100 || { reap "$pid"; chmod 600 "$status_file"; fail "an unreadable regular status was not reported"; } - [ "$(status_presentation_marker_offset "$marker" "$status_file")" = 0 ] \ - || { chmod 600 "$status_file"; fail "an unreadable regular status advanced its classification position"; } - ack_stopped_cycle "$state" || { chmod 600 "$status_file"; fail "could not acknowledge the unreadable regular-status wake"; } - touch "$state/.last-check" "$state/.last-heartbeat" - - watch_bg "$state" "$fakebin" "$out" - pid=$! - wait_poll_cycle "$state" "$pid" \ - || { reap "$pid"; chmod 600 "$status_file"; fail "an unchanged unreadable regular status reported again"; } - - chmod 600 "$status_file" - after_ident=$(_fm_open_decisions_file_ident "$status_file") - [ "$after_ident" = "$before_ident" ] || { reap "$pid"; fail "the permission-only recovery changed file identity"; } - wait_for_exit "$pid" 100 || { reap "$pid"; fail "readability recovery did not surface preserved content"; } - grep -Fx "signal: $status_file" "$out" >/dev/null \ - || fail "readability recovery did not use the actionable signal path: $(cat "$out")" - [ "$(status_presentation_marker_offset "$marker" "$status_file")" = "$(size_of "$status_file")" ] \ - || fail "readability recovery did not classify from the unadvanced position" - pass "permission recovery surfaces content from the unadvanced position" -} - -test_terminal_stale_surfaced() { - local dir state fakebin out drain_out capture_file window key pane_hash sig pid - dir=$(make_case terminal-stale); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - window="test:fm-done" - printf 'finished, awaiting review' > "$capture_file" - printf 'window=%s\nkind=ship\n' "$window" > "$state/done.meta" - printf 'done: PR https://example.test/pr/3\n' > "$state/done.status" - sig=$(seen_sig "$state/done.status"); printf '%s' "$sig" > "$state/.seen-done_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "finished, awaiting review") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not exit for a stale pane on a terminal status" - grep -Fx "stale: $window" "$out" >/dev/null || fail "watcher did not print the terminal stale wake" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the terminal stale failed" - grep "$(printf '\tstale\t')" "$drain_out" | grep -F "$window" >/dev/null || fail "terminal stale was not queued" - pass "a stale pane sitting on a terminal status is surfaced (queue + exit)" -} - -# --- stale pane, STALE terminal status overridden by an active run: absorbed --- -# Regression for the 2026-07 herdr false-surface incidents: a crew's own status -# log gets no new entry once firstmate hands it to a no-mistakes validation -# (AGENTS.md's sparse status-reporting contract), so the log keeps showing its -# pre-validation "done:" line as the LAST line for the run's entire (possibly -# many-minutes) duration. stale_is_terminal alone has no run-step awareness and -# would treat that leftover as still-current every time the pane goes quiet, -# immediately surfacing a crew that is actively validating. crew_is_provably_working -# must get a chance to override a captain-relevant-but-stale status line, exactly -# as it already does for a plain non-terminal one. -test_stale_terminal_status_overridden_by_active_run() { - local dir state fakebin out drain_out capture_file window key pane_hash sig pid - dir=$(make_case terminal-stale-overridden); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - window="test:fm-validating" - printf 'no-mistakes axi run: validating...' > "$capture_file" - printf 'window=%s\nkind=ship\n' "$window" > "$state/validating.meta" - # The crew reported done BEFORE firstmate triggered no-mistakes validation; - # this line never gets superseded by a newer status-log entry while the - # pipeline itself runs. - printf 'done: implementation complete, ready to validate\n' > "$state/validating.status" - sig=$(seen_sig "$state/validating.status"); printf '%s' "$sig" > "$state/.seen-validating_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "no-mistakes axi run: validating...") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' - - # Phase A: a high escalation threshold means the first sighting is absorbed, - # not surfaced, despite the captain-relevant "done:" status-log line. - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid"; fail "watcher exited for a stale terminal-looking status the run-step overrides (should absorb): $(cat "$out")" - fi - [ ! -s "$out" ] || fail "the overridden stale terminal status printed a wake reason during absorb" - [ ! -s "$state/.wake-queue" ] || fail "the overridden stale terminal status enqueued a wake during absorb" - [ "$(cat "$state/.stale-$key" 2>/dev/null || true)" = "$pane_hash" ] || fail "stale suppressor not advanced on absorb" - [ -s "$state/.stale-since-$key" ] || fail "stale-since escalation timer was not recorded on absorb" - [ ! -e "$state/.hb-surfaced-validating" ] || fail "an absorbed wake must not mark the status line as surfaced" - reap "$pid" - ack_stopped_cycle "$state" || fail "could not acknowledge the intentional phase-A watcher stop" - - # Phase B: backdate the idle timer past the threshold; the run genuinely - # wedges and the next poll escalates exactly like the non-terminal case. - echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not escalate an overridden stale terminal status past the threshold" - grep -F "stale: $window" "$out" >/dev/null || fail "escalation did not print a stale wake" - grep -F "possible wedge" "$out" >/dev/null || fail "escalation did not flag a possible wedge" - unset FM_FAKE_CREW_STATE - pass "a stale terminal-looking status is overridden and absorbed while a run is actively working, then wedge-escalated" -} - -# --- a parked keyed decision is surfaced once, not once per pane repaint ------- -# A parked decision is genuinely not working, so crew_is_provably_working rightly -# refuses to absorb its first stale pane. The separate repaint suppressor must key -# on the complete open-decision set so volatile footer changes alone stay silent. -# -# Fixture note: each phase primes .hash-/.count- so the FIRST poll already sees a -# stably stale pane, exactly as the terminal-stale cases above do. -prime_stale_pane() { # <state> <window> <pane-text> <capture-file> - local state=$1 window=$2 text=$3 capture=$4 key - printf '%s' "$text" > "$capture" - key=$(printf '%s' "$window" | tr ':/.' '___') - printf '%s' "$(hash_text "$text")" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" -} - -# Prime the .seen-* signal suppressor for a status file so only the STALE path -# under test can surface a wake (a status write otherwise fires its own signal). - -test_parked_decision_survives_pane_repaint() { - local dir state fakebin out drain_out capture window pid - dir=$(make_case parked-decision-repaint); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture="$dir/pane.txt" - window="test:fm-parked" - printf 'window=%s\nkind=ship\n' "$window" > "$state/parked.meta" - printf 'needs-decision [key=review-gate]: ship as-is or split the migration\n' \ - > "$state/parked.status" - # A live agent: this crew is waiting, not wedged. - export FM_FAKE_TMUX_CURRENT_COMMAND=claude - - # Phase A: the normal first sighting is the status signal. - printf '%s' 'awaiting decision · context 41%' > "$capture" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not surface a newly parked captain decision" - grep -F "signal: $state/parked.status" "$out" >/dev/null \ - || fail "the parked decision did not print its initial status signal" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the parked decision failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "parked.status" >/dev/null \ - || fail "the parked decision signal was not queued on its first sighting" - # Presented records stay durable until acknowledged, so phase B's "nothing new - # was queued" assertion needs phase A's own record retired first. This retires - # the record this phase just handled; it never suppresses a later one. - ack_stopped_cycle "$state" || fail "could not acknowledge the parked decision's first sighting" - - # Phase B: the pane repaints (a moving context percentage) while the SAME - # decision stays open. The hash changes; the decision does not. Nothing may - # surface. - : > "$out" - prime_stale_pane "$state" "$window" 'awaiting decision · context 39%' "$capture" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - if ! wait_live "$pid" 30; then - reap "$pid"; fail "a pane repaint re-surfaced an already-escalated parked decision: $(cat "$out")" - fi - [ ! -s "$out" ] || fail "a pane repaint printed a wake for an unchanged parked decision" - [ ! -s "$state/.wake-queue" ] || fail "a pane repaint queued a wake for an unchanged parked decision" - reap "$pid" - unset FM_FAKE_TMUX_CURRENT_COMMAND - pass "a parked keyed decision surfaces once, not once per pane repaint" -} - -test_resolved_decision_can_reopen_identically() { - local dir state fakebin out capture window pid - dir=$(make_case parked-decision-reopen); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture="$dir/pane.txt"; window="test:fm-reopen" - printf 'window=%s\nkind=ship\n' "$window" > "$state/reopen.meta" - printf 'needs-decision [key=review-gate]: ship as-is or split the migration\n' \ - > "$state/reopen.status" - prime_status_seen "$state" "$state/reopen.status" - export FM_FAKE_TMUX_CURRENT_COMMAND=claude - - prime_stale_pane "$state" "$window" 'awaiting decision · context 41%' "$capture" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not surface the original keyed decision" - FM_STATE_OVERRIDE="$state" "$DRAIN" > /dev/null 2>&1 || true - - printf 'resolved [key=review-gate]: migration will ship as-is\n' >> "$state/reopen.status" - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "the decision resolution did not surface" - grep -F "signal: $state/reopen.status" "$out" >/dev/null || fail "the resolution did not print a signal wake" - FM_STATE_OVERRIDE="$state" "$DRAIN" > /dev/null 2>&1 || true - - printf 'needs-decision [key=review-gate]: ship as-is or split the migration\n' \ - >> "$state/reopen.status" - prime_status_seen "$state" "$state/reopen.status" - prime_stale_pane "$state" "$window" 'awaiting decision again · context 38%' "$capture" - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "an identical decision did not surface after resolving and reopening" - grep -Fx "stale: $window" "$out" >/dev/null || fail "the identically reopened decision did not print a stale wake" - unset FM_FAKE_TMUX_CURRENT_COMMAND - pass "a resolved keyed decision can reopen identically and surface again" -} - -# The suppressor must be scoped to the open-decision set that was surfaced, -# never to the pane: a genuinely NEW decision on the same repainting pane is new -# work firstmate has not acted on and must wake normally. The status write's own -# signal wake is suppressed here so only the stale path can surface it. -test_new_keyed_decision_on_parked_pane_surfaces() { - local dir state fakebin out drain_out capture window pid - dir=$(make_case parked-decision-new-key); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture="$dir/pane.txt" - window="test:fm-parked2" - printf 'window=%s\nkind=ship\n' "$window" > "$state/parked2.meta" - printf 'needs-decision [key=review-gate]: ship as-is or split the migration\n' \ - > "$state/parked2.status" - prime_status_seen "$state" "$state/parked2.status" - export FM_FAKE_TMUX_CURRENT_COMMAND=claude - - prime_stale_pane "$state" "$window" 'awaiting decision · context 41%' "$capture" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not surface the first parked decision" - FM_STATE_OVERRIDE="$state" "$DRAIN" > /dev/null 2>&1 || true - - # A second, different decision opens on the same pane, which also repaints. - printf 'needs-decision [key=schema-choice]: single table or one per tenant\n' \ - >> "$state/parked2.status" - prime_status_seen "$state" "$state/parked2.status" - prime_stale_pane "$state" "$window" 'awaiting decision · context 38%' "$capture" - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "a NEW keyed decision on a parked pane did not surface" - grep -Fx "stale: $window" "$out" >/dev/null || fail "the new keyed decision did not print a stale wake" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the new keyed decision failed" - grep "$(printf '\tstale\t')" "$drain_out" | grep -F "$window" >/dev/null \ - || fail "the new keyed decision was not queued" - unset FM_FAKE_TMUX_CURRENT_COMMAND - pass "a new keyed decision on an already-parked pane still surfaces" -} - -# Suppressing the repeat surface must not cost wedge detection: a crew parked on -# an open decision whose agent then dies is a wedge suspect and must escalate -# through the shared timer. Only a confident dead verdict counts, so the live -# agent above stays absorbed while a shell-at-the-prompt endpoint escalates. -test_parked_decision_with_dead_agent_wedge_escalates() { - local dir state fakebin out capture window key pid - dir=$(make_case parked-decision-dead); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture="$dir/pane.txt" - window="test:fm-parked3" - key=$(printf '%s' "$window" | tr ':/.' '___') - printf 'window=%s\nkind=ship\n' "$window" > "$state/parked3.meta" - printf 'needs-decision [key=review-gate]: ship as-is or split the migration\n' \ - > "$state/parked3.status" - prime_status_seen "$state" "$state/parked3.status" - export FM_FAKE_TMUX_CURRENT_COMMAND=claude - - prime_stale_pane "$state" "$window" 'awaiting decision · context 41%' "$capture" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not surface the parked decision before the wedge case" - FM_STATE_OVERRIDE="$state" "$DRAIN" > /dev/null 2>&1 || true - - # The agent exits, leaving a bare shell at the endpoint, and the pane repaints - # once more. The decision is unchanged, so the repeat surface stays suppressed, - # but the wedge timer - backdated past the threshold - must still escalate. - export FM_FAKE_TMUX_CURRENT_COMMAND=zsh - prime_stale_pane "$state" "$window" 'awaiting decision · context 37%' "$capture" - echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "a parked crew whose agent died was not detected as a wedge suspect" - grep -F "stale: $window" "$out" >/dev/null || fail "the dead parked crew did not print a stale wake" - grep -F "possible wedge" "$out" >/dev/null || fail "the dead parked crew was not flagged as a possible wedge" - unset FM_FAKE_TMUX_CURRENT_COMMAND - pass "a parked crew that stops responding is still detected as a wedge suspect" -} - -# --- non-terminal stale, crew provably working: absorbed, then wedge-escalated --- -# A provably-working crew (an actively-running pipeline) legitimately sits on a -# static pane (e.g. waiting on CI), so a non-terminal stale is absorbed and only -# the wedge timer eventually escalates it - the low-churn behavior preserved. - -test_nonterminal_stale_provably_working_absorbed_then_escalated() { - local dir state fakebin out drain_out capture_file window key pane_hash sig pid - dir=$(make_case nonterminal-stale-working); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - window="test:fm-quiet" - printf 'idle building output' > "$capture_file" - printf 'window=%s\nkind=ship\n' "$window" > "$state/quiet.meta" - # Non-terminal status, and prime .seen-* so the signal scan does not pre-empt - # the stale path. - printf 'working: still compiling\n' > "$state/quiet.status" - sig=$(seen_sig "$state/quiet.status"); printf '%s' "$sig" > "$state/.seen-quiet_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "idle building output") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - # The crew's pipeline is actively running: a static pane is normal (waiting on CI). - export FM_FAKE_CREW_STATE='state: working · source: run-step · ci running' - - # Phase A: a high escalation threshold means the first sighting is absorbed. - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid"; fail "watcher exited for a fresh provably-working non-terminal stale (should absorb): $(cat "$out")" - fi - [ ! -s "$out" ] || fail "fresh provably-working stale printed a wake reason during absorb" - [ ! -s "$state/.wake-queue" ] || fail "fresh provably-working stale enqueued a wake during absorb" - [ "$(cat "$state/.stale-$key" 2>/dev/null || true)" = "$pane_hash" ] || fail "stale suppressor not advanced on absorb" - [ -s "$state/.stale-since-$key" ] || fail "stale-since escalation timer was not recorded on absorb" - reap "$pid" - ack_stopped_cycle "$state" || fail "could not acknowledge the intentional phase-A watcher stop" - - # Phase B: backdate the idle timer past the threshold; the next run escalates. - # (The subsequent-sight timer path does not re-read the crew state.) - echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not escalate a provably-working non-terminal stale past the threshold" - grep -F "stale: $window" "$out" >/dev/null || fail "escalation did not print a stale wake" - grep -F "possible wedge" "$out" >/dev/null || fail "escalation did not flag a possible wedge" - [ ! -e "$state/.stale-since-$key" ] || fail "stale-since timer was not cleared after escalation" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the wedge escalation failed" - grep "$(printf '\tstale\t')" "$drain_out" | grep -F "$window" >/dev/null || fail "wedge escalation was not queued" - pass "provably-working non-terminal stale is absorbed on first sight, then wedge-escalated past the threshold" -} - -# --- non-terminal stale, crew NOT provably working: surfaced immediately ------ -# The key requirement: a crew with no running pipeline that has gone quiet (and is -# not busy) has stopped - it may be done via interactive menus, waiting, or wedged. -# It must surface at once, never wait out the wedge timer, so these users (a -# non-no-mistakes crew, or any crew with no running pipeline) are never left hanging. - -test_nonterminal_stale_not_working_surfaced() { - local dir state fakebin out drain_out capture_file window key pane_hash sig pid - dir=$(make_case nonterminal-stale-stopped); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - window="test:fm-stopped" - printf 'idle prompt, finished' > "$capture_file" - printf 'window=%s\nkind=ship\n' "$window" > "$state/stopped.meta" - # Non-terminal status (the crew never wrote a captain-relevant verb), .seen-* - # primed so the signal scan does not pre-empt the stale path. - printf 'working: implementing\n' > "$state/stopped.status" - sig=$(seen_sig "$state/stopped.status"); printf '%s' "$sig" > "$state/.seen-stopped_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "idle prompt, finished") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - # No running pipeline; the pane is idle. NOT provably working. - export FM_FAKE_CREW_STATE='state: unknown · source: none · no current-state source available' - - # Even with a high wedge threshold, a not-provably-working stale surfaces at once. - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not surface a not-provably-working non-terminal stale at once" - grep -Fx "stale: $window" "$out" >/dev/null || fail "watcher did not print the immediate stale wake" - grep -F "possible wedge" "$out" >/dev/null && fail "an immediate stopped-crew stale was mislabeled a wedge" - [ "$(cat "$state/.stale-$key" 2>/dev/null || true)" = "$pane_hash" ] || fail "stale suppressor was not advanced on surface" - [ ! -e "$state/.stale-since-$key" ] || fail "stale-since timer should not be set when surfacing immediately" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the immediate stale failed" - grep "$(printf '\tstale\t')" "$drain_out" | grep -F "$window" >/dev/null || fail "immediate stale wake was not queued" - pass "a not-provably-working non-terminal stale is surfaced immediately (never left to wait out the timer)" -} - -# --- non-terminal stale, crew DECLARED a pause: absorbed, re-surfaced on a long -# cadence, never wedge-escalated ------------------------------------------ -# The live 2026-07-09/10 case: a crew intentionally held awaiting an upstream tool -# release (paused: ...) whose idle pane tripped repeated possible-wedge escalations -# all day. With the paused verb, its stale is absorbed like a working crew but never -# uses the wedge timer; it re-surfaces once past PAUSE_RESURFACE_SECS (anchored on -# the pause's own status-file age, so a churny idle pane cannot reset the cadence) -# for a recheck, so a forgotten pause cannot rot invisibly. -test_nonterminal_stale_paused_absorbed_then_resurfaced() { - local dir state fakebin out drain_out capture_file window key pane_hash sig pid back statusf - dir=$(make_case nonterminal-stale-paused); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - window="test:fm-held" - printf 'idle, holding for upstream' > "$capture_file" - printf 'window=%s\nkind=ship\n' "$window" > "$state/held.meta" - statusf="$state/held.status" - # A DECLARED pause (not captain-relevant), .seen-* primed so the signal scan does - # not pre-empt the stale path. - printf 'paused: holding for the upstream tool release\n' > "$statusf" - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-held_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "idle, holding for upstream") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - # crew_absorb_class reads the declared pause from fm-crew-state.sh. - export FM_FAKE_CREW_STATE='state: paused · source: status-log · holding for the upstream tool release' - - # Phase A: a fresh pause (status file just written) under a high re-surface - # threshold is absorbed - no wake, no wedge timer. - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_TMUX_CURRENT_COMMAND=zsh \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid"; fail "watcher exited for a fresh declared pause (should absorb): $(cat "$out")" - fi - [ ! -s "$out" ] || fail "fresh paused stale printed a wake reason during absorb" - [ ! -s "$state/.wake-queue" ] || fail "fresh paused stale enqueued a wake during absorb" - [ "$(cat "$state/.stale-$key" 2>/dev/null || true)" = "$pane_hash" ] || fail "stale suppressor not advanced on paused absorb" - [ -e "$state/.paused-$key" ] || fail "paused flag not recorded on absorb" - [ ! -e "$state/.stale-since-$key" ] || fail "a paused absorb must not start the wedge timer" - reap "$pid" - ack_stopped_cycle "$state" || fail "could not acknowledge the intentional paused phase-A stop" - - # Phase B: age the pause past the (now normal) threshold by backdating its - # status file, re-prime .seen-* to the new signature so the signal scan stays - # quiet, and confirm it re-surfaces as a paused recheck - never a wedge. - back=$(( $(date +%s) - 500 )) - if [ "$(uname)" = Darwin ]; then touch -mt "$(date -r "$back" '+%Y%m%d%H%M.%S')" "$statusf" - else touch -m -d "@$back" "$statusf"; fi - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-held_status" - : > "$out" - printf 'idle, holding for upstream (token 2)' > "$capture_file" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_TMUX_CURRENT_COMMAND=zsh \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not re-surface a declared pause past the threshold" - grep -F "stale: $window" "$out" >/dev/null || fail "re-surface did not print a stale wake" - grep -F "awaiting external" "$out" >/dev/null || fail "re-surface was not labeled a paused/awaiting-external recheck" - grep -F "possible wedge" "$out" >/dev/null && fail "a declared pause was mislabeled a possible wedge" - [ -e "$state/.paused-resurfaced-$key" ] || fail "the paused re-surface throttle marker was not recorded" - [ ! -e "$state/.stale-since-$key" ] || fail "a paused re-surface must not use the wedge timer" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the paused re-surface failed" - grep "$(printf '\tstale\t')" "$drain_out" | grep -F "$window" >/dev/null || fail "paused re-surface was not queued" - pass "a declared pause is absorbed on first sight, then re-surfaced as a recheck past the threshold, never wedge-escalated" -} - -# --- pause re-surface backoff ------------------------------------------------ -# The window widens per unchanged recheck and then stops widening, so a long -# healthy wait gets progressively cheap without ever becoming invisible. -test_pause_resurface_window_backs_off_and_caps() { - local w0 w1 w3 w9 wjunk wbig whuge - # shellcheck disable=SC2034 # Read by pause_resurface_window in the sourced fm-classify-lib.sh. - FM_PAUSE_RESURFACE_SECS=100 - w0=$(pause_resurface_window 0) - w1=$(pause_resurface_window 1) - w3=$(pause_resurface_window 3) - w9=$(pause_resurface_window 9) - wjunk=$(pause_resurface_window "") - # A cap large enough to overflow the shift must fail toward the base cadence, - # not into a negative window - which every age comparison reads as due, turning - # the backoff into a wake on every single poll. - # shellcheck disable=SC2034 # Read by pause_resurface_window in the sourced fm-classify-lib.sh. - FM_PAUSE_RESURFACE_MAX_STREAK=64 - wbig=$(pause_resurface_window 64) - # shellcheck disable=SC2034 # Read by pause_resurface_window in the sourced fm-classify-lib.sh. - FM_PAUSE_RESURFACE_SECS=3600 - # shellcheck disable=SC2034 # Read by pause_resurface_window in the sourced fm-classify-lib.sh. - FM_PAUSE_RESURFACE_MAX_STREAK=52 - whuge=$(pause_resurface_window 52) - unset FM_PAUSE_RESURFACE_SECS FM_PAUSE_RESURFACE_MAX_STREAK - [ "$w0" = 100 ] || fail "a first recheck must use the base window, got $w0" - [ "$w1" = 200 ] || fail "the second recheck must double the window, got $w1" - [ "$w3" = 800 ] || fail "the fourth recheck must be 8x the base window, got $w3" - [ "$w9" = 800 ] || fail "the window must stop widening at the cap, got $w9" - [ "$wjunk" = 100 ] || fail "a missing streak must fall back to the base window, got $wjunk" - [ "$wbig" -ge 100 ] || fail "an overflowing cap must never yield a window below the base, got $wbig" - [ "$wbig" = 86400 ] || fail "an overflowing cap must clamp to the bounded ceiling, got $wbig" - [ "$whuge" = 86400 ] || fail "an overflowing cap must clamp to the bounded ceiling, got $whuge" - pass "pause_resurface_window doubles per unchanged recheck, caps, and clamps a misconfigured cap" -} - -# The streak record is written by two supervisors through three helpers, so the -# reset-on-a-changed-wait cannot live in only one of them: a bump handed a wait -# that is not the one on record must start that new wait's streak, whether or not -# any caller reconciled the record first. -test_pause_streak_bump_reconciles_a_changed_wait() { - local f - f="$TMP_ROOT/streak-record" - printf '3\npaused: awaiting the upstream release\n' > "$f" - pause_streak_bump "$f" "paused: awaiting the upstream release" - [ "$(pause_streak_count "$f")" = 4 ] \ - || fail "a recheck of the same wait must continue its streak, got $(pause_streak_count "$f")" - pause_streak_bump "$f" "paused: awaiting the captain merge call" - [ "$(pause_streak_count "$f")" = 1 ] \ - || fail "a recheck of a changed wait must start over, got $(pause_streak_count "$f")" - pause_streak_sync "$f" "paused: awaiting the captain merge call" \ - && fail "the bump did not leave the changed wait on record" - rm -f "$f" - pass "pause_streak_bump reconciles the wait on record before counting a recheck" -} - -# The live 2026-08-04 case behind issue 47: three tasks correctly parked on one -# captain-owned merge decision re-surfaced on a fixed cadence, each recheck -# costing a supervision turn to confirm a wait that had not changed. The recheck -# must survive - a forgotten hold cannot rot invisibly - but an UNCHANGED wait -# must cost less each time. Phase E is the disconfirming half: nothing about the -# backoff may reach a crew that never declared a wait, which still absorbs on the -# wedge timer and still escalates as a possible wedge at the unchanged threshold; -# and phases C and D are the boundary between them - the widened cadence is -# earned by one unchanged wait and dies the moment that wait changes, whether it -# is replaced by a different wait or dropped entirely. -test_paused_resurface_backs_off_while_wedge_still_escalates() { - local dir state fakebin out capture_file window key pane_hash sig pid back statusf wakes - dir=$(make_case paused-resurface-backoff); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt" - window="test:fm-parked" - printf 'idle awaiting the merge decision' > "$capture_file" - printf 'window=%s\nkind=ship\n' "$window" > "$state/parked.meta" - statusf="$state/parked.status" - printf 'paused: awaiting the merge decision on the open PR\n' > "$statusf" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "idle awaiting the merge decision") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - export FM_FAKE_CREW_STATE='state: paused · source: status-log · awaiting the merge decision on the open PR' - - # Phase A: the wait is well past the base window, so the FIRST recheck fires. - back=$(( $(date +%s) - 500 )) - set_mtime "$back" "$statusf" - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-parked_status" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_TMUX_CURRENT_COMMAND=zsh \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "the first recheck of a long declared wait did not re-surface" - grep -F "awaiting external" "$out" >/dev/null || fail "the first recheck was not a paused recheck" - [ "$(pause_streak_count "$state/.paused-streak-$key")" = 1 ] \ - || fail "the first recheck did not record a re-surface streak" - # Each phase below asserts the rechecks IT produced. Presented records stay - # durable until acknowledged, so every phase retires its own before the next - # watcher arms; otherwise the next arm legitimately resurfaces the unhandled - # row and the phase never exercises its own cadence decision. - ack_stopped_cycle "$state" || fail "could not acknowledge the first recheck" - - # Phase B: the wait has not changed. One base window later is now too soon - - # the second recheck must wait for the DOUBLED window before firing again. - : > "$out" - set_mtime "$(( $(date +%s) - 300 ))" "$state/.paused-resurfaced-$key" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_TMUX_CURRENT_COMMAND=zsh \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - if ! wait_live "$pid" 30; then - reap "$pid"; fail "an unchanged wait re-surfaced again inside the widened window: $(cat "$out")" - fi - reap "$pid" - [ ! -s "$out" ] || fail "an unchanged wait printed a wake inside the widened window: $(cat "$out")" - # This phase deliberately queued nothing, so only the stopped cycle's own - # downtime marker needs retiring; the queue is left exactly as it is. - retire_downtime_marker "$state" || fail "could not retire the inside-window stop" - - # Past the widened window it DOES fire again, so the wait still cannot rot. - set_mtime "$(( $(date +%s) - 600 ))" "$state/.paused-resurfaced-$key" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_TMUX_CURRENT_COMMAND=zsh \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "a declared wait past its widened window did not re-surface" - grep -F "awaiting external" "$out" >/dev/null || fail "the widened-window recheck was not a paused recheck" - grep -F "possible wedge" "$out" >/dev/null && fail "a declared wait was mislabeled a possible wedge" - [ "$(pause_streak_count "$state/.paused-streak-$key")" = 2 ] \ - || fail "the second recheck did not widen the streak further" - wakes=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w { n++ } END { print n + 0 }' "$state/.wake-queue") - [ "$wakes" -eq 1 ] || fail "expected exactly 1 recheck from the widened window, got $wakes" - ack_stopped_cycle "$state" || fail "could not acknowledge the widened-window recheck" - - # Phase C: the widened cadence is earned by ONE wait, so a DIFFERENT declared - # wait cannot inherit it. The crew replaces its paused line with another; the - # new wait has stood for one base window, which is well inside the window the - # previous wait had widened to, and it must be rechecked anyway. - printf 'paused: awaiting the security review sign-off\n' > "$statusf" - set_mtime "$(( $(date +%s) - 300 ))" "$statusf" - set_mtime "$(( $(date +%s) - 300 ))" "$state/.paused-resurfaced-$key" - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-parked_status" - export FM_FAKE_CREW_STATE='state: paused · source: status-log · awaiting the security review sign-off' - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_TMUX_CURRENT_COMMAND=zsh \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 \ - || fail "a replaced wait inherited the previous wait's widened window instead of the base one" - grep -F "awaiting external" "$out" >/dev/null || fail "the replaced wait's recheck was not a paused recheck" - grep -F "possible wedge" "$out" >/dev/null && fail "a replaced declared wait was mislabeled a possible wedge" - [ "$(pause_streak_count "$state/.paused-streak-$key")" = 1 ] \ - || fail "the streak survived the wait that earned it being replaced" - wakes=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w { n++ } END { print n + 0 }' "$state/.wake-queue") - [ "$wakes" -eq 1 ] || fail "expected exactly 1 recheck from the replaced wait, got $wakes" - ack_stopped_cycle "$state" || fail "could not acknowledge the replaced wait's recheck" - - # Phase D: the same death by the other route. The moment the crew stops - # declaring a wait at all, the streak goes with the rest of the pause tracking, - # so this crew going quiet again later starts back at the base window and never - # inherits a cadence widened by something else. - printf 'working: resumed, the merge decision landed\n' > "$statusf" - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-parked_status" - export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_TMUX_CURRENT_COMMAND=zsh \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_live "$pid" 30 || true - reap "$pid" - [ ! -e "$state/.paused-streak-$key" ] \ - || fail "the widened cadence outlived the wait that earned it (streak $(cat "$state/.paused-streak-$key"))" - [ ! -e "$state/.paused-$key" ] || fail "pause tracking survived the crew resuming" - - # Phase E: a crew that declared NO wait is untouched by any of this. It is - # absorbed only while provably working, and once its idle time crosses the - # wedge threshold it still escalates as a possible wedge - at the unchanged - # FM_STALE_ESCALATE_SECS threshold, which no backoff may widen. - dir=$(make_case paused-backoff-wedge-control); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt" - window="test:fm-quiet" - printf 'idle, no declared wait' > "$capture_file" - printf 'window=%s\nkind=ship\n' "$window" > "$state/quiet.meta" - printf 'working: pushed the branch\n' > "$state/quiet.status" - sig=$(seen_sig "$state/quiet.status"); printf '%s' "$sig" > "$state/.seen-quiet_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "idle, no declared wait") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - printf '%s' "$pane_hash" > "$state/.stale-$key" - printf '%s\n' $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" - export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_TMUX_CURRENT_COMMAND=zsh \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 \ - FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "an undeclared idle crew past the wedge threshold did not escalate" - grep -F "possible wedge" "$out" >/dev/null || fail "the undeclared idle crew was not flagged a possible wedge" - [ ! -e "$state/.paused-streak-$key" ] || fail "the pause backoff leaked onto a crew that declared no wait" - pass "an unchanged declared wait rechecks on a widening cadence while a genuine wedge still escalates" -} - -# A captain-held crew can leave a stable backend endpoint after its agent exits. -# fm-crew-state then authoritatively reports stopped rather than paused, but the -# confirmed-dead agent plus the declared wait or captain-held transfer must retain -# bounded pause handling. -# A still-live agent under a CAPTAIN-HELD transfer is the disconfirming case: the -# crew declared nothing there, so it must still surface once, while the unchanged -# hash must not append the same wake on every watcher re-arm. -test_exited_declared_pause_is_bounded_but_live_captain_held_surfaces() { - local dir state fakebin out capture_file statusf window key pane_hash sig pid back round wakes bare - dir=$(make_case exited-declared-pause); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; statusf="$state/held.status" - window="test:fm-held" - printf 'idle bare shell after agent exit\n' > "$capture_file" - printf 'window=%s\nkind=ship\nharness=grok\nbackend=tmux\n' "$window" > "$state/held.meta" - printf 'paused: held per captain while an external decision is pending\n' > "$statusf" - back=$(( $(date +%s) - 500 )) - if [ "$(uname)" = Darwin ]; then touch -mt "$(date -r "$back" '+%Y%m%d%H%M.%S')" "$statusf" - else touch -m -d "@$back" "$statusf"; fi - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-held_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "idle bare shell after agent exit") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - - round=1 - while [ "$round" -le 6 ]; do - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_TMUX_CURRENT_COMMAND=zsh FM_FAKE_CREW_STATE='state: stopped · source: pane · bare shell' \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" >> "$out" & - pid=$! - if wait_poll_cycle "$state" "$pid"; then - reap "$pid" - elif kill -0 "$pid" 2>/dev/null; then - reap "$pid" - fail "dead-agent watcher round $round timed out before completing a poll cycle" - else - wait "$pid" || fail "dead-agent watcher round $round failed" - fi - round=$((round + 1)) - done - # A watcher that queues nothing never creates .wake-queue, so these counts - # read a path that may legitimately be absent. awk aborts on a missing file - # before END runs, which collapses the count to the empty string and turns the - # next comparison into an "integer expression expected" error - reported as a - # flood of an unprintable number of wakes instead of the real contract breach - # the grep below names. No queue means no wakes, per the drain-count read at - # the end of this file. - wakes=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w { n++ } END { print n + 0 }' "$state/.wake-queue" 2>/dev/null || echo 0) - bare=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w && $5 == "stale: " w { n++ } END { print n + 0 }' "$state/.wake-queue" 2>/dev/null || echo 0) - [ "$wakes" -le 1 ] || fail "dead-agent declared pause flooded $wakes stale wakes across six unchanged polls" - [ "$bare" -eq 0 ] || fail "dead-agent declared pause surfaced as $bare bare stopped-crew wakes" - grep -F "awaiting external" "$state/.wake-queue" >/dev/null \ - || fail "dead-agent declared pause did not use the bounded paused recheck" - - dir=$(make_case exited-captain-held); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; statusf="$state/held.status" - window="test:fm-held" - printf 'idle bare shell after captain-held transfer\n' > "$capture_file" - printf 'window=%s\nkind=ship\nharness=grok\nbackend=tmux\n' "$window" > "$state/held.meta" - printf 'captain-held [key=route]: tracked by held-decision-route\n' > "$statusf" - back=$(( $(date +%s) - 500 )) - if [ "$(uname)" = Darwin ]; then touch -mt "$(date -r "$back" '+%Y%m%d%H%M.%S')" "$statusf" - else touch -m -d "@$back" "$statusf"; fi - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-held_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "idle bare shell after captain-held transfer") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_TMUX_CURRENT_COMMAND=zsh FM_FAKE_CREW_STATE='state: stopped · source: pane · bare shell' \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "captain-held dead-agent pane did not re-surface on the bounded cadence" - grep -F "awaiting the captain" "$state/.wake-queue" >/dev/null \ - || fail "captain-held dead-agent pane surfaced as a stopped crew instead of a captain-owned recheck: $(cat "$state/.wake-queue")" - grep -F "awaiting external" "$state/.wake-queue" >/dev/null \ - && fail "captain-held dead-agent pane borrowed the pause verb's external-wait wording" - - # The disconfirming half. A captain-held transfer is firstmate's own record, not - # the crew declaring that this pane is idle on purpose, so the live-agent rule - # that the declared pause no longer pays is still in force here: surface once, - # then hold the same hash without re-appending on every re-arm. - dir=$(make_case alive-captain-held); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; statusf="$state/gate.status" - window="test:fm-gate" - printf 'idle under a captain-held transfer\n' > "$capture_file" - printf 'window=%s\nkind=ship\nharness=grok\nbackend=tmux\n' "$window" > "$state/gate.meta" - printf 'captain-held [key=route]: tracked by held-decision-route\n' > "$statusf" - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-gate_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "idle under a captain-held transfer") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - - # First sight must surface promptly so a live captain-held pane is not hidden - # behind the pause cadence. - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_TMUX_CURRENT_COMMAND=grok FM_FAKE_CREW_STATE='state: unknown · source: none · live pane under a captain-held transfer' \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" >> "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "live captain-held pane did not surface immediately" - ack_stopped_cycle "$state" || fail "could not acknowledge the immediate captain-held surface" - - # Re-arm with the stale timer already beyond the wedge threshold. This is the - # exact unchanged-hash fallback after the immediate surface: it must retain - # the pause cadence and discard any residual wedge timer instead of emitting - # a second possible-wedge wake. - printf '%s\n' $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_TMUX_CURRENT_COMMAND=grok FM_FAKE_CREW_STATE='state: unknown · source: none · live pane under a captain-held transfer' \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 FM_PAUSE_RESURFACE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" >> "$out" & - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid" - fail "live captain-held pane escalated on the wedge timer after its immediate surface: $(cat "$out")" - fi - [ -e "$state/.paused-$key" ] || { reap "$pid"; fail "live captain-held pane lost its pause cadence marker"; } - [ ! -e "$state/.stale-since-$key" ] || { reap "$pid"; fail "live captain-held pane retained the wedge timer"; } - reap "$pid" - wakes=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w { n++ } END { print n + 0 }' "$state/.wake-queue" 2>/dev/null || echo 0) - bare=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w && $5 == "stale: " w { n++ } END { print n + 0 }' "$state/.wake-queue" 2>/dev/null || echo 0) - [ "$wakes" -eq 0 ] || fail "acknowledged captain-held surface replayed $wakes wakes" - [ "$bare" -eq 0 ] || fail "acknowledged captain-held bare stale remained queued" - pass "exited declared-pause and captain-held panes use bounded pause cadence while a live captain-held pane still surfaces once" -} - -# Issue 142, the live half of the family issue 67 closed. The generated brief -# tells every crew to append `paused: <why>` and stop, and on every verified -# harness "stop" means end the turn, so the agent stays ALIVE at its prompt. The -# watcher used to honour a declared pause only for a confidently dead agent, which -# made the designed widening cadence unreachable for exactly the crew the brief -# creates: it re-surfaced a bare `stale: <window>` every few minutes for as long -# as the wait lasted (measured at 320s, 341s and 449s gaps on a task known to be -# waiting on a human). -# -# The discriminator is the artifact set, not log text. pause_streak_bump is the -# only writer of .paused-streak-<key> and is called only by handle_paused_stale, -# so a present streak file proves the designed path ran, while a missing one -# beside freshly stamped .paused-*/.paused-rechecked-*/.paused-resurfaced-* -# markers is the exact signature of surface_nonterminal_stale having written them -# instead - the field evidence from the two parked 2026-08-14 tasks. -# -# Phase C is the disconfirming half: a live pane that declared NO wait is -# untouched and still escalates as a possible wedge on the unchanged threshold. -test_live_declared_pause_is_absorbed_on_the_designed_cadence() { - local dir state fakebin out capture_file statusf window key pane_hash sig pid round wakes bare - dir=$(make_case live-declared-pause); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; statusf="$state/parked.status" - window="test:fm-parked" - printf 'idle at its prompt, parked on the hardware bench\n' > "$capture_file" - printf 'window=%s\nkind=ship\nharness=grok\nbackend=tmux\n' "$window" > "$state/parked.meta" - printf 'paused: awaiting the captain hardware bench run\n' > "$statusf" - set_mtime "$(( $(date +%s) - 500 ))" "$statusf" - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-parked_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "idle at its prompt, parked on the hardware bench") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - - # Phase A: the agent is ALIVE (pane_current_command matches the recorded - # harness) and the wait is past the base window, so this is a recheck on the - # designed cadence - annotated, streak-counted, and never a bare stale. - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_TMUX_CURRENT_COMMAND=grok \ - FM_FAKE_CREW_STATE='state: paused · source: status-log · awaiting the captain hardware bench run' \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 \ - FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "a live crew's declared pause never reached the designed recheck" - grep -F "awaiting external" "$out" >/dev/null \ - || fail "a live crew's declared pause did not use the annotated paused recheck: $(cat "$out")" - grep -F "possible wedge" "$out" >/dev/null && fail "a live crew's declared pause was mislabeled a possible wedge" - [ "$(pause_streak_count "$state/.paused-streak-$key")" = 1 ] \ - || fail "handle_paused_stale never ran for a live declared pause (no re-surface streak recorded)" - bare=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w && $5 == "stale: " w { n++ } END { print n + 0 }' "$state/.wake-queue") - [ "$bare" -eq 0 ] || fail "a live declared pause emitted $bare bare stale wakes" - - # Phase B: the wait has not changed, so the widened window now owns this pane. - # Six re-arms inside it, each with a DIFFERENT pane hash (a live agent repaints - # its prompt), must add no wake at all: the pre-fix leak was one bare wake per - # first-sighting of each new hash, which is what made it fire every few minutes. - round=1 - while [ "$round" -le 6 ]; do - printf 'idle at its prompt, parked on the hardware bench (repaint %s)\n' "$round" > "$capture_file" - printf '%s' "$(hash_text "$(cat "$capture_file")")" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_TMUX_CURRENT_COMMAND=grok \ - FM_FAKE_CREW_STATE='state: paused · source: status-log · awaiting the captain hardware bench run' \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 \ - FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" >> "$out" & - pid=$! - if wait_live "$pid" 15; then reap "$pid"; else wait "$pid" || true; fi - round=$((round + 1)) - done - wakes=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w { n++ } END { print n + 0 }' "$state/.wake-queue") - [ "$wakes" -eq 1 ] \ - || fail "a live declared pause flooded $wakes stale wakes across six repaints inside its widened window" - [ -e "$state/.paused-streak-$key" ] \ - || fail "a repainting live pane dropped the backoff streak its unchanged wait had earned" - - # Phase C: the disconfirming control. Same live agent, same silence, but no - # declared wait - it must still escalate as a possible wedge, unchanged. - dir=$(make_case live-undeclared-quiet); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt" - window="test:fm-quiet" - printf 'idle at its prompt, nothing declared\n' > "$capture_file" - printf 'window=%s\nkind=ship\nharness=grok\nbackend=tmux\n' "$window" > "$state/quiet.meta" - printf 'working: pushed the branch\n' > "$state/quiet.status" - sig=$(seen_sig "$state/quiet.status"); printf '%s' "$sig" > "$state/.seen-quiet_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "idle at its prompt, nothing declared") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - printf '%s' "$pane_hash" > "$state/.stale-$key" - printf '%s\n' $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_TMUX_CURRENT_COMMAND=grok \ - FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 \ - FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "a live crew that declared no wait stopped escalating past the wedge threshold" - grep -F "possible wedge" "$out" >/dev/null \ - || fail "a live crew that declared no wait was not flagged a possible wedge: $(cat "$out")" - [ ! -e "$state/.paused-streak-$key" ] || fail "the pause cadence leaked onto a crew that declared no wait" - pass "a live crew's declared pause is absorbed on the designed widening cadence while an undeclared quiet pane still wedges" -} - -# The ordering half of issue 142. surface_nonterminal_stale queued its wake FIRST -# and only afterwards read the status line, found the pause, and stamped the -# .paused-* markers - so the pause was recognised one step too late to suppress -# the wake it had just queued. -# -# The race is made deterministic through the real seam that produced it: the -# classifier reads the status line, then calls fm-crew-state.sh, and the crew is -# free to append its pause during that call. This fake does exactly that, so the -# classifier's own read predates the declaration and its verdict routes the pane -# to surface_nonterminal_stale with a paused status already on disk. Nothing but -# the read-before-queue ordering can save it: a bare stale here, with the three -# .paused-* markers stamped and no streak file, is the pre-fix signature exactly. -test_declared_pause_landing_mid_classification_never_emits_a_bare_stale() { - local dir state fakebin out capture_file statusf window key pane_hash sig pid bare - dir=$(make_case pause-lands-mid-classification); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; statusf="$state/racing.status" - window="test:fm-racing" - printf 'idle at its prompt\n' > "$capture_file" - printf 'window=%s\nkind=ship\nharness=grok\nbackend=tmux\n' "$window" > "$state/racing.meta" - printf 'working: running the bench sweep\n' > "$statusf" - set_mtime "$(( $(date +%s) - 500 ))" "$statusf" - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-racing_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "idle at its prompt") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - - # The crew declares its wait while the classifier is mid-read. Keep the status - # mtime old so the declared wait is immediately past its re-surface window, - # and leave the .seen-* signature matching so this write is not itself a signal. - cat > "$fakebin/fm-crew-state.sh" <<SH -#!/usr/bin/env bash -set -u -printf 'paused: awaiting the captain hardware bench run\n' > "$statusf" -$(declare -f set_mtime) -set_mtime "\$(( \$(date +%s) - 500 ))" "$statusf" -$(declare -f seen_sig) -printf '%s' "\$(seen_sig "$statusf")" > "$state/.seen-racing_status" -printf 'state: unknown · source: none · read before the wait was declared\n' -exit 0 -SH - chmod +x "$fakebin/fm-crew-state.sh" - - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_TMUX_CURRENT_COMMAND=grok \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 \ - FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "the pane that declared its wait mid-classification never surfaced at all" - bare=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w && $5 == "stale: " w { n++ } END { print n + 0 }' "$state/.wake-queue") - [ "$bare" -eq 0 ] \ - || fail "the wake was queued before the status was read: $bare bare stale wakes for an already-declared pause" - grep -F "awaiting external" "$state/.wake-queue" >/dev/null \ - || fail "a pause already on disk did not reach the annotated recheck: $(cat "$state/.wake-queue")" - [ -e "$state/.paused-streak-$key" ] \ - || fail "the three pause markers were stamped without the streak file: surface_nonterminal_stale wrote them, not handle_paused_stale" - pass "a pause already on disk is recognised before the wake is queued, never after" -} - -# A pause is honoured on the crew's declaration, but liveness keeps its RECOVERY -# job: the same parked pane whose agent later exits must still be reachable as a -# stopped crew rather than disappearing behind the wait it declared while alive. -test_live_declared_pause_still_recoverable_once_its_agent_dies() { - local dir state fakebin out capture_file statusf window key pane_hash sig pid - dir=$(make_case parked-then-dead); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; statusf="$state/parked.status" - window="test:fm-parked" - printf 'idle at its prompt, parked\n' > "$capture_file" - printf 'window=%s\nkind=ship\nharness=grok\nbackend=tmux\n' "$window" > "$state/parked.meta" - printf 'paused: awaiting the captain hardware bench run\n' > "$statusf" - set_mtime "$(( $(date +%s) - 500 ))" "$statusf" - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-parked_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "idle at its prompt, parked") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - - # Alive: absorbed on the designed cadence, streak recorded. - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_TMUX_CURRENT_COMMAND=grok \ - FM_FAKE_CREW_STATE='state: paused · source: status-log · awaiting the captain hardware bench run' \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 \ - FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "the live parked pane never reached its declared-wait recheck" - [ "$(pause_streak_count "$state/.paused-streak-$key")" = 1 ] \ - || fail "the live parked pane never reached the designed pause path" - ack_stopped_cycle "$state" || fail "could not acknowledge the live parked pane's recheck" - - # The agent then exits, leaving a bare shell on the same endpoint and the same - # declared wait. fm-crew-state falls back to stopped; the recovery signal - the - # crew-state read that reports a stopped crew, and the endpoint liveness behind - # it - must still be exercised, and the pane must stay on the bounded recheck - # rather than going silent. - printf 'bare shell after the agent exited\n' > "$capture_file" - printf '%s' "$(hash_text "$(cat "$capture_file")")" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - rm -f "$state/.paused-rechecked-$key" - set_mtime "$(( $(date +%s) - 900 ))" "$state/.paused-resurfaced-$key" - set_mtime "$(( $(date +%s) - 900 ))" "$statusf" - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-parked_status" - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_TMUX_CURRENT_COMMAND=zsh \ - FM_FAKE_CREW_STATE='state: stopped · source: pane · bare shell' \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 \ - FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "a parked pane whose agent died went silent instead of re-surfacing" - grep -F "awaiting external" "$out" >/dev/null \ - || fail "a parked pane whose agent died did not re-surface on the bounded recheck: $(cat "$out")" - [ "$(pause_streak_count "$state/.paused-streak-$key")" = 2 ] \ - || fail "the dead-agent recheck did not continue the same wait's streak" - pass "a parked pane whose agent later dies stays on the bounded recheck, so liveness keeps its recovery job" -} - -# A dead worker reaches handle_paused_stale rather than the live fallback above. -# When one declared wait directly replaces another, the existing -# throttle belongs to the old declaration and must not suppress the new wait's -# first inspection merely because its timestamp is still young. -test_absorbed_replacement_wait_does_not_inherit_the_old_throttle() { - local spec name initial replacement expected dir state fakebin out capture_file - local statusf window key sig back pid wakes - for spec in \ - 'paused-replacement|paused: waiting on validation run one|paused: waiting on validation run two|awaiting external' \ - 'captain-held-replacement|captain-held [key=route]: awaiting the routing call|captain-held [key=release]: awaiting the release call|awaiting the captain' - do - name=${spec%%|*}; spec=${spec#*|} - initial=${spec%%|*}; spec=${spec#*|} - replacement=${spec%%|*}; expected=${spec#*|} - dir=$(make_case "$name"); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; statusf="$state/held.status" - window="test:fm-held" - printf 'idle after agent exit\n' > "$capture_file" - printf 'window=%s\nkind=ship\nharness=grok\nbackend=tmux\n' "$window" > "$state/held.meta" - printf '%s\n' "$initial" > "$statusf" - back=$(( $(date +%s) - 500 )) - if [ "$(uname)" = Darwin ]; then touch -mt "$(date -r "$back" '+%Y%m%d%H%M.%S')" "$statusf" - else touch -m -d "@$back" "$statusf"; fi - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-held_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - printf '%s' "$(hash_text 'idle after agent exit')" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_TMUX_CURRENT_COMMAND=zsh FM_FAKE_CREW_STATE='state: stopped · source: pane · bare shell' \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" >> "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "[$name] initial declared wait did not re-surface" - ack_stopped_cycle "$state" || fail "[$name] could not acknowledge the initial declared wait" - - printf '%s\n' "$replacement" >> "$statusf" - # The fork rechecks a replacement at its own base deadline, rather than - # immediately. It must not inherit the prior wait's widened window. - set_mtime "$(( $(date +%s) - 300 ))" "$statusf" - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-held_status" - printf 'idle after replacement wait\n' > "$capture_file" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_TMUX_CURRENT_COMMAND=zsh FM_FAKE_CREW_STATE='state: stopped · source: pane · bare shell' \ - FM_WATCH_HANDLING_SUCCESSOR=1 \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" >> "$out" & - pid=$! - wait_for_exit "$pid" 100 \ - || { reap "$pid"; fail "[$name] replacement declared wait inherited the old throttle"; } - wakes=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w { n++ } END { print n + 0 }' \ - "$state/.wake-queue" 2>/dev/null || echo 0) - [ "$wakes" -eq 1 ] || fail "[$name] replacement declared wait produced $wakes wakes instead of one" - grep -F "$expected" "$state/.wake-queue" >/dev/null \ - || fail "[$name] replacement declared wait used the wrong recheck reason: $(cat "$state/.wake-queue")" - done - pass "absorbed paused and captain-held replacements each start their own re-surface cadence" -} - -# Run one watcher round against a parked-worker fixture, so a round differs only -# in the pane contents the case just wrote. Armed the way fm-watch-arm.sh arms a -# successor after firstmate handled a wake, because that is what a supervision -# turn actually does and it is the only arm that stays in the poll loop instead of -# re-announcing the previous round's downtime - without it a round exits on -# `check: rearm-resurface` before it ever reaches the stale path, and every -# absorb assertion below passes vacuously. A live agent (pane_current_command -# matching the recorded harness) on an idle pane is the exact population -# pause_state_class answers `none` for. -# <mode> `exit` requires the watcher to surface and exit; `absorb` requires it to -# survive whole poll cycles - enough to see the new hash, count it stable, and -# reach the stale path. Returns 1 when the watcher does the other thing. -parked_watch_round() { # <state> <fakebin> <out> <capture> <window> <exit|absorb> - local state=$1 fakebin=$2 out=$3 capture=$4 window=$5 mode=$6 pid cycles=0 - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture" \ - FM_FAKE_TMUX_CURRENT_COMMAND=grok \ - FM_FAKE_CREW_STATE='state: paused · source: status-log · parked' \ - FM_WATCH_HANDLING_SUCCESSOR=1 \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_PAUSE_RESURFACE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" >> "$out" & - pid=$! - if [ "$mode" = exit ]; then - wait_for_exit "$pid" 100 || { reap "$pid"; return 1; } - return 0 - fi - while [ "$cycles" -lt 4 ]; do - wait_poll_cycle "$state" "$pid" 300 || { reap "$pid"; return 1; } - cycles=$((cycles + 1)) - done - reap "$pid" - return 0 -} - -# A live captain-held pane earns one initial inspection and then a -# declaration-scoped throttle. Declared paused workers use the fork's widening -# cadence instead, covered by the live-declared-pause cases above. -test_live_declared_wait_churn_honors_the_resurface_throttle() { - local name status_line dir state fakebin out capture_file statusf window key - local sig round wakes bare text throttle replacement - name=captain-held-churn - status_line='captain-held [key=route]: awaiting the captain on the routing call' - dir=$(make_case "$name"); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; statusf="$state/parked.status" - window="test:fm-parked" - printf 'window=%s\nkind=ship\nharness=grok\nbackend=tmux\n' "$window" > "$state/parked.meta" - printf '%s\n' "$status_line" > "$statusf" - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-parked_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - throttle="$state/.paused-resurfaced-$key" - - # First sight of a parked-but-live worker must still surface: the state is - # inconclusive and firstmate has to look at it. - text='parked, elapsed 1s' - printf '%s' "$text" > "$capture_file" - printf '%s' "$(hash_text "$text")" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - parked_watch_round "$state" "$fakebin" "$out" "$capture_file" "$window" exit \ - || fail "[$name] first sight of a parked live worker did not surface" - ack_stopped_cycle "$state" || fail "[$name] could not acknowledge the first surface" - [ -e "$throttle" ] || fail "[$name] the first surface recorded no re-surface throttle" - - # The pane now churns while the SAME declared wait stands, each round fully - # handled as a real supervision turn would. Every one of these used to alarm. - round=2 - while [ "$round" -le 4 ]; do - printf 'parked, elapsed %ss' "$round" > "$capture_file" - parked_watch_round "$state" "$fakebin" "$out" "$capture_file" "$window" absorb \ - || fail "[$name] watcher exited during churn round $round instead of supervising through it" - wakes=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w { n++ } END { print n + 0 }' \ - "$state/.wake-queue" 2>/dev/null || echo 0) - [ "$wakes" -eq 0 ] \ - || fail "[$name] pane churn re-alarmed a parked worker $wakes time(s) inside the re-surface window" - [ -e "$throttle" ] || fail "[$name] pane churn cleared the re-surface throttle" - round=$((round + 1)) - done - - # A direct wait-to-wait transition starts a NEW declaration even though the - # same window remains parked. Its first sight must not inherit the previous - # declaration's throttle, or an unrelated replacement wait can stay silent - # for nearly the whole old cadence window. - replacement='captain-held [key=release]: awaiting the captain on the release call' - printf '%s\n' "$replacement" >> "$statusf" - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-parked_status" - printf 'replacement wait, elapsed 1s' > "$capture_file" - parked_watch_round "$state" "$fakebin" "$out" "$capture_file" "$window" exit \ - || fail "[$name] a replacement declared wait inherited the previous wait's re-surface throttle" - wakes=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w { n++ } END { print n + 0 }' \ - "$state/.wake-queue" 2>/dev/null || echo 0) - bare=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w && $5 == "stale: " w { n++ } END { print n + 0 }' \ - "$state/.wake-queue" 2>/dev/null || echo 0) - [ "$wakes" -eq 1 ] || fail "[$name] replacement declared wait produced $wakes first wakes instead of one" - [ "$bare" -eq 1 ] || fail "[$name] replacement declared wait changed the wake identity: $(cat "$state/.wake-queue")" - ack_stopped_cycle "$state" || fail "[$name] could not acknowledge the replacement wait's first surface" - - printf 'replacement wait, elapsed 2s' > "$capture_file" - parked_watch_round "$state" "$fakebin" "$out" "$capture_file" "$window" absorb \ - || fail "[$name] replacement wait re-alarmed inside its own re-surface window" - wakes=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w { n++ } END { print n + 0 }' \ - "$state/.wake-queue" 2>/dev/null || echo 0) - [ "$wakes" -eq 0 ] || fail "[$name] replacement wait re-alarmed $wakes time(s) inside its own re-surface window" - - # End of the window: the wait must re-surface exactly once, on the same plain - # identity as before, so absorbing churn never becomes silence. - set_mtime "$(( $(date +%s) - 2000 ))" "$throttle" - printf 'parked, elapsed 5s' > "$capture_file" - parked_watch_round "$state" "$fakebin" "$out" "$capture_file" "$window" exit \ - || fail "[$name] a parked worker did not re-surface once its re-surface window elapsed" - wakes=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w { n++ } END { print n + 0 }' \ - "$state/.wake-queue" 2>/dev/null || echo 0) - bare=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w && $5 == "stale: " w { n++ } END { print n + 0 }' \ - "$state/.wake-queue" 2>/dev/null || echo 0) - [ "$wakes" -eq 1 ] || fail "[$name] elapsed re-surface window produced $wakes wakes instead of one" - [ "$bare" -eq 1 ] || fail "[$name] elapsed re-surface changed the wake identity: $(cat "$state/.wake-queue")" - pass "a live captain-held worker surfaces once, absorbs pane churn for the whole re-surface window, then re-surfaces when it elapses" -} - -test_secondmate_paused_resurfaces_in_normal_mode() { - local dir state fakebin out capture_file statusf window key pane_hash sig pid back - dir=$(make_case secondmate-paused-resurface); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; statusf="$state/secondmate-held.status" - window="test:fm-secondmate-held" - printf 'idle awaiting external\n' > "$capture_file" - printf 'window=%s\nkind=secondmate\n' "$window" > "$state/secondmate-held.meta" - printf 'paused: awaiting the upstream release\n' > "$statusf" - back=$(( $(date +%s) - 500 )) - if [ "$(uname)" = Darwin ]; then touch -mt "$(date -r "$back" '+%Y%m%d%H%M.%S')" "$statusf" - else touch -m -d "@$back" "$statusf"; fi - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-secondmate-held_status" - key=$(printf '%s' "$window" | tr '.:/' '___') - pane_hash=$(hash_text "idle awaiting external") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - export FM_FAKE_CREW_STATE='state: paused · source: status-log · awaiting the upstream release' - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not re-surface a paused secondmate" - grep -F "stale: $window" "$out" >/dev/null || fail "paused secondmate did not emit a stale recheck" - grep -F "awaiting external" "$out" >/dev/null || fail "paused secondmate recheck omitted its external-wait reason" - grep -F "awaiting the captain" "$out" >/dev/null && fail "paused secondmate recheck named the captain instead of its external dependency" - grep -F "possible wedge" "$out" >/dev/null && fail "paused secondmate was mislabeled a wedge" - unset FM_FAKE_CREW_STATE - pass "a declared paused secondmate re-surfaces on the bounded normal-mode cadence" -} - -# A captain hold is the other declared wait, but unlike paused: it has no -# current-state mapping, so a held mate reports `unknown` rather than `paused`. -# The bounded re-surface must still reach it, or a mate's hold rots invisibly: -# nothing else re-reads a quiet mate's endpoint. -test_secondmate_captain_held_resurfaces_in_normal_mode() { - local dir state fakebin out capture_file statusf window key pane_hash sig pid back - dir=$(make_case secondmate-held-resurface); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; statusf="$state/secondmate-hold.status" - window="test:fm-secondmate-hold" - printf 'idle awaiting the captain\n' > "$capture_file" - printf 'window=%s\nkind=secondmate\n' "$window" > "$state/secondmate-hold.meta" - printf 'captain-held [key=route]: tracked by task-decision-route\n' > "$statusf" - back=$(( $(date +%s) - 500 )) - if [ "$(uname)" = Darwin ]; then touch -mt "$(date -r "$back" '+%Y%m%d%H%M.%S')" "$statusf" - else touch -m -d "@$back" "$statusf"; fi - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-secondmate-hold_status" - key=$(printf '%s' "$window" | tr '.:/' '___') - pane_hash=$(hash_text "idle awaiting the captain") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - export FM_FAKE_CREW_STATE='state: unknown · source: none · no current-state source available' - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not re-surface a captain-held secondmate" - grep -F "stale: $window" "$out" >/dev/null || fail "captain-held secondmate did not emit a stale recheck" - grep -F "awaiting the captain" "$out" >/dev/null || fail "captain-held secondmate recheck did not name the captain as the blocker: $(cat "$out")" - grep -F "awaiting external" "$out" >/dev/null && fail "captain-held secondmate recheck claimed an external wait" - grep -F "possible wedge" "$out" >/dev/null && fail "captain-held secondmate was mislabeled a wedge" - unset FM_FAKE_CREW_STATE - pass "a captain-held secondmate re-surfaces on the bounded normal-mode cadence" -} - -test_secondmate_nonpaused_stale_remains_suppressed() { - local dir state fakebin out capture_file statusf window key pane_hash sig pid - dir=$(make_case secondmate-stale-suppressed); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; statusf="$state/secondmate-working.status" - window="test:fm-secondmate-working" - printf 'idle while the parent supervises\n' > "$capture_file" - printf 'window=%s\nkind=secondmate\n' "$window" > "$state/secondmate-working.meta" - printf 'working: the parent supervises this secondmate\n' > "$statusf" - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-secondmate-working_status" - key=$(printf '%s' "$window" | tr '.:/' '___') - pane_hash=$(hash_text "idle while the parent supervises") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid"; fail "watcher surfaced an ordinary secondmate stale pane: $(cat "$out")" - fi - [ ! -s "$out" ] || { reap "$pid"; fail "ordinary secondmate stale pane printed a wake reason: $(cat "$out")"; } - reap "$pid" - pass "a non-paused secondmate retains normal stale suppression" -} - -test_secondmate_unpause_clears_pause_tracking() { - local dir state fakebin out statusf window key pid - dir=$(make_case secondmate-unpause-clears); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; statusf="$state/secondmate-resumed.status"; window="test:fm-secondmate-resumed" - printf 'window=%s\nkind=secondmate\n' "$window" > "$state/secondmate-resumed.meta" - printf 'working: upstream landed\n' > "$statusf" - printf '%s' "$(seen_sig "$statusf")" > "$state/.seen-secondmate-resumed_status" - key=${window//:/_} - key=${key//\//_} - key=${key//./_} - : > "$state/.paused-$key" - : > "$state/.paused-rechecked-$key" - : > "$state/.paused-resurfaced-$key" - : > "$state/.stale-$key" - : > "$state/.stale-since-$key" - : > "$state/.wedge-escalations-$key" - watch_bg "$state" "$fakebin" "$out" - pid=$! - wait_poll_cycle "$state" "$pid" || fail "watcher exited while reconciling a resumed secondmate: $(cat "$out")" - [ ! -e "$state/.paused-$key" ] || { reap "$pid"; fail "resumed secondmate retained the pause marker"; } - [ ! -e "$state/.stale-$key" ] || { reap "$pid"; fail "resumed secondmate retained stale tracking"; } - [ ! -e "$state/.wedge-escalations-$key" ] || { reap "$pid"; fail "resumed secondmate retained wedge tracking"; } - reap "$pid" - pass "a resumed secondmate clears pause and stale tracking before stale exemption" -} - -test_nonterminal_stale_pause_transitions_reclassify_unchanged_hash() { - local dir state fakebin out capture_file window key pane_hash sig pid i - dir=$(make_case nonterminal-stale-pause-transition); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-transition" - printf 'idle awaiting external\n' > "$capture_file" - printf 'window=%s\nkind=ship\n' "$window" > "$state/transition.meta" - printf 'paused: awaiting the upstream release\n' > "$state/transition.status" - sig=$(seen_sig "$state/transition.status"); printf '%s' "$sig" > "$state/.seen-transition_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "idle awaiting external") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '%s' "$pane_hash" > "$state/.stale-$key" - printf '1\n' > "$state/.count-$key" - printf '%s\n' $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" - export FM_FAKE_CREW_STATE='state: paused · source: status-log · awaiting the upstream release' - - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_TMUX_CURRENT_COMMAND=zsh \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - i=0 - while [ "$i" -lt 100 ] && kill -0 "$pid" 2>/dev/null; do - [ -e "$state/.paused-$key" ] && [ ! -e "$state/.stale-since-$key" ] && break - sleep 0.1 - i=$((i + 1)) - done - kill -0 "$pid" 2>/dev/null || { reap "$pid"; fail "a stale hash that entered pause was wedge-escalated: $(cat "$out")"; } - [ -e "$state/.paused-$key" ] || { reap "$pid"; fail "unchanged stale hash did not enter paused mode"; } - [ ! -e "$state/.stale-since-$key" ] || { reap "$pid"; fail "pause transition retained its wedge timer"; } - wait_poll_cycle "$state" "$pid" || { reap "$pid"; fail "a stale hash that entered pause was wedge-escalated: $(cat "$out")"; } - reap "$pid" - ack_stopped_cycle "$state" || fail "could not acknowledge the intentional entered-pause watcher stop" - - printf 'working: upstream landed, resuming\n' > "$state/transition.status" - sig=$(seen_sig "$state/transition.status"); printf '%s' "$sig" > "$state/.seen-transition_status" - FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - i=0 - while [ "$i" -lt 100 ] && kill -0 "$pid" 2>/dev/null; do - [ ! -e "$state/.paused-$key" ] && [ -s "$state/.stale-since-$key" ] && break - sleep 0.1 - i=$((i + 1)) - done - kill -0 "$pid" 2>/dev/null || { reap "$pid"; fail "a stale hash that left pause did not resume wedge tracking: $(cat "$out")"; } - [ ! -e "$state/.paused-$key" ] || { reap "$pid"; fail "unchanged stale hash retained paused mode after resume"; } - [ -s "$state/.stale-since-$key" ] || { reap "$pid"; fail "unchanged stale hash did not restart wedge tracking after resume"; } - wait_poll_cycle "$state" "$pid" || { reap "$pid"; fail "a stale hash that left pause did not resume wedge tracking: $(cat "$out")"; } - reap "$pid" - unset FM_FAKE_CREW_STATE - pass "unchanged stale hashes reclassify when a crew enters or leaves pause" -} - -test_nonterminal_paused_rechecks_authoritative_state() { - local dir state fakebin out capture_file window key pane_hash sig pid - dir=$(make_case nonterminal-paused-recheck); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-pause-recheck" - printf 'idle awaiting external\n' > "$capture_file" - printf 'window=%s\nkind=ship\n' "$window" > "$state/pause-recheck.meta" - printf 'paused: awaiting the upstream release\n' > "$state/pause-recheck.status" - sig=$(seen_sig "$state/pause-recheck.status"); printf '%s' "$sig" > "$state/.seen-pause-recheck_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "idle awaiting external") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '%s' "$pane_hash" > "$state/.stale-$key" - printf '1\n' > "$state/.count-$key" - : > "$state/.paused-$key" - export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' - - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid"; fail "an active run behind a declared pause surfaced instead of resuming wedge tracking: $(cat "$out")" - fi - # Authoritative state moves TRACKING to the wedge timer, which is what this - # case is about. The declaration itself is deliberately kept on record (issue - # 67): it is the fallback the escalation uses when the run yields no progress - # evidence, and it carries that cadence's re-surface throttle and backoff, so - # discarding it here would restart both on every poll. - [ -s "$state/.stale-since-$key" ] || { reap "$pid"; fail "authoritative active run did not resume wedge tracking"; } - [ -e "$state/.paused-$key" ] || { reap "$pid"; fail "authoritative active run discarded the declared wait it may still need"; } - reap "$pid" - unset FM_FAKE_CREW_STATE - pass "a declared pause is periodically rechecked against authoritative active-run state" -} - -test_paused_authoritative_working_preserves_wedge_timer() { - local dir state fakebin out capture_file window key pane_hash sig pid since - dir=$(make_case paused-working-preserves-wedge-timer); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-paused-working" - printf 'idle awaiting external\n' > "$capture_file" - printf 'window=%s\nkind=ship\n' "$window" > "$state/paused-working.meta" - printf 'paused: awaiting the upstream release\n' > "$state/paused-working.status" - sig=$(seen_sig "$state/paused-working.status"); printf '%s' "$sig" > "$state/.seen-paused-working_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "idle awaiting external") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '%s' "$pane_hash" > "$state/.stale-$key" - printf '1\n' > "$state/.count-$key" - : > "$state/.paused-$key" - export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' - - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_numeric_file "$state/.stale-since-$key" 30 || { reap "$pid"; fail "authoritative working state did not start wedge tracking"; } - since=$(cat "$state/.stale-since-$key") - sleep 2 - [ "$(cat "$state/.stale-since-$key" 2>/dev/null || true)" = "$since" ] \ - || { reap "$pid"; fail "repeat authoritative working recheck reset the wedge timer"; } - reap "$pid" - ack_stopped_cycle "$state" || fail "could not acknowledge the intentional authoritative-working stop" - - echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" - : > "$out" - # The verdict is the run's, not the pane's: with the declared wait still on - # record, only positive evidence that the run stopped or the agent died may - # raise the alarm here (issue 67). The no-evidence half is pinned by - # test_declared_wait_with_no_progress_evidence_rechecks_instead_of_wedging. - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_FAKE_RUN_PROGRESS='progress: stranded · test running, last activity 31m0s ago' \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "authoritative working state did not wedge-escalate past the threshold" - grep -F "possible wedge" "$out" >/dev/null || fail "authoritative working wedge escalation omitted its reason" - [ ! -e "$state/.stale-since-$key" ] || fail "wedge timer remained after authoritative working escalation" - unset FM_FAKE_CREW_STATE - pass "a paused status overridden by authoritative working preserves its wedge timer and escalates" -} - -# --- the wedge escalation consults the run's PROGRESS, not just its existence -- -# -# A worker that backgrounds a validation call and goes quiet was escalated as a -# possible wedge every threshold, five times in a row on one pane, while -# fm-crew-state reported "working · run-step · validating (running)" the whole -# time. `status: running` alone cannot separate that from a run that has -# stranded, so the escalation point reads bin/fm-run-progress.sh, whose classes -# these four cases pin from the escalation side without a declared wait: -# -# progressing -> held -# stranded -> still escalates, naming the step that stopped -# none -> escalates byte-identically to before this gate existed -# dead agent -> escalates however well its run is moving -# -# The reader's own parsing and threshold live in fm-run-progress.test.sh. - -# Fixture: a crew parked on a validation run, its wedge timer already backdated -# past the threshold, so the very next poll reaches the escalation decision. -# Its status line deliberately declares NO wait, so these four cases isolate the -# run-progress axis alone; the declared-wait axis has its own four cases below, -# over prime_declared_wait_at_threshold. Keeping both axes in one fixture is what -# made issue 67 invisible here - a declared wait rides a different branch of the -# stale path, and a fixture that carries one silently tests that branch instead. -prime_wedge_at_threshold() { # <state> <task> <window> <capture-file> - local state=$1 task=$2 window=$3 capture=$4 key - printf 'window=%s\nkind=ship\n' "$window" > "$state/$task.meta" - printf 'working: no-mistakes run under way, parked on the pipeline call\n' > "$state/$task.status" - prime_status_seen "$state" "$state/$task.status" - prime_stale_pane "$state" "$window" 'validating · esc to interrupt' "$capture" - key=$(printf '%s' "$window" | tr ':/.' '___') - printf '%s' "$(hash_text 'validating · esc to interrupt')" > "$state/.stale-$key" - echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" -} - -# Drive one watcher over that fixture with a fixed run-progress verdict. Extra -# `NAME=value` arguments are applied on top, through `env` rather than an -# assignment prefix (one arriving through "$@" is expanded too late to be -# recognized as an assignment). -run_wedge_watcher() { # <state> <fakebin> <window> <capture> <out> <progress-verdict> [env assignments...] - local state=$1 fakebin=$2 window=$3 capture=$4 out=$5 verdict=$6 - shift 6 - env "PATH=$fakebin:$PATH" "FM_FAKE_TMUX_WINDOW=$window" "FM_FAKE_TMUX_CAPTURE=$capture" \ - "FM_STATE_OVERRIDE=$state" "FM_CREW_STATE_BIN=$fakebin/fm-crew-state.sh" \ - 'FM_FAKE_CREW_STATE=state: working · source: run-step · validating (running)' \ - "FM_FAKE_RUN_PROGRESS=$verdict" \ - FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$@" "$WATCH" > "$out" & -} - -test_progressing_run_holds_the_wedge_escalation() { - local dir state fakebin out capture window key pid since - dir=$(make_case wedge-run-progressing); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture="$dir/pane.txt"; window="test:fm-validating" - key=$(printf '%s' "$window" | tr ':/.' '___') - export FM_FAKE_TMUX_CURRENT_COMMAND=claude - prime_wedge_at_threshold "$state" validating "$window" "$capture" - since=$(cat "$state/.stale-since-$key") - - run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" \ - 'progress: progressing · test running, last activity 7m4s ago (silent 424s, bound 1800s)' - pid=$! - if ! wait_live "$pid" 30; then - reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND - fail "a crew parked on a progressing validation run still wedge-escalated: $(cat "$out")" - fi - [ ! -s "$out" ] || { reap "$pid"; fail "the held escalation still printed a wake: $(cat "$out")"; } - # Held, not cleared: the timer restarts so the next look is a full window away - # rather than one poll away, which is what keeps the bounded run-progress read - # to once per window per pane. - [ -s "$state/.stale-since-$key" ] || { reap "$pid"; fail "the held escalation cleared the wedge timer"; } - [ "$(cat "$state/.stale-since-$key")" != "$since" ] \ - || { reap "$pid"; fail "the held escalation did not restart the wedge timer"; } - [ ! -e "$state/.wedge-escalations-$key" ] \ - || { reap "$pid"; fail "the held escalation still counted as an escalation"; } - reap "$pid" - FM_STATE_OVERRIDE="$state" "$DRAIN" 2>/dev/null | grep -F "$window" >/dev/null \ - && fail "the held escalation was queued" - unset FM_FAKE_TMUX_CURRENT_COMMAND - pass "a crew parked on a demonstrably progressing validation run holds its wedge escalation" -} - -test_stranded_run_still_wedge_escalates() { - local dir state fakebin out capture window key pid - dir=$(make_case wedge-run-stranded); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture="$dir/pane.txt"; window="test:fm-stranded" - key=$(printf '%s' "$window" | tr ':/.' '___') - export FM_FAKE_TMUX_CURRENT_COMMAND=claude - prime_wedge_at_threshold "$state" stranded "$window" "$capture" - - run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" \ - 'progress: stranded · test running, last activity 31m0s ago (silent 1860s, past the 1800s bound)' - pid=$! - wait_for_exit "$pid" 100 || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "a stranded validation run did not wedge-escalate"; } - grep -F "possible wedge" "$out" >/dev/null || fail "the stranded run's escalation dropped its wedge reason" - grep -F "validation run stranded: test running, last activity 31m0s ago" "$out" >/dev/null \ - || fail "the stranded run's escalation did not name the step that stopped: $(cat "$out")" - [ "$(cat "$state/.wedge-escalations-$key" 2>/dev/null || echo 0)" = 1 ] \ - || fail "the stranded run's escalation was not counted" - unset FM_FAKE_TMUX_CURRENT_COMMAND - pass "a crew whose validation run has stranded still wedge-escalates, naming the step" -} - -test_wedged_crew_with_no_run_escalates_unchanged() { - local dir state fakebin out drain_out capture window key pid - dir=$(make_case wedge-run-absent); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture="$dir/pane.txt" - window="test:fm-norun" - key=$(printf '%s' "$window" | tr ':/.' '___') - export FM_FAKE_TMUX_CURRENT_COMMAND=claude - prime_wedge_at_threshold "$state" norun "$window" "$capture" - - # The ORIGINAL purpose of this alarm: a quiet crew with no active run at all. - # `none` is what every no-evidence shape collapses to, so this pins that the - # gate cannot weaken it. - run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" \ - 'progress: none · no run attributed to this crew' - pid=$! - wait_for_exit "$pid" 100 || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "a wedged crew with no active run did not escalate"; } - grep -F "possible wedge" "$out" >/dev/null || fail "the no-run wedge escalation lost its reason" - grep -F "validation run stranded" "$out" >/dev/null \ - && fail "a crew with no run was described as having a stranded run" - [ ! -e "$state/.stale-since-$key" ] || fail "the no-run escalation left its wedge timer standing" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the no-run wedge failed" - grep "$(printf '\tstale\t')" "$drain_out" | grep -F "possible wedge" >/dev/null \ - || fail "the no-run wedge escalation was not queued" - unset FM_FAKE_TMUX_CURRENT_COMMAND - pass "a crew wedged with no active run escalates exactly as it does today" -} - -test_dead_agent_escalates_even_while_its_run_progresses() { - local dir state fakebin out capture window pid - dir=$(make_case wedge-run-dead-agent); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture="$dir/pane.txt"; window="test:fm-deadagent" - # A bare shell at the endpoint is the confident dead verdict. - export FM_FAKE_TMUX_CURRENT_COMMAND=zsh - prime_wedge_at_threshold "$state" deadagent "$window" "$capture" - - # The pipeline runs its own steps, so a run keeps advancing with nobody left - # to answer its next gate. That is a wedge, and it is exactly the shape "the - # run is fine" would otherwise hide. - run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" \ - 'progress: progressing · test running, last activity 10s ago (silent 10s, bound 1800s)' - pid=$! - wait_for_exit "$pid" 100 || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "a dead agent was absorbed because its run was progressing"; } - grep -F "possible wedge" "$out" >/dev/null || fail "the dead-agent escalation lost its wedge reason" - unset FM_FAKE_TMUX_CURRENT_COMMAND - pass "a confidently dead agent still escalates however well its validation run is moving" -} - -# The principle this pins, which is the whole reason the hold is capped: RUN -# PROGRESS IS EVIDENCE ABOUT THE RUN, NOT ABOUT THE WORKER. They are different -# subjects. A moving pipeline licenses a DELAY in alarming and never permanent -# silence, because the failure permanent silence would hide is a worker whose -# harness hung mid-turn while its pipeline kept executing its own steps quite -# happily - the endpoint reads alive, so the dead-agent short-circuit never -# fires, the run reports `progressing` on every look, and the pane would be held -# for the whole remaining run. That is silent, indefinite, and worse than the -# noise the hold exists to cut. So past FM_RUN_PROGRESS_HOLD_MAX the pane -# escalates REGARDLESS of how healthy its run looks. -test_progressing_run_escalates_anyway_past_the_hold_cap() { - local dir state fakebin out capture window key pid verdict - dir=$(make_case wedge-run-hold-cap); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture="$dir/pane.txt"; window="test:fm-holdcap" - key=$(printf '%s' "$window" | tr ':/.' '___') - export FM_FAKE_TMUX_CURRENT_COMMAND=claude - prime_wedge_at_threshold "$state" holdcap "$window" "$capture" - verdict='progress: progressing · test running, last activity 7m4s ago (silent 424s, bound 1800s)' - - # Phase A: below the cap, unchanged - held, and the hold is counted. - run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" "$verdict" FM_RUN_PROGRESS_HOLD_MAX=2 - pid=$! - if ! wait_live "$pid" 30; then - reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND - fail "a hold below the cap escalated: $(cat "$out")" - fi - [ "$(cat "$state/.wedge-holds-$key" 2>/dev/null || echo 0)" = 1 ] \ - || { reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the hold was not counted"; } - reap "$pid" - ack_stopped_cycle "$state" || fail "could not acknowledge the intentional phase-A hold stop" - - # Phase B: at the cap, with the run reporting the very same healthy verdict. - echo 2 > "$state/.wedge-holds-$key" - echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" - : > "$out" - run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" "$verdict" FM_RUN_PROGRESS_HOLD_MAX=2 - pid=$! - wait_for_exit "$pid" 100 || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "a pane held to the cap never escalated: $(cat "$out")"; } - grep -F "possible wedge" "$out" >/dev/null \ - || fail "the forced escalation dropped the possible-wedge marker: $(cat "$out")" - # It must stay INFORMATIVE rather than reading like a dead pane: the - # supervisor has to see "the run is still moving, this pane is not" straight - # off the wake. - grep -F "still progressing" "$out" >/dev/null \ - || fail "the forced escalation did not say the run was still moving: $(cat "$out")" - grep -F "test running, last activity 7m4s ago" "$out" >/dev/null \ - || fail "the forced escalation dropped the progress detail: $(cat "$out")" - [ "$(cat "$state/.wedge-escalations-$key" 2>/dev/null || echo 0)" = 1 ] \ - || fail "the forced escalation did not count toward demand-deep-inspection" - [ ! -e "$state/.wedge-holds-$key" ] || fail "the forced escalation did not reset the hold count" - ack_stopped_cycle "$state" || fail "could not acknowledge the forced escalation" - - # Phase C: and the count reset makes the cap a repeating check-in cadence, not - # a one-shot that then goes quiet forever - the next window holds again. - echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" - : > "$out" - run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" "$verdict" FM_RUN_PROGRESS_HOLD_MAX=2 - pid=$! - if ! wait_live "$pid" 30; then - reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND - fail "the window after a forced escalation did not hold again: $(cat "$out")" - fi - [ "$(cat "$state/.wedge-holds-$key" 2>/dev/null || echo 0)" = 1 ] \ - || { reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the hold count did not restart after the forced escalation"; } - reap "$pid" - unset FM_FAKE_TMUX_CURRENT_COMMAND - pass "consecutive run-progress holds are capped: past the cap the pane escalates anyway naming the still-moving run, and the count resets so the cadence repeats" -} - -# The busy-turn bound routes through the same wedge_timer_check, so the hold and -# its cap apply there DELIBERATELY, not incidentally. BUSY_TURN_MAX_SECS exists -# to bound a hung FOREGROUND call that a rendered busy footer would otherwise -# hide, and a crew driving `no-mistakes axi run` in the foreground IS such a -# call: busy for the whole pipeline with no completed turn. Whether such a pane -# reads busy or stale is only an artifact of whether its harness backgrounded -# the pipeline call, so holding for one and not the other would be arbitrary. -# The cap matters MORE here, because a busy pane has already waited a full -# BUSY_TURN_MAX_SECS before its first escalation. -test_busy_pane_progressing_run_holds_then_escalates_past_the_cap() { - local dir state fakebin out capture_file window key pane_hash sig pid verdict - dir=$(make_case busy-run-progress-hold); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-busy-validating" - printf 'Working...' > "$capture_file" - printf 'window=%s\nkind=ship\nharness=pi\n' "$window" > "$state/busy-validating.meta" - record_pi_busy "$state" busy-validating - printf 'working: no-mistakes run under way in the foreground\n' > "$state/busy-validating.status" - sig=$(seen_sig "$state/busy-validating.status"); printf '%s' "$sig" > "$state/.seen-busy-validating_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "Working...") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - # No completed turn ever recorded: age the spawn record past the busy bound. - touch -t 200001010000 "$state/busy-validating.meta" - export FM_FAKE_TMUX_CURRENT_COMMAND=claude - verdict='progress: progressing · test running, last activity 7m4s ago (silent 424s, bound 1800s)' - echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" - - # Phase A: past the busy bound AND past the wedge threshold, but the run is - # moving - held, exactly as the stale path holds. - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_RUN_PROGRESS_HOLD_MAX=2 FM_FAKE_RUN_PROGRESS="$verdict" \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - if ! wait_live "$pid" 30; then - reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND - fail "a busy pane on a progressing validation run escalated instead of holding: $(cat "$out")" - fi - [ "$(cat "$state/.wedge-holds-$key" 2>/dev/null || echo 0)" = 1 ] \ - || { reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the busy-path hold was not counted"; } - reap "$pid" - ack_stopped_cycle "$state" || fail "could not acknowledge the intentional busy-path phase-A stop" - - # Phase B: at the cap it escalates anyway, carrying the progress detail. - echo 2 > "$state/.wedge-holds-$key" - echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_RUN_PROGRESS_HOLD_MAX=2 FM_FAKE_RUN_PROGRESS="$verdict" \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "a busy pane held to the cap never escalated: $(cat "$out")"; } - grep -F "possible wedge" "$out" >/dev/null \ - || fail "the busy-path forced escalation dropped the possible-wedge marker: $(cat "$out")" - grep -F "still progressing" "$out" >/dev/null \ - || fail "the busy-path forced escalation did not say the run was still moving: $(cat "$out")" - grep -F "test running, last activity 7m4s ago" "$out" >/dev/null \ - || fail "the busy-path forced escalation dropped the progress detail: $(cat "$out")" - [ ! -e "$state/.wedge-holds-$key" ] || fail "the busy-path forced escalation did not reset the hold count" - unset FM_FAKE_TMUX_CURRENT_COMMAND - pass "a busy pane past its turn-age bound is held while its run is moving and escalates anyway past the hold cap" -} - -# --- a DECLARED wait, overridden by an active run, still counts at the alarm --- -# -# Issue 67, reproduced 2026-08-07: a crew that followed its brief and appended -# `paused:` before parking on a backgrounded pipeline call was wedge-escalated -# twice inside five minutes while fm-crew-state reported an actively running -# validation. The declared wait was consumed the moment authoritative state -# outranked it (correctly - a crew that declared a pause and then started a run -# IS working), and from there the pane was on the plain 240s wedge cadence with -# nothing left of the worker's own statement about its silence. -# -# The four cases below are the whole policy, and each is deliberately the -# opposite of one of the others, so no single change can satisfy them all by -# widening or narrowing absorption: -# -# progressing -> held (positive evidence the run is moving) -# stranded -> escalates (positive evidence the run stopped) -# dead agent -> escalates (nobody left to answer the next gate) -# none + live agent -> declared-wait recheck, NOT a wedge alarm -# -# Only the last one changes behavior. `none` is no evidence either way - no run -# attributed, a status read that could not complete, a run between steps - and -# an alarm needs a reason to fire rather than the absence of one, once the crew -# itself has said the silence is deliberate. bin/fm-supervise-daemon.sh's own -# stale recheck has always short-circuited a declared pause before the wedge -# escalation, so this is also what stops the two supervisors disagreeing about -# the same pane. - -# Fixture: an IDLE (non-busy) pane whose crew declared a wait and whose -# authoritative state is an active run, its wedge timer already past the -# threshold and its stale hash already classified, so the very next poll reaches -# the escalation decision through the declared-wait branch. Deliberately not -# prime_wedge_at_threshold: that fixture's pane text carries a busy signature and -# routes through the busy-turn-age path instead, which is why the stale path's -# declared-wait branch had no coverage of its own. -prime_declared_wait_at_threshold() { # <state> <task> <window> <capture-file> [<status-line>] - local state=$1 task=$2 window=$3 capture=$4 key - local status_line=${5:-'paused: no-mistakes run under way, parked on the pipeline call'} - printf 'window=%s\nkind=ship\n' "$window" > "$state/$task.meta" - printf '%s\n' "$status_line" > "$state/$task.status" - prime_status_seen "$state" "$state/$task.status" - prime_stale_pane "$state" "$window" 'idle, parked on the pipeline call' "$capture" - key=$(printf '%s' "$window" | tr ':/.' '___') - printf '%s' "$(hash_text 'idle, parked on the pipeline call')" > "$state/.stale-$key" - echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" -} - -test_declared_wait_with_no_progress_evidence_rechecks_instead_of_wedging() { - local dir state fakebin out drain_out capture window key pid - dir=$(make_case declared-wait-no-evidence); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture="$dir/pane.txt" - window="test:fm-declared-none" - key=$(printf '%s' "$window" | tr ':/.' '___') - export FM_FAKE_TMUX_CURRENT_COMMAND=claude - prime_declared_wait_at_threshold "$state" declared-none "$window" "$capture" - - # Phase A: the wait was declared moments ago, so its own re-surface window has - # not elapsed - the pane is absorbed outright, where before it alarmed. - run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" \ - 'progress: none · no run attributed to this crew' - pid=$! - if ! wait_live "$pid" 30; then - reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND - fail "a declared wait with no progress evidence still wedge-escalated: $(cat "$out")" - fi - [ ! -s "$out" ] || { reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the deferred pane still printed a wake: $(cat "$out")"; } - [ ! -e "$state/.wedge-escalations-$key" ] \ - || { reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the deferral was counted as a wedge escalation"; } - [ -e "$state/.paused-$key" ] \ - || { reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the pane was not handed back to the declared-wait cadence"; } - grep -F "deferred non-terminal stale (provably working after a declared wait) wedge escalation" \ - "$state/.watch-triage.log" >/dev/null \ - || { reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the deferral was not distinguishable in the triage log: $(cat "$state/.watch-triage.log")"; } - reap "$pid" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the deferral failed" - [ -s "$drain_out" ] \ - && { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the deferred pane queued a wake: $(cat "$drain_out")"; } - - # Phase B: absorbed is not silenced. Age the wait past its own re-surface - # window and it comes back as a recheck the supervisor can act on - never as a - # possible wedge - so a wait that stops being true cannot rot invisibly. - set_mtime $(( $(date +%s) - 500 )) "$state/declared-none.status" - prime_status_seen "$state" "$state/declared-none.status" - echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" - : > "$out" - run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" \ - 'progress: none · no run attributed to this crew' FM_PAUSE_RESURFACE_SECS=240 - pid=$! - wait_for_exit "$pid" 100 \ - || { reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "an aged declared wait never came back for a recheck: $(cat "$out")"; } - grep -F "awaiting external" "$out" >/dev/null \ - || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the re-surfaced wake was not a declared-wait recheck: $(cat "$out")"; } - grep -F "possible wedge" "$out" >/dev/null \ - && { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the recheck was raised as a possible wedge: $(cat "$out")"; } - unset FM_FAKE_TMUX_CURRENT_COMMAND - pass "a declared wait whose run yields no progress evidence is rechecked on its own cadence, never wedge-escalated" -} - -test_declared_wait_on_a_progressing_run_holds() { - local dir state fakebin out capture window key pid since - dir=$(make_case declared-wait-progressing); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture="$dir/pane.txt"; window="test:fm-declared-moving" - key=$(printf '%s' "$window" | tr ':/.' '___') - export FM_FAKE_TMUX_CURRENT_COMMAND=claude - prime_declared_wait_at_threshold "$state" declared-moving "$window" "$capture" - since=$(cat "$state/.stale-since-$key") - - # The reproduction's own numbers: `document` running, last activity ~3m ago, - # comfortably inside the stranded bound. - run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" \ - 'progress: progressing · document running, last activity 3m16s ago (silent 196s, bound 1800s)' - pid=$! - if ! wait_live "$pid" 30; then - reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND - fail "a declared wait inside an actively progressing run wedge-escalated: $(cat "$out")" - fi - [ ! -s "$out" ] || { reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the held declared wait still printed a wake: $(cat "$out")"; } - [ "$(cat "$state/.wedge-holds-$key" 2>/dev/null || echo 0)" = 1 ] \ - || { reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the declared wait's hold was not counted"; } - [ "$(cat "$state/.stale-since-$key")" != "$since" ] \ - || { reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the held escalation did not restart the wedge timer"; } - reap "$pid" - unset FM_FAKE_TMUX_CURRENT_COMMAND - pass "a declared wait inside an actively progressing validation run holds its wedge escalation" -} - -test_declared_wait_on_a_stranded_run_still_escalates() { - local dir state fakebin out capture window key pid - dir=$(make_case declared-wait-stranded); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture="$dir/pane.txt"; window="test:fm-declared-stranded" - key=$(printf '%s' "$window" | tr ':/.' '___') - export FM_FAKE_TMUX_CURRENT_COMMAND=claude - prime_declared_wait_at_threshold "$state" declared-stranded "$window" "$capture" - - # Positive evidence the run stopped outranks the crew's own statement that its - # silence is deliberate: a stranded step is exactly the case the declared wait - # must never hide. - run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" \ - 'progress: stranded · test running, last activity 31m0s ago (silent 1860s, past the 1800s bound)' - pid=$! - wait_for_exit "$pid" 100 \ - || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "a declared wait on a stranded run did not escalate: $(cat "$out")"; } - grep -F "possible wedge" "$out" >/dev/null \ - || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the stranded declared wait lost its wedge reason: $(cat "$out")"; } - grep -F "validation run stranded: test running, last activity 31m0s ago" "$out" >/dev/null \ - || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the stranded escalation did not name the step that stopped: $(cat "$out")"; } - [ "$(cat "$state/.wedge-escalations-$key" 2>/dev/null || echo 0)" = 1 ] \ - || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the stranded declared wait's escalation was not counted"; } - unset FM_FAKE_TMUX_CURRENT_COMMAND - pass "a declared wait whose validation run has stranded still wedge-escalates, naming the step" -} - -test_declared_wait_with_a_dead_agent_still_escalates() { - local dir state fakebin out capture window pid - dir=$(make_case declared-wait-dead-agent); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture="$dir/pane.txt"; window="test:fm-declared-dead" - # A bare shell at the endpoint is the confident dead verdict. - export FM_FAKE_TMUX_CURRENT_COMMAND=zsh - prime_declared_wait_at_threshold "$state" declared-dead "$window" "$capture" - - # No progress evidence AND nobody left to answer the run's next gate. The - # declared wait is a statement about a worker that is no longer there, so it - # must not buy the pane the recheck cadence. - run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" \ - 'progress: none · status read did not complete' - pid=$! - wait_for_exit "$pid" 100 \ - || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "a declared wait whose agent had died did not escalate: $(cat "$out")"; } - grep -F "possible wedge" "$out" >/dev/null \ - || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the dead-agent declared wait lost its wedge reason: $(cat "$out")"; } - unset FM_FAKE_TMUX_CURRENT_COMMAND - pass "a declared wait whose agent has confidently exited still wedge-escalates" -} - -# --- the triage log must distinguish the two provably-working absorptions ------ -# What hid issue 67 for two days: a first sighting that absorbed and STARTED the -# wedge timer, and a repeat poll that absorbed and ADVANCED an already-running -# timer toward an escalation, wrote the identical line. The log therefore showed -# a steady stream of absorptions while escalations kept arriving, agreeing with -# the intended behavior rather than the actual behavior. -test_provably_working_absorptions_are_distinguishable_in_the_triage_log() { - local dir state fakebin out capture window key pane_hash pid log - dir=$(make_case absorb-log-distinct); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture="$dir/pane.txt"; window="test:fm-logdistinct" - key=$(printf '%s' "$window" | tr ':/.' '___') - log="$state/.watch-triage.log" - export FM_FAKE_TMUX_CURRENT_COMMAND=claude - prime_declared_wait_at_threshold "$state" logdistinct "$window" "$capture" - pane_hash=$(hash_text 'idle, parked on the pipeline call') - # Phase A: an UNCLASSIFIED hash, so this poll is the first sighting - it - # absorbs and starts the timer. - rm -f "$state/.stale-$key" "$state/.stale-since-$key" - - run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" \ - 'progress: progressing · document running, last activity 3m16s ago' - pid=$! - if ! wait_live "$pid" 30; then - reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the first sighting was not absorbed: $(cat "$out")" - fi - reap "$pid" - grep -F "absorbed non-terminal stale (provably working, wedge timer started)" "$log" >/dev/null \ - || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the first sighting did not record that it STARTED the timer: $(cat "$log")"; } - # The first sighting is the FIRST absorption in the log; the polls behind it - # are already repeat polls of the same hash and rightly say so. - grep -F "absorbed non-terminal stale (provably working" "$log" | head -1 \ - | grep -F "wedge timer started" >/dev/null \ - || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the first absorption was not the timer-started event: $(cat "$log")"; } - ack_stopped_cycle "$state" || fail "could not acknowledge the intentional first-sighting stop" - - # Phase B: the same hash on a later poll, with the timer short of the - # threshold so nothing else can write a line - it absorbs and ADVANCES. - : > "$log"; : > "$out" - printf '%s' "$pane_hash" > "$state/.stale-$key" - date +%s > "$state/.stale-since-$key" - run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" \ - 'progress: progressing · document running, last activity 3m16s ago' - pid=$! - if ! wait_live "$pid" 30; then - reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the repeat poll was not absorbed: $(cat "$out")" - fi - reap "$pid" - grep -F "wedge timer advanced" "$log" >/dev/null \ - || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the repeat poll did not record that it ADVANCED the timer: $(cat "$log")"; } - grep -F "wedge timer started" "$log" >/dev/null \ - && { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the repeat poll re-used the first-sighting line"; } - unset FM_FAKE_TMUX_CURRENT_COMMAND - pass "starting the wedge timer and advancing it are distinguishable events in the triage log" -} - -# --- consecutive wedge escalations on the same pane demand deep inspection ---- -# Root cause of the PR #252 incident's ~20 minutes of unnoticed green: each -# wedge escalation fires, gets classified as "still validating" one poll later -# (the timer restarts, see wedge_timer_check), and repeats forever on a pane -# that never changes. A single escalation reason looks identical every round, -# so nothing in the payload itself signals "this has now happened N times in a -# row" - that judgment call was left entirely to the supervisor noticing the -# repetition on its own. This is the safety-net fix: past -# FM_WEDGE_DEMAND_INSPECT_COUNT consecutive escalations on the SAME pane, the -# wake reason itself carries a "demand-deep-inspection" marker. - -test_wedge_escalation_marks_demand_deep_inspection_after_threshold() { - local dir state fakebin out capture_file window key pane_hash sig pid n - dir=$(make_case wedge-escalation); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt" - window="test:fm-wedged" - printf 'idle building output' > "$capture_file" - printf 'window=%s\nkind=ship\n' "$window" > "$state/wedged.meta" - printf 'working: still monitoring ci\n' > "$state/wedged.status" - sig=$(seen_sig "$state/wedged.status"); printf '%s' "$sig" > "$state/.seen-wedged_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "idle building output") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - # The crew's pipeline is actively running: a static pane is normal (waiting on CI). - export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' - - # Priming round: first sighting of this stale hash classifies and absorbs it - # (establishing .stale-$key and starting the wedge timer) without going - # through wedge_timer_check at all - mirrors the existing wedge tests' Phase A. - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid"; fail "watcher exited on the priming round (should absorb): $(cat "$out")" - fi - reap "$pid" - ack_stopped_cycle "$state" || fail "could not acknowledge the intentional wedge priming stop" - - n=1 - while [ "$n" -le 3 ]; do - # Backdate the wedge timer past the threshold before each round, mirroring - # the existing wedge-escalation tests' Phase B (the subsequent-sight timer - # path does not re-read the crew state). - echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "watcher did not escalate on consecutive wedge round $n: $(cat "$out")" - grep -F "escalation $n" "$out" >/dev/null || fail "round $n did not report escalation count $n: $(cat "$out")" - if [ "$n" -lt 3 ]; then - grep -F "demand-deep-inspection" "$out" >/dev/null && fail "round $n escalated to demand-deep-inspection before the threshold: $(cat "$out")" - else - grep -F "demand-deep-inspection" "$out" >/dev/null || fail "round $n (threshold) did not demand deep inspection: $(cat "$out")" - fi - ack_stopped_cycle "$state" || fail "could not acknowledge wedge escalation round $n" - n=$((n + 1)) - done - [ "$(cat "$state/.wedge-escalations-$key" 2>/dev/null || echo 0)" = 3 ] || fail "escalation counter did not persist across consecutive rounds" - unset FM_FAKE_CREW_STATE - pass "consecutive wedge escalations on the same pane accumulate and demand deep inspection at the threshold" -} - -test_wedge_escalation_resets_when_pane_becomes_active() { - local dir state fakebin out capture_file window key pane_hash sig pid - dir=$(make_case wedge-escalation-reset); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt" - window="test:fm-wedged-reset" - printf 'idle building output' > "$capture_file" - printf 'window=%s\nkind=ship\n' "$window" > "$state/wedged-reset.meta" - printf 'working: still monitoring ci\n' > "$state/wedged-reset.status" - sig=$(seen_sig "$state/wedged-reset.status"); printf '%s' "$sig" > "$state/.seen-wedged-reset_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "idle building output") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - # Pre-seed one escalation as if a prior wedge round already fired. - printf '1\n' > "$state/.wedge-escalations-$key" - export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' - - # The pane content changes (the crew is active again): the hash no longer - # matches, so the watcher resets escalation bookkeeping instead of escalating. - printf 'new output, crew active again' > "$capture_file" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid"; fail "watcher exited on a fresh (changed) pane hash: $(cat "$out")" - fi - [ ! -e "$state/.wedge-escalations-$key" ] || fail "a changed pane hash did not reset the wedge-escalation counter" - reap "$pid" - unset FM_FAKE_CREW_STATE - pass "a pane becoming active again resets the consecutive wedge-escalation counter" -} - -# --- busy pane duration bound: a completed-turn age gate on top of busy ----- -# 2026-07 hibit-agent-focus-nonsteal-r1 incident: a busy pane (herdr "working" -# and/or the harness's rendered busy footer) is unconditional, unbounded proof -# of liveness in every existing classifier, so a genuinely hung foreground tool -# call behind a busy signature ran undetected for 25h. BUSY_TURN_MAX_SECS bounds -# how long a busy pane may run with no completed turn (state/<id>.turn-ended, or -# the task's spawn record before any turn completes); past the bound, panes -# without a declared external wait or verified captain-held transfer take the -# SAME wedge_timer_check already used for a provably-working non-busy stale. -# Escalation reuses the identical stale reason, escalation counter, and -# demand-deep-inspection marker - never an -# automatic interrupt or restart. - -test_busy_pane_below_turn_age_bound_is_absorbed() { - local dir state fakebin out capture_file window key sig pid - dir=$(make_case busy-below-turn-age); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-busy-fresh" - printf 'Working... (12.3s)' > "$capture_file" - printf 'window=%s\nkind=ship\nharness=pi\n' "$window" > "$state/busy-fresh.meta" - record_pi_busy "$state" busy-fresh - printf 'working: setup complete\n' > "$state/busy-fresh.status" - sig=$(seen_sig "$state/busy-fresh.status"); printf '%s' "$sig" > "$state/.seen-busy-fresh_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - touch "$state/busy-fresh.turn-ended" - prime_turnend_seen "$state/busy-fresh.turn-ended" - - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_BUSY_TURN_MAX_SECS=999 FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid"; fail "a busy pane below the turn-age bound was escalated: $(cat "$out")" - fi - [ ! -s "$out" ] || fail "a busy pane below the turn-age bound printed a wake reason" - [ ! -e "$state/.stale-since-$key" ] || fail "a busy pane below the turn-age bound started a wedge timer" - reap "$pid" - pass "a busy worker below the turn-age bound remains working with no escalation" -} - -test_busy_pane_stable_hash_escalates_past_turn_age_bound() { - local dir state fakebin out capture_file window key pane_hash sig pid - dir=$(make_case busy-stable-hash-turn-age); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-busy-stable" - printf 'Working...' > "$capture_file" - printf 'window=%s\nkind=ship\nharness=pi\n' "$window" > "$state/busy-stable.meta" - record_pi_busy "$state" busy-stable - printf 'working: setup complete\n' > "$state/busy-stable.status" - sig=$(seen_sig "$state/busy-stable.status"); printf '%s' "$sig" > "$state/.seen-busy-stable_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "Working...") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - # No completed turn ever recorded for this task: age the spawn record itself. - touch -t 200001010000 "$state/busy-stable.meta" - - # Phase A: past the bound, the stable-hash busy pane is absorbed but starts - # the wedge timer (mirrors the existing provably-working-stale Phase A/B). - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid"; fail "a stable-hash busy pane past the turn-age bound escalated before the wedge threshold: $(cat "$out")" - fi - [ -s "$state/.stale-since-$key" ] || fail "a stable-hash busy pane past the turn-age bound did not start a wedge timer" - reap "$pid" - ack_stopped_cycle "$state" || fail "could not acknowledge the intentional stable-hash phase-A stop" - - # Phase B: backdate the wedge timer past the threshold; the next poll escalates. - echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "a stable-hash busy pane did not wedge-escalate past the turn-age bound" - grep -F "stale: $window" "$out" >/dev/null || fail "busy turn-age escalation did not print the stale wake" - grep -F "possible wedge" "$out" >/dev/null || fail "busy turn-age escalation did not flag a possible wedge" - pass "a busy worker with a stable pane hash still escalates once its completed-turn age reaches the bound" -} - -# Regression fixture for the incident's actual masking condition: Pi's rendered -# elapsed-time footer changes every poll, so the pane hash never repeats and the -# watcher always takes the "new hash" branch, never the stable-hash one above. -test_busy_pane_changing_hash_escalates_past_turn_age_bound() { - local dir state fakebin out capture_file window key pid - dir=$(make_case busy-changing-hash-turn-age); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-busy-ticking" - printf 'Working... (3600.1s)' > "$capture_file" - printf 'window=%s\nkind=ship\nharness=pi\n' "$window" > "$state/busy-ticking.meta" - record_pi_busy "$state" busy-ticking - printf 'working: setup complete\n' > "$state/busy-ticking.status" - sig=$(seen_sig "$state/busy-ticking.status"); printf '%s' "$sig" > "$state/.seen-busy-ticking_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - touch -t 200001010000 "$state/busy-ticking.meta" - # No pre-seeded .hash-<key>: with a real ticking elapsed footer, every poll - # lands here (h != prev) - the reproduction's actual masking condition. - - # Phase A: first sight past the bound absorbs and starts the wedge timer, - # without ever needing the "genuinely stale" hash-match path. - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid"; fail "a changing-hash busy pane past the turn-age bound escalated before the wedge threshold: $(cat "$out")" - fi - [ -s "$state/.stale-since-$key" ] || fail "a changing-hash busy pane past the turn-age bound did not start a wedge timer" - reap "$pid" - ack_stopped_cycle "$state" || fail "could not acknowledge the intentional changing-hash phase-A stop" - - # Phase B: another tick (still a fresh, never-before-seen hash) plus a - # backdated wedge timer escalates exactly as the stable-hash case does. - printf 'Working... (3601.2s)' > "$capture_file" - echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "a changing-hash busy pane did not wedge-escalate past the turn-age bound" - grep -F "stale: $window" "$out" >/dev/null || fail "busy turn-age escalation (changing hash) did not print the stale wake" - grep -F "possible wedge" "$out" >/dev/null || fail "busy turn-age escalation (changing hash) did not flag a possible wedge" - pass "a busy worker whose pane hash changes every poll still escalates once its completed-turn age reaches the bound" -} - -test_busy_pane_turn_end_touch_resets_age() { - local dir state fakebin out capture_file window key pane_hash sig pid - dir=$(make_case busy-turn-end-resets-age); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-busy-reset" - printf 'Working...' > "$capture_file" - printf 'window=%s\nkind=ship\nharness=pi\n' "$window" > "$state/busy-reset.meta" - record_pi_busy "$state" busy-reset - printf 'working: setup complete\n' > "$state/busy-reset.status" - sig=$(seen_sig "$state/busy-reset.status"); printf '%s' "$sig" > "$state/.seen-busy-reset_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "Working...") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - # A wedge is already mid-escalation, as if several over-age polls already ran. - echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" - printf '1\n' > "$state/.wedge-escalations-$key" - # The worker's most recent turn just completed: touching turn-ended resets age. - touch "$state/busy-reset.turn-ended" - prime_turnend_seen "$state/busy-reset.turn-ended" - - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_BUSY_TURN_MAX_SECS=3600 FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid"; fail "a freshly completed turn on a busy pane was still escalated: $(cat "$out")" - fi - [ ! -s "$out" ] || fail "a freshly completed turn on a busy pane printed a wake reason" - [ ! -e "$state/.stale-since-$key" ] || fail "a freshly completed turn did not clear the wedge timer" - [ ! -e "$state/.wedge-escalations-$key" ] || fail "a freshly completed turn did not clear the escalation counter" - reap "$pid" - pass "touching a busy worker's completed-turn marker resets the age and prevents an old-age escalation" -} - -test_busy_pane_repeated_escalation_reaches_demand_deep_inspection() { - local dir state fakebin out capture_file window key pane_hash sig pid n - dir=$(make_case busy-turn-age-demand-inspect); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-busy-demand-inspect" - printf 'Working...' > "$capture_file" - printf 'window=%s\nkind=ship\nharness=pi\n' "$window" > "$state/busy-demand.meta" - record_pi_busy "$state" busy-demand - printf 'working: setup complete\n' > "$state/busy-demand.status" - sig=$(seen_sig "$state/busy-demand.status"); printf '%s' "$sig" > "$state/.seen-busy-demand_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "Working...") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - touch -t 200001010000 "$state/busy-demand.turn-ended" - prime_turnend_seen "$state/busy-demand.turn-ended" - - # Priming round: first sighting past the turn-age bound absorbs and starts - # the wedge timer, mirroring the existing provably-working wedge tests. - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid"; fail "priming round for busy turn-age escalation was not absorbed: $(cat "$out")" - fi - reap "$pid" - ack_stopped_cycle "$state" || fail "could not acknowledge the intentional busy-wedge priming stop" - - n=1 - while [ "$n" -le 3 ]; do - echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "busy turn-age escalation round $n did not escalate: $(cat "$out")" - grep -F "escalation $n" "$out" >/dev/null || fail "busy turn-age round $n did not report escalation count $n: $(cat "$out")" - if [ "$n" -lt 3 ]; then - grep -F "demand-deep-inspection" "$out" >/dev/null && fail "busy turn-age round $n escalated to demand-deep-inspection before the threshold: $(cat "$out")" - else - grep -F "demand-deep-inspection" "$out" >/dev/null || fail "busy turn-age round $n (threshold) did not demand deep inspection: $(cat "$out")" - fi - ack_stopped_cycle "$state" || fail "could not acknowledge busy turn-age escalation round $n" - n=$((n + 1)) - done - [ "$(cat "$state/.wedge-escalations-$key" 2>/dev/null || echo 0)" = 3 ] || fail "busy turn-age escalation counter did not persist across consecutive rounds" - pass "repeated busy turn-age escalations reuse the existing escalation counter and demand deep inspection at the threshold" -} - -# --- declared pause + busy pane: the busy-turn bound must honor the declaration -# A single foreground call can keep a declared external wait semantically busy -# past the completed-turn bound, bypassing the ordinary stale-pause path. -# This fixture pins all three halves of the contract: the declared pause is -# absorbed instead of wedged (A), it is still rechecked on the long -# PAUSE_RESURFACE_SECS cadence so a forgotten wait cannot rot invisibly (B), and -# lifting the declaration on the SAME busy over-age pane restores the wedge -# escalation, proving the discriminator is the worker's own declaration and not a -# blanket silencing of the escalator (C). -test_busy_declared_pause_is_rechecked_not_wedge_escalated() { - local dir state fakebin out capture_file window key sig pid statusf back - dir=$(make_case busy-declared-pause); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-review-scout" - statusf="$state/review-scout.status" - printf 'Working... (7200.4s) lavish-axi poll' > "$capture_file" - printf 'window=%s\nkind=scout\nharness=pi\n' "$window" > "$state/review-scout.meta" - record_pi_busy "$state" review-scout - printf 'paused: hosting the Lavish review, awaiting captain feedback\n' > "$statusf" - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-review-scout_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - # No completed turn for hours (the single blocking poll call): age the spawn - # record itself, exactly as the never-completed-a-turn fixtures above do. - touch -t 200001010000 "$state/review-scout.meta" - # No pre-seeded .hash-<key>: a live harness footer ticks, so every poll lands - # on the changed-hash branch - the review scout's real masking condition. - - # Phase A: past the bound, with the wedge threshold set as low as it goes, the - # declared pause is absorbed on the long cadence and never starts a wedge. - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ - FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=1 FM_PAUSE_RESURFACE_SECS=999 \ - FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_poll_cycle "$state" "$pid" || { reap "$pid"; fail "a declared pause on a busy review pane was escalated: $(cat "$out")"; } - reap "$pid" - [ ! -s "$out" ] || fail "a declared pause on a busy review pane printed a wake reason: $(cat "$out")" - [ -e "$state/.paused-$key" ] || fail "the busy-turn bound did not apply the declared-pause cadence" - [ ! -e "$state/.stale-since-$key" ] || fail "a declared pause on a busy pane started the wedge timer" - [ ! -e "$state/.wedge-escalations-$key" ] || fail "a declared pause on a busy pane incremented the escalation counter" - ack_stopped_cycle "$state" || fail "could not acknowledge the intentional declared-pause phase-A stop" - - # Phase B: age the pause past the (now normal) long cadence and let the pane - # settle on one stable hash, so the still-busy pane takes the repeat-hash - # branch whose pause bookkeeping the bound must not wipe. It re-surfaces once - # as a recheck, never as a wedge. - back=$(( $(date +%s) - 500 )) - if [ "$(uname)" = Darwin ]; then touch -mt "$(date -r "$back" '+%Y%m%d%H%M.%S')" "$statusf" - else touch -m -d "@$back" "$statusf"; fi - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-review-scout_status" - printf '%s' "$(hash_text "$(cat "$capture_file")")" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ - FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=1 FM_PAUSE_RESURFACE_SECS=240 \ - FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || { reap "$pid"; fail "a declared pause past the long cadence was never rechecked"; } - grep -F "awaiting external" "$out" >/dev/null || fail "the recheck was not labeled a declared-pause recheck: $(cat "$out")" - grep -F "possible wedge" "$out" >/dev/null && fail "a declared pause on a busy pane was mislabeled a possible wedge: $(cat "$out")" - [ -e "$state/.paused-resurfaced-$key" ] || fail "the declared-pause re-surface throttle was cleared by the busy-turn bound" - [ ! -e "$state/.stale-since-$key" ] || fail "a declared-pause recheck used the wedge timer" - ack_stopped_cycle "$state" || fail "could not acknowledge the declared-pause recheck" - - # Phase C: the pause is lifted on the SAME busy, over-age pane. Nothing else - # changes, so a still-absorbed pane here would mean the bound was silenced - # rather than taught the declaration. It must wedge-escalate exactly as before. - printf 'working: review closed, resuming the sweep\n' > "$statusf" - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-review-scout_status" - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ - FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=999 FM_PAUSE_RESURFACE_SECS=999 \ - FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_poll_cycle "$state" "$pid" || { reap "$pid"; fail "a lifted pause escalated before the wedge threshold: $(cat "$out")"; } - reap "$pid" - [ -s "$state/.stale-since-$key" ] || fail "a lifted pause did not restore the busy-turn wedge timer" - [ ! -e "$state/.paused-$key" ] || fail "a lifted pause left stale declared-pause bookkeeping behind" - ack_stopped_cycle "$state" || fail "could not acknowledge the intentional lifted-pause priming stop" - - echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ - FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=240 FM_PAUSE_RESURFACE_SECS=999 \ - FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || { reap "$pid"; fail "a lifted pause on an over-age busy pane no longer wedge-escalates"; } - grep -F "possible wedge" "$out" >/dev/null || fail "the restored busy-turn escalation did not flag a possible wedge: $(cat "$out")" - pass "a busy pane under a declared pause is rechecked on the long cadence, and lifting the pause restores the wedge escalation" -} - -# --- declared pause + busy pane + AWAY MODE: the bound must hand off, not decorate -# Away mode is daemon-owned: the watcher reverts to one-shot and lets the daemon -# classify. The busy-turn bound used to be the one stale path that ignored that, -# running the wedge timer under afk and handing the daemon a wake already decorated -# as a possible wedge. That decoration outranks the daemon's own pause verdict, so a -# crew that declared the wait itself was wedge-escalated once per -# FM_STALE_ESCALATE_SECS for as long as the wait lasted, with the escalation count -# climbing into demand-deep-inspection on a pane nobody needed to inspect. -# Phase A pins the handoff: the plain window identity, no wedge timer, no escalation -# counter, and no normal-mode pause bookkeeping (the daemon owns that in away mode). -# Phase B re-arms on the same unchanged pane and pins the one-shot: a second wake -# here is what the climbing ladder looked like. Phase C drives the discriminator -# apart on the SAME afk, busy, over-age pane - lifting the declaration restores the -# wedge escalation, so this is the worker's declaration being honored rather than -# away mode silencing the escalator. -test_afk_busy_declared_pause_hands_off_plain_stale() { - local dir state fakebin out capture_file window key sig pid statusf - dir=$(make_case afk-busy-declared-pause); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-afk-review-scout" - statusf="$state/afk-review-scout.status" - printf 'Working... (7200.4s) lavish-axi poll' > "$capture_file" - printf 'window=%s\nkind=scout\nharness=pi\n' "$window" > "$state/afk-review-scout.meta" - record_pi_busy "$state" afk-review-scout - printf 'paused: hosting the Lavish review, awaiting captain feedback\n' > "$statusf" - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-afk-review-scout_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - touch -t 200001010000 "$state/afk-review-scout.meta" - date '+%s' > "$state/.afk" - - # Phase A: past the bound, with the wedge threshold as low as it goes, the - # declaration is handed to the daemon undecorated instead of being wedge-timed. - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ - FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=1 FM_PAUSE_RESURFACE_SECS=999 \ - FM_POLL=0.2 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 150 || { reap "$pid"; fail "the away-mode busy-turn bound never handed the declared pause to the daemon"; } - grep -Fx "stale: $window" "$out" >/dev/null \ - || fail "the away-mode busy-turn bound did not hand off the plain window identity: $(cat "$out")" - grep -F "possible wedge" "$out" >/dev/null \ - && fail "away mode decorated a declared pause as a possible wedge: $(cat "$out")" - [ ! -e "$state/.stale-since-$key" ] \ - || fail "the away-mode handoff started the wedge timer on a declared pause" - [ ! -e "$state/.wedge-escalations-$key" ] \ - || fail "the away-mode handoff incremented the wedge escalation count on a declared pause" - [ ! -e "$state/.paused-$key" ] \ - || fail "the away-mode handoff recorded normal-mode pause tracking instead of leaving it to the daemon" - ack_stopped_cycle "$state" || fail "could not acknowledge the away-mode declared-pause handoff" - - # Phase B: re-arm on the same unchanged pane. The bound has already handed this - # stale hash off, so it must stay silent rather than re-waking the daemon - a - # second wake here is the escalation ladder the wedge timer used to climb. - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ - FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=1 FM_PAUSE_RESURFACE_SECS=999 \ - FM_POLL=0.2 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_poll_cycle "$state" "$pid" || { reap "$pid"; fail "the away-mode bound re-woke on an already-handed-off declared pause: $(cat "$out")"; } - reap "$pid" - [ ! -s "$out" ] || fail "the away-mode bound re-surfaced an already-handed-off declared pause: $(cat "$out")" - [ ! -e "$state/.wedge-escalations-$key" ] \ - || fail "re-arming on an unchanged declared pause started a wedge escalation ladder" - ack_stopped_cycle "$state" || fail "could not acknowledge the intentional away-mode re-arm stop" - - # Phase C: lift the declaration on the SAME afk, busy, over-age pane. Nothing else - # changes, so a wedge escalation here proves the declaration was the discriminator. - printf 'working: resumed the review write-up\n' > "$statusf" - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-afk-review-scout_status" - echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ - FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=240 FM_PAUSE_RESURFACE_SECS=999 \ - FM_POLL=0.2 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 150 || { reap "$pid"; fail "a lifted pause on an away-mode over-age busy pane no longer wedge-escalates"; } - grep -F "possible wedge" "$out" >/dev/null \ - || fail "the restored away-mode busy-turn escalation did not flag a possible wedge: $(cat "$out")" - pass "away mode hands a busy declared pause to the daemon as a plain stale, and lifting the declaration restores the wedge escalation" -} - -# --- declared pause + busy pane + AWAY MODE + a TICKING footer: one wake per declaration -# The static-pane case above cannot tell a hash-keyed one-shot from a -# declaration-keyed one, because its capture never changes between polls. The -# incident pane's harness footer ticks on every capture, so a one-shot keyed on the -# pane hash re-fires on every poll, and the daemon, which relaunches the watcher -# after each handled wake, is woken in a loop for the whole declared wait. This -# fixture's fake tmux renders a fresh footer on EVERY capture-pane and asserts that -# divergence outright on every re-arm (.hash-<key> moves, .count-<key> never -# climbs), so the one-wake assertion across five silent re-arms cannot pass -# vacuously on a pane that happened to sit still. Round 1 also starts from an -# undeclared wedge timer and escalation count, which the handoff must clear the -# way the normal-mode absorber does, so lifting the declaration later starts the -# wedge path from a fresh timer rather than resuming a stale count. -test_afk_busy_declared_pause_ticking_pane_hands_off_once() { - local dir state fakebin out drain_out window key sig pid statusf ticks round prev_hash cur_hash prev_ticks - dir=$(make_case afk-busy-declared-pause-ticking); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; window="test:fm-afk-ticking-scout" - statusf="$state/afk-ticking-scout.status"; ticks="$dir/ticks" - cat > "$fakebin/tmux" <<'SH' -#!/usr/bin/env bash -set -u -case "${1:-}" in - list-windows) - [ -n "${FM_FAKE_TMUX_WINDOW:-}" ] && printf '%s\n' "${FM_FAKE_TMUX_WINDOW#*:}" - exit 0 ;; - capture-pane) - n=$(( $(cat "$FM_FAKE_TMUX_TICKS" 2>/dev/null || echo 0) + 1 )) - echo "$n" > "$FM_FAKE_TMUX_TICKS" - printf 'Working... (%d.%ds) lavish-axi poll' "$(( 7200 + n ))" "$(( n % 10 ))" - exit 0 ;; - display-message) - case "$*" in - *pane_current_command*) printf '%s\n' "${FM_FAKE_TMUX_CURRENT_COMMAND:-}"; exit 0 ;; - esac ;; -esac -exit 1 -SH - chmod +x "$fakebin/tmux" - printf 'window=%s\nkind=scout\nharness=pi\n' "$window" > "$state/afk-ticking-scout.meta" - record_pi_busy "$state" afk-ticking-scout - printf 'paused: hosting the Lavish review, awaiting captain feedback\n' > "$statusf" - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-afk-ticking-scout_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - touch -t 200001010000 "$state/afk-ticking-scout.meta" - date '+%s' > "$state/.afk" - # An undeclared busy phase already ran the wedge timer and escalated twice - # before the crew declared the wait. - echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" - printf '2\n' > "$state/.wedge-escalations-$key" - date +%s > "$state/.writing-since-$key" - - # Round 1: the declaration is handed off once, undecorated, and the undeclared - # phase's wedge bookkeeping is cleared with it. - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_TICKS="$ticks" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ - FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=1 FM_PAUSE_RESURFACE_SECS=999 \ - FM_POLL=0.2 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 150 || { reap "$pid"; fail "the away-mode busy-turn bound never handed a ticking declared pause to the daemon"; } - grep -Fx "stale: $window" "$out" >/dev/null \ - || fail "the away-mode busy-turn bound did not hand off the plain window identity for a ticking pane: $(cat "$out")" - grep -F "possible wedge" "$out" >/dev/null \ - && fail "away mode decorated a ticking declared pause as a possible wedge: $(cat "$out")" - [ ! -e "$state/.stale-since-$key" ] \ - || fail "the away-mode handoff left the undeclared phase's wedge timer in place" - [ ! -e "$state/.wedge-escalations-$key" ] \ - || fail "the away-mode handoff left the undeclared phase's escalation count in place" - [ ! -e "$state/.writing-since-$key" ] \ - || fail "the away-mode handoff left the undeclared phase's write-deferral chain in place" - [ ! -e "$state/.paused-$key" ] \ - || fail "the away-mode handoff recorded normal-mode pause tracking on a ticking pane" - ack_stopped_cycle "$state" || fail "could not acknowledge the ticking declared-pause handoff" - - # Rounds 2-6: five consecutive re-arms on the same standing declaration. Every - # capture renders a new footer, so every poll lands on the changed-hash branch - - # the exact shape a hash-keyed one-shot re-fires on. Each round proves the pane - # really moved before it asserts silence, so the case cannot go vacuous. - round=2 - while [ "$round" -le 6 ]; do - prev_hash=$(cat "$state/.hash-$key" 2>/dev/null || true) - prev_ticks=$(cat "$ticks" 2>/dev/null || echo 0) - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_TICKS="$ticks" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ - FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=1 FM_PAUSE_RESURFACE_SECS=999 \ - FM_POLL=0.2 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_poll_cycle "$state" "$pid" || { reap "$pid"; fail "re-arm $round on a ticking declared pause re-woke the daemon: $(cat "$out")"; } - reap "$pid" - cur_hash=$(cat "$state/.hash-$key" 2>/dev/null || true) - [ "$(cat "$ticks" 2>/dev/null || echo 0)" -gt "$prev_ticks" ] \ - || fail "re-arm $round never captured the pane, so its silence proves nothing" - [ -n "$cur_hash" ] && [ "$cur_hash" != "$prev_hash" ] \ - || fail "re-arm $round saw the same pane hash as the round before, so it cannot tell a hash-keyed one-shot from a declaration-keyed one" - [ "$(cat "$state/.count-$key" 2>/dev/null || echo missing)" = 0 ] \ - || fail "re-arm $round settled on a stable hash instead of ticking on every poll" - [ ! -s "$out" ] || fail "re-arm $round re-surfaced a standing declared pause on a ticking pane: $(cat "$out")" - [ ! -e "$state/.stale-since-$key" ] \ - || fail "re-arm $round started the wedge timer on a standing declared pause" - [ ! -e "$state/.wedge-escalations-$key" ] \ - || fail "re-arm $round climbed the wedge escalation ladder on a standing declared pause" - ack_stopped_cycle "$state" || fail "could not acknowledge the intentional re-arm $round stop" - round=$((round + 1)) - done - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || true - grep "$(printf '\tstale\t')" "$drain_out" >/dev/null \ - && fail "the silent re-arms still queued a stale row for the standing declaration: $(cat "$drain_out")" - pass "away mode wakes the daemon once per declaration for a busy pane whose footer ticks on every capture" -} - -# Behavioral proof that the production default (no FM_BUSY_TURN_MAX_SECS override -# anywhere in this env) is 3600s: a completed turn 5 minutes old must not start a -# wedge timer, while one 66 minutes old must - bracketing the default around 3600 -# without waiting a literal hour. -test_busy_pane_default_turn_age_bound_is_3600s() { - local dir state fakebin out capture_file window key pane_hash sig pid - dir=$(make_case busy-default-turn-age); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-busy-default" - printf 'Working...' > "$capture_file" - printf 'window=%s\nkind=ship\nharness=pi\n' "$window" > "$state/busy-default.meta" - record_pi_busy "$state" busy-default - printf 'working: setup complete\n' > "$state/busy-default.status" - sig=$(seen_sig "$state/busy-default.status"); printf '%s' "$sig" > "$state/.seen-busy-default_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "Working...") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - - set_mtime $(( $(date +%s) - 300 )) "$state/busy-default.turn-ended" - prime_turnend_seen "$state/busy-default.turn-ended" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid"; fail "a 5-minute-old completed turn tripped the default busy-turn-age bound: $(cat "$out")" - fi - [ ! -e "$state/.stale-since-$key" ] || fail "a 5-minute-old completed turn started a wedge timer under the default bound" - reap "$pid" - ack_stopped_cycle "$state" || fail "could not acknowledge the intentional five-minute-bound stop" - - set_mtime $(( $(date +%s) - 4000 )) "$state/busy-default.turn-ended" - prime_turnend_seen "$state/busy-default.turn-ended" - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid"; fail "a 66-minute-old completed turn escalated before the wedge threshold under the default bound: $(cat "$out")" - fi - [ -s "$state/.stale-since-$key" ] || fail "a 66-minute-old completed turn did not start a wedge timer under the default bound (default is not 3600s)" - reap "$pid" - pass "the production default busy-turn-age bound is 3600s (5min under does not wedge, 66min over does)" -} - -test_nonterminal_stale_repairs_missing_or_corrupt_timer() { - local dir state fakebin out capture_file window key pane_hash sig pid since - dir=$(make_case nonterminal-stale-timer-repair); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt" - window="test:fm-quiet-timer" - printf 'idle building output' > "$capture_file" - printf 'window=%s\nkind=ship\n' "$window" > "$state/quiet-timer.meta" - printf 'working: still compiling\n' > "$state/quiet-timer.status" - sig=$(seen_sig "$state/quiet-timer.status"); printf '%s' "$sig" > "$state/.seen-quiet-timer_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "idle building output") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - printf '%s' "$pane_hash" > "$state/.stale-$key" - - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_numeric_file "$state/.stale-since-$key" 30 || { reap "$pid"; fail "matching stale suppressor with missing timer did not initialize stale-since"; } - if ! kill -0 "$pid" 2>/dev/null; then - wait "$pid" 2>/dev/null || true - fail "watcher exited while repairing a missing stale-since timer: $(cat "$out")" - fi - [ ! -s "$state/.wake-queue" ] || { reap "$pid"; fail "missing stale-since repair enqueued a wake"; } - reap "$pid" - ack_stopped_cycle "$state" || fail "could not acknowledge the intentional missing-timer repair stop" - - printf 'corrupt\n' > "$state/.stale-since-$key" - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_numeric_file "$state/.stale-since-$key" 30 || { reap "$pid"; fail "matching stale suppressor with corrupt timer did not repair stale-since"; } - since=$(cat "$state/.stale-since-$key" 2>/dev/null || true) - [ "$since" != "corrupt" ] || { reap "$pid"; fail "corrupt stale-since value was left in place"; } - [ ! -s "$state/.wake-queue" ] || { reap "$pid"; fail "corrupt stale-since repair enqueued a wake"; } - reap "$pid" - pass "matching non-terminal stale suppressors repair missing or corrupt stale-since timers" -} - -# --- quiet pane, worktree still being written: deferred, never wedge-escalated - -# The live 2026-08-14 case: one crew produced eight consecutive possible-wedge -# escalations in an afternoon, three of them demanding deep inspection, while it -# was demonstrably writing source, then tests, then documentation. The detector's -# two inputs (pane quietness, run step) cannot see that, so the pane looks frozen. -# Both halves of the contract are asserted on the SAME fixture, because the whole -# point is that only the worktree evidence differs: writing defers, silent -# escalates on the unchanged schedule. -# Every wait below is the file's standard one (wait_poll_cycle for an absorbing -# watcher, a 100-tick wait_for_exit for an escalating one), because the poll these -# tests assert on is the ONE poll that spawns the bounded worktree walk: on a -# loaded runner it outlives a fixed liveness budget, and a round reaped before it -# finished reports a lost deferral instead of the deferral under test. -test_wedge_escalation_deferred_while_worktree_is_written() { - local dir state fakebin out drain_out capture_file window key pane_hash sig pid wt back - dir=$(make_case wedge-worktree-writes); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - window="test:fm-writing"; wt="$dir/wt" - mkdir -p "$wt/src" - printf 'idle building output' > "$capture_file" - printf 'window=%s\nkind=ship\nworktree=%s\n' "$window" "$wt" > "$state/writing.meta" - printf 'working: implementing\n' > "$state/writing.status" - sig=$(seen_sig "$state/writing.status"); printf '%s' "$sig" > "$state/.seen-writing_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "idle building output") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - # Already-classified hash with an idle window that opened 500s ago, so the very - # first stale poll lands straight on the at-threshold wedge branch (this repeat - # path never re-reads crew state, so the worktree evidence is the only input - # that can change the outcome). - printf '%s' "$pane_hash" > "$state/.stale-$key" - back=$(( $(date +%s) - 500 )) - echo "$back" > "$state/.stale-since-$key" - set_mtime "$back" "$state/.stale-since-$key" - - # Phase A: the crew wrote a file after the idle window opened. Deferred. - printf 'int main(void) { return 0; }\n' > "$wt/src/main.c" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 \ - FM_PAUSE_RESURFACE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid"; fail "watcher wedge-escalated a quiet pane whose worktree was being written: $(cat "$out")" - fi - [ ! -s "$out" ] || { reap "$pid"; fail "a written-worktree deferral printed a wake reason: $(cat "$out")"; } - [ ! -s "$state/.wake-queue" ] || { reap "$pid"; fail "a written-worktree deferral enqueued a wake"; } - [ -e "$state/.writing-since-$key" ] || { reap "$pid"; fail "the write-deferral chain marker was not recorded"; } - [ ! -e "$state/.wedge-escalations-$key" ] || { reap "$pid"; fail "a deferral advanced the wedge escalation counter"; } - [ "$(cat "$state/.stale-since-$key" 2>/dev/null || echo 0)" -gt "$back" ] \ - || { reap "$pid"; fail "a deferral did not restart the idle timer, so the next window cannot re-probe"; } - reap "$pid" - ack_stopped_cycle "$state" || fail "could not acknowledge the intentional phase-A watcher stop" - - # Phase B: same fixture, same quiet pane, but nothing written during this idle - # window (the crew really is stalled). The unchanged schedule must still fire. - set_mtime "$(( $(date +%s) - 900 ))" "$wt/src/main.c" - echo "$back" > "$state/.stale-since-$key" - set_mtime "$back" "$state/.stale-since-$key" - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 \ - FM_PAUSE_RESURFACE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "a stalled crew that wrote nothing did not wedge-escalate on the existing schedule" - grep -F "stale: $window" "$out" >/dev/null || fail "the stalled-crew escalation did not print a stale wake" - grep -F "possible wedge" "$out" >/dev/null || fail "the stalled-crew escalation did not flag a possible wedge" - [ "$(cat "$state/.wedge-escalations-$key" 2>/dev/null || true)" = 1 ] || fail "the stalled-crew escalation was not counted" - [ ! -e "$state/.stale-since-$key" ] || fail "the idle timer was not cleared after a real escalation" - [ ! -e "$state/.writing-since-$key" ] || fail "the write-deferral chain outlived a real escalation" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the stalled-crew escalation failed" - grep "$(printf '\tstale\t')" "$drain_out" | grep -F "$window" >/dev/null || fail "the stalled-crew escalation was not queued" - pass "a quiet pane writing its own worktree is deferred, while one writing nothing still wedge-escalates on the unchanged schedule" -} - -# A deferral is not silence. A worktree can churn without real progress (a -# rewritten log, a build touching the same file), so the whole deferral chain ages -# and re-surfaces once per PAUSE_RESURFACE_SECS - the same bounded cadence a -# declared pause uses - labeled as a recheck rather than a wedge. -test_write_deferral_resurfaces_on_the_bounded_cadence() { - local dir state fakebin out drain_out capture_file window key pane_hash sig pid wt back - dir=$(make_case wedge-worktree-resurface); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - window="test:fm-churn"; wt="$dir/wt" - mkdir -p "$wt/src" - printf 'idle building output' > "$capture_file" - printf 'window=%s\nkind=ship\nworktree=%s\n' "$window" "$wt" > "$state/churn.meta" - printf 'working: implementing\n' > "$state/churn.status" - sig=$(seen_sig "$state/churn.status"); printf '%s' "$sig" > "$state/.seen-churn_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "idle building output") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - printf '%s' "$pane_hash" > "$state/.stale-$key" - back=$(( $(date +%s) - 500 )) - echo "$back" > "$state/.stale-since-$key" - set_mtime "$back" "$state/.stale-since-$key" - # This pane has been deferring on write evidence for 500s already. - : > "$state/.writing-since-$key" - set_mtime "$back" "$state/.writing-since-$key" - printf 'churn\n' > "$wt/src/main.c" - - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 \ - FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "a long-running write deferral never re-surfaced on the bounded cadence" - grep -F "stale: $window" "$out" >/dev/null || fail "the write-deferral recheck did not print a stale wake" - grep -F "writing its worktree" "$out" >/dev/null || fail "the write-deferral recheck was not labeled as such" - grep -F "possible wedge" "$out" >/dev/null && fail "a write-deferral recheck was mislabeled a possible wedge" - [ -e "$state/.writing-resurfaced-$key" ] || fail "the write-deferral re-surface throttle marker was not recorded" - [ ! -e "$state/.wedge-escalations-$key" ] || fail "a write-deferral recheck advanced the wedge escalation counter" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the write-deferral recheck failed" - grep "$(printf '\tstale\t')" "$drain_out" | grep -F "$window" >/dev/null || fail "the write-deferral recheck was not queued" - pass "a write deferral re-surfaces once on the bounded pause cadence, so a churning worktree cannot stay invisible" -} - -# The worktree recorded for a secondmate is a provisioned firstmate home, and that -# home runs its OWN supervision inside itself: its watcher beacon, pane hashes and -# heartbeats keep state/ churning whether or not the mate produced anything. Reading -# that as crew progress would quietly relax the kind-agnostic busy-turn backstop from -# the escalation cadence to the hourly recheck for work that produced nothing, so the -# probe must report no evidence and the unchanged schedule must still fire. -test_secondmate_home_supervision_churn_is_not_write_evidence() { - local dir state fakebin out drain_out capture_file window key sig pid home back - dir=$(make_case secondmate-home-churn); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - window="test:fm-mate"; home="$dir/mate-home" - mkdir -p "$home/state" - printf 'sm-mate\n' > "$home/.fm-secondmate-home" - printf 'Working... (12.3s)' > "$capture_file" - printf 'window=%s\nkind=ship\nharness=pi\nworktree=%s\n' "$window" "$home" > "$state/mate.meta" - record_pi_busy "$state" mate - # An ordinary crew recording a provisioned mate home is the route that actually - # reaches the probe: a kind=secondmate window of its own is triaged only under a - # declared pause, and a declared pause takes the bounded recheck cadence instead of - # the wedge timer. The home marker alone is what excludes the walk, so the exclusion - # is what this asserts. A busy pane is bounded by its completed-turn age; no turn - # ever completed here, so the spawn record itself is aged past the bound that routes - # it into the wedge timer. - printf 'working: implementing\n' > "$state/mate.status" - sig=$(seen_sig "$state/mate.status"); printf '%s' "$sig" > "$state/.seen-mate_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - set_mtime "$(( $(date +%s) - 4000 ))" "$state/mate.meta" - back=$(( $(date +%s) - 500 )) - echo "$back" > "$state/.stale-since-$key" - set_mtime "$back" "$state/.stale-since-$key" - # The only thing written since the idle window opened is the mate home's own - # supervision bookkeeping. - printf 'beat\n' > "$home/state/.last-watcher-beat" - - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_STALE_ESCALATE_SECS=240 FM_BUSY_TURN_MAX_SECS=1 FM_PAUSE_RESURFACE_SECS=999 \ - FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "a mate home's own supervision churn deferred an escalation it must not defer" - grep -F "stale: $window" "$out" >/dev/null || fail "the mate-home escalation did not print a stale wake" - grep -F "possible wedge" "$out" >/dev/null || fail "the mate-home escalation did not flag a possible wedge" - [ ! -e "$state/.writing-since-$key" ] || fail "a mate's provisioned home was probed as if it were a code tree" - [ "$(cat "$state/.wedge-escalations-$key" 2>/dev/null || true)" = 1 ] || fail "the mate escalation was not counted" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the mate escalation failed" - grep "$(printf '\tstale\t')" "$drain_out" | grep -F "$window" >/dev/null || fail "the mate escalation was not queued" - pass "a secondmate's own home supervision churn is not crew write evidence, so a pane recording that home keeps the unchanged escalation schedule" -} - -# A write deferral is a bounded chain, not a permanent one: its .writing-since -# marker ages the whole chain so a churning worktree still re-surfaces once per -# PAUSE_RESURFACE_SECS. That only holds while the chain belongs to the CURRENT quiet -# stretch, so every path that restarts the idle-window timer must drop it too. The -# reachable case is a pane that deferred on write evidence and later has its timer -# repaired: a long-finished chain would make the first deferral of the new window -# re-surface immediately instead of after a fresh window. -test_timer_repair_drops_a_finished_write_deferral_chain() { - local dir state fakebin out capture_file window key pane_hash sig pid wt back - dir=$(make_case wedge-write-chain-timer-repair); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt" - window="test:fm-chain-repair"; wt="$dir/wt" - mkdir -p "$wt/src" - printf 'idle building output' > "$capture_file" - printf 'window=%s\nkind=ship\nworktree=%s\n' "$window" "$wt" > "$state/chain-repair.meta" - printf 'working: implementing\n' > "$state/chain-repair.status" - sig=$(seen_sig "$state/chain-repair.status"); printf '%s' "$sig" > "$state/.seen-chain-repair_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "idle building output") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - printf '%s' "$pane_hash" > "$state/.stale-$key" - # A deferral chain left over from an earlier quiet stretch, already well past the - # bounded re-surface window. - back=$(( $(date +%s) - 5000 )) - : > "$state/.writing-since-$key" - set_mtime "$back" "$state/.writing-since-$key" - # The idle-window timer is corrupt, so this poll repairs it and opens a NEW quiet - # window without probing the worktree at all. - printf 'corrupt\n' > "$state/.stale-since-$key" - - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_STALE_ESCALATE_SECS=240 FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - # Watcher startup performs bounded recovery scans before its first stale poll; - # give this positive marker assertion the same loaded-runner budget as the - # suite's other startup-sensitive waits instead of failing after only 3s. - wait_numeric_file "$state/.stale-since-$key" 100 \ - || { reap "$pid"; fail "the corrupt idle-window timer was not repaired"; } - [ ! -e "$state/.writing-since-$key" ] \ - || { reap "$pid"; fail "an idle-window timer repair kept a finished write-deferral chain"; } - [ ! -s "$state/.wake-queue" ] || { reap "$pid"; fail "the idle-window timer repair enqueued a wake"; } - reap "$pid" - ack_stopped_cycle "$state" || fail "could not acknowledge the intentional timer-repair watcher stop" - - # The new quiet window now crosses the escalation threshold while the crew writes - # its worktree. That deferral must get a FRESH re-surface window rather than - # inheriting the finished chain's age. - back=$(( $(date +%s) - 500 )) - echo "$back" > "$state/.stale-since-$key" - set_mtime "$back" "$state/.stale-since-$key" - printf 'int main(void) { return 0; }\n' > "$wt/src/main.c" - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_STALE_ESCALATE_SECS=240 FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid" - fail "the first deferral of a new quiet window re-surfaced at once, so it inherited a finished chain: $(cat "$out")" - fi - [ ! -s "$out" ] || { reap "$pid"; fail "a fresh write deferral printed a wake reason: $(cat "$out")"; } - [ ! -s "$state/.wake-queue" ] || { reap "$pid"; fail "a fresh write deferral enqueued a wake"; } - [ -e "$state/.writing-since-$key" ] || { reap "$pid"; fail "the new deferral recorded no chain marker"; } - [ ! -e "$state/.writing-resurfaced-$key" ] \ - || { reap "$pid"; fail "a fresh write deferral spent its bounded re-surface on the first poll"; } - reap "$pid" - pass "an idle-window timer repair drops a finished write-deferral chain, so the next deferral gets a fresh re-surface window" -} - -# The same chain must not outlive either first-sight path through a captain-relevant -# status line, because both also open a new idle window: the provably-working absorb -# and the plain surface. -test_terminal_first_sight_drops_a_finished_write_deferral_chain() { - local dir state fakebin out capture_file window key pane_hash sig pid wt back - dir=$(make_case wedge-write-chain-first-sight); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; capture_file="$dir/pane.txt" - window="test:fm-chain-firstsight"; wt="$dir/wt" - mkdir -p "$wt/src" - printf 'no-mistakes axi run: validating...' > "$capture_file" - printf 'window=%s\nkind=ship\nworktree=%s\n' "$window" "$wt" > "$state/chain-first.meta" - printf 'done: implementation complete, ready to validate\n' > "$state/chain-first.status" - sig=$(seen_sig "$state/chain-first.status"); printf '%s' "$sig" > "$state/.seen-chain-first_status" - key=$(printf '%s' "$window" | tr ':/.' '___') - pane_hash=$(hash_text "no-mistakes axi run: validating...") - printf '%s' "$pane_hash" > "$state/.hash-$key" - printf '1\n' > "$state/.count-$key" - back=$(( $(date +%s) - 5000 )) - : > "$state/.writing-since-$key" - set_mtime "$back" "$state/.writing-since-$key" - export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' - - # First sight of this hash, absorbed because the active run outranks the stale - # captain-relevant line. The absorb opens a new idle window, so the finished chain - # must go with it. - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_STALE_ESCALATE_SECS=999 FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid"; fail "the overridden terminal status was not absorbed on first sight: $(cat "$out")" - fi - [ "$(cat "$state/.stale-$key" 2>/dev/null || true)" = "$pane_hash" ] \ - || { reap "$pid"; fail "the first-sight absorb did not advance the stale suppressor"; } - [ ! -e "$state/.writing-since-$key" ] \ - || { reap "$pid"; fail "the provably-working first-sight absorb kept a finished write-deferral chain"; } - reap "$pid" - ack_stopped_cycle "$state" || fail "could not acknowledge the intentional first-sight absorb stop" - - # Same pane, first sight again, but nothing overrides the status line now, so it - # surfaces. That path drops the idle-window timer, so it must drop the chain too. - rm -f "$state/.stale-$key" "$state/.stale-since-$key" - printf '1\n' > "$state/.count-$key" - : > "$state/.writing-since-$key" - set_mtime "$back" "$state/.writing-since-$key" - FM_FAKE_CREW_STATE='state: unknown · source: none · no run, no busy pane' - : > "$out" - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ - FM_STALE_ESCALATE_SECS=999 FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "a first-sight captain-relevant status was not surfaced" - grep -F "stale: $window" "$out" >/dev/null || fail "the first-sight surface did not print a stale wake" - [ ! -e "$state/.writing-since-$key" ] \ - || fail "the first-sight surface kept a finished write-deferral chain" - unset FM_FAKE_CREW_STATE - pass "both first-sight paths through a captain-relevant status drop a finished write-deferral chain with the idle window" -} - -# --- triage debug log stays size capped ------------------------------------- - -test_triage_log_size_cap_accepts_spaced_wc_counts() { - local dir state fakebin out status_file pid lines i - dir=$(make_case triage-log-spaced-wc); state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out" - i=1 - while [ "$i" -le 3000 ]; do - printf 'old line %04d\n' "$i" >> "$state/.watch-triage.log" - i=$((i + 1)) - done - cat > "$fakebin/wc" <<'SH' -#!/usr/bin/env bash -set -u -if [ "${1:-}" = "-c" ]; then - printf ' 999999\n' - exit 0 -fi -exit 127 -SH - chmod +x "$fakebin/wc" - status_file="$state/task.status" - printf 'working: compiling step 2\n' > "$status_file" - # Provably working so the no-verb signal is absorbed (which is what writes the - # triage log line under test). - export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' - PATH="$fakebin:$PATH" FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 FM_WATCH_TRIAGE_LOG_MAX_BYTES=1 "$WATCH" > "$out" & - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid"; fail "watcher exited for a benign signal while testing log capping: $(cat "$out")" - fi - i=0 - while [ "$i" -lt 30 ]; do - lines=$(awk 'END { print NR + 0 }' "$state/.watch-triage.log") - [ "$lines" -le 2000 ] && break - sleep 0.1 - i=$((i + 1)) - done - [ "$lines" -le 2000 ] || { reap "$pid"; fail "triage log was not capped when wc emitted a spaced byte count (lines=$lines)"; } - [ ! -s "$state/.wake-queue" ] || { reap "$pid"; fail "benign signal enqueued a wake while testing log capping"; } - reap "$pid" - pass "triage log capping handles wc byte counts with leading spaces" -} - -# --- process-event delivery ------------------------------------------------- -# A durably captured process-event result publishes an ordinary `check` wake on -# the durable queue. The watcher must deliver that queued wake proactively - -# print an actionable reason and exit into the same rewake path every other -# actionable wake uses - rather than leaving it to be found by a manual drain. - -# Run the runner against a case home. FM_ROOT_OVERRIDE (exported by the shared -# wake harness to keep the drain's tangle check inert) would otherwise point the -# runner at a root with no installed adapters, and the claim root must stay -# inside the case so nothing here can observe a real home's source ownership. -pe_case() { # <dir> <command>... - local dir=$1 - dir=$(cd "$dir" && pwd -P) || return 1 - shift - (unset FM_ROOT_OVERRIDE - FM_PROCEVENT_CLAIM_ROOT="$dir/claims" FM_HOME="$dir" "$ROOT/bin/fm-procevent.sh" "$@") -} - -# Capture one real process-event result into <dir>'s home, then retire the -# source so the fixture holds exactly the reported end state: one durably -# captured, unhandled, queued result and no remaining poll work. -seed_captured_procevent_result() { # <dir> - local dir=$1 i=0 - pe_case "$dir" register lavish delivery-src -- \ - /bin/sh -c 'printf "session:\n file: /a.html\n status: waiting\n"' >/dev/null || return 1 - pe_case "$dir" reconcile >/dev/null || return 1 - while [ "$i" -lt 100 ]; do - [ -s "$dir/state/.wake-queue" ] && break - sleep 0.1 - i=$((i + 1)) - done - pe_case "$dir" retire delivery-src >/dev/null || return 1 - [ -s "$dir/state/.wake-queue" ] -} - -# The watcher, scoped by FM_HOME rather than FM_STATE_OVERRIDE, so the -# per-cycle reconcile it launches resolves the same home's state. -procevent_watch_bg() { # <dir> <out> - local dir=$1 out=$2 - dir=$(cd "$dir" && pwd -P) || return 1 - PATH="$dir/fakebin:$PATH" FM_HOME="$dir" FM_PROCEVENT_CLAIM_ROOT="$dir/claims" \ - FM_CREW_STATE_BIN="$dir/fakebin/fm-crew-state.sh" \ - FM_POLL=0.2 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & -} - -test_procevent_captured_result_surfaces_proactively() { - local dir state out drain_out pid beacon_age - dir=$(make_case procevent-delivery); state="$dir/state" - out="$dir/watch.out"; drain_out="$dir/drain.out" - seed_captured_procevent_result "$dir" || fail "the fixture captured no process-event result" - grep -F "procevent lavish delivery-src 1" "$state/.wake-queue" >/dev/null \ - || fail "the captured result was never published to the durable queue" - - procevent_watch_bg "$dir" "$out" - pid=$! - wait_for_exit "$pid" 100 \ - || fail "a healthy watcher never surfaced a durably captured process-event result: $(cat "$out")" - grep -F "check:" "$out" >/dev/null \ - || fail "the process-event wake was not reported as an actionable check: $(cat "$out")" - grep -F "procevent:delivery-src:1" "$out" >/dev/null \ - || fail "the actionable reason did not name the queued result: $(cat "$out")" - beacon_age=$(FM_STATE_OVERRIDE="$state" bash -c \ - '. "$1/bin/fm-wake-lib.sh"; fm_path_age "$2"' _ "$ROOT" "$state/.last-watcher-beat") - [ "$beacon_age" -lt 60 ] || fail "the surfacing watcher was not a healthy one (beacon age ${beacon_age}s)" - - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the process-event wake failed" - grep "$(printf '\tcheck\t')" "$drain_out" | grep -F "procevent lavish delivery-src 1" >/dev/null \ - || fail "the process-event result was not queued for the drain that follows the wake" - pass "a captured process-event result wakes a healthy watcher proactively, with no manual drain" -} - -test_procevent_unacknowledged_result_redrains_until_handled() { - local dir state out replay_out replay_err pid before after sequence generation - dir=$(make_case procevent-redrain); state="$dir/state" - out="$dir/watch.out"; replay_out="$dir/replay.out"; replay_err="$dir/replay.err" - seed_captured_procevent_result "$dir" || fail "the fixture captured no process-event result" - - procevent_watch_bg "$dir" "$out" - pid=$! - wait_for_exit "$pid" 100 || fail "the first proactive wake never happened: $(cat "$out")" - FM_STATE_OVERRIDE="$state" "$DRAIN" >/dev/null 2>&1 || fail "drain after the first process-event wake failed" - - # An interrupted handler leaves the captured result durable. The successor - # must re-surface it through recovery, then its drain must print the same row. - : > "$out" - procevent_watch_bg "$dir" "$out" - pid=$! - wait_for_exit "$pid" 100 \ - || fail "an unacknowledged process-event result was not re-surfaced on re-arm: $(cat "$out")" - grep -F 'check: rearm-resurface' "$out" >/dev/null \ - || fail "the successor did not report recovery for the unacknowledged result: $(cat "$out")" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$replay_out" 2> "$replay_err" \ - || fail "the successor could not re-drain the unacknowledged process-event result" - grep "$(printf '\tcheck\t')" "$replay_out" | grep -F 'procevent lavish delivery-src 1' >/dev/null \ - || fail "the successor drain did not re-print the durable process-event row" - - pe_case "$dir" handled delivery-src 1 >/dev/null || fail "could not acknowledge the captured result" - sequence=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through \([0-9][0-9]*\) --recovery-generation [A-Za-z0-9._-][A-Za-z0-9._-]*$/\1/p' "$replay_err") - generation=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through [0-9][0-9]* --recovery-generation \([A-Za-z0-9._-][A-Za-z0-9._-]*\)$/\1/p' "$replay_err") - [ -n "$sequence" ] && [ -n "$generation" ] \ - || fail "the replay drain omitted its post-handling acknowledgement boundary" - FM_STATE_OVERRIDE="$state" "$DRAIN" --ack-through "$sequence" --recovery-generation "$generation" \ - || fail "completed process-event handling could not acknowledge the replay" - [ ! -s "$state/.wake-queue" ] || fail "acknowledged process-event replay remained durable" - - before=$(awk 'END { print NR + 0 }' "$state/.wake-queue" 2>/dev/null || echo 0) - : > "$out" - procevent_watch_bg "$dir" "$out" - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - fail "a handled process-event result woke the watcher: $(cat "$out")" - fi - reap "$pid" - after=$(awk 'END { print NR + 0 }' "$state/.wake-queue" 2>/dev/null || echo 0) - [ "$after" = "$before" ] || fail "a handled result was announced again ($before -> $after queued records)" - pass "an unacknowledged process-event result re-drains until handling is acknowledged" -} - -test_procevent_marker_keys_are_injective() { - local dir state out pid marker_count - dir=$(make_case procevent-marker-identity); state="$dir/state"; out="$dir/watch.out" - append_wake "$state" check "procevent:a.b:1" "check: procevent fixture a.b 1" - append_wake "$state" check "procevent:a_b:1" "check: procevent fixture a_b 1" - procevent_watch_bg "$dir" "$out" - pid=$! - wait_for_exit "$pid" 100 || fail "colliding-looking process-event keys were not surfaced" - grep -F "procevent:a.b:1" "$out" >/dev/null || fail "the dotted queue key was suppressed" - grep -F "procevent:a_b:1" "$out" >/dev/null || fail "the underscored queue key was suppressed" - marker_count=$(find "$state" -maxdepth 1 -name '.seen-procevent-*' -type f | awk 'END { print NR + 0 }') - [ "$marker_count" = 2 ] || fail "distinct queue keys produced $marker_count seen markers" - FM_STATE_OVERRIDE="$state" "$DRAIN" >/dev/null 2>&1 || fail "marker identity fixture drain failed" - pass "complete process-event queue keys map to distinct seen markers" -} - -install_marker_mv_fault() { # <dir> - local dir=$1 - REAL_MV=$(command -v mv) - export REAL_MV - cat > "$dir/fakebin/mv" <<'SH' -#!/usr/bin/env bash -dest=${!#} -case "$dest" in - */.seen-procevent-*) - case "${FM_MARKER_MV_MODE:-}" in - pause) - printf '1\n' > "$FM_MARKER_MV_READY" - while [ ! -e "$FM_MARKER_MV_RELEASE" ]; do sleep 0.02; done - ;; - kill-before) kill -KILL "$PPID"; exit 1 ;; - kill-after) "$REAL_MV" "$@" || exit; kill -KILL "$PPID"; exit 1 ;; - fail) exit 1 ;; - esac - ;; -esac -exec "$REAL_MV" "$@" -SH - chmod +x "$dir/fakebin/mv" -} - -test_procevent_surface_serializes_with_drain() { - local dir state out drain_out ready release pid drain_pid - dir=$(make_case procevent-drain-race); state="$dir/state"; out="$dir/watch.out" - drain_out="$dir/drain.out"; ready="$dir/marker-ready"; release="$dir/marker-release" - append_wake "$state" check "procevent:drain-race:1" "check: procevent fixture drain-race 1" - install_marker_mv_fault "$dir" - FM_MARKER_MV_MODE=pause FM_MARKER_MV_READY="$ready" FM_MARKER_MV_RELEASE="$release" \ - procevent_watch_bg "$dir" "$out" - pid=$! - wait_numeric_file "$ready" 100 || fail "the watcher never reached its marker commit boundary" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" & - drain_pid=$! - wait_live "$drain_pid" 10 || fail "a concurrent drain split the surfacing transition" - [ -s "$state/.wake-queue" ] || fail "the concurrent drain consumed the record before marker commit" - touch "$release" - wait "$pid" || fail "the paused watcher did not finish surfacing" - wait "$drain_pid" || fail "the concurrent drain failed after surfacing committed" - grep -F "procevent:drain-race:1" "$drain_out" >/dev/null \ - || fail "the serialized drain lost the process-event record" - pass "queue revalidation, proactive output, and marker commit serialize with drain" -} - -test_procevent_surface_crash_boundaries() { - local dir state out fifo pid reader marker exit_status replay_err sequence generation - dir=$(make_case procevent-output-fail); state="$dir/state"; out="$dir/watch.out"; fifo="$dir/output.fifo" - append_wake "$state" check "procevent:output-fail:1" "check: procevent fixture output-fail 1" - mkfifo "$fifo" - sh -c ': < "$1"' _ "$fifo" & reader=$! - PATH="$dir/fakebin:$PATH" FM_HOME="$dir" FM_PROCEVENT_CLAIM_ROOT="$dir/claims" \ - FM_CREW_STATE_BIN="$dir/fakebin/fm-crew-state.sh" FM_POLL=0.2 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$fifo" & - pid=$! - wait "$reader" || true - wait_for_exit "$pid" 100 - exit_status=$? - [ "$exit_status" -ne 124 ] || fail "the watcher survived a failed actionable output write" - marker=$(find "$state" -maxdepth 1 -name '.seen-procevent-*' -type f | head -1) - [ -z "$marker" ] || fail "failed output committed a suppression marker" - [ -s "$state/.wake-queue" ] || fail "failed output consumed the durable queue record" - procevent_watch_bg "$dir" "$out"; pid=$! - wait_for_exit "$pid" 100 || fail "the record was not replayable after output failure" - grep -F "procevent:output-fail:1" "$out" >/dev/null || fail "output failure lost proactive replay" - - dir=$(make_case procevent-before-marker); state="$dir/state"; out="$dir/watch.out" - append_wake "$state" check "procevent:before-marker:1" "check: procevent fixture before-marker 1" - install_marker_mv_fault "$dir" - FM_MARKER_MV_MODE=kill-before procevent_watch_bg "$dir" "$out"; pid=$! - wait_for_exit "$pid" 100 - exit_status=$? - [ "$exit_status" -ne 124 ] || fail "the watcher survived the injected pre-marker crash" - grep -F "procevent:before-marker:1" "$out" >/dev/null || fail "the pre-marker crash happened before output" - marker=$(find "$state" -maxdepth 1 -name '.seen-procevent-*' -type f | head -1) - [ -z "$marker" ] || fail "a pre-marker crash committed suppression" - procevent_watch_bg "$dir" "$out.replay"; pid=$! - wait_for_exit "$pid" 100 || fail "a pre-marker crash was not replayable" - - dir=$(make_case procevent-after-marker); state="$dir/state"; out="$dir/watch.out" - append_wake "$state" check "procevent:after-marker:1" "check: procevent fixture after-marker 1" - install_marker_mv_fault "$dir" - FM_MARKER_MV_MODE=kill-after procevent_watch_bg "$dir" "$out"; pid=$! - wait_for_exit "$pid" 100 - exit_status=$? - [ "$exit_status" -ne 124 ] || fail "the watcher survived the injected post-marker crash" - grep -F "procevent:after-marker:1" "$out" >/dev/null || fail "the post-marker crash lost actionable output" - marker=$(find "$state" -maxdepth 1 -name '.seen-procevent-*' -type f | head -1) - [ -n "$marker" ] || fail "the post-marker crash did not reach marker commit" - : > "$out.replay" - procevent_watch_bg "$dir" "$out.replay"; pid=$! - wait_for_exit "$pid" 100 \ - || fail "an unacknowledged delivered record was not re-surfaced on re-arm: $(cat "$out.replay")" - grep -F 'check: rearm-resurface' "$out.replay" >/dev/null \ - || fail "the successor did not recover the delivered-but-unacknowledged record: $(cat "$out.replay")" - replay_err="$out.replay.err" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$out.replay.drain" 2> "$replay_err" \ - || fail "post-marker successor drain failed" - grep "$(printf '\tcheck\t')" "$out.replay.drain" | grep -F 'procevent fixture after-marker 1' >/dev/null \ - || fail "post-marker successor did not re-drain the durable record" - sequence=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through \([0-9][0-9]*\) --recovery-generation [A-Za-z0-9._-][A-Za-z0-9._-]*$/\1/p' "$replay_err") - generation=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through [0-9][0-9]* --recovery-generation \([A-Za-z0-9._-][A-Za-z0-9._-]*\)$/\1/p' "$replay_err") - [ -n "$sequence" ] && [ -n "$generation" ] \ - || fail "post-marker replay omitted its post-handling acknowledgement boundary" - FM_STATE_OVERRIDE="$state" "$DRAIN" --ack-through "$sequence" --recovery-generation "$generation" \ - || fail "post-marker replay acknowledgement failed" - [ ! -s "$state/.wake-queue" ] || fail "post-marker acknowledgement left the durable record queued" - pass "surfacing failures replay until post-handling acknowledgement" -} - -test_procevent_marker_failure_exits_and_replays() { - local dir state out pid marker output_count - dir=$(make_case procevent-marker-failure); state="$dir/state"; out="$dir/watch.out" - append_wake "$state" check "procevent:marker-failure:1" "check: procevent fixture marker-failure 1" - install_marker_mv_fault "$dir" - FM_MARKER_MV_MODE=fail procevent_watch_bg "$dir" "$out" - pid=$! - wait_for_exit "$pid" 100 || fail "marker failure did not end the actionable watcher cycle successfully" - output_count=$(grep -Fc "procevent:marker-failure:1" "$out" || true) - [ "$output_count" = 1 ] || fail "marker failure printed the actionable reason $output_count times" - marker=$(find "$state" -maxdepth 1 -name '.seen-procevent-*' -type f | head -1) - [ -z "$marker" ] || fail "marker failure committed suppression" - [ ! -e "$state/.wake-queue.lock" ] && [ ! -L "$state/.wake-queue.lock" ] \ - || fail "marker failure left the queue lock held" - procevent_watch_bg "$dir" "$out.replay" - pid=$! - wait_for_exit "$pid" 100 || fail "marker failure did not leave the durable record replayable" - grep -F "procevent:marker-failure:1" "$out.replay" >/dev/null \ - || fail "marker failure lost the later proactive replay" - FM_STATE_OVERRIDE="$state" "$DRAIN" >/dev/null 2>&1 || fail "marker-failure fixture drain failed" - pass "marker failure exits through the shared wake owner, releases its lock, and replays later" -} - -# --- heartbeat: no-change absorbed, backstop surfaces a missed status -------- - -test_heartbeat_no_change_absorbed() { - local dir state fakebin out pid i sig - dir=$(make_case heartbeat-absorb); state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out" - printf 'working: routine heartbeat history\n' > "$state/routine.status" - sig=$(seen_sig "$state/routine.status"); printf '%s' "$sig" > "$state/.seen-routine_status" - # A quiet fleet with a fast heartbeat cadence. - PATH="$fakebin:$PATH" FM_STATE_OVERRIDE="$state" FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=1 "$WATCH" > "$out" & - pid=$! - if ! wait_poll_cycle "$state" "$pid"; then - reap "$pid"; fail "watcher exited for a no-change heartbeat (should absorb): $(cat "$out")" - fi - # The heartbeat fires on the first poll whose .last-heartbeat has aged past - # FM_HEARTBEAT, which need not be the first completed cycle, so wait for the - # absorbed heartbeat itself rather than assuming one cycle produced it. - i=0 - while [ "$i" -lt 200 ]; do - [ "$(cat "$state/.heartbeat-streak" 2>/dev/null || echo 0)" -ge 1 ] && break - kill -0 "$pid" 2>/dev/null || break - sleep 0.1 - i=$((i + 1)) - done - [ ! -s "$out" ] || fail "no-change heartbeat printed a wake reason: $(cat "$out")" - [ ! -s "$state/.wake-queue" ] || fail "no-change heartbeat enqueued a durable wake record" - [ "$(cat "$state/.heartbeat-streak" 2>/dev/null || echo 0)" -ge 1 ] || fail "heartbeat backoff streak did not advance while absorbing" - [ "$(status_presentation_marker_offset "$state/.hb-surfaced-routine" "$state/routine.status")" = \ - "$(size_of "$state/routine.status")" ] \ - || fail "routine heartbeat classification did not commit its captured endpoint" - reap "$pid" - pass "a heartbeat with no captain-relevant change is absorbed and backs off the cadence" -} - -test_heartbeat_backstop_surfaces_a_masked_status() { - local dir state fakebin out sig pid - dir=$(make_case heartbeat-masked); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out" - # Same miss as below, but the captain-relevant event is followed by a routine - # append, so its last line reads benign. The backstop must still catch it. - printf 'working: setup\nneeds-decision: pick A or B\nworking: tidying the branch\n' \ - > "$state/miss.status" - sig=$(seen_sig "$state/miss.status"); printf '%s' "$sig" > "$state/.seen-miss_status" - PATH="$fakebin:$PATH" FM_STATE_OVERRIDE="$state" FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=1 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 \ - || fail "heartbeat backstop missed a decision hidden behind a later working: line" - grep -Fx "heartbeat" "$out" >/dev/null || fail "backstop did not exit with a heartbeat wake" - [ "$(status_presentation_marker_offset "$state/.hb-surfaced-miss" "$state/miss.status")" = \ - "$(size_of "$state/miss.status")" ] \ - || fail "backstop did not record the masked status as surfaced through its end" - pass "the heartbeat backstop surfaces a captain event hidden behind a later routine append" -} - -test_heartbeat_backstop_surfaces_unsurfaced_status() { - local dir state fakebin out drain_out sig pid - dir=$(make_case heartbeat-backstop); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out" - # A captain-relevant status whose .seen-* signature ALREADY matches (so the - # per-poll signal scan stays quiet) but which was never surfaced (no - # .hb-surfaced-* marker). This stands in for a per-wake-path miss; the heartbeat - # fleet-scan backstop must catch it and wake firstmate. - printf 'done: PR https://example.test/pr/5\n' > "$state/miss.status" - sig=$(seen_sig "$state/miss.status"); printf '%s' "$sig" > "$state/.seen-miss_status" - PATH="$fakebin:$PATH" FM_STATE_OVERRIDE="$state" FM_POLL=1 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=1 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "heartbeat backstop did not surface an unsurfaced captain-relevant status" - grep -Fx "heartbeat" "$out" >/dev/null || fail "backstop did not exit with a heartbeat wake" - [ "$(status_presentation_marker_offset "$state/.hb-surfaced-miss" "$state/miss.status")" = \ - "$(size_of "$state/miss.status")" ] \ - || fail "backstop did not record the status as surfaced through its end (would re-fire next heartbeat)" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the backstop heartbeat failed" - grep "$(printf '\theartbeat\t')" "$drain_out" >/dev/null || fail "backstop heartbeat was not queued" - pass "heartbeat backstop fail-safe surfaces a captain-relevant status the per-wake path missed" -} - -# --- beacon stays fresh while absorbing ------------------------------------- - -test_beacon_stays_fresh_while_absorbing() { - local dir state fakebin out status_file pid m1 m2 now - dir=$(make_case beacon-fresh); state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out" - status_file="$state/task.status" - printf 'working: a\n' > "$status_file" - # Provably working so the working: notes are absorbed (the path that must keep the - # beacon fresh). - export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' - watch_bg "$state" "$fakebin" "$out" - pid=$! - # Wait on the beacon itself rather than a fixed liveness budget: the watcher's - # bounded startup can outlast a short wait, and reading an absent beacon would - # report a missing beacon that simply had not been written yet. - wait_poll_cycle "$state" "$pid" || { reap "$pid"; fail "watcher exited while absorbing the first benign signal"; } - m1=$(file_mtime "$state/.last-watcher-beat") - # A second benign signal keeps it absorbing; the beacon must keep advancing. - printf 'working: b\n' >> "$status_file" - wait_poll_cycle "$state" "$pid" || { reap "$pid"; fail "watcher exited while absorbing a second benign signal"; } - m2=$(file_mtime "$state/.last-watcher-beat") - now=$(date +%s) - if [ -z "$m1" ] || [ -z "$m2" ]; then - reap "$pid" - fail "watcher beacon missing while absorbing" - fi - [ "$m2" -ge "$m1" ] || { reap "$pid"; fail "beacon mtime regressed while absorbing"; } - [ "$(( now - m2 ))" -lt 10 ] || { reap "$pid"; fail "beacon went stale while absorbing (age $(( now - m2 ))s)"; } - [ ! -s "$state/.wake-queue" ] || { reap "$pid"; fail "absorbing benign signals enqueued a wake"; } - reap "$pid" - pass "the liveness beacon stays fresh while the watcher absorbs benign wakes (fm-guard never false-alarms)" -} - -# --- afk coherence: the daemon owns triage; the watcher does not double-triage --- - -test_afk_signal_records_heartbeat_endpoint() { - local dir state fakebin out status_file pid - dir=$(make_case afk-heartbeat-endpoint); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; status_file="$state/task.status" - printf 'needs-decision: choose release target\nworking: preparing both targets\n' > "$status_file" - date '+%s' > "$state/.afk" - export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' - watch_bg "$state" "$fakebin" "$out" - pid=$! - wait_for_exit "$pid" 100 || fail "afk watcher did not hand the actionable signal to the daemon" - [ "$(status_presentation_marker_offset "$state/.hb-surfaced-task" "$status_file")" = \ - "$(size_of "$status_file")" ] \ - || fail "afk signal did not record the endpoint handed to the daemon" - unset FM_FAKE_CREW_STATE - pass "an afk signal records its captured heartbeat endpoint" -} - -test_afk_present_reverts_watcher_to_one_shot() { - local dir state fakebin out drain_out status_file pid - dir=$(make_case afk-coherence); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out" - status_file="$state/task.status" - printf 'working: routine note\n' > "$status_file" - date '+%s' > "$state/.afk" # away mode: the supervise-daemon owns triage - # Set a PROVABLY-WORKING verdict: if afk failed to bypass the provably-working - # check, this no-verb signal would be absorbed (not surfaced). The test asserting - # a surface therefore also proves afk reverts to one-shot and skips the costly read. - export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' - watch_bg "$state" "$fakebin" "$out" - pid=$! - wait_for_exit "$pid" 100 || fail "with .afk present the watcher did not exit one-shot for a benign signal" - grep -F "signal: $status_file" "$out" >/dev/null || fail "afk-mode watcher did not surface the signal for the daemon" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the afk-mode signal failed" - grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$status_file" >/dev/null \ - || fail "afk-mode benign signal was not queued for the daemon to classify" - pass "with .afk present the watcher reverts to one-shot so the daemon owns triage (no double-triage)" -} - -# A paused pane can first appear as a changed hash. In AFK mode that initial path -# must still hand off the plain window identity to the daemon, rather than running -# the normal-mode pause re-surface and decorating the stale identity. -test_afk_paused_changed_pane_hands_off_plain_stale() { - local dir state fakebin out drain_out capture_file statusf window key sig pid back - dir=$(make_case afk-paused-changed-pane); state="$dir/state"; fakebin="$dir/fakebin" - out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" - window="test:fm-afk-held" - printf 'idle, awaiting upstream\n' > "$capture_file" - printf 'window=%s\nkind=ship\n' "$window" > "$state/afk-held.meta" - statusf="$state/afk-held.status" - printf 'paused: awaiting the upstream tool release\n' > "$statusf" - back=$(( $(date +%s) - 500 )) - if [ "$(uname)" = Darwin ]; then touch -mt "$(date -r "$back" '+%Y%m%d%H%M.%S')" "$statusf" - else touch -m -d "@$back" "$statusf"; fi - sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-afk-held_status" - date '+%s' > "$state/.afk" - key=$(printf '%s' "$window" | tr '.:/' '___') - - # Deliberately do not seed .hash-*: this is the changed-pane path that used to - # call handle_paused_stale before AFK's one-shot daemon handoff. - PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ - FM_FAKE_CREW_STATE='state: paused · source: status-log · awaiting the upstream tool release' \ - FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=240 FM_POLL=0.2 FM_SIGNAL_GRACE=1 \ - FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & - pid=$! - wait_for_exit "$pid" 100 || fail "AFK paused changed pane did not hand off a stale wake" - grep -Fx "stale: $window" "$out" >/dev/null || fail "AFK paused stale did not preserve its plain window identity: $(cat "$out")" - grep -F "awaiting external" "$out" >/dev/null && fail "AFK watcher decorated a stale identity instead of handing it to the daemon" - [ ! -e "$state/.paused-$key" ] || fail "AFK watcher recorded normal-mode pause tracking instead of handing off" - FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after AFK paused stale failed" - grep "$(printf '\tstale\t')" "$drain_out" | grep -F "stale: $window" >/dev/null \ - || fail "AFK paused stale was not queued with the plain window identity" - pass "AFK changed paused panes hand off plain stale identities for daemon-owned pause triage" -} +# shellcheck source=tests/watch-triage-helpers.sh +. "$(dirname "${BASH_SOURCE[0]}")/watch-triage-helpers.sh" test_status_span_actionable_classifier test_status_span_survives_a_later_routine_append @@ -5301,7 +14,6 @@ test_malformed_seen_signature_reads_the_whole_log test_stale_is_terminal_classifier test_classifier_primitives test_crew_is_provably_working_classifier -test_status_is_paused_classifier test_crew_absorb_class_classifier test_crew_wedge_progress_classifier test_crew_worktree_written_since_classifier @@ -5349,17 +61,12 @@ test_unreadable_status_reports_once_per_file_state test_permission_recovery_surfaces_preserved_status test_terminal_stale_surfaced test_stale_terminal_status_overridden_by_active_run -test_parked_decision_survives_pane_repaint -test_resolved_decision_can_reopen_identically -test_new_keyed_decision_on_parked_pane_surfaces -test_parked_decision_with_dead_agent_wedge_escalates test_nonterminal_stale_provably_working_absorbed_then_escalated test_progressing_run_holds_the_wedge_escalation test_stranded_run_still_wedge_escalates test_wedged_crew_with_no_run_escalates_unchanged test_dead_agent_escalates_even_while_its_run_progresses test_progressing_run_escalates_anyway_past_the_hold_cap -test_busy_pane_progressing_run_holds_then_escalates_past_the_cap test_declared_wait_with_no_progress_evidence_rechecks_instead_of_wedging test_declared_wait_on_a_progressing_run_holds test_declared_wait_on_a_stranded_run_still_escalates @@ -5367,33 +74,9 @@ test_declared_wait_with_a_dead_agent_still_escalates test_provably_working_absorptions_are_distinguishable_in_the_triage_log test_wedge_escalation_marks_demand_deep_inspection_after_threshold test_wedge_escalation_resets_when_pane_becomes_active -test_busy_pane_below_turn_age_bound_is_absorbed -test_busy_pane_stable_hash_escalates_past_turn_age_bound -test_busy_pane_changing_hash_escalates_past_turn_age_bound -test_busy_pane_turn_end_touch_resets_age -test_busy_pane_repeated_escalation_reaches_demand_deep_inspection -test_busy_pane_default_turn_age_bound_is_3600s -test_busy_declared_pause_is_rechecked_not_wedge_escalated -test_afk_busy_declared_pause_hands_off_plain_stale -test_afk_busy_declared_pause_ticking_pane_hands_off_once test_nonterminal_stale_not_working_surfaced -test_nonterminal_stale_paused_absorbed_then_resurfaced -test_pause_resurface_window_backs_off_and_caps -test_pause_streak_bump_reconciles_a_changed_wait -test_paused_resurface_backs_off_while_wedge_still_escalates -test_exited_declared_pause_is_bounded_but_live_captain_held_surfaces -test_live_declared_pause_is_absorbed_on_the_designed_cadence -test_declared_pause_landing_mid_classification_never_emits_a_bare_stale -test_live_declared_pause_still_recoverable_once_its_agent_dies -test_absorbed_replacement_wait_does_not_inherit_the_old_throttle -test_live_declared_wait_churn_honors_the_resurface_throttle -test_secondmate_paused_resurfaces_in_normal_mode +test_failed_wake_append_does_not_arm_the_captain_hold_throttle test_secondmate_captain_held_resurfaces_in_normal_mode -test_secondmate_nonpaused_stale_remains_suppressed -test_secondmate_unpause_clears_pause_tracking -test_nonterminal_stale_pause_transitions_reclassify_unchanged_hash -test_nonterminal_paused_rechecks_authoritative_state -test_paused_authoritative_working_preserves_wedge_timer test_nonterminal_stale_repairs_missing_or_corrupt_timer test_wedge_escalation_deferred_while_worktree_is_written test_write_deferral_resurfaces_on_the_bounded_cadence @@ -5413,5 +96,4 @@ test_heartbeat_backstop_surfaces_a_masked_status test_beacon_stays_fresh_while_absorbing test_afk_signal_records_heartbeat_endpoint test_afk_present_reverts_watcher_to_one_shot -test_afk_paused_changed_pane_hands_off_plain_stale printf '\nall fm-watch-triage tests passed\n' diff --git a/tests/lib.sh b/tests/lib.sh index e211b8b2089..e62c0369d3e 100644 --- a/tests/lib.sh +++ b/tests/lib.sh @@ -42,6 +42,11 @@ umask 022 # strips this to verify real refusal. export FM_GATE_REFUSE_BYPASS=1 +# Clear the task-worker marker bin/fm-spawn.sh exports into ship and scout +# panes. This suite builds git-init fixture repositories whose primary checkout +# it runs a copied bin/fm-test-run.sh in, and that runner refuses the primary +# under the marker. A case that verifies the refusal sets FM_TASK_ID itself. +unset FM_TASK_ID # Isolate the dashboard's event instrumentation and its store from the # developer's own host. Both live OUTSIDE any FM_HOME by design, so neither is # covered by the FM_*_OVERRIDE isolation every other suite relies on. @@ -154,6 +159,56 @@ FM_TEST_OWNER_IDENTITY=$(fm_test_pid_identity "$$") || { return 1 } +# --- process-event runner reaping ------------------------------------------- +# +# A process-event runner is detached into its own process group and reparents to +# init, so removing a fixture directory does not stop one: only sweeping the home +# that owns it does. Registration goes through a `$$`-keyed registry file for the +# same reason the temp roots do - a fixture home is almost always built inside a +# command substitution (`home=$(make_home x)`), and an array append there never +# reaches the caller, so a suite that tracked its homes in a shell array was +# silently tracking nothing and left every runner it started behind. +# +# The sweep is scoped to the exact home (and its claim root when the suite uses a +# private one). It never matches on a script or process name, which would reach +# into another home's live runners. + +FM_TEST_PROCEVENT_REGISTRY=$(mktemp "${TMPDIR:-/tmp}/.fm-test-procevent.$$.XXXXXX") || return 1 + +fm_test_track_procevent_home() { # <home> [claim-root] + [ -n "${1:-}" ] || return 1 + printf '%s\t%s\n' "$1" "${2-}" >> "$FM_TEST_PROCEVENT_REGISTRY" +} + +fm_test_reap_procevent_homes() { + local home claim_root seen=$'\n' + [ -f "$FM_TEST_PROCEVENT_REGISTRY" ] || return 0 + while IFS=$'\t' read -r home claim_root; do + [ -n "$home" ] || continue + case "$seen" in *$'\n'"$home"$'\n'*) continue ;; esac + seen+="$home"$'\n' + [ -d "$home/state/procevent" ] || continue + if [ -n "$claim_root" ]; then + FM_HOME="$home" FM_STATE_OVERRIDE="$home/state" FM_PROCEVENT_CLAIM_ROOT="$claim_root" \ + "$ROOT/bin/fm-procevent.sh" sweep-home >/dev/null 2>&1 || true + else + FM_HOME="$home" FM_STATE_OVERRIDE="$home/state" \ + "$ROOT/bin/fm-procevent.sh" sweep-home >/dev/null 2>&1 || true + fi + done < "$FM_TEST_PROCEVENT_REGISTRY" + rm -f "$FM_TEST_PROCEVENT_REGISTRY" +} + +# Ceiling on how long a fixture's blocking stub may keep polling. A stub that +# waits for a trigger file by re-running `sleep` is a high-frequency source of +# process spawns, and one that outlives its test - because the test was killed +# before any cleanup ran - is what turned leftover fixtures into a host-wide +# process storm. Every blocking stub this suite writes stops itself at this +# bound, so an escaped one is bounded in duration and cost on its own, before +# its owner's guard reaps it. +FM_TEST_STUB_MAX_BLOCK_SECONDS=${FM_TEST_STUB_MAX_BLOCK_SECONDS:-120} +export FM_TEST_STUB_MAX_BLOCK_SECONDS + fm_test_cleanup() { # Ordered before the directory removal below on purpose: a daemon left alive # can write its state back into a tree that is being unlinked, which fails the @@ -162,6 +217,7 @@ fm_test_cleanup() { fm_test_reap_descendants local d + fm_test_reap_procevent_homes for d in "${FM_TEST_CLEANUP_DIRS[@]:-}"; do # A case that hardened a directory to prove a read-only path leaves a tree # rm cannot unlink, and an aborted case never gets to restore it. Write @@ -324,6 +380,8 @@ fm_assert_no_user_event_store_leak() { # <snapshot-taken-before-the-suite> trap fm_test_cleanup EXIT trap 'fm_test_cleanup; exit 130' INT trap 'fm_test_cleanup; exit 143' TERM +trap 'fm_test_cleanup; exit 129' HUP +trap 'fm_test_cleanup; exit 131' QUIT # fm_test_reap_orphans: best-effort sweep for fixture roots left behind by a # prior run that was killed hard enough to skip the traps above (e.g. a @@ -367,6 +425,101 @@ if [ "${FM_TEST_SKIP_ORPHAN_REAP:-0}" != 1 ]; then fm_test_reap_orphans fi +# --- live-capability gate --------------------------------------------------- +# +# fm_live_gate <policy> <vars> [tool ...] +# +# The single gate every live-harness guard opens with, so "can this host run +# this guard for real, and should it?" is decided in one place instead of in +# two dozen hand-rolled env checks. It returns 0 when the guard should run, and +# otherwise ends the script with one runner-readable line: +# +# skip: live: <tool> absent this host cannot run the guard +# skip: live: disabled by <VAR>=0 an explicit local opt-out +# skip: live: opt-in; set <VAR>=1 to run a guard that spends model tokens +# +# <policy> is default-on for a guard that spends no model tokens, so it runs +# wherever its tools are installed - notably on the machine the product and its +# validation actually run on - and opt-in for a guard that submits prompts, +# which stays deliberate. <vars> is the guard's own control variable, or a +# comma-separated list when a guard has more than one entry point. +# +# Setting any of those variables to 1 (or FM_LIVE=1, for every guard at once) +# both turns the guard on and makes an absent tool a hard failure rather than a +# skip, which is how "run it after a harness upgrade" keeps proving the guard +# actually ran. Setting one to 0 (or FM_LIVE=0) turns it off; a guard's own +# variable wins over FM_LIVE. +# +# Sourcing this library also exports FM_GATE_REFUSE_BYPASS=1, which is what +# lets a live guard drive the real fm-spawn/fm-send/fm-teardown from inside a +# no-mistakes gate worktree instead of being refused by +# bin/fm-gate-refuse-lib.sh. + +fm_live_gate() { + local policy=$1 vars=$2 + shift 2 + local var value rest primary requested=0 disabled_by='' tool + local -a var_list=() + + case "$policy" in + default-on | opt-in) ;; + *) fail "fm_live_gate: unknown policy '$policy' (expected default-on or opt-in)" ;; + esac + + rest=$vars + while [ -n "$rest" ]; do + var=${rest%%,*} + if [ "$var" = "$rest" ]; then + rest='' + else + rest=${rest#*,} + fi + [ -n "$var" ] && var_list+=("$var") + done + [ "${#var_list[@]}" -gt 0 ] || fail "fm_live_gate: at least one control variable is required" + primary=${var_list[0]} + + for var in "${var_list[@]}"; do + value=${!var:-} + case "$value" in + 1) requested=1 ;; + 0) [ -n "$disabled_by" ] || disabled_by=$var ;; + esac + done + + if [ "$requested" -eq 0 ]; then + if [ -n "$disabled_by" ]; then + printf 'skip: live: disabled by %s=0\n' "$disabled_by" + exit 0 + fi + case "${FM_LIVE:-}" in + 0) + printf 'skip: live: disabled by FM_LIVE=0\n' + exit 0 + ;; + 1) requested=1 ;; + *) + if [ "$policy" = opt-in ]; then + printf 'skip: live: opt-in; set %s=1 to run\n' "$primary" + exit 0 + fi + ;; + esac + fi + + for tool in "$@"; do + command -v "$tool" >/dev/null 2>&1 && continue + if [ "$requested" -eq 1 ]; then + printf 'not ok - %s was requested but %s is not installed\n' "$primary" "$tool" >&2 + exit 1 + fi + printf 'skip: live: %s absent\n' "$tool" + exit 0 + done + + return 0 +} + # --- fakebin / PATH shims --------------------------------------------------- # # fm_fakebin <dir> creates <dir>/fakebin and echoes it; prepend it to PATH to @@ -583,3 +736,43 @@ assert_absent() { assert_present() { [ -e "$1" ] || fail "$2" } + +# fm_test_base_path_sans <base_path> <tool...>: returns the path to a single +# curated directory that resolves every tool <base_path> would have resolved, +# except the named ones. Some hosts have real system binaries (node, orca, +# ...) sitting in BASE_PATH; a fixture that simulates a tool as missing by +# omitting it from fakebin still falls through to that host binary via +# BASE_PATH, silently defeating the simulation. Dropping whole directories +# out of BASE_PATH is not a safe fix: on a usr-merged host /bin, /sbin, and +# /usr/sbin are symlinks that collapse to the same directory as /usr/bin, so +# dropping any one of them because it resolves the excluded tool drops every +# other tool a test still needs (git, awk, sed, ...) too. Building a curated +# directory instead hides only the named tool(s). Use only at the specific +# assertions that simulate a tool as absent - every other case keeps using +# bare BASE_PATH. +fm_test_base_path_sans() { + local base_path=$1 dir src entry name tool skip + shift + local tools=("$@") + dir=$(fm_test_tmproot fm-base-path-sans) || return 1 + local dirs + IFS=: read -ra dirs <<< "$base_path" + for src in "${dirs[@]}"; do + [ -d "$src" ] || continue + for entry in "$src"/*; do + [ -e "$entry" ] || [ -L "$entry" ] || continue + name=${entry##*/} + [ -e "$dir/$name" ] && continue + skip=0 + for tool in "${tools[@]}"; do + if [ "$name" = "$tool" ]; then + skip=1 + break + fi + done + [ "$skip" -eq 1 ] && continue + ln -s "$entry" "$dir/$name" 2>/dev/null || true + done + done + printf '%s\n' "$dir" +} diff --git a/tests/watch-triage-helpers.sh b/tests/watch-triage-helpers.sh new file mode 100644 index 00000000000..74f42743e87 --- /dev/null +++ b/tests/watch-triage-helpers.sh @@ -0,0 +1,5590 @@ +#!/usr/bin/env bash +# tests/watch-triage-helpers.sh - the always-on wake triage built into +# bin/fm-watch.sh and the shared classifier (bin/fm-classify-lib.sh). The watcher +# now absorbs the benign majority of wakes in bash and exits ONLY on an actionable +# wake, so firstmate's LLM re-arms once per actionable event instead of once per +# wake. These tests cover the classifier predicates as pure functions, then drive +# a real fm-watch.sh subprocess to assert the behavioral contract: +# provably-working no-verb wakes absorbed (no exit, no queue entry, suppressor +# advanced, beacon fresh), stopped-crew no-verb wakes surfaced (queue + exit), +# provably-working stale panes absorbed-then-escalated past the threshold, +# terminal-looking stale status lines overridden by an active run, the heartbeat +# backstop fail-safe, and afk coherence (no double-triage while the away-mode +# daemon owns supervision). +# +# Daemon-side classification/injection lives in fm-daemon.test.sh; watcher/lock +# liveness in fm-watcher-lock.test.sh; the durable-queue safety matrix in +# fm-wake-queue.test.sh. +set -u + +# shellcheck source=tests/wake-helpers.sh +. "$(dirname "${BASH_SOURCE[0]}")/wake-helpers.sh" +# shellcheck source=/dev/null +. "$ROOT/bin/fm-classify-lib.sh" + +WATCH="$ROOT/bin/fm-watch.sh" +DRAIN="$ROOT/bin/fm-wake-drain.sh" + +TMP_ROOT=$(fm_test_tmproot fm-watch-triage-tests) + +ack_stopped_cycle() { # <state> + local state=$1 err sequence generation + err="$state/.test-cycle-drain.err" + FM_STATE_OVERRIDE="$state" "$DRAIN" >/dev/null 2> "$err" || return 1 + sequence=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through \([0-9][0-9]*\) --recovery-generation [A-Za-z0-9._-][A-Za-z0-9._-]*$/\1/p' "$err") + generation=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through [0-9][0-9]* --recovery-generation \([A-Za-z0-9._-][A-Za-z0-9._-]*\)$/\1/p' "$err") + rm -f "$err" + [ -n "$sequence" ] && [ -n "$generation" ] || return 1 + FM_STATE_OVERRIDE="$state" "$DRAIN" --ack-through "$sequence" \ + --recovery-generation "$generation" +} + +# Common watcher knobs: tight poll/grace, no check or heartbeat cadence unless a +# test overrides them, so a test only exercises the path it targets. FM_CREW_STATE_BIN +# points at the case's hermetic fake fm-crew-state.sh (installed by make_case) so the +# absorb-only-when-provably-working triage reads a canned verdict; a test fixes that +# verdict via FM_FAKE_CREW_STATE in its environment before calling watch_bg. +watch_bg() { # <state> <fakebin> <out> [extra env assignments...] + local state=$1 fakebin=$2 out=$3 + shift 3 + PATH="$fakebin:$PATH" FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$@" "$WATCH" > "$out" & +} + +# Wait up to <limit> 0.1s ticks while <pid> stays alive; 0 if still alive, 1 if it died. +wait_live() { + local pid=$1 limit=${2:-30} i=0 + while [ "$i" -lt "$limit" ]; do + kill -0 "$pid" 2>/dev/null || return 1 + sleep 0.1 + i=$((i + 1)) + done + return 0 +} + +# Wait until <pid>'s watcher has completed a whole poll cycle, or exited first. +# A fixed wait_live budget only proves the process is still ALIVE: fm-watch.sh +# does bounded startup work (the recovery-marker snapshot, lock acquisition) +# before its first stale scan, so on a loaded +# machine a short fixed budget can reap a round before the cycle it asserts on +# ever ran - and then every "no wake, no marker" assertion passes vacuously +# while every "marker written" assertion fails spuriously. +# The liveness beacon is touched at the TOP of every poll, so this drops any +# beacon left by an earlier round, waits for THIS watcher to write a fresh one +# (some poll's top), then waits for that one to advance (the next poll's top) - +# and the whole cycle in between is what the caller's assertions describe. +# 0 if the watcher is still alive after a completed cycle, 1 if it exited. +wait_poll_cycle() { # <state> <pid> [limit-ticks] + local state=$1 pid=$2 limit=${3:-300} beat first now i=0 + beat="$state/.last-watcher-beat" + rm -f "$beat" + first="" + while [ "$i" -lt "$limit" ]; do + kill -0 "$pid" 2>/dev/null || return 1 + first=$(file_mtime "$beat") + [ -n "$first" ] && break + sleep 0.1 + i=$((i + 1)) + done + while [ "$i" -lt "$limit" ]; do + kill -0 "$pid" 2>/dev/null || return 1 + now=$(file_mtime "$beat") + if [ -n "$now" ] && [ "$now" != "$first" ]; then + return 0 + fi + sleep 0.1 + i=$((i + 1)) + done + return 1 +} + +# Every wait_for_exit budget in this file is 100 ticks (10s), not because any +# watcher takes that long to decide, but because fm-watch.sh does bounded +# startup work before its first poll: a tighter budget reaps the process while +# it is still starting and reports a spurious "did not surface" failure. A +# generous budget can only remove that false negative - a watcher that never +# exits still fails the assertion when the budget runs out. +wait_numeric_file() { + local file=$1 limit=${2:-30} i=0 value + while [ "$i" -lt "$limit" ]; do + value=$(cat "$file" 2>/dev/null || true) + case "$value" in + ''|*[!0-9]*) ;; + *) return 0 ;; + esac + sleep 0.1 + i=$((i + 1)) + done + return 1 +} + +# Portable mtime in epoch seconds. Platform-detected, never the `stat -f || stat -c` +# fallback (which writes a partial filesystem dump on Linux; see fm-watch.sh). +file_mtime() { + if [ "$(uname)" = Darwin ]; then stat -f %m "$1" 2>/dev/null; else stat -c %Y "$1" 2>/dev/null; fi +} + +# Set <file>'s mtime to exactly <epoch> seconds, for aging a busy-turn marker by +# a precise amount (touch -t takes a local-time stamp, not an epoch, on both +# platforms, so convert via BSD `date -r` or GNU `date -d @`). +set_mtime() { # <epoch> <file> + local epoch=$1 f=$2 stamp + if stamp=$(date -r "$epoch" +%Y%m%d%H%M.%S 2>/dev/null); then + touch -t "$stamp" "$f" + else + stamp=$(date -d "@$epoch" +%Y%m%d%H%M.%S) + touch -t "$stamp" "$f" + fi +} + +# Signature a primed .seen-* marker must hold so the per-poll signal scan does not +# fire on a pre-existing status (mirrors fm-watch.sh's stat_sig exactly). +seen_sig() { + local reported size ident + case "$1" in + *.status) + reported=$(status_observed_signature "$1") + size=$(size_of "$1") + ident=$(_fm_open_decisions_file_ident "$1") + printf 'v2\t%s\t%s@%s' "$reported" "$size" "$ident" + ;; + *) + if [ "$(uname)" = Darwin ]; then stat -f '%z:%Fm' "$1" 2>/dev/null; else stat -c '%s:%Y' "$1" 2>/dev/null; fi + ;; + esac +} + +# Prime <file>'s .seen-* suppressor to its CURRENT signature, so the per-poll +# no-verb signal scan (which watches every *.turn-ended for a size:mtime change) +# treats a just-created or just-backdated turn-ended marker as already seen. +# Busy-turn-age fixtures create/backdate turn-ended directly (there is no real +# harness touching it), so without this the marker's own first sighting would +# fire an unrelated "signal:" wake and mask the busy-turn-age assertion under +# test. Call again after any further touch/set_mtime on the same file. +prime_turnend_seen() { # <file> + local f=$1 base + base=$(basename "$f" | tr '.' '_') + printf '%s' "$(seen_sig "$f")" > "$(dirname "$f")/.seen-$base" +} + +record_pi_busy() { # <state-dir> <id> + local state=$1 id=$2 gen + gen=$("$ROOT/bin/fm-busy-event.sh" arm "$state" "$id") + "$ROOT/bin/fm-busy-event.sh" apply "$state" "$id" busy --gen "$gen" \ + --source pi-ext --event agent-start +} + +reap() { kill "$1" 2>/dev/null || true; wait "$1" 2>/dev/null || true; } + +# --- pure classifier predicates (fm-classify-lib.sh) ------------------------ + +size_of() { LC_ALL=C wc -c < "$1" | tr -d '[:space:]'; } + +test_status_span_actionable_classifier() { + local dir state offset + dir=$(make_case classify-signal); state="$dir/state" + printf 'working: step 1\nworking: step 2\n' > "$state/a.status" + status_span_has_actionable "$state/a.status" 0 && fail "benign working: span classified actionable" + printf 'working: x\nneeds-decision: pick A or B\n' > "$state/b.status" + status_span_has_actionable "$state/b.status" 0 || fail "captain-relevant span classified benign" + # A failure and a merge result are captain-relevant and must always wake. + printf 'failed: build broke on main\n' > "$state/d.status" + status_span_has_actionable "$state/d.status" 0 || fail "a failed: line was not actionable" + printf 'merged\n' > "$state/e.status" + status_span_has_actionable "$state/e.status" 0 || fail "a legacy merged line was not actionable" + # An offset past the whole log has nothing left to classify: an event already + # classified must not re-fire on the next append. + offset=$(size_of "$state/b.status") + status_span_has_actionable "$state/b.status" "$offset" \ + && fail "an already-classified needs-decision re-fired from its own end offset" + printf 'working: tidying up\n' >> "$state/b.status" + status_span_has_actionable "$state/b.status" "$offset" \ + && fail "a routine append after a classified decision was classified actionable" + # An unusable offset (absent, malformed, or past a truncated log) reads the + # whole file rather than losing the events it cannot account for. + status_span_has_actionable "$state/b.status" "" || fail "an empty offset did not read the whole log" + status_span_has_actionable "$state/b.status" "not-a-number" || fail "a malformed offset did not read the whole log" + status_span_has_actionable "$state/b.status" 99999 || fail "an offset past the log did not read the whole log" + pass "status_span_has_actionable: benign absorbed, captain events surfaced, classified events not re-fired" +} + +# The reported bug, at the classifier: an actionable event followed by a ROUTINE +# append must stay actionable, and must be reported as ITSELF rather than as the +# routine line that happens to sit last. +test_status_span_survives_a_later_routine_append() { + local dir state event + dir=$(make_case classify-masked); state="$dir/state" + printf 'working: setup\nneeds-decision: pick A or B\nworking: still tidying the branch\n' \ + > "$state/mask.status" + status_span_has_actionable "$state/mask.status" 0 \ + || fail "a needs-decision hidden behind a later working: line was classified routine" + event=$(status_span_first_actionable "$state/mask.status" 0) + [ "$event" = "needs-decision: pick A or B" ] \ + || fail "the span reported '$event' instead of the decision it found" + # The captain-reported shape: a finished release/install reported as done and + # then followed by routine cleanup chatter must still reach the captain. + printf 'working: publishing\ndone: release 1.4.0 published and installed\nworking: cleaning the build dir\nnote: cache pruned\n' \ + > "$state/release.status" + status_span_has_actionable "$state/release.status" 0 \ + || fail "a done: completion hidden behind later routine appends was classified routine" + event=$(status_span_first_actionable "$state/release.status" 0) + [ "$event" = "done: release 1.4.0 published and installed" ] \ + || fail "the span reported '$event' instead of the completion it found" + # A blocker is the away-mode shape of the same masking. + printf 'blocked: cannot reach the release host\npaused: waiting for release access\n' \ + > "$state/blocked.status" + status_span_has_actionable "$state/blocked.status" 0 \ + || fail "a blocked: event hidden behind a current wait was classified routine" + pass "an actionable event is not hidden by later routine appends, and is named as itself" +} + +# Closure is the one thing that may retire an event inside a span, and only +# through status_open_decisions' own open/closed rule. +test_status_span_respects_decision_closure() { + local dir state event open + dir=$(make_case classify-closure); state="$dir/state" + printf 'needs-decision [key=api]: pick A or B\nresolved [key=api]: took A\n' > "$state/closed.status" + status_span_has_actionable "$state/closed.status" 0 \ + && fail "a decision the same span already closed was still classified actionable" + # Reopening the SAME key after a close must survive: the close belongs to the + # earlier opening, not to the one that came after it. + printf 'needs-decision [key=api]: pick A or B\nresolved [key=api]: took A\nneeds-decision: [key=api] pick A or B\n' \ + > "$state/reopened.status" + event=$(status_span_first_actionable "$state/reopened.status" 0) \ + || fail "a decision reopened under a key that was closed earlier was classified routine" + [ "$event" = "needs-decision: [key=api] pick A or B" ] \ + || fail "the reopened key surfaced its closed opening instead of the live reopening: $event" + # A terminal event is never retired by a later closure line. + printf 'failed: build broke on main\nresolved [key=api]: unrelated\n' > "$state/term.status" + status_span_has_actionable "$state/term.status" 0 \ + || fail "a failed: event was retired by an unrelated closure" + # A live decision must survive a NEWER closure that belongs to another key. + printf 'needs-decision [key=api]: pick A or B\nneeds-decision [key=db]: pick a store\nresolved [key=db]: took sqlite\n' \ + > "$state/two.status" + event=$(status_span_first_actionable "$state/two.status" 0) \ + || fail "a still-open decision was retired by a newer closure under another key" + [ "$event" = "needs-decision [key=api]: pick A or B" ] \ + || fail "the span reported '$event' instead of the decision still open" + printf 'needs-decision [key=pending-reply-x]: unrelated request\nworking: awaiting reconciliation\n' \ + > "$state/rejected-reserved.status" + event=$(status_span_first_actionable "$state/rejected-reserved.status" 0) \ + || fail "a rejected reserved-key request was silently dropped" + [ "$event" = "reconciliation-required: needs-decision [key=pending-reply-x]: unrelated request" ] \ + || fail "a rejected reserved-key request was not labeled for reconciliation: $event" + open=$(status_open_decisions "$state/rejected-reserved.status") + [ -z "$open" ] \ + || fail "span classification treated a rejected reserved-key request as an open decision: $open" + pass "span classification retires closed decisions and surfaces rejected transitions for reconciliation" +} + +test_malformed_seen_signature_reads_the_whole_log() { + local dir state f marker offset + dir=$(make_case malformed-seen); state="$dir/state"; f="$state/task.status" + printf 'needs-decision: choose the release target\nworking: cleanup\n' > "$f" + marker="$state/.seen-task_status" + printf '40' > "$marker" + offset=$(bash -c '. "$1"; fm_wake_signal_seen_size "$2" "$3"' _ \ + "$ROOT/bin/fm-wake-lib.sh" "$state" "$f") + [ "$offset" = 0 ] \ + || fail "a digits-only malformed seen signature was accepted as an offset" + status_span_has_actionable "$f" "$offset" \ + || fail "a malformed seen signature skipped the actionable start of the log" + pass "a malformed seen signature causes the whole status log to be classified" +} + +test_stale_is_terminal_classifier() { + local dir state + dir=$(make_case classify-stale); state="$dir/state" + printf 'done: ready in branch fm/x\n' > "$state/term.status" + stale_is_terminal "sess:fm-term" "$state" || fail "terminal stale status not classified terminal" + fm_write_meta "$state/herdr-term.meta" "window=default:w1:p2" "backend=herdr" + printf 'done: ready in branch fm/herdr\n' > "$state/herdr-term.status" + stale_is_terminal "default:w1:p2" "$state" || fail "terminal herdr stale status not resolved through metadata" + printf 'working: compiling\n' > "$state/nonterm.status" + stale_is_terminal "sess:fm-nonterm" "$state" && fail "non-terminal stale classified terminal" + stale_is_terminal "sess:fm-missing" "$state" && fail "stale with no status classified terminal" + pass "stale_is_terminal: terminal status surfaces, non-terminal and no-status are benign" +} + +test_classifier_primitives() { + local dir state open activity + dir=$(make_case classify-primitives); state="$dir/state" + printf 'working: a\n\ndone: b\n\n' > "$state/x.status" + [ "$(last_status_line "$state/x.status")" = "done: b" ] || fail "last_status_line did not return the last non-blank line" + status_is_captain_relevant "done: b" || fail "done: not recognized as captain-relevant" + status_is_captain_relevant "needs-decision [key=q1]: b" || fail "keyed needs-decision not recognized as captain-relevant" + status_is_captain_relevant "working: b" && fail "working: wrongly recognized as captain-relevant" + # Incident regression: free-text "merged" inside a nonterminal working: line must + # not become captain-relevant (AFK false-terminal path). + status_is_captain_relevant \ + "working: stage 2 setup complete on PR #74 exact source branch rebased onto merged #76; task dates preserved" \ + && fail "working: ... merged #N wrongly recognized as captain-relevant" + status_is_captain_relevant "working: rebased onto predecessor #76" \ + && fail "working: predecessor prose wrongly recognized as captain-relevant" + status_is_captain_relevant "working: PR ready checks green merged ready in branch" \ + && fail "working: free-text tokens wrongly recognized as captain-relevant" + status_is_captain_relevant "done: PR https://x/pull/76 checks green" \ + || fail "genuine done: checks green not captain-relevant" + status_is_terminal_verb "done: PR https://x/pull/76 checks green" \ + || fail "done: not a terminal verb" + status_is_terminal_verb "working: rebased onto merged #76" \ + && fail "working: wrongly classed as terminal verb" + status_is_captain_relevant "merged" || fail "legacy bare merged free-text not captain-relevant" + status_is_captain_relevant "PR ready https://x/pull/2" \ + || fail "legacy bare PR ready free-text not captain-relevant" + [ "$(window_to_task "sess:fm-fix-login-k3")" = "fix-login-k3" ] || fail "window_to_task did not strip session+fm- prefix" + fm_write_meta "$state/herdr-task.meta" "window=default:w1:p2" "backend=herdr" + [ "$(window_to_task "default:w1:p2" "$state")" = "herdr-task" ] || fail "window_to_task did not resolve opaque backend target through metadata" + FM_CAPTAIN_RE='custom-verb:' status_is_captain_relevant "custom-verb: x" || fail "FM_CAPTAIN_RE override not honored" + FM_CAPTAIN_RE='custom-verb:' status_is_captain_relevant "done: x" && fail "FM_CAPTAIN_RE override did not replace the default verb set" + FM_CAPTAIN_RE='merged|custom-verb:' status_is_captain_relevant "working: rebased onto merged #76" \ + && fail "FM_CAPTAIN_RE override bypassed working: suppression" + FM_CAPTAIN_RE='checks green|custom-verb:' status_is_captain_relevant "paused: checks green pending approval" \ + && fail "FM_CAPTAIN_RE override bypassed paused: suppression" + FM_CAPTAIN_RE='custom-verb:' status_is_captain_relevant "custom-verb: x" \ + || fail "nonterminal suppression weakened custom bare-line behavior" + printf 'needs-decision: should docs mention [key=prose]?\nneeds-decision [key=q1]: real choice\nresolved: docs still mention [key=q1]\nneeds-decision [key=bad key]: malformed\n' > "$state/keys.status" + open=$(status_open_decisions "$state/keys.status") + printf '%s' "$open" | grep -F $'q1\t' >/dev/null \ + || fail "a key token in resolved note prose closed the keyed decision" + printf '%s' "$open" | grep -F $'prose\t' >/dev/null \ + && fail "a key token in note prose changed the decision key" + printf '%s' "$open" | grep -F $'bad key\t' >/dev/null \ + && fail "an invalid key slug entered the open-decision set" + cat > "$state/activity.status" <<'EOF' +working [key=phase7]: Phase 7 started +working [key=phase6]: Phase 6 started +working [key=legal]: reviewing legal dependency +done [key=phase6]: Phase 6 completed +resolved [key=phase7]: Phase 7 completed and moved to Done +paused [key=legal]: awaiting external counsel +resolved [key=legal]: legal item returned to the queue +working [key=phase8]: Phase 8 started +EOF + activity=$(status_open_activities "$state/activity.status") + printf '%s' "$activity" | grep -F $'phase8\tworking\tPhase 8 started' >/dev/null \ + || fail "the current keyed working phase was not retained" + printf '%s' "$activity" | grep -F $'phase7\t' >/dev/null \ + && fail "a keyed resolved event did not close the older working phase" + printf '%s' "$activity" | grep -F $'phase6\t' >/dev/null \ + && fail "a same-key terminal event did not supersede the older working phase" + printf '%s' "$activity" | grep -F $'legal\t' >/dev/null \ + && fail "a keyed resolved event did not close the declared pause" + printf 'working: legacy start\ndone: legacy completion\n' > "$state/legacy-activity.status" + [ -z "$(status_open_activities "$state/legacy-activity.status")" ] \ + || fail "a legacy terminal event did not supersede the default working phase" + pass "classifier primitives: keyed decisions and activity phases, captain relevance, window-to-task, and overrides" +} + +# crew_is_provably_working: the absorb-only-when-provably-working predicate. It is +# benign (absorb) ONLY when fm-crew-state.sh reports the crew as working from a +# current or bounded-degraded pipeline step (source run-step or +# run-step-degraded) or a busy pane (source pane); +# everything else - a stale working: status-log line, a finished/parked/failed run, +# an unknown/torn-down crew, or an empty id - is NOT provable, so it surfaces. The +# fake fm-crew-state.sh (FM_CREW_STATE_BIN) returns a canned verdict per case. +test_crew_is_provably_working_classifier() { + local dir fakebin + dir=$(make_case provably-working); fakebin="$dir/fakebin" + # Point the predicate at this case's hermetic fake and drive its verdict per case. + # export marks the var for the fake subprocess; it is unset again at the end so it + # cannot leak into a later test (every behavioral test sets its own verdict anyway). + export FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" + export FM_FAKE_CREW_STATE + FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' + crew_is_provably_working a || fail "active run-step not treated as provably working" + FM_FAKE_CREW_STATE='state: working · source: pane · harness busy' + crew_is_provably_working a || fail "busy pane not treated as provably working" + FM_FAKE_CREW_STATE='state: working · source: status-log · working: compiling' + ! crew_is_provably_working a || fail "stale status-log working: treated as provably working" + FM_FAKE_CREW_STATE='state: done · source: run-step · checks green' + ! crew_is_provably_working a || fail "finished run treated as provably working" + FM_FAKE_CREW_STATE='state: parked · source: run-step · parked at review' + ! crew_is_provably_working a || fail "parked run treated as provably working" + FM_FAKE_CREW_STATE='state: failed · source: run-step · run failed' + ! crew_is_provably_working a || fail "failed run treated as provably working" + FM_FAKE_CREW_STATE='state: abandoned · source: run-step · worker gone (dead)' + ! crew_is_provably_working a || fail "abandoned run treated as provably working" + FM_FAKE_CREW_STATE='state: unknown · source: none · worktree gone' + ! crew_is_provably_working a || fail "unknown crew treated as provably working" + FM_FAKE_CREW_STATE='state: working · source: run-step · x' + ! crew_is_provably_working "" || fail "empty id treated as provably working" + unset FM_FAKE_CREW_STATE + pass "crew_is_provably_working: only working+run-step/pane is provable; idle/finished/parked/failed/unknown surface" +} + +# status_is_paused: the shared pause verb test both consumers read (so neither +# hardcodes the literal). Matches only the verb before the first colon, so a reason +# that merely mentions "paused" does not false-match, and a genuine blocker stays a +# blocker. +test_status_is_paused_classifier() { + status_is_paused 'paused: holding for the upstream release' || fail "paused verb not recognized" + status_is_paused ' paused: waiting on a rate-limit reset' || fail "leading-space paused verb not recognized" + status_is_paused 'blocked: the build is paused upstream' && fail "a blocked line mentioning paused false-matched" + status_is_paused 'working: paused the animation loop' && fail "a working line mentioning paused false-matched" + status_is_paused 'done: shipped' && fail "done classified as paused" + status_is_paused '' && fail "empty line classified as paused" + # A pause is deliberately NOT captain-relevant: it is a stop-nagging signal, not + # work to keep surfacing. + status_is_captain_relevant 'paused: holding for the upstream release' && fail "paused is captain-relevant (should not be)" + status_is_paused_or_captain_held 'paused: holding for the upstream release' \ + || fail "declared pause not recognized by the bounded-idle classifier" + status_is_paused_or_captain_held 'captain-held [key=route]: tracked by task-decision-route' \ + || fail "captain-held transfer not recognized by the bounded-idle classifier" + status_is_paused_or_captain_held 'resolved [key=route]: captain answered' \ + && fail "resolved decision remained classed as captain-held" + # The two declarations share one cadence but block on different humans, so the + # combined predicate cannot be the only discriminator: a recheck has to know which + # verb it is naming. + status_is_captain_held 'captain-held [key=route]: tracked by task-decision-route' \ + || fail "captain-held verb not recognized" + status_is_captain_held 'paused: holding for the upstream release' \ + && fail "a declared pause matched the captain-held verb" + status_is_captain_held 'working: the captain-held backlog item is next' \ + && fail "a working line mentioning captain-held false-matched" + status_is_captain_held '' && fail "empty line classified as captain-held" + pass "status_is_paused: only the leading paused verb matches, paused is not captain-relevant, and the two declared-wait verbs stay separable" +} + +# crew_absorb_class: the single fm-crew-state.sh read that returns BOTH absorb +# reasons - working (active run/busy pane), paused (declared external wait), or none +# (surface it) - so the watcher's stale path gets both for one bounded call. +# crew_is_paused delegates to it exactly as crew_is_provably_working does. +test_crew_absorb_class_classifier() { + local dir fakebin + dir=$(make_case absorb-class); fakebin="$dir/fakebin" + export FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" + export FM_FAKE_CREW_STATE + FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' + [ "$(crew_absorb_class a)" = working ] || fail "active run-step not classed working" + FM_FAKE_CREW_STATE='state: working · source: pane · harness busy' + [ "$(crew_absorb_class a)" = working ] || fail "busy pane not classed working" + FM_FAKE_CREW_STATE='state: paused · source: status-log · awaiting upstream' + [ "$(crew_absorb_class a)" = paused ] || fail "declared pause not classed paused" + crew_is_paused a || fail "crew_is_paused did not recognize a paused verdict" + ! crew_is_provably_working a || fail "a paused crew was treated as provably working" + FM_FAKE_CREW_STATE='state: working · source: status-log · working: compiling' + [ "$(crew_absorb_class a)" = none ] || fail "stale working: status-log classed absorbable" + FM_FAKE_CREW_STATE='state: unknown · source: none · worktree gone' + [ "$(crew_absorb_class a)" = none ] || fail "unknown crew classed absorbable" + ! crew_is_paused a || fail "unknown crew classed paused" + FM_FAKE_CREW_STATE='state: abandoned · source: run-step · validating (running) · worker gone (dead)' + [ "$(crew_absorb_class a)" = none ] || fail "abandoned run-step classed absorbable" + ! crew_is_provably_working a || fail "abandoned run-step treated as provably working" + FM_FAKE_CREW_STATE='state: abandoned · source: run-step-degraded · worker gone (dead)' + [ "$(crew_absorb_class a)" = none ] || fail "abandoned degraded replay classed absorbable" + [ "$(crew_absorb_class "")" = none ] || fail "empty id not classed none" + unset FM_FAKE_CREW_STATE + pass "crew_absorb_class: working/paused/none from one read; crew_is_paused and crew_is_provably_working agree" +} + +# crew_wedge_progress is the single owner of the run-progress policy both +# supervisors apply, so it is pinned as a pure decision here, independently of +# either one's escalation path: only `progressing` may produce a run-progress +# hold, a confidently dead agent never does, and everything unrecognized +# collapses to `none`. +test_crew_wedge_progress_classifier() { + export FM_FAKE_RUN_PROGRESS + + FM_FAKE_RUN_PROGRESS='progress: progressing · test running, last activity 2m0s ago' + case "$(crew_run_progress a)" in + progressing*) ;; + *) fail "a progressing verdict was not passed through: $(crew_run_progress a)" ;; + esac + case "$(crew_wedge_progress a alive)" in + progressing*) ;; + *) fail "a live agent on a progressing run did not permit a hold" ;; + esac + # The pipeline runs its own steps, so a moving run proves nothing about a crew + # that is gone: a dead agent short-circuits without even paying for the read. + [ "$(crew_wedge_progress a dead)" = none ] \ + || fail "a confidently dead agent was absorbed by its progressing run" + case "$(crew_wedge_progress a unknown)" in + progressing*) ;; + *) fail "an inconclusive liveness verdict blocked a hold" ;; + esac + + FM_FAKE_RUN_PROGRESS='progress: stranded · test running, last activity 31m0s ago' + case "$(crew_wedge_progress a alive)" in + stranded*) ;; + *) fail "a stranded verdict was not passed through" ;; + esac + [ "$(run_progress_detail "$(crew_wedge_progress a alive)")" = "test running, last activity 31m0s ago" ] \ + || fail "run_progress_detail did not strip the class and separator" + + FM_FAKE_RUN_PROGRESS='progress: none · no run attributed to this crew' + [ "$(crew_wedge_progress a alive)" = none ] || fail "a no-evidence verdict was not none" + FM_FAKE_RUN_PROGRESS='progress: something-else · new class' + [ "$(crew_wedge_progress a alive)" = none ] || fail "an unrecognized class was not collapsed to none" + FM_FAKE_RUN_PROGRESS='total gibberish' + [ "$(crew_wedge_progress a alive)" = none ] || fail "unparseable reader output was not collapsed to none" + FM_FAKE_RUN_PROGRESS='progress: progressing · moving' + [ "$(crew_wedge_progress '' alive)" = none ] || fail "an unresolvable task id was not none" + [ "$(run_progress_detail none)" = "" ] || fail "a class-only line reported a detail" + + unset FM_FAKE_RUN_PROGRESS + pass "crew_wedge_progress: only a progressing run holds, a dead agent never does, everything else is none" +} + +# The wedge detector's third liveness input: writes inside the crew's own recorded +# worktree. Every negative outcome must report "no evidence" so the caller keeps +# its existing escalation schedule, and a supervisor-side git read (which touches +# .git, never tracked files) must not be able to fake a positive. +test_crew_worktree_written_since_classifier() { + local dir state anchor wt home statedir_wt + dir=$(make_case classify-worktree-writes); state="$dir/state" + anchor="$state/anchor"; wt="$dir/wt"; home="$dir/mate-home"; statedir_wt="$dir/wt-with-state" + mkdir -p "$wt/src" "$wt/.git/objects" + printf 'old\n' > "$wt/src/existing.c" + set_mtime "$(( $(date +%s) - 300 ))" "$wt/src/existing.c" + : > "$anchor" + set_mtime "$(( $(date +%s) - 120 ))" "$anchor" + + # No recorded worktree at all: absence of evidence, never a positive. + printf 'window=test:fm-a\nkind=ship\n' > "$state/a.meta" + ! crew_worktree_written_since a "$state" "$anchor" \ + || fail "a task with no recorded worktree reported write evidence" + # Recorded but gone (torn down): still no evidence. + printf 'window=test:fm-b\nkind=ship\nworktree=%s\n' "$dir/missing" > "$state/b.meta" + ! crew_worktree_written_since b "$state" "$anchor" \ + || fail "a torn-down worktree reported write evidence" + # Present, but nothing written since the anchor. + printf 'window=test:fm-c\nkind=ship\nworktree=%s\n' "$wt" > "$state/c.meta" + ! crew_worktree_written_since c "$state" "$anchor" \ + || fail "a quiet worktree reported write evidence" + # A missing anchor cannot be compared against: no evidence. + ! crew_worktree_written_since c "$state" "$state/absent-anchor" \ + || fail "a missing anchor reported write evidence" + # Only .git churn (what firstmate's own read-only git commands touch): pruned. + printf 'pack\n' > "$wt/.git/objects/fresh" + printf 'ref\n' > "$wt/.git/index" + ! crew_worktree_written_since c "$state" "$anchor" \ + || fail ".git churn alone reported write evidence (a supervisor read could fake liveness)" + # A real file written after the anchor: positive evidence. + printf 'new\n' > "$wt/src/new.c" + crew_worktree_written_since c "$state" "$anchor" \ + || fail "a file written after the anchor was not reported as write evidence" + # An empty id is never evidence. + ! crew_worktree_written_since "" "$state" "$anchor" || fail "an empty id reported write evidence" + + # A secondmate records a provisioned firstmate home, not a code tree, and such a + # home supervises itself: its own watcher beacon, pane hashes, and heartbeats keep + # its state/ churning whether or not the mate produced anything. + mkdir -p "$home/state" + printf 'sm-classify-1\n' > "$home/.fm-secondmate-home" + printf 'beat\n' > "$home/state/.last-watcher-beat" + printf 'window=remote:sm\nkind=secondmate\nworktree=%s\n' "$home" > "$state/sm.meta" + ! crew_worktree_written_since sm "$state" "$anchor" \ + || fail "a secondmate's own home supervision churn reported crew write evidence" + # The home marker alone is enough, even when the record does not say secondmate. + printf 'window=test:fm-sm2\nkind=ship\nworktree=%s\n' "$home" > "$state/sm2.meta" + ! crew_worktree_written_since sm2 "$state" "$anchor" \ + || fail "a marked firstmate home reported crew write evidence" + # But an ordinary worktree that merely holds a directory named state is real + # work: only the home is excluded, never a source directory of that name. + mkdir -p "$statedir_wt/state" + printf 'machine\n' > "$statedir_wt/state/machine.go" + printf 'window=test:fm-d\nkind=ship\nworktree=%s\n' "$statedir_wt" > "$state/d.meta" + crew_worktree_written_since d "$state" "$anchor" \ + || fail "a source directory named state was hidden from the write probe" + pass "crew_worktree_written_since: real writes are evidence; no worktree, no anchor, quiet trees, .git churn and a mate's own home are not" +} + +# FM_WORKTREE_WRITE_PRUNE is a skip list, so clearing it skips nothing and is the +# obvious way to widen the probe to the whole depth-bounded tree. An empty list must +# therefore widen the walk rather than report no evidence at all, which would +# quietly cost the wedge detector its third liveness input on a home that cleared +# the knob to get more coverage, not less. +test_empty_write_prune_widens_the_probe() { + local dir state anchor wt saved + dir=$(make_case classify-empty-write-prune); state="$dir/state" + anchor="$state/anchor"; wt="$dir/wt" + mkdir -p "$wt/src" "$wt/.git" + : > "$anchor" + set_mtime "$(( $(date +%s) - 120 ))" "$anchor" + printf 'window=test:fm-e\nkind=ship\nworktree=%s\n' "$wt" > "$state/e.meta" + saved=$FM_WORKTREE_WRITE_PRUNE + FM_WORKTREE_WRITE_PRUNE='' + # A quiet tree is still no evidence, so the caller's schedule is untouched. + ! crew_worktree_written_since e "$state" "$anchor" \ + || fail "an empty prune list reported write evidence for a quiet worktree" + printf 'new\n' > "$wt/src/new.c" + crew_worktree_written_since e "$state" "$anchor" \ + || fail "an empty prune list disabled the probe instead of widening it" + # Widened means nothing is skipped, including what the default list prunes. + set_mtime "$(( $(date +%s) - 900 ))" "$wt/src/new.c" + printf 'pack\n' > "$wt/.git/index" + crew_worktree_written_since e "$state" "$anchor" \ + || fail "an empty prune list still skipped a directory the default list prunes" + # Restoring the default prunes .git again, so a supervisor's own read-only git + # command still cannot fake liveness. + FM_WORKTREE_WRITE_PRUNE=$saved + ! crew_worktree_written_since e "$state" "$anchor" \ + || fail "the default prune list stopped keeping .git out of the probe" + pass "an empty FM_WORKTREE_WRITE_PRUNE widens the probe to the whole depth-bounded tree instead of disabling it" +} + +# The same widening, reached the way a home actually configures it: through the +# process ENVIRONMENT, not an in-process assignment made after the library was +# sourced. An empty exported value must survive as empty, because defaulting it with +# the colon form reads "explicitly cleared" as "never set" and hands the default skip +# list straight back to the one home that asked for a wider walk. +# shellcheck disable=SC2016 # single quotes are deliberate: the library path, state dir, and anchor expand inside the bash -c child, not here +test_empty_write_prune_from_the_environment_widens_the_probe() { + local dir state anchor wt + dir=$(make_case classify-empty-write-prune-env); state="$dir/state" + anchor="$state/anchor"; wt="$dir/wt" + mkdir -p "$wt/.git/objects" + : > "$anchor" + set_mtime "$(( $(date +%s) - 120 ))" "$anchor" + printf 'window=test:fm-wenv\nkind=ship\nworktree=%s\n' "$wt" > "$state/wenv.meta" + # The one thing written since the anchor sits exactly where the DEFAULT list prunes. + printf 'pack\n' > "$wt/.git/objects/fresh" + env -u FM_WORKTREE_WRITE_PRUNE \ + bash -c '. "$1"; crew_worktree_written_since wenv "$2" "$3"' _ \ + "$ROOT/bin/fm-classify-lib.sh" "$state" "$anchor" \ + && fail "the default skip list let .git churn count as write evidence" + FM_WORKTREE_WRITE_PRUNE='' \ + bash -c '. "$1"; crew_worktree_written_since wenv "$2" "$3"' _ \ + "$ROOT/bin/fm-classify-lib.sh" "$state" "$anchor" \ + || fail "an empty FM_WORKTREE_WRITE_PRUNE in the environment fell back to the default skip list instead of widening the probe" + pass "an empty FM_WORKTREE_WRITE_PRUNE exported into the environment prunes nothing, widening the probe" +} + +# The probe's walk runs synchronously inside the poll that was about to escalate, so +# it must be wall-clock bounded: -xdev keeps it out of a nested mount, but a worktree +# root that is ITSELF on a hung mount would otherwise stall the very supervisor that +# exists to notice a wedge. A fake find that never returns in time stands in for that +# mount. Hitting the bound must read as NO evidence, exactly like every other +# negative outcome, so the caller's escalation schedule is untouched. +test_worktree_write_probe_is_wall_clock_bounded() { + local dir state anchor wt slowbin fastbin started elapsed + dir=$(make_case classify-write-probe-bound); state="$dir/state" + anchor="$state/anchor"; wt="$dir/wt"; slowbin="$dir/slowbin"; fastbin="$dir/fastbin" + mkdir -p "$wt/src" "$slowbin" "$fastbin" + : > "$anchor" + set_mtime "$(( $(date +%s) - 120 ))" "$anchor" + printf 'window=test:fm-slow\nkind=ship\nworktree=%s\n' "$wt" > "$state/slow.meta" + # Both stand-ins report the same hit; only one of them takes longer than the bound + # to do it, so the prompt one shows what a positive outcome looks like and the + # bounded assertion below cannot pass merely because the fake failed. + cat > "$fastbin/find" <<'SH' +#!/usr/bin/env bash +set -u +printf '%s\n' "$1/hit" +SH + cat > "$slowbin/find" <<'SH' +#!/usr/bin/env bash +set -u +sleep 30 +printf '%s\n' "$1/hit" +SH + chmod +x "$fastbin/find" "$slowbin/find" + PATH="$fastbin:$PATH" \ + bash -c '. "$1"; crew_worktree_written_since slow "$2" "$3"' _ \ + "$ROOT/bin/fm-classify-lib.sh" "$state" "$anchor" \ + || fail "a walk that reported a hit inside its bound was not read as write evidence" + started=$(date +%s) + PATH="$slowbin:$PATH" FM_WORKTREE_WRITE_TIMEOUT=1 \ + bash -c '. "$1"; crew_worktree_written_since slow "$2" "$3"' _ \ + "$ROOT/bin/fm-classify-lib.sh" "$state" "$anchor" \ + && fail "a walk that outlived its bound was reported as write evidence" + elapsed=$(( $(date +%s) - started )) + [ "$elapsed" -lt 10 ] \ + || fail "the worktree write probe was not wall-clock bounded: one walk held the caller for ${elapsed}s" + pass "the worktree write probe is wall-clock bounded, and hitting the bound reads as no write evidence" +} + +# signal_crew_provably_working: a no-verb "signal:" wake is benign ONLY when EVERY +# task it references is provably working; if any crew has stopped, or no task can be +# resolved, it surfaces. Files map to ids by stripping .status / .turn-ended. +test_signal_crew_provably_working_classifier() { + local dir fakebin state + dir=$(make_case signal-provably-working); fakebin="$dir/fakebin"; state="$dir/state" + export FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" + export FM_FAKE_CREW_STATE_a='state: working · source: run-step · running' + export FM_FAKE_CREW_STATE_b='state: done · source: run-step · run passed' + signal_crew_provably_working "$state/a.status" "$state/a.turn-ended" \ + || fail "a single provably-working crew (status+turn-end) was not benign" + ! signal_crew_provably_working "$state/a.status" "$state/b.turn-ended" \ + || fail "a coalesced batch including a stopped crew was treated as benign" + ! signal_crew_provably_working "$state/b.turn-ended" \ + || fail "a stopped crew's bare turn-end was treated as benign" + ! signal_crew_provably_working "$state/a.meta" \ + || fail "a non-signal file resolved to a benign verdict" + ! signal_crew_provably_working \ + || fail "an empty signal file list was treated as benign" + unset FM_FAKE_CREW_STATE_a FM_FAKE_CREW_STATE_b + pass "signal_crew_provably_working: benign only when every referenced crew is provably working" +} + +test_secondmate_status_signal_never_absorbed_classifier() { + local dir fakebin state + dir=$(make_case secondmate-signal-classify); fakebin="$dir/fakebin"; state="$dir/state" + export FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" + # Even PROVABLY working, a secondmate's .status signal is its routed-reply + # channel and must surface; its bare turn-ended keeps the ordinary absorb. + export FM_FAKE_CREW_STATE_sm='state: working · source: run-step · running' + printf 'kind=secondmate\n' > "$state/sm.meta" + printf 'working: routed reply for the parent\n' > "$state/sm.status" + ! signal_crew_provably_working "$state/sm.status" \ + || fail "a working secondmate's status signal was treated as absorbable" + signal_crew_provably_working "$state/sm.turn-ended" \ + || fail "a working secondmate's bare turn-end lost its ordinary absorb" + # An ordinary crewmate with the same verdict stays absorbable: the rule is + # keyed on recorded kind, not on task naming or content guessing. + export FM_FAKE_CREW_STATE_crew='state: working · source: run-step · running' + printf 'kind=ship\n' > "$state/crew.meta" + printf 'working: progress\n' > "$state/crew.status" + signal_crew_provably_working "$state/crew.status" \ + || fail "the secondmate rule leaked onto an ordinary crewmate status" + unset FM_FAKE_CREW_STATE_sm FM_FAKE_CREW_STATE_crew + pass "a secondmate's status signal is never absorbed as provably working; crewmates are unaffected" +} + +# --- benign wakes are absorbed ONLY when the crew is provably working --------- + +test_provably_working_signal_absorbed() { + local dir state fakebin out status_file pid + dir=$(make_case provably-working-signal); state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out" + status_file="$state/task.status" + printf 'working: compiling step 2\n' > "$status_file" + # The crew's pipeline is in an actively-running step: positive evidence it is + # still working, so a no-verb working: signal is absorbed (the original low-churn + # case during a long validation). + export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' + watch_bg "$state" "$fakebin" "$out" + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid"; fail "watcher exited for a working: signal whose crew is provably working (should absorb): $(cat "$out")" + fi + [ ! -s "$out" ] || fail "provably-working signal printed a wake reason: $(cat "$out")" + [ ! -s "$state/.wake-queue" ] || fail "provably-working signal enqueued a durable wake record" + [ -s "$state/.seen-task_status" ] || fail "provably-working signal did not advance its .seen-* suppressor" + [ -e "$state/.last-watcher-beat" ] || fail "watcher beacon was not touched while absorbing" + reap "$pid" + pass "a no-verb signal whose crew is provably working is absorbed (no exit, no queue, suppressor advanced, beacon present)" +} + +test_turn_ended_provably_working_absorbed() { + local dir state fakebin out pid + dir=$(make_case turn-ended-working); state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out" + : > "$state/task.turn-ended" + # A busy pane is the second form of positive evidence (covers a queued + # continuation right after the turn-end). + export FM_FAKE_CREW_STATE='state: working · source: pane · harness busy' + watch_bg "$state" "$fakebin" "$out" + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid"; fail "watcher exited for a turn-end whose crew is provably working (should absorb): $(cat "$out")" + fi + [ ! -s "$out" ] || fail "provably-working turn-end printed a wake reason: $(cat "$out")" + [ ! -s "$state/.wake-queue" ] || fail "provably-working turn-end enqueued a durable wake record" + reap "$pid" + pass "a bare turn-end whose crew is provably working (busy pane) is absorbed" +} + +# --- a no-verb signal whose crew is NOT provably working SURFACES ------------- +# This is the swallowed-finish fix: a crew that finished (or stopped and waits) +# reports its final turn-end with no captain-relevant status and no running +# pipeline, so the wake must surface instead of being absorbed. + +test_turn_ended_not_working_surfaced() { + local dir state fakebin out drain_out pid + dir=$(make_case turn-ended-stopped); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out" + : > "$state/task.turn-ended" + # No running pipeline, no busy pane: the crew has stopped (e.g. it finished via + # an interactive menu and wrote no done: status). Default unknown verdict. + export FM_FAKE_CREW_STATE='state: unknown · source: none · no current-state source available' + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not surface a turn-end whose crew is not provably working" + grep -F "signal: $state/task.turn-ended" "$out" >/dev/null || fail "watcher did not print the surfaced turn-end signal" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the surfaced turn-end failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/task.turn-ended" >/dev/null || fail "surfaced turn-end was not queued" + pass "a bare turn-end whose crew is not provably working is surfaced (the swallowed-finish fix)" +} + +# --- bare turn-end, unverifiable harness: pane churn is the third proof -------- +# A harness whose semantic busy state has no verified source (codex) can never +# report working, so the two proofs above are unreachable for it and EVERY worker +# turn boundary woke firstmate. Pane content that changed since the previous poll +# is harness-independent positive evidence the crew is still executing - the same +# liveness input the stale backbone already trusts - so a bare turn-end from a +# churning pane is benign. The pane going quiet afterwards is still caught by that +# backbone, which is why this widens the proof rather than bounding the wake rate. + +# The pane-churn turn-end absorb is opt-in per home, so every case that exercises +# it (whether it expects an absorb or one of the guards that must still surface) +# points the watcher at a case-local config dir holding the flag. A case that must +# NOT have it points at an empty one, so no developer's real config can leak in. +churn_config() { # <dir> [off] + local cfg="$1/config" + mkdir -p "$cfg" + [ "${2:-}" = off ] || : > "$cfg/turnend-churn-absorb" + printf '%s\n' "$cfg" +} + +# Wait until the watcher records an absorbed wake matching <needle> in its triage +# log. 1 if the watcher exits first (i.e. it surfaced the wake instead), which is +# exactly the unfixed behavior this case exists to catch. Polls the log rather +# than a poll cycle so the assertion lands inside the FIRST poll, long before an +# unchanging fixture pane could reach the stale backbone. +wait_for_absorbed() { # <state> <pid> <needle> + local state=$1 pid=$2 needle=$3 i=0 + while [ "$i" -lt 100 ]; do + grep -Fq "$needle" "$state/.watch-triage.log" 2>/dev/null && return 0 + kill -0 "$pid" 2>/dev/null || return 1 + sleep 0.1 + i=$((i + 1)) + done + return 1 +} + +test_turn_ended_churning_pane_absorbed() { + local dir state fakebin out capture_file window key pid + dir=$(make_case turn-ended-churning); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt" + window="test:fm-codexer" + : > "$state/codexer.turn-ended" + printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codexer.meta" + printf 'apply_patch: writing bin/thing.sh' > "$capture_file" + key=$(printf '%s' "$window" | tr ':/.' '___') + # The previous poll recorded DIFFERENT pane content, so this poll's capture is + # churn: the crew rendered output between the two polls. + printf '%s' "$(hash_text 'reading the brief')" > "$state/.hash-$key" + printf '0\n' > "$state/.count-$key" + # The codex verdict verbatim: a verified dispatch adapter with no verified + # semantic busy source, so crew_is_provably_working can never be satisfied. + export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' + # A slow poll leaves the first cycle's absorb assertion many ticks clear of the + # stale backbone, which this static fixture pane would otherwise reach. + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_absorbed "$state" "$pid" "absorbed benign signal:" \ + || { reap "$pid"; fail "a bare turn-end from a churning pane was not absorbed: $(cat "$out")"; } + [ ! -s "$out" ] || fail "an absorbed churning-pane turn-end printed a wake reason: $(cat "$out")" + [ ! -s "$state/.wake-queue" ] || fail "an absorbed churning-pane turn-end enqueued a durable wake record" + [ -s "$state/.churn-since-$key" ] \ + || { reap "$pid"; fail "an absorbed churning-pane turn-end did not open a bounded deferral window"; } + reap "$pid" + unset FM_FAKE_CREW_STATE + pass "a bare turn-end from a pane that churned since the previous poll is absorbed" +} + +test_turn_ended_churn_resets_prior_stale_classification() { + local dir state fakebin out capture_file window key old_hash active_hash pid i + dir=$(make_case turn-ended-churn-resets-stale); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt" + window="test:fm-codexreturned" + : > "$state/codexreturned.turn-ended" + printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codexreturned.meta" + old_hash=$(hash_text 'idle prompt from an earlier turn') + active_hash=$(hash_text 'rendering a new turn') + printf 'rendering a new turn' > "$capture_file" + key=$(printf '%s' "$window" | tr ':/.' '___') + printf '%s' "$old_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + printf '%s' "$old_hash" > "$state/.stale-$key" + date +%s > "$state/.stale-since-$key" + export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=999 \ + FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_absorbed "$state" "$pid" "absorbed benign signal:" \ + || { reap "$pid"; fail "a churning turn-end with prior stale state was not absorbed: $(cat "$out")"; } + i=0 + while [ "$i" -lt 100 ] && [ "$(cat "$state/.hash-$key" 2>/dev/null || true)" != "$active_hash" ]; do + kill -0 "$pid" 2>/dev/null || { reap "$pid"; fail "watcher exited before recording the active pane"; } + sleep 0.1 + i=$((i + 1)) + done + [ "$(cat "$state/.hash-$key" 2>/dev/null || true)" = "$active_hash" ] \ + || { reap "$pid"; fail "watcher did not record the active pane after absorbing its turn-end"; } + + # The worker stops on bytes that happened to be stale in an earlier turn. + # This is a new quiet interval, so it must surface through ordinary staleness + # instead of inheriting the earlier interval's wedge timer. + printf 'idle prompt from an earlier turn' > "$capture_file" + wait_for_exit "$pid" 100 \ + || { reap "$pid"; fail "a stopped pane matching an earlier stale render waited for the wedge timeout"; } + grep -Fx "stale: $window" "$out" >/dev/null \ + || fail "the returned stale render did not surface through ordinary staleness" + grep -F "possible wedge" "$out" >/dev/null \ + && fail "the returned stale render inherited the earlier quiet interval's wedge classification" + unset FM_FAKE_CREW_STATE + pass "pane churn starts a fresh stale-classification interval before a stopped render returns" +} + +test_turn_ended_churn_resets_wedge_state_before_stale_poll() { + local dir state fakebin out capture_file capture_count window key pid + dir=$(make_case turn-ended-churn-resets-wedge); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; capture_count="$dir/capture.count" + window="test:fm-codexfreshinterval" + : > "$state/codexfreshinterval.turn-ended" + printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codexfreshinterval.meta" + printf 'rendering a new turn' > "$capture_file" + key=$(printf '%s' "$window" | tr ':/.' '___') + printf '%s' "$(hash_text 'idle output from the prior interval')" > "$state/.hash-$key" + printf '2\n' > "$state/.wedge-escalations-$key" + export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_TMUX_CAPTURE_COUNT_FILE="$capture_count" FM_FAKE_TMUX_CAPTURE_FAIL_AFTER=1 \ + FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_absorbed "$state" "$pid" "absorbed benign signal:" \ + || { reap "$pid"; fail "a churning turn-end was not absorbed before the stale-path capture failed: $(cat "$out")"; } + [ ! -e "$state/.wedge-escalations-$key" ] \ + || { reap "$pid"; fail "churn retained the prior quiet interval's wedge-escalation count"; } + [ ! -s "$state/.wake-queue" ] \ + || { reap "$pid"; fail "the absorbed churn fixture queued an unexpected wake"; } + reap "$pid" + unset FM_FAKE_CREW_STATE + pass "pane churn resets prior wedge escalation state before the stale-path poll" +} + +# The safety half: the same unverifiable harness, the same fixture, but the pane +# has NOT changed since the previous poll. There is no positive evidence, so the +# wake must still surface - a stopped worker is exactly what the turn-end marker +# earns its keep detecting, and widening the proof must not cost that. +test_turn_ended_still_pane_surfaced() { + local dir state fakebin out drain_out capture_file window key pid + dir=$(make_case turn-ended-still); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + window="test:fm-codexstopped" + : > "$state/codexstopped.turn-ended" + printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codexstopped.meta" + printf 'apply_patch: writing bin/thing.sh' > "$capture_file" + key=$(printf '%s' "$window" | tr ':/.' '___') + # The previous poll recorded THIS pane content: nothing rendered since. + printf '%s' "$(hash_text 'apply_patch: writing bin/thing.sh')" > "$state/.hash-$key" + printf '0\n' > "$state/.count-$key" + export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not surface a bare turn-end from an unchanged pane" + grep -F "signal: $state/codexstopped.turn-ended" "$out" >/dev/null \ + || fail "watcher did not print the surfaced still-pane turn-end signal" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the still-pane turn-end failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/codexstopped.turn-ended" >/dev/null \ + || fail "surfaced still-pane turn-end was not queued" + unset FM_FAKE_CREW_STATE + pass "a bare turn-end from a pane unchanged since the previous poll still surfaces" +} + +test_turn_ended_malformed_prior_hash_surfaced() { + local dir state fakebin out drain_out capture_file window key pid + dir=$(make_case turn-ended-malformed-hash); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + window="test:fm-codexmalformed" + : > "$state/codexmalformed.turn-ended" + printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codexmalformed.meta" + printf 'stopped after rendering this' > "$capture_file" + key=$(printf '%s' "$window" | tr ':/.' '___') + printf 'x' > "$state/.hash-$key" + printf '0\n' > "$state/.count-$key" + export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher absorbed a turn-end backed by a malformed prior hash" + grep -F "signal: $state/codexmalformed.turn-ended" "$out" >/dev/null \ + || fail "watcher did not print the surfaced malformed-hash turn-end" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ + || fail "drain after the malformed-hash turn-end failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/codexmalformed.turn-ended" >/dev/null \ + || fail "malformed-hash turn-end was not queued" + unset FM_FAKE_CREW_STATE + pass "a bare turn-end backed by a malformed prior hash surfaces" +} + +test_turn_ended_trailing_newline_prior_hash_surfaced() { + local dir state fakebin out drain_out capture_file window key pid + dir=$(make_case turn-ended-newline-hash); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + window="test:fm-codexnewline" + : > "$state/codexnewline.turn-ended" + printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codexnewline.meta" + printf 'rendered after the prior poll' > "$capture_file" + key=$(printf '%s' "$window" | tr ':/.' '___') + printf '%s\n' "$(hash_text 'the previous render')" > "$state/.hash-$key" + printf '0\n' > "$state/.count-$key" + export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher absorbed a turn-end backed by a newline-terminated prior hash" + grep -F "signal: $state/codexnewline.turn-ended" "$out" >/dev/null \ + || fail "watcher did not print the surfaced newline-hash turn-end" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ + || fail "drain after the newline-hash turn-end failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/codexnewline.turn-ended" >/dev/null \ + || fail "newline-hash turn-end was not queued" + [ ! -e "$state/.churn-since-$key" ] \ + || fail "a newline-terminated prior hash opened a deferral window" + unset FM_FAKE_CREW_STATE + pass "a bare turn-end backed by a newline-terminated prior hash surfaces" +} + +test_secondmate_turn_ended_churning_pane_surfaced() { + local dir state fakebin out drain_out capture_file window key pid + dir=$(make_case secondmate-turn-ended-churning); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + window="test:fm-mate-churning" + : > "$state/mate.turn-ended" + printf 'window=%s\nkind=secondmate\nharness=pi\n' "$window" > "$state/mate.meta" + printf 'working on the next routed item' > "$capture_file" + key=$(printf '%s' "$window" | tr ':/.' '___') + printf '%s' "$(hash_text 'waiting for work')" > "$state/.hash-$key" + printf '0\n' > "$state/.count-$key" + export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable' + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not surface a churning secondmate turn-end" + grep -F "signal: $state/mate.turn-ended" "$out" >/dev/null \ + || fail "watcher did not print the surfaced churning secondmate turn-end" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ + || fail "drain after the churning secondmate turn-end failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/mate.turn-ended" >/dev/null \ + || fail "churning secondmate turn-end was not queued" + unset FM_FAKE_CREW_STATE + pass "a churning secondmate turn-end surfaces without a stale resurface path" +} + +test_turn_ended_colliding_window_key_surfaced() { + local dir state fakebin out drain_out capture_file window colliding key pid + dir=$(make_case turn-ended-colliding-key); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + window="test:fm-a.b"; colliding="test:fm-a_b" + : > "$state/a.b.turn-ended" + printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/a.b.meta" + printf 'window=%s\nkind=ship\nharness=codex\n' "$colliding" > "$state/a_b.meta" + printf 'rendered after the prior poll' > "$capture_file" + key=$(printf '%s' "$window" | tr ':/.' '___') + printf '%s' "$(hash_text 'the other window pane')" > "$state/.hash-$key" + printf '0\n' > "$state/.count-$key" + export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not surface a turn-end with an ambiguous pane marker" + grep -F "signal: $state/a.b.turn-ended" "$out" >/dev/null \ + || fail "watcher did not print the surfaced ambiguous-marker turn-end" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ + || fail "drain after the ambiguous-marker turn-end failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/a.b.turn-ended" >/dev/null \ + || fail "ambiguous-marker turn-end was not queued" + unset FM_FAKE_CREW_STATE + pass "a turn-end whose marker key matches another recorded endpoint surfaces" +} + +test_turn_ended_duplicate_endpoint_records_surfaced() { + local dir state fakebin out drain_out capture_file window key pid + dir=$(make_case turn-ended-duplicate-endpoint); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + window="test:fm-shared" + : > "$state/first.turn-ended" + printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/first.meta" + printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/second.meta" + printf 'rendered after the prior poll' > "$capture_file" + key=$(printf '%s' "$window" | tr ':/.' '___') + printf '%s' "$(hash_text 'the previous render')" > "$state/.hash-$key" + printf '0\n' > "$state/.count-$key" + export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher absorbed a turn-end shared by two endpoint records" + grep -F "signal: $state/first.turn-ended" "$out" >/dev/null \ + || fail "watcher did not print the surfaced duplicate-endpoint turn-end" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ + || fail "drain after the duplicate-endpoint turn-end failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/first.turn-ended" >/dev/null \ + || fail "duplicate-endpoint turn-end was not queued" + [ ! -e "$state/.churn-since-$key" ] \ + || fail "duplicate endpoint records opened a deferral window" + unset FM_FAKE_CREW_STATE + pass "two metadata records sharing one endpoint make churn evidence ambiguous" +} + +test_turn_ended_mixed_positive_evidence_batch_absorbed() { + local dir state fakebin out capture_file first_window second_window first_key second_key pid + dir=$(make_case turn-ended-mixed-evidence); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt" + first_window="test:fm-first"; second_window="test:fm-second" + : > "$state/first.turn-ended" + : > "$state/second.turn-ended" + printf 'window=%s\nkind=ship\nharness=pi\n' "$first_window" > "$state/first.meta" + printf 'window=%s\nkind=ship\nharness=codex\n' "$second_window" > "$state/second.meta" + printf 'second task rendered after the prior poll' > "$capture_file" + first_key=$(printf '%s' "$first_window" | tr ':/.' '___') + second_key=$(printf '%s' "$second_window" | tr ':/.' '___') + printf '%s' "$(hash_text 'first task static pane')" > "$state/.hash-$first_key" + printf '%s' "$(hash_text 'second task previous render')" > "$state/.hash-$second_key" + printf '0\n' > "$state/.count-$first_key" + printf '0\n' > "$state/.count-$second_key" + export FM_FAKE_CREW_STATE_first='state: working · source: run-step · running' + export FM_FAKE_CREW_STATE_second='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOWS="$(printf 'fm-first\nfm-second')" \ + FM_FAKE_TMUX_CAPTURE="$capture_file" FM_FAKE_TMUX_FORBIDDEN_TARGET="$first_window" \ + FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_absorbed "$state" "$pid" "absorbed benign signal:" \ + || { reap "$pid"; fail "a mixed authoritative-and-churn batch was not absorbed: $(cat "$out")"; } + [ ! -s "$out" ] || fail "an absorbed mixed-evidence batch printed a wake reason: $(cat "$out")" + [ ! -s "$state/.wake-queue" ] || fail "an absorbed mixed-evidence batch enqueued a durable wake record" + [ ! -e "$state/.churn-since-$first_key" ] \ + || fail "an authoritatively working task opened a pane-churn deadline" + [ -s "$state/.churn-since-$second_key" ] \ + || fail "the churn-proven task did not open its bounded deferral window" + reap "$pid" + unset FM_FAKE_CREW_STATE_first FM_FAKE_CREW_STATE_second + pass "a batch may satisfy positive evidence independently per task" +} + +test_turn_ended_mixed_positive_evidence_batch_default_off() { + local dir state fakebin out drain_out capture_file first_window second_window first_key second_key pid + dir=$(make_case turn-ended-mixed-evidence-off); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + first_window="test:fm-firstoff"; second_window="test:fm-secondoff" + : > "$state/firstoff.turn-ended" + : > "$state/secondoff.turn-ended" + printf 'window=%s\nkind=ship\nharness=pi\n' "$first_window" > "$state/firstoff.meta" + printf 'window=%s\nkind=ship\nharness=codex\n' "$second_window" > "$state/secondoff.meta" + printf 'second task rendered after the prior poll' > "$capture_file" + first_key=$(printf '%s' "$first_window" | tr ':/.' '___') + second_key=$(printf '%s' "$second_window" | tr ':/.' '___') + printf '%s' "$(hash_text 'first task static pane')" > "$state/.hash-$first_key" + printf '%s' "$(hash_text 'second task previous render')" > "$state/.hash-$second_key" + printf '0\n' > "$state/.count-$first_key" + printf '0\n' > "$state/.count-$second_key" + export FM_FAKE_CREW_STATE_firstoff='state: working · source: run-step · running' + export FM_FAKE_CREW_STATE_secondoff='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOWS="$(printf 'fm-firstoff\nfm-secondoff')" \ + FM_FAKE_TMUX_CAPTURE="$capture_file" FM_CONFIG_OVERRIDE="$(churn_config "$dir" off)" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher absorbed a mixed-evidence batch without the opt-in flag" + grep -F "$state/firstoff.turn-ended" "$out" >/dev/null \ + || fail "watcher did not print the first default-off turn-end" + grep -F "$state/secondoff.turn-ended" "$out" >/dev/null \ + || fail "watcher did not print the second default-off turn-end" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ + || fail "drain after the default-off mixed-evidence batch failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/firstoff.turn-ended" >/dev/null \ + || fail "the first default-off turn-end was not queued" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/secondoff.turn-ended" >/dev/null \ + || fail "the second default-off turn-end was not queued" + [ ! -e "$state/.churn-since-$first_key" ] && [ ! -e "$state/.churn-since-$second_key" ] \ + || fail "the default-off mixed-evidence batch opened a deferral window" + unset FM_FAKE_CREW_STATE_firstoff FM_FAKE_CREW_STATE_secondoff + pass "per-task evidence composition stays off until the home opts in" +} + +test_status_and_turn_end_batch_never_uses_churn_evidence() { + local dir state fakebin out drain_out capture_file first_window second_window second_key pid + dir=$(make_case status-and-turn-ended-churn); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + first_window="test:fm-firststatus"; second_window="test:fm-secondturn" + printf 'working: authoritative task still running\n' > "$state/firststatus.status" + : > "$state/secondturn.turn-ended" + printf 'window=%s\nkind=ship\nharness=pi\n' "$first_window" > "$state/firststatus.meta" + printf 'window=%s\nkind=ship\nharness=codex\n' "$second_window" > "$state/secondturn.meta" + printf 'second task rendered after the prior poll' > "$capture_file" + second_key=$(printf '%s' "$second_window" | tr ':/.' '___') + printf '%s' "$(hash_text 'second task previous render')" > "$state/.hash-$second_key" + printf '0\n' > "$state/.count-$second_key" + export FM_FAKE_CREW_STATE_firststatus='state: working · source: run-step · running' + export FM_FAKE_CREW_STATE_secondturn='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOWS="$(printf 'fm-firststatus\nfm-secondturn')" \ + FM_FAKE_TMUX_CAPTURE="$capture_file" FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher absorbed a status-and-turn-end batch on churn evidence" + grep -F "$state/firststatus.status" "$out" >/dev/null \ + || fail "watcher did not print the status file from the surfaced mixed batch" + grep -F "$state/secondturn.turn-ended" "$out" >/dev/null \ + || fail "watcher did not print the turn-end from the surfaced mixed batch" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ + || fail "drain after the surfaced status-and-turn-end batch failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/firststatus.status" >/dev/null \ + || fail "the status file from the surfaced mixed batch was not queued" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/secondturn.turn-ended" >/dev/null \ + || fail "the turn-end from the surfaced mixed batch was not queued" + [ ! -e "$state/.churn-since-$second_key" ] \ + || fail "a status-bearing batch opened a pane-churn deadline" + unset FM_FAKE_CREW_STATE_firststatus FM_FAKE_CREW_STATE_secondturn + pass "a status-bearing batch never falls through to pane-churn evidence" +} + +# The opt-in half. Pane churn infers execution from rendered bytes rather than +# from a verdict the harness vouches for, so a home that has not asked for it must +# see exactly the pre-change triage: the same churning fixture that absorbs above +# surfaces here purely because the flag is absent. +test_turn_ended_churn_absorb_off_by_default() { + local dir state fakebin out drain_out capture_file window key pid + dir=$(make_case turn-ended-churn-default-off); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + window="test:fm-codexdefault" + : > "$state/codexdefault.turn-ended" + printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codexdefault.meta" + printf 'apply_patch: writing bin/thing.sh' > "$capture_file" + key=$(printf '%s' "$window" | tr ':/.' '___') + printf '%s' "$(hash_text 'reading the brief')" > "$state/.hash-$key" + printf '0\n' > "$state/.count-$key" + export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_CONFIG_OVERRIDE="$(churn_config "$dir" off)" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher absorbed a churning turn-end without the opt-in flag" + grep -F "signal: $state/codexdefault.turn-ended" "$out" >/dev/null \ + || fail "watcher did not print the surfaced default-off churning turn-end" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ + || fail "drain after the default-off churning turn-end failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/codexdefault.turn-ended" >/dev/null \ + || fail "default-off churning turn-end was not queued" + [ ! -e "$state/.churn-since-$key" ] \ + || fail "the default-off path opened a bounded deferral window" + unset FM_FAKE_CREW_STATE + pass "pane-churn turn-end absorb is off until a home opts in" +} + +# The bound. Churn and pane staleness read the same pane, so a pane that renders +# continuously (a clock, a spinner, a harness that leaves a background renderer +# alive after its agent yields) never reaches the staleness backbone's two +# identical hashes either. Without a bound on the churn absorb a worker that had +# genuinely stopped behind such a renderer would have no path left to surface at +# all, so an exhausted deferral window must surface and restart. +test_turn_ended_churn_absorb_bounded() { + local dir state fakebin out drain_out capture_file window key pid + dir=$(make_case turn-ended-churn-bounded); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + window="test:fm-codexclock" + : > "$state/codexclock.turn-ended" + printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codexclock.meta" + printf 'a background renderer that never stops' > "$capture_file" + key=$(printf '%s' "$window" | tr ':/.' '___') + printf '%s' "$(hash_text 'the previous frame')" > "$state/.hash-$key" + printf '0\n' > "$state/.count-$key" + # This endpoint has already been riding churn evidence longer than the bound. + printf '%s' "$(( $(date +%s) - 600 ))" > "$state/.churn-since-$key" + export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_CONFIG_OVERRIDE="$(churn_config "$dir")" FM_TURNEND_CHURN_ABSORB_SECS=60 \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 \ + || fail "a perpetually churning pane deferred its turn-end past the absorb bound" + grep -F "signal: $state/codexclock.turn-ended" "$out" >/dev/null \ + || fail "watcher did not print the turn-end surfaced by the exhausted absorb bound" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ + || fail "drain after the bounded churn turn-end failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/codexclock.turn-ended" >/dev/null \ + || fail "the turn-end surfaced by the exhausted absorb bound was not queued" + [ ! -e "$state/.churn-since-$key" ] \ + || fail "an exhausted deferral window was not restarted after surfacing" + unset FM_FAKE_CREW_STATE + pass "a perpetually churning pane surfaces once its bounded deferral window is spent" +} + +test_turn_ended_churn_timer_write_failure_surfaced() { + local dir state fakebin out drain_out capture_file window key pid + dir=$(make_case turn-ended-churn-timer-write-failure); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + window="test:fm-codextimer" + : > "$state/codextimer.turn-ended" + printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codextimer.meta" + printf 'rendered after the previous poll' > "$capture_file" + key=$(printf '%s' "$window" | tr ':/.' '___') + printf '%s' "$(hash_text 'the previous render')" > "$state/.hash-$key" + printf '0\n' > "$state/.count-$key" + mkdir "$state/.churn-since-$key" + export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" 2>/dev/null & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher absorbed a churning turn-end without recording its deadline" + grep -F "signal: $state/codextimer.turn-ended" "$out" >/dev/null \ + || fail "watcher did not print the turn-end whose churn deadline could not be recorded" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ + || fail "drain after the failed churn deadline write failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/codextimer.turn-ended" >/dev/null \ + || fail "turn-end with an unrecordable churn deadline was not queued" + unset FM_FAKE_CREW_STATE + pass "an unrecordable pane-churn deadline surfaces the turn-end" +} + +test_turn_ended_invalid_churn_bound_surfaced() { + local dir state fakebin out drain_out capture_file window key pid + dir=$(make_case turn-ended-invalid-churn-bound); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + window="test:fm-codexbound" + : > "$state/codexbound.turn-ended" + printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codexbound.meta" + printf 'rendered after the previous poll' > "$capture_file" + key=$(printf '%s' "$window" | tr ':/.' '___') + printf '%s' "$(hash_text 'the previous render')" > "$state/.hash-$key" + printf '0\n' > "$state/.count-$key" + export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_CONFIG_OVERRIDE="$(churn_config "$dir")" FM_TURNEND_CHURN_ABSORB_SECS=bogus \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" 2>/dev/null & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not surface a turn-end with an invalid churn bound" + grep -F "signal: $state/codexbound.turn-ended" "$out" >/dev/null \ + || fail "watcher terminated before printing the invalid-bound turn-end" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ + || fail "drain after the invalid churn bound failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/codexbound.turn-ended" >/dev/null \ + || fail "turn-end with an invalid churn bound was not queued" + [ ! -e "$state/.churn-since-$key" ] \ + || fail "an invalid churn bound opened a deferral window" + unset FM_FAKE_CREW_STATE + pass "an invalid pane-churn bound surfaces the turn-end" +} + +test_turn_ended_oversized_churn_bound_surfaced() { + local dir state fakebin out drain_out capture_file window key pid + dir=$(make_case turn-ended-oversized-churn-bound); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + window="test:fm-codexoversized" + : > "$state/codexoversized.turn-ended" + printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codexoversized.meta" + printf 'rendered after the previous poll' > "$capture_file" + key=$(printf '%s' "$window" | tr ':/.' '___') + printf '%s' "$(hash_text 'the previous render')" > "$state/.hash-$key" + printf '0\n' > "$state/.count-$key" + export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_CONFIG_OVERRIDE="$(churn_config "$dir")" FM_TURNEND_CHURN_ABSORB_SECS=999999999999999999999999999999999999 \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" 2>/dev/null & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not surface a turn-end with an oversized churn bound" + grep -F "signal: $state/codexoversized.turn-ended" "$out" >/dev/null \ + || fail "watcher terminated before printing the oversized-bound turn-end" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ + || fail "drain after the oversized churn bound failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/codexoversized.turn-ended" >/dev/null \ + || fail "turn-end with an oversized churn bound was not queued" + [ ! -e "$state/.churn-since-$key" ] \ + || fail "an oversized churn bound opened a deferral window" + unset FM_FAKE_CREW_STATE + pass "an oversized pane-churn bound surfaces the turn-end" +} + +test_turn_ended_invalid_churn_deadline_surfaced() { + local variant value dir state fakebin out drain_out capture_file window key marker pid + for variant in empty leading-zero nonnumeric future overflow; do + dir=$(make_case "turn-ended-invalid-churn-deadline-$variant") + state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + window="test:fm-codexdeadline" + : > "$state/codexdeadline.turn-ended" + printf 'window=%s\nkind=ship\nharness=codex\n' "$window" > "$state/codexdeadline.meta" + printf 'rendered after the previous poll' > "$capture_file" + key=$(printf '%s' "$window" | tr ':/.' '___') + marker="$state/.churn-since-$key" + printf '%s' "$(hash_text 'the previous render')" > "$state/.hash-$key" + printf '0\n' > "$state/.count-$key" + case "$variant" in + empty) value='' ;; + leading-zero) value=09 ;; + nonnumeric) value=bogus ;; + future) value=$(( $(date +%s) + 600 )) ;; + overflow) value=999999999999999999999999999999999999 ;; + esac + printf '%s' "$value" > "$marker" + export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" 2>/dev/null & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not surface a turn-end with a $variant churn deadline" + grep -F "signal: $state/codexdeadline.turn-ended" "$out" >/dev/null \ + || fail "watcher terminated before printing the $variant-deadline turn-end" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ + || fail "drain after the $variant churn deadline failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/codexdeadline.turn-ended" >/dev/null \ + || fail "turn-end with a $variant churn deadline was not queued" + [ "$(cat "$marker")" = "$value" ] \ + || fail "the $variant churn deadline was rewritten" + done + unset FM_FAKE_CREW_STATE + pass "invalid existing pane-churn deadlines surface without mutation" +} + +test_turn_ended_surfaced_batch_opens_no_partial_deadline() { + local dir state fakebin out drain_out capture_file first_window second_window first_key second_key pid + dir=$(make_case turn-ended-no-partial-churn-deadline); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + first_window="test:fm-codexfirst"; second_window="test:fm-codexsecond" + : > "$state/first.turn-ended" + : > "$state/second.turn-ended" + printf 'window=%s\nkind=ship\nharness=codex\n' "$first_window" > "$state/first.meta" + printf 'window=%s\nkind=ship\nharness=codex\n' "$second_window" > "$state/second.meta" + printf 'rendered after the previous poll' > "$capture_file" + first_key=$(printf '%s' "$first_window" | tr ':/.' '___') + second_key=$(printf '%s' "$second_window" | tr ':/.' '___') + printf '%s' "$(hash_text 'first previous render')" > "$state/.hash-$first_key" + printf '%s' "$(hash_text 'second previous render')" > "$state/.hash-$second_key" + printf '0\n' > "$state/.count-$first_key" + printf '0\n' > "$state/.count-$second_key" + printf 'bogus' > "$state/.churn-since-$second_key" + export FM_FAKE_CREW_STATE='state: unknown · source: pane · harness state unavailable (unknown codex-unverified)' + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOWS="$(printf 'fm-codexfirst\nfm-codexsecond')" \ + FM_FAKE_TMUX_CAPTURE="$capture_file" FM_CONFIG_OVERRIDE="$(churn_config "$dir")" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=3 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" 2>/dev/null & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher absorbed a batch containing an invalid churn deadline" + grep -F "$state/first.turn-ended" "$out" >/dev/null \ + || fail "watcher did not print the first turn-end from the surfaced batch" + grep -F "$state/second.turn-ended" "$out" >/dev/null \ + || fail "watcher did not print the second turn-end from the surfaced batch" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null \ + || fail "drain after the surfaced churn batch failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/first.turn-ended" >/dev/null \ + || fail "the first turn-end from the surfaced batch was not queued" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/second.turn-ended" >/dev/null \ + || fail "the second turn-end from the surfaced batch was not queued" + [ ! -e "$state/.churn-since-$first_key" ] \ + || fail "a surfaced batch opened a partial churn deadline" + [ "$(cat "$state/.churn-since-$second_key")" = bogus ] \ + || fail "the invalid churn deadline in a surfaced batch was rewritten" + unset FM_FAKE_CREW_STATE + pass "a surfaced batch opens no partial pane-churn deadline" +} + +test_working_note_not_working_surfaced() { + local dir state fakebin out drain_out status_file pid + dir=$(make_case working-note-stopped); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out" + status_file="$state/task.status" + printf 'working: compiling step 2\n' > "$status_file" + # A non-no-mistakes crew (no run) whose pane went idle: fm-crew-state falls back + # to the stale working: status-log line. That is NOT positive evidence, so the + # wake must surface - these users must never be left hanging. + export FM_FAKE_CREW_STATE='state: working · source: status-log · working: compiling step 2' + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not surface a working: note whose crew has no running pipeline and an idle pane" + grep -F "signal: $status_file" "$out" >/dev/null || fail "watcher did not print the surfaced working: signal" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the surfaced working: note failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$status_file" >/dev/null || fail "surfaced working: note was not queued" + [ -s "$state/.seen-task_status" ] || fail "surfaced working: note did not advance its .seen-* suppressor" + pass "a no-verb working: note whose crew is idle with no running pipeline is surfaced" +} + +test_secondmate_status_note_surfaced_despite_busy_agent() { + local dir state fakebin out drain_out pid + dir=$(make_case secondmate-note-surfaced); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out" + printf 'kind=secondmate\n' > "$state/mate.meta" + printf 'working: routed reply landed in the parent stream\n' > "$state/mate.status" + # Busy evidence that would absorb an ordinary crewmate's no-verb note must + # not absorb a secondmate's: its status stream is the routed-reply channel. + export FM_FAKE_CREW_STATE='state: working · source: run-step · running' + FM_CONFIG_OVERRIDE="$(churn_config "$dir")" watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || fail "watcher absorbed a busy secondmate's routed status note" + grep -F "signal: $state/mate.status" "$out" >/dev/null \ + || fail "watcher did not print the surfaced secondmate note" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the surfaced note failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$state/mate.status" >/dev/null \ + || fail "surfaced secondmate note was not queued" + pass "a secondmate's status note surfaces even while its own agent is busy" +} + +test_self_announced_close_does_not_rewake_but_next_note_does() { + local dir state fakebin out status_file pid rc + dir=$(make_case self-close-quiet); state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out" + status_file="$state/task.status" + printf 'needs-decision [key=k1]: pick one\n' > "$status_file" + prime_status_seen "$state" "$status_file" || fail "could not prime the announced baseline" + # The home's own bookkeeping close, written through the guarded + # self-announced append this home's answerers use. + rc=0 + FM_STATE_OVERRIDE="$state" bash -c ' + . "$1" + fm_wake_status_append_self_announced "$2" "$3" "resolved [key=k1]: answered: closed by this home" + ' _ "$ROOT/bin/fm-wake-lib.sh" "$state" "$status_file" || rc=$? + [ "$rc" -eq 0 ] || fail "the bookkeeping close was not self-announced (rc=$rc)" + export FM_FAKE_CREW_STATE='state: unknown · source: none · idle worker' + watch_bg "$state" "$fakebin" "$out" + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid"; fail "the home's own bookkeeping close re-woke its own watcher: $(cat "$out")" + fi + [ ! -s "$out" ] || { reap "$pid"; fail "self-announced close printed a wake reason: $(cat "$out")"; } + [ ! -s "$state/.wake-queue" ] || { reap "$pid"; fail "self-announced close enqueued a durable wake"; } + # A later, different note on the SAME task still wakes: dedup is keyed on the + # exact announced bytes, never on task identity. + printf 'needs-decision [key=k2]: a genuinely new decision\n' >> "$status_file" + wait_for_exit "$pid" 100 || fail "a later different note after a self-announced close was swallowed" + grep -F "signal: $status_file" "$out" >/dev/null \ + || fail "the later note did not surface as a signal" + pass "a self-announced close never wakes its own home, and the next real note still does" +} + +# --- actionable wakes are surfaced (queue + exit) --------------------------- + +test_actionable_signal_surfaced() { + local dir state fakebin out drain_out status_file pid + dir=$(make_case actionable-signal); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out" + status_file="$state/task.status" + printf 'working: setup\nneeds-decision: pick A or B\n' > "$status_file" + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not exit for an actionable needs-decision signal" + grep -F "signal: $status_file" "$out" >/dev/null || fail "watcher did not print the actionable signal reason" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the actionable signal failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$status_file" >/dev/null || fail "actionable signal was not queued" + [ -s "$state/.hb-surfaced-task" ] || fail "actionable signal did not record the surfaced marker" + pass "captain-relevant signal is surfaced (queue + exit) and marked surfaced" +} + +# A needs-decision status append surfaced through this actionable signal path +# must skip the Pi supervision branch and reach main directly +# (docs/pi-supervision-branch.md "Autonomy"). The row still +# queues as an ordinary signal-kind wake - fm-branch-dispatch.ts's +# scopeForUnreadWake tells it apart from a routine signal by this payload +# marker, not by kind. +test_needs_decision_signal_payload_marked_for_branch_exclusion() { + local dir state fakebin out status_file pid + dir=$(make_case needs-decision-payload); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out" + status_file="$state/task.status" + printf 'working: setup\nneeds-decision: pick A or B\n' > "$status_file" + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not exit for an actionable needs-decision signal" + grep -F "$(printf 'signal\ttask.status\tneeds-decision:')" "$state/.wake-queue" >/dev/null \ + || fail "a needs-decision signal row was not payload-marked for branch exclusion: $(cat "$state/.wake-queue")" + pass "a needs-decision signal row's queued payload is marked needs-decision: for branch exclusion" +} + +# A needs-decision whose key transition was rejected by the reserved-key +# vocabulary is reported as a "reconciliation-required: " wrapped event +# (fm-classify-lib.sh's status_span_first_actionable_record), but it is still a +# needs-decision signal that this path routes directly to main - the payload +# marker must not be fooled by that wrapper. +test_needs_decision_reconciliation_required_still_marked() { + local dir state fakebin out status_file pid + dir=$(make_case needs-decision-reconciliation); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out" + status_file="$state/task.status" + printf 'needs-decision [key=pending-reply-x]: unrelated request\nworking: awaiting reconciliation\n' \ + > "$status_file" + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not exit for a rejected-reserved-key needs-decision" + grep -F "$(printf 'signal\ttask.status\tneeds-decision:')" "$state/.wake-queue" >/dev/null \ + || fail "a reconciliation-required needs-decision row was not payload-marked for branch exclusion: $(cat "$state/.wake-queue")" + pass "a reconciliation-required needs-decision row's queued payload is still marked needs-decision:" +} + +# A captain-held declaration is itself actionable. Positive evidence that the +# crew is still working must not absorb the signal before its main-only marker +# can be delivered. +test_captain_held_signal_payload_marked_for_branch_exclusion() { + local dir state fakebin out status_file pid + dir=$(make_case captain-held-signal-payload); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out" + status_file="$state/task.status" + printf 'captain-held [key=route]: awaiting the captain\n' > "$status_file" + export FM_FAKE_CREW_STATE='state: working · source: run-step · still wrapping up' + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || fail "watcher absorbed a captain-held signal while the crew was still working" + grep -F "signal: $status_file" "$out" >/dev/null \ + || fail "a captain-held signal changed its wake reason: $(cat "$out")" + grep -F "$(printf 'signal\ttask.status\tneeds-decision:')" "$state/.wake-queue" >/dev/null \ + || fail "a captain-held signal was not payload-marked for branch exclusion: $(cat "$state/.wake-queue")" + pass "a captain-held signal stays actionable while the crew is still working" +} + +test_pending_reply_escalation_signal_payload_marked_for_branch_exclusion() { + local dir state fakebin out status_file pid corr + dir=$(make_case pending-reply-escalation-payload); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out" + status_file="$state/task.status" + corr=0123456789abcdef + printf 'blocked [key=pending-reply-%s]: pending-reply-missed: task=task pending-reply-id=%s request=finish report\n' \ + "$corr" "$corr" > "$status_file" + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not exit for a pending-reply escalation" + grep -F "$(printf 'signal\ttask.status\tneeds-decision:')" "$state/.wake-queue" >/dev/null \ + || fail "a pending-reply escalation was not payload-marked for branch exclusion: $(cat "$state/.wake-queue")" + pass "a pending-reply second-mate escalation is marked for main-only routing" +} + +test_ordinary_blocked_signal_payload_remains_branch_eligible() { + local dir state fakebin out status_file pid + dir=$(make_case ordinary-blocked-payload); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out" + status_file="$state/task.status" + printf 'blocked [key=dependency]: waiting for an upstream release\n' > "$status_file" + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not exit for an ordinary blocked event" + grep -F "$(printf 'signal\ttask.status\tsignal:')" "$state/.wake-queue" >/dev/null \ + || fail "an ordinary blocked event lost branch-eligible routing: $(cat "$state/.wake-queue")" + if grep -F "$(printf 'signal\ttask.status\tneeds-decision:')" "$state/.wake-queue" >/dev/null; then + fail "an ordinary blocked event was marked as a second-mate escalation" + fi + pass "an ordinary blocked event remains branch-eligible" +} + +# A routine (non-needs-decision) captain-relevant event must keep its ordinary +# payload: only a genuine needs-decision gets the exclusion marker. +test_routine_signal_payload_not_marked_needs_decision() { + local dir state fakebin out status_file pid + dir=$(make_case routine-signal-payload); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out" + status_file="$state/task.status" + printf 'working: setup\ndone: migration complete ; needs-decision: documented in follow-up\n' > "$status_file" + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not exit for an actionable done signal" + grep -F "$(printf 'signal\ttask.status\tneeds-decision:')" "$state/.wake-queue" >/dev/null \ + && fail "a routine done signal was incorrectly payload-marked needs-decision: $(cat "$state/.wake-queue")" + grep -F "$(printf 'signal\ttask.status\tsignal:')" "$state/.wake-queue" >/dev/null \ + || fail "a routine signal lost its ordinary payload: $(cat "$state/.wake-queue")" + pass "a routine event containing a needs-decision phrase keeps its ordinary payload, unmarked" +} + +# The reported bug, end to end through a real watcher: a crew reports something +# the captain must act on and then keeps appending routine progress, which is +# ordinary while the watcher lingers its signal grace window to coalesce a status +# write with the same turn's turn-end. Classifying only the last line reads the +# batch as routine, and because the crew IS provably working the no-verb fallback +# absorbs it too - the .seen-* suppressor then advances and nothing ever re-reads +# the event, so the work stalls with the captain never told. +test_actionable_signal_survives_a_later_routine_append() { + local dir state fakebin out drain_out status_file sig pid + dir=$(make_case actionable-masked); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out" + status_file="$state/task.status" + # Everything through "working: setup" was already classified, so this asserts + # the newly appended span, not merely a whole-file re-read. + printf 'working: setup\n' > "$status_file" + sig=$(seen_sig "$status_file"); printf '%s' "$sig" > "$state/.seen-task_status" + printf 'needs-decision: pick A or B\nworking: still tidying the branch\n' >> "$status_file" + # Positive evidence the crew is still working, so the no-verb fallback cannot + # rescue the wake: only reading the event itself can surface it. + export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 \ + || { reap "$pid"; fail "watcher absorbed a needs-decision hidden behind a later working: line"; } + grep -F "signal: $status_file" "$out" >/dev/null || fail "watcher did not print the actionable signal reason" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the masked signal failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$status_file" >/dev/null \ + || fail "the masked actionable signal was not queued" + unset FM_FAKE_CREW_STATE + pass "a captain event hidden behind a later routine append is still surfaced (queue + exit)" +} + +# The captain-reported completion shape of the same masking, end to end. +test_release_completion_survives_a_later_routine_append() { + local dir state fakebin out drain_out status_file sig pid + dir=$(make_case release-masked); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out" + status_file="$state/task.status" + printf 'working: publishing\n' > "$status_file" + sig=$(seen_sig "$status_file"); printf '%s' "$sig" > "$state/.seen-task_status" + printf 'done: release 1.4.0 published and installed\nworking: cleaning the build dir\n' >> "$status_file" + export FM_FAKE_CREW_STATE='state: working · source: pane · harness busy' + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 \ + || { reap "$pid"; fail "watcher absorbed a release/install completion hidden behind later cleanup chatter"; } + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the masked completion failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$status_file" >/dev/null \ + || fail "the masked completion was not queued" + unset FM_FAKE_CREW_STATE + pass "a finished release reported before routine cleanup chatter is still surfaced" +} + +# The other direction: the fix must not turn ordinary progress into wakes. +test_routine_appends_after_a_classified_event_stay_absorbed() { + local dir state fakebin out status_file sig pid + dir=$(make_case actionable-classified); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out" + status_file="$state/task.status" + # The decision is BEHIND the classified position, so only the new routine line + # is in the span. A supervisor that re-read the whole log would wake again here. + printf 'working: setup\nneeds-decision: pick A or B\n' > "$status_file" + sig=$(seen_sig "$status_file"); printf '%s' "$sig" > "$state/.seen-task_status" + printf 'working: still tidying the branch\n' >> "$status_file" + export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' + watch_bg "$state" "$fakebin" "$out" + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid"; fail "watcher re-surfaced a decision it had already classified: $(cat "$out")" + fi + [ ! -s "$state/.wake-queue" ] || fail "a routine append after a classified decision enqueued a wake" + reap "$pid" + unset FM_FAKE_CREW_STATE + pass "a routine append after an already-classified event is absorbed (no re-wake)" +} + +test_unreadable_status_reports_once_per_file_state() { + local dir state fakebin out status_file target marker sig pid + dir=$(make_case unreadable-status); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; status_file="$state/task.status"; target="$dir/missing-status-target" + ln -s "$target" "$status_file" + marker="$state/.seen-task_status" + + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || { reap "$pid"; fail "a dangling status symlink was not reported"; } + grep -Fx "signal: $status_file" "$out" >/dev/null \ + || fail "a dangling status symlink did not use the immediate signal path: $(cat "$out")" + sig=$(status_observed_signature "$status_file") + status_presentation_marker_reported_matches "$marker" "$sig" \ + || fail "the unreadable status report did not advance its wake signature" + [ "$(status_presentation_marker_offset "$marker" "$status_file")" = 0 ] \ + || fail "the unreadable status report advanced its classification position" + ack_stopped_cycle "$state" || fail "could not acknowledge the first unreadable-status wake" + touch "$state/.last-check" "$state/.last-heartbeat" + + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_poll_cycle "$state" "$pid" \ + || { reap "$pid"; fail "an unchanged unreadable status reported again after restart: $(cat "$out")"; } + reap "$pid" + + printf 'blocked: changed target state with a longer path\n' > "$dir/status-target-two-longer" + ln -snf "$dir/status-target-two-longer" "$status_file" + target="$dir/status-target-two-longer" + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || { reap "$pid"; fail "a changed unreadable status did not report again"; } + [ "$(status_presentation_marker_offset "$marker" "$status_file")" = 0 ] \ + || fail "a changed unreadable status advanced its classification position" + ack_stopped_cycle "$state" || fail "could not acknowledge the changed unreadable-status wake" + + rm -f "$status_file" + cp "$target" "$status_file" + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || { reap "$pid"; fail "a readable replacement did not surface preserved content"; } + [ "$(status_presentation_marker_offset "$marker" "$status_file")" = "$(size_of "$status_file")" ] \ + || fail "readable recovery did not classify content written before the failure" + pass "unreadable status reports are bounded without advancing classification" +} + +test_permission_recovery_surfaces_preserved_status() { + local dir state fakebin out status_file marker before_ident after_ident pid + dir=$(make_case permission-recovery); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; status_file="$state/task.status"; marker="$state/.seen-task_status" + printf 'blocked: release approval required\nworking: preserving context\n' > "$status_file" + before_ident=$(_fm_open_decisions_file_ident "$status_file") + chmod 000 "$status_file" + if [ -r "$status_file" ]; then + chmod 600 "$status_file" + pass "permission recovery skipped because permissions cannot deny reads" + return + fi + + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || { reap "$pid"; chmod 600 "$status_file"; fail "an unreadable regular status was not reported"; } + [ "$(status_presentation_marker_offset "$marker" "$status_file")" = 0 ] \ + || { chmod 600 "$status_file"; fail "an unreadable regular status advanced its classification position"; } + ack_stopped_cycle "$state" || { chmod 600 "$status_file"; fail "could not acknowledge the unreadable regular-status wake"; } + touch "$state/.last-check" "$state/.last-heartbeat" + + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_poll_cycle "$state" "$pid" \ + || { reap "$pid"; chmod 600 "$status_file"; fail "an unchanged unreadable regular status reported again"; } + + chmod 600 "$status_file" + after_ident=$(_fm_open_decisions_file_ident "$status_file") + [ "$after_ident" = "$before_ident" ] || { reap "$pid"; fail "the permission-only recovery changed file identity"; } + wait_for_exit "$pid" 100 || { reap "$pid"; fail "readability recovery did not surface preserved content"; } + grep -Fx "signal: $status_file" "$out" >/dev/null \ + || fail "readability recovery did not use the actionable signal path: $(cat "$out")" + [ "$(status_presentation_marker_offset "$marker" "$status_file")" = "$(size_of "$status_file")" ] \ + || fail "readability recovery did not classify from the unadvanced position" + pass "permission recovery surfaces content from the unadvanced position" +} + +test_terminal_stale_surfaced() { + local dir state fakebin out drain_out capture_file window key pane_hash sig pid + dir=$(make_case terminal-stale); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + window="test:fm-done" + printf 'finished, awaiting review' > "$capture_file" + printf 'window=%s\nkind=ship\n' "$window" > "$state/done.meta" + printf 'done: PR https://example.test/pr/3\n' > "$state/done.status" + sig=$(seen_sig "$state/done.status"); printf '%s' "$sig" > "$state/.seen-done_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "finished, awaiting review") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not exit for a stale pane on a terminal status" + grep -Fx "stale: $window" "$out" >/dev/null || fail "watcher did not print the terminal stale wake" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the terminal stale failed" + grep "$(printf '\tstale\t')" "$drain_out" | grep -F "$window" >/dev/null || fail "terminal stale was not queued" + pass "a stale pane sitting on a terminal status is surfaced (queue + exit)" +} + +# --- stale pane, STALE terminal status overridden by an active run: absorbed --- +# Regression for the 2026-07 herdr false-surface incidents: a crew's own status +# log gets no new entry once firstmate hands it to a no-mistakes validation +# (AGENTS.md's sparse status-reporting contract), so the log keeps showing its +# pre-validation "done:" line as the LAST line for the run's entire (possibly +# many-minutes) duration. stale_is_terminal alone has no run-step awareness and +# would treat that leftover as still-current every time the pane goes quiet, +# immediately surfacing a crew that is actively validating. crew_is_provably_working +# must get a chance to override a captain-relevant-but-stale status line, exactly +# as it already does for a plain non-terminal one. +test_stale_terminal_status_overridden_by_active_run() { + local dir state fakebin out drain_out capture_file window key pane_hash sig pid + dir=$(make_case terminal-stale-overridden); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + window="test:fm-validating" + printf 'no-mistakes axi run: validating...' > "$capture_file" + printf 'window=%s\nkind=ship\n' "$window" > "$state/validating.meta" + # The crew reported done BEFORE firstmate triggered no-mistakes validation; + # this line never gets superseded by a newer status-log entry while the + # pipeline itself runs. + printf 'done: implementation complete, ready to validate\n' > "$state/validating.status" + sig=$(seen_sig "$state/validating.status"); printf '%s' "$sig" > "$state/.seen-validating_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "no-mistakes axi run: validating...") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' + + # Phase A: a high escalation threshold means the first sighting is absorbed, + # not surfaced, despite the captain-relevant "done:" status-log line. + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid"; fail "watcher exited for a stale terminal-looking status the run-step overrides (should absorb): $(cat "$out")" + fi + [ ! -s "$out" ] || fail "the overridden stale terminal status printed a wake reason during absorb" + [ ! -s "$state/.wake-queue" ] || fail "the overridden stale terminal status enqueued a wake during absorb" + [ "$(cat "$state/.stale-$key" 2>/dev/null || true)" = "$pane_hash" ] || fail "stale suppressor not advanced on absorb" + [ -s "$state/.stale-since-$key" ] || fail "stale-since escalation timer was not recorded on absorb" + [ ! -e "$state/.hb-surfaced-validating" ] || fail "an absorbed wake must not mark the status line as surfaced" + reap "$pid" + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional phase-A watcher stop" + + # Phase B: backdate the idle timer past the threshold; the run genuinely + # wedges and the next poll escalates exactly like the non-terminal case. + echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not escalate an overridden stale terminal status past the threshold" + grep -F "stale: $window" "$out" >/dev/null || fail "escalation did not print a stale wake" + grep -F "possible wedge" "$out" >/dev/null || fail "escalation did not flag a possible wedge" + unset FM_FAKE_CREW_STATE + pass "a stale terminal-looking status is overridden and absorbed while a run is actively working, then wedge-escalated" +} + +test_explicit_decision_view_preserves_the_durable_fold() { + local dir file explicit retained + dir=$(make_case explicit-decision-view); file="$dir/state/view.status" + printf '%s\n' 'blocked: unkeyed dependency' 'needs-decision: [key=choice] explicit choice' > "$file" + retained=$(status_open_decisions "$file") + explicit=$(status_open_decisions "$file" --explicit-only) + assert_contains "$retained" $'default\tblocked\t' "unkeyed blocker must remain durably open" + assert_contains "$explicit" $'choice\tneeds-decision\t' "note-head explicit key must remain eligible" + assert_not_contains "$explicit" $'default\t' "implicit default must not suppress repaint alarms" + printf '%s\n' 'blocked [key=default]: explicitly named default' >> "$file" + explicit=$(status_open_decisions "$file" --explicit-only) + assert_contains "$explicit" $'default\tblocked\t' "literal default key is explicitly keyed" + printf '%s\n' 'blocked: replacement without a key' >> "$file" + explicit=$(status_open_decisions "$file" --explicit-only) + assert_not_contains "$explicit" $'default\t' "unkeyed replacement must lose explicit provenance" + assert_contains "$(status_open_decisions "$file")" 'replacement without a key' "filter must not close a durable blocker" + printf '%s\n' 'resolved: [key=choice] choice handled' >> "$file" + [ -z "$(status_open_decisions "$file" --explicit-only)" ] || fail "resolved explicit key remained eligible" + pass "explicit-key view preserves durable unkeyed retention and default-key provenance" +} + +# --- a parked keyed decision is surfaced once, not once per pane repaint ------- +# A parked decision is genuinely not working, so crew_is_provably_working rightly +# refuses to absorb its first stale pane. The separate repaint suppressor must key +# on the complete open-decision set so volatile footer changes alone stay silent. +# +# Fixture note: each phase primes .hash-/.count- so the FIRST poll already sees a +# stably stale pane, exactly as the terminal-stale cases above do. +prime_stale_pane() { # <state> <window> <pane-text> <capture-file> + local state=$1 window=$2 text=$3 capture=$4 key + printf '%s' "$text" > "$capture" + key=$(printf '%s' "$window" | tr ':/.' '___') + printf '%s' "$(hash_text "$text")" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" +} + +# Prime the .seen-* signal suppressor for a status file so only the STALE path +# under test can surface a wake (a status write otherwise fires its own signal). + +test_parked_decision_survives_pane_repaint() { + local dir state fakebin out drain_out capture window pid + dir=$(make_case parked-decision-repaint); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture="$dir/pane.txt" + window="test:fm-parked" + printf 'window=%s\nkind=ship\n' "$window" > "$state/parked.meta" + printf 'needs-decision [key=review-gate]: ship as-is or split the migration\n' \ + > "$state/parked.status" + # A live agent: this crew is waiting, not wedged. + export FM_FAKE_TMUX_CURRENT_COMMAND=claude + + # Phase A: the normal first sighting is the status signal. + printf '%s' 'awaiting decision · context 41%' > "$capture" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not surface a newly parked captain decision" + grep -F "signal: $state/parked.status" "$out" >/dev/null \ + || fail "the parked decision did not print its initial status signal" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the parked decision failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "parked.status" >/dev/null \ + || fail "the parked decision signal was not queued on its first sighting" + # Presented records stay durable until acknowledged, so phase B's "nothing new + # was queued" assertion needs phase A's own record retired first. This retires + # the record this phase just handled; it never suppresses a later one. + ack_stopped_cycle "$state" || fail "could not acknowledge the parked decision's first sighting" + + # Phase B: the pane repaints (a moving context percentage) while the SAME + # decision stays open. The hash changes; the decision does not. Nothing may + # surface. + : > "$out" + prime_stale_pane "$state" "$window" 'awaiting decision · context 39%' "$capture" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + if ! wait_live "$pid" 30; then + reap "$pid"; fail "a pane repaint re-surfaced an already-escalated parked decision: $(cat "$out")" + fi + [ ! -s "$out" ] || fail "a pane repaint printed a wake for an unchanged parked decision" + [ ! -s "$state/.wake-queue" ] || fail "a pane repaint queued a wake for an unchanged parked decision" + reap "$pid" + unset FM_FAKE_TMUX_CURRENT_COMMAND + assert_contains "$(status_open_decisions "$state/parked.status")" $'review-gate\t' \ + "acknowledging a wake must not close the keyed decision" + pass "a parked keyed decision surfaces once, not once per pane repaint" +} + +test_resolved_decision_can_reopen_identically() { + local dir state fakebin out capture window pid + dir=$(make_case parked-decision-reopen); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture="$dir/pane.txt"; window="test:fm-reopen" + printf 'window=%s\nkind=ship\n' "$window" > "$state/reopen.meta" + printf 'needs-decision [key=review-gate]: ship as-is or split the migration\n' \ + > "$state/reopen.status" + prime_status_seen "$state" "$state/reopen.status" + export FM_FAKE_TMUX_CURRENT_COMMAND=claude + + prime_stale_pane "$state" "$window" 'awaiting decision · context 41%' "$capture" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not surface the original keyed decision" + FM_STATE_OVERRIDE="$state" "$DRAIN" > /dev/null 2>&1 || true + + printf 'resolved [key=review-gate]: migration will ship as-is\n' >> "$state/reopen.status" + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "the decision resolution did not surface" + grep -F "signal: $state/reopen.status" "$out" >/dev/null || fail "the resolution did not print a signal wake" + FM_STATE_OVERRIDE="$state" "$DRAIN" > /dev/null 2>&1 || true + + printf 'needs-decision [key=review-gate]: ship as-is or split the migration\n' \ + >> "$state/reopen.status" + prime_status_seen "$state" "$state/reopen.status" + prime_stale_pane "$state" "$window" 'awaiting decision again · context 38%' "$capture" + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "an identical decision did not surface after resolving and reopening" + grep -Fx "stale: $window" "$out" >/dev/null || fail "the identically reopened decision did not print a stale wake" + unset FM_FAKE_TMUX_CURRENT_COMMAND + pass "a resolved keyed decision can reopen identically and surface again" +} + +# The suppressor must be scoped to the open-decision set that was surfaced, +# never to the pane: a genuinely NEW decision on the same repainting pane is new +# work firstmate has not acted on and must wake normally. The status write's own +# signal wake is suppressed here so only the stale path can surface it. +test_new_keyed_decision_on_parked_pane_surfaces() { + local dir state fakebin out drain_out capture window pid + dir=$(make_case parked-decision-new-key); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture="$dir/pane.txt" + window="test:fm-parked2" + printf 'window=%s\nkind=ship\n' "$window" > "$state/parked2.meta" + printf 'needs-decision [key=review-gate]: ship as-is or split the migration\n' \ + > "$state/parked2.status" + prime_status_seen "$state" "$state/parked2.status" + export FM_FAKE_TMUX_CURRENT_COMMAND=claude + + prime_stale_pane "$state" "$window" 'awaiting decision · context 41%' "$capture" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not surface the first parked decision" + FM_STATE_OVERRIDE="$state" "$DRAIN" > /dev/null 2>&1 || true + + # A second, different decision opens on the same pane, which also repaints. + printf 'needs-decision [key=schema-choice]: single table or one per tenant\n' \ + >> "$state/parked2.status" + prime_status_seen "$state" "$state/parked2.status" + prime_stale_pane "$state" "$window" 'awaiting decision · context 38%' "$capture" + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "a NEW keyed decision on a parked pane did not surface" + grep -Fx "stale: $window" "$out" >/dev/null || fail "the new keyed decision did not print a stale wake" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the new keyed decision failed" + grep "$(printf '\tstale\t')" "$drain_out" | grep -F "$window" >/dev/null \ + || fail "the new keyed decision was not queued" + unset FM_FAKE_TMUX_CURRENT_COMMAND + pass "a new keyed decision on an already-parked pane still surfaces" +} + +# Suppressing the repeat surface must not cost wedge detection: a crew parked on +# an open decision whose agent then dies is a wedge suspect and must escalate +# through the shared timer. Only a confident dead verdict counts, so the live +# agent above stays absorbed while a shell-at-the-prompt endpoint escalates. +test_parked_decision_with_dead_agent_wedge_escalates() { + local dir state fakebin out capture window key pid + dir=$(make_case parked-decision-dead); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture="$dir/pane.txt" + window="test:fm-parked3" + key=$(printf '%s' "$window" | tr ':/.' '___') + printf 'window=%s\nkind=ship\n' "$window" > "$state/parked3.meta" + printf 'needs-decision [key=review-gate]: ship as-is or split the migration\n' \ + > "$state/parked3.status" + prime_status_seen "$state" "$state/parked3.status" + export FM_FAKE_TMUX_CURRENT_COMMAND=claude + + prime_stale_pane "$state" "$window" 'awaiting decision · context 41%' "$capture" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not surface the parked decision before the wedge case" + FM_STATE_OVERRIDE="$state" "$DRAIN" > /dev/null 2>&1 || true + + # The agent exits, leaving a bare shell at the endpoint, and the pane repaints + # once more. The decision is unchanged, so the repeat surface stays suppressed, + # but the wedge timer - backdated past the threshold - must still escalate. + export FM_FAKE_TMUX_CURRENT_COMMAND=zsh + prime_stale_pane "$state" "$window" 'awaiting decision · context 37%' "$capture" + echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "a parked crew whose agent died was not detected as a wedge suspect" + grep -F "stale: $window" "$out" >/dev/null || fail "the dead parked crew did not print a stale wake" + grep -F "possible wedge" "$out" >/dev/null || fail "the dead parked crew was not flagged as a possible wedge" + unset FM_FAKE_TMUX_CURRENT_COMMAND + pass "a parked crew that stops responding is still detected as a wedge suspect" +} + +# --- non-terminal stale, crew provably working: absorbed, then wedge-escalated --- +# A provably-working crew (an actively-running pipeline) legitimately sits on a +# static pane (e.g. waiting on CI), so a non-terminal stale is absorbed and only +# the wedge timer eventually escalates it - the low-churn behavior preserved. + +test_nonterminal_stale_provably_working_absorbed_then_escalated() { + local dir state fakebin out drain_out capture_file window key pane_hash sig pid + dir=$(make_case nonterminal-stale-working); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + window="test:fm-quiet" + printf 'idle building output' > "$capture_file" + printf 'window=%s\nkind=ship\n' "$window" > "$state/quiet.meta" + # Non-terminal status, and prime .seen-* so the signal scan does not pre-empt + # the stale path. + printf 'working: still compiling\n' > "$state/quiet.status" + sig=$(seen_sig "$state/quiet.status"); printf '%s' "$sig" > "$state/.seen-quiet_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "idle building output") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + # The crew's pipeline is actively running: a static pane is normal (waiting on CI). + export FM_FAKE_CREW_STATE='state: working · source: run-step · ci running' + + # Phase A: a high escalation threshold means the first sighting is absorbed. + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid"; fail "watcher exited for a fresh provably-working non-terminal stale (should absorb): $(cat "$out")" + fi + [ ! -s "$out" ] || fail "fresh provably-working stale printed a wake reason during absorb" + [ ! -s "$state/.wake-queue" ] || fail "fresh provably-working stale enqueued a wake during absorb" + [ "$(cat "$state/.stale-$key" 2>/dev/null || true)" = "$pane_hash" ] || fail "stale suppressor not advanced on absorb" + [ -s "$state/.stale-since-$key" ] || fail "stale-since escalation timer was not recorded on absorb" + reap "$pid" + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional phase-A watcher stop" + + # Phase B: backdate the idle timer past the threshold; the next run escalates. + # (The subsequent-sight timer path does not re-read the crew state.) + echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not escalate a provably-working non-terminal stale past the threshold" + grep -F "stale: $window" "$out" >/dev/null || fail "escalation did not print a stale wake" + grep -F "possible wedge" "$out" >/dev/null || fail "escalation did not flag a possible wedge" + [ ! -e "$state/.stale-since-$key" ] || fail "stale-since timer was not cleared after escalation" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the wedge escalation failed" + grep "$(printf '\tstale\t')" "$drain_out" | grep -F "$window" >/dev/null || fail "wedge escalation was not queued" + pass "provably-working non-terminal stale is absorbed on first sight, then wedge-escalated past the threshold" +} + +# --- non-terminal stale, crew NOT provably working: surfaced immediately ------ +# The key requirement: a crew with no running pipeline that has gone quiet (and is +# not busy) has stopped - it may be done via interactive menus, waiting, or wedged. +# It must surface at once, never wait out the wedge timer, so these users (a +# non-no-mistakes crew, or any crew with no running pipeline) are never left hanging. + +test_nonterminal_stale_not_working_surfaced() { + local dir state fakebin out drain_out capture_file window key pane_hash sig pid + dir=$(make_case nonterminal-stale-stopped); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + window="test:fm-stopped" + printf 'idle prompt, finished' > "$capture_file" + printf 'window=%s\nkind=ship\n' "$window" > "$state/stopped.meta" + # Non-terminal status (the crew never wrote a captain-relevant verb), .seen-* + # primed so the signal scan does not pre-empt the stale path. + printf 'working: implementing\n' > "$state/stopped.status" + sig=$(seen_sig "$state/stopped.status"); printf '%s' "$sig" > "$state/.seen-stopped_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "idle prompt, finished") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + # No running pipeline; the pane is idle. NOT provably working. + export FM_FAKE_CREW_STATE='state: unknown · source: none · no current-state source available' + + # Even with a high wedge threshold, a not-provably-working stale surfaces at once. + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not surface a not-provably-working non-terminal stale at once" + grep -Fx "stale: $window" "$out" >/dev/null || fail "watcher did not print the immediate stale wake" + grep -F "possible wedge" "$out" >/dev/null && fail "an immediate stopped-crew stale was mislabeled a wedge" + [ "$(cat "$state/.stale-$key" 2>/dev/null || true)" = "$pane_hash" ] || fail "stale suppressor was not advanced on surface" + [ ! -e "$state/.stale-since-$key" ] || fail "stale-since timer should not be set when surfacing immediately" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the immediate stale failed" + grep "$(printf '\tstale\t')" "$drain_out" | grep -F "$window" >/dev/null || fail "immediate stale wake was not queued" + pass "a not-provably-working non-terminal stale is surfaced immediately (never left to wait out the timer)" +} + +# --- non-terminal stale, crew DECLARED a pause: absorbed, re-surfaced on a long +# cadence, never wedge-escalated ------------------------------------------ +# The live 2026-07-09/10 case: a crew intentionally held awaiting an upstream tool +# release (paused: ...) whose idle pane tripped repeated possible-wedge escalations +# all day. With the paused verb, its stale is absorbed like a working crew but never +# uses the wedge timer; it re-surfaces once past PAUSE_RESURFACE_SECS (anchored on +# the pause's own status-file age, so a churny idle pane cannot reset the cadence) +# for a recheck, so a forgotten pause cannot rot invisibly. +test_nonterminal_stale_paused_absorbed_then_resurfaced() { + local dir state fakebin out drain_out capture_file window key pane_hash sig pid back statusf + dir=$(make_case nonterminal-stale-paused); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + window="test:fm-held" + printf 'idle, holding for upstream' > "$capture_file" + printf 'window=%s\nkind=ship\n' "$window" > "$state/held.meta" + statusf="$state/held.status" + # A DECLARED pause (not captain-relevant), .seen-* primed so the signal scan does + # not pre-empt the stale path. + printf 'paused: holding for the upstream tool release\n' > "$statusf" + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-held_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "idle, holding for upstream") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + # crew_absorb_class reads the declared pause from fm-crew-state.sh. + export FM_FAKE_CREW_STATE='state: paused · source: status-log · holding for the upstream tool release' + + # Phase A: a fresh pause (status file just written) under a high re-surface + # threshold is absorbed - no wake, no wedge timer. + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_TMUX_CURRENT_COMMAND=zsh \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid"; fail "watcher exited for a fresh declared pause (should absorb): $(cat "$out")" + fi + [ ! -s "$out" ] || fail "fresh paused stale printed a wake reason during absorb" + [ ! -s "$state/.wake-queue" ] || fail "fresh paused stale enqueued a wake during absorb" + [ "$(cat "$state/.stale-$key" 2>/dev/null || true)" = "$pane_hash" ] || fail "stale suppressor not advanced on paused absorb" + [ -e "$state/.paused-$key" ] || fail "paused flag not recorded on absorb" + [ ! -e "$state/.stale-since-$key" ] || fail "a paused absorb must not start the wedge timer" + reap "$pid" + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional paused phase-A stop" + + # Phase B: age the pause past the (now normal) threshold by backdating its + # status file, re-prime .seen-* to the new signature so the signal scan stays + # quiet, and confirm it re-surfaces as a paused recheck - never a wedge. + back=$(( $(date +%s) - 500 )) + if [ "$(uname)" = Darwin ]; then touch -mt "$(date -r "$back" '+%Y%m%d%H%M.%S')" "$statusf" + else touch -m -d "@$back" "$statusf"; fi + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-held_status" + : > "$out" + printf 'idle, holding for upstream (token 2)' > "$capture_file" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_TMUX_CURRENT_COMMAND=zsh \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not re-surface a declared pause past the threshold" + grep -F "stale: $window" "$out" >/dev/null || fail "re-surface did not print a stale wake" + grep -F "awaiting external" "$out" >/dev/null || fail "re-surface was not labeled a paused/awaiting-external recheck" + grep -F "possible wedge" "$out" >/dev/null && fail "a declared pause was mislabeled a possible wedge" + [ -e "$state/.paused-resurfaced-$key" ] || fail "the paused re-surface throttle marker was not recorded" + [ ! -e "$state/.stale-since-$key" ] || fail "a paused re-surface must not use the wedge timer" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the paused re-surface failed" + grep "$(printf '\tstale\t')" "$drain_out" | grep -F "$window" >/dev/null || fail "paused re-surface was not queued" + pass "a declared pause is absorbed on first sight, then re-surfaced as a recheck past the threshold, never wedge-escalated" +} + +# --- pause re-surface backoff ------------------------------------------------ +# The window widens per unchanged recheck and then stops widening, so a long +# healthy wait gets progressively cheap without ever becoming invisible. +test_pause_resurface_window_backs_off_and_caps() { + local w0 w1 w3 w9 wjunk wbig whuge + # shellcheck disable=SC2034 # Read by pause_resurface_window in the sourced fm-classify-lib.sh. + FM_PAUSE_RESURFACE_SECS=100 + w0=$(pause_resurface_window 0) + w1=$(pause_resurface_window 1) + w3=$(pause_resurface_window 3) + w9=$(pause_resurface_window 9) + wjunk=$(pause_resurface_window "") + # A cap large enough to overflow the shift must fail toward the base cadence, + # not into a negative window - which every age comparison reads as due, turning + # the backoff into a wake on every single poll. + # shellcheck disable=SC2034 # Read by pause_resurface_window in the sourced fm-classify-lib.sh. + FM_PAUSE_RESURFACE_MAX_STREAK=64 + wbig=$(pause_resurface_window 64) + # shellcheck disable=SC2034 # Read by pause_resurface_window in the sourced fm-classify-lib.sh. + FM_PAUSE_RESURFACE_SECS=3600 + # shellcheck disable=SC2034 # Read by pause_resurface_window in the sourced fm-classify-lib.sh. + FM_PAUSE_RESURFACE_MAX_STREAK=52 + whuge=$(pause_resurface_window 52) + unset FM_PAUSE_RESURFACE_SECS FM_PAUSE_RESURFACE_MAX_STREAK + [ "$w0" = 100 ] || fail "a first recheck must use the base window, got $w0" + [ "$w1" = 200 ] || fail "the second recheck must double the window, got $w1" + [ "$w3" = 800 ] || fail "the fourth recheck must be 8x the base window, got $w3" + [ "$w9" = 800 ] || fail "the window must stop widening at the cap, got $w9" + [ "$wjunk" = 100 ] || fail "a missing streak must fall back to the base window, got $wjunk" + [ "$wbig" -ge 100 ] || fail "an overflowing cap must never yield a window below the base, got $wbig" + [ "$wbig" = 86400 ] || fail "an overflowing cap must clamp to the bounded ceiling, got $wbig" + [ "$whuge" = 86400 ] || fail "an overflowing cap must clamp to the bounded ceiling, got $whuge" + pass "pause_resurface_window doubles per unchanged recheck, caps, and clamps a misconfigured cap" +} + +# The streak record is written by two supervisors through three helpers, so the +# reset-on-a-changed-wait cannot live in only one of them: a bump handed a wait +# that is not the one on record must start that new wait's streak, whether or not +# any caller reconciled the record first. +test_pause_streak_bump_reconciles_a_changed_wait() { + local f + f="$TMP_ROOT/streak-record" + printf '3\npaused: awaiting the upstream release\n' > "$f" + pause_streak_bump "$f" "paused: awaiting the upstream release" + [ "$(pause_streak_count "$f")" = 4 ] \ + || fail "a recheck of the same wait must continue its streak, got $(pause_streak_count "$f")" + pause_streak_bump "$f" "paused: awaiting the captain merge call" + [ "$(pause_streak_count "$f")" = 1 ] \ + || fail "a recheck of a changed wait must start over, got $(pause_streak_count "$f")" + pause_streak_sync "$f" "paused: awaiting the captain merge call" \ + && fail "the bump did not leave the changed wait on record" + rm -f "$f" + pass "pause_streak_bump reconciles the wait on record before counting a recheck" +} + +# The live 2026-08-04 case behind issue 47: three tasks correctly parked on one +# captain-owned merge decision re-surfaced on a fixed cadence, each recheck +# costing a supervision turn to confirm a wait that had not changed. The recheck +# must survive - a forgotten hold cannot rot invisibly - but an UNCHANGED wait +# must cost less each time. Phase E is the disconfirming half: nothing about the +# backoff may reach a crew that never declared a wait, which still absorbs on the +# wedge timer and still escalates as a possible wedge at the unchanged threshold; +# and phases C and D are the boundary between them - the widened cadence is +# earned by one unchanged wait and dies the moment that wait changes, whether it +# is replaced by a different wait or dropped entirely. +test_paused_resurface_backs_off_while_wedge_still_escalates() { + local dir state fakebin out capture_file window key pane_hash sig pid back statusf wakes + dir=$(make_case paused-resurface-backoff); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt" + window="test:fm-parked" + printf 'idle awaiting the merge decision' > "$capture_file" + printf 'window=%s\nkind=ship\n' "$window" > "$state/parked.meta" + statusf="$state/parked.status" + printf 'paused: awaiting the merge decision on the open PR\n' > "$statusf" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "idle awaiting the merge decision") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + export FM_FAKE_CREW_STATE='state: paused · source: status-log · awaiting the merge decision on the open PR' + + # Phase A: the wait is well past the base window, so the FIRST recheck fires. + back=$(( $(date +%s) - 500 )) + set_mtime "$back" "$statusf" + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-parked_status" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_TMUX_CURRENT_COMMAND=zsh \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "the first recheck of a long declared wait did not re-surface" + grep -F "awaiting external" "$out" >/dev/null || fail "the first recheck was not a paused recheck" + [ "$(pause_streak_count "$state/.paused-streak-$key")" = 1 ] \ + || fail "the first recheck did not record a re-surface streak" + # Each phase below asserts the rechecks IT produced. Presented records stay + # durable until acknowledged, so every phase retires its own before the next + # watcher arms; otherwise the next arm legitimately resurfaces the unhandled + # row and the phase never exercises its own cadence decision. + ack_stopped_cycle "$state" || fail "could not acknowledge the first recheck" + + # Phase B: the wait has not changed. One base window later is now too soon - + # the second recheck must wait for the DOUBLED window before firing again. + : > "$out" + set_mtime "$(( $(date +%s) - 300 ))" "$state/.paused-resurfaced-$key" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_TMUX_CURRENT_COMMAND=zsh \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + if ! wait_live "$pid" 30; then + reap "$pid"; fail "an unchanged wait re-surfaced again inside the widened window: $(cat "$out")" + fi + reap "$pid" + [ ! -s "$out" ] || fail "an unchanged wait printed a wake inside the widened window: $(cat "$out")" + # This phase deliberately queued nothing, so only the stopped cycle's own + # downtime marker needs retiring; the queue is left exactly as it is. + retire_downtime_marker "$state" || fail "could not retire the inside-window stop" + + # Past the widened window it DOES fire again, so the wait still cannot rot. + set_mtime "$(( $(date +%s) - 600 ))" "$state/.paused-resurfaced-$key" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_TMUX_CURRENT_COMMAND=zsh \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "a declared wait past its widened window did not re-surface" + grep -F "awaiting external" "$out" >/dev/null || fail "the widened-window recheck was not a paused recheck" + grep -F "possible wedge" "$out" >/dev/null && fail "a declared wait was mislabeled a possible wedge" + [ "$(pause_streak_count "$state/.paused-streak-$key")" = 2 ] \ + || fail "the second recheck did not widen the streak further" + wakes=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w { n++ } END { print n + 0 }' "$state/.wake-queue") + [ "$wakes" -eq 1 ] || fail "expected exactly 1 recheck from the widened window, got $wakes" + ack_stopped_cycle "$state" || fail "could not acknowledge the widened-window recheck" + + # Phase C: the widened cadence is earned by ONE wait, so a DIFFERENT declared + # wait cannot inherit it. The crew replaces its paused line with another; the + # new wait has stood for one base window, which is well inside the window the + # previous wait had widened to, and it must be rechecked anyway. + printf 'paused: awaiting the security review sign-off\n' > "$statusf" + set_mtime "$(( $(date +%s) - 300 ))" "$statusf" + set_mtime "$(( $(date +%s) - 300 ))" "$state/.paused-resurfaced-$key" + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-parked_status" + export FM_FAKE_CREW_STATE='state: paused · source: status-log · awaiting the security review sign-off' + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_TMUX_CURRENT_COMMAND=zsh \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 \ + || fail "a replaced wait inherited the previous wait's widened window instead of the base one" + grep -F "awaiting external" "$out" >/dev/null || fail "the replaced wait's recheck was not a paused recheck" + grep -F "possible wedge" "$out" >/dev/null && fail "a replaced declared wait was mislabeled a possible wedge" + [ "$(pause_streak_count "$state/.paused-streak-$key")" = 1 ] \ + || fail "the streak survived the wait that earned it being replaced" + wakes=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w { n++ } END { print n + 0 }' "$state/.wake-queue") + [ "$wakes" -eq 1 ] || fail "expected exactly 1 recheck from the replaced wait, got $wakes" + ack_stopped_cycle "$state" || fail "could not acknowledge the replaced wait's recheck" + + # Phase D: the same death by the other route. The moment the crew stops + # declaring a wait at all, the streak goes with the rest of the pause tracking, + # so this crew going quiet again later starts back at the base window and never + # inherits a cadence widened by something else. + printf 'working: resumed, the merge decision landed\n' > "$statusf" + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-parked_status" + export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_TMUX_CURRENT_COMMAND=zsh \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_live "$pid" 30 || true + reap "$pid" + [ ! -e "$state/.paused-streak-$key" ] \ + || fail "the widened cadence outlived the wait that earned it (streak $(cat "$state/.paused-streak-$key"))" + [ ! -e "$state/.paused-$key" ] || fail "pause tracking survived the crew resuming" + + # Phase E: a crew that declared NO wait is untouched by any of this. It is + # absorbed only while provably working, and once its idle time crosses the + # wedge threshold it still escalates as a possible wedge - at the unchanged + # FM_STALE_ESCALATE_SECS threshold, which no backoff may widen. + dir=$(make_case paused-backoff-wedge-control); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt" + window="test:fm-quiet" + printf 'idle, no declared wait' > "$capture_file" + printf 'window=%s\nkind=ship\n' "$window" > "$state/quiet.meta" + printf 'working: pushed the branch\n' > "$state/quiet.status" + sig=$(seen_sig "$state/quiet.status"); printf '%s' "$sig" > "$state/.seen-quiet_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "idle, no declared wait") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + printf '%s' "$pane_hash" > "$state/.stale-$key" + printf '%s\n' $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_TMUX_CURRENT_COMMAND=zsh \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 \ + FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "an undeclared idle crew past the wedge threshold did not escalate" + grep -F "possible wedge" "$out" >/dev/null || fail "the undeclared idle crew was not flagged a possible wedge" + [ ! -e "$state/.paused-streak-$key" ] || fail "the pause backoff leaked onto a crew that declared no wait" + pass "an unchanged declared wait rechecks on a widening cadence while a genuine wedge still escalates" +} + +# A captain-held crew can leave a stable backend endpoint after its agent exits. +# fm-crew-state then authoritatively reports stopped rather than paused, but the +# confirmed-dead agent plus the declared wait or captain-held transfer must retain +# bounded pause handling. +# A still-live agent under a CAPTAIN-HELD transfer is the disconfirming case: the +# crew declared nothing there, so it must still surface once, while the unchanged +# hash must not append the same wake on every watcher re-arm. +test_exited_declared_pause_is_bounded_but_live_captain_held_surfaces() { + local dir state fakebin out capture_file statusf window key pane_hash sig pid back round wakes bare + dir=$(make_case exited-declared-pause); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; statusf="$state/held.status" + window="test:fm-held" + printf 'idle bare shell after agent exit\n' > "$capture_file" + printf 'window=%s\nkind=ship\nharness=grok\nbackend=tmux\n' "$window" > "$state/held.meta" + printf 'paused: held per captain while an external decision is pending\n' > "$statusf" + back=$(( $(date +%s) - 500 )) + if [ "$(uname)" = Darwin ]; then touch -mt "$(date -r "$back" '+%Y%m%d%H%M.%S')" "$statusf" + else touch -m -d "@$back" "$statusf"; fi + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-held_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "idle bare shell after agent exit") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + + round=1 + while [ "$round" -le 6 ]; do + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_TMUX_CURRENT_COMMAND=zsh FM_FAKE_CREW_STATE='state: stopped · source: pane · bare shell' \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" >> "$out" & + pid=$! + if wait_poll_cycle "$state" "$pid"; then + reap "$pid" + elif kill -0 "$pid" 2>/dev/null; then + reap "$pid" + fail "dead-agent watcher round $round timed out before completing a poll cycle" + else + wait "$pid" || fail "dead-agent watcher round $round failed" + fi + round=$((round + 1)) + done + # A watcher that queues nothing never creates .wake-queue, so these counts + # read a path that may legitimately be absent. awk aborts on a missing file + # before END runs, which collapses the count to the empty string and turns the + # next comparison into an "integer expression expected" error - reported as a + # flood of an unprintable number of wakes instead of the real contract breach + # the grep below names. No queue means no wakes, per the drain-count read at + # the end of this file. + wakes=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w { n++ } END { print n + 0 }' "$state/.wake-queue" 2>/dev/null || echo 0) + bare=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w && $5 == "stale: " w { n++ } END { print n + 0 }' "$state/.wake-queue" 2>/dev/null || echo 0) + [ "$wakes" -le 1 ] || fail "dead-agent declared pause flooded $wakes stale wakes across six unchanged polls" + [ "$bare" -eq 0 ] || fail "dead-agent declared pause surfaced as $bare bare stopped-crew wakes" + grep -F "awaiting external" "$state/.wake-queue" >/dev/null \ + || fail "dead-agent declared pause did not use the bounded paused recheck" + + dir=$(make_case exited-captain-held); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; statusf="$state/held.status" + window="test:fm-held" + printf 'idle bare shell after captain-held transfer\n' > "$capture_file" + printf 'window=%s\nkind=ship\nharness=grok\nbackend=tmux\n' "$window" > "$state/held.meta" + printf 'captain-held [key=route]: tracked by held-decision-route\n' > "$statusf" + back=$(( $(date +%s) - 500 )) + if [ "$(uname)" = Darwin ]; then touch -mt "$(date -r "$back" '+%Y%m%d%H%M.%S')" "$statusf" + else touch -m -d "@$back" "$statusf"; fi + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-held_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "idle bare shell after captain-held transfer") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_TMUX_CURRENT_COMMAND=zsh FM_FAKE_CREW_STATE='state: stopped · source: pane · bare shell' \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "captain-held dead-agent pane did not re-surface on the bounded cadence" + grep -F "awaiting the captain" "$state/.wake-queue" >/dev/null \ + || fail "captain-held dead-agent pane surfaced as a stopped crew instead of a captain-owned recheck: $(cat "$state/.wake-queue")" + grep -F "awaiting external" "$state/.wake-queue" >/dev/null \ + && fail "captain-held dead-agent pane borrowed the pause verb's external-wait wording" + + # The disconfirming half. A captain-held transfer is firstmate's own record, not + # the crew declaring that this pane is idle on purpose, so the live-agent rule + # that the declared pause no longer pays is still in force here: surface once, + # then hold the same hash without re-appending on every re-arm. + dir=$(make_case alive-captain-held); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; statusf="$state/gate.status" + window="test:fm-gate" + printf 'idle under a captain-held transfer\n' > "$capture_file" + printf 'window=%s\nkind=ship\nharness=grok\nbackend=tmux\n' "$window" > "$state/gate.meta" + printf 'captain-held [key=route]: tracked by held-decision-route\n' > "$statusf" + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-gate_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "idle under a captain-held transfer") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + + # First sight must surface promptly so a live captain-held pane is not hidden + # behind the pause cadence. + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_TMUX_CURRENT_COMMAND=grok FM_FAKE_CREW_STATE='state: unknown · source: none · live pane under a captain-held transfer' \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" >> "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "live captain-held pane did not surface immediately" + ack_stopped_cycle "$state" || fail "could not acknowledge the immediate captain-held surface" + + # Re-arm with the stale timer already beyond the wedge threshold. This is the + # exact unchanged-hash fallback after the immediate surface: it must retain + # the pause cadence and discard any residual wedge timer instead of emitting + # a second possible-wedge wake. + printf '%s\n' $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_TMUX_CURRENT_COMMAND=grok FM_FAKE_CREW_STATE='state: unknown · source: none · live pane under a captain-held transfer' \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 FM_PAUSE_RESURFACE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" >> "$out" & + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid" + fail "live captain-held pane escalated on the wedge timer after its immediate surface: $(cat "$out")" + fi + [ -e "$state/.paused-$key" ] || { reap "$pid"; fail "live captain-held pane lost its pause cadence marker"; } + [ ! -e "$state/.stale-since-$key" ] || { reap "$pid"; fail "live captain-held pane retained the wedge timer"; } + reap "$pid" + wakes=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w { n++ } END { print n + 0 }' "$state/.wake-queue" 2>/dev/null || echo 0) + bare=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w && $5 == "stale: " w { n++ } END { print n + 0 }' "$state/.wake-queue" 2>/dev/null || echo 0) + [ "$wakes" -eq 0 ] || fail "acknowledged captain-held surface replayed $wakes wakes" + [ "$bare" -eq 0 ] || fail "acknowledged captain-held bare stale remained queued" + pass "exited declared-pause and captain-held panes use bounded pause cadence while a live captain-held pane still surfaces once" +} + +# Issue 142, the live half of the family issue 67 closed. The generated brief +# tells every crew to append `paused: <why>` and stop, and on every verified +# harness "stop" means end the turn, so the agent stays ALIVE at its prompt. The +# watcher used to honour a declared pause only for a confidently dead agent, which +# made the designed widening cadence unreachable for exactly the crew the brief +# creates: it re-surfaced a bare `stale: <window>` every few minutes for as long +# as the wait lasted (measured at 320s, 341s and 449s gaps on a task known to be +# waiting on a human). +# +# The discriminator is the artifact set, not log text. pause_streak_bump is the +# only writer of .paused-streak-<key> and is called only by handle_paused_stale, +# so a present streak file proves the designed path ran, while a missing one +# beside freshly stamped .paused-*/.paused-rechecked-*/.paused-resurfaced-* +# markers is the exact signature of surface_nonterminal_stale having written them +# instead - the field evidence from the two parked 2026-08-14 tasks. +# +# Phase C is the disconfirming half: a live pane that declared NO wait is +# untouched and still escalates as a possible wedge on the unchanged threshold. +test_live_declared_pause_is_absorbed_on_the_designed_cadence() { + local dir state fakebin out capture_file statusf window key pane_hash sig pid round wakes bare + dir=$(make_case live-declared-pause); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; statusf="$state/parked.status" + window="test:fm-parked" + printf 'idle at its prompt, parked on the hardware bench\n' > "$capture_file" + printf 'window=%s\nkind=ship\nharness=grok\nbackend=tmux\n' "$window" > "$state/parked.meta" + printf 'paused: awaiting the captain hardware bench run\n' > "$statusf" + set_mtime "$(( $(date +%s) - 500 ))" "$statusf" + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-parked_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "idle at its prompt, parked on the hardware bench") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + + # Phase A: the agent is ALIVE (pane_current_command matches the recorded + # harness) and the wait is past the base window, so this is a recheck on the + # designed cadence - annotated, streak-counted, and never a bare stale. + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_TMUX_CURRENT_COMMAND=grok \ + FM_FAKE_CREW_STATE='state: paused · source: status-log · awaiting the captain hardware bench run' \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 \ + FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "a live crew's declared pause never reached the designed recheck" + grep -F "awaiting external" "$out" >/dev/null \ + || fail "a live crew's declared pause did not use the annotated paused recheck: $(cat "$out")" + grep -F "possible wedge" "$out" >/dev/null && fail "a live crew's declared pause was mislabeled a possible wedge" + [ "$(pause_streak_count "$state/.paused-streak-$key")" = 1 ] \ + || fail "handle_paused_stale never ran for a live declared pause (no re-surface streak recorded)" + bare=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w && $5 == "stale: " w { n++ } END { print n + 0 }' "$state/.wake-queue") + [ "$bare" -eq 0 ] || fail "a live declared pause emitted $bare bare stale wakes" + + # Phase B: the wait has not changed, so the widened window now owns this pane. + # Six re-arms inside it, each with a DIFFERENT pane hash (a live agent repaints + # its prompt), must add no wake at all: the pre-fix leak was one bare wake per + # first-sighting of each new hash, which is what made it fire every few minutes. + round=1 + while [ "$round" -le 6 ]; do + printf 'idle at its prompt, parked on the hardware bench (repaint %s)\n' "$round" > "$capture_file" + printf '%s' "$(hash_text "$(cat "$capture_file")")" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_TMUX_CURRENT_COMMAND=grok \ + FM_FAKE_CREW_STATE='state: paused · source: status-log · awaiting the captain hardware bench run' \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 \ + FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" >> "$out" & + pid=$! + if wait_live "$pid" 15; then reap "$pid"; else wait "$pid" || true; fi + round=$((round + 1)) + done + wakes=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w { n++ } END { print n + 0 }' "$state/.wake-queue") + [ "$wakes" -eq 1 ] \ + || fail "a live declared pause flooded $wakes stale wakes across six repaints inside its widened window" + [ -e "$state/.paused-streak-$key" ] \ + || fail "a repainting live pane dropped the backoff streak its unchanged wait had earned" + + # Phase C: the disconfirming control. Same live agent, same silence, but no + # declared wait - it must still escalate as a possible wedge, unchanged. + dir=$(make_case live-undeclared-quiet); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt" + window="test:fm-quiet" + printf 'idle at its prompt, nothing declared\n' > "$capture_file" + printf 'window=%s\nkind=ship\nharness=grok\nbackend=tmux\n' "$window" > "$state/quiet.meta" + printf 'working: pushed the branch\n' > "$state/quiet.status" + sig=$(seen_sig "$state/quiet.status"); printf '%s' "$sig" > "$state/.seen-quiet_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "idle at its prompt, nothing declared") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + printf '%s' "$pane_hash" > "$state/.stale-$key" + printf '%s\n' $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_TMUX_CURRENT_COMMAND=grok \ + FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 \ + FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "a live crew that declared no wait stopped escalating past the wedge threshold" + grep -F "possible wedge" "$out" >/dev/null \ + || fail "a live crew that declared no wait was not flagged a possible wedge: $(cat "$out")" + [ ! -e "$state/.paused-streak-$key" ] || fail "the pause cadence leaked onto a crew that declared no wait" + pass "a live crew's declared pause is absorbed on the designed widening cadence while an undeclared quiet pane still wedges" +} + +# The ordering half of issue 142. surface_nonterminal_stale queued its wake FIRST +# and only afterwards read the status line, found the pause, and stamped the +# .paused-* markers - so the pause was recognised one step too late to suppress +# the wake it had just queued. +# +# The race is made deterministic through the real seam that produced it: the +# classifier reads the status line, then calls fm-crew-state.sh, and the crew is +# free to append its pause during that call. This fake does exactly that, so the +# classifier's own read predates the declaration and its verdict routes the pane +# to surface_nonterminal_stale with a paused status already on disk. Nothing but +# the read-before-queue ordering can save it: a bare stale here, with the three +# .paused-* markers stamped and no streak file, is the pre-fix signature exactly. +test_declared_pause_landing_mid_classification_never_emits_a_bare_stale() { + local dir state fakebin out capture_file statusf window key pane_hash sig pid bare + dir=$(make_case pause-lands-mid-classification); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; statusf="$state/racing.status" + window="test:fm-racing" + printf 'idle at its prompt\n' > "$capture_file" + printf 'window=%s\nkind=ship\nharness=grok\nbackend=tmux\n' "$window" > "$state/racing.meta" + printf 'working: running the bench sweep\n' > "$statusf" + set_mtime "$(( $(date +%s) - 500 ))" "$statusf" + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-racing_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "idle at its prompt") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + + # The crew declares its wait while the classifier is mid-read. Keep the status + # mtime old so the declared wait is immediately past its re-surface window, + # and leave the .seen-* signature matching so this write is not itself a signal. + cat > "$fakebin/fm-crew-state.sh" <<SH +#!/usr/bin/env bash +set -u +printf 'paused: awaiting the captain hardware bench run\n' > "$statusf" +$(declare -f set_mtime) +set_mtime "\$(( \$(date +%s) - 500 ))" "$statusf" +$(declare -f seen_sig) +printf '%s' "\$(seen_sig "$statusf")" > "$state/.seen-racing_status" +printf 'state: unknown · source: none · read before the wait was declared\n' +exit 0 +SH + chmod +x "$fakebin/fm-crew-state.sh" + + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_TMUX_CURRENT_COMMAND=grok \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 \ + FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "the pane that declared its wait mid-classification never surfaced at all" + bare=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w && $5 == "stale: " w { n++ } END { print n + 0 }' "$state/.wake-queue") + [ "$bare" -eq 0 ] \ + || fail "the wake was queued before the status was read: $bare bare stale wakes for an already-declared pause" + grep -F "awaiting external" "$state/.wake-queue" >/dev/null \ + || fail "a pause already on disk did not reach the annotated recheck: $(cat "$state/.wake-queue")" + [ -e "$state/.paused-streak-$key" ] \ + || fail "the three pause markers were stamped without the streak file: surface_nonterminal_stale wrote them, not handle_paused_stale" + pass "a pause already on disk is recognised before the wake is queued, never after" +} + +# A pause is honoured on the crew's declaration, but liveness keeps its RECOVERY +# job: the same parked pane whose agent later exits must still be reachable as a +# stopped crew rather than disappearing behind the wait it declared while alive. +test_live_declared_pause_still_recoverable_once_its_agent_dies() { + local dir state fakebin out capture_file statusf window key pane_hash sig pid + dir=$(make_case parked-then-dead); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; statusf="$state/parked.status" + window="test:fm-parked" + printf 'idle at its prompt, parked\n' > "$capture_file" + printf 'window=%s\nkind=ship\nharness=grok\nbackend=tmux\n' "$window" > "$state/parked.meta" + printf 'paused: awaiting the captain hardware bench run\n' > "$statusf" + set_mtime "$(( $(date +%s) - 500 ))" "$statusf" + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-parked_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "idle at its prompt, parked") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + + # Alive: absorbed on the designed cadence, streak recorded. + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_TMUX_CURRENT_COMMAND=grok \ + FM_FAKE_CREW_STATE='state: paused · source: status-log · awaiting the captain hardware bench run' \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 \ + FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "the live parked pane never reached its declared-wait recheck" + [ "$(pause_streak_count "$state/.paused-streak-$key")" = 1 ] \ + || fail "the live parked pane never reached the designed pause path" + ack_stopped_cycle "$state" || fail "could not acknowledge the live parked pane's recheck" + + # The agent then exits, leaving a bare shell on the same endpoint and the same + # declared wait. fm-crew-state falls back to stopped; the recovery signal - the + # crew-state read that reports a stopped crew, and the endpoint liveness behind + # it - must still be exercised, and the pane must stay on the bounded recheck + # rather than going silent. + printf 'bare shell after the agent exited\n' > "$capture_file" + printf '%s' "$(hash_text "$(cat "$capture_file")")" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + rm -f "$state/.paused-rechecked-$key" + set_mtime "$(( $(date +%s) - 900 ))" "$state/.paused-resurfaced-$key" + set_mtime "$(( $(date +%s) - 900 ))" "$statusf" + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-parked_status" + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_TMUX_CURRENT_COMMAND=zsh \ + FM_FAKE_CREW_STATE='state: stopped · source: pane · bare shell' \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 \ + FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "a parked pane whose agent died went silent instead of re-surfacing" + grep -F "awaiting external" "$out" >/dev/null \ + || fail "a parked pane whose agent died did not re-surface on the bounded recheck: $(cat "$out")" + [ "$(pause_streak_count "$state/.paused-streak-$key")" = 2 ] \ + || fail "the dead-agent recheck did not continue the same wait's streak" + pass "a parked pane whose agent later dies stays on the bounded recheck, so liveness keeps its recovery job" +} + +# A dead worker reaches handle_paused_stale rather than the live fallback above. +# When one declared wait directly replaces another, the existing +# throttle belongs to the old declaration and must not suppress the new wait's +# first inspection merely because its timestamp is still young. +test_absorbed_replacement_wait_does_not_inherit_the_old_throttle() { + local spec name initial replacement expected dir state fakebin out capture_file + local statusf window key sig back pid wakes + for spec in \ + 'paused-replacement|paused: waiting on validation run one|paused: waiting on validation run two|awaiting external' \ + 'captain-held-replacement|captain-held [key=route]: awaiting the routing call|captain-held [key=release]: awaiting the release call|awaiting the captain' + do + name=${spec%%|*}; spec=${spec#*|} + initial=${spec%%|*}; spec=${spec#*|} + replacement=${spec%%|*}; expected=${spec#*|} + dir=$(make_case "$name"); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; statusf="$state/held.status" + window="test:fm-held" + printf 'idle after agent exit\n' > "$capture_file" + printf 'window=%s\nkind=ship\nharness=grok\nbackend=tmux\n' "$window" > "$state/held.meta" + printf '%s\n' "$initial" > "$statusf" + back=$(( $(date +%s) - 500 )) + if [ "$(uname)" = Darwin ]; then touch -mt "$(date -r "$back" '+%Y%m%d%H%M.%S')" "$statusf" + else touch -m -d "@$back" "$statusf"; fi + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-held_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + printf '%s' "$(hash_text 'idle after agent exit')" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_TMUX_CURRENT_COMMAND=zsh FM_FAKE_CREW_STATE='state: stopped · source: pane · bare shell' \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" >> "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "[$name] initial declared wait did not re-surface" + ack_stopped_cycle "$state" || fail "[$name] could not acknowledge the initial declared wait" + + printf '%s\n' "$replacement" >> "$statusf" + # The fork rechecks a replacement at its own base deadline, rather than + # immediately. It must not inherit the prior wait's widened window. + set_mtime "$(( $(date +%s) - 300 ))" "$statusf" + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-held_status" + printf 'idle after replacement wait\n' > "$capture_file" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_TMUX_CURRENT_COMMAND=zsh FM_FAKE_CREW_STATE='state: stopped · source: pane · bare shell' \ + FM_WATCH_HANDLING_SUCCESSOR=1 \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" >> "$out" & + pid=$! + wait_for_exit "$pid" 100 \ + || { reap "$pid"; fail "[$name] replacement declared wait inherited the old throttle"; } + wakes=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w { n++ } END { print n + 0 }' \ + "$state/.wake-queue" 2>/dev/null || echo 0) + [ "$wakes" -eq 1 ] || fail "[$name] replacement declared wait produced $wakes wakes instead of one" + grep -F "$expected" "$state/.wake-queue" >/dev/null \ + || fail "[$name] replacement declared wait used the wrong recheck reason: $(cat "$state/.wake-queue")" + done + pass "absorbed paused and captain-held replacements each start their own re-surface cadence" +} + +# Run one watcher round against a parked-worker fixture, so a round differs only +# in the pane contents the case just wrote. Armed the way fm-watch-arm.sh arms a +# successor after firstmate handled a wake, because that is what a supervision +# turn actually does and it is the only arm that stays in the poll loop instead of +# re-announcing the previous round's downtime - without it a round exits on +# `check: rearm-resurface` before it ever reaches the stale path, and every +# absorb assertion below passes vacuously. A live agent (pane_current_command +# matching the recorded harness) on an idle pane is the exact population +# pause_state_class answers `none` for. +# <mode> `exit` requires the watcher to surface and exit; `absorb` requires it to +# survive whole poll cycles - enough to see the new hash, count it stable, and +# reach the stale path. Returns 1 when the watcher does the other thing. +parked_watch_round() { # <state> <fakebin> <out> <capture> <window> <exit|absorb> + local state=$1 fakebin=$2 out=$3 capture=$4 window=$5 mode=$6 pid cycles=0 + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture" \ + FM_FAKE_TMUX_CURRENT_COMMAND=grok \ + FM_FAKE_CREW_STATE='state: paused · source: status-log · parked' \ + FM_WATCH_HANDLING_SUCCESSOR=1 \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_PAUSE_RESURFACE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" >> "$out" & + pid=$! + if [ "$mode" = exit ]; then + wait_for_exit "$pid" 100 || { reap "$pid"; return 1; } + return 0 + fi + while [ "$cycles" -lt 4 ]; do + wait_poll_cycle "$state" "$pid" 300 || { reap "$pid"; return 1; } + cycles=$((cycles + 1)) + done + reap "$pid" + return 0 +} + +# A live captain-held pane earns one initial inspection and then a +# declaration-scoped throttle. Declared paused workers use the fork's widening +# cadence instead, covered by the live-declared-pause cases above. +test_live_declared_wait_churn_honors_the_resurface_throttle() { + local name status_line dir state fakebin out capture_file statusf window key + local sig round wakes bare text throttle replacement + name=captain-held-churn + status_line='captain-held [key=route]: awaiting the captain on the routing call' + dir=$(make_case "$name"); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; statusf="$state/parked.status" + window="test:fm-parked" + printf 'window=%s\nkind=ship\nharness=grok\nbackend=tmux\n' "$window" > "$state/parked.meta" + printf '%s\n' "$status_line" > "$statusf" + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-parked_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + throttle="$state/.paused-resurfaced-$key" + + # First sight of a parked-but-live worker must still surface: the state is + # inconclusive and firstmate has to look at it. + text='parked, elapsed 1s' + printf '%s' "$text" > "$capture_file" + printf '%s' "$(hash_text "$text")" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + parked_watch_round "$state" "$fakebin" "$out" "$capture_file" "$window" exit \ + || fail "[$name] first sight of a parked live worker did not surface" + ack_stopped_cycle "$state" || fail "[$name] could not acknowledge the first surface" + [ -e "$throttle" ] || fail "[$name] the first surface recorded no re-surface throttle" + + # The pane now churns while the SAME declared wait stands, each round fully + # handled as a real supervision turn would. Every one of these used to alarm. + round=2 + while [ "$round" -le 4 ]; do + printf 'parked, elapsed %ss' "$round" > "$capture_file" + parked_watch_round "$state" "$fakebin" "$out" "$capture_file" "$window" absorb \ + || fail "[$name] watcher exited during churn round $round instead of supervising through it" + wakes=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w { n++ } END { print n + 0 }' \ + "$state/.wake-queue" 2>/dev/null || echo 0) + [ "$wakes" -eq 0 ] \ + || fail "[$name] pane churn re-alarmed a parked worker $wakes time(s) inside the re-surface window" + [ -e "$throttle" ] || fail "[$name] pane churn cleared the re-surface throttle" + round=$((round + 1)) + done + + # A direct wait-to-wait transition starts a NEW declaration even though the + # same window remains parked. Its first sight must not inherit the previous + # declaration's throttle, or an unrelated replacement wait can stay silent + # for nearly the whole old cadence window. + replacement='captain-held [key=release]: awaiting the captain on the release call' + printf '%s\n' "$replacement" >> "$statusf" + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-parked_status" + printf 'replacement wait, elapsed 1s' > "$capture_file" + parked_watch_round "$state" "$fakebin" "$out" "$capture_file" "$window" exit \ + || fail "[$name] a replacement declared wait inherited the previous wait's re-surface throttle" + wakes=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w { n++ } END { print n + 0 }' \ + "$state/.wake-queue" 2>/dev/null || echo 0) + bare=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w && $5 == "stale: " w { n++ } END { print n + 0 }' \ + "$state/.wake-queue" 2>/dev/null || echo 0) + [ "$wakes" -eq 1 ] || fail "[$name] replacement declared wait produced $wakes first wakes instead of one" + [ "$bare" -eq 1 ] || fail "[$name] replacement declared wait changed the wake identity: $(cat "$state/.wake-queue")" + ack_stopped_cycle "$state" || fail "[$name] could not acknowledge the replacement wait's first surface" + + printf 'replacement wait, elapsed 2s' > "$capture_file" + parked_watch_round "$state" "$fakebin" "$out" "$capture_file" "$window" absorb \ + || fail "[$name] replacement wait re-alarmed inside its own re-surface window" + wakes=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w { n++ } END { print n + 0 }' \ + "$state/.wake-queue" 2>/dev/null || echo 0) + [ "$wakes" -eq 0 ] || fail "[$name] replacement wait re-alarmed $wakes time(s) inside its own re-surface window" + + # End of the window: the wait must re-surface exactly once, on the same plain + # identity as before, so absorbing churn never becomes silence. + set_mtime "$(( $(date +%s) - 2000 ))" "$throttle" + printf 'parked, elapsed 5s' > "$capture_file" + parked_watch_round "$state" "$fakebin" "$out" "$capture_file" "$window" exit \ + || fail "[$name] a parked worker did not re-surface once its re-surface window elapsed" + wakes=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w { n++ } END { print n + 0 }' \ + "$state/.wake-queue" 2>/dev/null || echo 0) + bare=$(awk -F '\t' -v w="$window" '$3 == "stale" && $4 == w && $5 == "stale: " w { n++ } END { print n + 0 }' \ + "$state/.wake-queue" 2>/dev/null || echo 0) + [ "$wakes" -eq 1 ] || fail "[$name] elapsed re-surface window produced $wakes wakes instead of one" + [ "$bare" -eq 1 ] || fail "[$name] elapsed re-surface changed the wake identity: $(cat "$state/.wake-queue")" + pass "a live captain-held worker surfaces once, absorbs pane churn for the whole re-surface window, then re-surfaces when it elapses" +} + +# --- work the captain is already holding: pane churn must not re-alarm ------- +# The other record of a legitimate wait. The declared-wait bound above reads the +# status LINE, and a delivered task's line stays `done: PR ...` while the wait +# itself lives in the BACKLOG, written there by bin/fm-captain-hold.sh. No line +# predicate can see that record, so both stale alarms - the captain-relevant one +# and the inconclusive one - re-fired on every new pane hash for as long as the +# captain was deciding, which is the 2026-09 loop observed on delivered work +# awaiting their merge word. +# Pinned here, in both directions: while the call stands the first sight still +# alarms, further sights of the SAME call and status-log state are absorbed, and +# a new pane hash after the window's end alarms once more; and the identical +# fixture WITHOUT the hold keeps alarming on every hash, because a bound that +# swallowed an unheld delivery or blocker would be worse than the churn it removes. +# +# The backlog is real rather than a fixture file: bin/fm-captain-hold.sh is the +# only writer of a hold and tasks-axi the only reader, so a hand-written row +# would pin this test's idea of a hold instead of the one the watcher consults. +# +# Cost: every case below drives churn through ONE watcher process rather than +# relaunching per pane change. Watcher startup dominates a round here, and an +# absorbing watcher stays in its poll loop across churn in production anyway, so +# the cheaper shape is also the more faithful one. + +# The window key every hold fixture uses, derived the way fm-watch.sh derives it. +hold_key() { + printf '%s' test:fm-held-merge | tr ':/.' '___' +} + +# bin/fm-captain-hold.sh against a hold fixture's own home. +run_hold() { # <dir> <args...> + local dir=$1 + shift + FM_HOME="$dir" FM_STATE_OVERRIDE="$dir/state" FM_DATA_OVERRIDE="$dir/data" \ + FM_CONFIG_OVERRIDE="$dir/config" "$ROOT/bin/fm-captain-hold.sh" "$@" >/dev/null 2>&1 +} + +make_hold_home() { # <name> <status-line> <hold|nohold> + local name=$1 line=$2 hold=$3 dir state + dir=$(make_case "$name"); state="$dir/state" + mkdir -p "$dir/data" "$dir/config" + cp "$ROOT/.tasks.toml" "$dir/.tasks.toml" || return 1 + printf '## In flight\n\n## Queued\n\n## Done\n' > "$dir/data/backlog.md" + (cd "$dir" && tasks-axi add held-merge 'delivered work' --file data/backlog.md) >/dev/null 2>&1 \ + || return 1 + if [ "$hold" = hold ]; then + run_hold "$dir" hold held-merge --reason 'awaiting the captain on the merge' || return 1 + fi + printf 'window=test:fm-held-merge\nkind=ship\nharness=grok\nbackend=tmux\n' \ + > "$state/held-merge.meta" + printf '%s\n' "$line" > "$state/held-merge.status" + printf '%s' "$(seen_sig "$state/held-merge.status")" > "$state/.seen-held-merge_status" + printf '%s\n' "$dir" +} + +# Launch one watcher against a hold fixture, armed the way parked_watch_round +# arms one, plus the home the backlog read resolves against. The crew reads +# stopped: a delivered worker's agent has exited, and that is the population +# whose alarm the call must bound. The pid lands in HOLD_WATCH_PID rather than on +# stdout: a command substitution would background the watcher inside a subshell, +# leaving the caller unable to wait on or reap its own watcher. +HOLD_WATCH_PID= +hold_watch_launch() { # <dir> <out> <capture> + local dir=$1 out=$2 capture=$3 + PATH="$dir/fakebin:$PATH" FM_FAKE_TMUX_WINDOW=test:fm-held-merge \ + FM_FAKE_TMUX_CAPTURE="$capture" FM_FAKE_TMUX_CURRENT_COMMAND=zsh \ + FM_FAKE_CREW_STATE='state: stopped · source: pane · bare shell' \ + FM_WATCH_HANDLING_SUCCESSOR=1 \ + FM_HOME="$dir" FM_DATA_OVERRIDE="$dir/data" FM_CONFIG_OVERRIDE="$dir/config" \ + FM_STATE_OVERRIDE="$dir/state" FM_CREW_STATE_BIN="$dir/fakebin/fm-crew-state.sh" \ + FM_PAUSE_RESURFACE_SECS="${FM_HOLD_PAUSE_RESURFACE_SECS:-999}" FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" >> "$out" 2>&1 & + HOLD_WATCH_PID=$! +} + +# One sighting that must surface and exit the cycle. +hold_watch_surface() { # <dir> <out> <capture> <pane-text> + local dir=$1 out=$2 capture=$3 text=$4 + printf '%s\n' "$text" > "$capture" + hold_watch_launch "$dir" "$out" "$capture" + wait_for_exit "$HOLD_WATCH_PID" 100 || { reap "$HOLD_WATCH_PID"; return 1; } + return 0 +} + +# <count> successive pane changes driven through ONE watcher, each given three +# poll cycles: one to see the new hash, one to count it stable and classify, one +# to prove the classification held. The watcher must stay in the loop throughout. +hold_watch_churn() { # <dir> <out> <capture> <label> <count> + local dir=$1 out=$2 capture=$3 label=$4 count=$5 i=1 c + local state="$dir/state" + printf '%s 0\n' "$label" > "$capture" + hold_watch_launch "$dir" "$out" "$capture" + while [ "$i" -le "$count" ]; do + printf '%s %s\n' "$label" "$i" > "$capture" + c=0 + while [ "$c" -lt 3 ]; do + wait_poll_cycle "$state" "$HOLD_WATCH_PID" 300 \ + || { reap "$HOLD_WATCH_PID"; return 1; } + c=$((c + 1)) + done + i=$((i + 1)) + done + reap "$HOLD_WATCH_PID" + return 0 +} + +hold_stale_wakes() { # <state> + awk -F '\t' '$3 == "stale" && $4 == "test:fm-held-merge" { n++ } END { print n + 0 }' \ + "$1/.wake-queue" 2>/dev/null || echo 0 +} + +# Both status lines a held task really carries: the delivery that routes through +# the captain-relevant stale branch, and a worker line that routes through the +# inconclusive one. The hold is invisible to the status line in both, so both +# branches had the same blindness and both are covered. +test_open_captain_call_bounds_stale_churn() { + local spec name line dir state out capture throttle wakes + command -v tasks-axi >/dev/null 2>&1 \ + || { echo "skip: tasks-axi not found (captain-hold stale bound)"; return 0; } + for spec in \ + 'held-delivery|done: PR https://example.invalid/pull/1 checks green' \ + 'held-worker-line|working: still tidying the branch' + do + name=${spec%%|*}; line=${spec#*|} + dir=$(make_hold_home "$name" "$line" hold) \ + || fail "[$name] could not build a captain-held backlog fixture" + state="$dir/state"; out="$dir/watch.out"; capture="$dir/pane.txt" + throttle="$state/.paused-resurfaced-$(hold_key)" + + # First sight still alarms: the call bounds repetition, never the first look. + hold_watch_surface "$dir" "$out" "$capture" 'idle, elapsed 1s' \ + || fail "[$name] first sight of held work did not surface" + wakes=$(hold_stale_wakes "$state") + [ "$wakes" -eq 1 ] || fail "[$name] first sight produced $wakes wakes instead of one" + ack_stopped_cycle "$state" || fail "[$name] could not acknowledge the first surface" + + # The pane churns while the SAME call stands. Every one of these alarmed. + hold_watch_churn "$dir" "$out" "$capture" 'idle, tick' 2 \ + || fail "[$name] watcher exited during pane churn instead of supervising through it" + wakes=$(hold_stale_wakes "$state") + [ "$wakes" -eq 0 ] \ + || fail "[$name] pane churn re-alarmed held work $wakes time(s) inside the re-surface window" + + # After the window ends, the next new pane hash re-surfaces held work exactly + # once, so a forgotten call on a churning pane cannot hide behind the bound. + [ -e "$throttle" ] || fail "[$name] the absorbed churn recorded no re-surface cadence to elapse" + set_mtime "$(( $(date +%s) - 5000 ))" "$throttle" + hold_watch_surface "$dir" "$out" "$capture" 'idle, elapsed 9s' \ + || fail "[$name] held work did not re-surface once its re-surface window elapsed" + wakes=$(hold_stale_wakes "$state") + [ "$wakes" -eq 1 ] \ + || fail "[$name] elapsed re-surface window produced $wakes wakes instead of one" + done + pass "work under an open captain call surfaces once, absorbs pane churn, then re-surfaces when the window elapses" +} + + + +# The other half of the same bound, and the one that decides whether widening the +# wait was safe: the identical fixtures with NO hold must keep alarming on every +# new hash, on both branches. +test_stale_churn_without_a_captain_call_still_alarms() { + local spec name line dir state out capture round wakes + command -v tasks-axi >/dev/null 2>&1 \ + || { echo "skip: tasks-axi not found (unheld stale alarm)"; return 0; } + for spec in \ + 'unheld-delivery|done: PR https://example.invalid/pull/1 checks green' \ + 'unheld-blocker|blocked: cannot reach the release host' \ + 'unheld-worker-line|working: still tidying the branch' + do + name=${spec%%|*}; line=${spec#*|} + dir=$(make_hold_home "$name" "$line" nohold) \ + || fail "[$name] could not build an unheld backlog fixture" + state="$dir/state"; out="$dir/watch.out"; capture="$dir/pane.txt" + round=1 + while [ "$round" -le 2 ]; do + hold_watch_surface "$dir" "$out" "$capture" "idle, elapsed ${round}s" \ + || fail "[$name] an unheld stale window stopped alarming on round $round" + wakes=$(hold_stale_wakes "$state") + [ "$wakes" -eq 1 ] \ + || fail "[$name] round $round produced $wakes wakes instead of one" + ack_stopped_cycle "$state" || fail "[$name] could not acknowledge round $round" + if [ "$name" = unheld-blocker ]; then + assert_contains "$(status_open_decisions "$state/held-merge.status")" $'default\tblocked\t' \ + "unkeyed blocker must remain open after each acknowledged repaint wake" + fi + round=$((round + 1)) + done + done + pass "a stale window with no open captain call keeps alarming on every new hash" +} + + +# The cadence marker may never outlive the wake it claims to record. Recording it +# before publishing the durable wake turned a delayed alarm into a lost one: the +# append fails, the watcher exits with nothing queued, and the next sighting +# reads that fresh marker and absorbs the retry. An unwritable queue is the real +# failure, so it is the one this drives. +test_failed_wake_append_does_not_arm_the_captain_hold_throttle() { + local dir state out capture wakes rc + command -v tasks-axi >/dev/null 2>&1 \ + || { echo "skip: tasks-axi not found (failed wake append)"; return 0; } + dir=$(make_hold_home append-failure 'done: PR https://example.invalid/pull/1 checks green' hold) \ + || fail "could not build a captain-held backlog fixture" + state="$dir/state"; out="$dir/watch.out"; capture="$dir/pane.txt" + + # A directory where the queue file belongs: every append fails, whatever the + # caller does, so the watcher cannot publish the wake it just decided to send. + # Its exit code is read directly here because a refusing watcher exits NON-zero, + # which is the correct outcome and not the "surfaced" one hold_watch_surface means. + rm -f "$state/.wake-queue" + mkdir -p "$state/.wake-queue" + printf 'idle, elapsed 1s\n' > "$capture" + hold_watch_launch "$dir" "$out" "$capture" + wait_for_exit "$HOLD_WATCH_PID" 100 + rc=$? + rmdir "$state/.wake-queue" + [ "$rc" -ne 124 ] || fail "the watcher did not exit when its durable queue could not be written" + [ "$rc" -ne 0 ] || fail "the watcher reported success despite an unwritable durable queue" + [ -e "$state/.paused-resurfaced-$(hold_key)" ] \ + && fail "a wake that never reached the durable queue still armed the re-surface throttle" + + # The retry must alarm: nothing was ever delivered, so nothing may be absorbed. + hold_watch_surface "$dir" "$out" "$capture" 'idle, elapsed 2s' \ + || fail "the retry after a failed wake append was absorbed instead of alarming" + wakes=$(hold_stale_wakes "$state") + [ "$wakes" -eq 1 ] \ + || fail "the retry after a failed wake append produced $wakes wakes instead of one" + pass "a wake that never reached the durable queue arms no re-surface throttle" +} + +# The task id is not the captain call. A task can be answered with `--release` +# and held again as a genuinely different call with NO status append, and binding +# the throttle to the status-log signature alone let the second call inherit the +# first one's silence and absorbed its first sight. That is the one alarm this +# bound must never swallow: a delivery announced twice is noise, but a decision +# waiting on the captain that is never surfaced is invisible. +# Measured at base c499f84 this fixture alarms on every sighting, so the +# suppression was introduced by the bound itself rather than pre-existing. +test_reheld_captain_call_starts_its_own_resurface_window() { + local dir state out capture wakes + command -v tasks-axi >/dev/null 2>&1 \ + || { echo "skip: tasks-axi not found (re-held captain call)"; return 0; } + dir=$(make_hold_home reheld-call 'done: PR https://example.invalid/pull/1 checks green' hold) \ + || fail "could not build a captain-held backlog fixture" + state="$dir/state"; out="$dir/watch.out"; capture="$dir/pane.txt" + + hold_watch_surface "$dir" "$out" "$capture" 'idle, elapsed 1s' \ + || fail "first sight of the first captain call did not surface" + ack_stopped_cycle "$state" || fail "could not acknowledge the first call's surface" + hold_watch_churn "$dir" "$out" "$capture" 'idle, tick' 1 \ + || fail "the first call's churn was not absorbed" + [ "$(hold_stale_wakes "$state")" -eq 0 ] \ + || fail "the first call's churn re-alarmed inside its own window" + + # Answer and release, then re-hold: a second, distinct captain call on the same + # task id, with no status append, so the status signature cannot tell them apart. + printf 'go ahead\n' > "$dir/decision.txt" + run_hold "$dir" answer held-merge --decision-file "$dir/decision.txt" --release \ + || fail "could not record the captain's answer" + run_hold "$dir" hold held-merge --reason 'awaiting the captain a second time' \ + || fail "could not re-hold the task as a second captain call" + + hold_watch_surface "$dir" "$out" "$capture" 'idle, elapsed 3s' \ + || fail "the second captain call inherited the first call's silence" + wakes=$(hold_stale_wakes "$state") + [ "$wakes" -eq 1 ] \ + || fail "the second captain call produced $wakes first wakes instead of one" + pass "a released-then-re-held task is a distinct captain call whose first sight still alarms" +} + + + +test_secondmate_paused_resurfaces_in_normal_mode() { + local dir state fakebin out capture_file statusf window key pane_hash sig pid back + dir=$(make_case secondmate-paused-resurface); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; statusf="$state/secondmate-held.status" + window="test:fm-secondmate-held" + printf 'idle awaiting external\n' > "$capture_file" + printf 'window=%s\nkind=secondmate\n' "$window" > "$state/secondmate-held.meta" + printf 'paused: awaiting the upstream release\n' > "$statusf" + back=$(( $(date +%s) - 500 )) + if [ "$(uname)" = Darwin ]; then touch -mt "$(date -r "$back" '+%Y%m%d%H%M.%S')" "$statusf" + else touch -m -d "@$back" "$statusf"; fi + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-secondmate-held_status" + key=$(printf '%s' "$window" | tr '.:/' '___') + pane_hash=$(hash_text "idle awaiting external") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + export FM_FAKE_CREW_STATE='state: paused · source: status-log · awaiting the upstream release' + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not re-surface a paused secondmate" + grep -F "stale: $window" "$out" >/dev/null || fail "paused secondmate did not emit a stale recheck" + grep -F "awaiting external" "$out" >/dev/null || fail "paused secondmate recheck omitted its external-wait reason" + grep -F "awaiting the captain" "$out" >/dev/null && fail "paused secondmate recheck named the captain instead of its external dependency" + grep -F "possible wedge" "$out" >/dev/null && fail "paused secondmate was mislabeled a wedge" + unset FM_FAKE_CREW_STATE + pass "a declared paused secondmate re-surfaces on the bounded normal-mode cadence" +} + +# A captain hold is the other declared wait, but unlike paused: it has no +# current-state mapping, so a held mate reports `unknown` rather than `paused`. +# The bounded re-surface must still reach it, or a mate's hold rots invisibly: +# nothing else re-reads a quiet mate's endpoint. +test_secondmate_captain_held_resurfaces_in_normal_mode() { + local dir state fakebin out capture_file statusf window key pane_hash sig pid back + dir=$(make_case secondmate-held-resurface); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; statusf="$state/secondmate-hold.status" + window="test:fm-secondmate-hold" + printf 'idle awaiting the captain\n' > "$capture_file" + printf 'window=%s\nkind=secondmate\n' "$window" > "$state/secondmate-hold.meta" + printf 'captain-held [key=route]: tracked by task-decision-route\n' > "$statusf" + back=$(( $(date +%s) - 500 )) + if [ "$(uname)" = Darwin ]; then touch -mt "$(date -r "$back" '+%Y%m%d%H%M.%S')" "$statusf" + else touch -m -d "@$back" "$statusf"; fi + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-secondmate-hold_status" + key=$(printf '%s' "$window" | tr '.:/' '___') + pane_hash=$(hash_text "idle awaiting the captain") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + export FM_FAKE_CREW_STATE='state: unknown · source: none · no current-state source available' + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not re-surface a captain-held secondmate" + grep -F "stale: $window" "$out" >/dev/null || fail "captain-held secondmate did not emit a stale recheck" + grep -F "awaiting the captain" "$out" >/dev/null || fail "captain-held secondmate recheck did not name the captain as the blocker: $(cat "$out")" + grep -F "awaiting external" "$out" >/dev/null && fail "captain-held secondmate recheck claimed an external wait" + grep -F "possible wedge" "$out" >/dev/null && fail "captain-held secondmate was mislabeled a wedge" + unset FM_FAKE_CREW_STATE + pass "a captain-held secondmate re-surfaces on the bounded normal-mode cadence" +} + +test_secondmate_nonpaused_stale_remains_suppressed() { + local dir state fakebin out capture_file statusf window key pane_hash sig pid + dir=$(make_case secondmate-stale-suppressed); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; statusf="$state/secondmate-working.status" + window="test:fm-secondmate-working" + printf 'idle while the parent supervises\n' > "$capture_file" + printf 'window=%s\nkind=secondmate\n' "$window" > "$state/secondmate-working.meta" + printf 'working: the parent supervises this secondmate\n' > "$statusf" + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-secondmate-working_status" + key=$(printf '%s' "$window" | tr '.:/' '___') + pane_hash=$(hash_text "idle while the parent supervises") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid"; fail "watcher surfaced an ordinary secondmate stale pane: $(cat "$out")" + fi + [ ! -s "$out" ] || { reap "$pid"; fail "ordinary secondmate stale pane printed a wake reason: $(cat "$out")"; } + reap "$pid" + pass "a non-paused secondmate retains normal stale suppression" +} + +test_secondmate_unpause_clears_pause_tracking() { + local dir state fakebin out statusf window key pid + dir=$(make_case secondmate-unpause-clears); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; statusf="$state/secondmate-resumed.status"; window="test:fm-secondmate-resumed" + printf 'window=%s\nkind=secondmate\n' "$window" > "$state/secondmate-resumed.meta" + printf 'working: upstream landed\n' > "$statusf" + printf '%s' "$(seen_sig "$statusf")" > "$state/.seen-secondmate-resumed_status" + key=${window//:/_} + key=${key//\//_} + key=${key//./_} + : > "$state/.paused-$key" + : > "$state/.paused-rechecked-$key" + : > "$state/.paused-resurfaced-$key" + : > "$state/.stale-$key" + : > "$state/.stale-since-$key" + : > "$state/.wedge-escalations-$key" + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_poll_cycle "$state" "$pid" || fail "watcher exited while reconciling a resumed secondmate: $(cat "$out")" + [ ! -e "$state/.paused-$key" ] || { reap "$pid"; fail "resumed secondmate retained the pause marker"; } + [ ! -e "$state/.stale-$key" ] || { reap "$pid"; fail "resumed secondmate retained stale tracking"; } + [ ! -e "$state/.wedge-escalations-$key" ] || { reap "$pid"; fail "resumed secondmate retained wedge tracking"; } + reap "$pid" + pass "a resumed secondmate clears pause and stale tracking before stale exemption" +} + +test_nonterminal_stale_pause_transitions_reclassify_unchanged_hash() { + local dir state fakebin out capture_file window key pane_hash sig pid i + dir=$(make_case nonterminal-stale-pause-transition); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-transition" + printf 'idle awaiting external\n' > "$capture_file" + printf 'window=%s\nkind=ship\n' "$window" > "$state/transition.meta" + printf 'paused: awaiting the upstream release\n' > "$state/transition.status" + sig=$(seen_sig "$state/transition.status"); printf '%s' "$sig" > "$state/.seen-transition_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "idle awaiting external") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '%s' "$pane_hash" > "$state/.stale-$key" + printf '1\n' > "$state/.count-$key" + printf '%s\n' $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + export FM_FAKE_CREW_STATE='state: paused · source: status-log · awaiting the upstream release' + + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_TMUX_CURRENT_COMMAND=zsh \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + i=0 + while [ "$i" -lt 100 ] && kill -0 "$pid" 2>/dev/null; do + [ -e "$state/.paused-$key" ] && [ ! -e "$state/.stale-since-$key" ] && break + sleep 0.1 + i=$((i + 1)) + done + kill -0 "$pid" 2>/dev/null || { reap "$pid"; fail "a stale hash that entered pause was wedge-escalated: $(cat "$out")"; } + [ -e "$state/.paused-$key" ] || { reap "$pid"; fail "unchanged stale hash did not enter paused mode"; } + [ ! -e "$state/.stale-since-$key" ] || { reap "$pid"; fail "pause transition retained its wedge timer"; } + wait_poll_cycle "$state" "$pid" || { reap "$pid"; fail "a stale hash that entered pause was wedge-escalated: $(cat "$out")"; } + reap "$pid" + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional entered-pause watcher stop" + + printf 'working: upstream landed, resuming\n' > "$state/transition.status" + sig=$(seen_sig "$state/transition.status"); printf '%s' "$sig" > "$state/.seen-transition_status" + FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + i=0 + while [ "$i" -lt 100 ] && kill -0 "$pid" 2>/dev/null; do + [ ! -e "$state/.paused-$key" ] && [ -s "$state/.stale-since-$key" ] && break + sleep 0.1 + i=$((i + 1)) + done + kill -0 "$pid" 2>/dev/null || { reap "$pid"; fail "a stale hash that left pause did not resume wedge tracking: $(cat "$out")"; } + [ ! -e "$state/.paused-$key" ] || { reap "$pid"; fail "unchanged stale hash retained paused mode after resume"; } + [ -s "$state/.stale-since-$key" ] || { reap "$pid"; fail "unchanged stale hash did not restart wedge tracking after resume"; } + wait_poll_cycle "$state" "$pid" || { reap "$pid"; fail "a stale hash that left pause did not resume wedge tracking: $(cat "$out")"; } + reap "$pid" + unset FM_FAKE_CREW_STATE + pass "unchanged stale hashes reclassify when a crew enters or leaves pause" +} + +test_nonterminal_paused_rechecks_authoritative_state() { + local dir state fakebin out capture_file window key pane_hash sig pid + dir=$(make_case nonterminal-paused-recheck); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-pause-recheck" + printf 'idle awaiting external\n' > "$capture_file" + printf 'window=%s\nkind=ship\n' "$window" > "$state/pause-recheck.meta" + printf 'paused: awaiting the upstream release\n' > "$state/pause-recheck.status" + sig=$(seen_sig "$state/pause-recheck.status"); printf '%s' "$sig" > "$state/.seen-pause-recheck_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "idle awaiting external") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '%s' "$pane_hash" > "$state/.stale-$key" + printf '1\n' > "$state/.count-$key" + : > "$state/.paused-$key" + export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' + + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid"; fail "an active run behind a declared pause surfaced instead of resuming wedge tracking: $(cat "$out")" + fi + # Authoritative state moves TRACKING to the wedge timer, which is what this + # case is about. The declaration itself is deliberately kept on record (issue + # 67): it is the fallback the escalation uses when the run yields no progress + # evidence, and it carries that cadence's re-surface throttle and backoff, so + # discarding it here would restart both on every poll. + [ -s "$state/.stale-since-$key" ] || { reap "$pid"; fail "authoritative active run did not resume wedge tracking"; } + [ -e "$state/.paused-$key" ] || { reap "$pid"; fail "authoritative active run discarded the declared wait it may still need"; } + reap "$pid" + unset FM_FAKE_CREW_STATE + pass "a declared pause is periodically rechecked against authoritative active-run state" +} + +test_paused_authoritative_working_preserves_wedge_timer() { + local dir state fakebin out capture_file window key pane_hash sig pid since + dir=$(make_case paused-working-preserves-wedge-timer); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-paused-working" + printf 'idle awaiting external\n' > "$capture_file" + printf 'window=%s\nkind=ship\n' "$window" > "$state/paused-working.meta" + printf 'paused: awaiting the upstream release\n' > "$state/paused-working.status" + sig=$(seen_sig "$state/paused-working.status"); printf '%s' "$sig" > "$state/.seen-paused-working_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "idle awaiting external") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '%s' "$pane_hash" > "$state/.stale-$key" + printf '1\n' > "$state/.count-$key" + : > "$state/.paused-$key" + export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' + + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_numeric_file "$state/.stale-since-$key" 30 || { reap "$pid"; fail "authoritative working state did not start wedge tracking"; } + since=$(cat "$state/.stale-since-$key") + sleep 2 + [ "$(cat "$state/.stale-since-$key" 2>/dev/null || true)" = "$since" ] \ + || { reap "$pid"; fail "repeat authoritative working recheck reset the wedge timer"; } + reap "$pid" + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional authoritative-working stop" + + echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + : > "$out" + # The verdict is the run's, not the pane's: with the declared wait still on + # record, only positive evidence that the run stopped or the agent died may + # raise the alarm here (issue 67). The no-evidence half is pinned by + # test_declared_wait_with_no_progress_evidence_rechecks_instead_of_wedging. + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_FAKE_RUN_PROGRESS='progress: stranded · test running, last activity 31m0s ago' \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "authoritative working state did not wedge-escalate past the threshold" + grep -F "possible wedge" "$out" >/dev/null || fail "authoritative working wedge escalation omitted its reason" + [ ! -e "$state/.stale-since-$key" ] || fail "wedge timer remained after authoritative working escalation" + unset FM_FAKE_CREW_STATE + pass "a paused status overridden by authoritative working preserves its wedge timer and escalates" +} + +# --- the wedge escalation consults the run's PROGRESS, not just its existence -- +# +# A worker that backgrounds a validation call and goes quiet was escalated as a +# possible wedge every threshold, five times in a row on one pane, while +# fm-crew-state reported "working · run-step · validating (running)" the whole +# time. `status: running` alone cannot separate that from a run that has +# stranded, so the escalation point reads bin/fm-run-progress.sh, whose classes +# these four cases pin from the escalation side without a declared wait: +# +# progressing -> held +# stranded -> still escalates, naming the step that stopped +# none -> escalates byte-identically to before this gate existed +# dead agent -> escalates however well its run is moving +# +# The reader's own parsing and threshold live in fm-run-progress.test.sh. + +# Fixture: a crew parked on a validation run, its wedge timer already backdated +# past the threshold, so the very next poll reaches the escalation decision. +# Its status line deliberately declares NO wait, so these four cases isolate the +# run-progress axis alone; the declared-wait axis has its own four cases below, +# over prime_declared_wait_at_threshold. Keeping both axes in one fixture is what +# made issue 67 invisible here - a declared wait rides a different branch of the +# stale path, and a fixture that carries one silently tests that branch instead. +prime_wedge_at_threshold() { # <state> <task> <window> <capture-file> + local state=$1 task=$2 window=$3 capture=$4 key + printf 'window=%s\nkind=ship\n' "$window" > "$state/$task.meta" + printf 'working: no-mistakes run under way, parked on the pipeline call\n' > "$state/$task.status" + prime_status_seen "$state" "$state/$task.status" + prime_stale_pane "$state" "$window" 'validating · esc to interrupt' "$capture" + key=$(printf '%s' "$window" | tr ':/.' '___') + printf '%s' "$(hash_text 'validating · esc to interrupt')" > "$state/.stale-$key" + echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" +} + +# Drive one watcher over that fixture with a fixed run-progress verdict. Extra +# `NAME=value` arguments are applied on top, through `env` rather than an +# assignment prefix (one arriving through "$@" is expanded too late to be +# recognized as an assignment). +run_wedge_watcher() { # <state> <fakebin> <window> <capture> <out> <progress-verdict> [env assignments...] + local state=$1 fakebin=$2 window=$3 capture=$4 out=$5 verdict=$6 + shift 6 + env "PATH=$fakebin:$PATH" "FM_FAKE_TMUX_WINDOW=$window" "FM_FAKE_TMUX_CAPTURE=$capture" \ + "FM_STATE_OVERRIDE=$state" "FM_CREW_STATE_BIN=$fakebin/fm-crew-state.sh" \ + 'FM_FAKE_CREW_STATE=state: working · source: run-step · validating (running)' \ + "FM_FAKE_RUN_PROGRESS=$verdict" \ + FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$@" "$WATCH" > "$out" & +} + +test_progressing_run_holds_the_wedge_escalation() { + local dir state fakebin out capture window key pid since + dir=$(make_case wedge-run-progressing); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture="$dir/pane.txt"; window="test:fm-validating" + key=$(printf '%s' "$window" | tr ':/.' '___') + export FM_FAKE_TMUX_CURRENT_COMMAND=claude + prime_wedge_at_threshold "$state" validating "$window" "$capture" + since=$(cat "$state/.stale-since-$key") + + run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" \ + 'progress: progressing · test running, last activity 7m4s ago (silent 424s, bound 1800s)' + pid=$! + if ! wait_live "$pid" 30; then + reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND + fail "a crew parked on a progressing validation run still wedge-escalated: $(cat "$out")" + fi + [ ! -s "$out" ] || { reap "$pid"; fail "the held escalation still printed a wake: $(cat "$out")"; } + # Held, not cleared: the timer restarts so the next look is a full window away + # rather than one poll away, which is what keeps the bounded run-progress read + # to once per window per pane. + [ -s "$state/.stale-since-$key" ] || { reap "$pid"; fail "the held escalation cleared the wedge timer"; } + [ "$(cat "$state/.stale-since-$key")" != "$since" ] \ + || { reap "$pid"; fail "the held escalation did not restart the wedge timer"; } + [ ! -e "$state/.wedge-escalations-$key" ] \ + || { reap "$pid"; fail "the held escalation still counted as an escalation"; } + reap "$pid" + FM_STATE_OVERRIDE="$state" "$DRAIN" 2>/dev/null | grep -F "$window" >/dev/null \ + && fail "the held escalation was queued" + unset FM_FAKE_TMUX_CURRENT_COMMAND + pass "a crew parked on a demonstrably progressing validation run holds its wedge escalation" +} + +test_stranded_run_still_wedge_escalates() { + local dir state fakebin out capture window key pid + dir=$(make_case wedge-run-stranded); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture="$dir/pane.txt"; window="test:fm-stranded" + key=$(printf '%s' "$window" | tr ':/.' '___') + export FM_FAKE_TMUX_CURRENT_COMMAND=claude + prime_wedge_at_threshold "$state" stranded "$window" "$capture" + + run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" \ + 'progress: stranded · test running, last activity 31m0s ago (silent 1860s, past the 1800s bound)' + pid=$! + wait_for_exit "$pid" 100 || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "a stranded validation run did not wedge-escalate"; } + grep -F "possible wedge" "$out" >/dev/null || fail "the stranded run's escalation dropped its wedge reason" + grep -F "validation run stranded: test running, last activity 31m0s ago" "$out" >/dev/null \ + || fail "the stranded run's escalation did not name the step that stopped: $(cat "$out")" + [ "$(cat "$state/.wedge-escalations-$key" 2>/dev/null || echo 0)" = 1 ] \ + || fail "the stranded run's escalation was not counted" + unset FM_FAKE_TMUX_CURRENT_COMMAND + pass "a crew whose validation run has stranded still wedge-escalates, naming the step" +} + +test_wedged_crew_with_no_run_escalates_unchanged() { + local dir state fakebin out drain_out capture window key pid + dir=$(make_case wedge-run-absent); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture="$dir/pane.txt" + window="test:fm-norun" + key=$(printf '%s' "$window" | tr ':/.' '___') + export FM_FAKE_TMUX_CURRENT_COMMAND=claude + prime_wedge_at_threshold "$state" norun "$window" "$capture" + + # The ORIGINAL purpose of this alarm: a quiet crew with no active run at all. + # `none` is what every no-evidence shape collapses to, so this pins that the + # gate cannot weaken it. + run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" \ + 'progress: none · no run attributed to this crew' + pid=$! + wait_for_exit "$pid" 100 || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "a wedged crew with no active run did not escalate"; } + grep -F "possible wedge" "$out" >/dev/null || fail "the no-run wedge escalation lost its reason" + grep -F "validation run stranded" "$out" >/dev/null \ + && fail "a crew with no run was described as having a stranded run" + [ ! -e "$state/.stale-since-$key" ] || fail "the no-run escalation left its wedge timer standing" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the no-run wedge failed" + grep "$(printf '\tstale\t')" "$drain_out" | grep -F "possible wedge" >/dev/null \ + || fail "the no-run wedge escalation was not queued" + unset FM_FAKE_TMUX_CURRENT_COMMAND + pass "a crew wedged with no active run escalates exactly as it does today" +} + +test_dead_agent_escalates_even_while_its_run_progresses() { + local dir state fakebin out capture window pid + dir=$(make_case wedge-run-dead-agent); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture="$dir/pane.txt"; window="test:fm-deadagent" + # A bare shell at the endpoint is the confident dead verdict. + export FM_FAKE_TMUX_CURRENT_COMMAND=zsh + prime_wedge_at_threshold "$state" deadagent "$window" "$capture" + + # The pipeline runs its own steps, so a run keeps advancing with nobody left + # to answer its next gate. That is a wedge, and it is exactly the shape "the + # run is fine" would otherwise hide. + run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" \ + 'progress: progressing · test running, last activity 10s ago (silent 10s, bound 1800s)' + pid=$! + wait_for_exit "$pid" 100 || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "a dead agent was absorbed because its run was progressing"; } + grep -F "possible wedge" "$out" >/dev/null || fail "the dead-agent escalation lost its wedge reason" + unset FM_FAKE_TMUX_CURRENT_COMMAND + pass "a confidently dead agent still escalates however well its validation run is moving" +} + +# The principle this pins, which is the whole reason the hold is capped: RUN +# PROGRESS IS EVIDENCE ABOUT THE RUN, NOT ABOUT THE WORKER. They are different +# subjects. A moving pipeline licenses a DELAY in alarming and never permanent +# silence, because the failure permanent silence would hide is a worker whose +# harness hung mid-turn while its pipeline kept executing its own steps quite +# happily - the endpoint reads alive, so the dead-agent short-circuit never +# fires, the run reports `progressing` on every look, and the pane would be held +# for the whole remaining run. That is silent, indefinite, and worse than the +# noise the hold exists to cut. So past FM_RUN_PROGRESS_HOLD_MAX the pane +# escalates REGARDLESS of how healthy its run looks. +test_progressing_run_escalates_anyway_past_the_hold_cap() { + local dir state fakebin out capture window key pid verdict + dir=$(make_case wedge-run-hold-cap); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture="$dir/pane.txt"; window="test:fm-holdcap" + key=$(printf '%s' "$window" | tr ':/.' '___') + export FM_FAKE_TMUX_CURRENT_COMMAND=claude + prime_wedge_at_threshold "$state" holdcap "$window" "$capture" + verdict='progress: progressing · test running, last activity 7m4s ago (silent 424s, bound 1800s)' + + # Phase A: below the cap, unchanged - held, and the hold is counted. + run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" "$verdict" FM_RUN_PROGRESS_HOLD_MAX=2 + pid=$! + if ! wait_live "$pid" 30; then + reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND + fail "a hold below the cap escalated: $(cat "$out")" + fi + [ "$(cat "$state/.wedge-holds-$key" 2>/dev/null || echo 0)" = 1 ] \ + || { reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the hold was not counted"; } + reap "$pid" + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional phase-A hold stop" + + # Phase B: at the cap, with the run reporting the very same healthy verdict. + echo 2 > "$state/.wedge-holds-$key" + echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + : > "$out" + run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" "$verdict" FM_RUN_PROGRESS_HOLD_MAX=2 + pid=$! + wait_for_exit "$pid" 100 || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "a pane held to the cap never escalated: $(cat "$out")"; } + grep -F "possible wedge" "$out" >/dev/null \ + || fail "the forced escalation dropped the possible-wedge marker: $(cat "$out")" + # It must stay INFORMATIVE rather than reading like a dead pane: the + # supervisor has to see "the run is still moving, this pane is not" straight + # off the wake. + grep -F "still progressing" "$out" >/dev/null \ + || fail "the forced escalation did not say the run was still moving: $(cat "$out")" + grep -F "test running, last activity 7m4s ago" "$out" >/dev/null \ + || fail "the forced escalation dropped the progress detail: $(cat "$out")" + [ "$(cat "$state/.wedge-escalations-$key" 2>/dev/null || echo 0)" = 1 ] \ + || fail "the forced escalation did not count toward demand-deep-inspection" + [ ! -e "$state/.wedge-holds-$key" ] || fail "the forced escalation did not reset the hold count" + ack_stopped_cycle "$state" || fail "could not acknowledge the forced escalation" + + # Phase C: and the count reset makes the cap a repeating check-in cadence, not + # a one-shot that then goes quiet forever - the next window holds again. + echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + : > "$out" + run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" "$verdict" FM_RUN_PROGRESS_HOLD_MAX=2 + pid=$! + if ! wait_live "$pid" 30; then + reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND + fail "the window after a forced escalation did not hold again: $(cat "$out")" + fi + [ "$(cat "$state/.wedge-holds-$key" 2>/dev/null || echo 0)" = 1 ] \ + || { reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the hold count did not restart after the forced escalation"; } + reap "$pid" + unset FM_FAKE_TMUX_CURRENT_COMMAND + pass "consecutive run-progress holds are capped: past the cap the pane escalates anyway naming the still-moving run, and the count resets so the cadence repeats" +} + +# The busy-turn bound routes through the same wedge_timer_check, so the hold and +# its cap apply there DELIBERATELY, not incidentally. BUSY_TURN_MAX_SECS exists +# to bound a hung FOREGROUND call that a rendered busy footer would otherwise +# hide, and a crew driving `no-mistakes axi run` in the foreground IS such a +# call: busy for the whole pipeline with no completed turn. Whether such a pane +# reads busy or stale is only an artifact of whether its harness backgrounded +# the pipeline call, so holding for one and not the other would be arbitrary. +# The cap matters MORE here, because a busy pane has already waited a full +# BUSY_TURN_MAX_SECS before its first escalation. +test_busy_pane_progressing_run_holds_then_escalates_past_the_cap() { + local dir state fakebin out capture_file window key pane_hash sig pid verdict + dir=$(make_case busy-run-progress-hold); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-busy-validating" + printf 'Working...' > "$capture_file" + printf 'window=%s\nkind=ship\nharness=pi\n' "$window" > "$state/busy-validating.meta" + record_pi_busy "$state" busy-validating + printf 'working: no-mistakes run under way in the foreground\n' > "$state/busy-validating.status" + sig=$(seen_sig "$state/busy-validating.status"); printf '%s' "$sig" > "$state/.seen-busy-validating_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "Working...") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + # No completed turn ever recorded: age the spawn record past the busy bound. + touch -t 200001010000 "$state/busy-validating.meta" + export FM_FAKE_TMUX_CURRENT_COMMAND=claude + verdict='progress: progressing · test running, last activity 7m4s ago (silent 424s, bound 1800s)' + echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + + # Phase A: past the busy bound AND past the wedge threshold, but the run is + # moving - held, exactly as the stale path holds. + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_RUN_PROGRESS_HOLD_MAX=2 FM_FAKE_RUN_PROGRESS="$verdict" \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + if ! wait_live "$pid" 30; then + reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND + fail "a busy pane on a progressing validation run escalated instead of holding: $(cat "$out")" + fi + [ "$(cat "$state/.wedge-holds-$key" 2>/dev/null || echo 0)" = 1 ] \ + || { reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the busy-path hold was not counted"; } + reap "$pid" + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional busy-path phase-A stop" + + # Phase B: at the cap it escalates anyway, carrying the progress detail. + echo 2 > "$state/.wedge-holds-$key" + echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_RUN_PROGRESS_HOLD_MAX=2 FM_FAKE_RUN_PROGRESS="$verdict" \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "a busy pane held to the cap never escalated: $(cat "$out")"; } + grep -F "possible wedge" "$out" >/dev/null \ + || fail "the busy-path forced escalation dropped the possible-wedge marker: $(cat "$out")" + grep -F "still progressing" "$out" >/dev/null \ + || fail "the busy-path forced escalation did not say the run was still moving: $(cat "$out")" + grep -F "test running, last activity 7m4s ago" "$out" >/dev/null \ + || fail "the busy-path forced escalation dropped the progress detail: $(cat "$out")" + [ ! -e "$state/.wedge-holds-$key" ] || fail "the busy-path forced escalation did not reset the hold count" + unset FM_FAKE_TMUX_CURRENT_COMMAND + pass "a busy pane past its turn-age bound is held while its run is moving and escalates anyway past the hold cap" +} + +# --- a DECLARED wait, overridden by an active run, still counts at the alarm --- +# +# Issue 67, reproduced 2026-08-07: a crew that followed its brief and appended +# `paused:` before parking on a backgrounded pipeline call was wedge-escalated +# twice inside five minutes while fm-crew-state reported an actively running +# validation. The declared wait was consumed the moment authoritative state +# outranked it (correctly - a crew that declared a pause and then started a run +# IS working), and from there the pane was on the plain 240s wedge cadence with +# nothing left of the worker's own statement about its silence. +# +# The four cases below are the whole policy, and each is deliberately the +# opposite of one of the others, so no single change can satisfy them all by +# widening or narrowing absorption: +# +# progressing -> held (positive evidence the run is moving) +# stranded -> escalates (positive evidence the run stopped) +# dead agent -> escalates (nobody left to answer the next gate) +# none + live agent -> declared-wait recheck, NOT a wedge alarm +# +# Only the last one changes behavior. `none` is no evidence either way - no run +# attributed, a status read that could not complete, a run between steps - and +# an alarm needs a reason to fire rather than the absence of one, once the crew +# itself has said the silence is deliberate. bin/fm-supervise-daemon.sh's own +# stale recheck has always short-circuited a declared pause before the wedge +# escalation, so this is also what stops the two supervisors disagreeing about +# the same pane. + +# Fixture: an IDLE (non-busy) pane whose crew declared a wait and whose +# authoritative state is an active run, its wedge timer already past the +# threshold and its stale hash already classified, so the very next poll reaches +# the escalation decision through the declared-wait branch. Deliberately not +# prime_wedge_at_threshold: that fixture's pane text carries a busy signature and +# routes through the busy-turn-age path instead, which is why the stale path's +# declared-wait branch had no coverage of its own. +prime_declared_wait_at_threshold() { # <state> <task> <window> <capture-file> [<status-line>] + local state=$1 task=$2 window=$3 capture=$4 key + local status_line=${5:-'paused: no-mistakes run under way, parked on the pipeline call'} + printf 'window=%s\nkind=ship\n' "$window" > "$state/$task.meta" + printf '%s\n' "$status_line" > "$state/$task.status" + prime_status_seen "$state" "$state/$task.status" + prime_stale_pane "$state" "$window" 'idle, parked on the pipeline call' "$capture" + key=$(printf '%s' "$window" | tr ':/.' '___') + printf '%s' "$(hash_text 'idle, parked on the pipeline call')" > "$state/.stale-$key" + echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" +} + +test_declared_wait_with_no_progress_evidence_rechecks_instead_of_wedging() { + local dir state fakebin out drain_out capture window key pid + dir=$(make_case declared-wait-no-evidence); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture="$dir/pane.txt" + window="test:fm-declared-none" + key=$(printf '%s' "$window" | tr ':/.' '___') + export FM_FAKE_TMUX_CURRENT_COMMAND=claude + prime_declared_wait_at_threshold "$state" declared-none "$window" "$capture" + + # Phase A: the wait was declared moments ago, so its own re-surface window has + # not elapsed - the pane is absorbed outright, where before it alarmed. + run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" \ + 'progress: none · no run attributed to this crew' + pid=$! + if ! wait_live "$pid" 30; then + reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND + fail "a declared wait with no progress evidence still wedge-escalated: $(cat "$out")" + fi + [ ! -s "$out" ] || { reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the deferred pane still printed a wake: $(cat "$out")"; } + [ ! -e "$state/.wedge-escalations-$key" ] \ + || { reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the deferral was counted as a wedge escalation"; } + [ -e "$state/.paused-$key" ] \ + || { reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the pane was not handed back to the declared-wait cadence"; } + grep -F "deferred non-terminal stale (provably working after a declared wait) wedge escalation" \ + "$state/.watch-triage.log" >/dev/null \ + || { reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the deferral was not distinguishable in the triage log: $(cat "$state/.watch-triage.log")"; } + reap "$pid" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the deferral failed" + [ -s "$drain_out" ] \ + && { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the deferred pane queued a wake: $(cat "$drain_out")"; } + + # Phase B: absorbed is not silenced. Age the wait past its own re-surface + # window and it comes back as a recheck the supervisor can act on - never as a + # possible wedge - so a wait that stops being true cannot rot invisibly. + set_mtime $(( $(date +%s) - 500 )) "$state/declared-none.status" + prime_status_seen "$state" "$state/declared-none.status" + echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + : > "$out" + run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" \ + 'progress: none · no run attributed to this crew' FM_PAUSE_RESURFACE_SECS=240 + pid=$! + wait_for_exit "$pid" 100 \ + || { reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "an aged declared wait never came back for a recheck: $(cat "$out")"; } + grep -F "awaiting external" "$out" >/dev/null \ + || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the re-surfaced wake was not a declared-wait recheck: $(cat "$out")"; } + grep -F "possible wedge" "$out" >/dev/null \ + && { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the recheck was raised as a possible wedge: $(cat "$out")"; } + unset FM_FAKE_TMUX_CURRENT_COMMAND + pass "a declared wait whose run yields no progress evidence is rechecked on its own cadence, never wedge-escalated" +} + +test_declared_wait_on_a_progressing_run_holds() { + local dir state fakebin out capture window key pid since + dir=$(make_case declared-wait-progressing); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture="$dir/pane.txt"; window="test:fm-declared-moving" + key=$(printf '%s' "$window" | tr ':/.' '___') + export FM_FAKE_TMUX_CURRENT_COMMAND=claude + prime_declared_wait_at_threshold "$state" declared-moving "$window" "$capture" + since=$(cat "$state/.stale-since-$key") + + # The reproduction's own numbers: `document` running, last activity ~3m ago, + # comfortably inside the stranded bound. + run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" \ + 'progress: progressing · document running, last activity 3m16s ago (silent 196s, bound 1800s)' + pid=$! + if ! wait_live "$pid" 30; then + reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND + fail "a declared wait inside an actively progressing run wedge-escalated: $(cat "$out")" + fi + [ ! -s "$out" ] || { reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the held declared wait still printed a wake: $(cat "$out")"; } + [ "$(cat "$state/.wedge-holds-$key" 2>/dev/null || echo 0)" = 1 ] \ + || { reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the declared wait's hold was not counted"; } + [ "$(cat "$state/.stale-since-$key")" != "$since" ] \ + || { reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the held escalation did not restart the wedge timer"; } + reap "$pid" + unset FM_FAKE_TMUX_CURRENT_COMMAND + pass "a declared wait inside an actively progressing validation run holds its wedge escalation" +} + +test_declared_wait_on_a_stranded_run_still_escalates() { + local dir state fakebin out capture window key pid + dir=$(make_case declared-wait-stranded); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture="$dir/pane.txt"; window="test:fm-declared-stranded" + key=$(printf '%s' "$window" | tr ':/.' '___') + export FM_FAKE_TMUX_CURRENT_COMMAND=claude + prime_declared_wait_at_threshold "$state" declared-stranded "$window" "$capture" + + # Positive evidence the run stopped outranks the crew's own statement that its + # silence is deliberate: a stranded step is exactly the case the declared wait + # must never hide. + run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" \ + 'progress: stranded · test running, last activity 31m0s ago (silent 1860s, past the 1800s bound)' + pid=$! + wait_for_exit "$pid" 100 \ + || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "a declared wait on a stranded run did not escalate: $(cat "$out")"; } + grep -F "possible wedge" "$out" >/dev/null \ + || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the stranded declared wait lost its wedge reason: $(cat "$out")"; } + grep -F "validation run stranded: test running, last activity 31m0s ago" "$out" >/dev/null \ + || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the stranded escalation did not name the step that stopped: $(cat "$out")"; } + [ "$(cat "$state/.wedge-escalations-$key" 2>/dev/null || echo 0)" = 1 ] \ + || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the stranded declared wait's escalation was not counted"; } + unset FM_FAKE_TMUX_CURRENT_COMMAND + pass "a declared wait whose validation run has stranded still wedge-escalates, naming the step" +} + +test_declared_wait_with_a_dead_agent_still_escalates() { + local dir state fakebin out capture window pid + dir=$(make_case declared-wait-dead-agent); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture="$dir/pane.txt"; window="test:fm-declared-dead" + # A bare shell at the endpoint is the confident dead verdict. + export FM_FAKE_TMUX_CURRENT_COMMAND=zsh + prime_declared_wait_at_threshold "$state" declared-dead "$window" "$capture" + + # No progress evidence AND nobody left to answer the run's next gate. The + # declared wait is a statement about a worker that is no longer there, so it + # must not buy the pane the recheck cadence. + run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" \ + 'progress: none · status read did not complete' + pid=$! + wait_for_exit "$pid" 100 \ + || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "a declared wait whose agent had died did not escalate: $(cat "$out")"; } + grep -F "possible wedge" "$out" >/dev/null \ + || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the dead-agent declared wait lost its wedge reason: $(cat "$out")"; } + unset FM_FAKE_TMUX_CURRENT_COMMAND + pass "a declared wait whose agent has confidently exited still wedge-escalates" +} + +# --- the triage log must distinguish the two provably-working absorptions ------ +# What hid issue 67 for two days: a first sighting that absorbed and STARTED the +# wedge timer, and a repeat poll that absorbed and ADVANCED an already-running +# timer toward an escalation, wrote the identical line. The log therefore showed +# a steady stream of absorptions while escalations kept arriving, agreeing with +# the intended behavior rather than the actual behavior. +test_provably_working_absorptions_are_distinguishable_in_the_triage_log() { + local dir state fakebin out capture window key pane_hash pid log + dir=$(make_case absorb-log-distinct); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture="$dir/pane.txt"; window="test:fm-logdistinct" + key=$(printf '%s' "$window" | tr ':/.' '___') + log="$state/.watch-triage.log" + export FM_FAKE_TMUX_CURRENT_COMMAND=claude + prime_declared_wait_at_threshold "$state" logdistinct "$window" "$capture" + pane_hash=$(hash_text 'idle, parked on the pipeline call') + # Phase A: an UNCLASSIFIED hash, so this poll is the first sighting - it + # absorbs and starts the timer. + rm -f "$state/.stale-$key" "$state/.stale-since-$key" + + run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" \ + 'progress: progressing · document running, last activity 3m16s ago' + pid=$! + if ! wait_live "$pid" 30; then + reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the first sighting was not absorbed: $(cat "$out")" + fi + reap "$pid" + grep -F "absorbed non-terminal stale (provably working, wedge timer started)" "$log" >/dev/null \ + || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the first sighting did not record that it STARTED the timer: $(cat "$log")"; } + # The first sighting is the FIRST absorption in the log; the polls behind it + # are already repeat polls of the same hash and rightly say so. + grep -F "absorbed non-terminal stale (provably working" "$log" | head -1 \ + | grep -F "wedge timer started" >/dev/null \ + || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the first absorption was not the timer-started event: $(cat "$log")"; } + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional first-sighting stop" + + # Phase B: the same hash on a later poll, with the timer short of the + # threshold so nothing else can write a line - it absorbs and ADVANCES. + : > "$log"; : > "$out" + printf '%s' "$pane_hash" > "$state/.stale-$key" + date +%s > "$state/.stale-since-$key" + run_wedge_watcher "$state" "$fakebin" "$window" "$capture" "$out" \ + 'progress: progressing · document running, last activity 3m16s ago' + pid=$! + if ! wait_live "$pid" 30; then + reap "$pid"; unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the repeat poll was not absorbed: $(cat "$out")" + fi + reap "$pid" + grep -F "wedge timer advanced" "$log" >/dev/null \ + || { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the repeat poll did not record that it ADVANCED the timer: $(cat "$log")"; } + grep -F "wedge timer started" "$log" >/dev/null \ + && { unset FM_FAKE_TMUX_CURRENT_COMMAND; fail "the repeat poll re-used the first-sighting line"; } + unset FM_FAKE_TMUX_CURRENT_COMMAND + pass "starting the wedge timer and advancing it are distinguishable events in the triage log" +} + +# --- consecutive wedge escalations on the same pane demand deep inspection ---- +# Root cause of the PR #252 incident's ~20 minutes of unnoticed green: each +# wedge escalation fires, gets classified as "still validating" one poll later +# (the timer restarts, see wedge_timer_check), and repeats forever on a pane +# that never changes. A single escalation reason looks identical every round, +# so nothing in the payload itself signals "this has now happened N times in a +# row" - that judgment call was left entirely to the supervisor noticing the +# repetition on its own. This is the safety-net fix: past +# FM_WEDGE_DEMAND_INSPECT_COUNT consecutive escalations on the SAME pane, the +# wake reason itself carries a "demand-deep-inspection" marker. + +test_wedge_escalation_marks_demand_deep_inspection_after_threshold() { + local dir state fakebin out capture_file window key pane_hash sig pid n + dir=$(make_case wedge-escalation); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt" + window="test:fm-wedged" + printf 'idle building output' > "$capture_file" + printf 'window=%s\nkind=ship\n' "$window" > "$state/wedged.meta" + printf 'working: still monitoring ci\n' > "$state/wedged.status" + sig=$(seen_sig "$state/wedged.status"); printf '%s' "$sig" > "$state/.seen-wedged_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "idle building output") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + # The crew's pipeline is actively running: a static pane is normal (waiting on CI). + export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' + + # Priming round: first sighting of this stale hash classifies and absorbs it + # (establishing .stale-$key and starting the wedge timer) without going + # through wedge_timer_check at all - mirrors the existing wedge tests' Phase A. + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid"; fail "watcher exited on the priming round (should absorb): $(cat "$out")" + fi + reap "$pid" + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional wedge priming stop" + + n=1 + while [ "$n" -le 3 ]; do + # Backdate the wedge timer past the threshold before each round, mirroring + # the existing wedge-escalation tests' Phase B (the subsequent-sight timer + # path does not re-read the crew state). + echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not escalate on consecutive wedge round $n: $(cat "$out")" + grep -F "escalation $n" "$out" >/dev/null || fail "round $n did not report escalation count $n: $(cat "$out")" + if [ "$n" -lt 3 ]; then + grep -F "demand-deep-inspection" "$out" >/dev/null && fail "round $n escalated to demand-deep-inspection before the threshold: $(cat "$out")" + else + grep -F "demand-deep-inspection" "$out" >/dev/null || fail "round $n (threshold) did not demand deep inspection: $(cat "$out")" + fi + ack_stopped_cycle "$state" || fail "could not acknowledge wedge escalation round $n" + n=$((n + 1)) + done + [ "$(cat "$state/.wedge-escalations-$key" 2>/dev/null || echo 0)" = 3 ] || fail "escalation counter did not persist across consecutive rounds" + unset FM_FAKE_CREW_STATE + pass "consecutive wedge escalations on the same pane accumulate and demand deep inspection at the threshold" +} + +test_wedge_escalation_resets_when_pane_becomes_active() { + local dir state fakebin out capture_file window key pane_hash sig pid + dir=$(make_case wedge-escalation-reset); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt" + window="test:fm-wedged-reset" + printf 'idle building output' > "$capture_file" + printf 'window=%s\nkind=ship\n' "$window" > "$state/wedged-reset.meta" + printf 'working: still monitoring ci\n' > "$state/wedged-reset.status" + sig=$(seen_sig "$state/wedged-reset.status"); printf '%s' "$sig" > "$state/.seen-wedged-reset_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "idle building output") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + # Pre-seed one escalation as if a prior wedge round already fired. + printf '1\n' > "$state/.wedge-escalations-$key" + export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' + + # The pane content changes (the crew is active again): the hash no longer + # matches, so the watcher resets escalation bookkeeping instead of escalating. + printf 'new output, crew active again' > "$capture_file" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid"; fail "watcher exited on a fresh (changed) pane hash: $(cat "$out")" + fi + [ ! -e "$state/.wedge-escalations-$key" ] || fail "a changed pane hash did not reset the wedge-escalation counter" + reap "$pid" + unset FM_FAKE_CREW_STATE + pass "a pane becoming active again resets the consecutive wedge-escalation counter" +} + +# --- busy pane duration bound: a completed-turn age gate on top of busy ----- +# 2026-07 hibit-agent-focus-nonsteal-r1 incident: a busy pane (herdr "working" +# and/or the harness's rendered busy footer) is unconditional, unbounded proof +# of liveness in every existing classifier, so a genuinely hung foreground tool +# call behind a busy signature ran undetected for 25h. BUSY_TURN_MAX_SECS bounds +# how long a busy pane may run with no completed turn (state/<id>.turn-ended, or +# the task's spawn record before any turn completes); past the bound, panes +# without a declared external wait or verified captain-held transfer take the +# SAME wedge_timer_check already used for a provably-working non-busy stale. +# Escalation reuses the identical stale reason, escalation counter, and +# demand-deep-inspection marker - never an +# automatic interrupt or restart. + +test_busy_pane_below_turn_age_bound_is_absorbed() { + local dir state fakebin out capture_file window key sig pid + dir=$(make_case busy-below-turn-age); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-busy-fresh" + printf 'Working... (12.3s)' > "$capture_file" + printf 'window=%s\nkind=ship\nharness=pi\n' "$window" > "$state/busy-fresh.meta" + record_pi_busy "$state" busy-fresh + printf 'working: setup complete\n' > "$state/busy-fresh.status" + sig=$(seen_sig "$state/busy-fresh.status"); printf '%s' "$sig" > "$state/.seen-busy-fresh_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + touch "$state/busy-fresh.turn-ended" + prime_turnend_seen "$state/busy-fresh.turn-ended" + + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_BUSY_TURN_MAX_SECS=999 FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid"; fail "a busy pane below the turn-age bound was escalated: $(cat "$out")" + fi + [ ! -s "$out" ] || fail "a busy pane below the turn-age bound printed a wake reason" + [ ! -e "$state/.stale-since-$key" ] || fail "a busy pane below the turn-age bound started a wedge timer" + reap "$pid" + pass "a busy worker below the turn-age bound remains working with no escalation" +} + +test_busy_pane_stable_hash_escalates_past_turn_age_bound() { + local dir state fakebin out capture_file window key pane_hash sig pid + dir=$(make_case busy-stable-hash-turn-age); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-busy-stable" + printf 'Working...' > "$capture_file" + printf 'window=%s\nkind=ship\nharness=pi\n' "$window" > "$state/busy-stable.meta" + record_pi_busy "$state" busy-stable + printf 'working: setup complete\n' > "$state/busy-stable.status" + sig=$(seen_sig "$state/busy-stable.status"); printf '%s' "$sig" > "$state/.seen-busy-stable_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "Working...") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + # No completed turn ever recorded for this task: age the spawn record itself. + touch -t 200001010000 "$state/busy-stable.meta" + + # Phase A: past the bound, the stable-hash busy pane is absorbed but starts + # the wedge timer (mirrors the existing provably-working-stale Phase A/B). + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid"; fail "a stable-hash busy pane past the turn-age bound escalated before the wedge threshold: $(cat "$out")" + fi + [ -s "$state/.stale-since-$key" ] || fail "a stable-hash busy pane past the turn-age bound did not start a wedge timer" + reap "$pid" + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional stable-hash phase-A stop" + + # Phase B: backdate the wedge timer past the threshold; the next poll escalates. + echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "a stable-hash busy pane did not wedge-escalate past the turn-age bound" + grep -F "stale: $window" "$out" >/dev/null || fail "busy turn-age escalation did not print the stale wake" + grep -F "possible wedge" "$out" >/dev/null || fail "busy turn-age escalation did not flag a possible wedge" + pass "a busy worker with a stable pane hash still escalates once its completed-turn age reaches the bound" +} + +# Regression fixture for the incident's actual masking condition: Pi's rendered +# elapsed-time footer changes every poll, so the pane hash never repeats and the +# watcher always takes the "new hash" branch, never the stable-hash one above. +test_busy_pane_changing_hash_escalates_past_turn_age_bound() { + local dir state fakebin out capture_file window key pid + dir=$(make_case busy-changing-hash-turn-age); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-busy-ticking" + printf 'Working... (3600.1s)' > "$capture_file" + printf 'window=%s\nkind=ship\nharness=pi\n' "$window" > "$state/busy-ticking.meta" + record_pi_busy "$state" busy-ticking + printf 'working: setup complete\n' > "$state/busy-ticking.status" + sig=$(seen_sig "$state/busy-ticking.status"); printf '%s' "$sig" > "$state/.seen-busy-ticking_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + touch -t 200001010000 "$state/busy-ticking.meta" + # No pre-seeded .hash-<key>: with a real ticking elapsed footer, every poll + # lands here (h != prev) - the reproduction's actual masking condition. + + # Phase A: first sight past the bound absorbs and starts the wedge timer, + # without ever needing the "genuinely stale" hash-match path. + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid"; fail "a changing-hash busy pane past the turn-age bound escalated before the wedge threshold: $(cat "$out")" + fi + [ -s "$state/.stale-since-$key" ] || fail "a changing-hash busy pane past the turn-age bound did not start a wedge timer" + reap "$pid" + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional changing-hash phase-A stop" + + # Phase B: another tick (still a fresh, never-before-seen hash) plus a + # backdated wedge timer escalates exactly as the stable-hash case does. + printf 'Working... (3601.2s)' > "$capture_file" + echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "a changing-hash busy pane did not wedge-escalate past the turn-age bound" + grep -F "stale: $window" "$out" >/dev/null || fail "busy turn-age escalation (changing hash) did not print the stale wake" + grep -F "possible wedge" "$out" >/dev/null || fail "busy turn-age escalation (changing hash) did not flag a possible wedge" + pass "a busy worker whose pane hash changes every poll still escalates once its completed-turn age reaches the bound" +} + +test_busy_pane_turn_end_touch_resets_age() { + local dir state fakebin out capture_file window key pane_hash sig pid + dir=$(make_case busy-turn-end-resets-age); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-busy-reset" + printf 'Working...' > "$capture_file" + printf 'window=%s\nkind=ship\nharness=pi\n' "$window" > "$state/busy-reset.meta" + record_pi_busy "$state" busy-reset + printf 'working: setup complete\n' > "$state/busy-reset.status" + sig=$(seen_sig "$state/busy-reset.status"); printf '%s' "$sig" > "$state/.seen-busy-reset_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "Working...") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + # A wedge is already mid-escalation, as if several over-age polls already ran. + echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + printf '1\n' > "$state/.wedge-escalations-$key" + # The worker's most recent turn just completed: touching turn-ended resets age. + touch "$state/busy-reset.turn-ended" + prime_turnend_seen "$state/busy-reset.turn-ended" + + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_BUSY_TURN_MAX_SECS=3600 FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid"; fail "a freshly completed turn on a busy pane was still escalated: $(cat "$out")" + fi + [ ! -s "$out" ] || fail "a freshly completed turn on a busy pane printed a wake reason" + [ ! -e "$state/.stale-since-$key" ] || fail "a freshly completed turn did not clear the wedge timer" + [ ! -e "$state/.wedge-escalations-$key" ] || fail "a freshly completed turn did not clear the escalation counter" + reap "$pid" + pass "touching a busy worker's completed-turn marker resets the age and prevents an old-age escalation" +} + +test_busy_pane_repeated_escalation_reaches_demand_deep_inspection() { + local dir state fakebin out capture_file window key pane_hash sig pid n + dir=$(make_case busy-turn-age-demand-inspect); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-busy-demand-inspect" + printf 'Working...' > "$capture_file" + printf 'window=%s\nkind=ship\nharness=pi\n' "$window" > "$state/busy-demand.meta" + record_pi_busy "$state" busy-demand + printf 'working: setup complete\n' > "$state/busy-demand.status" + sig=$(seen_sig "$state/busy-demand.status"); printf '%s' "$sig" > "$state/.seen-busy-demand_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "Working...") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + touch -t 200001010000 "$state/busy-demand.turn-ended" + prime_turnend_seen "$state/busy-demand.turn-ended" + + # Priming round: first sighting past the turn-age bound absorbs and starts + # the wedge timer, mirroring the existing provably-working wedge tests. + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid"; fail "priming round for busy turn-age escalation was not absorbed: $(cat "$out")" + fi + reap "$pid" + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional busy-wedge priming stop" + + n=1 + while [ "$n" -le 3 ]; do + echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "busy turn-age escalation round $n did not escalate: $(cat "$out")" + grep -F "escalation $n" "$out" >/dev/null || fail "busy turn-age round $n did not report escalation count $n: $(cat "$out")" + if [ "$n" -lt 3 ]; then + grep -F "demand-deep-inspection" "$out" >/dev/null && fail "busy turn-age round $n escalated to demand-deep-inspection before the threshold: $(cat "$out")" + else + grep -F "demand-deep-inspection" "$out" >/dev/null || fail "busy turn-age round $n (threshold) did not demand deep inspection: $(cat "$out")" + fi + ack_stopped_cycle "$state" || fail "could not acknowledge busy turn-age escalation round $n" + n=$((n + 1)) + done + [ "$(cat "$state/.wedge-escalations-$key" 2>/dev/null || echo 0)" = 3 ] || fail "busy turn-age escalation counter did not persist across consecutive rounds" + pass "repeated busy turn-age escalations reuse the existing escalation counter and demand deep inspection at the threshold" +} + +# --- declared pause + busy pane: the busy-turn bound must honor the declaration +# A single foreground call can keep a declared external wait semantically busy +# past the completed-turn bound, bypassing the ordinary stale-pause path. +# This fixture pins all three halves of the contract: the declared pause is +# absorbed instead of wedged (A), it is still rechecked on the long +# PAUSE_RESURFACE_SECS cadence so a forgotten wait cannot rot invisibly (B), and +# lifting the declaration on the SAME busy over-age pane restores the wedge +# escalation, proving the discriminator is the worker's own declaration and not a +# blanket silencing of the escalator (C). +test_busy_declared_pause_is_rechecked_not_wedge_escalated() { + local dir state fakebin out capture_file window key sig pid statusf back + dir=$(make_case busy-declared-pause); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-review-scout" + statusf="$state/review-scout.status" + printf 'Working... (7200.4s) lavish-axi poll' > "$capture_file" + printf 'window=%s\nkind=scout\nharness=pi\n' "$window" > "$state/review-scout.meta" + record_pi_busy "$state" review-scout + printf 'paused: hosting the Lavish review, awaiting captain feedback\n' > "$statusf" + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-review-scout_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + # No completed turn for hours (the single blocking poll call): age the spawn + # record itself, exactly as the never-completed-a-turn fixtures above do. + touch -t 200001010000 "$state/review-scout.meta" + # No pre-seeded .hash-<key>: a live harness footer ticks, so every poll lands + # on the changed-hash branch - the review scout's real masking condition. + + # Phase A: past the bound, with the wedge threshold set as low as it goes, the + # declared pause is absorbed on the long cadence and never starts a wedge. + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ + FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=1 FM_PAUSE_RESURFACE_SECS=999 \ + FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_poll_cycle "$state" "$pid" || { reap "$pid"; fail "a declared pause on a busy review pane was escalated: $(cat "$out")"; } + reap "$pid" + [ ! -s "$out" ] || fail "a declared pause on a busy review pane printed a wake reason: $(cat "$out")" + [ -e "$state/.paused-$key" ] || fail "the busy-turn bound did not apply the declared-pause cadence" + [ ! -e "$state/.stale-since-$key" ] || fail "a declared pause on a busy pane started the wedge timer" + [ ! -e "$state/.wedge-escalations-$key" ] || fail "a declared pause on a busy pane incremented the escalation counter" + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional declared-pause phase-A stop" + + # Phase B: age the pause past the (now normal) long cadence and let the pane + # settle on one stable hash, so the still-busy pane takes the repeat-hash + # branch whose pause bookkeeping the bound must not wipe. It re-surfaces once + # as a recheck, never as a wedge. + back=$(( $(date +%s) - 500 )) + if [ "$(uname)" = Darwin ]; then touch -mt "$(date -r "$back" '+%Y%m%d%H%M.%S')" "$statusf" + else touch -m -d "@$back" "$statusf"; fi + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-review-scout_status" + printf '%s' "$(hash_text "$(cat "$capture_file")")" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ + FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=1 FM_PAUSE_RESURFACE_SECS=240 \ + FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || { reap "$pid"; fail "a declared pause past the long cadence was never rechecked"; } + grep -F "awaiting external" "$out" >/dev/null || fail "the recheck was not labeled a declared-pause recheck: $(cat "$out")" + grep -F "possible wedge" "$out" >/dev/null && fail "a declared pause on a busy pane was mislabeled a possible wedge: $(cat "$out")" + [ -e "$state/.paused-resurfaced-$key" ] || fail "the declared-pause re-surface throttle was cleared by the busy-turn bound" + [ ! -e "$state/.stale-since-$key" ] || fail "a declared-pause recheck used the wedge timer" + ack_stopped_cycle "$state" || fail "could not acknowledge the declared-pause recheck" + + # Phase C: the pause is lifted on the SAME busy, over-age pane. Nothing else + # changes, so a still-absorbed pane here would mean the bound was silenced + # rather than taught the declaration. It must wedge-escalate exactly as before. + printf 'working: review closed, resuming the sweep\n' > "$statusf" + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-review-scout_status" + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ + FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=999 FM_PAUSE_RESURFACE_SECS=999 \ + FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_poll_cycle "$state" "$pid" || { reap "$pid"; fail "a lifted pause escalated before the wedge threshold: $(cat "$out")"; } + reap "$pid" + [ -s "$state/.stale-since-$key" ] || fail "a lifted pause did not restore the busy-turn wedge timer" + [ ! -e "$state/.paused-$key" ] || fail "a lifted pause left stale declared-pause bookkeeping behind" + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional lifted-pause priming stop" + + echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ + FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=240 FM_PAUSE_RESURFACE_SECS=999 \ + FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || { reap "$pid"; fail "a lifted pause on an over-age busy pane no longer wedge-escalates"; } + grep -F "possible wedge" "$out" >/dev/null || fail "the restored busy-turn escalation did not flag a possible wedge: $(cat "$out")" + pass "a busy pane under a declared pause is rechecked on the long cadence, and lifting the pause restores the wedge escalation" +} + +# --- declared pause + busy pane + AWAY MODE: the bound must hand off, not decorate +# Away mode is daemon-owned: the watcher reverts to one-shot and lets the daemon +# classify. The busy-turn bound used to be the one stale path that ignored that, +# running the wedge timer under afk and handing the daemon a wake already decorated +# as a possible wedge. That decoration outranks the daemon's own pause verdict, so a +# crew that declared the wait itself was wedge-escalated once per +# FM_STALE_ESCALATE_SECS for as long as the wait lasted, with the escalation count +# climbing into demand-deep-inspection on a pane nobody needed to inspect. +# Phase A pins the handoff: the plain window identity, no wedge timer, no escalation +# counter, and no normal-mode pause bookkeeping (the daemon owns that in away mode). +# Phase B re-arms on the same unchanged pane and pins the one-shot: a second wake +# here is what the climbing ladder looked like. Phase C drives the discriminator +# apart on the SAME afk, busy, over-age pane - lifting the declaration restores the +# wedge escalation, so this is the worker's declaration being honored rather than +# away mode silencing the escalator. +test_afk_busy_declared_pause_hands_off_plain_stale() { + local dir state fakebin out capture_file window key sig pid statusf + dir=$(make_case afk-busy-declared-pause); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-afk-review-scout" + statusf="$state/afk-review-scout.status" + printf 'Working... (7200.4s) lavish-axi poll' > "$capture_file" + printf 'window=%s\nkind=scout\nharness=pi\n' "$window" > "$state/afk-review-scout.meta" + record_pi_busy "$state" afk-review-scout + printf 'paused: hosting the Lavish review, awaiting captain feedback\n' > "$statusf" + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-afk-review-scout_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + touch -t 200001010000 "$state/afk-review-scout.meta" + date '+%s' > "$state/.afk" + + # Phase A: past the bound, with the wedge threshold as low as it goes, the + # declaration is handed to the daemon undecorated instead of being wedge-timed. + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ + FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=1 FM_PAUSE_RESURFACE_SECS=999 \ + FM_POLL=0.2 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 150 || { reap "$pid"; fail "the away-mode busy-turn bound never handed the declared pause to the daemon"; } + grep -Fx "stale: $window" "$out" >/dev/null \ + || fail "the away-mode busy-turn bound did not hand off the plain window identity: $(cat "$out")" + grep -F "possible wedge" "$out" >/dev/null \ + && fail "away mode decorated a declared pause as a possible wedge: $(cat "$out")" + [ ! -e "$state/.stale-since-$key" ] \ + || fail "the away-mode handoff started the wedge timer on a declared pause" + [ ! -e "$state/.wedge-escalations-$key" ] \ + || fail "the away-mode handoff incremented the wedge escalation count on a declared pause" + [ ! -e "$state/.paused-$key" ] \ + || fail "the away-mode handoff recorded normal-mode pause tracking instead of leaving it to the daemon" + ack_stopped_cycle "$state" || fail "could not acknowledge the away-mode declared-pause handoff" + + # Phase B: re-arm on the same unchanged pane. The bound has already handed this + # stale hash off, so it must stay silent rather than re-waking the daemon - a + # second wake here is the escalation ladder the wedge timer used to climb. + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ + FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=1 FM_PAUSE_RESURFACE_SECS=999 \ + FM_POLL=0.2 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_poll_cycle "$state" "$pid" || { reap "$pid"; fail "the away-mode bound re-woke on an already-handed-off declared pause: $(cat "$out")"; } + reap "$pid" + [ ! -s "$out" ] || fail "the away-mode bound re-surfaced an already-handed-off declared pause: $(cat "$out")" + [ ! -e "$state/.wedge-escalations-$key" ] \ + || fail "re-arming on an unchanged declared pause started a wedge escalation ladder" + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional away-mode re-arm stop" + + # Phase C: lift the declaration on the SAME afk, busy, over-age pane. Nothing else + # changes, so a wedge escalation here proves the declaration was the discriminator. + printf 'working: resumed the review write-up\n' > "$statusf" + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-afk-review-scout_status" + echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ + FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=240 FM_PAUSE_RESURFACE_SECS=999 \ + FM_POLL=0.2 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 150 || { reap "$pid"; fail "a lifted pause on an away-mode over-age busy pane no longer wedge-escalates"; } + grep -F "possible wedge" "$out" >/dev/null \ + || fail "the restored away-mode busy-turn escalation did not flag a possible wedge: $(cat "$out")" + pass "away mode hands a busy declared pause to the daemon as a plain stale, and lifting the declaration restores the wedge escalation" +} + +# --- declared pause + busy pane + AWAY MODE + a TICKING footer: one wake per declaration +# The static-pane case above cannot tell a hash-keyed one-shot from a +# declaration-keyed one, because its capture never changes between polls. The +# incident pane's harness footer ticks on every capture, so a one-shot keyed on the +# pane hash re-fires on every poll, and the daemon, which relaunches the watcher +# after each handled wake, is woken in a loop for the whole declared wait. This +# fixture's fake tmux renders a fresh footer on EVERY capture-pane and asserts that +# divergence outright on every re-arm (.hash-<key> moves, .count-<key> never +# climbs), so the one-wake assertion across five silent re-arms cannot pass +# vacuously on a pane that happened to sit still. Round 1 also starts from an +# undeclared wedge timer and escalation count, which the handoff must clear the +# way the normal-mode absorber does, so lifting the declaration later starts the +# wedge path from a fresh timer rather than resuming a stale count. +test_afk_busy_declared_pause_ticking_pane_hands_off_once() { + local dir state fakebin out drain_out window key sig pid statusf ticks round prev_hash cur_hash prev_ticks + dir=$(make_case afk-busy-declared-pause-ticking); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; window="test:fm-afk-ticking-scout" + statusf="$state/afk-ticking-scout.status"; ticks="$dir/ticks" + cat > "$fakebin/tmux" <<'SH' +#!/usr/bin/env bash +set -u +case "${1:-}" in + list-windows) + [ -n "${FM_FAKE_TMUX_WINDOW:-}" ] && printf '%s\n' "${FM_FAKE_TMUX_WINDOW#*:}" + exit 0 ;; + capture-pane) + n=$(( $(cat "$FM_FAKE_TMUX_TICKS" 2>/dev/null || echo 0) + 1 )) + echo "$n" > "$FM_FAKE_TMUX_TICKS" + printf 'Working... (%d.%ds) lavish-axi poll' "$(( 7200 + n ))" "$(( n % 10 ))" + exit 0 ;; + display-message) + case "$*" in + *pane_current_command*) printf '%s\n' "${FM_FAKE_TMUX_CURRENT_COMMAND:-}"; exit 0 ;; + esac ;; +esac +exit 1 +SH + chmod +x "$fakebin/tmux" + printf 'window=%s\nkind=scout\nharness=pi\n' "$window" > "$state/afk-ticking-scout.meta" + record_pi_busy "$state" afk-ticking-scout + printf 'paused: hosting the Lavish review, awaiting captain feedback\n' > "$statusf" + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-afk-ticking-scout_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + touch -t 200001010000 "$state/afk-ticking-scout.meta" + date '+%s' > "$state/.afk" + # An undeclared busy phase already ran the wedge timer and escalated twice + # before the crew declared the wait. + echo $(( $(date +%s) - 500 )) > "$state/.stale-since-$key" + printf '2\n' > "$state/.wedge-escalations-$key" + date +%s > "$state/.writing-since-$key" + + # Round 1: the declaration is handed off once, undecorated, and the undeclared + # phase's wedge bookkeeping is cleared with it. + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_TICKS="$ticks" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ + FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=1 FM_PAUSE_RESURFACE_SECS=999 \ + FM_POLL=0.2 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 150 || { reap "$pid"; fail "the away-mode busy-turn bound never handed a ticking declared pause to the daemon"; } + grep -Fx "stale: $window" "$out" >/dev/null \ + || fail "the away-mode busy-turn bound did not hand off the plain window identity for a ticking pane: $(cat "$out")" + grep -F "possible wedge" "$out" >/dev/null \ + && fail "away mode decorated a ticking declared pause as a possible wedge: $(cat "$out")" + [ ! -e "$state/.stale-since-$key" ] \ + || fail "the away-mode handoff left the undeclared phase's wedge timer in place" + [ ! -e "$state/.wedge-escalations-$key" ] \ + || fail "the away-mode handoff left the undeclared phase's escalation count in place" + [ ! -e "$state/.writing-since-$key" ] \ + || fail "the away-mode handoff left the undeclared phase's write-deferral chain in place" + [ ! -e "$state/.paused-$key" ] \ + || fail "the away-mode handoff recorded normal-mode pause tracking on a ticking pane" + ack_stopped_cycle "$state" || fail "could not acknowledge the ticking declared-pause handoff" + + # Rounds 2-6: five consecutive re-arms on the same standing declaration. Every + # capture renders a new footer, so every poll lands on the changed-hash branch - + # the exact shape a hash-keyed one-shot re-fires on. Each round proves the pane + # really moved before it asserts silence, so the case cannot go vacuous. + round=2 + while [ "$round" -le 6 ]; do + prev_hash=$(cat "$state/.hash-$key" 2>/dev/null || true) + prev_ticks=$(cat "$ticks" 2>/dev/null || echo 0) + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_TICKS="$ticks" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_FAKE_CREW_STATE='state: working · source: pane · harness busy (pi-ext)' \ + FM_BUSY_TURN_MAX_SECS=1 FM_STALE_ESCALATE_SECS=1 FM_PAUSE_RESURFACE_SECS=999 \ + FM_POLL=0.2 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_poll_cycle "$state" "$pid" || { reap "$pid"; fail "re-arm $round on a ticking declared pause re-woke the daemon: $(cat "$out")"; } + reap "$pid" + cur_hash=$(cat "$state/.hash-$key" 2>/dev/null || true) + [ "$(cat "$ticks" 2>/dev/null || echo 0)" -gt "$prev_ticks" ] \ + || fail "re-arm $round never captured the pane, so its silence proves nothing" + [ -n "$cur_hash" ] && [ "$cur_hash" != "$prev_hash" ] \ + || fail "re-arm $round saw the same pane hash as the round before, so it cannot tell a hash-keyed one-shot from a declaration-keyed one" + [ "$(cat "$state/.count-$key" 2>/dev/null || echo missing)" = 0 ] \ + || fail "re-arm $round settled on a stable hash instead of ticking on every poll" + [ ! -s "$out" ] || fail "re-arm $round re-surfaced a standing declared pause on a ticking pane: $(cat "$out")" + [ ! -e "$state/.stale-since-$key" ] \ + || fail "re-arm $round started the wedge timer on a standing declared pause" + [ ! -e "$state/.wedge-escalations-$key" ] \ + || fail "re-arm $round climbed the wedge escalation ladder on a standing declared pause" + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional re-arm $round stop" + round=$((round + 1)) + done + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || true + grep "$(printf '\tstale\t')" "$drain_out" >/dev/null \ + && fail "the silent re-arms still queued a stale row for the standing declaration: $(cat "$drain_out")" + pass "away mode wakes the daemon once per declaration for a busy pane whose footer ticks on every capture" +} + +# Behavioral proof that the production default (no FM_BUSY_TURN_MAX_SECS override +# anywhere in this env) is 3600s: a completed turn 5 minutes old must not start a +# wedge timer, while one 66 minutes old must - bracketing the default around 3600 +# without waiting a literal hour. +test_busy_pane_default_turn_age_bound_is_3600s() { + local dir state fakebin out capture_file window key pane_hash sig pid + dir=$(make_case busy-default-turn-age); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt"; window="test:fm-busy-default" + printf 'Working...' > "$capture_file" + printf 'window=%s\nkind=ship\nharness=pi\n' "$window" > "$state/busy-default.meta" + record_pi_busy "$state" busy-default + printf 'working: setup complete\n' > "$state/busy-default.status" + sig=$(seen_sig "$state/busy-default.status"); printf '%s' "$sig" > "$state/.seen-busy-default_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "Working...") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + + set_mtime $(( $(date +%s) - 300 )) "$state/busy-default.turn-ended" + prime_turnend_seen "$state/busy-default.turn-ended" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid"; fail "a 5-minute-old completed turn tripped the default busy-turn-age bound: $(cat "$out")" + fi + [ ! -e "$state/.stale-since-$key" ] || fail "a 5-minute-old completed turn started a wedge timer under the default bound" + reap "$pid" + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional five-minute-bound stop" + + set_mtime $(( $(date +%s) - 4000 )) "$state/busy-default.turn-ended" + prime_turnend_seen "$state/busy-default.turn-ended" + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid"; fail "a 66-minute-old completed turn escalated before the wedge threshold under the default bound: $(cat "$out")" + fi + [ -s "$state/.stale-since-$key" ] || fail "a 66-minute-old completed turn did not start a wedge timer under the default bound (default is not 3600s)" + reap "$pid" + pass "the production default busy-turn-age bound is 3600s (5min under does not wedge, 66min over does)" +} + +test_nonterminal_stale_repairs_missing_or_corrupt_timer() { + local dir state fakebin out capture_file window key pane_hash sig pid since + dir=$(make_case nonterminal-stale-timer-repair); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt" + window="test:fm-quiet-timer" + printf 'idle building output' > "$capture_file" + printf 'window=%s\nkind=ship\n' "$window" > "$state/quiet-timer.meta" + printf 'working: still compiling\n' > "$state/quiet-timer.status" + sig=$(seen_sig "$state/quiet-timer.status"); printf '%s' "$sig" > "$state/.seen-quiet-timer_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "idle building output") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + printf '%s' "$pane_hash" > "$state/.stale-$key" + + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_numeric_file "$state/.stale-since-$key" 30 || { reap "$pid"; fail "matching stale suppressor with missing timer did not initialize stale-since"; } + if ! kill -0 "$pid" 2>/dev/null; then + wait "$pid" 2>/dev/null || true + fail "watcher exited while repairing a missing stale-since timer: $(cat "$out")" + fi + [ ! -s "$state/.wake-queue" ] || { reap "$pid"; fail "missing stale-since repair enqueued a wake"; } + reap "$pid" + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional missing-timer repair stop" + + printf 'corrupt\n' > "$state/.stale-since-$key" + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_STALE_ESCALATE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_numeric_file "$state/.stale-since-$key" 30 || { reap "$pid"; fail "matching stale suppressor with corrupt timer did not repair stale-since"; } + since=$(cat "$state/.stale-since-$key" 2>/dev/null || true) + [ "$since" != "corrupt" ] || { reap "$pid"; fail "corrupt stale-since value was left in place"; } + [ ! -s "$state/.wake-queue" ] || { reap "$pid"; fail "corrupt stale-since repair enqueued a wake"; } + reap "$pid" + pass "matching non-terminal stale suppressors repair missing or corrupt stale-since timers" +} + +# --- quiet pane, worktree still being written: deferred, never wedge-escalated - +# The live 2026-08-14 case: one crew produced eight consecutive possible-wedge +# escalations in an afternoon, three of them demanding deep inspection, while it +# was demonstrably writing source, then tests, then documentation. The detector's +# two inputs (pane quietness, run step) cannot see that, so the pane looks frozen. +# Both halves of the contract are asserted on the SAME fixture, because the whole +# point is that only the worktree evidence differs: writing defers, silent +# escalates on the unchanged schedule. +# Every wait below is the file's standard one (wait_poll_cycle for an absorbing +# watcher, a 100-tick wait_for_exit for an escalating one), because the poll these +# tests assert on is the ONE poll that spawns the bounded worktree walk: on a +# loaded runner it outlives a fixed liveness budget, and a round reaped before it +# finished reports a lost deferral instead of the deferral under test. +test_wedge_escalation_deferred_while_worktree_is_written() { + local dir state fakebin out drain_out capture_file window key pane_hash sig pid wt back + dir=$(make_case wedge-worktree-writes); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + window="test:fm-writing"; wt="$dir/wt" + mkdir -p "$wt/src" + printf 'idle building output' > "$capture_file" + printf 'window=%s\nkind=ship\nworktree=%s\n' "$window" "$wt" > "$state/writing.meta" + printf 'working: implementing\n' > "$state/writing.status" + sig=$(seen_sig "$state/writing.status"); printf '%s' "$sig" > "$state/.seen-writing_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "idle building output") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + # Already-classified hash with an idle window that opened 500s ago, so the very + # first stale poll lands straight on the at-threshold wedge branch (this repeat + # path never re-reads crew state, so the worktree evidence is the only input + # that can change the outcome). + printf '%s' "$pane_hash" > "$state/.stale-$key" + back=$(( $(date +%s) - 500 )) + echo "$back" > "$state/.stale-since-$key" + set_mtime "$back" "$state/.stale-since-$key" + + # Phase A: the crew wrote a file after the idle window opened. Deferred. + printf 'int main(void) { return 0; }\n' > "$wt/src/main.c" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 \ + FM_PAUSE_RESURFACE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid"; fail "watcher wedge-escalated a quiet pane whose worktree was being written: $(cat "$out")" + fi + [ ! -s "$out" ] || { reap "$pid"; fail "a written-worktree deferral printed a wake reason: $(cat "$out")"; } + [ ! -s "$state/.wake-queue" ] || { reap "$pid"; fail "a written-worktree deferral enqueued a wake"; } + [ -e "$state/.writing-since-$key" ] || { reap "$pid"; fail "the write-deferral chain marker was not recorded"; } + [ ! -e "$state/.wedge-escalations-$key" ] || { reap "$pid"; fail "a deferral advanced the wedge escalation counter"; } + [ "$(cat "$state/.stale-since-$key" 2>/dev/null || echo 0)" -gt "$back" ] \ + || { reap "$pid"; fail "a deferral did not restart the idle timer, so the next window cannot re-probe"; } + reap "$pid" + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional phase-A watcher stop" + + # Phase B: same fixture, same quiet pane, but nothing written during this idle + # window (the crew really is stalled). The unchanged schedule must still fire. + set_mtime "$(( $(date +%s) - 900 ))" "$wt/src/main.c" + echo "$back" > "$state/.stale-since-$key" + set_mtime "$back" "$state/.stale-since-$key" + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 \ + FM_PAUSE_RESURFACE_SECS=999 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "a stalled crew that wrote nothing did not wedge-escalate on the existing schedule" + grep -F "stale: $window" "$out" >/dev/null || fail "the stalled-crew escalation did not print a stale wake" + grep -F "possible wedge" "$out" >/dev/null || fail "the stalled-crew escalation did not flag a possible wedge" + [ "$(cat "$state/.wedge-escalations-$key" 2>/dev/null || true)" = 1 ] || fail "the stalled-crew escalation was not counted" + [ ! -e "$state/.stale-since-$key" ] || fail "the idle timer was not cleared after a real escalation" + [ ! -e "$state/.writing-since-$key" ] || fail "the write-deferral chain outlived a real escalation" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the stalled-crew escalation failed" + grep "$(printf '\tstale\t')" "$drain_out" | grep -F "$window" >/dev/null || fail "the stalled-crew escalation was not queued" + pass "a quiet pane writing its own worktree is deferred, while one writing nothing still wedge-escalates on the unchanged schedule" +} + +# A deferral is not silence. A worktree can churn without real progress (a +# rewritten log, a build touching the same file), so the whole deferral chain ages +# and re-surfaces once per PAUSE_RESURFACE_SECS - the same bounded cadence a +# declared pause uses - labeled as a recheck rather than a wedge. +test_write_deferral_resurfaces_on_the_bounded_cadence() { + local dir state fakebin out drain_out capture_file window key pane_hash sig pid wt back + dir=$(make_case wedge-worktree-resurface); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + window="test:fm-churn"; wt="$dir/wt" + mkdir -p "$wt/src" + printf 'idle building output' > "$capture_file" + printf 'window=%s\nkind=ship\nworktree=%s\n' "$window" "$wt" > "$state/churn.meta" + printf 'working: implementing\n' > "$state/churn.status" + sig=$(seen_sig "$state/churn.status"); printf '%s' "$sig" > "$state/.seen-churn_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "idle building output") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + printf '%s' "$pane_hash" > "$state/.stale-$key" + back=$(( $(date +%s) - 500 )) + echo "$back" > "$state/.stale-since-$key" + set_mtime "$back" "$state/.stale-since-$key" + # This pane has been deferring on write evidence for 500s already. + : > "$state/.writing-since-$key" + set_mtime "$back" "$state/.writing-since-$key" + printf 'churn\n' > "$wt/src/main.c" + + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_STALE_ESCALATE_SECS=240 \ + FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "a long-running write deferral never re-surfaced on the bounded cadence" + grep -F "stale: $window" "$out" >/dev/null || fail "the write-deferral recheck did not print a stale wake" + grep -F "writing its worktree" "$out" >/dev/null || fail "the write-deferral recheck was not labeled as such" + grep -F "possible wedge" "$out" >/dev/null && fail "a write-deferral recheck was mislabeled a possible wedge" + [ -e "$state/.writing-resurfaced-$key" ] || fail "the write-deferral re-surface throttle marker was not recorded" + [ ! -e "$state/.wedge-escalations-$key" ] || fail "a write-deferral recheck advanced the wedge escalation counter" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the write-deferral recheck failed" + grep "$(printf '\tstale\t')" "$drain_out" | grep -F "$window" >/dev/null || fail "the write-deferral recheck was not queued" + pass "a write deferral re-surfaces once on the bounded pause cadence, so a churning worktree cannot stay invisible" +} + +# The worktree recorded for a secondmate is a provisioned firstmate home, and that +# home runs its OWN supervision inside itself: its watcher beacon, pane hashes and +# heartbeats keep state/ churning whether or not the mate produced anything. Reading +# that as crew progress would quietly relax the kind-agnostic busy-turn backstop from +# the escalation cadence to the hourly recheck for work that produced nothing, so the +# probe must report no evidence and the unchanged schedule must still fire. +test_secondmate_home_supervision_churn_is_not_write_evidence() { + local dir state fakebin out drain_out capture_file window key sig pid home back + dir=$(make_case secondmate-home-churn); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + window="test:fm-mate"; home="$dir/mate-home" + mkdir -p "$home/state" + printf 'sm-mate\n' > "$home/.fm-secondmate-home" + printf 'Working... (12.3s)' > "$capture_file" + printf 'window=%s\nkind=ship\nharness=pi\nworktree=%s\n' "$window" "$home" > "$state/mate.meta" + record_pi_busy "$state" mate + # An ordinary crew recording a provisioned mate home is the route that actually + # reaches the probe: a kind=secondmate window of its own is triaged only under a + # declared pause, and a declared pause takes the bounded recheck cadence instead of + # the wedge timer. The home marker alone is what excludes the walk, so the exclusion + # is what this asserts. A busy pane is bounded by its completed-turn age; no turn + # ever completed here, so the spawn record itself is aged past the bound that routes + # it into the wedge timer. + printf 'working: implementing\n' > "$state/mate.status" + sig=$(seen_sig "$state/mate.status"); printf '%s' "$sig" > "$state/.seen-mate_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + set_mtime "$(( $(date +%s) - 4000 ))" "$state/mate.meta" + back=$(( $(date +%s) - 500 )) + echo "$back" > "$state/.stale-since-$key" + set_mtime "$back" "$state/.stale-since-$key" + # The only thing written since the idle window opened is the mate home's own + # supervision bookkeeping. + printf 'beat\n' > "$home/state/.last-watcher-beat" + + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_STALE_ESCALATE_SECS=240 FM_BUSY_TURN_MAX_SECS=1 FM_PAUSE_RESURFACE_SECS=999 \ + FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "a mate home's own supervision churn deferred an escalation it must not defer" + grep -F "stale: $window" "$out" >/dev/null || fail "the mate-home escalation did not print a stale wake" + grep -F "possible wedge" "$out" >/dev/null || fail "the mate-home escalation did not flag a possible wedge" + [ ! -e "$state/.writing-since-$key" ] || fail "a mate's provisioned home was probed as if it were a code tree" + [ "$(cat "$state/.wedge-escalations-$key" 2>/dev/null || true)" = 1 ] || fail "the mate escalation was not counted" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the mate escalation failed" + grep "$(printf '\tstale\t')" "$drain_out" | grep -F "$window" >/dev/null || fail "the mate escalation was not queued" + pass "a secondmate's own home supervision churn is not crew write evidence, so a pane recording that home keeps the unchanged escalation schedule" +} + +# A write deferral is a bounded chain, not a permanent one: its .writing-since +# marker ages the whole chain so a churning worktree still re-surfaces once per +# PAUSE_RESURFACE_SECS. That only holds while the chain belongs to the CURRENT quiet +# stretch, so every path that restarts the idle-window timer must drop it too. The +# reachable case is a pane that deferred on write evidence and later has its timer +# repaired: a long-finished chain would make the first deferral of the new window +# re-surface immediately instead of after a fresh window. +test_timer_repair_drops_a_finished_write_deferral_chain() { + local dir state fakebin out capture_file window key pane_hash sig pid wt back + dir=$(make_case wedge-write-chain-timer-repair); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt" + window="test:fm-chain-repair"; wt="$dir/wt" + mkdir -p "$wt/src" + printf 'idle building output' > "$capture_file" + printf 'window=%s\nkind=ship\nworktree=%s\n' "$window" "$wt" > "$state/chain-repair.meta" + printf 'working: implementing\n' > "$state/chain-repair.status" + sig=$(seen_sig "$state/chain-repair.status"); printf '%s' "$sig" > "$state/.seen-chain-repair_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "idle building output") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + printf '%s' "$pane_hash" > "$state/.stale-$key" + # A deferral chain left over from an earlier quiet stretch, already well past the + # bounded re-surface window. + back=$(( $(date +%s) - 5000 )) + : > "$state/.writing-since-$key" + set_mtime "$back" "$state/.writing-since-$key" + # The idle-window timer is corrupt, so this poll repairs it and opens a NEW quiet + # window without probing the worktree at all. + printf 'corrupt\n' > "$state/.stale-since-$key" + + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_STALE_ESCALATE_SECS=240 FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + # Watcher startup performs bounded recovery scans before its first stale poll; + # give this positive marker assertion the same loaded-runner budget as the + # suite's other startup-sensitive waits instead of failing after only 3s. + wait_numeric_file "$state/.stale-since-$key" 100 \ + || { reap "$pid"; fail "the corrupt idle-window timer was not repaired"; } + [ ! -e "$state/.writing-since-$key" ] \ + || { reap "$pid"; fail "an idle-window timer repair kept a finished write-deferral chain"; } + [ ! -s "$state/.wake-queue" ] || { reap "$pid"; fail "the idle-window timer repair enqueued a wake"; } + reap "$pid" + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional timer-repair watcher stop" + + # The new quiet window now crosses the escalation threshold while the crew writes + # its worktree. That deferral must get a FRESH re-surface window rather than + # inheriting the finished chain's age. + back=$(( $(date +%s) - 500 )) + echo "$back" > "$state/.stale-since-$key" + set_mtime "$back" "$state/.stale-since-$key" + printf 'int main(void) { return 0; }\n' > "$wt/src/main.c" + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_STALE_ESCALATE_SECS=240 FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid" + fail "the first deferral of a new quiet window re-surfaced at once, so it inherited a finished chain: $(cat "$out")" + fi + [ ! -s "$out" ] || { reap "$pid"; fail "a fresh write deferral printed a wake reason: $(cat "$out")"; } + [ ! -s "$state/.wake-queue" ] || { reap "$pid"; fail "a fresh write deferral enqueued a wake"; } + [ -e "$state/.writing-since-$key" ] || { reap "$pid"; fail "the new deferral recorded no chain marker"; } + [ ! -e "$state/.writing-resurfaced-$key" ] \ + || { reap "$pid"; fail "a fresh write deferral spent its bounded re-surface on the first poll"; } + reap "$pid" + pass "an idle-window timer repair drops a finished write-deferral chain, so the next deferral gets a fresh re-surface window" +} + +# The same chain must not outlive either first-sight path through a captain-relevant +# status line, because both also open a new idle window: the provably-working absorb +# and the plain surface. +test_terminal_first_sight_drops_a_finished_write_deferral_chain() { + local dir state fakebin out capture_file window key pane_hash sig pid wt back + dir=$(make_case wedge-write-chain-first-sight); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; capture_file="$dir/pane.txt" + window="test:fm-chain-firstsight"; wt="$dir/wt" + mkdir -p "$wt/src" + printf 'no-mistakes axi run: validating...' > "$capture_file" + printf 'window=%s\nkind=ship\nworktree=%s\n' "$window" "$wt" > "$state/chain-first.meta" + printf 'done: implementation complete, ready to validate\n' > "$state/chain-first.status" + sig=$(seen_sig "$state/chain-first.status"); printf '%s' "$sig" > "$state/.seen-chain-first_status" + key=$(printf '%s' "$window" | tr ':/.' '___') + pane_hash=$(hash_text "no-mistakes axi run: validating...") + printf '%s' "$pane_hash" > "$state/.hash-$key" + printf '1\n' > "$state/.count-$key" + back=$(( $(date +%s) - 5000 )) + : > "$state/.writing-since-$key" + set_mtime "$back" "$state/.writing-since-$key" + export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' + + # First sight of this hash, absorbed because the active run outranks the stale + # captain-relevant line. The absorb opens a new idle window, so the finished chain + # must go with it. + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_STALE_ESCALATE_SECS=999 FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid"; fail "the overridden terminal status was not absorbed on first sight: $(cat "$out")" + fi + [ "$(cat "$state/.stale-$key" 2>/dev/null || true)" = "$pane_hash" ] \ + || { reap "$pid"; fail "the first-sight absorb did not advance the stale suppressor"; } + [ ! -e "$state/.writing-since-$key" ] \ + || { reap "$pid"; fail "the provably-working first-sight absorb kept a finished write-deferral chain"; } + reap "$pid" + ack_stopped_cycle "$state" || fail "could not acknowledge the intentional first-sight absorb stop" + + # Same pane, first sight again, but nothing overrides the status line now, so it + # surfaces. That path drops the idle-window timer, so it must drop the chain too. + rm -f "$state/.stale-$key" "$state/.stale-since-$key" + printf '1\n' > "$state/.count-$key" + : > "$state/.writing-since-$key" + set_mtime "$back" "$state/.writing-since-$key" + FM_FAKE_CREW_STATE='state: unknown · source: none · no run, no busy pane' + : > "$out" + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" \ + FM_STALE_ESCALATE_SECS=999 FM_PAUSE_RESURFACE_SECS=240 FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "a first-sight captain-relevant status was not surfaced" + grep -F "stale: $window" "$out" >/dev/null || fail "the first-sight surface did not print a stale wake" + [ ! -e "$state/.writing-since-$key" ] \ + || fail "the first-sight surface kept a finished write-deferral chain" + unset FM_FAKE_CREW_STATE + pass "both first-sight paths through a captain-relevant status drop a finished write-deferral chain with the idle window" +} + +# --- triage debug log stays size capped ------------------------------------- + +test_triage_log_size_cap_accepts_spaced_wc_counts() { + local dir state fakebin out status_file pid lines i + dir=$(make_case triage-log-spaced-wc); state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out" + i=1 + while [ "$i" -le 3000 ]; do + printf 'old line %04d\n' "$i" >> "$state/.watch-triage.log" + i=$((i + 1)) + done + cat > "$fakebin/wc" <<'SH' +#!/usr/bin/env bash +set -u +if [ "${1:-}" = "-c" ]; then + printf ' 999999\n' + exit 0 +fi +exit 127 +SH + chmod +x "$fakebin/wc" + status_file="$state/task.status" + printf 'working: compiling step 2\n' > "$status_file" + # Provably working so the no-verb signal is absorbed (which is what writes the + # triage log line under test). + export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' + PATH="$fakebin:$PATH" FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 FM_WATCH_TRIAGE_LOG_MAX_BYTES=1 "$WATCH" > "$out" & + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid"; fail "watcher exited for a benign signal while testing log capping: $(cat "$out")" + fi + i=0 + while [ "$i" -lt 30 ]; do + lines=$(awk 'END { print NR + 0 }' "$state/.watch-triage.log") + [ "$lines" -le 2000 ] && break + sleep 0.1 + i=$((i + 1)) + done + [ "$lines" -le 2000 ] || { reap "$pid"; fail "triage log was not capped when wc emitted a spaced byte count (lines=$lines)"; } + [ ! -s "$state/.wake-queue" ] || { reap "$pid"; fail "benign signal enqueued a wake while testing log capping"; } + reap "$pid" + pass "triage log capping handles wc byte counts with leading spaces" +} + +# --- process-event delivery ------------------------------------------------- +# A durably captured process-event result publishes an ordinary `check` wake on +# the durable queue. The watcher must deliver that queued wake proactively - +# print an actionable reason and exit into the same rewake path every other +# actionable wake uses - rather than leaving it to be found by a manual drain. + +# Run the runner against a case home. FM_ROOT_OVERRIDE (exported by the shared +# wake harness to keep the drain's tangle check inert) would otherwise point the +# runner at a root with no installed adapters, and the claim root must stay +# inside the case so nothing here can observe a real home's source ownership. +pe_case() { # <dir> <command>... + local dir=$1 + dir=$(cd "$dir" && pwd -P) || return 1 + shift + (unset FM_ROOT_OVERRIDE + FM_PROCEVENT_CLAIM_ROOT="$dir/claims" FM_HOME="$dir" "$ROOT/bin/fm-procevent.sh" "$@") +} + +# Capture one real process-event result into <dir>'s home, then retire the +# source so the fixture holds exactly the reported end state: one durably +# captured, unhandled, queued result and no remaining poll work. +seed_captured_procevent_result() { # <dir> + local dir=$1 i=0 + pe_case "$dir" register lavish delivery-src -- \ + /bin/sh -c 'printf "session:\n file: /a.html\n status: waiting\n"' >/dev/null || return 1 + pe_case "$dir" reconcile >/dev/null || return 1 + while [ "$i" -lt 100 ]; do + [ -s "$dir/state/.wake-queue" ] && break + sleep 0.1 + i=$((i + 1)) + done + pe_case "$dir" retire delivery-src >/dev/null || return 1 + [ -s "$dir/state/.wake-queue" ] +} + +# The watcher, scoped by FM_HOME rather than FM_STATE_OVERRIDE, so the +# per-cycle reconcile it launches resolves the same home's state. +procevent_watch_bg() { # <dir> <out> + local dir=$1 out=$2 + dir=$(cd "$dir" && pwd -P) || return 1 + PATH="$dir/fakebin:$PATH" FM_HOME="$dir" FM_PROCEVENT_CLAIM_ROOT="$dir/claims" \ + FM_CREW_STATE_BIN="$dir/fakebin/fm-crew-state.sh" \ + FM_POLL=0.2 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & +} + +test_procevent_captured_result_surfaces_proactively() { + local dir state out drain_out pid beacon_age + dir=$(make_case procevent-delivery); state="$dir/state" + out="$dir/watch.out"; drain_out="$dir/drain.out" + seed_captured_procevent_result "$dir" || fail "the fixture captured no process-event result" + grep -F "procevent lavish delivery-src 1" "$state/.wake-queue" >/dev/null \ + || fail "the captured result was never published to the durable queue" + + procevent_watch_bg "$dir" "$out" + pid=$! + wait_for_exit "$pid" 100 \ + || fail "a healthy watcher never surfaced a durably captured process-event result: $(cat "$out")" + grep -F "check:" "$out" >/dev/null \ + || fail "the process-event wake was not reported as an actionable check: $(cat "$out")" + grep -F "procevent:delivery-src:1" "$out" >/dev/null \ + || fail "the actionable reason did not name the queued result: $(cat "$out")" + beacon_age=$(FM_STATE_OVERRIDE="$state" bash -c \ + '. "$1/bin/fm-wake-lib.sh"; fm_path_age "$2"' _ "$ROOT" "$state/.last-watcher-beat") + [ "$beacon_age" -lt 60 ] || fail "the surfacing watcher was not a healthy one (beacon age ${beacon_age}s)" + + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the process-event wake failed" + grep "$(printf '\tcheck\t')" "$drain_out" | grep -F "procevent lavish delivery-src 1" >/dev/null \ + || fail "the process-event result was not queued for the drain that follows the wake" + pass "a captured process-event result wakes a healthy watcher proactively, with no manual drain" +} + +test_procevent_unacknowledged_result_redrains_until_handled() { + local dir state out replay_out replay_err pid before after sequence generation + dir=$(make_case procevent-redrain); state="$dir/state" + out="$dir/watch.out"; replay_out="$dir/replay.out"; replay_err="$dir/replay.err" + seed_captured_procevent_result "$dir" || fail "the fixture captured no process-event result" + + procevent_watch_bg "$dir" "$out" + pid=$! + wait_for_exit "$pid" 100 || fail "the first proactive wake never happened: $(cat "$out")" + FM_STATE_OVERRIDE="$state" "$DRAIN" >/dev/null 2>&1 || fail "drain after the first process-event wake failed" + + # An interrupted handler leaves the captured result durable. The successor + # must re-surface it through recovery, then its drain must print the same row. + : > "$out" + procevent_watch_bg "$dir" "$out" + pid=$! + wait_for_exit "$pid" 100 \ + || fail "an unacknowledged process-event result was not re-surfaced on re-arm: $(cat "$out")" + grep -F 'check: rearm-resurface' "$out" >/dev/null \ + || fail "the successor did not report recovery for the unacknowledged result: $(cat "$out")" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$replay_out" 2> "$replay_err" \ + || fail "the successor could not re-drain the unacknowledged process-event result" + grep "$(printf '\tcheck\t')" "$replay_out" | grep -F 'procevent lavish delivery-src 1' >/dev/null \ + || fail "the successor drain did not re-print the durable process-event row" + + pe_case "$dir" handled delivery-src 1 >/dev/null || fail "could not acknowledge the captured result" + sequence=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through \([0-9][0-9]*\) --recovery-generation [A-Za-z0-9._-][A-Za-z0-9._-]*$/\1/p' "$replay_err") + generation=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through [0-9][0-9]* --recovery-generation \([A-Za-z0-9._-][A-Za-z0-9._-]*\)$/\1/p' "$replay_err") + [ -n "$sequence" ] && [ -n "$generation" ] \ + || fail "the replay drain omitted its post-handling acknowledgement boundary" + FM_STATE_OVERRIDE="$state" "$DRAIN" --ack-through "$sequence" --recovery-generation "$generation" \ + || fail "completed process-event handling could not acknowledge the replay" + [ ! -s "$state/.wake-queue" ] || fail "acknowledged process-event replay remained durable" + + before=$(awk 'END { print NR + 0 }' "$state/.wake-queue" 2>/dev/null || echo 0) + : > "$out" + procevent_watch_bg "$dir" "$out" + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + fail "a handled process-event result woke the watcher: $(cat "$out")" + fi + reap "$pid" + after=$(awk 'END { print NR + 0 }' "$state/.wake-queue" 2>/dev/null || echo 0) + [ "$after" = "$before" ] || fail "a handled result was announced again ($before -> $after queued records)" + pass "an unacknowledged process-event result re-drains until handling is acknowledged" +} + +test_procevent_marker_keys_are_injective() { + local dir state out pid marker_count + dir=$(make_case procevent-marker-identity); state="$dir/state"; out="$dir/watch.out" + append_wake "$state" check "procevent:a.b:1" "check: procevent fixture a.b 1" + append_wake "$state" check "procevent:a_b:1" "check: procevent fixture a_b 1" + procevent_watch_bg "$dir" "$out" + pid=$! + wait_for_exit "$pid" 100 || fail "colliding-looking process-event keys were not surfaced" + grep -F "procevent:a.b:1" "$out" >/dev/null || fail "the dotted queue key was suppressed" + grep -F "procevent:a_b:1" "$out" >/dev/null || fail "the underscored queue key was suppressed" + marker_count=$(find "$state" -maxdepth 1 -name '.seen-procevent-*' -type f | awk 'END { print NR + 0 }') + [ "$marker_count" = 2 ] || fail "distinct queue keys produced $marker_count seen markers" + FM_STATE_OVERRIDE="$state" "$DRAIN" >/dev/null 2>&1 || fail "marker identity fixture drain failed" + pass "complete process-event queue keys map to distinct seen markers" +} + +install_marker_mv_fault() { # <dir> + local dir=$1 + REAL_MV=$(command -v mv) + export REAL_MV + cat > "$dir/fakebin/mv" <<'SH' +#!/usr/bin/env bash +dest=${!#} +case "$dest" in + */.seen-procevent-*) + case "${FM_MARKER_MV_MODE:-}" in + pause) + printf '1\n' > "$FM_MARKER_MV_READY" + while [ ! -e "$FM_MARKER_MV_RELEASE" ]; do sleep 0.02; done + ;; + kill-before) kill -KILL "$PPID"; exit 1 ;; + kill-after) "$REAL_MV" "$@" || exit; kill -KILL "$PPID"; exit 1 ;; + fail) exit 1 ;; + esac + ;; +esac +exec "$REAL_MV" "$@" +SH + chmod +x "$dir/fakebin/mv" +} + +test_procevent_surface_serializes_with_drain() { + local dir state out drain_out ready release pid drain_pid + dir=$(make_case procevent-drain-race); state="$dir/state"; out="$dir/watch.out" + drain_out="$dir/drain.out"; ready="$dir/marker-ready"; release="$dir/marker-release" + append_wake "$state" check "procevent:drain-race:1" "check: procevent fixture drain-race 1" + install_marker_mv_fault "$dir" + FM_MARKER_MV_MODE=pause FM_MARKER_MV_READY="$ready" FM_MARKER_MV_RELEASE="$release" \ + procevent_watch_bg "$dir" "$out" + pid=$! + wait_numeric_file "$ready" 100 || fail "the watcher never reached its marker commit boundary" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" & + drain_pid=$! + wait_live "$drain_pid" 10 || fail "a concurrent drain split the surfacing transition" + [ -s "$state/.wake-queue" ] || fail "the concurrent drain consumed the record before marker commit" + touch "$release" + wait "$pid" || fail "the paused watcher did not finish surfacing" + wait "$drain_pid" || fail "the concurrent drain failed after surfacing committed" + grep -F "procevent:drain-race:1" "$drain_out" >/dev/null \ + || fail "the serialized drain lost the process-event record" + pass "queue revalidation, proactive output, and marker commit serialize with drain" +} + +test_procevent_surface_crash_boundaries() { + local dir state out fifo pid reader marker exit_status replay_err sequence generation + dir=$(make_case procevent-output-fail); state="$dir/state"; out="$dir/watch.out"; fifo="$dir/output.fifo" + append_wake "$state" check "procevent:output-fail:1" "check: procevent fixture output-fail 1" + mkfifo "$fifo" + sh -c ': < "$1"' _ "$fifo" & reader=$! + PATH="$dir/fakebin:$PATH" FM_HOME="$dir" FM_PROCEVENT_CLAIM_ROOT="$dir/claims" \ + FM_CREW_STATE_BIN="$dir/fakebin/fm-crew-state.sh" FM_POLL=0.2 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$fifo" & + pid=$! + wait "$reader" || true + wait_for_exit "$pid" 100 + exit_status=$? + [ "$exit_status" -ne 124 ] || fail "the watcher survived a failed actionable output write" + marker=$(find "$state" -maxdepth 1 -name '.seen-procevent-*' -type f | head -1) + [ -z "$marker" ] || fail "failed output committed a suppression marker" + [ -s "$state/.wake-queue" ] || fail "failed output consumed the durable queue record" + procevent_watch_bg "$dir" "$out"; pid=$! + wait_for_exit "$pid" 100 || fail "the record was not replayable after output failure" + grep -F "procevent:output-fail:1" "$out" >/dev/null || fail "output failure lost proactive replay" + + dir=$(make_case procevent-before-marker); state="$dir/state"; out="$dir/watch.out" + append_wake "$state" check "procevent:before-marker:1" "check: procevent fixture before-marker 1" + install_marker_mv_fault "$dir" + FM_MARKER_MV_MODE=kill-before procevent_watch_bg "$dir" "$out"; pid=$! + wait_for_exit "$pid" 100 + exit_status=$? + [ "$exit_status" -ne 124 ] || fail "the watcher survived the injected pre-marker crash" + grep -F "procevent:before-marker:1" "$out" >/dev/null || fail "the pre-marker crash happened before output" + marker=$(find "$state" -maxdepth 1 -name '.seen-procevent-*' -type f | head -1) + [ -z "$marker" ] || fail "a pre-marker crash committed suppression" + procevent_watch_bg "$dir" "$out.replay"; pid=$! + wait_for_exit "$pid" 100 || fail "a pre-marker crash was not replayable" + + dir=$(make_case procevent-after-marker); state="$dir/state"; out="$dir/watch.out" + append_wake "$state" check "procevent:after-marker:1" "check: procevent fixture after-marker 1" + install_marker_mv_fault "$dir" + FM_MARKER_MV_MODE=kill-after procevent_watch_bg "$dir" "$out"; pid=$! + wait_for_exit "$pid" 100 + exit_status=$? + [ "$exit_status" -ne 124 ] || fail "the watcher survived the injected post-marker crash" + grep -F "procevent:after-marker:1" "$out" >/dev/null || fail "the post-marker crash lost actionable output" + marker=$(find "$state" -maxdepth 1 -name '.seen-procevent-*' -type f | head -1) + [ -n "$marker" ] || fail "the post-marker crash did not reach marker commit" + : > "$out.replay" + procevent_watch_bg "$dir" "$out.replay"; pid=$! + wait_for_exit "$pid" 100 \ + || fail "an unacknowledged delivered record was not re-surfaced on re-arm: $(cat "$out.replay")" + grep -F 'check: rearm-resurface' "$out.replay" >/dev/null \ + || fail "the successor did not recover the delivered-but-unacknowledged record: $(cat "$out.replay")" + replay_err="$out.replay.err" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$out.replay.drain" 2> "$replay_err" \ + || fail "post-marker successor drain failed" + grep "$(printf '\tcheck\t')" "$out.replay.drain" | grep -F 'procevent fixture after-marker 1' >/dev/null \ + || fail "post-marker successor did not re-drain the durable record" + sequence=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through \([0-9][0-9]*\) --recovery-generation [A-Za-z0-9._-][A-Za-z0-9._-]*$/\1/p' "$replay_err") + generation=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through [0-9][0-9]* --recovery-generation \([A-Za-z0-9._-][A-Za-z0-9._-]*\)$/\1/p' "$replay_err") + [ -n "$sequence" ] && [ -n "$generation" ] \ + || fail "post-marker replay omitted its post-handling acknowledgement boundary" + FM_STATE_OVERRIDE="$state" "$DRAIN" --ack-through "$sequence" --recovery-generation "$generation" \ + || fail "post-marker replay acknowledgement failed" + [ ! -s "$state/.wake-queue" ] || fail "post-marker acknowledgement left the durable record queued" + pass "surfacing failures replay until post-handling acknowledgement" +} + +test_procevent_marker_failure_exits_and_replays() { + local dir state out pid marker output_count + dir=$(make_case procevent-marker-failure); state="$dir/state"; out="$dir/watch.out" + append_wake "$state" check "procevent:marker-failure:1" "check: procevent fixture marker-failure 1" + install_marker_mv_fault "$dir" + FM_MARKER_MV_MODE=fail procevent_watch_bg "$dir" "$out" + pid=$! + wait_for_exit "$pid" 100 || fail "marker failure did not end the actionable watcher cycle successfully" + output_count=$(grep -Fc "procevent:marker-failure:1" "$out" || true) + [ "$output_count" = 1 ] || fail "marker failure printed the actionable reason $output_count times" + marker=$(find "$state" -maxdepth 1 -name '.seen-procevent-*' -type f | head -1) + [ -z "$marker" ] || fail "marker failure committed suppression" + [ ! -e "$state/.wake-queue.lock" ] && [ ! -L "$state/.wake-queue.lock" ] \ + || fail "marker failure left the queue lock held" + procevent_watch_bg "$dir" "$out.replay" + pid=$! + wait_for_exit "$pid" 100 || fail "marker failure did not leave the durable record replayable" + grep -F "procevent:marker-failure:1" "$out.replay" >/dev/null \ + || fail "marker failure lost the later proactive replay" + FM_STATE_OVERRIDE="$state" "$DRAIN" >/dev/null 2>&1 || fail "marker-failure fixture drain failed" + pass "marker failure exits through the shared wake owner, releases its lock, and replays later" +} + +# --- heartbeat: no-change absorbed, backstop surfaces a missed status -------- + +test_heartbeat_no_change_absorbed() { + local dir state fakebin out pid i sig + dir=$(make_case heartbeat-absorb); state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out" + printf 'working: routine heartbeat history\n' > "$state/routine.status" + sig=$(seen_sig "$state/routine.status"); printf '%s' "$sig" > "$state/.seen-routine_status" + # A quiet fleet with a fast heartbeat cadence. + PATH="$fakebin:$PATH" FM_STATE_OVERRIDE="$state" FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=1 "$WATCH" > "$out" & + pid=$! + if ! wait_poll_cycle "$state" "$pid"; then + reap "$pid"; fail "watcher exited for a no-change heartbeat (should absorb): $(cat "$out")" + fi + # The heartbeat fires on the first poll whose .last-heartbeat has aged past + # FM_HEARTBEAT, which need not be the first completed cycle, so wait for the + # absorbed heartbeat itself rather than assuming one cycle produced it. + i=0 + while [ "$i" -lt 200 ]; do + [ "$(cat "$state/.heartbeat-streak" 2>/dev/null || echo 0)" -ge 1 ] && break + kill -0 "$pid" 2>/dev/null || break + sleep 0.1 + i=$((i + 1)) + done + [ ! -s "$out" ] || fail "no-change heartbeat printed a wake reason: $(cat "$out")" + [ ! -s "$state/.wake-queue" ] || fail "no-change heartbeat enqueued a durable wake record" + [ "$(cat "$state/.heartbeat-streak" 2>/dev/null || echo 0)" -ge 1 ] || fail "heartbeat backoff streak did not advance while absorbing" + [ "$(status_presentation_marker_offset "$state/.hb-surfaced-routine" "$state/routine.status")" = \ + "$(size_of "$state/routine.status")" ] \ + || fail "routine heartbeat classification did not commit its captured endpoint" + reap "$pid" + pass "a heartbeat with no captain-relevant change is absorbed and backs off the cadence" +} + +test_heartbeat_backstop_surfaces_a_masked_status() { + local dir state fakebin out sig pid + dir=$(make_case heartbeat-masked); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out" + # Same miss as below, but the captain-relevant event is followed by a routine + # append, so its last line reads benign. The backstop must still catch it. + printf 'working: setup\nneeds-decision: pick A or B\nworking: tidying the branch\n' \ + > "$state/miss.status" + sig=$(seen_sig "$state/miss.status"); printf '%s' "$sig" > "$state/.seen-miss_status" + PATH="$fakebin:$PATH" FM_STATE_OVERRIDE="$state" FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=1 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 \ + || fail "heartbeat backstop missed a decision hidden behind a later working: line" + grep -Fx "heartbeat" "$out" >/dev/null || fail "backstop did not exit with a heartbeat wake" + [ "$(status_presentation_marker_offset "$state/.hb-surfaced-miss" "$state/miss.status")" = \ + "$(size_of "$state/miss.status")" ] \ + || fail "backstop did not record the masked status as surfaced through its end" + pass "the heartbeat backstop surfaces a captain event hidden behind a later routine append" +} + +test_heartbeat_backstop_surfaces_unsurfaced_status() { + local dir state fakebin out drain_out sig pid + dir=$(make_case heartbeat-backstop); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out" + # A captain-relevant status whose .seen-* signature ALREADY matches (so the + # per-poll signal scan stays quiet) but which was never surfaced (no + # .hb-surfaced-* marker). This stands in for a per-wake-path miss; the heartbeat + # fleet-scan backstop must catch it and wake firstmate. + printf 'done: PR https://example.test/pr/5\n' > "$state/miss.status" + sig=$(seen_sig "$state/miss.status"); printf '%s' "$sig" > "$state/.seen-miss_status" + PATH="$fakebin:$PATH" FM_STATE_OVERRIDE="$state" FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=1 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "heartbeat backstop did not surface an unsurfaced captain-relevant status" + grep -Fx "heartbeat" "$out" >/dev/null || fail "backstop did not exit with a heartbeat wake" + [ "$(status_presentation_marker_offset "$state/.hb-surfaced-miss" "$state/miss.status")" = \ + "$(size_of "$state/miss.status")" ] \ + || fail "backstop did not record the status as surfaced through its end (would re-fire next heartbeat)" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the backstop heartbeat failed" + grep "$(printf '\theartbeat\t')" "$drain_out" >/dev/null || fail "backstop heartbeat was not queued" + pass "heartbeat backstop fail-safe surfaces a captain-relevant status the per-wake path missed" +} + +# --- beacon stays fresh while absorbing ------------------------------------- + +test_beacon_stays_fresh_while_absorbing() { + local dir state fakebin out status_file pid m1 m2 now + dir=$(make_case beacon-fresh); state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out" + status_file="$state/task.status" + printf 'working: a\n' > "$status_file" + # Provably working so the working: notes are absorbed (the path that must keep the + # beacon fresh). + export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' + watch_bg "$state" "$fakebin" "$out" + pid=$! + # Wait on the beacon itself rather than a fixed liveness budget: the watcher's + # bounded startup can outlast a short wait, and reading an absent beacon would + # report a missing beacon that simply had not been written yet. + wait_poll_cycle "$state" "$pid" || { reap "$pid"; fail "watcher exited while absorbing the first benign signal"; } + m1=$(file_mtime "$state/.last-watcher-beat") + # A second benign signal keeps it absorbing; the beacon must keep advancing. + printf 'working: b\n' >> "$status_file" + wait_poll_cycle "$state" "$pid" || { reap "$pid"; fail "watcher exited while absorbing a second benign signal"; } + m2=$(file_mtime "$state/.last-watcher-beat") + now=$(date +%s) + if [ -z "$m1" ] || [ -z "$m2" ]; then + reap "$pid" + fail "watcher beacon missing while absorbing" + fi + [ "$m2" -ge "$m1" ] || { reap "$pid"; fail "beacon mtime regressed while absorbing"; } + [ "$(( now - m2 ))" -lt 10 ] || { reap "$pid"; fail "beacon went stale while absorbing (age $(( now - m2 ))s)"; } + [ ! -s "$state/.wake-queue" ] || { reap "$pid"; fail "absorbing benign signals enqueued a wake"; } + reap "$pid" + pass "the liveness beacon stays fresh while the watcher absorbs benign wakes (fm-guard never false-alarms)" +} + +# --- afk coherence: the daemon owns triage; the watcher does not double-triage --- + +test_afk_signal_records_heartbeat_endpoint() { + local dir state fakebin out status_file pid + dir=$(make_case afk-heartbeat-endpoint); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; status_file="$state/task.status" + printf 'needs-decision: choose release target\nworking: preparing both targets\n' > "$status_file" + date '+%s' > "$state/.afk" + export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || fail "afk watcher did not hand the actionable signal to the daemon" + [ "$(status_presentation_marker_offset "$state/.hb-surfaced-task" "$status_file")" = \ + "$(size_of "$status_file")" ] \ + || fail "afk signal did not record the endpoint handed to the daemon" + unset FM_FAKE_CREW_STATE + pass "an afk signal records its captured heartbeat endpoint" +} + +test_afk_present_reverts_watcher_to_one_shot() { + local dir state fakebin out drain_out status_file pid + dir=$(make_case afk-coherence); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out" + status_file="$state/task.status" + printf 'working: routine note\n' > "$status_file" + date '+%s' > "$state/.afk" # away mode: the supervise-daemon owns triage + # Set a PROVABLY-WORKING verdict: if afk failed to bypass the provably-working + # check, this no-verb signal would be absorbed (not surfaced). The test asserting + # a surface therefore also proves afk reverts to one-shot and skips the costly read. + export FM_FAKE_CREW_STATE='state: working · source: run-step · validating (running)' + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || fail "with .afk present the watcher did not exit one-shot for a benign signal" + grep -F "signal: $status_file" "$out" >/dev/null || fail "afk-mode watcher did not surface the signal for the daemon" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after the afk-mode signal failed" + grep "$(printf '\tsignal\t')" "$drain_out" | grep -F "$status_file" >/dev/null \ + || fail "afk-mode benign signal was not queued for the daemon to classify" + pass "with .afk present the watcher reverts to one-shot so the daemon owns triage (no double-triage)" +} + +# A paused pane can first appear as a changed hash. In AFK mode that initial path +# must still hand off the plain window identity to the daemon, rather than running +# the normal-mode pause re-surface and decorating the stale identity. +test_afk_paused_changed_pane_hands_off_plain_stale() { + local dir state fakebin out drain_out capture_file statusf window key sig pid back + dir=$(make_case afk-paused-changed-pane); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out"; drain_out="$dir/drain.out"; capture_file="$dir/pane.txt" + window="test:fm-afk-held" + printf 'idle, awaiting upstream\n' > "$capture_file" + printf 'window=%s\nkind=ship\n' "$window" > "$state/afk-held.meta" + statusf="$state/afk-held.status" + printf 'paused: awaiting the upstream tool release\n' > "$statusf" + back=$(( $(date +%s) - 500 )) + if [ "$(uname)" = Darwin ]; then touch -mt "$(date -r "$back" '+%Y%m%d%H%M.%S')" "$statusf" + else touch -m -d "@$back" "$statusf"; fi + sig=$(seen_sig "$statusf"); printf '%s' "$sig" > "$state/.seen-afk-held_status" + date '+%s' > "$state/.afk" + key=$(printf '%s' "$window" | tr '.:/' '___') + + # Deliberately do not seed .hash-*: this is the changed-pane path that used to + # call handle_paused_stale before AFK's one-shot daemon handoff. + PATH="$fakebin:$PATH" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_CAPTURE="$capture_file" \ + FM_FAKE_CREW_STATE='state: paused · source: status-log · awaiting the upstream tool release' \ + FM_STATE_OVERRIDE="$state" FM_CREW_STATE_BIN="$fakebin/fm-crew-state.sh" FM_PAUSE_RESURFACE_SECS=240 FM_POLL=0.2 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out" & + pid=$! + wait_for_exit "$pid" 100 || fail "AFK paused changed pane did not hand off a stale wake" + grep -Fx "stale: $window" "$out" >/dev/null || fail "AFK paused stale did not preserve its plain window identity: $(cat "$out")" + grep -F "awaiting external" "$out" >/dev/null && fail "AFK watcher decorated a stale identity instead of handing it to the daemon" + [ ! -e "$state/.paused-$key" ] || fail "AFK watcher recorded normal-mode pause tracking instead of handing off" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$drain_out" 2>/dev/null || fail "drain after AFK paused stale failed" + grep "$(printf '\tstale\t')" "$drain_out" | grep -F "stale: $window" >/dev/null \ + || fail "AFK paused stale was not queued with the plain window identity" + pass "AFK changed paused panes hand off plain stale identities for daemon-owned pause triage" +}