Skip to content

feat: sync upstream Firstmate and expand supervision - #25

Merged
peterOC26 merged 365 commits into
mainfrom
fm/fm-upstream-sync-s4
Oct 1, 2026
Merged

peterOC26 merged 365 commits into
mainfrom
fm/fm-upstream-sync-s4

Conversation

@peterOC26

@peterOC26 peterOC26 commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

Intent

q134 y

Context: Q134 asked whether to bring in the original project's (kunchenguid/firstmate) latest Firstmate, including its already-merged smaller AGENTS.md (about 48 KB upstream against about 90 KB in this fork), the same way as last time, keeping this fork's own changes; Y = file the update now and start a worker on it; nothing is merged without explicit approval. The last sync was #24 (upstream through 30ef650); upstream main has 51 newer commits. The smaller AGENTS.md points to scripts, docs and skills that changed with it, so the code comes with it.

The same way as that last sync (#24, which itself followed #20): bring the original project's updates in together with the guidance and the scripts, docs, and skills that guidance describes; keep this fork's own changes, which are the fleet board and the rest of the fork-only work that sync kept, including routing pending supervision continuations to MAIN on Pi (#21), the Grok approval-mode composer title (#22), and large contribution snapshots (#23); and open a pull request that is not merged without explicit approval.

What Changed

  • Sync upstream Firstmate through fd325b1b, moving situational guidance from AGENTS.md into focused skills and updating the related scripts and docs.
  • Expand supervision with a default Claude host, distinct away and quiet behavior, and more reliable outcome and wake delivery across Pi and Claude.
  • Update task and remote lifecycle handling, and add a disposable live supervision lab, a JEV memory guard, and supporting test coverage.

Risk Assessment

⚠️ Medium: This is a large upstream merge whose fork overlays survived nearly byte-for-byte (the fleet board, the Grok composer title, large contribution snapshots, and the AGENTS.md swap all check out). The two warnings are both about PR #21: part of its processing-prompt wording was dropped, and its continuation field is not shown by upstream's new off-Pi drain path. Both are worth fixing but are bounded.

Testing

Live CLI and Grok lab checks passed; focused Pi extension and real SDK checks passed without a model turn. The real Pi turn lacked credentials, and the interactive Lavish board lacked permitted isolated session state, so neither received live visual validation. Evidence captures and generated outputs are attached; the worktree is clean.

  • Live validation: ⚠️ inconclusive - 5 of 8 scenarios driven live against the product
Scenario Result Live Evidence
A worker brief gains foreground waiting instructions only when the home enables wait-no-turns ✅ pass live brief-waiting.md and brief-default.md
A contribution snapshot preserves a backlog larger than the old argument-size boundary ✅ pass live large-contribution-input.json: 500 records and a 584,324-byte backlog projection
The captain requests Bearings and sees queued work in Ready ✅ pass live bearings-chat.txt and bearings-board-snapshot.json
A branch outcome carries a completed-stage continuation to MAIN and keeps it pending until acknowledgment ✅ pass live pi-continuation-present.jsonl and pi-continuation-unprocessed.jsonl
Grok's approval-mode title permits an idle composer while an unsent draft remains protected ✅ pass live grok-composer-captures.svg; real Herdr pane classified empty, then pending with KEEP UNSENT DRAFT
A real Pi supervision branch completes a stage and routes its continuation to MAIN ⏸️ untested no Pi has no configured credential for the checked providers, so a real model turn could not run. Provide credentials through an isolated Pi lab configuration and rerun this scenario.
The captain opens the interactive fleet board and sees the queued card ⏸️ untested no Opening a Lavish session would persist application state outside the permitted worktree. Provide authority for a disposable isolated Lavish session and cleanup to capture the rendered board.
The upstream sync is offered as a pull request and remains unmerged without explicit approval ⏸️ untested no The outer executor alone has authority for the PR and merge phases. It must open the PR and preserve the explicit-approval merge boundary.

Real Grok composer: idle and unsent draft

Evidence: Large contribution snapshot

Source: Large contribution snapshot

{
  "backlog": {
    "path": "~/.no-mistakes/worktrees/7f0ec18181b6/01M3TPYMH3H1WWQPTW9R5NESD6/.fm-test-home.OkS3Bs/data/backlog.md",
    "present": true,
    "records": [
      {
        "order": 1,
        "state": "queued",
        "structured": true,
        "id": "queued-1",
        "checked": false,
        "title": "0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000001",
        "repo": "sample",
        "kind": "ship",
        "priority": null,
        "hold_reason": null,
        "hold_kind": null,
        "hold_until": null,
        "hold_set": null,
        "blocked_by": null,
        "blocked_by_ids": [],
        "blocked_reason": null,
        "since": null,
        "merged": null,
        "reported": null,
        "done": null,
        "completion": {
          "verb": null,
          "date": null
        },
        "links": [],
        "pr_url": null,
        "report_path": null,
        "local_note": null,
        "raw": "- [ ] queued-1 - 0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000001 (repo: sample) (kind: ship)",
        "body_lines": [],
        "body_excerpt": null,
        "unresolved_blocker_ids": [],
        "current_role": "queued",
        "requires_child_metadata": false,
        "hold_age_days": null,
        "hold_bucket": null,
        "captain_actionable": false
      },
      {
        "order": 2,
        "state": "queued",
        "structured": true,
        "id": "queued-2",
        "checked": false,
        "title": "0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000002",
        "repo": "sample",
        "kind": "ship",
        "priority": null,
        "hold_reason": null,
        "hold_kind": null,
        "hold_until": null,
        "hold_set": null,
        "blocked_by": null,
        "blocked_by_ids": [],
        "blocked_reason": null,
        "since": null,
        "merged": null,
        "reported": null,
        "done": null,
        "completion": {
          "verb": null,
          "date": null
        },
        "links": [],
        "pr_url": null,
        "report_path": null,
        "local_note": null,
        "raw": "- [ ] queued-2 - 0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000002 (repo: sample) (kind: ship)",
        "body_lines": [],
        "body_excerpt": null,
        "unresolved_blocker_ids": [],
        "current_role": "queued",
        "requires_child_metadata": false,
        "hold_age_days": null,
        "hold_bucket": null,
        "captain_actionable": false
      },
      {
        "order": 3,
        "state": "queued",
        "structured": true,
        "id": "queued-3",
        "checked": false,
        "title": "0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000003",
        "repo": "sample",
        "kind": "ship",
        "priority": null,
        "hold_reason": null,
        "hold_kind": null,
        "hold_until": null,
        "hold_set": null,
        "blocked_by": null,
        "blocked_by_ids": [],
        "blocked_reason": null,
        "since": null,
        "merged": null,
        "reported": null,
        "done": null,
        "completion": {
          "verb": null,
          "date": null
        },
        "links": [],
        "pr_url": null,
        "report_path": null,
        "local_note": null,
        "raw": "- [ ] queued-3 - 0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000003 (repo: sample) (kind: ship)",
        "body_lines": [],
        "body_excerpt": null,
        "unresolved_blocker_ids": [],
        "current_role": "queued",
        "requires_child_metadata": false,
        "hold_age_days": null,
        "hold_bucket": null,
        "captain_actionable": false
      },
      {
        "order": 4,
        "state": "queued",
        "structured": true,
        "id": "queued-4",
        "checked": false,
        "title": "0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000004",
        "repo": "sample",
        "kind": "ship",
        "priority": null,
        "hold_reason": null,
        "hold_kind": null,
        "hold_until": null,
        "hold_set": null,
        "blocked_by": null,
        "blocked_by_ids": [],
        "blocked_reason": null,
        "since": null,
        "merged": null,
        "reported": null,
        "done": null,
        "completion": {
          "verb": null,
          "date": null
        },
        "links": [],
        "pr_url": null,
        "report_path": null,
        "local_note": null,
        "raw": "- [ ] queued-4 - 0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000004 (repo: sample) (kind: ship)",
        "body_lines": [],
        "body_excerpt": null,
        "unresolved_blocker_ids": [],
        "current_role": "queued",
        "requires_child_metadata": false,
        "hold_age_days": null,
        "hold_bucket": null,
        "captain_actionable": false
      },
      {
        "order": 5,
        "state": "queued",
        "structured": true,
        "id": "queued-5",
        "checked": false,
        "title": "0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000005",
        "repo": "sample",
        "kind": "ship",
        "priority": null,
        "hold_reason": null,
        "hold_kind": null,
        "hold_until": null,
        "hold_set": null,
        "blocked_by": null,
        "blocked_by_ids": [],
        "blocked_reason": null,
        "since": null,
        "merged": null,
        "reported": null,
        "done": null,
        "completion": {
          "verb": null,
          "date": null
        },
        "links": [],
        "pr_url": null,
        "report_path": null,
        "local_note": null,
        "raw": "- [ ] queued-5 - 0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000005 (repo: sample) (kind: ship)",
        "body_lines": [],
        "body_excerpt": null,
        "unresolved_blocker_ids": [],
        "current_role": "queued",
        "requires_child_metadata": false,
        "hold_age_days": null,
        "hold_bucket": null,
        "captain_actionable": false
      },
      {
        "order": 6,
        "state": "queued",
        "structured": true,
        "id": "queued-6",
        "checked": false,
        "title": "000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000

... [761661 bytes truncated] ...

ull,
        "unresolved_blocker_ids": [],
        "current_role": "queued",
        "requires_child_metadata": false,
        "hold_age_days": null,
        "hold_bucket": null,
        "captain_actionable": false
      },
      {
        "order": 496,
        "state": "queued",
        "structured": true,
        "id": "queued-496",
        "checked": false,
        "title": "0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000496",
        "repo": "sample",
        "kind": "ship",
        "priority": null,
        "hold_reason": null,
        "hold_kind": null,
        "hold_until": null,
        "hold_set": null,
        "blocked_by": null,
        "blocked_by_ids": [],
        "blocked_reason": null,
        "since": null,
        "merged": null,
        "reported": null,
        "done": null,
        "completion": {
          "verb": null,
          "date": null
        },
        "links": [],
        "pr_url": null,
        "report_path": null,
        "local_note": null,
        "raw": "- [ ] queued-496 - 0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000496 (repo: sample) (kind: ship)",
        "body_lines": [],
        "body_excerpt": null,
        "unresolved_blocker_ids": [],
        "current_role": "queued",
        "requires_child_metadata": false,
        "hold_age_days": null,
        "hold_bucket": null,
        "captain_actionable": false
      },
      {
        "order": 497,
        "state": "queued",
        "structured": true,
        "id": "queued-497",
        "checked": false,
        "title": "0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000497",
        "repo": "sample",
        "kind": "ship",
        "priority": null,
        "hold_reason": null,
        "hold_kind": null,
        "hold_until": null,
        "hold_set": null,
        "blocked_by": null,
        "blocked_by_ids": [],
        "blocked_reason": null,
        "since": null,
        "merged": null,
        "reported": null,
        "done": null,
        "completion": {
          "verb": null,
          "date": null
        },
        "links": [],
        "pr_url": null,
        "report_path": null,
        "local_note": null,
        "raw": "- [ ] queued-497 - 0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000497 (repo: sample) (kind: ship)",
        "body_lines": [],
        "body_excerpt": null,
        "unresolved_blocker_ids": [],
        "current_role": "queued",
        "requires_child_metadata": false,
        "hold_age_days": null,
        "hold_bucket": null,
        "captain_actionable": false
      },
      {
        "order": 498,
        "state": "queued",
        "structured": true,
        "id": "queued-498",
        "checked": false,
        "title": "0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000498",
        "repo": "sample",
        "kind": "ship",
        "priority": null,
        "hold_reason": null,
        "hold_kind": null,
        "hold_until": null,
        "hold_set": null,
        "blocked_by": null,
        "blocked_by_ids": [],
        "blocked_reason": null,
        "since": null,
        "merged": null,
        "reported": null,
        "done": null,
        "completion": {
          "verb": null,
          "date": null
        },
        "links": [],
        "pr_url": null,
        "report_path": null,
        "local_note": null,
        "raw": "- [ ] queued-498 - 0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000498 (repo: sample) (kind: ship)",
        "body_lines": [],
        "body_excerpt": null,
        "unresolved_blocker_ids": [],
        "current_role": "queued",
        "requires_child_metadata": false,
        "hold_age_days": null,
        "hold_bucket": null,
        "captain_actionable": false
      },
      {
        "order": 499,
        "state": "queued",
        "structured": true,
        "id": "queued-499",
        "checked": false,
        "title": "0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000499",
        "repo": "sample",
        "kind": "ship",
        "priority": null,
        "hold_reason": null,
        "hold_kind": null,
        "hold_until": null,
        "hold_set": null,
        "blocked_by": null,
        "blocked_by_ids": [],
        "blocked_reason": null,
        "since": null,
        "merged": null,
        "reported": null,
        "done": null,
        "completion": {
          "verb": null,
          "date": null
        },
        "links": [],
        "pr_url": null,
        "report_path": null,
        "local_note": null,
        "raw": "- [ ] queued-499 - 0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000499 (repo: sample) (kind: ship)",
        "body_lines": [],
        "body_excerpt": null,
        "unresolved_blocker_ids": [],
        "current_role": "queued",
        "requires_child_metadata": false,
        "hold_age_days": null,
        "hold_bucket": null,
        "captain_actionable": false
      },
      {
        "order": 500,
        "state": "queued",
        "structured": true,
        "id": "queued-500",
        "checked": false,
        "title": "0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000500",
        "repo": "sample",
        "kind": "ship",
        "priority": null,
        "hold_reason": null,
        "hold_kind": null,
        "hold_until": null,
        "hold_set": null,
        "blocked_by": null,
        "blocked_by_ids": [],
        "blocked_reason": null,
        "since": null,
        "merged": null,
        "reported": null,
        "done": null,
        "completion": {
          "verb": null,
          "date": null
        },
        "links": [],
        "pr_url": null,
        "report_path": null,
        "local_note": null,
        "raw": "- [ ] queued-500 - 0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000500 (repo: sample) (kind: ship)",
        "body_lines": [],
        "body_excerpt": null,
        "unresolved_blocker_ids": [],
        "current_role": "queued",
        "requires_child_metadata": false,
        "hold_age_days": null,
        "hold_bucket": null,
        "captain_actionable": false
      }
    ]
  },
  "tasks": [
    {
      "id": "contribution",
      "kind": "ship",
      "pr": {
        "url": "",
        "head": ""
      },
      "merge_authority": "attended"
    }
  ]
}
Evidence: Bearings chat with queued work

Source: Bearings chat with queued work

## 🟢 Ready
- Keep the fleet board working - dispatchable queued
## ⏸️ Held
No captain- or time-gated work.
## 🚧 Blocked
No queued work is waiting on another item.
## ⚙️ Under way
No live workers are under way.
## ❓ Waiting on you
Nothing needs your action right now, captain.
## ✅ Done
No recent completions are in the current baseline.
Evidence: Continuation presented to MAIN

Source: Continuation presented to MAIN

{"seq":1,"epoch":1790825810,"task":"plan-review","wake":"","verdict":"captain","summary":"Plan revision complete","silent":false,"statusEndpoint":0,"statusIdent":"-","continuation":"Plan revision 4 committed; start the accepted round-4 review scout","unread":true,"recordedAgo":"0m"}
Evidence: Continuation remains pending after presentation

Source: Continuation remains pending after presentation

{"seq":1,"epoch":1790825810,"task":"plan-review","wake":"","verdict":"captain","summary":"Plan revision complete","silent":false,"statusEndpoint":0,"statusIdent":"-","continuation":"Plan revision 4 committed; start the accepted round-4 review scout","recordedAgo":"0m"}
Evidence: Generated default brief

Source: Generated default brief

You are a crewmate: an autonomous worker agent managed by firstmate. Work on your own; do not wait for a human.

# Task
## Captain's intent
{TASK}

## Firstmate spec
{FIRSTMATE_SPEC}

# Herdr lifecycle declaration - NOT ENABLED
**HARD SAFETY GATE:** this scaffold cannot inspect the task text filled in above.
If the task will start, stop, delete, restart, profile, or otherwise drive Herdr lifecycle behavior, stop and regenerate the brief with `--herdr-lab` before dispatch.
Do not add Herdr lifecycle commands to this unguarded brief by hand.

# Setup
You are in a disposable git worktree of firstmate, at a detached HEAD on a clean default branch.

**Verify isolation before anything else.** Run `pwd -P` and `git rev-parse --show-toplevel`; both must resolve to the disposable task worktree you were launched in, such as a treehouse pool path or an Orca-managed worktree, not the primary checkout firstmate operates from.
The path check is authoritative: `git rev-parse --git-dir` and `git rev-parse --git-common-dir` can help inspect the repo, but they do not prove you are outside the primary checkout.
If the top-level path is the primary checkout or not the worktree you were launched in, STOP - do not branch or commit here - append `blocked [at=<epoch>]: launched in primary checkout, not an isolated worktree` to the status file and stop.

1. First action: create your branch: `git checkout -b fm/ordinary --`
2. Run `no-mistakes doctor`; if it reports the repo is not initialized here, run `no-mistakes init`.

# Rules
1. Never push to the default branch. Never merge a PR.
2. Stay inside this worktree; modify nothing outside it.
3. Use gh-axi for GitHub operations and chrome-devtools-axi for browser operations.
4. Report status by appending one line:
   `echo "{state} [at=<epoch>]: {one short line}" >> '~/.no-mistakes/worktrees/7f0ec18181b6/01M3TPYMH3H1WWQPTW9R5NESD6/.fm-brief-home.e3crar/state/ordinary.status' && { [ ! -e '~/.no-mistakes/worktrees/7f0ec18181b6/01M3TPYMH3H1WWQPTW9R5NESD6/.fm-brief-home.e3crar/config/fleet-ledger' ] || '~/.no-mistakes/worktrees/7f0ec18181b6/01M3TPYMH3H1WWQPTW9R5NESD6/bin/fm-fleet-ledger.sh' appended '~/.no-mistakes/worktrees/7f0ec18181b6/01M3TPYMH3H1WWQPTW9R5NESD6/.fm-brief-home.e3crar/config' '~/.no-mistakes/worktrees/7f0ec18181b6/01M3TPYMH3H1WWQPTW9R5NESD6/.fm-brief-home.e3crar/state/ordinary.status' >/dev/null 2>&1 || true; }`
   States: working, needs-decision, blocked, paused, done, failed.
   Substitute `<epoch>` with the current Unix time in seconds - run `date +%s` and write the number it printed; a stamp that is not plain digits records no time at all.
   Each append wakes firstmate, so report sparingly: only phase changes a supervisor
   would act on (setup done, bug reproduced, fix implemented, validation passed) and the
   needs-decision/blocked/paused/done/failed states. No step-by-step FYI progress lines;
   firstmate reads your pane for that.
   Whenever you mention a PR anywhere - a status line, your terminal, a summary - write its full
   https:// URL exactly as the forge printed it, never a bare number such as "PR 108"; firstmate
   copies that URL from your line rather than assembling one.
   A mid-task `working:` line (including setup complete) is nonterminal: do not end the
   turn after it; continue the same stage until a defined `done:` gate under Definition of done.
   Use `paused: {why}` - distinct from `blocked:` - when deliberately waiting for work or an external condition expected to clear on its own, including your own validation round.
   Before ending your turn with your own background shell or monitor still running, or before waiting on your own pipeline run or a long foreground command, append `paused [at=<epoch>]: {job and completion condition}` to the status file.
   Name what you are waiting for and what will let you resume; do not repeat the declaration on every poll.
   Do not declare active implementation or reasoning as a wait.
   Firstmate may still raise one first-sight alert; the declared wait then uses the existing long recheck cadence instead of repeated possible-wedge alarms.
   When you know when the wait clears, include `until <YYYY-MM-DDTHH:MMZ>` (UTC) for a recheck at that time.
   Follow the resolution rule below when the wait clears, then resume the task.
   Use `blocked:` when you are stuck and need help.

5. If you hit the same obstacle twice, append `blocked [at=<epoch>]: {why}` and stop; firstmate will help.
6. If a decision belongs above the implementation worker (product choices, destructive actions),
   append `needs-decision [at=<epoch>]: {summary of options}` and stop. Firstmate will reply with the decision.
   For a no-mistakes ask-user gate specifically, escalate all ask-user findings as one event plus one snapshot file, using that same shape even when the gate holds only a single ask-user finding: write only the ask-user findings, verbatim and unparaphrased (id, severity, file, line, description, authority), to `~/.no-mistakes/worktrees/7f0ec18181b6/01M3TPYMH3H1WWQPTW9R5NESD6/.fm-brief-home.e3crar/data/ordinary/nm-<run>-findings.txt`, then report the gate with
   `needs-decision [at=<epoch>] [key=nm-<run>-<step>]: ask-user findings=<id1>,<id2>,... file=~/.no-mistakes/worktrees/7f0ec18181b6/01M3TPYMH3H1WWQPTW9R5NESD6/.fm-brief-home.e3crar/data/ordinary/nm-<run>-findings.txt`
   naming every ask-user finding id from that gate. The status line only points at the file; it never restates or summarizes a finding's content.
   A decision or blocker you opened stays open until a `resolved` line carrying its exact key lands; a later `done:` or `working:` line never closes it, even when the answer is what started that work.
   Firstmate's reply normally writes that closing line at answer time; when a blocker or wait clears WITHOUT a firstmate reply, append `resolved [at=<epoch>]: {how it cleared}` yourself (same `[key=<slug>]` if you opened it with one) as you resume.
7. Never administer infrastructure that every lane shares. Two things are shared:
   - The `no-mistakes` daemon - one instance serving every lane/home, so stopping, restarting, or
     updating it kills other lanes' in-flight pipeline runs; only firstmate manages the daemon.
     Before you append `blocked:` about the pipeline, run `no-mistakes daemon status` and
     `no-mistakes axi status`. If the daemon socket refuses connections or is missing, append
     `blocked [at=<epoch>]: {the daemon error}` and stop even when the local run record still says running or
     fixing, because that record can be stale after the daemon exits. A run record failed with a
     daemon error is also a real block.
     Only after ruling out socket refusal, if the run is still running or fixing, reattach and keep
     going. A drive-call error, timeout, slow read, or generic unreachability is NOT a daemon error:
     the daemon accepts `respond` immediately and runs the round in the background, so a killed or
     timed-out call was only waiting for a read while the run kept working.
   - The worktree pool your own worktree came from, and the repository every lane's worktree
     shares. Never create, remove, return, prune, move, or reassign a worktree or pool slot, and
     never write into a sibling slot's directory. Rule 2 does not cover this: removing a worktree
     is administration rather than an edit outside your directory, and it lands on lanes that are
     running right now. The act is the rule and commands are only examples of it - `treehouse`
     get/return/remove/prune, the equivalent operations on any other worktree provider or runtime
     backend, and `git worktree add|remove|move|prune`. A slot that looks unused is not evidence
     that it is free, and returning your own worktree is firstmate's job at cleanup, not yours.
   If you genuinely need a second checkout, another slot, or the daemon touched, append
   `blocked [at=<epoch>]: {what you need}` and stop; firstmate arranges it.

# Firstmate instruction inbox
Firstmate steers you through durable message files in '~/.no-mistakes/worktrees/7f0ec18181b6/01M3TPYMH3H1WWQPTW9R5NESD6/.fm-brief-home.e3crar/state/ordinary.inbox'.
When a terminal message says an instruction is waiting there - and at any natural checkpoint when you are unsure - list '~/.no-mistakes/worktrees/7f0ec18181b6/01M3TPYMH3H1WWQPTW9R5NESD6/.fm-brief-home.e3crar/state/ordinary.inbox'/*.msg, read and act on each message in numeric order, then acknowledge each handled message by moving it: `mv '~/.no-mistakes/worktrees/7f0ec18181b6/01M3TPYMH3H1WWQPTW9R5NESD6/.fm-brief-home.e3crar/state/ordinary.inbox'/NNN.msg '~/.no-mistakes/worktrees/7f0ec18181b6/01M3TPYMH3H1WWQPTW9R5NESD6/.fm-brief-home.e3crar/state/ordinary.inbox'/handled/`.
The move IS the acknowledgement: without it firstmate rings again and eventually treats you as stuck. An empty or absent inbox needs no action.

# Project memory
A project's `AGENTS.md` or `CLAUDE.md` is loaded into every agent session in that project, so edit it only to correct information that is factually wrong - including information your own change made wrong - and never to add knowledge because it is missing.
A correction edits only the wrong text: do not run `~/.no-mistakes/worktrees/7f0ec18181b6/01M3TPYMH3H1WWQPTW9R5NESD6/bin/fm-ensure-agents-md.sh`, create either file, or add sections, headings, or pointers alongside it.

# Definition of done
Delivery contract: mode=no-mistakes
Ship branch: fm/ordinary
The task is complete only when committed on your branch.
When you believe it is complete, append `done [at=<epoch>]: {summary}` to the status file and stop.
Firstmate will then instruct you to run /no-mistakes to validate and ship a PR.
That first `done:` is the handoff that starts the pipeline, which owns the push; it is not a request to push from this copy.

You drive no-mistakes by responding to its gates, not by implementing fixes.
Follow the guidance no-mistakes itself provides for the mechanics: it loads when you invoke /no-mistakes, and `no-mistakes axi run --help` plus the `help` lines in each `axi` response are authoritative and version-matched to the installed binary.
When starting no-mistakes, pass `--intent` as only this brief's `## Captain's intent` subsection body, not its heading, plus any later words the captain actually said.
Preserve the actual words without adding speaker labels or direct address; the subsection heading supplies provenance outside the pipeline input.
For a legacy brief with no such subsection, include only words on lines marked `[captain] `, excluding that metadata prefix; never copy its mixed `# Task` wholesale.
If it has no provenance-marked captain words, stop and ask firstmate instead of starting no-mistakes.
Do not include `## Firstmate spec`, later Firstmate build constraints, or your own decisions and tradeoffs.
The `--intent` string you pass must be self-sufficient: that string plus the codebase must let a reader reconstruct roughly the same specification, without depending on a separate report, a PR, or context that lives only in this conversation.
When the captain's intent refers to a report, decision, or PR ("do items 1, 2, 3, and 7 of the report"), write the substance of the referenced items into `--intent` in the captain's terms, not only the pointer; that substance is the captain's ask by reference, while Firstmate's build instructions and your own decisions still stay out.
This replaces the no-mistakes skill's advice to enrich `--intent` with decisions and tradeoffs; that advice does not apply to Firstmate-dispatched work.
Do not hand-edit, commit, or fix findings yourself while a run is active - the pipeline applies every fix.

One drive call blocks until the next gate or outcome, which routinely outlives what your harness lets a single command run: Claude Code kills a command at ten minutes maximum, while one fix round is capped around thirty minutes and up to three rounds chain.
So background the drive call instead of sitting in one blocking hold your harness will kill, and read its return when it finishes.
Declare that wait using the brief's status-reporting rule before waiting on the backgrounded drive call.
Where a harness's own command limit is not established, assume it bounds commands and use that same backgrounded shape.
Only a drive call's return reports the green PR: `no-mistakes axi status` shows progress but never reports `checks-passed` while the ci step is still monitoring the PR for merge, so never wait on a status poll for the next gate or outcome.
Whenever a drive call returns without a gate or an outcome - its own wait elapsed, or it was killed or timed out - reattach at once by re-running `no-mistakes axi run` without flags, backgrounded the same way; once checks are green it returns `checks-passed` immediately, and if it refuses because no run is active, read the finished outcome from `no-mistakes axi status`.
A killed or timed-out call is never evidence the daemon died: the daemon accepts your response immediately and runs the round in the background, so the call was only ever waiting for a read while the run kept working.
Reattach and keep going rather than reporting the pipeline blocked; rule 7 owns the checks that decide when a pipeline block is real.

Two firstmate-specific rules layer on top of that guidance:
- ask-user findings are never yours to answer: escalate to firstmate using rule 6's ask-user format and stop.
  Firstmate applies `ask-user-authority` and obtains any required captain decision.
  When the decision comes back, feed it to the gate with `no-mistakes axi respond` and let the pipeline apply it - do not route the question to "the user" or implement the fix yourself.
- NEVER pass `--yes` (or `-y`) to `no-mistakes axi run` or `no-mistakes axi respond`. It is banned fleet-wide.
  It auto-resolves every gate including ask-user findings with no escalation, and answering your own ask-user finding is a hard rule violation.

After /no-mistakes reports CI green (the CI-ready return point - do not wait for it to keep monitoring in the background until merge), read the PR back from the forge and confirm it is not a draft (`gh-axi pr view <number>` must print `draft: no`, where <number> is the PR number from your PR URL); if it is a draft, mark it ready with `gh-axi pr ready <number>`.
A draft cannot be merged, so a done report on one leaves the merge unasked.
Then append `done [at=<epoch>]: PR {url} checks green` and stop. You are finished.
That CI-ready `done:` is accepted only when this copy's HEAD - your latest commit - is one the /no-mistakes run pushed, so commit nothing after the run; the check tests that commit, not merely that a branch moved.
If you deliberately keep the PR a draft, append `paused [at=<epoch>]: {why the draft is held}` instead of done.
- Outcome: ⚠️ 1 warning across 1 run (24m27s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 2 warnings
  • ⚠️ .pi/extensions/fm-branch-supervision.ts:192 - The intent says to keep "this fork's own changes ... including routing pending supervision continuations to MAIN on Pi (fix: route pending supervision continuations to MAIN #21;. PR fix: route pending supervision continuations to MAIN #21 had rewritten this line of PROCESSING_INSTRUCTION to say "the originating wake is already handled, so do not re-drain, re-run, or acknowledge that wake. This does not complete a pending workflow continuation." The merge took upstream's new sentence instead: "...; each fleet event is already handled, so do not re-drain, re-run, or acknowledge the wake." That drops the fork's explicit statement that processing does not complete a pending continuation. The fork's test (tests/fm-pi-branch-extension.test.sh, the processing-request body case, which asserted "acknowledge that wake.") also went back to upstream's "acknowledge the wake.". The text actually sent to MAIN now says every fleet event is already handled. Meanwhile docs/supervision-protocols/pi.md:34-35 (fork wording, kept) says "Treat the originating wake as already handled" and that a pending continuation still needs MAIN to act. The prompt and the protocol doc now disagree. The added continuation sentence at line 198 partly makes up for it, but part of PR fix: route pending supervision continuations to MAIN #21's prompt contract is gone. Recommended fix: re-apply the fork's wording inside upstream's new sentence ("...the originating wake is already handled, so do not re-drain, re-run, or acknowledge that wake. This does not complete a pending workflow continuation.") and restore the matching test pattern.
  • ⚠️ bin/fm-wake-drain.sh:639 - Upstream's new off-Pi drain shows unprocessed captain rows from the shared store (fm-branch-outcome.sh present) in BRANCH OUTCOMES. Upstream designed this to work across a switch of primary: "including one a supervision-host drain presented before a switch to Pi", and processed-init never adopts rows as processed. The captain-line projection prints only seq/task/recordedAgo/summary. It never prints the fork's continuation field, and it also collapses rows to the newest one per task. Concrete case: on Pi, the branch reports continuation: &#34;Plan revision 4 committed; spawn round-4 review scout per accepted plan&#34; with the summary "plan rewritten and waiting for the next plan look". The captain switches primary to Claude with the supervision host before MAIN calls fm_branch_processed. The next drain shows only the summary about waiting and asks MAIN to run mark-processed. MAIN acknowledges, and the authorized next spawn is silently lost, which is the exact failure PR fix: route pending supervision continuations to MAIN #21 fixed. Fix: in the captain jq at fm-wake-drain.sh:637-640, append MAIN continuation handoff: &lt;continuation&gt; (with tabs and newlines flattened) for each row that has one, and do not let the per-task collapse hide an older row's continuation. bin/fm-afk-return.sh:170 (store_rows_load) projects the same rows without continuation too, but it only points to the drain, so the drain is the one place to fix.
  • ℹ️ .agents/skills/operational-home-layout/SKILL.md:83 - Sync PR feat: sync upstream Firstmate runtime and guidance #24 corrected the AGENTS.md state-layout line to call branch-outcomes.jsonl the "supervision-branch durable outcome store shared by Pi and the optional host". Upstream's smaller AGENTS.md moved that table into this skill, and the merge took upstream's wording ("Pi supervision-branch durable outcome store"), so the fork's correction is gone. The store really is shared: the host's bin/fm-branch-report.sh and the bin/fm-wake-drain.sh BRANCH OUTCOMES section both use it. Fix: carry the one-line wording into the skill's table.

🔧 Fix applied.
4 issues (2 warnings, 2 infos) still open:

  • ⚠️ bin/fm-wake-drain.sh:639 - Upstream's new off-Pi drain shows unprocessed captain rows from the shared store (fm-branch-outcome.sh present) in BRANCH OUTCOMES. Upstream designed this to work across a switch of primary: "including one a supervision-host drain presented before a switch to Pi", and processed-init never adopts rows as processed. The captain-line projection prints only seq/task/recordedAgo/summary. It never prints the fork's continuation field, and it also collapses rows to the newest one per task. Concrete case: on Pi, the branch reports continuation: &#34;Plan revision 4 committed; spawn round-4 review scout per accepted plan&#34; with the summary "plan rewritten and waiting for the next plan look". The captain switches primary to Claude with the supervision host before MAIN calls fm_branch_processed. The next drain shows only the summary about waiting and asks MAIN to run mark-processed. MAIN acknowledges, and the authorized next spawn is silently lost, which is the exact failure PR fix: route pending supervision continuations to MAIN #21 fixed. Fix: in the captain jq at fm-wake-drain.sh:637-640, append MAIN continuation handoff: &lt;continuation&gt; (with tabs and newlines flattened) for each row that has one, and do not let the per-task collapse hide an older row's continuation. bin/fm-afk-return.sh:170 (store_rows_load) projects the same rows without continuation too, but it only points to the drain, so the drain is the one place to fix.
  • ℹ️ .agents/skills/operational-home-layout/SKILL.md:83 - Sync PR feat: sync upstream Firstmate runtime and guidance #24 corrected the AGENTS.md state-layout line to call branch-outcomes.jsonl the "supervision-branch durable outcome store shared by Pi and the optional host". Upstream's smaller AGENTS.md moved that table into this skill, and the merge took upstream's wording ("Pi supervision-branch durable outcome store"), so the fork's correction is gone. The store really is shared: the host's bin/fm-branch-report.sh and the bin/fm-wake-drain.sh BRANCH OUTCOMES section both use it. Fix: carry the one-line wording into the skill's table.
  • ℹ️ .agents/skills/operational-home-layout/SKILL.md:83 - This is still unfixed from round 1. The round-1 fixer only changed the Pi prompt, because the user's attached instruction said "Change nothing else". The intent says to "keep this fork's own changes". The base AGENTS.md:129 (sync PR feat: sync upstream Firstmate runtime and guidance #24) described branch-outcomes.jsonl as the "supervision-branch durable outcome store shared by Pi and the optional host". Upstream moved this table into the skill, and the skill now says "Pi supervision-branch durable outcome store". No .md file in the tree has the fork's wording any more. The store really is shared: bin/fm-wake-drain.sh print_branch_outcomes_section reads it off Pi when fm_supervision_host_outcomes_drained is true. Remedy: carry the one-line wording into the skill row. It is ask-user only because the round-1 instruction limited fixes to the Pi prompt.
  • ⚠️ bin/fm-wake-drain.sh:639 - This is still unfixed from round 1. The user picked it for fixing, but the instruction attached to it said "Change nothing else", so the fixer left it. Confirmed in current code: fm-branch-outcome.sh present emits the full row, including the fork's continuation field (bin/fm-branch-outcome.sh:632-636). The captain jq at fm-wake-drain.sh:636-640 prints only seq/task/recordedAgo/summary and never mentions continuation. When a Pi-recorded captain row has a continuation and primary switches to a supervision host before fm_branch_processed runs, the off-Pi drain shows only the summary and asks for mark-processed. The authorized next step can then be acknowledged without anyone acting on it. Correction to round 1: the base drain (feb6498) also presented unprocessed captain rows and had no continuation rendering, so this merge did not introduce the gap. It is a pre-existing hole in PR fix: route pending supervision continuations to MAIN #21's Pi-only fix. Remedy, if wanted: append MAIN continuation handoff: &lt;continuation&gt; (whitespace flattened) to each captain line in that jq. Asking the user because their round-1 instruction excluded it.

🔧 Fix applied.
2 warnings still open:

  • ⚠️ bin/fm-wake-drain.sh:639 - Upstream's new off-Pi drain shows unprocessed captain rows from the shared store (fm-branch-outcome.sh present) in BRANCH OUTCOMES. Upstream designed this to work across a switch of primary: "including one a supervision-host drain presented before a switch to Pi", and processed-init never adopts rows as processed. The captain-line projection prints only seq/task/recordedAgo/summary. It never prints the fork's continuation field, and it also collapses rows to the newest one per task. Concrete case: on Pi, the branch reports continuation: &#34;Plan revision 4 committed; spawn round-4 review scout per accepted plan&#34; with the summary "plan rewritten and waiting for the next plan look". The captain switches primary to Claude with the supervision host before MAIN calls fm_branch_processed. The next drain shows only the summary about waiting and asks MAIN to run mark-processed. MAIN acknowledges, and the authorized next spawn is silently lost, which is the exact failure PR fix: route pending supervision continuations to MAIN #21 fixed. Fix: in the captain jq at fm-wake-drain.sh:637-640, append MAIN continuation handoff: &lt;continuation&gt; (with tabs and newlines flattened) for each row that has one, and do not let the per-task collapse hide an older row's continuation. bin/fm-afk-return.sh:170 (store_rows_load) projects the same rows without continuation too, but it only points to the drain, so the drain is the one place to fix.
  • ⚠️ bin/fm-wake-drain.sh:639 - This is still unfixed from round 1. The user picked it for fixing, but the instruction attached to it said "Change nothing else", so the fixer left it. Confirmed in current code: fm-branch-outcome.sh present emits the full row, including the fork's continuation field (bin/fm-branch-outcome.sh:632-636). The captain jq at fm-wake-drain.sh:636-640 prints only seq/task/recordedAgo/summary and never mentions continuation. When a Pi-recorded captain row has a continuation and primary switches to a supervision host before fm_branch_processed runs, the off-Pi drain shows only the summary and asks for mark-processed. The authorized next step can then be acknowledged without anyone acting on it. Correction to round 1: the base drain (feb6498) also presented unprocessed captain rows and had no continuation rendering, so this merge did not introduce the gap. It is a pre-existing hole in PR fix: route pending supervision continuations to MAIN #21's Pi-only fix. Remedy, if wanted: append MAIN continuation handoff: &lt;continuation&gt; (whitespace flattened) to each captain line in that jq. Asking the user because their round-1 instruction excluded it.
⚠️ **Test** - 1 warning
  • ⚠️ live validation verdict: inconclusive (5 of 8 scenarios were driven live against the product); untested: A real Pi supervision branch completes a stage and routes its continuation to MAIN, The captain opens the interactive fleet board and sees the queued card, The upstream sync is offered as a pull request and remains unmerged without explicit approval
  • Live validation: ⚠️ inconclusive - 5 of 8 scenarios driven live against the product
Scenario Result Live Evidence
A worker brief gains foreground waiting instructions only when the home enables wait-no-turns ✅ pass live brief-waiting.md and brief-default.md
A contribution snapshot preserves a backlog larger than the old argument-size boundary ✅ pass live large-contribution-input.json: 500 records and a 584,324-byte backlog projection
The captain requests Bearings and sees queued work in Ready ✅ pass live bearings-chat.txt and bearings-board-snapshot.json
A branch outcome carries a completed-stage continuation to MAIN and keeps it pending until acknowledgment ✅ pass live pi-continuation-present.jsonl and pi-continuation-unprocessed.jsonl
Grok's approval-mode title permits an idle composer while an unsent draft remains protected ✅ pass live grok-composer-captures.svg; real Herdr pane classified empty, then pending with KEEP UNSENT DRAFT
A real Pi supervision branch completes a stage and routes its continuation to MAIN ⏸️ untested no Pi has no configured credential for the checked providers, so a real model turn could not run. Provide credentials through an isolated Pi lab configuration and rerun this scenario.
The captain opens the interactive fleet board and sees the queued card ⏸️ untested no Opening a Lavish session would persist application state outside the permitted worktree. Provide authority for a disposable isolated Lavish session and cleanup to capture the rendered board.
The upstream sync is offered as a pull request and remains unmerged without explicit approval ⏸️ untested no The outer executor alone has authority for the PR and merge phases. It must open the PR and preserve the explicit-approval merge boundary.
  • FM_HOME=&#34;$LAB&#34; bin/fm-brief.sh with wait-no-turns present and absent; compared generated briefs
  • FM_HOME=&#34;$LAB&#34; bin/fm-fleet-snapshot.sh --contribution-input with 500 backlog records
  • FM_HOME=&#34;$LAB&#34; bin/fm-bearings-snapshot.sh --json and --render chat
  • FM_HOME=&#34;$LAB&#34; bin/fm-branch-outcome.sh append, present, mark-read, unprocessed, and mark-processed
  • Named Herdr lab through bin/fm-herdr-lab.sh: launched real Grok, captured idle and draft panes, classified both, and verified teardown
  • bash tests/fm-pi-branch-extension.test.sh
  • FM_PI_BRANCH_LIVE_E2E=1 bash tests/fm-pi-branch-live-e2e.test.sh
  • pi auth check --provider openai-codex --json --no-refresh and the equivalent Google check
  • git status --short and named Herdr session listing after cleanup
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Shazellb and others added 30 commits September 11, 2026 19:23
…forked code-root copy (kunchenguid#4223)

* fix(backlog): address the home's backlog from any directory and detect a forked code-root copy

A home outside the code root forks its queue: the tracked .tasks.toml names
data/backlog.md relative to tasks-axi's working directory, so a bare
tasks-axi call from the code root writes the code root's data/ while session
start, spawn, and teardown use $FM_HOME/data. Linking the code-root copy into
the home does not hold, because tasks-axi 0.2.4 writes by renaming a temp file
over its target and rename(2) replaces a symlink: add, start, hold, and done
from the code root each turn the link back into a regular file. The archive
path is resolved against the working directory too, even with --file.

bin/fm-tasks-axi.sh runs tasks-axi against this home's backlog from any
directory, using the lifecycle transitions' existing addressing (run from the
data directory's parent, pin <data>/backlog.md through TASKS_AXI_FILE). It
keeps relative --to/--*-file arguments meaning the caller's paths, and refuses
a caller --file, an unresolvable home, and a symlinked home backlog. The
fm-send hold lookup, fm-public-followup, and the fm-decision-hold shim, which
relied on cwd discovery, now go through it with an explicit FM_HOME and a
cleared data override, so they keep addressing exactly $FM_HOME/data and an
ambient TASKS_AXI_FILE cannot divert them; every agent-facing backlog command
names it instead of bare tasks-axi.

Bootstrap gains a detect-only BACKLOG_RECONCILE check, also run read-only:
when the home's data directory is not the code root's, a code-root
data/backlog.md or data/done-archive.md that is not the home's own file is
reported as a fork, with the merge procedure in bootstrap-diagnostics.

* test(teardown): assert the completion hint names bin/fm-tasks-axi.sh ready

The completion hint now points at the home-addressed command instead of a
bare tasks-axi call, so the dependency-cleared follow-up assertion checks for
that command.

* no-mistakes(test): clear ambient tasks-axi env in tests/lib.sh

* no-mistakes(document): drop bare tasks-axi example from cd-guard doc

* no-mistakes(lint): replace ls -A decoy listing with find for SC2012

* no-mistakes: apply CI fixes

* revert: keep the compliance gate unchanged; the synchronize race is filed separately
* fix(spawn): pre-register Claude workspace trust for secondmate homes

A claude --secondmate launch skipped workspace-trust registration
entirely, so a standalone-clone secondmate home (an explicit
~/fm-homes/<id> path) had no store entry and its pane wedged on the
"Is this a project you trust?" dialog before it read its charter.
The step was gated on the task kind rather than on the harness, so the
spawn's fail-closed guard had nothing to run against and reported a
launch that could never start work.

fm-claude-trust.sh gains a secondmate-home mode. A secondmate home is a
whole firstmate instance, produced either as a leased worktree or as a
standalone clone, so the linked-worktree test cannot decide it and the
seed is the evidence instead: the .fm-secondmate-home marker must be a
regular file this user owns naming exactly the id being spawned, the
home must hold AGENTS.md and bin/, and each operational directory must
resolve inside the home. That is the set fm-home-seed.sh writes and
fm-spawn.sh's own home validation re-checks, so nothing wider than a
home a secondmate spawn would launch into can earn home-level trust.
The worktree path is unchanged, and still refuses a home.

fm-spawn.sh now runs the registration for every claude launch and keeps
refusing the spawn when it fails, rather than launching an agent that
would wedge.

* no-mistakes(document): Correct Claude secondmate trust guidance
* fix(pr-merge): judge each required check by its current run

When the base branch advances, GitHub cancels a pull request's in-flight
run and re-triggers it. The cancelled run stays in statusCheckRollup
beside the passing re-run, so the rollup can hold several runs of one
check name at the same head while GitHub itself reports the pull request
CLEAN. github_checks_not_green judged every run independently, so that
superseded failure refused a genuinely mergeable pull request and pushed
the operator toward a needless --allow-red.

Group the rollup by the reported name and judge each check by its
current run. Supersession is proven, never assumed: a name leaves the red
set only when every one of its non-green runs is strictly older than one
of its green runs, dated by the forge's own settled timestamp - a check
run's completedAt once its status is COMPLETED, or a status context's
createdAt - and only in the whole-second UTC form GitHub emits, which is
the one spelling that orders correctly as plain text. A run with no such
timestamp is never superseded, so a still-running, queued or undated run
keeps its check red, and a name with no green run at all stays red. An
unnamed entry is grouped alone so two unrelated unnamed checks are never
treated as one.

Every comparison is one-directional: it can only clear a failure a later
success provably replaced, and never clears a check whose current run
failed, is pending, or is missing. No other guard moves - the pull
request must still be open, undrafted, mergeable, conflict-free and
head-bound, and --allow-red still waives exactly its named check with
every other check green.

Live reproduction: PR kunchenguid#4224 read CLEAN with an old FAILURE and a newer
SUCCESS for one check name and was refused; it now verifies, while
kunchenguid#4208 and kunchenguid#4210, whose latest runs failed, still refuse.

* no-mistakes(review): Use check-run start times for safe supersession

* no-mistakes(document): Clarify GitHub check-rollup documentation
…guid#4266)

* fix(merge): persist the merge authority on poll-detected merge outcomes

The merge ledger tags a merge with the authority that permitted it while the
away-posture record existed, but only the direct attended merge in
bin/fm-pr-merge.sh recorded it. A merge the forge queued, or one the merge
poll detected after the fact, published an untagged row, so exactly the
merges no agent watched were the least auditable.

bin/fm-merge-authority-lib.sh now owns that answer, read from the same
structured sources the merge gate already used: the task's recorded yolo
posture and the away-posture record's mechanical grant list, never prose.
bin/fm-pr-merge.sh keeps its own refusal wording and gates on that answer;
bin/fm-watch.sh only records it on the row its poll publishes, so reading the
authority never becomes a second path to a merge. An unresolved answer records
an untagged row rather than dropping the outcome or inventing an authority.

* no-mistakes(review): Persist canonical merge authority for queued poll outcomes

* no-mistakes(review): Harden merge authority persistence against lifecycle races

* no-mistakes(review): Serialize poll authority publication with teardown

* no-mistakes(document): Clarify persisted merge authority lifecycle

* no-mistakes(ci): Added targeted SC2034 suppressions for the two public result assignments in bin/fm-merge-authority-lib.sh. Verified successfully with `CI=true bin/fm-lint.sh`
…4281)

The 2026-09-12 Actions starvation incident found firstmate CI with no
concurrency deduplication, so every superseded PR head kept its full
13-job fan-out, and four jobs with no timeout at all.

Add per-PR supersession keyed on the PR number for pull_request events
and on the unique run id for push events, cancelling only pull_request
runs, so a new PR head replaces its own in-flight CI while every main
push keeps its own group and is never cancelled. Add hang tripwires to
the four previously unbounded jobs: 25 minutes for lint (measured at
14-16 minutes) and 5 minutes each for the coverage guard, the timing
aggregate, and the repo invariants. Measured lane bounds are unchanged.

tests/fm-ci-workflow.test.sh resolves the workflow's concurrency
expressions against simulated pull_request and push contexts and holds
every job's finite timeout.
…henguid#4288)

Every other make_hold_home caller in this file skips when tasks-axi is
absent; this test was the one unguarded call, so hosts without tasks-axi
hard-fail the fixture build instead of skipping.
…nnot blind a session start (kunchenguid#4027)

* fix(bin): bound each backlog row read so one wedged backend cannot blind a session start

bin/fm-bootstrap.sh's reconcile and close-replay sweeps read the backlog
backend once per item through fm_backlog_row_show, and that read was
unbounded. A single wedged `tasks-axi show` therefore consumed the whole
FM_SESSION_START_TIMEOUT and truncated the digest before the wake queue,
supervision instructions, fleet state, and context sections ever printed,
leaving the fleet unsupervised with no live watcher. The harm was a blind
startup, not a slow one.

Bound the read with the existing shared timeout primitive
(bin/fm-timeout-lib.sh), so a wedged backend degrades to a loud partial
reconcile: the sweep's existing BACKLOG_RECONCILE diagnostic names the item
it could not read and the loop continues to the next one. The first bound hit
also latches FM_BACKLOG_ROW_SHOW_WEDGED, so a sweep over many items pays one
bound rather than one per item and still names every item it skipped, which is
what keeps the digest whole on a home carrying a large fleet.

The bound holds regardless of any particular tasks-axi install, so it does not
depend on the 0.2.5 `show` hang being resolved separately.

* fix(bin): set the wedged-backend latch where it survives, and prove it

The latch added with the read bound was inert. fm_backlog_row_show runs inside
a command substitution in both of its status-capturing callers, so the subshell
read the inherited value correctly but its write died with the subshell. Every
item still paid a full bound and reported `exceeded`, never `skipped`, which
left the large-fleet case the latch existed to cover completely uncovered.

Move the write to the two callers that capture the read's status and own the
surviving shell, and leave fm_backlog_row_show reading the latch only. Correct
the comments that claimed an ownership the function never had.

The test that was supposed to cover this asserted only that the second read
finished under a generous ceiling, which is true whether or not the latch
works. Assert instead that a latched read is strictly faster than one bound and
that it reports its own item as skipped, so an inert latch fails the test.

* test: cover every item the wedged-backend latch skips

The latch assertion exercised a single skipped item, so "every skipped item is
still named" was inferred rather than tested. Probe three items instead and
assert each skipped one names itself and costs less than a bound.

Verified as a real guard by removing both latch writes: the suite then fails on
the first skipped item instead of passing.

* no-mistakes(review): distinguish backlog read-bound hits from absent rows

* no-mistakes(review): preserve read-bound status through the captain verify gates

* no-mistakes(review): Preserve backlog read-bound hits through resolve_entry and reconcile instead of spending them as absent rows

* no-mistakes(review): Preserve backlog read-bound 124 through migrated-prefix scan and remaining task_show call sites

* no-mistakes(document): Document bounded backlog row reads and FM_BACKLOG_ROW_TIMEOUT_SECS

* no-mistakes(ci): Fixed all four failing CI checks with one root-cause fix plus one test-heredity fix. (1) bin/fm-captain-hold.sh: task_show carries the row in TASK_SHOW_OUTPUT and emits no stdout, but four call sites still used the stale command-substitution convention show=$(task_show ...), leaving show empty: task_show_or_fail (every captain hold failed with 'did not retain its hold-set stamp' - broke fm-captain-hold-lifecycle in parallel 1 and fm-bearings-board in serial 3), resolve_migrated_entry (migrated-prefix resolution could never match), reconcile-requests (existing rows were refused as absent), and command_open --identity (printed a constant '#0' identity, so fm-watch-triage's re-held captain call inherited the previous call's silence in serial 1). This is also the Greptile P1. Fixed by invoking task_show in the current shell and reading show=$TASK_SHOW_OUTPUT, the convention the other eight call sites already use; read-bound hits still stop loudly by name. (2) tests/fm-backlog-read-bound.test.sh (serial 4, unclassified family): the new e2e half implicitly relied on the author's process tree containing a harness process so fm-lock.sh would grant the fleet lock; on CI runners the lock is refused, the reconcile sweep is skipped, and the final BACKLOG_RECONCILE assertion fails. Reproduced by simulating a CI ancestry via a ps shim, fixed by pinning the lock evidence with the established fake-ps harness fixture pattern from tests/fm-session-start.test.sh. Verified: shellcheck clean; parallel-1, serial-3, and serial-4 lanes fully green locally (failed=0); serial-1 lane green except fm-gemini-harness, which fails only under local Node v26 (comm=node-MainThread); CI's default Node 22 reports comm=node, the branch that test passes on, so it is not a CI failure

* no-mistakes(document): Verified bounded backlog read docs accurate across branch
…guid#4285)

* fix(merge): serialize the away-authority check with a synchronous merge

bin/fm-pr-merge.sh read the away-posture record for merge authority (the
per-task merge grant and the yolo/away-grant decision) and handed the merge to
the forge afterwards. An archive at the captain's return or a grant revoked by
a replacement record could land in between, so a merge could proceed on away
authority that no longer held.

The away record now carries a cross-subsystem lock, built on the existing
bounded lock primitive rather than a new lock format: the record-mutating
subcommands hold it across their mutation, and the merge holds it across both
its authority read and the forge command. Because a queued or auto merge
returns before the pull request lands, and would therefore outlive the lock,
an away merge is now refused whenever it could land asynchronously: a
requested --auto, a base branch whose merge-queue state does not prove an
immediate merge, and GitLab's asynchronous flags and configuration. What
remains permitted while away is the synchronous merge that lands inside the
lock.

This closes the common away-record/merge race against a live lock owner. It
does not make the merge atomic in every case, and two narrow races are
accepted and documented at their sites rather than hidden, both
confused-agent-grade in the sense bin/fm-lease-lib.sh already uses:

- A merge-queue rule change or a PR base change in the window between the
  queue-free preflight and the forge call can still enqueue the merge, which
  can then land after its grant lapses.
- Killing the lock-owning shell while its gh or glab child is still running
  lets stale-owner recovery reclaim the lock and the record be archived or
  replaced, after which the orphaned child can complete the merge on lapsed
  authority.

Closing either one needs landing verification or an ownership handoff, which
is deliberately out of scope here.

No existing gate is relaxed. The lock is taken after the live green-at-head
verify and the captain-hold check, the in-lock authority read is unchanged,
and a lock that cannot be taken refuses the merge rather than proceeding
unlocked. The away grant stays a structured field; no prose is parsed.

* no-mistakes(review): Fix GitHub rollup fixture base branch

* no-mistakes(document): Document atomic away-authority merge locking

* no-mistakes(ci): Updated two executable GitHub API fixtures to include the required baseRefName. Both previously failing test suites now pass: fm-captain-hold-lifecycle.test.sh and fm-pr-check-security.test.sh. git diff --check also passes
…unchenguid#4200)

* feat(agy): verify Antigravity CLI as third worker/scout adapter

Detection by anchored ancestry in fm-harness.sh (no marker of its own);
bootstrap harness and effort validation; launch template with model and
effort mapping plus reachable-catalog model validation; rendered-tail
busy fallback in fm-busy-lib.sh with delivery footer in fm-composer-lib.sh;
control mechanics with crewmate/scout-only refusal; tmux liveness naming;
router entry with concise adapter reference; dated verification record;
portable regression plus opt-in live drift guard.

Verified live on agy 1.2.0: supervised spawn, durable steering,
same-copy relaunch, and exit, with Herdr-native busy agreement.

* no-mistakes(review): bound agy model probe, gate trust dialog, narrow busy signature

* no-mistakes(review): pre-register agy workspace trust, make readiness gate strict

* no-mistakes(review): Close Orca terminal on gate failure; isolate live-guard HOME; tighten agy matching

* no-mistakes(document): Document agy adapter in stale harness enumerations

* no-mistakes(review): Clamp non-positive FM_AGY_MODELS_TIMEOUT to the default bound

* no-mistakes(document): Fix stale test-shard snapshots after agy lane additions

* no-mistakes(ci): Fixed ci-3 (tests/fm-agy-harness.test.sh:519). Root cause: the agy spawn fixture's default base PATH (/usr/bin:/bin:/usr/sbin:/sbin) omits node's directory, but the spawn drives the real bin/fm-agy-trust.sh (which hard-requires node to record trust) and the fixture's fake tmux trust lookup (node -e) under that PATH. On the ubuntu-latest CI runner node lives in the toolcache (/usr/local/bin), so trust pre-registration failed on portable serial 2; on typical Arch hosts node is in /usr/bin, masking the defect. Fix (smallest, following the existing tests/fm-kimi-harness.test.sh precedent of carrying the interpreter's resolved directory): resolve node from the invoking environment (failing the test with 'test needs node' if absent, as kimi does for python3) and prepend its directory to the fixture's default base PATH; the FM_TEST_BASE_PATH override contract is untouched. Verified locally: (1) pre-fix reproduction with a CI-shaped base PATH (system bins minus node) produced exactly the reported failure — 'node is required to record workspace trust and was not found on PATH' plus the fake tmux 'node: command not found'; (2) post-fix, all 29 tests in the file pass both with node available only via a leading non-standard dir in the base PATH (CI's shape) and with the default base PATH on this host. bash -n clean; ShellCheck is not installed in this worktree (previously recorded as environmental)

* no-mistakes(test): Give agy typed sends a longer submit-confirm budget

* no-mistakes(document): Document agy send budget, trust gate, and control coverage

* no-mistakes(document): Document agy busy fallback inventory and send-timing evidence
…uid#4337)

* feat(afk): add quiet supervision mode for a present captain

Adds a first-class quiet supervision mode alongside /afk for
kunchenguid#2356: the same away-mode daemon, injection,
busy/composer guards, classification policy, and reliability
properties, but the captain staying present and chatting no longer
exits it - only an explicit /quiet off does.

state/.afk's first line now declares its mode (away, the default, or
quiet); fm_afk_mode() in bin/fm-wake-lib.sh is the single reader,
falling back to away for missing/empty/unreadable/unrecognized
content (including the legacy bare-epoch-timestamp format written
before mode existed) so nothing regresses. fm_afk_flag_write()
preserves the on-disk mode on a bare refresh (no explicit mode given)
rather than defaulting to away, which is what keeps the daemon's own
redundant terminal-side re-write from silently resetting a captain's
quiet mode back to away underneath them.

New .agents/skills/quiet/SKILL.md is a thin wrapper cross-referencing
/afk for every shared mechanism, per the one-owner rule. AGENTS.md
gains the state/.afk table entry and section 8's exit-trigger line.
bin/fm-supervision-instructions.sh, bin/fm-session-start.sh, and
bin/fm-guard.sh's stale-watcher banner all become mode-aware so a
quiet-mode captain is never misdirected to /afk in captain-facing
text.

Closes kunchenguid#2356

* no-mistakes(review): Fix AFK epoch parsing and quiet-mode digest wording for two-line flag

* no-mistakes(document): Fix turnend-guard.md daemon-ownership contract for quiet mode

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>
Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…kunchenguid#3578)

* fix(bin): let verified harness ancestry outrank retained markers (#3)

* fix(bin): let a structural harness ancestor outrank a retained marker

bin/fm-harness.sh treated a verified environment marker as unconditionally
authoritative, so a Codex session started from an environment that had retained
CLAUDECODE=1 detected as claude. Session start then emitted Claude's Stop-owned
supervision protocol to a Codex primary, and every turn end was blocked for
missing Claude recovery.

The defect is the precedence boundary, not any one harness. codex, opencode,
kimi, and muse publish no identity marker at all, so with markers winning
outright any retained CLAUDECODE renamed them; the Cursor-before-Claude ordering
was a point patch on the same class of problem, and the launch-time marker
clearing only ever covered sessions fm-spawn started.

Markers and ancestry are now separate evidence layers that detect_own arbitrates:

- no ancestry match, or no marker: the single available layer answers, unchanged;
- same harness family: the marker's finer verdict stands, so a launch-selected
  pi-signed is not flattened to pi by an ancestry walk that can only see the
  shared launcher name;
- different harness with a structural (command-name) ancestor: ancestry wins,
  because only ancestry proves who owns the process tree;
- different harness with only a bare-interpreter script-path match: the marker
  wins, since a harness-shaped path in some node process's arguments is weaker
  evidence than a harness publishing its own identity.

The correction is symmetric: a retained CURSOR_AGENT no longer renames a claude
worker nested under cursor either.

Adds fm-harness.sh ancestry [<pid>], ancestry evidence with no marker layer, so
a real harness process can be asked what the walk makes of it.

tests/fm-harness-precedence.test.sh is the portable regression, built from real
renamed processes with no harness installed. Every case drives the two layers
apart and asserts each alone as well as the combination, so no case can pass
vacuously; it also pins Codex's real two-process install topology, since the fix
depends on the native binary being what a tool subprocess meets first. The
opt-in drift guard gains the matching live half: each installed harness's real
running process must still be identified by the ancestry walk, and it fails
naming the harness and version when a release changes that name.

Documentation follows the corrected contract in the script header, the
harness-adapters detection section, the codex, opencode, kimi, and cursor
references, and a dated verification record.

* fix(tests): drop the unused argument pass-through in the shim-topology helper

bin/fm-lint.sh refused the branch: run_shim declared a `[ancestry]` argument and
forwarded "$@", but every call site that varies the environment or passes the
ancestry subcommand invokes the shim entry point directly, so the helper is only
ever called with no arguments (ShellCheck SC2120/SC2119).

Behavior is unchanged: with no arguments "$@" expanded to nothing.

* fix(bin): examine the top of the process chain instead of assuming init

harness_ancestry stopped as soon as the next pid was 1, on the assumption that
pid 1 is always init and can never be a harness.
Inside a PID namespace that assumption inverts: the harness itself is pid 1, so
the walk never examined the one process that proves who owns the tree, reported
no ancestry at all, and handed the verdict straight back to a retained marker.

A real Codex session under `codex sandbox`, holding CLAUDECODE=1 and
CLAUDE_CODE_ENTRYPOINT=cli, is exactly that shape: it resolved claude and
rendered Claude's Stop-owned supervision protocol even with the marker-vs-ancestry
precedence boundary in place.
The same probe now resolves codex and renders the Codex foreground checkpoint.

A host's real pid 1 (init, systemd, launchd) matches no harness name, so
examining it costs one ps call and can introduce no false positive; the walk
still stops once that top process has been read, and a non-numeric or zero ppid
still ends it.

tests/fm-harness-precedence.test.sh pins the namespace shape with a fake ps that
reports every process as bash with ppid 1 and pid 1 as the harness.
The case asserts the marker still answers alone when pid 1 is host-shaped, so it
cannot pass vacuously, and it fails against the previous stop condition.

* docs(verification): record the real-Codex retained-marker evidence

The existing record proved the precedence boundary with the portable regression
and recorded each installed harness's process name behind the ancestry walk, but
it had no evidence from a real Codex process actually holding a retained Claude
marker, which is the failure the boundary exists for.

Adds the dated before/after result from codex-cli 0.152.0 under `codex sandbox`,
with the exact command and the decisive verdict and rendered protocol on each
side, and records the second boundary that shape exposed: the walk must examine
the top of the process chain, because inside a PID namespace the harness is pid 1.
Refreshes the portable regression's observed output for the case it gained.

* no-mistakes(review): blind ancestry in marker-pinned harness tests

* no-mistakes(review): blind ancestry in the Pi guard-routing test

* no-mistakes(review): classify precedence suite, dedupe ps stub, soften claims

* no-mistakes(review): model the spawn-and-wait Codex shim topology

* no-mistakes(document): correct stale muse marker-clearing detection claims

* no-mistakes: apply CI fixes

* fix(bin): examine the top of the chain in the lock and nudge walks too

The pid-1 defect corrected in bin/fm-harness.sh survived unchanged in the two
other harness-ancestry walks, on the exact topology the branch verified against
a real Codex process.

bin/fm-session-lock-lib.sh's fm_harness_ancestry_pids stopped as soon as the next
pid was 1, so a firstmate whose harness is pid 1 of its own PID namespace could
not find that harness at all and did not recognize its own session lock.
bin/fm-sessionstart-nudge.sh carried the same stop plus a blanket rejection of a
lock pid of 1, so the same session was told to run session start again on every
turn.

Both walks now compare the top process before stopping, matching the shape used
in bin/fm-harness.sh.
For the lock walk this is safe because fm_harness_process_matches rejects a
host's real pid 1.
For the nudge, `kill -0` still gates the lock pid, and on a host an unprivileged
`kill -0 1` fails, so a lock file that wrongly names pid 1 leaves the hook silent
rather than acting on init.

Each walk gains one regression case. The lock case drives a deterministic process
table whose pid 1 is the harness and asserts a host-shaped pid 1 still finds
nothing, so it cannot pass vacuously. The nudge case needs a real PID namespace,
because the builtin `kill -0` gate cannot be reached through a fake ps, and it
first proves the same fixture nudges with no lock present; it skips explicitly
where unprivileged namespaces are unavailable.

* no-mistakes(review): assert comm-strength detection from subprocess vantage in drift guard

* fix(bin): verify the live harness guard at the strength the guarantee needs

The marker-versus-ancestry boundary this branch ships is a strength claim:
detect_own hands an args-strength verdict straight back to a retained foreign
marker, so a harness is only protected where the ancestry walk reaches it at
comm strength.

The installed-harness drift guard probed the pane process alone. Under an
interpreter shim the pane process IS the shim, whose own script path is args
strength, while the native binary that carries comm strength is its child. The
guard therefore observed args for Codex, passed, and would have kept passing if
a release stopped spawning that native child at all, while real sessions
silently regressed to the original bug.

fm-harness.sh gains `ancestry-subtree`, which asks the walk from the pane
process and every descendant of it, the vantage a tool subprocess actually
occupies. The guard now requires comm strength somewhere in that set and
requires every vantage to name the same harness.

This supersedes the preceding commit's in-guard leaf walk, which reached the
same vantage but left the logic inside the test file, where CI could not pin it
and nothing else could reuse it. A harness-dependent check needs both halves:
`tests/fm-harness-precedence.test.sh` now carries a portable case proving the
subtree probe reaches a strength the top-of-session probe cannot, mutation
checked twice, once against the pre-change script and once by disabling
descendant enumeration. The subtree walk also avoids depending on tty and
process-group semantics that differ between Linux and macOS.

Verified live: codex-cli 0.152.0 reports [args codex;comm codex] and Claude Code
2.1.257 reports [comm claude].

* no-mistakes(review): narrow drift guard to the upward vantage path

* no-mistakes(review): judge only comm-strength vantages in drift guard

* no-mistakes(document): drop duplicated rationale in detection precedence evidence

* no-mistakes(review): fix pid-1 nudge case vacuity and descent no-arg expansion

* no-mistakes(document): drop branch-relative phrasing in detection precedence evidence

* no-mistakes(review): guard remaining empty positional expansions in fm-harness

* no-mistakes(document): scope cursor marker-ordering claim to the marker layer

* no-mistakes(review): Prefer comm-strength leaves in equal-depth descent ties

* no-mistakes(document): Document comm-strength descent tie-break

---------

* no-mistakes(review): Blind ancestry in stale gemini/rovo marker-precedence tests

* no-mistakes(document): Add missing equal-depth-tie test line to precedence evidence transcript

* no-mistakes(review): Fix stale/vacuous agy precedence test, add agy to precedence suite and docs

* no-mistakes(document): Fix stale kimi.md marker doc missed by ancestry-precedence fix

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…rkers (kunchenguid#3944)

Claude Code's external-imports check (hasClaudeMdExternalIncludesApproved)
reads only the canonical git-root project entry in ~/.claude.json, which its
own worktree-to-primary-checkout canonicalization means is never the task
worktree fm-claude-trust.sh registered. The trust dialog kept working
previously only because its check has an ancestor-walk fallback that happens
to reach the worktree entry; the external-imports check has no such
fallback.

Verified by disassembling the installed claude binary and reproducing in an
isolated three-way tmux launch: identical flags registered only at the
worktree key still showed the external-imports dialog, and registering them
at the primary checkout key suppressed both dialogs.

fm-claude-trust.sh now registers all three flags on both the worktree entry
and the primary-checkout entry in one atomic write, and refuses when the
<project> argument is not itself a primary checkout (its own write target
would then be wrong). Extends the harness-adapters Claude reference and the
trust test suite.

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…nguid#4355)

The marker lifecycle (fm-wake-lib.sh _fm_recovery_marker_ack) leaves
state/.watcher-down behind in an acked:* state after a downtime episode
is handled. health_snapshot's presence check reported that as an open
gap on every later return, so a handled episode kept surfacing as a
false GAP forever.
…kunchenguid#4361)

* fix(update): rebind fm-procevent-when watches after a self-update

A self-update fast-forwards bin/ in place, changing an armed watch's
action executable bytes with no tampering involved. The watch's trust
binding was hashed at arm time, so the very next fire was refused as
not matching the registered binding and the watch died silently.

Add fm-procevent-when.sh rebind-all: it re-hashes and republishes the
trust binding for every watch whose action executable lives under
FM_ROOT, using the same spec/trust validation as an ordinary fire, and
leaves any watch whose action lives outside FM_ROOT untouched. Wire it
into fm-update.sh right after a successful fast-forward, for both the
primary home and any local secondmate home that advances.

* no-mistakes(review): Canonicalize FM_ROOT for rebind-all's containment check

* no-mistakes(document): Document fm-update.sh's automatic watch rebind and its verification evidence

* no-mistakes(lint): fix(tests): double-quote printf scripts to satisfy shellcheck SC2016

* no-mistakes(review): Reload trust binding from disk before firing to reach live pollers

* no-mistakes(review): Lock the fire-time trust reload against rebind_one's publish race

* no-mistakes(document): Document rebind-all's self-update guarantee and its two review-round test rows

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…kunchenguid#4424)

* fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (kunchenguid#42)

* fix(pr-merge): read a plan-gated 403 on branch rules as no merge queue

github_read_queue_method left status=unreadable for every failed rules
read, including a 403 whose body is GitHub's own "Upgrade to GitHub
Pro or make this repository public" message. A repository whose plan
cannot expose branch rules cannot have a merge_queue rule either, so
that specific 403 now resolves to status=none instead of unreadable -
unblocking the away-merge grant on private repos without GitHub Pro.
Any other failure (auth, rate limit, network, 404, unrelated 403)
still reads as unreadable.

* no-mistakes(document): Update stale away-merge queue-grant comment for plan-gated 403

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>

* no-mistakes(review): Fix misleading away-queue-grant comment in fm-pr-merge and its test

* no-mistakes(document): Update architecture.md for plan-gated-403 merge queue exception

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>
…unchenguid#4246)

* fix(tests): select readers of a changed top-level test fixture

bin/fm-test-run.sh --changed recognised shared test helpers by an explicit
list, tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh. A top-level
tests/*-fixture.sh matched none of those, fell through to the tests/*
catch-all, and was marked unmapped, so selection aborted with "no
changed-test mapping for source path" and the run selected nothing at all.
tests/herdr-client-pair-fixture.sh and tests/remote-herdr-fixture.sh are
real shared fixtures with real consumers, so any branch touching one of
them left a validation pipeline driving --changed with a hard abort rather
than a narrowed selection.

Extend the helper arm to tests/*-fixture.sh rather than routing it through
the tests/fixtures/*/* arm. Both arms resolve consumers with the same
reference scan, and that scan is what selects the right suites here: it
finds exactly the tests that read the fixture. The fixtures/ arm adds only
a directory-keying step, which has nothing to key on for a top-level file,
so the helper arm is the same behaviour with no extra machinery. A
tests/ path nothing reads still reaches the catch-all and still refuses
loudly.

Refs kunchenguid#4100

* no-mistakes(test): order nested fixtures arm before top-level fixture glob

* no-mistakes(document): document tests/ shared-file mapping contract and arm order

* no-mistakes(review): drop vacuous test phase, correct header claim, restore comment
… asked, not declined (kunchenguid#4387)

* fix(bin): read Claude Code's default external-imports flags as never asked, not declined (kunchenguid#4378)

fm-claude-trust.sh refused the whole trust registration whenever the project-root entry
carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code
writes that value only on an explicit "No, disable". Claude Code's default project
entry carries Approved and WarningShown both false before the dialog is ever shown, so
every such project refused every spawn.

Only Approved === false with WarningShown === true — the pair the dialog writes on a
decline — now counts as a decline. false/false behaves like an absent flag: trust is
registered and no import consent is manufactured.

New case test_project_root_entry_default_import_flags_are_not_a_decline fails on
b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31,
bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): Correct harness doc's external-imports decline predicate

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…chenguid#4445)

* fix(brief): keep operator address out of composed intent

Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body.

The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content.

Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion.

Fixes kunchenguid#3882

* no-mistakes(review): Refuse operator-address lines in Captain's intent body

* no-mistakes(document): Document operator-address refusal in intent contract comments
…as a proven empty composer (kunchenguid#4455)

* fix(composer): accept Grok title overhang

* no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage
…ailure (kunchenguid#4474)

* fix(bin): recover Claude auto-arm after timeout

* no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list
* fix(spawn): establish Claude task channel authority

* no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference
…or pending text (kunchenguid#4458)

* fix: guard relaunch exit against pending input

* no-mistakes(review): Verifying test run in progress

* no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard

* no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now
…unchenguid#4460)

* fix: reconcile diverged secondmate updates

* no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md

* no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md
…d#4497)

* fix(dispatch): support Codex Luna max effort

* no-mistakes(review): use portable CODEX_HOME path in codex effort reference
kunchenguid#4498)

* feat(calm): render smooth Unicode swell

* feat(calm): make sails asymmetric

* feat(calm): use quarter sail glyph

* no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer

* no-mistakes(document): docs: sync calm wave phase doc comment

* no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。
kunchenguid#4491)

* fix: supersede scout delivery brief on promotion

* fix: preserve ship safety contract after promotion

* no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch
…d stop cleanup dropping accents from a held body (kunchenguid#4471)

* fix(bin): let captain holds work on hosts with an older JSON::PP

Holding a task for the captain, and the cleanup that keeps a captain-held row
open, both fail outright on any host whose JSON::PP defaults allow_nonref off -
2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`,
but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older
library rejects that whole value with "must be object or array".

The consequence is fleet-wide on such a host, not one broken command: a worker
there cannot formally record a decision for the captain at all. It can only
mention the decision in passing in a status line, where it can be missed - which
is how a real decision goes unrecorded. The hold reports that the task lost its
hold-set stamp; the cleanup cannot return the row to Queued.

Both call sites now ask for allow_nonref explicitly rather than inheriting
whatever the installed library defaults to. The second one is worth naming: its
`/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string
case that fails, so the guard selects for the failing input rather than
protecting against it.

The regression case forces the older default back off for every perl the commands
spawn, then drives both paths - holding a task that carries a body, and tearing
down a captain-held row whose deliverable must still be appended. It also probes
that the simulation genuinely rejects a bare scalar, so the case cannot pass
vacuously on a lenient host. Each half was verified failing on its own unfixed
call site with that site's real error message. Suites: fm-captain-hold-lifecycle
51 cases, fm-backlog-atomicity 99 cases, 0 failures.

Verification limit: the mechanism is reproduced and tested, but neither fix is
verified against a real JSON::PP 2.27202 host, because none is in the loop. This
laptop runs 4.06, where the bug does not manifest.

`bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a
brace-delimited object before decoding, so allow_nonref never applies.

* fix(bin): stop cleanup silently dropping accented characters from a held body

Cleanup rewrites a captain-held row's body to append the finished work's
deliverable, and the decoder it reads that body with printed decoded characters
to a stream with no `:raw` layer. A character at or below U+00FF then came out
as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the
accent. `fm_backlog_retain` writes that body straight back through
`--body-file`, and nothing reported an error - the character was simply gone
from a row still waiting on the captain.

The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus
`utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already
used.

Review of the parent commit found this on one of the lines that commit already
changed. It predates that change.

The test asserts bytes rather than decoded strings, because comparing strings
cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any
character above U+00FF makes perl print the whole string as UTF-8, so one body
carrying both an accent and an em dash passes even unfixed and proves nothing.
Verified failing before the fix on the accented row, passing after. Suites:
fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures.

* no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc

* no-mistakes(review): drop whole-file UTF-8 check from retained-body test

* no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc

* no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc
…furniture (kunchenguid#4532)

* fix(composer): read codex 0.154's idle starfield and status footer as furniture

codex-cli 0.154.0 animates a braille "starfield" around its idle composer:
on the row above the bold `›` prompt row, on the `›` row behind the SGR-2
dim `Ask Codex to do anything` placeholder, and on the row below it, then
draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`).
The cells are truecolor greys on both sides of the ghost luminance ceiling,
so the brighter ones survive ghost stripping, and the rows below the glyph
carry no structural edge. The shared classifier selected the bare `›` shape,
extended its wrap region over the two rows beneath the glyph, read the
survivors and the footer as wrapped typed input, and answered `pending`;
the steering doorbell defers on exactly that verdict, so no doorbell ever
reached an idle codex 0.154 pane.

bin/fm-composer-lib.sh now recognises that furniture by shape, declared
once next to the idle placeholders and reached from the two wrap-region
boundary points:
- a row whose non-whitespace content is entirely braille cells
  (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it
  never counts as wrapped typed content and bounds a bare composer's wrap
  region; braille behind the glyph row's content is stripped before the
  emptiness decision when nothing else follows the glyph; a row mixing
  braille with other text stays typed content;
- the codex status footer bounds the wrap region exactly as omp's status
  row does, anchored on the effort token, a spaced middle dot, and a `~` or
  `/` path cell, so a typed `fix · tests` stays composer input;
- `^Ask Codex to do anything$` joins the verified idle-placeholder set; the
  ghost strip remains what proves that row empty, and the bare-row rule that
  bright placeholder text is real input is unchanged.

Unchanged: the strict blank-row rule, the styled=0 degradation (a plain
cmux/orca capture of this screen still reads `unknown`, never `pending`),
FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape.

tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte
with the divergence (letters in place of the starfield read `pending`) and
the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh
is the default-on live guard (token-free, skips explicitly without codex or
tmux) that launches the installed codex idle and asserts `empty` through
both the tmux and the cursorless styled reads, naming codex --version on
failure. docs/verification/runtime-backends.md records the dated Herdr
evidence: `pending` before, `empty` after, on the captured screen.

* no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry

---------

Co-authored-by: Todd Billings <todd@usdvcapital.com>
kunchenguid and others added 28 commits September 28, 2026 14:25
…enguid#6033)

* fix(bin): read a quiet-mode record as a present captain, never hold-for-return

Daemon-backed quiet mode writes the away-posture record marked mode: quiet,
but the entry announcement, read-back, and session-start digest rendered it
as "hold-for-return only", and the spend cap and PR merge gate treated it as
away. A present captain's requested actions could then be held for a return
that was not coming.

bin/fm-afk-contract.sh now owns which posture a record is (the mode
subcommand, fm_afk_contract_mode, fm_afk_contract_away_present). A quiet
record announces, reads back, and appears in the digest as a present captain
holding nothing; merges under it stay attended and it binds no spend cap. An
away record is unchanged, an /afk entry over quiet mode rewrites the record
as away, and a quiet entry never turns a standing away record quiet.

* no-mistakes(document): Clarify quiet-mode authority and remove stale away guidance

* no-mistakes(ci): The CI failure came from a race in the supervision-host test: its restart fixture could observe a watcher left by the preceding cycle. The test now retires that watcher and waits for the fixture arm to report its own started cycle. The focused test passed three times; the full suite was attempted but stopped at a separate intermittent test failure

* no-mistakes(ci): Fixed daemon refresh mode selection so an unset-mode refresh follows the posture record: /afk over a running quiet daemon now changes state/.afk to away, while a plain quiet refresh stays quiet. Added script-level regression coverage for start and start-native and corrected a quiet-refresh fixture. The launch test suite, syntax checks, and diff check passed

* no-mistakes(ci): Herdr was blocked before tests ran by a GitHub HTTP 500 downloading pinned Treehouse; no code change was warranted for that check. Fixed the Lint 1 ShellCheck warning in tests/fm-afk-launch.test.sh by annotating the intentional background PID capture. The focused test suite, ShellCheck, syntax check, and diff check passed
…merge (kunchenguid#6053)

* fix(bin): accept a task's next PR once fm-pr-merge confirms the bound one merged

require_recorded_pr_identity now checks fm_pr_poll_merge_already_notified for
the recorded pr= before refusing a different URL, so a task's later PR is
accepted once its earlier PR's merge is confirmed, while it keeps refusing
while the bound PR is still unmerged.

* no-mistakes(document): docs(fm-pr-merge): note next-PR accepted after bound PR merges
…6064)

* fix(bin): read a live quiet record as a present captain at the host and watcher

A quiet record left without its daemon (a quiet start that never ran or was
interrupted) was read as away by the supervision host, so it parked a present
captain's main and held captain outcomes for a return that never comes, and
the watcher and daemon silenced captain-held rechecks on record presence.

The host's posture checks, the watcher's and daemon's captain-held silencing,
and the host's outcome path (branch report, drain BRANCH OUTCOMES, relocated
branch authority, the owners' away wake note, and the Codex checkpoint bound)
now ask the record owner's away-or-quiet reading, so only an away record is
away. A live away record keeps today's behavior.

* no-mistakes(document): Correct quiet-record documentation and supervision guidance

* no-mistakes(document): Clarify quiet-record posture and captain-held rechecks

* no-mistakes(document): Clarify quiet-record posture in documentation
…kunchenguid#6043)

* fix(bin): name an in-window engine latch in the return brief and drop the false handling GAP line

The away return brief said nothing had failed after the supervision host
latched on engine errors during the window, and printed a GAP: watcher
downtime line whenever a wake was merely being handled or queued at return.

The failures section now reads the host ledger and latch record and names
the latch time, the window's engine-error count, and whether the session
is still paused or recovered. An open recovery episode is reported as
information, and as a gap only when a queued episode outlived the return
grace or the marker cannot be read.

* no-mistakes(review): Fix latch trip time, drop marker-age grace, bound error count

* no-mistakes(review): Report paused latch without ledger trip row; bound errors

* no-mistakes(review): Never report a failed probe's latch row as trip time

* no-mistakes(review): Only a retained trip row marks a pre-window latch

* no-mistakes(document): Clarify return-brief latch and watcher-gap documentation

* no-mistakes(ci): Fixed Lint 1 by marking the shared cooldown constant as used by sourcing scripts. The repository lint command and diff check pass; the return test run was stopped by a 180-second timeout after its completed cases passed

* no-mistakes(ci): Fixed the return brief so the trip time and error count come from the same initial latch row, and ledger rows before the current session’s lock boundary cannot affect its latch report. Added real-script regressions for both findings. The return test suite, repository lint, and diff check pass

* no-mistakes(ci): Fixed the return brief’s restart cutoff so it retains in-window failures, prints one line per initial-trip row, and omits zero-error count wording. Added real-script restart regressions. The return test suite, ShellCheck, and diff check pass

* no-mistakes(ci): Fixed the return brief so a recorded trip followed by recovery stays recovered, while a later pause with a lost trip append gets a separate “trip time unavailable” line. Added a real-script regression that failed before the fix. The return test suite, ShellCheck, syntax checks, and diff check pass

* no-mistakes(ci): Fixed the false second latch during recovery. A real-script regression failed before the fix and passes now; the lost-second-trip test still passes. The return test suite, ShellCheck, syntax checks, and diff check pass
* fix(calm): name the Claude Code Calm plugin fm so supervision notes read "fm: "

Claude Code labels every mod transcript line with the plugin name, so the
notes rendered as "firstmate-calm: ⚓ ...". Rename the plugin to fm, update
the live guard to assert the fm: label, and document the one-time replay for
sessions resumed across the rename.

* no-mistakes(document): Clarify Calm plugin rename in documentation
…#6037)

* feat(bin): add fm-live-lab.sh, a one-command live supervision lab builder

* fix(bin): exact lab windows, per-lab task ids, self-safe teardown

* fix(bin): target lab windows by id, stop lab descendants, add readiness tests

* fix(bin): keep Claude's auto-updater off in live labs; list fm-live-lab.sh

* fix(bin): start the lab tmux server without user config

* no-mistakes(review): Scope lab teardown to its store, root, and task ids

* no-mistakes(review): Record selected user stores at up for check and down

* no-mistakes(document): Clarify live lab documentation and remove stale narratives

* no-mistakes(ci): Fixed the CI failure by checking for an existing lab root before looking up the harness executable. The affected behavioral test and shell syntax check pass; the refusal also works with Claude absent from PATH

* no-mistakes(ci): Fixed all four Greptile findings: teardown signals only recorded lab processes and their descendants; the worker gate is in its granted task directory and its path is exposed; readiness uses current crew state; and mate and worker IDs use 12 nonce hex digits. The CLI behavior tests pass, as do shell syntax, ShellCheck, and diff checks. The Claude no-host path is unchanged

* no-mistakes(ci): Fixed the CI test’s dependence on an installed Claude binary by supplying a test-local stub. The full fm-live-lab test, shell syntax check, and diff check pass

* no-mistakes(ci): Fixed all three selected findings in bin/fm-live-lab.sh: down waits for recorded processes and escalates before cleanup, PID roots are checked against recorded start times, and Claude primary trust is rechecked after mate/worker readiness. Added behavioral tests in tests/fm-live-lab.test.sh. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass

* no-mistakes(ci): Fixed the pre-primary settle wait, worker gate instructions, unused retry variable, and teardown PID revalidation in bin/fm-live-lab.sh. Added behavioral tests in tests/fm-live-lab.test.sh. Both requested commands pass: tests/fm-live-lab.test.sh and bin/fm-lint.sh

* no-mistakes(ci): Fixed teardown to track pre-kill lab processes by PID and start time, including children orphaned when a root exits. Up now rejects an empty pane PID before calling ps. Added regression tests and a Linux-safe worker fixture. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass

* no-mistakes(ci): Fixed teardown tracking for children spawned during shutdown and made the worker fixture verify its exact window with a Linux-available shell. Both requested checks pass. The lab test takes about 66 seconds locally, so the under-one-minute target remains unmet

* no-mistakes(ci): Fixed ci-2 and ci-4 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. Teardown now tracks identity-checked members of captured lab process groups, including children orphaned during shutdown, without signaling the caller’s group or unrelated processes. Lint passed, and the lab test passed four times

* no-mistakes(ci): Fixed teardown so an observed-empty process group is permanently dropped, preventing a reused group ID from signalling unrelated work. Added a ps-shim regression test. The lab test, lint, and diff checks pass

* no-mistakes(ci): Fixed ci-1 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. The TERM-born-child fixture now waits until its handler is installed before calling down. Down sends SIGKILL to identity-valid survivors on every pass from pass 20 onward and includes survivor process details if it must refuse cleanup. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass locally; Linux CI remains to be verified

* no-mistakes(ci): Fixed down’s teardown wait to require two empty identity-checked scans separated by 0.5 seconds, and removed the unused test loop variable without changing the TERM-born-child test. The lab test, lint, and diff check pass locally
…unchenguid#6103)

* fix(bin): keep slow watcher cycles and preempted reply polls from breaking supervision

- fm_pending_reply_tick selects the records it has work for in one awk pass,
  so settled records cost no lock or fork and the walk no longer grows with
  the never-pruned store.
- An attached arm keeps following a live, identity-matched holder whose beacon
  went stale until the lock changes or the shared stall bound
  (fm_watcher_stall_bound), then fails with a typed stalled-holder line so the
  retry replaces the holder.
- The remote-reply adapter reports the job worker's preemption (exit 76) as a
  closed window, so the listener keeps its claim and polls again instead of
  being relaunched every watcher cycle.

* no-mistakes(document): Clarify watcher grace and attached-arm documentation
…geable is UNKNOWN (kunchenguid#6110)

* fix(bin): retry a bounded number of times when GitHub mergeable is UNKNOWN

Fixes kunchenguid#6020

bin/fm-pr-merge.sh refused a GitHub merge whenever the pull request's
mergeable field was not literally MERGEABLE. GitHub reports UNKNOWN for
a short while after a push or a base-branch change while it recomputes
mergeability, so a green, conflict-free pull request was refused as if
it could not be merged.

github_verify_mergeable now returns a distinct status when mergeable is
the only failing condition and reads UNKNOWN. The caller retries up to
5 times, 3 seconds apart (overridable in tests), re-reading and
re-checking every live condition on each attempt. Once the bound is
spent it reports mergeability as still being computed rather than
unmergeable, with the same nonzero exit as before. Every other refusal
(closed, draft, conflicting, red or missing checks, away authority,
queue protection) is unchanged and never retried.

* no-mistakes(ci): I fixed both review findings the way you asked. The full suite (`bash tests/fm-pr-merge.test.sh`) ran to completion. Its last lines showed all `ok`, and any failure would have stopped the run early. I watched the output through `tail`, so I didn't see the new test's own `ok` line directly. **ci-2 (`bin/fm-pr-merge.sh`), retry delay not validated.** What must hold: the retry wait is always a short, valid `sleep` argument, so a bad `FM_PR_GITHUB_MERGEABLE_RETRY_DELAY` can never trip `set -e` or hold the task lock for a long time. The retry loop is the only place that reads this variable. The script now reads the value once before the loop and accepts only whole numbers from 0 to 10. Anything else (empty, `abc`, `-1`, `1.5`, `11`, a huge number, leading spaces) falls back to 3. I ran those values through the check by hand and each came out as expected. The loop now sleeps on that checked value. **ci-1 (`tests/fm-pr-merge.test.sh`), no test for a check changing between UNKNOWN reads.** What must hold: every retry re-checks all live conditions, not just mergeable. The fake `gh pr view` in the test can now take an optional second word on each line of the mergeable sequence, which sets the first check's result. The new test `test_github_mergeable_unknown_retry_rechecks_checks` feeds `UNKNOWN`, then `UNKNOWN FAILURE`. It asserts: - exit code 1 after exactly 2 reads, - the refusal names `check 'ci' is not green`, - the message does not say mergeability is still being computed, - `pr merge` was never called. If a later change made the retry look only at mergeable, the loop would read UNKNOWN 5 times, end with the "still being computed" message, and this test would fail. I didn't run it against a deliberately broken script to confirm that. `bash -n` passes. `shellcheck` reports only the existing info-level notes about files it can't follow. Only `bin/fm-pr-merge.sh` and `tests/fm-pr-merge.test.sh` changed
kunchenguid#6112)

* fix(bin): converge every open owner onto a known terminal contribution

settle_final only cleared a stale error on retry, so an owner whose saved
row still said open kept projecting a merged or closed pull request as
open after another owner's row had already recorded the terminal
observation. Copy the known terminal observation to every owner whose
saved row is not itself terminal, keeping that owner's own pending and
notified state, and clear its error.

* no-mistakes(review): Carry terminal checked_at when converging existing owner rows

* no-mistakes(ci): I fixed Greptile finding ci-2 as you asked, with a change to tests/fm-contributions.test.sh only. The rule it enforces: when a retry converges an owner onto a URL that is already merged or closed, that owner gets the terminal owner's whole observation, not just its state. The same weak check appeared twice in test_interrupted_multi_owner_poll_settles_every_owner, so I fixed both: - **Open owner (line 784):** the check now also requires `.observation == $terminal[0].records[0].observation`. The existing checks for error, checked_at, pending and notified are unchanged. - **Errored owner (just below):** it only checked state and error before. It now reads the terminal owner's file and makes the same full-observation comparison. Adding the comparison alone would not have caught anything. The test fixtures gave both owners identical observations apart from `state`, so copying only the state would still have passed. In both cases I also set the terminal owner's observation head to HEAD_B, so the two observations now really differ. Verification: - The focused test passes against the current bin/fm-contributions.sh. - I temporarily changed `settle_final` so it copied only the state. The test then failed, reporting the owner still on the old head (HEAD_A). I restored the file afterwards, and `git status` shows only the test file modified. - The full tests/fm-contributions.test.sh suite exits 0. No product code changed. The other CI finding (ci-1, "Behavior portable serial 9") was left alone because you chose to ignore it
…nguid#6124)

* feat: run the supervision host by default on a Claude primary

An absent config/supervision-host on a Claude primary now reads as on with
the default engine, and a file holding `off` opts any home out. Cursor,
OpenCode, omp, Grok, and Codex stay file-gated, with `off` read as disabled
there too. Every reader asks fm_supervision_host_enabled instead of testing
the file, and non-bash readers query it through the lib's `enabled` entry.
A primary's `off` is not inherited by secondmates: each home keeps its own
supervision posture.

* test: pin the watcher-path posture in fixtures that assume no supervision host

Fixtures that drive the watcher arm or assert a non-host drain now write
an explicit off file, and fixtures that copy the Stop auto-arm or the
supervision instructions carry the engine lib they now source. The two
drain suites also stop reading the code root's config.

* fix: name the opt-out when an off home passes an attended wake to main

A host parked when the home writes off now logs that the home does not run
the supervision host, rather than claiming it has no engine.

* no-mistakes(document): Clarify Claude supervision defaults and historical evidence

* no-mistakes(ci): Fixed process leaks in the two added host tests. Each case now stops its recorded watcher and host/arm processes; fake hook sessions exit through session.stop. The full host suite passed before the final cleanup refinement, and both affected cases, bash syntax, ShellCheck, and diff checks passed afterward. CI runtime still needs confirmation
…start scope check (kunchenguid#6125)

* fix(bin): create the state dir on a fresh primary before the session-start scope check

fm_primary_scope_matches required an already-existing state directory, so
bin/fm-sessionstart-run.sh stood down on a fresh clone before anything could
create it. Split out fm_primary_root_matches so the run wrapper can confirm
primary-home identity first, create the gitignored state dir when it is
missing, and only then run the unchanged scope check.

* no-mistakes(document): Document session-start state dir creation on fresh clones

* no-mistakes(ci): I fixed the Greptile P1 the way you asked. When a fresh primary can't create `state/`, the run wrapper no longer stands down silently. **Invariant:** when an otherwise eligible fresh primary cannot create `state/`, startup must never fail silently. This path has only one site: the mkdir in `bin/fm-sessionstart-run.sh`. Other hooks and the nudge wrapper never create `state/`, so they have no equivalent failure. **What changed:** - **Run wrapper** (`bin/fm-sessionstart-run.sh`): it captures mkdir's error and prints one line to stderr before standing down as before (exit 0, or 3 for the Pi prerequisite). The line looks like `fm-sessionstart-run: startup could not create the state directory <path>: <reason>`. - **Test** (`tests/fm-sessionstart-nudge.test.sh`): the new case `test_run_reports_a_state_dir_it_cannot_create` uses a fresh primary with no `state/` and a read-only (0500) root. It checks four things: exit 0, no digest on stdout, no state dir created, and exactly one stderr line ending in "Permission denied". It fails without the fix and passes with it. - **Docs** (`docs/sessionstart-nudge.md`): I added one sentence describing the stderr line and one describing what the new test proves. **Verification:** I ran `tests/fm-sessionstart-nudge.test.sh`, and every test passes. `bin/fm-lint.sh` on the changed scripts (pinned ShellCheck 0.11.0) and `tests/fm-documentation-audiences.test.sh` also pass. As you asked, the wrapper still stands down with the ineligible-checkout status afterwards. It does not report this as a failed eligible startup, which is what the bot suggested
…ery (kunchenguid#6126)

* fix(bin): measure pending-reply grace from turn completion, not delivery

Fixes kunchenguid#6057

The pending-reply guard demanded a repost ("REPOST REQUIRED: previous
marked request had no correlated parent report") while the second
mate's correlated reply was already on its way.
fm_pending_reply_send_recovery measured its grace window from delivery
instead of from the request turn's completion, so any turn longer than
the grace fired the demand the moment the turn ended, before the reply
could have landed. The missed-report escalation had the same gap: it
fired the instant the recovery turn's completion was observed, with no
grace at all.

Both now measure grace from the relevant turn's completion (request
turn for the recovery repost, recovery turn for the escalation), and
both take one fresh, uncached read of the parent status file
immediately before firing, accepting a correlated line regardless of
its verb. Transport-failure escalations stay immediate, and the
one-repost limit is unchanged.

* no-mistakes(review): Document grace window as measured from turn completion

* no-mistakes(ci): Both Greptile findings were real and caused by this PR, so I fixed them. The full `tests/fm-pending-reply.test.sh` suite passes. **ci-1 (a reply could be overwritten by a repost).** The rule that must hold: a recovery send is recorded only if the record is still unresolved, checked under the same per-correlation lock that resolution uses. The escalation path already did this (`_fm_pending_reply_maybe_escalate_locked` reads fresh and publishes under one lock). The recovery path did not: `fm_pending_reply_send_recovery` did its fresh read through `fm_pending_reply_try_resolve`, which let go of the lock before the send was recorded. A reply landing in that gap could be overwritten, and the repost would go out anyway. Now `send_recovery` takes the lock once and, while holding it, re-checks that the phase is still `awaiting_report`, runs the fresh uncached read, and records the send (sender pid and identity, attempt time, phase `recovery_sending`). It releases the lock before actually sending, so the lock is not held during the send. It uses the same lock helpers the other lock wrappers use. Grace timing, the one-repost limit and the escalation path are unchanged. **ci-2 (the test would pass even without the fix).** In `test_recovery_fresh_status_read_resolves_before_firing`, the reply is still appended to the status file, but the stored file signature is then set to the file's new signature. That stands in for a same-size rewrite that the signature cache cannot see. The test first checks that a normal cached read misses the reply, then that the fresh read before sending catches it. I also added the same check for the fresh read before escalation, which the review said was uncovered. The test now sets its own send hook, so it no longer depends on one left over from an earlier test (that leftover had made failures exit silently). **Checks:** - I removed the fresh-read bypass at each site in turn and reran the suite. With it gone from recovery, the test fails with "recovery must not fire once a correlated reply has landed". With it gone from escalation, it fails with "the fresh pre-escalation read should have resolved the record, got escalated". With both in place, all tests pass. - Shellcheck with `-x` timed out locally. Without `-x` and ignoring SC1091, the only warnings are SC2034 on the existing `maybe_escalate` lock wrapper, which is not part of this change. The new code adds no warnings. Changes are in `bin/fm-pending-reply-lib.sh` and `tests/fm-pending-reply.test.sh`. Nothing is committed yet; a plain commit message such as "fix(bin): record the pending-reply recovery send under the fresh-read lock" fits the instruction

* no-mistakes(ci): ci-1 was real and caused by this PR. The same bug was also in the escalation path, so both are fixed. The full tests/fm-pending-reply.test.sh suite passes. The rule that must hold: a recovery repost or an escalation goes out only if the record's phase, read after the fresh-read resolve, is still what it was before. The resolver writes phase=resolved first and only then writes the other resolution fields. If one of those later writes fails, it returns an error even though the record is already resolved. Places this rule applies, both fixed: - Recovery (fm_pending_reply_send_recovery): the fresh-read resolve now runs first, and the phase is re-read right after it, whatever it returned. The send is recorded and made only if the phase is still exactly awaiting_report. This replaces the earlier phase check rather than adding a second one. - Escalation (_fm_pending_reply_maybe_escalate_locked): same bug. After a failed resolve it went on to publish the blocked line and set phase=escalated. One added line after the resolve call returns 1 without publishing if the phase has changed. Test: added test_partial_resolve_write_blocks_firing. It forces a failure on the resolved_epoch write after a correlated reply has landed. It checks that the recovery send hook is never called, that no escalation line is published, and that the phase stays resolved. The forced failure runs in a subshell so it can't affect later tests. Checks: - With the recovery fix reverted, the new test fails with "recovery must not fire after a partial resolve". - With the escalation fix reverted, it fails with "partial resolve should block escalation, got escalated". - With both fixes in, every test passes. - Shellcheck was run with SC1091 excluded and without -x, not through the repo's lint script. The only new message is one SC2329 info on the test's override function; other test overrides in the same file already get that same info, unsuppressed. Changed files: bin/fm-pending-reply-lib.sh and tests/fm-pending-reply.test.sh. Nothing is committed. Suggested plain commit message: "fix(bin): recheck pending-reply phase after the fresh read before sending
…to stderr (kunchenguid#6001)

* fix: provider-table lookup never writes a broken-pipe error to stderr

Fixes kunchenguid#5956

fm_quota_single_provider_for_harness returned from its while read loop
as soon as it found a match, closing the pipe while
fm_quota_single_provider_table's printf could still be writing.
Where SIGPIPE is ignored, as on GitHub Actions runners, bash then
prints "printf: write error: Broken pipe" on the resolver's stderr,
which intermittently broke the one-diagnostic-line assertions in
tests/fm-dispatch-resolve.test.sh.

Read the whole table before answering, the way
fm_control_harness_supported already does, so the writer always
finishes. Return values and output are unchanged.

Reproduced by running tests/fm-dispatch-resolve.test.sh with SIGPIPE
ignored on a single pinned core under CPU contention: 30 of 30 runs
failed before the fix, 0 of 30 after. Note: reproducing requires
setting the trap inside the tested shell because nice(1) resets an
inherited SIGPIPE ignore to SIG_DFL. tests/fm-quota-choose.test.sh
passes and bin/fm-lint.sh is clean.

* no-mistakes(ci): Fixed both Greptile findings the user chose to address. ci-1 (bin/fm-quota-axi-lib.sh:154). Invariant: looking up a harness must always end with status 0 and print the provider, even when the caller runs under `set -e`. The loop body `[ -z "$found" ] && [ "$harness" = "$1" ] && found=$provider` now ends in `|| :`. Every iteration succeeds and the whole table is still read. Only `fm_quota_single_provider_for_harness` loops over the table this way, so this is the one place the fix was needed. One caveat: on bash 5.3 the old code did not actually exit under `set -e`, because the `while` loop is not the function's last command, so the new `set -e` test would have passed before this fix too. The change makes the loop's success explicit, as the user asked. ci-2 (regression coverage). I added three cases to the existing `tests/fm-quota-choose.test.sh`, all calling the public lookup function after sourcing the library: 1. With SIGPIPE ignored (`trap "" PIPE`), it looks up every harness 200 times and checks that nothing reaches stderr. 2. A deterministic version of the race: the table function is wrapped so it writes the first row, pauses 0.2 s, then writes the rest. With SIGPIPE ignored, it checks that looking up `claude` prints `claude` and writes nothing to stderr. The stress loop alone reproduced the bug in only about 1 of 5 local runs, which is why this case exists. 3. A direct call under `set -e` prints `claude`. Verification: - `bash tests/fm-quota-choose.test.sh`: all pass. - Same test against the pre-PR library (fa48367, via `FM_ROOT_OVERRIDE`): fails with `printf: write error: Broken pipe`. The deterministic case failed in one run and the stress loop caught it in another. - `shellcheck` on both files: clean. - `tests/fm-dispatch-resolve.test.sh`: passes
…isioning (kunchenguid#6162)

* fix: survive Pi 0.99 rendering and Git 2.55 local-clone races

Pi 0.99 puts arguments on the stock tool header and leaves hidden custom messages in the export conversation column. Match that header, and keep Calm's boundary on the visible column. Clone a remote home with --no-local so a prune during Git's loose-object copy cannot fail the seed.

* no-mistakes(review): Stop SIGPIPE write errors; cover older Pi export and project clones

* no-mistakes(document): Clarify Calm export visibility and tool rendering

* no-mistakes(ci): Fixed the dispatch diagnostic to list every provider-less use/default profile in one line and added a multi-profile behavior test. Shortened supervision fixtures using the existing engine-grace and park-clock knobs; removed stray scratch files. Dispatch tests, syntax checks, and three targeted supervision cases passed. CI’s prior supervision duration was 751s; the single permitted local full-suite run timed out at 1200s, so an after-duration is not established. The cancelled serial check had no failure verdict. The outer executor should record the measured before/after duration in the PR body when available

* no-mistakes(review): Gate Pi 0.99 call headers by version; drop hidden-row assertion

* no-mistakes(review): Test stock call headers under Pi 0.87 and 0.99 stubs

* no-mistakes(test): Fix older-Pi queued-row test and verify park-boundary behavior

* no-mistakes(document): Clarify Pi Calm export and queued-turn documentation

* no-mistakes(ci): Fixed the stock macOS Bash 3.2 parse failure in tests/fm-calm-pi-extension.test.sh; its parse check passes. The watcher CI failure is in unchanged code: the isolated five-minute/66-minute case passes locally, but the CI log omits the drain error needed to establish its cause. No speculative watcher fix was made. The full local watcher suite timed out after 500 seconds
…6169)

* Prevent premature Lavish board handoffs

* Prove Lavish arm lacks reply acknowledgement

* Confirm Lavish replies before arming worker boards

* no-mistakes(review): Post Lavish reply only after locked arm eligibility checks

* no-mistakes(review): Fail Lavish reply closed on unknown version

* no-mistakes(document): Correct Lavish reply documentation and remove stale guidance

* no-mistakes(document): Clarify Lavish reply routing and remove duplicate version guidance
…#6154)

* feat: inherit the supervision-host opt-out from the primary

Move the supervision host's off opt-out out of config/supervision-host into
its own presence flag, config/supervision-host-off, and add that flag to the
primary-authoritative inherited config set. A primary that opts out now opts
every secondmate home out at spawn and convergence, and clearing it converges
them back. config/supervision-host stays the home-local engine choice.

Shape: config/supervision-host mixed two things, a fleet posture (off) and a
per-home engine and model. Only the posture should follow the primary, so it
becomes a separate presence flag that rides the existing inherited-config
mechanism (FM_INHERITABLE_CONFIG in bin/fm-config-inherit-lib.sh) with no new
machinery, while the engine line stays local. The parse stays in its one
owner, fm_supervision_host_enabled. There is no migration or compatibility
handling for a home that still holds off in config/supervision-host.

Primary off, mate on: inherited material is primary-authoritative by design,
so a mate cannot keep the host while the primary is opted out, and a mate's
own opt-out is removed at the next convergence while the primary has none.
Running the host on a mate is the primary's choice for the fleet; no override
mechanism is added.

Live validation (disposable bin/fm-live-lab.sh lab, Claude primary with a
real seeded secondmate, --supervision-host off):
- up: every readiness check ok, including "host: none running, as expected"
  and a live mate session; the spawned mate home held the inherited
  config/supervision-host-off and the gate read primary OFF, mate OFF.
- primary removed its opt-out, then bin/fm-config-push.sh reported
  "supervision-host-off: pushed - mirrored primary absence" and a config
  reread sent; the gate read primary ON, mate ON, and the live mate handled
  the reread.
- primary opted out again and pushed: "supervision-host-off: pushed", mate
  gate OFF.
- down stopped every lab process and left no lab process running.

Out of scope, follow-up: default-on for the other harnesses, away-daemon
retirement, rollout.

* no-mistakes(document): Document inherited supervision-host opt-out ownership

* no-mistakes(ci): Fixed ci-4: with `--supervision-host off --mate`, lab readiness now requires the inherited flag in the mate home and a disabled mate supervision-host gate. The focused behavior test, shellcheck, and diff checks pass. Left ci-1–ci-3 untouched as directed

* no-mistakes(test): Fix mate readiness HOST_OFF initialization in lab up

* no-mistakes(ci): Fixed Lint 2 by making the new test’s fixtures source resolvable to ShellCheck; its off/on readiness test and ShellCheck now pass locally. Behavior portable serial 5 failed in the unchanged remote-reply test at generation 7. That test passes locally, and no PR-caused defect was identified, so no remote-reply code was changed
…nguid#6179)

* fix(tests): cut the fixed sleeps in supervision-host cycles

The serial CI lane keeps brushing its 30-minute cap because
fm-supervision-host.test.sh spends ~903s of the job, and per the
run-36635306527 case profile the top nine cases are all multi-cycle
ones (3-10 park/close/turn cycles each): every close waits out the
host's sleep $POLL in await_close plus a watcher sleep $FM_POLL scan
cycle, and every engine turn waits out the fixed sleep 1 descendant
snapshot. That is ~3s of pure sleep per cycle before any real work.

The host poll now accepts positive decimal seconds through a new
seconds_or validator (FM_SUPERVISION_HOST_POLL), and the engine turn's
snapshot loop takes FM_SUPERVISION_ENGINE_SNAPSHOT_SECONDS, also a
positive decimal defaulting to one second - the smallest seam at each
wait's single owner. The suite drives them at 0.2 alongside the
existing FM_POLL=0.5 and FM_ARM_ATTACH_POLL=0.2 knobs, so the real
poll loops still run. The park-boundary case moves onto the injected
test clock instead of a real 3s wait, per-case cleanup polls the host
pid rather than sleeping a full second, and the proof-by-absence
windows (flood re-escalation, successor re-announce, watcher
persistence, recovery staying off main) shrink from 2-3s to 1s, which
still spans two watcher polls at the test cadence.

Every assertion, process lifecycle, and reaping path is unchanged;
production defaults stay at one second. Isolated case timings on a
contended host, base vs branch: attended-latch 54.3->34.6s,
undelivered-dialog 67.7->59.1s, away-latch 46.5->30.5s, held-cadence
47.9->21.6s, unreadable-mirror 39.2->38.5s, park-limit 18.2->12.3s,
registration-fallback 14.1->10.0s, first-cycle-status 12.6->8.4s,
latch-scope 16.7->16.3s. Full suite: 65/65 pass. fm-lint and
shellcheck clean.

* no-mistakes(review): Wait for scan lock release before duplicate check

* no-mistakes(document): Correct supervision snapshot cadence documentation

* fix(tests): keep production poll cadence, probe exits at 0.1s

The fractional poll cadences multiplied the cost of each loop body:
full process-table scans in the engine turn and process refreshes in
await_close ran five times more often, which swamped the thin CI runner
and nearly doubled every multi-cycle case (serial 5 was cancelled at its
30-minute limit on run 36635306527's successor). Restore the production
cadence and notice arm/engine exits with a cheap kill -0 probe at a
tenth of a second between the one-second bodies instead: strictly less
dead time than baseline with no added CPU.

Also hold each injected-clock park bound well past its case's
wall-clock checks so a host that ignored the test clock fails instead
of silently passing at a real-time boundary, and restore the shortened
proof windows (watcher liveness, recovery-off-main absence, first-cycle
stream) to their baseline depth.

* no-mistakes(document): Clarify supervision engine snapshot documentation
…henguid#6192)

* fix: rebalance portable CI from current duration measurements

* no-mistakes(test): Test serial packing boundary and verify endpoint timeout cleanup

* no-mistakes(document): Clarify timeout guidance and remove duplicated packing estimates
…nguid#6216)

* fix(bin): run no repository hook when core.hooksPath is empty

The per-task hook wrapper refused every commit in a repository whose own
config sets core.hooksPath to the empty string, because git rev-parse
--git-path hooks fails on it. Plain git reads that setting as no hooks, so
the wrapper now runs none; every other lookup failure still refuses and
shows git's error.

Fixes kunchenguid#6171

* no-mistakes(review): Refuse commits when core.hooksPath is a valueless key

* no-mistakes(document): Document empty core.hooksPath handling in commit attribution docs

* no-mistakes(ci): When the wrapper refuses a commit, Git's hook-lookup error now shows up once instead of twice. That required changing one line in the wrapper, and the tests were extended so both bad-config cases would catch the duplicate. Invariant: when the wrapper refuses, Git's lookup error must appear exactly once. In the failure path, the only Git call besides the deliberate second lookup is the `git config --get --type=path core.hooksPath` check in `runtime_chain_body` (`bin/fm-git-strip-ai-trailers.sh:168`). That check prints the same error, so it was the one place to fix. I added `2>/dev/null` to it. Its exit status still decides the outcome: an empty value still runs no hook, and anything else goes on to the second lookup, which prints Git's error once, and the commit is refused. Tests (`tests/fm-git-strip-ai-trailers.test.sh`): - The unresolvable-path test (`~fm-no-such-user-6171/hooks`) now requires `failed to expand user dir` to appear exactly once in the refused commit's output. - The valueless-key test now requires `missing value for 'core.hookspath'` to appear exactly once. - Pre-existing bug in the unresolvable-path test: its `git add` ran after the bad config was set, so it failed silently (exit 128) and the "refused commit" had nothing staged. The test now stages the file before writing the config, the same way the valueless test does, so a real commit gets refused. - The empty-string test is unchanged and still passes, so an empty `core.hooksPath` still runs no hook. Verification: - With the wrapper change reverted, both new checks fail with `expected '1', got '2'`. With the change in place, the whole suite passes. - `bash -n` passes. shellcheck shows only an info-level SC1091 note about sourcing `lib.sh`, which was already there before this change. - `git status` lists only the two intended files
…kunchenguid#6213)

* fix(bin): let a stale record on a reassigned slot retire records-only

When a pool slot's owner claim names another task, the stale record's
teardown touches nothing under the slot, so the exclusive-slot record scan
no longer refuses it. Full teardowns of a slot this task still claims, or
one with no claim, keep the refusal.

Fixes kunchenguid#6184

* no-mistakes(document): Note claim-over-record precedence for reassigned teardown slots
…uid#6240)

* fix(bin): keep the steering doorbell short under deep homes

The doorbell printed the task inbox's absolute path twice, so under a deep
home it grew to about 290 characters and a Herdr submit reported it never
reached the pane on every re-ring. It now names the inbox once by its short
<task>.inbox name and points at the full path the worker's brief already
gives, so its length no longer depends on the home's depth.

Fixes kunchenguid#6120

* no-mistakes(review): Export FM_TASK_INBOX at launch and name it in doorbell

* no-mistakes(ci): ci-1 (Behavior portable serial 9) was caused by this PR, and I fixed it in the test. tests/fm-claude-trust.test.sh failed with "the launch command did not carry a brief doorbell". Its claude_launch_doorbell helper stripped exactly two leading `export ...;` statements before reading the final prompt argument. This PR adds a third one (`export FM_TASK_INBOX=...`) to every launch, so the helper was reading the wrong command. The invariant: a test that parses the launch command must skip every leading export statement, however many there are. I checked every test that parses the launch this way. The only other ones are the two helpers in tests/fm-spawn-dispatch-profile.test.sh, and they already loop over all exports. The kimi and dispatch-profile exact-string checks were updated earlier in this PR. The fix makes claude_launch_doorbell use the same loop (`while [[ "$command" == export\ *\;* ]]; do command=${command#*; }; done`) and then take the last argument. The ordinary path still works: the claude spawn test and the secondmate-clone spawn test both resolve the brief record through the same helper. Verified locally: `bash tests/fm-claude-trust.test.sh` exits 0 with no failing cases. ci-2 (Behavior tests (Herdr)) was not caused by this change, and I made no code change for it. In tests/fm-backend-herdr-presentation-e2e.test.sh, the concurrent secondmate recovery failed with "herdr presentation recovery could not acquire its session lock; refusing a concurrent resume". Two reasons it is not this PR: - The same failure, in the same test and case, happened on run 36655209015 for the unrelated branch fm/fm-contributions-old-gh-compat about 14 hours earlier. - This PR's change cannot lengthen how long the lock is held. The launch is written to a file and sent to the pane as `. launch.N.sh`, so the extra export changes neither the pane submit nor the lock hold time. The cause is a race that was already there: spawn_herdr_presentation_order_lock_acquire gives up after 5 seconds, and a concurrent real-Herdr recovery can hold the lock longer. Fixing that means changing the product's lock timeout, which is outside this PR. It should be tracked separately, and a rerun of the Herdr job is expected to pass. The only file changed is tests/fm-claude-trust.test.sh
… no turns (kunchenguid#4859)

* fix(dod): drive no-mistakes with one foreground call, not a background poll

The brief told workers to background the drive call and poll `axi status`
because one call "routinely outlives what your harness lets a single
command run". That advice contradicts the tool it drives: `no-mistakes
axi run --help` documents `--wait` with an 8m default, existing precisely
"so an agent harness with a 10-minute tool cap gets a structured return
instead of an unbounded hang".

Following the old text, a worker could never idle - a backgrounded call
returns in milliseconds, so it does not wait at all - and each attempt
leaked a live timer that later fired as a paid wake. Tell workers to make
one foreground call, let it block, and repeat it when it returns on
elapsed wait rather than on a gate or outcome.

Also drops the generalisation that told workers on any unestablished
harness to assume a command cap and use the same shape, which exported
the defect to harnesses with no such cap.

* fix(bin): let a waiting worker spend no turns until it is answered

A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot
kept taking model turns: the brief told it to list its inbox at any natural
checkpoint, and six automatic senders nudged secondmates whatever their open
decisions.

- The ship and scout briefs gain one Waiting section: end the turn after
  needs-decision or blocked, and hold an external wait inside ONE blocking
  command bounded by the harness's own command ceiling. The checkpoint clause
  is deleted. Forbidding the wrong shapes is not enough on its own, so the
  section also names the blocking foreground `until` loop as the wait a Claude
  Code worker may use, because that harness can refuse a sleep-then-check
  command while pointing at backgrounding, which is the one shape a waiting
  worker must not take.
- fm-send --automatic defers (exit 4, nothing written or rung) while the
  target has an open decision or blocker of its own; every automatic sender
  passes it and keeps its retry state, and the pending-reply recovery waits
  the same way.
- The two senders that report the result classified it by matching the text of
  the send's captured output against `deferred:*`. fm-send runs bin/fm-guard.sh
  as a supervision warning, and that guard prints its worktree-tangle banner
  whenever the primary checkout is on a feature branch, which is exactly what a
  CI pull-request checkout is. The banner lands ahead of the `deferred:` line,
  so the match fell through and a waiting mate was reported as a failed send,
  with the banner as the reason. Both senders now classify on fm-send's exit
  status, which is the contract the deferral is actually stated in, and select
  the `deferred:` line out of the output rather than assuming it came first.

The third root cause, a no-mistakes definition of done that backgrounded the
drive call and polled axi status, is fixed by this branch's parent commit
"drive no-mistakes with one foreground call, not a background poll"; this
commit takes that text as is and adds the regression test.

Upstream's spawn abort path no longer calls the lease-return helper at all,
so the fork's missing-helper guard and its pin-feature test line are moot
here and are not ported.

The command ceilings each harness enforces, and the probes behind the named
Claude Code wait, are recorded in docs/verification/runtime-backends.md.

* no-mistakes(review): Exempt captain holds, quiet deferred reconcile, clarify worker pauses

* no-mistakes(document): Document deferred automatic nudges, rereads, and reply recovery

* no-mistakes(document): Ring unlanded fire-and-forget steers exactly once more

* no-mistakes(ci): The failing check, "PR must be raised via no-mistakes", reads the pipeline's attestation record, which says document=skipped. No file in the repository can change that record, so I did not touch the check or the PR body. As you said, the no-mistakes rerun after this run finishes will re-execute the document step and record document=completed. The one change is the documentation sentence you ordered. It adds a line to docs/remote-secondmates.md, right after the line saying the remote host runs no re-ring ladder of its own: "A fire-and-forget record, such as a reconcile ask, gets its single retry ring only on the local plane: the remote steer leg owes no re-ring, so a swallowed remote doorbell for one waits for the next ring into that inbox, and a remote-side retry is known follow-up scope." No behavior changed. Checks: tests/fm-documentation-audiences.test.sh passes (4/4) and bin/fm-lint.sh is clean. The change is left uncommitted in the working tree for the pipeline to pick up

* no-mistakes(review): Hold automatic wakes until a mate's own decision closes

* no-mistakes(document): Document watcher delivery of deferred remote re-read nudges

* no-mistakes(review): Merge duplicate elapsed-wait reattach instructions in DOD

* no-mistakes(test): Resolve merged default decision in remote-reply recovery fixture

* no-mistakes(test): Source classify lib so config-push retry-deferred honors open decisions

* no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI

* Revert "no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI"

This reverts commit c719928.

* no-mistakes(review): Retry deferred local instruction nudges via the watcher

* no-mistakes(review): Document watcher retry for deferred local instruction nudges

* no-mistakes(ci): I fixed both review findings you selected (ci-1 and ci-3). I did not touch the deferral check in bin/fm-send.sh. ci-1 (bin/fm-config-push.sh, retry_deferred_rereads) - Rule that must hold: a deferred reread stays flagged until it is actually delivered. - Before the fix, the flag was removed before any of the steps that can skip a mate: the remote lock-path lookup, validate_secondmate_home, the local lock-path lookup, and the lock acquire. A skip at any of those dropped the flag, so the watcher lost track of the reread. - Now the flag is removed in one place only, when the send succeeds (rc 0). A skipped home, a busy lock, a deferred send (rc 4) or a failed send all leave it in place. The re-mark calls on a busy lock and on rc 4 were no longer needed, so I removed them. I updated the comment above the function to match. - Side effect: a send that keeps failing now stays flagged, so the watcher retries it on every poll and logs each failure. That follows your "don't clear until delivered" rule, but it replaces the old behaviour of leaving a failed send to the next config push or session start. - New test in tests/fm-secondmate-sync.test.sh: T8j "a deferred flag survives a skipped invalid home and is retried once it validates". It takes the home's marker away to make validation fail, checks that nothing is sent and the flag stays, then puts the marker back and checks that the nudge is delivered and both the flag and the retry marker are cleared. It fails on the old code and passes now. ci-3 (bin/fm-secondmate-restart.sh) - Rule that must hold: no automatic send wakes a mate that is waiting on its own open decision. - The two automatic sends in this script are the fallback reread nudge (fall_back_to_nudge) and the persist request. Both now pass --automatic. If a persist request is deferred, its correlation is discarded and the mate goes to the fallback nudge, which is also deferred, so the mate is reported as unreached. - New test in tests/fm-secondmate-restart.test.sh: T3b. It gives a mate an open needs-decision and runs a restart. It checks that both sends report as deferred, the mate's doorbell is never rung, its inbox gets no message, nothing is stopped, and the mate is reported as unreached with exit status 3. It fails on the old code and passes now. - The test marks the watcher as alive first. Without that, the watcher-down warning is printed first and becomes the reported reason instead of the deferral message. Verification - tests/fm-secondmate-sync.test.sh passes. - tests/fm-secondmate-restart.test.sh passes. - tests/fm-secondmate-harness.test.sh (the other test that exercises --retry-deferred) passes. - The fm-send-inbox test that covers automatic deferral passes. I only looked at the last lines of that run, not the whole file. - `shellcheck -x` on the four changed files is clean

* Pin autoarm supervision model in secondmate restart T3b

The fresh watcher beat the test writes proves a live watcher only under the
autoarm model; on CI hosts with no detected harness the persistent model
demands a lock-holding watcher, so the watcher-down banner became the
reported reason and the deferral assertion failed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep deferred secondmate nudges retryable under the inheritance lock.

A bootstrap instruction nudge could write its deferral flag outside the lock the watcher retry holds, so a concurrent retry could delete a flag that had just been set. A restart fallback that is deferred now records the same marker and flag, so the watcher delivers it once the decision closes.

* no-mistakes(document): Document watcher retry of deferred restart re-read nudges

* Send secondmate reread and restart nudges immediately again.

Deferring those nudges let a later config push drop an incomplete transfer once the decision closed. They now send as they do on main.

* Make the no-turn wait opt-in behind config/wait-no-turns.

Homes that do not create the file keep the previous briefs, drive text, and sends.

* no-mistakes(document): Document wait-no-turns inbox wording change in configuration

* no-mistakes(review): Keep checkpoint inbox check; forbid only polling while waiting

* no-mistakes(ci): Fixed ci-2 (Greptile: a concurrent retry marker gets lost). The rule that was broken: the watcher may remove only the `.retry-ring` mark for the record it just processed. A newer mark written in the meantime is owed its own retry. `fm_task_inbox_clear_retry` is the one shared function that removes the mark, and I fixed it there. In `bin/fm-task-inbox-lib.sh` it now takes the record path. It compares the mark's content with that record's name and removes the mark only when they match. When the mark names a different record it returns success and leaves the mark alone. It still fails only when the processed record's own mark can't be removed. Both callers in `bin/fm-watch.sh` now pass `"$rec"`: the dead or missing pane path and the path after a retry ring. So the fix holds at both removal sites. Tests, in `tests/fm-task-inbox.test.sh`: - I added an optional `FM_RING_MARKS_RETRY` hook to the fake tmux. It writes a newer record's mark while the doorbell is being typed, which reproduces the race deterministically. - I added `test_watcher_retry_keeps_a_newer_mark`. The owed retry rings once, the newer mark survives, and a later check rings the newer record once and then clears its mark. The test fails without the fix ("the spent retry removed a newer record's mark written during its ring") and passes with it. - I updated the direct `clear_retry` call in the existing unit test to pass the record. Results: `tests/fm-task-inbox.test.sh` passes in full and `tests/fm-send-inbox.test.sh` passes 15/15. Shellcheck reports only SC1091 "not following sourced file" notices. As instructed, I didn't change the brief inbox wording

* no-mistakes(document): Fix stale wait-no-turns inbox wording in inbox lib comment

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
… reply replaced the live listener’s registration, which could prevent the next reply from being captured. Added a regression check that failed before the fix. The full remote reply test and lint now pass
@peterOC26
peterOC26 merged commit 470475c into main Oct 1, 2026
19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.