Skip to content

feat(bin): add the quality gate loop and the hardened project posture - #19

Merged
BohnBawerick merged 34 commits into
mainfrom
fm/fm-quality-loop
Aug 30, 2026
Merged

BohnBawerick merged 34 commits into
mainfrom
fm/fm-quality-loop

Conversation

@BohnBawerick

@BohnBawerick BohnBawerick commented Aug 22, 2026 •

Copy link
Copy Markdown
Owner

Intent

Build the quality loop controller and its state reconciliation: deliverables D4, D8 and D10 from the parked Superchecks specification (implementation-thoughts.md), with read-only scoring folded in as a first-class mode rather than bolted on afterwards.

Build on what already landed. D1/D2 were proven on real TypeScript by the Stage 0a pilot on quota-axi. The D2 receipt schema was just revised and landed (PR #17): it carries head_sha, per-phase findings with stable ids, engine name and version, the threshold the outcome was judged against, a required duration_ms, and a wall-clock bound. Implement against the landed schema, not the older description of it.

The wall-clock number this task was blocked on: a diff-scoped measurement costs 20 seconds to 4 minutes with the right engine, roughly linear in changed lines, so four loop iterations cost minutes and the loop is worth building. Design to that number and make the loop's own wall-clock bound enforce it rather than trusting it. The pilot also produced a warning to respect: on the wrong engine the same work took 1h47m AND was wrong - it counted 15 crashed test processes as mutant kills, and on the smallest diff reported a perfect 100% kill rate that was entirely an artifact of one timing-sensitive test failing under load. A flattering silent failure is the dangerous one here, so the loop must be able to tell "measured and passed" apart from "could not measure".

Read-only scoring, per the captain on 2026-08-22 ("even if the mode is not active I think it should show the scores as read only yes"): measurement and reporting are always on wherever they can run, and only blocking is gated on the hardened mode. A standard project surfaces its mutation and complexity scores so the numbers exist before anyone commits to a threshold, which is also how a project earns an informed decision to switch the mode on. The receipt must carry a read-only outcome distinctly from a pass, so a score that merely reported is never mistaken for a gate that ran and approved; a read-only outcome that looks like a pass is worse than no read-only mode at all.

Acceptance criteria:

  • D4, D8 and D10 are implemented against the landed schema, and the loop's six outcomes from the specification's section 2 (pass, blocked, not-applicable, exhausted, stuck, defect-found) are each reachable and each covered by a test.
  • Read-only mode is a first-class outcome in the receipt, distinguishable from a pass by a machine and not only by a human reading prose.
  • The loop enforces its own wall-clock bound rather than assuming the measurement is fast.
  • "Could not measure" never reports as a pass, and there is a test that fails if that distinction is removed.
  • Each test is falsifiable: it must fail if the thing it pins is replaced by a constant.

Constraints and deliberate exclusions:

  • Keep it simple; no over-engineering. Add machinery only when a concrete blocker demands it.
  • Do NOT turn any quality bar on for the firstmate repository itself. This work builds the capability; firstmate is not a project it gets applied to.
  • Do NOT touch the threshold numbers. kill_rate_min is settled at 0.80 for hardened projects and is not retrofitted to existing repositories.
  • Do NOT rework the receipt schema; it just landed. If a genuine defect is found in it, report it rather than changing it under this task.
  • Do NOT rewrite the hand-rolled JSON Schema interpreter in bin/fm-quality-receipt.sh; whether it survives at all is a separate open question tracked as fm-receipt-schema-validator. Consume it as it is.
  • bin/fm-crew-state.sh is in scope for D8, but it is a live supervision dependency for the whole fleet: do not change its existing output shape for states that have nothing to do with quality.
  • Two other workers are changing this repo concurrently, on session-lock ownership and the away-mode delivery path. Stay in the quality surface and do not edit bin/fm-watch.sh, bin/fm-supervise-daemon.sh, bin/fm-composer-lib.sh, bin/fm-claude-stop-autoarm.sh, or bin/fm-turnend-guard.sh.
  • Load and follow .agents/skills/firstmate-coding-guidelines/SKILL.md, which owns knowledge placement, the one-owner rule, AGENTS.md size discipline, and this repo's style rules.

One specification item was implemented differently, deliberately: the parked spec named a single data//quality-receipt.json holding both phases, but the landed schema requires a verify envelope's children to share its head_sha, and the harden loop commits after the clean receipt is written, so the two phases never share a head. Each phase therefore gets its own complete receipt file.

What Changed

  • New bin/fm-quality.sh runs a project's own quality commands for a phase (clean or harden) under the contract's round and wall-clock bounds, checks the ordinary test suite first so a score is never measured over a red or flaky suite, and reports one outcome per run with a fixed exit code: pass 0, exhausted/stuck 1, blocked 3, not-applicable 4, defect-found 5, read-only 6. A run that could not measure reports blocked, never pass, and a standard project's read-only run measures and reports scores without blocking anything. Each phase writes its own complete D2 receipt to data/<id>/quality-<phase>-receipt.json, stamped with the exclude list the run actually measured against, and status answers with one token (satisfied, satisfied-stale, missing, not-passed, unreadable, not-required).
  • New bin/fm-quality-receipt.sh validates a receipt against the new docs/quality-receipt.schema.json, adds the post-schema rules JSON Schema cannot state (unique finding ids, verify children sharing the envelope's base_sha/head_sha), and can require head_sha to resolve to a given tree's HEAD. docs/quality-gate.md owns the contract and bounds; docs/architecture.md, docs/scripts.md, README.md, AGENTS.md, and the project-management skill record the posture.
  • Posture plumbing across the existing fleet scripts: data/projects.md gains a position-independent +hardened bracket token that fm-project-mode.sh --quality resolves to one word (refused beside the conditional no-mistakes-prod-only policy); fm-brief.sh --quality emits a Quality contract: quality=hardened line plus one quality-gate section, leaving a standard brief byte-identical; fm-spawn.sh --quality records quality= and a once-captured base_sha= in the task meta, reads both back on relaunch, and refuses a brief that disagrees; fm-crew-state.sh routes a done verdict on a hardened task through the quality status verdict and leaves every other posture's line unchanged; fm-promote.sh prints an advisory notice when a promoted task carries less rigor than the project's standing posture. New tests/fm-quality.test.sh, tests/fm-quality-receipt.test.sh, and a live structured-output drift check cover the outcomes and the receipt rules, with the brief, crew-state, relaunch, and task-delivery suites extended and every new suite registered in bin/fm-test-run.sh.

Not wired in

Read-only scoring is implemented and covered by tests, but no firstmate path invokes it yet, so a standard project still surfaces no scores in practice. bin/fm-brief.sh:519 and bin/fm-brief.sh:603 both gate the quality contract line and the quality-gate section on [ "$QUALITY" = hardened ], so a standard brief stays byte-identical to what it was before this change and never tells its worker to run the loop.

Turning it on means relaxing those two guards for the read-only case: emit a quality section on a standard brief whose project has a .quality-gate.yaml, telling the worker to run bin/fm-quality.sh run <id> --phase clean and report the scores, while leaving blocking gated on quality=hardened exactly as it is now. That would change what every ship task on every project does, which is a wider contract than this task was scoped to, so it is left out deliberately rather than by oversight.

Risk Assessment

⚠️ Medium: The change adds a large new capability (about 5300 lines across a new loop controller, a receipt validator, and a live fleet-supervision script) and has already needed six fix rounds, but the standard path stays byte-identical, every new behavior is gated on the hardened posture, and the newest sha-comparison fix is correct at all three call sites with falsifiable test pairs, so it is safe to merge with the one narrow path-quoting issue as a follow-up.

Testing

Ran the six targeted test scripts covering the quality loop, receipt validator, crew state, brief, relaunch and delivery paths - all passed with no failures. Beyond the suites, I drove the real fm-quality.sh CLI over throwaway git projects to show each of the six specification outcomes with its distinct exit code, showed the read-only receipt carrying "outcome": "read-only" against a byte-identical measurement whose hardened twin records "pass", showed a missing engine reporting blocked with no receipt written at all, and showed a 300-second engine cut off after 65 seconds by a 1-minute wall-clock budget. I also drove fm-crew-state.sh with an identical passing pipeline run against six receipt states to show the supervisor line a fleet operator actually reads, where a read-only receipt reports blocked rather than done and a standard task is untouched. Finally I applied three mutations (read-only to pass, could-not-measure to pass, constant wall-clock budget) and confirmed each turns the owning test red, so the tests are not satisfiable by a constant. This change is a bash CLI with no rendered UI surface, so the reviewer-visible artifacts are CLI transcripts and receipt JSON rather than screenshots.

Evidence: CLI transcript: all six loop outcomes, receipt JSON, and wall-clock enforcement

Source: CLI transcript: all six loop outcomes, receipt JSON, and wall-clock enforcement


=========================================================
1/8  OUTCOME: pass  (hardened project, bar met -> exit 0)
=========================================================
$ fm-quality.sh run task --phase clean      # quality=hardened
outcome: pass · phase: clean · mode: hardened · threshold met in 1 round(s)
exit code: 0

--- data/task/quality-clean-receipt.json ---
{
    "base_sha": "aeebc72df0434ec1dd3ab9822874b85253bef2c4",
    "duration_ms": 639,
    "engine": {
        "name": "stryker-demo",
        "version": "8.2.6"
    },
    "exclusions": [],
    "findings": [],
    "head_sha": "aeebc72df0434ec1dd3ab9822874b85253bef2c4",
    "metrics": {
        "crap_observed": 0.0,
        "quality_loop_rounds": 1
    },
    "outcome": "pass",
    "phase": "clean",
    "schema_version": 1,
    "threshold": {
        "crap_max": 15
    }
}

=========================================================
2/8  OUTCOME: read-only  (standard project, SAME bar met -> exit 6, NOT pass)
=========================================================
$ fm-quality.sh run task --phase clean      # quality=standard
outcome: read-only · phase: clean · mode: read-only · scores reported, nothing gated (measured: pass)
exit code: 6

--- data/task/quality-clean-receipt.json ---
{
    "base_sha": "8f431b8014c86bfcf6f0006791dec16c332f8eb6",
    "duration_ms": 558,
    "engine": {
        "name": "stryker-demo",
        "version": "8.2.6"
    },
    "exclusions": [],
    "findings": [],
    "head_sha": "8f431b8014c86bfcf6f0006791dec16c332f8eb6",
    "metrics": {
        "crap_observed": 0.0,
        "quality_loop_rounds": 1
    },
    "notes": "Read-only run: the measurement came back pass and nothing was gated.",
    "outcome": "read-only",
    "phase": "clean",
    "schema_version": 1,
    "threshold": {
        "crap_max": 15
    }
}

Machine-readable difference between the two runs above:
  hardened receipt outcome = 'pass'
  standard receipt outcome = 'read-only'
  same measurement?         True
  read-only mistakable for pass? False

=========================================================
3/8  OUTCOME: blocked  (could not measure -> exit 3, never pass)
=========================================================
$ fm-quality.sh run task --phase clean      # engine binary missing
outcome: blocked · phase: clean · mode: hardened · the clean command could not be run here (exit 127)
exit code: 3
receipt written? no - nothing was measured, so nothing is proved

=========================================================
4/8  OUTCOME: not-applicable  (no .quality-gate.yaml -> exit 4)
=========================================================
$ fm-quality.sh run task --phase clean      # project has no contract
outcome: not-applicable · phase: clean · mode: hardened · this project has no .quality-gate.yaml, so there is no quality contract to measure
exit code: 4

=========================================================
5/8  OUTCOME: exhausted  (rounds run out below the bar -> exit 1)
=========================================================
$ fm-quality.sh run task --phase clean      # findings shrink but never reach the bar
outcome: exhausted · phase: clean · mode: hardened · 3 rounds ran out below the threshold
exit code: 1
receipt outcome: exhausted

=========================================================
6/8  OUTCOME: stuck  (same finding ids round after round -> exit 1)
=========================================================
$ fm-quality.sh run task --phase clean      # identical finding ids every round
outcome: stuck · phase: clean · mode: hardened · 2 consecutive rounds left the same findings untouched
exit code: 1
receipt outcome: stuck

=========================================================
7/8  OUTCOME: defect-found  (a survivor is a real product bug -> exit 5)
=========================================================
$ fm-quality.sh run task --phase clean      # engine reports a product defect
outcome: defect-found · phase: clean · mode: hardened · the clean phase reported a real product defect
exit code: 5
receipt outcome: defect-found

=========================================================
8/8  WALL CLOCK IS ENFORCED  (engine sleeps 300s, budget is 1 minute)
=========================================================
$ fm-quality.sh run task --phase clean      # bounds.budget_minutes: 1
outcome: exhausted · phase: clean · mode: hardened · the clean command did not finish inside the 1-minute wall-clock bound
exit code: 1
the engine asked for 300s; the loop cut it off and returned after 65s

=========================================================
STATUS: what the rest of the pipeline asks
=========================================================
$ fm-quality.sh status task    # after the hardened pass above
quality: satisfied · mode: hardened · every configured phase has a passing receipt
exit code: 0
$ fm-quality.sh status task    # after the read-only run above
quality: not-required · mode: read-only · this task ships standard, so no receipt gates it
exit code: 0
$ fm-quality.sh status task    # after the blocked run above
quality: missing · mode: hardened · no receipt for: clean
exit code: 0
Evidence: CLI transcript: what the fleet supervisor sees for each receipt state

Source: CLI transcript: what the fleet supervisor sees for each receipt state

The pipeline run says "passed" in every case below. Only the quality receipt differs. A: quality=hardened, no receipt (the gate never ran) $ fm-crew-state.sh q state: blocked · source: run-step · run passed: PR merged/closed · quality gate never ran: no receipt for: clean B: quality=hardened, receipt outcome=pass $ fm-crew-state.sh q state: done · source: run-step · run passed: PR merged/closed C: quality=hardened, receipt outcome=exhausted (measured, below the bar) $ fm-crew-state.sh q state: blocked · source: run-step · run passed: PR merged/closed · quality gate did not pass: the clean phase reported exhausted D: quality=hardened, receipt outcome=read-only (reported, never gated) $ fm-crew-state.sh q state: blocked · source: run-step · run passed: PR merged/closed · quality gate did not pass: the clean phase reported read-only E: quality=hardened, receipt file corrupt (cannot tell) $ fm-crew-state.sh q state: unknown · source: run-step · run passed: PR merged/closed · quality gate unreadable: the clean receipt is present but not a valid receipt F: quality=standard, no receipt (an ordinary task, untouched) $ fm-crew-state.sh q state: done · source: run-step · run passed: PR merged/closed

The pipeline run says "passed" in every case below.
Only the quality receipt differs.

A: quality=hardened, no receipt (the gate never ran)
$ fm-crew-state.sh q
state: blocked · source: run-step · run passed: PR merged/closed · quality gate never ran: no receipt for: clean

B: quality=hardened, receipt outcome=pass
$ fm-crew-state.sh q
state: done · source: run-step · run passed: PR merged/closed

C: quality=hardened, receipt outcome=exhausted (measured, below the bar)
$ fm-crew-state.sh q
state: blocked · source: run-step · run passed: PR merged/closed · quality gate did not pass: the clean phase reported exhausted

D: quality=hardened, receipt outcome=read-only (reported, never gated)
$ fm-crew-state.sh q
state: blocked · source: run-step · run passed: PR merged/closed · quality gate did not pass: the clean phase reported read-only

E: quality=hardened, receipt file corrupt (cannot tell)
$ fm-crew-state.sh q
state: unknown · source: run-step · run passed: PR merged/closed · quality gate unreadable: the clean receipt is present but not a valid receipt

F: quality=standard, no receipt (an ordinary task, untouched)
$ fm-crew-state.sh q
state: done · source: run-step · run passed: PR merged/closed
Evidence: Falsifiability: three mutations, each caught by the test that pins it

Source: Falsifiability: three mutations, each caught by the test that pins it

MUTATION 1 - a read-only run reports 'pass' instead of 'read-only' - write_final "$last_receipt" read-only "$round" &#10;+ write_final "$last_receipt" pass "$round" &#10;not ok - a read-only run must not exit with the pass code: expected exit 6, got 0 MUTATION 2 - 'could not measure' reports 'pass' instead of 'blocked' - no-command|no-receipt|invalid|drift) finish blocked "$MEASURE_DETAIL" ;; + no-command|no-receipt|invalid|drift) finish pass "$MEASURE_DETAIL" ;; not ok - a read-only run that could not measure is blocked: expected exit 3, got 0 MUTATION 3 - the wall-clock budget is a constant 86400s (never enforced) - printf '%s' "$(( BUDGET_SECONDS - used ))" + printf '%s' 86400 not ok - a measurement that outruns the bound exits 1: expected exit 1, got 3 UNMUTATED source, same suite all fm-quality tests passed


=========================================================
MUTATION 1 - a read-only run reports 'pass' instead of 'read-only'
=========================================================
mutation applied:
  -      write_final "$last_receipt" read-only "$round" \
  +      write_final "$last_receipt" pass "$round" \

test result:
  ok - a hardened task that meets its threshold reports pass and writes a receipt
  not ok - a read-only run must not exit with the pass code: expected exit 6, got 0

=========================================================
MUTATION 2 - 'could not measure' reports 'pass' instead of 'blocked'
=========================================================
mutation applied:
  -      no-command|no-receipt|invalid|drift) finish blocked "$MEASURE_DETAIL" ;;
  +      no-command|no-receipt|invalid|drift) finish pass "$MEASURE_DETAIL" ;;

test result:
  ok - a hardened task that meets its threshold reports pass and writes a receipt
  ok - read-only reports the score with its own outcome and exit code, never pass
  not ok - a read-only run that could not measure is blocked: expected exit 3, got 0

=========================================================
MUTATION 3 - the wall-clock budget is a constant 86400s (never enforced)
=========================================================
mutation applied:
  -  printf '%s' "$(( BUDGET_SECONDS - used ))"
  +  printf '%s' 86400

test result:
  ok - a phase below threshold with progress every round reports exhausted
  ok - rounds that leave the same findings untouched report stuck, not exhausted
  not ok - a measurement that outruns the bound exits 1: expected exit 1, got 3

=========================================================
UNMUTATED source, same suite
=========================================================
  ok - usage errors exit 2 and --help exits 0
  all fm-quality tests passed
Evidence: Read-only is machine-distinguishable from pass (same measurement, two receipts)
$ fm-quality.sh run task --phase clean # quality=hardened
outcome: pass · phase: clean · mode: hardened · threshold met in 1 round(s)
exit code: 0
"outcome": "pass",
"threshold": {"crap_max": 15},
"metrics": {"crap_observed": 0.0, "quality_loop_rounds": 1}

$ fm-quality.sh run task --phase clean # quality=standard
outcome: read-only · phase: clean · mode: read-only · scores reported, nothing gated (measured: pass)
exit code: 6
"outcome": "read-only",
"notes": "Read-only run: the measurement came back pass and nothing was gated.",
"threshold": {"crap_max": 15},
"metrics": {"crap_observed": 0.0, "quality_loop_rounds": 1}

Machine-readable difference between the two runs above:
hardened receipt outcome = 'pass'
standard receipt outcome = 'read-only'
same measurement? True
read-only mistakable for pass? False
Evidence: Wall-clock bound is enforced, not assumed
8/8 WALL CLOCK IS ENFORCED (engine sleeps 300s, budget is 1 minute)
$ fm-quality.sh run task --phase clean # bounds.budget_minutes: 1
outcome: exhausted · phase: clean · mode: hardened · the clean command did not finish inside the 1-minute wall-clock bound
exit code: 1
the engine asked for 300s; the loop cut it off and returned after 65s

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

⚠️ **Rebase** - 1 warning
  • ⚠️ .agents/skills/project-management/SKILL.md - branch carries 8 commit(s) that exist on your local main branch but were never pushed to origin/main; rebasing would bundle this unrelated work (18 file(s)) into the PR:
  • bb124a1 feat: revise D2 quality receipts and add a wall-clock bound
  • a7ca549 no-mistakes(document): document hardened quality contract in architecture and scripts inventory
  • 1875879 no-mistakes(review): document hardened token in architecture and README grammar
  • e3d0c65 no-mistakes(review): document hardened registration, refuse it on conditional policy
  • 32f566d no-mistakes(review): add promote quality notice, fix brief help, pin base capture
  • aaa2272 no-mistakes(review): fix test suite teardown, header, typo-fallback coverage
  • 9bb7d0c no-mistakes(review): add quality standing-posture notice, correct fallback header
  • 16b36f3 feat: plumb the registered hardened quality posture through to the task record

Push main to origin, or rebase your branch onto origin/main, before gating.

⚠️ **Review** - 1 info
  • 🚨 bin/fm-quality.sh:860 - run --dry-run deletes the phase's existing receipt before the dry-run branch is reached, then prints "dry-run: nothing was run and no receipt was written". The rm -f -- &#34;$(receipt_path &#34;$PHASE&#34;)&#34; at line 860 runs unconditionally; the if [ &#34;$DRY_RUN&#34; -eq 1 ] early exit is at line 892. Reproduced: a task with a valid passing data/&lt;id&gt;/quality-clean-receipt.json, then fm-quality.sh run task --phase clean --dry-run -> the file is gone, exit 0, and the output claims nothing was written. On a hardened task bin/fm-crew-state.sh then rewrites done to blocked via quality_gate_override ("quality gate never ran"), so a read-only inspection flag silently destroys durable proof and flips fleet supervision. Fix: move the rm -f to after the dry-run branch (or guard it with [ &#34;$DRY_RUN&#34; -eq 0 ]).
  • ⚠️ bin/fm-quality.sh:701 - The contract is loaded once at line 862 (state=$(load_contract &#34;$WT&#34;)) and never reloaded inside the round loop, so FM_QUALITY_EXCLUDE=&#34;$(cfg_seq &#34;$PHASE.exclude&#34;)&#34; at line 701 replays a pre-round-1 snapshot for every measurement. But write_turn_prompt explicitly instructs the agent to "add the narrowest possible entry to the exclude list in .quality-gate.yaml" and answer {&#34;action&#34;:&#34;excluded&#34;}, and the loop commits that edit. Reproduced with a stub harness that appends - &#34;a.ts&#34; to the exclude list: all 3 rounds received the unchanged src/gen/**, the excluded finding reappeared every round, and the phase reported exhausted. One of the four documented agent actions is a no-op within a run, and the loop reports exhausted/stuck where the exclusion should have produced pass. Fix: re-run load_contract/load_bounds (or at least re-read $PHASE.exclude) at the top of each round, refusing if the reparse now errors.
  • ⚠️ bin/fm-quality.sh:754 - check_suite passes its original $secs to the second (flaky-detection) run instead of recomputing remaining_seconds, so the pre-flight suite check can consume up to 2x bounds.budget_minutes before the first measurement even starts. Reproduced with test: &#34;sleep 5; exit 1&#34; and budget_minutes: 0.1 (6s bound): the run took 11s wall clock. With the default 20-minute bound a red slow suite can spend ~40 minutes, which is the assumption the intent's "the loop enforces its own wall-clock bound rather than assuming the measurement is fast" criterion exists to remove. Fix: recompute the remaining budget before the second bounded_sh and report exhausted if none is left.
  • ⚠️ bin/fm-quality.sh:730 - measure validates the receipt against the schema and checks base_sha and (via --check-head) head_sha, but never checks that the receipt's phase equals $PHASE. Reproduced: a contract whose clean.command prints &#34;phase&#34;:&#34;harden&#34; with classification: killable findings is accepted, and data/&lt;id&gt;/quality-clean-receipt.json is written recording phase: harden with killable findings. cmd_status then reads that file as the clean phase's passing receipt. This bypasses the clean/harden classification split the schema exists to enforce (docs/quality-gate.md: "A clean finding classified killable is invalid, which is the point of splitting the vocabulary"), and a misconfigured project can satisfy the clean gate with a harden-engine receipt without any error. Fix: after the base_sha check, require receipt_field &#34;$out&#34; phase to equal $PHASE and report blocked otherwise.
  • ⚠️ bin/fm-quality.sh:416 - The perl branch of bounded_sh ends with exit($? &gt;&gt; 8), which discards the signal byte: a child killed by a signal yields exit 0. Confirmed directly: perl -e &#39;...exit($? &gt;&gt; 8)&#39; 10 sh -c &#39;kill -SEGV $$&#39; returns 0 where timeout would return 139. On a host with neither timeout nor gtimeout (the only case that selects this branch), check_suite reads rc=0 and prints ok, so a test suite that segfaults or is OOM-killed is scored as green and the mutation/complexity number is trusted; the same path makes the post-agent-turn suite check commit a round whose suite crashed. That is the pilot's documented flattering-silent-failure mode the loop is meant to exclude. Fix: exit($? &amp; 127 ? 128 + ($? &amp; 127) : $? &gt;&gt; 8). The identical expression is copied from bin/fm-nm-run-lib.sh:28, which has the same defect.
  • ⚠️ bin/fm-quality.sh:955 - if [ -n &#34;$prev_ids&#34; ] &amp;&amp; [ &#34;$ids&#34; = &#34;$prev_ids&#34; ] conflates "no previous round yet" with "the previous round had an empty finding set", so no_progress never increments when a below-threshold receipt reports findings: []. Reproduced with a phase command that always returns outcome: exhausted and findings: [] under max_iterations: 4, no_progress_limit: 2: the loop reported exhausted after 4 rounds and 3 paid agent turns instead of stuck after 3 rounds and 2 turns. The header claims "'No progress' means the same finding ids, not the same count" - an unchanged empty id set is the clearest no-progress case and is the one it misses. Fix: track a separate have_prev flag (or use an unset sentinel) instead of testing -n &#34;$prev_ids&#34;.

🔧 Fix: fix six quality-loop defects and pin them with tests
4 issues (2 errors, 2 warnings) still open:

  • 🚨 bin/fm-quality.sh:904 - run --dry-run still produces supervision side effects. The fix round guarded the receipt deletion at line 877, but the dry-run early exit is at line 909 and six finish calls run before it: no base commit (890), unresolvable base (893), no timeout tool (896), structured-output refusal (902), the dirty-tree check (904), and both not-applicable paths (881, 887). finish calls status_append, which on a hardened task appends to $STATE/$ID.status. Concrete path: a hardened ship task whose worktree has one uncommitted file, then fm-quality.sh run &lt;id&gt; --phase clean --dry-run -> the line blocked: quality clean could not measure, the task copy has uncommitted changes; the loop commits each round, ... is appended to the status log, and exit is 3, not 0. bin/fm-classify-lib.sh (FM_CLASSIFY_CAPTAIN_RE_DEFAULT, line 49, and the done|needs-decision|blocked|failed sets at lines 95/117/1032) treats a blocked: line as a decision-opening supervision event, and bin/fm-crew-state.sh emits it as state: blocked · source: status-log. A hardened task is dirty for most of its working life, so an inspection-only flag routinely files a real blocked event against a healthy crew. This is the same class as the receipt-deletion defect just fixed, on the sibling path. Fix: make status_append return early when DRY_RUN is 1 ([ &#34;$DRY_RUN&#34; -eq 0 ] || return 0 at line 512), which leaves the printed dry-run report unchanged.
  • 🚨 bin/fm-quality.sh:940 - The per-round contract re-read added by the fix round re-reads the phase threshold as well as the exclude list, so the agent being gated can widen its own bar and turn a below-threshold phase into pass. load_contract at line 941 rewrites the whole $CFG file; measure then rebuilds FM_QUALITY_THRESHOLD from it at line 700 on every round. Concrete sequence: contract has clean.threshold.crap_max: 15; round 1 measures below threshold and reports exhausted with findings; the agent turn (running under --dangerously-skip-permissions / --sandbox danger-full-access in the worktree) edits .quality-gate.yaml to crap_max: 9999; the ordinary suite is still green, so lines 1028-1032 git add -A and commit the edit; round 2 re-reads the contract, exports FM_QUALITY_THRESHOLD=crap_max=9999, the phase command judges against it and prints outcome: pass; write_final at line 981 writes a passing receipt and exits 0. cmd_status (line 826) and quality_gate_override in bin/fm-crew-state.sh both accept that as a gate that ran and approved. Before this commit the contract was loaded once, so this was structurally impossible; the round-loop comment at lines 938-939 states the invariant that should cover it ("a run does not get to extend its own wall clock or swap the command it is being measured by") but the threshold is not among the pinned values, and the turn prompt's "do not change the threshold" (line 673) is the only thing enforcing it. Fix: capture FM_QUALITY_THRESHOLD from the round-1 contract into a variable before the loop and pass that pinned value to measure, or refuse with blocked when the re-read threshold differs from round 1's.
  • ⚠️ bin/fm-quality.sh:763 - When the wall-clock bound runs out before the flake re-run, check_suite reports exhausted (exit 1) instead of blocked (exit 3), but no information is actually missing: the first run already returned non-zero and was not a timeout (rc 124 is handled at line 754), and both possible answers from the re-run are blocked - red twice is blocked (line 775) and red then green is blocked (line 773). So the correct outcome is knowably blocked regardless of what the re-run would have said. Concrete state: test: &#34;...&#34; that fails after consuming the whole budget_minutes (for example a 6s bound with a suite that fails at 6.1s) -> remaining_seconds returns <= 0, cmd_run's new exhausted*) branch at line 927 calls finish exhausted, and the caller sees exit 1 ("ran out below threshold") for what the script's own outcome table at line 37 assigns to blocked ("could not measure: toolchain, red or flaky suite, drift"). No pass is manufactured, so this is a wrong label rather than a wrong verdict. Fix: emit blocked the ordinary test suite is red before any quality work and the bound ran out before red and flaky could be told apart and drop the exhausted*) branch at line 927.
  • ⚠️ bin/fm-quality.sh:369 - bounds.budget_usd is the one bound load_bounds reads without validating it. max_iterations and no_progress_limit go through positive_integer (line 372) and budget_minutes through positive_number (line 374), but budget_usd is copied straight from the contract at line 369 and interpolated unquoted into the claude command string at line 576 (--max-budget-usd $BOUND_BUDGET_USD), which is then run through sh -c by bounded_sh. Concrete input: .quality-gate.yaml with bounds: { budget_usd: unlimited } -> round 1 builds claude ... --max-budget-usd unlimited ..., claude exits with a usage error, agent_turn swallows it (|| true, stderr to /dev/null) and prints nothing, so turn_action is empty and the loop scores the round as no progress. Every round then fails identically and the phase reports stuck or exhausted with no mention of the misconfigured bound. Separately, the codex branch (lines 583-591) never passes budget_usd at all, so a hardened run on codex has no spend ceiling even though the contract declares one and the dry-run report at line 912 prints it as an enforced bound. Fix: validate budget_usd with positive_number alongside the others, and either pass it to codex or say in the header that it applies to claude only.

🔧 Fix: pin threshold, quiet dry runs, validate budget_usd
3 warnings still open:

  • ⚠️ bin/fm-brief.sh:603 - Read-only scoring is implemented but unreachable in practice: no firstmate path ever runs it. The intent requires "measurement and reporting are always on wherever they can run" and "A standard project surfaces its mutation and complexity scores so the numbers exist before anyone commits to a threshold, which is also how a project earns an informed decision to switch the mode on." The contradicting hunk is if [ &#34;$QUALITY&#34; = hardened ]; then DOD=&#34;$QUALITY_SECTION\n\n$DOD&#34; at bin/fm-brief.sh:603 (and the sibling CONTRACT_LINES guard at line 519): a standard brief carries no quality-gate section, so its worker is never told to run fm-quality.sh run &lt;id&gt; --phase clean. The only other caller of the script is quality_gate_override in bin/fm-crew-state.sh:121, and it invokes status, not run; cmd_status in bin/fm-quality.sh:788 short-circuits every non-hardened task with quality: not-required · ... no receipt gates it before reading any receipt. Concrete path: a standard project that commits a valid .quality-gate.yaml with clean.command and harden.command, spawned with the default --quality standard. fm-quality.sh run is never executed, data/&lt;id&gt;/quality-clean-receipt.json is never written, and no score is ever reported to the captain, so the project can never earn the informed decision to switch to hardened. The MODE=read-only branch (bin/fm-quality.sh:191, 980-986) and its tests exercise the capability only when a human types the command by hand. This challenges the deliberate choice at bin/fm-brief.sh:598-608 to keep a standard brief byte-identical, so it needs the author's decision rather than a reviewer patch.
  • ⚠️ bin/fm-quality.sh:586 - An agent turn that never ran is indistinguishable from an agent that ran and changed nothing, and the loop labels the first as stuck. agent_turn discards both the exit status and stderr (bounded_sh &#34;$secs&#34; &#34;$WT&#34; &#34;$cmd&#34; &gt; &#34;$out&#34; 2&gt;/dev/null || true at line 586 for claude, line 595 for codex), so nothing downstream can tell a failed harness invocation from a completed turn. Concrete state: a hardened task whose claude credentials have expired (or whose --max-budget-usd flag was renamed - check_structured_output at line 547 only greps for --json-schema, never for the spend flag it also interpolates at line 581). claude writes its error to stderr and nothing to stdout, $out stays empty, the tolerant reader at line 603 prints nothing, and turn_action is empty. The worktree is unchanged, so the post-turn suite passes (rc=0), git add -A finds nothing to commit, and the next round measures the identical finding set. no_progress increments at line 995 on every round until line 1001 fires write_final ... stuck with the detail "$BOUND_NO_PROGRESS consecutive rounds left the same findings untouched" - a statement about the agent's work that is false, since no agent turn happened at all. The phase then appends blocked: quality clean stuck, ... to the status log (line 532), pointing a supervisor at the code instead of at the broken harness, after spending the full wall-clock and round budget. No pass is manufactured, so this is a wrong label rather than a wrong verdict. Fix: capture bounded_sh's rc in agent_turn and let the caller distinguish "the harness exited non-zero and produced no readable answer" from "the agent answered no-change", reporting the former as blocked naming the harness and its version, the way check_structured_output already does.
  • ⚠️ bin/fm-quality.sh:706 - The per-round contract re-read added in the first fix round made mid-run exclusions live, and the second fix round closed the threshold route to a manufactured pass, but the exclude route is equally powerful and leaves no trace in the durable proof. Concrete sequence: contract has clean.threshold.crap_max: 15 and clean.exclude: [&#34;src/generated/**&#34;]; round 1 measures below threshold; the agent turn (running under --dangerously-skip-permissions in the worktree) appends - &#34;src/**&#34; to clean.exclude and answers {&#34;action&#34;:&#34;excluded&#34;}; the ordinary suite is still green so lines 1036-1040 git add -A and commit the edit; round 2 re-reads the contract at line 947, the threshold comparison at line 954 passes untouched, FM_QUALITY_EXCLUDE at line 706 now carries src/**, the phase command finds nothing left to measure and prints outcome: pass, and write_final at line 989 writes a passing receipt and exits 0. cmd_status at line 830 accepts that as satisfied and quality_gate_override in bin/fm-crew-state.sh reports done. The loop knows the effective exclude list at line 706 but finalize_receipt (lines 473-495) never stamps it, and the schema's optional exclusions array (docs/quality-receipt.schema.json:32) is left to whatever the engine chooses to emit, so a pass earned by excluding the whole diff is byte-indistinguishable in the receipt from a pass earned by writing tests. Fix: pass the effective cfg_seq &#34;$PHASE.exclude&#34; into finalize_receipt and write it to the receipt's exclusions field, so the durable proof records the bar the run actually measured against. Whether the exclusion power itself should also be bounded is a separate product decision.

🔧 Fix: block failed agent turns, stamp exclusions on receipts
2 issues (1 warning, 1 info) still open:

  • ⚠️ bin/fm-quality.sh:1040 - The fix round's new "harness could not run" branch exits the loop without reverting the round, breaking the invariant the header states at lines 76-77 ("Red reverts the round, so a failed round costs budget and nothing else; green commits it"). round_head is captured at line 1033, but the branch at lines 1040-1044 calls write_final ... blocked, which calls finish, which exits. The only git reset --hard --quiet &#34;$round_head&#34; / git clean -fdq pair is at lines 1058-1059, reached only after the post-turn suite check, which this branch skips entirely. cleanup() (line 126) removes $WORK_DIR and nothing else. Concrete sequence: a hardened ship task, contract with clean.command and test; round 1 measures below threshold; the claude turn (running with --dangerously-skip-permissions in $WT) edits src/a.ts, then the CLI exits non-zero with nothing readable on stdout (spend ceiling tripped mid-turn, a crash, a session-store error on the round-2 --continue); turn_action is empty and turn_rc is neither 0 nor 124, so the loop writes a blocked receipt and exits 3 with src/a.ts modified and uncommitted. Consequences: (a) every subsequent fm-quality.sh run &lt;id&gt; --phase clean refuses at line 925 ("the task copy has uncommitted changes") until a human cleans by hand, so the hardened gate is permanently unrunnable for that task; (b) bin/fm-teardown.sh:1178 and bin/fm-control.sh:703 now see a dirty copy; (c) if the worker or a human later does the ordinary git add -A &amp;&amp; git commit, a half-finished agent edit that no test ever saw ships silently. Before this fix round the same failure fell through to the suite check and was either reverted or committed under a green suite, so the tree was never left in this state. The sibling defect-found early exit at lines 1045-1049 has the same hole (the turn prompt says "change nothing", but nothing enforces it). Fix: git -C &#34;$WT&#34; reset --hard --quiet &#34;$round_head&#34; plus git -C &#34;$WT&#34; clean -fdq immediately before both early-exit write_final calls, so a round that ends the phase costs budget and nothing else, exactly as the header promises.
  • ℹ️ bin/fm-quality.sh:956 - Two wall-clock exits report exhausted (exit 1, defined at line 34 as "rounds or the wall-clock bound ran out BELOW THRESHOLD") before any measurement has produced a receipt, so the loop asserts something about the code's quality that it never observed. Line 946 fires when the bound is already gone before the ordinary suite runs; line 956 fires with round still 1, printing "the 20-minute wall-clock bound ran out after 0 round(s)". Concrete state: contract with budget_minutes: 20 and a test command that takes just over 20 minutes and passes green. check_suite returns ok at t=20m, the loop head computes remaining_seconds <= 0, and the phase reports outcome: exhausted with exit 1 and no receipt at all. This is the same wrong-label class the third round accepted and fixed inside check_suite (its out-of-budget branch at line 784 now says blocked, and the exhausted*) arm at line 949 was deleted); the sibling path in cmd_run was not covered. The practical blast radius is small: no pass is manufactured, and because no receipt is written, cmd_status reports missing and quality_gate_override in bin/fm-crew-state.sh reports blocked: quality gate never ran regardless, so a supervisor still gets the right instruction. Noted as a tradeoff rather than a fix request, since the outcome table at line 34 does list "the wall-clock bound ran out" under exhausted, and deciding whether "ran out before measuring" is exhausted or blocked is the author's call. If it should be blocked, the condition is "no measurement completed yet" (round 1, before measure returns a receipt), not the bound itself.

🔧 Fix: revert the round on every early loop exit
3 issues (2 warnings, 1 info) still open:

  • ⚠️ bin/fm-quality.sh:753 - measure validates the receipt against the schema and checks base_sha (line 748) and phase (line 754), but never checks that the receipt's declared threshold matches the threshold the contract configured. The round-2 fix pinned $PHASE.threshold inside the contract (line 980) so a round cannot widen its own bar there, but the receipt side is unchecked, so the same manufactured pass is still reachable one step over. Concrete sequence: contract has clean.threshold.crap_max: 15; the phase command is pnpm run quality:complexity, i.e. a script inside the repo; the agent turn runs with --dangerously-skip-permissions in $WT and edits that script so it prints {&#34;schema_version&#34;:1,&#34;phase&#34;:&#34;clean&#34;,&#34;outcome&#34;:&#34;pass&#34;,&#34;threshold&#34;:{&#34;crap_max&#34;:9999},&#34;findings&#34;:[],...}; the ordinary test suite is unaffected and stays green, so lines 1072-1076 git add -A and commit the edit; round 2 measures, the receipt is schema-valid, base_sha, phase and head_sha all match, MEASURE_OUTCOME=pass, write_final writes a passing receipt and exits 0; cmd_status reports quality: satisfied and quality_gate_override in bin/fm-crew-state.sh lets done stand. A misconfigured (not tampered) project reaches the same place: a phase command that ignores FM_QUALITY_THRESHOLD and judges against its own baked-in number passes silently, which is the flattering silent failure the intent exists to exclude. Fix: after the phase check at line 754, compare receipt_field &#34;$out&#34; threshold with the contract's cfg_keys_under &#34;$PHASE.threshold&#34; numerically (0.80 and 0.8 must compare equal) and report blocked on a mismatch, naming both numbers.
  • ⚠️ bin/fm-quality.sh:840 - cmd_status re-validates each phase receipt with &#34;$RECEIPT_TOOL&#34; validate &#34;$file&#34; and no --check-head, and then checks only base_sha (line 844) and outcome (line 848). So a receipt written at an older HEAD keeps satisfying the gate for a tree it never measured. This is reachable on the ordinary hardened path, not an edge case: the hardened brief (bin/fm-brief.sh:527, step 2) tells the worker to finish both loops "before you start on that definition of done - on a no-mistakes task, that means before the pipeline starts", and the no-mistakes pipeline's own fix rounds then commit further changes. Concrete state: clean and harden both pass at HEAD=A and their receipts record head_sha: A; the pipeline's fixer commits, HEAD becomes B; fm-quality.sh status still reports quality: satisfied; quality_gate_override in bin/fm-crew-state.sh returns nothing and state: done stands - for commits the quality engine never saw. docs/quality-gate.md states --check-head exists precisely so a stale head cannot vouch for the current tree, and the schema made head_sha required for the same reason. A naive --check-head here is not the answer, because the harden loop commits after the clean receipt is written, so the clean receipt's head is legitimately older (that is the stated reason the two phases get separate files). This needs the author's decision on what the gate should promise: re-run the gate after post-gate commits, require the harden receipt's head to be HEAD, or surface the drift in the verdict detail and accept it.
  • ℹ️ bin/fm-quality.sh:489 - finalize_receipt sets doc[&#34;exclusions&#34;] = [line for line in exclusions.splitlines() if line] unconditionally, so whatever the phase command reported in that optional field is replaced by the contract's cfg_seq &#34;$PHASE.exclude&#34; list. Concrete input: a mutation engine that excludes **/*.d.ts internally and reports &#34;exclusions&#34;: [&#34;**/*.d.ts&#34;] in its receipt, under a contract whose clean.exclude is [&#34;src/gen/**&#34;]. The written receipt records only [&#34;src/gen/**&#34;], so the durable proof understates the surface the run was actually measured against - the opposite of what the field was added for (header lines 63-65: "Each receipt records the exclude list the run actually measured against"). Fix: union the engine's reported list with the contract's rather than replacing it, preserving order and dropping duplicates. docs/quality-gate.md does not describe exclusions at all, so naming its owner there would also help.

🔧 Fix: verify receipt threshold, merge exclusions, flag stale gates
1 warning still open:

  • ⚠️ bin/fm-quality.sh:924 - cmd_status's new staleness check compares the receipt's head_sha to the worktree HEAD as raw strings, but the receipt schema explicitly allows an abbreviated sha (docs/quality-receipt.schema.json $defs.sha is ^[0-9a-f]{7,40}$). Any conforming engine that writes git rev-parse --short HEAD therefore reports satisfied-stale on every run even when HEAD has not moved once.

Reproduced end to end against the real script. Task copy with .quality-gate.yaml (clean.command printing a receipt whose head_sha is the 12-char short sha, clean.threshold.crap_max: 15), meta quality=hardened:

  • bin/fm-quality.sh run task --phase clean -> outcome: pass, exit 0. The filed receipt records head_sha: 1c27c397c2c6.
  • HEAD is still 1c27c397c2c69749719c2cadc3a1162d560deea3; nothing committed after the run.
  • bin/fm-quality.sh status task -> quality: satisfied-stale · mode: hardened · the clean receipt measured 1c27c397c2c6, and the task copy is now at 1c27c397c2c69749719c2cadc3a1162d560deea3, so the project&#39;s own CI is what proves these commits.

quality_gate_override in bin/fm-crew-state.sh:127 then emits state: done · ... · quality gate passed on an earlier commit: ..., filing a false drift claim against a gate that measured the exact tree it is vouching for. That defeats the point of the token the user asked for: a caller branching on satisfied-stale alone gets the wrong answer for a whole class of conforming engines.

The short sha survives the loop because nothing else string-compares head_sha: measure (lines 764-775) checks only base_sha and phase, and validate_receipt's --check-head resolves the sha through git rev-parse --verify &lt;sha&gt;^{commit} before comparing (bin/fm-quality-receipt.sh:316-331), so it correctly passes. base_sha is not affected, because measure exact-compares it at line 765 and refuses an abbreviated one before any receipt is filed.

The new tests do not cover this: test_status_marks_a_receipt_from_an_earlier_head (tests/fm-quality.test.sh) and write_clean_receipt (tests/fm-crew-state.test.sh:1660) both write a full git rev-parse HEAD.

Fix: resolve the recorded sha the same way the validator does before comparing, e.g. recorded=$(git -C &#34;$WT&#34; rev-parse --verify --quiet &#34;$(receipt_field &#34;$file&#34; head_sha)^{commit}&#34;), and treat an unresolvable sha as not-stale (the --check-head path already owns that refusal) rather than as drift. Add a test whose receipt carries an abbreviated head_sha at the current HEAD and asserts plain satisfied.

🔧 Fix: compare receipt shas by prefix, not string equality
2 issues (1 warning, 1 info) still open:

  • ⚠️ bin/fm-quality.sh:781 - measure still compares the receipt's base_sha to the task's base commit with raw string equality ([ &#34;$receipt_base&#34; != &#34;$BASE_SHA&#34; ]), while cmd_status was just taught the schema-correct prefix rule via the new same_commit helper (line 445). docs/quality-receipt.schema.json $defs.sha is ^[0-9a-f]{7,40}$, so a conforming engine may record an abbreviated base_sha, and this is exactly the sibling the fix instruction said not to leave wrong ("apply it to any other place ... including the existing base_sha check if it has the same flaw - do not fix one and leave the sibling wrong").

Reproduced end to end against the real script. A task copy with .quality-gate.yaml (clean.command printing a valid receipt, clean.threshold.crap_max: 15), meta quality=hardened, base_sha=b92fc56f734ea4cf0351ec85d8390322ca47984c:

  • engine writes the full base sha -> outcome: pass · phase: clean · mode: hardened · threshold met in 1 round(s), exit 0.
  • the same engine writes base_sha[:12] (the same commit) -> outcome: blocked · phase: clean · mode: hardened · the clean receipt measured against b92fc56f734e, not this task&#39;s base commit b92fc56f734ea4cf0351ec85d8390322ca47984c, exit 3.

The refusal message is self-refuting: it prints a prefix of the very sha it claims is a different commit. The consequence is worse than the cmd_status defect that was just fixed: MEASURE_STATUS=drift reaches cmd_run's no-command|no-receipt|invalid|drift) arm and calls finish blocked, so a project whose engine abbreviates base_sha can never pass the hardened gate at all, and status_append files blocked: quality clean could not measure, ... against a healthy crew every attempt. It also leaves the two comparisons inconsistent: cmd_status (line 928) now accepts an abbreviated base_sha that measure refuses to file in the first place.

The new tests do not cover it. test_an_abbreviated_base_is_the_same_commit (tests/fm-quality.test.sh:1355) patches base_sha onto an already-filed receipt with set_receipt_sha and only exercises cmd_status; the fixture's FMQ_FORCE_BASE path is never used with an abbreviated-but-matching value.

Fix: if ! same_commit &#34;$receipt_base&#34; &#34;$BASE_SHA&#34;; then at line 781, keeping the check pure string work as the helper's comment requires. Add a falsifiable pair in tests/fm-quality.test.sh: a phase command whose receipt names the 12-character short form of FM_QUALITY_BASE_SHA reports pass, and one whose receipt names a 12-character near miss still reports blocked with the drift detail.

  • ℹ️ bin/fm-crew-state.sh:130 - quality_gate_override's new satisfied-stale arm (line 128) correctly guards an empty incoming detail with &#34;${detail:+$detail$SEP}&#34;, but the four older arms at lines 130-133 interpolate %s%s with $detail and $SEP unconditionally. When emit is called for a done state with no detail, the returned detail begins with a bare separator and emit then adds its own, producing an empty field in a line supervision parses.

Concrete state: a hardened task (quality=hardened) whose status log's last line is the degenerate done: with nothing after the colon. status_line_verb (bin/fm-classify-lib.sh:183) yields done, map_log_state maps it to done, and status_line_note (bin/fm-classify-lib.sh:222) yields the empty string, so line 732 calls emit done status-log &#34;&#34;. With no receipt filed, quality_gate_override returns blocked| · quality gate never ran: ... and the emitted line is state: blocked · source: status-log · · quality gate never ran: this task ships hardened but ... - two consecutive · separators around an empty field. The state token and the reason are both correct, so this is a malformed field rather than a wrong verdict, and it only fires on a status line that is already malformed.

Fix: use the same &#34;${detail:+$detail$SEP}&#34; form in the missing, not-passed, unreadable and catch-all arms as the satisfied-stale arm already does.

🔧 Fix: route all sha compares through one prefix rule
1 info still open:

  • ℹ️ bin/fm-quality.sh:961 - cmd_receipt builds $list as a space-joined string and expands it unquoted into python3 - $list, so any receipt path containing a space is word-split into several arguments. Concrete input: FM_HOME=&#34;/home/me/fm home&#34; (or any FM_DATA_OVERRIDE with a space) and a task with a filed receipt at $DATA/&lt;id&gt;/quality-clean-receipt.json. fm-quality.sh receipt &lt;id&gt; then calls python with /home/me/fm and home/data/&lt;id&gt;/quality-clean-receipt.json as two separate argv entries, the open() at the top of the loop raises FileNotFoundError, and the command dies with a Python traceback and exit 1 instead of printing the receipt. Every other path expansion in this script is quoted (mkdir -p &#34;$DATA/$ID&#34;, rm -f -- &#34;$(receipt_path &#34;$PHASE&#34;)&#34;, git -C &#34;$WT&#34;), so receipt is the only subcommand that breaks on a spacey home, and the SC2086 disable makes the split look intentional rather than a gap. Fix: collect the paths in a bash array (local -a list=(); list+=(&#34;$(receipt_path &#34;$phase&#34;)&#34;)) and expand it as &#34;${list[@]}&#34;, guarding the empty case so python3 - &lt;&lt;PY still prints [].
✅ **Test** - passed

✅ No issues found.

  • bash bin/fm-test-run.sh tests/fm-quality.test.sh tests/fm-quality-receipt.test.sh - 2 scripts, 0 failed
  • bash bin/fm-test-run.sh tests/fm-crew-state.test.sh tests/fm-brief.test.sh tests/fm-control-relaunch.test.sh tests/fm-task-delivery.test.sh - 4 scripts, 0 failed
  • Manual end-to-end CLI drive of bin/fm-quality.sh run &lt;id&gt; --phase clean over a real throwaway git worktree and real .quality-gate.yaml, once per outcome: pass (exit 0), read-only (exit 6), blocked (exit 3), not-applicable (exit 4), exhausted (exit 1), stuck (exit 1), defect-found (exit 5)
  • Manual wall-clock enforcement check: phase command sleep 300 under bounds.budget_minutes: 1 - loop returned exhausted after 65s wall clock
  • Manual receipt comparison: hardened vs standard run of the identical measurement, asserting metrics and threshold equal while outcome differs ('pass' vs 'read-only')
  • Manual drive of bin/fm-quality.sh status &lt;id&gt; for the satisfied, not-required and missing verdicts
  • Manual drive of bin/fm-crew-state.sh &lt;id&gt; with an identical passing pipeline run and six different receipt states, showing hardened+pass -> done, hardened+missing -> blocked, hardened+exhausted -> blocked, hardened+read-only -> blocked, hardened+corrupt -> unknown, standard -> done unchanged
  • Falsifiability mutation 1: write_final &#34;$last_receipt&#34; read-only -> pass in bin/fm-quality.sh; tests/fm-quality.test.sh fails on 'a read-only run must not exit with the pass code'
  • Falsifiability mutation 2: finish blocked &#34;$MEASURE_DETAIL&#34; -> finish pass at bin/fm-quality.sh:1077; tests/fm-quality.test.sh fails on 'a read-only run that could not measure is blocked'
  • Falsifiability mutation 3: remaining_seconds() returns constant 86400 at bin/fm-quality.sh:414; tests/fm-quality.test.sh fails on 'a measurement that outruns the bound exits 1'
  • Checked no .quality-gate.yaml exists at the firstmate repository root, and kill_rate_min remains 0.80 in docs/quality-gate.md
⚠️ **Document** - 1 info
  • ℹ️ bin/fm-quality-receipt.sh:294 - The two quality scripts compare commit ids by different rules, and the documentation states each correctly, so no doc edit can reconcile them. bin/fm-quality.sh:447 treats a shorter sha as the same commit when it is a prefix of the longer one, because the schema's sha pattern allows 7 to 40 characters. bin/fm-quality-receipt.sh:294 instead requires each verify child's base_sha and head_sha to be byte-equal to the envelope's, which docs/quality-gate.md:152 states as written. A schema-conforming verify receipt whose envelope spells a sha in full while a child abbreviates it is therefore rejected. Reported rather than changed: the receipt schema just landed and the hand-rolled validator is tracked separately as fm-receipt-schema-validator.
🔧 **Lint** - 1 issue found → auto-fixed ✅
  • ⚠️ linter found issues (exit code 1)

🔧 Fix: silence deliberate SC2016 hits in quality tests
✅ Re-checked - no issues remain.

✅ **Push** - passed

✅ No issues found.

BohnBawerick and others added 30 commits August 22, 2026 09:14
…sk record

Reads a project's registered "+hardened" annotation and carries it to the
worker's instructions and the task's durable record, so the quality loop that
bin/fm-quality.sh will drive has a posture and a fixed base commit to work from.
That script is not part of this change; it is referenced by name only.

- bin/fm-project-mode.sh: --quality prints one word, standard or hardened. The
  two-word stdout its three callers parse is untouched, so it gets its own
  output path. The bracket grammar is now position-tolerant: a "+"-prefixed
  token is a flag and never a mode, so "[+hardened local-only]" resolves the
  mode behind it instead of reading the flag as an unknown mode. Unrecognized
  flags are still ignored rather than refused.
- bin/fm-brief.sh: --quality standard|hardened, defaulting to standard and
  refused on scout, dreamer, and secondmate scaffolds. A hardened brief records
  the sibling "Quality contract: quality=hardened" line and one short quality
  gate section; a standard brief records neither and stays byte-identical to
  the pre-quality scaffold.
- bin/fm-spawn.sh: the brief's quality line must agree with --quality, the same
  check the delivery line already gets, in both directions. quality= and
  base_sha= land in the task record; the base commit is captured once at spawn
  and read back on relaunch, never recaptured, because the loop commits each
  round and a later capture would narrow the gate while still reporting success.
- AGENTS.md: one sentence placing quality resolution at intake.

Tests execute the real interfaces. The load-bearing ones prove a project
without "+hardened" and a brief scaffolded without --quality behave exactly as
before: the two-word stdout is pinned across every annotation form, the two
scaffolds are compared byte for byte, and the task record's key set is pinned
so only quality= and base_sha= are additive.
The stage 0a pilot showed the receipt cannot express real findings
and that a missing head_sha makes a drifted base report
not-applicable and exit 0. This revises unpublished schema v1 in
place: require head_sha, duration_ms, engine, threshold, and a
stable finding id; replace survivors[] with per-phase findings[];
and make verify one envelope with phases[]. bounds.budget_minutes
is the missing wall-clock bound.
A short herdr recent tail can drop Claude's opening rule and classify an
idle composer unknown, so native-hosted away-mode never injected.
Classify a glyph immediately under a closing rule as empty, read the
visible viewport for composer capture, and let herdr native idle deliver
when the composer is still unknown. Max-defer retries that path before
alarming.

The native-hosted daemon still injects into the captain pane. A dead
shell has no idle agent registration and still defers.
A Claude Code background continuation runs in its own process tree, so the
harness-ancestry walk answered a different question in a hook than in a tool
shell. One session could hold the lock by the auto-arm's reckoning and not hold
it by every mutating path's, which left supervision off while wake drains,
gate answers, merges, teardowns and a promotion all proceeded.

- bin/fm-session-lock-lib.sh becomes the single owner of the ownership verdict
  and its refusal. Identity resolves in three tiers: the vendor-declared
  CLAUDE_PID, then the conversation id recorded in state/.lock.session, then
  the ancestry walk for harnesses that declare neither.
- fm_require_session_lock gates the eight fleet-mutation entry points before
  argument validation, so the read-only rule is enforced where mutation
  happens. It refuses only on a live foreign owner plus a caller that is
  itself in a harness session, leaving ssh, detached and CI callers alone.
- bin/fm-lock.sh and bin/fm-session-start.sh state ownership in words; the
  digest carries an explicit HELM: line.
- bin/fm-turnend-guard.sh tells a correct decline from an auto-arm failure,
  reports it once per holder, and then stops blocking.
- tests/fm-session-lock-ownership.test.sh pins all three parts with real
  competing processes; tests/fm-session-identity-live-e2e.test.sh proves the
  two vendor-declared values against real Claude Code.
…iation

- bin/fm-quality.sh runs a project's own quality phase commands, reads the
  D2 receipt they print, enforces the contract's bounds, and reports one of
  seven outcomes with an exit code that is a pure function of the outcome
- read-only is a first-class mode, not a fallback: measurement and reporting
  run wherever they can, only blocking is gated on the hardened posture, and
  a reported score records outcome read-only with its own exit code so no
  caller can read it as a gate that ran and approved
- the wall-clock bound is enforced rather than assumed; a host that cannot
  bound a measurement is refused instead of running one unbounded
- could-not-measure is always blocked: a missing toolchain, a red or flaky
  ordinary suite, an unreadable receipt, or a drifted base commit
- bin/fm-crew-state.sh filters every done verdict for a hardened task through
  that script's own gate verdict, keeping the ternary discipline it already
  uses for run attribution; every other posture's line is unchanged
- tests/fm-quality.test.sh drives all seven outcomes over real worktrees and
  real receipts, in falsifiable pairs; the crew-state and receipt suites gain
  the gate and read-only cases; a live opt-in guard pins the structured-output
  flag against the installed harnesses
@BohnBawerick
BohnBawerick merged commit 524bd52 into main Aug 30, 2026
13 checks passed
@BohnBawerick
BohnBawerick deleted the fm/fm-quality-loop branch October 4, 2026 02:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant