feat(bin): add the quality gate loop and the hardened project posture - #19
Merged
Merged
Conversation
…sk record Reads a project's registered "+hardened" annotation and carries it to the worker's instructions and the task's durable record, so the quality loop that bin/fm-quality.sh will drive has a posture and a fixed base commit to work from. That script is not part of this change; it is referenced by name only. - bin/fm-project-mode.sh: --quality prints one word, standard or hardened. The two-word stdout its three callers parse is untouched, so it gets its own output path. The bracket grammar is now position-tolerant: a "+"-prefixed token is a flag and never a mode, so "[+hardened local-only]" resolves the mode behind it instead of reading the flag as an unknown mode. Unrecognized flags are still ignored rather than refused. - bin/fm-brief.sh: --quality standard|hardened, defaulting to standard and refused on scout, dreamer, and secondmate scaffolds. A hardened brief records the sibling "Quality contract: quality=hardened" line and one short quality gate section; a standard brief records neither and stays byte-identical to the pre-quality scaffold. - bin/fm-spawn.sh: the brief's quality line must agree with --quality, the same check the delivery line already gets, in both directions. quality= and base_sha= land in the task record; the base commit is captured once at spawn and read back on relaunch, never recaptured, because the loop commits each round and a later capture would narrow the gate while still reporting success. - AGENTS.md: one sentence placing quality resolution at intake. Tests execute the real interfaces. The load-bearing ones prove a project without "+hardened" and a brief scaffolded without --quality behave exactly as before: the two-word stdout is pinned across every annotation form, the two scaffolds are compared byte for byte, and the task record's key set is pinned so only quality= and base_sha= are additive.
…ture and scripts inventory
The stage 0a pilot showed the receipt cannot express real findings and that a missing head_sha makes a drifted base report not-applicable and exit 0. This revises unpublished schema v1 in place: require head_sha, duration_ms, engine, threshold, and a stable finding id; replace survivors[] with per-phase findings[]; and make verify one envelope with phases[]. bounds.budget_minutes is the missing wall-clock bound.
A short herdr recent tail can drop Claude's opening rule and classify an idle composer unknown, so native-hosted away-mode never injected. Classify a glyph immediately under a closing rule as empty, read the visible viewport for composer capture, and let herdr native idle deliver when the composer is still unknown. Max-defer retries that path before alarming. The native-hosted daemon still injects into the captain pane. A dead shell has no idle agent registration and still defers.
… proven idle composers
…erdr composer read
A Claude Code background continuation runs in its own process tree, so the harness-ancestry walk answered a different question in a hook than in a tool shell. One session could hold the lock by the auto-arm's reckoning and not hold it by every mutating path's, which left supervision off while wake drains, gate answers, merges, teardowns and a promotion all proceeded. - bin/fm-session-lock-lib.sh becomes the single owner of the ownership verdict and its refusal. Identity resolves in three tiers: the vendor-declared CLAUDE_PID, then the conversation id recorded in state/.lock.session, then the ancestry walk for harnesses that declare neither. - fm_require_session_lock gates the eight fleet-mutation entry points before argument validation, so the read-only rule is enforced where mutation happens. It refuses only on a live foreign owner plus a caller that is itself in a harness session, leaving ssh, detached and CI callers alone. - bin/fm-lock.sh and bin/fm-session-start.sh state ownership in words; the digest carries an explicit HELM: line. - bin/fm-turnend-guard.sh tells a correct decline from an auto-arm failure, reports it once per holder, and then stops blocking. - tests/fm-session-lock-ownership.test.sh pins all three parts with real competing processes; tests/fm-session-identity-live-e2e.test.sh proves the two vendor-declared values against real Claude Code.
…tity, and turn-end notices
…iation - bin/fm-quality.sh runs a project's own quality phase commands, reads the D2 receipt they print, enforces the contract's bounds, and reports one of seven outcomes with an exit code that is a pure function of the outcome - read-only is a first-class mode, not a fallback: measurement and reporting run wherever they can, only blocking is gated on the hardened posture, and a reported score records outcome read-only with its own exit code so no caller can read it as a gate that ran and approved - the wall-clock bound is enforced rather than assumed; a host that cannot bound a measurement is refused instead of running one unbounded - could-not-measure is always blocked: a missing toolchain, a red or flaky ordinary suite, an unreadable receipt, or a drifted base commit - bin/fm-crew-state.sh filters every done verdict for a hardened task through that script's own gate verdict, keeping the ternary discipline it already uses for run attribution; every other posture's line is unchanged - tests/fm-quality.test.sh drives all seven outcomes over real worktrees and real receipts, in falsifiable pairs; the crew-state and receipt suites gain the gate and read-only cases; a live opt-in guard pins the structured-output flag against the installed harnesses
BohnBawerick
force-pushed
the
fm/fm-quality-loop
branch
from
August 22, 2026 11:03
32b1dae to
524bd52
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Build the quality loop controller and its state reconciliation: deliverables D4, D8 and D10 from the parked Superchecks specification (implementation-thoughts.md), with read-only scoring folded in as a first-class mode rather than bolted on afterwards.
Build on what already landed. D1/D2 were proven on real TypeScript by the Stage 0a pilot on quota-axi. The D2 receipt schema was just revised and landed (PR #17): it carries head_sha, per-phase findings with stable ids, engine name and version, the threshold the outcome was judged against, a required duration_ms, and a wall-clock bound. Implement against the landed schema, not the older description of it.
The wall-clock number this task was blocked on: a diff-scoped measurement costs 20 seconds to 4 minutes with the right engine, roughly linear in changed lines, so four loop iterations cost minutes and the loop is worth building. Design to that number and make the loop's own wall-clock bound enforce it rather than trusting it. The pilot also produced a warning to respect: on the wrong engine the same work took 1h47m AND was wrong - it counted 15 crashed test processes as mutant kills, and on the smallest diff reported a perfect 100% kill rate that was entirely an artifact of one timing-sensitive test failing under load. A flattering silent failure is the dangerous one here, so the loop must be able to tell "measured and passed" apart from "could not measure".
Read-only scoring, per the captain on 2026-08-22 ("even if the mode is not active I think it should show the scores as read only yes"): measurement and reporting are always on wherever they can run, and only blocking is gated on the hardened mode. A standard project surfaces its mutation and complexity scores so the numbers exist before anyone commits to a threshold, which is also how a project earns an informed decision to switch the mode on. The receipt must carry a read-only outcome distinctly from a pass, so a score that merely reported is never mistaken for a gate that ran and approved; a read-only outcome that looks like a pass is worse than no read-only mode at all.
Acceptance criteria:
Constraints and deliberate exclusions:
One specification item was implemented differently, deliberately: the parked spec named a single data//quality-receipt.json holding both phases, but the landed schema requires a verify envelope's children to share its head_sha, and the harden loop commits after the clean receipt is written, so the two phases never share a head. Each phase therefore gets its own complete receipt file.
What Changed
bin/fm-quality.shruns a project's own quality commands for a phase (cleanorharden) under the contract's round and wall-clock bounds, checks the ordinary test suite first so a score is never measured over a red or flaky suite, and reports one outcome per run with a fixed exit code:pass0,exhausted/stuck1,blocked3,not-applicable4,defect-found5,read-only6. A run that could not measure reportsblocked, neverpass, and a standard project'sread-onlyrun measures and reports scores without blocking anything. Each phase writes its own complete D2 receipt todata/<id>/quality-<phase>-receipt.json, stamped with the exclude list the run actually measured against, andstatusanswers with one token (satisfied,satisfied-stale,missing,not-passed,unreadable,not-required).bin/fm-quality-receipt.shvalidates a receipt against the newdocs/quality-receipt.schema.json, adds the post-schema rules JSON Schema cannot state (unique finding ids, verify children sharing the envelope'sbase_sha/head_sha), and can requirehead_shato resolve to a given tree's HEAD.docs/quality-gate.mdowns the contract and bounds;docs/architecture.md,docs/scripts.md,README.md,AGENTS.md, and the project-management skill record the posture.data/projects.mdgains a position-independent+hardenedbracket token thatfm-project-mode.sh --qualityresolves to one word (refused beside the conditionalno-mistakes-prod-onlypolicy);fm-brief.sh --qualityemits aQuality contract: quality=hardenedline plus one quality-gate section, leaving a standard brief byte-identical;fm-spawn.sh --qualityrecordsquality=and a once-capturedbase_sha=in the task meta, reads both back on relaunch, and refuses a brief that disagrees;fm-crew-state.shroutes adoneverdict on a hardened task through the quality status verdict and leaves every other posture's line unchanged;fm-promote.shprints an advisory notice when a promoted task carries less rigor than the project's standing posture. Newtests/fm-quality.test.sh,tests/fm-quality-receipt.test.sh, and a live structured-output drift check cover the outcomes and the receipt rules, with the brief, crew-state, relaunch, and task-delivery suites extended and every new suite registered inbin/fm-test-run.sh.Not wired in
Read-only scoring is implemented and covered by tests, but no firstmate path invokes it yet, so a standard project still surfaces no scores in practice.
bin/fm-brief.sh:519andbin/fm-brief.sh:603both gate the quality contract line and the quality-gate section on[ "$QUALITY" = hardened ], so a standard brief stays byte-identical to what it was before this change and never tells its worker to run the loop.Turning it on means relaxing those two guards for the read-only case: emit a quality section on a standard brief whose project has a
.quality-gate.yaml, telling the worker to runbin/fm-quality.sh run <id> --phase cleanand report the scores, while leaving blocking gated onquality=hardenedexactly as it is now. That would change what every ship task on every project does, which is a wider contract than this task was scoped to, so it is left out deliberately rather than by oversight.Risk Assessment
Testing
Ran the six targeted test scripts covering the quality loop, receipt validator, crew state, brief, relaunch and delivery paths - all passed with no failures. Beyond the suites, I drove the real fm-quality.sh CLI over throwaway git projects to show each of the six specification outcomes with its distinct exit code, showed the read-only receipt carrying "outcome": "read-only" against a byte-identical measurement whose hardened twin records "pass", showed a missing engine reporting blocked with no receipt written at all, and showed a 300-second engine cut off after 65 seconds by a 1-minute wall-clock budget. I also drove fm-crew-state.sh with an identical passing pipeline run against six receipt states to show the supervisor line a fleet operator actually reads, where a read-only receipt reports blocked rather than done and a standard task is untouched. Finally I applied three mutations (read-only to pass, could-not-measure to pass, constant wall-clock budget) and confirmed each turns the owning test red, so the tests are not satisfiable by a constant. This change is a bash CLI with no rendered UI surface, so the reviewer-visible artifacts are CLI transcripts and receipt JSON rather than screenshots.
Evidence: CLI transcript: all six loop outcomes, receipt JSON, and wall-clock enforcement
Source: CLI transcript: all six loop outcomes, receipt JSON, and wall-clock enforcement
Evidence: CLI transcript: what the fleet supervisor sees for each receipt state
Source: CLI transcript: what the fleet supervisor sees for each receipt state
The pipeline run says "passed" in every case below. Only the quality receipt differs. A: quality=hardened, no receipt (the gate never ran) $ fm-crew-state.sh q state: blocked · source: run-step · run passed: PR merged/closed · quality gate never ran: no receipt for: clean B: quality=hardened, receipt outcome=pass $ fm-crew-state.sh q state: done · source: run-step · run passed: PR merged/closed C: quality=hardened, receipt outcome=exhausted (measured, below the bar) $ fm-crew-state.sh q state: blocked · source: run-step · run passed: PR merged/closed · quality gate did not pass: the clean phase reported exhausted D: quality=hardened, receipt outcome=read-only (reported, never gated) $ fm-crew-state.sh q state: blocked · source: run-step · run passed: PR merged/closed · quality gate did not pass: the clean phase reported read-only E: quality=hardened, receipt file corrupt (cannot tell) $ fm-crew-state.sh q state: unknown · source: run-step · run passed: PR merged/closed · quality gate unreadable: the clean receipt is present but not a valid receipt F: quality=standard, no receipt (an ordinary task, untouched) $ fm-crew-state.sh q state: done · source: run-step · run passed: PR merged/closedEvidence: Falsifiability: three mutations, each caught by the test that pins it
Source: Falsifiability: three mutations, each caught by the test that pins it
MUTATION 1 - a read-only run reports 'pass' instead of 'read-only' - write_final "$last_receipt" read-only "$round" + write_final "$last_receipt" pass "$round" not ok - a read-only run must not exit with the pass code: expected exit 6, got 0 MUTATION 2 - 'could not measure' reports 'pass' instead of 'blocked' - no-command|no-receipt|invalid|drift) finish blocked "$MEASURE_DETAIL" ;; + no-command|no-receipt|invalid|drift) finish pass "$MEASURE_DETAIL" ;; not ok - a read-only run that could not measure is blocked: expected exit 3, got 0 MUTATION 3 - the wall-clock budget is a constant 86400s (never enforced) - printf '%s' "$(( BUDGET_SECONDS - used ))" + printf '%s' 86400 not ok - a measurement that outruns the bound exits 1: expected exit 1, got 3 UNMUTATED source, same suite all fm-quality tests passedEvidence: Read-only is machine-distinguishable from pass (same measurement, two receipts)
Evidence: Wall-clock bound is enforced, not assumed
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
.agents/skills/project-management/SKILL.md- branch carries 8 commit(s) that exist on your local main branch but were never pushed to origin/main; rebasing would bundle this unrelated work (18 file(s)) into the PR:Push main to origin, or rebase your branch onto origin/main, before gating.
bin/fm-quality.sh:860-run --dry-rundeletes the phase's existing receipt before the dry-run branch is reached, then prints "dry-run: nothing was run and no receipt was written". Therm -f -- "$(receipt_path "$PHASE")"at line 860 runs unconditionally; theif [ "$DRY_RUN" -eq 1 ]early exit is at line 892. Reproduced: a task with a valid passingdata/<id>/quality-clean-receipt.json, thenfm-quality.sh run task --phase clean --dry-run-> the file is gone, exit 0, and the output claims nothing was written. On a hardened taskbin/fm-crew-state.shthen rewritesdonetoblockedviaquality_gate_override("quality gate never ran"), so a read-only inspection flag silently destroys durable proof and flips fleet supervision. Fix: move therm -fto after the dry-run branch (or guard it with[ "$DRY_RUN" -eq 0 ]).bin/fm-quality.sh:701- The contract is loaded once at line 862 (state=$(load_contract "$WT")) and never reloaded inside the round loop, soFM_QUALITY_EXCLUDE="$(cfg_seq "$PHASE.exclude")"at line 701 replays a pre-round-1 snapshot for every measurement. Butwrite_turn_promptexplicitly instructs the agent to "add the narrowest possible entry to the exclude list in .quality-gate.yaml" and answer{"action":"excluded"}, and the loop commits that edit. Reproduced with a stub harness that appends- "a.ts"to the exclude list: all 3 rounds received the unchangedsrc/gen/**, the excluded finding reappeared every round, and the phase reportedexhausted. One of the four documented agent actions is a no-op within a run, and the loop reportsexhausted/stuckwhere the exclusion should have producedpass. Fix: re-runload_contract/load_bounds(or at least re-read$PHASE.exclude) at the top of each round, refusing if the reparse now errors.bin/fm-quality.sh:754-check_suitepasses its original$secsto the second (flaky-detection) run instead of recomputingremaining_seconds, so the pre-flight suite check can consume up to 2xbounds.budget_minutesbefore the first measurement even starts. Reproduced withtest: "sleep 5; exit 1"andbudget_minutes: 0.1(6s bound): the run took 11s wall clock. With the default 20-minute bound a red slow suite can spend ~40 minutes, which is the assumption the intent's "the loop enforces its own wall-clock bound rather than assuming the measurement is fast" criterion exists to remove. Fix: recompute the remaining budget before the secondbounded_shand reportexhaustedif none is left.bin/fm-quality.sh:730-measurevalidates the receipt against the schema and checksbase_shaand (via--check-head)head_sha, but never checks that the receipt'sphaseequals$PHASE. Reproduced: a contract whoseclean.commandprints"phase":"harden"withclassification: killablefindings is accepted, anddata/<id>/quality-clean-receipt.jsonis written recordingphase: hardenwithkillablefindings.cmd_statusthen reads that file as the clean phase's passing receipt. This bypasses the clean/harden classification split the schema exists to enforce (docs/quality-gate.md: "A clean finding classifiedkillableis invalid, which is the point of splitting the vocabulary"), and a misconfigured project can satisfy the clean gate with a harden-engine receipt without any error. Fix: after the base_sha check, requirereceipt_field "$out" phaseto equal$PHASEand reportblockedotherwise.bin/fm-quality.sh:416- Theperlbranch ofbounded_shends withexit($? >> 8), which discards the signal byte: a child killed by a signal yields exit 0. Confirmed directly:perl -e '...exit($? >> 8)' 10 sh -c 'kill -SEGV $$'returns 0 wheretimeoutwould return 139. On a host with neithertimeoutnorgtimeout(the only case that selects this branch),check_suitereads rc=0 and printsok, so a test suite that segfaults or is OOM-killed is scored as green and the mutation/complexity number is trusted; the same path makes the post-agent-turn suite check commit a round whose suite crashed. That is the pilot's documented flattering-silent-failure mode the loop is meant to exclude. Fix:exit($? & 127 ? 128 + ($? & 127) : $? >> 8). The identical expression is copied from bin/fm-nm-run-lib.sh:28, which has the same defect.bin/fm-quality.sh:955-if [ -n "$prev_ids" ] && [ "$ids" = "$prev_ids" ]conflates "no previous round yet" with "the previous round had an empty finding set", sono_progressnever increments when a below-threshold receipt reportsfindings: []. Reproduced with a phase command that always returnsoutcome: exhaustedandfindings: []undermax_iterations: 4, no_progress_limit: 2: the loop reportedexhaustedafter 4 rounds and 3 paid agent turns instead ofstuckafter 3 rounds and 2 turns. The header claims "'No progress' means the same finding ids, not the same count" - an unchanged empty id set is the clearest no-progress case and is the one it misses. Fix: track a separatehave_prevflag (or use an unset sentinel) instead of testing-n "$prev_ids".🔧 Fix: fix six quality-loop defects and pin them with tests
4 issues (2 errors, 2 warnings) still open:
bin/fm-quality.sh:904-run --dry-runstill produces supervision side effects. The fix round guarded the receipt deletion at line 877, but the dry-run early exit is at line 909 and sixfinishcalls run before it: no base commit (890), unresolvable base (893), no timeout tool (896), structured-output refusal (902), the dirty-tree check (904), and both not-applicable paths (881, 887).finishcallsstatus_append, which on a hardened task appends to$STATE/$ID.status. Concrete path: a hardened ship task whose worktree has one uncommitted file, thenfm-quality.sh run <id> --phase clean --dry-run-> the lineblocked: quality clean could not measure, the task copy has uncommitted changes; the loop commits each round, ...is appended to the status log, and exit is 3, not 0. bin/fm-classify-lib.sh (FM_CLASSIFY_CAPTAIN_RE_DEFAULT, line 49, and thedone|needs-decision|blocked|failedsets at lines 95/117/1032) treats ablocked:line as a decision-opening supervision event, and bin/fm-crew-state.sh emits it asstate: blocked · source: status-log. A hardened task is dirty for most of its working life, so an inspection-only flag routinely files a real blocked event against a healthy crew. This is the same class as the receipt-deletion defect just fixed, on the sibling path. Fix: makestatus_appendreturn early whenDRY_RUNis 1 ([ "$DRY_RUN" -eq 0 ] || return 0at line 512), which leaves the printed dry-run report unchanged.bin/fm-quality.sh:940- The per-round contract re-read added by the fix round re-reads the phase threshold as well as the exclude list, so the agent being gated can widen its own bar and turn a below-threshold phase intopass.load_contractat line 941 rewrites the whole$CFGfile;measurethen rebuildsFM_QUALITY_THRESHOLDfrom it at line 700 on every round. Concrete sequence: contract hasclean.threshold.crap_max: 15; round 1 measures below threshold and reportsexhaustedwith findings; the agent turn (running under--dangerously-skip-permissions/--sandbox danger-full-accessin the worktree) edits.quality-gate.yamltocrap_max: 9999; the ordinary suite is still green, so lines 1028-1032git add -Aand commit the edit; round 2 re-reads the contract, exportsFM_QUALITY_THRESHOLD=crap_max=9999, the phase command judges against it and printsoutcome: pass;write_finalat line 981 writes a passing receipt and exits 0.cmd_status(line 826) andquality_gate_overridein bin/fm-crew-state.sh both accept that as a gate that ran and approved. Before this commit the contract was loaded once, so this was structurally impossible; the round-loop comment at lines 938-939 states the invariant that should cover it ("a run does not get to extend its own wall clock or swap the command it is being measured by") but the threshold is not among the pinned values, and the turn prompt's "do not change the threshold" (line 673) is the only thing enforcing it. Fix: captureFM_QUALITY_THRESHOLDfrom the round-1 contract into a variable before the loop and pass that pinned value tomeasure, or refuse withblockedwhen the re-read threshold differs from round 1's.bin/fm-quality.sh:763- When the wall-clock bound runs out before the flake re-run,check_suitereportsexhausted(exit 1) instead ofblocked(exit 3), but no information is actually missing: the first run already returned non-zero and was not a timeout (rc 124 is handled at line 754), and both possible answers from the re-run areblocked- red twice isblocked(line 775) and red then green isblocked(line 773). So the correct outcome is knowablyblockedregardless of what the re-run would have said. Concrete state:test: "..."that fails after consuming the wholebudget_minutes(for example a 6s bound with a suite that fails at 6.1s) ->remaining_secondsreturns <= 0, cmd_run's newexhausted*)branch at line 927 callsfinish exhausted, and the caller sees exit 1 ("ran out below threshold") for what the script's own outcome table at line 37 assigns toblocked("could not measure: toolchain, red or flaky suite, drift"). No pass is manufactured, so this is a wrong label rather than a wrong verdict. Fix: emitblocked the ordinary test suite is red before any quality work and the bound ran out before red and flaky could be told apartand drop theexhausted*)branch at line 927.bin/fm-quality.sh:369-bounds.budget_usdis the one boundload_boundsreads without validating it.max_iterationsandno_progress_limitgo throughpositive_integer(line 372) andbudget_minutesthroughpositive_number(line 374), butbudget_usdis copied straight from the contract at line 369 and interpolated unquoted into the claude command string at line 576 (--max-budget-usd $BOUND_BUDGET_USD), which is then run throughsh -cbybounded_sh. Concrete input:.quality-gate.yamlwithbounds: { budget_usd: unlimited }-> round 1 buildsclaude ... --max-budget-usd unlimited ..., claude exits with a usage error,agent_turnswallows it (|| true, stderr to /dev/null) and prints nothing, soturn_actionis empty and the loop scores the round as no progress. Every round then fails identically and the phase reportsstuckorexhaustedwith no mention of the misconfigured bound. Separately, the codex branch (lines 583-591) never passesbudget_usdat all, so a hardened run on codex has no spend ceiling even though the contract declares one and the dry-run report at line 912 prints it as an enforced bound. Fix: validatebudget_usdwithpositive_numberalongside the others, and either pass it to codex or say in the header that it applies to claude only.🔧 Fix: pin threshold, quiet dry runs, validate budget_usd
3 warnings still open:
bin/fm-brief.sh:603- Read-only scoring is implemented but unreachable in practice: no firstmate path ever runs it. The intent requires "measurement and reporting are always on wherever they can run" and "A standard project surfaces its mutation and complexity scores so the numbers exist before anyone commits to a threshold, which is also how a project earns an informed decision to switch the mode on." The contradicting hunk isif [ "$QUALITY" = hardened ]; then DOD="$QUALITY_SECTION\n\n$DOD"at bin/fm-brief.sh:603 (and the siblingCONTRACT_LINESguard at line 519): a standard brief carries no quality-gate section, so its worker is never told to runfm-quality.sh run <id> --phase clean. The only other caller of the script isquality_gate_overridein bin/fm-crew-state.sh:121, and it invokesstatus, notrun;cmd_statusin bin/fm-quality.sh:788 short-circuits every non-hardened task withquality: not-required · ... no receipt gates itbefore reading any receipt. Concrete path: a standard project that commits a valid.quality-gate.yamlwithclean.commandandharden.command, spawned with the default--quality standard.fm-quality.sh runis never executed,data/<id>/quality-clean-receipt.jsonis never written, and no score is ever reported to the captain, so the project can never earn the informed decision to switch to hardened. The MODE=read-only branch (bin/fm-quality.sh:191, 980-986) and its tests exercise the capability only when a human types the command by hand. This challenges the deliberate choice at bin/fm-brief.sh:598-608 to keep a standard brief byte-identical, so it needs the author's decision rather than a reviewer patch.bin/fm-quality.sh:586- An agent turn that never ran is indistinguishable from an agent that ran and changed nothing, and the loop labels the first asstuck.agent_turndiscards both the exit status and stderr (bounded_sh "$secs" "$WT" "$cmd" > "$out" 2>/dev/null || trueat line 586 for claude, line 595 for codex), so nothing downstream can tell a failed harness invocation from a completed turn. Concrete state: a hardened task whose claude credentials have expired (or whose--max-budget-usdflag was renamed -check_structured_outputat line 547 only greps for--json-schema, never for the spend flag it also interpolates at line 581). claude writes its error to stderr and nothing to stdout,$outstays empty, the tolerant reader at line 603 prints nothing, andturn_actionis empty. The worktree is unchanged, so the post-turn suite passes (rc=0),git add -Afinds nothing to commit, and the next round measures the identical finding set.no_progressincrements at line 995 on every round until line 1001 fireswrite_final ... stuckwith the detail "$BOUND_NO_PROGRESS consecutive rounds left the same findings untouched" - a statement about the agent's work that is false, since no agent turn happened at all. The phase then appendsblocked: quality clean stuck, ...to the status log (line 532), pointing a supervisor at the code instead of at the broken harness, after spending the full wall-clock and round budget. No pass is manufactured, so this is a wrong label rather than a wrong verdict. Fix: capturebounded_sh's rc inagent_turnand let the caller distinguish "the harness exited non-zero and produced no readable answer" from "the agent answered no-change", reporting the former asblockednaming the harness and its version, the way check_structured_output already does.bin/fm-quality.sh:706- The per-round contract re-read added in the first fix round made mid-run exclusions live, and the second fix round closed the threshold route to a manufactured pass, but the exclude route is equally powerful and leaves no trace in the durable proof. Concrete sequence: contract hasclean.threshold.crap_max: 15andclean.exclude: ["src/generated/**"]; round 1 measures below threshold; the agent turn (running under--dangerously-skip-permissionsin the worktree) appends- "src/**"toclean.excludeand answers{"action":"excluded"}; the ordinary suite is still green so lines 1036-1040git add -Aand commit the edit; round 2 re-reads the contract at line 947, the threshold comparison at line 954 passes untouched,FM_QUALITY_EXCLUDEat line 706 now carriessrc/**, the phase command finds nothing left to measure and printsoutcome: pass, andwrite_finalat line 989 writes a passing receipt and exits 0.cmd_statusat line 830 accepts that as satisfied andquality_gate_overridein bin/fm-crew-state.sh reportsdone. The loop knows the effective exclude list at line 706 butfinalize_receipt(lines 473-495) never stamps it, and the schema's optionalexclusionsarray (docs/quality-receipt.schema.json:32) is left to whatever the engine chooses to emit, so a pass earned by excluding the whole diff is byte-indistinguishable in the receipt from a pass earned by writing tests. Fix: pass the effectivecfg_seq "$PHASE.exclude"intofinalize_receiptand write it to the receipt'sexclusionsfield, so the durable proof records the bar the run actually measured against. Whether the exclusion power itself should also be bounded is a separate product decision.🔧 Fix: block failed agent turns, stamp exclusions on receipts
2 issues (1 warning, 1 info) still open:
bin/fm-quality.sh:1040- The fix round's new "harness could not run" branch exits the loop without reverting the round, breaking the invariant the header states at lines 76-77 ("Red reverts the round, so a failed round costs budget and nothing else; green commits it").round_headis captured at line 1033, but the branch at lines 1040-1044 callswrite_final ... blocked, which callsfinish, whichexits. The onlygit reset --hard --quiet "$round_head"/git clean -fdqpair is at lines 1058-1059, reached only after the post-turn suite check, which this branch skips entirely.cleanup()(line 126) removes$WORK_DIRand nothing else. Concrete sequence: a hardened ship task, contract withclean.commandandtest; round 1 measures below threshold; the claude turn (running with--dangerously-skip-permissionsin$WT) editssrc/a.ts, then the CLI exits non-zero with nothing readable on stdout (spend ceiling tripped mid-turn, a crash, a session-store error on the round-2--continue);turn_actionis empty andturn_rcis neither 0 nor 124, so the loop writes ablockedreceipt and exits 3 withsrc/a.tsmodified and uncommitted. Consequences: (a) every subsequentfm-quality.sh run <id> --phase cleanrefuses at line 925 ("the task copy has uncommitted changes") until a human cleans by hand, so the hardened gate is permanently unrunnable for that task; (b) bin/fm-teardown.sh:1178 and bin/fm-control.sh:703 now see a dirty copy; (c) if the worker or a human later does the ordinarygit add -A && git commit, a half-finished agent edit that no test ever saw ships silently. Before this fix round the same failure fell through to the suite check and was either reverted or committed under a green suite, so the tree was never left in this state. The siblingdefect-foundearly exit at lines 1045-1049 has the same hole (the turn prompt says "change nothing", but nothing enforces it). Fix:git -C "$WT" reset --hard --quiet "$round_head"plusgit -C "$WT" clean -fdqimmediately before both early-exitwrite_finalcalls, so a round that ends the phase costs budget and nothing else, exactly as the header promises.bin/fm-quality.sh:956- Two wall-clock exits reportexhausted(exit 1, defined at line 34 as "rounds or the wall-clock bound ran out BELOW THRESHOLD") before any measurement has produced a receipt, so the loop asserts something about the code's quality that it never observed. Line 946 fires when the bound is already gone before the ordinary suite runs; line 956 fires withroundstill 1, printing "the 20-minute wall-clock bound ran out after 0 round(s)". Concrete state: contract withbudget_minutes: 20and atestcommand that takes just over 20 minutes and passes green.check_suitereturnsokat t=20m, the loop head computesremaining_seconds<= 0, and the phase reportsoutcome: exhaustedwith exit 1 and no receipt at all. This is the same wrong-label class the third round accepted and fixed insidecheck_suite(its out-of-budget branch at line 784 now saysblocked, and theexhausted*)arm at line 949 was deleted); the sibling path incmd_runwas not covered. The practical blast radius is small: no pass is manufactured, and because no receipt is written,cmd_statusreportsmissingandquality_gate_overridein bin/fm-crew-state.sh reportsblocked: quality gate never ranregardless, so a supervisor still gets the right instruction. Noted as a tradeoff rather than a fix request, since the outcome table at line 34 does list "the wall-clock bound ran out" underexhausted, and deciding whether "ran out before measuring" is exhausted or blocked is the author's call. If it should be blocked, the condition is "no measurement completed yet" (round 1, beforemeasurereturns a receipt), not the bound itself.🔧 Fix: revert the round on every early loop exit
3 issues (2 warnings, 1 info) still open:
bin/fm-quality.sh:753-measurevalidates the receipt against the schema and checksbase_sha(line 748) andphase(line 754), but never checks that the receipt's declaredthresholdmatches the threshold the contract configured. The round-2 fix pinned$PHASE.thresholdinside the contract (line 980) so a round cannot widen its own bar there, but the receipt side is unchecked, so the same manufactured pass is still reachable one step over. Concrete sequence: contract hasclean.threshold.crap_max: 15; the phase command ispnpm run quality:complexity, i.e. a script inside the repo; the agent turn runs with--dangerously-skip-permissionsin$WTand edits that script so it prints{"schema_version":1,"phase":"clean","outcome":"pass","threshold":{"crap_max":9999},"findings":[],...}; the ordinary test suite is unaffected and stays green, so lines 1072-1076git add -Aand commit the edit; round 2 measures, the receipt is schema-valid,base_sha,phaseandhead_shaall match,MEASURE_OUTCOME=pass,write_finalwrites a passing receipt and exits 0;cmd_statusreportsquality: satisfiedandquality_gate_overridein bin/fm-crew-state.sh letsdonestand. A misconfigured (not tampered) project reaches the same place: a phase command that ignoresFM_QUALITY_THRESHOLDand judges against its own baked-in number passes silently, which is the flattering silent failure the intent exists to exclude. Fix: after the phase check at line 754, comparereceipt_field "$out" thresholdwith the contract'scfg_keys_under "$PHASE.threshold"numerically (0.80 and 0.8 must compare equal) and reportblockedon a mismatch, naming both numbers.bin/fm-quality.sh:840-cmd_statusre-validates each phase receipt with"$RECEIPT_TOOL" validate "$file"and no--check-head, and then checks onlybase_sha(line 844) andoutcome(line 848). So a receipt written at an older HEAD keeps satisfying the gate for a tree it never measured. This is reachable on the ordinary hardened path, not an edge case: the hardened brief (bin/fm-brief.sh:527, step 2) tells the worker to finish both loops "before you start on that definition of done - on a no-mistakes task, that means before the pipeline starts", and the no-mistakes pipeline's own fix rounds then commit further changes. Concrete state: clean and harden both pass at HEAD=A and their receipts recordhead_sha: A; the pipeline's fixer commits, HEAD becomes B;fm-quality.sh statusstill reportsquality: satisfied;quality_gate_overridein bin/fm-crew-state.sh returns nothing andstate: donestands - for commits the quality engine never saw. docs/quality-gate.md states--check-headexists precisely so a stale head cannot vouch for the current tree, and the schema madehead_sharequired for the same reason. A naive--check-headhere is not the answer, because the harden loop commits after the clean receipt is written, so the clean receipt's head is legitimately older (that is the stated reason the two phases get separate files). This needs the author's decision on what the gate should promise: re-run the gate after post-gate commits, require the harden receipt's head to be HEAD, or surface the drift in the verdict detail and accept it.bin/fm-quality.sh:489-finalize_receiptsetsdoc["exclusions"] = [line for line in exclusions.splitlines() if line]unconditionally, so whatever the phase command reported in that optional field is replaced by the contract'scfg_seq "$PHASE.exclude"list. Concrete input: a mutation engine that excludes**/*.d.tsinternally and reports"exclusions": ["**/*.d.ts"]in its receipt, under a contract whoseclean.excludeis["src/gen/**"]. The written receipt records only["src/gen/**"], so the durable proof understates the surface the run was actually measured against - the opposite of what the field was added for (header lines 63-65: "Each receipt records the exclude list the run actually measured against"). Fix: union the engine's reported list with the contract's rather than replacing it, preserving order and dropping duplicates. docs/quality-gate.md does not describeexclusionsat all, so naming its owner there would also help.🔧 Fix: verify receipt threshold, merge exclusions, flag stale gates
1 warning still open:
bin/fm-quality.sh:924-cmd_status's new staleness check compares the receipt'shead_shato the worktree HEAD as raw strings, but the receipt schema explicitly allows an abbreviated sha (docs/quality-receipt.schema.json$defs.shais^[0-9a-f]{7,40}$). Any conforming engine that writesgit rev-parse --short HEADtherefore reportssatisfied-staleon every run even when HEAD has not moved once.Reproduced end to end against the real script. Task copy with
.quality-gate.yaml(clean.commandprinting a receipt whosehead_shais the 12-char short sha,clean.threshold.crap_max: 15), metaquality=hardened:bin/fm-quality.sh run task --phase clean->outcome: pass, exit 0. The filed receipt recordshead_sha: 1c27c397c2c6.1c27c397c2c69749719c2cadc3a1162d560deea3; nothing committed after the run.bin/fm-quality.sh status task->quality: satisfied-stale · mode: hardened · the clean receipt measured 1c27c397c2c6, and the task copy is now at 1c27c397c2c69749719c2cadc3a1162d560deea3, so the project's own CI is what proves these commits.quality_gate_overridein bin/fm-crew-state.sh:127 then emitsstate: done · ... · quality gate passed on an earlier commit: ..., filing a false drift claim against a gate that measured the exact tree it is vouching for. That defeats the point of the token the user asked for: a caller branching onsatisfied-stalealone gets the wrong answer for a whole class of conforming engines.The short sha survives the loop because nothing else string-compares
head_sha:measure(lines 764-775) checks onlybase_shaandphase, andvalidate_receipt's--check-headresolves the sha throughgit rev-parse --verify <sha>^{commit}before comparing (bin/fm-quality-receipt.sh:316-331), so it correctly passes.base_shais not affected, becausemeasureexact-compares it at line 765 and refuses an abbreviated one before any receipt is filed.The new tests do not cover this:
test_status_marks_a_receipt_from_an_earlier_head(tests/fm-quality.test.sh) andwrite_clean_receipt(tests/fm-crew-state.test.sh:1660) both write a fullgit rev-parse HEAD.Fix: resolve the recorded sha the same way the validator does before comparing, e.g.
recorded=$(git -C "$WT" rev-parse --verify --quiet "$(receipt_field "$file" head_sha)^{commit}"), and treat an unresolvable sha as not-stale (the--check-headpath already owns that refusal) rather than as drift. Add a test whose receipt carries an abbreviated head_sha at the current HEAD and asserts plainsatisfied.🔧 Fix: compare receipt shas by prefix, not string equality
2 issues (1 warning, 1 info) still open:
bin/fm-quality.sh:781-measurestill compares the receipt'sbase_shato the task's base commit with raw string equality ([ "$receipt_base" != "$BASE_SHA" ]), whilecmd_statuswas just taught the schema-correct prefix rule via the newsame_commithelper (line 445).docs/quality-receipt.schema.json$defs.shais^[0-9a-f]{7,40}$, so a conforming engine may record an abbreviatedbase_sha, and this is exactly the sibling the fix instruction said not to leave wrong ("apply it to any other place ... including the existing base_sha check if it has the same flaw - do not fix one and leave the sibling wrong").Reproduced end to end against the real script. A task copy with
.quality-gate.yaml(clean.commandprinting a valid receipt,clean.threshold.crap_max: 15), metaquality=hardened,base_sha=b92fc56f734ea4cf0351ec85d8390322ca47984c:outcome: pass · phase: clean · mode: hardened · threshold met in 1 round(s), exit 0.base_sha[:12](the same commit) ->outcome: blocked · phase: clean · mode: hardened · the clean receipt measured against b92fc56f734e, not this task's base commit b92fc56f734ea4cf0351ec85d8390322ca47984c, exit 3.The refusal message is self-refuting: it prints a prefix of the very sha it claims is a different commit. The consequence is worse than the
cmd_statusdefect that was just fixed:MEASURE_STATUS=driftreachescmd_run'sno-command|no-receipt|invalid|drift)arm and callsfinish blocked, so a project whose engine abbreviatesbase_shacan never pass the hardened gate at all, andstatus_appendfilesblocked: quality clean could not measure, ...against a healthy crew every attempt. It also leaves the two comparisons inconsistent:cmd_status(line 928) now accepts an abbreviatedbase_shathatmeasurerefuses to file in the first place.The new tests do not cover it.
test_an_abbreviated_base_is_the_same_commit(tests/fm-quality.test.sh:1355) patchesbase_shaonto an already-filed receipt withset_receipt_shaand only exercisescmd_status; the fixture'sFMQ_FORCE_BASEpath is never used with an abbreviated-but-matching value.Fix:
if ! same_commit "$receipt_base" "$BASE_SHA"; thenat line 781, keeping the check pure string work as the helper's comment requires. Add a falsifiable pair in tests/fm-quality.test.sh: a phase command whose receipt names the 12-character short form ofFM_QUALITY_BASE_SHAreportspass, and one whose receipt names a 12-character near miss still reportsblockedwith the drift detail.bin/fm-crew-state.sh:130-quality_gate_override's newsatisfied-stalearm (line 128) correctly guards an empty incoming detail with"${detail:+$detail$SEP}", but the four older arms at lines 130-133 interpolate%s%swith$detailand$SEPunconditionally. Whenemitis called for adonestate with no detail, the returned detail begins with a bare separator andemitthen adds its own, producing an empty field in a line supervision parses.Concrete state: a hardened task (
quality=hardened) whose status log's last line is the degeneratedone:with nothing after the colon.status_line_verb(bin/fm-classify-lib.sh:183) yieldsdone,map_log_statemaps it todone, andstatus_line_note(bin/fm-classify-lib.sh:222) yields the empty string, so line 732 callsemit done status-log "". With no receipt filed,quality_gate_overridereturnsblocked| · quality gate never ran: ...and the emitted line isstate: blocked · source: status-log · · quality gate never ran: this task ships hardened but ...- two consecutive·separators around an empty field. The state token and the reason are both correct, so this is a malformed field rather than a wrong verdict, and it only fires on a status line that is already malformed.Fix: use the same
"${detail:+$detail$SEP}"form in themissing,not-passed,unreadableand catch-all arms as thesatisfied-stalearm already does.🔧 Fix: route all sha compares through one prefix rule
1 info still open:
bin/fm-quality.sh:961-cmd_receiptbuilds$listas a space-joined string and expands it unquoted intopython3 - $list, so any receipt path containing a space is word-split into several arguments. Concrete input:FM_HOME="/home/me/fm home"(or anyFM_DATA_OVERRIDEwith a space) and a task with a filed receipt at$DATA/<id>/quality-clean-receipt.json.fm-quality.sh receipt <id>then calls python with/home/me/fmandhome/data/<id>/quality-clean-receipt.jsonas two separate argv entries, theopen()at the top of the loop raises FileNotFoundError, and the command dies with a Python traceback and exit 1 instead of printing the receipt. Every other path expansion in this script is quoted (mkdir -p "$DATA/$ID",rm -f -- "$(receipt_path "$PHASE")",git -C "$WT"), soreceiptis the only subcommand that breaks on a spacey home, and the SC2086 disable makes the split look intentional rather than a gap. Fix: collect the paths in a bash array (local -a list=();list+=("$(receipt_path "$phase")")) and expand it as"${list[@]}", guarding the empty case sopython3 - <<PYstill prints[].✅ **Test** - passed
✅ No issues found.
bash bin/fm-test-run.sh tests/fm-quality.test.sh tests/fm-quality-receipt.test.sh- 2 scripts, 0 failedbash bin/fm-test-run.sh tests/fm-crew-state.test.sh tests/fm-brief.test.sh tests/fm-control-relaunch.test.sh tests/fm-task-delivery.test.sh- 4 scripts, 0 failedManual end-to-end CLI drive ofbin/fm-quality.sh run <id> --phase cleanover a real throwaway git worktree and real.quality-gate.yaml, once per outcome: pass (exit 0), read-only (exit 6), blocked (exit 3), not-applicable (exit 4), exhausted (exit 1), stuck (exit 1), defect-found (exit 5)Manual wall-clock enforcement check: phase commandsleep 300underbounds.budget_minutes: 1- loop returnedexhaustedafter 65s wall clockManual receipt comparison: hardened vs standard run of the identical measurement, assertingmetricsandthresholdequal whileoutcomediffers ('pass' vs 'read-only')Manual drive ofbin/fm-quality.sh status <id>for the satisfied, not-required and missing verdictsManual drive ofbin/fm-crew-state.sh <id>with an identical passing pipeline run and six different receipt states, showing hardened+pass -> done, hardened+missing -> blocked, hardened+exhausted -> blocked, hardened+read-only -> blocked, hardened+corrupt -> unknown, standard -> done unchangedFalsifiability mutation 1:write_final "$last_receipt" read-only->passin bin/fm-quality.sh;tests/fm-quality.test.shfails on 'a read-only run must not exit with the pass code'Falsifiability mutation 2:finish blocked "$MEASURE_DETAIL"->finish passat bin/fm-quality.sh:1077;tests/fm-quality.test.shfails on 'a read-only run that could not measure is blocked'Falsifiability mutation 3:remaining_seconds()returns constant 86400 at bin/fm-quality.sh:414;tests/fm-quality.test.shfails on 'a measurement that outruns the bound exits 1'Checked no.quality-gate.yamlexists at the firstmate repository root, andkill_rate_minremains 0.80 in docs/quality-gate.mdbin/fm-quality-receipt.sh:294- The two quality scripts compare commit ids by different rules, and the documentation states each correctly, so no doc edit can reconcile them. bin/fm-quality.sh:447 treats a shorter sha as the same commit when it is a prefix of the longer one, because the schema's sha pattern allows 7 to 40 characters. bin/fm-quality-receipt.sh:294 instead requires each verify child's base_sha and head_sha to be byte-equal to the envelope's, which docs/quality-gate.md:152 states as written. A schema-conforming verify receipt whose envelope spells a sha in full while a child abbreviates it is therefore rejected. Reported rather than changed: the receipt schema just landed and the hand-rolled validator is tracked separately as fm-receipt-schema-validator.🔧 **Lint** - 1 issue found → auto-fixed ✅
🔧 Fix: silence deliberate SC2016 hits in quality tests
✅ Re-checked - no issues remain.
✅ **Push** - passed
✅ No issues found.