Skip to content

feat: stamp task milestones, gate promptly and report review coverage - #48

Merged
doitdigital0495 merged 8 commits into
mainfrom
fm/dev-speed-fm-tooling
Oct 2, 2026
Merged

doitdigital0495 merged 8 commits into
mainfrom
fm/dev-speed-fm-tooling

Conversation

@doitdigital0495

@doitdigital0495 doitdigital0495 commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

Intent

The captain asked to implement all changes proposed in the report-development speed analysis (~/.firstmate/data/geris-report-dev-speed-analysis/report.md, section 3). This task covers the two recommendations that live in firstmate tooling: recommendation 3 (gate promptly, avoid stale runs) and recommendation 5 (instrumentation of milestone timestamps). Recommendations 1, 2 and 4 go to a separate fabric_monorepo task.
3. Firstmate/tooling: gate promptly, avoid stale runs. Start after local proof; proposal: dispatch ≤15m. Isolate browser session/tab selection and unique E2E ports. Escalate base drift ≤30m instead of leaving obsolete runs overnight; preserve every managed fix via authorized custody transition. Fix stale coverage reporting; never restart shared no-mistakes daemon or silently skip coverage.
5. Instrumentation: stamp actual brief, local proof, gate start/end/round count, PR creation/merge, deployed SHA, data-ready snapshot, first live proof, UAT merge/deploy/acceptance. Keep status schema strict. Measure next week brief→accepted DEV and accepted DEV→accepted UAT medians, failures and rework count. Proposed targets above are not evidence of savings.
Authorize a Firstmate-side reader in bin/ (in scope) that reports review coverage from the existing status/outcome records; do not touch the external daemon. Proceed and ship.

What Changed

  • Adds bin/fm-task-milestones.sh. It can stamp, list, time, archive and check deadlines for task milestones: stamp, records, durations, coverage, deadlines, archive. Milestones go in as strict milestone [at=] [name=] ... lines in the existing status stream, and status_milestone_record in bin/fm-classify-lib.sh parses them: a closed set of names and fields, full SHAs, forge URLs required for PR and merge, and acceptance counts only with result=passed. The classifier ignores milestone lines when it works out a worker's state or captain relevance. fm-brief.sh stamps a brief milestone when it creates a scaffold. Teardown archives milestone lines to data/<id>/milestones.status so phases stamped after teardown still measure against the brief.

  • Adds fm_nm_review_coverage to bin/fm-nm-run-lib.sh. It reads the no-mistakes run record read-only (state.sqlite, mode=ro) and returns covered only when the run's head and its review-approved head both match the given SHA. Otherwise it returns stale, or unverified if the record can't be read. The daemon is never restarted or written to.

  • The no-mistakes definition of done in bin/fm-dod-lib.sh changes. Briefs and scout promotion now tell workers to:

    • start /no-mistakes themselves within 15 minutes of local proof
    • stamp gate start and end with run id and round count
    • stamp gate-obsolete on base drift and escalate within 30 minutes, keeping managed fixes through an authorized custody transition
    • check review coverage through the new reader
    • use task-owned browser sessions and unique E2E ports

    fm-watch.sh wakes Firstmate once per overdue dispatch or escalation deadline on no-mistakes tasks. Tests, the test-run family mapping, AGENTS.md and docs/scripts.md are updated to match.

🤖 Generated with Claude Code

Risk Assessment

⚠️ Medium: The behavioural change (worker starts its own gate, strict milestone schema kept out of state classification) is bounded and fails closed. But milestones are stored where teardown deletes them, and the coverage reader keys on fields the existing records don't appear to emit, so the required brief->DEV->UAT measurement and the coverage reporting don't work in the normal flow.

Testing

I ran the milestones behavior test (green) and the watcher triage file; its two milestone-deadline cases passed, and the remaining unrelated cases were still running when I stopped reading. I then drove the real CLI end-to-end in a throwaway FM_HOME: brief scaffold stamp, local proof, the 15-minute dispatch deadline, the 30-minute obsolete-run escalation and its clearing, strict-schema refusals, the archive step with late DEV/UAT stamps and readback of both durations, and review coverage read from this run's real daemon record. Everything driven live passed. The watcher deadline wakes, the full fm-teardown.sh path and a watcher under a live Herdr session were not driven live; I exercised them only through the watcher's fake-pane behavior test and the archive subcommand. The temp home was removed.

  • Live validation: ✅ go - 6 of 8 scenarios driven live against the product
Scenario Result Live Evidence
Scaffolding a ship brief with fm-brief.sh stamps the actual brief milestone in the status file ✅ pass live live-cli-transcript.txt section 1: milestone [name=brief] [at=...]: ship brief scaffold created, and records prints the brief row
A local proof with no gate-start reports gate dispatch overdue at 15 minutes and stays quiet before that ✅ pass live live-cli-transcript.txt section 2: empty at 899s, 'gate dispatch overdue' at 1000s, cleared after gate-start
An obsolete gate run escalates at 30 minutes and clears once a successor run starts ✅ pass live live-cli-transcript.txt section 2: empty at 850s, 'obsolete gate escalation due: run run-A' at 1800s, empty after run-B gate-start
Adversarial: the strict schema refuses a merge without a URL, a short SHA, a waived-result acceptance, and a stamp for a task that was never briefed (no orphan file) ✅ pass live live-cli-transcript.txt section 3: exit=1 for each, orphan file exists=no
Milestones survive the teardown archive; later DEV/UAT stamps append to the archive and durations report brief->accepted DEV and DEV->accepted UAT ✅ pass live live-cli-transcript.txt section 4: status file not recreated, archive holds all lines, durations row live-demo 1 2 2 unknown ... (gate rounds unknown because run-A never ended, as designed)
The review coverage reader reports covered/stale/unverified from the real no-mistakes daemon record without touching the daemon ✅ pass live live-cli-transcript.txt section 5: covered at target SHA, stale at base SHA, unverified for an unknown run and a missing NM_HOME (read-only sqlite)
The watcher wakes Firstmate once for an overdue dispatch or an obsolete run (any tag order, malformed records fail closed) without the worker; fresh tasks and non-gated modes stay quiet ⏸️ untested no The prior payload established this only through tests/fm-watch-triage.test.sh with fake pane endpoints, not a live result. A live run needs a provisioned fm-lab-* session via bin/fm-herdr-lab.sh with…
The real fm-teardown.sh run archives the task's milestones to data/<id>/milestones.status before retiring the status file ⏸️ untested no The prior payload established no live result: only the archive subcommand that teardown invokes was driven. A full teardown needs a live task window/worktree; provide it by running it in a bin/fm-he…
Evidence: Live CLI transcript: brief stamp, deadlines, schema refusals, archive + late stamps, durations, real-daemon coverage

Source: Live CLI transcript: brief stamp, deadlines, schema refusals, archive + late stamps, durations, real-daemon coverage

## 1. real fm-brief.sh scaffold in throwaway FM_HOME
scaffolded: /tmp/fm-live-Xf42/data/live-demo/brief.md (ship, mode=no-mistakes; replace {TASK}, {ASKS}, and {FIRSTMATE_SPEC})
--- status file after scaffold:
milestone [name=brief] [at=1790928898]: ship brief scaffold created
--- records:
brief	1790928898	-	-	-	-	-

## 2. gate-dispatch and obsolete-run deadlines
deadlines (local proof 1000s old, no gate-start):
gate dispatch overdue: no gate-start 15 minutes after local proof at 1790927898
deadlines at 899s (fresh):
(empty=ok)
deadlines after gate-start:
(empty=ok)
stamp exit=0 (obsolete at older ts than start)
deadlines 850s after obsolete:
(empty=ok)
deadlines 1800s after obsolete:
obsolete gate escalation due: run run-A obsolete 30 minutes since 1790928048
deadlines after successor run-B:
(empty=ok)
adversarial: hand-written at-first tag order line + malformed short sha

## 3. adversarial: strict schema rejects
fm-task-milestones: invalid milestone; see --help
exit=1
fm-task-milestones: invalid milestone; see --help
exit=1
fm-task-milestones: invalid milestone; see --help
exit=1
fm-task-milestones: no status file or milestone archive for this task
orphan stamp exit=1 exists=no

## 4. merge, teardown archive, late DEV/UAT stamps, durations
status file recreated after teardown? no
--- durable archive data/live-demo/milestones.status:
milestone [name=brief] [at=1790928898]: ship brief scaffold created
milestone [name=local-proof] [at=1790927898] [sha=e317f47c51144db87b9dfe371ac2519b508e8c6f]: tests/fm-task-milestones.test.sh green
milestone [name=gate-start] [at=1790927998] [sha=e317f47c51144db87b9dfe371ac2519b508e8c6f] [run=run-A]: gate dispatched
milestone [name=gate-obsolete] [at=1790927048] [sha=e317f47c51144db87b9dfe371ac2519b508e8c6f] [run=run-A]: base moved
milestone [name=gate-obsolete] [at=1790928048] [sha=e317f47c51144db87b9dfe371ac2519b508e8c6f] [run=run-A]: base moved
milestone [name=gate-start] [at=1790928798] [sha=e317f47c51144db87b9dfe371ac2519b508e8c6f] [run=run-B]: successor run
milestone [name=gate-end] [at=1790928848] [sha=e317f47c51144db87b9dfe371ac2519b508e8c6f] [run=run-B] [rounds=2] [result=passed]: passed
milestone [name=pr-created] [sha=e317f47c51144db87b9dfe371ac2519b508e8c6f] [at=1790928899]: https://github.com/doitdigital0495/firstmate/pull/999
milestone [name=merge] [sha=e317f47c51144db87b9dfe371ac2519b508e8c6f] [at=1790928899]: https://github.com/doitdigital0495/firstmate/pull/999
milestone [name=dev-deploy] [sha=e317f47c51144db87b9dfe371ac2519b508e8c6f] [at=1790928899]: late phase observed
milestone [name=data-ready] [sha=e317f47c51144db87b9dfe371ac2519b508e8c6f] [snapshot=snap1] [at=1790928899]: late phase observed
milestone [name=first-live-proof] [sha=e317f47c51144db87b9dfe371ac2519b508e8c6f] [snapshot=snap1] [result=passed] [at=1790928899]: late phase observed
milestone [name=dev-accepted] [sha=e317f47c51144db87b9dfe371ac2519b508e8c6f] [snapshot=snap1] [result=passed] [at=1790928899]: late phase observed
milestone [name=uat-merge] [sha=e317f47c51144db87b9dfe371ac2519b508e8c6f] [at=1790928901]: https://github.com/doitdigital0495/firstmate/pull/1000
milestone [name=uat-deploy] [sha=e317f47c51144db87b9dfe371ac2519b508e8c6f] [at=1790928901]: late phase observed
milestone [name=uat-data-ready] [sha=e317f47c51144db87b9dfe371ac2519b508e8c6f] [snapshot=snap1] [at=1790928901]: late phase observed
milestone [name=uat-live-proof] [sha=e317f47c51144db87b9dfe371ac2519b508e8c6f] [snapshot=snap1] [result=passed] [at=1790928901]: late phase observed
milestone [name=uat-accepted] [sha=e317f47c51144db87b9dfe371ac2519b508e8c6f] [snapshot=snap1] [result=passed] [at=1790928901]: late phase observed
--- durations:
task	brief_to_accepted_dev_seconds	accepted_dev_to_accepted_uat_seconds	gate_runs	gate_rounds	rework_rounds	failed_gates
live-demo	1	2	2	unknown	unknown	unknown

## 5. review coverage from REAL no-mistakes daemon run record (read-only) for this run
this run @ target e317f47c51144db87b9dfe371ac2519b508e8c6f: covered
this run @ base sha: stale
unknown run: unverified
missing NM_HOME: unverified
Evidence: Milestones behavior test log

Source: Milestones behavior test log

ok - strict milestone schema refuses malformed epochs, duplicates, extra/missing fields and false acceptance without writing
ok - readback distinguishes deploy/data/live/acceptance and deduplicates gate rounds and failures by run
ok - missing, malformed, reversed, partial, waived and mismatched phase evidence stays unknown
ok - milestone records never hide worker states, open decisions or terminal outcomes
ok - review coverage is covered only when the run record reviewed its exact current head
ok - teardown archive keeps milestones durable and late stamps append there without an orphan status log
ok - dispatch and obsolete-run targets are readable at exact 15/30-minute boundaries
ok - writers stamp actual clock and successful brief creation exactly once
EXIT=0
Evidence: Watcher triage test log (milestone deadline cases at lines 19-20)

Source: Watcher triage test log (milestone deadline cases at lines 19-20)

ok - status_span_has_actionable: benign absorbed, captain events surfaced, classified events not re-fired
ok - an actionable event is not hidden by later routine appends, and is named as itself
ok - span classification retires closed decisions and surfaces rejected transitions for reconciliation
ok - a malformed seen signature causes the whole status log to be classified
ok - stale_is_terminal: terminal status surfaces, non-terminal and no-status are benign
ok - classifier primitives: keyed decisions and activity phases, captain relevance, window-to-task, and overrides
ok - crew_is_provably_working: only working+run-step/pane is provable; idle/finished/parked/failed/unknown surface
ok - status_is_paused: only the leading paused verb matches, paused is not captain-relevant, and the two declared-wait verbs stay separable
ok - crew_absorb_class: working/paused/none from one read; crew_is_paused and crew_is_provably_working agree
ok - crew_worktree_written_since: real writes are evidence; no worktree, no anchor, quiet trees, .git churn and a mate's own home are not
ok - an empty FM_WORKTREE_WRITE_PRUNE widens the probe to the whole depth-bounded tree instead of disabling it
ok - an empty FM_WORKTREE_WRITE_PRUNE exported into the environment prunes nothing, widening the probe
ok - the worktree write probe is wall-clock bounded, and hitting the bound reads as no write evidence
ok - signal_crew_provably_working: benign only when every referenced crew is provably working
ok - a secondmate's status signal is never absorbed as provably working; crewmates are unaffected
ok - a no-verb signal whose crew is provably working is absorbed (no exit, no queue, suppressor advanced, beacon present)
ok - a bare turn-end whose crew is provably working (busy pane) is absorbed
ok - a bare turn-end whose crew is not provably working is surfaced (the swallowed-finish fix)
ok - an overdue gate dispatch wakes Firstmate from the watcher once, without the worker
ok - obsolete runs and unreadable milestones wake Firstmate; fresh and non-gated tasks stay quiet, whatever the tag order
ok - a bare turn-end from a pane that churned since the previous poll is absorbed
ok - pane churn starts a fresh stale-classification interval before a stopped render returns
ok - pane churn resets prior wedge escalation state before the stale-path poll
ok - a churning turn-end inside an already-open deferral window is absorbed without re-marking
ok - a bare turn-end from a pane unchanged since the previous poll still surfaces
ok - a bare turn-end backed by a malformed prior hash surfaces
ok - a bare turn-end backed by a newline-terminated prior hash surfaces
ok - a churning secondmate turn-end surfaces without a stale resurface path
ok - a turn-end whose marker key matches another recorded endpoint surfaces
ok - two metadata records sharing one endpoint make churn evidence ambiguous
ok - a batch may satisfy positive evidence independently per task
ok - per-task evidence composition stays off until the home opts in
ok - a status-bearing batch never falls through to pane-churn evidence
ok - pane-churn turn-end absorb is off until a home opts in
ok - a perpetually churning pane surfaces once its bounded deferral window is spent
ok - an unrecordable pane-churn deadline surfaces the turn-end
ok - an invalid pane-churn bound surfaces the turn-end
ok - an oversized pane-churn bound surfaces the turn-end
ok - invalid existing pane-churn deadlines surface without mutation
ok - a surfaced batch opens no partial pane-churn deadline
ok - a no-verb working: note whose crew is idle with no running pipeline is surfaced
ok - a secondmate's status note surfaces even while its own agent is busy
ok - a secondmate blocker wakes despite busy evidence and later unrelated appends
ok - a self-announced close never wakes its own home, and the next real note still does
ok - a close after OPEN DECISIONS fold never wakes its own home, and the next real note still does
ok - a close after OPEN DECISIONS fold still surfaces a worker failure inside the folded span
ok - a close after OPEN DECISIONS fold still surfaces unlisted secondmate lines inside the folded span
ok - captain-relevant signal is surfaced (queue + exit) and marked surfaced
ok - a needs-decision signal row's queued payload is marked needs-decision: for branch exclusion
ok - a reconciliation-required needs-decision row's queued payload is still marked needs-decision:
ok - a captain-held signal stays actionable while the crew is still working
ok - a pending-reply second-mate escalation is marked for main-only routing
ok - an ordinary blocked event remains branch-eligible
ok - a routine event containing a needs-decision phrase keeps its ordinary payload, unmarked
ok - a captain event hidden behind a later routine append is still surfaced (queue + exit)
ok - a finished release reported before routine cleanup chatter is still surfaced
ok - a routine append after an already-classified event is absorbed (no re-wake)
ok - unreadable status reports are bounded without advancing classification
ok - permission recovery surfaces content from the unadvanced position
ok - a stale pane sitting on a terminal status is surfaced (queue + exit)
ok - a stale terminal-looking status is overridden and absorbed while a run is actively working, then wedge-escalated
ok - provably-working non-terminal stale is absorbed on first sight, then wedge-escalated past the threshold
ok - consecutive wedge escalations on the same pane accumulate and demand deep inspection at the threshold
ok - a pane becoming active again resets the consecutive wedge-escalation counter
ok - a record whose endpoint is dead or missing reports itself once and is never re-escalated
ok - a live wedged agent, an unattributable one, and an unreadable endpoint escalate unchanged
ok - the once-only gone report re-arms when the endpoint comes back, and reports a later death again
ok - a second death after a same-window relaunch reports in full without a live probe, and an unchanged dead pane stays silent
ok - a successor's byte-identical dead display reports in full, and the same incarnation still absorbs
ok - a busy worker below the turn-age bound remains working with no escalation
ok - a busy worker with a stable pane hash still escalates once its completed-turn age reaches the bound
ok - a busy worker whose pane hash changes every poll still escalates once its completed-turn age reaches the bound
ok - touching a busy worker's completed-turn marker resets the age and prevents an old-age escalation
ok - native progress resets busy age without a completed turn or notification
ok - repeated busy turn-age escalations reuse the existing escalation counter and demand deep inspection at the threshold
ok - the production default busy-turn-age bound is 3600s (5min under does not wedge, 66min over does)
ok - a busy pane under a declared pause is rechecked on the long cadence, and lifting the pause restores the wedge escalation
ok - away mode hands a busy declared pause to the daemon as a plain stale, and lifting the declaration restores the wedge escalation
ok - away mode wakes the daemon once per declaration for a busy pane whose footer ticks on every capture
ok - a not-provably-working non-terminal stale is surfaced immediately (never left to wait out the timer)
ok - a declared pause is absorbed on first sight, then re-surfaced as a recheck past the threshold, never wedge-escalated
ok - exited declared-pause and captain-held panes use bounded pause cadence while a live decision gate still surfaces once
ok - absorbed paused and captain-held replacements each start their own re-surface cadence
ok - a parked live worker surfaces once, absorbs pane churn for the whole re-surface window, then re-surfaces when it elapses
ok - a live paused worker stays absorbed until its declared time, then rechecks
Terminated
~/.no-mistakes/worktrees/cf62ac6bfd91/01M3XN1FX4Y825QBEGKG6GJXEE/bin/fm-wake-lib.sh: line 581: [: : integer expression expected
Evidence: Real daemon coverage results
this run @ target e317f47: covered
this run @ base 837f39d: stale
unknown run: unverified
missing NM_HOME: unverified

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 1 info
  • 🚨 bin/fm-task-milestones.sh:125 - Milestones are only stored in $STATE/&lt;id&gt;.status, and teardown deletes that file (bin/fm-teardown.sh:945/3851 -> status_retire_presentation_task, bin/fm-classify-lib.sh:1647 rm -f &#34;$state/$task.status&#34;). In a normal no-mistakes ship, Firstmate tears the task down after the PR merges. DEV deploy, data-ready, first live proof, dev-accepted and every UAT phase happen after that. The DoD tells Firstmate or the release/proof owner to stamp those 'after your delivery stop point using the same helper' (bin/fm-dod-lib.sh:213). Concrete sequence: brief stamped at scaffold, gate stamps by the worker, merge, then teardown removes the status file. A later stamp ... dev-accepted re-creates an orphan status log (stamp only checks that the parent dir exists, line ~81) with no brief record. durations then prints unknown for brief_to_accepted_dev and accepted_dev_to_accepted_uat for every task, so the intent's REQUIRED measurement ('Measure next week brief->accepted DEV and accepted DEV->accepted UAT medians, failures and rework count') cannot be produced in the intended flow. The re-created file also shows up as an 'Orphan status log' in fm-session-start.sh:871. The same orphan effect applies to briefs scaffolded but never spawned, because fm-brief.sh:266 stamps before spawn and only teardown retires the file. The remedy is a durable milestone record that survives teardown (for example data/<id>/ or an archive step in teardown), which adds new durable state, so it needs authorization.
  • ⚠️ bin/fm-nm-run-lib.sh:296 - fm_nm_review_coverage reports a verdict only when the captured AXI output contains head_sha, review_coverage, review_coverage_run and review_coverage_head_sha. The existing no-mistakes readers in this repo use id, branch, head, outcome and status (bin/fm-teardown.sh:1904-1911, bin/fm-nm-run-lib.sh:339), never head_sha, and no code or doc outside the new test fixture shows the daemon emitting the review_coverage* fields. With real axi status or outcome output, the head_sha count check returns unverified on every run. The status-milestone path can only become covered if a worker stamps review-coverage result=covered, and the brief tells the worker to take that verdict from this same reader, so the loop never yields covered. The intent asks for 'a Firstmate-side reader in bin/ ... that reports review coverage from the existing status/outcome records'. This reader reads fields the existing records apparently lack, so every gated task will be reported 'not covered' and escalated. It fails closed, but it carries no information. Deciding which real record fields prove coverage is a product and contract decision.
  • ⚠️ bin/fm-task-milestones.sh:166 - The deadlines subcommand (15-minute dispatch and 30-minute obsolete-run escalation) has no caller: no watcher, wake path or brief references it (rg finds it only in the script, its test and the docs row). The failure the intent targets, obsolete runs left overnight, happens when the worker is idle. Escalation still depends on that same worker stamping gate-obsolete and escalating by itself (bin/fm-dod-lib.sh:261), so this reader enforces nothing. Either wire it into Firstmate supervision so stale runs are escalated without the worker, or remove it as an unrequired component.
  • ℹ️ bin/fm-dod-lib.sh:353 - fm_dod_block recomputes root, helper and status (lines 352-354), and fm_ship_milestone_contract already computes the same values (lines 308-310). Two copies of the helper and status-path rule can drift apart. Have fm_ship_milestone_contract be the only owner (for example also print or export the coverage line) and drop the duplicate locals.

🔧 Fix applied.
5 issues (1 error, 3 warnings, 1 info) still open:

  • 🚨 bin/fm-task-milestones.sh:125 - Milestones are only stored in $STATE/&lt;id&gt;.status, and teardown deletes that file (bin/fm-teardown.sh:945/3851 -> status_retire_presentation_task, bin/fm-classify-lib.sh:1647 rm -f &#34;$state/$task.status&#34;). In a normal no-mistakes ship, Firstmate tears the task down after the PR merges. DEV deploy, data-ready, first live proof, dev-accepted and every UAT phase happen after that. The DoD tells Firstmate or the release/proof owner to stamp those 'after your delivery stop point using the same helper' (bin/fm-dod-lib.sh:213). Concrete sequence: brief stamped at scaffold, gate stamps by the worker, merge, then teardown removes the status file. A later stamp ... dev-accepted re-creates an orphan status log (stamp only checks that the parent dir exists, line ~81) with no brief record. durations then prints unknown for brief_to_accepted_dev and accepted_dev_to_accepted_uat for every task, so the intent's REQUIRED measurement ('Measure next week brief->accepted DEV and accepted DEV->accepted UAT medians, failures and rework count') cannot be produced in the intended flow. The re-created file also shows up as an 'Orphan status log' in fm-session-start.sh:871. The same orphan effect applies to briefs scaffolded but never spawned, because fm-brief.sh:266 stamps before spawn and only teardown retires the file. The remedy is a durable milestone record that survives teardown (for example data/<id>/ or an archive step in teardown), which adds new durable state, so it needs authorization.
  • ⚠️ bin/fm-task-milestones.sh:166 - The deadlines subcommand (15-minute dispatch and 30-minute obsolete-run escalation) has no caller: no watcher, wake path or brief references it (rg finds it only in the script, its test and the docs row). The failure the intent targets, obsolete runs left overnight, happens when the worker is idle. Escalation still depends on that same worker stamping gate-obsolete and escalating by itself (bin/fm-dod-lib.sh:261), so this reader enforces nothing. Either wire it into Firstmate supervision so stale runs are escalated without the worker, or remove it as an unrequired component.
  • ℹ️ bin/fm-dod-lib.sh:353 - fm_dod_block recomputes root, helper and status (lines 352-354), and fm_ship_milestone_contract already computes the same values (lines 308-310). Two copies of the helper and status-path rule can drift apart. Have fm_ship_milestone_contract be the only owner (for example also print or export the coverage line) and drop the duplicate locals.
  • ⚠️ bin/fm-watch.sh:2694 - This is the watcher wake that fix round 1 added. It runs deadlines for every task whose status contains a local-proof milestone, whatever the task's delivery mode. The shared milestone contract (bin/fm-dod-lib.sh:225, fm_ship_milestone_contract) is rendered into the direct-PR and local-only DoD blocks too (bin/fm-dod-lib.sh:368,380), and it tells those workers to stamp local proof. Those modes never produce a gate-start. Concrete sequence: a direct-PR worker stamps local-proof, then works on its PR for 15 minutes. The watcher queues a check: &lt;id&gt; gate dispatch overdue wake, and Firstmate chases a gate that should never run. The deadlines dispatch rule in bin/fm-task-milestones.sh:186 has the same mode-blindness. Fix: run the dispatch check only when the task meta records mode=no-mistakes (fm-promote writes mode= into meta, line 287), or skip the dispatch rule for other modes.
  • ⚠️ bin/fm-task-milestones.sh:177 - deadlines aborts as soon as any milestone line in the file is malformed, because records returns 1 on the first bad line. The watcher (bin/fm-watch.sh:2695) discards that failure with 2&gt;/dev/null || due=, so the 15-minute and 30-minute escalation is silently turned off for that task. Concrete sequence: a worker writes a milestone by hand instead of using the helper, for example milestone [name=gate-start] [sha=abc1234] ... with a short SHA. Then a valid gate-obsolete stamp never produces the 30-minute wake. This contradicts the intent's 'escalate base drift ≤30m'. Fail closed instead: either deadlines skips malformed records and still evaluates the valid ones, or the watcher raises a check wake that names the malformed milestone instead of swallowing the error. This is fix round 1 code.

🔧 No changes applied.
7 issues (1 error, 4 warnings, 2 infos) still open:

  • 🚨 bin/fm-task-milestones.sh:125 - Milestones are only stored in $STATE/&lt;id&gt;.status, and teardown deletes that file (bin/fm-teardown.sh:945/3851 -> status_retire_presentation_task, bin/fm-classify-lib.sh:1647 rm -f &#34;$state/$task.status&#34;). In a normal no-mistakes ship, Firstmate tears the task down after the PR merges. DEV deploy, data-ready, first live proof, dev-accepted and every UAT phase happen after that. The DoD tells Firstmate or the release/proof owner to stamp those 'after your delivery stop point using the same helper' (bin/fm-dod-lib.sh:213). Concrete sequence: brief stamped at scaffold, gate stamps by the worker, merge, then teardown removes the status file. A later stamp ... dev-accepted re-creates an orphan status log (stamp only checks that the parent dir exists, line ~81) with no brief record. durations then prints unknown for brief_to_accepted_dev and accepted_dev_to_accepted_uat for every task, so the intent's REQUIRED measurement ('Measure next week brief->accepted DEV and accepted DEV->accepted UAT medians, failures and rework count') cannot be produced in the intended flow. The re-created file also shows up as an 'Orphan status log' in fm-session-start.sh:871. The same orphan effect applies to briefs scaffolded but never spawned, because fm-brief.sh:266 stamps before spawn and only teardown retires the file. The remedy is a durable milestone record that survives teardown (for example data/<id>/ or an archive step in teardown), which adds new durable state, so it needs authorization.
  • ⚠️ bin/fm-task-milestones.sh:166 - The deadlines subcommand (15-minute dispatch and 30-minute obsolete-run escalation) has no caller: no watcher, wake path or brief references it (rg finds it only in the script, its test and the docs row). The failure the intent targets, obsolete runs left overnight, happens when the worker is idle. Escalation still depends on that same worker stamping gate-obsolete and escalating by itself (bin/fm-dod-lib.sh:261), so this reader enforces nothing. Either wire it into Firstmate supervision so stale runs are escalated without the worker, or remove it as an unrequired component.
  • ℹ️ bin/fm-dod-lib.sh:353 - fm_dod_block recomputes root, helper and status (lines 352-354), and fm_ship_milestone_contract already computes the same values (lines 308-310). Two copies of the helper and status-path rule can drift apart. Have fm_ship_milestone_contract be the only owner (for example also print or export the coverage line) and drop the duplicate locals.
  • ⚠️ bin/fm-watch.sh:2694 - This is the watcher wake that fix round 1 added. It runs deadlines for every task whose status contains a local-proof milestone, whatever the task's delivery mode. The shared milestone contract (bin/fm-dod-lib.sh:225, fm_ship_milestone_contract) is rendered into the direct-PR and local-only DoD blocks too (bin/fm-dod-lib.sh:368,380), and it tells those workers to stamp local proof. Those modes never produce a gate-start. Concrete sequence: a direct-PR worker stamps local-proof, then works on its PR for 15 minutes. The watcher queues a check: &lt;id&gt; gate dispatch overdue wake, and Firstmate chases a gate that should never run. The deadlines dispatch rule in bin/fm-task-milestones.sh:186 has the same mode-blindness. Fix: run the dispatch check only when the task meta records mode=no-mistakes (fm-promote writes mode= into meta, line 287), or skip the dispatch rule for other modes.
  • ⚠️ bin/fm-task-milestones.sh:177 - deadlines aborts as soon as any milestone line in the file is malformed, because records returns 1 on the first bad line. The watcher (bin/fm-watch.sh:2695) discards that failure with 2&gt;/dev/null || due=, so the 15-minute and 30-minute escalation is silently turned off for that task. Concrete sequence: a worker writes a milestone by hand instead of using the helper, for example milestone [name=gate-start] [sha=abc1234] ... with a short SHA. Then a valid gate-obsolete stamp never produces the 30-minute wake. This contradicts the intent's 'escalate base drift ≤30m'. Fail closed instead: either deadlines skips malformed records and still evaluates the valid ones, or the watcher raises a check wake that names the malformed milestone instead of swallowing the error. This is fix round 1 code.
  • ⚠️ bin/fm-task-milestones.sh:63 - This is in fix round 2 code (milestones-lost-at-teardown). archive_for finds the durable archive from the caller's environment (FM_DATA_OVERRIDE, else FM_HOME/data, else the script's repo root). It does not use the FILE argument the DoD hands out. When that lookup misses, stamp quietly falls back to the status path and recreates the orphan status log the fix was meant to stop (lines 95 and 105-108). Concrete sequence: the Geris home has no bin/ of its own, so its briefs render the helper as ~/.firstmate/bin/fm-task-milestones.sh and the status file as ~/.firstmate-geris/state/&lt;id&gt;.status. Teardown archives to ~/.firstmate-geris/data/&lt;id&gt;/milestones.status. A release or proof owner then runs the exact rendered command after teardown without FM_HOME=~/.firstmate-geris exported. archive_for resolves to ~/.firstmate/data/&lt;id&gt;/..., which does not exist. stamp appends dev-accepted to a new ~/.firstmate-geris/state/&lt;id&gt;.status that has no brief record. durations then prints unknown, and session-start lists the file as an orphan. The same env-derived lookup is used by records, durations and deadlines (line 82). The fix-round test pins FM_DATA_OVERRIDE (tests/fm-task-milestones.test.sh:9), so it cannot catch this. It also calls archive directly and never runs the real teardown. Fix: derive the archive from FILE itself ($(dirname FILE)/../data/&lt;id&gt;/milestones.status unless an override is set). Also make stamp refuse when neither the status file nor the archive exists, instead of creating a new status log.
  • ℹ️ tests/fm-watch-triage.test.sh:2142 - This is the fix round 2 test for deadlines-unwired. It only covers the 15-minute dispatch wake through the watcher. The user's fix instruction asked for proof that the 30-minute gate-obsolete escalation is raised without the worker, and that a fresh task produces no wake. Neither is checked at the watcher level. The deadlines subcommand logic itself is exercised in tests/fm-task-milestones.test.sh. Add a watcher case with a 31-minute-old gate-obsolete stamp, and one with a fresh local-proof that must not wake.

🔧 Fix applied.
3 issues (2 warnings, 1 info) still open:

  • ℹ️ bin/fm-dod-lib.sh:353 - fm_dod_block recomputes root, helper and status (lines 352-354), and fm_ship_milestone_contract already computes the same values (lines 308-310). Two copies of the helper and status-path rule can drift apart. Have fm_ship_milestone_contract be the only owner (for example also print or export the coverage line) and drop the duplicate locals.
  • ⚠️ bin/fm-watch.sh:2694 - This is the watcher wake that fix round 1 added. It runs deadlines for every task whose status contains a local-proof milestone, whatever the task's delivery mode. The shared milestone contract (bin/fm-dod-lib.sh:225, fm_ship_milestone_contract) is rendered into the direct-PR and local-only DoD blocks too (bin/fm-dod-lib.sh:368,380), and it tells those workers to stamp local proof. Those modes never produce a gate-start. Concrete sequence: a direct-PR worker stamps local-proof, then works on its PR for 15 minutes. The watcher queues a check: &lt;id&gt; gate dispatch overdue wake, and Firstmate chases a gate that should never run. The deadlines dispatch rule in bin/fm-task-milestones.sh:186 has the same mode-blindness. Fix: run the dispatch check only when the task meta records mode=no-mistakes (fm-promote writes mode= into meta, line 287), or skip the dispatch rule for other modes.
  • ⚠️ bin/fm-watch.sh:2697 - Fix round 2 added a grep pre-filter in the watcher, and it only matches milestone lines where name is the first tag (^milestone \[name=(local-proof|gate-obsolete)\]). The strict parser accepts tags in any order, and the schema comment documents at first (bin/fm-classify-lib.sh:417: milestone [at=&lt;epoch&gt;] [name=&lt;name&gt;] ...). Concrete sequence: a worker or release owner writes a schema-valid line by hand, such as milestone [at=1759400000] [name=gate-obsolete] [sha=&lt;40hex&gt;] [run=r1]: base moved. records/deadlines parse that line correctly, but the grep never matches it, so the watcher never runs deadlines and the 30-minute obsolete-run escalation silently never fires. This is a sibling of the fail-closed deadline fix from round 2 (deadlines-fail-open-on-malformed), which that round left behind. Fix: loosen the pre-filter to ^milestone .*\[name=(local-proof|gate-obsolete)\], or drop it and rely on deadlines printing nothing when no relevant record exists.

🔧 Fix applied.
2 issues (1 warning, 1 info) still open:

  • ℹ️ bin/fm-dod-lib.sh:353 - fm_dod_block recomputes root, helper and status (lines 352-354), and fm_ship_milestone_contract already computes the same values (lines 308-310). Two copies of the helper and status-path rule can drift apart. Have fm_ship_milestone_contract be the only owner (for example also print or export the coverage line) and drop the duplicate locals.
  • ⚠️ bin/fm-task-milestones.sh:84 - Round 1 wired deadlines into the watcher (bin/fm-watch.sh:2696), and round 2 widened its pre-filter. It now runs on every poll (POLL=15s) for every no-mistakes task whose status holds a local-proof or gate-obsolete line, for the rest of that task's life, because the line is never removed. records forks a command-substitution subshell for status_line_verb on every status line. status_milestone_record adds another fork per milestone line, with nested $(status_line_at_epoch ...) forks inside it. Concrete case: a few gated tasks with append-only status files of a few hundred lines cost several hundred forks per task, every cycle, long after the deadline is satisfied. That adds up to seconds of watcher cycle time. Fix: use the existing out-var form status_line_verb &#34;$line&#34; verb at line 84 and at the identical archive loop at line 123, so non-milestone lines cost no fork. Optionally skip calling deadlines once the stored deadline-marker state shows nothing is pending.

🔧 Fix applied.
1 info still open:

  • ℹ️ bin/fm-dod-lib.sh:353 - fm_dod_block recomputes root, helper and status (lines 352-354), and fm_ship_milestone_contract already computes the same values (lines 308-310). Two copies of the helper and status-path rule can drift apart. Have fm_ship_milestone_contract be the only owner (for example also print or export the coverage line) and drop the duplicate locals.
✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 6 of 8 scenarios driven live against the product
Scenario Result Live Evidence
Scaffolding a ship brief with fm-brief.sh stamps the actual brief milestone in the status file ✅ pass live live-cli-transcript.txt section 1: milestone [name=brief] [at=...]: ship brief scaffold created, and records prints the brief row
A local proof with no gate-start reports gate dispatch overdue at 15 minutes and stays quiet before that ✅ pass live live-cli-transcript.txt section 2: empty at 899s, 'gate dispatch overdue' at 1000s, cleared after gate-start
An obsolete gate run escalates at 30 minutes and clears once a successor run starts ✅ pass live live-cli-transcript.txt section 2: empty at 850s, 'obsolete gate escalation due: run run-A' at 1800s, empty after run-B gate-start
Adversarial: the strict schema refuses a merge without a URL, a short SHA, a waived-result acceptance, and a stamp for a task that was never briefed (no orphan file) ✅ pass live live-cli-transcript.txt section 3: exit=1 for each, orphan file exists=no
Milestones survive the teardown archive; later DEV/UAT stamps append to the archive and durations report brief->accepted DEV and DEV->accepted UAT ✅ pass live live-cli-transcript.txt section 4: status file not recreated, archive holds all lines, durations row live-demo 1 2 2 unknown ... (gate rounds unknown because run-A never ended, as designed)
The review coverage reader reports covered/stale/unverified from the real no-mistakes daemon record without touching the daemon ✅ pass live live-cli-transcript.txt section 5: covered at target SHA, stale at base SHA, unverified for an unknown run and a missing NM_HOME (read-only sqlite)
The watcher wakes Firstmate once for an overdue dispatch or an obsolete run (any tag order, malformed records fail closed) without the worker; fresh tasks and non-gated modes stay quiet ⏸️ untested no The prior payload established this only through tests/fm-watch-triage.test.sh with fake pane endpoints, not a live result. A live run needs a provisioned fm-lab-* session via bin/fm-herdr-lab.sh with…
The real fm-teardown.sh run archives the task's milestones to data/<id>/milestones.status before retiring the status file ⏸️ untested no The prior payload established no live result: only the archive subcommand that teardown invokes was driven. A full teardown needs a live task window/worktree; provide it by running it in a bin/fm-he…
  • bash tests/fm-task-milestones.test.sh (all 8 behavior cases ok)
  • bash tests/fm-watch-triage.test.sh - the cases test_milestone_deadline_surfaced_once and test_milestone_deadline_cases passed (log lines 19-20)
  • Manual: FM_HOME=&lt;tmp&gt; bin/fm-brief.sh live-demo repo --mode no-mistakes, then checked the status file and records
  • Manual: fm-task-milestones.sh stamp/deadlines at 899s, 1000s, 850s and 1800s, before and after gate-start and the successor run
  • Manual adversarial: merge without a URL, short SHA, acceptance with a waived result, stamp for a task that was never briefed
  • Manual: archive + remove the status file + 9 late DEV/UAT stamps + durations
  • Manual: fm-task-milestones.sh coverage 01M3XN1FX4Y825QBEGKG6GJXEE &lt;target|base&gt; against the real ~/.no-mistakes/state.sqlite (read-only), plus an unknown run and a missing NM_HOME
✅ **Document** - passed

✅ No issues found.

🔧 **Lint** - 1 issue found → auto-fixed ✅
  • ⚠️ linter found issues (exit code 1)

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ **Push** - passed

✅ No issues found.

…. The rule that broke: every test fixture that runs the real bin/fm-teardown.sh from a hand-built fake bin/ must include every sibling script teardown calls. This PR made teardown call bin/fm-task-milestones.sh (the milestone archive step). tests/fm-gotmp.test.sh builds that fake bin/ in two places, make_fake_root and the separate fixture in test_teardown_skips_gracefully_without_tasktmp, and neither had the new script, so teardown exited non-zero. Fix: symlink the real bin/fm-task-milestones.sh into both fixtures. No other test builds a per-file fake bin/ like this. Verified: tests/fm-gotmp.test.sh failed before the fix ("teardown exited non-zero with a valid tasktmp") and passes after it; shellcheck is clean. ci-1 (Behavior portable serial 5): tests/fm-backend.test.sh failed in its symlinked-prefix spawn case with "could not resolve the shared Treehouse project lock". The function behind that error (fm_treehouse_project_lock_path in bin/fm-wake-lib.sh) and its caller bin/fm-spawn.sh are not changed by this PR. The test passed on the base commit's CI, and it passes locally both on its own and in a clean checkout after replaying shard 5's exact test order up to fm-backend. I could not reproduce the failure and nothing in this diff reaches that path, so I made no code change for it. Re-running the shard should settle it; the cause on that CI runner is unconfirmed. In that local replay, fm-secondmate-restart and fm-startup-network also exited 1. Both passed in CI on this shard, so this looks like a difference in the local environment; I did not investigate it
…ng test-isolation bug. The PR's new shard order exposed it; the PR's code did not cause it. Fixed in the test fixture only; no production code changed. Root cause: in tests/fm-backend.test.sh, the spawn cases never set FM_HOME. fm-spawn then falls back to the checkout itself as the home, and fm_treehouse_project_lock_path (bin/fm-wake-lib.sh:1213) requires that home to have a state/ directory. state/ is untracked, so a fresh CI checkout does not have it. The symlinked-prefix case (test_spawn_symlinked_project_prefix_avoids_false_refusal) runs before the FM_HOME='' refusal cases, and those are what happen to create $ROOT/state. So it hit "could not resolve the shared Treehouse project lock". It only passed before because of what earlier tests in the old shard 6 left behind. The earlier rounds could not reproduce it because the local shell exports FM_HOME=~/.firstmate, the operator's real home. That masked the bug and also leaked the real home into the test. Reproducer: with the checkout's state/ removed, `env -u FM_HOME bash tests/fm-backend.test.sh` fails the same way as CI ("not ok - fm-spawn.sh should succeed for a project reached through a symlinked prefix ... could not resolve the shared Treehouse project lock"). Invariant: every fm-spawn call in this suite that reaches the Treehouse project lock must run against an isolated FM_HOME that has a state/ directory. Four places need it: run_spawn_case, used by both symlink cases, and the three inline spawns in test_spawn_default_backend_writes_no_meta_field, test_spawn_explicit_backend_flag_beats_autodetect_herdr_env and test_spawn_autodetect_nesting_resolves_tmux_silently. Those three only passed because of test order. Fix: one throwaway SPAWN_FM_HOME ($TMP_ROOT/fm-home with state/) is created next to SPAWN_HOME and passed as FM_HOME at all four sites. The FM_HOME='' refusal tests exit before the lock and are unchanged. Verified: - The test now passes with FM_HOME unset and no checkout state/ (rc=0, 28 ok). - It also passes with the operator's FM_HOME exported (rc=0). - `shellcheck -x tests/fm-backend.test.sh` is clean. - Only tests/fm-backend.test.sh changed
@doitdigital0495
doitdigital0495 merged commit 50897df into main Oct 2, 2026
21 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants