Skip to content

feat(bin): unify design tasks into one planning conversation - #272

Merged
HelloWorldSungin merged 4 commits into
mainfrom
fm/fm-unified-kun-matt-planning
Sep 14, 2026
Merged

HelloWorldSungin merged 4 commits into
mainfrom
fm/fm-unified-kun-matt-planning

Conversation

@HelloWorldSungin

@HelloWorldSungin HelloWorldSungin commented Sep 11, 2026 •

Copy link
Copy Markdown
Owner

Intent

The captain approved one planning conversation rather than stacked Kun, Matt and ADR workflows. Use Kun research and visual proposals where useful; selectively use Matt dependency-aware questioning and domain-modeling. Retain short committed ADRs only for consequential architectural tradeoffs, not routine configuration changes. Preserve existing decisions and in-flight designs.

What Changed

  • fm-brief.sh --design now writes a brief for one planning conversation instead of stacked Kun, Matt and ADR passes. The worker researches facts first and may use lavish-axi for a visual proposal, but that review is optional. It uses the dispatch-pinned grilling skill only to ask the next unblocked decision, one keyed question at a time, and uses domain-modeling for terms. It must not import grill-with-docs, to-spec, to-tickets, wayfinder, implement or CONTEXT.md. A short ADR is written only when domain-modeling's ADR bar is met. If the change turns out to be routine configuration, the worker stops with a needs-decision instead of writing an ADR.
  • fm-spawn.sh's pinned-skills preamble now says to use those skills inside the one conversation. The design-profile skill, AGENTS.md and the docs now describe Design as an ADR task for consequential architectural tradeoffs. They say routine configuration changes are ships, and that no design interview or visual review goes in front of authorized, well-specified implementation. Live tasks keep their existing briefs and keyed decision inventory.
  • Messages from fm-design-skills.sh, fm-brief.sh and fm-spawn.sh, plus related comments, now call the plugin install "captain-owned" instead of pointing to an auto-updating /plugin action. tests/fm-brief.test.sh, tests/fm-design-skills.test.sh and tests/fm-outcome-manifest.test.sh were updated to match the new brief and error wording.

🤖 Generated with Claude Code

Risk Assessment

✅ Low: Most of the change is prompt, doc and error-message wording. The only runtime effect is new text in the generated design brief and in the design-skill refusal messages. It meets the stated intent: one planning conversation, optional visual proposals, grilling used selectively, ADRs held to the domain-modeling ADR bar, and in-flight briefs left untouched. The tests check that generated output.

Testing

I ran the design and ship scaffolding CLIs in an isolated FM_HOME against the real installed Matt plugin. The design brief carries the one-conversation contract: optional lavish-axi, one unblocked grilling question at a time, domain-modeling's ADR bar with no ceremonial ADRs, and no /plugin updater. The ship brief has none of the planning stack. A missing plugin is refused with captain-owned wording and no brief is written, and an existing in-flight brief is left byte-identical. The targeted tests for design dispatch pinning, relaunch, manifest provenance and the harness matrix all passed, but they use stub harnesses, so those scenarios are not counted as live. Real worker sessions were not launched, so how a model actually follows the brief is still untested. Everything is CLI or text output, so there are no screenshots. The worktree was left clean.

  • Live validation: ✅ go - 5 of 8 scenarios driven live against the product
Scenario Result Live Evidence
Captain scaffolds a design task with the real installed plugin and the brief describes one planning conversation: optional lavish-axi visual, one unblocked grilling question at a time, domain-modeling… ✅ pass live 01-design-scaffold-real-plugin.txt, 02-design-brief.md
The ADR bar and design-tree wording in the brief match what the dispatched real plugin skills actually define (domain-modeling's 3 criteria; grilling's frontier batching is overridden by one keyed que… ✅ pass live grep of ~/.claude/plugins/cache/mattpocock/mattpocock-skills/1.2.3 domain-modeling and grilling SKILL.md against 02-design-brief.md
Adversarial: a well-specified ship brief does not pick up grilling, lavish-axi, the planning conversation, or ADR requirements ✅ pass live 03-ship-refusal-inflight.txt (all planning-stack matches=0)
Adversarial: with no plugin registry, design scaffolding and the resolver refuse with captain-owned refresh wording, write no brief, and never suggest a worker install ✅ pass live 03-ship-refusal-inflight.txt (rc=1, no data dir)
Adversarial: an in-flight design brief generated under the old interactive-ADR contract is not rewritten by a new scaffold ✅ pass live 03-ship-refusal-inflight.txt (sha1 unchanged, 'already exists' rc=1)
fm-spawn --design resolves the plugin once, pins exact skill paths, and adds the one-conversation line and captain-owned blocker to the worker-facing brief ⏸️ untested no The prior payload did not establish a live result: it was covered only by tests/fm-design-skills.test.sh with a stub harness. A live check needs a real design worker launched through fm-spawn (tmux pl…
Relaunching an in-flight design worker keeps its dispatch pin, inbox, and uncommitted work on claude, codex, and pi ⏸️ untested no The prior payload did not establish a live result: it was covered only by tests/fm-control-relaunch.test.sh with stub harnesses. A live check needs a running design worker in a real fleet session to r…
A real design worker holds a single conversation (no separate Kun/Matt workflows) and declines a ceremonial ADR for routine configuration ⏸️ untested no The prior payload only verified the generated prompt (02-design-brief.md). Checking model behavior needs a live, billed design agent session run by the captain on a real design ask.
Evidence: Real plugin resolve + design scaffold transcript
$ bin/fm-design-skills.sh check
mattpocock design skills ready: version=1.2.3 updated=2026-08-19T00:30:39.351Z path=~/.claude/plugins/cache/mattpocock/mattpocock-skills/1.2.3
rc=0

$ bin/fm-design-skills.sh resolve
{
  "schema": "fm-design-skills.v1",
  "plugin": "mattpocock-skills@mattpocock",
  "install_path": "~/.claude/plugins/cache/mattpocock/mattpocock-skills/1.2.3",
  "version": "1.2.3",
  "last_updated": "2026-08-19T00:30:39.351Z",
  "skills": {
    "grilling": "~/.claude/plugins/cache/mattpocock/mattpocock-skills/1.2.3/skills/productivity/grilling/SKILL.md",
    "domain_modeling": "~/.claude/plugins/cache/mattpocock/mattpocock-skills/1.2.3/skills/engineering/domain-modeling/SKILL.md",
    "ask_matt": "~/.claude/plugins/cache/mattpocock/mattpocock-skills/1.2.3/skills/engineering/ask-matt/SKILL.md"
  }
}
rc=0

$ FM_HOME=$H bin/fm-brief.sh design-live sample --design --mode no-mistakes
scaffolded: /tmp/tmp.afhNKB9RXp/data/design-live/brief.md (design, mode=no-mistakes; replace {TASK} and {FIRSTMATE_SPEC})
rc=0
Evidence: Generated design brief (worker-facing prompt)
You are a crewmate: an autonomous worker agent managed by firstmate. Work on your own; do not wait for a human.

# Task
## Captain's intent
{TASK}

## Firstmate spec
{FIRSTMATE_SPEC}

<!-- firstmate-task-branch=fm/design-live -->
# Herdr lifecycle declaration - NOT ENABLED
**HARD SAFETY GATE:** this scaffold cannot inspect the task text filled in above.
If the task will start, stop, delete, restart, profile, or otherwise drive Herdr lifecycle behavior, stop and regenerate the brief with `--herdr-lab` before dispatch.
Do not add Herdr lifecycle commands to this unguarded brief by hand.

# Design profile
This is one DESIGN planning conversation whose only tracked project deliverable is one short architectural decision record.
Do not implement the resulting design or make unrelated project changes.
Do not create or modify any other tracked project file, including `AGENTS.md` or `CLAUDE.md`.
Do not stack a separate Kun workflow, a separate Matt workflow, and this ADR as three passes.

Read and follow `~/.no-mistakes/worktrees/bc8432f7c9f8/01M294Z727X91MCT3V8QZ0EVJP/.agents/skills/design-profile/SKILL.md` before beginning the conversation.
At dispatch Firstmate prepends the exact `grilling` and `domain_modeling` paths from the resolver call whose plugin release it records for this task.
Read only those dispatch-pinned paths, never resolve the plugin again from this worker, and stop with the binding's blocker if either exact file is unavailable.
This direct file-binding contract is identical on Claude, Codex, and Pi and does not depend on harness-specific skill-command spelling.
Never install, update, copy, vendor, pin, or modify that plugin from this task.
Plugin lifecycle is captain-owned outside this repository; do not add a competing updater here.
Use those skills selectively for the design tree, the next unblocked question, and domain terms.
Do not import grill-with-docs, to-spec, to-tickets, wayfinder, implement, or CONTEXT.md from that plugin.
Do not create or update `CONTEXT.md`, even if a dependency instructs you to do so.
Record every resolved term only in the ADR so it remains the sole tracked project deliverable.

Investigate factual questions from repository evidence before asking for a decision.
If an ambiguous choice is clearer as a diagram or interactive proposal, you may use lavish-axi in this same conversation.
Do not require a visual review, and do not start a separate visual-review workflow for a well-specified ADR ask.
Ask exactly one decision question at a time, with one stable key, the evidence, and your recommended answer.
Choose that question as the next unblocked decision on the design tree; do not batch the whole frontier.
Append `needs-decision [key=<stable-slug>]: {one question} Recommendation: {answer and evidence}`, then stop and wait.
Never batch questions, answer on behalf of firstmate, or proceed while the current key is unresolved.
When an answer arrives, append `resolved [key=<same-stable-slug>]: {decision returned by firstmate}` and `working: continuing the design interview` in the same breath, then capture the decision in the ADR.
State the converged decision back to firstmate before drafting the ADR.

Write a short ADR only when the dispatched domain-modeling skill's ADR bar is met.
If the conversation shows a routine configuration change rather than a consequential architectural tradeoff, append `needs-decision` rather than padding a ceremonial ADR.
Use an existing project ADR convention when one exists.
Otherwise use `docs/adr/NNNN-<slug>.md`, incrementing the highest existing number.
The ADR must stand alone with context, decision, rationale, relevant alternatives, and non-obvious consequences.

# Setup
You are in a disposable git worktree of sample, at a detached HEAD on a clean default branch.

**Verify isolation before anything else.** Run `pwd -P` and `git rev-parse --show-toplevel`; both must resolve to the disposable task worktree you were launched in, such as a treehouse pool path or an Orca-managed worktree, not the primary checkout firstmate operates from.
The path check is authoritative: `git rev-parse --git-dir` and `git rev-parse --git-common-dir` can help inspect the repo, but they do not prove you are outside the primary checkout.
If the top-level path is the primary checkout or not the worktree you were launched in, STOP - do not branch or commit here - append `blocked: launched in primary checkout, not an isolated worktree` to the status file and stop.

1. First action: create your branch: `git checkout -b fm/design-live`
2. Run `no-mistakes doctor`; if it reports the repo is not initialized here, run `no-mistakes init`.

# Rules
1. Never push to the default branch. Never merge a PR.
2. Stay inside this worktree; modify nothing outside it.
3. Use gh-axi for GitHub operations and chrome-devtools-axi for browser operations.
4. Report status by appending one line:
   `echo "{state}: {one short line}" >> '/tmp/tmp.afhNKB9RXp/state/design-live.status'`
   States: working, needs-decision, blocked, paused, done, failed.
   Each append wakes firstmate, so report sparingly: only phase changes a supervisor
   would act on (setup done, bug reproduced, fix implemented, validation passed) and the
   needs-decision/blocked/paused/done/failed states. No step-by-step FYI progress lines;
   firstmate reads your pane for that.
   Your LATEST line is your entire visible state, so never leave a stale or stateless one standing.
   Whenever you mention a PR anywhere - a status line, your terminal, a summary - write its full
   https:// URL exactly as the forge printed it, never a bare number such as "PR 108"; firstmate
   copies that URL from your line rather than assembling one.
   A mid-task `working:` line (including setup complete) is nonterminal: do not end the
   turn after it; continue the same stage until a stopping point defined under Definition of done.
   Choose the verb by what clears the wait, not by whether you are idle.
   `paused:` is for a bounded external wait expected to clear on its own - an upstream release,
   a rate-limit reset, a scheduled window, or a backgrounded call you are parked on - for example
   `paused: rate limit resets at 06:00 UTC`.
   Firstmate then leaves your idle pane alone and rechecks it on a long cadence instead of treating
   it as a possible wedge.
   `blocked:` is when firstmate must act before you can continue - for example
   `blocked: implemented and committed, ready to validate` when implementation is done and you
   need firstmate's validation trigger, or `blocked: needs firstmate to steer past repeated failure`.
   Wrong-verb cost: `blocked:` on a self-clearing wait is a cheap extra wake; `paused:` on a
   wait that needs firstmate can idle you for an hour under away mode before anyone rechecks.
   Park-and-resume pairing: whenever you background a pipeline call and go idle, append
   `paused:` BEFORE going idle and `working:` as soon as it returns - otherwise a spent
   `needs-decision:` stays standing and firstmate reads you as still waiting on a decision it
   already answered.
   Never poll with `pgrep -f` or `pkill -f` on a pattern that appears in your own command line - the wait matches itself and cannot exit.
   Wait on the actual PID, or run the command in the foreground; when killing, kill by PID.
5. If you hit the same obstacle twice, append `blocked: {why}` and stop; firstmate will help.
6. Every design question follows the one-at-a-time Design profile contract above.
   Firstmate owns the answer or escalation; you are instructed by firstmate and the captain is not in the loop.
   Use the same stable key on the `needs-decision` event and the `resolved` event that closes it, and do not continue while that key is unanswered.
   Name firstmate in `resolved:` lines, PR bodies, and commits unless the decision text explicitly states the captain was consulted.
   For a no-mistakes ask-user gate specifically, escalate all ask-user findings as one event plus one snapshot file, using that same shape even when the gate holds only a single ask-user finding: write only the ask-user findings, verbatim and unparaphrased (id, severity, file, line, description, authority), to `/tmp/tmp.afhNKB9RXp/data/design-live/nm-<run>-findings.txt`, then report the gate with
   `needs-decision [key=nm-<run>-<step>]: ask-user findings=<id1>,<id2>,... file=/tmp/tmp.afhNKB9RXp/data/design-live/nm-<run>-findings.txt`
   naming every ask-user finding id from that gate. The status line only points at the file; it never restates or summarizes a finding's content.
   `resolved:` carries NO state, so it must never be your last line: append the next state line
   (normally `working:`) in the same breath. A trailing `resolved:` makes you read as no state at
   all - invisible to firstmate and indistinguishable from a dead worker, which is worse than stale.
7. Never stop, restart, or update the shared `no-mistakes` daemon - it is one instance serving
   every lane/home, so restarting it kills other lanes' in-flight pipeline runs; only firstmate
   manages the daemon.
   Before you append `blocked:` about the pipeline, run `no-mistakes daemon status` and
   `no-mistakes axi status`. If the daemon socket refuses connections or is missing, append
   `blocked: {the daemon error}` and stop even when the local run record still says running or
   fixing, because that record can be stale after the daemon exits. A run record failed with a
   daemon error is also a real block.
   Only after ruling out socket refusal, if the run is still running or fixing, reattach and keep
   going. A drive-call error, timeout, slow read, or generic unreachability is NOT a daemon error:
   the daemon accepts `respond` immediately and runs the round in the background, so a killed or
   timed-out call was only waiting for a read while the run kept working.

# Firstmate instruction inbox
Firstmate steers you through durable message files in '/tmp/tmp.afhNKB9RXp/state/design-live.inbox'.
When a terminal message says an instruction is waiting there - and at any natural checkpoint when you are unsure - list '/tmp/tmp.afhNKB9RXp/state/design-live.inbox'/*.msg, read and act on each message in numeric order, then acknowledge each handled message by moving it: `mv '/tmp/tmp.afhNKB9RXp/state/design-live.inbox'/NNN.msg '/tmp/tmp.afhNKB9RXp/state/design-live.inbox'/handled/`.
The move IS the acknowledgement: without it firstmate rings again and eventually treats you as stuck. An empty or absent inbox needs no action.

# Definition of done
Delivery contract: mode=no-mistakes
Before reporting the ADR ready, read and follow `~/.no-mistakes/worktrees/bc8432f7c9f8/01M294Z727X91MCT3V8QZ0EVJP/.agents/skills/captain-hold-lifecycle/SKILL.md` and pass its shared completion gate for every unresolved decision surfaced by the interview or ADR.
Inspect the branch diff and confirm the ADR is the only worker-authored tracked project change.
The final status summary must name the ADR path and concisely state the decisions taken.
This ADR ships through **no-mistakes**: `done:` means the PR is open with its checks green.
A clean local ADR commit is NOT done, and neither is your own test run passing - this task has exactly one `done:` line and it is the last one, `done: PR {url} checks green`.
The ADR is ready for validation only when committed on your branch.
When the ADR is complete and committed, append `paused: ADR complete and committed, ready to validate` and stop there; that handoff is a defined stopping point and a declared wait, and firstmate will then instruct you to run /no-mistakes to validate and ship the ADR PR.

You drive no-mistakes by responding to its gates, not by applying fixes.
Follow the guidance no-mistakes itself provides for the mechanics: it loads when you invoke /no-mistakes, and `no-mistakes axi run --help` plus the `help` lines in each `axi` response are authoritative and version-matched to the installed binary.
When starting no-mistakes, pass `--intent` as only this brief's `## Captain's intent` subsection plus any later words the captain actually said.
For a legacy brief with no such subsection, include only words explicitly labeled `Captain:`, `Captain's words:`, `Captain's ask:`, or `Captain's intent:`; never copy its mixed `# Task` wholesale. If it has no provenance-marked captain words, stop and ask firstmate instead of starting no-mistakes.
Do not include `## Firstmate spec`, later Firstmate build constraints, or your own decisions and tradeoffs.
The `--intent` string you pass must be self-sufficient: that string plus the codebase must let a reader reconstruct roughly the same specification, without depending on a separate report, a PR, or context that lives only in this conversation.
When the captain's intent refers to a report, decision, or PR ("do items 1, 2, 3, and 7 of the report"), write the substance of the referenced items into `--intent` in the captain's terms, not only the pointer; that substance is the captain's ask by reference, while Firstmate's build instructions and your own decisions still stay out.
This replaces the no-mistakes skill's advice to enrich `--intent` with decisions and tradeoffs; that advice does not apply to Firstmate-dispatched work.
Do not hand-edit, commit, or apply findings yourself while a run is active - the pipeline applies every fix.
While you sit parked on a backgrounded `axi run` or `axi respond` call, rule 4's park-and-resume pairing applies: append `paused:` before you go idle and `working:` when the call returns.

One drive call blocks until the next gate or outcome, which routinely outlives what your harness lets a single command run: Claude Code kills a command at ten minutes maximum, while one fix round is capped around thirty minutes and up to three rounds chain.
So background the drive call and poll `no-mistakes axi status` from a separate call instead of sitting in one blocking hold your harness will kill.
Where a harness's own command limit is not established, assume it bounds commands and use that same background-and-poll shape.
A killed or timed-out call is never evidence the daemon died: the daemon accepts your response immediately and runs the round in the background, so the call was only ever waiting for a read while the run kept working.
Reattach and keep going rather than reporting the pipeline blocked; rule 7 owns the checks that decide when a pipeline block is real.

Two firstmate-specific rules layer on top of that guidance:
- ask-user findings are never yours to answer: escalate to firstmate using rule 6's ask-user format and stop.
  Firstmate applies `ask-user-authority` and obtains any required captain decision.
  When the decision comes back, feed it to the gate with `no-mistakes axi respond` and let the pipeline apply it - do not route the question to "the user" or apply the fix yourself.
- NEVER pass `--yes` (or `-y`) to `no-mistakes axi run` or `no-mistakes axi respond`. It is banned fleet-wide.
  It auto-resolves every gate including ask-user findings with no escalation, and answering your own ask-user finding is a hard rule violation.

After /no-mistakes reports CI green (the CI-ready return point - do not wait for it to keep monitoring in the background until merge), append `done: PR {url} checks green` and stop. You are finished.
Evidence: Ship brief, missing-plugin refusal, in-flight preservation transcript
$ FM_HOME=$H bin/fm-brief.sh ship-live sample --mode no-mistakes
scaffolded: /tmp/tmp.afhNKB9RXp/data/ship-live/brief.md (ship, mode=no-mistakes; replace {TASK} and {FIRSTMATE_SPEC})
rc=0
129 /tmp/tmp.afhNKB9RXp/data/ship-live/brief.md
grilling                 matches=0
domain-modeling          matches=0
lavish-axi               matches=0
planning conversation    matches=0
Design profile           matches=0
ADR                      matches=0
needs-decision           matches=5

$ FM_HOME=$H FM_MATTPOCOCK_PLUGIN_REGISTRY=/nonexistent.json bin/fm-brief.sh design-missing sample --design --mode no-mistakes
error: mattpocock plugin registry is unavailable at /nonexistent.json; refresh the captain-owned plugin install. Workers must not install or copy it
error: --design requires the captain-owned mattpocock grilling and domain-modeling skills; do not install or copy them from a worker
rc=1
ls: cannot access '/tmp/tmp.afhNKB9RXp/data/design-missing': No such file or directory

$ FM_MATTPOCOCK_PLUGIN_REGISTRY=/nonexistent.json bin/fm-design-skills.sh check
error: mattpocock plugin registry is unavailable at /nonexistent.json; refresh the captain-owned plugin install. Workers must not install or copy it
rc=1

# in-flight design task: brief generated under the previous contract must not be rewritten
160bd445177801c8ce3ee9adbc3910d6e326281a  /tmp/tmp.afhNKB9RXp/data/design-inflight/brief.md
error: /tmp/tmp.afhNKB9RXp/data/design-inflight/brief.md already exists
rc=1
160bd445177801c8ce3ee9adbc3910d6e326281a  /tmp/tmp.afhNKB9RXp/data/design-inflight/brief.md
This is an interactive DESIGN task whose only tracked project deliverable is one architectural decision record.
Evidence: Targeted design/brief/relaunch/manifest test output
=== tests/fm-design-skills.test.sh
ok - design skill resolver reads the active plugin capabilities without modifying them
ok - design skill resolver refuses an incomplete plugin without working around it
ok - design skill resolver reports the captain-owned dependency action
ok - resolve and check keep the v1 schema, grilling and domain_modeling keys, and check line
ok - resolve reports the ask-matt router path
ok - resolve looks up an extra named skill by directory name
ok - resolve refuses an unknown skill by name rather than emitting an empty key
ok - each design dispatch records the plugin release it resolved at that dispatch
ok - one resolver result supplies both recorded provenance and pinned brief paths
ok - a missing pinned path refuses instead of resolving another release
ok - a non-design dispatch neither resolves nor records the design plugin
ok - a design dispatch refuses rather than launching an untraceable interview

all fm-design-skills tests passed
rc=0
=== tests/fm-brief.test.sh
ok - fm-brief.sh: the documented {TASK} and {FIRSTMATE_SPEC} fills cannot corrupt the Herdr safety gate
ok - fm-brief.sh: Herdr lab contract covers scouts and rejects secondmate misuse
ok - fm-brief.sh: --no-projects scaffolds a project-less charter and guards misuse
ok - fm-brief.sh: marked requests avoid generic acknowledgements and preserve material reporting
ok - fm-brief.sh: relative directory inputs ignore CDPATH, render stable absolute charter paths, or fail loudly
ok - fm-brief.sh: custom pause verb renders in every scaffold
ok - fm-brief.sh: investigation and visual-review completions load the shared decision policy
ok - fm-brief: scout and secondmate code paths still scaffold well-formed briefs
ok - fm-brief.sh: the brain instruction appears only for a home that has one, in every scaffold
ok - fm-brief.sh: a found scaffold search embeds labeled rows and prints the nearest prior work
ok - fm-brief.sh: an empty, failed, never-started, or overrunning scaffold search leaves the instruction-only section
ok - fm-brief.sh: the embed and advisory gates reject untrusted or incomplete protocol states
ok - fm-brief.sh: the embed honors its byte cap, charters run no search, and --query is validated
ok - fm-brief.sh: firstmate-repo persona guidance is git-common-dir gated and split by kind
ok - fm-brief.sh: the home-root candidate needs a home-root registry declaration, not merely registration
ok - design DOD no-mistakes: exact rendered text authorizes no implementation
ok - design DOD direct-PR: exact rendered text authorizes no implementation
ok - design DOD local-only: exact rendered text authorizes no implementation
ok - fm-brief.sh: design profile is ADR-only and resolves identically across supported harnesses
ok - fm-brief.sh: linked homes retain firstmate guidance without registry prose or a marker
ok - fm-brief.sh: resolved: lines and PR/commit attribution guidance is present in all brief variants
ok - fm-brief.sh: all continuation mode and kind combinations agree
ok - fm-brief.sh: --continue-branch rejects invalid and protected names

all fm-brief tests passed
rc=0
=== tests/fm-control-relaunch.test.sh
ok - fm-control relaunch: a launch failure after the stop keeps the prior record and reports the real state
ok - fm-control relaunch: unpublished rollback keeps concurrent durable metadata
ok - fm-control relaunch: post-publication failure keeps the new durable record
ok - fm-control relaunch: partial stop reconciles actual agent state
ok - fm-control relaunch: an already-stopped agent recovers - idempotent exit, replacement launched into the same endpoint
ok - fm-control relaunch: failed journal replacement preserves durable phase
ok - fm-spawn relaunch: prepublication abort removes replacement state
ok - fm-control relaunch: the checkpoint records the exact unlanded work it preserved
ok - fm-control relaunch: a secondmate's child work is accounted for and its charter is left alone
ok - fm-control relaunch: a secondmate home that is not this secondmate's is refused
ok - fm-control relaunch: unreadable and untraversable child state fails checkpoint
ok - fm-control relaunch: two control actions on one task serialize instead of interleaving
ok - fm-spawn relaunch: direct entry participates in lifecycle serialization
ok - fm-promote: promotion participates in lifecycle serialization
ok - fm-spawn --relaunch: refuses to launch a second agent into a live endpoint
ok - fm-spawn --relaunch: symlinked records refuse before inspection
ok - fm-spawn --relaunch: keeps its early meta lock continuous
ok - fm-spawn --relaunch: pending closes refuse before replacement begins
ok - fm-spawn --relaunch: every identity axis comes from the record, and a contradicting flag refuses
ok - fm-spawn --relaunch: an unrecorded task is refused
ok - fm-spawn --relaunch: refuses to start a replacement outside the copy holding the work
ok - relaunch re-reads the backlog item instead of blindly re-running the transition
ok - relaunch heals an item that drifted out of In flight while the task stayed live

all fm-control-relaunch tests passed
rc=0
=== tests/fm-outcome-manifest.test.sh
ok - the manifest composes dispatch, outcome, PR, attribution, and provenance from live records
ok - the manifest carries an optional provider receipt and refuses to publish its malformed values
ok - design is a valid durable outcome manifest kind
ok - a design task's manifest records the plugin release it was dispatched against
ok - a task that never resolved the plugin records no release rather than a guess
ok - a manifest written before the design plugin record stays valid in durable history
ok - secondmate manifests retain nullable title and provider fields in durable history
ok - a manifest written before the usage session map stays valid in durable history
ok - manifest publication and reads require the complete single-root contract
ok - work-item references round-trip across forges and hosts, upsert by URL, and stay renderable unenriched
ok - the work-item store refuses relative, non-http, traversing, and malformed URLs
ok - work-item mutations preserve invalid single- and multi-root durable stores
ok - each forge vocabulary maps onto the normalized state, review, check, and mergeability enumerations
ok - the PR observation retains same-PR failures and rejects stale cache identity
ok - GitLab review normalization follows current approval state and degrades safely
ok - the portable fallback bounds forge calls when GNU timeout is unavailable
ok - durable history orders newest first, bounds with disclosure, and surfaces unreadable records
ok - a task, its attribution, and its work items stay in durable history after the volatile records are gone
ok - a history read that loses its records refuses, and says what it counted, instead of reporting a fleet that finished nothing
ok - neither the manifest nor the snapshot emits credential, brief, or payload content
ok - durable free text is capped and collapsed to a single line

all fm-outcome-manifest tests passed
rc=0
Evidence: Design dispatch matrix output
ok - unresolvable relative spawn overrides fail with named diagnostics
ok - watermark failure preserves recoverable metadata without launching
ok - design matrix codex 1/5: dispatch schema selects codex/gpt-5.5/xhigh
ok - design matrix codex 2/5: fm-harness.sh fallback resolves codex
ok - design matrix codex 3/5: fm-launch-lib.sh constructs both commands
ok - design matrix codex 4/5: fm-spawn.sh validates design and records all axes
ok - design matrix codex 5/5: harness-adapters axes render gpt-5.5/xhigh
ok - design matrix pi 1/5: dispatch schema selects pi/anthropic/claude-sonnet-5/xhigh
ok - design matrix pi 2/5: fm-harness.sh fallback resolves pi
ok - design matrix pi 3/5: fm-launch-lib.sh constructs both commands
ok - design matrix pi 4/5: fm-spawn.sh validates design and records all axes
ok - design matrix pi 5/5: harness-adapters axes render anthropic/claude-sonnet-5/xhigh
ok - design matrix claude 1/5: dispatch schema selects claude/claude-sonnet-5/xhigh
ok - design matrix claude 2/5: fm-harness.sh fallback resolves claude
ok - design matrix claude 3/5: fm-launch-lib.sh constructs both commands
ok - design matrix claude 4/5: fm-spawn.sh validates design and records all axes
ok - design matrix claude 5/5: harness-adapters axes render claude-sonnet-5/xhigh
all fm-spawn-dispatch-profile tests passed

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 1 info
  • ⚠️ tests/fm-brief.test.sh:1572 - The new ship-brief check runs fm-brief.sh ship-no-planning-stack sample-ship ... &gt;/dev/null 2&gt;&amp;1 and never checks its exit code or that the brief exists. assert_no_grep is ! grep -F pat file, and grep returns 2 on a missing file, so if this scaffold ever fails (bad arg, new required env, refusal) all three assert_no_grep checks (grilling, lavish-axi, planning conversation) pass without testing anything. Capture rc and expect_code 0, or assert_grep/[ -f ] the brief before the negative checks.
  • ℹ️ bin/fm-spawn.sh:1977 - The change rewords plugin provenance from 'auto-updating, captain uses /plugin' to 'captain-owned install that can change between tasks', but two places keep the old framing: the comment at bin/fm-spawn.sh:1977 ('the captain has that plugin on auto-update') and docs/scripts.md:45 ('captain-installed mattpocock plugin'). Align them with the new wording so the plugin lifecycle is described one way everywhere.

🔧 Fix applied.
✅ Re-checked - no issues remain.

  • ℹ️ bin/fm-design-skills-lib.sh:28 - This change rewords the plugin as a 'captain-owned install that can change between tasks' everywhere else, and round 1 already fixed the same leftover wording in fm-spawn.sh and docs/scripts.md. The comment on adopt_relaunch_design_skills still says 'a later plugin auto-update must not silently rebind the interview'. Change it to something like 'a later change to the captain-owned install must not silently rebind the interview' so the plugin lifecycle is described one way.
✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 4 of 6 scenarios driven live against the product
Scenario Result Live Evidence
Captain scaffolds a --design task and gets one planning conversation brief (optional visual proposal, selective Matt questioning/terms, no stacked Kun/Matt/ADR passes) ✅ pass live design-brief.md Design profile section; cli-transcript.txt rc=0
Design brief keeps ADRs for consequential tradeoffs only and sends routine configuration back as needs-decision ✅ pass live design-brief.md: 'Write a short ADR only when ... ADR bar is met' / 'rather than padding a ceremonial ADR'
Well-specified ship task brief does not inherit grilling, lavish-axi, or the planning conversation ✅ pass live cli-transcript.txt grep section '(none)'; ship-brief.md
Adversarial: design scaffold with the plugin registry missing refuses, writes no brief, and tells workers not to install it ✅ pass live cli-transcript.txt blocked-design rc=1, brief absent
Design dispatch via fm-spawn pins the resolved skill paths and tells the worker to use them inside one conversation ⏸️ untested no The prior payload didn't establish a live result. It only ran tests/fm-design-skills.test.sh with a fake harness. A live run needs the shared no-mistakes daemon and a treehouse worktree pool to launch…
Existing ADRs and in-flight designs are preserved ⏸️ untested no The prior payload didn't establish a live result. This was only checked by reading the diff (git diff --name-status 668b61f 989f42f), not by running the product.
  • bin/fm-design-skills.sh check against the real captain-installed mattpocock plugin (v1.2.3)
  • FM_HOME=&lt;tmp&gt; bin/fm-brief.sh unified-design sample --design --mode no-mistakes and inspected the rendered Design profile section
  • FM_HOME=&lt;tmp&gt; bin/fm-brief.sh plain-ship sample --mode no-mistakes then grep for grilling/lavish-axi/planning conversation/domain-modeling (none found)
  • FM_MATTPOCOCK_PLUGIN_REGISTRY=/nonexistent bin/fm-brief.sh blocked-design sample --design (refused rc=1, no brief written)
  • FM_MATTPOCOCK_PLUGIN_REGISTRY=/nonexistent bin/fm-design-skills.sh resolve (captain-owned refusal wording)
  • bash tests/fm-design-skills.test.sh (includes the fm-spawn design dispatch: pinned paths, one-conversation wording, refusal)
  • bash tests/fm-brief.test.sh
  • git diff --name-status 668b61f1 989f42f4 to confirm no docs/adr or in-flight design files changed

✅ No issues found.

  • Live validation: ✅ go - 5 of 8 scenarios driven live against the product
Scenario Result Live Evidence
Captain scaffolds a design task with the real installed plugin and the brief describes one planning conversation: optional lavish-axi visual, one unblocked grilling question at a time, domain-modeling… ✅ pass live 01-design-scaffold-real-plugin.txt, 02-design-brief.md
The ADR bar and design-tree wording in the brief match what the dispatched real plugin skills actually define (domain-modeling's 3 criteria; grilling's frontier batching is overridden by one keyed que… ✅ pass live grep of ~/.claude/plugins/cache/mattpocock/mattpocock-skills/1.2.3 domain-modeling and grilling SKILL.md against 02-design-brief.md
Adversarial: a well-specified ship brief does not pick up grilling, lavish-axi, the planning conversation, or ADR requirements ✅ pass live 03-ship-refusal-inflight.txt (all planning-stack matches=0)
Adversarial: with no plugin registry, design scaffolding and the resolver refuse with captain-owned refresh wording, write no brief, and never suggest a worker install ✅ pass live 03-ship-refusal-inflight.txt (rc=1, no data dir)
Adversarial: an in-flight design brief generated under the old interactive-ADR contract is not rewritten by a new scaffold ✅ pass live 03-ship-refusal-inflight.txt (sha1 unchanged, 'already exists' rc=1)
fm-spawn --design resolves the plugin once, pins exact skill paths, and adds the one-conversation line and captain-owned blocker to the worker-facing brief ⏸️ untested no The prior payload did not establish a live result: it was covered only by tests/fm-design-skills.test.sh with a stub harness. A live check needs a real design worker launched through fm-spawn (tmux pl…
Relaunching an in-flight design worker keeps its dispatch pin, inbox, and uncommitted work on claude, codex, and pi ⏸️ untested no The prior payload did not establish a live result: it was covered only by tests/fm-control-relaunch.test.sh with stub harnesses. A live check needs a running design worker in a real fleet session to r…
A real design worker holds a single conversation (no separate Kun/Matt workflows) and declines a ceremonial ADR for routine configuration ⏸️ untested no The prior payload only verified the generated prompt (02-design-brief.md). Checking model behavior needs a live, billed design agent session run by the captain on a real design ask.
  • bin/fm-design-skills.sh check and resolve against the real installed mattpocock-skills 1.2.3 registry
  • FM_HOME=&lt;tmp&gt; bin/fm-brief.sh design-live sample --design --mode no-mistakes with the real plugin, then read the generated brief.md
  • Checked that the real plugin's domain-modeling SKILL.md defines the 3-part ADR bar the brief cites, and that grilling's 'ask whole frontier' is overridden by the brief
  • FM_HOME=&lt;tmp&gt; bin/fm-brief.sh ship-live sample --mode no-mistakes and counted grilling/lavish-axi/planning conversation/ADR in the brief
  • FM_MATTPOCOCK_PLUGIN_REGISTRY=/nonexistent.json bin/fm-brief.sh design-missing ... --design and bin/fm-design-skills.sh check refusal paths
  • Pre-seeded an in-flight design brief under the old contract and re-ran fm-brief.sh design-inflight --design to confirm it is not rewritten (sha1 unchanged)
  • bash tests/fm-design-skills.test.sh (drives fm-spawn.sh --design dispatch pinning with stub harness)
  • bash tests/fm-brief.test.sh
  • bash tests/fm-control-relaunch.test.sh (design relaunch keeps pin/inbox/work on claude, codex, pi)
  • bash tests/fm-outcome-manifest.test.sh
  • bash tests/fm-spawn-dispatch-profile.test.sh (design matrix across harnesses)
✅ **Document** - passed

✅ No issues found.

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

✅ No issues found.

Sungin Kim and others added 4 commits September 14, 2026 02:54
Keep Kun research and optional visuals, dependency-aware questioning, and
short ADRs in a single design path instead of stacked workflows, and leave
plugin lifecycle to the captain-owned install outside this repo.

Co-authored-by: Cursor <cursoragent@cursor.com>
@HelloWorldSungin
HelloWorldSungin force-pushed the fm/fm-unified-kun-matt-planning branch from f2dad68 to b4f82e8 Compare September 14, 2026 03:02
@HelloWorldSungin
HelloWorldSungin merged commit 2026a11 into main Sep 14, 2026
17 checks passed
@HelloWorldSungin
HelloWorldSungin deleted the fm/fm-unified-kun-matt-planning branch September 14, 2026 03:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant