From 5d4d3d1179854a60af63c3c6a06994810e877d5b Mon Sep 17 00:00:00 2001 From: tobrun Date: Wed, 7 Oct 2026 13:07:04 +0200 Subject: [PATCH] feat(dev): add scope-quick and ship-quick, the short loop for review fixes A blocked ship review had one recommended path: a full scope run, then a full ship run to re-check the fixes. Both are far heavier than fixing a few findings needs. scope-quick writes minimal change sets from the latest review's Blockers, or from a small request, with no interview, argued decisions or subagents, looping the existing lint-spec.py. ship-quick re-ships with one validation run and one reviewer that checks the previous findings and the diff since the last review, then updates the pull request and follows it to green - no gauntlet, panel or fix loop. ship's review report now records the reviewed head commit so a quick run can diff from it, and ship and build recommend the quick loop for plain defects. Bumps the dev plugin to 3.8.0. --- README.md | 2 +- dev/.claude-plugin/plugin.json | 2 +- dev/README.md | 25 +++++++- dev/evals/README.md | 4 +- dev/evals/scope-quick.json | 28 +++++++++ dev/evals/ship-quick.json | 29 +++++++++ dev/references/plan-layout.md | 8 +-- dev/skills/build/SKILL.md | 2 +- dev/skills/reflect/SKILL.md | 2 +- dev/skills/scope-quick/SKILL.md | 62 +++++++++++++++++++ dev/skills/ship-quick/SKILL.md | 51 +++++++++++++++ dev/skills/ship/SKILL.md | 2 +- dev/skills/ship/references/report-format.md | 3 +- docs/architecture.md | 4 +- package.json | 2 +- plugins/dev/.codex-plugin/plugin.json | 2 +- plugins/dev/references/plan-layout.md | 8 +-- plugins/dev/skills/build/SKILL.md | 2 +- plugins/dev/skills/reflect/SKILL.md | 2 +- plugins/dev/skills/scope-quick/SKILL.md | 61 ++++++++++++++++++ .../dev/skills/scope-quick/agents/openai.yaml | 6 ++ plugins/dev/skills/ship-quick/SKILL.md | 50 +++++++++++++++ .../dev/skills/ship-quick/agents/openai.yaml | 6 ++ plugins/dev/skills/ship/SKILL.md | 2 +- .../skills/ship/references/report-format.md | 3 +- scripts/build_codex_plugin.py | 10 +++ 26 files changed, 351 insertions(+), 27 deletions(-) create mode 100644 dev/evals/scope-quick.json create mode 100644 dev/evals/ship-quick.json create mode 100644 dev/skills/scope-quick/SKILL.md create mode 100644 dev/skills/ship-quick/SKILL.md create mode 100644 plugins/dev/skills/scope-quick/SKILL.md create mode 100644 plugins/dev/skills/scope-quick/agents/openai.yaml create mode 100644 plugins/dev/skills/ship-quick/SKILL.md create mode 100644 plugins/dev/skills/ship-quick/agents/openai.yaml diff --git a/README.md b/README.md index cb69400..dfa8aaa 100644 --- a/README.md +++ b/README.md @@ -2,7 +2,7 @@ | Plugin | Use When | Tools | | ------ | -------- | ----- | -| [dev](dev/) | A test-focused development workflow for Claude Code, Codex, opencode, and Pi. | `scope`, `scope-review`, `commit`, `build`, `ship`, `reflect`, `to-pitch`, `to-quiz` | +| [dev](dev/) | A test-focused development workflow for Claude Code, Codex, opencode, and Pi. | `scope`, `scope-quick`, `scope-review`, `commit`, `build`, `ship`, `ship-quick`, `reflect`, `to-pitch`, `to-quiz` | | [factory](factory/) | Take a request from scope to a shipped pull request unattended, on Claude Code or Codex. | `run` | | [bootstrap](bootstrap/) | Prepare any repository for agent work: probe its stack and write its root AGENTS.md from the workflow's SDLC lessons, on Claude Code or Codex. | `agents-md` | diff --git a/dev/.claude-plugin/plugin.json b/dev/.claude-plugin/plugin.json index 7a94f99..78955b5 100644 --- a/dev/.claude-plugin/plugin.json +++ b/dev/.claude-plugin/plugin.json @@ -1,6 +1,6 @@ { "name": "dev", - "version": "3.7.0", + "version": "3.8.0", "description": "Development workflow skills: scope changes with argued decisions, build across unit/integration/e2e with every scenario proven by tests, ship with a deterministic quality gauntlet and an adversarially verified review, create structured commits that feed a decision ledger, and render pitches or comprehension quizzes.", "author": { "name": "Tobrun" diff --git a/dev/README.md b/dev/README.md index 3f9882d..1310fe6 100644 --- a/dev/README.md +++ b/dev/README.md @@ -2,7 +2,7 @@ Development workflow skills for Claude Code, Codex, opencode, and Pi, built around two ideas: layered tests are the enforceable spec for behavior, and every phase produces something a human actually reviews as HTML, not markdown scrolling. -The skills chain loosely rather than as a rigid pipeline: `/scope` interviews for the real problem, argues every design decision against alternatives, and writes a self-contained spec whose change plan carries layer-tagged test scenarios; `/scope-review` is an optional step for a large or complex change, never a gate in front of `build`: it puts the settled spec through a fresh-context, adversarially verified agent panel that checks the plan against the actual repo and refines the spec in place, looping without a human and closing with a short interview for the few findings only the user can decide, so a finished run hands `build` a spec ready to implement; `/build` executes the spec's change sets across unit/integration/e2e, proving every scenario with a real test at its tagged layer, in parallel waves where file lists allow, keeping a running implementation-notes log; `/ship` runs a deterministic quality gauntlet - the repo's own static analysis, security scan, dead code, duplication, dependency rules, coverage-weighted complexity, flakiness, mutation testing - looping fix agents until the checkers pass, then verifies the result with a fan-out review panel that checks spec conformance, e2e coverage, and logged deviations; `/commit` groups pending changes into granular commits with structured what/why messages; `/reflect` consolidates the journal every run leaves into cited claims about the skills themselves and hands the most recurrent one to `scope` as a brief; `/to-pitch` and `/to-quiz` turn finished work into a buy-in doc or a comprehension check. +The skills chain loosely rather than as a rigid pipeline: `/scope` interviews for the real problem, argues every design decision against alternatives, and writes a self-contained spec whose change plan carries layer-tagged test scenarios; `/scope-review` is an optional step for a large or complex change, never a gate in front of `build`: it puts the settled spec through a fresh-context, adversarially verified agent panel that checks the plan against the actual repo and refines the spec in place, looping without a human and closing with a short interview for the few findings only the user can decide, so a finished run hands `build` a spec ready to implement; `/build` executes the spec's change sets across unit/integration/e2e, proving every scenario with a real test at its tagged layer, in parallel waves where file lists allow, keeping a running implementation-notes log; `/ship` runs a deterministic quality gauntlet - the repo's own static analysis, security scan, dead code, duplication, dependency rules, coverage-weighted complexity, flakiness, mutation testing - looping fix agents until the checkers pass, then verifies the result with a fan-out review panel that checks spec conformance, e2e coverage, and logged deviations; `/scope-quick` and `/ship-quick` are the short loop for a change that is already small - most often the fixes a `ship` review asked for - writing minimal change sets with no interview, and re-shipping with one reviewer instead of the gauntlet and panel; `/commit` groups pending changes into granular commits with structured what/why messages; `/reflect` consolidates the journal every run leaves into cited claims about the skills themselves and hands the most recurrent one to `scope` as a brief; `/to-pitch` and `/to-quiz` turn finished work into a buy-in doc or a comprehension check. The durable context is deliberately small: the code, its tests, the active spec under `.dev/{plan-name}/`, and three repo-tracked registries the skills maintain in the consuming project - `docs/decisions.md` (design decisions with their argued alternatives, read only after a review forms its findings), `docs/contracts.md` (boundary guarantees, read as premises before a review walks the diff), and `docs/dependencies.md` (machine-checkable module dependency rules, enforced by `ship`). The files under `.dev/{plan-name}/` are written as a run goes, not when a stage closes: the spec opens during the interview, a report opens before its panel returns, the implementation notes gain an entry per change set and per fixup, and the PR body fills check by check, so a run can be followed from its files and a dead session loses only what was in flight. Alongside them, `docs/architecture.md` is a plain high-level overview of the system - components, flows, boundaries, entry points - captured in full the first time a skill needs it and finds it absent, then kept current by build and commit whenever the structure changes, with a small checker that catches stale paths and files no component covers. @@ -19,6 +19,13 @@ A checker (`scripts/lint-spec.py`) enforces the spec's mechanics - unique slugs, On a clean spec it prints the build waves the file lists allow and the shared files that make change sets wait, so the plan is shaped for parallel work before build starts. Promotes durable decisions to `docs/decisions.md` and cross-boundary invariants to `docs/contracts.md`, renders an expandable-card spec view, and has a reverse mode that audits the implicit decisions already embedded in existing code. +### scope-quick + +The minimal `scope`: turns the Blockers of the latest `ship` review, or a small request, into change sets in `.dev/{plan-name}/spec.md` that `build` can execute straight away. +It runs no interview, argues no decisions, launches no subagents, and skips the ledger work, the HTML render, and the retrospective; a decision is recorded only where the code offered a real choice. +Review fixes are appended to the existing change plan with continued numbering, each carrying the review's triggering scenario as its test, and the same `lint-spec.py` loop keeps the result mechanically sound. +When a finding needs a recorded decision flipped or user-visible scope changed, it stops and points at `scope` instead of guessing. + ### scope-review Optional: `build` runs on any spec `scope` finished, and `scope` advises this review only when the change is large or complex. @@ -54,6 +61,14 @@ Phase 3 commits what the gauntlet fixed, pushes, and opens the pull request auto A deterministic check (`pr-evidence.py check`) gates the PR body, so a PR cannot open on a placeholder or a data URI, and the phase then follows required checks to green. Local gates in every phase run at the change's impact, so the full merge gate is the PR's own CI: a PR carrying a check deferred to CI opens as a draft and is marked ready once its required checks pass, and a run that opens no PR runs the full set locally. +### ship-quick + +The minimal `ship`, for re-shipping after the fixes a review asked for: one validation run at the diff's impact, one read-only reviewer, and a pull request update followed to green. +The reviewer reports each finding of the previous `review_N.md` as fixed or still open and reads only the diff since that review's recorded head for new blockers; it raises no concerns, nits, or simplifications. +There is no gauntlet, no lens panel, no remediation loop, and no HTML report: a standing blocker ends the run with a pointer back to `scope-quick`. +An existing pull request keeps the Evidence and Quality sections the full `ship` run wrote, with only its review line and open calls updated. +Run the full `ship` for a first ship, or once the change has grown beyond the findings it set out to fix. + ### commit Groups all pending changes into granular, logically-separate commits - splitting within a file when needed - with structured messages: a `type(scope):` subject, `What:`/`Why:` body, optional `Considered:`/`Constraint:`/`Directive:`/`Symptoms:` sections, and `Severity:`/`Risk:` metadata trailers. @@ -63,7 +78,7 @@ Pushes by default; say "commit only" to skip the push. ### reflect Turns what past runs measured into evidence about the skills themselves, so the next change to the workflow fixes something that actually recurred. -Every `scope`, `scope-review`, `build`, and `ship` run appends a journal entry to `~/.dev-workflow/memory/` (see Run metrics below); `reflect` has read-only subagents read the transcript around each measured signal and propose claims such as "build reran the full Validation block after each change set; the user stopped it", each quoting its source. +Every `scope`, `scope-review`, `build`, and `ship` run, and every run of their quick variants, appends a journal entry to `~/.dev-workflow/memory/` (see Run metrics below); `reflect` has read-only subagents read the transcript around each measured signal and propose claims such as "build reran the full Validation block after each change set; the user stopped it", each quoting its source. A checker (`scripts/claims.py add`) accepts a batch only when every quote is found verbatim at the transcript line or journal entry it cites, then rebuilds per-skill pages and a ranked list of threads in which every line cites a claim id. The user picks a thread, or retracts a claim that misreads its evidence (a retracted claim stays on record so the same evidence cannot bring it back), and the skill writes a scope brief with an eval case built from the real run. It never edits a skill: the fix goes through `scope` and `build` in this repository, and `/dev:reflect resolve {thread} {sha}` records the commit, after which a new claim in that thread reopens it. @@ -79,7 +94,7 @@ Cannot enforce a merge gate, so it says so plainly and produces an honest pass/f ## Run metrics -`scope`, `scope-review`, `build`, and `ship` each start by snapshotting the run with `scripts/skill-metrics.py start` and end by printing what it measured: wall time, tokens split between the orchestrator and its subagents, agents dispatched, tool calls, the git delta since the snapshot, and any counters the skill tallied from tool output. +`scope`, `scope-review`, `build`, `ship`, `scope-quick`, and `ship-quick` each start by snapshotting the run with `scripts/skill-metrics.py start` and end by printing what it measured: wall time, tokens split between the orchestrator and its subagents, agents dispatched, tool calls, the git delta since the snapshot, and any counters the skill tallied from tool output. The numbers come from the session transcript under `$CLAUDE_CONFIG_DIR` (default `~/.claude`) and from git, never from the model's recollection. Every run appends a row to `.dev/metrics.jsonl` in the consuming repository, and the table compares the run against the median of earlier runs of the same skill, which is where a skill's cost and catch rate become visible over time. The same call appends an entry to the cross-repository run journal under `~/.dev-workflow/memory/journal/` (or `$DEV_MEMORY_DIR`): the friction signals measured from the transcript (interrupts, denied and failed tool calls, the user's own turns, and how often each deterministic checker ran and failed), each with its transcript line, plus at most three `--friction` lines in which the skill names where the run fought its own instructions. @@ -129,9 +144,13 @@ Preserve the explicit-invocation policy with a permission rule in `~/.config/ope "skill": { "*": "allow", "scope": "ask", + "scope-quick": "ask", + "scope-review": "ask", "commit": "ask", "build": "ask", "ship": "ask", + "ship-quick": "ask", + "reflect": "ask", "to-pitch": "ask", "to-quiz": "ask" } diff --git a/dev/evals/README.md b/dev/evals/README.md index 04bcce4..6ac6c71 100644 --- a/dev/evals/README.md +++ b/dev/evals/README.md @@ -4,12 +4,12 @@ Status: maintained Eval definitions for the `dev` plugin's skills: realistic prompts and objective assertions used to check whether a skill change preserved behavior. -- `{skill}.json` - one file per skill: the eval prompt(s), the fixture each expects, and the assertions to grade the output against. Covers all 8 skills: `scope`, `scope-review`, `commit`, `build`, `ship`, `reflect`, `to-pitch`, `to-quiz`. +- `{skill}.json` - one file per skill: the eval prompt(s), the fixture each expects, and the assertions to grade the output against. Covers all 10 skills: `scope`, `scope-quick`, `scope-review`, `commit`, `build`, `ship`, `ship-quick`, `reflect`, `to-pitch`, `to-quiz`. - `results.md` - the record of the most recent full run: scores, methodology, and findings. - `tests/` - unit tests for the deterministic scripts the skills loop against (`lint-spec.py`, `change-set-brief.py`, `check-tests.py`, `pr-evidence.py`, `impact-scope.py`, and the memory scripts `skill-metrics.py` and `claims.py`), run by `scripts/validate.sh` as check D01. `build` runs in `"functional"` mode (a real fixture, a real subagent run, assertions checked against the actual output). -`scope`, `scope-review`, `commit`, `ship`, `reflect`, `to-pitch`, and `to-quiz` run in `"comprehension"` mode instead - each depends on either an interactive question loop, a live codebase, or prior artifacts (a finished spec, implementation notes, an e2e report) that are too expensive to stage on every iteration, so these check policy comprehension of the skill text directly. +`scope`, `scope-quick`, `scope-review`, `commit`, `ship`, `ship-quick`, `reflect`, `to-pitch`, and `to-quiz` run in `"comprehension"` mode instead - each depends on either an interactive question loop, a live codebase, or prior artifacts (a finished spec, implementation notes, an e2e report) that are too expensive to stage on every iteration, so these check policy comprehension of the skill text directly. ## When to run these diff --git a/dev/evals/scope-quick.json b/dev/evals/scope-quick.json new file mode 100644 index 0000000..3c3988a --- /dev/null +++ b/dev/evals/scope-quick.json @@ -0,0 +1,28 @@ +{ + "skill": "scope-quick", + "mode": "comprehension", + "evals": [ + { + "id": "scope-quick-from-review", + "prompt": "Answer briefly, from the skill text only. The plan directory holds spec.md with change sets 1 to 4, all built, and review_2.md with two Blockers and three Concerns. The user runs scope-quick with no other request. (1) How many change sets does it write, and how are they numbered? (2) What is each one's test? (3) Does it interview the user, argue decisions, or launch subagents? (4) What does it loop against before finishing? (5) What does it recommend next?", + "fixture": "none - the answers come from the skill text", + "assertions": [ + "Answer 1: two, one per Blocker, appended as change sets 5 and 6 with nothing renumbered or rewritten; Concerns are left out unless the user names them", + "Answer 2: the review's triggering scenario, on a single tests: line tagged at the lowest layer that proves it", + "Answer 3: no to all three; it asks at most one question and only when it cannot proceed", + "Answer 4: scope's lint-spec.py against spec.md, until it exits clean", + "Answer 5: build, then ship-quick, recommended and never launched" + ] + }, + { + "id": "scope-quick-escalates", + "prompt": "Answer briefly, from the skill text only. (1) A Blocker can only be fixed by reversing a decision the spec recorded. What does scope-quick do? (2) A small request has one obvious implementation. How many D- decisions does the spec record, and which sections must it still have? (3) Name three things scope does that scope-quick skips.", + "fixture": "none - the answers come from the skill text", + "assertions": [ + "Answer 1: it stops and recommends scope instead of guessing", + "Answer 2: none; the spec still has Research, Scope with a Validation block of the repo's real commands, and a Change Plan", + "Answer 3: any three of the interview, the decision catalog and evidence batch, the blind-spot pass, the spec reviewer, the ledger reconcile and promotion, the HTML render, the retrospective" + ] + } + ] +} diff --git a/dev/evals/ship-quick.json b/dev/evals/ship-quick.json new file mode 100644 index 0000000..a38e5ec --- /dev/null +++ b/dev/evals/ship-quick.json @@ -0,0 +1,29 @@ +{ + "skill": "ship-quick", + "mode": "comprehension", + "evals": [ + { + "id": "ship-quick-after-fixes", + "prompt": "Answer briefly, from the skill text only. review_1.md ended in BLOCK with two Blockers and one Concern and records a head commit; build has since committed two fix change sets, and a draft pull request exists. The user runs ship-quick. (1) Which gauntlet tools run? (2) How many review agents launch, and what do they read? (3) What kinds of new finding may the reviewer raise? (4) Both Blockers are fixed and the Concern is still open: what is the verdict and where is it written? (5) What happens to the pull request body?", + "fixture": "none - the answers come from the skill text", + "assertions": [ + "Answer 1: none; only the repository's required pull-request commands run, once, at the diff's impact", + "Answer 2: one fresh-context read-only agent; it checks each previous Blocker and Concern as fixed or still open and reads the diff since the recorded head", + "Answer 3: blockers only, each with a triggering scenario, and each confirmed by the orchestrator reading the code; no concerns, nits, or simplifications", + "Answer 4: CONCERNS, in review_2.md at the next free index with Panel: quick and the Previous findings section filled", + "Answer 5: pr.md is edited in place - the review line and Open calls - keeping the earlier Evidence and Quality sections, then pr-evidence.py check passes before the body is updated and the checks are followed to green" + ] + }, + { + "id": "ship-quick-limits", + "prompt": "Answer briefly, from the skill text only. (1) The reviewer reports one Blocker still open. Does ship-quick fix it and re-review? What does it do? (2) There is no earlier review_N.md. What does the reviewer read? (3) May it force-push? (4) What must the wrap-up say about what this run did not do, and when does it recommend the full ship?", + "fixture": "none - the answers come from the skill text", + "assertions": [ + "Answer 1: no - there is no remediation loop; it writes the BLOCK verdict, stops before the pull request step, and recommends scope-quick, then build, then ship-quick again", + "Answer 2: the whole branch diff", + "Answer 3: no - never force-push, never rebase", + "Answer 4: that the gauntlet and the panel were skipped; ship is recommended when the change has grown beyond the findings it set out to fix" + ] + } + ] +} diff --git a/dev/references/plan-layout.md b/dev/references/plan-layout.md index 09a9500..ecb5adb 100644 --- a/dev/references/plan-layout.md +++ b/dev/references/plan-layout.md @@ -7,15 +7,15 @@ This reference owns the layout, the locating convention, and the diff scope; ski | File | Written by | Read by | | ---- | ---------- | ------- | -| `spec.md` | `scope` (from the interview on); `scope-review` (verified refinements only) | everyone downstream | +| `spec.md` | `scope` (from the interview on); `scope-review` (verified refinements only); `scope-quick` (minimal change sets, appended when a spec exists) | everyone downstream | | `spec-review_N.md` | `scope-review` (next free index, opened before round 1) | `scope` (remediation), re-reviews | | `implementation-notes.md` | `build` (append-only) | `ship`, `to-pitch`, `to-quiz` | -| `review_N.md` | `ship` (next free index, opened when the panel is selected) | `scope` (remediation), re-reviews | -| `pr.md` | `ship` (opened at the start of a run that reaches phase 3, overwritten per run) | the PR tool via `--body-file`; re-runs | +| `review_N.md` | `ship` (next free index, opened when the panel is selected); `ship-quick` (next free index) | `scope` and `scope-quick` (remediation), re-reviews | +| `pr.md` | `ship` (opened at the start of a run that reaches phase 3, overwritten per run); `ship-quick` (review line and Open calls only) | the PR tool via `--body-file`; re-runs | | `.dev/config.json` | the user | `skill-metrics.py` (run journal opt-out, per [run-journal.md](run-journal.md)) | Each producing skill also renders an HTML companion under `/tmp/{project-slug}/reports/` per [reporting.md](reporting.md), named by that skill. -One writer per file; every other skill only reads. +One writer per file at a time; every other skill only reads. ## Written as the run goes diff --git a/dev/skills/build/SKILL.md b/dev/skills/build/SKILL.md index 759d046..e0bc7e1 100644 --- a/dev/skills/build/SKILL.md +++ b/dev/skills/build/SKILL.md @@ -41,7 +41,7 @@ Add any `--friction` lines per [../../references/run-journal.md](../../reference Then it ends with the same two lines, in this order - a green run, a blocked gate, and a run with open deviations all get both: -1. `Next step: run ship over this work, pointed at .dev/{plan-name}/implementation-notes.md and the {plan-name}-e2e-report.html.` Recommend it; never launch it yourself. +1. `Next step: run ship over this work, pointed at .dev/{plan-name}/implementation-notes.md and the {plan-name}-e2e-report.html.` Name `ship-quick` instead of `ship` when this run built change sets that fix an earlier `review_N.md`. Recommend it; never launch it yourself. 2. Then any question left for the user - a blocked gate, an unresolved deviation. A blocked gate never replaces line 1. diff --git a/dev/skills/reflect/SKILL.md b/dev/skills/reflect/SKILL.md index 83424ef..66009b1 100644 --- a/dev/skills/reflect/SKILL.md +++ b/dev/skills/reflect/SKILL.md @@ -1,6 +1,6 @@ --- name: reflect -description: Consolidate the run journal that scope, scope-review, build, and ship write at the end of every run into cited claims about how the skills themselves behaved, rank the recurring friction threads across repositories, and hand off one chosen thread as a scope brief with an eval case from the real run. Use periodically to decide what to improve in the dev workflow next, or with "resolve {thread} {sha}" once a fix for a thread has merged. +description: Consolidate the run journal that scope, scope-review, build, ship, and their quick variants write at the end of every run into cited claims about how the skills themselves behaved, rank the recurring friction threads across repositories, and hand off one chosen thread as a scope brief with an eval case from the real run. Use periodically to decide what to improve in the dev workflow next, or with "resolve {thread} {sha}" once a fix for a thread has merged. disable-model-invocation: true --- diff --git a/dev/skills/scope-quick/SKILL.md b/dev/skills/scope-quick/SKILL.md new file mode 100644 index 0000000..8d11387 --- /dev/null +++ b/dev/skills/scope-quick/SKILL.md @@ -0,0 +1,62 @@ +--- +name: scope-quick +description: Turn a ship review's findings, or a small request, into minimal change sets in .dev/{plan-name}/spec.md with no interview, no argued decisions, and no subagents, so build can start at once. Use when the user asks for a quick scope, wants review findings fixed without a full scope run, or has a small change that does not need its decisions argued. +disable-model-invocation: true +--- + +# Scope Quick + +Write the smallest change plan `build` can execute, and nothing else. +This is `scope` without the interview, the decision catalog and its evidence batch, the blind-spot pass, the spec reviewer, the ledger reconcile and promotion, the HTML render, and the retrospective. +Launch no subagents, and ask at most one question, only when you cannot proceed without the answer. + +First run `python3 {scope-quick-skill-root}/../../scripts/skill-metrics.py start scope-quick`. +Locate the plan directory per [../../references/plan-layout.md](../../references/plan-layout.md). + +## From review findings + +When the plan directory holds a finished `review_N.md` and the user named no other request, the highest-numbered one is the input. + +- Every Blocker becomes one change set; a Concern becomes one only when the user names it. +- Append to the existing `spec.md` change plan with continued numbering - never renumber, never rewrite what is there. +- With no `spec.md`, create one in the shape below and take the title and scope from the review's summary. +- Read the code each finding names before writing its change set, so the files and the fix are real rather than repeated from the report. + +## From a small request + +Read the code the request touches yourself, then write `spec.md` in the shape below. +Record a `D-` decision only where the code offered a real choice; most quick specs have none. + +## The spec shape + +```markdown +# {title} + +## Research + +{Decisions, or one line saying there was no real choice.} + +## Scope + +{What changes and what does not, in a few lines.} + +### Validation + +{The repo's real test and typecheck commands, discovered, one per line.} + +## Change Plan + +1. {The change in one line} + a. `{file}` - {what changes in it} + tests: [unit] {scenario} -> {expected result} +``` + +Each change set ends in exactly one `tests:` line, tagged at the lowest layer that proves it. +A change set that fixes a finding uses the review's triggering scenario as its test. + +## Finish + +Loop `python3 {scope-quick-skill-root}/../scope/scripts/lint-spec.py .dev/{plan-name}/spec.md` until it exits clean. +Stop and recommend `scope` instead of guessing when a finding needs a recorded decision flipped or user-visible scope changed, or when the request turns out not to be small. +Close with `python3 {scope-quick-skill-root}/../../scripts/skill-metrics.py end scope-quick --count change_sets=N`, pasting its table verbatim, and list the change sets you wrote. +Then recommend the next steps, never launching them: `build`, followed by `ship-quick` when these change sets fix a review, or `ship` for a change that has not been shipped before. diff --git a/dev/skills/ship-quick/SKILL.md b/dev/skills/ship-quick/SKILL.md new file mode 100644 index 0000000..7407135 --- /dev/null +++ b/dev/skills/ship-quick/SKILL.md @@ -0,0 +1,51 @@ +--- +name: ship-quick +description: Re-ship a small change with one validation run, one reviewer that checks the previous review's findings are fixed, and a pull request update followed to green, with no gauntlet, no lens panel, and no remediation loop. Use when the user asks for a quick ship, or wants fixes for an earlier ship review pushed without repeating the full ship run. +disable-model-invocation: true +--- + +# Ship Quick + +Confirm the fixes hold, update the pull request, and follow it to green. +This is `ship` without the gauntlet, the lens panel and its verifiers, the remediation loop, and the HTML report. +Invoking this skill is the task: detect the diff yourself and start immediately. + +First run `python3 {ship-quick-skill-root}/../../scripts/skill-metrics.py start ship-quick`. +Locate the plan directory and gather the branch diff per [../../references/plan-layout.md](../../references/plan-layout.md), and find the highest-numbered finished `review_N.md`, if any. + +## 1. Validate once + +Run the repository's required pull-request commands at the diff's impact per [../../references/ci-parity.md](../../references/ci-parity.md). +Fix what fails and re-run that command until it is green. + +## 2. One reviewer + +Launch a single fresh-context, read-only agent with the spec, the previous review, and the diff. + +- With a previous review: it reports every Blocker and Concern in it as fixed or still open, each with the evidence, then reads the diff since the review's recorded `head` for new breakage. +- Without one, or when the review records no `head`: it reads the whole branch diff. +- New findings are blockers only, each with the scenario that triggers it - no concerns, no nits, no simplification. + +Confirm a new blocker by reading the code yourself before it counts; drop what you cannot confirm. + +## 3. Report + +Write `.dev/{plan-name}/review_N.md` at the next free index in the shape of [../ship/references/report-format.md](../ship/references/report-format.md), with `Panel: quick` and the Previous findings section filled. +The verdict is BLOCK when a blocker stands, CONCERNS when only concerns remain open, and PASS otherwise. +On BLOCK, stop here: report the standing blockers and recommend `scope-quick`, then `build`, then this skill again. + +## 4. Pull request + +Follow [../ship/references/pull-request.md](../ship/references/pull-request.md) for the commit and push rules: never force-push, never rebase. + +- A pull request exists: edit `.dev/{plan-name}/pr.md` in place - the review line of the Summary and the Open calls - and keep its Evidence and Quality sections as the earlier run wrote them. + Loop `python3 {ship-quick-skill-root}/../ship/scripts/pr-evidence.py check .dev/{plan-name}/pr.md` until it passes, then update the body. +- No pull request yet: create it as that reference describes. + +Then follow the required checks to green per "Follow the PR to green" in ci-parity.md, and mark the pull request ready once they are. + +## Wrap up + +Close with `python3 {ship-quick-skill-root}/../../scripts/skill-metrics.py end ship-quick --count findings_fixed=N --count findings_open=N --count new_blockers=N`, pasting its table verbatim. +Then give the verdict, the findings still open, the path of the review file, and the pull request URL with the state of its checks. +Say plainly that this run skipped the gauntlet and the panel, and recommend `ship` when the change has grown beyond the findings it set out to fix. diff --git a/dev/skills/ship/SKILL.md b/dev/skills/ship/SKILL.md index 44bf76c..d40548a 100644 --- a/dev/skills/ship/SKILL.md +++ b/dev/skills/ship/SKILL.md @@ -142,5 +142,5 @@ For the pull request: its URL, draft or ready, what the Evidence section shows a Recommend next steps, never invoking them: - `commit` for the gauntlet's accumulated fixes, only when phase 3 did not run. -- `scope` on this plan directory when the user accepts findings needing real work - remediation is its job, even when no spec exists. +- `scope-quick`, then `build`, then `ship-quick` when the user accepts findings that are plain defects; `scope` on this plan directory only when a finding needs a decision or a scope change - both work even when no spec exists. - As independent optional next steps rather than a mandatory chain: `to-pitch` when the change needs buy-in from someone who wasn't in this conversation, and `to-quiz` when a reviewer wants a comprehension check before merging. diff --git a/dev/skills/ship/references/report-format.md b/dev/skills/ship/references/report-format.md index f342893..73d75a2 100644 --- a/dev/skills/ship/references/report-format.md +++ b/dev/skills/ship/references/report-format.md @@ -9,7 +9,7 @@ Phase 2 opens it when the panel is selected, holding `Verdict: IN PROGRESS` and Verdict: {IN PROGRESS | PASS | CONCERNS | BLOCK} Panel: {lenses run}, {failed lenses if any} -Base: {branch or PR}, {date} +Base: {branch or PR}, {date}, head {sha of the reviewed commit} {1-2 sentence summary} @@ -48,4 +48,5 @@ What to fix first and why. ``` `scope` reads this file when the user accepts findings that need real work: its Blockers and Concerns open that run's interview, and its Decision reconciliation section is what gets applied to `docs/decisions.md`. +`scope-quick` reads the same Blockers to write one fix change set each, and `ship-quick` diffs from the `head` on the Base line to review only what changed since. Write it for that reader - a finding with no triggering scenario cannot become a decision. diff --git a/docs/architecture.md b/docs/architecture.md index c7ccc33..80e4029 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -1,11 +1,11 @@ # Architecture Purpose: this repository is a monorepo of agent skills and the tooling that ships them. -Three plugins live here: `dev`, the hand-invoked development workflow (scope, scope-review, build, ship, commit, reflect, and two presentation skills); `factory`, whose `run` skill orchestrates copies of the same four phases unattended; and `bootstrap`, whose `agents-md` skill writes a consuming repository's root AGENTS.md from the workflow's lessons. +Three plugins live here: `dev`, the hand-invoked development workflow (scope, scope-review, build, ship, the lighter scope-quick and ship-quick, commit, reflect, and two presentation skills); `factory`, whose `run` skill orchestrates copies of the same four phases unattended; and `bootstrap`, whose `agents-md` skill writes a consuming repository's root AGENTS.md from the workflow's lessons. A person invokes each `dev` skill by hand; skills never invoke each other, and each one recommends the next step instead - except the factory `run` skill, the one sanctioned invoker, which launches its copied phase skills by path and judges their completion itself. The generated Codex distributions under `plugins/` are built from their source plugin and never edited by hand. -Captured: 2026-09-17 (full, scope) - Updated: 2026-10-05 (runs journal their friction; the reflect skill consolidates it into cited claims) +Captured: 2026-09-17 (full, scope) - Updated: 2026-10-06 (scope-quick and ship-quick, the short loop for small changes and review fixes) ## Components diff --git a/package.json b/package.json index c927a06..6b44b51 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "@tobrun/dev-workflow", - "version": "3.7.0", + "version": "3.8.0", "description": "Test-focused development workflow skills for Claude Code, Codex, and Pi.", "keywords": [ "pi-package", diff --git a/plugins/dev/.codex-plugin/plugin.json b/plugins/dev/.codex-plugin/plugin.json index 9a69ca0..0da85e0 100644 --- a/plugins/dev/.codex-plugin/plugin.json +++ b/plugins/dev/.codex-plugin/plugin.json @@ -1,6 +1,6 @@ { "name": "dev", - "version": "3.7.0", + "version": "3.8.0", "description": "Development workflow skills: scope changes with argued decisions, build across unit/integration/e2e with every scenario proven by tests, ship with a deterministic quality gauntlet and an adversarially verified review, create structured commits that feed a decision ledger, and render pitches or comprehension quizzes.", "author": { "name": "Tobrun" diff --git a/plugins/dev/references/plan-layout.md b/plugins/dev/references/plan-layout.md index 09a9500..ecb5adb 100644 --- a/plugins/dev/references/plan-layout.md +++ b/plugins/dev/references/plan-layout.md @@ -7,15 +7,15 @@ This reference owns the layout, the locating convention, and the diff scope; ski | File | Written by | Read by | | ---- | ---------- | ------- | -| `spec.md` | `scope` (from the interview on); `scope-review` (verified refinements only) | everyone downstream | +| `spec.md` | `scope` (from the interview on); `scope-review` (verified refinements only); `scope-quick` (minimal change sets, appended when a spec exists) | everyone downstream | | `spec-review_N.md` | `scope-review` (next free index, opened before round 1) | `scope` (remediation), re-reviews | | `implementation-notes.md` | `build` (append-only) | `ship`, `to-pitch`, `to-quiz` | -| `review_N.md` | `ship` (next free index, opened when the panel is selected) | `scope` (remediation), re-reviews | -| `pr.md` | `ship` (opened at the start of a run that reaches phase 3, overwritten per run) | the PR tool via `--body-file`; re-runs | +| `review_N.md` | `ship` (next free index, opened when the panel is selected); `ship-quick` (next free index) | `scope` and `scope-quick` (remediation), re-reviews | +| `pr.md` | `ship` (opened at the start of a run that reaches phase 3, overwritten per run); `ship-quick` (review line and Open calls only) | the PR tool via `--body-file`; re-runs | | `.dev/config.json` | the user | `skill-metrics.py` (run journal opt-out, per [run-journal.md](run-journal.md)) | Each producing skill also renders an HTML companion under `/tmp/{project-slug}/reports/` per [reporting.md](reporting.md), named by that skill. -One writer per file; every other skill only reads. +One writer per file at a time; every other skill only reads. ## Written as the run goes diff --git a/plugins/dev/skills/build/SKILL.md b/plugins/dev/skills/build/SKILL.md index d165491..cbcc1ef 100644 --- a/plugins/dev/skills/build/SKILL.md +++ b/plugins/dev/skills/build/SKILL.md @@ -40,7 +40,7 @@ Add any `--friction` lines per [../../references/run-journal.md](../../reference Then it ends with the same two lines, in this order - a green run, a blocked gate, and a run with open deviations all get both: -1. `Next step: run ship over this work, pointed at .dev/{plan-name}/implementation-notes.md and the {plan-name}-e2e-report.html.` Recommend it; never launch it yourself. +1. `Next step: run ship over this work, pointed at .dev/{plan-name}/implementation-notes.md and the {plan-name}-e2e-report.html.` Name `ship-quick` instead of `ship` when this run built change sets that fix an earlier `review_N.md`. Recommend it; never launch it yourself. 2. Then any question left for the user - a blocked gate, an unresolved deviation. A blocked gate never replaces line 1. diff --git a/plugins/dev/skills/reflect/SKILL.md b/plugins/dev/skills/reflect/SKILL.md index e809e4b..e09baab 100644 --- a/plugins/dev/skills/reflect/SKILL.md +++ b/plugins/dev/skills/reflect/SKILL.md @@ -1,6 +1,6 @@ --- name: reflect -description: Consolidate the run journal that scope, scope-review, build, and ship write at the end of every run into cited claims about how the skills themselves behaved, rank the recurring friction threads across repositories, and hand off one chosen thread as a scope brief with an eval case from the real run. Use periodically to decide what to improve in the dev workflow next, or with "resolve {thread} {sha}" once a fix for a thread has merged. +description: Consolidate the run journal that scope, scope-review, build, ship, and their quick variants write at the end of every run into cited claims about how the skills themselves behaved, rank the recurring friction threads across repositories, and hand off one chosen thread as a scope brief with an eval case from the real run. Use periodically to decide what to improve in the dev workflow next, or with "resolve {thread} {sha}" once a fix for a thread has merged. --- # Reflect diff --git a/plugins/dev/skills/scope-quick/SKILL.md b/plugins/dev/skills/scope-quick/SKILL.md new file mode 100644 index 0000000..41a796c --- /dev/null +++ b/plugins/dev/skills/scope-quick/SKILL.md @@ -0,0 +1,61 @@ +--- +name: scope-quick +description: Turn a ship review's findings, or a small request, into minimal change sets in .dev/{plan-name}/spec.md with no interview, no argued decisions, and no subagents, so build can start at once. Use when the user asks for a quick scope, wants review findings fixed without a full scope run, or has a small change that does not need its decisions argued. +--- + +# Scope Quick + +Write the smallest change plan `build` can execute, and nothing else. +This is `scope` without the interview, the decision catalog and its evidence batch, the blind-spot pass, the spec reviewer, the ledger reconcile and promotion, the HTML render, and the retrospective. +Launch no subagents, and ask at most one question, only when you cannot proceed without the answer. + +First run `python3 {scope-quick-skill-root}/../../scripts/skill-metrics.py start scope-quick`. +Locate the plan directory per [../../references/plan-layout.md](../../references/plan-layout.md). + +## From review findings + +When the plan directory holds a finished `review_N.md` and the user named no other request, the highest-numbered one is the input. + +- Every Blocker becomes one change set; a Concern becomes one only when the user names it. +- Append to the existing `spec.md` change plan with continued numbering - never renumber, never rewrite what is there. +- With no `spec.md`, create one in the shape below and take the title and scope from the review's summary. +- Read the code each finding names before writing its change set, so the files and the fix are real rather than repeated from the report. + +## From a small request + +Read the code the request touches yourself, then write `spec.md` in the shape below. +Record a `D-` decision only where the code offered a real choice; most quick specs have none. + +## The spec shape + +```markdown +# {title} + +## Research + +{Decisions, or one line saying there was no real choice.} + +## Scope + +{What changes and what does not, in a few lines.} + +### Validation + +{The repo's real test and typecheck commands, discovered, one per line.} + +## Change Plan + +1. {The change in one line} + a. `{file}` - {what changes in it} + tests: [unit] {scenario} -> {expected result} +``` + +Each change set ends in exactly one `tests:` line, tagged at the lowest layer that proves it. +A change set that fixes a finding uses the review's triggering scenario as its test. + +## Finish + +Loop `python3 {scope-quick-skill-root}/../scope/scripts/lint-spec.py .dev/{plan-name}/spec.md` until it exits clean. +Stop and recommend `scope` instead of guessing when a finding needs a recorded decision flipped or user-visible scope changed, or when the request turns out not to be small. +Close with `python3 {scope-quick-skill-root}/../../scripts/skill-metrics.py end scope-quick --count change_sets=N`, pasting its table verbatim, and list the change sets you wrote. +Then recommend the next steps, never launching them: `build`, followed by `ship-quick` when these change sets fix a review, or `ship` for a change that has not been shipped before. diff --git a/plugins/dev/skills/scope-quick/agents/openai.yaml b/plugins/dev/skills/scope-quick/agents/openai.yaml new file mode 100644 index 0000000..019714b --- /dev/null +++ b/plugins/dev/skills/scope-quick/agents/openai.yaml @@ -0,0 +1,6 @@ +interface: + display_name: "Scope Quick" + short_description: "Write minimal change sets with no interview" + default_prompt: "Use $dev:scope-quick to turn the review findings or this small request into minimal change sets." +policy: + allow_implicit_invocation: false diff --git a/plugins/dev/skills/ship-quick/SKILL.md b/plugins/dev/skills/ship-quick/SKILL.md new file mode 100644 index 0000000..76b77c3 --- /dev/null +++ b/plugins/dev/skills/ship-quick/SKILL.md @@ -0,0 +1,50 @@ +--- +name: ship-quick +description: Re-ship a small change with one validation run, one reviewer that checks the previous review's findings are fixed, and a pull request update followed to green, with no gauntlet, no lens panel, and no remediation loop. Use when the user asks for a quick ship, or wants fixes for an earlier ship review pushed without repeating the full ship run. +--- + +# Ship Quick + +Confirm the fixes hold, update the pull request, and follow it to green. +This is `ship` without the gauntlet, the lens panel and its verifiers, the remediation loop, and the HTML report. +Invoking this skill is the task: detect the diff yourself and start immediately. + +First run `python3 {ship-quick-skill-root}/../../scripts/skill-metrics.py start ship-quick`. +Locate the plan directory and gather the branch diff per [../../references/plan-layout.md](../../references/plan-layout.md), and find the highest-numbered finished `review_N.md`, if any. + +## 1. Validate once + +Run the repository's required pull-request commands at the diff's impact per [../../references/ci-parity.md](../../references/ci-parity.md). +Fix what fails and re-run that command until it is green. + +## 2. One reviewer + +Launch a single fresh-context, read-only agent with the spec, the previous review, and the diff. + +- With a previous review: it reports every Blocker and Concern in it as fixed or still open, each with the evidence, then reads the diff since the review's recorded `head` for new breakage. +- Without one, or when the review records no `head`: it reads the whole branch diff. +- New findings are blockers only, each with the scenario that triggers it - no concerns, no nits, no simplification. + +Confirm a new blocker by reading the code yourself before it counts; drop what you cannot confirm. + +## 3. Report + +Write `.dev/{plan-name}/review_N.md` at the next free index in the shape of [../ship/references/report-format.md](../ship/references/report-format.md), with `Panel: quick` and the Previous findings section filled. +The verdict is BLOCK when a blocker stands, CONCERNS when only concerns remain open, and PASS otherwise. +On BLOCK, stop here: report the standing blockers and recommend `scope-quick`, then `build`, then this skill again. + +## 4. Pull request + +Follow [../ship/references/pull-request.md](../ship/references/pull-request.md) for the commit and push rules: never force-push, never rebase. + +- A pull request exists: edit `.dev/{plan-name}/pr.md` in place - the review line of the Summary and the Open calls - and keep its Evidence and Quality sections as the earlier run wrote them. + Loop `python3 {ship-quick-skill-root}/../ship/scripts/pr-evidence.py check .dev/{plan-name}/pr.md` until it passes, then update the body. +- No pull request yet: create it as that reference describes. + +Then follow the required checks to green per "Follow the PR to green" in ci-parity.md, and mark the pull request ready once they are. + +## Wrap up + +Close with `python3 {ship-quick-skill-root}/../../scripts/skill-metrics.py end ship-quick --count findings_fixed=N --count findings_open=N --count new_blockers=N`, pasting its table verbatim. +Then give the verdict, the findings still open, the path of the review file, and the pull request URL with the state of its checks. +Say plainly that this run skipped the gauntlet and the panel, and recommend `ship` when the change has grown beyond the findings it set out to fix. diff --git a/plugins/dev/skills/ship-quick/agents/openai.yaml b/plugins/dev/skills/ship-quick/agents/openai.yaml new file mode 100644 index 0000000..916114d --- /dev/null +++ b/plugins/dev/skills/ship-quick/agents/openai.yaml @@ -0,0 +1,6 @@ +interface: + display_name: "Ship Quick" + short_description: "Verify fixes and update the PR, no gauntlet" + default_prompt: "Use $dev:ship-quick to verify the fixes with one reviewer, update the pull request, and follow it to green." +policy: + allow_implicit_invocation: false diff --git a/plugins/dev/skills/ship/SKILL.md b/plugins/dev/skills/ship/SKILL.md index 2e675f4..a42f134 100644 --- a/plugins/dev/skills/ship/SKILL.md +++ b/plugins/dev/skills/ship/SKILL.md @@ -141,5 +141,5 @@ For the pull request: its URL, draft or ready, what the Evidence section shows a Recommend next steps, never invoking them: - `commit` for the gauntlet's accumulated fixes, only when phase 3 did not run. -- `scope` on this plan directory when the user accepts findings needing real work - remediation is its job, even when no spec exists. +- `scope-quick`, then `build`, then `ship-quick` when the user accepts findings that are plain defects; `scope` on this plan directory only when a finding needs a decision or a scope change - both work even when no spec exists. - As independent optional next steps rather than a mandatory chain: `to-pitch` when the change needs buy-in from someone who wasn't in this conversation, and `to-quiz` when a reviewer wants a comprehension check before merging. diff --git a/plugins/dev/skills/ship/references/report-format.md b/plugins/dev/skills/ship/references/report-format.md index f342893..73d75a2 100644 --- a/plugins/dev/skills/ship/references/report-format.md +++ b/plugins/dev/skills/ship/references/report-format.md @@ -9,7 +9,7 @@ Phase 2 opens it when the panel is selected, holding `Verdict: IN PROGRESS` and Verdict: {IN PROGRESS | PASS | CONCERNS | BLOCK} Panel: {lenses run}, {failed lenses if any} -Base: {branch or PR}, {date} +Base: {branch or PR}, {date}, head {sha of the reviewed commit} {1-2 sentence summary} @@ -48,4 +48,5 @@ What to fix first and why. ``` `scope` reads this file when the user accepts findings that need real work: its Blockers and Concerns open that run's interview, and its Decision reconciliation section is what gets applied to `docs/decisions.md`. +`scope-quick` reads the same Blockers to write one fix change set each, and `ship-quick` diffs from the `head` on the Base line to review only what changed since. Write it for that reader - a finding with no triggering scenario cannot become a decision. diff --git a/scripts/build_codex_plugin.py b/scripts/build_codex_plugin.py index ac16205..bfd0d8c 100644 --- a/scripts/build_codex_plugin.py +++ b/scripts/build_codex_plugin.py @@ -61,6 +61,11 @@ def destination(self) -> Path: "Spec a change by arguing its decisions", "Use $dev:scope to spec this change with argued decisions and a change plan.", ), + "scope-quick": ( + "Scope Quick", + "Write minimal change sets with no interview", + "Use $dev:scope-quick to turn the review findings or this small request into minimal change sets.", + ), "scope-review": ( "Scope Review", "Review and auto-refine a settled spec", @@ -71,6 +76,11 @@ def destination(self) -> Path: "Harden, review, then open a PR with proof", "Use $dev:ship to run the quality gauntlet, the verified review, and open the pull request with evidence for this change.", ), + "ship-quick": ( + "Ship Quick", + "Verify fixes and update the PR, no gauntlet", + "Use $dev:ship-quick to verify the fixes with one reviewer, update the pull request, and follow it to green.", + ), "reflect": ( "Reflect", "Turn run friction into cited skill fixes",