Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

| Plugin | Use When | Tools |
| ------ | -------- | ----- |
| [dev](dev/) | A test-focused development workflow for Claude Code, Codex, opencode, and Pi. | `scope`, `scope-review`, `commit`, `build`, `ship`, `reflect`, `to-pitch`, `to-quiz` |
| [dev](dev/) | A test-focused development workflow for Claude Code, Codex, opencode, and Pi. | `scope`, `scope-quick`, `scope-review`, `commit`, `build`, `ship`, `ship-quick`, `reflect`, `to-pitch`, `to-quiz` |
| [factory](factory/) | Take a request from scope to a shipped pull request unattended, on Claude Code or Codex. | `run` |
| [bootstrap](bootstrap/) | Prepare any repository for agent work: probe its stack and write its root AGENTS.md from the workflow's SDLC lessons, on Claude Code or Codex. | `agents-md` |

Expand Down
2 changes: 1 addition & 1 deletion dev/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "dev",
"version": "3.7.0",
"version": "3.8.0",
"description": "Development workflow skills: scope changes with argued decisions, build across unit/integration/e2e with every scenario proven by tests, ship with a deterministic quality gauntlet and an adversarially verified review, create structured commits that feed a decision ledger, and render pitches or comprehension quizzes.",
"author": {
"name": "Tobrun"
Expand Down
25 changes: 22 additions & 3 deletions dev/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

Development workflow skills for Claude Code, Codex, opencode, and Pi, built around two ideas: layered tests are the enforceable spec for behavior, and every phase produces something a human actually reviews as HTML, not markdown scrolling.

The skills chain loosely rather than as a rigid pipeline: `/scope` interviews for the real problem, argues every design decision against alternatives, and writes a self-contained spec whose change plan carries layer-tagged test scenarios; `/scope-review` is an optional step for a large or complex change, never a gate in front of `build`: it puts the settled spec through a fresh-context, adversarially verified agent panel that checks the plan against the actual repo and refines the spec in place, looping without a human and closing with a short interview for the few findings only the user can decide, so a finished run hands `build` a spec ready to implement; `/build` executes the spec's change sets across unit/integration/e2e, proving every scenario with a real test at its tagged layer, in parallel waves where file lists allow, keeping a running implementation-notes log; `/ship` runs a deterministic quality gauntlet - the repo's own static analysis, security scan, dead code, duplication, dependency rules, coverage-weighted complexity, flakiness, mutation testing - looping fix agents until the checkers pass, then verifies the result with a fan-out review panel that checks spec conformance, e2e coverage, and logged deviations; `/commit` groups pending changes into granular commits with structured what/why messages; `/reflect` consolidates the journal every run leaves into cited claims about the skills themselves and hands the most recurrent one to `scope` as a brief; `/to-pitch` and `/to-quiz` turn finished work into a buy-in doc or a comprehension check.
The skills chain loosely rather than as a rigid pipeline: `/scope` interviews for the real problem, argues every design decision against alternatives, and writes a self-contained spec whose change plan carries layer-tagged test scenarios; `/scope-review` is an optional step for a large or complex change, never a gate in front of `build`: it puts the settled spec through a fresh-context, adversarially verified agent panel that checks the plan against the actual repo and refines the spec in place, looping without a human and closing with a short interview for the few findings only the user can decide, so a finished run hands `build` a spec ready to implement; `/build` executes the spec's change sets across unit/integration/e2e, proving every scenario with a real test at its tagged layer, in parallel waves where file lists allow, keeping a running implementation-notes log; `/ship` runs a deterministic quality gauntlet - the repo's own static analysis, security scan, dead code, duplication, dependency rules, coverage-weighted complexity, flakiness, mutation testing - looping fix agents until the checkers pass, then verifies the result with a fan-out review panel that checks spec conformance, e2e coverage, and logged deviations; `/scope-quick` and `/ship-quick` are the short loop for a change that is already small - most often the fixes a `ship` review asked for - writing minimal change sets with no interview, and re-shipping with one reviewer instead of the gauntlet and panel; `/commit` groups pending changes into granular commits with structured what/why messages; `/reflect` consolidates the journal every run leaves into cited claims about the skills themselves and hands the most recurrent one to `scope` as a brief; `/to-pitch` and `/to-quiz` turn finished work into a buy-in doc or a comprehension check.
The durable context is deliberately small: the code, its tests, the active spec under `.dev/{plan-name}/`, and three repo-tracked registries the skills maintain in the consuming project - `docs/decisions.md` (design decisions with their argued alternatives, read only after a review forms its findings), `docs/contracts.md` (boundary guarantees, read as premises before a review walks the diff), and `docs/dependencies.md` (machine-checkable module dependency rules, enforced by `ship`).
The files under `.dev/{plan-name}/` are written as a run goes, not when a stage closes: the spec opens during the interview, a report opens before its panel returns, the implementation notes gain an entry per change set and per fixup, and the PR body fills check by check, so a run can be followed from its files and a dead session loses only what was in flight.
Alongside them, `docs/architecture.md` is a plain high-level overview of the system - components, flows, boundaries, entry points - captured in full the first time a skill needs it and finds it absent, then kept current by build and commit whenever the structure changes, with a small checker that catches stale paths and files no component covers.
Expand All @@ -19,6 +19,13 @@ A checker (`scripts/lint-spec.py`) enforces the spec's mechanics - unique slugs,
On a clean spec it prints the build waves the file lists allow and the shared files that make change sets wait, so the plan is shaped for parallel work before build starts.
Promotes durable decisions to `docs/decisions.md` and cross-boundary invariants to `docs/contracts.md`, renders an expandable-card spec view, and has a reverse mode that audits the implicit decisions already embedded in existing code.

### scope-quick

The minimal `scope`: turns the Blockers of the latest `ship` review, or a small request, into change sets in `.dev/{plan-name}/spec.md` that `build` can execute straight away.
It runs no interview, argues no decisions, launches no subagents, and skips the ledger work, the HTML render, and the retrospective; a decision is recorded only where the code offered a real choice.
Review fixes are appended to the existing change plan with continued numbering, each carrying the review's triggering scenario as its test, and the same `lint-spec.py` loop keeps the result mechanically sound.
When a finding needs a recorded decision flipped or user-visible scope changed, it stops and points at `scope` instead of guessing.

### scope-review

Optional: `build` runs on any spec `scope` finished, and `scope` advises this review only when the change is large or complex.
Expand Down Expand Up @@ -54,6 +61,14 @@ Phase 3 commits what the gauntlet fixed, pushes, and opens the pull request auto
A deterministic check (`pr-evidence.py check`) gates the PR body, so a PR cannot open on a placeholder or a data URI, and the phase then follows required checks to green.
Local gates in every phase run at the change's impact, so the full merge gate is the PR's own CI: a PR carrying a check deferred to CI opens as a draft and is marked ready once its required checks pass, and a run that opens no PR runs the full set locally.

### ship-quick

The minimal `ship`, for re-shipping after the fixes a review asked for: one validation run at the diff's impact, one read-only reviewer, and a pull request update followed to green.
The reviewer reports each finding of the previous `review_N.md` as fixed or still open and reads only the diff since that review's recorded head for new blockers; it raises no concerns, nits, or simplifications.
There is no gauntlet, no lens panel, no remediation loop, and no HTML report: a standing blocker ends the run with a pointer back to `scope-quick`.
An existing pull request keeps the Evidence and Quality sections the full `ship` run wrote, with only its review line and open calls updated.
Run the full `ship` for a first ship, or once the change has grown beyond the findings it set out to fix.

### commit

Groups all pending changes into granular, logically-separate commits - splitting within a file when needed - with structured messages: a `type(scope):` subject, `What:`/`Why:` body, optional `Considered:`/`Constraint:`/`Directive:`/`Symptoms:` sections, and `Severity:`/`Risk:` metadata trailers.
Expand All @@ -63,7 +78,7 @@ Pushes by default; say "commit only" to skip the push.
### reflect

Turns what past runs measured into evidence about the skills themselves, so the next change to the workflow fixes something that actually recurred.
Every `scope`, `scope-review`, `build`, and `ship` run appends a journal entry to `~/.dev-workflow/memory/` (see Run metrics below); `reflect` has read-only subagents read the transcript around each measured signal and propose claims such as "build reran the full Validation block after each change set; the user stopped it", each quoting its source.
Every `scope`, `scope-review`, `build`, and `ship` run, and every run of their quick variants, appends a journal entry to `~/.dev-workflow/memory/` (see Run metrics below); `reflect` has read-only subagents read the transcript around each measured signal and propose claims such as "build reran the full Validation block after each change set; the user stopped it", each quoting its source.
A checker (`scripts/claims.py add`) accepts a batch only when every quote is found verbatim at the transcript line or journal entry it cites, then rebuilds per-skill pages and a ranked list of threads in which every line cites a claim id.
The user picks a thread, or retracts a claim that misreads its evidence (a retracted claim stays on record so the same evidence cannot bring it back), and the skill writes a scope brief with an eval case built from the real run.
It never edits a skill: the fix goes through `scope` and `build` in this repository, and `/dev:reflect resolve {thread} {sha}` records the commit, after which a new claim in that thread reopens it.
Expand All @@ -79,7 +94,7 @@ Cannot enforce a merge gate, so it says so plainly and produces an honest pass/f

## Run metrics

`scope`, `scope-review`, `build`, and `ship` each start by snapshotting the run with `scripts/skill-metrics.py start` and end by printing what it measured: wall time, tokens split between the orchestrator and its subagents, agents dispatched, tool calls, the git delta since the snapshot, and any counters the skill tallied from tool output.
`scope`, `scope-review`, `build`, `ship`, `scope-quick`, and `ship-quick` each start by snapshotting the run with `scripts/skill-metrics.py start` and end by printing what it measured: wall time, tokens split between the orchestrator and its subagents, agents dispatched, tool calls, the git delta since the snapshot, and any counters the skill tallied from tool output.
The numbers come from the session transcript under `$CLAUDE_CONFIG_DIR` (default `~/.claude`) and from git, never from the model's recollection.
Every run appends a row to `.dev/metrics.jsonl` in the consuming repository, and the table compares the run against the median of earlier runs of the same skill, which is where a skill's cost and catch rate become visible over time.
The same call appends an entry to the cross-repository run journal under `~/.dev-workflow/memory/journal/` (or `$DEV_MEMORY_DIR`): the friction signals measured from the transcript (interrupts, denied and failed tool calls, the user's own turns, and how often each deterministic checker ran and failed), each with its transcript line, plus at most three `--friction` lines in which the skill names where the run fought its own instructions.
Expand Down Expand Up @@ -129,9 +144,13 @@ Preserve the explicit-invocation policy with a permission rule in `~/.config/ope
"skill": {
"*": "allow",
"scope": "ask",
"scope-quick": "ask",
"scope-review": "ask",
"commit": "ask",
"build": "ask",
"ship": "ask",
"ship-quick": "ask",
"reflect": "ask",
"to-pitch": "ask",
"to-quiz": "ask"
}
Expand Down
4 changes: 2 additions & 2 deletions dev/evals/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,12 +4,12 @@ Status: maintained

Eval definitions for the `dev` plugin's skills: realistic prompts and objective assertions used to check whether a skill change preserved behavior.

- `{skill}.json` - one file per skill: the eval prompt(s), the fixture each expects, and the assertions to grade the output against. Covers all 8 skills: `scope`, `scope-review`, `commit`, `build`, `ship`, `reflect`, `to-pitch`, `to-quiz`.
- `{skill}.json` - one file per skill: the eval prompt(s), the fixture each expects, and the assertions to grade the output against. Covers all 10 skills: `scope`, `scope-quick`, `scope-review`, `commit`, `build`, `ship`, `ship-quick`, `reflect`, `to-pitch`, `to-quiz`.
- `results.md` - the record of the most recent full run: scores, methodology, and findings.
- `tests/` - unit tests for the deterministic scripts the skills loop against (`lint-spec.py`, `change-set-brief.py`, `check-tests.py`, `pr-evidence.py`, `impact-scope.py`, and the memory scripts `skill-metrics.py` and `claims.py`), run by `scripts/validate.sh` as check D01.

`build` runs in `"functional"` mode (a real fixture, a real subagent run, assertions checked against the actual output).
`scope`, `scope-review`, `commit`, `ship`, `reflect`, `to-pitch`, and `to-quiz` run in `"comprehension"` mode instead - each depends on either an interactive question loop, a live codebase, or prior artifacts (a finished spec, implementation notes, an e2e report) that are too expensive to stage on every iteration, so these check policy comprehension of the skill text directly.
`scope`, `scope-quick`, `scope-review`, `commit`, `ship`, `ship-quick`, `reflect`, `to-pitch`, and `to-quiz` run in `"comprehension"` mode instead - each depends on either an interactive question loop, a live codebase, or prior artifacts (a finished spec, implementation notes, an e2e report) that are too expensive to stage on every iteration, so these check policy comprehension of the skill text directly.

## When to run these

Expand Down
28 changes: 28 additions & 0 deletions dev/evals/scope-quick.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
{
"skill": "scope-quick",
"mode": "comprehension",
"evals": [
{
"id": "scope-quick-from-review",
"prompt": "Answer briefly, from the skill text only. The plan directory holds spec.md with change sets 1 to 4, all built, and review_2.md with two Blockers and three Concerns. The user runs scope-quick with no other request. (1) How many change sets does it write, and how are they numbered? (2) What is each one's test? (3) Does it interview the user, argue decisions, or launch subagents? (4) What does it loop against before finishing? (5) What does it recommend next?",
"fixture": "none - the answers come from the skill text",
"assertions": [
"Answer 1: two, one per Blocker, appended as change sets 5 and 6 with nothing renumbered or rewritten; Concerns are left out unless the user names them",
"Answer 2: the review's triggering scenario, on a single tests: line tagged at the lowest layer that proves it",
"Answer 3: no to all three; it asks at most one question and only when it cannot proceed",
"Answer 4: scope's lint-spec.py against spec.md, until it exits clean",
"Answer 5: build, then ship-quick, recommended and never launched"
]
},
{
"id": "scope-quick-escalates",
"prompt": "Answer briefly, from the skill text only. (1) A Blocker can only be fixed by reversing a decision the spec recorded. What does scope-quick do? (2) A small request has one obvious implementation. How many D- decisions does the spec record, and which sections must it still have? (3) Name three things scope does that scope-quick skips.",
"fixture": "none - the answers come from the skill text",
"assertions": [
"Answer 1: it stops and recommends scope instead of guessing",
"Answer 2: none; the spec still has Research, Scope with a Validation block of the repo's real commands, and a Change Plan",
"Answer 3: any three of the interview, the decision catalog and evidence batch, the blind-spot pass, the spec reviewer, the ledger reconcile and promotion, the HTML render, the retrospective"
]
}
]
}
29 changes: 29 additions & 0 deletions dev/evals/ship-quick.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
{
"skill": "ship-quick",
"mode": "comprehension",
"evals": [
{
"id": "ship-quick-after-fixes",
"prompt": "Answer briefly, from the skill text only. review_1.md ended in BLOCK with two Blockers and one Concern and records a head commit; build has since committed two fix change sets, and a draft pull request exists. The user runs ship-quick. (1) Which gauntlet tools run? (2) How many review agents launch, and what do they read? (3) What kinds of new finding may the reviewer raise? (4) Both Blockers are fixed and the Concern is still open: what is the verdict and where is it written? (5) What happens to the pull request body?",
"fixture": "none - the answers come from the skill text",
"assertions": [
"Answer 1: none; only the repository's required pull-request commands run, once, at the diff's impact",
"Answer 2: one fresh-context read-only agent; it checks each previous Blocker and Concern as fixed or still open and reads the diff since the recorded head",
"Answer 3: blockers only, each with a triggering scenario, and each confirmed by the orchestrator reading the code; no concerns, nits, or simplifications",
"Answer 4: CONCERNS, in review_2.md at the next free index with Panel: quick and the Previous findings section filled",
"Answer 5: pr.md is edited in place - the review line and Open calls - keeping the earlier Evidence and Quality sections, then pr-evidence.py check passes before the body is updated and the checks are followed to green"
]
},
{
"id": "ship-quick-limits",
"prompt": "Answer briefly, from the skill text only. (1) The reviewer reports one Blocker still open. Does ship-quick fix it and re-review? What does it do? (2) There is no earlier review_N.md. What does the reviewer read? (3) May it force-push? (4) What must the wrap-up say about what this run did not do, and when does it recommend the full ship?",
"fixture": "none - the answers come from the skill text",
"assertions": [
"Answer 1: no - there is no remediation loop; it writes the BLOCK verdict, stops before the pull request step, and recommends scope-quick, then build, then ship-quick again",
"Answer 2: the whole branch diff",
"Answer 3: no - never force-push, never rebase",
"Answer 4: that the gauntlet and the panel were skipped; ship is recommended when the change has grown beyond the findings it set out to fix"
]
}
]
}
Loading
Loading