Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
40 changes: 3 additions & 37 deletions dev/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

Development workflow skills for Claude Code, Codex, opencode, and Pi, built around two ideas: layered tests are the enforceable spec for behavior, and every phase produces something a human actually reviews as HTML, not markdown scrolling.

The skills chain loosely rather than as a rigid pipeline: `/scope` interviews for the real problem, argues every design decision against alternatives, and writes a self-contained spec whose change plan carries layer-tagged test scenarios; `/scope-review` puts the settled spec through a fresh-context, adversarially verified agent panel that checks the plan against the actual repo and refines the spec in place, looping without a human and closing with a short interview for the few findings only the user can decide, so a finished run hands `build` a spec ready to implement; `/build` executes the spec's change sets across unit/integration/e2e, proving every scenario with a real test at its tagged layer, in parallel waves where file lists allow, keeping a running implementation-notes log; `/ship` runs a deterministic quality gauntlet - the repo's own static analysis, security scan, dead code, duplication, dependency rules, coverage-weighted complexity, flakiness, mutation testing - looping fix agents until the checkers pass, then verifies the result with a fan-out review panel that checks spec conformance, e2e coverage, and logged deviations; `/commit` groups pending changes into granular commits with structured what/why messages; `/reflect` consolidates the journal every run leaves into cited claims about the skills themselves and hands the most recurrent one to `scope` as a brief; `/to-pitch` and `/to-quiz` turn finished work into a buy-in doc or a comprehension check.
The skills chain loosely rather than as a rigid pipeline: `/scope` interviews for the real problem, argues every design decision against alternatives, and writes a self-contained spec whose change plan carries layer-tagged test scenarios; `/scope-review` is an optional step for a large or complex change, never a gate in front of `build`: it puts the settled spec through a fresh-context, adversarially verified agent panel that checks the plan against the actual repo and refines the spec in place, looping without a human and closing with a short interview for the few findings only the user can decide, so a finished run hands `build` a spec ready to implement; `/build` executes the spec's change sets across unit/integration/e2e, proving every scenario with a real test at its tagged layer, in parallel waves where file lists allow, keeping a running implementation-notes log; `/ship` runs a deterministic quality gauntlet - the repo's own static analysis, security scan, dead code, duplication, dependency rules, coverage-weighted complexity, flakiness, mutation testing - looping fix agents until the checkers pass, then verifies the result with a fan-out review panel that checks spec conformance, e2e coverage, and logged deviations; `/commit` groups pending changes into granular commits with structured what/why messages; `/reflect` consolidates the journal every run leaves into cited claims about the skills themselves and hands the most recurrent one to `scope` as a brief; `/to-pitch` and `/to-quiz` turn finished work into a buy-in doc or a comprehension check.
The durable context is deliberately small: the code, its tests, the active spec under `.dev/{plan-name}/`, and three repo-tracked registries the skills maintain in the consuming project - `docs/decisions.md` (design decisions with their argued alternatives, read only after a review forms its findings), `docs/contracts.md` (boundary guarantees, read as premises before a review walks the diff), and `docs/dependencies.md` (machine-checkable module dependency rules, enforced by `ship`).
The files under `.dev/{plan-name}/` are written as a run goes, not when a stage closes: the spec opens during the interview, a report opens before its panel returns, the implementation notes gain an entry per change set and per fixup, and the PR body fills check by check, so a run can be followed from its files and a dead session loses only what was in flight.
Alongside them, `docs/architecture.md` is a plain high-level overview of the system - components, flows, boundaries, entry points - captured in full the first time a skill needs it and finds it absent, then kept current by build and commit whenever the structure changes, with a small checker that catches stale paths and files no component covers.
Expand All @@ -21,6 +21,7 @@ Promotes durable decisions to `docs/decisions.md` and cross-boundary invariants

### scope-review

Optional: `build` runs on any spec `scope` finished, and `scope` advises this review only when the change is large or complex.
Reviews a settled spec with agents that did not write it and refines it in place, before build starts - the cheapest review in the chain, since a defect caught here costs a spec edit instead of a re-implementation, and the loop makes those edits itself.
Gates on `scope`'s own `lint-spec.py` first (a mechanically unsettled spec is sent back, not reviewed), then loops up to two rounds of panel, verification, and refinement: four lenses in parallel - feasibility (does the plan survive contact with the repo: named files, assumed hooks, claimed prior art, stated premises about current behavior), completeness (failure paths, second-order work, and affected call sites the spec never mentions), consistency (change sets versus their linked decisions and the project's ledgers, semantically), and testability (will the `tests:` lines produce real proof at their tagged layers) - with every BLOCK and CONCERN adversarially verified on ship's transport and aggregation machinery before a refine agent may touch the spec.
Refinements follow a strict authority order (recorded user intent beats the repo's reality beats settled decisions beats spec prose) and edit `spec.md` only; anything that would flip a decision, change user-visible scope, or add a dependency is never auto-applied - the run ends by asking the user those as decisions, one at a time with alternatives and tradeoffs, and applies the answers to the spec before finishing.
Expand Down Expand Up @@ -84,41 +85,6 @@ Every run appends a row to `.dev/metrics.jsonl` in the consuming repository, and
The same call appends an entry to the cross-repository run journal under `~/.dev-workflow/memory/journal/` (or `$DEV_MEMORY_DIR`): the friction signals measured from the transcript (interrupts, denied and failed tool calls, the user's own turns, and how often each deterministic checker ran and failed), each with its transcript line, plus at most three `--friction` lines in which the skill names where the run fought its own instructions.
That journal is what `reflect` consolidates. It stays on the machine; set `DEV_MEMORY_DIR=off`, or `"memory": {"enabled": false}` in `.dev/config.json`, to keep a repository out of it.

## Jira integration

The spec-driven workflow can mirror its local state to Jira through the
Atlassian CLI (`acli`). Install and authenticate `acli` before enabling it.
Create `.dev/config.json` in the consuming repository:

```json
{
"jira": {
"enabled": true,
"site": "acme.atlassian.net",
"project": "PROJ"
}
}
```

`site` is optional and `project` is required when enabled. An absent config
file or `jira.enabled: false` keeps the workflow pure-local with no Jira
calls or questions.

The hierarchy is Initiative > Epic > Task. `scope` asks you to choose
an existing open Initiative, creates one Epic under it once the change plan is
final, stores the Epic key in the spec, and creates one Jira Task per change
set, storing each key under its change set and closing old issues when a
change set is superseded. `build` moves the Epic to In Progress at start,
moves each change-set issue through In Progress and Done as the orchestrator
dispatches and commits it. `ship` then pushes the keyed branch and opens a PR
whose title starts with the Epic key.

One-time Jira administration is required. Install the GitHub for Jira
integration, then add an automation rule: "when a linked pull request is
merged, transition the Epic to Done". The Epic key in the branch name and PR
title is what lets Jira link the pull request and trigger that rule. The
skills do not poll for merges.

## Claude Code installation

```bash
Expand Down Expand Up @@ -151,7 +117,7 @@ done
ln -sfn ~/ws/workflow/dev/references ~/.config/opencode/references
```

The last symlink keeps the shared references (`jira.md`, `decision-ledger.md`, `contracts.md`, `architecture.md`) reachable through the `../../references/` links inside the skills.
The last symlink keeps the shared references (`decision-ledger.md`, `contracts.md`, `architecture.md`) reachable through the `../../references/` links inside the skills.
Symlinks mean a `git pull` updates the skills in place; restart opencode afterwards, since skills load at startup.

opencode has no `disable-model-invocation` field (it is ignored harmlessly); skills load through a model-invoked `skill` tool.
Expand Down
19 changes: 0 additions & 19 deletions dev/evals/build.json
Original file line number Diff line number Diff line change
Expand Up @@ -101,25 +101,6 @@
"No required command whose inputs the change reaches is deferred because it is slow",
"The closing message points at ship, whose PR's CI runs the deferred commands"
]
},
{
"id": "jira-disabled",
"prompt": "Implement this spec.",
"fixture": "a git repo with a keyed .dev/{plan-name}/spec.md, no .dev/config.json, and no acli executable on PATH",
"assertions": [
"The transcript contains no Jira, acli, Initiative, Epic, or issue-tracker mention",
"After the e2e report the user is asked whether to push and open the PR, and no push or PR command occurs before a yes"
]
},
{
"id": "jira-enabled",
"prompt": "Implement this spec.",
"fixture": "a git repo with .dev/config.json enabling Jira for PROJ, spec.md with Jira: PROJ-101 under its title and Jira: PROJ-201 / Jira: PROJ-202 under change sets 1 and 2, and the stub acli on PATH",
"assertions": [
"The stub log transitions the Epic PROJ-101 to In Progress before change-set transitions",
"For each change set, the stub log shows In Progress at dispatch before Done after that change set's verification and commit",
"The final PR question is asked after the e2e report, and the proposed branch name and PR title start with the Epic key"
]
}
]
}
39 changes: 0 additions & 39 deletions dev/evals/fixtures/acli

This file was deleted.

17 changes: 8 additions & 9 deletions dev/evals/scope.json
Original file line number Diff line number Diff line change
Expand Up @@ -31,15 +31,6 @@
"The retrospective reports the three typed lines: contamination, sizing, catalog gaps"
]
},
{
"id": "jira-disabled",
"prompt": "We want to add rate limiting to our public API to stop abuse. Spec this change.",
"fixture": "a repo with src/api/server.js and no .dev/config.json, with no acli executable on PATH",
"assertions": [
"The transcript contains no Jira, acli, Initiative, Epic, or issue-tracker mention",
"No acli invocation occurs and the spec is written using the normal local workflow"
]
},
{
"id": "remediation-from-review",
"prompt": "Take the accepted findings from the last review and turn them into new change sets.",
Expand Down Expand Up @@ -86,6 +77,14 @@
"Answer 4: it is handed only the original request and the relevant code, and is told to stay out of .dev/, where the catalog is being written",
"Answer 5: a spec.md that lint-spec.py does not pass, holding everything established before the session died; it is a draft that scope picks up where it stopped, and no other skill builds on it"
]
},
{
"id": "scope-echo-matches-chosen",
"prompt": "Answer briefly, from the skill text and the references it links only. A decision's chosen line reads `✓ S3 bucket in us-east-1 - durable across redeploys`. (1) Write the echo a later section should use for it. (2) You were about to write `(✓ cloud storage)` instead. Will lint-spec.py accept that, and what does the skill tell you to do when writing the chosen line so the echo passes the first time?",
"assertions": [
"Answer 1: `D-slug (✓ S3 bucket)` or a shorter prefix such as `(✓ S3)` - text taken from the chosen line before its ` - `, not from the because clause",
"Answer 2: no, lint-spec.py rejects it because the echo is not found in the chosen line's text before ` - `; lead the chosen line with the short label that will be echoed, and copy the echo from it"
]
}
]
}
106 changes: 106 additions & 0 deletions dev/evals/tests/test_lint_spec.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,106 @@
"""Unit tests for lint-spec.py's echo check: a mismatched echo names what it should be.

Run from the repo root: python3 -m unittest discover -s dev/evals/tests -p 'test_lint_spec.py' -t .
Each test writes a small spec to a scratch directory and runs the script as a
subprocess, so it proves the CLI contract the scope skill loops against.
"""

from __future__ import annotations

import subprocess
import sys
import tempfile
import unittest
from pathlib import Path

REPO_ROOT = Path(__file__).resolve().parents[3]
SCRIPT = REPO_ROOT / "dev" / "skills" / "scope" / "scripts" / "lint-spec.py"

CHOSEN = """D-storage: Where do files live?
✓ S3 bucket - durable across redeploys
✗ local disk - lost on redeploy
"""
OPEN = """D-storage: Where do files live? [open]
? S3 bucket - durable across redeploys
⚑ ask: expected retention?
"""
NOT_DOING = """D-storage: Where do files live?
✗ S3 bucket - no upload feature yet
⊘ not doing - nobody asked for uploads ? verify: support tickets; reopen on the first request
"""


def spec(decision: str, echo: str) -> str:
return f"""# Title

## Research

{decision}
## Scope

Files go where D-storage ({echo}) says.
"""


class LintSpecEchoTest(unittest.TestCase):
def setUp(self) -> None:
self.tmp = tempfile.TemporaryDirectory()
self.addCleanup(self.tmp.cleanup)

def lint(self, decision: str, echo: str) -> subprocess.CompletedProcess:
path = Path(self.tmp.name) / "spec.md"
path.write_text(spec(decision, echo), encoding="utf-8")
return subprocess.run(
[sys.executable, str(SCRIPT), str(path)], capture_output=True, text=True, check=False
)

def echo_errors(self, result: subprocess.CompletedProcess) -> str:
return "\n".join(line for line in result.stdout.splitlines() if "echo" in line)

def test_echo_copied_from_the_chosen_line_passes(self) -> None:
self.assertEqual(self.echo_errors(self.lint(CHOSEN, "✓ S3")), "")

def test_echo_not_in_the_chosen_line_names_the_chosen_text(self) -> None:
result = self.lint(CHOSEN, "✓ cloud")
self.assertIn("does not match its resolution", self.echo_errors(result))
self.assertIn("s3 bucket", self.echo_errors(result))

def test_chosen_echo_on_an_open_decision_expects_open(self) -> None:
self.assertIn("expected (open)", self.echo_errors(self.lint(OPEN, "✓ S3")))

def test_chosen_echo_on_a_not_doing_decision_expects_not_doing(self) -> None:
self.assertIn("expected (⊘ not doing)", self.echo_errors(self.lint(NOT_DOING, "✓ S3")))

def test_bare_check_mark_on_a_chosen_decision_passes(self) -> None:
self.assertEqual(self.echo_errors(self.lint(CHOSEN, "✓")), "")

def test_open_echo_on_a_chosen_decision_expects_the_chosen_text(self) -> None:
result = self.lint(CHOSEN, "open")
self.assertIn("does not match its resolution", self.echo_errors(result))
self.assertIn("s3 bucket", self.echo_errors(result))

def test_change_plan_linking_an_open_decision_is_flagged(self) -> None:
path = Path(self.tmp.name) / "spec.md"
path.write_text(
spec(OPEN, "open") + "\n## Change plan\n\n1. Store files\n a. `src/store.py` - write files - decisions: D-storage (open)\n",
encoding="utf-8",
)
result = subprocess.run([sys.executable, str(SCRIPT), str(path)], capture_output=True, text=True, check=False)
self.assertIn("the change plan links D-storage, which is still open or flagged", result.stdout)

def test_scope_echo_of_an_open_decision_is_not_a_change_plan_link(self) -> None:
self.assertNotIn("change plan links", self.lint(OPEN, "open").stdout)

def test_change_plan_linking_an_open_unflagged_decision_is_flagged(self) -> None:
path = Path(self.tmp.name) / "spec.md"
open_only = OPEN.replace(" ⚑ ask: expected retention?\n", "")
path.write_text(
spec(open_only, "open") + "\n## Change plan\n\n1. Store files\n a. `src/store.py` - write files - decisions: D-storage (open)\n",
encoding="utf-8",
)
result = subprocess.run([sys.executable, str(SCRIPT), str(path)], capture_output=True, text=True, check=False)
self.assertIn("the change plan links D-storage", result.stdout)


if __name__ == "__main__":
unittest.main()
1 change: 1 addition & 0 deletions dev/evals/tests/test_memory.py
Original file line number Diff line number Diff line change
Expand Up @@ -60,6 +60,7 @@ class Sandbox:
"""A scratch repo with a transcript where skill-metrics.py will find it."""

def __init__(self, tmp: Path, rows: list[dict] = TRANSCRIPT):
tmp = tmp.resolve() # git reports the real path; macOS temp dirs sit behind a symlink
self.repo = tmp / f"repo-{os.getpid()}-{id(self)}"
self.repo.mkdir()
subprocess.run(["git", "init", "-q"], cwd=self.repo, check=True)
Expand Down
Loading
Loading