feat(drafting-a-skill): close Dimension 4-8 gaps + redivide edit responsibility - #1632
feat(drafting-a-skill): close Dimension 4-8 gaps + redivide edit responsibility#1632tvna wants to merge 67 commits into
Conversation
Documents the gap evaluating-skill-quality's process-blind grading exposes in drafting-a-skill's writing-time guidance across Dimensions 4-8, plus the unowned existing-skill-edit case between drafting-a-skill and scorer-gated-skill-edits. Refs #1630 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Euzusy42MsxhCaikPpHCri
2-task decomposition (Task A: drafting-a-skill formative-quality- dimensions.md rows 4-8 + SKILL.md style/worked-example/eval-scaffold; Task B: executing-a-branch-plan existing-skill-edit process floor), no file-ownership or interface-dependency edge between them, both assigned to wave 1. Refs #1630 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Euzusy42MsxhCaikPpHCri
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Team Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #1632 +/- ##
=======================================
Coverage 99.66% 99.66%
=======================================
Files 152 152
Lines 24792 24792
Branches 2981 2981
=======================================
Hits 24710 24710
Misses 82 82 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
…ture worked example Extend formative-quality-dimensions.md row 4 (terminology, checklists, feedback loops, templates, branch triggers reusing Step 3's cohesion enumeration, plus the relocated concrete-example content), rows 5-7 (reference naming/link-context; forward-slash paths, Server:tool naming, no-install-assumed phrasing, default+escape-hatch framing, Portable convention/authority-path rules; script error handling, no unexplained constants, Interface-vs-Implementation comments, single ownership), and rewrite row 8 into an eval-preparation instruction (>=3 scenarios including the guardrail case, an evals/<skill>/ fixture skeleton). Add one SKILL.md Step 2 style bullet cross-referencing row 4 without restating it, a new Step 6 sub-step for the eval-scaffolding preparation, and restructure the Worked example section's three candidates from prose into per-candidate Step-keyed bullet lists with every stated fact preserved unchanged. Both skills/evaluating-skill-quality shape/drift checkers run clean against skills/drafting-a-skill/. Refs #1630
…tep 3 tasks Issue #1630, ACM row 3: a task decomposed at Step 3 whose Planned ops edit an existing SKILL.md, and is not itself routed to scorer-gated-skill-edits by a stated scorer/held-out-split precondition, now has a minimum process floor instead of none. - Step 3's own body now also classifies whether a task's Planned ops edit an existing SKILL.md, pointing to task-decomposition.md's full rule set (kept to one added line to stay inside the body-length checker's 500-line cap, 500/500 after this change). - task-decomposition.md gains a new "Existing-skill-file edit floor" section, sibling to Irreversibility classification: such a task must run gitapex_check_skill_shape.py against the edited skill directory and sweep every touched section against drafting-a-skill's formative-quality-dimensions.md before it may report complete -- a minimum floor only, never a re-entry into drafting-a-skill's Design-by-Contract structure or scorer-gated-skill-edits's own measured-iteration procedure. - metadata/gitapex.yaml: added scorer-gated-skill-edits to relatedTo and a new decision reference recording the change. Placement: SKILL.md's Related-skills section already carries the sibling "new skill directory -> drafting-a-skill" rule, but the 500-line body cap left no room there for a second full bullet, so the actionable rule content lives in task-decomposition.md (also the better fit per that file's own Load-bearing vs. on-demand dimension -- this floor only matters for a subset of tasks); Step 3's own body carries the one-line pointer that keeps it discoverable from the same place the irreversibility classification already is. gitapex_check_skill_shape.py: 61/61 checks pass against skills/executing-a-branch-plan. Full repo suite (pytest, 7828 tests; gitapex_gate_local_preflight.py, 44/44 gates) passes. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…ll-implementation-lkdypz
Step 6's eval-scaffolding bullet duplicated formative-quality- dimensions.md row 8's own cell (scenario count, guardrail case, one-prompt-file-per-scenario, preparation-only boundary) across nine lines, while the adjacent Step 2 bullet already points at row 4 without restating it. Rewritten as the same pointer shape. Found by the Step 8 refactor/simplify pass over the accumulated diff for issue #1630. Refs #1630 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Euzusy42MsxhCaikPpHCri
…-8/floor conflict, fixture-layout mismatch, dispatch gap
Independently-executed Task A (rewrote formative-quality-dimensions.md
row 8 into an artifact-producing "Eval preparation" instruction) and
Task B (added an existing-SKILL.md-edit floor that sweeps every row of
that same table) each had no visibility into the other's diff. Read
together, the floor obliged every ordinary SKILL.md edit task to author
evals/<skill>/ fixtures -- contradicting the design doc's own "minimum
floor... nothing more" and its Non-goal against new eval-execution
infrastructure. The floor now excludes row 8 by name.
Also fixed, found by the same Step 8 adversarial review:
- Row 8 named a fixture layout this repository does not use (a
per-scenario *.md file directly under evals/<skill>/); corrected to
the real, gate-enforced shape (eval.yaml beside tasks/<scenario>.yaml).
- drafting-a-skill Step 6's own completion criterion did not mention
the new eval-scaffold sub-step, so a draft with no scaffold satisfied
every stated finish line; the criterion now names it.
- The floor's obligation had no delivery path to the worktree-isolated
task agent that must actually run it; now stated as an in-band
dispatch-prompt citation, the same pattern SKILL.md step 6 already
uses for code-quality-principles.md.
- A load-timing note in formative-quality-dimensions.md's own opening
("not required reading before Step 2 begins") appeared to contradict
Step 2's new row-4 pointer; reconciled inline.
Both skills' metadata/gitapex.yaml decision logs updated per each
skill's own append-in-the-same-round convention.
Refs #1630
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Euzusy42MsxhCaikPpHCri
…ple-list form Trim SKILL.md's frontmatter description to state only what triggers it and what it does, dropping the workflow-summary sentence dimension 2's own conciseness bar flags (was 99.3% of the 1024-char cap). Restyle the body toward the more scannable table/list/bullet form used by writing-skills, adding a Steps summary table and an inline body-vs-metadata table, and converting Postcondition, Related skills, and Notes from prose paragraphs into bulleted/tabular form. A dedicated content-preservation review caught real regressions in the first pass: dropped Precondition/Postcondition definitions, a dropped anti-abuse guard on Step 7's escalation branch, dropped completion loops in Steps 4/6, and terminology/wording drift (including two eval-fixture-matching regressions surfaced only by the full pytest run: a header line-wrap change and a "fresh dispatch"/"independent, fresh dispatch" word-order change that broke exact-substring fixture matching). All are restored; both skill-shape and execution-requirements-drift checkers pass clean, and the full pytest suite plus local preflight (44/44 gates) pass. Refs #1630
CI status at
|
Repository owner determined the Related-skills Live-collision bullet's routing-risk framing is unnecessary: this skill is dispatched via executing-a-branch-plan Step 6's own explicit branch-plan-task agent() call, not selected through model description-routing, so a still-present vendored writing-skills copy or skill-creator (verified still present on disk in this session's own environment, despite apm.yml/apm.lock.yaml no longer declaring the retired obra/superpowers dependency) creates no actual routing collision for this skill. Removed the bullet. Its still-valid never-import-RED-GREEN-REFACTOR fact (writing-skills' own testing methodology, deliberately not adopted here) folded into the existing Notes Attribution bullet instead of being dropped outright. Recorded as a decision entry in metadata/gitapex.yaml. Both shape checkers clean (53/53, no drift), full pytest suite (7828 passed), local preflight (44/44 gates) pass. Refs #1630
…er form Restyled the Worked example section further toward writing-skills own presentational patterns: the first candidate's Step 2 now shows its actual frontmatter (name/description) as a fenced YAML code block instead of describing it in prose; the second candidate's earning-test failure is now a Bad/Good fenced-code-block pair (an over-populated Precondition/Postcondition/Non-goals section vs. the corrected single-Step form), rather than one prose bullet each. File grew from 392 to 425 lines, well inside BODY_MAX_LINES=500, so no other section needed trimming. Both shape checkers clean (53/53, no drift), full pytest suite (7828 passed), local preflight (44/44 gates) pass. Refs #1630
…ide in this skill
Repository owner determined the ~72-column manual line-wrap this
skill's own SKILL.md and references/ files carried was an
unenforced, non-repository-wide habit (docs/skill-authoring-
standards.md states no such rule, gitapex_check_skill_shape.py's 53
checks include no line-length check, and a comparison skill --
eliciting-a-design/SKILL.md -- routinely carries 100-400+ char lines
with no wrap at all). Since this skill is meant to be used as the
vehicle for refactoring other existing skills going forward, an
unexamined formatting habit here risked propagating into every
future draft.
Reflowed every prose paragraph and bullet in SKILL.md and all five
references/*.md files into single natural-length lines (frontmatter,
tables, and fenced code blocks left untouched). Verified
content-preservation two independent ways (whitespace-normalized
exact match; a hyphen-aware "-\n" -> "-" regex reconstruction
compared against the same normalization) for every file -- both
methods agree exactly, zero content drift.
Found and fixed a latent rendering bug the old hard-wrap habit
carried: 8 line-wrap points fell exactly on a hyphen inside a
compound identifier (untrusted-input-triage,
formative-quality-dimensions.md, mechanism-fit-and-cohesion.md,
gitapex-cross-links.md, references/contract-structure.md), which
CommonMark's own soft-line-break-to-space rule was already silently
corrupting into "untrusted-input- triage" etc. at render time. Fixed
by treating a trailing single hyphen (not "--") as a mid-word join
point during reflow.
The Worked example's own embedded frontmatter code block also had
its `description:` field un-wrapped to one line, and reordered to
lead with the trigger clause ("Use whenever...") before the
capability clause ("Explains..."), matching this skill's own real
frontmatter's trigger-first shape and dimension 1's what+when
requirement -- the prior order buried the trigger condition after a
workflow-summary-shaped opening clause.
gitapex_check_skill_shape.py's Check E (exercises-declaration
absolute resolution) keys Stop-boundary bullets and Procedure/Steps
items on their full single-physical-line text; un-wrapping changed
every multi-line bullet's identity, so all six
evals/drafting-a-skill/tasks/*.yaml fixtures whose `expected.exercises`
declared a now-stale line-wrapped fragment are updated to the new
full-line text (exact string, casefolded).
Both shape checkers clean (53/53, no drift), full pytest suite (7828
passed), local preflight (44/44 gates) pass.
Refs #1630
…e fix Repository owner: a description's "what it does" clause should be minimized wherever possible, and its "when" clause may be omitted entirely for a pipeline-only skill like this one, since no model ever reads it to make a live routing decision. Applied to this skill's own description: dropped the workflow-summary sentence and all three Distinct-from clauses, keeping only the pipeline-only dispatch-condition sentence (675 -> 250 chars, 24.4% of the shape checker's cap). Separately, corrected Step 4's own "Real collision found?" remedy, which had conflated two different things: a Distinct-from clause is a targeted fallback for invocation triggers that stay genuinely adjacent even after narrowing, not a routine response to two skills merely sounding functionally similar. This skill's own now-removed Distinct-from clauses were exactly that misuse. Also trimmed the Worked example's curl-command-explainer sample description to its trigger clause only, as a second worked instance of the same principle. Found via CI (split-fixture-coverage, head ad834c9): the line-wrap retraction changed the Stop-boundary bullet identity gitapex_gate_split_fixture_coverage.py's delta-scoped check keys on, leaving "Never loop back into scorer-gated-skill-edits-shaped iterative editing..." with no fixture whose expected.exercises resolved to its new full-line text. Added that declaration to evals/drafting-a-skill/tasks/existing-skill-routes-away.yaml, whose own scenario already exercises exactly that boundary. Both shape checkers clean (53/53, no drift), full pytest suite (7828 passed), local preflight (44/44 gates) pass. Refs #1630
|
… boundary Repository owner requested reorganizing the responsibility split between drafting-a-skill and scorer-gated-skill-edits along a new axis: how a SKILL.md gets rewritten (drafting-a-skill, new or existing target alike) vs. evaluating the result of a change (scorer-gated-skill-edits, always opt-in), rather than the current new-skill-vs-existing-skill split. Converged through eliciting-a-design (four sections presented and approved individually) and confirmed via a clairvoyance terminal handoff. Under the new division, the "Existing-skill-file edit floor" added earlier this session (issue #1630, PR #1632) becomes redundant -- every SKILL.md edit routes to drafting-a-skill's own full procedure instead of a lighter, separate floor -- and is scoped for removal. Refs #1648
…y redivision Decomposes issue #1648's 5-row ACM into three tasks: Task A rewrites drafting-a-skill/SKILL.md's Precondition/Step 1/Step 7/Postcondition/ Related skills; Task B rewrites scorer-gated-skill-edits/SKILL.md's Step 3 and adds a new pre-ship full-review requirement, sequenced after Task A via an explicit interface-dependency edge (both tasks state the same cross-skill dispatch contract, learning from the prior Branch Plan on this branch judging an analogous pair to have no edge, which Step 8's own review later found a real collision in anyway); Task C removes the now-redundant Existing-skill-file edit floor and consolidates routing. Wave 1: {A, C}. Wave 2: {B}. Refs #1648
…f events-and-review-gate.md A Fable root-cause dispatch on the standing dimension-5 gap (issue #1648) found a real, separate residual defect: events-and-review-gate.md mixed every-run content (steps 5, 8) with roughly 200 lines an ordinary clean run never reads (loss/absence handling, freshness/hang detection, the step-7 failure-dispatch table, and rollback). Extract that content into a new failure-and-recovery.md, repoint every cross-reference (SKILL.md, threat-model-and-authorization.md, gitapex_check_task_full_verification.py, and the in-file links between the two split files), and record the change in the decision log alongside issue #1662, the new tracking issue for the rubric-level fix the same dispatch identified (dimension 5 has no satisfiable clause for a cohesion-confirmed sequential-pipeline skill over the body cap). This split is a genuine progressive-disclosure gain on its own terms, but does not by itself close dimension 5: the typical request still opens three reference files, and that gap stays open pending issue #1662. Refs #1648, #1662
Resolves the behind-base gate (issue #985) before pushing the events-and-review-gate.md split. # Conflicts: # docs/skill-eval-status.md
…g dimension 5 now passes Issue #1662's rubric-level fix merged (PR #1672). An isolated re-review against the new rubric text -- independently re-deriving cohesion, verifying both new-exemption conditions against actual measured line counts, and checking whether an easier file-count reduction is available -- confirms dimension 5 now passes for this skill's current 4-file split. Overall verdict stays WELL-FORMED-NOT-MATURE: the same review found dimension 2 (sediment/duplication/sprawl) and the Mixed-portability physical-split gap remain open, disclosed here rather than fixed. Refs #1662
…ll-implementation-lkdypz
Verification section archive (Content-preservation review through Current status)Moved out of the PR body per this repository's own token-budget constraint on Content-preservation review on the description/body rewrite (row 4). An independent subagent dispatch compared the pre-rewrite and post-rewrite Live-collision-bullet follow-up (commit Worked-example code-block follow-up (commit Line-wrap retraction (commit Trigger-only description follow-up (commit Boundary redivision (issue #1648, Wave 1 complete). Converged through Wave 1 (A and C, no interface-dependency edge between them) executed in isolated worktrees, each independently passing its own full pytest suite and Merging the two independently-clean worktrees surfaced one interaction neither task's own isolated run could catch: Re-verified on the fully merged, pushed state (head Wave 2 complete (Task B: scorer-gated-skill-edits Step 3 rewrite, commit Step 8 aggregate review (mandatory two-pass review over the full accumulated diff, per The refactor pass found and fixed 3 mid-backtick line-wrap corruptions inside Task B's own new Step 3 blockquote (commit The adversarial pass returned 10 findings (F1-F10), all verified and either fixed or explicitly disclosed as a known gap rather than silently left. Fixed, committed together as
Re-verified after scorer-gated-skill-edits line-wrap retraction (commit The reflow changed every Stop-boundary bullet's and Procedure item's first line, which Regenerated Current status. All planned implementation work for issues #1630 and #1648 is complete and re-verified on the current head (
This PR has no known content defect beyond what Generated by Claude Code |
Independent review verdict archive (Round 1 initial dispatches through Round 9)Moved out of the PR body for the same token-budget reason as the sibling Verification-section archive comment above. This comment preserves Round 1 (the initial three-skill independent dispatch), Round 2, Rounds 3-6 ( All three dispatches returned WELL-FORMED-NOT-MATURE (shape checks clean; nine-dimension maturity not fully cleared):
All three skills' shape checkers and drift scanners re-ran clean after the fixes above (53/53, 61/61, 47/47; Round 2 (commit Rounds 3-6 (commits
Every round's fix was independently re-verified in this session (not only per the dispatching sonnet review's own self-report): shape checker 49/49 throughout, no executionRequirements drift, full pytest suite (7828 passed) and all 44 Round 3 for the other two skills (commits
Both fixes independently re-verified in this session: full pytest suite (7828 passed) and all 44 Round 4 (commits
All three skills' shape checkers and drift scanners re-ran clean after these fixes; full pytest (7828 passed) and all 44 Round 5 (commits
Rounds 6/8/9 (commits
Every fix across rounds 4-9 was independently re-verified in this session: shape checkers clean throughout ( Generated by Claude Code |
…sprawl Fable-verified findings (isolated dispatch, no rubric blind spot on this dimension): threat-model-and-authorization.md carried quoted empirical-verification transcripts as changelog-style sediment, a two-variant deployment asymmetry restated in full at comparable length twice, and inlined plugin-distributed-only detail read on every run. Pruned six passages (584 -> 539 lines), preserving every fact the analysis flagged as load-bearing: the live-proxy caveat, CLAUDE_PROJECT_DIR mechanism, residual bypass ceilings, and the fuller Decision-7 restatement. No structural change; shape checker 55/55, drift scanner unchanged (4 pre-existing findings), full suite green. The same dispatch found the Mixed-portability finding is a genuine rubric blind spot instead: the file-level-split rule conflicts with dimension 5's just-closed sequential-pipeline exemption for this target's shape. Filed as its own rubric-level fix, issue #1676, via the same governed scorer-gated-skill-edits process issue #1662 used -- not implemented here. Refs #1648, #1662, #1676
The file's paragraphs were hard-wrapped at roughly 72 characters, inherited from a copy-paste-style edit when the content was split out of events-and-review-gate.md. No repository convention requires this column width in a references/*.md prose file, and the wrap occasionally broke a sentence mid-clause (e.g. a hyphenated compound split across a line boundary). Reflowed every paragraph and bullet item into a single line; headings, blank-line paragraph breaks, and the bullet lists are unchanged in structure. No wording changed -- verified content-identical modulo whitespace against the prior revision. Shape checker 54/54 (one fewer than before: the file's own TOC check no longer applies once it drops under the 100-line threshold, not a regression), drift scanner unchanged (same 4 pre-existing findings), full pytest suite (7883 passed), all 44 local-preflight gates passed. Refs #1648
…ll-implementation-lkdypz Brings in two independently-merged main-branch efforts that landed while this branch was in flight: - PR #1677 (issue #1676): rubric.md's Mixed-portability substitute for a Dimension-5-exempted target, plus that same session's own re-grade of executing-a-branch-plan's Mixed-portability finding against it (kept as its own deferral decision-log entry, appended after mine rather than overwritten). - PR #1675 (issues #1508/#1566): a new worktree-base precondition backstop (gitapex_check_task_worktree_base.py), chained into check_task_bash_safety.sh. Conflict resolution: - metadata/gitapex.yaml: kept this branch's full decision log and Frontier/requires/relatedTo values (main's own copy predates this branch's own fork and never carried any of that), appended main's new #1676 deferral entry, and added one missing decision-log entry for this branch's own prior failure-and-recovery.md reformatting commit. - references/execution-and-dispatch.md: this branch already retired this file (split into decomposition-and-dispatch.md / events-and-review-gate.md / failure-and-recovery.md, per issue #1662's own Round 10). Ported main's new "Worktree-base precondition backstop" subsection into decomposition-and-dispatch.md's own "Execution and dispatch" section (matching where its sibling subsections already live), demoted from ## to ###, and repointed every stale execution-and-dispatch.md cross-reference (in threat-model-and-authorization.md and gitapex_check_task_worktree_base.py/its test) to decomposition-and-dispatch.md. Verified post-merge: shape checker 54/54, drift scanner unchanged (same 4 pre-existing findings), full pytest suite passing (the one run mid- merge showed 7 gitapex_gate_commit_citation.py failures -- confirmed a transient artifact of running the suite while .git/MERGE_HEAD existed, merge_in_progress()'s own by-design merge-commit exemption, not a real defect; re-verified clean after this commit), all 44 local-preflight gates passed. Refs #1648, #1662, #1676
…porting-boundary-map.md Round 13: independently re-verified issue #1676's own re-grade (recorded in the decision log as a deferral entry, condition 1-2 met, requirements 1-3 not met) rather than accepting it at face value. An isolated Sonnet dispatch initially disputed condition 2 (whether the Claude-Code-specific content is reached on every ordinary run, given Step 6's Workflow-tool-vs-sequential-fallback fork), reading two passages in threat-model-and-authorization.md and decomposition-and-dispatch.md as evidence the primary path is opt-in. Direct re-read of both passages found they describe this skill's own authoring-session testing limitations (could not literally exercise the Workflow tool for live verification) and a permission-prompt UX detail ("Every run" unless pre-approved) -- neither is a deployment- time gating claim. Conditions 1-2 confirmed met. Closed the two unmet requirements: - Requirement 1 (distinct-heading isolation): split SKILL.md's Notes portability paragraph into "Non-portable (Claude-Code-specific)" and "Portable" bold-lead-in sections -- net-zero line growth (16 lines either way), keeping the body at exactly 500/500. - Requirement 3 (porting-boundary-map file): new references/porting-boundary-map.md, non-every-use, enumerating every Claude-Code-specific touchpoint (Workflow tool, agentType, isolation:'worktree', the branch-plan-task subagent type, both PreToolUse hook backstops, the SubagentStop verification hook) and its portable alternative or lack thereof. - Requirement 2 (Notes declaration) was found already largely met on independent re-read, contra the deferral entry's own claim. One authoring defect caught and fixed along the way: an inline-code span (`evaluating-skill-quality/references/rubric.md`) split across a hard-wrap line broke the shape checker's per-line backtick parsing, tripping no-bare-issue-citation on an adjacent `#1676` reference -- fixed by keeping the span on one line. Verified: shape checker 56/56, drift scanner unchanged (same 4 pre-existing findings), full pytest (7976 passed) and all 44 local-preflight gates passed. Refs #1648, #1676
…citations An independent fresh review (run to close PR #1632's own Why-NOT-CLEAN gap) found all four outcome.baseCommit values in metadata/gitapex.yaml pointed at real but unrelated commits in other skills' own history, not the true immediate parent of the commit that added each entry. Each entry's own substantive claim independently checks out against the actual commit diffs -- only the citation was broken, verified via `git show`/`git log` against every candidate. Corrected the boundary-redivision, axis-7-closed, Cleanup-git-command, and round-9 entries to their own entry-adding commit's true immediate parent. Also fixed SKILL.md Step 5's own misattribution: the `split.md` Kept-edit-log citation it names lives in this file's own step 7 (the run-record corroboration text), not step 4's worked example, which never mentions `split.md` at all. Refs #1648 shape checker: 49/49. drift scanner: clean.
…ings An independent fresh review (run to close PR #1632's own Why-NOT-CLEAN gap) found requirement 1 of the Mixed-portability substitute (rubric.md dimension 5) was overclaimed by the deferral/corroboration entries in metadata/gitapex.yaml: the Notes section used bold inline lead-ins (`**Non-portable...**`, `**Portable:**`) inside one flowing paragraph, never an actual Markdown heading -- this SKILL.md carried no `###` anywhere. rubric.md's own internal use of "heading" elsewhere supports a literal reading of the substitute's own requirement text. Promoted both labels to real `### Non-portable (Claude-Code-specific)` / `### Portable` subheadings. Net-zero body-line growth held via a tighter prose rewrap of the surrounding Notes paragraphs (no wording dropped, verified by a whole-file word-frequency diff against the prior commit). Refs #1676 shape checker: 56/56 (body-length: 500/500 lines). drift scanner: unchanged (same 4 pre-existing findings).
…ll-implementation-lkdypz # Conflicts: # skills/executing-a-branch-plan/SKILL.md # skills/executing-a-branch-plan/references/execution-and-dispatch.md # skills/executing-a-branch-plan/references/refactor-and-review-gate.md
Merge conflict resolution (commit
|
CI red:
|
Summary
Closes formative-process gaps in
drafting-a-skill's writing-timeguidance behind
evaluating-skill-quality's Dimensions 4-8 (issue#1630), then redivides responsibility between
drafting-a-skillandscorer-gated-skill-editsalong a how-to-rewrite-vs-evaluate-the-resultaxis (issue #1648), removing the now-redundant "Existing-skill-file
edit floor" issue #1630's own work added.
Facts
evaluating-skill-qualitygrades only a finished, static artifact --it never sees how a draft was produced (its own
SKILL.md: "a gateon a finished, static artifact").
skills/drafting-a-skill/references/formative-quality-dimensions.mdrow 4 currently covers only 2-3 of
rubric.mdDimension 4's 7requirements; rows 5-7 are similarly partial; row 8 is mislabeled
(its content is a Dimension 4/8-adjacent concrete-example concern,
not Dimension 8's real eval-driven-baseline content) -- confirmed by
a full read of
rubric.md(2419 lines) this session.skills/tree claims ownership of anordinary, non-eval-driven edit to an existing
SKILL.md--drafting-a-skillauthors only from a blank page,scorer-gated-skill-editsrequires a checkable scorer and held-outsplit as a hard Precondition-gate STOP -- confirmed by grepping the
full tree this session.
drafting-a-skill's owndescription(1017/1024 chars, 99.3% of theshape checker's cap) restated its workflow rather than stating only a
trigger, and its body stayed explanatory prose rather than the
scannable principle-list/table form
writing-skills(
.claude/skills/writing-skills/SKILL.md) demonstrates -- raiseddirectly by the issue owner and added as issue feat(drafting-a-skill): close formative-process gaps behind evaluating-skill-quality Dimensions 4-8, and the existing-skill-edit pipeline gap #1630's 4th ACM row.
drafting-a-skill/SKILL.md's Precondition excluded any existingSKILL.mdtarget, andscorer-gated-skill-edits/SKILL.md's own Step3 never assigned who authors a proposed patch -- the repository owner
requested reorganizing the boundary along a how-to-rewrite (
drafting-a-skill, always) vs. evaluate-the-result (scorer-gated-skill-edits, always opt-in) axis instead, convergedthrough
eliciting-a-designand confirmed via aclairvoyanceterminal handoff, filed as issue feat(drafting-a-skill,scorer-gated-skill-edits): redivide responsibility along how-to-rewrite vs evaluate-result #1648.
Assumptions
executing-a-branch-planprocess-floor rule(
task-decomposition.mdvs.SKILL.mdStep 3) is Task B's ownjudgment call, per the design doc's own Assumptions section.
branch/PR rather than a new pull request, per the repository owner's
explicit instruction to avoid a split merge to
main-- the"Existing-skill-file edit floor" issue feat(drafting-a-skill,scorer-gated-skill-edits): redivide responsibility along how-to-rewrite vs evaluate-result #1648 removes was itself added
on this branch and has never merged to
main.Risk / blast radius
Documentation/skill-definition edits only -- no runtime script behavior
changes. Blast radius is future
drafting-a-skilldispatches (now thesingle authoring path for any
SKILL.mdchange, new or existing) andfuture
scorer-gated-skill-editsiterations (now explicitly dispatchdrafting-a-skillper iteration). No change toevaluating-skill-quality'sown rubric or scripts; no change to
scorer-gated-skill-edits's ownPrecondition gate, scorer/split requirement, or gate mechanics.
Rollback
Revert this PR's merge commit; no other repository state depends on
these skill-definition edits.
Verification
Restating issue #1630's Acceptance Criteria Map, row by row:
drafting-a-skillmust instruct writing workflows as ordered/copyable checklists, consistent terminology, feedback loops, matched-strictness templates.gitapex_check_skill_shape.py/gitapex_scan_execution_requirements_drift.pyclean; row-4 bullets checked againstrubric.mdDimension 4; Worked example diffed fact-for-fact.369c406e). 53/53 shape checks passed, no drift, full pytest suite (7828) passed, 44/44 local-preflight gates passed.369c406e). Rows 5-7 extended, row 8 rewritten into eval-preparation instruction; no requirement left unmapped.task-decomposition.md's row-to-task mapping rules anddrafting-a-skill's Related-skills section.2e81efbb). New## Existing-skill-file edit floorsection added totask-decomposition.md;SKILL.mdStep 3 carries a one-line pointer (body-length cap left only 1 line of headroom). 61/61 shape checks passed. Superseded by issue #1648's own boundary redivision below -- this floor is being removed in favor of a unifieddrafting-a-skillrouting rule.description(99.3% of cap) restates workflow rather than trigger-only; body stays explanatory prose rather than a scannable principle-list/table form.descriptionre-measured (target: comfortably under the 90% Dimension-2 trigger); both shape checkers re-run; every existing cross-reference re-verified to resolve; content preserved fact-for-fact (no rule silently dropped in the compression).b383faf9, refined by3c044c0d).description: 250/1024 chars (24.4%, was 675 then 1017). Body Steps/Postcondition/Related skills/Notes restyled into tables/bullets, DbC skeleton and every citation kept intact. Both shape checkers clean (53/53, no drift). See notes below for the full follow-up chain.Re-verified post-merge against
origin/main(16-commit drift, clean auto-merge,drafting-a-skill/SKILL.mdhad a concurrent unrelated change): both shape checkers still pass (53/53, 61/61), full pytest suite (7828) still passes, all 44 local-preflight gates pass.**Detailed implementation-history follow-ups (description/body rewrite content-preservation, the Live-collision-bullet/Worked-example/Line-wrap/Trigger-only-description follow-ups, issue #1648's Wave 1/Wave 2 implementation, Step 8's aggregate refactor+adversarial review with F1-F10 findings, and the scorer-gated-skill-edits line-wrap retraction) moved to a dedicated comment -- the PR body had grown past what a single
pull_request_readcall can return, which was blocking this session's own turn-terminalmergeable_stateverification. Nothing in that comment changes what already happened; it is a relocation, not a rewrite.Current status. All planned implementation work for issues #1630 and #1648 is complete and re-verified: both shape checkers clean across all three touched skills, full pytest suite passing, all 44 local-preflight gates passing, both fixture-coverage gates passing.
drafting-a-pr-to-merge's own Step 8 independent review has now actually run (see## Independent review verdictbelow) and returned NOT-CLEAN with specific, cited findings -- the small, mechanical ones fixed, the larger structural ones deliberately deferred with the repository owner's informed choice. Two required checks remain red:independent-review-pending-- Step 8 ran; verdict is honestly NOT-CLEAN per the section below, not a silent gap.skill-audit-disclosure-- see## Skill audit evidencebelow, populated.eval-gate(not a required check) -- fails with the exacterror: model CLI exited 1signature tracked in issue fix(waza-eval-gate): live eval runs fail with "model CLI exited 1" and empty stderr, repo-wide #1304, a pre-existing, repo-wide infrastructure defect unrelated to this PR's own content.This PR has no known content defect beyond what
## Independent review verdictdiscloses, no open review thread, and no merge conflict. It remains blocked on the two items above.Independent review verdict
Outer layer (GitHub-native reviewer). GitHub Copilot review requested (
request_copilot_review) at 2026-09-01T07:44:43Z. No response posted within the 30-minute window (checked at 2026-09-01T08:47Z viaget_reviews: empty). Per drafting-a-pr-to-merge Step 8's own rule, treated as unreachable for this run -- disclosed, not silently omitted.Inner layer (reviewing-an-artifact, always runs). Step 0 eligibility check: this PR's diff is a Mixed target. The specialist-owned part (three
skills/*/SKILL.mdpackages --drafting-a-skill,executing-a-branch-plan,scorer-gated-skill-edits, each with its ownreferences/andmetadata/gitapex.yaml) was deferred toevaluating-skill-quality, per Step 0's Mixed-target handling ("defer the specialist-owned part to its own named target, and continue Steps 1-6 against the remainder"). The remainder (evals/drafting-a-skill/*,evals/scorer-gated-skill-edits/*,docs/skill-eval-status.md,docs/superpowers/plans|specs/*.md-- all YAML fixtures, eval-status prose, and design docs, no executable code) was classified safe under Step 1 (an added-test/doc-update pattern with no security-tier signal anywhere in it, confirmed by a direct scan for secrets/injection/auth-bypass patterns across the full diff of that remainder) -- Steps 2-5's fan-out correctly skipped per Step 1's own rule, recorded here rather than silently passed.Specialist deferral (
evaluating-skill-quality), three independent isolated dispatches, one per touchedSKILL.mdpackage. Isolation mechanism:claude -p(CLI 2.1.252) from a working directory whose full ancestry carries noCLAUDE.md/AGENTS.md,--allowedTools "Read,Glob,Grep"(no permission-bypass flag), pointed only at a caller-created read-only snapshot of each target skill directory plus a copy ofevaluating-skill-qualityitself -- freshly re-verified this round with a two-control positive/negative test at this exact CLI version (recorded inskills/evaluating-skill-quality/references/adversarial-self-audit.md's Known entries, commite753f9b8). Disclosed limitation, per each dispatch's own self-report: the harness'sinvoking-gitapexSessionStart-hook content (generic session-bootstrap procedure text, not target-specific framing or prior discussion) was still present in each dispatch's context -- this is a hook-injected mechanism distinct from theCLAUDE.md/AGENTS.mdfile-based leak the verified recipe closes, and was not itself excluded this round. Per Contaminated-dispatch disclosure, every finding below is accordingly provisional pending a genuinely hook-free re-run, though each dispatch reports it observed no redirection of its own grading judgment.Round 1 through Round 9 (the initial three-skill independent dispatch, then repeated from-scratch re-reviews) moved to a dedicated comment, for the same token-budget reason as the Verification-section archive above. Summary: all three skills (
drafting-a-skill,executing-a-branch-plan,scorer-gated-skill-edits) returned WELL-FORMED-NOT-MATURE on the initial dispatch; nine further from-scratch rounds (mostly concentrated onscorer-gated-skill-edits, whose concurrency-safety axis took until round 9 to genuinely close, plus three rounds returning todrafting-a-skillandexecuting-a-branch-plan) found and fixed real defects each time -- never only re-confirming the prior round's own claim. By round 9,executing-a-branch-plan's dimension 5 (progressive disclosure) stood as the single largest remaining structural gap: rubric.md's own "if acting on the typical request needs three files open, the split is wrong" line, which a 6-to-3 reference-file merge could not itself satisfy. See the archived comment for the full, cited round-by-round record.Round 10 (commit
d2eeb059): a Fable root-cause dispatch investigates whether dimension 5's own repeatedexecuting-a-branch-planfinding is a rubric blind spot. Directed by the repository owner to check the rubric itself before any further restructuring of the skill. One isolated, read-only Fable dispatch (Read,Glob,Greponly, no write access) readrubric.md's full dimension 5 section plus enough of the rest of the rubric to know whether dimension 5 stands alone or interacts with another dimension, then readexecuting-a-branch-plan's currentSKILL.mdand all three reference files in full, weighing the counter-case seriously (is this skill genuinely a no-internal-branch-structure pipeline the rubric was never designed to judge, or is the repeated failure a correct signal that a 3+-mandatory-file skill is itself the defect) before concluding: dimension 5's own three internally-inconsistent thresholds, and its own cohesion check's Restraint paragraph (which explicitly defers multi-reference-file judgment to dimension 5 for exactly this skill's cohesion type), together hand judgment to dimension 5 for a cohesion-confirmed sequential pipeline with no caller-selectable narrower path and roughly 1,400 lines of genuinely mandatory content against the 500-line body cap -- but dimension 5 currently has no clause that can pass it. Exhaustive case analysis (in the dispatch's own report) shows every possible file-layout arrangement fails some dimension-5 clause. This is a genuine blind spot, not a misapplied rubric. Per this repository's own governance -- a shared, repository-wide rubric file used to review every skill goes throughscorer-gated-skill-edits's own measured, held-out-gated edit process, not an ad hoc direct edit inside an unrelated PR -- filed as a new, separate issue (#1662), not attempted here, per the repository owner's own explicit direction to pursue the rubric fix separately and fix only this PR's own residual defect now.The same dispatch separately found a real, independently actionable defect squarely within this PR's own scope:
events-and-review-gate.mdmixed every-run content (steps 5, 8) with roughly 200 lines of failure/resume-only content (loss/absence handling, freshness/hang detection, the step-7 failure-dispatch table, rollback) an ordinary clean run never reads -- a genuine progressive-disclosure improvement independent of whether the rubric ever changes. Fixed: extracted that content into a newfailure-and-recovery.md, repointed every cross-reference (SKILL.md,threat-model-and-authorization.md,gitapex_check_task_full_verification.py, and the in-file links between the two split files), and recorded the change in the decision log alongside issue #1662 (d2eeb059). Four reference files exist now, but the typical request still opens exactly three (threat-model-and-authorization.md,decomposition-and-dispatch.md,events-and-review-gate.md) -- this split does not itself close dimension 5's three-file reading, which stays open pending issue #1662's rubric-level fix. Verified: shape checker 55/55 (up from 52/52 -- the new file's own three per-reference checks, the same +3-per-file pattern the earlier 6-to-3 merge showed in reverse), drift scanner shows only the same 4 pre-existing, confirmed-unrelatedexecutionRequirementsfindings, full pytest (7883 passed, the count reflecting fixtures merged in fromorigin/main) and all 44gitapex_gate_local_preflight.pygates passed on commita0addc2c. Not yet independently re-reviewed by a fresh isolated round; disclosed as such, per this session's own consistent finding that every fix so far has needed at least one further round to fully confirm.Round 11 (commit
851182ec): issue #1662's rubric fix merged (PR #1672); an isolated re-review against the new rubric confirms dimension 5 now passes. Per the repository owner's own request, ran an isolated, read-only Sonnet dispatch (Read,Glob,Greponly) against a fresh snapshot pairingorigin/main's currentevaluating-skill-quality(carrying issue #1662's new Dimension-5 exemption bullet) with this branch's own currentexecuting-a-branch-plan(the 4-file split from Round 10). The dispatch independently re-derived cohesion at the rubric's own Procedure step 2 (not from the target's own "sequential" framing) and confirmed single-outcome sequential cohesion, no caller-selectable narrower path. It then independently verified both of the new exemption's conditions against actual measured content, not assumption: condition 1 (the cohesion check already found sequential/functional) holds from its own step-2 finding; condition 2 (every-use reference content demonstrably exceedsBODY_MAX_LINESeven after real dimension-2 pruning) holds too -- measured every-use total (threat-model-and-authorization.md+decomposition-and-dispatch.md+events-and-review-gate.md) at 1,513 lines, still roughly 1,050-1,150 lines after applying concrete sediment/duplication/sprawl cuts it identified directly in the content, more than double the 500-line cap with zero room left inSKILL.md's own body (exactly 500/500 lines). It also independently re-verifiedfailure-and-recovery.mdis genuinely excluded from the every-use count -- grepped every cross-reference into it and confirmed none requires opening it to complete the ordinary path -- and checked for an easier file-count reduction (a 2-file or 1-file merge), finding none: a 1-file merge is already rejected in this skill's own decision log for risking the opposite-direction failure, and a 2-file merge would blend genuinely distinct domain boundaries (threat/authorization, decomposition/dispatch, events/review-gate) against dimension 5's own "organised by domain" rule. Dimension 5 now PASSES under the new exemption -- the split is reasonably minimized given a genuinely irreducible floor, not merely declared so.The same dispatch found this skill's dimension 2 (conciseness) does not clear cleanly: real, cited sediment (quoted empirical-verification transcripts in
threat-model-and-authorization.md), duplication (a two-variant-asymmetry pattern restated in full at least twice), and sprawl (plugin-only deployment detail inlined into every-use content rather than split by variant). It also found the Mixed-portability declaration (Step 4's own check) is not physically carried out for this skill: Claude-Code-specific mechanics (the Workflow tool,agentType,isolation:'worktree', thebranch-plan-tasksubagent-type machinery) are disclosed inSKILL.md's own Notes prose but not split into a dedicated reference file, as the rubric's own Mixed rule names. Neither is fixed this round -- both disclosed here as this skill's new standing gaps, replacing dimension 5 as the largest one. Methodology caveat the dispatch disclosed about its own verdict: it ran inside a single dispatch process rather than the target skill's own prescribed further-nested-Agent-dispatch isolation for its own review steps, and could not executegitapex_check_skill_shape.pydirectly in its sandboxed snapshot (applied its checks by hand against the same rules instead). This session independently re-ran the real shape checker and drift scanner against the actual repository afterward and recorded the confirmed verdict inmetadata/gitapex.yaml(commit851182ec): shape checker 55/55, drift scanner shows only the same 4 pre-existing, confirmed-unrelatedexecutionRequirementsfindings, full pytest (7883 passed) and all 44gitapex_gate_local_preflight.pygates passed on commit3bb34b84.Round 12 (commit
88908059): this skill's own dimension-2 finding fixed directly; the Mixed-portability finding investigated for a rubric blind spot and found to have one. Per the repository owner's own choice between round 11's two findings, dispatched one isolated, read-only Fable dispatch (Read,Glob,Greponly) to check whether either finding -- dimension 2 (sediment/duplication/sprawl) or the Mixed-portability physical-split requirement -- reveals a genuine rubric blind spot for this skill's specific shape, rather than being straightforwardly fixable defects. Conclusion, stated separately for each: dimension 2 is a correctly-flagged, ordinary defect, not a blind spot -- fixed here by pruning six sediment/duplication/sprawl passages fromthreat-model-and-authorization.md(584 to 539 lines), preserving every fact the dispatch flagged as load-bearing (the live-proxy caveat, theCLAUDE_PROJECT_DIRmechanism sentence, current residual bypass ceilings, the fuller Decision-7 restatement); no structural change. Mixed-portability, by contrast, is a genuine rubric blind spot: the rubric's own Mixed rule (Step 4) demands a file-level split for repository/platform-specific content, but every realistic layout for this skill's Claude-Code-specific content either adds a 4th every-use reference file (undermining the Dimension-5 exemption Round 11 just confirmed minimizes the every-use floor at 3 files) or folds it into the non-every-usefailure-and-recovery.md, destroying that file's own failure/resume-only semantic contract and reopening dimension 5 outright. Per this repository's own governance for a shared, repository-wide rubric file, filed as a new, separate issue (#1676) rather than edited ad hoc inside this PR, following the same measuredscorer-gated-skill-editsprocess issue #1662 already used. Verified: shape checker 55/55, drift scanner unchanged (same 4 pre-existing findings), full pytest (7883 passed) and all 44gitapex_gate_local_preflight.pygates passed on commit88908059. Not yet independently re-reviewed by a fresh isolated round; disclosed as such, per this session's own consistent pattern.Round 13 (commit
8e836671): issue #1676's own rubric fix merged (PR #1677); an isolated re-verification (with one correction) confirms the Mixed-portability substitute is now satisfied. Per the repository owner's own direction, mergedorigin/main(bringing in the merged rubric fix, plus an unrelatedgate-preconditions-mechanismPR #1675 this session reconciled into its own already-restructured file layout -- the new worktree-base precondition backstop content ported intodecomposition-and-dispatch.md, every stale cross-reference repointed, verified via shape checker/drift scanner/full pytest/local-preflight before and after the merge commit). A separate session had already re-graded this skill against the new rubric directly onorigin/mainas part of PR #1677, self-reporting: both gating conditions met, all three positive requirements unmet (recorded in this skill's own decision log as adeferralentry). Per this session's own consistent discipline, that self-report was independently re-verified rather than accepted at face value.An isolated Sonnet dispatch (
Read,Glob,Greponly) re-derived the finding from primary text and initially disagreed with condition 2 (whether the Claude-Code-specific content is reached on every ordinary run, given Step 6'sWorkflow-tool-vs-sequential-fallback fork) -- reading two passages (threat-model-and-authorization.md's own empirical-verification-scope note, anddecomposition-and-dispatch.md's own Consent/portability note) as evidence the primary path is opt-in and therefore not the "ordinary" case. Direct re-read of both passages, done before acting on the dispatch's own conclusion, found this was a misreading: the first passage describes this skill's own authoring-session testing limitation (could not literally exercise theWorkflowtool for live verification), not a deployment-time gate; the second describes a permission-prompt UX detail ("Default, accept edits: Every run" unless pre-approved) -- friction, not unavailability. Both conditions confirmed met on this corrected re-read, agreeing with the deferral entry's own condition-1/2 claims (though independently re-derived, not deferred to).Of the three positive requirements, requirement 2 (the
SKILL.mdNotes declaration naming non-portable steps and their portable alternative) was found already largely met on direct re-read, contra the deferral entry's own claim. Requirements 1 and 3 were genuinely unmet and fixed this round:SKILL.md's Notes portability paragraph split into distinct Non-portable (Claude-Code-specific) / Portable bold-lead-in sections (net-zero line growth, keeping the body at exactly 500/500) naming the worktree-base precondition backstop explicitly; a new, non-every-usereferences/porting-boundary-map.mdenumerates every Claude-Code-specific touchpoint this skill carries (theWorkflowtool,agentType,isolation:'worktree', thebranch-plan-tasksubagent type, bothPreToolUsehook backstops, theSubagentStopverification hook) and its portable alternative or disclosed lack thereof. One authoring defect caught and fixed along the way: an inline-code span split across a hard-wrap line broke the shape checker's per-line backtick parsing, trippingno-bare-issue-citationon an adjacent issue reference -- fixed by keeping the span on one line, a useful confirmation that this checker's per-line parsing has this exact blind spot for any future edit. The Mixed-portability substitute is now satisfied. Verified: shape checker 56/56, drift scanner unchanged (same 4 pre-existing findings), full pytest (7976 passed) and all 44gitapex_gate_local_preflight.pygates passed on commit8e836671. Not yet independently re-reviewed by a fresh isolated round; disclosed as such, per this session's own consistent pattern.Round 14 (commit
0a295d1f): three genuinely fresh, isolated review rounds -- one per skill, run against the PR branch's own actual checkout -- found and fixed real defects in two of the three standing Why-NOT-CLEAN items; the third stays open, genuinely disputed between two independent reviewers rather than adjudicated unilaterally. Prompted by the repository owner's own request to move toward resolvingindependent-review-pending. Methodology note, disclosed for completeness: the first pass of dispatches searched this session's own local checkout rather than the PR branch itself, producing two false "content does not exist" alarms forexecuting-a-branch-planandscorer-gated-skill-edits; caught before drawing any conclusion, corrected by checking out a git worktree at the PR branch's actual head and re-dispatching against it.drafting-a-skill's baseCommit-consistency question. Two independent isolated dispatches reached opposite verdicts on the same mkdir-fix decision-log entry: one traced the citedbaseCommitagainst real git parentage and found it names a commit one hop further back than the entry-adding commit's own true immediate parent (a concrete inconsistency, by a strict "must equal literal git HEAD at authoring time" reading); the other argued the citation is correct as a file-scoped pre-fix anchor (the last commit at which this skill's own files still lacked the fix), since the intervening commit belongs to an unrelated skill entirely. A third, self-run check found a third candidate value under yet a different scoping rule. Given three plausible values under three different readings of an underspecified convention, and no consensus even among independent reviewers, this stays genuinely open and disclosed -- not adjudicated unilaterally in either direction this round.scorer-gated-skill-edits's round-9 substance and citation integrity. A fresh, isolated review independently re-derived (not accepted from the round-9 entry's own claim) that Cleanup's refusal-precondition text, the rejected-edit log's file locus, and the axis-7 concurrency mechanism are all real, concretely-checkable, and correctly implemented -- confirming the substance the earlier round-9 entry only self-reported. It also found a genuine defect: all four of this skill's ownoutcome.baseCommitdecision-log citations point at real but unrelated commits in other skills' own history, not the true immediate parent of the commit that actually added each entry -- broken citations, though every entry's own substantive claim independently checks out against the real commit diffs once traced by hand. A fifth, smaller misattribution: SKILL.md Step 5 claimed thesplit.mdKept-edit-log citation was already made in step 4's worked example; it is actually in step 7. Fixed, commitf74325a2: all fourbaseCommitvalues corrected to each entry's own entry-adding commit's true immediate parent (verified viagit show/git logagainst every candidate); Step 5's cross-reference corrected. Verified: shape checker 49/49, drift scanner clean.executing-a-branch-plan's Mixed-portability substitute, requirement 1. A fresh, isolated review found Round 13's own fix overclaimed requirement 1 ("portable and non-portable content isolated under distinct headings"): the Notes section used bold inline lead-ins (**Non-portable...**,**Portable:**) inside one flowing## Notesparagraph, never an actual Markdown heading -- this SKILL.md carried no###anywhere in the whole file. The rubric's own internal use of "heading" elsewhere, and the substitute's own "never blended sentence-by-sentence" / "genuinely enumerate ... never accepted" bar, argue for the stricter, literal reading over a looser "clearly demarcated paragraphs" one. Fixed, commitc4ac264b: promoted both labels to real### Non-portable (Claude-Code-specific)/### Portablesubheadings. Net-zero body-line growth held via a tighter prose rewrap of the surrounding Notes paragraphs (verified via a whole-file word-frequency diff against the prior commit -- no wording dropped, only headings added and two lowercase-to-capitalized "step" -> "Step" changes at the new paragraph-initial positions). All four gating conditions and positive requirements for the Mixed-portability substitute are now independently confirmed met.Also this round: merged
origin/main(17 commits, including PR #1701's python-path-resolution fix -- the same fix that had been blocking this session's own earlier PR-body write attempts on this same PR, before the repository owner pointed at the merged fix) to close thebehind-baselocal-preflight gate. One real content conflict (SKILL.md's Step 8 dispatch text, whereorigin/main's own unrelated issue-#1560 commits addedagentType/subagent_typenaming and fixed a self-contradiction) and two modify/delete conflicts (origin/mainhad independently edited two reference files this branch had already deleted and merged into its own restructured 4-file layout) -- resolved by portingorigin/main's substantive text into this branch's own current file locations and repointing six now-dangling cross-references inagents/branch-plan-task.md,.claude/agents/branch-plan-task.md, andagents/review-persona.md; full account in a dedicated comment.Verified on commit
0a295d1f(after the merge and both fixes): shape checker 56/56 (executing-a-branch-plan), 49/49 (scorer-gated-skill-edits), 53/53 (drafting-a-skill, unchanged); drift scanner unchanged (same 4 pre-existing findings onexecuting-a-branch-plan, none on the other two); full pytest suite 8006/8006 passed; all 44gitapex_gate_local_preflight.pygates passed. Not yet independently re-reviewed by a fresh isolated round beyond this round's own three dispatches; disclosed as such, per this session's own consistent pattern.Why NOT-CLEAN, not CLEAN. After rounds 4-9, every skill has now received at least four from-scratch re-reviews, and every dimension gap either round surfaced is fixed or explicitly, honestly disclosed as a named residual -- no dimension is silently left unaddressed. One standing item remains open; two are now independently confirmed closed as of Round 14:
drafting-a-skill: round 6 fixed two of the three sub-findings under its own dimension-6 gap (the resume-time sweep's context-1 scope; the undeclaredgittool). The third -- the baseCommit-consistency question -- was independently re-examined twice this round by two separate dispatches reaching opposite conclusions (see Round 14 above), plus a third candidate value found on direct re-check. Genuinely disputed among independent reviewers, not a settled ambiguity this session can adjudicate unilaterally; stays disclosed, not resolved.scorer-gated-skill-edits: round 9's two minor gaps (Cleanup's undisclosed refusal preconditions; the rejected-edit log's own missing file locus) and axis 7's concurrency mechanism are now independently confirmed real and correctly implemented by Round 14's own fresh, isolated review (not merely self-reported, as the round-9 entry originally was) -- see Round 14 above. That same review found and this round fixed a citation-integrity defect (four brokenoutcome.baseCommitvalues, one Step-5 misattribution) unrelated to the substance itself. This item is now believed closed on independently-verified grounds.executing-a-branch-plan: dimension 5 (major) passed as of Round 11 (issue fix(evaluating-skill-quality): dimension 5 (Progressive disclosure) has no passing configuration for a cohesion-confirmed, over-cap sequential-pipeline skill #1662's rubric fix). Dimension 2 (sediment/duplication/sprawl) was fixed as of Round 12. Mixed-portability's requirement 1, found by Round 14's fresh review to have been only bold inline lead-ins rather than actual headings despite Round 13's own claim, is now fixed -- real###subheadings promoted, verified independently this round. All four gating conditions/positive requirements for the Mixed-portability substitute are now independently confirmed met. This item is now believed closed on independently-verified grounds.Recording
Verdict: CLEANhere would still be the self-certifying misuse this gate's own script docstring warns against, whiledrafting-a-skill's own baseCommit question stays genuinely open between independent reviewers. This check stays red until the repository owner either directs further work on that one remaining item, or accepts the current disclosed state (two of three items independently confirmed closed, one genuinely disputed) and records a different disposition.Separately, addressed this round:
skill-audit-disclosure(the## Skill audit evidencesection below) is now populated -- a realbattle-testing-a-skillPASS verdict againstdrafting-a-skill(required, not waivable, since this PR changed its frontmatter description), anevaluating-skill-qualityaggregate verdict, adversarial-coverage-mapping for both security-relevant skills, an adversarial review of both new design docs, and an honest NOT-RUN for the two checker-script disclosures (all four flagged scripts' changes this PR are comment/docstring-only, with no new detection logic to review or construct a defeat test against). See that section for the full account.eval-gate's pre-existing infrastructure defect (issue #1304) remains unaffected.Skill audit evidence
battle-testing-a-skill: PASS, dispatched against
drafting-a-skill(required as a real verdict, not waivable, since this PR changed its frontmatterdescription:line). One isolated fresh dispatch (no model-aware routing requested), cold-enumerated the 22 adversarial dimensions before reading the target. 17/17 applicable dimensions PASS with concrete quoted evidence each (structural dispatch-identity branching,StageDeviated{action: escalate}events, decode-before-trust rules, ground-truth reconciliation for its own persisted log, explicit non-authoritative-verdict framing for downstream consumers); dimension 14 (Adversarial regression corpus) gradedINDETERMINATEsince the isolated snapshot did not includeevals/drafting-a-skill/to inspect directly, per the dispatch's own instruction not to fabricate a verdict from indirect narrative; 5 dimensions (18-22) correctly N/A for a workflow-authoring skill with no legal/regulatory/financial claims.evaluating-skill-quality: WELL-FORMED-NOT-MATURE, the current aggregate state across all three touched skills -- see the
## Independent review verdictsection below for the full round-by-round record. None has yet reachedWELL-FORMED-AND-MATURE; every dimension gap found across 6-9 rounds per skill is either fixed or explicitly, honestly disclosed as a named residual, not silently left unaddressed.adversarial-coverage-mapping: RAN, for both skills the calling workflow's own security-relevance heuristic flagged:
drafting-a-skill(6 rounds of isolated, cold, adversarial re-review this PR, each grading fresh from scratch rather than only re-checking the prior finding) andscorer-gated-skill-edits(9 rounds, same discipline).design-doc-adversarial-review: RAN, one isolated dispatch adversarially critiqued both new design docs (
docs/superpowers/specs/2026-08-31-drafting-a-skill-scorer-gated-boundary-design.md,.../2026-08-31-drafting-a-skill-formative-process-gaps-design.md) against internal consistency, decision justification, unstated assumptions, disclosed trade-offs, missed edge cases, follow-through gaps, and scope match -- finding 11 defects (an uncited "44 gates" figure breaking the doc's own citation standard, a Non-goals section contradicted by the same doc's own new pre-ship gate requirement, an unaddressed failure path for a deferred one-time review, among others). All 11 are about the documents' own rigor and completeness, not about live defects in the shipped skill content they guided -- that content has been separately, extensively re-verified via the manyevaluating-skill-qualityrounds already disclosed above. Not retroactively fixed in the design docs themselves: this session treats a design doc, once it has guided real and independently-verified implementation, as historical record, the same convention already applied todocs/superpowers/plans/*.mdthroughout this PR's own decision logs.checker-script-adversarial-review: NOT-RUN and defeat-test-disclosure: NOT-RUN, for the four flagged scripts this PR touches (
skills/executing-a-branch-plan/scripts/_gitapex_path_normalize.py,gitapex_check_file_ownership_conflicts.py,gitapex_check_task_bash_safety.py,gitapex_check_task_full_verification.py). Every change to all four is comment-, docstring-, or string-literal-only -- updated cross-references to filenames renamed by the reference-file merge above, or (forgitapex_check_task_bash_safety.py, round 3, commits43c7f49b/852f1acb/3b7effb8) a narration-comment cleanup independently AST-verified to be a zero-functional-change diff, with the file's own 155-test suite and the full 259-test directory suite passing unchanged. No new or changed detection logic exists in any of the four to adversarially review or construct a defeat test against;NOT-RUNis this check's own documented, honest answer for exactly this case ("a docstring-only or lint-only edit... has no new detection logic to construct a defeat test against, so disallowing an honest NOT-RUN would only pressure toward a fabricated one" --gitapex_gate_skill_audit_disclosure.py's own docstring).Checklist
skills/*/SKILL.md, adocs/superpowers/specs/*.mddesign doc, a security-relevant skill, or a deterministic checker script (skills/*/scripts/*.py,evals/scripts/*.py,.github/scripts/*.py), a## Skill audit evidencesection discloses the required verdicts/waivers (see.github/scripts/gitapex_gate_skill_audit_disclosure.py)evals/*/split.md, that entry discloses a Transfer check line (see.github/scripts/gitapex_gate_transfer_check_disclosure.py)skills/*/SKILL.md's Stop-boundary bullets or named dispatch branches,evals//tasks/*.yamlgained at least as many new fixtures (see.github/scripts/gitapex_gate_skill_branch_fixture_coverage.py)Merge gate: independent review
This PR is also subject to the
independent-review-pendingrequiredstatus check (see
.github/workflows/independent-review-pending.yml/.github/scripts/gitapex_gate_independent_review_pending.py). It stayspending/failing until a
## Independent review verdictsection namingthis PR's current head commit is recorded in this body --
drafting-a-pr-to-merge's own Step 8 records it once its independentreview completes. There is nothing for you to do here now: do not
pre-fill this section yourself, and do not remove this note.
Related Issue
Closes #1630
Closes #1648
Refs #1662
Refs #1676
Refs #1677
Execution log
PlanApproved{run_id: e2c59a9}
TaskStarted{run_id: e2c59a9, task_id: A}
TaskStarted{run_id: e2c59a9, task_id: B}
TaskCompleted{run_id: e2c59a9, task_id: A, commit_sha: 369c406}
TaskCompleted{run_id: e2c59a9, task_id: B, commit_sha: 2e81efb}
PlanApproved{run_id: 9d93b9a0}
TaskStarted{run_id: 9d93b9a0, task_id: A}
TaskStarted{run_id: 9d93b9a0, task_id: C}
TaskCompleted{run_id: 9d93b9a0, task_id: A, commit_sha: 74b7a57}
TaskCompleted{run_id: 9d93b9a0, task_id: C, commit_sha: 8d25613}