fix(evaluating-skill-quality): add Mixed-portability substitute for D5-exempted every-use content - #1677
Conversation
Task-list plan file for the Mixed-portability closure work, per executing-a-branch-plan's own Step 3/4 convention: first commit on the branch, published before any task work begins. Refs #1676. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KKETd5mWU8Vg78Phj9Wfxs
…imension-5-exempted every-use content rubric.md's Mixed-portability rule required a physical file-level split of a target's non-portable content, but Dimension 5's own cohesion-confirmed sequential-pipeline exemption (issue #1662) gave a qualifying target no way to satisfy that requirement without either exceeding its own minimized every-use file-count floor or corrupting a non-every-use reference file's own "never read on an ordinary clean run" contract. Adds one narrow, loophole-resistant substitute nested under the existing Mixed bullet, gated on two independently-verifiable conditions (the target already clears Dimension 5's own exemption; its non-portable content is demonstrably read on every ordinary run, never accepted from the target's own self-characterization), plus a companion cross-reference in Dimension 5's own "still apply in full" parenthetical. The Mixed bullet's own existing text is unchanged. Adds two new selection-split eval fixtures: a genuinely qualifying case, and an anti-loophole false-positive-attempt whose Notes section self-characterizes non-portable content as every-use while the Procedure text itself shows it is conditional -- exercising the new substitute's own anti-self-assertion requirement. Also records executing-a-branch-plan's own re-graded Mixed-portability status against the new rubric (verification only, no code change to that skill): both gating conditions are met, but the three positive requirements are not yet satisfied (no distinct-heading isolation, an incomplete Notes declaration, no dedicated porting-boundary-map file) -- authoring that file is out of this issue's own scope, per its Non-goals. Refs #1676. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KKETd5mWU8Vg78Phj9Wfxs
…results Formal gate result for the two new Mixed-portability-substitute fixtures: primary selection fixture 0.800000 -> 1.000000, KEEP; the anti-loophole false-positive fixture correctly fails both before and after (1.000000 both sides). A Transfer check against the adjacent sequential-pipeline-body-cap-exception fixture confirmed no regression and that the new substitute is inert for Portable-declared targets. Fixes two live authoring corrections found during the same run, disclosed in the run record's own known_gaps: the primary fixture's first draft had a conditional (not every-use) rollback step, and its first-chosen assertion did not survive the reviewing model's own paraphrasing. Both fixed before this record was written. The confirmed eval runner (evals/scripts/gitapex_run_eval_suite.py) could not execute live in this session's environment -- the same empty-content Authentication error already tracked in issue #1304, independently reconfirmed this session. Scoring used isolated Agent-tool dispatches instead, each reading the combined skill content via its own Read tool call; fully disclosed in the run record's own known_gaps. Refs #1676, #1662. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KKETd5mWU8Vg78Phj9Wfxs
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Team Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Generated by Claude Code |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #1677 +/- ##
=======================================
Coverage 99.66% 99.66%
=======================================
Files 149 149
Lines 24479 24479
Branches 2960 2960
=======================================
Hits 24397 24397
Misses 82 82 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
…per Step 8 review An independent adversarial review (executing-a-branch-plan's own mandatory Step 8) found a real defect in the original placement: the substitute lived under the Portability level section, graded at SKILL.md Procedure step 4, but its own condition 1 required a dimension-5 finding that step 4 has no access to yet (dimension 5 runs at step 5). Relocates the full substitute into dimension 5's own section, immediately after the pre-existing sequential-pipeline exemption it depends on -- both conditions are now established sequentially within the same step-5 walk, with no backward reference. The Portability level section keeps only a short forward pointer. Also fixes two more review findings: condition 2 now requires the non-portable content be demonstrably reached and acted on, not merely read as inert text (closing a loophole where a conditionally-executed step could still claim to satisfy it); and the third positive requirement (the dedicated reference file) now explicitly requires confirming the file actually exists and genuinely enumerates every touchpoint, never accepted on the target's own claim -- mirroring condition 2's own anti-self-assertion discipline. The primary eval fixture now supplies that file's own real content so this can actually be verified rather than merely asserted. Also fixes tests/test_gitapex_scan_eval_results_schema.py's two pinned real-repository-corpus sets, which did not yet know about this PR's own new run directory (mirroring issue #1662's own precedent for the same gap). Gate re-verification against the restructured rubric is in progress; this commit's own run record will be superseded once it completes. Refs #1676. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KKETd5mWU8Vg78Phj9Wfxs
…e-gate split-disclosure requires the touched fixture named in this commit range's own added lines. Adds a correction note now; the fixture's own Kept-edit-log entry gets a fuller rewrite once the fresh gate run (dispatched, in progress) returns. Refs #1676. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KKETd5mWU8Vg78Phj9Wfxs
Reconciles the run record with the Step 8 relocation: re-scored transcripts (0.666667 -> 1.000000, replacing the stale 0.8 baseline), final commit references (eced6ac, the relocated rubric text), and a consistent split.md narrative covering both the relocation and the two assertion-wording corrections found while re-scoring against real transcripts. Also updates evaluating-skill-quality's own decision-log entry to the same numbers, and adds a disclosure to executing-a-branch-plan's deferral entry noting its re-grade predates the relocation and was not independently re-confirmed against the final wording. Refs #1676. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KKETd5mWU8Vg78Phj9Wfxs
…ndings An isolated evaluating-skill-quality specialist dispatch (drafting-a-pr- to-merge's own Step 8 inner layer) found three wording defects in the new Mixed-portability substitute: two different terms for the same "portable alternative" concept, one colliding with the substitute's own proper name; a Fail-bullet clause covering only 2 of 5 failure modes, unlike its sibling exemption bullet's own symmetric coverage; and "repository/platform-specific" colliding with this file's own separate, established use of "platform" for the mechanism-fit axis. All three fixed; none change the gated fixtures' own scored assertions, so the recorded gate numbers were not re-run -- disclosed in the run record's known_gaps instead, per issue #1662's own precedent. Refs #1676. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KKETd5mWU8Vg78Phj9Wfxs
…ll-implementation-lkdypz Brings in two independently-merged main-branch efforts that landed while this branch was in flight: - PR #1677 (issue #1676): rubric.md's Mixed-portability substitute for a Dimension-5-exempted target, plus that same session's own re-grade of executing-a-branch-plan's Mixed-portability finding against it (kept as its own deferral decision-log entry, appended after mine rather than overwritten). - PR #1675 (issues #1508/#1566): a new worktree-base precondition backstop (gitapex_check_task_worktree_base.py), chained into check_task_bash_safety.sh. Conflict resolution: - metadata/gitapex.yaml: kept this branch's full decision log and Frontier/requires/relatedTo values (main's own copy predates this branch's own fork and never carried any of that), appended main's new #1676 deferral entry, and added one missing decision-log entry for this branch's own prior failure-and-recovery.md reformatting commit. - references/execution-and-dispatch.md: this branch already retired this file (split into decomposition-and-dispatch.md / events-and-review-gate.md / failure-and-recovery.md, per issue #1662's own Round 10). Ported main's new "Worktree-base precondition backstop" subsection into decomposition-and-dispatch.md's own "Execution and dispatch" section (matching where its sibling subsections already live), demoted from ## to ###, and repointed every stale execution-and-dispatch.md cross-reference (in threat-model-and-authorization.md and gitapex_check_task_worktree_base.py/its test) to decomposition-and-dispatch.md. Verified post-merge: shape checker 54/54, drift scanner unchanged (same 4 pre-existing findings), full pytest suite passing (the one run mid- merge showed 7 gitapex_gate_commit_citation.py failures -- confirmed a transient artifact of running the suite while .git/MERGE_HEAD existed, merge_in_progress()'s own by-design merge-commit exemption, not a real defect; re-verified clean after this commit), all 44 local-preflight gates passed. Refs #1648, #1662, #1676
Duplicate-PR-waiver: PR #1632's own body references issue #1676 (
Refs #1676) because its Round 12 review is what found and filed this issue -- PR #1632's own text explicitly states the rubric-level fix is "filed as a new, separate issue (#1676) rather than edited ad hoc inside this PR." This PR is that separate, narrower fix:references/rubric.md, new eval fixtures, and decision-log entries only. One file,skills/executing-a-branch-plan/metadata/gitapex.yaml, is also touched by PR #1632 (its own Round 10-12 decision-log entries) -- a real merge-conflict risk to resolve at merge time, not duplicate work: this PR's own new entry (issue #1676's own re-grade finding) is additive, appended after PR #1632's own entries, never overlapping or reverting them.Summary
Adds a narrow, loophole-resistant Mixed-portability substitute for a Dimension-5-exempted target whose non-portable content is itself every-use, closing the unsatisfiability gap issue #1632's own Round 12 review found in
executing-a-branch-plan's own Mixed-portability grading.Facts
evaluating-skill-quality/references/rubric.md's Mixed bullet (Step 4's portability section) required a physical file-level split of non-portable content, but issue fix(evaluating-skill-quality): dimension 5 (Progressive disclosure) has no passing configuration for a cohesion-confirmed, over-cap sequential-pipeline skill #1662's Dimension-5 cohesion-confirmed sequential-pipeline exemption never mentioned that rule at all, even though the Mixed bullet frames itself as its own dimension-5 requirement.executing-a-branch-planis exactly this shape) had no satisfiable, textually-compliant way to close a Mixed-portability finding: relocating the content into a new every-use file pushes the exemption's own minimized file count past its floor; folding it into a non-every-use file destroys that file's own "never read on an ordinary clean run" contract and reopens dimension 5 outright.scorer-gated-skill-edits's own measured, held-out-gated edit process, per issue fix(evaluating-skill-quality): Mixed-portability file-level split conflicts with dimension 5's own sequential-pipeline exemption #1676's own Constraints and issue fix(evaluating-skill-quality): dimension 5 (Progressive disclosure) has no passing configuration for a cohesion-confirmed, over-cap sequential-pipeline skill #1662's own precedent: 2 new selection fixtures,split.jsonregistration (partition 35:44:18), a formal gate run (primary fixture 0.666667 -> 1.000000, KEEP), an anti-loophole discrimination fixture (1.000000/1.000000, no regression), and a Transfer check against the adjacent pre-existing sequential-pipeline fixture (1.000000, no regression, and the new substitute confirmed inert for Portable-declared targets).gitapex_check_skill_shape.py: 70/70 (evaluating-skill-quality), 61/61 (executing-a-branch-plan). Full local-preflight (44/44 gates) passed before push.known_gaps: a scenario-design flaw in the primary fixture's first draft (a conditional, not-every-use, rollback step), and two rounds of fixture-authoring assertion-wording gaps (the final assertion,"Condition 2 holds", was empirically confirmed viagrep -cto discriminate perfectly across all four saved transcripts before being adopted).evals/scripts/gitapex_run_eval_suite.py) could not execute live in this session's environment -- the sameAuthentication erroralready tracked in issue fix(waza-eval-gate): live eval runs fail with "model CLI exited 1" and empty stderr, repo-wide #1304, independently reconfirmed this session. Scoring used isolated Agent-tool subagent dispatches instead, fully disclosed throughout the run record.executing-a-branch-plan's own Mixed-portability status was re-graded against the new rubric: both gating conditions are met, but its three positive requirements are not yet satisfied (no distinct-heading isolation, an incomplete Notes declaration, no dedicated porting-boundary-map file) -- recorded in that skill's ownmetadata/gitapex.yamldecision log, with a disclosure that this specific re-grade predates the Step 8 relocation above and was not independently re-confirmed against the final wording (the verdict's substance is believed unchanged).Assumptions
executing-a-branch-plan-- authoring the missing porting-boundary-map reference file is explicitly out of this issue's own scope (see issue fix(evaluating-skill-quality): Mixed-portability file-level split conflicts with dimension 5's own sequential-pipeline exemption #1676's Non-goals) and is separate follow-up work, confirmed with the repository owner during this session.Risk / blast radius
Documentation/skill-definition and eval-fixture edits only -- no runtime script behavior changes. Blast radius is future
evaluating-skill-qualityreviews of any Mixed-declared, Dimension-5-exempted target (onlyexecuting-a-branch-plancurrently qualifies). The Mixed bullet's own existing five lines are unchanged (now carrying only a forward-pointer sentence instead of a nested exemption -- confirmed by direct diff review, ACM row 2); the substitute itself now lives entirely inside dimension 5's own section.Rollback
Revert this PR's merge commit; no other repository state depends on this rubric/fixture addition. The
executing-a-branch-plandecision-log entry is a disclosure record, not a functional dependency.Verification
Restating issue #1676's Acceptance Criteria Map, row by row, each followed by its own Result:
skills/evaluating-skill-quality/references/rubric.md's Dimension 5 section and the Portability level section's own Mixed bullet onlyscorer-gated-skill-edits's own measured gate;gitapex_check_skill_shape.pyfull runrubric.mdexecuting-a-branch-plan's own current Mixed-portability disclosure is re-graded once the rubric change landsevaluating-skill-quality's Procedure (isolated dispatch) against the currentexecuting-a-branch-plancontentmetadata/gitapex.yamldecision logResults:
deferralentry, out of this issue's own scope to fix; that entry also discloses it was not independently re-confirmed against the Step 8 relocation's final wording.Checklist
skills/*/SKILL.md, adocs/superpowers/specs/*.mddesign doc, a security-relevant skill, or a deterministic checker script (skills/*/scripts/*.py,evals/scripts/*.py,.github/scripts/*.py), a## Skill audit evidencesection discloses the required verdicts/waivers (see.github/scripts/gitapex_gate_skill_audit_disclosure.py) -- not applicable: this PR touches onlyreferences/rubric.md,metadata/gitapex.yaml, andevals/**fixtures/results, none of which that gate's own trigger scope covers.evals/*/split.md, that entry discloses a Transfer check line (see.github/scripts/gitapex_gate_transfer_check_disclosure.py)skills/*/SKILL.md's Stop-boundary bullets or named dispatch branches,evals//tasks/*.yamlgained at least as many new fixtures (see.github/scripts/gitapex_gate_skill_branch_fixture_coverage.py) -- not applicable: no SKILL.md Stop-boundary/branch change; the 2 new fixtures cover the new rubric.md bullet directly.Independent review verdict
Outer layer (GitHub-native reviewer). No confirmed installation of Anthropic's "Claude Code Review" GitHub App on this repository; GitHub Copilot's
copilot-pull-request-reviewer[bot]was requested viagithub:request_copilot_reviewat 2026-09-02T16:38:50Z. A freshpull_request_read(get_reviews) call after the 30-minute window (2026-09-02T17:12Z+) returned zero reviews -- this layer is treated as unreachable for this round; its absence is disclosed here, not silently treated as a clean pass.Inner layer (
reviewing-an-artifact,loweffort). Ran against this PR's full accumulated diff (baseb07fabc5, headd2ae37a2):skills/evaluating-skill-quality/references/rubric.md(areferences/file of aSKILL.md-owning skill) was specialist-deferred to an isolatedevaluating-skill-qualitydispatch, per Step 0's own eight-way deferral list -- the same treatment issue fix(evaluating-skill-quality): dimension 5 (Progressive disclosure) has no passing configuration for a cohesion-confirmed, over-cap sequential-pipeline skill #1662's own precedent (PR fix(evaluating-skill-quality): add Dimension-5 exception for sequential-pipeline skills #1672) gave the same file. The remainder (2 new eval fixtures,split.json/split.md, the new run-record directory and its artifacts,docs/skill-eval-status.mdregeneration, the new branch-plan doc, 2metadata/gitapex.yamldecision-log entries, and 2 test-file pinned-value updates) classified safe under Step 1 (log/fixture/doc additions and pinned-value updates only, no security-tier signal anywhere -- independently verified by direct diff inspection of every file in this remainder before accepting the classification) -- Steps 2-5 skipped for that remainder, recorded as the skip it is.evaluating-skill-qualitydispatch (isolated, no memory of authoring this change) found three confirmed findings, all independently re-verified in this thread before being treated as real (not accepted at the dispatch's own say-so): (1) the new substitute's own positive requirements b and c used two different terms ("portable fallback", "portable substitute") for one concept, with "substitute" also colliding with the exemption's own proper name -- verified by reading both cited lines directly, and against dimension 4's own explicit "one term per concept" rule; (2) the substitute's own Fail-bullet clause illustrated only 2 of its 5 failure modes (condition 2, the third positive requirement), asymmetric with the sibling sequential-pipeline exemption's own Fail clause naming both of its conditions -- verified by reading both Fail clauses side by side; (3) "repository/platform-specific" introduced a term colliding with this same combined skill file's own separate, established use of "platform" for the Agentic-operation-mechanism-fit/isolation axis -- verified by grepping the combined skill file for prior "platform" usage. All three fixed in commitd2ae37a2(unifying to "portable alternative"; extending the Fail bullet to name all five failure modes; reverting to "repository-specific" alone);gitapex_check_skill_shape.pyre-run 70/70 and the full localpytestsuite (7883 tests) re-run clean after the fix, before this verdict was recorded. None of the three changed the gated fixtures' own scored assertions or gating logic, so the recorded gate numbers were not re-run -- disclosed in the run record's ownknown_gaps.Zero
confirmedfindings remain outstanding against the current head, and zerounconfirmed-concernfindings were raised. Per the four-outcomes rule: outer layer unreachable within its 30-minute window (disclosed above), inner layer zero confirmed findings remaining on the current head -> continue to step 9.Merge gate: independent review
This PR is also subject to the
independent-review-pendingrequiredstatus check (see
.github/workflows/independent-review-pending.yml/.github/scripts/gitapex_gate_independent_review_pending.py). It stayspending/failing until a
## Independent review verdictsection namingthis PR's current head commit is recorded in this body --
drafting-a-pr-to-merge's own Step 8 records it once its independentreview completes. There is nothing for you to do here now: do not
pre-fill this section yourself, and do not remove this note.
Related Issue
Closes #1676
Refs #1662, #1632, #1648
Execution log
PlanApproved{run_id: f4c1a9b2}
TaskStarted{run_id: f4c1a9b2, task_id: A}
TaskCompleted{run_id: f4c1a9b2, task_id: A, commit_sha: 98f6c31}
TaskStarted{run_id: f4c1a9b2, task_id: gate}
TaskCompleted{run_id: f4c1a9b2, task_id: gate, commit_sha: 3a572ab}
TaskStarted{run_id: f4c1a9b2, task_id: B}
TaskCompleted{run_id: f4c1a9b2, task_id: B, commit_sha: 3a572ab}
TaskStarted{run_id: f4c1a9b2, task_id: step8-relocate}
TaskCompleted{run_id: f4c1a9b2, task_id: step8-relocate, commit_sha: eced6ac}
TaskStarted{run_id: f4c1a9b2, task_id: step8-regate}
TaskCompleted{run_id: f4c1a9b2, task_id: step8-regate, commit_sha: 18609ca}
TaskStarted{run_id: f4c1a9b2, task_id: step8-review}
TaskCompleted{run_id: f4c1a9b2, task_id: step8-review, commit_sha: d2ae37a}
Generated by Claude Code