Repository navigation
feat: add manual human evaluation skill - #24
Conversation
Add a private, fixed-source random evaluation workflow and blinded quiz for current local Behavior Diff summaries. Refs DRC-4795 Signed-off-by: Kent Huang <kent@infuseai.io>
Independent subagent reviewsReviewed committed diff Standards — APPROVEMaterial findings: none. No evidence-backed violation of AGENTS.md, CODING_GUIDELINES.md, or the evaluation protocol was found across the changed artifacts and relevant consumers. Decisions required: none. Optional suggestions: none. Spec — APPROVEMaterial findings: none. Traced the skill, lifecycle helpers, quiz, documentation, tests, and trial/extraction consumers against DRC-4795. No missing, incorrect, or unauthorized supported behavior was demonstrated. Decisions required: none. Optional suggestions: none. Scope auditAdded the authorized repository-local manual evaluation skill, private sampling/freeze/provenance lifecycle, read-only Claude launcher, blinded quiz and scoring, documentation, and deterministic checks. No mechanisms removed. Production summary generation and plugin payload remain unchanged. VerificationGitHub Actions Format and Unit both passed. Local verification also passed the four existing required checks, 27 new tests, formatting checks, and synthetic CLI plus desktop/mobile browser smoke. No new paid trials. Summary: 0 Standards findings; 0 Spec findings; no blocking issue. |
Summary
Verification
Review
Independent Standards and Spec reviews approved the implementation before publication. Fresh reviews of this committed PR are underway.
Refs DRC-4795.