Skip to content

feat: add manual human evaluation skill - #24

Merged
kentwelcome merged 1 commit into
mainfrom
feat/manual-human-evaluation
Oct 5, 2026
Merged

kentwelcome merged 1 commit into
mainfrom
feat/manual-human-evaluation

Conversation

@kentwelcome

Copy link
Copy Markdown
Contributor

Summary

  • Add a repository-local manual skill sampling five random eligible skill changes from DataRecce/recce-team against current local Behavior Diff code.
  • Freeze synthetic fixtures and questions before consent-gated trials; preserve provenance, failures, and first submissions.
  • Provide a private loopback quiz with unchanged generated summary excerpts, confidence/sufficiency responses, scoring, and post-submission evidence.
  • Add maintainer documentation and deterministic-only CI checks. No production summary wording, plugin payload, or version changes.

Verification

  • 15 workflow tests and 12 quiz tests passed.
  • Existing hooks, decisions self-check, live-report contract, and release-workflow checks passed.
  • Ruff, Docker shfmt, and git diff whitespace checks passed.
  • Actual fixed-source init and guarded Claude --version smoke passed without inference.
  • Synthetic freeze/build/serve/results smoke and desktop/mobile browser verification passed, including blinding, draft persistence, scoring, and full-report reveal.
  • No new paid trials; private source and evaluation artifacts remain outside the repository.

Review

Independent Standards and Spec reviews approved the implementation before publication. Fresh reviews of this committed PR are underway.

Refs DRC-4795.

Add a private, fixed-source random evaluation workflow and blinded quiz for current local Behavior Diff summaries.

Refs DRC-4795

Signed-off-by: Kent Huang <kent@infuseai.io>
@kentwelcome

Copy link
Copy Markdown
Contributor Author

Independent subagent reviews

Reviewed committed diff 1f6a0c5...ec001ff in two parallel, read-only subagents under REVIEWER_GUIDELINES.md. No private evaluation artifacts or live model calls were used.

Standards — APPROVE

Material findings: none. No evidence-backed violation of AGENTS.md, CODING_GUIDELINES.md, or the evaluation protocol was found across the changed artifacts and relevant consumers. Decisions required: none. Optional suggestions: none.

Spec — APPROVE

Material findings: none. Traced the skill, lifecycle helpers, quiz, documentation, tests, and trial/extraction consumers against DRC-4795. No missing, incorrect, or unauthorized supported behavior was demonstrated. Decisions required: none. Optional suggestions: none.

Scope audit

Added the authorized repository-local manual evaluation skill, private sampling/freeze/provenance lifecycle, read-only Claude launcher, blinded quiz and scoring, documentation, and deterministic checks. No mechanisms removed. Production summary generation and plugin payload remain unchanged.

Verification

GitHub Actions Format and Unit both passed. Local verification also passed the four existing required checks, 27 new tests, formatting checks, and synthetic CLI plus desktop/mobile browser smoke. No new paid trials.

Summary: 0 Standards findings; 0 Spec findings; no blocking issue.

@kentwelcome
kentwelcome merged commit ad143e5 into main Oct 5, 2026
2 checks passed
@kentwelcome
kentwelcome deleted the feat/manual-human-evaluation branch October 5, 2026 09:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant