Conversation
…main Two sender workflows, copied from lightcone-bench's senders/: - benchmark-on-comment: a maintainer comment "/benchmark" on a PR (author association OWNER/MEMBER/COLLABORATOR) dispatches lightcone-bench's benchmark workflow with stack_ref = the PR head sha, task lightcone-cli/snae-build, and acks with a rocket reaction. Expensive runs happen only on human judgment. - benchmark-baseline: every push to main dispatches the same with save_baseline, so the stored "current main" numbers in lightcone-bench track this repo. Both need one repository secret, BENCH_DISPATCH_TOKEN: a fine-grained PAT (or App token) with Actions read/write on LightconeResearch/lightcone-bench. The built-in GITHUB_TOKEN cannot dispatch into another repo. snae-build is lightcone-cli's own eval (evals/prompt.md + evals/tasks/snae) run through Harbor: oracle-gated, agent-agnostic, with the skills axis measurable. Results, traces and baselines live in lightcone-bench. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VcKmGNUek7oWEPv8zLqo9Q Signed-off-by: dkn16 <dkn16@foxmail.com>
The skills axis is held fixed (agent-skills main) on both the /benchmark runs and the baseline refresh, so the comparison stays matched; varying the skills is the agent-skills repo's own sender's job. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VcKmGNUek7oWEPv8zLqo9Q Signed-off-by: dkn16 <dkn16@foxmail.com>
…drafts benchmark-on-pr.yml dispatches lightcone-bench for every non-draft PR — opened, pushed to, reopened, or marked ready for review — against its head sha, with the skills from agent-skills main installed. Fork PRs are skipped (they do not receive this repo's secrets). benchmark-on-comment.yml remains the on-demand path: draft PRs, or a re-run. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VcKmGNUek7oWEPv8zLqo9Q Signed-off-by: dkn16 <dkn16@foxmail.com>
✅ Eval
lc statusConfusion & pain points (Claude analysis)Confusion & pain points
Full trace: |
Both PR workflows now run .github/scripts/bench_watch.sh after dispatching: it finds the lightcone-bench run (by the PR sha in the run's title), posts one comment on the PR with the link, then checks every 10 minutes and edits that comment with the outcome when the run ends. A new push replaces the watcher, so a PR never accumulates comments. Uses the dispatch token to read runs and the job's GITHUB_TOKEN (pull-requests: write) to comment. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VcKmGNUek7oWEPv8zLqo9Q Signed-off-by: dkn16 <dkn16@foxmail.com>
|
lightcone-bench — benchmark of |
Adds three small workflows that hand a PR, or a merge to
main, to lightcone-bench, where this repo's own eval (evals/prompt.md+evals/tasks/snae) runs through Harbor as the tasklightcone-cli/snae-build. Nothing runs in this repo; each file is one API call.What they do
benchmark-on-pr.yml— every non-draft PR is benchmarked automatically against its head commit: on open, on each push, on reopen, and when a draft is marked ready. Fork PRs are skipped (they don't receive this repo's secrets).benchmark-on-comment.yml— a maintainer (OWNER / MEMBER / COLLABORATOR) comments/benchmarkon a PR for the same run on demand: for a draft PR, or a re-run. Reacts 🚀 as the ack.benchmark-baseline.yml— every push tomaindispatches the same run withsave_baseline, so the bench's stored "current main" numbers track this repo.All three pass
use_skills, so the agent runs with the skills from agent-skillsmaininstalled. The skills axis is held fixed here; varying it is the agent-skills repo's own job.What a run does
In lightcone-bench: the task's image is built with this repo pinned to the requested sha, the oracle gate runs the scripted reference solution and must score 1.0 (no LLM, so a PR that breaks the CLI contract fails here, visibly), then each agent leg runs the task and the report shows per-task rewards, the delta against the stored main baseline, and a trace analysis. Today the opencode leg runs; claude-code and codex join once their API keys are set in lightcone-bench (they skip with a notice until then). Results live in the dispatched run's job summary under lightcone-bench's Actions tab; posting them back onto the PR is a planned follow-up.
Setup (one secret)
BENCH_DISPATCH_TOKENin this repo's Actions secrets: a fine-grained PAT (or App token) with repository access toLightconeResearch/lightcone-benchand Actions: read and write. The built-inGITHUB_TOKENcannot dispatch into another repository.Relation to
eval.ymleval.ymlruns Claude Code on every PR inside this repo; this is the same task run by the bench, across agents, with an oracle gate and stored baselines. They don't interact. Whether the bench replaceseval.ymlis a separate decision.Testing
benchmark-on-pr.ymlruns from the PR's own branch, so once the secret exists, this PR benchmarks itself on its next push. The/benchmarkcomment and the baseline trigger only take effect after merge: GitHub servesissue_commentandpushworkflows frommain. The bench side was validated by dispatching it directly with this repo's current main sha, which is exactly what the senders send: oracle 1.0, opencode leg ran, report produced. The first runs will showNEW — no baselinefor the skills variant until the first merge tomainrefreshes it.🤖 Generated with Claude Code
https://claude.ai/code/session_01VcKmGNUek7oWEPv8zLqo9Q