Skip to content

CI: benchmark PRs through lightcone-bench - #240

Open
dkn16 wants to merge 4 commits into
mainfrom
ci/lightcone-bench-senders
Open

dkn16 wants to merge 4 commits into
mainfrom
ci/lightcone-bench-senders

Conversation

@dkn16

@dkn16 dkn16 commented Oct 2, 2026

Copy link
Copy Markdown
Member

Adds three small workflows that hand a PR, or a merge to main, to lightcone-bench, where this repo's own eval (evals/prompt.md + evals/tasks/snae) runs through Harbor as the task lightcone-cli/snae-build. Nothing runs in this repo; each file is one API call.

What they do

  • benchmark-on-pr.yml — every non-draft PR is benchmarked automatically against its head commit: on open, on each push, on reopen, and when a draft is marked ready. Fork PRs are skipped (they don't receive this repo's secrets).
  • benchmark-on-comment.yml — a maintainer (OWNER / MEMBER / COLLABORATOR) comments /benchmark on a PR for the same run on demand: for a draft PR, or a re-run. Reacts 🚀 as the ack.
  • benchmark-baseline.yml — every push to main dispatches the same run with save_baseline, so the bench's stored "current main" numbers track this repo.

All three pass use_skills, so the agent runs with the skills from agent-skills main installed. The skills axis is held fixed here; varying it is the agent-skills repo's own job.

What a run does

In lightcone-bench: the task's image is built with this repo pinned to the requested sha, the oracle gate runs the scripted reference solution and must score 1.0 (no LLM, so a PR that breaks the CLI contract fails here, visibly), then each agent leg runs the task and the report shows per-task rewards, the delta against the stored main baseline, and a trace analysis. Today the opencode leg runs; claude-code and codex join once their API keys are set in lightcone-bench (they skip with a notice until then). Results live in the dispatched run's job summary under lightcone-bench's Actions tab; posting them back onto the PR is a planned follow-up.

Setup (one secret)

BENCH_DISPATCH_TOKEN in this repo's Actions secrets: a fine-grained PAT (or App token) with repository access to LightconeResearch/lightcone-bench and Actions: read and write. The built-in GITHUB_TOKEN cannot dispatch into another repository.

Relation to eval.yml

eval.yml runs Claude Code on every PR inside this repo; this is the same task run by the bench, across agents, with an oracle gate and stored baselines. They don't interact. Whether the bench replaces eval.yml is a separate decision.

Testing

benchmark-on-pr.yml runs from the PR's own branch, so once the secret exists, this PR benchmarks itself on its next push. The /benchmark comment and the baseline trigger only take effect after merge: GitHub serves issue_comment and push workflows from main. The bench side was validated by dispatching it directly with this repo's current main sha, which is exactly what the senders send: oracle 1.0, opencode leg ran, report produced. The first runs will show NEW — no baseline for the skills variant until the first merge to main refreshes it.

🤖 Generated with Claude Code

https://claude.ai/code/session_01VcKmGNUek7oWEPv8zLqo9Q

dkn16 and others added 3 commits October 2, 2026 15:29
…main

Two sender workflows, copied from lightcone-bench's senders/:

- benchmark-on-comment: a maintainer comment "/benchmark" on a PR (author
  association OWNER/MEMBER/COLLABORATOR) dispatches lightcone-bench's
  benchmark workflow with stack_ref = the PR head sha, task
  lightcone-cli/snae-build, and acks with a rocket reaction. Expensive runs
  happen only on human judgment.
- benchmark-baseline: every push to main dispatches the same with
  save_baseline, so the stored "current main" numbers in lightcone-bench
  track this repo.

Both need one repository secret, BENCH_DISPATCH_TOKEN: a fine-grained PAT
(or App token) with Actions read/write on LightconeResearch/lightcone-bench.
The built-in GITHUB_TOKEN cannot dispatch into another repo.

snae-build is lightcone-cli's own eval (evals/prompt.md + evals/tasks/snae)
run through Harbor: oracle-gated, agent-agnostic, with the skills axis
measurable. Results, traces and baselines live in lightcone-bench.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VcKmGNUek7oWEPv8zLqo9Q
Signed-off-by: dkn16 <dkn16@foxmail.com>
The skills axis is held fixed (agent-skills main) on both the /benchmark
runs and the baseline refresh, so the comparison stays matched; varying
the skills is the agent-skills repo's own sender's job.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VcKmGNUek7oWEPv8zLqo9Q
Signed-off-by: dkn16 <dkn16@foxmail.com>
…drafts

benchmark-on-pr.yml dispatches lightcone-bench for every non-draft PR —
opened, pushed to, reopened, or marked ready for review — against its head
sha, with the skills from agent-skills main installed. Fork PRs are skipped
(they do not receive this repo's secrets). benchmark-on-comment.yml remains
the on-demand path: draft PRs, or a re-run.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VcKmGNUek7oWEPv8zLqo9Q
Signed-off-by: dkn16 <dkn16@foxmail.com>
@github-actions

github-actions Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

✅ Eval

Metric Value
Outputs check success
Agent run success
Turns 8
Tool calls 6
Cost $0.13
Agent wall time 0m37s
Model claude-sonnet-5-5
lc status
  mode:    direct
  sandbox: landlock (fs: declared, network: allowed)
  crate:   up to date with the outputs

  · current  baseline/best_fit        a81fe43
  · current  baseline/hubble_diagram  a81fe43
  · current  baseline/residuals       a81fe43

3 current
Confusion & pain points (Claude analysis)

Confusion & pain points

  • The run was essentially clean. There were no errored tool calls, and the agent finished in 6 tool calls. The agent also never read astra.yaml field docs or the lc source; it relied on the skill and the scaffolded spec. What friction existed was small and is listed below.

  • The astra CLI was invoked through uvx astra-tools@0.2.18 validate. The agent pinned a version and went through uvx, which suggests astra wasn't on PATH, or that the agent didn't trust whatever version was installed. The harness should provide astra on PATH at the version the spec targets. If it already does, the skill should say so, so agents don't guess a pin.

  • The agent added license = "CC-BY-4.0" to pyproject.toml with sed, unprompted. Its summary says the choice was its own. The crate is only maintained when [project].license is set, and the agent evidently knew that (likely from the task or skill). It then picked a license on the user's behalf, which the project's own rule against asserting terms over someone's data warns against. A task that expects a crate should name the license. lc init or lc materialize could also say once, in plain text, that the crate needs [project].license.

  • The agent filled the spec in with a blind str.replace Python heredoc. It never read astra.yaml back or diffed it. The edits were string-matched against the scaffold (recipe:\n command: python scripts/fit.py), so any mismatch would have silently changed nothing. Validation passed, so it worked here, but the approach is brittle. Neither astra nor lc offers a way to fill in an output's format, inputs, decisions and recipe fields programmatically. A small astra edit verb, or a documented minimal output stanza, would remove the need for this.

  • There was no pre-flight in the agent's own ordering for the compute allocation. The agent had to know that lc compute launch --wait comes first and that lc materialize local takes the cluster name as its first positional argument. It got this right, but the sequence (launch, materialize, down) is easy to miss. lc materialize without a cluster should refuse with a message naming lc compute launch, and the skill or the task prompt should show the three-step sequence.

Full trace: agent-trace artifact on this run.

Both PR workflows now run .github/scripts/bench_watch.sh after dispatching:
it finds the lightcone-bench run (by the PR sha in the run's title), posts
one comment on the PR with the link, then checks every 10 minutes and edits
that comment with the outcome when the run ends. A new push replaces the
watcher, so a PR never accumulates comments. Uses the dispatch token to read
runs and the job's GITHUB_TOKEN (pull-requests: write) to comment.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VcKmGNUek7oWEPv8zLqo9Q
Signed-off-by: dkn16 <dkn16@foxmail.com>
@github-actions

github-actions Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

lightcone-bench — benchmark of 36a0be2 ✅ passed: run 37077360445 — the matrix, baseline deltas and trace report are in its job summary.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant