Skip to content

feat(agent-cli): add Prime Agent and package agent runtimes - #54

Merged
abhinav-pola merged 5 commits into
mainfrom
devin/1787681826-prime-agent-benchmark-harness
Aug 26, 2026
Merged

feat(agent-cli): add Prime Agent and package agent runtimes#54
abhinav-pola merged 5 commits into
mainfrom
devin/1787681826-prime-agent-benchmark-harness

Conversation

@abhinav-pola

@abhinav-pola abhinav-pola commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

TL;DR

Adds Prime Agent as an Ori-backed harness and makes the default Claude Code, Pi, and Prime Agent image paths consume one pinned, checksum-verified runtime artifact without cold npm or NVM dependency resolution.

What changed?

  • Added prime-agent to the shared Ori configuration for Terminal Bench, DeepSWE, and SWE-Atlas, including its launch contract and structured telemetry parser.
  • Pinned Claude Code 2.1.235, Pi Coding Agent 0.84.2, Prime Agent 0.8.0, Node.js 22.23.2, upstream artifacts, and the complete npm dependency graph.
  • Added a deterministic combined runtime builder and published agent-runtime-v1 with SHA256SUMS.
  • Changed all three default image paths to download, verify, and extract the combined runtime, expose stable executable paths, and run the selected harness version check.
  • Preserved npm and NVM installation as an explicit fallback for custom package overrides, including Prime Agent's required bootstrap environment.
  • Added colocated coverage for default and override image construction, launch contracts, parser boundaries, generation IDs, usage, cost, assistant messages, turns, tool calls, and failures.

Why?

ECO-2361 needs comparable model harnesses with traceable OpenRouter telemetry. Packaging every currently supported harness removes npm availability and transitive dependency drift from default cold benchmark image construction.

How to test

bun run package-agent-runtime /tmp/agent-runtime
sha256sum /tmp/agent-runtime/agent-runtime-linux-x64-v1.tar.zst
bun run format:check
bun run check
bun run typecheck
bun test
bun run build

The archive checksum should be 3b4ab1055d41257071860235867f31b5934e7e9c08fa57644d2888d435293bf3. Generated default images should report Claude Code 2.1.235, Pi 0.84.2, Prime Agent 0.8.0, and Node.js v22.23.2, with no npm install or nvm install in their default image steps.

Proof
  • Two clean packaging runs produced the same archive SHA-256, which also matches the published agent-runtime-v1 SHA256SUMS asset.
  • Fresh generated Ubuntu 20.04 images placed Claude, Pi, Prime Agent, Node, npm, npx, uv, and Python on their expected runtime paths and passed version checks.
  • The real generated-image runAgentCli path completed file-writing smoke runs for Claude and Pi, returned successful exits and OpenRouter generation IDs, and produced the requested files. Prime Agent's equivalent generated-image launch was verified earlier in this PR.
  • Final validation: 1,378 tests passed, with formatting, lint/check, Effect diagnostics, typecheck, and build passing. Existing repository lint warnings and Effect informational messages remain unchanged.
  • Modal verifier scoring was not run. The ECO-2361 matrix has not started.

Benchmark impact

New runs may select agent: "prime-agent". Claude and Pi defaults now use pinned versions instead of @latest, so new results are reproducible but may differ from historical runs that resolved another package version. Datasets, prompts, scorers, routing, and existing launch contracts are unchanged.

Reviewer focus

  • Integrity and reproducibility of the combined runtime artifact.
  • No npm or NVM dependency resolution in default image construction.
  • Custom override compatibility and Prime Agent bootstrap requirements.
  • JSONL boundaries that exclude local session and intermediate update identifiers from OpenRouter generation IDs.

Checklist

  • Tests cover changed behavior
  • Public API or configuration changes are backward compatible, or the break is documented
  • Benchmark changes do not modify dataset provenance or licensing
  • No credentials, private results, or restricted dataset contents are included
  • Packaging and configuration surfaces are self-describing

Link to Devin session: https://openrouter.devinenterprise.com/sessions/ea6550fd84034992b6e49bb3c72ec2a6
Requested by: @abhinav-pola


Open in Devin Review

Co-Authored-By: Abhinav Pola <abhinav.pola@openrouter.ai>
@devin-ai-integration

Copy link
Copy Markdown
Contributor
Original prompt from Abhinav

SYSTEM:
<latest_message>
Abhinav Pola (U090K0G7JF3) [ts=1787674351.190119]: @Devin!benchmark_pipeline i want to continue making progress on ECO-2361. whats the current status?
</latest_message>

=== BEGIN THREAD HISTORY (in #intern-benchy) ===
Abhinav Pola (U090K0G7JF3) [ts=1787674351.190119]: @Devin!benchmark_pipeline i want to continue making progress on ECO-2361. whats the current status?
=== END THREAD HISTORY ===
Channel ID: C0BAKP8P5C3
Thread URL: https://openrouter.slack.com/archives/C0BAKP8P5C3/p1787674351190119?thread_ts=1787674351.190119&amp;cid=C0BAKP8P5C3

The <latest_message> is the message that you should use to guide your goals + task for this session, and you should use the rest of the slack thread as context.
A [ts=...] marker on a Slack message is that message's timestamp. To act on a specific message with the slack tool (e.g. adding an emoji reaction via the reaction command), pass that value as timestamp along with the Channel ID — no extra lookup call is needed.

@playbook:playbook-34832ba4a2a84d78a6432390e3b4d963

@devin-ai-integration

Copy link
Copy Markdown
Contributor

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR that start with 'DevinAI' or '@devin'.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Devin Review: No Issues Found

Devin Review analyzed this PR and found no potential bugs to report.

View in Devin Review to see 1 additional finding.

Open in Devin Review

Comment thread src/benchmarks/agent-cli/harness.ts Outdated
@devin-ai-integration

Copy link
Copy Markdown
Contributor

Live runtime verification passed for the Prime Agent harness at b38bc95.

Runtime and attribution proof

One paid OpenRouter sample cost $0.002224. It used this PR's getOriHarness("prime-agent"), agentImageBuildSteps, and runAgentCli without source changes.

  • The pinned tarball passed SHA-256 verification, prime-agent --version returned 0.8.0, and the harness installed Ori 0.10.2 inside the container.
  • The generated script launched ori prime-agent in JSON mode with explicit --reasoning-effort low, no session persistence, no endpoint pin, and no max-token override.
  • The run exited 0, recorded 2 turns and 1 tool call, and wrote the requested OK file.
  • The parser collected exactly the assistant message_end.responseId values gen-1787684077-IhnmGKl8Q8UFTGFyr5y3 and gen-1787684082-oPMfbg3MWvx6a9TOiQF3. Both resolve through GET /api/v1/generation with HTTP 200 and matching IDs.
  • Session UUIDs, tool-call IDs, user messages, empty IDs, and synthetic message_update IDs were excluded. Synthetic message_update cost was also excluded.
  • Parsed cost $0.002224 matched the generation-record sum and API-key usage delta within $1e-9. Parsed token totals also matched the generation records at 8,408 input, 56 output, and 8,464 total.
Boundary

Modal sandbox orchestration and Terminal Bench verifier/reward scoring remain untested because Modal credentials were unavailable and the requester chose to skip them. Local Docker exercised the generated image, Ori bootstrap, Prime Agent run, parser, OpenRouter lookup, and billing reconciliation.

The --reasoning-effort low flag is proven in the launched command. Its provider-side effect is not observable with openai/gpt-4.1-mini, which returned zero reasoning tokens.

Written by Devin

devin-ai-integration Bot and others added 3 commits August 25, 2026 20:21
Co-Authored-By: Abhinav Pola <abhinav.pola@openrouter.ai>
Co-Authored-By: Abhinav Pola <abhinav.pola@openrouter.ai>
Co-Authored-By: Abhinav Pola <abhinav.pola@openrouter.ai>
@devin-ai-integration devin-ai-integration Bot changed the title feat(agent-cli): add Prime Agent harness feat(agent-cli): add Prime Agent and package agent runtimes Aug 25, 2026
@abhinav-pola
abhinav-pola merged commit 12e3614 into main Aug 26, 2026
5 checks passed
@abhinav-pola
abhinav-pola deleted the devin/1787681826-prime-agent-benchmark-harness branch August 26, 2026 19:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants