An open source library for evaluating model performance across different tasks.
This project contains the @primer/agent-eval library and CLI for evaluating agent performance across different tasks. This is useful for establishing
benchmarks for the capabilities of your project (like a design system) or for running experiments to understand the best way to improve model outcomes for the task at hand.
To do so, we create a variety of scenarios that describe the task, the starting workspace, and the checks or judges that evaluate the result. These scenarios are then used in benchmarks and experiments to evaluate model performance.
To learn more about how to use this library, visit @primer/agent-eval or install the agent-eval skill to get started.
Install the agent-eval skill to help your agent set up and run evaluations:
npx skills add primer/agent-eval --skill agent-evalThe skill includes a getting-started guide and references for evaluation methodology, domain models, and the CLI. The skill and runtime package are installed separately.
Install the agent-eval-investigate skill to investigate experiment or benchmark results, explain observed model behavior (such as why a tool was or was not called), and suggest setup improvements:
npx skills add primer/agent-eval --skill agent-eval-investigateUse /agent-eval-investigate with a result bundle, the relevant scenario or trial,
and the behavior you want to understand. The skill can also offer to apply a
proposed change and rerun a focused comparison with your approval.
We love collaborating with folks inside and outside of GitHub and welcome contributions! If you're interested, check out our contributing docs for more info on how to get started.
Licensed under the MIT License.