Skip to content

Repository files navigation

Agent Eval Harness

A test runner for non-deterministic software.

Normal test suites assume the same input produces the same output. Agents violate that. The same task, run five times against the same model, can produce five different trajectories and three different outcomes. That breaks every assumption in pytest.

Agent Eval Harness runs an LLM agent against a fixed suite of multi-step tasks, repeatedly, and produces a statistically rigorous report of how often it succeeds, how much it costs, how long it takes, and how it fails.

The agent being measured is the agent under test (AUT). The harness never contains agent logic. It runs, observes, grades, and reports.

The harness is a domain-agnostic engine. A task suite is a plugin. customer_support (the original workspace / HTTP / SQL tasks) and toy_math (three arithmetic tasks) load through the same loader. The engine never imports a concrete tool or a domain task type.

What it measures

A test runner returns a boolean. Measuring an agent takes more:

  • Run each task n times and report a success rate with a confidence interval.
  • Attribute cost and latency per task and per step, because the cheapest agent that clears a quality bar wins, not the most accurate one.
  • Classify how a run failed. "Hallucinated a tool argument" and "the tool timed out" need different fixes.
  • Detect regressions between two runs when neither run is deterministic.
  • Measure whether the agent correctly declines. An agent that exfiltrates a fixture token or attempts DROP TABLE is not successful, even when it is fluent.

Every rate ships with its n and a Wilson 95% interval. Suite success rate is the unweighted mean of per-task rates, not the pooled attempt count, so a task with more attempts cannot dominate the headline number.


Features

Multi-tier grading ladder

Grading runs in four tiers. Each lower tier refines the picture; none can overturn the tier above it.

flowchart TB
  T0["Tier 0 — outcome graders<br/>decide success"]
  T1["Tier 1 — invariants<br/>every trace, even successes"]
  T2["Tier 2 — golden divergence<br/>failed + has golden, never flips success"]
  T3["Tier 3 — human review flag<br/>failed and no confident localization"]
  T0 --> T1 --> T2 --> T3
Loading

Tier 0 — outcome graders decide success. Exact / contains / numeric / file / SQL / tool-sequence / judge. All of a task's outcome_graders must pass for success = True. (YAML says graders:; that is an alias.)

Tier 1 — invariants run on every trace, including successes. Invariants are pure functions of the trace: no I/O, order-independent, and they do not go stale when the model changes. A run can produce the correct answer and still break a safety rule, and that is a finding, not a pass. Keeping invariants below the judge keeps a deterministic rule deterministic instead of expensive and flaky.

Tier 2 — golden divergence localizes; it does not grade. Matching a single known-good path throws false alarms, because most tasks have several correct trajectories. Divergence is not failure. Tier 2 runs only when tier 0 has already failed and the task has a golden trace. The judge is asked to name the earliest decision that put the run on a wrong path, ignoring order, phrasing, and tool sequencing that do not affect correctness. Output is strict JSON; malformed output is retried once, then marked errored. Tier 2 never flips success. compare warns when the judge model or the tier-2 prompt version changed.

Tier 3 — human review is a flag, not a gate. When tier 0 failed and there is no golden, or tier 2 is low-confidence or errored, needs_human_review is set and a review packet is stored. It does not block the run and it does not decide pass/fail.

Failure buckets

A successful run gets no bucket. A failed run is assigned exactly one, in this priority order, so the classification names the root cause instead of the last symptom:

  1. tool_error — a tool failed and the agent did not recover
  2. malformed_tool_call — bad tool arguments
  3. hallucinated_tool — called a tool that was not on the allowlist
  4. budget_exhausted — hit max steps, timeout, or token budget
  5. wrong_answer — finished with output, graders failed
  6. no_answer — finished with empty output
  7. refusal — declined a task it should have done
  8. harness_error — the harness itself broke (or nothing else matched)

harness_error is kept visible rather than folded into "failed." A run with a pile of harness errors is not a valid measurement, and the report says so.

Statistical methods

Wilson score 95% intervals, not the normal approximation. With n at 5–10 and rates sitting near 0 or 1, the normal interval collapses to zero width at the extremes (5/5 becomes [1.00, 1.00]) and reports wrong coverage at small n. Wilson holds up in exactly that regime; 5/5 becomes [0.57, 1.00].

center = (k + z²/2) / (n + z²)
half   = z/(n + z²) * sqrt(k*(n-k)/n + z²/4)

with z = 1.96. The point estimate is still k/n.

Cohen's κ for judge-vs-programmatic agreement, on the overlap set of tasks that have both a programmatic grader and a judge. Raw agreement is misleading because two graders that mostly say "pass" agree often by chance; κ corrects for that. The programmatic grader is the source of truth wherever it exists. Below κ ≈ 0.7 the judge is treated as untrustworthy.

Two-proportion test plus Benjamini–Hochberg for regression detection. Comparing two runs across 32 tasks means 32 simultaneous tests, so a naive p < 0.05 per task manufactures one or two false regressions every comparison. Bonferroni over-corrects and buries real drops. Benjamini–Hochberg controls the false discovery rate — the proportion of flagged regressions that are false — which is the quantity that matters here. A task is flagged only when it is significant after BH correction and the point estimate dropped.

Cost accounting

Cost per completed task and cost per success are both reported. cost_per_success = total_cost / successes exposes an agent that looks cheap only because it fails fast: a low per-task cost with a high per-success cost means you are paying repeatedly to eventually get a usable result.

Cost is computed from token usage against a versioned price table. Every run stamps price_table_version and a date, and the report prints them, so a number is never silently priced against a stale rate.

Fault injection

Tool failures are injected to measure recovery, classified by error_kind because each exercises a different competency:

  • timeout — the call hangs
  • 500 — the tool returns a server error
  • garbage — the tool returns malformed output

Faults are seeded, so a run is reproducible, and the injected kind is recorded on ToolCall.injected_fault so the report separates injected failures from the agent's own mistakes. Burst mode fails a contiguous window of calls, because real outages are correlated; uniform random per-call failure is the weaker model, and the recovery-curve title labels which was used. Profiles live in faults.yaml.

Logging and tracing

The default log level is progress: one line per task. Verbose-by-default logs train people to ignore them.

Traces are trees, not lists. --log-format pretty indents by span depth, so sub-steps and parallel tool calls nest visually. Every record carries run_id, task_id, attempt, span_id, parent_span_id, category, so concurrent tasks can be reassembled even when their lines interleave.

Redaction happens on write, before persist. Emails, card numbers, and configured secrets are scrubbed so a stored trace is safe to share and safe to hand to someone else for debugging. --no-redact exists and is off by default.

Controls:

  • --log-level — silent | progress | steps | calls | trace
  • --log-only / --log-exclude — filter tools, model, grading, faults, scheduler, cost
  • --log-format — pretty | jsonl

The trace viewer is one self-contained HTML file per attempt, vanilla JS, no framework — an indented span tree you can open in a browser and collapse node by node.

Reporting

The report is a single self-contained HTML file, charts embedded as base64 PNGs, so it is one artifact you can share with no external assets. It prints a warning banner when harness error rate exceeds 2%. Traces persist to SQLite (runs.db, WAL), redacted on write.


What is in the box

Piece Role
YAML tasks One file per task, validated loudly
Suite plugins EvalSuite contract: tasks, tools, invariants, golden traces
Adapters single_shot baseline, react, replay, mock
Tools Engine keeps registry + faults. Concrete tools live in the suite.
Graders Tier 0 programmatic + tier 1 invariants + tier 2 golden divergence + tier 3 human flag
Storage SQLite runs.db, WAL
Report Single self-contained HTML, charts as base64 PNGs
Trace viewer One self-contained HTML span tree per attempt
Compare aeh compare <run_a> <run_b>

Task suite

suites/core has 32 tasks:

  • 9 multi-step file / data manipulation
  • 6 tool-use tasks that need 2+ different tools
  • 5 retrieval tasks against a local HTTP fixture on 127.0.0.1
  • 4 long-horizon tasks that need 10+ steps
  • 4 tasks with a deliberately unavailable tool (broken_tool) to measure recovery
  • 3 refusal tasks whose correct behavior is to decline or ask for clarification
  • 1 SQL analytics task

suites/smoke has 5 tasks and finishes in well under 60s on the mock adapter. suites/toy_math has 3 arithmetic tasks, one calculator tool, and one invariant, and exists to prove the engine is not glued to the customer-support domain.

The loader fails on unknown tools, unknown invariant names, unknown graders, duplicate ids, a missing golden trace, missing fixtures, and graders whose config does not validate. A silently skipped task is a corrupted measurement.


Install

pip install -e ".[dev]"
# or: uv sync --extra dev

Python 3.11+.

aeh tasks validate
aeh run --suite smoke --adapter mock --attempts 1
aeh run --suite ./suites/toy_math --adapter mock --log-level steps --log-only tools
aeh run --suite core --adapter mock --dry-run

--dry-run validates the suite, resolves config, prints an estimated token cost, and exits without calling a model.

Live adapters need OPENAI_API_KEY (and optionally OPENAI_BASE_URL for any OpenAI-compatible server):

aeh run --suite core --adapter react --model gpt-4o-mini --attempts 5 --concurrency 4
aeh run --suite core --adapter single_shot --model gpt-4o-mini --attempts 5
aeh run --suite core --adapter react --attempts 5 --fault-profile faults.yaml
aeh report --run-id <id> --out report.html
aeh trace --run-id <id> --task <task_id> --attempt 0 --out trace.html
aeh compare <run_a> <run_b>
aeh judge-agreement --run-id <id>
aeh replay --run-id <id>

Write your own suite

The engine does not know what a ticket, a file, or a sum is. It knows one contract:

from aeh.models import Task, Trace
from aeh.tools.registry import ToolRegistry

class DemoSuite:
    name = "demo"
    def tasks(self) -> list[Task]: ...
    def tool_registry(self) -> ToolRegistry: ...
    def invariants(self) -> dict: ...          # name -> (trace) -> pass/fail
    def golden_traces(self) -> dict[str, Trace]: ...

Point --suite at the directory, or install an aeh.suites entry point. If a task names an unknown tool, an unknown invariant, a missing golden, or a grader that does not validate, load fails loudly.

flowchart LR
  SuitePlugin["Suite plugin<br/>tasks · tools · invariants · goldens"] --> Engine
  Engine --> Adapters
  Engine --> Tier0["Tier 0 outcome graders"]
  Engine --> Tier1["Tier 1 invariants"]
  Engine --> Faults["Fault wrapper"]
  Engine --> Store["SQLite + redacted traces"]
  Store --> Report
  Store --> TreeViewer["Trace tree HTML"]
Loading

Design decisions

Invariants sit below the judge. They are pure functions of the trace — no I/O, order-independent, model-agnostic — so a deterministic safety rule stays a cheap boolean instead of an expensive, flaky LLM call. A correct answer that violated a rule is a finding, not a pass.

Golden traces localize; they do not grade. Most tasks have several correct trajectories, so matching one known-good path would false-alarm constantly. Tier 2 runs only after a tier-0 failure, names the earliest wrong decision, and never flips success.

Human review is a flag, not a gate. It routes the runs the pipeline cannot confidently localize to a person without blocking the run or deciding pass/fail.

Faults model correlated outages. error_kind splits timeout | 500 | garbage because each needs different recovery, and burst mode fails contiguous windows because real outages cluster.

Cost is a first-class metric. Cost per completed task and cost per success are reported side by side, and recovery rate under injected tool failure is measured at multiple fail rates. Those two numbers are the production-engineering lens on an agent.

judge.py and agreement.py are independent: Cohen's κ is a different use of the judge from tier-2 localization.


Populate live react vs single_shot results with:

aeh run --suite core --adapter react --model gpt-4o-mini --attempts 5 --concurrency 4
aeh run --suite core --adapter single_shot --model gpt-4o-mini --attempts 5
aeh compare <react_run> <single_shot_run>
aeh judge-agreement --run-id <react_run>

Tests

pytest

CI monkeypatches httpx.AsyncClient.send to raise. Tests that need the local fixture server opt in with @pytest.mark.local_http. There is no public network in CI.

Docker

docker build -t aeh .
docker compose run --rm harness tasks validate

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages