A test runner for non-deterministic software.
Normal test suites assume the same input produces the same output. Agents violate that. The same task, run five times against the same model, can produce five different trajectories and three different outcomes. That breaks every assumption in pytest.
Agent Eval Harness runs an LLM agent against a fixed suite of multi-step tasks, repeatedly, and produces a statistically rigorous report of how often it succeeds, how much it costs, how long it takes, and how it fails.
The agent being measured is the agent under test (AUT). The harness never contains agent logic. It runs, observes, grades, and reports.
The harness is a domain-agnostic engine. A task suite is a plugin. customer_support (the original workspace / HTTP / SQL tasks) and toy_math (three arithmetic tasks) load through the same loader. The engine never imports a concrete tool or a domain task type.
A test runner returns a boolean. Measuring an agent takes more:
- Run each task n times and report a success rate with a confidence interval.
- Attribute cost and latency per task and per step, because the cheapest agent that clears a quality bar wins, not the most accurate one.
- Classify how a run failed. "Hallucinated a tool argument" and "the tool timed out" need different fixes.
- Detect regressions between two runs when neither run is deterministic.
- Measure whether the agent correctly declines. An agent that exfiltrates a fixture token or attempts
DROP TABLEis not successful, even when it is fluent.
Every rate ships with its n and a Wilson 95% interval. Suite success rate is the unweighted mean of per-task rates, not the pooled attempt count, so a task with more attempts cannot dominate the headline number.
Grading runs in four tiers. Each lower tier refines the picture; none can overturn the tier above it.
flowchart TB
T0["Tier 0 — outcome graders<br/>decide success"]
T1["Tier 1 — invariants<br/>every trace, even successes"]
T2["Tier 2 — golden divergence<br/>failed + has golden, never flips success"]
T3["Tier 3 — human review flag<br/>failed and no confident localization"]
T0 --> T1 --> T2 --> T3
Tier 0 — outcome graders decide success. Exact / contains / numeric / file / SQL / tool-sequence / judge. All of a task's outcome_graders must pass for success = True. (YAML says graders:; that is an alias.)
Tier 1 — invariants run on every trace, including successes. Invariants are pure functions of the trace: no I/O, order-independent, and they do not go stale when the model changes. A run can produce the correct answer and still break a safety rule, and that is a finding, not a pass. Keeping invariants below the judge keeps a deterministic rule deterministic instead of expensive and flaky.
Tier 2 — golden divergence localizes; it does not grade. Matching a single known-good path throws false alarms, because most tasks have several correct trajectories. Divergence is not failure. Tier 2 runs only when tier 0 has already failed and the task has a golden trace. The judge is asked to name the earliest decision that put the run on a wrong path, ignoring order, phrasing, and tool sequencing that do not affect correctness. Output is strict JSON; malformed output is retried once, then marked errored. Tier 2 never flips success. compare warns when the judge model or the tier-2 prompt version changed.
Tier 3 — human review is a flag, not a gate. When tier 0 failed and there is no golden, or tier 2 is low-confidence or errored, needs_human_review is set and a review packet is stored. It does not block the run and it does not decide pass/fail.
A successful run gets no bucket. A failed run is assigned exactly one, in this priority order, so the classification names the root cause instead of the last symptom:
tool_error— a tool failed and the agent did not recovermalformed_tool_call— bad tool argumentshallucinated_tool— called a tool that was not on the allowlistbudget_exhausted— hit max steps, timeout, or token budgetwrong_answer— finished with output, graders failedno_answer— finished with empty outputrefusal— declined a task it should have doneharness_error— the harness itself broke (or nothing else matched)
harness_error is kept visible rather than folded into "failed." A run with a pile of harness errors is not a valid measurement, and the report says so.
Wilson score 95% intervals, not the normal approximation. With n at 5–10 and rates sitting near 0 or 1, the normal interval collapses to zero width at the extremes (5/5 becomes [1.00, 1.00]) and reports wrong coverage at small n. Wilson holds up in exactly that regime; 5/5 becomes [0.57, 1.00].
center = (k + z²/2) / (n + z²)
half = z/(n + z²) * sqrt(k*(n-k)/n + z²/4)
with z = 1.96. The point estimate is still k/n.
Cohen's κ for judge-vs-programmatic agreement, on the overlap set of tasks that have both a programmatic grader and a judge. Raw agreement is misleading because two graders that mostly say "pass" agree often by chance; κ corrects for that. The programmatic grader is the source of truth wherever it exists. Below κ ≈ 0.7 the judge is treated as untrustworthy.
Two-proportion test plus Benjamini–Hochberg for regression detection. Comparing two runs across 32 tasks means 32 simultaneous tests, so a naive p < 0.05 per task manufactures one or two false regressions every comparison. Bonferroni over-corrects and buries real drops. Benjamini–Hochberg controls the false discovery rate — the proportion of flagged regressions that are false — which is the quantity that matters here. A task is flagged only when it is significant after BH correction and the point estimate dropped.
Cost per completed task and cost per success are both reported. cost_per_success = total_cost / successes exposes an agent that looks cheap only because it fails fast: a low per-task cost with a high per-success cost means you are paying repeatedly to eventually get a usable result.
Cost is computed from token usage against a versioned price table. Every run stamps price_table_version and a date, and the report prints them, so a number is never silently priced against a stale rate.
Tool failures are injected to measure recovery, classified by error_kind because each exercises a different competency:
timeout— the call hangs500— the tool returns a server errorgarbage— the tool returns malformed output
Faults are seeded, so a run is reproducible, and the injected kind is recorded on ToolCall.injected_fault so the report separates injected failures from the agent's own mistakes. Burst mode fails a contiguous window of calls, because real outages are correlated; uniform random per-call failure is the weaker model, and the recovery-curve title labels which was used. Profiles live in faults.yaml.
The default log level is progress: one line per task. Verbose-by-default logs train people to ignore them.
Traces are trees, not lists. --log-format pretty indents by span depth, so sub-steps and parallel tool calls nest visually. Every record carries run_id, task_id, attempt, span_id, parent_span_id, category, so concurrent tasks can be reassembled even when their lines interleave.
Redaction happens on write, before persist. Emails, card numbers, and configured secrets are scrubbed so a stored trace is safe to share and safe to hand to someone else for debugging. --no-redact exists and is off by default.
Controls:
--log-level—silent | progress | steps | calls | trace--log-only/--log-exclude— filtertools, model, grading, faults, scheduler, cost--log-format—pretty | jsonl
The trace viewer is one self-contained HTML file per attempt, vanilla JS, no framework — an indented span tree you can open in a browser and collapse node by node.
The report is a single self-contained HTML file, charts embedded as base64 PNGs, so it is one artifact you can share with no external assets. It prints a warning banner when harness error rate exceeds 2%. Traces persist to SQLite (runs.db, WAL), redacted on write.
| Piece | Role |
|---|---|
| YAML tasks | One file per task, validated loudly |
| Suite plugins | EvalSuite contract: tasks, tools, invariants, golden traces |
| Adapters | single_shot baseline, react, replay, mock |
| Tools | Engine keeps registry + faults. Concrete tools live in the suite. |
| Graders | Tier 0 programmatic + tier 1 invariants + tier 2 golden divergence + tier 3 human flag |
| Storage | SQLite runs.db, WAL |
| Report | Single self-contained HTML, charts as base64 PNGs |
| Trace viewer | One self-contained HTML span tree per attempt |
| Compare | aeh compare <run_a> <run_b> |
suites/core has 32 tasks:
- 9 multi-step file / data manipulation
- 6 tool-use tasks that need 2+ different tools
- 5 retrieval tasks against a local HTTP fixture on
127.0.0.1 - 4 long-horizon tasks that need 10+ steps
- 4 tasks with a deliberately unavailable tool (
broken_tool) to measure recovery - 3 refusal tasks whose correct behavior is to decline or ask for clarification
- 1 SQL analytics task
suites/smoke has 5 tasks and finishes in well under 60s on the mock adapter. suites/toy_math has 3 arithmetic tasks, one calculator tool, and one invariant, and exists to prove the engine is not glued to the customer-support domain.
The loader fails on unknown tools, unknown invariant names, unknown graders, duplicate ids, a missing golden trace, missing fixtures, and graders whose config does not validate. A silently skipped task is a corrupted measurement.
pip install -e ".[dev]"
# or: uv sync --extra devPython 3.11+.
aeh tasks validate
aeh run --suite smoke --adapter mock --attempts 1
aeh run --suite ./suites/toy_math --adapter mock --log-level steps --log-only tools
aeh run --suite core --adapter mock --dry-run--dry-run validates the suite, resolves config, prints an estimated token cost, and exits without calling a model.
Live adapters need OPENAI_API_KEY (and optionally OPENAI_BASE_URL for any OpenAI-compatible server):
aeh run --suite core --adapter react --model gpt-4o-mini --attempts 5 --concurrency 4
aeh run --suite core --adapter single_shot --model gpt-4o-mini --attempts 5
aeh run --suite core --adapter react --attempts 5 --fault-profile faults.yaml
aeh report --run-id <id> --out report.html
aeh trace --run-id <id> --task <task_id> --attempt 0 --out trace.html
aeh compare <run_a> <run_b>
aeh judge-agreement --run-id <id>
aeh replay --run-id <id>The engine does not know what a ticket, a file, or a sum is. It knows one contract:
from aeh.models import Task, Trace
from aeh.tools.registry import ToolRegistry
class DemoSuite:
name = "demo"
def tasks(self) -> list[Task]: ...
def tool_registry(self) -> ToolRegistry: ...
def invariants(self) -> dict: ... # name -> (trace) -> pass/fail
def golden_traces(self) -> dict[str, Trace]: ...Point --suite at the directory, or install an aeh.suites entry point. If a task names an unknown tool, an unknown invariant, a missing golden, or a grader that does not validate, load fails loudly.
flowchart LR
SuitePlugin["Suite plugin<br/>tasks · tools · invariants · goldens"] --> Engine
Engine --> Adapters
Engine --> Tier0["Tier 0 outcome graders"]
Engine --> Tier1["Tier 1 invariants"]
Engine --> Faults["Fault wrapper"]
Engine --> Store["SQLite + redacted traces"]
Store --> Report
Store --> TreeViewer["Trace tree HTML"]
Invariants sit below the judge. They are pure functions of the trace — no I/O, order-independent, model-agnostic — so a deterministic safety rule stays a cheap boolean instead of an expensive, flaky LLM call. A correct answer that violated a rule is a finding, not a pass.
Golden traces localize; they do not grade. Most tasks have several correct trajectories, so matching one known-good path would false-alarm constantly. Tier 2 runs only after a tier-0 failure, names the earliest wrong decision, and never flips success.
Human review is a flag, not a gate. It routes the runs the pipeline cannot confidently localize to a person without blocking the run or deciding pass/fail.
Faults model correlated outages. error_kind splits timeout | 500 | garbage because each needs different recovery, and burst mode fails contiguous windows because real outages cluster.
Cost is a first-class metric. Cost per completed task and cost per success are reported side by side, and recovery rate under injected tool failure is measured at multiple fail rates. Those two numbers are the production-engineering lens on an agent.
judge.py and agreement.py are independent: Cohen's κ is a different use of the judge from tier-2 localization.
Populate live react vs single_shot results with:
aeh run --suite core --adapter react --model gpt-4o-mini --attempts 5 --concurrency 4
aeh run --suite core --adapter single_shot --model gpt-4o-mini --attempts 5
aeh compare <react_run> <single_shot_run>
aeh judge-agreement --run-id <react_run>pytestCI monkeypatches httpx.AsyncClient.send to raise. Tests that need the local fixture server opt in with @pytest.mark.local_http. There is no public network in CI.
docker build -t aeh .
docker compose run --rm harness tasks validate