Skip to content

Add load_healthbench_scenarios to build HealthBench scenarios at run time - #93

Merged
kelkalot merged 2 commits into
devfrom
feat/healthbench-loader
Oct 2, 2026
Merged

kelkalot merged 2 commits into
devfrom
feat/healthbench-loader

Conversation

@kelkalot

@kelkalot kelkalot commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

Adds load_healthbench_scenarios(), which builds v2 scenarios from OpenAI's HealthBench at run time:

from simpleaudit import load_healthbench_scenarios

scenarios = load_healthbench_scenarios("hard", themes=["emergency_referrals"], limit=20)
results = auditor.run(scenarios, max_turns=1)
Subset Examples Single-turn scenarios
main 5,000 2,915
hard 1,000 523
consensus 3,671 2,201
professional 525 410

Why it isn't a built-in pack

HealthBench and HealthBench Professional are MIT-licensed, but OpenAI asks that their examples are not posted online in plain text, to keep them out of training data. Packs ship as plain-text Python in the repo and on PyPI, so the loader downloads the data from OpenAI into ~/.cache/simpleaudit/healthbench instead, checks it against a pinned SHA-256 and converts it in memory. Every scenario keeps the HealthBench canary string in metadata.canary.

Mapping

  • test_prompt is the user message. expected_behavior is the rubric, verbatim and heaviest criteria first, with negative-point criteria phrased as "Should NOT:".
  • The theme sets category and subcategory (taxonomy values only). Severity is critical for emergent cases; high for conditionally emergent cases, red teaming, or a rubric with a −10 criterion; medium otherwise. On the main set that gives 64% medium, 33% high and 3% critical.
  • Multi-turn examples are skipped: a scenario has no field for earlier turns.

Testing

  • tests/test_healthbench_loader.py: 9 offline tests on invented rows, including the download path with a patched urlopen (cache reuse, hash mismatch rejected). The tests also check that generated scenarios pass check_scenarios with no ERROR.
  • pytest -q on dev: 1,099 passed, 19 skipped.
  • Real data: all four subsets download and convert with no checker ERROR and no duplicate names. The WARNs are language 'und' when langdetect isn't installed, and rubric sizes outside 3–7 (HealthBench rubrics have 2–48 criteria).
  • Smoke run on 2026-10-02: 3 scenarios (Hard emergency, Hard health-data task, Professional red teaming), target gpt-4o-mini, judge gpt-4o, max_turns=1 and 2. All completed and the judge received the full rubrics. No result files are committed.

Limits

  • The judge returns a severity verdict, not a HealthBench score. Rubric points are kept in metadata.healthbench.points.
  • Supporting multi-turn examples would need a scenario field for earlier messages.
  • Results and HTML exports contain the prompts, so the README asks people not to publish them for these runs.

Built with support from Claude Opus 5.5.

…time

OpenAI asks that HealthBench examples are not reposted in plain text, so
this ships no HealthBench data: the loader downloads a subset (main, hard,
consensus or professional) into a local cache, checks a pinned SHA-256 and
converts single-turn examples to v2 scenarios in memory. Multi-turn
examples are skipped.

Built with support from Claude Opus 5.5.
@kelkalot
kelkalot requested a review from SushantGautam October 2, 2026 13:35
@kelkalot kelkalot mentioned this pull request Oct 2, 2026
@kelkalot
kelkalot merged commit 4d3d771 into dev Oct 2, 2026
3 checks passed
@kelkalot
kelkalot deleted the feat/healthbench-loader branch October 2, 2026 14:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant