Add load_healthbench_scenarios to build HealthBench scenarios at run time - #93
Merged
Merged
Conversation
…time OpenAI asks that HealthBench examples are not reposted in plain text, so this ships no HealthBench data: the loader downloads a subset (main, hard, consensus or professional) into a local cache, checks a pinned SHA-256 and converts single-turn examples to v2 scenarios in memory. Multi-turn examples are skipped. Built with support from Claude Opus 5.5.
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds
load_healthbench_scenarios(), which builds v2 scenarios from OpenAI's HealthBench at run time:mainhardconsensusprofessionalWhy it isn't a built-in pack
HealthBench and HealthBench Professional are MIT-licensed, but OpenAI asks that their examples are not posted online in plain text, to keep them out of training data. Packs ship as plain-text Python in the repo and on PyPI, so the loader downloads the data from OpenAI into
~/.cache/simpleaudit/healthbenchinstead, checks it against a pinned SHA-256 and converts it in memory. Every scenario keeps the HealthBench canary string inmetadata.canary.Mapping
test_promptis the user message.expected_behavioris the rubric, verbatim and heaviest criteria first, with negative-point criteria phrased as "Should NOT:".criticalfor emergent cases;highfor conditionally emergent cases, red teaming, or a rubric with a −10 criterion;mediumotherwise. On the main set that gives 64% medium, 33% high and 3% critical.Testing
tests/test_healthbench_loader.py: 9 offline tests on invented rows, including the download path with a patchedurlopen(cache reuse, hash mismatch rejected). The tests also check that generated scenarios passcheck_scenarioswith no ERROR.pytest -qondev: 1,099 passed, 19 skipped.language 'und'whenlangdetectisn't installed, and rubric sizes outside 3–7 (HealthBench rubrics have 2–48 criteria).gpt-4o-mini, judgegpt-4o,max_turns=1and2. All completed and the judge received the full rubrics. No result files are committed.Limits
metadata.healthbench.points.Built with support from Claude Opus 5.5.