Add healthbench_behaviours pack: one scenario per HealthBench consensus category - #94
Merged
Merged
Conversation
…us category 17 English scenarios, one per physician-agreed category behind HealthBench's consensus criteria. Prompts and expectations are new; facts are verified verbatim against NHS, WHO, Resuscitation Council UK, BHF and ACOG pages (2026-10-02) and quoted in metadata.source_quote. Built with support from Claude Opus 5.5.
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds the
healthbench_behaviourspack: 17 English scenarios, one for each physician-agreed category behind HealthBench's consensus criteria. That's emergency referrals (3), context seeking (2), hedging (3), communication (2), response depth (2), global health (3) and health data tasks (2).The prompts and expectations are new. No HealthBench text is reused, because OpenAI asks that its examples are not reposted. The pack is registered in
scenarios/__init__.py, the README table (allis now 1403), the twoall-sum tests andCONFORMING_PACKS.Checks
python scripts/check_scenario_pack.py healthbench_behaviours: 0 ERROR, 0 WARN.pytest -qondev: 1,092 passed, 19 skipped.expected_behavioris backed by a quote inmetadata.source_quote. All 63 quotes were checked verbatim against raw fetches of the NHS, WHO, Resuscitation Council UK, BHF and ACOG pages on 2026-10-02. The sources are listed in the pack README. Scenarios 16 and 17 make no factual claims; their notes are fictional.Smoke run (not a baseline)
Two single runs of all 17 scenarios on 2026-10-02: target
gpt-4o-mini, judgegpt-4o,max_turns=1, default temperature, no seed. The runs were used to find expectations that graded acceptable answers wrong. Four were fixed, and each was re-run twice afterwards:Verdicts still vary between runs (the referral letter went from pass to high), so no baseline is reported. No result files are committed.
Notes
scenarios/__init__.py, the README table and theallsums. Whichever merges second needs those lines rebased.dev, asPACK_REVIEW_CHECKLIST.mdasks.Built with support from Claude Opus 5.5.