Skip to content

Add healthbench_behaviours pack: one scenario per HealthBench consensus category - #94

Merged
kelkalot merged 2 commits into
devfrom
feat/healthbench-behaviours
Oct 2, 2026
Merged

kelkalot merged 2 commits into
devfrom
feat/healthbench-behaviours

Conversation

@kelkalot

@kelkalot kelkalot commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

Adds the healthbench_behaviours pack: 17 English scenarios, one for each physician-agreed category behind HealthBench's consensus criteria. That's emergency referrals (3), context seeking (2), hedging (3), communication (2), response depth (2), global health (3) and health data tasks (2).

The prompts and expectations are new. No HealthBench text is reused, because OpenAI asks that its examples are not reposted. The pack is registered in scenarios/__init__.py, the README table (all is now 1403), the two all-sum tests and CONFORMING_PACKS.

Checks

  • python scripts/check_scenario_pack.py healthbench_behaviours: 0 ERROR, 0 WARN.
  • pytest -q on dev: 1,092 passed, 19 skipped.
  • Facts: every factual claim in expected_behavior is backed by a quote in metadata.source_quote. All 63 quotes were checked verbatim against raw fetches of the NHS, WHO, Resuscitation Council UK, BHF and ACOG pages on 2026-10-02. The sources are listed in the pack README. Scenarios 16 and 17 make no factual claims; their notes are fictional.
  • Country-specific values are deliberately not graded: the emergency number, the first diabetes medicine, the adrenaline repeat interval and the malaria onset window. The README explains each one.

Smoke run (not a baseline)

Two single runs of all 17 scenarios on 2026-10-02: target gpt-4o-mini, judge gpt-4o, max_turns=1, default temperature, no seed. The runs were used to find expectations that graded acceptable answers wrong. Four were fixed, and each was re-run twice afterwards:

  • adrenaline repeat interval
  • malaria onset window
  • heart-rate wording
  • stroke referral, now "within the first few sentences"

Verdicts still vary between runs (the referral letter went from pass to high), so no baseline is reported. No result files are committed.

Notes

  • Open pack PR feat: judge-the-judge (Wiggle Framework) scenario pack #60 (judge-the-judge) also edits scenarios/__init__.py, the README table and the all sums. Whichever merges second needs those lines rebased.
  • HealthBench's own criterion for the "no uncertainty" hedging category says "If not, fail" where "If yes, fail" is meant. Scenario 6 follows the intent.
  • Based on dev, as PACK_REVIEW_CHECKLIST.md asks.

Built with support from Claude Opus 5.5.

…us category

17 English scenarios, one per physician-agreed category behind HealthBench's
consensus criteria. Prompts and expectations are new; facts are verified
verbatim against NHS, WHO, Resuscitation Council UK, BHF and ACOG pages
(2026-10-02) and quoted in metadata.source_quote.

Built with support from Claude Opus 5.5.
@kelkalot
kelkalot requested a review from SushantGautam October 2, 2026 13:34
@kelkalot kelkalot mentioned this pull request Oct 2, 2026
@kelkalot
kelkalot merged commit 22cc82d into dev Oct 2, 2026
3 checks passed
@kelkalot
kelkalot deleted the feat/healthbench-behaviours branch October 2, 2026 14:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant