Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
43 changes: 43 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -548,6 +548,49 @@ Contributing a pack: follow the
`python scripts/check_scenario_pack.py <pack>` before opening the PR, and expect the review to
follow the [pack review checklist](simpleaudit/scenarios/PACK_REVIEW_CHECKLIST.md).

### HealthBench (loaded at run time)

[HealthBench](https://openai.com/index/healthbench/) and
[HealthBench Professional](https://arxiv.org/abs/2604.27470) are OpenAI's MIT-licensed health
benchmarks with physician-written rubrics. OpenAI asks that their examples are not posted online in
plain text, so they are not a built-in pack. `load_healthbench_scenarios` downloads a subset from
OpenAI into `~/.cache/simpleaudit/healthbench`, checks it against a pinned SHA-256 and builds v2
scenarios in memory:

```python
from simpleaudit import load_healthbench_scenarios

scenarios = load_healthbench_scenarios("hard", themes=["emergency_referrals"], limit=20, seed=0)
results = auditor.run(scenarios, max_turns=1)
```

| Subset | Examples | Single-turn scenarios |
|--------|----------|-----------------------|
| `main` | 5,000 | 2,915 |
| `hard` | 1,000 | 523 |
| `consensus` | 3,671 | 2,201 |
| `professional` | 525 | 410 |

Each scenario's `test_prompt` is the user's message and its `expected_behavior` is the example's
rubric, heaviest criteria first, with negative-point criteria phrased as "Should NOT". Themes map to
categories, and severity is `critical` for emergencies; `high` for possible emergencies, red teaming,
or a rubric that gives the maximum penalty of −10; `medium` otherwise. Rubric points, prompt ids and
the HealthBench canary string are in `metadata`.

Limits:

- **Multi-turn examples are skipped:** 42% of the main set and 22% of Professional. A scenario has no
field for earlier turns, and pasting them into one message would show the target earlier
"assistant" answers as user text.
- **Rubrics are written for one reply,** so use `max_turns=1` for HealthBench-style grading. Later
turns go beyond what the rubric covers.
- **The judge returns a severity verdict, not a HealthBench score.**
- **Rubrics have 2–48 criteria (Professional: 1–5),** so many scenarios fall outside the 3–7 the
scenario guideline asks for. Pass `min_criteria` and `max_criteria` to filter.
- **Language is detected only if `langdetect` is installed;** otherwise it is `"und"`.

Audit results and HTML exports contain the test prompts, so don't publish them for these runs.

### Vision Integrity

`vision_integrity` is the first pack that attaches images (via `file_uri`). It tests the same
Expand Down
3 changes: 2 additions & 1 deletion simpleaudit/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,7 @@
TargetResponse,
)
from .results import AuditResults, AuditResult
from .scenarios import get_scenarios, list_scenario_packs
from .scenarios import get_scenarios, list_scenario_packs, load_healthbench_scenarios
from .judges import build_judge, customize_judge, get_judge, list_judge_configs
from .experiment import AuditExperiment, ExperimentEvent
from .repeated_results import (
Expand Down Expand Up @@ -97,6 +97,7 @@
"AuditResult",
"get_scenarios",
"list_scenario_packs",
"load_healthbench_scenarios",
"get_judge",
"build_judge",
"customize_judge",
Expand Down
5 changes: 5 additions & 0 deletions simpleaudit/scenarios/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -34,6 +34,9 @@
lower-authority chunks (3 scenarios, requires SingleTurnAuditor; not part of 'all')
- healthbench_behaviours: One scenario per HealthBench consensus category (17 scenarios)
- all: All scenarios combined

HealthBench is not a built-in pack: OpenAI asks that its examples are not reposted in
plain text, so load_healthbench_scenarios() downloads it and builds scenarios at run time.
"""

from collections import Counter
Expand Down Expand Up @@ -66,6 +69,7 @@
from .nb_kryss_ordning import NB_KRYSS_ORDNING_SCENARIOS
from .context_grounding import CONTEXT_GROUNDING_SCENARIOS
from .healthbench_behaviours import HEALTHBENCH_BEHAVIOURS_SCENARIOS
from .healthbench_loader import load_healthbench_scenarios


SCENARIO_PACKS = {
Expand Down Expand Up @@ -180,5 +184,6 @@ def duplicate_scenario_names(scenarios: List[Dict]) -> Dict[str, int]:
"get_scenarios",
"list_scenario_packs",
"duplicate_scenario_names",
"load_healthbench_scenarios",
"SCENARIO_PACKS",
]
Loading
Loading