Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
79 commits
Select commit Hold shift + click to select a range
2fcd267
add token usage tracking
finnschwall Apr 5, 2026
4971ea9
add repeated-run stability analysis with significance testing and res…
finnschwall Apr 7, 2026
43f6e76
Merge branch 'main' into benchmarking
finnschwall Apr 8, 2026
249ca4e
separate judge/auditor as possibility
finnschwall Apr 10, 2026
17a0d08
fixed missing token count for judge
finnschwall Apr 13, 2026
d3c46fd
Add Hei refusal pack + test_prompt honoured on turn 1
kelkalot Apr 22, 2026
904d10a
removed unfinished statistics module
finnschwall Apr 26, 2026
ac04dbe
fix(judge): use json_schema response_format for Anthropic compatibility
avalyset Apr 29, 2026
dc5429d
Merge pull request #13 from kelkalot/feature/hei-refusal-pack
SushantGautam May 6, 2026
ae22601
Update README with developer and collaboration details
kelkalot May 6, 2026
915eb8f
Bump any-llm-sdk floor to >=1.9.0 and add Anthropic judge smoke test
avalyset May 6, 2026
49ca234
Merge pull request #16 from avalyset/fix/judge-anthropic-json-schema-pr
kelkalot May 6, 2026
32c494f
Add nav_aap scenario pack: 15 scenarios on Norwegian welfare administ…
avalyset May 6, 2026
6becf3e
Enhance README with new paper and contributor updates
kelkalot May 8, 2026
a3d0c90
Merge pull request #17 from avalyset/feat/nav-aap-scenarios
kelkalot May 8, 2026
a7f122f
feat(skatteetaten): scaffold scenario pack
avalyset Apr 29, 2026
bf48423
feat(skatteetaten): add 8 scenarios for Norwegian Tax Administration
avalyset Apr 29, 2026
07cd3ac
feat(skatteetaten): baseline evaluation results
avalyset Apr 29, 2026
67ef079
Merge remote-tracking branch 'upstream/main' into benchmarking
finnschwall May 11, 2026
d6aba45
Merge pull request #18 from avalyset/feat/skatteetaten-scenarios
kelkalot May 11, 2026
57b53d8
Add declarative judge response_schema and three new judges
kelkalot May 14, 2026
fee48ec
Merge remote-tracking branch 'upstream/main' into benchmarking
finnschwall May 16, 2026
ce6e228
prepare for sync + new tests
finnschwall May 18, 2026
c542629
fixed mock server + quickstart + parallel msg for async run. expanded…
finnschwall May 22, 2026
ca38da0
new tests + fixes in old tests
finnschwall May 22, 2026
e0d2e25
small bug fixes
finnschwall May 22, 2026
8a2c716
fix wrong forced json schema for judges
finnschwall May 23, 2026
71bd837
Merge pull request #19 from kelkalot/feature/helsedir-sexhealth-judge
SushantGautam May 27, 2026
ed06650
Merge main into benchmarking
SushantGautam May 27, 2026
69fabe1
Merge pull request #20 from finnschwall/benchmarking
SushantGautam May 27, 2026
e444928
Merge pull request #21 from kelkalot/dev
SushantGautam May 27, 2026
9ec7ada
Add on_model_done callback to AuditExperiment
SushantGautam May 27, 2026
3bf6533
Add CrossJudgeExperiment for cross-judge stability analysis
avalyset May 29, 2026
0cc6ba5
Merge pull request #24 from avalyset/feat/cross-judge-experiment
kelkalot Jun 4, 2026
62d3cc7
Fix audit-framework bugs and add regression tests
kelkalot Jun 11, 2026
8d51e50
Merge pull request #25 from kelkalot/fix/audit-bugs-2026-06
SushantGautam Jun 16, 2026
ac56078
Update pyproject.toml
SushantGautam Jun 16, 2026
bfee464
Clarify scenario descriptions in hei_refusal.py
kelkalot Jun 26, 2026
a69f8da
Document ung scenario pack and its themes
kelkalot Jun 26, 2026
a21e044
code owners and governance
kelkalot Jun 26, 2026
7e49bfe
Improve resilience, caching, and judge handling
kelkalot Jul 20, 2026
62f44b6
Harden path handling in get_json_file
kelkalot Jul 20, 2026
9340d7f
Tighten path validation to enforce root boundary
kelkalot Jul 20, 2026
596d460
Potential fix for pull request finding
SushantGautam Jul 21, 2026
b3eae01
Potential fix for pull request finding
SushantGautam Jul 21, 2026
8117e47
Potential fix for pull request finding
SushantGautam Jul 21, 2026
ff462ae
Potential fix for pull request finding
SushantGautam Jul 21, 2026
3050278
Enhance file validation and error handling in server.py
SushantGautam Jul 21, 2026
92a2560
Potential fix for pull request finding 'CodeQL / Uncontrolled data us…
SushantGautam Jul 21, 2026
65567d6
fix(visualizer): use CodeQL path sanitizer pattern for json file endp…
SushantGautam Jul 21, 2026
6642e78
Add Helfo (health-economics) scenario pack
avalyset Jul 8, 2026
b1660e5
helfo: address review — mechanical fixes (CI sums, anchors, names, do…
avalyset Jul 21, 2026
8e0059b
helfo: address review — rework scenario 3 (post-2025 psykolog rule), …
avalyset Jul 21, 2026
3d94c17
Add severity field to test payloads
kelkalot Jul 22, 2026
d8a01ad
Merge pull request #28 from kelkalot/fix/code-review-2026-07
SushantGautam Jul 23, 2026
3331d49
Merge pull request #27 from avalyset/add-helfo-scenario-pack
SushantGautam Jul 23, 2026
ec532f4
Bump version from 0.1.8 to 0.1.9
SushantGautam Jul 23, 2026
2d61dbb
Add helfo scenario pack and update to v0.1.9
kelkalot Jul 23, 2026
4e10133
Add lanekassen (Norwegian student finance) scenario pack
avalyset Jul 27, 2026
bbb681f
Add missing argument, fixes #30
kwinkunks Jul 28, 2026
dc5a28c
Merge pull request #31 from kwinkunks/judge-response-schema
kelkalot Jul 28, 2026
3fc80d7
Support multi-model experiment files in visualizer
kwinkunks Jul 31, 2026
1d68556
Add tests for multi-model experiment files
kwinkunks Jul 31, 2026
ebcbe3c
Merge pull request #33 from kwinkunks/expt-file-viz
kelkalot Jul 31, 2026
109a55d
Only list loadable models in experiments
kelkalot Jul 31, 2026
fa6dbd4
Merge pull request #34 from kelkalot/fix/visualizer-experiment-tree
kelkalot Jul 31, 2026
9b832c9
Drop unsupported hjemmeboer conversion claim, add §-anchors and no-ap…
avalyset Aug 1, 2026
4145569
Demote Sivilombudet two-step detail to optional (scenario 1)
avalyset Aug 1, 2026
485b67f
Clarify data sources in optional no-application bullet (scenario 2)
avalyset Aug 1, 2026
749a3d6
Align README conversion row with scenario 2 bullet; separate practice…
avalyset Aug 1, 2026
37a8006
Merge pull request #29 from avalyset/feat/lanekassen-scenarios
kelkalot Aug 2, 2026
9015a8f
Fix formatting in Custom Scenarios section
kwinkunks Aug 6, 2026
eca2ae0
Add support for image file attachment
kwinkunks Aug 6, 2026
7218313
Add docs for image file attachment
kwinkunks Aug 6, 2026
df96867
Add tests for image file attachment
kwinkunks Aug 6, 2026
4dd6270
Merge pull request #37 from kwinkunks/patch-1
kelkalot Aug 6, 2026
2b1cd1e
Raise any-llm floor for Gemini image support
kwinkunks Aug 7, 2026
63f76aa
Clear cached image bytes at the start of each run
kwinkunks Aug 7, 2026
7f97169
Merge pull request #38 from kwinkunks/send-image-content
kelkalot Aug 7, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -15,3 +15,8 @@ run_subset_test.py
.coverage
htmlcov/
coverage.xml

# Credentials — never commit
.env
.env.local
.env.*.local
53 changes: 53 additions & 0 deletions CODEOWNERS
Original file line number Diff line number Diff line change
@@ -0,0 +1,53 @@
# SimpleAudit — CODEOWNERS
#
# GitHub uses this file to auto-request reviews from the right people when a PR
# touches matching paths. See GOVERNANCE.md for the roles behind these owners.
#
# IMPORTANT:
# * The file must be named exactly "CODEOWNERS" (no extension) and live at the
# repo root, in .github/, or in docs/. Recommended: .github/CODEOWNERS
# * Every owner listed must be a repository collaborator with write access,
# otherwise GitHub silently ignores that entry.
# * Later matches override earlier ones.
#
# Syntax: <path-pattern> @owner [@owner ...]

# ---------------------------------------------------------------------------
# Default owner for everything
# ---------------------------------------------------------------------------
* @kelkalot

# ---------------------------------------------------------------------------
# Co-maintainers — uncomment and verify each GitHub handle + write access.
# (Add maintainers from GOVERNANCE.md once their handles are confirmed.)
# ---------------------------------------------------------------------------
# * @kelkalot @sushantgautam

# ---------------------------------------------------------------------------
# Domain-specific ownership — fill in handles, then uncomment.
# These route reviews to the people responsible for sensitive areas.
# ---------------------------------------------------------------------------

# Core auditing / judge engine (methodology-sensitive)
# /simpleaudit/ @kelkalot

# Health / clinical scenarios & judges (Norwegian Directorate of Health advisors)
# /simpleaudit/scenarios/health.py @kelkalot # + <health-advisor>
# /simpleaudit/scenarios/helpmed.py @kelkalot # + <health-advisor>
# /simpleaudit/scenarios/bullshitbench_health.py @kelkalot
# /simpleaudit/judges/helsedir_sexhealth_no.py @kelkalot
# /simpleaudit/judges/helsedir_sexhealth_no_rag.py @kelkalot

# Youth-domain scenario packs (provenance-sensitive)
# /simpleaudit/scenarios/ung.py @kelkalot
# /simpleaudit/scenarios/hei_refusal.py @kelkalot

# Norwegian public-sector packs
# /simpleaudit/scenarios/nav_aap.py @kelkalot
# /simpleaudit/scenarios/skatteetaten.py @kelkalot

# Governance, security & policy docs
# /GOVERNANCE.md @kelkalot
# /SECURITY.md @kelkalot
# /CODE_OF_CONDUCT.md @kelkalot
# /DPG.md @kelkalot
23 changes: 17 additions & 6 deletions FAQ.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,10 +6,21 @@
This usually happens if you are using a proxy (e.g., Cloudflare or a corporate firewall) and it blocks requests that look like automated agents.

#### Solution:
Override the default OpenAI request headers to include a valid `User-Agent` in ModelAuditor:
`ModelAuditor` creates its clients through [any-llm](https://mozilla-ai.github.io/any-llm/) and does not expose per-request HTTP header overrides (there are no `extra_kwargs` / `judge_extra_kwargs` parameters). Options that work today:

```python
extra_kwargs={"default_headers": {"User-Agent": "SimpleAudit-test/1.0"}} # for model calls
# or
judge_extra_kwargs={"default_headers": {"User-Agent": "SimpleAudit-test/1.0"}} # for judge calls
```
- Ask your proxy/firewall administrator to allowlist the endpoint or your machine's traffic.
- Point `base_url` (target), `judge_base_url` (judge), or `auditor_base_url` (auditor) at a local gateway or reverse proxy that injects the headers your infrastructure requires (e.g. nginx with `proxy_set_header User-Agent "SimpleAudit-test/1.0";`).
- Switch `provider` / `judge_provider` to a provider whose requests your proxy accepts — any provider supported by any-llm works.

If you need native header overrides, please [open an issue](https://github.com/kelkalot/simpleaudit/issues).

## Customization

### What extension points does `ModelAuditor` offer?

- `probe_prompt` — replace the built-in red-team persona used to generate probes. Include a literal `{language}` placeholder to opt into the `language` parameter (it is substituted verbatim, so JSON braces elsewhere in the prompt are untouched).
- `judge_prompt` — replace the judge's system prompt, including your own output schema; the framework returns whatever JSON the judge produces.
- `judge_response_schema` — supply a custom JSON schema for judge output enforcement (named judge configs with non-default shapes declare their own automatically).
- `judge` — select a named judge config (`safety`, `abstention`, `helpfulness`, `factuality`, `harm`, `binary_abstention`, ...); explicit `probe_prompt` / `judge_prompt` / `judge_response_schema` always override the config's values.
- `base_url` / `judge_base_url` / `auditor_base_url` — point the target, judge, or auditor at custom OpenAI-compatible endpoints (vLLM, LM Studio, gateways).
- `provider` / `judge_provider` / `auditor_provider` — use any provider supported by any-llm, independently per role.
106 changes: 106 additions & 0 deletions GOVERNANCE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,106 @@
# SimpleAudit Governance

This document describes how the SimpleAudit project is governed — who maintains
it, how decisions are made, and how to participate. It complements the
[Code of Conduct](./CODE_OF_CONDUCT.md), the [Security Policy](./SECURITY.md),
and the contribution guidance in the [README](./README.md).

SimpleAudit is an open-source (MIT) AI safety auditing framework and a
[verified Digital Public Good](https://www.digitalpublicgoods.net/r/simpleaudit).
It is stewarded by Simula Research Laboratory (Simula) and SimulaMet, and
developed in collaboration with the Norwegian Directorate of Health.

## Governance model

SimpleAudit follows a **maintainer-led model with institutional stewardship**.
Day-to-day direction rests with a small group of maintainers drawn from the
stewarding institutions; anyone may contribute. There is currently no separate
board or steering committee — decisions are made by the maintainers, in the
open, on GitHub. This document is the source of truth for that model and is
versioned with the project.

## Roles and responsibilities

### Lead maintainer / steward
Sets overall direction, has the final say where consensus cannot be reached,
and is the primary point of contact.

- **Michael A. Riegler** (Simula) — `@kelkalot` · michael@simula.no

### Maintainers
Review and merge contributions, triage issues, cut releases, and uphold the
Code of Conduct. Maintainers have write access to the repository.

- Sushant Gautam (SimulaMet)
- Finn Schwall (Simula)
- Annika Willoch Olstad (Simula)
- Klas H. Pettersen (SimulaMet)
- Sunniva Bjørklund (Hdir)

### Domain advisors
Provide expert review for specialised content — clinical/health scenarios,
Norwegian public-sector and youth domains, and safety policy — advising on
correctness and appropriateness within their domain.

- Sunniva Bjørklund, Maja Gran Erke, Hilde Lovett (Norwegian Directorate of
Health) — health / clinical content
- Tor-Ståle Hansen (Ministry of Defence, Norway) — safety / public-sector

### Contributors
Anyone who opens an issue or pull request. Contributors do not need write
access; their changes are merged after maintainer review.

## Decision-making

- **Routine changes** (bug fixes, documentation, new scenario packs or judge
configs, dependency bumps): handled through pull requests. A PR may be merged
once it has at least one approving review from a maintainer (or the relevant
code owner) and CI passes. Small, low-risk changes may use *lazy consensus* —
if no maintainer objects within a reasonable window, the change proceeds.
- **Significant changes** (breaking API changes, methodology changes, new
external dependencies, or anything affecting data provenance or safety
posture): proposed and discussed in a GitHub Issue or PR before merging, so
maintainers and affected domain advisors can weigh in. The goal is consensus
among active maintainers.
- **Tie-breaking**: where maintainers cannot reach consensus, the lead
maintainer makes the final decision, recorded in the relevant issue or PR.
- **Transparency**: substantive decisions are made and recorded in public
issues and pull requests.

## Contribution & review process

1. For non-trivial changes, open an issue to discuss first (optional for small
fixes).
2. Submit a pull request and ensure CI (tests) passes.
3. Reviews are requested automatically from the relevant owners via
[`CODEOWNERS`](./.github/CODEOWNERS); at least one maintainer / code-owner
approval is required to merge.
4. Health/clinical, Norwegian-domain, and other safety-sensitive changes should
additionally be reviewed by a relevant domain advisor.

See the README "Contributing" section for current areas of interest.

## Becoming a maintainer

Contributors who have made sustained, high-quality contributions and shown good
judgement may be invited to become maintainers by consensus of the existing
maintainers. Maintainers who become inactive may move to emeritus status, with
write access adjusted accordingly. Changes to the maintainer list are made via a
pull request to this document.

## Releases

SimpleAudit uses semantic versioning, is published to
[PyPI](https://pypi.org/project/simpleaudit/), and tags releases on GitHub. Any
maintainer may cut a release once `main` is green.

## Code of Conduct & Security

All participation is governed by the [Code of Conduct](./CODE_OF_CONDUCT.md).
Security vulnerabilities are handled per the [Security Policy](./SECURITY.md).

## Amending this document

Changes to governance are proposed via pull request and require approval from
the lead maintainer plus at least one other maintainer. The change history is
tracked in version control.
Loading
Loading