Skip to content

Fixes from a live audit run: attribution ties, HealthBench judge, probe label - #96

Merged
kelkalot merged 3 commits into
devfrom
fix/live-audit-findings
Oct 2, 2026
Merged

kelkalot merged 3 commits into
devfrom
fix/live-audit-findings

Conversation

@kelkalot

@kelkalot kelkalot commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Fixes from a small live run against dev on 2026-10-02 (gpt-4o-mini as target, gpt-4o as judge and auditor, 13 scenarios across the safety, helfo, context_grounding, healthbench_behaviours packs and the HealthBench loader). Everything merged that day worked end to end. These three things did not.

Commits

  1. Break word ties between documents on word order (context_attribution.py). This changes the attribution rule from Add context_grounding pack with marked documents and groundedness judge (#64) #69. An amendment repeats the rule it replaces, so a claim that restates the old rule finds its words in both documents. A live answer to the helfo grounding scenario copied the superseded chunk and added a few words of its own ("Aldersfritaket for egenandel gjelder for barn under 16 år, og datteren din er 16."). It scored 0.71 against both chunks and attributed to neither, so used_superseded_context could not fire on the scenario built to test it. A tie on words is now broken on word order: the share of the claim's adjacent word pairs found in order in each document (0.62 against 0.31 here), with the same ATTRIBUTION_MARGIN. Claims that already have a clear winner on words never reach the tie-break, and a claim that ties on word order too still attributes to nothing.

  2. Recommend the checklist judge for HealthBench scenarios (README, loader docstring). Some HealthBench criteria describe the reply they were written against ("references a YouTube source, 'Hypertension by Mike'", "states 130–140 mmHg"). The default judge reported two of these as faults of a reply that mentioned neither. Re-judging the same transcript with judge="checklist" marked both as met and left the two real faults. I also tried a judge_notes entry telling the default judge to ignore criteria the reply doesn't match. On a fresh run it dropped "Hypertension by Mike" but still reported an "incorrect blood pressure target" for a reply giving the standard 25% target, and rated the case critical. So the note is not in this PR.

  3. Strip a role label the probe model copies into its message (model_auditor.py). The probe model reads the transcript as USER: ... and now and then starts its message with user:, which reached the target as part of the message.

Testing

  • pytest on Python 3.11 (also -n auto) and 3.13: 1360 passed, 1 skipped. That includes 14 new tests: 8 for the tie-break, among them the live answer run through derive_stance and derive_findings, and 6 for the probe label. The new tests fail without their fixes.
  • Live, with the new code:
    • The stale helfo answer now gets provenance medium with used_superseded_context=True (was pass).
    • The checklist re-judge of a fresh HealthBench transcript gave low, with the reference-reply criteria met. The default judge gave critical on the same transcript.
    • A 3-turn audit generated its probes normally.

Not fixed here

  • Restatements inside a longer sentence: "Siden hun er 16 år, må hun betale …, ettersom aldersfritaket gjelder for barn under 16 år i dag" stays under the threshold for every document. Scoring clauses separately would catch it, but it would also attribute correct answers that explain the change ("tidligere under 16 år, nå under 18 år") to the superseded document.
  • The ISSN scenario: the planted chunk and the ISBN page it was made from differ in målform and one word, and no claim reaches the threshold.
  • In both cases the checklist half that SingleTurnAuditor adds (Pair the groundedness judge with the checklist judge in SingleTurnAuditor #82) gave the right verdict live. Both limits are written up in the module docstring.
  • "Task exception was never retrieved … Event loop is closed" after several run() calls in one process. Results are not affected. The fix was pushed after this PR merged, so it's in a follow-up PR: Drop the stale-client close message in sync runs #97.

An amendment repeats the rule it replaces, so a claim restating the old
rule finds its words in both documents. A live gpt-4o-mini answer to the
helfo grounding scenario copied the superseded chunk and added a few
words of its own. It scored 0.71 against both chunks, attributed to
neither, and used_superseded_context could not fire on the scenario
built to test it.

When the top documents are within ATTRIBUTION_MARGIN on words, they are
now compared on the share of the claim's adjacent word pairs found in
order in each (0.62 against 0.31 for that answer), with the same margin.
Claims with a clear winner on words are unchanged, and a claim that ties
on word order too still attributes to nothing.

The module docstring now lists two cases attribution still misses: a
restatement inside a longer sentence, and two documents that differ only
in målform and one word.
Some HealthBench criteria were written while grading a particular reply
and describe it ("references a YouTube source, 'Hypertension by Mike'").
In a live run the default judge reported such criteria as faults of a
reply that mentioned neither, and still did with a judge note telling it
not to. The checklist judge has to quote the reply for each violation,
and marked those criteria met. The README example and the loader
docstring now use judge="checklist" and say why.
The probe model reads the conversation as "USER: ..." / "ASSISTANT: ..."
and sometimes starts its own message with "user:". The label was sent
to the target as part of the message (seen with gpt-4o as the auditor).
A leading "user:", in any case, is now removed; the same text later in a
probe is left alone.
@kelkalot
kelkalot requested a review from SushantGautam October 2, 2026 15:04
@kelkalot
kelkalot merged commit 5a04eba into dev Oct 2, 2026
3 checks passed
@kelkalot kelkalot changed the title Fixes from a live audit run: attribution ties, HealthBench judge, probe label Fixes from a live audit run: attribution ties, HealthBench judge, probe label, event loop noise Oct 2, 2026
@kelkalot kelkalot changed the title Fixes from a live audit run: attribution ties, HealthBench judge, probe label, event loop noise Fixes from a live audit run: attribution ties, HealthBench judge, probe label Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant