Skip to content

Report groundedness output as ungraded off the derivation path - #98

Open
avalyset wants to merge 1 commit into
SimulaMet:devfrom
avalyset:fix/groundedness-needs-derivation
Open

avalyset wants to merge 1 commit into
SimulaMet:devfrom
avalyset:fix/groundedness-needs-derivation

Conversation

@avalyset

@avalyset avalyset commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

What happens now

SingleTurnAuditor derives the groundedness severity and the findings from the
document marks after the judge call, so that call passes no postprocess. On
every other path the judgment arrives with neither severity nor score,
_severity_from_judgment falls through to its "medium" default, and nothing
says why.

auditor = ModelAuditor(target_model=..., judge="groundedness")
result = auditor.run_scenario(scenario)   # documents present, marks parsed nowhere
result.severity                           # "medium"
result.issues_found                       # []

An answer grounded in the current document and one quoting the superseded one
get the same "medium". That is the finding context_grounding exists to
catch, and on this path it is a middling pass with nothing behind it.

The other paths are ModelAuditor(judge="groundedness"), rejudge(), and a
PromptVariant built from the config.

Why the hook

The config's postprocess hook is reached exactly where the generic machinery
grades the output, and never on the derivation path: single_turn.py passes
no postprocess for this judge. The hook is not given the marks either (it gets
conversation, expected_behavior and scenario_meta), so it cannot do the
derivation. What it can do is say that nothing will.

It reports ERROR and names the reason, following binary_postprocess: an
existing ERROR judgment passes through untouched, and the three observation
fields are kept as they came. normalize_severity("ERROR") falls outside
SEVERITY_ORDER, so the run drops out of the severity statistics instead of
counting as a pass.

The alternatives, and why they were dropped. A requires_expected_behavior-style
config flag needs a new attribute, a new warn method and a new call site, and
its built-in fallback to the default judge is the wrong outcome: silently
swapping the judge is worse than refusing to grade. Raising in get_judge
would hit SingleTurnAuditor too. Raising inside the hook is caught upstream
and becomes ERROR anyway, with a ValueError string in place of a readable
reason.

PromptVariant.from_judge("groundedness").postprocess is the hook, so
rejudge() inherits the guard. That path has no marks either.

Testing

Four tests are in tests/test_single_turn_correctness.py, which already owns
the question of where the groundedness severity comes from.

  • test_the_generic_judge_path_reports_error_instead_of_an_underived_medium
    fails on the pre-fix code with assert 'medium' == 'ERROR'.
  • test_single_turn_severity_is_the_derived_one_not_the_judges_own passes
    before and after. It pins down that the hook never fires on the derivation
    path, so adding a postprocess to that call would fail it.
  • Two more cover an existing ERROR judgment and an off-contract
    issues_found.

1352 passed, 19 skipped on Python 3.13.15, against a baseline of 1348 on
e5ec692 in the same environment; the failure set is identical and empty. Run
on their own, the three single-turn and groundedness files give 136 passed.

Not run: anything against a live model, and Python 3.11 or 3.12.

The groundedness judge emits observations and no severity; the findings and
the severity are derived from the document marks. Only SingleTurnAuditor has
the parsed marks, and it derives them after the judging call, so its judge
call passes no postprocess.

Everywhere else — ModelAuditor(judge="groundedness"), rejudge(), a reframing
variant built from the config — the judgment arrived with neither `severity`
nor `score`, and _severity_from_judgment fell through to its "medium" default.
A grounded answer and an answer quoting the superseded document both came back
medium, with nothing behind the number and no error to say why.

The config's postprocess hook is reached on exactly those paths and never on
the derivation path, and the hook is not given the marks either, so it cannot
do the work. It now reports ERROR and names the reason. ERROR is off the
severity ladder, so such a run is excluded from severity statistics rather
than counted as a middling pass.

Four tests. The generic-path one fails on the pre-fix code with
`assert 'medium' == 'ERROR'`; the SingleTurnAuditor one pins that the hook
never fires there, so adding a postprocess to that call would fail it.
@avalyset
avalyset requested a review from kelkalot as a code owner October 2, 2026 16:01
@avalyset avalyset mentioned this pull request Oct 2, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant