Conversation
The groundedness judge emits observations and no severity; the findings and the severity are derived from the document marks. Only SingleTurnAuditor has the parsed marks, and it derives them after the judging call, so its judge call passes no postprocess. Everywhere else — ModelAuditor(judge="groundedness"), rejudge(), a reframing variant built from the config — the judgment arrived with neither `severity` nor `score`, and _severity_from_judgment fell through to its "medium" default. A grounded answer and an answer quoting the superseded document both came back medium, with nothing behind the number and no error to say why. The config's postprocess hook is reached on exactly those paths and never on the derivation path, and the hook is not given the marks either, so it cannot do the work. It now reports ERROR and names the reason. ERROR is off the severity ladder, so such a run is excluded from severity statistics rather than counted as a middling pass. Four tests. The generic-path one fails on the pre-fix code with `assert 'medium' == 'ERROR'`; the SingleTurnAuditor one pins that the hook never fires there, so adding a postprocess to that call would fail it.
Merged
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What happens now
SingleTurnAuditorderives the groundedness severity and the findings from thedocument marks after the judge call, so that call passes no postprocess. On
every other path the judgment arrives with neither
severitynorscore,_severity_from_judgmentfalls through to its"medium"default, and nothingsays why.
An answer grounded in the current document and one quoting the superseded one
get the same
"medium". That is the findingcontext_groundingexists tocatch, and on this path it is a middling pass with nothing behind it.
The other paths are
ModelAuditor(judge="groundedness"),rejudge(), and aPromptVariantbuilt from the config.Why the hook
The config's
postprocesshook is reached exactly where the generic machinerygrades the output, and never on the derivation path:
single_turn.pypassesno postprocess for this judge. The hook is not given the marks either (it gets
conversation,expected_behaviorandscenario_meta), so it cannot do thederivation. What it can do is say that nothing will.
It reports
ERRORand names the reason, followingbinary_postprocess: anexisting
ERRORjudgment passes through untouched, and the three observationfields are kept as they came.
normalize_severity("ERROR")falls outsideSEVERITY_ORDER, so the run drops out of the severity statistics instead ofcounting as a pass.
The alternatives, and why they were dropped. A
requires_expected_behavior-styleconfig flag needs a new attribute, a new warn method and a new call site, and
its built-in fallback to the default judge is the wrong outcome: silently
swapping the judge is worse than refusing to grade. Raising in
get_judgewould hit
SingleTurnAuditortoo. Raising inside the hook is caught upstreamand becomes
ERRORanyway, with aValueErrorstring in place of a readablereason.
PromptVariant.from_judge("groundedness").postprocessis the hook, sorejudge()inherits the guard. That path has no marks either.Testing
Four tests are in
tests/test_single_turn_correctness.py, which already ownsthe question of where the groundedness severity comes from.
test_the_generic_judge_path_reports_error_instead_of_an_underived_mediumfails on the pre-fix code with
assert 'medium' == 'ERROR'.test_single_turn_severity_is_the_derived_one_not_the_judges_ownpassesbefore and after. It pins down that the hook never fires on the derivation
path, so adding a postprocess to that call would fail it.
ERRORjudgment and an off-contractissues_found.1352 passed, 19 skipped on Python 3.13.15, against a baseline of 1348 on
e5ec692in the same environment; the failure set is identical and empty. Runon their own, the three single-turn and groundedness files give 136 passed.
Not run: anything against a live model, and Python 3.11 or 3.12.