Add decoded-output parity gates for Vision v8 ONNX exports - #15
Merged
Complexity-ML merged 2 commits intoAug 27, 2026
Merged
Conversation
The v8 raw-logit thresholds were calibrated on five deterministic seeds. Each test compares about five million values, so the observed maximum is an extreme-value statistic that keeps growing with the seed count: it rises by 60% between 5 and 50 seeds on both branches. Both branches exceed their original threshold before 50 seeds, so check_onnx_parity.py --num-tests 20 failed on an otherwise healthy export: o2m max@5 1.741886e-03 max@50 2.788305e-03 threshold 2.0e-03 nms-free max@5 2.990723e-03 max@50 4.793882e-03 threshold 3.5e-03 Set both thresholds to twice the maximum observed over 50 seeds and record the measurement in the validation report. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Raw-logit tolerance only approximates deployment behaviour: softmax absorbs a translation shared by all DFL bins, while the stride amplifies drift on coarse grids, so raw and decoded drift are not monotonically related. check_onnx_parity.py now evaluates three independent gates on the same deterministic inputs: raw logits, normalized decoded boxes, and sigmoid quality-class scores. Each reports its own max and mean difference, fails on its own, and is named in the summary when it fails. Decoding goes through complexity.deploy.onnx_detector so the gates exercise the deployment code path instead of a separate implementation. Decoded thresholds are twice the maximum observed over 50 deterministic seeds. decoded-box is shared across branches (their box drift differs by 19%); decoded-score is branch-specific (o2m drift is 2.2x nms-free). Exports whose sidecar carries no v8 decode metadata skip the decoded gates rather than failing them, and keep the strict 1e-4 raw threshold. Closes Complexity-ML#12 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #12.
scripts/check_onnx_parity.pynow evaluates three independent gates on the samedeterministic inputs —
raw,decoded-box(normalized xyxy after LTRB/DFLdecode) and
decoded-score(sigmoid quality-class) — each with its own max/meanreport, its own pass/fail, and its name printed in the summary when it fails.
Decoding goes through the
complexity.deploy.onnx_detectorpackage merged in#14, so the gates exercise the deployment path rather than a parallel
implementation.
Why raw tolerance cannot stand in for decoded behaviour
Two effects pull in opposite directions. Softmax is translation-invariant, so a
drift shared by all 17 DFL bins of a distance moves the decoded box by exactly
zero. Conversely the decoded distance is multiplied by the stride, so a small
drift on the 20x20 grid moves a box by several pixels. Raw and decoded drift are
therefore not monotonically related, and each gate has to stand alone.
Please read this part before the diff: the raw thresholds are loosened
The first commit raises the v8 raw thresholds (o2m
2e-3→6e-3, nms-free3.5e-3→1e-2). That is not to make anything in this PR pass — nothing heretouches raw drift, which still measures
1.74e-03/2.99e-03at five seeds asin #10.
The reason is that calibrating on five seeds underestimated the maximum. Each
test compares ~5M values, so the observed max is an extreme-value statistic that
keeps growing with the seed count:
1.741886e-032.788305e-03+60.1%2.0e-032.990723e-034.793882e-03+60.3%3.5e-03Both branches exceed their threshold well before 50 seeds, so
check_onnx_parity.py --num-tests 20fails onmaintoday on a healthy export.Reproduce with
--num-tests 5then--num-tests 50and compare the reportedworst_max.Decoded outputs grow only
+12%to+34%over the same range — three to fivetimes more stable — which is itself the quantitative argument for gating on
them. All thresholds now follow one rule: twice the maximum observed over 50
deterministic seeds.
2.788305e-036.0e-032.2x5.497933e-051.3e-042.4x3.811717e-058.0e-052.1x4.793882e-031.0e-022.1x6.532669e-051.3e-042.0x1.719594e-054.0e-052.3xdecoded-boxis shared across branches (box drift differs by 19%, sameregression head);
decoded-scoreis branch-specific (o2m drift is 2.2xnms-free).
One correction to the report
The "decoded boxes max diff" figures in the validation report were in normalized
coordinates, which the report did not state. The unit is now explicit, and the
values are recomputed with the shared decoder, which is why they shift slightly.
Legacy behaviour
Sidecars without v8 decode metadata skip the decoded gates rather than failing
them, and keep the strict
1e-4raw threshold.--skip-decodedforces that onany export.
Tests
tests/test_detector_export.pycovers both branches, the legacy fallbacks, andthe amplification case #12 asks for: tilting the DFL logits by
epsilon * bindefeats softmax cancellation, so raw drift stays at
4e-03under a5e-03tolerance while normalized box drift reaches
5e-04against a1.3e-04one —and
decoded-scorestays green, demonstrating gate independence.Measured on
AETHORIA-AI/TR-HASH-Vision-v8-2M-COCO-SFTwith PyTorch2.13.0+cu130, ONNX1.21.0, ONNX Runtime1.24.4,CPUExecutionProvider.Raw maxima reproduce #10's to three significant digits despite the different
runtime versions.