Skip to content

Add decoded-output parity gates for Vision v8 ONNX exports - #15

Merged
Complexity-ML merged 2 commits into
Complexity-ML:mainfrom
IIIllllIlIlllII:feat/onnx-decoded-parity
Aug 27, 2026
Merged

Complexity-ML merged 2 commits into
Complexity-ML:mainfrom
IIIllllIlIlllII:feat/onnx-decoded-parity

Conversation

@IIIllllIlIlllII

Copy link
Copy Markdown
Contributor

Closes #12.

scripts/check_onnx_parity.py now evaluates three independent gates on the same
deterministic inputs — raw, decoded-box (normalized xyxy after LTRB/DFL
decode) and decoded-score (sigmoid quality-class) — each with its own max/mean
report, its own pass/fail, and its name printed in the summary when it fails.
Decoding goes through the complexity.deploy.onnx_detector package merged in
#14, so the gates exercise the deployment path rather than a parallel
implementation.

Why raw tolerance cannot stand in for decoded behaviour

Two effects pull in opposite directions. Softmax is translation-invariant, so a
drift shared by all 17 DFL bins of a distance moves the decoded box by exactly
zero. Conversely the decoded distance is multiplied by the stride, so a small
drift on the 20x20 grid moves a box by several pixels. Raw and decoded drift are
therefore not monotonically related, and each gate has to stand alone.

Please read this part before the diff: the raw thresholds are loosened

The first commit raises the v8 raw thresholds (o2m 2e-3 → 6e-3, nms-free
3.5e-3 → 1e-2). That is not to make anything in this PR pass — nothing here
touches raw drift, which still measures 1.74e-03 / 2.99e-03 at five seeds as
in #10.

The reason is that calibrating on five seeds underestimated the maximum. Each
test compares ~5M values, so the observed max is an extreme-value statistic that
keeps growing with the seed count:

Branch max@5 max@50 growth old threshold
o2m 1.741886e-03 2.788305e-03 +60.1% 2.0e-03
nms-free 2.990723e-03 4.793882e-03 +60.3% 3.5e-03

Both branches exceed their threshold well before 50 seeds, so
check_onnx_parity.py --num-tests 20 fails on main today on a healthy export.
Reproduce with --num-tests 5 then --num-tests 50 and compare the reported
worst_max.

Decoded outputs grow only +12% to +34% over the same range — three to five
times more stable — which is itself the quantitative argument for gating on
them. All thresholds now follow one rule: twice the maximum observed over 50
deterministic seeds.

Branch Gate max@50 threshold headroom
o2m raw 2.788305e-03 6.0e-03 2.2x
o2m decoded-box 5.497933e-05 1.3e-04 2.4x
o2m decoded-score 3.811717e-05 8.0e-05 2.1x
nms-free raw 4.793882e-03 1.0e-02 2.1x
nms-free decoded-box 6.532669e-05 1.3e-04 2.0x
nms-free decoded-score 1.719594e-05 4.0e-05 2.3x

decoded-box is shared across branches (box drift differs by 19%, same
regression head); decoded-score is branch-specific (o2m drift is 2.2x
nms-free).

One correction to the report

The "decoded boxes max diff" figures in the validation report were in normalized
coordinates, which the report did not state. The unit is now explicit, and the
values are recomputed with the shared decoder, which is why they shift slightly.

Legacy behaviour

Sidecars without v8 decode metadata skip the decoded gates rather than failing
them, and keep the strict 1e-4 raw threshold. --skip-decoded forces that on
any export.

Tests

tests/test_detector_export.py covers both branches, the legacy fallbacks, and
the amplification case #12 asks for: tilting the DFL logits by epsilon * bin
defeats softmax cancellation, so raw drift stays at 4e-03 under a 5e-03
tolerance while normalized box drift reaches 5e-04 against a 1.3e-04 one —
and decoded-score stays green, demonstrating gate independence.

Measured on AETHORIA-AI/TR-HASH-Vision-v8-2M-COCO-SFT with PyTorch
2.13.0+cu130, ONNX 1.21.0, ONNX Runtime 1.24.4, CPUExecutionProvider.
Raw maxima reproduce #10's to three significant digits despite the different
runtime versions.

IIIllllIlIlllII and others added 2 commits August 27, 2026 13:45
The v8 raw-logit thresholds were calibrated on five deterministic seeds. Each
test compares about five million values, so the observed maximum is an
extreme-value statistic that keeps growing with the seed count: it rises by 60%
between 5 and 50 seeds on both branches.

Both branches exceed their original threshold before 50 seeds, so
check_onnx_parity.py --num-tests 20 failed on an otherwise healthy export:

  o2m       max@5 1.741886e-03  max@50 2.788305e-03  threshold 2.0e-03
  nms-free  max@5 2.990723e-03  max@50 4.793882e-03  threshold 3.5e-03

Set both thresholds to twice the maximum observed over 50 seeds and record the
measurement in the validation report.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Raw-logit tolerance only approximates deployment behaviour: softmax absorbs a
translation shared by all DFL bins, while the stride amplifies drift on coarse
grids, so raw and decoded drift are not monotonically related.

check_onnx_parity.py now evaluates three independent gates on the same
deterministic inputs: raw logits, normalized decoded boxes, and sigmoid
quality-class scores. Each reports its own max and mean difference, fails on its
own, and is named in the summary when it fails. Decoding goes through
complexity.deploy.onnx_detector so the gates exercise the deployment code path
instead of a separate implementation.

Decoded thresholds are twice the maximum observed over 50 deterministic seeds.
decoded-box is shared across branches (their box drift differs by 19%);
decoded-score is branch-specific (o2m drift is 2.2x nms-free).

Exports whose sidecar carries no v8 decode metadata skip the decoded gates
rather than failing them, and keep the strict 1e-4 raw threshold.

Closes Complexity-ML#12

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@Complexity-ML
Complexity-ML merged commit 3ac54d5 into Complexity-ML:main Aug 27, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add decoded-output parity gates for Vision v8 ONNX exports

2 participants