Context
PR #10 calibrated branch-specific raw-logit tolerances for Vision v8 because CPU ONNX Runtime drift is concentrated in fine-grid regression logits while decoded box and class-score drift remains much smaller.
Raw-logit tolerance is still only a proxy for deployment behavior. The validation report explicitly recommends checking decoded outputs.
Scope
- Extend
scripts/check_onnx_parity.py to compare decoded outputs in addition to raw logits.
- Decode LTRB/DFL regression using the exported model metadata and detector grid configuration.
- Compare normalized boxes and sigmoid quality-class scores for both
o2m and nms-free.
- Report raw-logit, decoded-box, and decoded-score max/mean differences separately.
- Keep deterministic inputs and make each gate fail independently.
- Add tests for both branches, including a case where raw drift is tolerated but decoded drift exceeds its limit.
Acceptance criteria
- Both v8 branches have documented decoded parity thresholds with numerical justification.
- The CLI output clearly identifies which parity gate failed.
- Existing legacy exports retain their current strict raw-logit behavior when decoded metadata is unavailable.
tests/test_detector_export.py covers the new behavior.
- The detector-export workflow passes on CPU ONNX Runtime.
Related: #10
Context
PR #10 calibrated branch-specific raw-logit tolerances for Vision v8 because CPU ONNX Runtime drift is concentrated in fine-grid regression logits while decoded box and class-score drift remains much smaller.
Raw-logit tolerance is still only a proxy for deployment behavior. The validation report explicitly recommends checking decoded outputs.
Scope
scripts/check_onnx_parity.pyto compare decoded outputs in addition to raw logits.o2mandnms-free.Acceptance criteria
tests/test_detector_export.pycovers the new behavior.Related: #10