Context
Vision v8 ONNX exports expose raw tensors shaped [batch, 34000, 148]: 68 LTRB/DFL regression logits followed by 80 quality-class logits.
Consumers currently need to implement the deployment contract themselves:
- O2M: decode plus NMS.
- NMS-free: decode plus confidence filtering.
A shared reference implementation would reduce integration errors and make end-to-end benchmarking meaningful.
Scope
- Add a small framework module and CLI example that performs preprocessing, ONNX Runtime inference, decode, and post-processing.
- Support both exported branches from their metadata sidecars.
- Implement DFL box decoding once and share it between branches.
- Apply configurable confidence filtering to both branches.
- Apply class-aware NMS only for O2M.
- Return a stable detection schema with normalized and pixel-space boxes, class IDs, and scores.
- Measure model latency and post-processing latency separately.
Acceptance criteria
- Synthetic unit tests cover grid mapping, DFL decode, confidence filtering, and O2M NMS.
- End-to-end outputs are compared against the PyTorch detector post-processing on fixed inputs.
- NMS-free execution never invokes NMS.
- The CLI accepts provider selection and reports the provider actually used.
- Documentation includes runnable CPU and CUDA examples.
Related: #10
Context
Vision v8 ONNX exports expose raw tensors shaped
[batch, 34000, 148]: 68 LTRB/DFL regression logits followed by 80 quality-class logits.Consumers currently need to implement the deployment contract themselves:
A shared reference implementation would reduce integration errors and make end-to-end benchmarking meaningful.
Scope
Acceptance criteria
Related: #10