This guide runs exported TR-Hash Vision v8 detector models with ONNX Runtime. The ONNX file returns raw predictions; the deployment pipeline performs the same preprocessing, DFL decode, confidence filtering, and branch-specific postprocessing documented in the contract.
Install the framework with its ONNX export/runtime dependencies:
python -m pip install -e ".[export]"CUDA and TensorRT execution providers still require the matching NVIDIA runtime libraries for the installed ONNX Runtime build.
Each deployed model needs two files:
- the
.onnxmodel exported byscripts/export_onnx.py; - the JSON sidecar written next to it by the exporter.
Do not rename one without passing both paths to the CLI.
python scripts/onnx_detect.py \
--model tr_hash_v8_o2m.onnx \
--metadata tr_hash_v8_o2m.json \
--image sample.jpg \
--provider cpu \
--prettyThe output is JSON:
{
"provider_used": "CPUExecutionProvider",
"branch_type": "o2m",
"timing": {
"preprocess_ms": 0.0,
"inference_ms": 0.0,
"postprocess_ms": 0.0
},
"detections": []
}Actual timings and detections depend on the image and hardware.
python scripts/onnx_detect.py \
--model tr_hash_v8_nms_free.onnx \
--metadata tr_hash_v8_nms_free.json \
--image sample.jpg \
--provider cuda \
--pretty--provider cuda requests CUDAExecutionProvider first and falls back to
CPUExecutionProvider. The JSON provider_used field reports what ONNX Runtime
actually selected after session creation.
TensorRT can be requested with:
python scripts/onnx_detect.py --model model.onnx --metadata model.json --image sample.jpg --provider tensorrtThat expands to TensorRT, CUDA, then CPU fallback.
o2mexports run confidence filtering followed by class-aware NMS.nms-freeexports run confidence filtering and top-k score selection only.--iou-thresholdonly affectso2m; the CLI warns if it is passed for an NMS-free export.--conf-thresholdoverrides the default confidence threshold for either branch.
The default thresholds match the PyTorch detector path:
- confidence threshold:
0.25; - O2M IoU threshold:
0.45; - max detections:
300.
Each detection contains:
box_norm: normalizedxyxyrelative to the square model input;box_pixel: restoredxyxyin original source-image pixels;class_id: integer class index;score: sigmoid quality-class score.
See the deployment contract for the exact tensor layout, grid mapping, DFL decode formula, preprocessing, and sidecar validation rules.