Skip to content

Add reproducible COCO accuracy gates for Vision v8 checkpoints - #21

Merged
Complexity-ML merged 4 commits into
Complexity-ML:mainfrom
milliyin:vision-v8-coco-accuracy-gates
Sep 1, 2026
Merged

Complexity-ML merged 4 commits into
Complexity-ML:mainfrom
milliyin:vision-v8-coco-accuracy-gates

Conversation

@milliyin

Copy link
Copy Markdown
Contributor

Issue: #18

Body

Adds a reproducible COCO accuracy-gating path for Vision v8 checkpoint publication.

This PR includes:

  • deterministic COCO evaluation reporting for native PyTorch
  • ONNX Runtime COCO evaluation using the shared deployment preprocessing/decode/postprocess pipeline
  • dataset pinning via annotation SHA-256 and sorted image-list SHA-256
  • runtime metadata in reports, including framework commit, checkpoint/model hash, Python, OS, PyTorch, ONNX Runtime, CUDA, and TensorRT info
  • JSON and Markdown reports suitable for release/model-card artifacts
  • configurable gate checks for malformed reports, absolute metric floors, baseline regressions, and repeated-run determinism
  • a GitHub Actions workflow with lightweight PR fixture tests and an explicit full-COCO manual evaluation job
  • optional upload of gated reports to an existing GitHub Release

Validation run locally:

python -m ruff check complexity\deploy\onnx_detector\pipeline.py complexity\deploy\onnx_detector\session.py scripts\check_vision_v8_coco_report.py scripts\evaluate_tr_hash_coco.py scripts\evaluate_onnx_coco.py tests\test_coco_release_evaluation.py tests\test_onnx_detector_core.py tests\test_vision_v8_coco_accuracy_gate.py

python -m pytest -q tests\test_coco_release_evaluation.py tests\test_onnx_detector_core.py tests\test_vision_v8_coco_accuracy_gate.py

python -m pytest -q tests\test_onnx_detector_skeleton.py tests\test_onnx_detector_core.py tests\test_onnx_detector_metadata.py tests\test_onnx_detector_pipeline.py tests\test_onnx_detect_cli.py

Note:

  • Full COCO val2017 evaluation still needs to be run on a runner with the pinned dataset and checkpoint/artifacts available.

@Complexity-ML Complexity-ML left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Request changes before merge.

Blocking findings:

  1. The release dataset is not actually pinned: annotations_sha256 and image_list_sha256 are both null, so the checker accepts any non-empty self-reported hashes instead of matching canonical COCO val2017 hashes.
  2. The gate does not require the complete configured branch set. A report containing only o2m-nms passes even though nms-free is configured and the documented contract requires both branches. The ONNX workflow also maps both to auto, evaluating only one sidecar branch.
  3. Non-finite metrics bypass the gate. A report using NaN for gated metrics passes both threshold and repeated-run checks because no numeric/finiteness validation is performed.
  4. Provider-specific determinism tolerances are configured but ignored; the checker always uses metric_tolerance.

The gate should also validate release-critical metadata such as checkpoint/model hash, backend, seed, protocol.release_eligible, and requested versus actual ONNX provider.

Local validation: Ruff passed; 19 gate/release tests passed; 18 additional ONNX tests passed. Counter-tests confirmed that both a NaN report and a report missing the configured nms-free branch currently pass. The full COCO evaluation remains skipped.

Please add negative regression tests for pinned hashes, required branches, non-finite/out-of-range metrics, release metadata, and backend-specific tolerances.

@milliyin

milliyin commented Sep 1, 2026

Copy link
Copy Markdown
Contributor Author

Thanks, fixed in the latest push.

Changes included:

  • pinned canonical COCO val2017 annotation and image-list hashes;
  • required both configured branches in release reports;
  • rejected non-finite and out-of-range gated metrics;
  • enforced release metadata: backend, seed, release_eligible, checkpoint/model hashes, and ONNX requested vs actual provider;
  • applied provider-specific determinism tolerances;
  • fixed the ONNX workflow so branch=both evaluates O2M and NMS-free sidecars separately, merges them, then gates the combined report;
  • added negative regression tests for the reported cases.

Local validation:

  • python -m pytest -q tests\test_vision_v8_coco_accuracy_gate.py tests\test_coco_release_evaluation.py
    24 passed
  • ruff check scripts/check_vision_v8_coco_report.py scripts/merge_vision_v8_coco_reports.py tests/test_vision_v8_coco_accuracy_gate.py tests/test_coco_release_evaluation.py
    All checks passed

@Complexity-ML Complexity-ML left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The follow-up commit fixes the four previously reported gate issues, and the targeted lint/tests pass. Two release-path blockers remain before merge:

  1. The native PyTorch evaluator requires a checkpoint directory (provenance.json, config.json, and weights), but _checkpoint_sha256() only hashes regular files. It therefore writes checkpoint_sha256: null, which the gate always rejects. Please hash the actual loaded ema.safetensors or model.safetensors file (or a deterministic checkpoint manifest) and add a directory-checkpoint regression test.

  2. The workflow exposes o2m-nms and nms-free as valid single-branch dispatch choices, while the checked-in gate configuration always requires both branches. Every single-branch dispatch therefore fails by construction. For the publication gate, please either require both or select an explicit matching policy for single-branch runs.

Local validation on bbac89a: Ruff passed and 31 targeted tests passed. GitHub fixture-gate, onnx-parity, and GitGuardian are green; full-coco-eval remains manual/skipped.

@milliyin

milliyin commented Sep 1, 2026

Copy link
Copy Markdown
Contributor Author

Thanks, fixed in the latest push.

Changes:

  • directory checkpoints now hash the loaded weights file: ema.safetensors, then model.safetensors;
  • release workflow now always evaluates both branches;
  • added regression coverage for directory checkpoint hashing;
  • updated docs.

Validation:

  • 25 gate tests passed
  • 32 workflow-covered tests passed
  • Ruff passed

Full COCO eval remains manual/skipped.

@Complexity-ML Complexity-ML left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The two previously requested release-path fixes are correct: native directory checkpoints now hash the exact weights selected by the loader, and publication runs now require both branches. Local Ruff and all 32 targeted tests pass.

One blocking workflow syntax error was introduced, however. Lines 60-61 of .github/workflows/vision-v8-coco-accuracy.yml still contain the orphaned entries - o2m-nms and - nms-free beneath runner.default. This makes the YAML invalid, so GitHub does not register or run the fixture-gate job on this commit; only onnx-parity and GitGuardian are visible.

Please remove those two stale lines, push the correction, and confirm that fixture-gate returns and passes.

@Complexity-ML
Complexity-ML merged commit 174dd08 into Complexity-ML:main Sep 1, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants