Add reproducible COCO accuracy gates for Vision v8 checkpoints - #21
Conversation
Complexity-ML
left a comment
There was a problem hiding this comment.
Request changes before merge.
Blocking findings:
- The release dataset is not actually pinned:
annotations_sha256andimage_list_sha256are bothnull, so the checker accepts any non-empty self-reported hashes instead of matching canonical COCO val2017 hashes. - The gate does not require the complete configured branch set. A report containing only
o2m-nmspasses even thoughnms-freeis configured and the documented contract requires both branches. The ONNX workflow also mapsbothtoauto, evaluating only one sidecar branch. - Non-finite metrics bypass the gate. A report using
NaNfor gated metrics passes both threshold and repeated-run checks because no numeric/finiteness validation is performed. - Provider-specific determinism tolerances are configured but ignored; the checker always uses
metric_tolerance.
The gate should also validate release-critical metadata such as checkpoint/model hash, backend, seed, protocol.release_eligible, and requested versus actual ONNX provider.
Local validation: Ruff passed; 19 gate/release tests passed; 18 additional ONNX tests passed. Counter-tests confirmed that both a NaN report and a report missing the configured nms-free branch currently pass. The full COCO evaluation remains skipped.
Please add negative regression tests for pinned hashes, required branches, non-finite/out-of-range metrics, release metadata, and backend-specific tolerances.
|
Thanks, fixed in the latest push. Changes included:
Local validation:
|
Complexity-ML
left a comment
There was a problem hiding this comment.
The follow-up commit fixes the four previously reported gate issues, and the targeted lint/tests pass. Two release-path blockers remain before merge:
-
The native PyTorch evaluator requires a checkpoint directory (
provenance.json,config.json, and weights), but_checkpoint_sha256()only hashes regular files. It therefore writescheckpoint_sha256: null, which the gate always rejects. Please hash the actual loadedema.safetensorsormodel.safetensorsfile (or a deterministic checkpoint manifest) and add a directory-checkpoint regression test. -
The workflow exposes
o2m-nmsandnms-freeas valid single-branch dispatch choices, while the checked-in gate configuration always requires both branches. Every single-branch dispatch therefore fails by construction. For the publication gate, please either requirebothor select an explicit matching policy for single-branch runs.
Local validation on bbac89a: Ruff passed and 31 targeted tests passed. GitHub fixture-gate, onnx-parity, and GitGuardian are green; full-coco-eval remains manual/skipped.
|
Thanks, fixed in the latest push. Changes:
Validation:
Full COCO eval remains manual/skipped. |
Complexity-ML
left a comment
There was a problem hiding this comment.
The two previously requested release-path fixes are correct: native directory checkpoints now hash the exact weights selected by the loader, and publication runs now require both branches. Local Ruff and all 32 targeted tests pass.
One blocking workflow syntax error was introduced, however. Lines 60-61 of .github/workflows/vision-v8-coco-accuracy.yml still contain the orphaned entries - o2m-nms and - nms-free beneath runner.default. This makes the YAML invalid, so GitHub does not register or run the fixture-gate job on this commit; only onnx-parity and GitGuardian are visible.
Please remove those two stale lines, push the correction, and confirm that fixture-gate returns and passes.
Issue: #18
Body
Adds a reproducible COCO accuracy-gating path for Vision v8 checkpoint publication.
This PR includes:
Validation run locally:
Note: