Conversation
Opset 23 allows output_dtype to differ from x_scale, but the CPU kernel uses the scale type to access the output buffer. Dispatch on the output tensor type and keep scale and output separately typed through per-tensor, per-axis, and blocked dequantization. Related to microsoft#32719. QuantizeLinear and CUDA are separate work. Authored with Codex assistance. Signed-off-by: TANGBUDU <tangbudu@gmail.com>
|
Azure Pipelines: There may be pipelines that require an authorized user to comment /azp run to run. |
Contributor
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The implementation is consistent across supported paths and has comprehensive focused tests with no actionable issues found.
Review effort: Balanced
Findings: None
What changed in this PR
Enables CPU DequantizeLinear to use scale and output types independently for opsets 23–25.
Changes:
- Dispatches computation using both scale and output element types.
- Adds mixed-type coverage across quantization granularities and input formats.
| File | Description |
|---|---|
onnxruntime/core/providers/cpu/quantization/quantize_linear.cc |
Separates scale and output template types. |
onnxruntime/test/providers/cpu/tensor/quantize_linear_test.cc |
Adds mixed-type CPU regression tests. |
💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
From opset 23 the output has its own type constraint T3. Register it with the supported float and float16 outputs so kernel matching validates the output type, as QuantizeLinear already does. Opsets 19 to 22 keep T1/T2, where the scale and output share T2.
TANGBUDU
added a commit
to TANGBUDU/onnxruntime
that referenced
this pull request
Sep 25, 2026
The opset-23 registration accepts independent input and scale types, but the kernel reads scale with the input type and ignores precision. Dispatch their storage types independently and round division in the requested FLOAT or FLOAT16 precision, defaulting to the scale type. Retain the existing FP32 computation and pre-opset-23 paths. Cover mixed types, rounding boundaries, saturation, blocked indexing, and packed output tails with CPU provider tests. Related to microsoft#32719. CPU DequantizeLinear is handled in microsoft#32728. Signed-off-by: TANGBUDU <tangbudu@gmail.com>
Scott McKay (skottmckay)
previously approved these changes
Sep 25, 2026
Regenerate the DequantizeLinear 23, 24 and 25+ rows to list the T3 output type added to the CPU kernel registrations.
Contributor
Author
|
Scott McKay (@skottmckay) Pushed a docs-only commit to regenerate OperatorKernels.md for the new T3 constraint, which the doc validation check was failing on. Could you approve the workflows and re-approve? Thanks! |
Scott McKay (skottmckay)
approved these changes
Sep 30, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Dispatch CPU
DequantizeLinearon the output tensor type and keep the scale type independent. This covers per-tensor, per-axis, and blocked quantization without converting the scale buffer.Motivation and Context
Since opset 23,
output_dtypecan differ from the scale type. A float scale with float16 output currently fails with a tensor type mismatch; the reverse combination fails as well.Addresses the CPU DequantizeLinear part of #32719. QuantizeLinear and CUDA support remain separate work.
Implemented with Codex assistance.