[Vulkan] Stop clamping per-row reductions and propagate NaN - #23205
Closed
mergennachin wants to merge 1 commit into
Closed
mergennachin wants to merge 1 commit into
mergennachin wants to merge 1 commit into
Conversation
convert.glslh guarded an fp16 clamp with `#if T == float16_t`, but neither side is a macro, so the preprocessor compares 0 with 0 and the clamp was always on. Per-row buffer sum, mean, amax and amin therefore saturated fp32 and int results at 65504. The clamp is removed. Single-dimension texture and per-row buffer amax and amin now propagate NaN, argmax and argmin return the index of the first NaN as ATen does, and mean divides in the accumulator type. FP16 texture output uses explicit nearest-even conversion so rounding and overflow agree with ATen across drivers. Authored with OpenAI Codex; split planned with Claude Code.
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23205
Note: Links to docs will display an error until the docs builds have been completed. ✅ No FailuresAs of commit 000b4ae with merge base a318382 ( This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This was referenced Sep 28, 2026
mergennachin
added a commit
that referenced
this pull request
Oct 3, 2026
convert.glslh guarded an fp16 clamp with `#if T == float16_t`, but neither side is a macro, so the preprocessor compares 0 with 0 and the clamp was always on. Per-row buffer sum, mean, amax and amin therefore saturated fp32 results at 65504. The clamp also affected generated int32 shader variants, which are now tested through direct backend lowering; integer reductions remain excluded by the partitioner. The clamp is removed. Single-dimension texture and per-row buffer amax and amin now propagate NaN, argmax and argmin return the index of the first NaN as ATen does, and mean divides in the accumulator type. FP16 texture output uses explicit nearest-even conversion so rounding and overflow agree with ATen across drivers. Exact FP16 texture tests cover even/odd ties, exponent carry, the normal/subnormal and zero/subnormal boundaries, and sums below, at and above the ±65520 overflow threshold. Comparisons are exact against ATen for both signs and include zero signs. Part 5/15 of the Vulkan transformer and operator-conformance stack. #23243 has landed; review against main. Integration PR: #23254. Validation: All three focused review tests pass on MoltenVK. The FP16 tests cover 18 signed cases with exact ATen comparisons, actual FP16 texture output, and zero-sign checks. The int32 test lowers amax/amin directly to Vulkan, verifies INT32 buffer inputs and outputs, and matches ATen beyond both ±65504 limits; the same inputs also match portable CPU execution. The previously rebased stack passed 51 native tests with one expected SwiftShader-only skip; the review updates add test coverage. Lintrunner and `git diff --check` pass. Vulkan CI is rerunning on the updated heads. Recreates #23205 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code.
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Superseded by #23244, part 5/15 of the ghstack replacement series. The current integration PR is #23254. This PR is closed in favor of the replacement; its previous description and review history are retained.
convert.glslh guarded an fp16 clamp with
#if T == float16_t, but neither side is a macro, so the preprocessor compares 0 with 0 and the clamp was always on. Per-row buffer sum, mean, amax and amin therefore saturated fp32 and int results at 65504. The clamp is removed. Single-dimension texture and per-row buffer amax and amin now propagate NaN, argmax and argmin return the index of the first NaN as ATen does, and mean divides in the accumulator type. FP16 texture output uses explicit nearest-even conversion so rounding and overflow agree with ATen across drivers.Part 5/15 of the Vulkan transformer and operator-conformance stack. Depends on #23204; review against the selected base branch. Integration PR: #23162.
Validation: 3 passed, 2 warnings in 57.33s. Lintrunner and git diff --check pass. Native tests use MoltenVK with portable CPU kernels; hardware and SwiftShader CI are pending.
Authored with OpenAI Codex; split planned with Claude Code.
cc @SS-JIA @manuelcandales @digantdesai @cbilgin