Skip to content

[Vulkan] Stop clamping per-row reductions and propagate NaN - #23205

Closed
mergennachin wants to merge 1 commit into
mergennachin/vulkan-23156-04-gelufrom
mergennachin/vulkan-23156-05-reduction-contract
Closed

mergennachin wants to merge 1 commit into
mergennachin/vulkan-23156-04-gelufrom
mergennachin/vulkan-23156-05-reduction-contract

Conversation

@mergennachin

@mergennachin mergennachin commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Superseded by #23244, part 5/15 of the ghstack replacement series. The current integration PR is #23254. This PR is closed in favor of the replacement; its previous description and review history are retained.


convert.glslh guarded an fp16 clamp with #if T == float16_t, but neither side is a macro, so the preprocessor compares 0 with 0 and the clamp was always on. Per-row buffer sum, mean, amax and amin therefore saturated fp32 and int results at 65504. The clamp is removed. Single-dimension texture and per-row buffer amax and amin now propagate NaN, argmax and argmin return the index of the first NaN as ATen does, and mean divides in the accumulator type. FP16 texture output uses explicit nearest-even conversion so rounding and overflow agree with ATen across drivers.

Part 5/15 of the Vulkan transformer and operator-conformance stack. Depends on #23204; review against the selected base branch. Integration PR: #23162.

Validation: 3 passed, 2 warnings in 57.33s. Lintrunner and git diff --check pass. Native tests use MoltenVK with portable CPU kernels; hardware and SwiftShader CI are pending.

Authored with OpenAI Codex; split planned with Claude Code.

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

convert.glslh guarded an fp16 clamp with `#if T == float16_t`, but neither side is a macro, so the preprocessor compares 0 with 0 and the clamp was always on. Per-row buffer sum, mean, amax and amin therefore saturated fp32 and int results at 65504. The clamp is removed. Single-dimension texture and per-row buffer amax and amin now propagate NaN, argmax and argmin return the index of the first NaN as ATen does, and mean divides in the accumulator type. FP16 texture output uses explicit nearest-even conversion so rounding and overflow agree with ATen across drivers.

Authored with OpenAI Codex; split planned with Claude Code.
@pytorch-bot pytorch-bot Bot added the module: vulkan Issues related to the Vulkan delegate and code under backends/vulkan/ label Sep 28, 2026
@pytorch-bot

pytorch-bot Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23205

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit 000b4ae with merge base a318382 (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 28, 2026
mergennachin added a commit that referenced this pull request Oct 3, 2026
convert.glslh guarded an fp16 clamp with `#if T == float16_t`, but
neither side is a macro, so the preprocessor compares 0 with 0 and the
clamp was always on. Per-row buffer sum, mean, amax and amin therefore
saturated fp32 results at 65504. The clamp also affected generated int32
shader variants, which are now tested through direct backend lowering;
integer reductions remain excluded by the partitioner. The clamp is
removed. Single-dimension texture and per-row buffer amax and amin now
propagate NaN, argmax and argmin return the index of the first NaN as
ATen does, and mean divides in the accumulator type. FP16 texture output
uses explicit nearest-even conversion so rounding and overflow agree
with ATen across drivers. Exact FP16 texture tests cover even/odd ties,
exponent carry, the normal/subnormal and zero/subnormal boundaries, and
sums below, at and above the ±65520 overflow threshold. Comparisons are
exact against ATen for both signs and include zero signs.

Part 5/15 of the Vulkan transformer and operator-conformance stack.
#23243 has landed; review against main. Integration PR: #23254.

Validation: All three focused review tests pass on MoltenVK. The FP16
tests cover 18 signed cases with exact ATen comparisons, actual FP16
texture output, and zero-sign checks. The int32 test lowers amax/amin
directly to Vulkan, verifies INT32 buffer inputs and outputs, and
matches ATen beyond both ±65504 limits; the same inputs also match
portable CPU execution. The previously rebased stack passed 51 native
tests with one expected SwiftShader-only skip; the review updates add
test coverage. Lintrunner and `git diff --check` pass. Vulkan CI is
rerunning on the updated heads.

Recreates #23205 through ghstack. Prior review discussion remains on
that PR.

Authored with OpenAI Codex; split planned with Claude Code.

This branch was successfully deployed

1 active deployment
cadence — 000b4aea Deployed Sep 28, 2026 by mergennachin via hifi-op-test / hifi4 #30191
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. module: vulkan Issues related to the Vulkan delegate and code under backends/vulkan/

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant