Skip to content

[Vulkan] Keep reductions the reduce shaders cannot handle on CPU - #23206

Closed
mergennachin wants to merge 1 commit into
mergennachin/vulkan-23156-05-reduction-contractfrom
mergennachin/vulkan-23156-06-reduction-guards
Closed

mergennachin wants to merge 1 commit into
mergennachin/vulkan-23156-05-reduction-contractfrom
mergennachin/vulkan-23156-06-reduction-guards

Conversation

@mergennachin

@mergennachin mergennachin commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Superseded by #23245, part 6/15 of the ghstack replacement series. The current integration PR is #23254. This PR is closed in favor of the replacement; its previous description and review history are retained.


The general reduce implementation handles one or two reduced dims, but the partitioner accepted empty dim lists and full sum/mean reductions over N-D inputs, which then failed to build or computed the wrong reduction. Texture storage also folds batch into channels, so 4D reductions over dim 0, or over dim 1 with batch above 1, cannot be computed. These cases now stay on CPU.

Part 6/15 of the Vulkan transformer and operator-conformance stack. Depends on #23205; review against the selected base branch. Integration PR: #23162.

Validation: 2 passed, 2 warnings in 70.88s (0:01:10). Lintrunner and git diff --check pass. Native tests use MoltenVK with portable CPU kernels; hardware and SwiftShader CI are pending.

Authored with OpenAI Codex; split planned with Claude Code.

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

The general reduce implementation handles one or two reduced dims, but the partitioner accepted empty dim lists and full sum/mean reductions over N-D inputs, which then failed to build or computed the wrong reduction. Texture storage also folds batch into channels, so 4D reductions over dim 0, or over dim 1 with batch above 1, cannot be computed. These cases now stay on CPU.

Authored with OpenAI Codex; split planned with Claude Code.
@pytorch-bot pytorch-bot Bot added the module: vulkan Issues related to the Vulkan delegate and code under backends/vulkan/ label Sep 28, 2026
@pytorch-bot

pytorch-bot Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23206

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit d0ff92f with merge base a318382 (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 28, 2026
mergennachin added a commit that referenced this pull request Oct 4, 2026
)

The general reduce implementation handles one or two reduced dims, but
the partitioner accepted empty dim lists and full sum/mean reductions
over N-D inputs, which then failed to build or computed the wrong
reduction. Texture storage also folds batch into channels, so 4D
reductions over dim 0, or over dim 1 with batch above 1, cannot be
computed. These cases now stay on CPU.

`argmax` and `argmin` are restricted to contiguous buffers and
last-dimension reductions, matching `ArgReduce.cpp`. A missing or `None`
dim is supported only for rank-1 inputs; rank-2 full reductions,
including `[1, N]`, fall back to CPU for either `keepdim` value.

Part 6/15 of the Vulkan transformer and operator-conformance stack.
Depends on #23244; review against the selected base branch. Integration
PR: #23254.

Validation: Four focused native tests passed on MoltenVK with portable
CPU kernels. The new test covers 28 argmax/argmin cases: rank-2
`dim=None`, non-last dimensions, rank-1 full reductions, and
positive/negative last dimensions, with both keepdim settings. It checks
actual Vulkan buffer storage for supported cases, CPU fallback for
unsupported cases, and exact agreement with ATen. Existing 4D reduction,
reduction-dimension fallback, and first-NaN tests also passed. The new
test failed on the reported `[1, N]` case before the fix. Lintrunner and
`git diff --check` pass; CI is rerunning.

Recreates #23206 through ghstack. Prior review discussion remains on
that PR.

Authored with OpenAI Codex; split planned with Claude Code.


cc @SS-JIA @manuelcandales @digantdesai @cbilgin

This branch was successfully deployed

1 active deployment
cadence — d0ff92f9 Deployed Sep 28, 2026 by mergennachin via hifi-op-test / hifi4 #30194
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. module: vulkan Issues related to the Vulkan delegate and code under backends/vulkan/

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant