[Vulkan] Make expand_copy resizable and fully delegate dynamic transformer blocks - #23254
Draft
mergennachin wants to merge 8 commits into
Draft
mergennachin wants to merge 8 commits into
mergennachin wants to merge 8 commits into
Conversation
…ghstack [ghstack-poisoned]
Contributor
Author
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23254
Note: Links to docs will display an error until the docs builds have been completed. ❌ 1 New FailureAs of commit 0f9c2dd with merge base 2c85103 ( NEW FAILURE - The following job has failed:
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This was referenced Sep 29, 2026
This was referenced Sep 29, 2026
[ghstack-poisoned]
mergennachin
added a commit
that referenced
this pull request
Oct 3, 2026
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers. Fixes #23156. This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 | Hardware Vulkan CI | | 2 | #23241 | Scalar cache type and signed-zero keys | | 3 | #23242 | Vulkan-local signed-zero serialization | | 4 | #23243 | GELU modes and view kwargs | | 5 | #23244 | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 | Reduction dimension guards | | 7 | #23246 | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename. Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed. NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: f02cc78 ghstack-comment-id: 5892899249 Pull-Request: #23254
[ghstack-poisoned]
mergennachin
added a commit
that referenced
this pull request
Oct 3, 2026
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers. Fixes #23156. This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 | Hardware Vulkan CI | | 2 | #23241 | Scalar cache type and signed-zero keys | | 3 | #23242 | Vulkan-local signed-zero serialization | | 4 | #23243 | GELU modes and view kwargs | | 5 | #23244 | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 | Reduction dimension guards | | 7 | #23246 | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename. Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed. NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: 2d9c5b5 ghstack-comment-id: 5892899249 Pull-Request: #23254
mergennachin
added a commit
that referenced
this pull request
Oct 3, 2026
…iews (#23243) The Vulkan gelu kernel always used the tanh approximation, so exact GELU, PyTorch's default, was delegated with a different result. It now evaluates an erf-based form unless approximate='tanh' is requested. SqueezeUnsqueezeInputs copied the original node's kwargs onto the view_copy nodes it inserts, which broke gelu on inputs with singleton dimensions. This also adds test_vulkan_dynamic.py, which lowers with require_dynamic_shapes and runs each program across several input shapes; later changes extend it. Part 4 of the original 15-PR Vulkan transformer and operator-conformance stack. Parts 1–3 (#23240, #23241, and #23242) have landed in main. This is now the first of the twelve remaining PRs, based on main. Integration PR: #23254. Validation: Rebased onto main at `0b3d26d8c1f9` after parts 1–3 landed. `git range-diff` confirms all twelve remaining patches match their prior versions. The Release Vulkan/portable runtime rebuilt successfully, and the complete native regression suite passed **51 tests with one expected SwiftShader-only skip** on MoltenVK. Coverage includes the dynamic tests, graph builder, serialization, and PT2E quantized linear with downcasting disabled. Lintrunner and `git diff --check` passed. Recreates #23204 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code.
…ension kernel; rebase after PR 23243 landed [ghstack-poisoned]
mergennachin
added a commit
that referenced
this pull request
Oct 3, 2026
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers. Fixes #23156. This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 | Hardware Vulkan CI | | 2 | #23241 | Scalar cache type and signed-zero keys | | 3 | #23242 | Vulkan-local signed-zero serialization | | 4 | #23243 | GELU modes and view kwargs | | 5 | #23244 | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 | Reduction dimension guards | | 7 | #23246 | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename. Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed. NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: 26d2d76 ghstack-comment-id: 5892899249 Pull-Request: #23254
[ghstack-poisoned]
mergennachin
added a commit
that referenced
this pull request
Oct 3, 2026
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers. Fixes #23156. This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 | Hardware Vulkan CI | | 2 | #23241 | Scalar cache type and signed-zero keys | | 3 | #23242 | Vulkan-local signed-zero serialization | | 4 | #23243 | GELU modes and view kwargs | | 5 | #23244 | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 | Reduction dimension guards | | 7 | #23246 | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename. Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed. NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: 3a6574c ghstack-comment-id: 5892899249 Pull-Request: #23254
… clamp [ghstack-poisoned]
mergennachin
added a commit
that referenced
this pull request
Oct 3, 2026
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers. Fixes #23156. This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 | Hardware Vulkan CI | | 2 | #23241 | Scalar cache type and signed-zero keys | | 3 | #23242 | Vulkan-local signed-zero serialization | | 4 | #23243 | GELU modes and view kwargs | | 5 | #23244 | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 | Reduction dimension guards | | 7 | #23246 | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename. Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed. NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: c9b5269 ghstack-comment-id: 5892899249 Pull-Request: #23254
…l buffers [ghstack-poisoned]
mergennachin
added a commit
that referenced
this pull request
Oct 3, 2026
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers. Fixes #23156. This is part 15 of the original ghstack series. Parts 1–4 (#23240 through #23243) have landed in main; the remaining eleven PRs run from #23244 through #23254. Review each PR against its selected base branch. This PR contains the final expansion and transformer changes; part 10 owns the scalar-tensor issue #23158. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 (landed) | Hardware Vulkan CI | | 2 | #23241 (landed) | Scalar cache type and signed-zero keys | | 3 | #23242 (landed) | Vulkan-local signed-zero serialization | | 4 | #23243 (landed) | GELU modes and view kwargs | | 5 | #23244 | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 | Reduction and arg-reduction dimension/storage guards | | 7 | #23246 | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim with keepdim=True | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. On native FP16 devices, the scalar-tensor, full texture, single-dimension texture reduction, and binary scalar shaders round their FP16 outputs to nearest-even. The binary scalar shaders also preserve the requested FP16 dtype when storage is emulated with FP32. Other operators retain their existing FP32 intermediate behavior on those devices. PR4 introduces `test_vulkan_dynamic.py` and its CI/Buck references. Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. Validation: Rebased onto main at `2c8510343e47` after parts 1–4 landed. `git range-diff` confirms all eleven remaining patches are unchanged by this rebase. On the previous base (`0b3d26d8c1f9`), the Release Vulkan/portable runtime rebuilt successfully and the native suite passed **51 tests with one expected SwiftShader-only skip**, before subsequent review tests were added. Coverage includes the dynamic tests, graph builder, serialization, and PT2E quantized linear with downcasting disabled. The two FP16 texture rounding tests pass 18 exact ATen cases covering both signs, signed zero, even/odd ties, exponent carry, the normal/subnormal and zero/subnormal boundaries, and values around the ±65520 overflow threshold. A separate test directly lowers int32 amax/amin to Vulkan buffers and checks exact ATen agreement beyond ±65504; the same inputs also match portable CPU execution. The arg-reduction guard update then passed four focused native tests, including 28 argmax/argmin cases checking Vulkan buffer execution and portable CPU fallback against ATen. The final rebase preserved these tested Vulkan sources; only upstream CI configuration and documentation changed. Lintrunner and `git diff --check` passed. Delegation of `any.dim` requires `keepdim=True`. Omitting `keepdim` or setting it to `False` keeps the reduction on CPU, avoiding a new bool-buffer requirement for surrounding texture delegates. The default/explicit-false regression test verifies texture storage and exact ATen results across changing shapes. Five focused native tests pass on MoltenVK, including existing `keepdim=True` execution and the full dynamic transformer tests; lintrunner and `git diff --check` also pass. SwiftShader CI will verify the updated heads. Vulkan CI will run again on the rebased ghstack heads. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Regression coverage includes scalar exponents 2.0001 and 2049 and a two-node FP16 mul.Scalar chain. The power and chained-multiplication tests compare with ATen in both texture and buffer storage. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: afcd265 ghstack-comment-id: 5892899249 Pull-Request: #23254
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes
expand_copyresizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers.Fixes #23156.
This is part 15 of the original ghstack series. Parts 1–4 (#23240 through #23243) have landed in main; the remaining eleven PRs run from #23244 through #23254. Review each PR against its selected base branch. This PR contains the final expansion and transformer changes; part 10 owns the scalar-tensor issue #23158.
The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. On native FP16 devices, the scalar-tensor, full texture, single-dimension texture reduction, and binary scalar shaders round their FP16 outputs to nearest-even. The binary scalar shaders also preserve the requested FP16 dtype when storage is emulated with FP32. Other operators retain their existing FP32 intermediate behavior on those devices. PR4 introduces
test_vulkan_dynamic.pyand its CI/Buck references.Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at
a31838280f9309f04af1b375147375b702d29339found no new regressions or unresolved comparisons.Validation uses the shared
backends/test/harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference usestexture_limits=(1,1,1), with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision:3b8c778c99766a8b4d0d04563ae0b16cbb276829, seed 0.Validation: Rebased onto main at
2c8510343e47after parts 1–4 landed.git range-diffconfirms all eleven remaining patches are unchanged by this rebase. On the previous base (0b3d26d8c1f9), the Release Vulkan/portable runtime rebuilt successfully and the native suite passed 51 tests with one expected SwiftShader-only skip, before subsequent review tests were added. Coverage includes the dynamic tests, graph builder, serialization, and PT2E quantized linear with downcasting disabled. The two FP16 texture rounding tests pass 18 exact ATen cases covering both signs, signed zero, even/odd ties, exponent carry, the normal/subnormal and zero/subnormal boundaries, and values around the ±65520 overflow threshold. A separate test directly lowers int32 amax/amin to Vulkan buffers and checks exact ATen agreement beyond ±65504; the same inputs also match portable CPU execution. The arg-reduction guard update then passed four focused native tests, including 28 argmax/argmin cases checking Vulkan buffer execution and portable CPU fallback against ATen. The final rebase preserved these tested Vulkan sources; only upstream CI configuration and documentation changed. Lintrunner andgit diff --checkpassed.Delegation of
any.dimrequireskeepdim=True. Omittingkeepdimor setting it toFalsekeeps the reduction on CPU, avoiding a new bool-buffer requirement for surrounding texture delegates. The default/explicit-false regression test verifies texture storage and exact ATen results across changing shapes. Five focused native tests pass on MoltenVK, including existingkeepdim=Trueexecution and the full dynamic transformer tests; lintrunner andgit diff --checkalso pass. SwiftShader CI will verify the updated heads.Vulkan CI will run again on the rebased ghstack heads.
Recreates #23162 through ghstack. Prior review discussion remains on that PR.
Regression coverage includes scalar exponents 2.0001 and 2049 and a two-node FP16 mul.Scalar chain. The power and chained-multiplication tests compare with ATen in both texture and buffer storage.
Authored with OpenAI Codex; split planned with Claude Code.
cc @SS-JIA @manuelcandales @digantdesai @cbilgin