[Vulkan] Honor GELU approximate='none' and keep kwargs off inserted views - #23243
Conversation
…ghstack [ghstack-poisoned]
…ghstack [ghstack-poisoned]
…ghstack [ghstack-poisoned]
…ghstack [ghstack-poisoned]
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23243
Note: Links to docs will display an error until the docs builds have been completed. ⏳ No Failures, 33 PendingAs of commit 0f3e873 with merge base 0b3d26d ( This comment was automatically generated by Dr. CI and updates every 15 minutes. |
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers. Fixes #23156. This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 | Hardware Vulkan CI | | 2 | #23241 | Scalar cache type and signed-zero keys | | 3 | #23242 | Vulkan-local signed-zero serialization | | 4 | #23243 | GELU modes and view kwargs | | 5 | #23244 | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 | Reduction dimension guards | | 7 | #23246 | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename. Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed. NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: c6a7e94 ghstack-comment-id: 5892899249 Pull-Request: #23254
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers. Fixes #23156. This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 | Hardware Vulkan CI | | 2 | #23241 | Scalar cache type and signed-zero keys | | 3 | #23242 | Vulkan-local signed-zero serialization | | 4 | #23243 | GELU modes and view kwargs | | 5 | #23244 | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 | Reduction dimension guards | | 7 | #23246 | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename. Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed. NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: f8d24d9 ghstack-comment-id: 5892899249 Pull-Request: #23254
[ghstack-poisoned]
This PR needs a
|
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers. Fixes #23156. This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 | Hardware Vulkan CI | | 2 | #23241 | Scalar cache type and signed-zero keys | | 3 | #23242 | Vulkan-local signed-zero serialization | | 4 | #23243 | GELU modes and view kwargs | | 5 | #23244 | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 | Reduction dimension guards | | 7 | #23246 | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename. Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed. NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: f02cc78 ghstack-comment-id: 5892899249 Pull-Request: #23254
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The implementation matches GELU semantics and includes focused coverage across modes, dtypes, storage types, and dynamic shapes.
Review effort: Balanced
Findings: None
What changed in this PR
Adds correct Vulkan handling for exact and tanh GELU modes while preventing GELU kwargs from leaking into inserted view operations.
Changes:
- Adds an erf-based exact GELU shader and runtime mode selection.
- Fixes kwargs handling for squeeze/unsqueeze views.
- Adds dynamic, storage, dtype, and singleton-dimension regression coverage.
| File | Description |
|---|---|
backends/vulkan/test/test_vulkan_dynamic.py |
Adds dynamic and singleton GELU tests. |
backends/vulkan/test/targets.bzl |
Registers the new test target. |
backends/vulkan/test/op_tests/cases.py |
Expands GELU operator coverage. |
backends/vulkan/runtime/graph/ops/impl/UnaryOp.cpp |
Selects the correct GELU shader. |
backends/vulkan/runtime/graph/ops/glsl/unary_op.yaml |
Defines the exact GELU variant. |
backends/vulkan/runtime/graph/ops/glsl/activations.h |
Implements erf-based GELU. |
backends/vulkan/_passes/squeeze_unsqueeze_inputs.py |
Removes invalid kwargs from inserted views. |
.github/workflows/vulkan.yml |
Runs dynamic tests on hardware. |
.github/workflows/pull.yml |
Runs dynamic tests in pull CI. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers. Fixes #23156. This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 | Hardware Vulkan CI | | 2 | #23241 | Scalar cache type and signed-zero keys | | 3 | #23242 | Vulkan-local signed-zero serialization | | 4 | #23243 | GELU modes and view kwargs | | 5 | #23244 | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 | Reduction dimension guards | | 7 | #23246 | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename. Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed. NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: 2d9c5b5 ghstack-comment-id: 5892899249 Pull-Request: #23254
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers. Fixes #23156. This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 | Hardware Vulkan CI | | 2 | #23241 | Scalar cache type and signed-zero keys | | 3 | #23242 | Vulkan-local signed-zero serialization | | 4 | #23243 | GELU modes and view kwargs | | 5 | #23244 | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 | Reduction dimension guards | | 7 | #23246 | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename. Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed. NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: 26d2d76 ghstack-comment-id: 5892899249 Pull-Request: #23254
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers. Fixes #23156. This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 | Hardware Vulkan CI | | 2 | #23241 | Scalar cache type and signed-zero keys | | 3 | #23242 | Vulkan-local signed-zero serialization | | 4 | #23243 | GELU modes and view kwargs | | 5 | #23244 | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 | Reduction dimension guards | | 7 | #23246 | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename. Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed. NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: 3a6574c ghstack-comment-id: 5892899249 Pull-Request: #23254
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers. Fixes #23156. This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 | Hardware Vulkan CI | | 2 | #23241 | Scalar cache type and signed-zero keys | | 3 | #23242 | Vulkan-local signed-zero serialization | | 4 | #23243 | GELU modes and view kwargs | | 5 | #23244 | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 | Reduction dimension guards | | 7 | #23246 | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename. Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed. NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: c9b5269 ghstack-comment-id: 5892899249 Pull-Request: #23254
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers. Fixes #23156. This is part 15 of the original ghstack series. Parts 1–4 (#23240 through #23243) have landed in main; the remaining eleven PRs run from #23244 through #23254. Review each PR against its selected base branch. This PR contains the final expansion and transformer changes; part 10 owns the scalar-tensor issue #23158. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 (landed) | Hardware Vulkan CI | | 2 | #23241 (landed) | Scalar cache type and signed-zero keys | | 3 | #23242 (landed) | Vulkan-local signed-zero serialization | | 4 | #23243 (landed) | GELU modes and view kwargs | | 5 | #23244 | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 | Reduction and arg-reduction dimension/storage guards | | 7 | #23246 | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim with keepdim=True | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. On native FP16 devices, the scalar-tensor, full texture, single-dimension texture reduction, and binary scalar shaders round their FP16 outputs to nearest-even. The binary scalar shaders also preserve the requested FP16 dtype when storage is emulated with FP32. Other operators retain their existing FP32 intermediate behavior on those devices. PR4 introduces `test_vulkan_dynamic.py` and its CI/Buck references. Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. Validation: Rebased onto main at `2c8510343e47` after parts 1–4 landed. `git range-diff` confirms all eleven remaining patches are unchanged by this rebase. On the previous base (`0b3d26d8c1f9`), the Release Vulkan/portable runtime rebuilt successfully and the native suite passed **51 tests with one expected SwiftShader-only skip**, before subsequent review tests were added. Coverage includes the dynamic tests, graph builder, serialization, and PT2E quantized linear with downcasting disabled. The two FP16 texture rounding tests pass 18 exact ATen cases covering both signs, signed zero, even/odd ties, exponent carry, the normal/subnormal and zero/subnormal boundaries, and values around the ±65520 overflow threshold. A separate test directly lowers int32 amax/amin to Vulkan buffers and checks exact ATen agreement beyond ±65504; the same inputs also match portable CPU execution. The arg-reduction guard update then passed four focused native tests, including 28 argmax/argmin cases checking Vulkan buffer execution and portable CPU fallback against ATen. The final rebase preserved these tested Vulkan sources; only upstream CI configuration and documentation changed. Lintrunner and `git diff --check` passed. Delegation of `any.dim` requires `keepdim=True`. Omitting `keepdim` or setting it to `False` keeps the reduction on CPU, avoiding a new bool-buffer requirement for surrounding texture delegates. The default/explicit-false regression test verifies texture storage and exact ATen results across changing shapes. Five focused native tests pass on MoltenVK, including existing `keepdim=True` execution and the full dynamic transformer tests; lintrunner and `git diff --check` also pass. SwiftShader CI will verify the updated heads. Vulkan CI will run again on the rebased ghstack heads. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Regression coverage includes scalar exponents 2.0001 and 2049 and a two-node FP16 mul.Scalar chain. The power and chained-multiplication tests compare with ATen in both texture and buffer storage. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: afcd265 ghstack-comment-id: 5892899249 Pull-Request: #23254
convert.glslh guarded an fp16 clamp with `#if T == float16_t`, but neither side is a macro, so the preprocessor compares 0 with 0 and the clamp was always on. Per-row buffer sum, mean, amax and amin therefore saturated fp32 results at 65504. The clamp also affected generated int32 shader variants, which are now tested through direct backend lowering; integer reductions remain excluded by the partitioner. The clamp is removed. Single-dimension texture and per-row buffer amax and amin now propagate NaN, argmax and argmin return the index of the first NaN as ATen does, and mean divides in the accumulator type. FP16 texture output uses explicit nearest-even conversion so rounding and overflow agree with ATen across drivers. Exact FP16 texture tests cover even/odd ties, exponent carry, the normal/subnormal and zero/subnormal boundaries, and sums below, at and above the ±65520 overflow threshold. Comparisons are exact against ATen for both signs and include zero signs. Part 5/15 of the Vulkan transformer and operator-conformance stack. #23243 has landed; review against main. Integration PR: #23254. Validation: All three focused review tests pass on MoltenVK. The FP16 tests cover 18 signed cases with exact ATen comparisons, actual FP16 texture output, and zero-sign checks. The int32 test lowers amax/amin directly to Vulkan, verifies INT32 buffer inputs and outputs, and matches ATen beyond both ±65504 limits; the same inputs also match portable CPU execution. The previously rebased stack passed 51 native tests with one expected SwiftShader-only skip; the review updates add test coverage. Lintrunner and `git diff --check` pass. Vulkan CI is rerunning on the updated heads. Recreates #23205 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code.
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows. Fixes #23156. This is part 15 of the original ghstack series. Parts 1–5 (#23240 through #23244) have landed in main; the remaining ten PRs run from #23245 through #23254. Review each PR against its selected base branch. This PR contains the final expansion and transformer changes; part 10 owns the scalar-tensor issue #23158. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 (landed) | Hardware Vulkan CI | | 2 | #23241 (landed) | Scalar cache type and signed-zero keys | | 3 | #23242 (landed) | Vulkan-local signed-zero serialization | | 4 | #23243 (landed) | GELU modes and view kwargs | | 5 | #23244 (landed) | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 | Reduction and arg-reduction dimension/storage guards | | 7 | #23246 | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim textures with either keepdim setting | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. On native FP16 devices, the scalar-tensor, full texture, single-dimension texture reduction, and binary scalar shaders round their FP16 outputs to nearest-even. The binary scalar shaders also preserve the requested FP16 dtype when storage is emulated with FP32. Other operators retain their existing FP32 intermediate behavior on those devices. Coverage includes scalar exponents 2.0001 and 2049 and a two-node FP16 mul.Scalar chain in both texture and buffer storage. PR4 introduces `test_vulkan_dynamic.py` and its CI/Buck references. In #23253, `any.dim` now supports both `keepdim` settings on textures. With `keepdim=False`, a reduction into a temporary texture is followed by a GPU view/repack into the squeezed output shape. Both nodes resize dynamically, and empty reduced axes produce false. This adds one temporary texture and one GPU dispatch without introducing a bool-buffer requirement for texture models. Unsupported scalar inputs and 4D batch/channel axes use CPU fallback. Current validation: Rebased onto main at `903cef063774` after parts 1–5 landed, preserving the ten existing code patches before adding the texture implementation in #23253. The Release Vulkan/portable runtime builds. Eleven focused native tests pass on Apple M1 Pro / MoltenVK, covering dynamic and chained any reductions, actual texture storage, scalar outputs, singleton and empty dimensions, growth after a zero-length dimension, transformer integration, supported and unsupported 4D reductions, special values, and FP16 rounding. An additional 27 FACTO-generated bool any.dim cases and six boundary cases match ATen and portable kernels exactly through the shared `backends/test/` harness: 23 run on Vulkan textures and 10 exercise expected CPU fallback. Lintrunner and `git diff --check` pass. SwiftShader and hardware CI will run on the updated ghstack heads. Earlier validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covered all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. The prior full native run passed 51 tests with one expected SwiftShader-only skip. Subsequent review coverage added exact FP16 rounding boundaries, int32 buffer amax/amin range checks against both ATen and portable kernels, and 28 argmax/argmin execution and fallback cases. These are prior validation results, not a repeat of the full suite for this update. Validation uses Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: 980827d ghstack-comment-id: 5892899249 Pull-Request: #23254
The Vulkan gelu kernel always used the tanh approximation, so exact GELU, PyTorch's default, was delegated with a different result. It now evaluates an erf-based form unless approximate='tanh' is requested. SqueezeUnsqueezeInputs copied the original node's kwargs onto the view_copy nodes it inserts, which broke gelu on inputs with singleton dimensions. This also adds test_vulkan_dynamic.py, which lowers with require_dynamic_shapes and runs each program across several input shapes; later changes extend it.
Part 4 of the original 15-PR Vulkan transformer and operator-conformance stack. Parts 1–3 (#23240, #23241, and #23242) have landed in main. This is now the first of the twelve remaining PRs, based on main. Integration PR: #23254.
Validation: Rebased onto main at
0b3d26d8c1f9after parts 1–3 landed.git range-diffconfirms all twelve remaining patches match their prior versions. The Release Vulkan/portable runtime rebuilt successfully, and the complete native regression suite passed 51 tests with one expected SwiftShader-only skip on MoltenVK. Coverage includes the dynamic tests, graph builder, serialization, and PT2E quantized linear with downcasting disabled. Lintrunner andgit diff --checkpassed.Recreates #23204 through ghstack. Prior review discussion remains on that PR.
Authored with OpenAI Codex; split planned with Claude Code.