Skip to content

Commit c16dd0a

Browse files
committed
[ET-VK][reduction] Work around Adreno 8x4 dispatch failure
## Ulterior Motive Restore ET-VK operator correctness on Galaxy S25 devices used by Eureka CI. ## Rationale **What**: Make `reduce_nworkers()` select four workers when scaling would select eight. **Why**: Galaxy S25's Adreno Vulkan driver skips alternating work groups for the resulting `8x4` local size. Reduction outputs retain uninitialized data for those groups. ## Details | Behavior | Before | After | | --- | --- | --- | | Short reductions selecting eight workers | Dispatch `8x4`; eight generated tests fail | Dispatch `4x4`; all generated tests pass | | Larger reductions | Scale to 16, 32, or 64 workers | Unchanged | Authored with Codex. Differential Revision: [D122891447](https://our.internmc.facebook.com/intern/diff/D122891447/) ghstack-source-id: 440132327 Pull-Request: pytorch#23325
1 parent fdd5140 commit c16dd0a

1 file changed

Lines changed: 3 additions & 1 deletion

File tree

‎backends/vulkan/runtime/graph/ops/impl/Reduce.cpp‎

Lines changed: 3 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -119,7 +119,9 @@ uint32_t reduce_nworkers(
119119
while (nworkers * 2u <= cap && nworkers < extent) {
120120
nworkers *= 2u;
121121
}
122-
return nworkers;
122+
// Some Adreno drivers skip alternating work groups for an 8x4 local size.
123+
// Keep short reductions on the portable four-worker geometry.
124+
return nworkers == 8u ? 4u : nworkers;
123125
}
124126

125127
GlobalWorkGrid reduce_gwg_impl(

0 commit comments

Comments
 (0)