Skip to content

Commit d612ca4

Browse files
Compare against a float threshold in sdpa_bitwise_mask_gen
Summary: The float path compares each mask element against the `double` threshold the op schema declares. `DLA_V130_1p0v2` has a single-precision FPU but no double-precision hardware, so every element promoted through two soft-float library calls: call8 __extendsfdf2 # float -> double call8 __ltdf2 # double < double Sixteen library calls per output byte, for a comparison the core can do in one `olt.s`. Rounding the threshold to float once outside the loop leaves eight `olt.s` in the loop body and nothing else. Rounding to nearest can only move the boundary for a threshold that is not exactly representable in float. Both producers pass `0.0`, which is exact, and the HiFi TIE kernel already rounds the same way, so the packed bytes are unchanged and the two kernels stay byte-compatible. The signature is untouched. This version answers review feedback on the benchmark this change ships. The 16 KB snapshot buffer moves off the stack to the heap, since these cores cannot be relied on to have that much task stack; the speedup division guards a zero cycle count; and the three-quarter regression bound is computed in 64-bit instead of truncating twice, which is the same bound arithmetically, so the guard still fires exactly where it did. Two comments that claimed more than they covered are corrected: the causal mask yields one mixed output byte per row rather than a mixed byte everywhere, and the byte-for-byte check holds because the threshold is exactly representable, not because rounding is neutral in general. The kernel comment now states that precondition and its limit rather than only the conclusion. No behaviour change; the ISS figures below are from a rerun on the edited benchmark and move by under 1%. The test plan's reproduction steps are corrected in this version too. They described a stack that no longer exists -- D120841844 and D121265174 have landed, and D120855219 is not being landed -- and they invoked a `compare_op_etdump` binary that ships with D120855219 and so will not exist. The baseline arm is now master plus D120855874, and the ETDump comparison is described rather than delegated to that tool. No measurement changes. Reviewed By: mcremon-meta Differential Revision: D120937330 Pull Request resolved: pytorch#23165
1 parent 674a7ea commit d612ca4

1 file changed

Lines changed: 11 additions & 1 deletion

File tree

‎backends/cadence/generic/operators/op_sdpa_bitwise_mask_gen.cpp‎

Lines changed: 11 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -69,10 +69,20 @@ ::executorch::aten::Tensor& sdpa_bitwise_mask_gen_out(
6969
} else if (dtype == ::executorch::aten::ScalarType::Float) {
7070
const float* __restrict mask_data = mask.mutable_data_ptr<float>();
7171
uint8_t* __restrict out_data = out.mutable_data_ptr<uint8_t>();
72+
// Compare in float. Xtensa DSP cores here have a single-precision FPU but
73+
// no double-precision hardware, so `float < double` promotes via two
74+
// soft-float library calls (__extendsfdf2, __ltdf2) on every element.
75+
//
76+
// This assumes the threshold is exactly representable in float, which is
77+
// what AOT emits and what every producer passes today (all pass 0.0). A
78+
// threshold that is not representable rounds to nearest and can move the
79+
// boundary by up to half an ULP. The HiFi TIE kernel already rounds the
80+
// same way, so the two kernels agree either way.
81+
const float f_threshold = static_cast<float>(threshold);
7282
for (int64_t i = 0, out_index = 0; i < numel; i += 8, out_index++) {
7383
uint8_t packed_mask = 0;
7484
for (int64_t j = 0; j < 8; j++) {
75-
packed_mask |= (mask_data[i + j] < threshold) << j;
85+
packed_mask |= (mask_data[i + j] < f_threshold) << j;
7686
}
7787
out_data[out_index] = packed_mask;
7888
}

0 commit comments

Comments
 (0)