Skip to content

Commit 5b829f9

Browse files
authored
Enable use_sdpa_with_kv_cache in the qwen3_5 example config (pytorch#22390)
Fixes pytorch#22045. `use_sdpa_with_kv_cache: False` in this config costs roughly 2x decode. @JakeStevens confirmed in the issue there is no known reason to keep it off, and the lfm2 config — the same hybrid attention/conv architecture — already sets it `True`. Measured on Qwen3.5-2B, XNNPACK, 8da4w + 8-bit embedding, the same checkpoint exported twice with only this line different (Mac arm64, token-by-token prefill then greedy decode, same prompt, 32 new tokens): | | decode | |---|---| | `use_sdpa_with_kv_cache: False` | 8.20 tok/s | | `use_sdpa_with_kv_cache: True` | **16.64 tok/s** | That it is the same model rather than a faster different one: three of four prompts decode token-identically (the fourth diverges inside `<think>`, a near-tie either way), first-step logits corr 0.995 with the same top-1, and at 338 tokens of context corr 0.990 with the same top-1 and top-2. File size is unchanged (1483 MB both). cc @mergennachin @iseeyuan @lucylq @helunwencser @tarun292 @kimishpatel @jackzhxng
1 parent 0c6ecea commit 5b829f9

1 file changed

Lines changed: 1 addition & 1 deletion

File tree

‎examples/models/qwen3_5/config/qwen3_5_xnnpack_fp32.yaml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -7,7 +7,7 @@ base:
77

88
model:
99
use_kv_cache: True
10-
use_sdpa_with_kv_cache: False
10+
use_sdpa_with_kv_cache: True
1111
enable_dynamic_shape: False
1212
dtype_override: fp32
1313

0 commit comments

Comments
 (0)