Found while validating 53_casual_attention with scripts/run_challenge.py (investigating #123).
A correct CUDA kernel passes every functional test but fails the performance test on numerics alone:
M = 5000, d = 128
Mismatch in 'output'
Expected: [-17.6672, -7.2163, -85.2337, 98.9976, 60.2540, ..., 38.1356, -6.5884, -33.2429, -97.7734, 27.5228]
Got: [-17.6672, -7.2163, -85.2337, 98.9976, 60.2540, ..., 38.1356, -6.5884, -33.2429, -97.7734, 27.5228]
Max abs diff: 0.04596710205078125
Note the printed values are identical to 4 decimals — this is not a logic error, it's float32 reassociation amplified by softmax.
Cause
generate_performance_test draws Q, K, V ~ U(-100, 100) at d = 128, while atol = rtol = 1e-5:
M, d = 5000, 128
"Q": RandTensor((M, d), -100.0, 100.0),
Logits are Q·Kᵀ/sqrt(d), so std ≈ (100²/3)·sqrt(128)/sqrt(128) ≈ 3333. Two consequences:
- A 128-term fp32 dot product of that magnitude carries ~
2⁻²⁴ · 3333 · sqrt(128) ≈ 0.025 absolute uncertainty purely from summation order. Any custom kernel necessarily sums in a different order than PyTorch's cuBLAS call in reference_impl.
exp() amplifies that. Simulating 200 rows of 5000 logits at σ = 3333, perturbing by 0.05, moves the softmax-weighted output by up to 0.011 — and over the real 5000 rows the observed 0.046 is entirely expected.
So the test demands agreement to 1e-5 on a quantity whose implementation-order-dependent spread is ~5e-2 — about 3.5 orders of magnitude tighter than achievable.
This challenge is the outlier
Every other softmax-attention challenge in the repo keeps logits well-conditioned, and most use looser tolerances:
| challenge |
perf input scale |
atol/rtol |
| 6_softmax_attention |
U(-0.1, 0.1) |
1e-4 |
| 12_multi_head_attention |
U(-10, 10) |
1e-5 |
| 26_multi_head_cross_attention |
randn |
1e-4 |
| 80_grouped_query_attention |
randn |
1e-4 |
| 92_decaying_causal_attention |
randn |
1e-3 |
| 112_attention_with_sinks |
U(-1, 1), same M=5000, d=128 |
1e-5 |
| 53_casual_attention |
U(-100, 100) |
1e-5 |
112_attention_with_sinks is the direct analogue — identical M = 5000, d = 128 — and uses U(-1, 1). Note that 53's own large_matrices functional test already uses uniform_(-0.1, 0.1), so only the perf test has the extreme range.
Suggested fix
Bring the perf-test range in line with 112_attention_with_sinks:
"Q": RandTensor((M, d), -1.0, 1.0),
and update the challenge.html constraint bullet, which currently documents [-100.0, 100.0]. Raising atol is not a sufficient alternative — the observed 0.046 error exceeds even 1e-3; it would need ~1e-1.
Not sending a PR for this one since it changes a published constraint and the effective difficulty of a live challenge — flagging for a maintainer call. Happy to open one on request.
Separately, #331 fixes an unrelated wrong expected value in this challenge's Example 2.
🤖 Generated with Claude Code
Found while validating
53_casual_attentionwithscripts/run_challenge.py(investigating #123).A correct CUDA kernel passes every functional test but fails the performance test on numerics alone:
Note the printed values are identical to 4 decimals — this is not a logic error, it's float32 reassociation amplified by softmax.
Cause
generate_performance_testdrawsQ, K, V ~ U(-100, 100)atd = 128, whileatol = rtol = 1e-5:Logits are
Q·Kᵀ/sqrt(d), sostd ≈ (100²/3)·sqrt(128)/sqrt(128) ≈ 3333. Two consequences:2⁻²⁴ · 3333 · sqrt(128) ≈ 0.025absolute uncertainty purely from summation order. Any custom kernel necessarily sums in a different order than PyTorch's cuBLAS call inreference_impl.exp()amplifies that. Simulating 200 rows of 5000 logits atσ = 3333, perturbing by 0.05, moves the softmax-weighted output by up to 0.011 — and over the real 5000 rows the observed 0.046 is entirely expected.So the test demands agreement to
1e-5on a quantity whose implementation-order-dependent spread is ~5e-2— about 3.5 orders of magnitude tighter than achievable.This challenge is the outlier
Every other softmax-attention challenge in the repo keeps logits well-conditioned, and most use looser tolerances:
U(-0.1, 0.1)U(-10, 10)randnrandnrandnU(-1, 1), sameM=5000, d=128U(-100, 100)112_attention_with_sinksis the direct analogue — identicalM = 5000, d = 128— and usesU(-1, 1). Note that 53's ownlarge_matricesfunctional test already usesuniform_(-0.1, 0.1), so only the perf test has the extreme range.Suggested fix
Bring the perf-test range in line with
112_attention_with_sinks:and update the
challenge.htmlconstraint bullet, which currently documents[-100.0, 100.0]. Raisingatolis not a sufficient alternative — the observed 0.046 error exceeds even1e-3; it would need ~1e-1.Not sending a PR for this one since it changes a published constraint and the effective difficulty of a live challenge — flagging for a maintainer call. Happy to open one on request.
Separately, #331 fixes an unrelated wrong expected value in this challenge's Example 2.
🤖 Generated with Claude Code