Skip to content

53_casual_attention performance test is numerically unpassable (logits ~1e5 vs atol=1e-5) #332

Description

@claude

Found while validating 53_casual_attention with scripts/run_challenge.py (investigating #123).

A correct CUDA kernel passes every functional test but fails the performance test on numerics alone:

M = 5000, d = 128
Mismatch in 'output'
Expected: [-17.6672, -7.2163, -85.2337, 98.9976, 60.2540, ..., 38.1356, -6.5884, -33.2429, -97.7734, 27.5228]
Got:      [-17.6672, -7.2163, -85.2337, 98.9976, 60.2540, ..., 38.1356, -6.5884, -33.2429, -97.7734, 27.5228]
Max abs diff: 0.04596710205078125

Note the printed values are identical to 4 decimals — this is not a logic error, it's float32 reassociation amplified by softmax.

Cause

generate_performance_test draws Q, K, V ~ U(-100, 100) at d = 128, while atol = rtol = 1e-5:

M, d = 5000, 128
"Q": RandTensor((M, d), -100.0, 100.0),

Logits are Q·Kᵀ/sqrt(d), so std ≈ (100²/3)·sqrt(128)/sqrt(128) ≈ 3333. Two consequences:

  1. A 128-term fp32 dot product of that magnitude carries ~2⁻²⁴ · 3333 · sqrt(128) ≈ 0.025 absolute uncertainty purely from summation order. Any custom kernel necessarily sums in a different order than PyTorch's cuBLAS call in reference_impl.
  2. exp() amplifies that. Simulating 200 rows of 5000 logits at σ = 3333, perturbing by 0.05, moves the softmax-weighted output by up to 0.011 — and over the real 5000 rows the observed 0.046 is entirely expected.

So the test demands agreement to 1e-5 on a quantity whose implementation-order-dependent spread is ~5e-2 — about 3.5 orders of magnitude tighter than achievable.

This challenge is the outlier

Every other softmax-attention challenge in the repo keeps logits well-conditioned, and most use looser tolerances:

challenge perf input scale atol/rtol
6_softmax_attention U(-0.1, 0.1) 1e-4
12_multi_head_attention U(-10, 10) 1e-5
26_multi_head_cross_attention randn 1e-4
80_grouped_query_attention randn 1e-4
92_decaying_causal_attention randn 1e-3
112_attention_with_sinks U(-1, 1), same M=5000, d=128 1e-5
53_casual_attention U(-100, 100) 1e-5

112_attention_with_sinks is the direct analogue — identical M = 5000, d = 128 — and uses U(-1, 1). Note that 53's own large_matrices functional test already uses uniform_(-0.1, 0.1), so only the perf test has the extreme range.

Suggested fix

Bring the perf-test range in line with 112_attention_with_sinks:

"Q": RandTensor((M, d), -1.0, 1.0),

and update the challenge.html constraint bullet, which currently documents [-100.0, 100.0]. Raising atol is not a sufficient alternative — the observed 0.046 error exceeds even 1e-3; it would need ~1e-1.

Not sending a PR for this one since it changes a published constraint and the effective difficulty of a live challenge — flagging for a maintainer call. Happy to open one on request.

Separately, #331 fixes an unrelated wrong expected value in this challenge's Example 2.

🤖 Generated with Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions