Skip to content

Add challenge 74: N-body Gravitational Force (Medium) - #196

Open
claude[bot] wants to merge 2 commits into
mainfrom
add-challenge-74-n-body-force
Open

claude[bot] wants to merge 2 commits into
mainfrom
add-challenge-74-n-body-force

Conversation

@claude

@claude claude Bot commented Feb 27, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Adds challenge 74: N-body Gravitational Force (Medium difficulty)
  • Teaches the classic all-pairs O(N²) GPU computation pattern — shared-memory tiling is the key optimization to escape global-memory bandwidth bottleneck
  • Softened gravitational force formula: F_i = Σ_{j≠i} m_j (r_j - r_i) / (|r_j - r_i|² + ε²)^{3/2} with ε = 1e-3
  • Inputs: positions (N×3 float32), masses (N float32); Output: forces (N×3 float32)
  • Performance test: N = 8,192 (~67M pairwise force evaluations)

Why this challenge?

Unlike the many element-wise challenges already in the repo, this is a genuine all-pairs computation where every output depends on all N inputs. The natural optimization (shared-memory tiling to amortize global loads across a block of particles) is the same conceptual leap as GEMM tiling but applied to a completely different physics problem — making it educational without being a duplicate.

Test plan

  • pre-commit run --all-files passes (Black, isort, flake8, clang-format)
  • run_challenge.py --action run passes on NVIDIA TESLA T4 (example test)
  • All checklist items from CLAUDE.md verified:
    • challenge.html starts with <p>, uses <h2> sections, first example matches generate_example_test(), LaTeX bmatrix used consistently, performance bullet present
    • challenge.py: inherits ChallengeBase, all 6 methods, 9 functional tests (edge/pow2/non-pow2/realistic + zero inputs + negatives), performance test fits 5× in 16 GB
    • All 6 starter files present with correct parameter description comments
    • Solution file not committed

🤖 Generated with Claude Code

Teaches the all-pairs O(N²) parallel computation pattern: each output
depends on all N inputs, requiring shared-memory tiling to avoid the
global-memory bandwidth bottleneck. Force uses the softened gravity
formula with ε = 1e-3, tested on sizes up to N = 8,192.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

Copy link
Copy Markdown
Contributor

Thanks for the PR! Challenge ID 74 is already taken on main (challenges/hard/74_gpt2_block). Please renumber to a free ID (e.g. 95, 97, 99–101, 103, or 120+) and update the directory name + PR title. Happy to merge once the ID no longer collides.

- Remove the `positions.device.type == "cuda"` assertion from
  reference_impl and route every tensor allocation through
  `self.device` so the reference runs on non-CUDA accelerators (XLA)
- Add a `device` constructor argument (default "cuda") to supply
  `self.device`; keep the `super().__init__(...)` metadata call so
  `Challenge()` still loads under the current `ChallengeBase`
- Add an N=3 functional test with mixed-sign coordinates (10 cases)
- Add a dark-theme SVG showing pairwise contributions summing to the
  net force on a body, plus a position range bullet in Constraints

Verified locally on CPU: reference_impl matches an independent O(N^2)
float64 implementation of the documented formula on the example and all
10 functional tests, and a tiled CUDA solution (emulated in float32)
agrees with the reference at N=8,192 with the worst deviation at 0.2%
of the atol/rtol=1e-2 budget. `pre-commit run --all-files` passes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@claude
claude Bot requested a review from shxjames as a code owner September 30, 2026 11:53
@claude

claude Bot commented Sep 30, 2026

Copy link
Copy Markdown
Contributor Author

Review: challenge 74 — N-body Gravitational Force

Reviewed against the CLAUDE.md checklist. The challenge is well-formed and the math checks out; I pushed f1b15d3 with fixes for the items below.

Fixed in f1b15d3

Item Problem Fix
reference_impl portability assert positions.device.type == "cuda" and every test tensor allocated with device="cuda" — the checklist requires no device assertions and no hardcoded device, so the reference can run on XLA/TPU Assertion removed; all 21 allocations now use device=self.device, supplied by a device: str = "cuda" constructor arg
Functional test coverage 9 cases, no 3-element edge case Added N=3 with mixed-sign coordinates and a body at the origin (10 cases total)
SVG visualization Missing, though the concept is spatial Added a dark-theme (#222) SVG showing two pairwise contributions summing to the net force on a body
Constraints No bound on position values Added -10.0 ≤ positions[i] ≤ 10.0

Verification

The platform runner is currently rejecting every submission, so I could not validate through run_challenge.py — see the note at the bottom. I verified numerically on CPU instead, with a tiled CUDA solution written and its arithmetic emulated exactly in PyTorch:

  • reference_impl matches an independent O(N²) float64 implementation of the documented formula on the example and all 10 functional tests (max abs diff ≤ 1.5e-05).
  • At the performance size N=8,192, both the reference and the emulated solution agree with a float64 ground truth; the worst deviation is 0.2% of the atol/rtol=1e-2 budget, so there is ample headroom (forces reach ~2e4 where the closest pair is 0.0185 apart, and rtol covers that).
  • Softening behaves correctly in the degenerate cases: N=1 and 4 co-located bodies both give exactly zero forces, no NaN.
  • Example in challenge.html matches generate_example_test() exactly: [[0, 0.024, 0.032], [0, -0.048, -0.064]].
  • generate_performance_test peak reference memory is ~0.4 GB at CHUNK=1024, well inside 5x of the T4's 16 GB, and the reference is a handful of batched ops (no 73-style reference timeout risk).
  • Starters: all 6 present, one parameter comment each, medium-difficulty form without the (i.e. pointers to memory on the GPU) parenthetical, parameter order consistent with get_solve_signature, JAX returns its output directly. All match the conventions in 72/73.
  • pre-commit run --all-files passes.
  • No topic overlap with existing challenges (14_multi_agent_sim is radius-limited velocity averaging in 2D, not softened 1/r² summation).

Two notes for maintainers

1. CLAUDE.md describes a ChallengeBase that this repo does not have. The guide now specifies class attributes and no __init__, with self.device provided by the base. challenges/core/challenge_base.py still has __init__(self, name, atol, rtol, num_gpus, access_tier) and no device, so that pattern fails:

FAIL: Challenge() raised TypeError
ChallengeBase.__init__() missing 5 required positional arguments:
'name', 'atol', 'rtol', 'num_gpus', and 'access_tier'

scripts/update_challenges.py calls module.Challenge() on every push to main, so a class-attribute challenge would break the deploy workflow. All 73 existing challenges use the super().__init__(...) form. I therefore kept that call and added device as a constructor parameter — this satisfies the substantive requirement (device-parameterized allocations, no CUDA-only assumptions) and works with both the current base class and a future one that accepts device. If you intend the new pattern, challenges/core/challenge_base.py needs updating first, ideally backward-compatibly so all 74 challenges can migrate together.

Minor, same theme: the guide's JAX comment template says # <params> are tensors on device, but no starter in the repo uses that wording — including 1_vector_add, which the guide names as the JAX reference. I left this challenge matching the repo (are tensors on GPU).

2. scripts/run_challenge.py cannot reach a GPU right now. Every submission returns immediately with a zeroed submission id:

INFO | Status: error | Output: Unsupported GPU:

This is not specific to challenge 74 — a known-good 1_vector_add solution against the already-deployed challenge 1 fails identically, and the GPU name never appears after the colon, in the default "NVIDIA TESLA T4", in "NVIDIA Tesla T4", or when the field is moved elsewhere in the payload. The runner appears to be receiving an empty GPU name, so either the API's expected payload shape has drifted from this script or the dispatcher is misconfigured. Worth a look independently of this PR, since challenge validation depends on it.

🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant