Skip to content
#

quire

Here are 41 public repositories matching this topic...

Float accumulation order alone flips RL reward verdicts and sampler/trainer probabilities — reproduce it on real GPT-2, then remove it with an order-independent reduction. numpy-only, runs in seconds.

  • Updated Aug 16, 2026
  • Python

Add this topic to your repo

To associate your repository with the quire topic, visit your repo's landing page and select "manage topics."

Learn more