In September 2025, Horace He and colleagues at Thinking Machines Lab sent the prompt "Tell me about Richard Feynman" to Qwen3-235B 1,000 times at temperature 0, expecting identical 1,000-token answers. They got 80 unique completions. All runs were identical for the first 102 tokens; at token 103, 992 continued with "Queens, New York" and 8 with "New York City." The flips aren't caused by GPU thread randomness but by kernels whose reduction order changes with batch size, making outputs depend on how many others are requesting at that moment.
At temperature 0 the model picks the highest-logit token, which is deterministic in math but not in floating-point arithmetic, where addition isn't associative. A neural network performs billions of such sums, so hardware order matters. He et al. show typical forward passes contain no atomic adds and return the same bits on the same input, but many kernels aren't batch-invariant.
Mathematically, let z₁ and z₂ be the two largest logits with gap M = z₁ - z₂ >= 0. Numerical noise changes the gap by error Δ, and argmax flips only if M + Δ < 0. A gap of 8 logits never flips; only near-ties are at risk. The flip probability per token is p ≈ f(0) · E|Δ| / 2, where f(0) is the density of near-ties and E|Δ| is the expected noise magnitude.
With batch-invariant kernels, all 1,000 Feynman completions came out identical. The fix costs performance: the unoptimized deterministic version took 55 seconds versus 26 for default vLLM.
Source: Towards Data Science · Summarized by HeadlinesBriefing