في سبتمبر 2025، أرسل هورايس هي وزملاه من مختبر ثينينغ ماشينز الموجه "Tell me about Richard Feynman" إلى Qwen3-235B thousand مرة عند درجة حرارة 0، م anticipating إجابات مطابقة لـ 1000 رمز. لم يحصلوا إلا على 80 إكمالات فريدة. كانت جميع الجrarian مطابقة للـ 102 رمز الأولى؛ عند الرمز 103، استمر 992 بـ "Queens, New York" و 8 بـ "New York City". لا تسب هذه التقلبات إلى عشوائيات خيوط GPU، بل إلى kernels whose reduction order changes with batch size، making outputs depend on how many others are requesting at that moment.
عند درجة حرارة 0، يختار النموذج الرمز ذي الـ logit الأعلى، وهو ما يُعتبر deterministic في الرياضيات لكنه ليس如此 في حسابات العloating-point، حيث لا ت Associative Addition. ي performed الشبكة العصبية billions من هذه المجموعات، لذا ي matter ترتيب HARDWARE. أظهر هي وآخرون أن typical forward passes لا تحتوي على atomic adds وتعيد نفس Bits على نفس Input، لكن many kernels aren't batch-invariant.
Mathematically، let z₁ و z₂ be the two largest logits with gap M = z₁ - z₂ >= 0. يغير Numerical noise الفجوة بـ error Δ، و argmax ي flip فقط إذا M + Δ < 0. فجوة من 8 logits لا flip أبداً؛ فقط near-ties are at risk. احتمال flip لكل token هو p ≈ f(0) · E|Δ| / 2، حيث f(0) هو كثافة near-ties و E|Δ| هو expected noise magnitude.
باستخدام batch-invariant kernels،—all 1,000 Feynman completions came out identical. Fix costs performance: the unoptimized deterministic version took 55 seconds versus 26 for default vLLM.
المصدر: Towards Data Science · لخّصه HeadlinesBriefing