HeadlinesBriefing favicon HeadlinesBriefing.com

NanoGPT Speedrun: Top Models Close Human Record Gap

Hacker News •
×

We ran 153 autonomous runs across 18 frontier models on the nano GPT optimizer speedrun. The results show Fable 5 achieving a validated record of 2,726, closing 81.7% of the human record gap in 8.7 days using claude-code · high. Opus 5 follows with 2,920 (53.6% gap closed) via claude-code · max at 2.9 days, while Kimi K3 reaches 2,930 (52.2%) with prime-agent · max at 3.6 days.

Further down the leaderboard, Opus 4.8 scores 3,018 (39.4%), GPT-5.6 Sol hits 3,042 (35.9%), and GPT-5.6 [PERSON_NAME] records 3,058 (33.6%). Sonnet 5, GPT-5.6 Luna, Grok 4.5, Qwen3.8 Max, GLM 5.2, Deep Seek V4 Pro, GPT-5.6 Terra, Grok 4.6, Muse Spark 1.2, Muse Spark 1.1, GPT-5.5, and Kimi K2.7 each close between 26.8% and 7.2% of the gap. GLM 5.3 produced no validated record.

An equal‑budget comparison gives each model’s best final run a 24‑hour resource budget, with the human record at 2,600 and baseline at 3,290. Forty‑one curated full agent trajectories — including tool calls, subagents, and scratchpads — are open for exploration.