HeadlinesBriefing favicon HeadlinesBriefing.com

Kimi K3 trails Fable 5 on AA-Briefcase benchmark

Hacker News •
×

Kimi (Moonshot AI) released Kimi K3, a 2.8T parameter model scoring 57 on the Artificial Analysis Intelligence Index, comparable to Opus 4.8 and GPT-5.5.

On the AA‑Briefcase agentic knowledge work benchmark, Kimi K3 achieves an Elo of 1543, second only to Claude Fable 5 (1574) and ahead of GPT‑5.6 Sol (1501), Claude Sonnet 5 (1388) and Opus 4.8 (1347). This is a +727 gain over Kimi K2.6.

Running Kimi K3 costs about $10.57 per task, among the most expensive, driven by token pricing and an average of 83 turns. It averages 56.4 minutes per task, roughly 2.5× slower than Fable 5.

Analytical quality Elo is 1754, close to Fable 5’s 1744, while presentation Elo is 1471, lower than GPT‑5.6 Sol’s 1660 and Opus 4.8’s 1492. Rubric pass rate is 51%, second to Fable 5’s 56%.