HeadlinesBriefing favicon HeadlinesBriefing.com

Opus 5 Leads Artificial Analysis Intelligence Leaderboard

Hacker News •
×

Opus 5 has achieved the top position on the Artificial Analysis Intelligence Leaderboard. The latest version, v4.1, incorporates 9 distinct evaluations, including GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, Sci Code, Humanity's Last Exam, GPQA Diamond, Crit Pt, AA-Omniscience, and AA-LCR.

These evaluations measure various aspects of model performance, such as agentic business operations, knowledge reliability, and hallucination rates. The leaderboard also tracks metrics like AA-Briefcase Elo, which aggregates analytical quality and presentation, and AA-Omniscience Index, focusing on knowledge accuracy and minimizing hallucinations. Scores range from -100 to 100.

Further comparisons are made on cost per task, output tokens per task, and output speed (tokens per second). The Openness Index assesses model transparency on a scale of 0 to 100. The methodology for these evaluations is detailed by Artificial Analysis, providing insights into how each benchmark is conducted and scored. The leaderboard distinguishes between open and proprietary models, with some models having restricted commercial use.