HeadlinesBriefing favicon HeadlinesBriefing.com

Artificial Analysis Intelligence Index v4.2 Released

Hacker News •
×

Artificial Analysis has released Intelligence Index v4.2, an interim update accelerating elements of the planned v5 release to keep pace with rapid frontier model advances. The update introduces more complex, realistic tasks and increases private, held-out test sets to 40% of Index weighting — double the v4.1 figure — to prevent gaming. Two major evaluations are added: AA-Briefcase, an in-house agentic knowledge work benchmark with private test sets testing multi-week projects across thousands of source files, and Surge AI's GDP.pdf, evaluating long-context document reasoning across 4,592 PDF pages with 1,275 expert-authored atomic criteria.

Grading infrastructure upgrades include improved scoring accuracy in AA-LCR v1.1, stabilized Elo ratings for GDPval-AA v2 and AA-Briefcase, and more robust grading sandboxes for Sci Code. Key results show Anthropic's Claude Fable 5.1 leading the Index, followed by Open AI's GPT-6 Astra with a 4-point gain over GPT-5.6 Sol. Meta ranks third, followed by Space XAI, Moonshot/Kimi, Z. AI, and Google.

The Cost per Task Pareto frontier is shared by Anthropic, Open AI, Meta, and Z. AI, while GPT-6 Astra dominates the output token efficiency frontier near the intelligence frontier.