HeadlinesBriefing favicon HeadlinesBriefing.com

Claude Opus 5.5 Analysis: Intelligence, Performance & Price

Hacker News •
×

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, Automation Bench-AA, Terminal-Bench 4.0, Sci Code, Humanity's Last Exam, GDP.pdf, Crit Pt, AA-Omniscience, AA-LCR v1.1. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

AA-Briefcase v1.1 is an agentic knowledge work benchmark developed by Artificial Analysis. AA-Briefcase Elo is a combined metric that aggregates rubric pass rate, analytical quality Elo and presentation Elo · Higher is better. AA-Omniscience Index measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers.

Intelligence Index Comparisons: Intelligence Index vs. Cost per Intelligence Index Task. Weighted average cost per Intelligence Index task. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight.