HeadlinesBriefing favicon HeadlinesBriefing.com

Mercury 2.5 Intelligence, Performance & Price Analysis

Hacker News •
×

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, Automation Bench-AA, Terminal-Bench 4.0, Sci Code, Humanity's Last Exam, GDP.pdf, Crit Pt, AA-Omniscience, AA-LCR v1.1. The index measures open weights versus proprietary models, with labels for commercial use restrictions. AA-Briefcase v1.1 is an agentic knowledge work benchmark combining rubric pass rate, analytical quality Elo, and presentation Elo.

AA-Omniscience Index scores range from -100 to 100, measuring knowledge reliability and hallucination with no penalty for refusal. The Cost per Intelligence Index Task calculates weighted average USD cost from input, cache hit, cache write, reasoning, and answer token prices divided by task count. Model intelligence generally translates across use cases, though specific evaluations may be more relevant for certain applications.

The methodology includes a breakdown of each evaluation and how they are run.