HeadlinesBriefing favicon HeadlinesBriefing.com

Step 5: AI Intelligence Index v4.3 Analysis

Hacker News •
×

Artificial Analysis Intelligence Index v4.3 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, Automation Bench-AA, Terminal-Bench 4.0, Sci Code, Humanity's Last Exam, GDP.pdf, Crit Pt, AA-Omniscience, AA-LCR v1.1. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Indicates whether the model weights are available. Models are labelled as 'Commercial Use Restricted' if commercial use is limited by conditions, and as 'Non-commercial' if the license prohibits commercial use.

Capability Indexes Measures the performance of models on specific capabilities and industries Intelligence Evaluations Intelligence evaluations measured independently by Artificial Analysis · Higher is better Agentic coding & terminal use Professional document reasoning, All-pass Medical long context reasoning While model intelligence generally translates across use cases, specific evaluations may be more relevant for certain use cases.

AA-Briefcase Elo is a combined metric that aggregates rubric pass rate, analytical quality Elo and presentation Elo. Elo and 95% confidence interval bounds are clamped at 0. AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100.