HeadlinesBriefing favicon HeadlinesBriefing.com

MiMo-V2.6-Pro - Intelligence, Performance & Price Analysis | Artificial Analysis

Hacker News •
×

Artificial Analysis released Intelligence Index v4.3.2, incorporating ten evaluations including AA-Briefcase v1.1, GDPval-AA v2.1, Automation Bench-AA, Terminal-Bench 4.0, Sci Code, Humanity's Last Exam, GDP.pdf, Crit Pt, AA-Omniscience, and AA-LCR v1.1. The updated benchmark measures model performance across agentic knowledge work, coding, and document reasoning. The Intelligence Index compares weighted average cost per task in USD against capability scores.

Models are labeled 'Commercial Use Restricted' or 'Non-commercial' based on licensing terms. The AA-Briefcase v1.1 benchmark aggregates rubric pass rate, analytical quality Elo, and presentation Elo. The AA-Omniscience Index measures knowledge reliability and hallucination on a scale from -100 to 100.

The Openness Index assesses model openness from 0 to 100. These standardized metrics allow for direct comparison of AI capabilities across diverse use cases and price points.