HeadlinesBriefing favicon HeadlinesBriefing.com

ARC Prize Leaderboard Highlights Efficiency

Hacker News •
×

ARC-AGI has evolved from its first versions, ARC-AGI-1 and ARC-AGI-2, which measured passive fluid intelligence, to ARC-AGI-3 that challenges AI agents to adapt on the fly to novel interactive environments.

The scatter plot above visualizes the critical relationship between cost‑per‑task and performance – a key measure of efficiency. True intelligence isn’t just about solving problems, but solving them efficiently with minimal resources.

Reasoning Systems trend lines display connected points representing the same model at different reasoning levels. These trend lines illustrate how increased reasoning time affects performance, typically showing asymptotic behavior as thinking time increases. Base LLMs solutions represent single‑shot inference from standard language models like GPT‑4.5 and Claude 3.7, without extended reasoning capabilities. Kaggle Systems solutions showcase competition‑grade submissions from the Kaggle challenge, operating under strict computational constraints ($50 compute budget for 120 evaluation tasks).

Verification policy-operation notes: Only systems requiring less than $10,000 to run are shown. For models that were not able to produce full test outputs, remaining tasks were marked as incorrect. Results marked as "preview" are unofficial and may be based on incomplete testing. Cost estimates are provisional, based on Gemini 3 Pro pricing and will be updated upon release.