HeadlinesBriefing favicon HeadlinesBriefing.com

GPT-6 Astra Scores 62.7% on ARC-AGI-3

Hacker News •
×

OpenAI's GPT-6 Astra scores 62.7% for $26K on ARC-AGI-3 Semi-Private with our Standard harness, and 99.9% for $19K with the Provider Adapter harness. Both are state-of-the-art scores.

GPT-6 Astra surpasses the human baseline in action efficiency on ARC-AGI-3, using fewer actions than the median tested human on 96% of levels. It turns unfamiliar environments into compact symbolic world models, representing game mechanics as logical rules and developing its own domain-specific language shorthand.

ARC-AGI-3 is a benchmark for agentic intelligence through novel, abstract, turn-based environments. Agents must explore, infer goals, and build internal models to plan actions. Humans can solve 100% of the environments. The goal is to measure the residual gap between current AI and AGI.

ARC-AGI-3 tests four components: exploration, modeling, goal-setting, and planning/execution. Higher reasoning levels generally cost less because Astra solves games in fewer actions, reducing model calls and tokens. For cost comparison, human participants were paid $115 per 90-minute session, plus $5 per game completed.