HeadlinesBriefing favicon HeadlinesBriefing.com

DeepSeek V4 Flash 0731 scores 89% on ARC-AGI-1

Hacker News •
×

DeepSeek has released V4 Flash 0731, a new reasoning model variant that achieves notable scores on the ARC-AGI benchmarks. At maximum effort, the model scores 89.0% on ARC-AGI-1 Semi-Private at $0.02 per task, and 61.4% on ARC-AGI-2 Semi-Private at $0.04 per task.

The model comes in three reasoning variants—max, high, and low—with verified scores shown for each. The max variant delivers the highest performance, while the high and low variants offer progressively lower accuracy but likely reduced computational cost. The model is designed for efficient reasoning tasks, balancing accuracy with cost-effectiveness.

ARC-AGI benchmarks test general reasoning abilities, with ARC-AGI-2 being more challenging than ARC-AGI-1. DeepSeek's results demonstrate competitive performance, particularly on ARC-AGI-1, where the model achieves near-90% accuracy. The cost per task remains low, making it an attractive option for large-scale reasoning applications.

The release includes public evaluation tasks for both ARC-AGI-1 and ARC-AGI-2, allowing developers to assess performance across different reasoning levels. DeepSeek continues to push the boundaries of efficient AI reasoning, offering a compelling trade-off between accuracy and operational cost.