HeadlinesBriefing favicon HeadlinesBriefing.com

Frontier Models Physics Benchmarks Broken

Hacker News •
×

Low reported scores on leading physics benchmarks, including those in the Artificial Analysis Intelligence Index (2026), suggest frontier language models still struggle with advanced physics. Yet this does not always align with domain experts' experiences.

We revisit these findings by evaluating frontier models on six widely used physics benchmarks and auditing them with experts, focusing on text-only problems with verifiable final answers. Faculty and graduate researchers review problem statements, reference solutions, and model responses to distinguish genuine model errors from grader errors, incorrect reference solutions, and ambiguous questions.

Most audited cases initially evaluated as incorrect reflect benchmarking issues rather than errors in the models' physics reasoning. Experts then correct erroneous reference solutions and repair or exclude flawed questions. GPT-5.6-Sol's measured mean@4 rises from 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark, while its corrected pass@4 reaches 94.4% on the 54 retained Crit Pt challenges.

Corrected scores are computed on retained evaluation subsets after expert review. Scores on audited subsets of UGPhysics, PRISM-Physics, and PHYBench also rise substantially. These findings suggest current benchmarks substantially understate frontier models' ability to solve well-posed physics problems. Near-saturation on these closed-ended tasks highlights the need for more demanding, expert-validated evaluations.