HeadlinesBriefing favicon HeadlinesBriefing.com

Anthropic Unveils Conceptual Reasoning Index

Hacker News •
×

Anthropic and Redwood Research introduce the Conceptual Reasoning Index (CRI), a suite of three benchmarks designed to gauge AI models’ ability to reason about high‑risk, non‑empirical questions. The CRI aggregates scores from LMCA (60%), ACCo RD (20%), and DTBench (20%) to give a single metric from 0 to 100. LMCA tests model judgments of expert‑rated arguments across philosophy, decision theory, and AI risk topics, while ACCo RD measures logical consistency in probability and preference estimates. DTBench presents 407 multiple‑choice questions probing models’ decision‑theoretic reasoning.

The CRI is currently the only public metric that rewards correct argument evaluation, consistency, and decision‑theory skill—areas where traditional supervised training falls short. Early results show Anthropic’s best models topping the chart, with scores reflecting a ceiling near 91 given human inter‑rater noise. Future updates will add new benchmarks and adjust weights as the field evolves.

Researchers can request access to the LMCA dataset via a form on conceptualreasoning.ai. The CRI will be updated continuously as new models and benchmarks are released.