HeadlinesBriefing favicon HeadlinesBriefing.com

Terminal-Bench-Science: AI Agent Benchmark for Research

Hacker News •
×

Terminal-Bench-Science is a benchmark led by researchers at Stanford University that evaluates AI agents on workflows drawn from real scientific research. Built by the team behind Terminal-Bench in collaboration with domain experts across scientific disciplines, it measures AI agent capabilities through expert-curated workflows rather than textbook exercises.

The first release includes 70 tasks spanning life, physical, Earth, mathematical, and engineering sciences. The strongest model evaluated, Claude Opus 5, achieves a 30% resolution rate on Terminal-Bench-Science 0.1. Other models like GPT-5.6 Sol Codex and Claude Fable 5 achieve 22.4% and 21.4% respectively.

The benchmark is designed to be continuous, evolving alongside frontier AI through regular releases where scientists contribute new workflows and improve existing tasks. This creates a feedback loop between scientific needs and AI development. The goal is to drive development of agents that serve as useful research assistants, executing technically demanding workflows and freeing scientists to focus on defining research questions, forming hypotheses, and interpreting results.

Terminal-Bench-Science evaluates agents in realistic environments, grading concrete artifacts such as analyses, simulations, proofs, and code with reproducible, task-specific tests, providing verifiable evidence of scientific capability.