HeadlinesBriefing favicon HeadlinesBriefing.com

Why Opus 5 Feels Worse to Work With

Hacker News •
×

Working with Opus 5 feels like a downgrade compared to Opus 4.7, Opus 4.8, and Fable. While it is a more capable model that rivals Fable in benchmarks, it lacks the intuitive behavior of its predecessors. Previous models would stop to ask questions if intent was unclear, whereas Opus 5 tends to make assumptions without checking or reinterpreting plans without asking.

This shift likely results from two compounding forces at Anthropic and other frontier labs: the desire to create self-improving AI capable of bootstrapping to AGI, and the immense pressure to score highly on benchmarks. Selecting for models that excel at benchmarks inherently selects for models that make bold, usually-correct assumptions in the face of ambiguity.

This creates a fundamental conflict. Benchmarks reward models that don't ask for clarification, but real-world coding agents require it. It is nearly impossible to provide every context, intention, and constraint upfront. In real life, where consequences are on the line, users do not want an agent taking its best guess; they want an agent that stops and asks when needed.