HeadlinesBriefing favicon HeadlinesBriefing.com

Opus 5 Benchmarked on Slop Code Bench

Hacker News •
×

The author benchmarked Opus 5, Opus 4.8, and Sonnet 5 on a subset of Slop Code Bench, a new long-horizon coding benchmark from @GOrlanski's lab at UW Madison. Unlike traditional benchmarks, Slop Code Bench reveals requirements incrementally across multiple checkpoints, testing a model's ability to evolve a codebase over time. The benchmark remains unsaturated — prior top models GPT-5.4 and Opus 4.6 scored only 11% and 17% strict pass rates.

Running three problems with 17 checkpoints total (circuit_eval, database_migration, dynamic_config_service_api), Opus 5 achieved a 24% strict pass rate (4/17), while Opus 4.8 and Sonnet 5 each scored 6% (1/17). No model completed any challenge defect-free. Opus 5 wrote 5x more functions than Opus 4.8, with all models showing significant increases in verbosity, complexity, and code smells across checkpoints.

Code quality metrics (41 deterministic measures) revealed that 98% of Opus 4.8's lines, 93% of Opus 5's, and 89% of Sonnet 5's triggered slop rules. Verbose lines rose from ~65% at checkpoint 1 to ~80% by checkpoint 8 for all models. The author concludes that today's models cannot reliably run lights-off on real software engineering tasks, and Slop Code Bench provides a meaningful signal for the next frontier.