HeadlinesBriefing favicon HeadlinesBriefing.com

livenerf: Pelacakan Benchmark Kemampuan Model Setelah Rilis

Hacker News •
×

livenerf is a deterministic, append-only benchmark designed to detect whether a frontier model gets quieter worse after launch. It addresses reports that Anthropic "nerfs" models post-release through quantization, routing, or effort changes. Starting with Claude Opus 5.5 (released 2026-09-22), the tool runs daily on a Claude Max subscription via headless Claude Code, using frozen prompts, pinned CLI (2.1.280), and exact graders to ensure reproducibility.

Built on Inspect from the UK AI Security Institute, its statistics follow Anthropic's Adding Error Bars to Evals methodology. The 30-day series began 2026-09-24, with days 1–10 as baseline. As of 2026-09-29, 6 of 30 days are collected, all running 90 samples on harness hash 461391b6fce64167.

The panel of 78 questions—selected from 2,336 GPQA Diamond, MMLU-Pro, competition-math, and AIME 2025–26 items—shows Opus 5.5 at ~93% accuracy initially. It can detect ~7.5-point accuracy shifts per 10-day window, but cannot distinguish Opus 5 from Opus 5.5 at 99% confidence. Lower effort reduces output tokens by 62% and accuracy by 8.3 points.

Eight answer keys appear wrong; 30 questions are ambiguous. Serving path issues cause some refusals, which are logged and excluded.