HeadlinesBriefing favicon HeadlinesBriefing.com

Minecraft neu nachbauen ist kein Benchmark

Hacker News •
×

GPT Astra released a couple of days ago and, inevitably, within the hour my entire feed was the same five things: recreating Minecraft in one prompt, painting themselves in MS Paint, the pelican riding a bicycle as an SVG, a ball bouncing in a rotating box with believable gravity, and an SVG game controller. On paper these look like harder, more visual problems for a model to solve, there’s a reason they’re as big as they are. I’ve started calling them demo-benchmarks, visual and understandable enough for everyone to get but finite enough for the next model to be “perfect” on.

That’s the problem, these tests can’t tell you how good a model is anymore because it’s trivial for labs to optimise for exactly these tests by the next release. It’s not really their fault either, honestly I’d say it’s dumb if they didn’t - nothing sells a launch like a pelican or a 3D game controller the timeline can’t stop quoting. A fixed, famous target and eight weeks of runway is a solved pelican, these tests never change and anything that never changes can be overfit.

Every launch cycle proves it again. A test you can perfect on a schedule measures preparation instead of capability, to me that’s anti the very definition of a benchmark, it should be a hard test, something very hard to perfect.