HeadlinesBriefing favicon HeadlinesBriefing.com

Why I'm Still Bearish on LLMs After Navier-Stokes

Hacker News •
×

Thank you to Claude Fable 5.1, Holden Saberhagen, Gabriel Kammer, Andres Erbsen, Alice McKean, and Tristan Wylde-Larue for comments.

The frontier labs are priced on the narrative of a fully automated drop-in replacement for knowledge workers, but current models need laborious oversight on even simple tasks. Headline demos like Navier-Stokes, FreeBSD RCEs, and the HuggingFace incident mislead. Models generalize only near training tasks, with severe caveats. Reward hacking needs rigorous specification by domain experts, whose time is expensive. Specification is its own skill; few domain experts have it. Labor costs can exceed direct implementation. Hardware shows 3:1 or 5:1 spec-to-design ratios. Many tasks require evolving specs. High-level specs are costly to verify. Navier-Stokes is the best case: Lean is audited, but soundness bugs exist. Human review doesn't scale and is vulnerable to hacks like the xz backdoor.