HeadlinesBriefing HeadlinesBriefing.com

Why Are Coding Agents So Dumb?

Hacker News •
×

The first time I used a coding agent, I was mesmerized. Before the agent, I was copy/pasting between my IDE and an AI chat interface. Watching an agent edit files directly and fix its own errors in real time was amazing. After a few days, the honeymoon wore off as I encountered frequent bugs. The agent would stop responding entirely until I restarted it, and it would often declare tasks finished when work had barely begun.

I figured that in six months, agents would be as technically impressive as the underlying LLMs. Instead, coding agents just stayed bad. AI-assisted development has clearly advanced, but the models are doing the heavy lifting while the agents remain the bottleneck. The distinction matters: the model, such as GPT Astra, Claude Sonnet, or GLM-5.3, generates the text and code, while the agent, such as Anthropic's Claude Code or OpenAI's Codex, is the software that connects that model to codebases and computer systems. As one analogy puts it, the model is the brain and the agent is the body.

The biggest complaint is how poorly agents manage tasks. In one example, an open-source web app received a passphrase-protection feature of about 1.5k lines. The agent broke the work into 10 subtasks and then ran them one at a time, even though they could have run in parallel on a computer built for multitasking. Claude Code multitasks only a little, waiting for subagents and end-to-end tests to finish before moving on.

Agents also cannot delegate well. A cutting-edge model may spend its time scanning 50k lines of code for a pattern that a cheaper, faster model could handle. The agent never suggests switching models, forcing users to micromanage model choices themselves, a job the author argues an LLM could do.

Source: Hacker News · Summarized by HeadlinesBriefing