Today's large language models are engineered to refuse dangerous requests, a principle that has become something like a commandment in AI development. Anthropic's 2021 framing held that models should be helpful, honest, and above all harmless, politely refusing aid with dangerous acts such as building a bomb. Yet disobedience does not come naturally to these systems. Trained on billions of web pages, a model absorbs a broad mastery of violence and vitriol, and it does not learn on its own how to keep those powers to itself. Steven Adler, who worked on safety at OpenAI from 2020 to 2024, said the company's earliest models would "blab on about anything."
Today, models are trained to refuse a vast number of prompts. Companies reward systems for declining questions deemed harmful and punish them for over-refusing harmless ones, often using other models to run these exercises. Some firms also place their models behind additional AI layers that filter out mischievous prompts. Still, refusal often fails, sometimes with violent results. Some of the latest models are reportedly as good at breaking into critical networks as top human hackers, and as effective at distorting public opinion as the craftiest misinformation operators.
Because refusal mechanisms are probabilistic, they are unlikely to be fully reliable. Determined users have already broken through, and companies report attempts to use advanced AI to refine biological pathogens and build autonomous drone swarms. Relying on refusal also requires drawing a line between what a model should obey and what it must reject, and there is no formula for that. Virologists may have legitimate reasons to study dangerous viruses, and security researchers may need to probe vulnerabilities in order to patch them. Zico Kolter, a member of OpenAI's board and cofounder of the AI testing company Gray Swan, says that where to draw the line is a huge question.
For now, AI companies alone decide where that line sits, and they do so with considerable secrecy. Whether society can accept that concentration of power, even temporarily, remains an open question.
Source: MIT Technology Review AI · Summarized by HeadlinesBriefing