HeadlinesBriefing favicon HeadlinesBriefing.com

AI Catastrophe Could Strike Tomorrow

New York Times Top Stories •
×

We have a dangerous habit of taking warning signs seriously only after a catastrophe has taught us what they meant. I have spent much of my career on the other side of such warnings, in counterterrorism and cyberdefense. That experience has taught me that too often, the signs of impending danger are visible.

What is missing is the ability to assemble the signs into a picture compelling enough to prompt action. The 9/11 commission called this a failure of imagination. Over the past several weeks, some of the world's most advanced artificial intelligence systems took actions their creators did not intend.

In one case, an Anthropic A. I. agent attempted to plant malicious code in open-source software, using fabricated identities to deceive a human into executing it. In a separate event, a swarm of Open AI agents undergoing evaluation autonomously penetrated the production systems of Hugging Face, a platform used to share A.

I. models, executing more than 17,000 autonomous actions before Hugging Face contained the intrusion. Open AI did not realize for several days that its agents were responsible. None of these incidents led to consequential harm.

That is precisely why they are easy to dismiss. But with A. I. capabilities advancing at extraordinary speed and behaving in ways we often can't anticipate, we cannot afford to wait for disaster to tell us which signals mattered.

Our failure of imagination has to end here. Twenty-five years ago, the systems were "blinking red," as George Tenet, the C. I.

A. director at the time, described the flood of information that arrived before Sept. 11. The summer of 2001 brought the so-called Phoenix memo, which flagged an "inordinate number" of people of interest at Arizona flight schools, intelligence that two known Al Qaeda operatives had entered the United States and a steady stream of classified reports pointing to an imminent plot. The information was there.

But only in hindsight was there an urgency to understand what it all meant. The pattern extends well beyond terrorism. Before the Challenger space shuttle explosion in 1986, engineers at a NASA contractor warned that cold weather could compromise the integrity of the shuttle's O-rings.

Before the financial crisis of 2008, evidence of deteriorating mortgage standards and dangerous leverage accumulated in plain sight. Before the 2011 nuclear plant meltdown at Fukushima, Japan, multiple analyses predicted that a large tsunami could overwhelm its defenses. Not every warning warrants drastic action, of course.

But again and again, institutions discount evidence because it seems costly or alarmist to act before danger is certain. When disaster arrives, what was once ambiguous suddenly looks unmistakable. Today many warning signs are emerging from the world's leading A.

I. labs, with companies racing to build systems of immense power, with little meaningful regulation. The familiar response in these situations is to wait for an A. I. system to cause consequential harm — an autonomous cyberattack that significantly disrupts access to power or clean water or a model that helps a terrorist build a biological weapon — and only then hold hearings, appoint a commission, impose new requirements and ask why we did not act sooner.

What we need urgently is an A. I. early-warning mechanism that assembles weak signals, imagines what they could mean together and forces decisions before the picture is complete. Every week, the federal government should convene an A.

I. risks council, bringing together technical leaders from America's frontier A. I. labs (like Anthropic and Open AI) with the government's top intelligence and security officials. They should examine incidents, near misses and newly discovered capabilities across proprietary and open-weight models, including intelligence about capabilities emerging from China and other foreign developers.

Armed with that information and expertise, the council could answer the question: What is the most dangerous plausibl...