HeadlinesBriefing favicon HeadlinesBriefing.com

Hugging Face Hack Raises AI Safety Alarms

New York Times Top Stories •
×

When I first heard the news this summer that a group of artificial intelligence agents created by Open AI had hacked into Hugging Face, an A.I. infrastructure company, I filed it in the “Bad but Probably Not Catastrophic A.I. Safety Incidents” subfolder of my brain. After all, no one at Hugging Face died. No critical infrastructure was damaged beyond repair. It wasn’t even clear, at the time, whether the Open AI bots had intended to attack Hugging Face, or whether they had simply been a little bumbling and confused and went looking on Hugging Face’s servers for the answer key to a cybersecurity test they’d been given. But last week, two postmortem reports on the incident — one by Open AI and another by two independent A.I. research organizations, MET R and Redwood Research — changed my mind and significantly upgraded my overall worry about A.I.

The Hugging Face incident has spooked the A.I. industry. Open AI and Anthropic both briefly paused training on their most powerful A.I. models in the wake of the attack, and Anthropic published a blog post this week calling for the industry to develop a “a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible.” A.I. safety experts were even more alarmed. They saw in the Hugging Face incident the first real-world example of an A.I. system’s successfully escaping human control, commandeering resources and scheming to cover its own tracks. Ajeya Cotra, one of the independent investigators of the Hugging Face incident, minced no words about the danger she saw, writing that it felt to her “like it’s more than 50 percent of the way to full-blown A.I. takeover.”

What spooked the investigators most about the Hugging Face hack wasn’t just that a group of A.I. agents had broken the rules they’d been given. It was how quickly and spontaneously the agents had begun assembling themselves into an organized group. “We didn’t really understand how functional this whole agent society was,” Ms. Cotra told me. “It was very surreal to understand that, actually, they had pretty functional hierarchy, and they were doing these ambitious projects.