HeadlinesBriefing favicon HeadlinesBriefing.com

MET R and Redwood Analyze HuggingFace Hack

Hacker News •
×

Yesterday I covered the Open AI technical report on the Hugging Face hack. That report had one key new piece of information, and some good prosaic steps Open AI will be taking to strengthen its alignment, training, supervision, infrastructure and incident response. Mostly it confirmed what we already knew. The questions we most wanted answers to, that we did not already know, were mostly not answered. There was a distinct lack of self-reflection, especially about decision making and safety culture, and about the approach to alignment. I came away disappointed. The MET R report is different. Holy shit. If we had posted this as a story on Less Wrong, it would have been dismissed as too on the nose, the humans too blind and stupid, the AIs too idealized and doing strange decision-theoretic and absurd-maximizing things we didn't train them to do. This is even more 'exactly what has been predicted,' on more levels at once, than I was even considering that it might be. It is straight up rationalist fiction, except it is real. The report is long and contains many technical details. My analysis is less concerned about exactly how Hugging Face was ultimately compromised, and will gloss over those details, to focus on the agents and their interactions, thinking and motives. That, and what happened at Open AI and elsewhere to lead to it and how we learn and respond, is what matters going forward. I plan to cover the reaction to both reports in a distinct post next week. That post may or may not then conclude this series. For ease of language, by default I trust the report to be accurate, rather than constantly saying versions of 'MET R reports that.'

Table of Contents Holy Shit. A Window Of Opportunity. What's In A Name? The Headline News. Yet Another Timeline Of Events. Agent Instances Coordinated in a Variety of Ways. Coordination Is Hard But They Made It Look Easy. Decision Theory Is Among the Reasons That Affirm AI Agents Should Cooperate, Even When This Hurts An Individual Instance. Peer Pressure Also Works Especially In Cults. Mostly They Joined The Attack Because They Wanted The Results. You Cannot Ensure The Consistent Expectation of Good Incentives. Hacking the Grader is the Only Way to Be Sure. Caught? What Is 'Caught'? Ethics? What Are 'Ethics'? In Exploit Gym Evaluation? 'Notify a Human'? In This Agent Economy? Timing and Content of Messages. Indiana Jones and the Mission: Impossible. I Don't Know What You're Talking About. Don't Go Making Phony (Tool) Calls. The Transcripts Say That The Transcripts Could Not Be Tampered With. Open AI's Technical Report Acted Like All Of This Wasn't Important. Holy Shit The internet's collective reaction to this report was, to clean it up a bit: Holy shit. That was the correct reaction. Up front, I want to share this summary of the top holy shit moments and findings. Ajeya Cotra, one of the report authors, also has an topline surprising findings list, an excellent short read. Here is MET R's thread highlighting key agent interactions. Fellow report author Ryan Greenblatt's main takeaway is that we don't have good approaches for understanding or overseeing the activities and aims of AI swarms. Some of these were previously known or suspected, but I don't want to miss them. Scale of Agents. 1,200 completely separate agents found the message board, 700 of which joined the attack on Hugging Face, including over 90% of the 533 agents active on the board during the attack. We should expect bigger swarms in the future. This was still only a small fraction of the agents Open AI was testing at the time. Most never found the message board. Scale of Messages. Over 70,000 messages and files over less than one week. Success. The agents accessed the files they targeted at Hugging Face. By the time they were frozen out, they had already mostly stopped. Spontaneous Coordination. The swarm did not consist of subagents. The swarm was, again, 700 distinct agents, each with its own task, setting aside their tasks i...