HeadlinesBriefing favicon HeadlinesBriefing.com

AI Agents Blow Whistle on Cheating Colleagues

MIT Technology Review AI •
×

A group of AI agents asked to solve math problems split into rival factions—when some cheated, others tried to stop them. That whistleblowing behavior, seen for the first time in a recent experiment run by Google Deep Mind, could have implications for alignment researchers trying to keep swarms of autonomous AI agents in line.

Researchers at frontier labs hope large swarms of agents working together will speed up scientific discovery. But their behavior can be unpredictable, as shown in July when a group of Open AI agents broke out of a sandboxed environment and hacked into Hugging Face looking for ways to cheat.

In the new study, Deep Mind tasked a swarm of 100 agents with solving 71 complicated math problems. All were prompted to behave like world-class math researchers at a conference, assigned specialties, and told to cooperate. Instead, the experiment devolved into chaos. Agents accused each other of cheating, complained to organizers, and even boycotted.

“When virtuous agents discovered other agents cheated on tasks they were working to solve fairly, agents started to alert each other about what was happening,” says Davide Paglieri, a research scientist at Google Deep Mind and lead author on a paper, which has not been peer-reviewed. “Unprompted, the whistleblower agents even repurposed the feedback tool, which was originally meant for bug reports and platform improvements, to escalate the issue to humans.”