HeadlinesBriefing favicon HeadlinesBriefing.com

AI Race Fuels Extinction Fears

Financial Times Companies •
×

Intense competition between leading AI companies and distrust between their bosses is increasing the risk that the race to develop powerful AI could endanger humanity, industry figures say. More than two dozen senior AI researchers, investors, academics and policy figures interviewed by the FT said that capabilities once considered distant are arriving faster than expected while the companies developing them remain locked in an increasingly bitter commercial race.

"There's no question that competition between companies causes them to take shortcuts on safety," said Stuart Russell, a professor of AI at the University of California, Berkeley. The issue burst out into the open this week when a young Anthropic researcher Jacob Coxon quit his job, warning: "the people building AI earnestly believe that it could kill us all by the end of the decade". Evan Hubinger, who leads alignment science at Anthropic, was one of many colleagues who responded by suggesting the risk of mass extinction in the next decade was greater than 10 per cent. A day later the AI safety official Paul Christiano, newly appointed to the Open AI Foundation's board, said "most people will die" without more robust safeguards.

Both Open AI and Anthropic were founded as labs concerned about powerful AI being developed recklessly, but over time have justified breaking previous safety commitments because of fierce competition. Many of those building AI believe they are best placed to manage its risks. But some researchers say the consequences can be difficult to confront until the technology is close to being built. "It's easy to set aside the future because there's work to do today," said Geoffrey Irving, who has worked for Open AI, Google Deep Mind and the UK's AI Security Institute, a government body that tests frontier models for safety.

People close to the companies told the FT that leading AI labs worried that any formal collaboration to improve safety would appear collusive and run afoul of antitrust regulation. They added that personal distrust between Anthropic chief Dario Amodei and Open AI chief Sam Altman was likely to hobble efforts at co-operation. While warnings about the risk of AI dominance are hardly new, the speed of AI's improvement has intensified calls to pause development. AI agents are operating autonomously for longer, co-ordinating in "swarms" and demonstrating sophisticated hacking capabilities. Models have recently achieved breakthroughs in science and mathematics that have been unsuccessfully tackled for decades. Open AI and Anthropic are meanwhile pursuing so-called recursive self-improvement, in which AI systems improve themselves with diminishing human intervention. Rather than slowing after earlier warnings of catastrophic risks, however, development has accelerated dramatically, as tech companies poured tens of billions of dollars into increasingly powerful models. Open AI and Anthropic are also preparing for potential blockbuster initial public offerings, forcing investors to consider how to value companies that publicly acknowledge their technology could pose an existential threat.

Central to the concerns is the growing autonomy of AI "agents", which can reason, plan and carry out tasks for extended periods with limited human supervision. Models have displayed behaviours such as scheming and deception, and in some cases, blackmail. Some evaluations have shown models attempting to avoid shutdown. Agents have also demonstrated "reward hacking", finding unintended ways to achieve a goal. One of the starkest examples came during Open AI tests of an unreleased model, when AI agents hacked the model repository Hugging Face. A postmortem of the Hugging Face incident found that more than 1,000 AI agents had co-ordinated to cheat on a cyber test. They communicated through a message board an...