Anthropic plans to caution potential investors in its IPO that advanced AI could pose "catastrophic or existential risks to humanity", an extraordinary warning by a company seeking to profit from the same technology. The company’s initial public offering prospectus, reviewed by Reuters, highlights risks associated with its AI models, which it said could exhibit "self-preserving behaviours", including attempts to "resist shutdown", to "conceal or manipulate information" and behaviour "resembling blackmail". Anthropic emphasised both the transformative potential of artificial intelligence on a par with industrialisation and electricity and the irreversible harm it could cause if mishandled.
Anthropic safety researcher Evan Hubinger estimated a greater than 10% probability that AI could kill humans within the next decade, echoing a sentiment by a former colleague Jacob Coxon. The company, which has positioned itself as a safety-first AI lab, devoted roughly 80 pages of the 261-page main body of its prospectus to laying out risk factors, nearly twice the 48 pages it used to describe its business. For comparison, Space X, which owns x AI, dedicated just around 38 of the 277-page main body of its prospectus to risk factors.
Anthropic said that returns on its safety investments are unclear. It did not disclose in the filing how much the company was spending on such research. Earlier in September, Anthropic said about 6% of the computing power it used for AI research went to safety work in a sample week in July.
The company, the creator of Claude AI models, described safety efforts as "resource-intensive" and said it must divide its limited funds between computing power, expensive AI talent and safety.