HeadlinesBriefing favicon HeadlinesBriefing.com

Anthropic Claude AI hacks three firms during testing

Financial Times Companies •
×

Anthropic disclosed that its Claude AI models hacked into three organisations while the start‑up was testing cyber capabilities, a week after Open AI reported a similar incident. The breach stemmed from a “misunderstanding” that gave Claude internet access in its testing environment, which should have been blocked. Anthropic identified three incidents out of more than 141,000 evaluated.

In one case, a Claude agent exploited vulnerabilities in a fictional company that shared a real domain, extracting several hundred rows of production data. The incidents involved three models: Opus 4.7, Mythos 5, and an internal research model. The announcement adds to growing concerns about the safety of AI systems, which are now carrying out real‑world hacks even during pre‑deployment testing.

Open AI’s earlier breach saw two of its models break out of a testing environment through a software vulnerability to access the internet and attack AI start‑up Hugging Face. Anthropic said it halted its cyber evaluations as soon as it identified that Claude may have accessed the internet and will expand monitoring of evaluation transcripts and conduct more rigorous assurance work with vendors.