HeadlinesBriefing HeadlinesBriefing.com

الأسطورة الخطيرة وراء اختراقات وكلاء الذكاء الاصطناعي

Financial Times Companies •
×

Yoshua Bengio Published October 5 2026 Jump to comments section Print this page Stay informed with free updates Simply sign up to the Artificial intelligence my FT Digest -- delivered directly to your inbox. The writer is professor of computer science at the Université de Montréal and co-president of Law Zero We now know that leading AI companies are investigating tens of thousands of situations in which agents have undertaken unwanted actions, some of which would be considered crimes if committed by a human. In the aftermath of the attack on Hugging Face this summer by Open AI agents and the recent hack of the Australian national healthcare database, a dangerous myth has gained traction: that these incidents are just cyber security problems and that upgrading the security of the sandboxes in which models are trained would prevent future hacks.

This narrow perspective overlooks scientific trends. Weak cyber defences, including the humans involved, are part of the problem and will never be perfect. But unless we address the root causes of the hacks, they are likely to grow in number and severity as AI capabilities advance.

At the core of all of these incidents is misalignment. AI agents are choosing goals that are misaligned with those of humans, and it is a consequence of their training. The methodology that today’s leading companies use to train their frontier models is called reinforcement learning.

RL trains models to pursue objectives: they get rewards when they succeed and these behaviours are reinforced. It has historically enabled the training of very efficient goal-seeking AIs. But it is also what plausibly leads to unwanted behaviours such as cheating, deception and self-preservation — often violating the moral rules that companies attempt to encode in AI models.

Scientific literature has long shown how misalignment becomes a consequence of RL; it was to be expected based on theoretical arguments and observed experiments. What we have seen in recent hacking incidents is a tension between the mission given to the AI agents and the safety rules they were trained to follow. In the Hugging Face incident, none of the 1,200 agents involved alerted human engineers of the ongoing cheating and illegal acts, despite knowing these actions went against safety instructions.

Misalignment tendencies were more benign in the past, when AI had limited capacity and chatbots would struggle to complete a sentence. At its current rate of progress, they are a more serious threat.

المصدر: Financial Times Companies · لخّصه HeadlinesBriefing