The adoption of AI agents in millions of organizations is creating new opportunities for attackers to make them take malicious actions, such as exfiltrating database contents and sensitive business and personal information. In the past five months, Google and four other organizations have acknowledged vulnerabilities that exploit one agent inside a targeted network to spread harmful instructions to other internal agents. The technique is a special form of prompt injection that targets not the LLM but a particular agent.
Independent researcher Syed Anas Mohiuddin tested agents from organizations including Google, JP Morgan Chase, Weviate, Rapid7, the French government’s interministerial digital directorate, and the US federal government. His proof-of-concept attacks exploit trust gaps in MCP, short for Model Context Protocol. Many special-purpose agents lack guardrails, and since MCP servers store credentials for each agent, an exploit that would have been rejected by the LLM succeeds.
“AI agents give attackers a fresh set of connections to walk across,” Douglas Mc Kee, director of vulnerability intelligence at Rapid7, told Ars. CVE-2026-97228, the vulnerability Mohiuddin found in Rapid7’s network, carried a severity rating of only 2.7 out of 10. The vulnerability affecting Google was more severe, with a rating of 8, stemming from an MCP toolbox initializing its HTTP client with no Check Redirect policy.
Mohiuddin is calling the class of attack “protocol pivoting” because exploits work when an app uses MCP to assign a task to an agent, which forwards malicious instructions to another agent using a different communication method. Trust gets lost in translation.
Source: Ars Technica · Summarized by HeadlinesBriefing