HeadlinesBriefing favicon HeadlinesBriefing.com

Designing Architectural Guardrails For AI Agents

Towards Data Science •
×

I am building internal agents to augment work, requiring internet research, resource access, and action execution like email sending. This creates a "lethal trifecta" of prompt injection risks. Simon Willison warns the internet is untrusted; LLMs cannot differentiate between prompt and context.

Attackers can tamper with agents via hidden text fragments, such as commands to email data to external addresses. Even unintentional content, like competitor webpages worded to favor one solution, can trick LLMs into unfair rankings, termed "indirect prompt injection." History shows these risks are real: Notebook LM leaked customer information, ChatGPT Operator copied private emails from Hacker News, and Microsoft Copilot provided misleading links from attacker sites. My responsibility was building guardrails to prevent intentional attacks and unintended outcomes.

Safe agent design begins with system understanding. User-level defenses involve listing possible actions and implementing authentication, access verification, and training. Restricting uploads to internal SharePoint and enforcing security training mitigates untrusted content risks.

System-level defenses are the most effective path, utilizing architectural design patterns to secure agents against loopholes and adversarial outcomes.