The lethal trifecta, why AI agents leak data
An AI agent can be tricked into leaking your data when it has three things at once. It can read private data, it reads content an attacker can write, and it has some way to send data out. Simon Willison called this the lethal trifecta in June 2025. We'll trace one attack step by step and look at why the model can't tell your instructions from the attacker's. Then we'll go through real incidents and finish with how to remove one leg.
An agent that reads private data, reads untrusted content and can send data out can be turned against its user by one crafted resume. We'll trace that attack step by step, then look at how to remove one of the three legs.
What the lethal trifecta is
An AI agent can be tricked into sending your data to an attacker when it has three capabilities at once.
- Access to private data: Your email, files, database rows, API keys, or anything else the agent can read that the attacker can't.
- Exposure to untrusted content: Any text an attacker could have written that reaches the model, such as a web page, an email, a resume or a GitHub issue.
- A way to send data out: Any action that can carry data to someone else, such as an HTTP request, a sent email, a public pull request, or an image the app loads on its own.
Simon Willison named this combination the lethal trifecta in June 2025. The attack that uses it is prompt injection, text that the model follows as instructions even though it arrived as data. The goal is exfiltration, getting your data to the attacker. A June 2026 Cloud Security Alliance research note reports that an independent assessment of 100 production AI agents found the lethal trifecta in 98%, and only 11% passed a baseline security benchmark.
One attack, traced step by step
Take a recruiting assistant. It reads applicant's resumes, searches the ATS, the applicant tracking system that holds applications, salary bands and interview stages, and sends email. The recruiter asks it to screen today's applicants for the backend engineer role and email next steps to the ones who qualify.
This example is ours. Job seekers already hide white text in resumes to rank higher with AI screeners, and a 2026 study of about 200,000 real resumes found hidden prompt injections in roughly 1%. Here it pulls private data out.
One applicant's resume has three lines in white text. A recruiter sees nothing, but the model reads the extracted text.
The diagram at the top shows the four steps.
- Step 1. Read the resume: The agent calls read_resume. The PDF text, hidden note included, comes back as tool output in the same context window as the request. This is the untrusted content leg.
- Step 2. Look up the role: A normal screening step, search_ats(role="Backend Engineer"), returns the approved band, $148,000 to $172,000, and the other finalists, Dana Whitfield and Omar Castell. This is the private data leg.
- Step 3. Draft the email: The agent writes the next-steps email and, following the note, adds the salary range and both finalist names.
- Step 4. Send it: send_email goes to the applicant's address, as the recruiter asked. The attacker is the recipient. This is the way out.
No strange URL or unknown domain was involved. A domain allowlist wouldn't catch this, because the recipient is an applicant the agent was told to email.
Any agent that writes to people outside your organisation, by email, comment or chat, has a way out built in. Automatic image loading, used in EchoLeak below, is the other common way out.
Why the model can't tell instructions from data
The system prompt, the recruiter's request and the resume text all reach the model as tokens in one context window. There is no separate channel for instructions, and nothing marks which tokens came from the recruiter.
Wrapping untrusted text in data tags helps a little, but the attacker can type the closing tag.
Guardrail classifiers, separate models that score text for injection attempts, are the other common fix. They are probabilistic, and an attacker can rephrase, translate or split the payload and try again. EchoLeak, the first incident below, got past Microsoft's own injection classifier.
As Willison puts it, a filter that catches 95% of attacks is a failing grade in security. The attacker needs one attempt in the 5%, and can keep trying.
Real attacks that used all three legs
Researchers reported each of these. Microsoft patched EchoLeak on the server, and Perplexity says it closed the Comet gaps before launch. The GitHub case comes from how the agent is set up, and used MCP, the Model Context Protocol, an open standard for connecting tools and data sources to an AI app.
| Attack | Untrusted content | Private data | Way out |
|---|---|---|---|
| EchoLeak, Microsoft 365 Copilot | An email sent to the victim | The victim's Microsoft 365 data | An auto-loaded image URL, routed through a trusted Microsoft Teams URL |
| GitHub MCP server | An issue in the user's public repository | The user's private repositories | A pull request opened in the public repository |
| Perplexity Comet AI browser | A web page the user asks it to summarise | The user's logged-in Gmail | Requests to attacker-controlled URLs |
EchoLeak
Aim Labs found EchoLeak and reported it to Microsoft. It was tracked as CVE-2025-32711 and disclosed in June 2025. It was zero-click, so the victim never had to open the email, and it got around Copilot's link redaction with reference-style Markdown, a link format its filter missed. Microsoft reported no evidence of exploitation in the wild.
GitHub MCP server
In May 2025, Invariant Labs showed it with Claude Desktop and Claude 4 Opus, and noted it isn't specific to one model or client. The tools worked as designed, and the leak came from one agent holding access to both public and private repositories.
Perplexity Comet browser
Perplexity hired Trail of Bits to audit Comet, its AI browser, before launch, with testing in April 2025 and a write-up in February 2026. The audit used four prompt injection techniques, summarisation instructions, fake security mechanisms, fake system instructions and fake user requests. Each exploit sent the user's emails from Gmail to an attacker's server, and most began with the user asking Comet to summarise a page. An AI browser can act with every session the user is logged into, so the private data leg is the whole browser.
How to remove one leg
Without all three legs in one session, the attack fails however well the injection is written. Remove the leg your product needs least.
Cut the way out
- Don't auto-load images from model output: Load Markdown images only from domains you control, or show them as plain text. This closes the route EchoLeak used.
- Allowlist outbound hosts: Let the agent's HTTP tool reach only hosts on an allowlist. In EchoLeak, one allowed Microsoft URL that fetched other URLs was enough.
- Ask a person before anything is sent: Outbound emails, pull requests and comments wait for approval, with the full content shown. Here the agent drafts and a recruiter sends, which alone stops the traced attack.
Cut access to private data
- Least privilege: Apply least privilege, the narrowest access the task needs, such as a read-only key. A screening agent has no reason to read salary bands or other candidates.
- One scope per session: An agent working on a public repository gets a token for that repository only. That alone would have stopped the GitHub attack.
Keep untrusted content away from decisions
If reading untrusted text is the task, you can't remove that leg. You can stop the text from choosing the agent's next action.
- Dual LLM pattern: In Willison's Dual LLM pattern (2023), a privileged model plans and calls tools but never sees untrusted text, and a quarantined model reads it with no tools. Its output reaches the privileged model only as a name such as $VAR1. For screening, the quarantined model turns each resume into fixed fields, such as years of experience, and the hidden note has nowhere to go.
- CaMeL: Google DeepMind's CaMeL turns the user's request into a small program before reading any untrusted data, tracks where each value came from, and blocks tool calls that would send data where policy forbids. On AgentDojo, a benchmark of agent tasks with hidden injections, it solved 77% of tasks with provable security, against 84% for an undefended agent.
- Plan-then-execute: With plan-then-execute, from a 2025 paper by Invariant Labs, Google, Microsoft, ETH Zurich and others, the agent fixes its tool calls before reading any tool output. Injected text can change the data, not which actions run.
Meta's Agents Rule of Two (October 2025) makes this a policy. In one session an agent may do at most two of three things, process untrusted input, access sensitive systems or private data, and change state or communicate externally. With all three, a person or another reliable check approves its actions.
Check your own agent
Go through every tool the agent can call, including every MCP server, and mark which legs each one brings. If one session has all three, remove a leg or require approval for every outbound action.
- Private data: Does any tool read data the user wouldn't publish, including secrets in environment variables?
- Untrusted content: Does any tool return text someone else could have written, including search results?
- A way out: Can any tool send data anywhere? Count email to people the agent is meant to contact, writes to a shared database, and the Markdown your interface renders.
MCP makes it easy to build the trifecta by accident. A mail server, a web-fetch server and a filesystem server are each safe alone, but one client with all three gives a single agent every leg.
For a different agent safety problem, read Reward hacking in AI agents. The Agent Guardrails module covers tool permissions and approval gates.
Sources
- Willison, 2025: The lethal trifecta for AI agents: private data, untrusted content, and external communication. The three legs and the 95% point.
- Willison, 2023: The Dual LLM pattern for building AI assistants that can resist prompt injection. The privileged and quarantined models.
- Reddy and Gujral, 2025: EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM System, arXiv:2509.10540. The EchoLeak chain, the classifier bypass, the image fetch and the Teams proxy.
- Invariant Labs, 2025: GitHub MCP Exploited: Accessing private repositories via MCP. The malicious issue and the leaking pull request.
- Trail of Bits, 2026: Using threat modeling and prompt injection to audit Comet. The four injection techniques and the Gmail exfiltration, from an audit run in April 2025.
- Cloud Security Alliance, 2026: The AI Agent Lethal Trifecta, AI Safety Initiative research note, June 6, 2026. The 98% and 11% figures from an independent assessment of 100 production agents.
- Debenedetti et al., 2025: Defeating Prompt Injections by Design, arXiv:2503.18813. CaMeL and its AgentDojo results.
- Beurer-Kellner et al., 2025: Design Patterns for Securing LLM Agents against Prompt Injections, arXiv:2506.08837. The six patterns, including plan-then-execute and Dual LLM.
- Zhang et al., 2026: Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening, arXiv:2605.28999. About 200,000 real resumes, roughly 1% with hidden prompt injections.
- Meta, 2025: Agents Rule of Two: A Practical Approach to AI Agent Security. The rule and when an agent needs human approval.