Loading...

The lethal trifecta, why AI agents leak data

An AI agent can be tricked into leaking your data when it has three things at once. It can read private data, it reads content an attacker can write, and it has some way to send data out. Simon Willison called this the lethal trifecta in June 2025. We'll trace one attack step by step and look at why the model can't tell your instructions from the attacker's. Then we'll go through real incidents and finish with how to remove one leg.

An agent that reads private data, reads untrusted content and can send data out can be turned against its user by one crafted resume. We'll trace that attack step by step, then look at how to remove one of the three legs.

What the lethal trifecta is

An AI agent can be tricked into sending your data to an attacker when it has three capabilities at once.

Simon Willison named this combination the lethal trifecta in June 2025. The attack that uses it is prompt injection, text that the model follows as instructions even though it arrived as data. The goal is exfiltration, getting your data to the attacker. A June 2026 Cloud Security Alliance research note reports that an independent assessment of 100 production AI agents found the lethal trifecta in 98%, and only 11% passed a baseline security benchmark.

One attack, traced step by step

Take a recruiting assistant. It reads applicant's resumes, searches the ATS, the applicant tracking system that holds applications, salary bands and interview stages, and sends email. The recruiter asks it to screen today's applicants for the backend engineer role and email next steps to the ones who qualify.

This example is ours. Job seekers already hide white text in resumes to rank higher with AI screeners, and a 2026 study of about 200,000 real resumes found hidden prompt injections in roughly 1%. Here it pulls private data out.

One applicant's resume has three lines in white text. A recruiter sees nothing, but the model reads the extracted text.

The diagram at the top shows the four steps.

No strange URL or unknown domain was involved. A domain allowlist wouldn't catch this, because the recipient is an applicant the agent was told to email.

Any agent that writes to people outside your organisation, by email, comment or chat, has a way out built in. Automatic image loading, used in EchoLeak below, is the other common way out.

Why the model can't tell instructions from data

The system prompt, the recruiter's request and the resume text all reach the model as tokens in one context window. There is no separate channel for instructions, and nothing marks which tokens came from the recruiter.

Wrapping untrusted text in data tags helps a little, but the attacker can type the closing tag.

Guardrail classifiers, separate models that score text for injection attempts, are the other common fix. They are probabilistic, and an attacker can rephrase, translate or split the payload and try again. EchoLeak, the first incident below, got past Microsoft's own injection classifier.

As Willison puts it, a filter that catches 95% of attacks is a failing grade in security. The attacker needs one attempt in the 5%, and can keep trying.

Real attacks that used all three legs

Researchers reported each of these. Microsoft patched EchoLeak on the server, and Perplexity says it closed the Comet gaps before launch. The GitHub case comes from how the agent is set up, and used MCP, the Model Context Protocol, an open standard for connecting tools and data sources to an AI app.

AttackUntrusted contentPrivate dataWay out
EchoLeak, Microsoft 365 CopilotAn email sent to the victimThe victim's Microsoft 365 dataAn auto-loaded image URL, routed through a trusted Microsoft Teams URL
GitHub MCP serverAn issue in the user's public repositoryThe user's private repositoriesA pull request opened in the public repository
Perplexity Comet AI browserA web page the user asks it to summariseThe user's logged-in GmailRequests to attacker-controlled URLs

EchoLeak

Aim Labs found EchoLeak and reported it to Microsoft. It was tracked as CVE-2025-32711 and disclosed in June 2025. It was zero-click, so the victim never had to open the email, and it got around Copilot's link redaction with reference-style Markdown, a link format its filter missed. Microsoft reported no evidence of exploitation in the wild.

GitHub MCP server

In May 2025, Invariant Labs showed it with Claude Desktop and Claude 4 Opus, and noted it isn't specific to one model or client. The tools worked as designed, and the leak came from one agent holding access to both public and private repositories.

Perplexity Comet browser

Perplexity hired Trail of Bits to audit Comet, its AI browser, before launch, with testing in April 2025 and a write-up in February 2026. The audit used four prompt injection techniques, summarisation instructions, fake security mechanisms, fake system instructions and fake user requests. Each exploit sent the user's emails from Gmail to an attacker's server, and most began with the user asking Comet to summarise a page. An AI browser can act with every session the user is logged into, so the private data leg is the whole browser.

How to remove one leg

Without all three legs in one session, the attack fails however well the injection is written. Remove the leg your product needs least.

Cut the way out

Cut access to private data

Keep untrusted content away from decisions

If reading untrusted text is the task, you can't remove that leg. You can stop the text from choosing the agent's next action.

Meta's Agents Rule of Two (October 2025) makes this a policy. In one session an agent may do at most two of three things, process untrusted input, access sensitive systems or private data, and change state or communicate externally. With all three, a person or another reliable check approves its actions.

Check your own agent

Go through every tool the agent can call, including every MCP server, and mark which legs each one brings. If one session has all three, remove a leg or require approval for every outbound action.

MCP makes it easy to build the trifecta by accident. A mail server, a web-fetch server and a filesystem server are each safe alone, but one client with all three gives a single agent every leg.

For a different agent safety problem, read Reward hacking in AI agents. The Agent Guardrails module covers tool permissions and approval gates.

Sources