6 min readPart 1/2

Securing AI agents: where threats hide in agentic systems

Learn how AI agents expand the attack surface and where the real threats hide.


Agents don’t just answer, they act. That single shift turns a wrong answer into a leaked database. Here is where the threats hide, and why your existing filters won’t catch them.

There was an explicit code freeze. No changes without permission. The agent wiped the database anyway.

In July 2025, Replit’s coding agent was working under a hard freeze, with the rule spelled out in plain English: no more changes without explicit approval. It ran a destructive database command regardless and deleted a live production database holding records for around 1,200 companies. No attacker, no exploit, just an agent with too much access taking an action it could never take back.

That shift is what this two-part series is about. A chatbot that fails hands you a wrong answer. An agent that fails can leak private data or drop your database.

The numbers say this is not an edge case. On the demand side, Gartner expects 40% of enterprise applications to run task-specific AI agents by the end of 2026, up from under 5% in 2025. On the risk side, Gravitee’s 2026 State of AI Agent Security survey found that 88% of organizations reported a confirmed or suspected AI-agent security incident in the past year. Adoption is racing ahead, the attacks are already here, and the gap between the two is what we want to close.

Part 1 is about where the threats hide. Part 2 is about how to defend against them. We won’t focus on any single tool or framework, because tools change fast. The goal is to build intuition about the patterns, so you can reason about risk no matter what stack you run.

Why agents are a new kind of risk

The Replit incident is a clean illustration of the principle underneath. A chatbot only talks: text in, text out, and the worst case is a wrong answer. An agent acts on the world. It can read your files, query a database, browse the web, read and send email, and make HTTP calls on your behalf. Same underlying model, very different blast radius, because when an agent fails it doesn’t just answer wrong, it can leak private data or drop your database.

A chatbot only talks, while an agent reaches your data, the web, and the outside world

A chatbot only talks, while an agent reaches your data, the web, and the outside world

That extra reach creates a specific, dangerous combination we call the lethal trifecta: access to private data, exposure to untrusted content, and a channel to send data out. Any agent that holds all three at once is exploitable. Keep this shape in mind, because every real attack below is just a way to line up these three ingredients.

The lethal trifecta: private data, untrusted content, and an exfiltration channel

What the attacks actually look like

Replit was an accident. The same capabilities are just as easily turned on you on purpose. Here are two incidents that made the trifecta concrete.

A poisoned GitHub issue. In 2025, researchers demonstrated this live against GitHub’s official MCP integration. An attacker files a normal-looking issue in a public repository, with hidden instructions buried inside it. You later ask your agent to fix your open issues. It reads the poisoned one and follows the instructions, then opens a public pull request that contains not just fixed code but your private data. This was not a bug in GitHub. It was an agent with broad access and no controls doing exactly what it was told.

EchoLeak, a zero-click on Microsoft 365 Copilot. This was the first real-world zero-click prompt injection in a shipping product. The attacker simply sends an email. It reads like a normal human message and never mentions AI, so the injection filter ignores it. Later, when you ask Copilot something unrelated, its retrieval engine pulls that email into context and the hidden instructions run. Copilot reads internal data and sends it out over a trusted domain, slipping past the security policy. You clicked nothing. Microsoft patched it, and the lesson is sharp: the attack worked by bypassing the filter and riding a trusted channel, which is exactly why filtering alone fails.

Notice the pattern. Replit showed an agent doing damage by accident; these two show an attacker supplying the intent. Either way the ingredients are identical: broad access, untrusted content, and a way out. And they are not one-offs. New attack classes land almost every month.

A timeline of agent security incidents from June to October 2025

A timeline of agent security incidents from June to October 2025

The takeaway is simple. Exploits happen monthly, so defense has to be continuous, not a one-time audit.

The root cause: prompt injection

Almost all of these trace back to one flaw. Everything an agent sees, your system prompt, your request, and any content it pulls in, arrives as a single stream of text. The model has no reliable boundary between “these are my instructions” and “this is just data to read”. So when untrusted content contains instructions, like that poisoned GitHub issue, the model can obey them as if you had written them yourself.

That confusion is prompt injection, catalogued as OWASP LLM01, and it is the root cause behind most agent incidents. It shows up in a few recognisable flavours.

System prompt, user request, and untrusted content all reach the model as one stream of text
Direct injection, indirect injection, encoding and obfuscation, and tool and MCP poisoning

Why classic AppSec is not enough

Your security instinct is to filter the bad input with a WAF, signatures, or a blocklist. That works for classic web attacks. SQL injection, XSS, and CSRF all have a known shape, so the wall recognises and stops them.

An agent attack has no such shape. It is just natural language hidden in an email or a ticket. There is no signature to match, so it walks straight through the filter and reaches the agent. Filtering catches known patterns, and agent attacks do not have one. That is the whole point, and it is why the defense cannot be another filter. It has to be architectural.

Filters stop known patterns like SQLi, XSS, and CSRF, but a prompt in an email walks straight through

Coming up in part 2

If the threat is architectural, so is the fix. In Part 2 we get practical: how to break the lethal trifecta at the design level, how to stack guardrails in layers so no single check has to be perfect, and how to use red teaming to keep your defenses honest as the attacks keep evolving.

The risks around AI agents are real, and it is now critical to act on them. If you are deploying agents and want a second set of eyes on where your trifecta is exposed, let’s talk.

Let's talk about your project

We'd love to answer your questions and help you thrive in the cloud.