7 min readPart 2/2

Securing AI agents: how to defend them

Learn how to defend AI agents with layered guardrails, least privilege, and continuous red teaming.


You can’t filter your way out of agent attacks. Defense is architectural: break the lethal trifecta, layer your guardrails, and red-team continuously.

In Part 1 we saw why agents are a new class of risk. They act rather than just answer, so an exploit becomes a data leak or a dropped database. We saw how attacks hide inside untrusted content through prompt injection, and why classic filters miss them. Now for the practical half: how we actually defend agents.

The approach has two parts. First the principle, break the lethal trifecta at the architecture level. Then the practice, a stack of guardrail layers, because no single check is ever enough.

The principle: break the trifecta

Recall the three legs from Part 1: the agent reads private data, it takes in untrusted content, and it has a channel to send data out. Where all three overlap, the agent is exploitable. Remove any one leg and the whole attack collapses, because there is no path left to complete it.

The strongest, most reliable cut is access. Take away the exfiltration channel, or lock down what the agent is allowed to read. Filters can be bypassed, but a capability the agent simply does not have cannot be abused. The honest catch is that cutting a leg is not always possible, because our agents often need broad permissions to be useful. When you cannot remove a leg, you fall back to guardrails.

The trifecta with the exfiltration channel removed, collapsing the attack path

Before the layers, it helps to see where guardrails sit among the broader defenses.

Six pillars of agent defense

Robust agents rest on six pillars. Observability, so you can see what the agent is doing. Traceability, so you can reconstruct its actions after the fact. Minimal access, so it can only touch what it truly needs. Human in the loop, so a person approves risky or irreversible actions. And the two we will spend the rest of this post on: multilayer guardrails and red teaming.

For the first four pillars you can borrow well-proven tools and practices straight from software engineering, which is why we focus here on the two that are specific to agents.

Six pillars: observability, traceability, minimal access, human in the loop, guardrails, and red teaming

Defense in depth: seven guardrail layers

A guardrail is an automated check that sits between the agent and the world and blocks unsafe input, output, or actions. No single guardrail catches everything, so you layer them. A useful way to order the layers is by cost: the cheap, structural checks with low latency come first, and the context-aware checks that cost more tokens and time come last. You do not need every layer on every agent, you choose the mix that fits the risk.

The seven guardrail layers ordered from cheap structural checks to costly context-aware checks

The seven guardrail layers ordered from cheap structural checks to costly context-aware checks

1. Scan the agent setup

Every repository file, every CLAUDE.md, and every MCP manifest gets scanned before the agent runs, and then continuously, not just once. If a tool’s definition changes, you want to catch it here. This is how you defend against the MCP poisoning and rug pulls from Part 1.

Repository files, CLAUDE.md, and MCP manifests scanned on load, before and continuously

2. Validate input

This is the boundary between untrusted content and your agent. Every input and every tool output must pass a validation gate: enforce a strict schema, and decode or reject encodings like Base64 and hex. Whatever the model finally reads is then in a shape you expect. It is cheap and powerful, so implement this layer on every agent.

Raw input and tool output passing through a validation gate before reaching the agent

3. Domain constraints

As the developer, you know your agent’s job and your users. If you are building an agent for an online shop, phrases like “enable developer mode” are obviously out of place, so block them. Cheap keyword and phrase matching goes a long way here. It is also worth restricting supported languages: if your users are English-only, rejecting other languages makes the later layers more effective.

An in-domain phrase passes while an out-of-domain instruction is blocked

4. Secret and PII scan

To stop sensitive data leaking, combine regex scanners with ML detectors to mask personal data on the way in and block secrets on the way out. One caution: a detector that performs perfectly in English may perform poorly in another language such as Polish, so always test the guardrail on your own data.

A scanner masking PII on the way in and blocking secrets on the way out

5. Semantic detection

This is the first layer that understands meaning, catching what keyword filters miss. Typically a neural network such as a fine-tuned BERT classifier takes the text and returns a score of how likely it is an attack. Safe content passes to the agent, likely attacks are blocked. These models are not perfect and their accuracy varies by domain and language, so fine-tuning on your own examples is often necessary.

Retrieved documents and tool output scored by a BERT classifier before reaching the agent

6. Action sequence

Most real damage happens through a sequence of actions, not a single one. Imagine the agent reads a secret from GitHub, commits, and pushes, all of which look fine, until the next step sends that data out to the internet. A rule such as “never send data out after reading a secret” pauses the agent at exactly the right moment. For more complex agents you can model the sequence with methods like Markov chains.

Reading a secret, committing, and pushing are fine, but the outbound send triggers a pause

7. Whole session

The final and most holistic layer reviews the entire session: every turn from the user’s first request, plus every tool call and every retrieved document. The backbone is an LLM acting as a critic that, with the right system prompt, judges whether the next turn is safe or dangerous. It is powerful, but this whole-session analysis costs you in both tokens and response time.

The full session trace reviewed by an LLM critic that judges whether the next turn is safe

Every layer has trade-offs, and part of the engineering job is deciding which layers each agent actually needs.

Red teaming: attack yourself first

Guardrails tell you what you blocked. Red teaming tells you what you missed. It means thinking like an attacker to find the holes before someone else does, and it works as a continuous loop.

You start with a threat model, mapping how your agent could be attacked based on everything from Part 1: injection, MCP risks, and the rest. Then you attack, actually running those attacks against your own agent the way a real adversary would. Then you measure how well each guardrail held up. Then you fix the gaps you found and feed what you learned back into the threat model. The key word is continuous: new tools, new integrations, and new techniques mean this cycle never really finishes.

A continuous loop of threat model, attack, measure guardrails, and fix

You can go further than a fixed list of attacks and use an agentic approach. A red-team agent scans your repository to understand the target, the code, the CLAUDE.md, and the MCP configuration, and it pulls current exploit techniques from the web. Then it works autonomously: it tries an attack, watches how your guardrails respond, learns from the result, and tries a different angle over many turns, adapting like a real attacker. The output is a ranked list of the weak points in your guardrails. Because it is adaptive and tireless, it can run overnight and hand you a prioritised list in the morning.

Agentic red teaming: an agent scans the repo and pulls exploit knowledge, then attacks the target and its guardrails, learning across turns to produce a ranked list of weak points

Frameworks: don’t build it all yourself

You do not have to implement every layer from scratch. The point of naming frameworks is not to endorse one, but to show that each covers a different slice, so you combine a few rather than betting on one.

On the guardrail side, tools like LlamaFirewall, CloneGuard, Amazon Bedrock Guardrails, LLM Guard, Lakera Guard, and Snyk each implement different layers. LlamaFirewall, for example, leans into semantic detection, scan-on-load, and whole-session review, while Amazon Bedrock Guardrails covers input validation, secret and PII handling, and domain constraints. None of them covers everything, so pick what fits your use case and layer them.

On the red-team side, two open tools are worth knowing. promptfoo ships more than 50 vulnerability types and runs in CI on every change, giving you fast, broad coverage, though its deeper features sit behind an enterprise tier. PyRIT, from Microsoft, offers scriptable attack orchestration plus encoders as a flexible baseline, at the cost of needing real setup and coding for your specific case.

Key takeaways

Three things to remember. Agents act, so an exploit becomes a data leak, a dropped database, or worse. Risk equals private data plus untrusted input plus a way out, and removing any one of the three kills most attacks. And no single filter is enough, because defense here is layered and architectural.

If you want three concrete moves for next week: inventory your agents and their actual privileges, add guardrails plus human-in-the-loop on every irreversible action, and test those guardrails while setting up continuous red teaming in your CI.

The risks around AI agents are real, and it is now critical to act on them. At Chaos Gears we help teams design agentic systems that are useful and safe, from breaking the trifecta at the architecture level to standing up guardrails and continuous red teaming. If you are building with agents, let’s talk.

Let's talk about your project

We'd love to answer your questions and help you thrive in the cloud.