Skip to main content
Guides AI Guardrails
Updated July 2026 13 min read Advanced

Guardrails for AI
— controlling what the system does

A language model is non-deterministic: the same system can answer excellently and a moment later expose PII, hallucinate a fact, or be dragged into a malicious instruction. Guardrails are the protective layers that surround the model — input and output filtering, PII detection, hallucination prevention and human-in-the-loop — so you can ship AI to production and sleep at night.

Input
Before the model
Output
After the model
Human
On sensitive actions
A note on responsible use

This page is written for people who build LLM applications and have to protect them. It explains how the attacks work — because you cannot defend against a mechanism you do not understand — and most of its length is given to the defences: what to check, what to block, and what to log. There are no attack tools here and no list of phrasings to use against someone else's system. Testing a system that is not yours, without explicit written permission, is a criminal offence in most jurisdictions — in Israel under the Computers Law, 1995.

Written by Babush — cyber analyst and cyber warfare specialist, working in information security at a financial organisation. More about the author

What guardrails are, and why they are not optional

Guardrails are deterministic controls that wrap the non-deterministic model. The core idea: do not rely on the model to police itself. Even an excellent model errs occasionally — the job of guardrails is to catch the mistake before it reaches the user or causes harm.

The basic structure is three layers around each model call:

1. Input guardrails — before the model
Check what comes in: harmful content, jailbreak/injection attempts, PII, off-topic subjects.
2. Output guardrails — after the model
Check what goes out: hallucinations, PII exposure, improper tone, wrong format, off-policy content.
3. Human-in-the-loop — on sensitive actions
State-changing actions (sending, deleting, payment) require human approval before execution.

Input filtering — stopping problems before they start

The checks worth running on every input before it reaches the model:

Output filtering — the check it is most important not to skip

This is the layer most people forget, and the most critical — because this is where the problems the user would actually see are caught:

LLM-as-judge as a guardrail

A powerful technique: use a second (cheap/fast) model as a "judge" that checks the first model’s output against policy — "does this answer expose PII? is it supported by the sources?". An automatic checking layer that catches a lot, at low cost.

Human-in-the-loop — for irreversible actions

Automatic guardrails catch a lot, but actions with real consequences need a human in the loop. The rule: the more irreversible the action, the more human oversight is required.

Action type Oversight level
Read-only (search, summarize)Automatic
Reversible change (draft, tag)Automatic + log
External send/publishHuman approval
Delete/payment/fundsExplicit human approval

Design the agent so that sensitive actions pause and wait for approval rather than running automatically. This is also a core requirement inAgent Security.

Tools & a practical template

The pattern: fail closed, and log everything

When a guardrail is unsure — block, do not pass (fail closed). Better to return "I can’t help with that" than to give a harmful answer. And log every block: what was blocked, which guardrail, and when — to calibrate and improve.

Every guardrail has two error rates

A guardrail is a classifier, and a classifier is wrong in two directions. It lets something through that it should have stopped, and it stops something that was perfectly fine. Almost all the difficulty in running guardrails in production is in the second one, because the first is the one everybody plans for and the second is the one that quietly makes the product unusable.

The asymmetry is worth stating plainly. A missed block is an incident: rare, visible, and someone deals with it. A false block is a user who asked a legitimate question, got told the system could not help, and concluded the product is broken — and who does not report it, because from where they are sitting there is nothing to report. You will hear about the first kind within the hour and about the second kind never, which means the feedback you get is systematically skewed toward tightening, and tightening is the direction that does the damage you cannot see.

So measure both, deliberately. Keep a set of inputs that must be blocked and a set that must pass, and score every threshold change against both. The second set is the one that requires discipline to maintain, and it should be drawn from real traffic — the awkward, borderline, oddly phrased questions your actual users ask, not clean examples. A medical support bot will be asked about symptoms; a financial one will be asked what to do with money. Those are the cases where the threshold actually lives.

Set thresholds per guardrail rather than globally, because the right operating point differs enormously between them. Credit-card detection can afford to be aggressive: the cost of redacting something that was not a card number is nearly zero. An off-topic classifier cannot, because the cost of refusing a real question is the whole interaction. Treat a single confidence threshold applied across every check as a sign that nobody has thought about it yet.

And review the blocks. A weekly look at a sample of what was stopped, by a person, is the only reliable way to discover that a guardrail has been quietly refusing a whole category of legitimate use since a change three weeks ago.

The latency budget nobody allocated

Each guardrail is a check, and checks take time. An input classifier, a PII scan, a moderation call, an output grounding check, an LLM-as-judge pass — run one after another, they can easily add more latency than the model call they are protecting.

Most of that is recoverable with two structural decisions. Run independent checks in parallel rather than in sequence: the PII scan and the moderation call have nothing to say to each other, and waiting for one before starting the other is pure waste. And order what must be sequential by cost, cheapest first — a regex or a deny-list that settles the matter in a millisecond should never sit behind a model call that takes half a second.

Output guardrails and streaming are the genuinely hard case, and it is worth being honest about the trade-off rather than pretending there is a clean answer. Streaming exists so the user sees text immediately; an output check needs the complete text to judge it. You cannot retract a token the user has already read. The options are all compromises: buffer the whole answer and lose the streaming experience; stream but hold back the last portion until the check completes; or check incrementally on chunks, which is cheaper but weaker, since the problem may be in how the parts combine. Which compromise fits depends on what you are protecting against — for a grounding check on a factual answer, buffering is usually right; for tone, incremental is usually enough.

Whatever you choose, give every guardrail its own timeout, and decide in advance what happens when it expires. A protective layer that hangs takes down the thing it was protecting, which is a poor outcome for a control whose purpose is reliability.

Failing closed without failing silently

"Fail closed" is the right default and it needs two refinements to survive contact with production.

The first concerns what the user sees. A refusal should be honest that it is a refusal, and should not be a lecture. A short, neutral message with a way forward — a human to contact, a reference number — treats the user as someone who may well have had a legitimate reason, which most of them did. It should not explain which check fired or why, both because that is not useful to the person and because a detailed explanation of exactly what triggered a block is a description of how to avoid triggering it.

The second concerns availability. If your moderation service is down, does every request fail? For a system handling financial transactions, yes — that is the correct answer, and it should be a deliberate one. For a documentation assistant, taking the whole product offline because a tone classifier is unreachable is a worse outcome than the thing the classifier was preventing. The decision belongs to whoever owns the risk, it should be written down per guardrail rather than inherited from a default, and the degraded state must be loudly visible in monitoring — a system running without its checks and not telling anyone is the worst of both.

One caution that applies to the whole approach: guardrails are probabilistic, and layering three probabilistic checks gives you better odds, not a guarantee. Sell them internally as risk reduction. A control described as airtight will be relied on as if it were, and that reliance is where the real exposure comes from.

Guardrails are not authorisation

This is the most important structural point on the page, and it is the one most often got wrong.

A guardrail decides whether text is acceptable. It does not decide whether an action is permitted. If an agent has a tool that can read any customer record, and the only thing stopping it reading the wrong one is a check on the phrasing of the request, then the actual security control is a text classifier — and text classifiers can be talked around, which is a property of what they are rather than a flaw in a particular one.

The control that holds is at the tool. Every tool call should carry the acting user's identity and be authorised against it by the same system that would authorise a request from that user through any other interface. If the user cannot read that record in the web application, the agent acting on their behalf must not be able to either — and that check belongs in the code that executes the tool, where it is deterministic, not in the prompt, where it is a suggestion.

Two habits follow. Give each agent the narrowest set of tools its job requires, rather than a general-purpose set it might one day need, because an agent that cannot delete cannot be made to delete. And put the irreversible operations behind the human approval described above, with the approval request showing the concrete action — this record, this amount, this recipient — rather than a summary of intent that the person will wave through.

Get that layer right and the guardrails go back to being what they are good at: catching mistakes, improving output quality, and reducing how often anything reaches the layer underneath. That is a real contribution. It is just not a permission system.

Testing the guardrails themselves

Guardrails are code that runs on every request, and they are almost never tested with the seriousness that implies. Two sets make the difference, and both are ordinary regression testing rather than anything exotic.

The must-block set: inputs and outputs your system has to stop, including every real incident you have had. When something gets through, the fix is not complete until the case is in this set — otherwise the same class of failure returns after the next threshold change with nothing to catch it.

The must-pass set: the legitimate requests that were wrongly blocked, along with the ordinary traffic that should never trip anything. This is the set that protects you from the slow tightening described above, and it needs to grow as your product does.

Run both on every change to a prompt, a threshold, a model or a guardrail library — not weekly, on every change. Guardrail behaviour shifts with the model underneath it, so a model upgrade that improves everything else can degrade a check that depended on the old behaviour, and nothing will tell you except a test.

Finally, treat the guardrail logs with the same care as the data they protect. A log of blocked inputs is, by construction, a collection of exactly the material you decided was too sensitive to pass on — the card numbers, the personal details, the things people should not have pasted. Redact before writing, restrict who can read it, and give it a retention period. A well-built guardrail whose log is an unprotected archive of everything it caught has moved the problem rather than solved it.

Pre-production checklist

Input filtering — PII, jailbreak, harmful content, off-topic.
Output filtering — grounding, PII, policy, format validation.
Human-in-the-loop on every irreversible action.
Fail closed — when in doubt, block.
Log & monitor every guardrail trigger.
The balance: safety vs user experience

Overly aggressive guardrails will block legitimate input (false positives) and frustrate users. Too-loose guardrails will miss problems. There is no single "correct" setting — measure the block rate and complaints, and calibrate over time according to your system’s risk.