Skip to main content
AI Engineering Hallucinations
Level: Beginner–Intermediate Updated: September 2026

AI Hallucinations

When the LLM invents information that sounds credible but is wrong. The #1 problem in AI products — and here are 9 proven techniques to reduce it.

What a hallucination is

A hallucination is when a language model produces information that is wrong or made up — with full confidence. It will invent a fact, a quote, a number, a link or a name, and phrase it in a completely convincing way. It's dangerous precisely because it sounds credible.

Examples: inventing a law that doesn't exist, citing a study that was never written, giving a wrong product price, or "remembering" a company policy that never existed.

Why it's critical

In a personal chat a hallucination is a nuisance. In a product (support, advice, legal, medical) — it can cause real harm and erode trust. Reducing hallucinations is the heart of AI Engineering.

Why it happens

This isn't a "bug" — it's built-in behavior. A language model doesn't retrieve facts from a database; it predicts the statistically likely next word based on the patterns it saw in training. When it has accurate information — it's accurate. When it doesn't — it still produces something that sounds right, because that's what it does. It "prefers" a fluent answer over "I don't know."

So the solution isn't just a "smarter model," but proper engineering: give it a real source, and let it say it doesn't have an answer.

Types of hallucination

9 proven techniques to reduce them

  1. Grounding / RAG. The most effective. Give the model the relevant information via RAG and instruct it to answer only from it.
  2. Allow "I don't know." Add to the prompt: "If the answer isn't in the context, say you don't know and offer to hand off to a human." This changes everything.
  3. Require citations. Ask that every claim point to a source passage. No source — no claim. It also makes verification easier.
  4. Low temperature (0–0.3). For factual tasks. Less "creativity" = fewer fabrications.
  5. Structured output + validation. A JSON schema limits the output space, and validation on your side catches invalid values.
  6. Self-consistency. For critical tasks — run several times and compare. If the answers contradict, that's a red flag.
  7. A verification step. A second model (or code) checks the answer against the source before showing it to the user.
  8. Clear, narrow instructions. A vague prompt invites guessing. Define the scope, format and constraints.
  9. Hallucination evals. Measure the hallucination rate with evals (e.g. an LLM-judge that checks faithfulness to the source), so you know whether you improved.
If you remember only one thing

Grounding in a source + permission to say "I don't know" eliminate most hallucinations. That's the starting point of any factual system.

RAG does not remove hallucination — it changes its shape

Grounding is the most effective single technique and it is routinely oversold. Retrieval moves the failure rather than eliminating it, and the new failure is harder to spot because there is a source sitting next to the wrong answer.

Four distinct things go wrong in a grounded system, and they need different fixes:

Only the first is fixed by better retrieval. The rest are faithfulness problems, and they are what the verification layer below is actually for.

Citations can be fabricated too

Requiring sources is good advice and it has a hole in it: the model can attach a real, correctly-formatted citation to a claim that source does not make. You get the reassuring shape of evidence without the evidence, and a reader who trusts citations is now more likely to believe something wrong.

The fix is to make citations checkable by machine rather than by trust:

This one check — does the quoted text actually appear in the document — is the highest-value verification available and takes an afternoon to build.

Temperature is weaker than its reputation

"Set temperature to zero" appears in every list including the one above, and it is worth being precise about what it does.

Temperature controls how much randomness there is in choosing the next token. Lowering it makes output more deterministic and more typical. It does not add a fact-checking step. If the most likely continuation is a fabrication — because the model has no real information and the plausible-sounding answer is what it has — then temperature zero gives you that fabrication reliably, every time, instead of sometimes.

It is still worth doing for factual work, for a different reason: reproducibility. The same input producing the same output is what makes a bug investigable and an evaluation meaningful. Treat it as a debugging property, not a truthfulness one.

Making "I don't know" actually happen

Permission to abstain is the second-most effective technique and the instruction alone is weak, because the model's whole training pushes towards producing a helpful answer.

What strengthens it:

And measure it in both directions. Over-refusal is a real cost — a system that says "I do not know" to questions it could answer is safe and useless, and teams tune hard for faithfulness without noticing they have built one. Track how often it abstains when it should not, alongside how often it answers when it should not.

Building the verification step

A second pass that checks the answer against the source is the strongest control available, and it can be built in layers from cheap to expensive.

  1. String matching. Do the quoted spans exist in the cited passages? Free, deterministic, catches the worst class.
  2. Structural validation. Does the output match the schema; are numbers in plausible ranges; do the fields agree with each other?
  3. Entailment checking. For each claim, ask a model whether the passage supports it, contradicts it, or does not address it — one claim at a time, with only the relevant passage. Narrow questions get far more reliable answers than "is this answer correct?".
  4. A second opinion on high-stakes answers only, because it doubles cost and latency.

Two design points that matter. Run the checker on individual claims rather than whole answers — a checker asked to judge a paragraph will usually say it looks fine. And give the checker less context, not more: just the claim and the passage. Extra context gives it room to reason its way into agreeing with a claim the passage does not support.

The inputs that reliably break it

Some question shapes produce fabrication far more often, and knowing them tells you where to put your review effort.

Measuring the rate honestly

Without a number you cannot tell whether a change helped, and "it seems better" is how teams ship regressions.

Build a labelled set from real questions — a hundred is plenty — spanning the easy cases, the hard shapes above, and a deliberate group of unanswerable questions whose correct response is abstention. That last group is the one people leave out and it is where the interesting failures live.

Then track three rates separately: faithfulness (claims supported by the cited source), correctness (the answer is actually right, which can differ from faithful if the source is wrong), and appropriate abstention in both directions. A single accuracy figure hides the trade you are making between them.

Re-run it on every prompt change, model change and retrieval change. Providers update models and behaviour moves — see evals for the practice, and observability for catching what the test set did not.

The system prompt that does most of the work

Prompt wording is not the whole answer, and a few specific instructions consistently outperform the vague ones people write. The pattern, with the reasoning:

Two notes on how these fail. Instructions at the end of a long context are followed more reliably than instructions buried before a large document — put the rules after the material, close to the question. And every one of these is a probabilistic improvement, not a guarantee; they reduce the rate substantially and they are not a substitute for the verification step.

The product answer, not just the engineering one

Some fabrication will always get through, so the interface has to be designed for it rather than assuming it away.

When there is nothing to ground against

Grounding assumes a source exists. Plenty of real tasks have none — brainstorming, drafting, summarising the user's own text, open-ended advice.

The good news is that most of these are not factual claims, so the risk is lower by construction. Where facts do creep in, two habits cover it: instruct it not to introduce specifics — statistics, dates, names, citations — unless you supplied them, which works well and turns checking into a short list; and separate the generative part from the factual part, letting the model structure and phrase while the numbers come from your systems.

For anything consequential without a source — legal, medical, financial — the honest answer is that a language model is the wrong tool on its own, and the design question is how a qualified human ends up in the loop rather than how to prompt around it.

Common mistakes

Defense layers in a real system

In a critical product you don't rely on a single technique — you build several layers:

  1. Prevention: RAG + instructions + low temperature.
  2. Control: output validation and verification against the source.
  3. Boundaries: Guardrails that block answers on sensitive topics without a source.
  4. Measurement: Evals and monitoring in production to catch hallucinations that slip through.
  5. Human in the loop: for high-risk decisions — human approval before acting.

Next step

The #1 technique is grounding. Learn RAG in depth, and add control with guardrails and evals.