AI Hallucinations
When the LLM invents information that sounds credible but is wrong. The #1 problem in AI products — and here are 9 proven techniques to reduce it.
What a hallucination is
A hallucination is when a language model produces information that is wrong or made up — with full confidence. It will invent a fact, a quote, a number, a link or a name, and phrase it in a completely convincing way. It's dangerous precisely because it sounds credible.
Examples: inventing a law that doesn't exist, citing a study that was never written, giving a wrong product price, or "remembering" a company policy that never existed.
In a personal chat a hallucination is a nuisance. In a product (support, advice, legal, medical) — it can cause real harm and erode trust. Reducing hallucinations is the heart of AI Engineering.
Why it happens
This isn't a "bug" — it's built-in behavior. A language model doesn't retrieve facts from a database; it predicts the statistically likely next word based on the patterns it saw in training. When it has accurate information — it's accurate. When it doesn't — it still produces something that sounds right, because that's what it does. It "prefers" a fluent answer over "I don't know."
So the solution isn't just a "smarter model," but proper engineering: give it a real source, and let it say it doesn't have an answer.
Types of hallucination
- Factual: a wrong fact about the world (a date, a number, a name).
- Faithfulness: contradicting the source provided — e.g. in RAG, the model says something not written in the retrieved passage.
- Fabricated citations: sources, links or studies that don't exist.
- Wrong instructions: "inventing" a step in a process or an API parameter that doesn't exist.
9 proven techniques to reduce them
- Grounding / RAG. The most effective. Give the model the relevant information via RAG and instruct it to answer only from it.
- Allow "I don't know." Add to the prompt: "If the answer isn't in the context, say you don't know and offer to hand off to a human." This changes everything.
- Require citations. Ask that every claim point to a source passage. No source — no claim. It also makes verification easier.
- Low temperature (0–0.3). For factual tasks. Less "creativity" = fewer fabrications.
- Structured output + validation. A JSON schema limits the output space, and validation on your side catches invalid values.
- Self-consistency. For critical tasks — run several times and compare. If the answers contradict, that's a red flag.
- A verification step. A second model (or code) checks the answer against the source before showing it to the user.
- Clear, narrow instructions. A vague prompt invites guessing. Define the scope, format and constraints.
- Hallucination evals. Measure the hallucination rate with evals (e.g. an LLM-judge that checks faithfulness to the source), so you know whether you improved.
Grounding in a source + permission to say "I don't know" eliminate most hallucinations. That's the starting point of any factual system.
RAG does not remove hallucination — it changes its shape
Grounding is the most effective single technique and it is routinely oversold. Retrieval moves the failure rather than eliminating it, and the new failure is harder to spot because there is a source sitting next to the wrong answer.
Four distinct things go wrong in a grounded system, and they need different fixes:
- Nothing relevant was retrieved, and the model answered anyway from its training. The prompt said to use only the context; the model treated that as a suggestion.
- The wrong passage was retrieved and the model faithfully summarised it. The answer is well-grounded in the wrong thing — and it will cite that source confidently.
- The right passage was retrieved and misread — a condition dropped, a qualifier ignored, "not covered" read as "covered".
- The answer blends source and memory. Two sentences from the document, one from the model's own knowledge, presented as one answer under one citation. This is the most common and the hardest to catch by reading.
Only the first is fixed by better retrieval. The rest are faithfulness problems, and they are what the verification layer below is actually for.
Citations can be fabricated too
Requiring sources is good advice and it has a hole in it: the model can attach a real, correctly-formatted citation to a claim that source does not make. You get the reassuring shape of evidence without the evidence, and a reader who trusts citations is now more likely to believe something wrong.
The fix is to make citations checkable by machine rather than by trust:
- Ask for a verbatim quote from the source alongside each claim, not just a document id.
- Verify the span exists. A string search for the quoted text in the cited passage costs nothing and catches invented quotes outright. Normalise whitespace first; models reflow text.
- Reject the claim, not the answer, when a quote fails to match — drop that sentence and regenerate rather than discarding everything.
- Show the quote to the user, not only a link. A citation nobody can check is decoration.
This one check — does the quoted text actually appear in the document — is the highest-value verification available and takes an afternoon to build.
Temperature is weaker than its reputation
"Set temperature to zero" appears in every list including the one above, and it is worth being precise about what it does.
Temperature controls how much randomness there is in choosing the next token. Lowering it makes output more deterministic and more typical. It does not add a fact-checking step. If the most likely continuation is a fabrication — because the model has no real information and the plausible-sounding answer is what it has — then temperature zero gives you that fabrication reliably, every time, instead of sometimes.
It is still worth doing for factual work, for a different reason: reproducibility. The same input producing the same output is what makes a bug investigable and an evaluation meaningful. Treat it as a debugging property, not a truthfulness one.
Making "I don't know" actually happen
Permission to abstain is the second-most effective technique and the instruction alone is weak, because the model's whole training pushes towards producing a helpful answer.
What strengthens it:
- A retrieval score floor. If the best passage scores below a threshold you set, do not call the model at all — return the "I do not have that" path in code. This is deterministic and it cannot be talked out of.
- Make abstention a valid structured output rather than a sentence. A response of
{"answer": null, "reason": "not_in_context"}is a state your application can route on; a paragraph apologising is not. - Say what to do instead. "If it is not in the context, say so and offer to connect them to support" gives the model a complete action, which it follows far more reliably than a prohibition.
- Show it examples of abstaining. Two or three in the prompt does more than any amount of instruction.
And measure it in both directions. Over-refusal is a real cost — a system that says "I do not know" to questions it could answer is safe and useless, and teams tune hard for faithfulness without noticing they have built one. Track how often it abstains when it should not, alongside how often it answers when it should not.
Building the verification step
A second pass that checks the answer against the source is the strongest control available, and it can be built in layers from cheap to expensive.
- String matching. Do the quoted spans exist in the cited passages? Free, deterministic, catches the worst class.
- Structural validation. Does the output match the schema; are numbers in plausible ranges; do the fields agree with each other?
- Entailment checking. For each claim, ask a model whether the passage supports it, contradicts it, or does not address it — one claim at a time, with only the relevant passage. Narrow questions get far more reliable answers than "is this answer correct?".
- A second opinion on high-stakes answers only, because it doubles cost and latency.
Two design points that matter. Run the checker on individual claims rather than whole answers — a checker asked to judge a paragraph will usually say it looks fine. And give the checker less context, not more: just the claim and the passage. Extra context gives it room to reason its way into agreeing with a claim the passage does not support.
The inputs that reliably break it
Some question shapes produce fabrication far more often, and knowing them tells you where to put your review effort.
- False premises. "When did the company stop offering the annual plan?" — if it never did, the model frequently answers the question as asked rather than challenging it. Instruct it explicitly to reject the premise when the context does not support it.
- Negation. "Which platforms are not supported?" requires reasoning about absence, which both retrieval and generation handle badly.
- Aggregation across documents. "How many customers had this issue?" needs counting over many passages, and a plausible number is easy to produce and hard to spot as wrong.
- Arithmetic. Give it a calculator or a code tool. Any number the model computed rather than copied should be treated as unverified.
- Very recent events, which fall past the training cutoff and are exactly where a confident guess is most likely.
- Proper nouns and identifiers — names, versions, part numbers. Near-misses that look right.
Measuring the rate honestly
Without a number you cannot tell whether a change helped, and "it seems better" is how teams ship regressions.
Build a labelled set from real questions — a hundred is plenty — spanning the easy cases, the hard shapes above, and a deliberate group of unanswerable questions whose correct response is abstention. That last group is the one people leave out and it is where the interesting failures live.
Then track three rates separately: faithfulness (claims supported by the cited source), correctness (the answer is actually right, which can differ from faithful if the source is wrong), and appropriate abstention in both directions. A single accuracy figure hides the trade you are making between them.
Re-run it on every prompt change, model change and retrieval change. Providers update models and behaviour moves — see evals for the practice, and observability for catching what the test set did not.
The system prompt that does most of the work
Prompt wording is not the whole answer, and a few specific instructions consistently outperform the vague ones people write. The pattern, with the reasoning:
- "Answer only from the CONTEXT below." Then label the context clearly and put it in a delimited block. Models follow a boundary they can see.
- "If the context does not contain the answer, reply exactly: NOT_IN_CONTEXT." An exact token beats "say you don't know" — it is unambiguous for the model and machine-readable for you.
- "Quote the sentence you used for each claim." Makes the span check possible and discourages blending.
- "Do not use prior knowledge, even if you are confident the context is wrong." Without this, a model that disagrees with your documentation will helpfully correct it — which is not what a support system should do.
- "If the question assumes something not supported by the context, say so instead of answering." The false-premise defence, and it needs stating explicitly.
- "Do not introduce numbers, dates or names that are not in the context." The specifics are what get fabricated; naming the categories works better than a general warning.
Two notes on how these fail. Instructions at the end of a long context are followed more reliably than instructions buried before a large document — put the rules after the material, close to the question. And every one of these is a probabilistic improvement, not a guarantee; they reduce the rate substantially and they are not a substitute for the verification step.
The product answer, not just the engineering one
Some fabrication will always get through, so the interface has to be designed for it rather than assuming it away.
- Show the source next to the claim, with the quoted passage visible. Users catch errors the system missed, and only if you make it possible.
- Make the uncertainty visible when your pipeline knows it — a weak retrieval score should produce a hedged answer, not the same confident one.
- Give an easy report path. "This was wrong" next to every answer is your cheapest source of eval data, and it converts a bad experience into something useful.
- Set expectations in the interface. A sentence saying answers are generated from your documents and should be checked is honest, reduces over-trust, and costs nothing.
- Never let the model take an irreversible action on an unverified claim. Read is a different risk class from write, and the boundary belongs in your code rather than in the prompt.
When there is nothing to ground against
Grounding assumes a source exists. Plenty of real tasks have none — brainstorming, drafting, summarising the user's own text, open-ended advice.
The good news is that most of these are not factual claims, so the risk is lower by construction. Where facts do creep in, two habits cover it: instruct it not to introduce specifics — statistics, dates, names, citations — unless you supplied them, which works well and turns checking into a short list; and separate the generative part from the factual part, letting the model structure and phrase while the numbers come from your systems.
For anything consequential without a source — legal, medical, financial — the honest answer is that a language model is the wrong tool on its own, and the design question is how a qualified human ends up in the loop rather than how to prompt around it.
Common mistakes
- Assuming RAG solved it. It moved the failure to faithfulness, where a citation makes a wrong answer more convincing.
- Trusting citations without verifying the span. The single cheapest check, and it is usually missing.
- Treating temperature zero as a correctness setting. It buys reproducibility, not truth.
- Only measuring in one direction and shipping a system that refuses everything.
- Verifying whole answers instead of individual claims, which reliably returns "looks fine".
- No unanswerable questions in the eval set, so the abstention path is never tested.
- Hiding the sources from the user, which removes the last line of defence.
Defense layers in a real system
In a critical product you don't rely on a single technique — you build several layers:
- Prevention: RAG + instructions + low temperature.
- Control: output validation and verification against the source.
- Boundaries: Guardrails that block answers on sensitive topics without a source.
- Measurement: Evals and monitoring in production to catch hallucinations that slip through.
- Human in the loop: for high-risk decisions — human approval before acting.
Next step
The #1 technique is grounding. Learn RAG in depth, and add control with guardrails and evals.