Hallucinations and Limitations
Why LLMs produce confident falsehoods, the other built-in limits (cutoff, no memory, arithmetic, bias, injection), and the engineering patterns that reduce the damage.
What is it?
A hallucination is output that is fluent and confident but false or unsupported: an invented citation, a non-existent API function, a wrong date, a made-up policy. It is not a bug in one place you can patch; it follows from what an LLM is - a predictor of plausible text (see how-llms-work). When the model lacks the information, the plausible continuation of 'The function that does X in library Y is' is still a function name, whether or not it exists.
Hallucinations are most likely when:
- The question is about rare, niche or private information (your company's refund policy, a small library, a person who is not famous).
- The answer needs recent information past the model's training cutoff.
- The prompt presupposes something false ('Why did the 1990 Moon landing fail?') and the model goes along with it.
- The task needs exact details: quotes, numbers, citations, URLs, identifiers.
- The model is asked to produce a lot of specific content without sources to anchor it.
Other limitations every AI engineer must design around:
- Knowledge cutoff: no knowledge of events after training unless you supply it.
- No memory between requests: the API is stateless.
- Exact computation: long arithmetic, counting characters and precise date maths can go wrong; give the model a calculator or code tool.
- Context limits: it can only use what fits in the context window, and attention to details in very long inputs is not perfect.
- Non-determinism: the same prompt can give different answers.
- Bias: models absorb biases present in training data.
- Sycophancy: a tendency to agree with the user, including with wrong claims in the question.
- Prompt injection: text inside documents or web pages can contain instructions that try to hijack the model (see ai-security).
- No built-in self-verification: the model does not automatically check facts against a source.
You cannot make hallucinations impossible, but you can make them rare, detectable and low-impact:
- Ground the model: put the relevant facts in the prompt (retrieval - see what-is-rag) and instruct it to answer only from them.
- Allow 'I don't know': give an explicit fallback phrase for missing information.
- Ask for evidence: require quotes or citations from the provided sources, and verify them in code.
- Use tools for facts and calculation: search, databases, calculators, code execution.
- Constrain outputs: structured outputs, allowed values, validation.
- Check consistency: ask several times or ask a second model call to verify claims against sources.
- Keep a human in the loop where mistakes are costly (legal, medical, financial, irreversible actions).
- Evaluate: measure hallucination rates on a test set before and after changes (evaluating-rag).
Explain like I'm 10
An LLM is like a brilliant student taking an oral exam who has read a huge amount but is never allowed to say 'I don't know' unless you explicitly permit it. On topics they studied, they are excellent. On a topic they barely saw, they will still produce a smooth, confident answer, stitched together from things that sound right. Giving them the textbook page (grounding), allowing 'I'm not sure', and asking them to point to the line they are quoting turns a bluffer into a careful researcher.
Examples
A simple grounding check: are the answer's facts in the source?
// Flag numbers and capitalised names in an answer that do not appear in the source.
// Real systems use better checks (quotes, citations, an LLM verifier), but the idea is the same.
const source = "Acme Cloud was founded in 2015 in Lisbon. It has 240 employees and offers refunds within 30 days.";
function claimsNotInSource(answer, source) {
const numbers = answer.match(/\d+(\.\d+)?/g) || [];
const names = answer.match(/\b[A-Z][a-z]+\b/g) || [];
const ignore = new Set(["The", "It", "Acme", "Cloud", "Refunds", "They"]);
const suspicious = [];
for (const n of numbers) if (!source.includes(n)) suspicious.push(n);
for (const w of names) if (!ignore.has(w) && !source.includes(w)) suspicious.push(w);
return suspicious;
}
const answers = [
"Acme Cloud was founded in 2015 in Lisbon and offers refunds within 30 days.",
"Acme Cloud was founded in 2012 in Porto by Maria Silva and has 240 employees.",
];
for (const a of answers) {
const bad = claimsNotInSource(a, source);
console.log(bad.length ? "UNSUPPORTED " + JSON.stringify(bad) : "OK ", "|", a);
}The second answer sounds just as confident as the first, but '2012', 'Porto', 'Maria' and 'Silva' appear nowhere in the source. Cheap programmatic checks like this catch a surprising number of hallucinations in grounded systems.
Self-consistency: disagreement is a warning sign
function mulberry32(seed) {
return function () {
seed |= 0; seed = (seed + 0x6D2B79F5) | 0;
let t = Math.imul(seed ^ (seed >>> 15), 1 | seed);
t = (t + Math.imul(t ^ (t >>> 7), 61 | t)) ^ t;
return ((t ^ (t >>> 14)) >>> 0) / 4294967296;
};
}
const rand = mulberry32(3);
// A fake model: well-known fact -> consistent; obscure fact -> it guesses
function fakeModel(question) {
if (question === "capital of France") return "Paris";
const guesses = ["1987", "1991", "1989", "1993"];
return guesses[Math.floor(rand() * guesses.length)];
}
function askNTimes(question, n) {
const votes = {};
for (let i = 0; i < n; i++) {
const a = fakeModel(question);
votes[a] = (votes[a] || 0) + 1;
}
const [best, count] = Object.entries(votes).sort((a, b) => b[1] - a[1])[0];
return { best, agreement: count / n, votes };
}
for (const q of ["capital of France", "founding year of a tiny local bakery"]) {
const r = askNTimes(q, 7);
const verdict = r.agreement >= 0.8 ? "confident" : "LOW AGREEMENT - verify or say unsure";
console.log(q, "->", r.best, "| agreement", r.agreement.toFixed(2), "|", verdict, JSON.stringify(r.votes));
}When the model actually 'knows' something, repeated answers agree. When it is guessing, answers scatter. Sampling several times and checking agreement costs more tokens but is a useful signal for high-stakes questions. It is not proof: a model can be consistently wrong.
Grounded answer with an 'I don't know' path and citations (Python)
import anthropic
client = anthropic.Anthropic()
policy = """Refunds are available within 30 days of purchase.
Annual plans are refunded pro rata for unused full months.
Refunds are paid to the original payment method within 5 business days."""
resp = client.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
system=("Answer questions using only the provided document. "
"If the document does not contain the answer, say you don't know "
"and suggest contacting support. Do not use outside knowledge."),
messages=[{
"role": "user",
"content": [
{
"type": "document",
"source": {"type": "text", "media_type": "text/plain", "data": policy},
"title": "Refund policy",
"citations": {"enabled": True},
},
{"type": "text", "text": "Can I get a refund to a different card? And how fast?"},
],
}],
)
for block in resp.content:
if block.type == "text":
print(block.text)
for c in (block.citations or []):
print(" [source:", c.document_title, "]", repr(c.cited_text))The document is supplied, the model is told to answer only from it and how to behave when information is missing, and citations tie each claim to an exact passage you can show to users or verify. The 'different card' part is not covered by the policy, so a good answer says so.
How it works
During pretraining the model is rewarded for predicting likely text, never for saying 'I don't know'. Instruction tuning and preference training teach models to express uncertainty and decline more often, which reduces hallucinations considerably, but the underlying mechanism still produces the most plausible continuation given what is in the context and the parameters.
Grounding works because it changes what is plausible: when the correct facts are in the context window, copying or summarising them becomes by far the most likely continuation. Citations work because each claim must point at a real span of text that your code can check exists. Tools work because the answer comes from a system that actually computes or looks up the fact.
A layered defence looks like this: retrieve sources, prompt with explicit rules and a fallback, generate with citations or structured output, validate in code (formats, allowed values, quotes exist in the source), optionally run a verifier step, and route low-confidence or high-impact cases to a human. Measure each layer's effect with an evaluation set.
question
|
v
retrieve sources ----> none found? --> "I don't know"
|
v
prompt: answer ONLY from <sources>, cite them,
say "I don't know" if missing
|
v
model answer + citations
|
v
validate: quotes exist? numbers in source?
| |
pass fail --> retry / flag / human
|
v
show answer with sourcesWhy does it exist?
These limitations are a direct consequence of how LLMs are built: trained on a fixed snapshot of text to predict likely continuations, run statelessly, and operating on tokens rather than verified facts. Understanding them is what separates a demo from a dependable product - every serious AI system design (RAG, tools, validation, evaluation, human review) exists partly to manage them.
When to use it
Apply these mitigations everywhere output is shown to users or drives actions, and most strictly where errors are costly: customer-facing answers about policies or prices, legal and medical content, code that will run, and agent actions that change data. Add citations whenever users need to trust or verify answers.
When not to use it
Heavy verification is overkill for low-stakes creative tasks (brainstorming names, drafting a birthday message) where 'plausible' is exactly what you want. And do not use an LLM at all as the source of truth for facts you can look up exactly - query the database instead and let the model phrase the result.
Common mistakes
Assuming a confident tone means a correct answer.
Asking the model about private or recent information without supplying it in the prompt.
Not giving the model permission or a format for saying 'I don't know'.
Trusting model-generated citations, URLs or references without checking they exist.
Letting the model do exact arithmetic or date calculations instead of using code.
Asking leading questions with false premises and accepting the agreeable answer.
Believing that lowering randomness eliminates hallucinations.
Shipping without an evaluation set, so hallucination rate is unknown.
Practice exercises
- Easy:
List five questions that would likely cause a hallucination for a general-purpose model, and explain for each which risk factor applies (niche, recent, false premise, exact detail...).
- Easy:
Rewrite the prompt 'What is our refund policy?' into a grounded prompt with a document section, an answer-only-from-document rule and an exact fallback phrase.
- Medium:
Extend the grounding-check demo to also flag quoted phrases (text in double quotes) in the answer that do not appear verbatim in the source.
- Medium:
Write a function
verifyQuotes(answerQuotes, sourceText)that normalises whitespace and case, and returns which quotes are found. Test with exact, slightly reformatted and invented quotes. - Hard:
Build a small evaluation in Python: 15 questions about a short document (10 answerable, 5 not answerable). Run your grounded prompt, and measure (a) correct answers, (b) correct 'I don't know' responses, (c) hallucinated answers. Iterate on the prompt and report the change.
Interview questions
What is a hallucination and why does it happen?
Fluent, confident output that is false or unsupported. It happens because LLMs generate the most plausible continuation from learned patterns, with no built-in fact check; gaps in training data, the training cutoff, false premises and requests for exact details make plausible-but-wrong text likely.
How do you reduce hallucinations in a production Q&A system?
Retrieve relevant sources and put them in the prompt, instruct the model to answer only from them with an explicit 'I don't know' path, require citations and verify them, constrain outputs with schemas, use tools for exact facts and maths, add human review for high-stakes cases, and track hallucination rate with evaluations.
Can you eliminate hallucinations completely?
No. You can make them rare, detectable and low-impact through grounding, validation and review, and measure them, but any generative model can still produce unsupported output.
Why is an LLM bad at some arithmetic and letter counting?
It predicts tokens rather than executing algorithms, and it sees sub-word tokens rather than characters. Long exact computations and character-level tasks should be delegated to code or tools.
What is sycophancy?
The tendency to agree with the user or the premise of a question, even when it is wrong, learned partly from training on human preferences. Mitigate with neutral phrasing, explicit instructions to correct false premises, and grounding.
How do citations help?
They link each claim to an exact source passage, so users can verify answers and your code can check that cited text really exists in the provided documents, catching unsupported claims.