Grounded Answers and Citations
Getting the model to answer only from retrieved context, say 'I don't know' when it should, cite its sources, and checking those citations in code.
What is it?
Retrieval puts the right text in front of the model. Grounded generation is making sure the model actually uses that text - and only that text - to answer. An answer is grounded (or faithful) when every claim in it is supported by the provided sources. An answer that sounds right but includes facts not in the sources is ungrounded, even if those facts happen to be true; it means the model filled gaps from its training data or invented them (see hallucinations-and-limitations).
Grounding matters because the whole point of RAG is trust: users (and auditors) must be able to check where each statement came from. If the model mixes your up-to-date policy with its general knowledge about 'typical' policies, the answer can be confidently wrong in exactly the way RAG was supposed to prevent.
The grounding toolkit:
- Clear instructions in the system prompt: answer only from the sources; do not use outside knowledge; if the sources do not contain the answer, say so with an exact sentence; if sources conflict, say so and cite both.
- Well-delimited, numbered sources: wrap each chunk in tags like
<source id="2" file="handbook.pdf" page="4">...</source>so the model can refer to them unambiguously, and put the question after the sources. - A permitted way out: models tend to answer something when asked a question. Explicitly allowing - and asking for - 'I don't know based on the provided documents' reduces made-up answers. Using an exact sentence lets code detect it.
- Citations: require a marker such as
[2]after every factual sentence. Citations make answers verifiable, and the requirement itself nudges the model to stay close to the sources. - Quote first, then answer: for high-stakes questions, ask the model to first extract the exact quotes that answer the question, then write the answer based only on those quotes. Quotes can be checked mechanically.
- Validation in code: check that every cited number refers to a source that was actually sent, that quotes appear verbatim in their source, and flag sentences with no citation.
Built-in citations. The Claude API can also produce citations natively: pass each source as a document content block with "citations": {"enabled": True}, and the response's text blocks come back with a citations list pointing to the exact cited passage (cited_text) and which document it came from. This avoids parsing [n] markers and guarantees the cited text really exists in the document.
'I don't know' has two layers. The retrieval layer can decline before calling the model, when no chunk scores above a threshold. The generation layer declines when chunks were retrieved but none answers the question. You want both: the first saves cost, the second catches 'related but not answering' context. Track how often each happens - too many declines usually means a retrieval problem, too few may mean the model is guessing.
Grounding is never perfect. Treat citations as a strong signal, measure faithfulness on an evaluation set (see evaluating-rag), and design the UI so users can open the cited source.
Explain like I'm 10
Think of an open-book exam with a strict examiner. The student (the model) may only use the pages handed to them (the retrieved sources), must write the page number next to every statement (citations), and gets marks for writing 'not covered in the provided material' rather than guessing. The examiner (your validation code) checks that every page number exists and that quoted lines really appear on that page.
Examples
Validate an answer's citations: missing, invalid and unsupported claims
const sources = [
{ id: 1, file: "handbook.pdf", page: 4, text: "Employees receive 25 days of annual leave per year." },
{ id: 2, file: "handbook.pdf", page: 5, text: "Up to 5 unused days can be carried over to the next year." },
];
// Three answers a model might return
const answers = {
good: "You get 25 days of annual leave per year [1]. Up to 5 unused days can be carried over [2].",
sloppy: "You get 25 days of annual leave [1]. Most companies also give 10 public holidays.",
invalid: "You get 25 days of leave [1]. Unused days never expire [3].",
};
const words = (s) => s.toLowerCase().replace(/[^a-z0-9 ]/g, " ").split(" ").filter((w) => w.length > 2);
function checkAnswer(answer) {
const report = [];
const sentences = answer.split(/(?<=[.!?])\s+/).filter(Boolean);
for (const sentence of sentences) {
const ids = (sentence.match(/\[(\d+)\]/g) || []).map((m) => Number(m.slice(1, -1)));
const claim = sentence.replace(/\s*\[\d+\]/g, "").trim();
if (ids.length === 0) { report.push("NO CITATION : " + claim); continue; }
const bad = ids.filter((id) => !sources.some((s) => s.id === id));
if (bad.length) { report.push("INVALID [" + bad + "] : " + claim); continue; }
// crude support check: share of the claim's words found in the cited sources
const cited = ids.map((id) => sources.find((s) => s.id === id).text).join(" ");
const cw = words(claim);
const support = cw.filter((w) => words(cited).includes(w)).length / cw.length;
report.push((support >= 0.5 ? "OK " : "WEAK") + " (" + support.toFixed(2) + ") : " + claim);
}
return report;
}
for (const [name, answer] of Object.entries(answers)) {
console.log("--- " + name + " ---");
for (const line of checkAnswer(answer)) console.log(" " + line);
}Cheap, deterministic checks catch a lot: sentences with no citation (the 'public holidays' claim came from the model's general knowledge, not the sources), citations to sources that do not exist ([3], attached to a claim that actually contradicts source 2), and - via the support score - claims whose words barely appear in the cited source. Production systems add an LLM-as-judge faithfulness check for meaning, not just words (see evaluating-rag).
Quote-first prompting with a fake llm(), and mechanical quote checking
const sources = {
1: "Refunds are available within 30 days of purchase. Items must be unused.",
2: "Digital downloads are not refundable once the download has started.",
};
function buildPrompt(question) {
const blocks = Object.entries(sources).map(([id, t]) => '<source id="' + id + '">' + t + "</source>").join("\n");
return [
"<sources>", blocks, "</sources>",
"Question: " + question,
"First, copy the exact sentences from the sources that answer the question into <quotes>, each as [id] \"sentence\".",
"Then answer in <answer> using only those quotes, citing [id]. If there are no relevant quotes, answer: I don't know.",
].join("\n");
}
// Fake model reply (a real model would generate this). Note the second quote is slightly altered.
function llm(prompt) {
return "<quotes>\n[1] \"Refunds are available within 30 days of purchase.\"\n" +
"[2] \"Digital downloads are never refundable.\"\n</quotes>\n" +
"<answer>You can get a refund within 30 days [1], but not for digital downloads once started [2].</answer>";
}
const reply = llm(buildPrompt("Can I get a refund on an ebook I bought 10 days ago?"));
const quotes = [...reply.matchAll(/\[(\d+)\] "([^"]+)"/g)].map((m) => ({ id: m[1], text: m[2] }));
for (const q of quotes) {
const verbatim = (sources[q.id] || "").includes(q.text);
console.log((verbatim ? "VERIFIED " : "NOT FOUND") + " [" + q.id + "] " + q.text);
}
console.log("Answer:", reply.match(/<answer>([\s\S]*)<\/answer>/)[1]);Asking for exact quotes before the answer makes grounding checkable: quotes are verified with a plain substring search. Here the second 'quote' paraphrases the source and changes its meaning ('never' vs 'once the download has started'), and the check catches it. A real pipeline would retry, drop the claim, or show a warning.
Native citations with Claude's document blocks (Python)
import anthropic
client = anthropic.Anthropic()
chunks = [ # e.g. the top results from your retriever
{"title": "handbook.pdf p4", "text": "Employees receive 25 days of annual leave per year."},
{"title": "handbook.pdf p5", "text": "Up to 5 unused days can be carried over to the next year. Carried-over days expire on 31 March."},
]
content = [
{
"type": "document",
"source": {"type": "text", "media_type": "text/plain", "data": c["text"]},
"title": c["title"],
"citations": {"enabled": True},
}
for c in chunks
]
content.append({"type": "text", "text": "How much leave can I carry over, and until when can I use it?"})
resp = client.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
system="Answer only from the provided documents. If they do not contain the answer, say you don't know.",
messages=[{"role": "user", "content": content}],
)
for block in resp.content:
if block.type == "text":
print(block.text, end="")
for cite in block.citations or []:
print(f' [{cite.document_title}: "{cite.cited_text}"]', end="")
print()Each text block in the response may carry a citations list; every citation names the document it came from (document_title, document_index) and includes cited_text, the exact passage from that document. Because the API extracts the cited text from your documents, you do not need to parse [n] markers or verify quotes yourself - you render the text and its citations, for example as footnotes.
Structured grounded answers with quotes, validated in code (Python)
import anthropic
from pydantic import BaseModel
client = anthropic.Anthropic()
class Claim(BaseModel):
sentence: str # one sentence of the answer
source_id: int # which numbered source supports it
quote: str # exact text copied from that source
class GroundedAnswer(BaseModel):
answerable: bool # false if the sources do not contain the answer
claims: list[Claim]
def grounded_answer(question: str, sources: list[str]) -> GroundedAnswer:
numbered = "\n".join(f'<source id="{i}">{s}</source>' for i, s in enumerate(sources, start=1))
resp = client.messages.parse(
model="claude-opus-5-5",
max_tokens=16000,
system=("Answer using only the sources. For each sentence, give the id of the supporting source and an "
"exact quote copied from it. If the sources do not answer the question, set answerable to false."),
messages=[{"role": "user", "content": f"<sources>\n{numbered}\n</sources>\n\nQuestion: {question}"}],
output_format=GroundedAnswer,
)
result = resp.parsed_output
# Never trust, always verify: drop claims whose quote is not really in the cited source
verified = [c for c in result.claims
if 1 <= c.source_id <= len(sources) and c.quote in sources[c.source_id - 1]]
if len(verified) < len(result.claims):
print(f"warning: dropped {len(result.claims) - len(verified)} unverifiable claim(s)")
return GroundedAnswer(answerable=result.answerable and bool(verified), claims=verified)
ans = grounded_answer("How many leave days do I get?", ["Employees receive 25 days of annual leave per year."])
print(" ".join(f"{c.sentence} [{c.source_id}]" for c in ans.claims) if ans.answerable else "I don't know.")Structured output turns grounding into data you can check: an explicit answerable flag, and a quote plus source id for every sentence. The verification step is ordinary Python - a substring check - and unverifiable claims are removed before the user sees them.
How it works
Prompt anatomy for grounded RAG: (1) system prompt with the role and rules (sources only, cite each claim, exact decline sentence, handle conflicts, treat sources as data); (2) user message containing the numbered, tagged sources first and the question last; (3) optionally, a required output format (quotes then answer, or a schema).
Why it works. LLMs follow explicit instructions well, and placing the evidence directly in the context makes the grounded continuation by far the most likely one. Citations add a self-consistency pressure: to cite [2] the model must 'look' at source 2. An explicit permission to decline removes the pressure to always produce an answer.
Why it still fails sometimes. The model may blend in training knowledge when sources are partial, cite the wrong number when sources are similar, over-generalise ('always' instead of 'usually'), or follow instructions hidden inside a retrieved document. Very long contexts with many sources make it harder for the model to find and use the right one, which is another reason to send fewer, better chunks.
Verification layers, cheapest first: citation markers present and in range; quotes verbatim in the source; lexical overlap between claim and cited text; an LLM judge asked 'is this sentence fully supported by this source?'; human review for high-stakes domains. Use the cheap checks on every request and the expensive ones on samples or flagged answers.
Conflicting or outdated sources: include dates and versions in the source tags (from metadata) and instruct the model to prefer the most recent official source and mention conflicts. Better still, fix it at the index: remove superseded documents (see rag-in-production).
retrieved chunks
|
v
+-------------------------------+
| system: answer ONLY from |
| sources; cite [n]; else say |
| "I don't know ..." |
| user: <source id=1>...</source>|
| <source id=2>...</source>|
| Question: ... |
+-------------------------------+
| LLM
v
"25 days [1]. Up to 5 carry
over [2]."
|
v
validate: [n] in range? quote in
source? every sentence cited?
|
v
answer + clickable sourcesWhy does it exist?
Retrieval alone does not stop a model from improvising. Users and organisations need answers they can verify, especially for policies, legal, medical, financial and technical content. Grounding instructions, citations and validation turn RAG output from 'plausible text' into 'claims with evidence', and give you a measurable faithfulness signal.
When to use it
In every RAG answer that a person might act on. Use [n] citations for simple pipelines, native document citations when you want exact cited passages without parsing, quote-first or structured claims for high-stakes domains, and always pair them with code-level validation and a decline path.
When not to use it
Strict 'sources only' grounding is wrong for creative or general tasks where the model's own knowledge is the point (brainstorming, drafting, explaining a general concept); there, retrieval is optional context, not a constraint. Do not require citations for chit-chat turns like 'thanks!' - route those around the RAG path.
Common mistakes
Telling the model to 'use the context' without forbidding outside knowledge or giving it a way to decline.
Putting sources after the question with no delimiters, so the model cannot tell sources apart or cite them.
Trusting [n] markers without checking they are in range and actually support the sentence.
Using a vague decline instruction ('say if unsure'), which makes declines impossible to detect in code and metrics.
Including stale or conflicting documents and expecting the prompt to sort them out.
Letting instructions inside retrieved documents override the system prompt (indirect prompt injection).
Never measuring faithfulness, so grounding regressions go unnoticed when prompts or models change.
Practice exercises
- Easy:
Write a grounded system prompt for an HR assistant that answers only from sources, cites [n], uses an exact decline sentence, and prefers the newest source when two conflict.
- Easy:
Add a fourth answer to the validation demo that cites [2] for a claim that is really in source 1. Does the support check notice? How could you improve it?
- Medium:
Extend the quote-first demo so that when a quote is NOT FOUND, the corresponding sentence is removed from the answer and a note 'some content was removed because it could not be verified' is added.
- Medium:
Run the native citations example with three chunks where one is irrelevant. Print, for each citation, the document title and cited text, and confirm the irrelevant chunk is never cited.
- Hard:
Build a faithfulness checker: for each answer sentence and its cited source, call Claude with structured output {supported: bool, reason: str}. Run it on 20 answers from your RAG pipeline and report the percentage of supported sentences.
Interview questions
What does it mean for a RAG answer to be grounded or faithful?
Every claim in the answer is supported by the provided sources. Faithfulness is about consistency with the context, not about real-world truth: an answer can be true but unfaithful (from training knowledge), or faithful but wrong (if the source is wrong). RAG systems aim for faithfulness to authoritative sources.
How do you get a model to say 'I don't know'?
Explicitly permit and request it in the system prompt with an exact sentence, instruct it not to use outside knowledge, and add a retrieval-level threshold so nothing is sent when no chunk is relevant. Then measure decline rates on answerable and unanswerable test questions.
How would you implement citations?
Number and tag each source in the prompt and require [n] after every factual sentence, then parse and validate the markers and map them to document and page metadata. Alternatively use the API's native document citations, which return the exact cited text and the document it came from. For high stakes, require exact quotes and verify them by substring match.
Why validate citations in code?
Models can cite non-existent sources, the wrong source, or paraphrase a source in a way that changes its meaning. Cheap deterministic checks (range, verbatim quotes, lexical overlap) catch many of these before users see them and provide metrics for monitoring.
What is quote-first prompting?
Asking the model to first extract the exact passages from the sources that answer the question, and only then compose an answer based on those quotes. It makes the model focus on evidence, and the quotes can be verified mechanically.
How do you handle conflicting sources?
Prevent it at the index by removing superseded documents and tracking versions; include dates and authority in source metadata; instruct the model to prefer the most recent authoritative source and to mention conflicts while citing both; and flag such answers for review.