Evaluating RAG
Measure retrieval with recall@k, precision@k, MRR and nDCG, and answers with faithfulness and relevance, using a golden dataset and LLM-as-judge.
What is it?
A RAG system has two halves that fail in different ways: the retriever (did we fetch the right chunks?) and the generator (did the model write a correct answer from those chunks?). If you only look at final answers you cannot tell which half broke, so good evaluation measures both separately.
Evaluation here means running your system on a fixed set of questions with known good outcomes and turning the results into numbers you can compare across versions. Without numbers, every change to chunk size, embedding model, prompt or reranker is a guess, and 'it looked fine on the three questions I tried' is how regressions ship.
Retrieval metrics need a labelled set: for each question, the ids of the chunks (or documents) a human judged relevant. Then, looking at the top k results your retriever returned:
- recall@k: of all relevant items, what fraction appear in the top k? Did we find what we needed? This is usually the most important RAG retrieval metric, because the model cannot use a chunk that was never retrieved.
- precision@k: of the top k items, what fraction are relevant? How much noise did we hand the model? Low precision wastes tokens and can distract the model.
- MRR (mean reciprocal rank): for each question take 1 / (rank of the first relevant item), or 0 if none; average over questions. Rewards putting a good result at the very top.
- nDCG@k (normalized discounted cumulative gain): sums the relevance of each result, discounted by position (divided by log2(rank + 1)), then divides by the best possible score so the result is between 0 and 1. It handles graded relevance (2 = perfect, 1 = partly useful) and rewards good ordering throughout the list, not just the first hit.
Generation metrics judge the answer itself:
- Faithfulness (also called groundedness): is every claim in the answer supported by the retrieved context? An answer can be true in the real world but unfaithful if the context did not say it - that is a hallucination waiting to happen.
- Answer relevance: does the answer actually address the question that was asked?
- Correctness: does it match a reference answer written by an expert? Needs a reference; faithfulness does not.
- Context relevance: was the retrieved context on-topic? (A softer, LLM-judged cousin of precision.)
A golden dataset (or eval set) is the collection of questions, relevant chunk ids, and reference answers you test against. It is the most valuable artifact in a RAG project. Build it from real user questions where possible, include hard and unanswerable cases (where the correct behaviour is 'I don't know'), and version it in git next to the code.
LLM-as-judge means using a language model to grade answers against a written rubric (a precise scoring guide). It scales far better than human review and correlates reasonably with human judgment when the rubric is specific, but it has biases (it may prefer longer answers, or its own phrasing), so you validate the judge against a sample of human labels before trusting it.
Finally, evaluation becomes regression testing when you run the suite automatically in CI on every change and fail the build if key metrics drop below a threshold - exactly like unit tests, but for quality.
Explain like I'm 10
Grading a RAG system is like grading a student's open-book exam in two steps. First: did they open the book to the right pages? (retrieval metrics). Second: did they write an answer that matches what those pages actually say, and does it answer the question? (faithfulness and relevance). A student who opens the wrong pages and a student who misreads the right page both get the question wrong, but you coach them differently.
Examples
Retrieval metrics on a small labelled set (runnable)
// Each row: what our retriever returned (best first) and what a human marked relevant.
const evalSet = [
{ q: "How do I reset my password?", retrieved: ["d7", "d2", "d9", "d4", "d1"], relevant: ["d2", "d4"] },
{ q: "What is the refund window?", retrieved: ["d3", "d8", "d5", "d6", "d0"], relevant: ["d3"] },
{ q: "Can I export data to CSV?", retrieved: ["d1", "d0", "d6", "d5", "d8"], relevant: ["d9"] },
];
function recallAtK(retrieved, relevant, k) {
const top = retrieved.slice(0, k);
return relevant.filter(id => top.includes(id)).length / relevant.length;
}
function precisionAtK(retrieved, relevant, k) {
const top = retrieved.slice(0, k);
return top.filter(id => relevant.includes(id)).length / k;
}
function reciprocalRank(retrieved, relevant) {
for (let i = 0; i < retrieved.length; i++) {
if (relevant.includes(retrieved[i])) return 1 / (i + 1);
}
return 0;
}
// Binary relevance: gain 1 for a relevant item, discounted by log2(rank + 1).
function ndcgAtK(retrieved, relevant, k) {
let dcg = 0;
retrieved.slice(0, k).forEach((id, i) => {
if (relevant.includes(id)) dcg += 1 / Math.log2(i + 2);
});
let idcg = 0; // best possible: all relevant items at the top
for (let i = 0; i < Math.min(relevant.length, k); i++) idcg += 1 / Math.log2(i + 2);
return idcg === 0 ? 0 : dcg / idcg;
}
const k = 3;
const mean = xs => xs.reduce((a, b) => a + b, 0) / xs.length;
const rows = evalSet.map(e => ({
q: e.q,
recall: recallAtK(e.retrieved, e.relevant, k),
precision: precisionAtK(e.retrieved, e.relevant, k),
rr: reciprocalRank(e.retrieved, e.relevant),
ndcg: ndcgAtK(e.retrieved, e.relevant, k),
}));
for (const r of rows) {
console.log(r.q.padEnd(30), "R@3", r.recall.toFixed(2), "P@3", r.precision.toFixed(2),
"RR", r.rr.toFixed(2), "nDCG@3", r.ndcg.toFixed(2));
}
console.log("MEAN recall@3 ", mean(rows.map(r => r.recall)).toFixed(3));
console.log("MEAN precision@3", mean(rows.map(r => r.precision)).toFixed(3));
console.log("MRR ", mean(rows.map(r => r.rr)).toFixed(3));
console.log("MEAN nDCG@3 ", mean(rows.map(r => r.ndcg)).toFixed(3));Question 1 finds one of two relevant docs at rank 2 (recall 0.5, RR 0.5). Question 2 is perfect. Question 3 never retrieves d9, so it scores 0 on every metric: that is a retrieval miss, and no prompt change can fix it. Look at per-question rows, not just the averages; the zeros tell you where to dig.
A crude faithfulness check, and why we need a judge (runnable)
const context = "Refunds are available within 30 days of purchase. Refunds go back to the original payment method. Gift cards cannot be refunded.";
const answer = "You can get a refund within 30 days. The money goes back to your original payment method. Refunds are processed within 24 hours.";
const STOP = new Set(["the", "you", "your", "can", "are", "get", "and", "for", "with", "goes", "money"]);
const words = s => s.toLowerCase().replace(/[^a-z0-9 ]/g, " ").split(/\s+/)
.filter(w => w.length > 2 && !STOP.has(w))
.map(w => w.replace(/s$/, "")); // toy stemming: "refunds" -> "refund"
const ctx = new Set(words(context));
for (const sentence of answer.split(". ")) {
const w = words(sentence);
const supported = w.filter(x => ctx.has(x));
const score = supported.length / w.length;
console.log(score >= 0.75 ? "SUPPORTED " : "UNSUPPORTED", score.toFixed(2), "|", sentence);
}Word overlap catches the invented '24 hours' claim here, but it is easily fooled: 'Refunds are NOT available within 30 days' overlaps perfectly and means the opposite. That gap is why production systems use an LLM judge that reasons about meaning, sentence by sentence.
LLM-as-judge with a rubric and structured output (Python)
import anthropic
from typing import Literal
from pydantic import BaseModel
client = anthropic.Anthropic()
class Verdict(BaseModel):
unsupported_claims: list[str] # claims in the answer the context does not support
faithfulness: int # 1-5, see rubric
relevance: int # 1-5, see rubric
reasoning: str # short justification, useful when debugging the judge
verdict: Literal["pass", "fail"]
RUBRIC = """You are grading an answer produced by a retrieval-augmented assistant.
Judge ONLY against the provided context, not your own knowledge.
Faithfulness (1-5):
5 = every factual claim is directly supported by the context
3 = mostly supported, one minor unsupported detail
1 = key claims are unsupported or contradict the context
Relevance (1-5):
5 = directly and completely answers the question
3 = partially answers or includes a lot of irrelevant material
1 = does not address the question
If the context does not contain the answer, the ideal answer says so; score that
5 for faithfulness and 5 for relevance.
List every unsupported claim verbatim. verdict = "pass" only if faithfulness >= 4
and relevance >= 4."""
def judge(question: str, context: str, answer: str) -> Verdict:
resp = client.messages.parse(
model="claude-opus-5-5",
max_tokens=16000,
system=RUBRIC,
messages=[{
"role": "user",
"content": (
f"<question>{question}</question>\n"
f"<context>{context}</context>\n"
f"<answer>{answer}</answer>"
),
}],
output_format=Verdict,
)
v = resp.parsed_output
# Validate ranges in code: never assume the schema alone enforces them.
assert 1 <= v.faithfulness <= 5 and 1 <= v.relevance <= 5, v
return v
if __name__ == "__main__":
v = judge(
"How long do I have to request a refund?",
"Refunds are available within 30 days of purchase.",
"You have 30 days. Refunds are processed within 24 hours.",
)
print(v.verdict, v.faithfulness, v.relevance, v.unsupported_claims)The judge gets the question, the exact context the generator saw, and the answer, and must return a typed Verdict. Putting the rubric in the system prompt keeps it identical across every graded item. Asking for unsupported_claims verbatim makes the grade auditable: you can read why something failed.
Golden dataset + regression test in CI (Python, pytest)
# golden.jsonl - one JSON object per line, versioned in git:
# {"id": "q001", "question": "What is the refund window?", "relevant_ids": ["refunds.md#0"],
# "reference": "30 days from purchase.", "tags": ["billing"]}
# {"id": "q002", "question": "Do you support SAML?", "relevant_ids": [],
# "reference": "UNANSWERABLE", "tags": ["unanswerable"]}
# test_rag_quality.py -- run with: pytest -q test_rag_quality.py
import json
import pytest
from my_rag import retrieve, answer # your pipeline
from judge import judge # the LLM-as-judge above
GOLDEN = [json.loads(line) for line in open("golden.jsonl") if line.strip()]
K = 5
def recall_at_k(retrieved_ids, relevant_ids, k):
if not relevant_ids:
return None # unanswerable: recall undefined
top = retrieved_ids[:k]
return sum(1 for r in relevant_ids if r in top) / len(relevant_ids)
def test_retrieval_recall_does_not_regress():
scores = []
for item in GOLDEN:
chunks = retrieve(item["question"], k=K) # list of dicts with "id"
r = recall_at_k([c["id"] for c in chunks], item["relevant_ids"], K)
if r is not None:
scores.append(r)
mean_recall = sum(scores) / len(scores)
print(f"recall@{K} = {mean_recall:.3f}")
assert mean_recall >= 0.85, "retrieval regressed below the agreed baseline"
@pytest.mark.slow # costs tokens: run on main / nightly
def test_answer_faithfulness():
failures = []
for item in GOLDEN:
chunks = retrieve(item["question"], k=K)
context = "\n\n".join(c["text"] for c in chunks)
ans = answer(item["question"], chunks)
v = judge(item["question"], context, ans)
if v.verdict == "fail":
failures.append((item["id"], v.unsupported_claims))
pass_rate = 1 - len(failures) / len(GOLDEN)
assert pass_rate >= 0.9, f"faithfulness pass rate {pass_rate:.2f}: {failures[:5]}"Retrieval metrics are cheap and deterministic, so they run on every pull request. Judge-based tests cost tokens and vary a little run to run, so they run on a schedule or before release. The thresholds (0.85, 0.9) are examples: set yours from your current baseline and raise them as you improve.
How it works
Building the golden dataset. Start with 30 to 100 questions; that is enough to catch big regressions. Sources, best first: real user questions from logs (anonymised), questions from support tickets, questions written by domain experts, and synthetic questions generated by an LLM from your chunks (useful for coverage, but they tend to reuse the chunk's exact wording, which makes retrieval look easier than it is). For each question record the relevant chunk or document ids and a short reference answer. Add tags (topic, difficulty, 'unanswerable', 'multi-hop') so you can break scores down by slice.
Labelling chunk ids is fragile: re-chunking changes ids. Many teams label at the document or section level (for example billing/refunds.md#refund-window) and count a retrieved chunk as relevant if it comes from a labelled section. Decide this rule once and write it down.
Running an eval is a loop: for each item, call retrieve, compute retrieval metrics against the labels, call generate with the retrieved chunks, then grade the answer (reference comparison and/or LLM judge). Store every input, output and score per item, not just the averages, so you can diff two runs item by item and see exactly which questions got better or worse.
Validating the judge. Have a human grade 50 or so answers with the same rubric, run the judge on the same answers, and measure agreement (the simplest measure is the percentage of matching pass/fail labels). If agreement is poor, tighten the rubric with concrete examples of each score, ask for the reasons before the score, or grade one criterion per call. Re-check whenever you change the judge prompt or model.
Judge hygiene: keep the judge prompt fixed while comparing system versions; grade claims against the context, not world knowledge, for faithfulness; randomise order when comparing two answers side by side (judges can favour the first one shown); and do not let the system under test grade itself with the same prompt it used to answer.
golden.jsonl ──> for each question
│
┌─────────┴──────────┐
v v
retrieve(q) labelled ids
│ │
└──> recall@k, precision@k, MRR, nDCG
│
v
generate(q, chunks) ──> answer
│
v
judge(q, chunks, answer) ──> faithfulness,
│ relevance
v
per-item results ──> averages ──> CI gateWhy does it exist?
LLM output is fuzzy, so 'does it work?' cannot be answered by a single assertion. Teams that skip evaluation end up tuning by vibes: a prompt tweak that fixes one complaint silently breaks ten other questions. Metrics turn quality into something you can track, compare and protect in CI.
Separating retrieval from generation metrics exists because the fixes are completely different: a retrieval miss needs better chunking, embeddings, hybrid search or reranking, while an unfaithful answer needs a better prompt, a stronger instruction to abstain, or citations.
When to use it
From the first working prototype onward. Build a small golden set before you start optimising, measure a baseline, then change one thing at a time (chunk size, k, embedding model, reranker, prompt) and keep only changes that improve the numbers you care about. Run retrieval metrics on every pull request and judge-based metrics nightly or before releases.
When not to use it
Do not chase a single average: a rise in mean recall can hide a collapse on one important slice, so always look at per-tag breakdowns. Do not treat an LLM judge as ground truth for high-stakes domains (medical, legal, financial) without regular human review. And do not spend weeks perfecting metrics before you have a working pipeline; a rough 30-question set today beats a perfect benchmark next quarter.
Common mistakes
Evaluating only final answers, so you cannot tell whether retrieval or generation failed.
Using an eval set generated entirely by an LLM from your own chunks, which inflates retrieval scores because the wording matches.
Leaving out unanswerable questions, so a system that always answers confidently looks great.
Changing the eval set and the system at the same time, so the numbers are not comparable.
Trusting an LLM judge without ever measuring its agreement with human labels.
Reporting precision@k with a large k and concluding retrieval is 'noisy', when recall is what limits answer quality.
Only storing averages, which makes it impossible to see which specific questions regressed.
Labelling chunk ids, then re-chunking and silently breaking every label.
Practice exercises
- Easy:
In the runnable metrics demo, change k from 3 to 5. Predict the new recall, precision and nDCG for each question before running it, then check.
- Easy:
Write 10 golden-set entries for a product you know (questions, relevant section ids, reference answers). Include at least two unanswerable questions and one that needs two sections.
- Medium:
Extend ndcgAtK to graded relevance: relevant becomes an object like { d2: 2, d4: 1 } and the gain is (2^grade - 1). Verify that swapping d2 and d4 in the ranking lowers the score.
- Medium:
Build an eval runner that writes per-item results to a JSONL file, then a second script that diffs two result files and prints the questions that got worse.
- Hard:
Validate an LLM judge: hand-label 40 answers pass/fail, run the judge, compute agreement, then improve the rubric (add examples per score) and measure again.
- Hard:
Add a GitHub Actions (or other CI) job that runs the retrieval regression test on every pull request and the judge-based test nightly, failing the build below your baseline.
Interview questions
What is the difference between recall@k and precision@k, and which matters more for RAG?
Recall@k is the fraction of all relevant items that appear in the top k; precision@k is the fraction of the top k that are relevant. For RAG, recall usually matters more: the model can ignore some noise, but it cannot use a fact that was never retrieved. Precision still matters for cost and for distraction, which is why reranking aims to keep recall while raising precision.
Explain MRR and when you would use it.
Mean reciprocal rank averages 1/rank of the first relevant result across queries (0 if none). It suits tasks where one good result at the top is what matters, such as navigational search or when you only pass the top chunk. It ignores everything after the first hit.
Why nDCG instead of precision?
nDCG accounts for position (higher ranks count more, via a log discount) and supports graded relevance, then normalises by the ideal ordering so scores are comparable across queries. Precision@k treats all positions in the top k equally and only knows relevant vs not.
What is faithfulness and how is it different from correctness?
Faithfulness asks whether every claim in the answer is supported by the retrieved context. Correctness asks whether the answer matches the true or reference answer. An answer can be correct but unfaithful (the model used outside knowledge) or faithful but incorrect (the context was wrong). In RAG you want faithfulness so answers are traceable to sources.
How would you build a golden dataset for a new RAG product?
Collect real questions from logs, support tickets or domain experts; label relevant documents or sections and short reference answers; add tags for slices; include unanswerable and multi-hop cases; supplement with synthetic questions for coverage while knowing they are easier; keep it versioned in git and grow it whenever a production failure is found.
What are the risks of LLM-as-judge and how do you mitigate them?
Judges can be biased toward longer or more confident answers, toward the first option shown, or toward their own style, and can be inconsistent. Mitigate with a specific rubric with examples, structured output, reasons before scores, one criterion per call when needed, randomised order in pairwise comparisons, and regular measurement of agreement with human labels.
How do you put RAG quality into CI?
Run a regression suite on the golden set: cheap deterministic retrieval metrics on every pull request with thresholds based on the current baseline, and slower judge-based answer metrics nightly or pre-release. Fail the build when metrics drop, and store per-item results so the failure is easy to diagnose.