What is RAG and Why It Exists

Retrieval-Augmented Generation: fetch relevant passages from your own data and give them to the model, so answers are current, private-data-aware and citable - compared with fine-tuning and long context.

What is it?

Retrieval-Augmented Generation (RAG) is a pattern for answering questions with an LLM using your information. Instead of hoping the model memorised the answer during training, you retrieve the most relevant pieces of your documents at question time and augment the prompt with them, then the model generates an answer based on that material.

It solves the problems from hallucinations-and-limitations directly:

  • Private data: the model never saw your internal wiki, contracts or support tickets. RAG supplies them.
  • Freshness: update a document and the next answer uses it - no retraining.
  • Grounding: the model answers from supplied text, which greatly reduces made-up answers.
  • Citations: you know which passages were used, so you can show sources and users can verify.
  • Access control: you can retrieve only documents the current user is allowed to see.

A RAG system has two phases:

  • Indexing (offline, ahead of time): load documents, split them into chunks (passages of a few hundred tokens), compute an embedding for each chunk, and store chunk text + vector + metadata (source, title, date, permissions) in a vector store.
  • Querying (online, per question): embed the question, find the top-k most similar chunks, optionally rerank them, put the best ones into a prompt with clear instructions ('answer only from these sources, cite them, say you don't know otherwise'), and call the LLM.

RAG vs the alternatives:

  • Fine-tuning (further training the model on your data) is good at changing behaviour, style and format, but it is a poor way to add facts: it is slow and costly to update, cannot cite sources, cannot enforce per-user permissions, and the model can still misremember. Use RAG for knowledge, fine-tuning (rarely) for behaviour (see fine-tuning-vs-rag).
  • Long context (just paste all the documents into the prompt) is perfectly reasonable when the material is small enough to fit comfortably and does not change much - it is the simplest solution and avoids retrieval errors. Prompt caching can make repeated long prompts cheaper. But it does not scale to large or growing collections, costs more per request, adds latency, and focus can suffer when the relevant part is a tiny fraction of a huge prompt. Many systems combine both: retrieve, then give the model generous context.
  • Tools / agents: for structured data (orders, accounts) the model can call an API or database query instead of retrieving text; agentic RAG lets the model decide when and what to search.

RAG quality depends mostly on retrieval: if the right chunk is not retrieved, the model cannot use it (and may guess). Most of the RAG section of this subject is about doing retrieval well: loading, chunking, embedding, hybrid search, reranking and evaluation.

Explain like I'm 10

A closed-book exam versus an open-book exam. A plain LLM sits a closed-book exam: it answers from memory and sometimes misremembers. RAG turns it into an open-book exam where a librarian first fetches the three most relevant pages and puts them on the desk, and the rules say 'answer from these pages and write down which page you used'. Fine-tuning is making the student study for weeks - great for learning the exam's style, but slow, and they still might misremember a detail. Long context is dumping the whole library on the desk: fine if it is a small shelf, unworkable if it is a building.

Examples

A complete RAG loop, offline (toy retrieval + fake LLM)

// 1) INDEXING: chunks of our 'knowledge base'
const chunks = [
  { id: "refunds#1", text: "Refunds are available within 30 days of purchase for monthly plans." },
  { id: "refunds#2", text: "Annual plans are refunded pro rata for unused full months." },
  { id: "security#1", text: "Two-factor authentication can be enabled under Settings > Security." },
  { id: "billing#1", text: "Invoices are emailed on the first business day of each month." },
];

// Toy 'embedding': word counts (real systems use an embedding model)
const tokenize = s => s.toLowerCase().split(/\W+/).filter(w => w.length > 2);
const vocab = [...new Set(chunks.flatMap(c => tokenize(c.text)))];
const embed = s => { const w = tokenize(s); return vocab.map(v => w.filter(x => x === v).length); };
const dot = (a, b) => a.reduce((s, x, i) => s + x * b[i], 0);
const cosine = (a, b) => { const d = Math.sqrt(dot(a, a) * dot(b, b)); return d ? dot(a, b) / d : 0; };
const index = chunks.map(c => ({ ...c, vector: embed(c.text) }));

// 2) RETRIEVAL
function retrieve(question, k = 2) {
  const q = embed(question);
  return index
    .map(c => ({ ...c, score: cosine(q, c.vector) }))
    .filter(c => c.score > 0)
    .sort((a, b) => b.score - a.score)
    .slice(0, k);
}

// 3) AUGMENT: build the prompt
function buildPrompt(question, hits) {
  const sources = hits.map(h => "<source id='" + h.id + "'>" + h.text + "</source>").join(" ");
  return "<sources>" + sources + "</sources> Answer only from the sources and cite ids. " +
    "If they don't contain the answer, say you don't know. Question: " + question;
}

// 4) GENERATE: a fake LLM that 'reads' the sources in the prompt
function fakeLLM(prompt) {
  const found = [...prompt.matchAll(/<source id='([^']+)'>([^<]+)<\/source>/g)];
  if (found.length === 0) return "I don't know - the knowledge base has nothing on that.";
  return found.map(m => m[2] + " [" + m[1] + "]").join(" ");
}

for (const question of ["Can I get refunds on annual plans?", "How do I enable two-factor authentication?", "What is your office address?"]) {
  const hits = retrieve(question);
  console.log("Q:", question);
  console.log("   retrieved:", hits.map(h => h.id + " (" + h.score.toFixed(2) + ")").join(", ") || "nothing");
  console.log("   A:", fakeLLM(buildPrompt(question, hits)));
}

Every real RAG system has these four steps. Swap the word-count embedding for a real embedding model, the array for a vector database, and the fake LLM for an API call, and you have a production architecture. Note the third question: nothing relevant is retrieved, so the system says it does not know instead of inventing an address.

Minimal real RAG in Python (in-memory, sentence-transformers + Claude)

# pip install anthropic sentence-transformers numpy
import numpy as np
import anthropic
from sentence_transformers import SentenceTransformer

chunks = [
    {"id": "refunds#1", "text": "Refunds are available within 30 days of purchase for monthly plans."},
    {"id": "refunds#2", "text": "Annual plans are refunded pro rata for unused full months."},
    {"id": "security#1", "text": "Two-factor authentication can be enabled under Settings > Security."},
    {"id": "billing#1", "text": "Invoices are emailed on the first business day of each month."},
]

# --- Indexing (do once, store the vectors) ---
embedder = SentenceTransformer("all-MiniLM-L6-v2")
vectors = embedder.encode([c["text"] for c in chunks], normalize_embeddings=True)

# --- Retrieval ---
def retrieve(question: str, k: int = 2):
    q = embedder.encode([question], normalize_embeddings=True)[0]
    scores = vectors @ q
    best = np.argsort(-scores)[:k]
    return [chunks[i] for i in best]

# --- Augment + generate ---
client = anthropic.Anthropic()

def answer(question: str) -> str:
    hits = retrieve(question)
    sources = "\n".join(f'<source id="{h["id"]}">{h["text"]}</source>' for h in hits)
    resp = client.messages.create(
        model="claude-opus-5-5",
        max_tokens=16000,
        system=("Answer using only the sources provided. Cite source ids in square brackets. "
                "If the sources do not contain the answer, say you don't know."),
        messages=[{"role": "user", "content": f"<sources>\n{sources}\n</sources>\n\nQuestion: {question}"}],
    )
    return next(b.text for b in resp.content if b.type == "text")

print(answer("I'm on the annual plan - can I get money back if I cancel?"))

Notice the question shares few words with the right chunk ('money back' vs 'refunded'), which semantic embeddings handle. The rest of the RAG section upgrades each part: a persistent vector database, better chunking, hybrid search, reranking, and citations.

Choosing between long context, RAG and fine-tuning

Need                                          -> Start with
--------------------------------------------- -> ---------------------
A few documents that fit easily in context    -> Long context (+ caching)
Large / growing / frequently updated corpus   -> RAG
Answers must cite sources                     -> RAG (or long context + citations)
Different users may see different documents   -> RAG with permission filters
Live structured data (orders, balances)       -> Tools / API calls
Consistent style, tone or output format       -> Prompting first, fine-tuning rarely
Teach the model new facts                     -> RAG, not fine-tuning

A quick decision table. In practice, start with the simplest thing that works (often prompting plus long context), and move to RAG as the corpus grows.

How it works

Indexing pipeline: documents -> loaders (extract clean text from PDFs, HTML, etc.) -> chunker (split into passages, often with some overlap) -> embedding model -> vector store (vector + text + metadata). This runs whenever documents change.

Query pipeline: question -> (optional) rewrite the question into a better search query -> embed -> similarity search for the top-k chunks (often 20-50 candidates) -> (optional) rerank to the best few -> build a prompt with sources in tags and grounding instructions -> LLM -> answer with citations -> (optional) validate citations.

Why the pieces matter: chunk too large and retrieval gets vague and the prompt fills with irrelevant text; too small and chunks lose context. Pure vector search misses exact terms; keyword search misses paraphrases; hybrid search combines both. Rerankers read the question and each candidate together for a more accurate relevance judgement. The prompt's instructions decide whether the model admits ignorance or guesses.

Evaluating RAG separates two questions: did retrieval find the right chunks (retrieval metrics like recall@k), and did the model answer faithfully from them (faithfulness, answer relevance)? Debugging always starts by checking which one failed (evaluating-rag).

INDEXING (offline)
 docs -> load -> chunk -> embed -> [vector store]
                                    vector+text+meta

QUERYING (per question)
 question -> embed -> search top-k ---^
                         |
                     rerank (optional)
                         v
  prompt = instructions + <sources> + question
                         |
                         v
                   LLM generates
                         v
             answer + citations [id]

Why does it exist?

LLMs know a lot, but not your information, not anything after their training cutoff, and not reliably the exact details that matter in business answers. Retraining a model whenever a document changes is impractical. RAG exists because it is the cheapest, most up-to-date and most verifiable way to connect a general model to specific knowledge: update the index, not the model.

When to use it

Use RAG for question answering over documentation, policies, knowledge bases, tickets, contracts, research papers and code - especially when the content is large, changes often, is private, needs per-user permissions, or answers must cite sources.

When not to use it

If all your material comfortably fits in the context window and rarely changes, skip the retrieval machinery and include it directly (with prompt caching). If the answer lives in structured data, query the database through a tool instead of embedding rows as text. If you need a different style or format, try prompting before anything else. And RAG cannot fix poor source documents - if the docs are wrong or missing, so are the answers.

Common mistakes

  • Building a complex RAG pipeline when the documents would simply fit in the prompt.

  • Thinking fine-tuning is the way to 'teach the model our documents'.

  • Blaming the LLM for bad answers when retrieval never returned the right chunk.

  • Not instructing the model to answer only from sources and to say when it doesn't know.

  • Ignoring metadata such as source, date and permissions, then being unable to cite or filter.

  • Retrieving too few chunks (missing the answer) or far too many (diluting the prompt).

  • Never re-indexing, so answers go stale as documents change.

  • Shipping without a set of test questions to measure retrieval and answer quality.

Practice exercises

  1. Easy:

    In your own words, describe the indexing phase and the query phase of RAG, and which one runs per question.

  2. Easy:

    Add three chunks to the offline RAG demo (for example about pricing) and ask questions that should retrieve them.

  3. Medium:

    In the offline demo, add a minScore threshold to retrieve and show a question where a weak, irrelevant match is retrieved without it and correctly rejected with it.

  4. Medium:

    For each scenario, choose long context, RAG, tools or fine-tuning and justify: (a) a 10-page employee handbook, (b) 50,000 support articles updated daily, (c) 'what is my order status?', (d) always replying in the company's formal tone.

  5. Hard:

    Build the Python RAG example over a real folder of Markdown files: split each file into paragraphs, embed and store them (in memory or numpy files), retrieve top-4, and answer with cited file names. Write 10 test questions and record which ones retrieved the right paragraph.

Interview questions

What is RAG?

Retrieval-Augmented Generation: at question time, retrieve relevant passages from an external knowledge source (usually via embedding similarity), insert them into the prompt, and have the LLM generate an answer grounded in them, ideally with citations.

Why use RAG instead of fine-tuning to add company knowledge?

RAG updates instantly by re-indexing, supports citations and per-user permissions, and grounds answers in exact text. Fine-tuning is costly to repeat, cannot cite, does not enforce access control, and is better suited to changing behaviour or style than to adding facts.

When is long context better than RAG?

When the relevant material is small enough to fit comfortably, changes rarely, and is needed as a whole; it avoids retrieval misses and complexity, and prompt caching reduces repeat cost. RAG wins for large, growing, frequently updated or permissioned corpora.

What are the main components of a RAG pipeline?

Document loading and cleaning, chunking, embedding, a vector store with metadata, retrieval (semantic, keyword or hybrid), optional reranking, prompt construction with grounding instructions, generation, and citation/answer validation, plus evaluation and re-indexing.

A RAG system gives a wrong answer. How do you debug it?

First check retrieval: were the correct chunks in the retrieved set? If not, look at chunking, embeddings, query phrasing, hybrid search or k. If they were retrieved, check the prompt and the model's faithfulness: instructions, ordering, too much irrelevant context, or missing 'I don't know' guidance.

Does RAG eliminate hallucinations?

No, but it reduces them substantially. The model can still misread sources, combine them wrongly or fall back on training knowledge, so you still need grounding instructions, citations, validation and evaluation.