Advanced RAG Techniques
Multi-query, HyDE, parent-child (small-to-big), contextual retrieval, metadata filtering, query routing and the GraphRAG idea.
What is it?
Basic RAG embeds the user's question, finds the nearest chunks and pastes them into the prompt. It breaks in predictable ways: the question is phrased differently from the documents, the best chunk is too small to be useful (or too big to be found), a chunk loses meaning when cut out of its document, the right answer lives in the wrong product version, or the question needs facts spread across many documents. Each advanced technique below targets one of those failures. Measure before and after with your eval set; none is free.
- Multi-query retrieval: ask an LLM to rewrite the question into several variants (synonyms, sub-questions, different vocabulary), retrieve for each, and merge the lists with reciprocal rank fusion (RRF). Fixes vocabulary mismatch and multi-part questions. Costs one extra LLM call and several retrievals.
- HyDE (Hypothetical Document Embeddings): ask an LLM to write a hypothetical answer to the question, then embed that answer and search with it. A fake answer often looks more like the real documents than the short question does. The hypothetical text may contain made-up facts; that is fine because it is only used for searching, never shown to the user.
- Parent-child / small-to-big: index small chunks (precise matches) but return the larger parent section they belong to (enough context to answer). Each small chunk stores its parent id.
- Contextual retrieval: before embedding a chunk, prepend a short LLM-written sentence that situates it in its document ('This chunk is from the ACME Q2 2024 earnings report, revenue section'). A chunk that says 'Revenue grew 3%' becomes findable by a query about ACME's Q2 revenue. You can prepend the same context before building the BM25 index too.
- Metadata filtering: store structured fields (product, version, language, date, access group) with each chunk and filter on them before or during vector search. Stops the right-looking chunk from the wrong version winning.
- Query routing: classify the question first and send it to the right place: a specific index, a SQL database, a web search, or no retrieval at all for small talk.
- GraphRAG: extract entities and relationships from documents into a knowledge graph, often with LLM-written summaries of clusters of related entities. Retrieval can then follow relationships (who reports to whom, which service depends on which) and answer global questions ('what are the main themes across all incident reports?') that no single chunk answers. Powerful, but expensive to build and maintain; reach for it only when relationship or corpus-wide questions dominate.
These techniques combine with what you already know: hybrid search (BM25 + vectors) and reranking usually give the biggest win for the least effort, so add those first. Then use your eval breakdown to pick the technique that fixes your most common failure.
Explain like I'm 10
Think of a librarian helping you. Multi-query: they search the catalogue under several phrasings of your question. HyDE: they imagine what the perfect book would say and look for books that sound like that. Small-to-big: they find the exact sentence, then hand you the whole chapter. Contextual retrieval: every photocopied page is stamped with the book title and chapter. Metadata filtering: they only look in the 2024 shelf. Routing: they send tax questions to the tax desk. GraphRAG: they keep a wall chart of who is connected to whom.
Examples
Multi-query and HyDE with reciprocal rank fusion (runnable)
// Tiny bag-of-words "embedding" + cosine so everything runs offline.
const STOP = new Set(["the","a","an","is","to","of","and","in","for","on","how","do","i","my","it","can","what","with","be","you","your","are"]);
const tokens = s => s.toLowerCase().replace(/[^a-z0-9 ]/g, " ").split(/\s+/).filter(w => w && !STOP.has(w));
function embed(text) { const v = {}; for (const t of tokens(text)) v[t] = (v[t] || 0) + 1; return v; }
function cosine(a, b) {
let dot = 0, na = 0, nb = 0;
for (const k in a) { na += a[k] * a[k]; if (b[k]) dot += a[k] * b[k]; }
for (const k in b) nb += b[k] * b[k];
return na && nb ? dot / Math.sqrt(na * nb) : 0;
}
const docs = [
{ id: "billing-1", text: "Invoices are emailed on the first day of each month." },
{ id: "billing-2", text: "To change the card we charge, open Billing and choose Update payment method." },
{ id: "account-1", text: "Reset your password from the login page using Forgot password." },
{ id: "billing-3", text: "Refunds return money to the original payment method within 30 days." },
];
const index = docs.map(d => ({ ...d, vec: embed(d.text) }));
const search = (q, k = 3) => index.map(d => ({ id: d.id, s: cosine(embed(q), d.vec) }))
.sort((x, y) => y.s - x.s).slice(0, k).map(d => d.id);
// Fake LLM: canned rewrites and a canned hypothetical answer.
const llm = {
rewrite: q => [q, "update payment method card", "change billing card details"],
hypothetical: q => "Open Billing and choose Update payment method to change the card we charge.",
};
function rrf(lists, k = 60) { // reciprocal rank fusion
const score = {};
for (const list of lists) list.forEach((id, rank) => { score[id] = (score[id] || 0) + 1 / (k + rank + 1); });
return Object.entries(score).sort((a, b) => b[1] - a[1]).map(([id]) => id);
}
const question = "How can I switch which credit gets charged each month?";
console.log("plain :", search(question));
console.log("multi-query:", rrf(llm.rewrite(question).map(q => search(q))));
console.log("HyDE :", search(llm.hypothetical(question)));The user's words ('switch', 'charged', 'each month') miss the document's wording ('change the card we charge', 'payment method') and accidentally match the invoice document, so plain search puts billing-1 first. The rewrites and the hypothetical answer use document-like vocabulary, so billing-2 rises to the top. Real embeddings handle synonyms far better than this toy, but the same mismatch happens with jargon, abbreviations and product names.
Parent-child retrieval and contextual chunk headers (runnable)
const STOP = new Set(["the","a","an","is","to","of","and","in","for","on","by","was","what","did","how","much"]);
const tokens = s => s.toLowerCase().replace(/[^a-z0-9 ]/g, " ").split(/\s+/).filter(w => w && !STOP.has(w));
const overlap = (q, text) => { const t = new Set(tokens(text)); const qt = tokens(q); return qt.filter(w => t.has(w)).length / qt.length; };
// Two parent sections from two different reports.
const parents = {
"acme-q2#revenue": { doc: "ACME Corp Q2 2024 report", section: "Revenue",
text: "Revenue grew 3% compared to the previous quarter. Growth came from the services line. Hardware sales were flat." },
"globex-q2#revenue": { doc: "Globex Q2 2024 report", section: "Revenue",
text: "Revenue grew 9% compared to the previous quarter. Growth came from new markets." },
};
// Small child chunks = single sentences, each pointing to its parent.
const children = [];
for (const [pid, p] of Object.entries(parents)) {
p.text.split(". ").forEach((s, i) => children.push({ id: pid + "/" + i, parent: pid, text: s,
// Contextual retrieval: prepend a situating header before indexing.
contextual: p.doc + ", " + p.section + " section: " + s }));
}
const q = "How much did ACME revenue grow in Q2?";
const rank = field => children.map(c => ({ c, s: overlap(q, c[field]) })).sort((a, b) => b.s - a.s);
const plain = rank("text");
console.log("plain top-2 :", plain.slice(0, 2).map(r => r.c.id + " (" + r.s.toFixed(2) + ")"));
const ctx = rank("contextual");
console.log("contextual :", ctx.slice(0, 2).map(r => r.c.id + " (" + r.s.toFixed(2) + ")"));
// Small-to-big: match on the child, hand the model the whole parent section.
const best = ctx[0].c;
console.log("send to LLM :", parents[best.parent].doc + " / " + parents[best.parent].section + ": " + parents[best.parent].text);Without context, 'Revenue grew 3%' and 'Revenue grew 9%' look identical to the query, so plain retrieval ties and may pick Globex. With the document name prepended, the ACME chunk wins. Small-to-big then returns the whole revenue section so the model also sees where the growth came from.
Metadata filtering and query routing (runnable)
const chunks = [
{ id: "v1-auth", product: "api", version: "v1", text: "Authenticate with an API key in the X-Key header." },
{ id: "v2-auth", product: "api", version: "v2", text: "Authenticate with an OAuth bearer token in the Authorization header." },
{ id: "app-auth", product: "mobile", version: "v2", text: "Sign in to the mobile app with your email." },
];
// Router: decide WHERE a question should go before retrieving anything.
function route(question) {
const q = question.toLowerCase();
if (/^(hi|hello|thanks)/.test(q)) return { target: "none" }; // no retrieval needed
if (/how many|count|total|average/.test(q)) return { target: "sql" }; // numbers live in a DB
const filter = {};
const v = q.match(/\bv(\d)\b/); if (v) filter.version = "v" + v[1];
if (q.includes("mobile") || q.includes("app")) filter.product = "mobile";
else if (q.includes("api")) filter.product = "api";
return { target: "docs", filter };
}
function retrieve(question, filter) {
// Filter FIRST, then rank only what passes (here every survivor is returned).
return chunks.filter(c => Object.entries(filter).every(([k, v]) => c[k] === v)).map(c => c.id);
}
for (const q of ["Hello!", "How many tickets were opened last week?",
"How do I authenticate to the API v2?", "How do I authenticate to the API?"]) {
const r = route(q);
console.log(q.padEnd(42), "->", r.target, r.filter ? JSON.stringify(r.filter) + " " + JSON.stringify(retrieve(q, r.filter)) : "");
}The router sends greetings nowhere and analytics questions to SQL. For docs questions it extracts filters, so 'API v2' can only return the v2 chunk. Note the last question: with no version stated, both versions come back, which is correct; the answer should mention both or ask which version. In production the router is usually a small, cheap LLM call with structured output instead of regexes.
Contextual retrieval with Claude and prompt caching (Python)
import anthropic
client = anthropic.Anthropic()
SITUATE_PROMPT = """Here is a chunk from the document above:
<chunk>
{chunk}
</chunk>
Write one or two short sentences that situate this chunk within the overall
document (what document, which section, what entity and time period it is about)
to improve search retrieval of the chunk. Answer only with that context."""
def contextualize(document: str, chunks: list[str]) -> list[str]:
out = []
for chunk in chunks:
resp = client.messages.create(
model="claude-opus-5-5",
max_tokens=1000,
# The whole document is the stable prefix: cache it so each chunk call
# after the first reads it from cache instead of paying full input price.
system=[{"type": "text", "text": "<document>\n" + document + "\n</document>",
"cache_control": {"type": "ephemeral"}}],
messages=[{"role": "user", "content": SITUATE_PROMPT.format(chunk=chunk)}],
output_config={"effort": "low"}, # simple task: keep it fast and cheap
)
context = "".join(b.text for b in resp.content if b.type == "text").strip()
out.append(context + "\n\n" + chunk) # embed AND BM25-index this text
print("cache read tokens:", resp.usage.cache_read_input_tokens)
return out
# Then embed as usual:
# from sentence_transformers import SentenceTransformer
# model = SentenceTransformer("all-MiniLM-L6-v2")
# vectors = model.encode(contextualize(doc_text, chunks), normalize_embeddings=True)
# Store the ORIGINAL chunk text for display/citation, the contextualized text for search.One LLM call per chunk is the cost of contextual retrieval, paid once at ingestion. Caching the document as the system prefix makes every call after the first much cheaper; the printed cache_read_input_tokens confirms the cache is working. Keep both versions of the text: contextualized for searching, original for showing and citing.
How it works
Multi-query + RRF. RRF gives each document a score of 1 / (k + rank) in every list it appears in (k is a constant, commonly 60, that softens the advantage of rank 1) and sums them. It needs only ranks, not comparable scores, so it can merge lists from different queries or even different retrievers. Run the per-variant searches in parallel to limit latency.
HyDE replaces the query vector with the vector of a generated answer. It helps most when questions are short and documents are long and formal; it can hurt when the model's guess is confidently off-topic (then it retrieves the wrong area entirely). A common safeguard is to search with both the question and the hypothetical document and fuse the results.
Parent-child needs two stores: a vector index of children with a parent_id in metadata, and a key-value lookup (a dict, a table, a document store) from parent id to full text. After retrieval, deduplicate parents: several children often point to the same parent.
Contextual retrieval is an ingestion-time change. For each chunk, an LLM reads the whole document plus the chunk and writes a short situating context; you index context + chunk for both embeddings and BM25. Combining it with hybrid search and reranking compounds the gains. Because every chunk call shares the same document prefix, prompt caching makes it affordable.
Metadata filtering works best as a pre-filter inside the vector database (Chroma where={...}, a SQL WHERE with pgvector), so you still get k results from the allowed subset. Post-filtering (retrieve k, then drop) can leave you with zero results. Access-control filters must always be enforced server-side, never left to the model.
Routing is a classifier in front of retrieval. Give it a fixed set of routes and structured output, log its decisions, and add routing accuracy to your eval set (label the expected route per question).
GraphRAG pipelines typically: extract entities and relations from each chunk with an LLM; merge duplicates into a graph; detect communities (clusters) and summarise each; then at query time either walk the graph from entities in the question (local questions) or combine community summaries (global questions). Indexing costs many LLM calls, and the graph must be updated as documents change.
question
│
v
[router] ──> sql / web / no-retrieval
│ docs
v
[rewrite: multi-query | HyDE]
│ several queries
v
[hybrid search + metadata pre-filter]
│ children (contextualized chunks)
v
[RRF fuse] ──> [rerank] ──> [child -> parent]
│
v
LLM + citationsWhy does it exist?
Naive RAG's ceiling is set by retrieval: if the right passage is not in the top k, the answer is wrong or 'I don't know'. These techniques exist because real corpora have vocabulary gaps, ambiguous fragments, many versions of the same page, and questions that span documents. Each technique adds a little cost to recover recall or precision in one of those situations.
When to use it
After your baseline (hybrid search + reranking) is measured and you know your failure modes from the eval breakdown. Vocabulary mismatch: multi-query or HyDE. Chunks found but too thin: small-to-big. Ambiguous chunks ('the company', 'this release'): contextual retrieval. Wrong version or product: metadata filtering. Mixed question types: routing. Relationship-heavy or corpus-wide questions: consider GraphRAG.
When not to use it
Do not stack every technique by default: each adds latency, cost and moving parts, and some interact badly (HyDE plus aggressive filters can return nothing). Skip query-time LLM rewrites when latency budgets are tight; prefer ingestion-time improvements like contextual retrieval, which cost nothing per query. Skip GraphRAG unless your eval shows relationship or global questions are common and failing.
Common mistakes
Adding advanced techniques before measuring a baseline, so you cannot tell whether they helped.
Showing the HyDE hypothetical answer to users, or using it as context: it may be fiction.
Returning children instead of parents in small-to-big, or returning the same parent several times.
Displaying or citing the contextualized text instead of the original chunk.
Post-filtering by metadata after retrieving k results, leaving too few or zero chunks.
Relying on the model to respect access permissions instead of filtering in the database.
Running multi-query searches sequentially, multiplying latency instead of running them in parallel.
Building a knowledge graph for a corpus where simple hybrid search already scores well.
Practice exercises
- Easy:
In the multi-query demo, add a fourth document about 'credit limits' and check whether multi-query or HyDE pulls it in wrongly. Explain why.
- Easy:
For each technique in this topic, write one sentence naming the failure it fixes and one naming its main cost.
- Medium:
Implement small-to-big with Chroma: index sentence-level children with parent_id metadata, keep parents in a dict, deduplicate parents after retrieval, and compare recall@5 against fixed-size chunks on your eval set.
- Medium:
Replace the regex router with an LLM router using structured output (a route enum plus optional filters). Add the expected route to 20 eval questions and measure routing accuracy.
- Hard:
Implement contextual retrieval for a small corpus, index both embeddings and BM25 on the contextualized text, and compare recall@k and answer faithfulness against your baseline. Report the ingestion cost in tokens with and without prompt caching.
Interview questions
What is HyDE and when can it backfire?
Hypothetical Document Embeddings: generate a plausible answer to the question and search with its embedding, because answers resemble documents more than questions do. It backfires when the model's guess is wrong about the domain or terminology, steering retrieval to the wrong area; mitigate by fusing results from the original query and the hypothetical document.
Explain parent-child (small-to-big) retrieval.
Index small chunks for precise matching, each with a pointer to a larger parent section. Retrieve on the small chunks, then pass the parent sections (deduplicated) to the model so it has enough surrounding context. It separates the unit of search from the unit of reading.
What problem does contextual retrieval solve?
Chunks lose meaning when separated from their document: 'revenue grew 3%' does not say which company or quarter. Contextual retrieval prepends an LLM-written situating sentence to each chunk before embedding and keyword indexing, so the chunk can be matched by queries that mention the missing context. It is an ingestion-time cost, made cheaper with prompt caching of the document.
Why is reciprocal rank fusion popular for merging result lists?
It only uses ranks, so it can combine lists whose scores are not comparable (BM25 vs cosine, or different query variants). Documents that rank well in several lists rise to the top, and it has essentially one parameter. It is simple, robust and fast.
Pre-filtering vs post-filtering on metadata?
Pre-filtering restricts the search to allowed items so you still get k results from the right subset; post-filtering retrieves k and then drops items, which can leave few or none. Pre-filtering is preferred, and mandatory for access control.
When would you consider GraphRAG?
When important questions depend on relationships between entities across documents, or ask for corpus-wide summaries ('main themes in all reports'), which chunk retrieval handles poorly. It costs many LLM calls to build and maintain, so justify it with eval results showing those question types fail with hybrid search plus reranking.