Reranking
Two-stage retrieval: fetch many candidates cheaply, then reorder them with a slower, smarter cross-encoder or LLM so the best chunks reach the prompt.
What is it?
First-stage retrieval (vector search, BM25 or hybrid) is built for speed over millions of chunks, so it uses shortcuts. Embedding search compresses the query and each chunk into separate vectors before they ever meet, then compares the vectors. It never reads the question and the chunk together. That is why its top results are often 'about the right topic' without actually answering the question.
Reranking adds a second stage: take the first stage's top candidates (say 20-50), score each one with a slower but more accurate model that reads the query and the chunk together, sort by that score, and keep the best few (say 3-8) for the prompt. This is called two-stage retrieval or retrieve-then-rerank: a wide, cheap net followed by a careful sort.
Bi-encoder vs cross-encoder. The embedding model used for search is a bi-encoder: it encodes query and document independently (two separate passes), which is what lets you pre-compute all document vectors once. A cross-encoder takes the query and one document concatenated as a single input and outputs one relevance score. Because the transformer's attention can compare every query word with every document word, it notices things a bi-encoder misses: whether the passage actually contains the requested number, whether it is about the right product version, whether a 'not' flips the meaning. The cost: nothing can be pre-computed, so you run the model once per (query, candidate) pair at query time. Fine for 50 candidates, impossible for 5 million.
Kinds of rerankers:
- Cross-encoder models - small open models such as
cross-encoder/ms-marco-MiniLM-L-6-v2(trained on the MS MARCO web search dataset of real queries and relevant passages) run locally with thesentence-transformerslibrary. Hosted reranking APIs also exist from several embedding providers. - LLM rerankers - ask an LLM to score or order the candidates ('rate each passage 0-10 for how well it answers the question'). Very flexible (you can describe what 'relevant' means in your domain) but slower and more expensive; use for small candidate sets or high-value queries.
- Heuristic or business-rule boosts - add points for recency, authoritative sources, or the user's own team's documents. Often combined with a model score.
What reranking buys you. Better precision in the few chunks you actually send: the prompt gets shorter (cheaper, faster) and more relevant (better answers, fewer hallucinations from distracting context). It also gives you a more meaningful score for thresholding ('no candidate scored above X, so say I don't know').
What it cannot do: a reranker only reorders what the first stage found. If the right chunk is not in the candidate set, reranking cannot recover it - so the first stage should be tuned for recall (fetch generously) and the reranker for precision.
Explain like I'm 10
Hiring for a job: a recruiter skims 2,000 CVs for keywords and similar job titles in a few seconds each and shortlists 30 (first-stage retrieval). Then the hiring manager actually reads those 30 CVs carefully against the job description and picks the best 5 to interview (reranking). The manager would never have time to read all 2,000, and the recruiter's skim alone would put some poor fits in the top 5.
Examples
Two-stage retrieval: cheap first stage, joint-scoring reranker
const chunks = [
"Annual leave helps employees rest, and how leave days are booked is up to each team.",
"Our leave policy was updated in 2025 after an employee survey.",
"Each staff member is entitled to 25 days of annual leave per calendar year.",
"Sick leave does not count towards annual leave.",
"Parental leave is 16 weeks at full pay.",
"The canteen is open from 8am to 3pm.",
];
const query = "how many days of annual leave do employees get";
const words = (s) => s.toLowerCase().replace(/[^a-z0-9 ]/g, " ").split(" ").filter((w) => w.length > 2);
// Stage 1 (fast, independent): fraction of query words present - like a crude bi-encoder/BM25
function stage1(q, k) {
const qw = new Set(words(q));
return chunks.map((text, id) => ({ id, text, s1: words(text).filter((w) => qw.has(w)).length / qw.size }))
.sort((a, b) => b.s1 - a.s1).slice(0, k);
}
// Stage 2 (slower, reads query AND chunk together): a hand-written stand-in for a cross-encoder
function crossScore(q, text) {
const qw = words(q), tw = words(text);
let score = qw.filter((w) => tw.includes(w)).length; // coverage
for (let i = 0; i < qw.length - 1; i++) { // phrase matches
if (text.toLowerCase().includes(qw[i] + " " + qw[i + 1])) score += 2;
}
if (/^how (many|much)/.test(q) && /[0-9]/.test(text)) score += 4; // asks for a number, has one
if (/does not|not /.test(text.toLowerCase())) score -= 1; // negation is risky evidence
return score;
}
const candidates = stage1(query, 5);
console.log("Stage 1 top-5 (fast, approximate):");
candidates.forEach((c, i) => console.log(" " + (i + 1) + ". [" + c.s1.toFixed(2) + "] " + c.text));
const reranked = candidates.map((c) => ({ ...c, s2: crossScore(query, c.text) }))
.sort((a, b) => b.s2 - a.s2).slice(0, 2);
console.log("After reranking, top-2 sent to the LLM:");
reranked.forEach((c, i) => console.log(" " + (i + 1) + ". [" + c.s2 + "] " + c.text));Stage 1 ranks a vague chunk at the top because it repeats many of the query's words ('how', 'leave', 'days', 'employees') without answering anything, while the real answer uses different words ('staff member', 'entitled'). The reranker reads query and chunk together, rewards the phrase 'annual leave', and notices the question asks 'how many' while the answer chunk contains a number - so the chunk that actually answers the question moves to first place. A real cross-encoder learns these signals (and many more) from training data instead of hand-written rules.
LLM as a reranker (offline, with a fake llm())
// Fake LLM: in reality you would send this prompt to a model and parse its JSON reply
function llm(prompt) {
const passages = prompt.split("\n").filter((l) => l.startsWith("[")); // "[i] text"
const scores = passages.map((line) => {
const i = Number(line.slice(1, line.indexOf("]")));
const text = line.toLowerCase();
const s = (text.includes("refund") ? 4 : 0) + (text.includes("days") ? 3 : 0) + (text.includes("digital") ? -2 : 0);
return { index: i, score: Math.max(0, Math.min(10, s + 2)) };
});
return JSON.stringify({ judgements: scores });
}
function llmRerank(question, candidates, keep) {
const prompt = [
"Rate how well each passage answers the question, from 0 (irrelevant) to 10 (directly answers).",
"Reply with JSON: {\"judgements\": [{\"index\": number, \"score\": number}]}",
"Question: " + question,
...candidates.map((c, i) => "[" + i + "] " + c),
].join("\n");
const reply = JSON.parse(llm(prompt));
const valid = reply.judgements.filter((j) => Number.isInteger(j.index) && j.index >= 0 && j.index < candidates.length);
return valid.sort((a, b) => b.score - a.score).slice(0, keep)
.map((j) => ({ score: j.score, text: candidates[j.index] }));
}
const candidates = [
"Digital downloads cannot be refunded once opened.",
"Our customer service team is available 24/7.",
"Refunds are accepted within 30 days of purchase with a receipt.",
"Shipping takes 3 to 5 days.",
];
console.log(llmRerank("How many days do I have to get a refund?", candidates, 2));A listwise LLM reranker: number the passages, ask for scores as JSON, validate the indexes the model returns (never trust them blindly), sort and keep the best. In production use structured output (see structured-output) so the reply is guaranteed to match the schema.
Cross-encoder reranking with sentence-transformers (Python)
# pip install sentence-transformers chromadb
import chromadb
from sentence_transformers import SentenceTransformer, CrossEncoder
embedder = SentenceTransformer("all-MiniLM-L6-v2") # bi-encoder (stage 1)
reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2") # cross-encoder (stage 2)
col = chromadb.PersistentClient(path="./db").get_or_create_collection("docs")
def search(question: str, fetch_k: int = 30, keep: int = 5, min_score: float | None = None):
# Stage 1: cast a wide net
q = embedder.encode([question], normalize_embeddings=True).tolist()
res = col.query(query_embeddings=q, n_results=fetch_k)
candidates = list(zip(res["documents"][0], res["metadatas"][0]))
if not candidates:
return []
# Stage 2: score each (question, chunk) pair jointly
scores = reranker.predict([(question, doc) for doc, _ in candidates])
ranked = sorted(zip(scores, candidates), key=lambda x: x[0], reverse=True)
results = [{"score": float(s), "text": doc, "meta": meta} for s, (doc, meta) in ranked[:keep]]
if min_score is not None:
results = [r for r in results if r["score"] >= min_score]
return results
for r in search("How many vacation days do new employees get?"):
print(f"{r['score']:7.3f} {r['meta'].get('source')} {r['text'][:80]}")Load both models once at startup, not per request. The cross-encoder outputs one relevance score per pair (higher is more relevant); the scale depends on the model, so compare scores within a query and calibrate any min_score threshold on your own data. Batch the pairs as shown - predict() processes the list efficiently, and a GPU makes it much faster.
LLM reranking with Claude and structured output (Python)
import anthropic
from pydantic import BaseModel
client = anthropic.Anthropic()
class Judgement(BaseModel):
index: int
score: int # 0-10
class Ranking(BaseModel):
judgements: list[Judgement]
def llm_rerank(question: str, passages: list[str], keep: int = 5) -> list[tuple[int, str]]:
numbered = "\n".join(f'<passage index="{i}">{p}</passage>' for i, p in enumerate(passages))
resp = client.messages.parse(
model="claude-opus-5-5",
max_tokens=16000,
output_config={"effort": "low"}, # a simple judging task: keep it fast
messages=[{"role": "user", "content":
f"Question: {question}\n\n{numbered}\n\n"
"Score every passage from 0 (irrelevant) to 10 (directly answers the question)."}],
output_format=Ranking,
)
judged = [j for j in resp.parsed_output.judgements if 0 <= j.index < len(passages)]
judged.sort(key=lambda j: j.score, reverse=True)
return [(j.score, passages[j.index]) for j in judged[:keep]]LLM reranking lets you describe relevance in words ('prefer official policy documents over chat messages'), at the cost of one model call per query. Keep the candidate list short, validate indexes, and fall back to the first-stage order if the call fails.
How it works
Stage 1 (recall): hybrid or vector search returns fetch_k candidates, typically 20-100. Cost is roughly constant in corpus size thanks to the index.
Stage 2 (precision): for each candidate, build the pair (query, chunk text) and run the reranker. A cross-encoder feeds [query] [SEP] [chunk] through a transformer and a small output layer produces a single relevance number. Cost is linear in fetch_k and in chunk length, which is why you rerank tens of candidates, not thousands, and why chunks should not be huge.
Selection: sort by reranker score, optionally drop candidates below a calibrated threshold, deduplicate near-identical chunks, and keep the top keep (often 3-8) for the prompt. Some systems also put the best chunks at the start and end of the context, because models can pay less attention to the middle of very long contexts.
Latency budget: a small cross-encoder on a CPU can take tens to hundreds of milliseconds for a few dozen short pairs; a GPU or hosted reranker is much faster. Measure on your hardware and reduce fetch_k or chunk length if needed.
Training data note: MS MARCO-trained cross-encoders are good general-purpose web-search rerankers. For specialised domains (legal, medical, code), evaluate them on your own questions; domain-specific or fine-tuned rerankers, or an LLM reranker, may do better.
question
|
v
+------------------------+ millions of chunks
| stage 1: hybrid / ANN | fast, approximate
+------------------------+
| top 30 candidates
v
+------------------------+ 30 x (question, chunk)
| stage 2: cross-encoder | slow, accurate,
| or LLM reranker | reads pair together
+------------------------+
| sorted, threshold
v
top 5 --> prompt --> LLMWhy does it exist?
Scalable search must pre-compute document representations, which caps how well it can judge relevance; accurate relevance judgement needs to read query and document together, which cannot scale to the whole corpus. Two-stage retrieval gets the best of both: the index finds a manageable shortlist, the reranker makes the final call.
When to use it
Add a reranker when the right chunk is usually somewhere in the top 20-50 but often not in the top 3-5 (measure recall@k for both), when you want to send fewer chunks to cut cost and latency, or when you need a trustworthy score for 'nothing relevant found'. It is one of the highest-impact, lowest-effort upgrades to a basic RAG pipeline.
When not to use it
Skip it when first-stage retrieval already puts the answer at rank 1-3 almost always, when the latency budget is extremely tight, or when the corpus is so small that you can send all candidates to the LLM anyway. Do not use a reranker to compensate for a first stage that never finds the right chunk - fix recall first.
Common mistakes
Reranking only the top 5 from stage 1, which leaves nothing to reorder - fetch 20-50 candidates.
Expecting the reranker to find chunks the first stage missed.
Loading the cross-encoder model on every request instead of once at startup.
Treating raw cross-encoder scores as probabilities or reusing a threshold across different models without calibration.
Reranking very long chunks, which is slow and may be truncated by the reranker's input limit.
Trusting LLM reranker output without validating indexes and handling failures.
Not measuring: adding a reranker without checking it improves precision@k on your own questions.
Practice exercises
- Easy:
In the two-stage demo, change stage1 to return only the top 2. What does the reranker now return, and why does that show the importance of fetch_k?
- Easy:
Add a chunk 'Employees do not receive extra leave days for overtime.' to the demo. Where does it land before and after reranking?
- Medium:
Extend crossScore with a recency boost: chunks containing a year get +1 if the year is 2025 or later. Discuss the risk of mixing business rules into relevance scores.
- Medium:
Run the Python cross-encoder example on your own indexed documents. For 10 questions, print the stage-1 rank and the reranked rank of the chunk that really answers each question.
- Hard:
Build an evaluation comparing three pipelines on 30 questions: vector top-5, vector top-30 + cross-encoder top-5, and hybrid top-30 + cross-encoder top-5. Report precision@5 and average latency for each.
Interview questions
What is the difference between a bi-encoder and a cross-encoder?
A bi-encoder embeds query and document separately into vectors and compares them with cosine similarity, so document vectors can be pre-computed and indexed - fast and scalable. A cross-encoder processes the query and document together in one forward pass and outputs a relevance score, so attention can compare them directly - more accurate, but it must run per pair at query time, so it is used to rerank a small candidate set.
Why use two-stage retrieval?
Because high recall over a huge corpus needs a fast index, while high precision needs a slow model that reads query and document together. The first stage fetches a broad candidate set cheaply; the second reorders it accurately; the prompt gets a few highly relevant chunks.
How do you choose fetch_k and the number of chunks to keep?
fetch_k should be large enough that recall@fetch_k is high on your evaluation set (often 20-50) but small enough for the reranker's latency budget. The number kept is tuned for answer quality, context cost and latency, often 3-8. Measure both rather than guessing.
When would you use an LLM as a reranker?
When relevance is subtle or domain-specific and can be described in instructions, when candidate sets are small, or for high-value queries where cost and latency are acceptable. Use structured output for scores, validate the response, and keep a fallback to the first-stage order.
Can a reranker fix poor retrieval?
Only partially. It can fix bad ordering among candidates but cannot surface documents the first stage never returned. If the relevant chunk is missing from the candidates, improve the first stage (hybrid search, better chunking, query rewriting, a larger fetch_k).
What are the latency implications of reranking?
Cross-encoder cost grows with the number of candidates and their length, run sequentially or in batches per query. On CPU this can add noticeable latency; mitigations are smaller models, GPUs, hosted rerankers, fewer and shorter candidates, and caching results for repeated queries.