Conversational RAG
Making RAG work in a chat: follow-up questions, rewriting them into standalone search queries, managing chat history, and deciding when to retrieve at all.
What is it?
A single-question RAG pipeline breaks the moment users start chatting. Turn 1: 'How many days of annual leave do employees get?' - works. Turn 2: 'And what about contractors?' - the retriever embeds 'And what about contractors?', which says nothing about leave, and returns chunks about contractor invoicing. Turn 3: 'Can I carry it over?' - it refers to leave, two turns ago. Follow-up questions are full of references to earlier turns (pronouns like it and that, ellipsis like 'what about X?') that only make sense with the conversation history.
Conversational RAG solves two separate problems:
- Retrieving correctly for a follow-up. The search query must be standalone: understandable without the chat. The standard technique is query rewriting (also called query condensing or contextualisation): before retrieving, ask an LLM to rewrite the latest message, given the recent history, into a self-contained question - 'How many days of annual leave do contractors get?'. You search with the rewritten question but still answer the user's actual message.
- Answering in context. The LLM API is stateless (see calling-an-llm-api): you must send the conversation history on every turn so the model knows what 'it' and 'that' mean and can keep a consistent tone. The freshly retrieved sources go into the current turn.
Why not just embed the whole conversation? Concatenating all previous turns into the search query drags in old topics: after discussing leave and then switching to expenses, the query still 'smells' of leave. Embedding only the last message loses references. Rewriting gets the best of both - it pulls in exactly the context that the latest message depends on and nothing else. A cheap fallback is to embed the last user message plus the previous user message, but rewriting is more reliable.
Managing history. Every turn adds tokens, and every token is resent each turn, so cost and latency grow with conversation length, and eventually you hit the context window (see tokens-and-context-windows). Common strategies:
- Do not keep old retrieved sources in history. Store each past turn as just the user's question and the assistant's answer. Sources for the current question are retrieved fresh. This alone saves most of the tokens.
- Sliding window: keep only the last N turns verbatim.
- Summary memory: when history gets long, replace the oldest turns with an LLM-written summary ('The user is a contractor in Germany asking about leave and expenses').
- Keep the rewrite input small: the rewriter usually needs only the last few turns.
Deciding whether to retrieve. Not every message needs a search: 'thanks!', 'make that shorter', or 'translate your answer into Spanish' should be answered from the conversation alone. A rewriting step can also output a flag such as needs_retrieval, so you skip the search for chit-chat and edits. This is a small step towards agentic RAG, where the model decides when and what to search (see agentic-rag).
Citations across turns. Source numbers restart every turn, because each turn has its own retrieved set. Store the citations with each answer in your application (not in the prompt) so the UI can still show sources for older messages.
Explain like I'm 10
Imagine a librarian with a colleague in the archive who has never heard your conversation. You tell the librarian 'and what about contractors?'. A bad librarian shouts exactly that down to the archive and gets random files about contractors. A good librarian translates it first: 'Find the annual leave policy for contractors', because they remember what you were talking about. Then they answer you in the flow of the conversation. Query rewriting is that translation step.
Examples
Why follow-ups break retrieval, and how rewriting fixes it
const chunks = [
"Employees receive 25 days of annual leave per year.",
"Contractors are not entitled to paid annual leave; they invoice for days worked.",
"Contractors must submit invoices by the 5th of each month; late contractor invoices are paid next cycle.",
"Unbilled hours carry over to the next invoice.",
"Up to 5 unused days of annual leave carry over to the next year.",
];
const stop = new Set(["and", "what", "about", "the", "how", "many", "do", "get", "can", "it", "of", "per", "a", "i", "they", "to"]);
const words = (s) => s.toLowerCase().replace(/[^a-z0-9 ]/g, " ").split(" ")
.filter((w) => w && !stop.has(w)).map((w) => (w.length > 3 && w.endsWith("s") ? w.slice(0, -1) : w));
function search(query) { // toy retriever: count shared words
const q = new Set(words(query));
return chunks.map((c) => ({ c, s: words(c).filter((w) => q.has(w)).length }))
.sort((a, b) => b.s - a.s)[0].c;
}
// Fake rewriter: a real system sends history + latest message to an LLM with a rewrite instruction
function rewrite(history, latest) {
const topic = "annual leave"; // what the LLM would infer from history
if (/what about (\w+)/i.test(latest)) return "How many days of " + topic + " do " + latest.match(/what about (\w+)/i)[1] + " get?";
if (/\bit\b/i.test(latest)) return latest.replace(/\bit\b/i, "unused " + topic);
return latest;
}
const history = [{ role: "user", content: "How many days of annual leave do employees get?" },
{ role: "assistant", content: "Employees receive 25 days per year [1]." }];
for (const latest of ["And what about contractors?", "Can I carry it over?"]) {
const standalone = rewrite(history, latest);
console.log("User: " + latest);
console.log(" raw search -> " + search(latest));
console.log(" rewritten query : " + standalone);
console.log(" rewritten search-> " + search(standalone));
}Searching with the raw follow-up 'And what about contractors?' finds the invoicing chunk, because the only content word is 'contractors' and that chunk mentions it twice. The rewritten, standalone question finds the chunk that actually answers it. 'Can I carry it over?' is similar: 'it' carries no meaning for a retriever, so the raw search matches a chunk about unbilled hours carrying over. The fake rewriter uses rules only to stay offline; a real LLM handles arbitrary phrasing.
A chat loop with fresh sources per turn and bounded history
const MAX_HISTORY_TURNS = 2; // keep the last 2 user/assistant pairs verbatim
const approxTokens = (s) => Math.ceil(s.length / 4); // rough rule of thumb for English
function retrieve(q) { return ["(top chunks for: " + q + ")"]; } // stand-in retriever
function rewrite(history, msg) { return history.length ? msg + " [made standalone]" : msg; }
function llm(messages) { return "Answer to: " + messages[messages.length - 1].content.split("Question: ")[1]; }
let summary = "";
const history = []; // stores ONLY questions and answers, never old sources
function chat(userMessage) {
const query = rewrite(history, userMessage);
const sources = retrieve(query);
// Sources go into the CURRENT user turn only
const current = { role: "user", content: "<sources>" + sources.join(" ") + "</sources>\nQuestion: " + userMessage };
const messages = [...history, current];
const system = "Answer from the sources. Conversation summary so far: " + (summary || "(none)");
const answer = llm(messages);
history.push({ role: "user", content: userMessage }, { role: "assistant", content: answer });
while (history.length > MAX_HISTORY_TURNS * 2) { // fold the oldest pair into the summary
const [q, a] = history.splice(0, 2);
summary += " User asked: " + q.content + " Assistant said: " + a.content + "."; // an LLM would summarise
}
const sent = approxTokens(system) + messages.reduce((n, m) => n + approxTokens(m.content), 0);
console.log("turn: " + JSON.stringify(userMessage) + " | messages sent: " + messages.length + " | ~tokens: " + sent);
return answer;
}
["How many leave days do employees get?", "What about contractors?", "Can they carry it over?",
"Thanks! And how do I submit expenses?", "What is the deadline for that?"].forEach(chat);
console.log("history kept verbatim:", history.length / 2, "turns");
console.log("summary:", summary.trim());History holds only questions and answers, so each turn sends a bounded number of messages no matter how long the conversation runs; older turns survive as a summary in the system prompt. Sources are retrieved fresh for the current question and never accumulate. In a real app the summary is written by an LLM and messages strictly alternate user/assistant, starting with a user message.
Conversational RAG with Claude: rewrite, decide, retrieve, answer (Python)
import anthropic
from pydantic import BaseModel
from common import embed, get_collection # from the building-a-rag-pipeline project
client = anthropic.Anthropic()
MODEL = "claude-opus-5-5"
NO_ANSWER = "I don't know based on the provided documents."
class Rewrite(BaseModel):
needs_retrieval: bool # false for chit-chat, thanks, or requests to reformat the last answer
standalone_question: str # the latest message rewritten to make sense on its own
def rewrite(history: list[dict], message: str) -> Rewrite:
recent = "\n".join(f"{m['role']}: {m['content']}" for m in history[-6:]) # last 3 turns is plenty
resp = client.messages.parse(
model=MODEL, max_tokens=16000,
output_config={"effort": "low"},
messages=[{"role": "user", "content":
f"<conversation>\n{recent}\n</conversation>\n<latest>{message}</latest>\n\n"
"Rewrite the latest message as a standalone search question, resolving pronouns and references "
"using the conversation. Set needs_retrieval to false if it can be answered from the conversation alone."}],
output_format=Rewrite,
)
return resp.parsed_output
def retrieve(question: str, k: int = 5) -> list[dict]:
res = get_collection().query(query_embeddings=embed([question]), n_results=k)
return [{"text": d, "meta": m} for d, m in zip(res["documents"][0], res["metadatas"][0])]
class ChatSession:
def __init__(self):
self.history: list[dict] = [] # plain question/answer turns only
def ask(self, message: str) -> str:
r = rewrite(self.history, message) if self.history else Rewrite(needs_retrieval=True, standalone_question=message)
if r.needs_retrieval:
hits = retrieve(r.standalone_question)
sources = "\n".join(f'<source id="{i}" file="{h["meta"]["source"]}">{h["text"]}</source>'
for i, h in enumerate(hits, start=1))
content = f"<sources>\n{sources}\n</sources>\n\nQuestion: {message}"
else:
content = message
resp = client.messages.create(
model=MODEL, max_tokens=16000,
system=("You are an HR assistant in an ongoing conversation. Use the conversation to understand the "
"question. Answer factual questions ONLY from the sources in the latest message and cite [n]. "
f"If they do not contain the answer, say: {NO_ANSWER}"),
messages=self.history[-10:] + [{"role": "user", "content": content}], # last 5 turns + current
)
answer = "".join(b.text for b in resp.content if b.type == "text")
self.history += [{"role": "user", "content": message}, {"role": "assistant", "content": answer}]
return answer
chat = ChatSession()
for q in ["How many days of annual leave do employees get?", "And contractors?", "Thanks, can you make that shorter?"]:
print("User:", q)
print("Assistant:", chat.ask(q), "\n")Three ideas in one class: a cheap structured rewrite step that also decides whether retrieval is needed; retrieval with the standalone question; and an answer call that sends trimmed history plus the current message with fresh sources. History slices are taken in pairs (even numbers) so the messages still alternate user/assistant and start with a user turn. Old sources are never stored in history.
How it works
Per-turn flow: (1) receive the user message; (2) if there is history, call the rewriter with the last few turns and the message to get a standalone question and a needs-retrieval flag; (3) if needed, retrieve (and rerank) with the standalone question; (4) build the messages list: trimmed history (questions and answers only) followed by the current user message containing the sources and the original message; (5) call the LLM; (6) store the plain question and answer, plus the citations in your app's database.
Rewriter prompt essentials: include only recent turns, tell it to resolve pronouns and references, keep the user's intent and constraints (names, dates, product versions), output only the rewritten question (structured output makes this robust), and return the message unchanged if it is already standalone. Use a low effort setting or a smaller model: it is a simple task that adds latency to every turn.
Why answer with the original message, not the rewrite? The rewrite is for the retriever. The answering model has the full history and the user's exact words, including tone and requests like 'briefly' that a rewrite may drop.
Token growth. Without trimming, turn n resends all n-1 earlier turns, so total tokens across a conversation grow roughly quadratically. With a sliding window plus summary and no stored sources, each turn costs about the same. Prompt caching (see cost-and-latency) can further cut the cost of the repeated prefix: system prompt and early history.
Multi-question messages. 'What is the leave policy and how do I claim expenses?' needs two searches. A rewriter can return a list of standalone questions, retrieve for each and merge the results - the multi-query idea from advanced-rag-techniques.
history (Q/A only) + new message
|
v
+------------------+
| rewrite (LLM) | -> standalone question
| | -> needs_retrieval?
+------------------+
| yes | no
v |
retrieve + rerank |
| |
v v
messages = trimmed history
+ user: <sources> + original message
|
v
LLM -> answer [n]
|
store Q + A (+ citations in app DB)Why does it exist?
People do not talk in standalone search queries; they ask follow-ups, use pronouns and refine. A RAG system that only handles single, fully specified questions feels broken in a chat UI. Query rewriting and history management make retrieval robust to natural conversation while keeping cost bounded.
When to use it
Any chat interface on top of RAG: support assistants, internal knowledge bots, documentation chat. Add query rewriting as soon as users can ask follow-ups, history trimming once conversations get long, and a needs-retrieval decision when many messages are chit-chat or edits of previous answers.
When not to use it
For single-shot search boxes or batch question answering there is no history to manage; skip the rewrite call and its latency. If almost every message is standalone (for example a search-like UI), measure whether rewriting actually improves retrieval before paying for it on every turn.
Common mistakes
Retrieving with the raw follow-up message, so 'what about contractors?' searches for nothing useful.
Concatenating the whole conversation into the search query, pulling in stale topics.
Keeping every turn's retrieved sources in history, which explodes token usage and confuses citation numbers.
Sending history that does not alternate user/assistant correctly or starts with an assistant message after trimming.
Letting the rewriter change the user's intent or drop constraints such as dates, versions or names.
Running retrieval for 'thanks' or 'make it shorter', wasting time and sometimes replacing a good answer with an unrelated one.
Forgetting that source numbers restart each turn when showing citations for older messages.
Practice exercises
- Easy:
Write three follow-up questions that would break naive retrieval for a product-documentation bot, and the standalone rewrite you would want for each.
- Easy:
In the chat loop demo, set MAX_HISTORY_TURNS to 1 and to 10. How do 'messages sent' and '~tokens' change across the five turns?
- Medium:
Add a needsRetrieval check to the chat loop demo: messages like 'thanks', 'shorter please' or 'translate that' skip retrieval and are answered from history.
- Medium:
Write the rewrite prompt for the Python example so that it returns a list of standalone questions when the user asks two things at once, and update ChatSession to retrieve for each and merge the results.
- Hard:
Build a 15-conversation test set (each 3-4 turns) for your documents. Measure, per follow-up turn, whether the answering chunk is retrieved with (a) the raw message, (b) last two user messages concatenated, (c) LLM rewriting. Report recall@5 for each.
Interview questions
Why does naive RAG fail on follow-up questions?
Follow-ups depend on earlier turns through pronouns and ellipsis, so the latest message alone lacks the content needed for retrieval. Embedding 'what about contractors?' does not encode that the topic is annual leave, so the retriever returns the wrong chunks.
What is query rewriting in conversational RAG?
An LLM step that takes the recent conversation and the latest message and produces a standalone question that makes sense without the history. Retrieval uses the rewritten question; the final answer still uses the original message and the history.
How do you manage chat history in a RAG chatbot?
Store only questions and answers, not old retrieved sources; keep a sliding window of recent turns; summarise older turns into a short summary; ensure the trimmed messages still alternate correctly; and consider prompt caching for the stable prefix. Citations are stored in the application database rather than in the prompt.
How do you decide whether to retrieve on a given turn?
Classify the message: questions needing facts trigger retrieval, while chit-chat, thanks, and edits of the previous answer (shorten, translate, reformat) do not. This can be part of the rewriter's structured output, or a rule-based pre-filter. Agentic RAG generalises this by letting the model call a search tool when it decides it needs one.
Why answer with the original message rather than the rewritten query?
The rewrite is optimised for search and may lose nuances like tone, formatting requests or exact wording. The answering model has the full history and can interpret the original message correctly; the rewrite only needs to make retrieval work.
What are the latency and cost implications of conversational RAG?
An extra rewrite call per turn adds latency and cost, mitigated by a low effort setting or smaller model and by skipping it when there is no history. History adds input tokens every turn, which trimming, summaries and caching keep bounded.