Agentic RAG
Retrieval as a tool: the model decides when to search, what to search for, and when it has enough to answer.
What is it?
Classic RAG (retrieval-augmented generation) is a fixed pipeline: take the user's question, retrieve the top-k chunks, put them in the prompt, generate an answer. It always retrieves exactly once, with the user's words as the query. That works well for direct questions, but fails when the first search misses, when a question needs facts from several places, or when no retrieval is needed at all.
Agentic RAG turns retrieval into a tool inside an agent loop. The model decides:
- Whether to retrieve at all ('What is 2 + 2?' needs no search; 'What is our refund policy?' does).
- What to search for: it can rewrite a vague question into precise queries, or split a comparison into one query per item.
- Where to search, when you offer several tools: a docs index, a ticket database, a SQL tool, the web.
- Whether the results are enough: if the hits are irrelevant or point to another document, it searches again with a new query (iterative retrieval), then answers.
Multi-hop questions are where this shines. 'Who approves refunds for the product our biggest customer uses?' needs: find the biggest customer, find their product, find that product's refund approver. A single similarity search for the whole question rarely finds all three pieces; an agent can follow the chain.
Designing the retrieval tool well is most of the work:
- Describe the corpus in the tool description ('Searches the internal HR and IT policy handbook. Not for customer data.').
- Return results with source ids (file, section, URL) so the model can cite them, and short enough snippets to keep context small.
- Offer useful parameters -
k, a metadata filter (source,date_after) - but keep the schema small. - Return an explicit 'no results' message so the model knows to rephrase rather than invent.
The system prompt should set the rules: search before answering questions about internal facts, cite sources by id, try a different query when results look irrelevant, stop after a few searches, and say 'I could not find this' instead of guessing. Combined with a max-steps guard, this keeps cost bounded.
The trade-off: agentic RAG is slower and costs more per question (several model calls instead of one), and it is less predictable. Many production systems use a hybrid - a fast fixed pipeline for most questions, and an agentic path for complex ones, chosen by a router (see Workflow Patterns).
Explain like I'm 10
Classic RAG is a librarian who, whatever you ask, fetches the five books whose titles best match your sentence and hands them over. Agentic RAG is a research assistant who reads your question, decides whether books are needed, looks something up, notices a footnote pointing to another book, fetches that one too, and comes back with an answer plus the page numbers.
Examples
Iterative retrieval with a fake model (runnable)
const docs = [
{ id: "refunds.md", text: "Refunds are allowed within 30 days of purchase. Return shipping costs are covered by the shipping policy." },
{ id: "shipping.md", text: "Shipping policy: we pay return shipping for damaged items; otherwise the customer pays 5 EUR." },
{ id: "careers.md", text: "We are hiring backend engineers in Lisbon." },
];
// Retrieval tool: simple keyword overlap scoring (real systems use embeddings/BM25).
function searchDocs(query, k) {
const words = query.toLowerCase().split(" ").filter((w) => w.length > 3);
return docs
.map((d) => ({ id: d.id, text: d.text, score: words.filter((w) => d.text.toLowerCase().includes(w)).length }))
.filter((d) => d.score > 0)
.sort((a, b) => b.score - a.score)
.slice(0, k);
}
// Fake model: decides whether/what to search based on what it has seen.
function model(question, observations) {
if (!question.toLowerCase().includes("refund") && !question.toLowerCase().includes("return")) {
return { answer: "No retrieval needed: 12 * 12 = 144." };
}
if (observations.length === 0) return { search: "refund policy purchase" };
const seen = observations.map((o) => o.id);
if (seen.includes("refunds.md") && !seen.includes("shipping.md")) {
return { search: "shipping policy damaged items" }; // follow the pointer in refunds.md
}
return { answer: "You can get a refund within 30 days [refunds.md]. Return shipping is free for damaged items, otherwise it costs 5 EUR [shipping.md]." };
}
function agenticRag(question, maxSearches) {
const observations = [];
for (let i = 0; i <= maxSearches; i++) {
const step = model(question, observations);
if (step.answer) return step.answer;
if (i === maxSearches) break;
const hits = searchDocs(step.search, 1).filter((h) => !observations.some((o) => o.id === h.id));
console.log(" search #" + (i + 1) + ": '" + step.search + "' -> " + (hits.map((h) => h.id).join(", ") || "nothing new"));
observations.push(...hits);
}
return "I could not find enough information to answer.";
}
console.log("Q1:", agenticRag("Can I return a broken kettle and who pays shipping?", 3));
console.log("Q2:", agenticRag("What is 12 times 12?", 3));For Q1 the first search finds refunds.md, which points to the shipping policy, so the model searches again before answering with citations. For Q2 it decides no retrieval is needed. A fixed pipeline would have searched once for both.
A search_docs tool backed by Chroma + sentence-transformers
# pip install anthropic chromadb sentence-transformers
import chromadb
from sentence_transformers import SentenceTransformer
embedder = SentenceTransformer("all-MiniLM-L6-v2") # 384-dimensional vectors
db = chromadb.PersistentClient(path="./db")
col = db.get_or_create_collection("docs") # filled earlier by your ingest script
SEARCH_TOOL = {
"name": "search_docs",
"description": (
"Semantic search over the company handbook (HR, IT, travel and expense policies). "
"Returns up to k passages, each starting with its source id in [brackets]. "
"Use short, specific queries; search again with different wording if results look "
"irrelevant. Cite source ids in your answer. Not for customer or order data."
),
"input_schema": {
"type": "object",
"properties": {
"query": {"type": "string", "description": "A focused search query"},
"k": {"type": "integer", "minimum": 1, "maximum": 8, "description": "Number of passages, default 4"},
},
"required": ["query"],
},
}
def search_docs(query: str, k: int = 4) -> str:
q = embedder.encode([query], normalize_embeddings=True)[0].tolist()
res = col.query(query_embeddings=[q], n_results=min(int(k), 8))
passages = []
for doc, meta, dist in zip(res["documents"][0], res["metadatas"][0], res["distances"][0]):
passages.append(f"[{meta['source']}] (distance {dist:.3f}) {doc[:800]}")
return "\n\n".join(passages) or "No results. Try different keywords."
SYSTEM = """You answer questions about the company handbook.
- For any question about policies, call search_docs before answering; do not rely on memory.
- If results are irrelevant, rephrase and search again (at most 4 searches).
- Answer only from retrieved passages and cite them like [travel.md].
- If the passages do not contain the answer, say you could not find it."""
# Plug SEARCH_TOOL + search_docs into the agent loop from Build an Agent from Scratch:
# tools=[SEARCH_TOOL], system=SYSTEM, and route "search_docs" in run_tool.The tool description tells the model what the corpus contains and how to use it; the result format puts source ids first so citations are easy. Lower distance means closer. Everything else is the same agent loop.
Classic vs agentic RAG, step by step
Question: "Compare the travel per-diem for Berlin and Tokyo."
Classic RAG (1 retrieval, 1 LLM call)
retrieve("Compare the travel per-diem for Berlin and Tokyo") -> top 4 chunks
-> maybe both cities, maybe only one -> answer (may guess the missing one)
Agentic RAG (several tool calls, several LLM calls)
model: search_docs("per diem Berlin") -> [travel.md#europe]
model: search_docs("per diem Tokyo") -> [travel.md#asia]
model: answer with both figures, citing travel.md#europe and travel.md#asiaSplitting a comparison into one query per item is a typical agentic move that a single similarity search cannot make.
How it works
Agentic RAG = the agent loop + a retrieval tool + rules in the system prompt. Each search result enters the context as a tool_result, so later steps can read earlier findings, notice gaps and refine queries. The model's final answer is grounded in the accumulated passages.
You can mix retrieval methods as separate tools (semantic search, keyword search, SQL over structured data, web search) or as one tool doing hybrid search internally. Fewer, well-described tools usually work better than many overlapping ones.
Cost control: cap the number of searches in the prompt and in code, keep passages short, and deduplicate passages already seen. Evaluate with the same metrics as RAG (did it retrieve the right sources, is the answer faithful) plus agent metrics (number of steps, unnecessary searches).
question
|
v
+--> model: need info? --no--> answer directly
| | yes
| v
| search_docs(query') ---> [vector / keyword index]
| |
| passages + source ids (tool_result)
| |
| enough? --no--> rewrite query ----+
| | yes |
| v |
| answer with [citations] |
+-------------------------------------+
(max searches guard)Why does it exist?
Fixed RAG pipelines make one retrieval with the user's raw wording, which breaks on vague questions, multi-part comparisons, multi-hop chains and questions that need no retrieval at all. Giving the model control over retrieval lets it reformulate, iterate and choose sources - the same way a person researches.
When to use it
Use agentic RAG for complex or multi-hop questions, research-style tasks, corpora spread across several sources, and assistants where many questions need no retrieval. It also suits conversational assistants where the model must decide if a follow-up needs a new search.
When not to use it
For simple FAQ-style lookups with strict latency or cost requirements, a fixed RAG pipeline (possibly with query rewriting and reranking) is faster, cheaper and easier to evaluate. Start there, measure failures, and add the agentic path only for the questions that need it.
Common mistakes
Vague tool descriptions that do not say what the corpus contains, so the model searches for things that are not there.
Returning passages without source ids, making citations impossible.
No limit on the number of searches, letting the agent loop on an unanswerable question.
Returning long passages so a few searches overflow the context.
Not telling the model it may answer without searching, causing pointless retrieval on every turn.
Not telling the model to say 'not found', so it fills gaps with guesses.
Evaluating only the final answer and ignoring whether the right documents were retrieved.
Practice exercises
- Easy:
In the runnable demo, set maxSearches to 1 and explain the new output for Q1.
- Medium:
Add a fourth document to the demo that the shipping policy points to (for example a 'damage claims' page) and make the fake model follow a three-hop chain.
- Medium:
Write tool descriptions for two retrieval tools over different corpora (product docs vs internal runbooks) that make it obvious which one to use.
- Hard:
Build the Python search_docs agent over 20 of your own documents. Create 10 questions (some multi-hop, some needing no retrieval) and log how many searches each takes and whether the cited sources are correct.
- Hard:
Implement a router: questions classified as simple go through a fixed one-shot RAG pipeline, complex ones through the agentic loop. Compare cost and accuracy against agentic-only.
Interview questions
What is agentic RAG and how does it differ from classic RAG?
Classic RAG always retrieves once with the user's query and then generates. Agentic RAG exposes retrieval as a tool in an agent loop, so the model decides whether to retrieve, writes its own queries, can search repeatedly or across sources, and stops when it has enough evidence.
What is a multi-hop question?
A question whose answer requires chaining facts from different documents, where each step depends on the previous one (find X, then use X to find Y). Single-shot similarity search rarely retrieves all pieces; iterative retrieval can.
How do you keep agentic RAG costs under control?
Limit searches in both the prompt and code, keep passages short, deduplicate results, use a router so simple questions take a cheap fixed path, and monitor steps per question.
What should a retrieval tool return?
A small number of concise passages, each with a stable source identifier (and optionally a score), plus an explicit message when nothing is found, so the model can cite sources and decide whether to search again.
How would you evaluate an agentic RAG system?
Retrieval metrics (did it find the relevant sources), answer faithfulness and relevance, citation correctness, plus trajectory metrics: number of searches, unnecessary or repeated searches, and cost and latency per question.