RAG in Production
Keeping a RAG system correct, fresh, secure, fast and observable: incremental re-indexing, deletions, permissions, caching, latency budgets and monitoring.
What is it?
A RAG demo answers questions about a folder you indexed once. A production RAG system serves many users, over documents that change every day, some of which most users must never see, under a latency budget, with someone on call when answers go wrong. The pipeline is the same; what changes is everything around it.
1. Freshness and incremental re-indexing. Documents are edited, added and deleted continuously. Re-embedding the entire corpus on every change is slow and expensive, so production systems index incrementally: detect what changed and only process that. The standard tool is a content hash (a fingerprint such as SHA-256 of the text) stored with each document or chunk. On each sync, compare hashes: new hash means new or changed content to (re-)embed, same hash means skip, and a document that no longer exists at the source means delete its chunks. Changes are detected by scheduled syncs (every few minutes or hours), by change events from the source system (webhooks, database change streams), or both. Track the time from 'document changed' to 'searchable' as a freshness metric.
Deletions are a correctness and compliance issue. If a policy is withdrawn but its chunks remain in the index, the assistant keeps confidently quoting it. If a user's data must be erased, it must be erased from the vector store too. Make sure every chunk can be traced to its source document (metadata) and deleted by it.
2. Re-indexing everything safely. Some changes require rebuilding the whole index: a new embedding model, a new chunking strategy. Do it blue/green: build a new collection or table alongside the live one, run your evaluation set against it, then switch traffic with a config change and keep the old one for rollback. Store the embedding model name and a pipeline version in metadata so you always know how a vector was made.
3. Access control (permissions). RAG can leak information: if an intern's question retrieves a chunk from the salary spreadsheet, the model will happily summarise it. Rules:
- Filter at retrieval time, in code, using permission metadata copied from the source system (which groups or users may read each document). Never rely on the prompt ('do not reveal salaries') - prompts are not a security boundary.
- Keep permissions in sync with the source: when someone loses access to a document, the index must reflect it quickly.
- Multi-tenant apps (one system, many customers) should isolate tenants strictly: a mandatory tenant filter on every query, or a separate collection or database per tenant.
- Treat retrieved text as untrusted input: documents can contain hidden instructions (indirect prompt injection), so give the model no more power than the user has (see ai-security).
4. Caching. Several layers can be cached: embeddings of chunks (keyed by content hash, so unchanged chunks are never re-embedded), query embeddings and retrieval results for repeated questions, whole answers for exact repeat questions (keyed by normalised question + user permissions + index version, and invalidated when the index changes), and prompt caching on the LLM side for the large stable prefix of every request such as the system prompt and instructions (see cost-and-latency). Never share cached answers between users with different permissions.
5. Latency. A typical request: query rewrite (LLM call, if conversational) -> embedding -> vector/hybrid search -> rerank -> answer generation. Set a budget per step and measure each. The answer generation usually dominates; stream it so users see text immediately. Load models and clients once at startup, run independent steps (vector and keyword search) in parallel, keep the number and size of chunks small, and use a lower effort setting or a smaller model for helper steps like rewriting.
6. Monitoring and feedback. For every request, log the question, rewritten query, retrieved chunk ids and scores, reranker scores, prompt version, model, token usage, latency per step, the answer, its citations, and the user's feedback (thumbs up/down). Watch aggregate signals: the 'I don't know' rate, the share of answers without valid citations, retrieval score distributions (dropping scores often means content gaps or a broken ingest), cost per request, and p95 latency. Turn real failures into new evaluation cases and run the evaluation suite whenever the prompt, model, chunking or index changes (see evaluating-rag and observability-and-debugging).
7. Reliability. LLM APIs and embedding services can be slow or rate-limited: use timeouts, retries with exponential backoff (the official SDKs retry some errors automatically), and graceful fallbacks ('search results without an AI summary'). Run ingestion as a background job with a queue so a huge upload does not block queries.
Explain like I'm 10
A demo RAG is a home bookshelf. Production RAG is a public library: new books arrive daily and must be catalogued quickly (incremental indexing), recalled books must be pulled from the shelves (deletions), the rare-manuscripts room needs a key card checked at the door, not a polite sign (access control), popular questions get a ready-made answer sheet at the front desk (caching), and the head librarian tracks which questions went unanswered so the collection can improve (monitoring).
Examples
Incremental re-indexing with content hashes: only embed what changed
// FNV-1a: a tiny, fast, deterministic string hash (use SHA-256 in production)
function hash(s) {
let h = 0x811c9dc5;
for (let i = 0; i < s.length; i++) { h ^= s.charCodeAt(i); h = Math.imul(h, 0x01000193) >>> 0; }
return h.toString(16);
}
const chunk = (text) => text.split("\n\n").map((t) => t.trim()).filter(Boolean); // paragraph chunks
const index = new Map(); // chunkId -> { hash, source, text } (stand-in for the vector store)
let embedCalls = 0;
const embed = (text) => { embedCalls++; return [text.length]; }; // stand-in embedding
function sync(sourceFiles) {
const stats = { embedded: 0, skipped: 0, deleted: 0 };
const seen = new Set();
for (const [path, text] of Object.entries(sourceFiles)) {
chunk(text).forEach((piece, i) => {
const id = path + "#" + i, h = hash(piece);
seen.add(id);
const existing = index.get(id);
if (existing && existing.hash === h) { stats.skipped++; return; } // unchanged: no work
index.set(id, { hash: h, source: path, text: piece, vector: embed(piece) });
stats.embedded++;
});
}
for (const id of [...index.keys()]) { // gone at the source
if (!seen.has(id)) { index.delete(id); stats.deleted++; }
}
return stats;
}
let files = {
"leave.md": "Employees get 25 days of leave.\n\nUp to 5 days carry over.\n\nCarry-over expires 31 March.",
"expenses.md": "Submit receipts within 30 days.\n\nMeals are capped at 40 per day.",
"old-policy.md": "Leave is 22 days.",
};
console.log("initial sync: ", sync(files), "| chunks:", index.size);
console.log("no changes: ", sync(files), "| chunks:", index.size);
files["leave.md"] = files["leave.md"].replace("25 days", "26 days"); // one paragraph edited
delete files["old-policy.md"]; // a withdrawn document
files["remote.md"] = "Remote work up to 3 days a week."; // a new document
console.log("after edits: ", sync(files), "| chunks:", index.size);
console.log("total embedding calls:", embedCalls, "(a full rebuild every time would have used",
6 + 6 + 6, ")");
console.log("old policy still searchable?", [...index.values()].some((c) => c.source === "old-policy.md"));The second sync does no work at all; the third embeds only the edited paragraph and the new file, and deletes the withdrawn policy so it can no longer be quoted. Note the limitation of position-based ids (path#index): inserting a paragraph near the top shifts every later id and re-embeds them; production systems often use ids derived from the content hash or from stable section ids.
Permission-filtered retrieval and a safe answer cache
const chunks = [
{ text: "The office opens at 8am.", groups: ["everyone"] },
{ text: "Salary bands for 2025: L1 40k-50k, L2 50k-65k.", groups: ["hr"] },
{ text: "Production database failover runbook: run failover.sh.", groups: ["eng"] },
];
let indexVersion = 1;
let clock = 0; // fake time in seconds
const cache = new Map();
const TTL = 300;
let llmCalls = 0;
function retrieve(question, user) {
const allowed = new Set(["everyone", ...user.groups]);
const visible = chunks.filter((c) => c.groups.some((g) => allowed.has(g))); // filter BEFORE ranking
const q = question.toLowerCase().split(" ");
return visible.filter((c) => q.some((w) => w.length > 3 && c.text.toLowerCase().includes(w)));
}
function generate(question, hits) {
llmCalls++;
return hits.length ? "Based on the documents: " + hits.map((h) => h.text).join(" ") : "I don't know based on the documents you can access.";
}
function ask(question, user) {
const normalized = question.toLowerCase().replace(/[^a-z0-9 ]/g, "").trim();
// The key includes permissions and index version, so users never see answers they could not get themselves
const key = indexVersion + "|" + [...user.groups].sort().join(",") + "|" + normalized;
const hit = cache.get(key);
if (hit && clock - hit.at < TTL) return "(cached) " + hit.answer;
const answer = generate(question, retrieve(question, user));
cache.set(key, { answer: answer, at: clock });
return answer;
}
const alice = { name: "alice", groups: ["hr"] }, bob = { name: "bob", groups: ["eng"] };
console.log("alice:", ask("What are the salary bands?", alice));
console.log("bob: ", ask("What are the salary bands?", bob));
clock += 60;
console.log("alice:", ask("what are the SALARY bands", alice));
chunks[1].text = "Salary bands for 2026: L1 42k-52k, L2 52k-68k.";
indexVersion++; // re-index bumps the version: old cache entries are ignored
console.log("alice:", ask("What are the salary bands?", alice));
console.log("LLM calls:", llmCalls);Permissions are enforced by filtering chunks before ranking, in code - Bob cannot retrieve the salary chunk however he phrases the question. The cache key includes the user's groups and the index version, so a cached HR answer is never served to Bob and an answer built from outdated content disappears as soon as the index changes.
Incremental ingest for the Chroma project: hashes, upserts and deletions (Python)
import hashlib
from pathlib import Path
from common import embed, get_collection # from building-a-rag-pipeline
from ingest import load, clean, chunk, SUPPORTED
def sha256(text: str) -> str:
return hashlib.sha256(text.encode("utf-8")).hexdigest()
def sync(folder: str = "docs", access_group: str = "everyone") -> None:
col = get_collection()
paths = {p.as_posix(): p for p in Path(folder).rglob("*") if p.suffix.lower() in SUPPORTED}
# 1. Files deleted at the source: remove their chunks
indexed = {m["path"] for m in col.get(include=["metadatas"])["metadatas"]}
for gone in indexed - paths.keys():
col.delete(where={"path": gone})
print("deleted", gone)
# 2. New or changed files: compare a whole-file hash stored on every chunk
for key, path in sorted(paths.items()):
pages = list(load(path))
file_hash = sha256("".join(text for _, text in pages))
existing = col.get(where={"path": key}, include=["metadatas"], limit=1)["metadatas"]
if existing and existing[0].get("file_hash") == file_hash:
continue # unchanged: no embedding cost
ids, texts, metas = [], [], []
for page, raw in pages:
for n, piece in enumerate(chunk(clean(raw))):
ids.append(f"{key}::p{page}::c{n}")
texts.append(piece)
metas.append({"source": path.name, "path": key, "page": page, "chunk": n,
"file_hash": file_hash, "access_group": access_group,
"embed_model": "all-MiniLM-L6-v2", "pipeline_version": 1})
col.delete(where={"path": key}) # drop stale chunks (file may have shrunk)
for i in range(0, len(texts), 100):
col.upsert(ids=ids[i:i + 100], documents=texts[i:i + 100],
embeddings=embed(texts[i:i + 100]), metadatas=metas[i:i + 100])
print(f"re-indexed {key}: {len(texts)} chunks")
def retrieve_for_user(question: str, user_groups: list[str], k: int = 20):
"""Permission filter applied inside the vector search, never in the prompt."""
return get_collection().query(
query_embeddings=embed([question]),
n_results=k,
where={"access_group": {"$in": ["everyone", *user_groups]}},
)
if __name__ == "__main__":
sync("docs")Run sync on a schedule (or on change events) instead of a full rebuild. Unchanged files cost one hash computation; changed files are re-chunked and re-embedded; files removed at the source lose their chunks. Every chunk records its access group, embedding model and pipeline version, so permission filtering, audits and blue/green rebuilds are possible. For large corpora, track known files and hashes in a small database table rather than reading all metadata from the vector store.
Production query path: prompt caching, streaming, timing and structured logs (Python)
import json
import time
import logging
import anthropic
logging.basicConfig(level=logging.INFO)
client = anthropic.Anthropic(max_retries=3, timeout=600.0) # create once at startup
# Large, stable instructions go first and are cached; per-question sources come after.
SYSTEM_BLOCKS = [{
"type": "text",
"text": open("prompts/rag_system_v3.txt").read(), # rules, style guide, examples
"cache_control": {"type": "ephemeral"},
}]
PROMPT_VERSION = "rag_system_v3"
def answer(question: str, user_id: str, hits: list[dict]) -> str:
t0 = time.perf_counter()
sources = "\n".join(f'<source id="{i}">{h["text"]}</source>' for i, h in enumerate(hits, start=1))
parts = []
with client.messages.stream(
model="claude-opus-5-5",
max_tokens=64000,
system=SYSTEM_BLOCKS,
messages=[{"role": "user", "content": f"<sources>\n{sources}\n</sources>\n\nQuestion: {question}"}],
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True) # in a web app: send to the client
parts.append(text)
final = stream.get_final_message()
print()
logging.info(json.dumps({
"event": "rag_answer",
"user": user_id,
"question": question,
"chunk_ids": [h["id"] for h in hits],
"prompt_version": PROMPT_VERSION,
"input_tokens": final.usage.input_tokens,
"output_tokens": final.usage.output_tokens,
"cache_read_tokens": final.usage.cache_read_input_tokens, # > 0 means the cached prefix was reused
"stop_reason": final.stop_reason,
"latency_s": round(time.perf_counter() - t0, 3),
"declined": "I don't know" in "".join(parts),
}))
return "".join(parts)Prompt caching works on prefixes: the stable system prompt comes first with cache_control, and the per-question sources come after it, so they do not invalidate the cache (very short prompts are below the minimum cacheable size and are not cached). Streaming gets the first words to the user quickly. One structured log line per request (chunk ids, prompt version, tokens, cache reads, stop reason, latency, decline flag) is the raw material for dashboards, alerts and debugging. Avoid logging sensitive document text or personal data unless your data policy allows it.
How it works
Ingestion service: a scheduled job or event consumer lists changed documents from each source (file storage, wiki, ticketing system), fetches content and permissions, hashes it, and pushes changed documents onto a queue. Workers load, clean, chunk, embed (in batches, with retries) and upsert chunks with metadata including source id, hash, access groups, timestamps, embedding model and pipeline version; they delete chunks of removed documents. Metrics: documents processed, failures, freshness lag.
Query service: authenticate the user and resolve their groups; optionally rewrite the question; check the answer cache (key = index version + permissions + normalised question); run vector and keyword search in parallel with a mandatory permission (and tenant) filter; fuse and rerank; build the grounded prompt with a cached stable prefix; stream the answer; validate citations; log the trace; store feedback.
Versioning: treat the prompt, chunking settings, embedding model and LLM as versioned configuration. Every logged answer records the versions used, so a quality change can be traced to the change that caused it, and the evaluation suite runs before any version goes live.
Capacity: embedding during ingestion is batch work and can run on cheaper or slower hardware; query-time work is latency-sensitive. Vector indexes need memory proportional to vectors x dimensions (plus graph links for HNSW), which drives instance sizing; compression or smaller-dimension embeddings reduce it.
Safety: besides permission filtering, scan or sanitise ingested content where appropriate, mark retrieved text as data in the prompt, do not give the answering model tools that can act on behalf of the user unless needed, and keep personal data out of logs or redact it (see ai-security).
SOURCES (wiki, drive, tickets)
| change events / scheduled sync
v
[ queue ] -> workers: hash? changed -> load,
chunk, embed, upsert(+acl, ver)
removed -> delete chunks
|
v
+--------------+
user -> auth --> | vector store | <- filter:
groups | +--------------+ tenant+acl
v |
cache? -> rewrite -> search -> rerank
| |
| prompt (cached prefix) + sources
v v
logs/metrics <---- stream answer + citations
feedback -> eval set -> regression testsWhy does it exist?
The failure modes of RAG in production are rarely about the model: they are stale content, leaked documents, slow responses, runaway cost and silent quality regressions. Incremental indexing, permission filtering, caching, latency budgets and monitoring exist to turn a demo into a system people can rely on and engineers can operate.
When to use it
As soon as a RAG system has real users, changing documents or any confidential content. Prioritise in this order: permission filtering (security), deletions and freshness (correctness), logging and evaluation (you cannot improve what you cannot see), then caching and latency optimisation (cost and experience).
When not to use it
For a personal prototype over public, static documents, most of this is overkill: a one-off ingest script and a simple query function are enough. Do not build complex multi-layer caching before measuring that repeated questions are common, and do not add a separate queue and worker fleet if a nightly re-index of a small corpus takes a few minutes.
Common mistakes
Rebuilding the entire index on every change, making ingestion slow and expensive, or never re-indexing at all.
Forgetting deletions, so withdrawn or erased documents keep being retrieved and quoted.
Enforcing permissions in the prompt instead of filtering retrieval in code.
Sharing an answer cache across users with different permissions, or not invalidating it when the index changes.
Switching the embedding model in place instead of building a new index alongside and cutting over after evaluation.
Logging nothing about retrieval (chunk ids, scores, versions), making bad answers impossible to debug.
Ignoring latency per step and discovering in production that reranking or query rewriting doubled response time.
Logging full document text and personal data without considering privacy and retention rules.
Practice exercises
- Easy:
In the incremental demo, insert a new first paragraph into expenses.md. How many chunks are re-embedded and why? Propose an id scheme that avoids this.
- Easy:
List every field you would log for one RAG request and, for each, one question it would help you answer when debugging.
- Medium:
Extend the cache demo with a 'tenant' field on users and chunks so that users from different companies can never retrieve each other's chunks or share cache entries.
- Medium:
Add timing to the Python query path for each step (embedding, search, rerank, LLM time-to-first-token, total). Run 20 questions and report p50 and p95 for each step.
- Hard:
Implement a blue/green re-index for the Chroma project: build a new collection 'docs_v2' with a different chunk size, run an evaluation set on both, and switch a config value to the new collection only if recall@5 does not drop.
- Hard:
Design (on paper) the ingestion pipeline for a company wiki with 200,000 pages, per-page permissions and hundreds of edits per hour: sources of change events, queue, workers, idempotency, deletion handling, freshness SLO and monitoring.
Interview questions
How do you keep a RAG index fresh without re-embedding everything?
Incremental indexing: detect changes via change events or scheduled syncs, compare content hashes stored with each document or chunk, re-chunk and re-embed only new or changed content, upsert by deterministic ids, and delete chunks for removed documents. Track freshness lag as a metric.
How do you enforce document permissions in RAG?
Copy access-control information from the source system into chunk metadata, keep it in sync, and apply a mandatory filter in the retrieval query based on the authenticated user's identity and groups (and tenant). Never rely on the prompt to hide content, because anything in the context can end up in the answer.
How would you migrate to a new embedding model?
Build a new index alongside the existing one by re-embedding all chunks with the new model, record the model in metadata, evaluate retrieval quality on a test set, then switch query traffic (which must also use the new model for queries) via configuration, keeping the old index for rollback until confident.
What would you cache in a RAG system?
Chunk embeddings keyed by content hash, query embeddings and retrieval results for frequent queries, full answers for exact repeats (with keys including permissions and index version, plus a TTL), and the LLM's stable prompt prefix with prompt caching. Each cache needs a clear invalidation rule.
Where does latency come from in a RAG request, and how do you reduce it?
Query rewriting and answer generation (LLM calls) usually dominate, then reranking, embedding and search. Reduce it by streaming the answer, running searches in parallel, loading models once, limiting candidates and chunk size, using lower effort or smaller models for helper steps, caching, and skipping rewriting when there is no history.
What do you monitor in a production RAG system?
Per request: retrieved chunk ids and scores, prompt and model versions, tokens and cost, latency per step, stop reason, citations and user feedback. Aggregates: decline rate, uncited-answer rate, retrieval score distributions, error rates, p95 latency, cost per request and freshness lag. Failures feed the evaluation set, which runs on every change.
Why are deletions important?
Stale chunks keep being retrieved and quoted after documents are withdrawn or corrected, causing wrong answers, and retaining data that must be erased can violate legal obligations. Every chunk must be traceable to its source document so it can be removed when the source is.