Fine-tuning vs RAG vs Prompting
A decision guide: what prompting, retrieval and fine-tuning each change, what they cost, and which problem each actually solves.
What is it?
When a model does not do what you want, there are three broad levers. They solve different problems, and choosing the wrong one wastes weeks.
- Prompting changes the instructions and examples the model sees on each call: clearer task description, output format, few-shot examples, XML structure, a persona. Instant to change, free to iterate, fully reversible. It should always be tried first and pushed hard.
- RAG (retrieval-augmented generation) changes the knowledge available on each call by fetching relevant text from your data and putting it in the prompt. It is the right tool when the model lacks facts: private documents, fresh information, large or frequently changing corpora, and anything needing citations or per-user permissions.
- Fine-tuning changes the model's weights by training further on your examples (input, ideal output pairs). It teaches behaviour: a consistent style or format, a narrow task done reliably with a short prompt, a specialised output structure, or matching a larger model's quality on one task with a smaller, cheaper model (distillation).
The single most important rule: fine-tuning is a poor way to add knowledge. A model fine-tuned on your documents may learn their style and some facts, but it cannot cite sources, cannot be updated without retraining, cannot respect per-user permissions, and may still hallucinate details confidently. Knowledge belongs in retrieval; behaviour belongs in prompts or, when prompts are not enough, fine-tuning.
A related option is long context: if your whole corpus is small enough, put it directly in the prompt (with prompt caching to control cost). This avoids retrieval errors entirely, at the price of more tokens per call. It is often the right answer for a handful of documents.
Some terms you will meet: supervised fine-tuning (SFT) trains on input/output pairs; LoRA / PEFT (parameter-efficient fine-tuning) trains small adapter weights instead of the whole model, which makes fine-tuning open-weight models affordable; preference tuning trains on pairs of better/worse answers. Fine-tuning availability for hosted models varies by provider and model, so check your provider's documentation before planning around it.
These levers combine: a production system often uses careful prompting, RAG for knowledge, and sometimes a fine-tuned small model for one high-volume step like classification.
Explain like I'm 10
Imagine hiring a smart new employee. Prompting is giving them a clear brief and a few examples of good work. RAG is giving them access to the company wiki and telling them to look things up and cite the page. Fine-tuning is months of on-the-job training that changes their habits. You would never send someone on a training course just to memorise this week's price list; you would give them the price list.
Examples
A decision helper encoding the guide (runnable)
function recommend(need) {
const out = [];
// Always start here; most problems end here.
out.push("prompting (clear instructions, format, examples)");
if (need.missingKnowledge) {
if (need.corpusTokens <= need.contextBudgetTokens && !need.perUserPermissions)
out.push("long context + prompt caching (corpus fits)");
else out.push("RAG (retrieve relevant chunks per question)");
}
if (need.citationsRequired || need.changesOften) {
if (!out.some(x => x.startsWith("RAG")) && !out.some(x => x.startsWith("long"))) out.push("RAG");
}
const promptingFailed = need.evalScoreWithBestPrompt < need.targetScore;
if (promptingFailed && need.isBehaviourProblem && need.labelledExamples >= 500)
out.push("fine-tuning (behaviour/format/narrow task; re-check with evals)");
if (need.highVolumeNarrowTask && need.labelledExamples >= 500)
out.push("consider distilling into a smaller fine-tuned model for cost");
return out;
}
const scenarios = {
"Support bot over 2,000 help articles": { missingKnowledge: true, corpusTokens: 3e6, contextBudgetTokens: 1e5,
citationsRequired: true, changesOften: true, evalScoreWithBestPrompt: 0.9, targetScore: 0.85 },
"Answer from one 30-page policy PDF": { missingKnowledge: true, corpusTokens: 2e4, contextBudgetTokens: 1e5,
evalScoreWithBestPrompt: 0.9, targetScore: 0.85 },
"Strict house style for release notes": { missingKnowledge: false, evalScoreWithBestPrompt: 0.7, targetScore: 0.9,
isBehaviourProblem: true, labelledExamples: 1200 },
"Classify 5M tickets/day into 12 labels": { missingKnowledge: false, evalScoreWithBestPrompt: 0.93, targetScore: 0.9,
highVolumeNarrowTask: true, labelledExamples: 20000 },
};
for (const [name, need] of Object.entries(scenarios)) console.log(name, "\n ->", recommend(need).join("\n -> "));The thresholds (500 examples, the context budget) are illustrative, not rules; the structure is the point. Knowledge gaps lead to retrieval or long context. Fine-tuning appears only when the best prompt still misses the target on a behaviour problem and you have enough good examples, or when a high-volume narrow task could run on something cheaper.
The escalation ladder, measured at every rung (Python)
import anthropic
from evals import run_eval # your harness: returns a score on the golden set
client = anthropic.Anthropic()
def baseline(q):
r = client.messages.create(model="claude-opus-5-5", max_tokens=16000,
messages=[{"role": "user", "content": q}])
return "".join(b.text for b in r.content if b.type == "text")
STYLE = """You write answers for our customer help centre.
<rules>
- Start with the direct answer in one sentence.
- Then at most 3 short steps as a numbered list.
- Never promise timelines that are not in the provided documents.
</rules>
<example>
Q: How do I change my email?
A: You can change your email in Settings > Account.
1. Open Settings.
2. Choose Account > Email.
3. Confirm via the link we send.
</example>"""
def prompted(q):
r = client.messages.create(model="claude-opus-5-5", max_tokens=16000, system=STYLE,
messages=[{"role": "user", "content": q}])
return "".join(b.text for b in r.content if b.type == "text")
def with_rag(q):
chunks = retrieve(q, k=5) # your retriever
context = "\n\n".join(f"<doc id='{c['id']}'>{c['text']}</doc>" for c in chunks)
r = client.messages.create(model="claude-opus-5-5", max_tokens=16000, system=STYLE,
messages=[{"role": "user", "content":
f"<documents>\n{context}\n</documents>\n\nAnswer using only the documents. "
f"If they do not contain the answer, say so.\n\nQ: {q}"}])
return "".join(b.text for b in r.content if b.type == "text")
for name, fn in [("baseline", baseline), ("prompting", prompted), ("prompting+RAG", with_rag)]:
print(name, run_eval(fn))
# Only if a behaviour metric is still below target after this ladder do you
# collect training examples and evaluate a fine-tuned model on the SAME eval set.Each rung is measured on the same golden set, so you know what each change bought. Most teams find that prompting plus RAG reaches their target; fine-tuning becomes a data-backed decision rather than a first instinct.
What fine-tuning data looks like (generic chat-format JSONL)
{"messages": [{"role": "system", "content": "Rewrite commit messages as release notes."}, {"role": "user", "content": "fix: null check in invoice export"}, {"role": "assistant", "content": "Fixed: exporting an invoice with no line items no longer fails."}]}
{"messages": [{"role": "system", "content": "Rewrite commit messages as release notes."}, {"role": "user", "content": "feat: add CSV export to reports"}, {"role": "assistant", "content": "New: you can now export any report as a CSV file."}]}Each line is one training example in a common chat format; exact field names depend on the provider or training library you use. Quality matters more than volume: hundreds of consistent, correct examples beat thousands of noisy ones. Hold out a test split and evaluate against the prompted baseline before switching.
How it works
Prompting works because instruction-tuned models are trained to follow instructions and imitate examples in context (in-context learning). Nothing about the model changes; the next call without the prompt behaves as before.
RAG works because the model can read and use information placed in its context window. Knowledge stays in your data store, so updating it is just re-indexing a document; access control is a filter at retrieval time; every claim can be traced to a source chunk.
Fine-tuning runs additional training steps: the model sees your input, predicts the output token by token, and its weights are nudged to make your ideal outputs more likely. With LoRA, the original weights are frozen and only small low-rank adapter matrices are learned, which needs far less memory and lets you keep several adapters for different tasks. The result is a new model version you must host or deploy, version, evaluate and eventually retrain when the base model improves.
Costs compared. Prompting: longer prompts cost tokens on every call (mitigated by caching). RAG: an ingestion pipeline, a vector store, retrieval latency and maintenance. Fine-tuning: data collection and labelling (usually the biggest cost), training runs, hosting, and a re-evaluation every time you change the data or the base model. Fine-tuning can lower per-call cost when it lets a smaller model replace a larger one or removes long few-shot prompts.
Risks of fine-tuning: overfitting to the training examples, degrading general abilities outside the trained task (sometimes called catastrophic forgetting), baking errors from bad labels into the weights, and losing the benefit of base-model upgrades until you retrain.
What is wrong?
│
├─ unclear task / format ──> PROMPTING (first, always)
│
├─ model lacks facts ──┬─ small corpus ─> long context
│ (private, fresh, │ + caching
│ cited, permissioned)└─ large/changing ─> RAG
│
└─ behaviour still off after best prompt,
many good examples, narrow task ──> FINE-TUNE
(re-check on the same evals)Why does it exist?
Teams routinely reach for fine-tuning to 'teach the model our docs', spend weeks building a dataset, and end up with a model that still hallucinates, cannot cite and is stale on day one. This decision guide exists to match each lever to the problem it actually solves, and to make the choice with eval data instead of intuition.
When to use it
Use this guide at the start of any LLM feature and whenever quality plateaus. Concretely: prompting for every task; RAG or long context for knowledge; fine-tuning for stable, narrow, high-volume behaviour problems where you have hundreds or more high-quality examples and the best prompt measurably falls short, or to distil a capable model's behaviour into a cheaper one.
When not to use it
Do not fine-tune to add or update facts, to enforce permissions, or because a prompt 'feels long'. Do not build RAG for a corpus that fits comfortably in the prompt with caching. Do not skip the prompting rung: without a strong prompted baseline you cannot tell whether a fine-tune helped.
Common mistakes
Fine-tuning on documents to inject knowledge, then wondering why it still hallucinates and cannot cite.
Skipping a serious prompting effort and an eval baseline before investing in fine-tuning.
Building a vector database for a few documents that would fit in the context window.
Training on noisy, inconsistent examples and baking those errors into the weights.
Forgetting that a fine-tuned model must be re-evaluated and possibly retrained when the base model or data changes.
Treating the three levers as exclusive instead of combining them.
Practice exercises
- Easy:
Classify each as a knowledge or behaviour problem and pick a lever: (a) bot doesn't know this week's prices, (b) answers are too long, (c) model must output a strict internal XML format, (d) answers must cite policy sections.
- Easy:
Add two scenarios of your own to the decision helper and check its recommendations against your intuition. Where do they disagree?
- Medium:
Run the escalation ladder on a small task: baseline, improved prompt, prompt plus few-shot examples, prompt plus RAG. Record scores and token counts for each rung.
- Medium:
Write 30 high-quality training examples for a narrow task (for example commit message to release note) and a held-out test set of 10. Score the prompted baseline on the test set.
- Hard:
Write a one-page decision memo for a real or imagined product: the failure modes observed, the eval results at each rung, and a justified recommendation for or against fine-tuning, including ongoing maintenance cost.
Interview questions
Why is fine-tuning a poor way to add knowledge?
Knowledge learned into weights cannot be cited, is hard to update without retraining, cannot be filtered per user, and is recalled unreliably, so the model may still hallucinate details. RAG keeps knowledge in a store you can update instantly, filter by permission and cite.
When is fine-tuning the right choice?
For behaviour rather than knowledge: consistent style or format, a narrow task done reliably with short prompts, or distilling a large model's quality on one task into a smaller, cheaper model. Only after a strong prompted baseline falls short on evals, and only with enough high-quality examples.
What is LoRA?
Low-Rank Adaptation, a parameter-efficient fine-tuning method: the base model's weights are frozen and small low-rank matrices are trained and added to certain layers. It needs much less memory and storage than full fine-tuning and allows swapping adapters per task.
When would you choose long context over RAG?
When the relevant corpus is small enough to fit comfortably in the prompt, does not need per-user filtering, and is reused often enough that prompt caching makes it affordable. It eliminates retrieval misses and is simpler to build.
How would you justify a fine-tuning project to your team?
Show the eval baseline for the best prompt (and prompt plus RAG), the specific behaviour metric still below target, the available labelled data, the expected gain or cost saving, and the maintenance plan for retraining and re-evaluation. Then compare the fine-tuned model on the same held-out eval set.