RAG vs fine-tuning: when to use which
Teams reach for fine-tuning far more often than the problem requires, usually because one question was never asked out loud: does the model lack knowledge, or does it lack behaviour? Retrieval fixes the first. Fine-tuning fixes the second. Choosing wrong costs months.
The question behind the question
Take one failing example and decide which category it falls into:
- Knowledge. The model does not know a fact, or knows an old version of it: a price list, an internal policy, a support ticket, a document published last week.
- Behaviour. The model knows enough but produces the wrong shape: verbose where you need JSON, wrong tone, missing tool calls, weak classification on your taxonomy.
Most "it is wrong about our product" complaints are knowledge problems. Most "we cannot ship this output" complaints are behaviour problems.
What fine-tuning actually changes
Fine-tuning continues training on your examples and moves the weights. It is supervised learning over pairs you write and curate — a few hundred to a few thousand for a narrow task — usually through LoRA or QLoRA adapters that train a small fraction of the parameters instead of the whole model. Vendor documentation is explicit that it teaches form rather than facts.
What improves reliably: output format, tone, domain vocabulary, tool-call syntax, classification boundaries, refusal style. What does not: storing facts. A fact baked into weights cannot be cited, cannot be updated without retraining, cannot be filtered per user, and returns confidently when it goes stale.
What retrieval actually changes
Retrieval — RAG, in the paper that named it — leaves the weights alone and changes what the model sees at answer time. Every stage can fail on its own:
- Chunking. Split on structure — headings, sections, table rows — not fixed character counts. Around 200 to 800 tokens per chunk with slight overlap suits most prose. Carry metadata with each chunk: source URL, document version, tenant and access rules.
- Embeddings. The model used at query time must be the one used at index time. Swap embedding models and the whole corpus has to be re-embedded.
- Search. Dense vectors miss exact identifiers such as error codes and part numbers; keyword search misses paraphrase. Run both and merge the results.
- Reranking. A cross-encoder over the top 50 to 100 candidates, keeping the best 3 to 8, usually improves answers more than any change to the generator.
- Grounding. Pass numbered sources, require citations, and return "I don't know" when retrieval comes back empty instead of letting the model answer from its weights.
def answer(question, tenant_id, pool=50, keep=5):
q = embed(question, model=EMBED_MODEL) # same model as indexing
hits = index.search(q, k=pool, filter={"tenant_id": tenant_id})
if not hits:
return "I don't know based on the available documents."
docs = rerank(question, hits)[:keep] # cross-encoder, not cosine
context = "\n\n".join(f"[{i+1}] {d.text}" for i, d in enumerate(docs))
reply = llm("Answer using only the sources below and cite them as [n].\n\n"
+ context + "\n\nQ: " + question)
return reply, [d.source_url for d in docs]
Why fine-tuning is a poor substitute for retrieval
The failure is predictable. Fine-tune on a corpus of support tickets and the model will discuss your product in a plausible voice — including the feature you deprecated last month, because that knowledge is frozen at training time. From there:
- Changing one policy means retraining, re-validating and redeploying.
- Deleting one record means the same, and there is no way to demonstrate the fact is gone from the weights.
- Per-user permissions cannot be enforced: everyone gets whatever the weights contain.
- Nothing can be cited, so reviewers cannot check an answer.
- Out-of-scope questions degrade instead of abstaining, because the model has been pushed toward your distribution.
Retrieval turns each of those into a cheap operation: reindex one document, delete one chunk, pre-filter by tenant.
Cost, maintenance and privacy
| Dimension | Fine-tuning | Retrieval |
|---|---|---|
| Time to first useful result | Weeks, mostly spent curating examples | Days, mostly spent on ingestion and evals |
| Main cost driver | Labelled data, training runs, a dedicated endpoint | Indexing and storage, plus extra tokens and latency per answer |
| Freshness | Frozen at the training run | As fresh as the last index job |
| Deleting one fact | Retrain, with no proof of removal | Delete the chunk |
| Per-user access | Not possible | Filter inside the query |
| Changing base model | Re-tune the adapter | Usually nothing to redo |
Privacy deserves its own sentence. Retrieval still sends retrieved chunks to whichever model answers, so a hosted API sees your documents; embedding and reranking locally keeps the index in your infrastructure. Fine-tuning on sensitive data is harder to undo: weights can memorise and reproduce training text, vendors may retain training data, and individual records cannot be revoked. Do not fine-tune on anything you would refuse to leak.
The hybrid that usually wins
- Retrieve first for anything that must be current, citable or permission-aware.
- Fine-tune a LoRA for house style: the exact JSON envelope, the mandated greeting, the approved refusal wording, your label set. A few hundred good examples are enough.
- Fine-tune the retriever instead of the generator when your vocabulary is unusual — statute citations, SKUs, error codes. An embedding model trained on your query and document pairs often beats a larger generator.
- Distil rather than train from scratch: label examples with a strong model, fine-tune a small one for the narrow task, then serve it cheaply.
Evaluating, then deciding
The two halves need different scoreboards, and "it looked good in the demo" is not one.
- Retrieval: collect 100 to 300 real questions with the document that should answer each. Track recall@k, MRR or nDCG, and how often unanswerable questions are correctly refused.
- Generation: faithfulness — does every claim appear in the cited chunk — plus citation accuracy and task success. Judge models are useful for triage, not for certification.
- Fine-tuning: a held-out set scored on exact match, schema validity and human ratings, alongside a general-capability regression check, because narrow tuning can quietly damage unrelated behaviour.
Then run the checklist in order:
- Try a better prompt and a handful of examples first. Formatting complaints often dissolve there.
- If answers depend on documents that change, or must be cited or access-filtered, use retrieval.
- If retrieval returns nothing useful, fix ingestion, chunking and reranking. Fine-tuning cannot retrieve what was never indexed.
- If the model knows the facts but will not hold the format, tone or labels, fine-tune.
- If both are true, do both and keep them separate so each can be scored on its own.
- If the task is small and stable, a prompted small model with a strict schema may beat both. Measure before you train.
Last updated 19 Sep 2026