Teams often frame the choice as RAG or fine-tuning, as if the two were competing ways to make a model smarter about their domain. They are not. Retrieval-augmented generation changes what the model can see at answer time: it fetches relevant documents and puts them in the prompt. Fine-tuning changes how the model behaves: it updates weights so the model follows a format, adopts a style, or performs a task more reliably. One supplies knowledge on demand; the other shapes behaviour. Most wrong decisions come from using one to fix a problem that belongs to the other.

This article gives a decision procedure that starts from evidence rather than preference. You label your current failures, classify each as a knowledge gap or a behaviour gap, choose the cheapest lever for each class, then measure every candidate on the same evaluation set. It covers the mechanics and costs of both approaches, a scoring function you can adapt, an evaluation harness, hybrids such as RAFT, a worked example and the failure modes of each path. For the basics of retrieval pipelines, start with how RAG bridges the knowledge gap.

Advertisement

What each approach actually changes

A RAG system has an indexing path and a query path. Documents are split into chunks, embedded and stored in a vector or hybrid index. At query time the system retrieves the top chunks, optionally reranks them, and inserts them into the prompt with an instruction to answer from them. Nothing in the model changes. Updating knowledge means re-indexing a document, which takes seconds, and every answer can cite the chunk it came from. Access control can be enforced at retrieval time, so two users asking the same question see only the documents they are allowed to see; see ACL-aware retrieval.

Fine-tuning trains the model further on examples. Supervised fine-tuning (SFT) on input and output pairs is the common case; preference methods such as DPO adjust which of two outputs the model favours. Parameter-efficient methods such as LoRA train small adapter matrices instead of all weights, which makes training cheap and lets one server host many adapters, as described in multi-LoRA serving. Once trained, the behaviour is baked in: no extra tokens in the prompt, no retrieval step, but also no citations, no per-user filtering, and a retraining cycle for every change.

Knowledge versus behaviour

The most useful distinction is between two kinds of gap. A knowledge gap is a failure where the model lacks a fact: your refund policy, last week's release notes, a customer's contract terms. A behaviour gap is a failure where the model has what it needs but does the wrong thing with it: wrong output format, wrong tone, ignoring a constraint, poor reasoning on a task shape specific to your domain, too long, too short.

Retrieval is the natural fix for knowledge gaps. Research supports this. Ovadia and colleagues (2023, Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs) found that retrieval outperformed unsupervised fine-tuning for injecting both previously seen and new knowledge, and that models struggle to absorb new facts through fine-tuning unless the facts are presented in many variations. Gekhman and colleagues (2024, Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?) found that examples carrying new facts are learned more slowly than examples consistent with what the model already knows, and that as they are learned the model's tendency to hallucinate increases. The practical reading: fine-tuning is a poor and risky way to teach facts.

Fine-tuning is the natural fix for behaviour gaps that prompting cannot close. If a model must emit a strict JSON schema, follow a house style across thousands of outputs, classify into your own taxonomy, or handle a task format it rarely saw in pre-training, a few hundred to a few thousand good examples often do what pages of instructions cannot. It also moves instructions out of the prompt, which saves tokens on every request.

A third class sits between them: grounding gaps, where the right document was retrieved but the model ignored it, mixed it with its own beliefs, or was distracted by irrelevant chunks. That is a retrieval-quality problem first and a behaviour problem second, and it is where hybrids earn their place.

Classify the gap first, then pick the cheapest lever that closes itFailing eval casesfrom real trafficWhat is missing?label each failurefacts it never sawfacts it saw, ignoredbehaviour, format, styleKnowledge gapRAGRetrieval / grounding gapfix retrieval, then RAFTBehaviour gapprompt, then fine-tuneSame eval set, same judge: baseline vs prompt vs RAG vs fine-tune vs bothship the cheapest arm that clears the barMost products end up with RAG for facts and a small fine-tune (or none) for behaviour; the labels tell you which.
Label each failing case, then route it: missing facts to retrieval, ignored facts to retrieval fixes and grounding training, and format or style failures to prompting and then fine-tuning. Every arm is measured on the same set.
Advertisement

The decision procedure

  1. Build an evaluation set from real traffic. Two to five hundred representative questions with reference answers or acceptance criteria, tagged by type. Without it every later step is opinion.
  2. Exhaust prompting. A clear system prompt, a few examples, and structured output constraints where the API offers them. Many apparent fine-tuning needs disappear here, and this baseline is what both alternatives must beat.
  3. Label the remaining failures as knowledge, grounding or behaviour. An hour of manual labelling on fifty failures is usually enough to see the split.
  4. Apply hard constraints. Facts that change weekly, a requirement for citations, or per-user permissions each force retrieval regardless of the split. A strict latency budget with no room for a retrieval round trip, or an offline device with no index, pushes towards fine-tuning or a smaller model.
  5. Choose the cheapest lever per class and measure each candidate arm on the same set with the same judge.
  6. Ship the cheapest arm that clears the bar, and keep the evaluation set as a regression suite.

The scoring function below encodes the same logic. Its thresholds are placeholders; set them from your own evaluation, not from this article.

from dataclasses import dataclass

@dataclass
class Need:
    knowledge_changes_per_month: int   # docs added or edited
    needs_citations: bool              # answers must point at a source
    per_user_access_control: bool      # users see different documents
    behaviour_failures_pct: float      # share of eval failures that are format/style/skill
    knowledge_failures_pct: float      # share that are missing or stale facts
    labelled_examples: int             # high-quality input/output pairs available
    p95_latency_budget_ms: int
    prompt_fixes_tried: bool

def recommend(n: Need) -> list[str]:
    plan = []
    if not n.prompt_fixes_tried:
        return ["prompting and few-shot examples first; re-measure"]
    if (n.knowledge_failures_pct >= 0.3 or n.knowledge_changes_per_month > 0
            or n.needs_citations or n.per_user_access_control):
        plan.append("RAG for knowledge")
    if n.behaviour_failures_pct >= 0.3:
        if n.labelled_examples >= 500:
            plan.append("LoRA fine-tune for behaviour")
        else:
            plan.append("collect examples; fine-tune when >= 500 clean pairs")
    if "RAG for knowledge" in plan and n.p95_latency_budget_ms < 300:
        plan.append("budget retrieval latency; consider caching or a smaller reranker")
    return plan or ["no change: failures are noise or eval bugs"]

Costs, as a model rather than a price list

Prices move monthly, so reason with variables. For RAG, the fixed cost is building and maintaining the ingestion pipeline and index; the variable cost is embedding new documents, storing vectors, the retrieval call, and the extra prompt tokens on every request. If each request adds k chunks of c tokens, you pay for roughly k times c more input tokens per call, forever, and you add retrieval latency, typically a vector search plus an optional reranker.

For fine-tuning, the fixed cost is the labelled data (usually the largest cost, because it needs domain experts), the training runs, and evaluation; the recurring cost is retraining whenever the base model, the task or the data changes, plus serving a custom model. Hosted fine-tuning often prices inference on tuned models differently from the base model, and self-hosting means paying for capacity whether or not it is used. Against that, a fine-tune can shorten every prompt by removing long instructions and examples, and can let a smaller model do a job that previously needed a larger one, which is the same economics as distillation.

A useful break-even question: tokens saved per request times requests per month, against training plus data plus the price difference of serving the tuned model. At low volume, prompting plus RAG almost always wins; at very high volume with a stable task, a fine-tuned smaller model can win decisively.

Evaluating the arms side by side

The decision is only as good as the comparison. Run every candidate arm on the same evaluation set, judged the same way, and break results down by failure tag, because an arm that fixes formatting while breaking facts can look neutral in an average.

ARMS = {
    "base":      lambda q: llm(SYSTEM, q),
    "prompted":  lambda q: llm(SYSTEM + FEW_SHOT, q),
    "rag":       lambda q: llm(SYSTEM, q, context=retrieve(q, k=5)),
    "ft":        lambda q: llm_ft(SYSTEM, q),
    "ft+rag":    lambda q: llm_ft(SYSTEM, q, context=retrieve(q, k=5)),
}

def run(eval_set, judge):
    rows = []
    for case in eval_set:                       # each case: question, reference, tags
        for arm, fn in ARMS.items():
            out = fn(case.question)
            rows.append({
                "arm": arm, "tag": case.tag,    # e.g. 'fact-new', 'fact-old', 'format'
                "correct": judge.correct(out, case.reference),
                "grounded": judge.supported_by_context(out) if "rag" in arm else None,
                "format_ok": validate_schema(out),
                "latency_ms": out.latency_ms, "cost": out.cost,
            })
    return summarise(rows, by=["arm", "tag"])  # compare arms per failure tag

Three measurements matter beyond correctness. Groundedness for the RAG arms: is each claim supported by retrieved context? Retrieval recall: did the right chunk appear in the top k at all? If not, no amount of generation tuning will help. And regressions on general ability for the fine-tuned arms, because a tune that fixes your format can degrade unrelated skills. Metric definitions and judge design are covered in RAG evaluation metrics.

Using both: hybrids

The approaches compose. The most common production shape is RAG for knowledge plus a light fine-tune for behaviour: the tuned model knows the output schema and tone, and retrieval supplies the facts. A more targeted hybrid trains the model to use retrieved context well. RAFT (Zhang and colleagues, 2024, RAFT: Adapting Language Model to Domain Specific RAG) trains on questions paired with the relevant document plus distractors, with some examples holding only distractors, with answers that quote the relevant passage in a chain of reasoning, so the model learns to find and cite the right evidence and ignore noise. It addresses grounding gaps directly.

Two cautions. First, fine-tune with the same retrieval format you will serve, including realistic noise; a model tuned on perfect context learns to trust whatever it is given. Second, a hybrid has two moving parts, so changes to the index, the chunker or the retriever can silently shift the distribution the tuned model was trained on. Re-run the evaluation after any change to either.

A worked example

A software company builds a support assistant. Its knowledge base has 4,000 articles, about 60 of which change each week. Answers must link to the source article, enterprise customers see private articles about their own deployments, and the assistant must return a JSON object with an answer, a confidence label and a list of article IDs for the ticketing system.

The team builds 300 evaluation questions from last quarter's tickets. A prompted base model with no retrieval fails 58 percent of them. Labelling fifty failures shows 36 missing or stale facts, 6 where the policy summary already in the system prompt held the answer but was ignored, and 8 malformed or chatty outputs. Weekly changes, citations and per-customer visibility all force retrieval, so RAG is the first arm. Adding hybrid retrieval with access filtering cuts failures sharply, and the residual failures shift towards format errors and grounding errors.

The team then tries structured output enforcement, which removes most of the malformed JSON, and a short instruction to answer only from the provided articles. The remaining grounding failures cluster around long articles with many similar sections, so they improve chunking and add a reranker rather than training. They hold a LoRA fine-tune in reserve and never need it. The decision was RAG plus prompting, reached in two weeks and backed by numbers on the team's own data; the numbers here are illustrative, and yours will differ.

Failure modes

  • Fine-tuning to teach facts. The model learns the phrasing of the training answers, recalls facts unreliably, and confidently states stale ones after the documents change.
  • RAG with poor retrieval. Generation is blamed for failures caused by the right chunk never being retrieved. Measure retrieval recall separately before touching the model.
  • Context stuffing. Retrieving twenty chunks to be safe increases cost and latency and gives the model more distractors.
  • Training on unrepresentative examples. A tune built from synthetic or cleaned-up examples fails on real, messy inputs.
  • Stale adapters. The base model is upgraded but the adapter, trained against the old weights, is not retrained; results degrade or the adapter cannot load.
  • No regression suite. A change that helps one tag hurts another, and nobody notices because only the average was tracked.

What to do next

  1. Collect two to five hundred real questions with reference answers and tag them by type.
  2. Measure a well-prompted baseline, then label fifty failures as knowledge, grounding or behaviour.
  3. If facts change, citations are required or users have different permissions, build retrieval first and measure its recall on its own.
  4. Fix format and style with prompting and structured output before considering a fine-tune; fine-tune only with several hundred clean, representative examples.
  5. Compare every arm on the same set by failure tag, including cost and latency, and ship the cheapest arm that clears the bar.
  6. Keep the evaluation set as a regression suite and re-run it on every change to the model, the index, the chunker or the prompt.
Key takeaway: RAG supplies knowledge at answer time and fine-tuning shapes behaviour in the weights; choose by classifying your real failures, not by preference. Use retrieval when facts change, need citations or depend on who is asking; use prompting and then fine-tuning for format, style and task skill; combine them, and consider grounding-focused training such as RAFT, when retrieved facts are ignored. Decide with one evaluation set, measured by failure tag, for every arm.