Fine-Tuning vs RAG vs Prompting
The Decision Framework: Knowledge, Behaviour, or Format?
Every "customise the model" request decomposes into one question: what are you actually trying to change? If the gap is knowledge — the model doesn't know your documents, your environment, your current data — the answer is retrieval (RAG), because knowledge injected at query time stays current and auditable. If the gap is behaviour or style — output format, domain terminology, consistent tone, reliable adherence to a complex spec — start with prompting and escalate to fine-tuning only when prompting demonstrably plateaus. If the gap is capability — reasoning the base model simply cannot do — neither RAG nor fine-tuning will save you; pick a stronger model. Most real deployments layer techniques: a fine-tuned or well-prompted model with RAG on top. The framework's value is preventing the classic failure — fine-tuning to teach facts, which is slow, stale on arrival, and worse than retrieval at faithful recall.
- Knowledge gap → RAG; behaviour gap → prompting, then fine-tuning; capability gap → better base model
- Fine-tuning is poor at adding facts — weights are a lossy, unupdatable store compared to retrieval
- The techniques compose — this is a layering decision, not a three-way either/or
Prompting First: Cheaper Than You Think, Further Than You Think
Prompting is the highest-leverage, lowest-commitment adaptation, and modern frontier models follow far more nuanced instructions than the folk wisdom from earlier generations suggests. A structured system prompt with role, constraints, output schema, and a handful of well-chosen few-shot examples routinely closes gaps teams assumed needed training. Long context expands what prompting can carry — style guides, glossaries, API references, worked examples — and prompt caching makes a large static prefix economically sane by persisting its KV cache across requests, so the marginal cost of a rich prompt collapses. Prompting also iterates in minutes and rolls back instantly, while any training loop iterates in days. The discipline that makes this real is evaluation: a fixed test set scored per prompt version, so "better" is a measurement, not an impression.
LoRA and QLoRA: Fine-Tuning Without the Price Tag
Full fine-tuning updates every weight — enormous GPU memory, a full model copy per variant, and real risk of catastrophic forgetting. LoRA (low-rank adaptation) observes that fine-tuning changes are approximately low-rank, so it freezes the base model and trains small paired matrices alongside existing weight matrices — typically well under one percent of parameters. The adapter is megabytes, swappable at runtime, and one base model can serve many adapters for different customers or tasks. QLoRA pushes accessibility further: quantise the frozen base to 4-bit, train LoRA adapters in higher precision on top — capable open models become fine-tunable on a single workstation GPU. Quality on targeted behaviours approaches full fine-tuning for most practical cases, which is why adapter methods are the default and full fine-tuning the exception requiring justification.
- LoRA: freeze the base, train low-rank additions — a fraction of a percent of the parameters
- Adapters are small and hot-swappable: one base model, many cheap specialisations
- QLoRA: 4-bit frozen base + trained adapters — serious fine-tuning on single-GPU budgets
- Primary sources: "LoRA: Low-Rank Adaptation of Large Language Models" and "QLoRA: Efficient Finetuning of Quantized LLMs"
When Fine-Tuning Wins — and What It Really Costs
Fine-tuning earns its keep in specific situations: rigid output formats that must hold at very high reliability; deep domain language (legal, medical, niche codebases) where prompting stays subtly off; distilling a frontier model's behaviour into a small, cheap model for high-volume narrow tasks; and trimming long prompt scaffolding at scale, where baked-in behaviour undercuts per-request prompt tokens. The honest cost isn't the training run — it's everything around it: hundreds to thousands of clean labelled examples, an eval harness to prove improvement without regression, and a retraining pipeline for every base-model update, forever. That last item is the silent killer: base models improve fast, and each upgrade restarts your tuning work. The bar is simple — fine-tune when measured prompting-plus-RAG performance genuinely can't meet requirements, and the workload's scale amortises a permanent maintenance commitment.
- Strong cases: strict format reliability, deep domain style, distillation into smaller models, prompt-token economics at scale
- Data is the real cost — curation and labelling dwarf GPU spend
- Every base-model upgrade restarts the cycle — fine-tuning is a subscription, not a purchase
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.