LLM Techniques
Four Ways to Extend a Model
Once a model exists, there are four main ways to change or extend its behaviour without rebuilding it. These are not mutually exclusive — enterprise products typically layer several. When customers ask "can we customise it for our environment?" they are describing one of these four, even if they don't know which one yet.
- Prompt engineering: shapes the model's input at runtime — fastest and cheapest; no retraining required
- RAG (Retrieval-Augmented Generation): injects external knowledge at query time — the standard pattern for grounding answers in private or up-to-date data
- Fine-tuning: retrains on domain examples — used when behaviour gaps can't be fixed through prompting alone
- RLHF (Reinforcement Learning from Human Feedback): aligns model outputs using human preference signals — applied during model development by the vendor, not post-deployment by the customer
RAG: The Enterprise Standard
RAG is the dominant enterprise pattern: embed the user query, find the most relevant chunks from a private knowledge base via vector search, inject those chunks into the LLM prompt as context, generate an answer grounded in retrieved data. For security: lets a SecOps copilot answer about your environment, your runbooks, your past incidents — without retraining anything. Cheaper than fine-tuning. Updates instantly when source data changes. Provides citations. Keeps proprietary data out of the base model.
Where RAG Fails
If retrieval is bad, generation is bad. Chunk strategy, embedding model choice, and reranking matter enormously. A common pilot failure is shipping naive RAG and blaming the LLM when retrieval was the actual problem. Ask vendors: what retrieval and reranking strategy does the product use? "We use RAG" is not sufficient — it is the beginning of the conversation, not the answer.
Prompt Engineering: The Right Default
You change behavior by changing instructions. Cheap, instant, reversible. Modern frontier models follow nuanced prompts well. Should be the default for 80%+ of use cases. Most enterprise customisation requests — tone, format, scope, persona — are solvable through prompt engineering. Escalate to RAG when the gap is knowledge; escalate to fine-tuning when the gap is behaviour that prompting genuinely cannot close.
Fine-Tuning vs RAG: Making the Call
Fine-tuning changes behavior by further training on examples. Expensive, slower to update, requires ML infrastructure. The diagnostic question: is the gap knowledge (what the model knows) or behaviour (how it responds)? Knowledge gaps → RAG. Behaviour gaps → fine-tuning or prompting. Most "we need a fine-tuned model on our security data" requests are actually RAG requests in disguise. RLHF is a training-time technique vendors apply during model development — customers cannot do it post-deployment.
The Honest Take
Most "we need a fine-tuned model on our security data" requests are actually RAG requests in disguise. Fine-tuning is the right answer when you need behavioral consistency that prompting cannot reliably achieve — far rarer than prospects assume. A structured prompting sprint of four weeks often eliminates the need for fine-tuning entirely. When a prospect insists on fine-tuning, ask: have you already ruled out RAG plus prompt engineering for this specific gap?
Compare sentences and see which the model considers close.
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.