LLM Governance & Safety — When Each Layer Applies
Safety Is Not One Thing
When customers ask "is this AI safe?" they are actually asking several different questions at once. Safety in LLMs is built across three separate phases — before the model ships, while it runs in production, and continuously as it evolves. Each phase catches different failure modes. No single layer is enough on its own.
- Training time: safety and alignment baked into the model before it is ever deployed
- Deployment time: guardrails applied at runtime to constrain what the live model can do
- Production time: observability to know what the model is actually doing at scale
- Continuous evals: validation that runs across all phases to catch regressions and drift
Training Time — Safety Built In Before the Model Ships
The first line of defence happens during training itself. Before a model is released, the team running training shapes its behaviour and probes it for failure modes. By the time it reaches customers, these properties are baked into the weights.
- Instruction tuning (SFT): trains the model to follow instructions and be helpful
- Constitutional AI (CAI): a principles-based self-critique loop where the model evaluates its own outputs against a set of rules during training
- RLHF: shapes behaviour through human preference feedback — humans rank model outputs and the model learns to prefer higher-rated responses
- Red-teaming: adversarial probing by humans trying to find failure modes before release
Deployment Time — Guardrails on the Live Model
Once a model is deployed, a second layer of controls wraps around it at runtime. These are not inside the model — they sit between the user and the model, filtering what goes in and what comes out on every single request.
- Output filters: block unsafe, harmful, or off-policy content at runtime — the last check before a response reaches the user
- Scope and policy limits: define what the model is and is not allowed to do in this specific deployment
- Prompt injection defence: reduces the odds of adversarial inputs steering model behaviour — it limits risk rather than creating a hard boundary, so pair it with least privilege
- PII and data privacy: detects and redacts sensitive personal data in both inputs and outputs
Production Time — Knowing What the Model Is Doing at Scale
Deploying safely is not a one-time event. Once a model is live and handling real traffic, you need visibility into every interaction. Production observability catches the problems that training and deployment guardrails missed.
- Prompt and response tracing: a full audit trail of every interaction — essential for incident investigation and compliance
- Cost and token tracking: per-request spend, token usage, and budget alerts
- Output drift detection: flags quality degradation and behaviour shifts over time
- Latency and performance: response time, throughput, and SLA monitoring
Continuous Evals — Validation That Never Stops
Evals are automated tests that run across all three phases. They are the feedback loop that tells you whether the model is still doing what you expect — and whether any change made things better or worse. In a production system, evals run on every update.
- Benchmarks: standardised capability tests used to compare models and track capability over versions
- Human preference eval: humans rank and compare model outputs — the most direct signal for actual quality
- LLM-as-judge: a model scores and ranks other models' outputs — scales human evaluation to volumes no human team can cover
- Eval harnesses: automated regression suites that run on every model update
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.