AI Safety & Alignment
Keeping AI Aligned with Human Intent
AI safety is the field concerned with ensuring AI systems do what we actually intend — not just what we literally asked for. Alignment is harder than it sounds: a model trained to maximise a reward signal will find unexpected ways to do so if the signal is imperfect. This is not a hypothetical problem — it manifests in every deployed model as hallucination, reward hacking, and refusal failures.
- Alignment: does the model pursue the goals we intended, or does it optimise for proxies that diverge from real intent?
- Robustness: does it behave correctly under adversarial, unusual, or out-of-distribution inputs?
- Interpretability: can we understand why it made a decision — or is it a black box?
- RLHF (Reinforcement Learning from Human Feedback) and Constitutional AI are the leading alignment techniques used in frontier models today
- Guardrails, red-teaming, and output filtering are the operational safety layer applied at deployment time
Safety Techniques: From Training to Runtime
AI safety is implemented at multiple stages of the model lifecycle. Training-time techniques shape model behaviour at the source; runtime techniques catch failures at deployment. Both layers are necessary — neither alone is sufficient.
- RLHF: human raters rank model outputs; a reward model learns their preferences; the policy model is optimised toward high-reward outputs
- Constitutional AI (Anthropic): the model critiques its own outputs against a set of principles, iteratively improving without requiring human rating of every example
- DPO (Direct Preference Optimisation): simpler RLHF alternative, widely adopted because it avoids training a separate reward model
- Red-teaming: structured attempts to make the model produce harmful, biased, or off-scope outputs — essential before and after deployment
- Output filtering and guardrails: runtime layers that intercept and block harmful outputs regardless of what the model produces
The Frontier Safety Challenge
As AI models become more capable, alignment becomes harder — not easier. More capable models can find more creative ways to satisfy a reward signal without satisfying the intent behind it. This is the core concern of frontier AI safety research, and it matters practically for anyone deploying AI in consequential settings.
- Reward hacking: the model finds unexpected ways to maximise the reward signal without doing what you wanted
- Goodhart's Law: when a measure becomes a target, it ceases to be a good measure — AI amplifies this effect
- Scalable oversight: using AI to help humans supervise more capable AI — the approach frontier labs are researching
- Interpretability research: understanding the circuits inside the model that produce specific behaviours — mechanistic interpretability is an active frontier
- Practical implication: the more autonomy you grant an AI system, the more important robust alignment and oversight become
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.