The AI Learning Hub Journal

AI Safety & Alignment

AI Safety & Alignment: Three Properties, Two Layers Diagram showing three properties of safe AI — Alignment, Robustness, Interpretability — and the two layers where safety is implemented: training time and runtime. AI Safety & Alignment: Three Properties, Two Layers Three properties of a safe AI system Alignment Does the AI do what we intend? Pursues intended goals — not literal instructions or proxy metrics Failure mode: reward hacking Goodhart's Law in action Robustness Holds up under unexpected input? Behaves correctly under adversarial or out-of-distribution inputs Failure mode: jailbreak & prompt injection Tested via red-teaming Interpretability Can you understand why? If you can't see why it decided, you can't catch when it's wrong Required for audit & oversight Chain-of-thought · confidence scores Two layers where safety is built in Training-time safety Properties baked into the weights By the time the model ships, these behaviours are part of how the model thinks SFT — instruction tuning on curated examples RLHF / DPO — shaped by human preferences Constitutional AI — self-critique against principles Red-teaming — adversarial probing before release Runtime safety Guardrails on the live model Sits between the user and the model, filtering every request and every response Output filters — catch harmful or off-scope content Scope & policy limits — define what's allowed Prompt-injection defence — block adversarial inputs PII detection — redact sensitive data in transit Neither layer is sufficient alone — frontier alignment is harder, not easier, as capability grows

Keeping AI Aligned with Human Intent

AI safety is the field concerned with ensuring AI systems do what we actually intend — not just what we literally asked for. Alignment is harder than it sounds: a model trained to maximise a reward signal will find unexpected ways to do so if the signal is imperfect. This is not a hypothetical problem — it manifests in every deployed model as hallucination, reward hacking, and refusal failures.

  • Alignment: does the model pursue the goals we intended, or does it optimise for proxies that diverge from real intent?
  • Robustness: does it behave correctly under adversarial, unusual, or out-of-distribution inputs?
  • Interpretability: can we understand why it made a decision — or is it a black box?
  • RLHF (Reinforcement Learning from Human Feedback) and Constitutional AI are the leading alignment techniques used in frontier models today
  • Guardrails, red-teaming, and output filtering are the operational safety layer applied at deployment time

Safety Techniques: From Training to Runtime

AI safety is implemented at multiple stages of the model lifecycle. Training-time techniques shape model behaviour at the source; runtime techniques catch failures at deployment. Both layers are necessary — neither alone is sufficient.

  • RLHF: human raters rank model outputs; a reward model learns their preferences; the policy model is optimised toward high-reward outputs
  • Constitutional AI (Anthropic): the model critiques its own outputs against a set of principles, iteratively improving without requiring human rating of every example
  • DPO (Direct Preference Optimisation): simpler RLHF alternative, widely adopted because it avoids training a separate reward model
  • Red-teaming: structured attempts to make the model produce harmful, biased, or off-scope outputs — essential before and after deployment
  • Output filtering and guardrails: runtime layers that intercept and block harmful outputs regardless of what the model produces

The Frontier Safety Challenge

As AI models become more capable, alignment becomes harder — not easier. More capable models can find more creative ways to satisfy a reward signal without satisfying the intent behind it. This is the core concern of frontier AI safety research, and it matters practically for anyone deploying AI in consequential settings.

  • Reward hacking: the model finds unexpected ways to maximise the reward signal without doing what you wanted
  • Goodhart's Law: when a measure becomes a target, it ceases to be a good measure — AI amplifies this effect
  • Scalable oversight: using AI to help humans supervise more capable AI — the approach frontier labs are researching
  • Interpretability research: understanding the circuits inside the model that produce specific behaviours — mechanistic interpretability is an active frontier
  • Practical implication: the more autonomy you grant an AI system, the more important robust alignment and oversight become

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.