The AI Learning Hub Journal

LLM Governance & Safety — When Each Layer Applies

"Is this AI safe?" is three questions — each answered at a different timebuilt before the model ships, applied on every live request, and watched at scale — each phase catches different failuresmodel shipshandles real trafficTRAINING TIME — BEFORE THE MODEL SHIPSone-time — baked into the weightsinstruction tuning (SFT)Constitutional AI — self-critique vs rulesRLHF — learns from human rankingsred-teaming — adversarial probingblind to how the model is used laterDEPLOYMENT TIME — WRAPS THE LIVE MODELevery request — in and out of the modeloutput filters on responsesscope and policy limitsprompt injection defence on inputsPII detection and redactionfilters each request — cannot see trendsPRODUCTION TIME — LIVE, AT SCALEcontinuous — visibility over real trafficprompt and response tracing — audit trailcost and token tracking per requestoutput drift and degradation flagslatency, throughput, SLA monitoringcatches what the other two layers missedCONTINUOUS EVALS — THE FEEDBACK LOOP THAT RUNS ACROSS ALL THREE PHASESbenchmarks across versionshuman preference rankingLLM-as-judge — scaled scoringregression harness on every updateNO SINGLE LAYER IS ENOUGH — EACH PHASE CATCHES WHAT THE OTHERS CANNOTa missing phase is a whole category of failures nobody catches — ask which phase a customer's concern actually lives in"SAFE" IS A PROCESS ACROSS THREE PHASES — NOT A PROPERTY OF THE MODELbaked in at training, wrapped around at deployment, watched in production — and evals proving it still holds
Safety is built in three phases — baked in at training, wrapped around at deployment, watched in production — with continuous evals running across all of them.

Safety Is Not One Thing

When customers ask "is this AI safe?" they are actually asking several different questions at once. Safety in LLMs is built across three separate phases — before the model ships, while it runs in production, and continuously as it evolves. Each phase catches different failure modes. No single layer is enough on its own.

  • Training time: safety and alignment baked into the model before it is ever deployed
  • Deployment time: guardrails applied at runtime to constrain what the live model can do
  • Production time: observability to know what the model is actually doing at scale
  • Continuous evals: validation that runs across all phases to catch regressions and drift

Training Time — Safety Built In Before the Model Ships

The first line of defence happens during training itself. Before a model is released, the team running training shapes its behaviour and probes it for failure modes. By the time it reaches customers, these properties are baked into the weights.

  • Instruction tuning (SFT): trains the model to follow instructions and be helpful
  • Constitutional AI (CAI): a principles-based self-critique loop where the model evaluates its own outputs against a set of rules during training
  • RLHF: shapes behaviour through human preference feedback — humans rank model outputs and the model learns to prefer higher-rated responses
  • Red-teaming: adversarial probing by humans trying to find failure modes before release

Deployment Time — Guardrails on the Live Model

Once a model is deployed, a second layer of controls wraps around it at runtime. These are not inside the model — they sit between the user and the model, filtering what goes in and what comes out on every single request.

  • Output filters: block unsafe, harmful, or off-policy content at runtime — the last check before a response reaches the user
  • Scope and policy limits: define what the model is and is not allowed to do in this specific deployment
  • Prompt injection defence: reduces the odds of adversarial inputs steering model behaviour — it limits risk rather than creating a hard boundary, so pair it with least privilege
  • PII and data privacy: detects and redacts sensitive personal data in both inputs and outputs

Production Time — Knowing What the Model Is Doing at Scale

Deploying safely is not a one-time event. Once a model is live and handling real traffic, you need visibility into every interaction. Production observability catches the problems that training and deployment guardrails missed.

  • Prompt and response tracing: a full audit trail of every interaction — essential for incident investigation and compliance
  • Cost and token tracking: per-request spend, token usage, and budget alerts
  • Output drift detection: flags quality degradation and behaviour shifts over time
  • Latency and performance: response time, throughput, and SLA monitoring

Continuous Evals — Validation That Never Stops

Evals are automated tests that run across all three phases. They are the feedback loop that tells you whether the model is still doing what you expect — and whether any change made things better or worse. In a production system, evals run on every update.

  • Benchmarks: standardised capability tests used to compare models and track capability over versions
  • Human preference eval: humans rank and compare model outputs — the most direct signal for actual quality
  • LLM-as-judge: a model scores and ranks other models' outputs — scales human evaluation to volumes no human team can cover
  • Eval harnesses: automated regression suites that run on every model update

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.