The AI Learning Hub Journal

Testing in Production: A/B and Guardrails

Offline blocks harm — online decides whether you were rightyour case mix is not the traffic mix and your rubric is not the user — so test live, behind guardrails agreed in advanceWHY OFFLINE AND ONLINE DISAGREE — STRUCTURALLY, NOT ACCIDENTALLYyour case mix is notthe live traffic mixrubric quality is not whatusers actually preferproduction carries contextthe harness never seescaching, retries, fallbackssit outside the eval pathA FAIR COMPARISON ON LIVE TRAFFIClive trafficrandomise byuser or sessionA — current versionB — the candidatefix the primary metric and the smallest effect worthacting on before launch; segment the results after —an aggregate win can hide a segment that got worseno peeking — stopping on a favourable read manufactures winsGUARDRAILS — AGREED STOPSname what must not degradeand set thresholds first:error rate · tail latency ·cost per success · refusal ·escalation · safety checksthe stop is mechanical — in themoment there is always a storyROLL OUT PROGRESSIVELY — GUARDRAILS EVALUATED AT EVERY STAGE5%25%50%100%rollback tested before you need it —an unactionable guardrail is a dashboardTWO CURVES THAT DISTORT EARLY READINGSnovelty — inflates the first days, then decayslearning — depresses a real improvement until users adjustrun long enough to see the curve, not the pointsplit returning from first-time users — only one feels noveltyFIX THE METRIC AND THE STOPS BEFORE LAUNCH — THEN LET THEM DECIDEin the moment there is always a plausible story for a bad number
Test live with A/B splits randomised by user, behind guardrail metrics whose stopping thresholds were agreed before launch — and read the curve, not the first days.

Why Offline Ever Disagrees With Online

An offline suite measures a fixed set of inputs against your definition of good. Production measures whatever users actually send against whether they got what they came for, and the two diverge for structural reasons rather than accidental ones. Your case mix is not the traffic mix, and it under-represents the messy, truncated, ambiguous inputs that dominate real usage. Your graders encode a definition of quality that users may not share, since people frequently prefer shorter, faster and blunter than a rubric rewards. Production carries context your harness lacks: the prior conversation, the state of the account, whatever the user was doing before they asked. And the deployed system includes caching, truncation, retries and fallbacks that the eval path often bypasses entirely. Offline evals exist to iterate quickly and block obvious harm; online measurement exists to find out whether you were right.

  • The case mix is not the traffic mix, and real inputs are messier than yours
  • Rubric quality and user-perceived quality are related but not the same thing
  • Production carries context and infrastructure the eval harness bypasses
  • Offline blocks harm and speeds iteration; online decides whether it worked

Running a Fair Comparison on Live Traffic

A live comparison is only evidence if it is built to be one. Randomise at the unit that matches the effect you are measuring — usually the user or the session rather than the individual request, because a user bounced between two versions mid-conversation gets an incoherent experience and gives you a contaminated measurement. Decide the primary metric and the smallest effect worth acting on before you start, then work out how long that takes to detect at your traffic volume; a test that cannot reach a conclusion is worse than no test, because it will be read anyway. Segment the results the same way you slice offline evals, since an aggregate win driven entirely by one segment while another degrades is a decision rather than a result. And resist stopping the moment the number looks favourable, because peeking until significance appears manufactures wins reliably.

  • Randomise at user or session level so the experience stays coherent
  • Fix the primary metric and the minimum meaningful effect before launching
  • Segment results — an aggregate win can hide a segment that got worse
  • Do not stop early on a favourable reading; repeated peeking manufactures wins

Guardrail Metrics and How a Rollout Stops

Alongside the metric you hope will improve, define the ones that must not get worse and give them thresholds that halt the rollout. The usual set is error rate, latency at the tail, cost per successful task, refusal rate, escalation to a human, and any safety check you run over live traffic — plus, for anything with a clinical, financial or legal dimension, the specific harm whose first confirmed instance ends the experiment outright. Guardrails only work when the stopping decision is agreed in advance and enforced mechanically, because in the moment there is always a plausible explanation for a bad number and a strong incentive to accept it. Roll out progressively behind them, evaluating the guardrails at each stage rather than only at the end, and test the rollback path before you need it. A guardrail nobody can act on within minutes is a dashboard.

  • Name the metrics that must not degrade and set their thresholds before launch
  • Error rate, tail latency, cost per success, refusal, escalation, and safety checks
  • Automate the stop — in the moment there is always a story for a bad number
  • Progressive rollout with a tested rollback, or the guardrail is decorative

Novelty, Learning, and the Win That Does Not Last

Two time-dependent effects distort early online readings in opposite directions. Novelty: a visibly different behaviour or interface draws engagement simply because it is new, and the lift decays over days or weeks, so a change measured only in its first days is partly measuring curiosity. Learning: users who had adapted to the previous behaviour perform worse at first with a better one, so a genuine improvement can read as a regression until people adjust. Both argue for running long enough to see the curve rather than the point, and for separating returning users from first-time ones, since only one of those groups can experience novelty at all. This is the most common way an offline win loses online: the eval measured a static task, the users were adapting, and the number moved for reasons the harness had no way to represent.

  • Novelty inflates early readings and learning effects depress them; both decay
  • Run long enough to see a curve, not a single point
  • Separate returning users from new ones — only one group can feel novelty
  • An offline win can lose online because the harness cannot represent adaptation

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.