Testing in Production: A/B and Guardrails
Why Offline Ever Disagrees With Online
An offline suite measures a fixed set of inputs against your definition of good. Production measures whatever users actually send against whether they got what they came for, and the two diverge for structural reasons rather than accidental ones. Your case mix is not the traffic mix, and it under-represents the messy, truncated, ambiguous inputs that dominate real usage. Your graders encode a definition of quality that users may not share, since people frequently prefer shorter, faster and blunter than a rubric rewards. Production carries context your harness lacks: the prior conversation, the state of the account, whatever the user was doing before they asked. And the deployed system includes caching, truncation, retries and fallbacks that the eval path often bypasses entirely. Offline evals exist to iterate quickly and block obvious harm; online measurement exists to find out whether you were right.
- The case mix is not the traffic mix, and real inputs are messier than yours
- Rubric quality and user-perceived quality are related but not the same thing
- Production carries context and infrastructure the eval harness bypasses
- Offline blocks harm and speeds iteration; online decides whether it worked
Running a Fair Comparison on Live Traffic
A live comparison is only evidence if it is built to be one. Randomise at the unit that matches the effect you are measuring — usually the user or the session rather than the individual request, because a user bounced between two versions mid-conversation gets an incoherent experience and gives you a contaminated measurement. Decide the primary metric and the smallest effect worth acting on before you start, then work out how long that takes to detect at your traffic volume; a test that cannot reach a conclusion is worse than no test, because it will be read anyway. Segment the results the same way you slice offline evals, since an aggregate win driven entirely by one segment while another degrades is a decision rather than a result. And resist stopping the moment the number looks favourable, because peeking until significance appears manufactures wins reliably.
- Randomise at user or session level so the experience stays coherent
- Fix the primary metric and the minimum meaningful effect before launching
- Segment results — an aggregate win can hide a segment that got worse
- Do not stop early on a favourable reading; repeated peeking manufactures wins
Guardrail Metrics and How a Rollout Stops
Alongside the metric you hope will improve, define the ones that must not get worse and give them thresholds that halt the rollout. The usual set is error rate, latency at the tail, cost per successful task, refusal rate, escalation to a human, and any safety check you run over live traffic — plus, for anything with a clinical, financial or legal dimension, the specific harm whose first confirmed instance ends the experiment outright. Guardrails only work when the stopping decision is agreed in advance and enforced mechanically, because in the moment there is always a plausible explanation for a bad number and a strong incentive to accept it. Roll out progressively behind them, evaluating the guardrails at each stage rather than only at the end, and test the rollback path before you need it. A guardrail nobody can act on within minutes is a dashboard.
- Name the metrics that must not degrade and set their thresholds before launch
- Error rate, tail latency, cost per success, refusal, escalation, and safety checks
- Automate the stop — in the moment there is always a story for a bad number
- Progressive rollout with a tested rollback, or the guardrail is decorative
Novelty, Learning, and the Win That Does Not Last
Two time-dependent effects distort early online readings in opposite directions. Novelty: a visibly different behaviour or interface draws engagement simply because it is new, and the lift decays over days or weeks, so a change measured only in its first days is partly measuring curiosity. Learning: users who had adapted to the previous behaviour perform worse at first with a better one, so a genuine improvement can read as a regression until people adjust. Both argue for running long enough to see the curve rather than the point, and for separating returning users from first-time ones, since only one of those groups can experience novelty at all. This is the most common way an offline win loses online: the eval measured a static task, the users were adapting, and the number moved for reasons the harness had no way to represent.
- Novelty inflates early readings and learning effects depress them; both decay
- Run long enough to see a curve, not a single point
- Separate returning users from new ones — only one group can feel novelty
- An offline win can lose online because the harness cannot represent adaptation
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.