The AI Learning Hub Journal

A Pilot That Proves Something

A pilot that can fail is the only pilot that provescriteria written after exposure to results are indistinguishable from criteria fitted to themBASELINE FIRSTtwo to four weeks measuringthe team you have: medians,90th percentiles, volumes,escalation + reopen rates —before the AI touches anythingCRITERIA IN WRITINGnumeric targets, measurementmethod, and time windowfixed before day one —including a quality floor,not just a speed targetRUN IT HONESTLYlive alerts from your ownqueue, a window longenough to include normalchaos, and a control queueor team where feasibleDECIDE AGAINST THEMcriteria met — a decision;missed — a walk-away;kill criteria written inadvance are the onlyones that ever fire"the AI triages in 4 minutes" means nothing if nobody recorded that analysts averaged 6THREE FAILURE MODES THAT MAKE PILOTS MISLEADdemo datavendor-supplied or sanitized alertsinstead of your broken live queuecherry-picked periodsix quiet weeks in August provelittle about the rest of the yearno control grouptool plus reorg at once — nothingattributes the improvement"THE ANALYSTS LIKED IT" IS A PRECONDITION, NOT A RESULTnovelty and survey framing inflate early enthusiasm predictably — sentiment is an adoption signal onlythe pairing that matters: analysts like it AND the measured numbers movedDECIDE WHAT KILLS THE PILOT BEFORE IT RUNS — DECIDED AFTER, NOTHING EVER KILLS ITA PILOT THAT CANNOT FAIL IS A ROLLOUT WEARING A PILOT'S BADGEwithout a baseline, the result is a number with nothing to stand next to
Measure the baseline, write success and kill criteria before day one, and run on live data with a control — sentiment alone is not a result.

Write the Success Criteria Before the Pilot Starts

The defining property of a real pilot is that its success criteria exist in writing before the first alert flows through the system. "Reduce median phishing triage time from 18 minutes to under 8, measured across six weeks, without escalation quality dropping below current levels" is a criterion. "See how it performs" is not — it is a plan to decide afterwards whether whatever happened counts as success, and afterwards, with a vendor relationship in motion and internal sponsors invested, the answer is always yes. Criteria written after exposure to results are indistinguishable from criteria fitted to results. The pilot that cannot fail teaches you nothing, and everyone in the room quietly knows it.

  • Numeric targets, measurement method, and time window fixed before day one
  • Include a quality floor, not just a speed target — fast and wrong is worse than slow
  • Post-hoc criteria always get fitted to whatever the results turned out to be
  • If no conceivable result would count as failure, it is a rollout wearing a pilot's badge

Baseline First: Measure the Team You Have

A pilot compares two states, and most teams only measure one of them. Before deployment, spend two to four weeks instrumenting current operations: median and 90th-percentile time-to-triage per alert category, daily volumes, escalation rates, how often closed alerts are reopened. This is unglamorous and teams skip it because the AI is already licensed and everyone wants to see it run. But without the baseline, the pilot ends in a number with nothing to stand next to — "the AI triages in 4 minutes" means nothing if nobody recorded that analysts averaged 6. The baseline also surfaces problems AI will not fix, like alerts nobody was reviewing at all. Those belong in the honest accounting too.

  • Two to four weeks of baseline measurement before the AI touches anything
  • Capture distributions, not just averages — the 90th percentile is where pain lives
  • Record escalation and reopen rates as your quality reference points
  • Baselining often reveals process debt no AI purchase will address

Why "The Analysts Liked It" Is Not an Outcome

Analyst sentiment matters — a tool the team refuses to use fails regardless of its benchmarks — but it is a precondition, not a result. New tools benefit from novelty, from relief at attention finally paid to the tier-one grind, and from the basic human tendency to be agreeable in a feedback survey the vendor helped write. None of that tells you whether triage got faster or whether escalations got better. Sentiment is also the metric most easily harvested and least easily audited, which is exactly why weak pilots lean on it. Collect it, take it seriously as an adoption signal, and refuse to let it substitute for the operational numbers. A pilot that reports only satisfaction scores has reported nothing.

  • Sentiment is an adoption precondition, not evidence of operational improvement
  • Novelty and survey framing inflate early enthusiasm predictably
  • If the final readout leads with satisfaction scores, ask what is being buried
  • The pairing that matters: analysts like it AND the measured numbers moved

Failure Modes, and the Kill Criteria That Guard Against Them

Three failure modes account for most pilots that mislead. Demo data: the evaluation runs on vendor-supplied or sanitized alerts instead of your live queue, with its broken log sources and ambiguous internal traffic. Cherry-picked periods: six quiet weeks in August prove little about the year. No control group: if the whole team gets the tool while workflows are simultaneously reorganized, nothing attributes the improvement. Run on live data, across a representative window, with a comparison queue or team where feasible. Then write the kill criteria: the specific results — missed true positives above a threshold, triage quality below the floor — that end the project. Deciding what kills the pilot after the pilot means nothing kills the pilot.

  • Live alerts from your own queue, or the results describe someone else's SOC
  • Representative time window — long enough to include normal chaos
  • A control queue or team turns "things improved" into "the AI improved things"
  • Kill criteria written in advance are the only ones that ever fire

Try It Yourself

The hardest part of a pilot is agreeing, in advance, what result would make you walk away. Write that down now, while nothing is at stake.

◆ Try it yourself

For a tool you are piloting or about to pilot, fill in the scaffold below and get one other person to agree to it in writing before the pilot starts. Date it.

The pilot: [tool, scope, how long]

Baseline we measured FIRST: [the current number, measured before the tool arrives]

Success means: [a specific number or threshold, not a feeling]

We kill it if: [the result that ends the project — decided now, not later]

Who decides, and when: [name, date]

What we will NOT count as evidence: [demo data, a good week, positive impressions]
How you'll know it worked
  • The kill criterion is specific enough that it could actually be met
  • The baseline is a number you measured, not one you assumed
  • A second person has agreed to the whole thing in writing, before the start

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.