Metrics That Survive Contact
The Four Metrics Worth Tracking
Four measurements capture most of what AI-assisted triage changes. Time-to-triage: how long from alert arrival to a dispositioned decision, tracked as a distribution per alert category. False-positive burden: how many analyst-minutes per day go to alerts that end in benign dispositions. Escalation quality: of the cases sent up a tier, what fraction the receiving tier judges worth the escalation — measured by asking them, not by assuming. Analyst hours returned: time genuinely freed for hunting, detection engineering, or tuning, verified by where the hours actually went. Each is measurable, each ties to something the SOC exists to do, and each — this is the catch — can be gamed. The next slides are about the gaming.
- Time-to-triage as a distribution per category, not one blended average
- False-positive burden in analyst-minutes, the unit fatigue is denominated in
- Escalation quality judged by the receiving tier, not the sending one
- Hours returned counts only if you can show where the hours went
Goodhart's Law, SOC Edition
When a measure becomes a target, it stops being a good measure — and each of the four fails in its own way once it becomes the number leadership watches. Optimize time-to-triage and the system learns that closing fast scores well, so dispositions get quicker and shallower; speed without a paired quality floor measures haste. Optimize false-positive burden and the cheapest path is dismissing more aggressively, converting analyst fatigue into silent misses. Optimize escalation quality and the AI escalates only sure things, and the interesting ambiguous cases — the ones tier three most needs to see — die in the queue. Optimize hours returned and the hours evaporate into unmeasured slack. Every metric needs a counter-metric watching its failure direction.
- Fast triage without a quality check rewards shallow dispositions
- FP burden drops fastest by dismissing more — including things that were real
- An AI graded on escalation precision learns to hide the ambiguous cases
- Pair every headline metric with the counter-metric that catches its abuse
The Seduction of Alert-Volume Reduction
"Reduced alert volume by 70%" is the most quoted number in this product category and, alone, close to meaningless. Volume reduction has a legitimate form — deduplication, correlation of related alerts into cases, suppression of known-benign patterns with documented rules. It also has an illegitimate form: raising thresholds and dismissing harder, which produces an identical headline number. The metric cannot distinguish a smarter funnel from a narrower one. Any reduction claim should arrive with its method attached — what exactly was suppressed, deduplicated, or dismissed, and under what logic — and with evidence about what the discarded volume contained. Which requires actually looking at the discarded volume.
- Deduplication, correlation, and documented suppression are real reduction
- Raised thresholds and aggressive dismissal produce the same headline number
- Demand the method behind any reduction percentage, not just the percentage
- A reduction claim without analysis of the discarded alerts is unverified
Measuring the Misses: Sample What It Dismissed
The alerts the AI escalates get human review by definition; the alerts it dismisses are where the unmeasured risk pools. The only honest instrument is sampling: pull a random slice of AI-dismissed alerts on a schedule — daily during the pilot, weekly in steady state — and have an analyst work them cold, without seeing the AI's reasoning first, so the review is independent rather than an exercise in agreeing with a confident summary. Track the disagreement rate and treat movement in it as an alarm, because it is also your early-warning system for drift: telemetry changes, new attacker tradecraft, a model update. This review time is a permanent cost of running the system. Budget it, or the miss rate becomes a number nobody knows.
- Random-sample dismissed alerts on a fixed schedule, forever, not just during the pilot
- Reviewers work the sample blind to the AI's disposition and reasoning
- Rising disagreement is a drift alarm before it is a miss statistic
- Sampling cost is part of the system's true operating price — budget it up front
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.