The AI Learning Hub Journal

Reading a Demo Skeptically

The demo is theirs — the evidence has to be yoursnothing in a polished demo predicts behaviour on your telemetry and your naming conventionsWHAT THE HIGHLIGHT REEL SHOWSgolden alerts, chosen because the product shines on themlog schemas the vendor built the parsers forentity names the model has seen a thousand timesa 90-second triage shown as an edited cutweeks of "environment preparation" forecasts the deploymentWHAT YOU ASK TO SEEone of your real, sanitised alerts — run livenatural-language query against your field namesan alert the product gets wrong, and the recoverydegraded modes: source outages, timeouts, conflictsa vendor comfortable demoing failure has measured itPIN DOWN WHAT "AGENTIC" MEANS HERE — MAKE THE WORD OPERATIONALwhich decisions are madewith no human, exactly?approval gates configurableper action class?can autonomy be shrunkafter deployment?one full investigationtrace, end to endno producible trace means autonomy that is theatre, or real and unauditable — both end the conversationTHE PILOT IS THE REAL DEMO — TIME-BOXED, IN YOUR ENVIRONMENTsuccess criteria writtenbefore it startsa defined finish lineand a walk-awayyour alert mix and volume,not a suggested subsetstaffed honestly — unusedpilots prove nothingbeware the perpetual proof-of-value: months of drift, your integration effort, unearned deal momentumBRING YOUR OWN DATA, ASK TO SEE A MISS, AND MAKE THE PILOT THE DECISION POINTthe unhappy path is where analysts live during real incidents — a demo that never goes there has shown you nothing
Demo data is curated to succeed — bring your own sanitised alert, ask to see a miss, make the vendor operationalise agentic, and let a time-boxed pilot decide.

Demo Data Is Not Your Data

Every polished demo runs on curated data: golden alerts chosen because the product handles them beautifully, log schemas the vendor built the parsers for, entity names the model has seen a thousand times. None of that predicts behaviour on your telemetry — your custom field names, your legacy log sources, your naming conventions no model has trained on. Natural-language query features are the classic case: flawless against the vendor's demo schema, brittle against yours. So bring your own material. Ask in advance to run one of your real, sanitised alerts through the product live. A vendor who welcomes that has confidence in the product; a vendor who needs weeks of "environment preparation" first is telling you what deployment will cost.

  • Golden alerts predict nothing — insist on one of your own sanitised alerts, live
  • Test natural-language query against your field names and log sources, not the demo schema
  • Note what was pre-loaded: parsers, enrichments, and integrations you would have to build
  • Reluctance to touch your data in a demo forecasts the deployment experience

Ask to See a Miss

A demo that only shows successes is a highlight reel. The most informative request you can make: show me an alert this product gets wrong. Watch the reaction. A mature vendor has examples ready — they know their false-positive patterns, they can show a wrong triage verdict and walk through how an analyst catches it, and they will show real response latency rather than an edited cut. An immature vendor treats the question as hostile. While you are there, probe the seams: ask what happens when a data source goes dark mid-investigation, when the model times out, when two components disagree. The unhappy path is where your analysts will live during a real incident; a demo that never goes there has not shown you the product.

  • "Show me a miss" is the single highest-signal demo request — note whether examples exist
  • Real latency matters: a 90-second triage shown as an edited cut hides the analyst experience
  • Probe degraded modes: source outages, timeouts, conflicting outputs
  • A vendor comfortable demoing failure has measured it; discomfort means they have not

Pin Down What "Agentic" Means Here

"Agentic" now appears on every datasheet, describing everything from a chat window to genuine multi-step investigation. Make the vendor operationalise it: which decisions does the system make without a human, exactly? What tools can it invoke, and what stops it invoking them wrongly? Where are the approval gates, and are they configurable per action type or all-or-nothing? Can you shrink its autonomy after deployment without ripping it out? Ask to see the audit trail for one full agent investigation — every retrieval, every tool call, every decision point. If the vendor cannot produce that trace, the honest reading is that either the autonomy is theatre, or it is real and unauditable. Both should end the conversation.

  • Ask for the concrete decision list: what does it decide alone, and what does it queue for approval?
  • Approval gates should be configurable per action class, not a single on/off switch
  • Autonomy you cannot shrink post-deployment is a one-way door — avoid it
  • One full investigation trace, end to end, is the proof: no trace, no trust

The Pilot Is the Real Demo

No demo, however honest, substitutes for the product running on your data, your log sources, and your alert volume. Insist on a time-boxed pilot in your environment with success criteria written down before it starts: which alert classes, what false-positive tolerance, what time-to-verdict target, measured how, and by whom. Agree in advance what happens at the end — criteria met means a decision, criteria missed means a walk-away, and the vendor knows both. Beware the perpetual proof-of-value that drifts for months without a finish line; it costs your team integration effort while creating deal momentum the product has not earned. And staff it honestly on your side: a pilot nobody uses proves nothing either way.

  • Success criteria written before the pilot starts, not negotiated after results arrive
  • Time-box it and define the walk-away — open-ended pilots are a sales tactic, not an evaluation
  • Measure on your alert mix and volume, not the subset the vendor suggests
  • Budget your own analyst time: an unused pilot generates no evidence

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.