The AI Learning Hub Journal

The Questions That Cut Through

Four question families that cut through the marketingpress until you get architecture, numbers, and specifics — evasion is data too1 · THE MODEL, NOT THE MARKETINGWhich model is under the hood — and who builds and trains it?Where does inference actually run?Frozen or continuously retrained — on what cadence?Classical ML with a chat interface? Fine — price it as suchadjectives instead of architecture are hiding something structural2 · FOLLOW YOUR DATA — CUSTODY BEFORE CAPABILITYDoes anything we send you train your models — contractually?Genuine tenant isolation, or logical separation?What is retained, for how long — and can we set it?Where does processing happen — can we pin the region?certifications describe control audits, not model-pipeline flows3 · ASK ABOUT FAILURE BEFORE FEATURESWhen the model is uncertain, does it say so?What does a false negative look like operationally?"Describe your last significant false negative"How are regressions detected — and disclosed?failure specifics signal maturity; failure denial is a warning4 · THE VENDOR'S OWN SOC IS THE TELL"What does your own SOC use this for, exactly?"What does your own team not trust it to do?What is the internal analyst override rate?A reference call with a practitioner, not a championthat texture is impossible to fake — generalities are the flagGET THE COMMITMENTS INTO THE CONTRACT, NOT THE SLIDE DECK — POLICIES CHANGE UNILATERALLYTHE ANSWERS DETERMINE DATA EXPOSURE, DETECTION FRESHNESS, AND DRIFT AFTER YOU SIGNa vendor who cannot say which model, whose data, and what failure looks like is hiding the structure
Open with the model, settle data custody, ask about failure, and make the vendor's own SOC testify — the vague answer is itself the finding.

Start With the Model, Not the Marketing

Every AI security product deserves the same opening question: what model is under the hood, and who built it? A vendor running a frontier model via API has different data-flow, latency, and cost characteristics than one running a fine-tuned open-weights model in their own infrastructure — and both differ from a classical ML pipeline wearing an "AI-powered" label. None of these is automatically better, but a vendor who cannot or will not tell you which one they are is hiding something structural. Ask whether the model is frozen or continuously retrained, and on what cadence. The answers determine everything downstream: your data exposure, your detection freshness, and how the product will drift after you sign.

  • Ask: which model, who trains it, and where does inference actually run?
  • A vague answer to "what model?" is itself diagnostic — press until you get architecture, not adjectives
  • Frozen models mean stale detections; get the retraining cadence in writing
  • Classical ML with a chat interface is fine — but it should be priced and evaluated as such

Follow Your Data

Before capability, settle custody. Four questions, in order: Does anything we send you train your models — and is the answer contractual, or a policy page that can change? Is our tenant genuinely isolated, or logically separated in shared infrastructure? What is retained, for how long, and can we set that retention? Where does processing happen, and can we pin it to a region? Good vendors answer all four crisply, because their security-mature customers have already forced them to. Evasive vendors reroute to certifications — but an audit badge tells you an assessor reviewed their controls, not what happens to your alert data inside a model pipeline. Get the data-handling commitments into the contract, not the slide deck.

  • Training use, isolation, retention, residency — settle all four before any capability discussion
  • Contractual language beats policy pages: policies change unilaterally, contracts do not
  • Certifications describe control audits, not model-pipeline data flows — ask the specific question anyway
  • Ask what leaves your tenant during enrichment: third-party lookups can quietly export your indicators

Ask About Failure Before Features

The feature list tells you what the product does when it works. Due diligence is about what it does when it fails — because in security, the failure modes are the risk. Ask: when the model is uncertain, does it say so, or does it produce a confident answer regardless? What does a false negative look like operationally — silence, or a low-confidence flag a human still sees? When the product makes a wrong call, how do you find out: does the vendor detect and disclose regressions, or do you discover them in an incident review? A vendor with real production maturity has watched their product fail and can describe it in specifics. A vendor who insists failure basically does not happen has either not measured it or is not telling you.

  • Uncertainty signalling is a product feature — its absence means confident wrong answers reach analysts
  • "Describe your last significant false negative" separates measured products from marketed ones
  • Ask how regressions are detected and disclosed — silence after a bad model update is the worst case
  • Failure specifics are a maturity signal; failure denial is a warning

The Vendor's Own SOC Is the Tell

One question cuts through more marketing than any RFP: does your own security team run this product in production, and what do they use it for? A vendor whose SOC genuinely lives on their own AI triage can tell you which alert classes it handles autonomously, where their analysts still override it, and what they turned off. That texture is impossible to fake. Follow with the inverse: what does your own team not trust it to do? An honest answer here — and there always is one — maps the product's real boundary far more accurately than the capability matrix. Vendors selling a product their own practitioners route around are selling you the same workaround future. Ask to speak to their SOC lead, not their sales engineer.

  • "What does your own SOC use this for?" — specifics are unfakeable, generalities are a red flag
  • The inverse question matters more: what does your own team not trust it to do?
  • Ask for the internal override rate: how often do their analysts overrule the product?
  • Request a reference call with a practitioner, not a champion curated by sales

Try It Yourself

These questions only work if you actually send them. Pick a product and send them this week — the answers, or the silence, will tell you more than another demo would.

◆ Try it yourself

Choose one AI security product your organisation already runs or is currently being pitched. Send the five questions below to the vendor, and write down each answer verbatim — including the ones you do not get.

1. Which model or models sit behind this feature, and who operates them?
2. Is our data used to train or improve any model? Where is that stated contractually?
3. What does the product do when it is not confident — block, flag, or stay silent?
4. Show us a case it got wrong, and what happened next.
5. What does your own security team do with this product internally?
How you'll know it worked
  • Every answer is written down in the vendor's words, not your paraphrase
  • You can name at least one question the vendor would not answer directly
  • Someone else on your team could read the answers and reach the same view you did

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.