The AI Learning Hub Journal

Transparency & Explainability

What we can and cannot see inside a modelOPAQUEOBSERVABLELAYERWHAT YOU CAN HONESTLY CLAIMRaw weights and activationsbillions of numbers, no direct human readingNothing you can state in words.This is the opaque floor.Feature attribution and probeswhich inputs and features moved the outputDirectional evidence about influence,not a complete account of the cause.Chain-of-thoughtthe reasoning the model states out loudA partial window only — the text canbe plausible and still be unfaithful.Model cards and documentationtraining scope, intended use, known limitsProvenance and scope, as declaredby whoever built the system.Evals and behavioural testingmeasured behaviour on held-out casesThe strongest evidence available —and it is about behaviour only.THE HONEST BOUNDARYBehaviour can be measured. Internal reasoning can only be partially inspected — never assume the stated reason is the real one.Every layer above the floor is an approximation — useful, incomplete, and worth saying so out loud
Opacity decreases as you move outward — but the outer layers measure behaviour, never the reasoning itself

Why Deep Models Are Opaque

A deep neural network is billions of numerical weights, tuned by an optimisation process no human directed step-by-step. There is no line of code that says "reject this loan application" — there is a cascade of matrix multiplications that produces a score. This is fundamentally different from traditional software, where a developer can trace any decision back to an explicit rule. The opacity is not a bug or laziness; it is a structural property of how these systems learn. And it collides directly with a basic expectation of consequential decisions: that someone can explain why. When a model denies credit, flags a transaction, or screens out a resume, "the weights produced that output" satisfies neither the affected person, the regulator, nor often the organisation that deployed it.

  • A trained model is billions of learned weights, not a set of human-readable rules — there is no "if" statement to point to
  • The same property that makes deep learning powerful — learning representations humans didn't specify — is what makes it opaque
  • Opacity matters most where decisions are consequential: credit, hiring, healthcare, justice, security
  • Even the model's builders cannot fully explain individual outputs — this is a research frontier, not an engineering oversight

What Explainability Actually Offers Today

Explainability research has produced real but partial tools. Feature attribution methods (such as SHAP and LIME) estimate which inputs most influenced a specific prediction — useful for spotting when a model leans on a suspicious signal, but an approximation, not a ground-truth account. Model cards and system cards document what a model was trained on, evaluated against, and intended for. Chain-of-thought output from reasoning models offers a readable window into a model's working — genuinely useful for review, but with a known limit: the written reasoning does not always faithfully reflect the computation that produced the answer, so it should be treated as evidence, not proof. Mechanistic interpretability — reverse-engineering the circuits inside models — is advancing but remains far from explaining frontier models end to end.

  • Feature attribution (SHAP, LIME): estimates which inputs drove a prediction — an approximation, useful for catching suspicious signals
  • Model cards and system cards: document training data scope, evaluation results, intended use, and known limitations
  • Chain-of-thought: a readable partial window — but models can produce reasoning text that does not faithfully match their actual computation
  • Mechanistic interpretability: promising research, not yet a practical tool for explaining production frontier models

What You Can Demand — and What Nobody Can Provide Yet

The practical skill is knowing the difference between reasonable and impossible transparency demands. Reasonable: documentation (model cards, intended-use statements, known limitations), evaluation results relevant to your use case (including disaggregated performance across subgroups), data provenance answers (what categories of data trained this system, and is my data used for training?), audit trails of decisions, and a clear account of where humans review outputs. Not yet possible from any vendor, however sincere: a complete causal explanation of why a large model produced one specific output, or a guarantee that a model will never behave unexpectedly. A vendor who promises full explainability of a deep model is overclaiming — while a vendor who cannot produce even a model card or evaluation results has not done the basic work. Calibrated buyers ask for everything in the first list and are suspicious of anyone promising the second.

  • Reasonable to demand: model cards, evaluation results on your use case, data provenance, audit trails, human oversight points
  • Not available from anyone yet: complete causal explanations of individual frontier-model outputs
  • Red flag in both directions: vendors promising "fully explainable AI" and vendors who cannot produce basic documentation
  • For high-stakes narrow decisions, a simpler interpretable model is sometimes the right call — explainability can be a reason to choose less capable technology

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.