The AI Learning Hub Journal

Presenting a Run So a Human Can Judge It

Show the evidence beside the story, never the story alonethe agent's explanation is generated text about a process it has no privileged access to — fluent, and possibly a reconstructionTHE REVIEW SURFACETHE NARRATIVE — AN AIDreads well; may be accurateor a plausible reconstruction— the text alone cannottell you whichpersuasive out of proportionto its evidential valueTHE ARTEFACTS — THE BASIS FOR THE DECISIONTHE DIFFwhat will actually change,shown as a deltaEXACT PARAMETERSof the pending action, nota summary of themCITED SOURCESthe passages relied on,with paths back to themVERIFICATION RESULTSwhat was checked andwhat it returnedORIGIN, TAGGED —from your own recordsfrom an inbound email — EXTERNALfrom stored memory — datedan action derived from an inbound email is a different proposition from one derived from your recordsWHERE THE SUMMARY AND THE TRACE DISAGREE, THE TRACE IS THE FACTFOR LARGE CHANGES, SHOW THE SHAPE BEFORE THE DETAIL — OUTLIERS FIRSThow many records, of what type, in which systems — with the outliers surfaced firstpeople spot the thing that does not belong far better than they confirm a hundred that doTHE EXPLANATION IS SHOWN BESIDE THE RECORD, NEVER IN PLACE OF ITa fluent rationale convinces through prose quality — give the reviewer things they can check against their own knowledge
A reviewer judges diffs, parameters, sources and verification results — the generated explanation sits beside that record and never replaces it

The Explanation Is Model Output

When an agent explains why it did something, that explanation is generated text describing a process the model does not have privileged access to. It is often accurate, it is sometimes a plausible reconstruction, and there is no reliable way to tell which from the text alone. This matters because a fluent rationale is persuasive out of proportion to its evidential value, and a reviewer reading a well-written justification is being convinced by prose quality rather than by evidence. The design response is not to hide the explanation, which is genuinely useful, but to place it alongside the record of what actually happened and never in place of it. Where the two disagree — the summary says the record was checked and no read appears in the trace — the trace is the fact. Show the summary as an aid to comprehension, and the artefacts as the basis for the decision.

  • A generated rationale can be an accurate account or a plausible reconstruction
  • Fluency is persuasive out of proportion to evidential value
  • Place the explanation beside the record of what happened, never instead of it
  • Where narrative and trace disagree, the trace is the fact

Show Artefacts and Deltas, Not Narrative

What a reviewer can actually judge is concrete: the diff of what will change, the exact parameters of the pending action, the specific source passages the conclusion rests on with a path back to them, and the verification results with what was checked and what it returned. These are things a person can check quickly against their own knowledge, which is the only kind of review that survives at any volume. Narrative summaries invite a different mode — reading for plausibility — which is exactly the mode that misses the wrong account number sitting inside a paragraph that reads perfectly. Where the change is large, show its shape before its detail: how many records, of what type, in which systems, with the outliers surfaced first. Reviewers are far better at spotting a thing that does not belong than at confirming that a hundred things all do.

  • Diffs, exact parameters, cited sources with paths back, verification results
  • Narrative invites reading for plausibility, which is how a wrong value gets through
  • For large changes show shape before detail, and surface outliers first
  • People spot the thing that does not belong far better than they confirm a hundred that do

Make Provenance Visible

The most useful single addition to a review surface is where this came from. A proposed action derived from a record in your own system and one derived from a sentence in an inbound email are different propositions, and the reviewer cannot distinguish them unless the interface does. Tag the material that drove the decision with its origin, and mark external or untrusted origins distinctly. This overlaps with the security treatment of the same problem and is worth doing on comprehension grounds alone: reviewers make better decisions when they can see what the agent was reading. The same applies to memory — an action taken because of a stored preference should show the preference and when it was recorded, since stale memory produces confident actions that make sense only against a fact that stopped being true months ago.

  • Show where the driving material came from and mark external origins distinctly
  • An action derived from an inbound email is a different proposition from one derived from your records
  • Surface which stored memories influenced the decision, and when they were written
  • Provenance improves ordinary review quality, not only adversarial detection

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.