The AI Learning Hub Journal

Review at Volume

Review at volume — the claim, the check, and what the hours are forreading work that consumed juniors for weeks now compresses into hours — volume is where the tools earn their keepcontract populationsconfirmation responsesboard minutescorrespondenceweeks → hoursthe tool does the readingand with the speed arrives the most seductive sentence in the brochureTHE FALSE COMFORT OF COVERAGE“we reviewed one hundred per cent of the population”reviewed to what standard?skimmed at a shallow standard,the population was filtered, not revieweda shallow test of everything can be weaker evidence than a deep test of a sampleand a flagged anomaly is a question, not a finding, until a person has followed it back to the source documentsTHE TOOL’S READING IS SOMEONE ELSE’S WORK — REVIEW IT THE WAY YOU WOULD A JUNIOR’Ssample what it summarised —read the documents yourself,in full, against its outputsample what it flagged —count the flags that dissolveon inspectionsample what it did not flag —the silent misses are wherethe audit risk actually livestrack error types — wrong party,wrong date, missed clause,invented confidencelet the observed error rate set how much you lean on the tool next timethe rigorous method behind this — evaluation sets, regression testing, quality over time — is the Does Your AI Actually Work? courseTHE UPSIDE — THE HOURS WERE READING HOURS, AND THEY COME BACK FOR JUDGEMENTfollowing the anomaliesthat deserve followingsitting longer with theestimate that does notsmell righthaving the difficultconversation with management,properly preparedfee pressure tempts firms to bank the hours instead — the judgement work is why they were freedJUDGEMENT WAS NEVER THE PART TO AUTOMATE — IT WAS THE PART THERE WAS NEVER TIME FOR
Coverage of the haystack is not assurance about any particular straw — sample the tool’s work the way you would a junior’s.

The False Comfort of Coverage

Volume is where the tools earn their keep: contract populations, confirmation responses, board minutes, correspondence — reading work that consumed juniors for weeks now compresses into hours. It is also where the most seductive sentence in the brochure lives: 'we reviewed one hundred per cent of the population.' Read as an auditor, the claim dissolves under one question — reviewed to what standard? A tool that skims everything at a shallow standard has not reviewed the population; it has filtered it, and coverage of the haystack is not assurance about any particular straw. The canon caveat from module one applies with full force here: a shallow test of everything can be weaker evidence than a deep test of a sample, and a flagged anomaly is a question, not a finding, until a person has followed it back to the source documents.

  • Volume reading — contracts, confirmations, minutes, correspondence — compresses from weeks to hours
  • 'One hundred per cent reviewed' dissolves under one question: reviewed to what standard?
  • Shallow coverage of everything can be weaker evidence than a deep sample
  • A flagged anomaly is a question, not a finding, until followed to source

Testing the Machine's Reading

The tool's reading is work performed by someone else, and the profession already knows what to do with that: review it. Take a sample of documents the tool summarised and read them yourself, in full, against its output. Take a sample of what it flagged and see how many flags dissolve on inspection; take a sample of what it did not flag and see what it missed, because the silent misses are where the audit risk actually lives. Track the kinds of error you find — wrong party, wrong date, missed clause, invented confidence — the way you would track a junior's, and let the observed error rate set how much you lean on the tool next time. The rigorous method behind this — evaluation sets, regression testing, measuring quality over time — is the subject of this site's Does Your AI Actually Work? course.

  • The tool's reading is someone else's work: review it the way you would a junior's
  • Sample what it summarised, what it flagged, and — above all — what it did not flag
  • Track error types and let the observed rate set how far you lean next time
  • Silent misses, not noisy false flags, are where the audit risk actually lives

The Hours That Judgement Needs

This module has been a list of disciplines, so it should end on what the disciplines buy. The hours that volume review used to consume were mostly not judgement hours; they were reading hours — necessary, honest, and spent before the interesting questions could even be asked. When a tool does the reading and a person verifies the tool, those hours come back, and the honest version of this technology's promise is what they come back for: following the anomalies that deserve following, sitting longer with the estimate that does not smell right, having the difficult conversation with management properly prepared. None of that appears in a productivity metric, which is why firms under fee pressure are tempted to bank the hours instead. The judgement was never the part anyone wanted to automate. It was the part there was never enough time for.

  • The hours volume review consumed were mostly reading hours, not judgement hours
  • They come back for following anomalies, sitting with estimates, preparing hard conversations
  • Fee pressure tempts firms to bank the hours; the judgement work is why they were freed
  • Judgement was never the part to automate — it was the part there was never time for

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.