Before the Firm Relies on a Tool
Four Things the Firm Should Know First
Before any engagement team leans on a tool, someone at the firm should be able to answer four questions about it. What was it tested on — whose documents, in which languages, at what quality of scan, and how closely that population resembles the engagements it will actually meet? How does it err — what does it miss, what does it invent, how often, and whether its errors cluster in particular document types rather than spreading evenly? How is it versioned — what changes between releases, on whose schedule, and whether the firm is told before behaviour shifts? And who answers for it — a named person inside the firm who owns the decision to approve it, not a vendor's support desk. None of these questions is technical beyond a practitioner's reach. They are the questions a firm would ask before relying on any expert whose work feeds audit evidence.
- Tested on what: whose documents, which languages, what scan quality, and how close to your engagements
- How it errs: what it misses, what it invents, how often, and where the errors cluster
- How it versions: what changes, when, and whether the firm hears about it before behaviour shifts
- Who answers: a named owner inside the firm, never a vendor's support desk
The Approved List Is a Control
From the engagement team's side, an approved-tool list reads as bureaucracy — a queue between them and something useful. From the firm's side it is a control, and it is worth being precise about what it controls. The list is the record that the validation questions were asked before reliance began, and the mechanism by which the answers reach every engagement rather than only the team that asked. It is also the boundary that keeps client data out of tools nobody assessed. Two failure modes are predictable. A list nobody maintains goes quietly stale while AI features arrive inside software the firm already licenses. And a prohibition with no permitted alternative produces unapproved use that nobody monitors — the worst of both positions, because the firm carries the risk without the visibility. Exception requests are the health signal: a list nobody ever asks to extend is being routed around.
- The list records that validation happened before reliance began, and carries the answers to every team
- It is the boundary that keeps client data out of tools the firm has never assessed
- An unmaintained list goes stale while AI features arrive inside software already licensed
- Prohibition without a permitted route produces unmonitored use; exception requests are the health signal
Validation Is Its Own Discipline
What validation actually involves is a discipline of its own, and its shape is familiar to any auditor. Sample the tool's reading against a person's reading of the same documents, on material that resembles the firm's engagements rather than the vendor's demonstration set. Measure what it missed and what it invented, separately, because the two errors do different damage in an audit file. Repeat the exercise when the version changes, because last quarter's results describe last quarter's tool. The method for doing this properly — building evaluation sets, regression testing, measuring against a baseline — is the territory of this site's course Does Your AI Actually Work?, and it transfers to audit tools with little adaptation. The audit-specific point is what happens to the results: the validation evidence is itself a record the firm should keep, because the inspector's question — why did the firm approve this? — deserves a written answer.
- Sample the tool's reading against a person's, on documents like yours, not the vendor's demonstration set
- Measure misses and inventions separately — they do different damage in an audit file
- Revalidate on version change: last quarter's results describe last quarter's tool
- Keep the validation evidence — the firm's approval decision deserves a written answer too
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.