The AI Learning Hub Journal

The Honest Map of Task Suitability

Legal tasks by how safely AI can touch themsupervision required rises as you go downTASKWHERE IT GOES WRONGSUPERVISION REQUIREDSummarisation and first draftsinternal material, reworked before it leavesAI can carry real weight hereErrors surface during the editingyou were going to do anywayNORMAL REVIEWread it as you would a junior draftDocument review and issue spottingnarrowing a large set down for a humanAI can carry real weight hereSilent misses and confident falseflags — both need samplingSUPERVISED AND SAMPLEDcheck a slice, not just the hitsResearch, authority and citationanything you would put in front of a courtAI can carry real weight hereFluent text can point at nothing,or at something it misstatesVERIFY EVERY SOURCEunverified means unusableAdvice, strategy and judgementwhat the client should actually dothe line AI does not crossThe duty, the licence and theliability all sit with a personNOT DELEGABLEa tool can inform it, never make itNothing here says do not use it — it says decide, in advance, who is checking what
Suitability tracks how visible the failure is — the tasks where a mistake hides longest are the ones that need a human hardest

Strong: Summarisation and First-Draft Generation

Language models are genuinely good at compressing text you already have and at producing a structured first draft from instructions you supply. Summarising a long deposition transcript, condensing a client email chain into a chronology, turning bullet instructions into a first-pass letter, producing a plain-English explanation of a clause for a client — these play to what the technology actually does. The common thread is that the source material is in front of the model and the human reviewer already knows roughly what the answer should look like. That second condition is what makes these tasks safe: an error is visible to the reviewer because the reviewer has the ground truth. Where you cannot check the output against something you already hold, the task has quietly moved into a different risk category.

  • Best fit: the source text is supplied by you and the reviewer can spot an error on sight
  • Summaries still drop or distort emphasis — read the summary against the source before it goes anywhere external
  • First drafts save typing, not thinking; the structure is a starting point, not a position
  • If you could not detect a wrong answer, the task is not in this category however routine it feels

Strong With Supervision: Document Review and Issue Spotting

Reviewing a large document set for relevance, flagging clauses that deviate from a playbook, spotting missing provisions, surfacing candidate issues in a contract — these work well as a first pass that narrows human attention. The critical word is first pass. The model produces recall and prioritisation, not a conclusion. Two failure directions matter and they are not symmetric: false positives cost review time and are self-correcting, while false negatives are silent and may never be discovered. A system that flags eighty per cent of the risky clauses looks excellent in a demo and is dangerous if anyone treats the unflagged remainder as cleared. Supervision here means sampling what the tool did not flag, not only checking what it did.

  • The output is prioritisation and recall, not a legal conclusion — treat it as triage
  • False negatives are the dangerous direction: silent, unmeasured, and easy to mistake for a clean set
  • Sample the unflagged material deliberately; reviewing only the flags measures nothing about coverage
  • Demo performance on clean documents rarely survives contact with real, messy production sets

Dangerous Without Verification: Research and Citation

Legal research is where generative models are least trustworthy and most convincing. A model asked for authority on a point will produce authority-shaped text: a case name in the right format, a plausible court, a confident parenthetical, a holding that fits the argument. Any or all of it may be fabricated, and none of it looks fabricated. Retrieval-backed research tools that search a real corpus reduce this substantially, but they do not remove it — the model can still misdescribe a real case, cite a real case for a proposition it does not support, or surface authority that has since been overruled. Module 3 covers this in full. For now, the rule is simply that no citation reaches a document until a human has opened the actual source.

  • Fabricated authority is formatted exactly like real authority — plausibility is not a signal here
  • Retrieval-backed tools reduce the risk substantially but do not eliminate misdescription or superseded law
  • A parenthetical that fits your argument suspiciously well is a prompt to verify, not to relax
  • Nothing gets cited until someone has read the actual source document — no exceptions for time pressure

Not Delegable: Advice, Judgement, and Strategy

Some work cannot move to a model regardless of how good the model gets, because the thing being produced is professional judgement exercised by a person who is accountable for it. Advising a client on whether to settle, deciding what to argue and what to concede, assessing a witness, weighing reputational and commercial factors against legal ones, choosing how candid to be with a tribunal — these are the practice of law, not text generation. A model can help you prepare: it can lay out considerations, stress-test an argument, or draft the memo that records your reasoning. What it cannot do is be the source of the judgement. Where a tool appears to offer conclusions rather than material, treat that appearance as a product-design choice and not as a capability.

  • Judgement, advice, and strategy stay with the accountable human — the model prepares, it does not decide
  • Useful support role: surfacing considerations, stress-testing arguments, drafting the record of your reasoning
  • A confident recommendation from a tool is a formatting choice, not evidence that the tool is competent to advise
  • Unauthorised-practice and accountability questions arise fastest exactly where this line is blurred

Try It Yourself

The four categories only earn their keep when you apply them to your own workload. This takes about twenty minutes and no tools at all.

◆ Try it yourself

List ten tasks you actually did last month, described generically — no client names, no matter details, nothing confidential. Sort each into one of the four categories from this lesson, and for every task in the top two categories write the specific check that would catch an error. This is a paper exercise: do not put client material into any AI tool.

Task (described generically):
Category: strong / strong with supervision / dangerous without verification / not delegable
How an error would show itself:
The check that would catch it:
Who performs that check:
How you'll know it worked
  • At least one task changed category once you asked whether you could detect a wrong answer
  • Every task in the top two categories has a named check and a named checker
  • You can say in one sentence why each not-delegable task stays there

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.