The AI Learning Hub Journal

When the AI Inside Your App Can Be Wrong

A wrong answer from your AI feature looks exactly like a right onea broken button looks broken — a wrong summary just looks like a summaryTHE REST OF YOUR APP — BREAKS LOUDLYBook nowa broken button doesnothing — you notice itthe moment you try itfailure announces itselfTHE AI FEATURE — WRONG QUIETLYwell written · specific · plausiblenothing on the screenmarks it as wrong — andyour user asked becausethey cannot check itriskiest exactly where answers cannot be checkedCHECK IT THE BORING WAY — REAL EXAMPLES, ANSWERS WRITTEN FIRSTcollect 10–20 real exampleswrite the good answer firstcompare by eye, one at a timeredo after each changean afternoon of dull work — and honestly how the professional version starts tooDESIGN FOR BEING WRONG — WHERE THE OUTPUT GOES MATTERS MOSTa draft a person approves —never an automatic actionshow where the answer camefrom, so it can be checkedsay on screen: AI wrote this,and it may be wrongTRUST IT LESS THAN THE REST OF YOUR APP — IT FAILS LOOKING FINEkeep the output as a draft, show the source, and re-run your examples after every instruction change
An AI feature fails looking fine — check it against real examples you wrote answers for, and keep its output as a draft a person approves.

Confidently Wrong Looks Exactly Like Right

If your app uses AI to summarise, answer, classify or draft, you have added a component that fails differently from everything else in software. A broken button looks broken. A wrong summary looks like a summary. It will be well written, plausible and specific, and nothing on screen distinguishes it from a correct one. This matters most where users cannot check the answer themselves, which is usually why they are asking. Nothing here says avoid AI features. It says that "it gave a good answer when I tried it" is even weaker evidence here than elsewhere, because this failure mode is specifically shaped to look fine.

  • Wrong output arrives looking exactly like correct output
  • Confidence in the writing carries no information about accuracy
  • Riskiest where users cannot verify the answer — usually the reason they asked
  • Trying it a few times is even weaker evidence than usual here

Check It the Boring Way

You do not need statistics to get useful evidence. Collect ten or twenty real examples: actual customer emails, actual documents, actual questions. Write down what a good answer would be for each. Then run them through and compare, by eye, one at a time. Do it again whenever you change the instructions you give the AI, because those changes have effects you will not predict. This is an afternoon of dull work, and it is genuinely how the professional version starts too. The rigorous version differs in scale and measurement, not in the underlying idea: compare output against answers you decided on in advance.

  • Ten to twenty real examples with the answer you would accept, written first
  • Compare by eye, one at a time — this is enough for a prototype
  • Re-run them whenever you change the AI's instructions
  • Same idea as professional evaluation, just smaller and without the maths

Design for Being Wrong

The most effective protection is not better prompting. It is where you put the output. Keep AI results as drafts a person approves, rather than actions taken automatically. Show the source, so someone can check it. Make it easy to correct and easy to report. Never let an unreviewed AI output send an email to a customer, change a record, or make a decision about a person. And say plainly in the interface that the answer came from AI and may be wrong. Users who know that read differently, and they catch things you never will. Designing for wrongness is cheaper and far more reliable than trying to eliminate it.

  • Draft-for-approval rather than automatic action, wherever a person is affected
  • Show sources so the answer can be checked without trusting it
  • Say in the interface that it is AI-generated and may be wrong
  • Designing around wrongness beats trying to eliminate it

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.