The AI Learning Hub Journal

Task Completion vs Output Quality

Task completion and output quality are not the same measurementone run, two questions, and two entirely different sets of instrumentsOne run of the systemDID IT COMPLETE THE TASK?an outcome you can recomputeTYPICAL CHECKSDid it return the expected structure?Was the right tool called, with the right input?Did the record actually get created?Did it finish inside the time budget?INSTRUMENTATIONAssertions, schema checks, exit codes, replayed tracesSHAPE OF THE ANSWERPass or fail, or a score anyone can recomputeWAS THE OUTPUT ANY GOOD?a judgement someone has to makeTYPICAL CHECKSIs the summary faithful to the source?Is the tone right for the person reading it?Is the answer useful, not just correct?Would a reviewer be happy to send this out?INSTRUMENTATIONA written rubric, human raters, or a model acting as judgeSHAPE OF THE ANSWERA graded judgement, only as good as rater agreementReport them as one blended number and both of these failures vanish into the averageCOMPLETED, AND STILL BADEvery assertion green. The answer was fluent, confidentlywrong, and nothing in the suite was grading that at all.GOOD, AND STILL FAILEDThe text was excellent and the rubric loved it. The toolcall never fired, so nothing ever reached the user.Instrument them separately, report them separately, and let each one fail on its own terms
Two questions, two instruments — blend them into one number and you can no longer tell which half broke

Two Different Questions

Evaluation metrics divide into two families that answer different questions and are constantly conflated. Task completion asks whether the job got done: was the refund issued, does the generated code compile and pass its tests, did the extraction capture every required field, is the ticket now in the right state. It is objective, usually checkable in code, and maps directly to user value. Output quality asks how good the artefact is on dimensions with no crisp definition: clarity, tone, faithfulness to sources, appropriate hedging, structure. It needs rubrics and judges. Most systems need both, but they should never be blended into a single score — a beautifully written answer that fails the task is a failure, and a terse answer that succeeds is a success with a style note.

  • Task completion is binary and checkable; quality is graded and subjective
  • Report them separately — a blended score hides which one moved
  • When they conflict, completion is the primary metric
  • Many teams measure only quality because it is easier to prompt a judge for

Decompose Before You Measure

An end-to-end pass rate tells you the system failed but not where, which makes it a poor debugging instrument on its own. Decompose the pipeline and measure each stage against its own ground truth: did retrieval surface the document that contains the answer, did the router classify the intent correctly, did the model select the right tool with well-formed arguments, was the tool result interpreted correctly, does the final output satisfy the task. Stage metrics localise faults instantly and prevent the classic misdiagnosis where a retrieval failure is blamed on the model and answered with an expensive upgrade that changes nothing. Keep the end-to-end number as the headline — stage metrics that improve while the end-to-end rate stays flat are a warning that you are optimising a stage that was never the bottleneck.

  • Per-stage ground truth: retrieval hit, intent class, tool choice, argument validity, final outcome
  • Most "the model is not good enough" reports are retrieval or context failures
  • Stage gains with flat end-to-end results mean you fixed a non-bottleneck

Binary Beats Five-Point

When defining what counts as success, prefer several independent binary checks to one graded scale. Humans and models both apply binary criteria far more consistently — the difference between a three and a four on a five-point scale is exactly where inter-rater agreement collapses, and it varies by rater, by mood, and by what they graded immediately before. Decompose the quality you care about into concrete yes-or-no criteria: does it cite a source for every factual claim, does it stay within scope, does it avoid asserting anything absent from the provided context, does it follow the required format. You get a percentage from the criteria that pass, a diagnostic breakdown of which one failed, and a definition of good that survives being handed to a new reviewer.

  • Several binary criteria beat one Likert scale for consistency and diagnosis
  • Mid-scale distinctions are where rater agreement falls apart
  • A per-criterion breakdown tells you what to fix; a score of 3.6 does not
  • Reserve partial credit for cases where the task genuinely has degrees

Counting Partial Credit Honestly

Some tasks genuinely admit partial success — extracting seven of nine fields, completing four of five steps — and the temptation is to average everything into a tidy percentage. Be careful, because averaging assumes the components are interchangeable and they usually are not. Missing an optional field and missing the identifier that keys the record are not the same event, and a mean treats them identically. Where components differ in consequence, weight them or, better, define a hard requirement set that must all pass for the case to count as completed, with the remainder reported as a separate quality dimension. State the rule explicitly in the case definition, because implicit partial-credit conventions are among the most common reasons two teams compute different numbers from the same runs.

  • Averaging components assumes they matter equally — they rarely do
  • Define a must-pass core and grade the remainder separately
  • Write the partial-credit rule into the case; implicit conventions diverge

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.