Task Completion vs Output Quality
Two Different Questions
Evaluation metrics divide into two families that answer different questions and are constantly conflated. Task completion asks whether the job got done: was the refund issued, does the generated code compile and pass its tests, did the extraction capture every required field, is the ticket now in the right state. It is objective, usually checkable in code, and maps directly to user value. Output quality asks how good the artefact is on dimensions with no crisp definition: clarity, tone, faithfulness to sources, appropriate hedging, structure. It needs rubrics and judges. Most systems need both, but they should never be blended into a single score — a beautifully written answer that fails the task is a failure, and a terse answer that succeeds is a success with a style note.
- Task completion is binary and checkable; quality is graded and subjective
- Report them separately — a blended score hides which one moved
- When they conflict, completion is the primary metric
- Many teams measure only quality because it is easier to prompt a judge for
Decompose Before You Measure
An end-to-end pass rate tells you the system failed but not where, which makes it a poor debugging instrument on its own. Decompose the pipeline and measure each stage against its own ground truth: did retrieval surface the document that contains the answer, did the router classify the intent correctly, did the model select the right tool with well-formed arguments, was the tool result interpreted correctly, does the final output satisfy the task. Stage metrics localise faults instantly and prevent the classic misdiagnosis where a retrieval failure is blamed on the model and answered with an expensive upgrade that changes nothing. Keep the end-to-end number as the headline — stage metrics that improve while the end-to-end rate stays flat are a warning that you are optimising a stage that was never the bottleneck.
- Per-stage ground truth: retrieval hit, intent class, tool choice, argument validity, final outcome
- Most "the model is not good enough" reports are retrieval or context failures
- Stage gains with flat end-to-end results mean you fixed a non-bottleneck
Binary Beats Five-Point
When defining what counts as success, prefer several independent binary checks to one graded scale. Humans and models both apply binary criteria far more consistently — the difference between a three and a four on a five-point scale is exactly where inter-rater agreement collapses, and it varies by rater, by mood, and by what they graded immediately before. Decompose the quality you care about into concrete yes-or-no criteria: does it cite a source for every factual claim, does it stay within scope, does it avoid asserting anything absent from the provided context, does it follow the required format. You get a percentage from the criteria that pass, a diagnostic breakdown of which one failed, and a definition of good that survives being handed to a new reviewer.
- Several binary criteria beat one Likert scale for consistency and diagnosis
- Mid-scale distinctions are where rater agreement falls apart
- A per-criterion breakdown tells you what to fix; a score of 3.6 does not
- Reserve partial credit for cases where the task genuinely has degrees
Counting Partial Credit Honestly
Some tasks genuinely admit partial success — extracting seven of nine fields, completing four of five steps — and the temptation is to average everything into a tidy percentage. Be careful, because averaging assumes the components are interchangeable and they usually are not. Missing an optional field and missing the identifier that keys the record are not the same event, and a mean treats them identically. Where components differ in consequence, weight them or, better, define a hard requirement set that must all pass for the case to count as completed, with the remainder reported as a separate quality dimension. State the rule explicitly in the case definition, because implicit partial-credit conventions are among the most common reasons two teams compute different numbers from the same runs.
- Averaging components assumes they matter equally — they rarely do
- Define a must-pass core and grade the remainder separately
- Write the partial-credit rule into the case; implicit conventions diverge
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.