Does Your AI Actually Work?
Testing AI systems properly — building eval sets, judging quality, and catching regressions before your users do
Read it in the library →Why Evals Are the Job
Why vibes-based development collapses at scale, evals as the unit test of probabilistic systems, eval-driven development as a workflow, and the trap of optimising a metric that has drifted from user value.
Building Eval Sets
Sourcing real cases, curating golden datasets, choosing metrics, writing rubrics both humans and models can apply, running LLM judges without fooling yourself, and knowing when the sample is too small to conclude anything.
Benchmarks, Regression, and Production
What public benchmarks can and cannot tell you, contamination, regression suites that survive model upgrades, CI for prompts and agents, testing around non-determinism, and closing the loop from production traces back into the eval set.
Evaluating What You Actually Build
Evaluation applied to the systems teams really ship: separating retrieval failure from generation failure, grading agent outcomes against trajectories, human review that scales without exhausting reviewers, online testing behind guardrails, choosing on the quality-cost-latency frontier, and the habits that keep an eval suite alive.
Every lesson is free, with no sign-up. Reading happens in the library, where your progress is saved on your device.