The AI Learning Hub Journal
◆ Free course · 4 modules · 24 lessons

Does Your AI Actually Work?

Testing AI systems properly — building eval sets, judging quality, and catching regressions before your users do

Read it in the library →
Module 1 · 5 lessons

Why Evals Are the Job

Why vibes-based development collapses at scale, evals as the unit test of probabilistic systems, eval-driven development as a workflow, and the trap of optimising a metric that has drifted from user value.

  1. Vibes Don't Scale
  2. Evals as the Unit Test of Probabilistic Systems
  3. The Cost of Not Having Them
  4. Eval-Driven Development
  5. What Good Looks Like — and Goodhart's Trap
Module 2 · 6 lessons

Building Eval Sets

Sourcing real cases, curating golden datasets, choosing metrics, writing rubrics both humans and models can apply, running LLM judges without fooling yourself, and knowing when the sample is too small to conclude anything.

  1. Sourcing Real Cases
  2. Golden Datasets and Curation
  3. Task Completion vs Output Quality
  4. Rubrics Humans and Models Can Both Apply
  5. LLM-as-Judge, Done Properly
  6. Agreement, Significance, and Sample Size
Module 3 · 7 lessons

Benchmarks, Regression, and Production

What public benchmarks can and cannot tell you, contamination, regression suites that survive model upgrades, CI for prompts and agents, testing around non-determinism, and closing the loop from production traces back into the eval set.

  1. Reading Public Benchmarks Honestly
  2. Regression Suites That Survive Model Upgrades
  3. CI for Prompts and Agents
  4. Testing Around Non-Determinism
  5. Observability and Closing the Loop
  6. Cost and Latency as First-Class Metrics
  7. Safety Evals: Testing for Harm
Module 4 · 6 lessons

Evaluating What You Actually Build

Evaluation applied to the systems teams really ship: separating retrieval failure from generation failure, grading agent outcomes against trajectories, human review that scales without exhausting reviewers, online testing behind guardrails, choosing on the quality-cost-latency frontier, and the habits that keep an eval suite alive.

  1. Testing a System That Answers From Your Documents
  2. Testing Agents That Take Several Steps
  3. Human Review That Scales
  4. Testing in Production: A/B and Guardrails
  5. Quality Against Cost and Latency
  6. Making Evals a Team Habit

Every lesson is free, with no sign-up. Reading happens in the library, where your progress is saved on your device.