The AI Learning Hub Journal
◆ Foundations

Vibes Don't Scale

Where "it seems better" stops being an answerwhat one person can hold in their head is flat; what the system needs checked is notCONFIDENCEWhat informaljudgement coversWhat you need toknow to be sureWhere vibes stopbeing enoughHIGHLOWvibes are fine hereand not fine herea handfulhundredscases, users, reviewers and weeks — all of them going upWHAT BREAKSCases multiplyTen examples you can holdin your head at once. Twohundred of them you cannot.Memory fadesYou cannot recall how lastmonth's version handledthis exact awkward input.Reviewers disagreeTwo people read the sameoutput and rate itdifferently, with no rubric.Regressions hideA change fixes what youlooked at and quietlybreaks what you did not.Past the crossing you are not comparing systems any more, you are comparing memoriesThe crossing is not a milestone you notice — it is one you find out you passedWrite the handful of cases down while they still fit in your head, and score them the same way twiceThe moment two people disagree about whether it improved, you needed a number yesterday
Informal judgement is a fine first eval and a terrible fifth one — the crossing arrives sooner than expected

The Demo Is Not the System

Every AI feature starts the same way. Someone writes a prompt, tries five or six inputs, and the output looks good. That impression — real, honest, and completely uninformative — is the entire evidence base for most shipped AI systems. It works for a while because early usage is narrow and the builder is also the tester. It stops working the moment the input distribution widens beyond what one person imagined, because the failures that matter are not the ones you thought to try. Vibes measure the intersection of your imagination and your patience. Users occupy the complement of that set: the long tail of phrasings, edge-case documents, hostile inputs, and multi-step tasks where a plausible-looking answer is quietly wrong.

  • You test the cases you thought of; users find the ones you didn't
  • Fluent output is a poor accuracy signal — the model is optimised to sound right
  • Manual spot-checks scale linearly with effort while input variety scales combinatorially

The Regression You Never Saw

The specific failure mode that ends vibes-based development is the silent regression. You tighten a system prompt to fix one complaint and quietly break three behaviours nobody re-tested. You swap a model for a cheaper one, confirm your favourite five prompts still work, and lose accuracy on the twelve percent of traffic that involves tables. You add a tool and the agent starts preferring it in situations where it is wrong. None of this produces an exception, a stack trace, or a failing build. It produces slightly worse outputs, distributed across users who mostly do not report them. In deterministic software a regression announces itself; in a probabilistic system it hides inside variance until someone senior notices the product feels worse than it did last quarter.

  • Prompt edits have non-local effects — fixing one behaviour perturbs others
  • Nothing crashes: the failure surface is quality, not availability
  • Users under-report degradation; they route around it or churn
  • Without a baseline you cannot distinguish a regression from a bad day

Three Questions That Expose the Gap

You do not need a maturity model to know whether a team has an evaluation problem. Three questions do it. First: what percentage of real tasks does the system complete correctly right now, and how confident are you in that number? Second: if a better model shipped this morning, how long until you could tell whether it improves your product? Third: when you changed the prompt last week, what got worse? A team with evals answers all three in minutes, with numbers and confidence intervals. A team without them answers with anecdotes, an estimate of several weeks, and silence. The gap between those two teams is not talent or model access — it is measurement infrastructure, and it compounds monthly.

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.