Vibes Don't Scale
The Demo Is Not the System
Every AI feature starts the same way. Someone writes a prompt, tries five or six inputs, and the output looks good. That impression — real, honest, and completely uninformative — is the entire evidence base for most shipped AI systems. It works for a while because early usage is narrow and the builder is also the tester. It stops working the moment the input distribution widens beyond what one person imagined, because the failures that matter are not the ones you thought to try. Vibes measure the intersection of your imagination and your patience. Users occupy the complement of that set: the long tail of phrasings, edge-case documents, hostile inputs, and multi-step tasks where a plausible-looking answer is quietly wrong.
- You test the cases you thought of; users find the ones you didn't
- Fluent output is a poor accuracy signal — the model is optimised to sound right
- Manual spot-checks scale linearly with effort while input variety scales combinatorially
The Regression You Never Saw
The specific failure mode that ends vibes-based development is the silent regression. You tighten a system prompt to fix one complaint and quietly break three behaviours nobody re-tested. You swap a model for a cheaper one, confirm your favourite five prompts still work, and lose accuracy on the twelve percent of traffic that involves tables. You add a tool and the agent starts preferring it in situations where it is wrong. None of this produces an exception, a stack trace, or a failing build. It produces slightly worse outputs, distributed across users who mostly do not report them. In deterministic software a regression announces itself; in a probabilistic system it hides inside variance until someone senior notices the product feels worse than it did last quarter.
- Prompt edits have non-local effects — fixing one behaviour perturbs others
- Nothing crashes: the failure surface is quality, not availability
- Users under-report degradation; they route around it or churn
- Without a baseline you cannot distinguish a regression from a bad day
Three Questions That Expose the Gap
You do not need a maturity model to know whether a team has an evaluation problem. Three questions do it. First: what percentage of real tasks does the system complete correctly right now, and how confident are you in that number? Second: if a better model shipped this morning, how long until you could tell whether it improves your product? Third: when you changed the prompt last week, what got worse? A team with evals answers all three in minutes, with numbers and confidence intervals. A team without them answers with anecdotes, an estimate of several weeks, and silence. The gap between those two teams is not talent or model access — it is measurement infrastructure, and it compounds monthly.
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.