Backtest Overfitting and the Illusion of Alpha
How a Strategy Becomes Beautiful
Nobody sets out to overfit. The process that produces it is ordinary and looks like diligence. A researcher has an idea, tests it, and the result is unremarkable. So a parameter is adjusted. A different lookback window is tried. Two loss-making years are set aside as unrepresentative. A filter is added that happens to drop the trades that went wrong. The universe is narrowed to the names where the effect is cleaner. Each step has a defensible rationale, and after enough of them the equity curve is smooth and the drawdowns are shallow. What has been found is not a market regularity. It is the particular shape of one historical sample, learned extremely well.
- Overfitting is produced by ordinary, individually defensible decisions, not by bad faith
- Parameter tuning, window choice, sample exclusions, filters and universe narrowing all fit the same history
- The smoother the backtested curve, the more selection has usually gone into producing it
- What is learned is the shape of one sample, not a relationship that will persist
Multiple Testing, in Plain Terms
Test enough strategies against the same history and some will look excellent by chance alone. This is not a subtle statistical point; it is arithmetic. A significance threshold appropriate for a single pre-specified hypothesis becomes meaningless once a hundred variants have been tried, because the threshold was designed to be crossed occasionally by noise. The practical difficulty is that the number of trials is rarely recorded and often not even known: variants abandoned mentally, ideas that died in a notebook, configurations inherited from a colleague who left, and every earlier version of the same idea all count as trials, and none of them appear in the write-up.
- Enough variants against one history guarantees that some will look excellent from noise alone
- A threshold calibrated for one pre-specified hypothesis means nothing after a hundred quiet attempts
- The trial count includes abandoned ideas and inherited configurations, and is almost never recorded
- A reported performance figure is the maximum of an unknown number of draws, not a typical one
"We Validated It" Often Means "We Kept Trying"
Out-of-sample testing is the standard reassurance, and it stops working the moment the out-of-sample period is used more than once. A strategy that fails the holdout, is revised, and is retested has consumed the holdout: the second test is in-sample, whatever it is called in the memo. Over a research programme running for years, the same market history is used repeatedly by the same team, so there is often no genuinely unseen data left anywhere in the building. This is why the useful question about a backtest is not "was it validated out of sample?" but "how many things were tried, by how many people, against this same history, before this one reached me?"
- A holdout is spent the first time a failing strategy is revised and tested against it again
- Across a long research programme the same history is reused until nothing is genuinely unseen
- Ask how many variants were tried and by whom, not whether an out-of-sample test was run
- Forward-dated live or paper trading is the only sample the researcher could not have fitted
What Honest Evaluation Looks Like
Much of this is fixable with process rather than mathematics. Record the number of configurations tested and report it with the result, because a strategy that survived one attempt and one that survived four hundred deserve different scepticism. Pre-specify the rule, the universe, the horizon and the acceptance criterion before anything is run. Prefer statistics that explicitly discount for the number of trials — deflated performance measures and multiple-testing adjustments exist for exactly this — over an unadjusted headline. Keep one genuinely untouched period, held by somebody other than the researcher, and spend it once. And treat economic rationale as a filter: an unexplainable signal is more likely a coincidence found by search.
- Record and report the number of configurations tried — it changes how the headline should be read
- A criterion written once the result is known is not a criterion — the timestamp is what makes it one
- Use trial-count-aware statistics and multiple-testing adjustments rather than a raw performance figure
- Hold one untouched period outside the researcher's control and spend it exactly once
See how many test cases you need before a result means anything.
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.