The AI Learning Hub Journal

Backtest Overfitting and the Illusion of Alpha

The strategy got beautiful because it was selected, not because it is realnobody sets out to overfit — the process that produces it is ordinary and looks like diligence HOW A STRATEGY BECOMES BEAUTIFULthe idea is tested — the result is unremarkablea parameter is adjusteda different lookback window is triedtwo loss-making years are set asidea filter drops the trades that went wrongthe universe is narrowed to cleaner namesthe curve is smooth and the drawdowns shallowthe smoother the backtested curve, the more selection went into producing it A REPORTED FIGURE IS THE MAXIMUM OF UNKNOWN DRAWS enough variants against one history and some look excellent by chancethe one that reached youbacktested performance of every variant tried, including the ones nobody recorded →a threshold calibrated to be crossed occasionally by noise means little “WE VALIDATED IT OUT OF SAMPLE” OFTEN MEANS “WE KEPT TRYING” it fails the holdout it is revised it is tested on the same holdout the holdout is now spent the second test is in-sampleacross a long programme the same history is reused by the same team until nothing in the building is genuinely unseenforward-dated live or paper trading is the only sample the researcher could not have fitted WHAT HONEST EVALUATION LOOKS LIKE — MOSTLY PROCESS, NOT MATHEMATICSrecord and report how manyconfigurations were tried —it changes how to read itpre-specify rule, universe,horizon and criterion — thetimestamp is what makes it oneuse trial-count-awarestatistics and multiple-testing adjustmentshold one untouched periodoutside the researcher’scontrol, and spend it onceand treat economic rationale as a filter — an unexplainable signal is more likely a coincidence found by search A STRATEGY THAT SURVIVED ONE ATTEMPT AND ONE THAT SURVIVED FOUR HUNDRED DESERVE DIFFERENT SCEPTICISM the trial count includes abandoned ideas and inherited configurations, and it is almost never recorded
Ask how many things were tried against this same history, not whether an out-of-sample test was run.

How a Strategy Becomes Beautiful

Nobody sets out to overfit. The process that produces it is ordinary and looks like diligence. A researcher has an idea, tests it, and the result is unremarkable. So a parameter is adjusted. A different lookback window is tried. Two loss-making years are set aside as unrepresentative. A filter is added that happens to drop the trades that went wrong. The universe is narrowed to the names where the effect is cleaner. Each step has a defensible rationale, and after enough of them the equity curve is smooth and the drawdowns are shallow. What has been found is not a market regularity. It is the particular shape of one historical sample, learned extremely well.

  • Overfitting is produced by ordinary, individually defensible decisions, not by bad faith
  • Parameter tuning, window choice, sample exclusions, filters and universe narrowing all fit the same history
  • The smoother the backtested curve, the more selection has usually gone into producing it
  • What is learned is the shape of one sample, not a relationship that will persist

Multiple Testing, in Plain Terms

Test enough strategies against the same history and some will look excellent by chance alone. This is not a subtle statistical point; it is arithmetic. A significance threshold appropriate for a single pre-specified hypothesis becomes meaningless once a hundred variants have been tried, because the threshold was designed to be crossed occasionally by noise. The practical difficulty is that the number of trials is rarely recorded and often not even known: variants abandoned mentally, ideas that died in a notebook, configurations inherited from a colleague who left, and every earlier version of the same idea all count as trials, and none of them appear in the write-up.

  • Enough variants against one history guarantees that some will look excellent from noise alone
  • A threshold calibrated for one pre-specified hypothesis means nothing after a hundred quiet attempts
  • The trial count includes abandoned ideas and inherited configurations, and is almost never recorded
  • A reported performance figure is the maximum of an unknown number of draws, not a typical one

"We Validated It" Often Means "We Kept Trying"

Out-of-sample testing is the standard reassurance, and it stops working the moment the out-of-sample period is used more than once. A strategy that fails the holdout, is revised, and is retested has consumed the holdout: the second test is in-sample, whatever it is called in the memo. Over a research programme running for years, the same market history is used repeatedly by the same team, so there is often no genuinely unseen data left anywhere in the building. This is why the useful question about a backtest is not "was it validated out of sample?" but "how many things were tried, by how many people, against this same history, before this one reached me?"

  • A holdout is spent the first time a failing strategy is revised and tested against it again
  • Across a long research programme the same history is reused until nothing is genuinely unseen
  • Ask how many variants were tried and by whom, not whether an out-of-sample test was run
  • Forward-dated live or paper trading is the only sample the researcher could not have fitted

What Honest Evaluation Looks Like

Much of this is fixable with process rather than mathematics. Record the number of configurations tested and report it with the result, because a strategy that survived one attempt and one that survived four hundred deserve different scepticism. Pre-specify the rule, the universe, the horizon and the acceptance criterion before anything is run. Prefer statistics that explicitly discount for the number of trials — deflated performance measures and multiple-testing adjustments exist for exactly this — over an unadjusted headline. Keep one genuinely untouched period, held by somebody other than the researcher, and spend it once. And treat economic rationale as a filter: an unexplainable signal is more likely a coincidence found by search.

  • Record and report the number of configurations tried — it changes how the headline should be read
  • A criterion written once the result is known is not a criterion — the timestamp is what makes it one
  • Use trial-count-aware statistics and multiple-testing adjustments rather than a raw performance figure
  • Hold one untouched period outside the researcher's control and spend it exactly once
◆ See it for yourself
Open Eval Confidence in the library →

See how many test cases you need before a result means anything.

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.