The AI Learning Hub Journal

Leakage and Look-Ahead Bias

Information that would not have been therefinancial databases store a current best view of history, not what was known on each date AS-REPORTED IS WHAT YOU HAD · AS-RESTATED IS WHAT TURNED OUT TO BE TRUE the decision date AVAILABLE AT THE DECISION· as-reported figures, at their publication date· prices as they stood, before later adjustment· features rebuilt as of this date, missingness included ARRIVES ONLY AFTERWARDS· a restatement published afterwards· index membership as it stands today· a corporate action announced later· the outcome the label already encodesused in training anyway — performance looks excellent, and it cannot be reproduced in production WHERE IT HIDES IN FINANCIAL DATA — CONSISTENT ENOUGH TO MAKE A CHECKLIST· fundamentals dated to the period, not to publication· prices adjusted for corporate actions announced later· universes built from today’s membership — survivorship renamed· delisted, defaulted or acquired entities dropped from the sample· target leakage: collections flags, write-off codes, closure reasons· scaling, imputation and feature engineering before the split SPLITTING TIME HONESTLY TRAIN — THE PAST ONLY PURGE EMBARGO TESTrandom train-test splits are the default in most tooling and are wrong for anything with a time dimensionthe purge and embargo gap stops overlapping label horizons and slow-moving features crossing the boundarykeep the same borrower, account, household or issuer out of both sides — or the model recognises the entity, not the pattern EXCELLENT RESULTS THAT COLLAPSE IN PRODUCTION ARE THE SIGNATURE the cause sits upstream of the model, in how the data was assembled and how it was split
Leakage means training on information that would not have been there — and the result will not reproduce live.

Information That Would Not Have Been There

Leakage is the presence, in training data, of information that would not have been available at the moment the decision is made. The model uses it, performance looks excellent, and the result cannot be reproduced in production because the information arrives after the point of use — or never arrives at all. Look-ahead bias is the time-ordered form of the same defect, and it is endemic in financial data, because most financial databases are maintained as a current best view of history rather than as a record of what was known on each date. The distinction that decides whether a backtest is honest is between as-reported data, which is what you had, and as-restated data, which is what turned out to be true.

  • Leakage means training on information unavailable at decision time; the result will not reproduce live
  • Financial databases usually store a current best view of history, not what was known on each date
  • As-reported versus as-restated is the distinction that decides whether a backtest is honest
  • Excellent results that collapse in production are the signature, and the cause sits upstream of the model

Where It Hides in Financial Data

The hiding places are consistent enough to make a checklist. Fundamental data timestamped to the period it describes rather than to its publication date. Universes and index constituents built from today's membership, which is survivorship bias wearing a different name. Prices adjusted for corporate actions announced later. Ratings, spreads or benchmark series that have since been revised. Delisted, defaulted or acquired entities quietly dropped from the sample. Target leakage in credit and fraud data, where a collections flag, a write-off code or an account closure reason is populated only because the outcome already happened. And features scaled, normalised or imputed across the whole history before the split.

  • Fundamentals dated to the period rather than to publication, and prices adjusted for later corporate actions
  • Universes built from current membership, and samples missing delisted, defaulted or acquired entities
  • Target leakage: collections flags, write-off codes and closure reasons that exist only because the outcome did
  • Scaling, imputation and feature engineering performed across the full history before the split

Splitting Time Honestly

Random train-test splits are wrong for anything with a time dimension, and they are the default in most tooling. Split chronologically, train only on the past, and leave a gap between the training window and the evaluation window so that overlapping label horizons and slow-moving features cannot carry information across the boundary; purging and embargoing are the standard names for that gap. Keep the same entity out of both sides wherever records are correlated, because the same borrower, household, account or issuer appearing in training and test lets the model recognise the entity rather than the pattern. And rebuild the feature set as of the decision date, missingness included, since a field being absent is itself informative.

  • Split chronologically — random splits are the tooling default and are wrong for time-ordered data
  • Leave a purge and embargo gap so overlapping label horizons cannot cross the boundary
  • Keep the same borrower, account, household or issuer out of both sides of the split
  • Rebuild features as of the decision date, missingness included — an absent field carries information

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.