Leakage and Look-Ahead Bias
Information That Would Not Have Been There
Leakage is the presence, in training data, of information that would not have been available at the moment the decision is made. The model uses it, performance looks excellent, and the result cannot be reproduced in production because the information arrives after the point of use — or never arrives at all. Look-ahead bias is the time-ordered form of the same defect, and it is endemic in financial data, because most financial databases are maintained as a current best view of history rather than as a record of what was known on each date. The distinction that decides whether a backtest is honest is between as-reported data, which is what you had, and as-restated data, which is what turned out to be true.
- Leakage means training on information unavailable at decision time; the result will not reproduce live
- Financial databases usually store a current best view of history, not what was known on each date
- As-reported versus as-restated is the distinction that decides whether a backtest is honest
- Excellent results that collapse in production are the signature, and the cause sits upstream of the model
Where It Hides in Financial Data
The hiding places are consistent enough to make a checklist. Fundamental data timestamped to the period it describes rather than to its publication date. Universes and index constituents built from today's membership, which is survivorship bias wearing a different name. Prices adjusted for corporate actions announced later. Ratings, spreads or benchmark series that have since been revised. Delisted, defaulted or acquired entities quietly dropped from the sample. Target leakage in credit and fraud data, where a collections flag, a write-off code or an account closure reason is populated only because the outcome already happened. And features scaled, normalised or imputed across the whole history before the split.
- Fundamentals dated to the period rather than to publication, and prices adjusted for later corporate actions
- Universes built from current membership, and samples missing delisted, defaulted or acquired entities
- Target leakage: collections flags, write-off codes and closure reasons that exist only because the outcome did
- Scaling, imputation and feature engineering performed across the full history before the split
Splitting Time Honestly
Random train-test splits are wrong for anything with a time dimension, and they are the default in most tooling. Split chronologically, train only on the past, and leave a gap between the training window and the evaluation window so that overlapping label horizons and slow-moving features cannot carry information across the boundary; purging and embargoing are the standard names for that gap. Keep the same entity out of both sides wherever records are correlated, because the same borrower, household, account or issuer appearing in training and test lets the model recognise the entity rather than the pattern. And rebuild the feature set as of the decision date, missingness included, since a field being absent is itself informative.
- Split chronologically — random splits are the tooling default and are wrong for time-ordered data
- Leave a purge and embargo gap so overlapping label horizons cannot cross the boundary
- Keep the same borrower, account, household or issuer out of both sides of the split
- Rebuild features as of the decision date, missingness included — an absent field carries information
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.