Data Leakage: What It Is and How to Avoid It
Data leakage happens when information that wouldn't actually be available at prediction time somehow makes its way into your training process. It's one of the most common and most dangerous mistakes in machine learning, precisely because it doesn't look like a mistake — it looks like a great model, right up until it's deployed and quietly falls apart.
Two main flavours:
- Target leakage — a feature contains information that's a direct or indirect proxy for the label, in a way that wouldn't be known at prediction time. Classic example: predicting whether a patient has a disease, using a feature like "was prescribed medication X" — the prescription often only happens because the diagnosis was already made.
- Train-test contamination — information from the test set leaks into training, usually by accident. Common causes: scaling, encoding, or imputing on the full dataset before splitting; computing aggregate features (like an average) across the whole dataset instead of within each fold; or duplicate/near-duplicate rows ending up split across both the training and test sets.
Why is this important? A model suffering from data leakage will show unusually high, almost too-good-to-be-true accuracy during evaluation, and then perform noticeably worse in production — because in production, the leaked information genuinely isn't available anymore. This is one of the most common reasons a model that looked excellent on paper quietly fails once it's actually deployed, and it can be embarrassing to discover after the fact rather than before.
How to prevent it:
- Split first, transform second. Always split your data into train/validation/test before fitting any scaler, encoder, imputer, or feature selector — fit those only on the training data, then apply the fitted transformation to the rest.
- Be suspicious of "too good" results. If your model's accuracy seems unusually high for the problem, that's a signal to check for leakage before celebrating.
- Think about timing. For every feature, ask: would this value genuinely have been known at the moment you needed to make the prediction? This catches a lot of target leakage before it happens.
- Watch time series carefully. A random train/test split can put "future" rows in training and "past" rows in testing, letting the model implicitly learn from the future. Use a time-based split instead for temporal data.
- Check for duplicates across splits. Near-identical rows ending up on both sides of a split can let the model effectively memorise test answers.
Where it connects: leakage can sneak in through almost every step covered elsewhere in this section — imputing missing values, scaling, target encoding, feature selection, and PCA are all places where "fit on everything first" is a tempting shortcut that quietly introduces leakage.
Where to go deeper: the Wikipedia article on Leakage (machine learning) gives a solid overview of the different types and their causes, with further references for deeper reading.