← Back to Resources

Data Leakage: What It Is and How to Avoid It

Practical Resources - AI Engineering

Data leakage happens when information that wouldn't actually be available at prediction time somehow makes its way into your training process. It's one of the most common and most dangerous mistakes in machine learning, precisely because it doesn't look like a mistake — it looks like a great model, right up until it's deployed and quietly falls apart.


Two main flavours:

Why is this important? A model suffering from data leakage will show unusually high, almost too-good-to-be-true accuracy during evaluation, and then perform noticeably worse in production — because in production, the leaked information genuinely isn't available anymore. This is one of the most common reasons a model that looked excellent on paper quietly fails once it's actually deployed, and it can be embarrassing to discover after the fact rather than before.

How to prevent it:

Where it connects: leakage can sneak in through almost every step covered elsewhere in this section — imputing missing values, scaling, target encoding, feature selection, and PCA are all places where "fit on everything first" is a tempting shortcut that quietly introduces leakage.

Where to go deeper: the Wikipedia article on Leakage (machine learning) gives a solid overview of the different types and their causes, with further references for deeper reading.