Handling Outliers
An outlier is a data point that sits far outside the pattern of the rest of your data. Sometimes it's a genuine, important edge case; sometimes it's a data entry error or a broken sensor. Treating one as if it were the other is one of the more common ways a model quietly ends up wrong.
Step one: figure out why it's there. Before removing or transforming anything, ask whether the outlier is:
- An error — a typo (age = 999), a unit mismatch, a broken sensor reading. These should generally be corrected or removed.
- A rare but real event — a genuine fraud transaction, an extreme weather event, a celebrity's unusually large purchase. These are often exactly what you care most about predicting, and removing them can quietly delete the signal your model most needs.
Common ways to detect outliers:
- Z-score — flag points more than roughly 3 standard deviations from the mean. Simple, but assumes roughly normal data and is itself sensitive to extreme values pulling the mean and standard deviation around.
- IQR (interquartile range) method — flag points below Q1 − 1.5×IQR or above Q3 + 1.5×IQR. More robust than the Z-score approach and the basis for the whiskers on a standard box plot.
- Visual inspection — box plots, scatter plots, and histograms will often make outliers obvious at a glance, and are a good first step before reaching for a formula. See the Exploratory Data Analysis resource.
- Model-based detection — for higher-dimensional data where simple thresholds don't capture "unusual," methods like isolation forests or local outlier factor can flag points that are unusual across several features at once, not just one at a time.
What to do once you've found them:
- Remove confirmed errors, once you've verified they aren't real.
- Cap / winsorize extreme-but-real values to a reasonable maximum or minimum, keeping the row but limiting its influence.
- Transform the whole feature (e.g. a log transform) to shrink the relative influence of extreme values without deleting them.
- Use a robust technique instead of removing anything — e.g. robust scaling (see Feature Scaling and Normalization) or a model that's naturally less sensitive to extreme values, like tree-based methods.
- Leave them in, deliberately, if they represent exactly the rare cases your model needs to learn to handle.
Pitfalls to avoid: don't reach for automatic outlier removal as a default step in every pipeline. In fraud detection, medical diagnosis, and many other domains, the outliers are the point. Always ask what an outlier being removed would mean for the real-world problem before deleting it.
Where to go deeper: scikit-learn's novelty and outlier detection documentation covers isolation forests, local outlier factor, and other model-based approaches with examples.