Feature Scaling and Normalization
A lot of algorithms silently assume that all your numeric features live on roughly the same scale. If one feature ranges from 0 to 1 and another ranges from 0 to 1,000,000, many models will implicitly treat the second feature as far more important — not because it actually is, but purely because of its units.
The main techniques:
- Standardization (Z-score scaling) — subtract the mean and divide by the standard deviation, so each feature has mean 0 and standard deviation 1. The default choice for most algorithms that assume roughly normal-ish data, like linear/logistic regression, SVMs, and neural networks.
- Min-max scaling — rescale values into a fixed range, usually [0, 1]. Useful when you need bounded inputs (e.g. some neural network activation functions) but sensitive to outliers, since a single extreme value compresses everything else into a narrow band.
- Robust scaling — scale using the median and interquartile range instead of the mean and standard deviation. Much less sensitive to outliers than the two methods above, and a good default when your data has extreme values you don't want to remove.
- Log transformation — not strictly scaling, but often used alongside it for heavily right-skewed data (like income or word counts), to pull in a long tail before scaling the rest.
Which models actually need this? Distance-based and gradient-based algorithms care a lot: k-nearest-neighbours, k-means clustering, SVMs, PCA, and neural networks trained with gradient descent all perform noticeably better, and often converge faster, with scaled inputs. Tree-based models (decision trees, random forests, gradient boosting) are an important exception — they split on thresholds one feature at a time, so scaling generally doesn't change their behaviour at all.
Pitfalls to avoid:
- Fit your scaler on the training data only, then apply that same fitted scaler to the validation/test data. Fitting it on the full dataset before splitting leaks information about the test set's distribution into training.
- Remember to apply the exact same scaling at inference time in production, using the parameters learned from training — not values recalculated on whatever new data shows up.
- Scaling doesn't fix a badly skewed or bimodal distribution on its own; it just changes the range. Combine it with transformations like log-scaling when the underlying shape of the data is the actual problem.
Where to go deeper: scikit-learn's preprocessing documentation covers all of the scalers above with code examples and visual comparisons of how each one reshapes a distribution.