← Back to Resources

Feature Scaling and Normalization

Practical Resources - AI Engineering

A lot of algorithms silently assume that all your numeric features live on roughly the same scale. If one feature ranges from 0 to 1 and another ranges from 0 to 1,000,000, many models will implicitly treat the second feature as far more important — not because it actually is, but purely because of its units.


The main techniques:

Which models actually need this? Distance-based and gradient-based algorithms care a lot: k-nearest-neighbours, k-means clustering, SVMs, PCA, and neural networks trained with gradient descent all perform noticeably better, and often converge faster, with scaled inputs. Tree-based models (decision trees, random forests, gradient boosting) are an important exception — they split on thresholds one feature at a time, so scaling generally doesn't change their behaviour at all.

Pitfalls to avoid:

Where to go deeper: scikit-learn's preprocessing documentation covers all of the scalers above with code examples and visual comparisons of how each one reshapes a distribution.