Regularization
Regularization is a family of techniques that deliberately constrain a model, trading a small amount of training accuracy for better generalisation to new data. It's one of the main tools for fighting the overfitting side of the bias-variance tradeoff.
Common techniques:
- L2 regularization (Ridge / weight decay) — adds a penalty proportional to the square of the model's weights, discouraging any single weight from becoming too large. Tends to shrink weights smoothly towards zero without necessarily eliminating any.
- L1 regularization (Lasso) — adds a penalty proportional to the absolute value of the weights. Tends to push some weights to exactly zero, which effectively performs a form of automatic feature selection as a side effect.
- Elastic Net — a blend of L1 and L2, useful when you want some feature selection but also the stability L2 provides.
- Dropout (neural networks) — randomly "turns off" a fraction of neurons during each training step, preventing the network from relying too heavily on any single pathway and forcing it to learn more robust, redundant representations.
- Early stopping — stop training once validation performance stops improving (or starts getting worse), even if training performance is still improving. Simple, cheap, and one of the most broadly useful regularization tools available.
- Data augmentation — not usually filed under "regularization" by name, but it has a similar effect: exposing the model to more variation makes it harder for it to simply memorise the training set. See the Data Augmentation resource.
How much regularization is right? This is itself a hyperparameter to tune (see Hyperparameter Tuning), not a fixed rule. Too little and you're still overfitting; too much and you push the model back towards underfitting by constraining it more than the data actually calls for. Watch the gap between training and validation performance as you adjust the regularization strength — that gap narrowing (without both getting worse) is the sign you're moving in the right direction.
Why is this important? A model that performs beautifully on training data but poorly in the real world is arguably worse than a slightly less accurate model that generalises reliably — the whole point of building a model is for it to work on data it hasn't seen yet. Regularization is one of the most direct, well-understood levers for closing that gap.
Where to go deeper: scikit-learn's documentation on Ridge and Lasso regression gives concrete, code-backed examples of L1 vs. L2 regularization in practice.