Bias-Variance Tradeoff, Overfitting and Underfitting
Almost every problem you'll hit while training a model traces back to one of two failure modes: the model is too simple to capture the real pattern (underfitting), or it's captured the training data too precisely, noise and all (overfitting). Understanding this tradeoff makes debugging a badly performing model far less mysterious.
Bias is the error that comes from a model being too simple to represent the true underlying pattern — think of trying to fit a straight line through data that's actually curved. High-bias models underfit: they perform poorly on both the training data and new data, because they never really learned the pattern in the first place.
Variance is the error that comes from a model being too sensitive to the specific training data it saw, including its noise and quirks. High-variance models overfit: they perform very well on training data (sometimes near-perfectly) and noticeably worse on new, unseen data, because they've effectively memorised training examples rather than learning generalisable patterns.
The tradeoff: as you increase a model's complexity (more features, deeper trees, more layers), bias tends to go down and variance tends to go up. The goal isn't to eliminate either one entirely — it's to find the sweet spot where their combined effect on real-world error is smallest.
How to tell which one you're dealing with:
- Underfitting — both training and validation error are high and similar to each other. The model isn't even doing well on the data it learned from.
- Overfitting — training error is low, but validation error is noticeably higher. The gap between the two is the tell.
- Plotting a learning curve (error vs. amount of training data or training time) is the most direct way to diagnose which regime you're in.
What to do about each:
- If underfitting: use a more expressive model, add more/better features, reduce regularization, or train for longer.
- If overfitting: get more training data, simplify the model, add regularization, use early stopping, or apply data augmentation.
Why is this important? Almost every decision covered elsewhere in this section — model choice, regularization, hyperparameter tuning, ensembling — is ultimately a way of managing this one tradeoff. Once you can look at a training/validation gap and immediately know whether you're overfitting or underfitting, most model debugging becomes a lot less like guesswork.
Where to go deeper: IBM's explainer on the bias-variance tradeoff covers the underlying theory and prediction error decomposition in more technical depth, with clear visuals.