Evaluation Metrics for Regression
When you're predicting a number rather than a category — house prices, temperatures, delivery times — you need a different set of metrics from classification. They all measure some version of "how far off were the predictions," but they disagree on how to treat those errors.
The main options:
- MAE (Mean Absolute Error) — the average of the absolute difference between predicted and actual values. Easy to interpret directly in the original units (e.g. "off by $12,000 on average"), and treats all errors proportionally — a $10 error counts exactly ten times as much as a $1 error.
- MSE (Mean Squared Error) — the average of the squared differences. Squaring penalises large errors much more heavily than small ones, which is useful when big mistakes are disproportionately costly, but it's less intuitive to interpret directly (the units are squared).
- RMSE (Root Mean Squared Error) — the square root of MSE, which brings the units back to the original scale while keeping MSE's tendency to penalise large errors more. Probably the single most commonly reported regression metric.
- R² (coefficient of determination) — the proportion of variance in the target that the model explains, on a scale roughly from 0 to 1 (it can go negative for a model worse than just predicting the mean). Useful for a quick sense of overall model quality, but doesn't tell you anything about the actual size of errors in real units.
- MAPE (Mean Absolute Percentage Error) — error expressed as a percentage of the actual value, useful when you want a scale-independent metric (e.g. comparing performance across products with very different price ranges). Breaks down badly when actual values can be at or near zero, since you end up dividing by something tiny or zero.
Which should you use? If large errors are especially costly (e.g. underestimating how much stock to order for a big event), RMSE or MSE better reflects that. If you want a metric that's robust to a few extreme outliers and easy to explain to a non-technical stakeholder, MAE is usually the friendlier choice. Reporting more than one metric is common and often useful — a low RMSE with a high MAE can be a sign that most predictions are decent but a few are wildly off.
Why is this important? The choice of metric shapes what "improving the model" even means — optimising for RMSE will lead you to different modelling decisions than optimising for MAE, because they disagree about how much a handful of large errors should matter. Choose the metric that reflects the real-world cost of being wrong, not just the one that's easiest to compute.
Where to go deeper: scikit-learn's regression metrics documentation covers all of the above with formulas and usage notes.