← Back to Resources

Evaluation Metrics for Regression

Practical Resources - AI Engineering

When you're predicting a number rather than a category — house prices, temperatures, delivery times — you need a different set of metrics from classification. They all measure some version of "how far off were the predictions," but they disagree on how to treat those errors.


The main options:

Which should you use? If large errors are especially costly (e.g. underestimating how much stock to order for a big event), RMSE or MSE better reflects that. If you want a metric that's robust to a few extreme outliers and easy to explain to a non-technical stakeholder, MAE is usually the friendlier choice. Reporting more than one metric is common and often useful — a low RMSE with a high MAE can be a sign that most predictions are decent but a few are wildly off.

Why is this important? The choice of metric shapes what "improving the model" even means — optimising for RMSE will lead you to different modelling decisions than optimising for MAE, because they disagree about how much a handful of large errors should matter. Choose the metric that reflects the real-world cost of being wrong, not just the one that's easiest to compute.

Where to go deeper: scikit-learn's regression metrics documentation covers all of the above with formulas and usage notes.