Evaluation Metrics for Classification
"Accuracy" is the metric everyone reaches for first, and it's often the wrong one. Picking the right evaluation metric matters just as much as picking the right model — the wrong metric can make a genuinely bad model look great, or a genuinely good one look mediocre.
Why accuracy can mislead you: imagine a dataset where 95% of examples are the negative class (e.g. detecting a rare disease, or fraud). A model that just predicts "negative" every single time gets 95% accuracy while being completely useless. This is exactly the scenario covered in Handling Class Imbalance, and it's why the metrics below usually give a far more honest picture.
The core building blocks:
- Precision — of everything the model predicted as positive, what fraction actually was? High precision means few false alarms.
- Recall (sensitivity) — of everything that actually was positive, what fraction did the model catch? High recall means few missed cases.
- F1 score — the harmonic mean of precision and recall, useful as a single number when you care about both and don't want to favour one over the other by default.
- Confusion matrix — a simple table of true positives, false positives, true negatives, and false negatives. Often the most informative single artifact, since every other metric here is really just a summary of it.
Which matters more, precision or recall? It depends entirely on the cost of each type of mistake. In spam detection, a false positive (a real email marked as spam) is often worse than a false negative (a spam email that slips through) — so precision matters more. In cancer screening, missing a real case (a false negative) is usually far worse than a false alarm — so recall matters more. There's no universally "correct" answer; it's a judgement call tied to the real-world consequences of each error.
Threshold-independent metrics:
- ROC-AUC — plots true positive rate against false positive rate across all possible classification thresholds, and summarises it as a single number between 0.5 (no better than random) and 1 (perfect). Useful for comparing models independent of any one chosen threshold, but can be overly optimistic on heavily imbalanced datasets.
- Precision-Recall AUC — similar idea, but plots precision against recall. Generally a more informative choice than ROC-AUC when the positive class is rare, since it doesn't get inflated by the (often huge) number of true negatives.
Where to go deeper: scikit-learn's model evaluation documentation covers all of the metrics above with formulas and code, and Google's Machine Learning Crash Course chapter on classification is a clear, visual introduction if you want the intuition built up from scratch.