← Back to Resources

Evaluation Metrics for Classification

Practical Resources - AI Engineering

"Accuracy" is the metric everyone reaches for first, and it's often the wrong one. Picking the right evaluation metric matters just as much as picking the right model — the wrong metric can make a genuinely bad model look great, or a genuinely good one look mediocre.


Why accuracy can mislead you: imagine a dataset where 95% of examples are the negative class (e.g. detecting a rare disease, or fraud). A model that just predicts "negative" every single time gets 95% accuracy while being completely useless. This is exactly the scenario covered in Handling Class Imbalance, and it's why the metrics below usually give a far more honest picture.

The core building blocks:

Which matters more, precision or recall? It depends entirely on the cost of each type of mistake. In spam detection, a false positive (a real email marked as spam) is often worse than a false negative (a spam email that slips through) — so precision matters more. In cancer screening, missing a real case (a false negative) is usually far worse than a false alarm — so recall matters more. There's no universally "correct" answer; it's a judgement call tied to the real-world consequences of each error.

Threshold-independent metrics:

Where to go deeper: scikit-learn's model evaluation documentation covers all of the metrics above with formulas and code, and Google's Machine Learning Crash Course chapter on classification is a clear, visual introduction if you want the intuition built up from scratch.