← Back to Resources

Learning Curves and Diagnosing Model Problems

Practical Resources - AI Engineering

When a model isn't performing well, it's tempting to just start randomly trying things — a bigger model, more data, different features. A learning curve gives you a much faster, evidence-based way to figure out what's actually wrong before you spend time on the wrong fix.


What a learning curve is: a plot of training and validation error (or accuracy) against the amount of training data used (or, in a related version, against training time/epochs). It shows you not just how the model is doing, but how it's trending — which tells you a lot about what would actually help.

How to read one:

Why this matters more than it seems: without this kind of diagnosis, it's very easy to spend days collecting more data when the real problem was an underfitting model that needed more capacity — or to spend days building a bigger model when the real problem was noisy, insufficient data. A learning curve turns "something's wrong, let's guess" into "here's specifically what's wrong."

Related diagnostics worth knowing:

Where to go deeper: scikit-learn's learning curve documentation includes ready-to-use code for plotting both learning and validation curves, with example plots showing each of the patterns described above.