Monitoring Models in Production
Deploying a model isn't the finish line — it's the point where a new kind of work starts. Unlike most software, a machine learning model can silently get worse over time without a single line of its code changing, purely because the world it's making predictions about has moved on. Monitoring is how you find out before your users, or your grade, do.
What to monitor — three broad layers:
- System health — the same things you'd monitor for any deployed service: latency, error rates, uptime, resource usage (CPU/memory/GPU). If the model is too slow or crashing, nothing else about it matters yet.
- Data and prediction health — is the incoming data still shaped like what the model was trained on? Are predictions still distributed roughly the way you'd expect? See Data and Concept Drift for the specific patterns to watch for here.
- Business/task performance — the metrics that actually matter for the model's purpose (accuracy, precision/recall, conversion rate, user engagement) — where possible, tracked against real outcomes as they become available, not just predictions in isolation.
The tricky part: you often don't have ground truth immediately. In many real systems, you don't find out whether a prediction was actually correct until much later (or not at all) — you might know a loan default happened months after the prediction, or never directly know if a recommendation was "correct." This is exactly why monitoring input data distributions and prediction distributions matters: they're available immediately and can act as an early warning sign, even before you have the labels needed to measure accuracy directly.
Setting up useful alerts:
- Alert on meaningful changes, not noise — a threshold that's too sensitive trains people to ignore alerts entirely, which defeats the purpose.
- Segment your monitoring where it makes sense (by user group, region, input type) — an average metric can look perfectly healthy while performance has quietly collapsed for one important subgroup.
- Track trends over time, not just point-in-time snapshots — a slow, steady decline is often more dangerous (and easier to miss) than a sudden obvious break.
Why is this important? A model that isn't monitored can fail silently for a long time, and by the time someone notices (usually because a person complains, or a downstream number looks off), real damage may already be done. Monitoring turns "we hope it's still working" into "we know it's still working," which matters a great deal once real users or real decisions depend on the model's output.
Where to go deeper: the open-source Evidently AI library is built specifically for ML monitoring — data drift, prediction drift, and model quality reports — and its documentation is a good practical entry point even for a smaller student-scale project.