← Back to Resources

A/B Testing and Canary Deployments for ML

Practical Resources - AI Engineering

A new model scoring better on your offline validation set is encouraging, but it isn't proof it will actually perform better for real users in the real system. Offline metrics and real-world impact don't always agree — which is why teams rarely flip a switch and send 100% of traffic to a brand new model on day one.


Canary deployment — roll the new model out to a small slice of traffic first (say, 5%), while the rest continues to use the existing model. Watch the new version closely for errors, latency problems, or anything unexpected, and gradually increase its share of traffic if it looks healthy. This is primarily a risk-management technique: it limits the blast radius if something is wrong with the new model, rather than exposing every user to a broken deployment at once.

A/B testing — split traffic between two (or more) versions deliberately and consistently, specifically to measure which one performs better on metrics that matter, using statistical methods to determine whether an observed difference is real or just noise. This is primarily a measurement technique: the goal isn't just to roll out safely, it's to actually learn which model is better and by how much.

How they differ in practice: a canary deployment is usually short-lived and asymmetric (a small percentage on the new version, ramping up over time), focused on catching problems quickly. An A/B test is often run for longer, with a more even and carefully controlled split, focused on getting a statistically meaningful answer to "which one is actually better." In practice, many teams do both together — start as a canary to de-risk the rollout, then let it run long enough to also serve as a valid A/B test.

Practical considerations:

Why is this important? Offline evaluation metrics (see Evaluation Metrics for Classification) are a proxy for real-world value, not a guarantee of it. Controlled, gradual rollouts are how teams close that gap safely, catching both technical failures and cases where a model that looks better on paper doesn't actually translate into a better outcome for real users.

Where to go deeper: Google Cloud's ML on Google Cloud best practices guide touches on staged rollout strategies for models in production, and Martin Fowler's classic article on Canary Release explains the general pattern this section adapts from standard software deployment.