Scaling Inference
A model that comfortably serves ten requests a minute during testing can behave completely differently under real load — thousands of requests a minute, unpredictable spikes, users on the other side of the world. Scaling inference is about keeping a model fast and reliable as demand grows, without the cost growing out of control alongside it.
Core techniques:
- Horizontal scaling / load balancing — run multiple copies of your model server behind a load balancer that distributes incoming requests across them. Lets you handle more traffic by adding more instances, and improves reliability, since one instance failing doesn't take the whole service down.
- Batching — instead of running the model once per individual request, group several incoming requests together and run them through the model in a single batch. Many models (especially neural networks on GPUs) are dramatically more efficient per-request when processing a batch than processing one input at a time, though batching adds a small amount of latency while waiting to collect a batch.
- Caching — if the same input (or a very similar one) is requested repeatedly, store and reuse the previous prediction instead of recomputing it. Especially effective when a meaningful share of requests repeat (a popular product recommendation, a common search query).
- Asynchronous / queued processing — for workloads that don't need an instant response, put requests in a queue and process them as capacity allows, smoothing out traffic spikes rather than trying to handle every burst instantly.
- Autoscaling — automatically add or remove server instances based on current load, so you're not paying for peak capacity around the clock when most traffic is much lower.
Reducing the cost per prediction, not just adding more servers: scaling isn't only about horizontal capacity — making each individual prediction cheaper and faster matters just as much, and is often more cost-effective than simply running more copies of an expensive model. See Model Compression and Optimization for techniques like quantization and distillation that reduce a model's compute footprint directly.
Why is this important? A model that works perfectly in a demo can fail in ways that have nothing to do with its accuracy once it meets real traffic — timeouts, dropped requests, ballooning cloud bills. These are genuinely different problems from the modelling work covered elsewhere in this section, and they're just as capable of making a project fail in practice.
Where to go deeper: the Kubernetes autoscaling documentation is a good practical starting point once you're deploying at a scale where manually managing server instances stops being reasonable.