Model Serving Basics
A trained model sitting in a notebook isn't useful to anyone else yet. Model serving is the process of making that model available so other software (or people) can actually send it inputs and get predictions back — the bridge between "we built a model" and "the model does something for someone."
Batch vs. real-time inference:
- Batch inference — run the model on a large set of inputs all at once, on a schedule (e.g. nightly), and store the results somewhere for later use. Good fit when predictions don't need to be instant — recommending products for tomorrow's homepage, generating a weekly report. Usually cheaper and simpler to build than real-time serving.
- Real-time (online) inference — the model responds to individual requests as they arrive, typically within milliseconds to a few seconds. Necessary when a prediction is needed immediately as part of a live interaction — fraud detection at the moment of a transaction, a chatbot response, a live recommendation while someone is browsing.
Choosing between them is mostly a question of latency requirements and cost: if "good enough by tomorrow morning" is actually good enough, batch inference is almost always simpler and cheaper to run.
Exposing a model as a service: for real-time inference, the model is usually wrapped behind an API so other systems can call it without needing to know anything about how it works internally.
- REST APIs — the most common approach, using standard HTTP requests (usually JSON in, JSON out). Simple, widely understood, and easy to test with tools every developer already knows.
- gRPC — a faster, more efficient protocol better suited to high-throughput or low-latency internal services, at the cost of being less human-readable and slightly more complex to set up.
Frameworks like FastAPI or Flask (Python) make it straightforward to wrap a model in a simple REST API for a student project; dedicated model-serving tools like TensorFlow Serving, TorchServe, or BentoML add production-oriented features (batching, versioning, monitoring hooks) on top once you need them.
Why is this important? A model's accuracy on a validation set means very little to the person actually using the system if it never reaches them. Serving is where a lot of practical engineering concerns show up for the first time — latency, reliability, how to handle a malformed request — that don't exist at all while you're just running experiments in a notebook.
Where to go deeper: the FastAPI documentation is a clear, practical starting point for wrapping a Python model in a REST API, and Google Cloud's ML serving best practices guide covers batch vs. real-time trade-offs in more depth.