← Back to Resources

Model Serving Basics

Practical Resources - AI Engineering

A trained model sitting in a notebook isn't useful to anyone else yet. Model serving is the process of making that model available so other software (or people) can actually send it inputs and get predictions back — the bridge between "we built a model" and "the model does something for someone."


Batch vs. real-time inference:

Choosing between them is mostly a question of latency requirements and cost: if "good enough by tomorrow morning" is actually good enough, batch inference is almost always simpler and cheaper to run.

Exposing a model as a service: for real-time inference, the model is usually wrapped behind an API so other systems can call it without needing to know anything about how it works internally.

Frameworks like FastAPI or Flask (Python) make it straightforward to wrap a model in a simple REST API for a student project; dedicated model-serving tools like TensorFlow Serving, TorchServe, or BentoML add production-oriented features (batching, versioning, monitoring hooks) on top once you need them.

Why is this important? A model's accuracy on a validation set means very little to the person actually using the system if it never reaches them. Serving is where a lot of practical engineering concerns show up for the first time — latency, reliability, how to handle a malformed request — that don't exist at all while you're just running experiments in a notebook.

Where to go deeper: the FastAPI documentation is a clear, practical starting point for wrapping a Python model in a REST API, and Google Cloud's ML serving best practices guide covers batch vs. real-time trade-offs in more depth.