Responsible AI in Production
A model can be technically excellent by every metric covered in Model Selection, Training and Evaluation and still cause real harm once it's making decisions that affect real people. This resource is a starting checklist for the questions worth asking once a model leaves the notebook, not a complete treatment of a genuinely large field.
Fairness: a model can perform well on average while performing meaningfully worse for a specific subgroup — a pattern that's easy to miss if you only ever look at aggregate metrics. Where it's appropriate and you have the data, evaluate performance separately across relevant groups, not just overall. Bias often enters not through deliberate intent but through training data that reflects historical inequities, or through under-representation of certain groups in the data itself.
Explainability: for many applications, especially ones with real consequences for people (lending, hiring, healthcare, criminal justice), it matters not just that a model is accurate, but that its decisions can be explained. Simpler models (linear/logistic regression, single decision trees) are naturally more interpretable; for more complex models, tools like SHAP or LIME can approximate an explanation for individual predictions after the fact. Explainability requirements are also frequently a genuine legal obligation in regulated domains, not just a nice-to-have.
Security: deployed ML systems have their own specific attack surface beyond normal application security. A few worth knowing about:
- Adversarial examples — inputs deliberately, sometimes subtly, crafted to fool a model into a wrong prediction.
- Data poisoning — an attacker deliberately corrupting training data (especially relevant if a model retrains on data that includes user-submitted content) to bias the model's future behaviour.
- Model extraction/inversion — an attacker repeatedly querying a deployed model to reconstruct a close approximation of it, or to infer sensitive information about the data it was trained on.
- Standard practices still matter too: input validation, rate limiting, and access control on your model's API are basic but genuinely important defences.
Privacy: models can sometimes memorise and inadvertently leak details of their training data, which matters a great deal if that data includes anything personal or sensitive. Be deliberate about what data a model is trained on, and be aware that "the model doesn't store the raw data" isn't automatically the same as "the model can't leak information about the data."
Why is this important? These considerations rarely show up in a standard accuracy metric, which is exactly why they're easy to overlook under deadline pressure. Thinking about them explicitly — even briefly, even just as a short section in your project documentation — is often what separates a project that's genuinely thoughtful about its real-world impact from one that technically works but hasn't considered who it affects and how.
Where to go deeper: Google's Responsible AI Practices guide is a practical, non-academic starting point covering fairness, interpretability, privacy, and security together, with links to more specialised tools for each area.