Tips on Handling Missing Values
Real-world datasets almost always have gaps: a sensor that dropped a reading, a survey question someone skipped, a field that simply wasn't tracked before a certain date. How you handle those gaps can change your model's performance and behaviour more than most people expect — and doing it carelessly is a common, avoidable source of bugs and biased results.
Step one: understand why the data is missing. This matters more than which technique you eventually use. Missingness is usually grouped into three types:
- Missing Completely at Random (MCAR) — the gap has nothing to do with any variable, observed or not (e.g. a random sensor glitch). Safest case to handle.
- Missing at Random (MAR) — the chance of missingness depends on other observed variables (e.g. older survey respondents are less likely to answer an income question). Still manageable, but naive fixes can introduce bias if you ignore the relationship.
- Missing Not at Random (MNAR) — the missingness depends on the missing value itself (e.g. people with very high incomes are less likely to disclose them). The hardest and most dangerous case — filling these gaps naively can systematically distort your model.
Common ways to handle it:
- Deletion — drop rows or columns with missing values. Simple, but can throw away a lot of useful data if missingness is common, and can bias your dataset if the missing rows aren't random.
- Simple imputation — fill gaps with the mean, median, or mode of the column. Fast and often a reasonable baseline, but it can quietly shrink the variance of a feature and weaken its relationship with the target.
- Model-based imputation — predict the missing value from the other features, using something like k-nearest-neighbours imputation or a regression/iterative imputer. More accurate than simple imputation, more computationally expensive, and still an estimate rather than a fact.
- Add a "was missing" indicator column — alongside whatever imputation you use, add a binary flag marking which rows had the value filled in. This lets the model learn that missingness itself might carry information (especially useful for MAR/MNAR data), rather than hiding that signal entirely.
- Category of its own — for categorical features, sometimes the cleanest option is just treating "missing" as its own valid category rather than trying to guess a real one.
Pitfalls to avoid:
- Never compute the mean/median/mode (or fit any imputation model) on the full dataset before splitting into train and test — that leaks test-set information into training, similar to the mistakes covered in the Data Leakage resource. Fit the imputer on the training set only, then apply it to the test set.
- Don't assume missing always means "impute it away." Sometimes missingness is the most informative signal in the dataset (e.g. a loan applicant who leaves the income field blank), and papering over it can remove real predictive power.
- Always check how much is missing per column before deciding on a strategy. A column that's 90% empty is usually better dropped than imputed — you'd mostly be inventing data at that point.
Where to go deeper: scikit-learn's imputation module documentation covers simple, KNN, and iterative imputers with working examples, and is a good practical starting point beyond the theory above.