Train/Validation/Test Splits and Cross-Validation
If you evaluate your model on the same data it was trained on, you're not measuring how well it works — you're measuring how well it memorised. Splitting your data properly is how you get an honest estimate of how your model will perform on data it's never seen.
The three-way split:
- Training set — the data the model actually learns from.
- Validation set — used during development to tune hyperparameters and compare models. You can look at this repeatedly while iterating, but that also means it slowly stops being a fully unbiased measure the more you use it to make decisions.
- Test set — touched once, at the very end, to get a final, honest estimate of performance. If you tune anything based on test set results, it's no longer a fair test — it's effectively become a second validation set.
A common split for a reasonably sized dataset is around 70/15/15 or 80/10/10, though the right proportions depend on how much data you have overall — with very large datasets, even a small percentage held out for validation/test can be plenty of examples.
Cross-validation: with a smaller dataset, a single validation split can give a noisy, unreliable estimate of performance just by chance of which rows ended up where. K-fold cross-validation addresses this by splitting the training data into k roughly equal parts, training k times (each time holding out a different part as validation), and averaging the results. This gives a more stable estimate and uses your data more efficiently, at the cost of training the model k times instead of once.
Variants worth knowing:
- Stratified k-fold — preserves the class proportions in each fold, important for classification with imbalanced classes (see Handling Class Imbalance).
- Time-series split — for temporal data, folds must respect chronological order (train on the past, validate on the future) rather than being shuffled randomly, or you risk leaking future information into training.
- Group k-fold — used when multiple rows belong to the same underlying entity (e.g. multiple images of the same patient) and you need to make sure that entity's data doesn't end up split across both train and validation.
Why is this important? Getting this wrong is one of the most common causes of a model that looks great in development and performs poorly in the real world — closely related to the mistakes covered in the Data Leakage resource. A trustworthy evaluation setup is the foundation everything else in this section builds on; there's no point comparing models or tuning hyperparameters carefully if the number you're optimising against is misleading in the first place.
Where to go deeper: scikit-learn's cross-validation documentation covers standard k-fold, stratified, time-series, and group splitting strategies with code examples for each.