Reproducibility in ML Experiments
A few weeks into a project, most teams end up with a folder of models named things like final_model_v3_ACTUALLY_final.pkl, no memory of which hyperparameters produced their best result, and no way to reliably reproduce it. Reproducibility is the discipline of avoiding that, and it's far easier to build the habit early than to reconstruct the history after the fact.
Sources of randomness to control: most ML pipelines have several points where randomness sneaks in — the train/test split, weight initialization in a neural network, dropout, data shuffling during training, and some algorithms' internal randomness (random forests, for example). Setting a fixed random seed at each of these points is what makes a run repeatable. Note that even with a fixed seed, results can still differ slightly across different hardware or library versions — "reproducible" in practice usually means "close enough to draw the same conclusions," not bit-for-bit identical.
What to track for every experiment:
- The exact code version used (a git commit hash is ideal).
- The dataset version (data changes over time too — know which snapshot you trained on).
- All hyperparameters, including ones left at their default value.
- The random seed(s) used.
- The resulting metrics, on both validation and (sparingly) test data.
- Library and framework versions — a model trained with one version of a library can behave subtly differently on another.
Tools that help: experiment tracking tools like MLflow, Weights & Biases, and DVC automate most of the tracking above, logging hyperparameters, metrics, and artifacts for every run so you can compare experiments side by side later rather than digging through old notebook cells. Even a shared spreadsheet with the fields above, updated consistently, is far better than nothing.
Why is this important? Beyond the obvious practical benefit (you can actually find and reproduce your best model later), reproducibility is directly relevant to how your project gets evaluated — a supervisor can far more easily verify and trust results that come with a clear record of how they were produced. It also connects to Effective Teamwork: a shared, logged experiment history means a teammate can pick up and understand your modelling work without a long verbal explanation.
Where to go deeper: the MLflow documentation is a practical, widely used starting point for experiment tracking, model logging, and comparing runs, with a free and open-source core.