Hyperparameter Tuning
Hyperparameters are the settings you choose before training starts — learning rate, tree depth, number of layers, regularization strength — as opposed to parameters the model learns on its own from data. Picking good hyperparameters can be the difference between a mediocre model and a genuinely strong one, using the exact same algorithm and data.
Main search strategies:
- Grid search — try every combination of a predefined set of values for each hyperparameter. Exhaustive and simple to reason about, but the number of combinations explodes quickly as you add more hyperparameters or more values per parameter.
- Random search — sample random combinations instead of trying every one. Perhaps counterintuitively, this is often more efficient than grid search for the same computational budget, because it explores a wider range of values for the hyperparameters that actually matter, rather than wasting evaluations on fine-grained values of ones that don't.
- Bayesian optimization — use the results of previous trials to intelligently choose which combination to try next, rather than searching blindly. Converges to good hyperparameters in fewer trials than grid or random search, at the cost of more complexity to set up. Tools like Optuna and Hyperopt implement this.
- Manual tuning — guided by experience and intuition about what each hyperparameter does. Still common for quick iteration, especially early in a project, but doesn't scale well and is easy to bias unconsciously.
Practical advice:
- Always tune against your validation set (or cross-validation), never your test set — the test set exists to give you one honest, final number, and tuning against it defeats that purpose.
- Start with a coarse search over a wide range, then narrow in around promising regions, rather than guessing a fine-grained range up front.
- Not all hyperparameters matter equally. Learning rate, for instance, is usually one of the most impactful for neural networks — it's often worth allocating more of your search budget there than to less sensitive settings.
- Log every trial's hyperparameters and results, even informally. See the Reproducibility resource — it's very easy to lose track of which combination actually produced your best result.
Why is this important? The same model architecture can perform dramatically differently depending purely on hyperparameter choices. Skipping this step and just using default values leaves real performance on the table; over-investing in it too early (before your features and model choice are solid) wastes time tuning something you're about to change anyway. Get the earlier stages — data, features, model choice — roughly right first, then tune.
Where to go deeper: the Optuna documentation is a good practical starting point for Bayesian hyperparameter optimization in Python, with a gentle learning curve and good integration with common ML libraries.