← Back to Resources

Ensemble Methods

Practical Resources - AI Engineering

Instead of picking one model and hoping it's the best possible fit, ensemble methods combine several models together, on the idea that their combined predictions can be more accurate and more stable than any single one alone — especially when the individual models make different kinds of mistakes.


Bagging (Bootstrap Aggregating) — train many versions of the same model on different random samples of the training data (sampled with replacement), then average their predictions (regression) or take a majority vote (classification). This mainly reduces variance, which is why it works especially well with high-variance models like unpruned decision trees. Random Forests are the classic example: many decision trees, each trained on a different data sample and a different random subset of features, combined together.

Boosting — train models sequentially, where each new model focuses specifically on correcting the mistakes of the ones before it. This mainly reduces bias, and tends to produce very strong predictive performance, though it's more prone to overfitting than bagging if left unchecked and needs careful regularization and tuning. Gradient boosting libraries like XGBoost, LightGBM, and CatBoost are consistently among the strongest performers on structured/tabular data, and are a very common choice in real-world and competition settings alike.

Stacking — train several different models (potentially very different types — a linear model, a tree-based model, a neural network), then train a further "meta-model" on top of their outputs to learn how best to combine them. Can squeeze out extra performance by leveraging different models' different strengths, at the cost of significantly more complexity to build, tune, and maintain.

Why ensembles tend to work: if individual models make somewhat different, somewhat independent errors, combining them tends to cancel out some of those errors rather than compounding them — similar in spirit to why averaging several independent guesses tends to beat any one guess. The less correlated the individual models' mistakes are, the more an ensemble typically helps.

Trade-offs to keep in mind:

Where to go deeper: the XGBoost documentation's introduction to boosted trees is a clear, practical explanation of gradient boosting specifically, and scikit-learn's ensemble methods documentation covers bagging, boosting, and stacking together with a consistent framing.