Visual explainer
Ensemble Stacking
Averaging models with fixed weights can't know which one to trust when. Stacking trains a meta-model on their predictions instead, learning the weighting from data.
Fixed weights fail because they treat every prediction equally. When you average a tree, a linear model, and a ridge regression, the result is completely rigid. You need a way to learn when to trust a tree's complex rules and when to trust a linear model's simple trend, relying on their individual expertise.
Out-of-fold Predictions
To train a judge, the evidence must be honest. If base models predict on data they were trained on, a deep tree will just spit back memorised answers and look perfect. Out-of-fold predictions ensure the meta-model only sees how base models perform on genuinely unseen data, preventing a silent overfit.
Training the Meta-Model
The meta-model treats the out-of-fold base predictions as its new input features. It learns a combination rule — often just a straightforward logistic regression — optimising for the final correct answer instead of blindly averaging. The heavy lifting is already done; this layer just learns the blend.
Learned Trust
Instead of a fixed arithmetic average, the stack acts like a board of medical specialists. It learns from historical outcomes that the tree is right in sudden cold snaps and the linear model is right in mild weather, trusting the right tool for each specific case as it arrives.
Where It Breaks
If you stack three gradient-boosted trees with similar hyperparameters, they will make the exact same mistakes on the exact same rows. The meta-model has no diversity of opinion to exploit, giving you one model's predictive performance while forcing you to pay three times the inference cost.
The Quick Version
- Fixed averages can't adapt to different models' varying strengths.
- Stacking trains a meta-model on base model predictions.
- Base predictions must be out-of-fold to prevent leakage.
- The meta-model learns when to trust which base model case by case.
- Stacking identical models adds compute cost without improving accuracy.
What to Read Next
- Ensemble MethodsCombining multiple diverse models to produce a single, more accurate and robust prediction.
- Generative Adversarial NetworksLearn how a game between two networks—a generator forging data and a discriminator catching fakes—allows us to synthesize sharp, highly realistic images.
- Gradient BoostingChain weak learners so each one corrects only the mistakes of the one before it — an additive process that converts many shallow models into a powerful ensemble.