Skip to content
AI360Xpert
Core ML

Visual explainer

Ensemble Stacking

Averaging models with fixed weights can't know which one to trust when. Stacking trains a meta-model on their predictions instead, learning the weighting from data.

A simple average gives equal weight to all models, even when one model is clearly wrong for a specific case.
A simple average gives equal weight to all models, even when one model is clearly wrong for a specific case.

Fixed weights fail because they treat every prediction equally. When you average a tree, a linear model, and a ridge regression, the result is completely rigid. You need a way to learn when to trust a tree's complex rules and when to trust a linear model's simple trend, relying on their individual expertise.

Out-of-fold Predictions

Base models predict on held-out folds so they cannot use memorised training data.
Base models predict on held-out folds so they cannot use memorised training data.

To train a judge, the evidence must be honest. If base models predict on data they were trained on, a deep tree will just spit back memorised answers and look perfect. Out-of-fold predictions ensure the meta-model only sees how base models perform on genuinely unseen data, preventing a silent overfit.

Training the Meta-Model

The out-of-fold predictions become the input features for a simple meta-model.
The out-of-fold predictions become the input features for a simple meta-model.

The meta-model treats the out-of-fold base predictions as its new input features. It learns a combination rule — often just a straightforward logistic regression — optimising for the final correct answer instead of blindly averaging. The heavy lifting is already done; this layer just learns the blend.

Learned Trust

The meta-model learns to trust the tree for complex regions and the linear model for simple trends.
The meta-model learns to trust the tree for complex regions and the linear model for simple trends.

Instead of a fixed arithmetic average, the stack acts like a board of medical specialists. It learns from historical outcomes that the tree is right in sudden cold snaps and the linear model is right in mild weather, trusting the right tool for each specific case as it arrives.

Where It Breaks

Stacking identical models gives the meta-model nothing to learn.
Stacking identical models gives the meta-model nothing to learn.

If you stack three gradient-boosted trees with similar hyperparameters, they will make the exact same mistakes on the exact same rows. The meta-model has no diversity of opinion to exploit, giving you one model's predictive performance while forcing you to pay three times the inference cost.

The Quick Version

  • Fixed averages can't adapt to different models' varying strengths.
  • Stacking trains a meta-model on base model predictions.
  • Base predictions must be out-of-fold to prevent leakage.
  • The meta-model learns when to trust which base model case by case.
  • Stacking identical models adds compute cost without improving accuracy.