Visual explainer
SHAP & LIME
Explainability methods that approximate black-box models locally to attribute predictions to individual features.
Modern machine learning models—like deep neural networks or large tree ensembles—are powerful black boxes. They make highly accurate predictions but hide their internal reasoning.
LIME (Local Interpretable Model-agnostic Explanations) addresses this by zooming in. While a model's global decision boundary is hopelessly non-linear, the local region immediately surrounding any single data point can be accurately approximated by a fully transparent linear surrogate.
Additive Feature Forces
Where LIME builds a local surrogate, SHAP (SHapley Additive exPlanations) uses game theory to assign exact credit to each feature.
SHAP values decompose the model's prediction into an additive sum. Starting from the dataset's overall average prediction (the base value), each feature acts as an independent force pushing the prediction up or down until they arrive precisely at the model's final score.
The Perturbation Problem
Both methods measure feature importance by systematically masking or perturbing input data to see how the output changes. However, this creates a major vulnerability when features are strongly correlated.
Masking one feature while holding its correlated partner fixed creates impossible combinations—like retaining a senior executive's salary while artificially zeroing out their age. This forces the model to evaluate synthetic scenarios far outside its training distribution, resulting in explanations that cannot be trusted.
The Quick Version
- The Black Box: Complex models predict accurately but hide their underlying logic.
- Local Surrogates: LIME fits a simple, understandable line to a complex curve locally.
- Force Plot: SHAP balances individual feature impacts so they sum perfectly to the prediction.
- Failure Mode: Strongly correlated features produce unrealistic synthetic data that breaks the approximation.