Haar Cascades
Haar cascades detect objects with chains of dirt-cheap light-versus-dark rectangle tests that throw away boring windows in a few steps and inspect only the survivors.
Why Does This Exist?
Template matching compares raw pixels and dies on lighting changes. Classical object detection needs a feature that survives lighting and runs fast enough to test thousands of windows per image. Faces gave the perfect target: eyes are darker than the forehead and the nose bridge is brighter than the nostrils, in almost any light.
Haar-like features turn those observations into rectangle arithmetic: sum the pixels in the dark rectangle, subtract the light one, and you have one number describing one contrast. One test is weak, but thousands of them, arranged so the cheap ones run first, made realtime face detection possible on 2001 hardware. The full system built on them is Viola-Jones; this page covers the features and the cascade trick.
Think of It Like This
Airport security lanes
Imagine an airport where every passenger first walks past a camera that checks one thing: are you carrying anything metallic? Almost everyone passes in one second. Only the beeping few go to the bag scanner, and only the suspicious remainder get the full pat-down.
Haar cascades are those lanes. Early stages are the metal detector: two rectangle tests that reject obvious non-faces instantly. Later stages are the pat-down: hundreds of tests reserved for windows that already look face-like. Average cost stays tiny because almost nobody reaches the pat-down.
Where it stops: real security adapts to new threats, but a trained cascade is frozen. A contrast pattern it never saw in training sails through or gets stopped wrongly, every time.
How It Actually Works
1. Haar-like features
A feature is two to four adjacent rectangles with / weights. Its value is the weighted sum of pixel intensities: . Edge features catch boundaries (dark-light), line features catch bars (light-dark-light), and four-rectangle features catch diagonals. They measure local contrast, not absolute brightness, so a face in shadow and one in sunlight give similar values.
2. The integral image makes them free
Summing rectangles naively costs one add per pixel. The integral image stores at each point the total of everything above and left of it, and then any rectangle sum is four lookups: for corners (top-left), , , (bottom-right). Worked small: for the image holding through in reading order, the top-left block sums to , and the integral image returns that 12 with four array reads no matter how big the rectangle grows. Every Haar feature therefore costs constant time.
3. The cascade spends effort where it matters
AdaBoost picks the best features and builds each stage to have near-perfect recall with modest precision: a stage may keep of faces while letting of background through. Chain twenty such stages and background survival is , effectively zero, while faces survive every stage. A window that fails stage 3 never pays for stages 4 through 20, so the average window costs a handful of features instead of thousands.
Watch Out For
Frontal, upright, lit: pick two out of three
Haar cascades memorize the contrasts of their training pose. Profile faces, tilted heads, and heavy occlusion break the exact rectangle patterns, and detection collapses rather than degrading gracefully. If your camera angle or subject pose varies, this is the wrong tool; modern detectors in object detection architectures handle pose far better.
minNeighbors is doing real work
The OpenCV detector fires overlapping boxes around each hit, and minNeighbors sets how many overlapping boxes a hit needs to survive. Too low and every curtain fold is a face; too high and real faces vanish. Tune it on your own footage, because the default was tuned on someone else's.
The Quick Version
- Haar features are weighted rectangle differences measuring local contrast, robust to lighting.
- The integral image reduces any rectangle sum to four array lookups, so features cost constant time.
- A block of a -to- image sums to 12 in four reads, at any scale.
- Cascade stages keep all faces and reject much background each, so average cost is a few features.
- Frozen to training pose: profile and tilted faces fail, which is why deep detectors replaced cascades.