Classical Object Detection
Before neural networks, detectors found objects by sliding a small window over the image and asking a classifier at every stop whether something interesting sits inside.
Why Does This Exist?
Say you want to find every pedestrian in a photo. A classifier can answer "person or not" for one cropped patch, but your photo is not one patch. It is hundreds of thousands of possible patches at different positions and sizes, and the pedestrian could be in any of them.
Classical object detection is the machinery built around that problem before deep learning: pick a window size, slide it across the image in small steps, describe each window with hand-made features, and let a classifier vote. Then repeat at every scale, because a pedestrian far away is a small window and one close up is a big one. Everything in this family, from template matching to Haar cascades, Viola-Jones, HOG with SVM, and deformable part models, is an answer to the same cost squeeze: there are far more windows than objects, so most of the work is rejecting empty windows cheaply.
This page maps the family. Each member gets its own page for the mechanism; start from object detection if the task itself (boxes plus labels) is new to you.
Think of It Like This
Checking every seat in a dark theater
You lost your keys in a theater with 3,000 seats and the lights are off. You cannot scan the whole room at once, so you walk row by row with a flashlight, glancing at each seat just long enough to rule it out. Most seats take half a second. The three seats with a jacket on them get a proper search.
That is the sliding-window deal. The flashlight glance is a cheap classifier rejecting sky, road, and wall. The proper search is the expensive classifier reserved for windows that survived. Classical detection research is mostly about making the glance cheaper and better, because the glance runs thousands of times per image and the search runs a handful.
Where the analogy stops: a theater has fixed seats, but objects come in any size, so the detector must re-walk the whole theater once per scale.
How It Actually Works
Every classical detector shares one pipeline with three moving parts.
1. Propose windows
Fix a window size, say pixels for pedestrians, and slide it over the image with a stride of a few pixels. Then shrink the image and repeat, building a pyramid of scales. A photo with stride 8 holds positions across and down at one scale, which is windows, and the pyramid multiplies that several times over. Nearly all of them are background.
2. Describe each window with features
Raw pixels shift with lighting and clothing, so each window is converted into something stabler: pixel comparisons (template matching), light-dark rectangle differences (Haar cascades), edge-direction histograms (HOG features), or a root plus movable parts (deformable part models). The feature is the whole bet: if it cannot tell person from lamppost, no classifier downstream saves you.
3. Classify, then prune
A trained classifier, usually AdaBoost or a linear SVM, scores every window. Because neighboring windows overlap heavily, one pedestrian fires dozens of overlapping boxes, so a final pruning step merges them. That step is non-maximum suppression: keep the strongest box, delete the rest that overlap it too much.
The family in one pass
- Template matching: the oldest idea, correlate a patch against the image. Exact but brittle.
- Haar cascades: rectangle features plus a cascade of cheap rejectors. Fast enough for realtime faces.
- Viola-Jones: the full system, Haar plus AdaBoost plus cascade, that made face detection practical.
- HOG plus SVM: edge-direction histograms that owned pedestrian detection until CNNs.
- Deformable part models: a root filter plus movable part filters, the last classical champion before R-CNN.
Watch Out For
The scale tax
A detector tuned at one size is blind at others, so every scale of the pyramid pays the full window count again. Beginners test on one image size, ship, and then watch small or close-up objects vanish. If your recall changes with object size, the pyramid step, not the classifier, is the first suspect.
Windows are not objects
Sliding windows score rectangles, not things. A window half-covering two pedestrians can outscore a tight window on one, and heavy overlap needs non-maximum suppression to clean up. If your output shows stacked boxes on one person, that is a pruning bug, not a detection triumph.
The Quick Version
- Classical detection slides a fixed window over positions and scales, then classifies each one.
- One image needs 3,285 classifier calls per scale at stride 8, so cheap rejection matters most.
- Hand-made features carry the whole approach: templates, Haar rectangles, HOG histograms, or part filters.
- Overlapping hits get merged by non-maximum suppression into one box per object.
- Deep detectors replaced the hand-made features, but the propose-describe-prune shape survives in them.