Deformable Part Models
Deformable part models score an object as a whole plus its movable parts, so a person reads as a torso with arms and legs allowed to be somewhere nearby, not glued in place.
Why Does This Exist?
Rigid templates like HOG features describe one fixed silhouette. That works for upright pedestrians but breaks the moment arms swing, legs stride, or a cat curls up: the silhouette changes while the object stays the same. You could train one template per pose, but poses are endless.
Deformable part models (DPM), developed by Felzenszwalb, Girshick, McAllester, and Ramanan, split the difference. One coarse root filter scores the whole object, several fine part filters score pieces like head and limbs, and each part may sit away from its ideal spot at a price. A striding pedestrian still scores high because the parts match well and only pay small displacement penalties. DPM won the PASCAL VOC detection challenges from 2007 to 2011 and was the last classical method standing before R-CNN took over in 2012.
Think of It Like This
A police sketch with movable features
A witness describes a suspect: round face, and a scar that sits near the left cheek, give or take. The detective does not demand the scar at exact coordinates. A scar two centimeters off is still a strong match with a small doubt discount; a scar on the forehead barely counts.
The root filter is the round face. The part filters are the scar and other features. The doubt discount is the deformation cost: quadratic in how far the part drifted from its anchor. Total score is whole plus parts minus drift, and the suspect with everything slightly shifted still gets arrested.
How It Actually Works
1. Root plus parts
The model stores a coarse HOG-like root template for the whole window and higher-resolution part templates (typically 6 to 8), each with an anchor position relative to the root. At test time each part is allowed to displace to the position where it matches best.
2. Score minus deformation
The detection score at a root placement is:
is the root filter response, is part 's best response, is how far part moved from its anchor, holds displacement features , and holds the learned deformation weights penalizing drift. Worked small: root , two parts scoring and after moving, with deformation costs and : total . A rigid template would have scored the shifted limbs near zero instead.
3. Latent training
Part positions are not labeled in training data, so training treats them as hidden (latent) variables: alternate between finding each part's best placement and retraining the filters with those placements frozen, using a latent SVM. Mixture components (separate root-plus-parts for front, side, and seated views) handle viewpoint on top of deformation.
Watch Out For
Parts drift onto background texture
Weakly supervised parts sometimes anchor onto background that correlates with the class, like wheels locking onto road stripes instead of the car. The detector then fails on any new background. If accuracy collapses outside the training scene style, inspect where the parts actually fire before blaming the classifier.
Slow by modern standards
Each part searches a neighborhood at double resolution, so DPM runs seconds per image while YOLO runs in milliseconds. DPM is worth studying for the part-deformation idea, which survives inside modern pose and detection models, not for shipping new detectors.
The Quick Version
- DPM scores root plus movable parts minus a quadratic displacement penalty.
- Root with parts and and costs , totals even with shifted limbs.
- Part positions train as latent variables since datasets never label them.
- Mixture components add viewpoint handling on top of deformation.
- VOC champion 2007 to 2011, then displaced by R-CNN; the deformation idea lives on.