Top-Down Pose Estimation
Top-down pose detects each person first, crops them large, and estimates one clean skeleton per crop for maximum accuracy.
Why Does This Exist?
Bottom-up speed comes with grouping errors, and OpenPose swaps limbs in crowds. Top-down methods (SimpleBaseline, Xiao et al., 2018; HRNet, Sun et al., 2019) buy accuracy with compute: run a person detector, enlarge each box, and feed normalized crops to a single-person head. Each head sees one centred, scale-normalized body, so heatmaps stay sharp even for small figures.
Think of It Like This
Portrait booths at a busy fair
Instead of sketching the crowd from a tower, the fair builds one booth per visitor: frame them, seat them centred, good lighting, full canvas. Each artist draws one sitter well. The cost is booths and artists per visitor, and a missed framing ruins that portrait.
Where it stops: fifty visitors need fifty sittings, so rush hour queues grow linearly.
How It Actually Works
A detector (Faster R-CNN or similar) outputs person boxes, expanded by about and cropped with aspect ratio . The pose head (ResNet-deconv in SimpleBaseline, HRNet in the accuracy leader) predicts heatmaps per crop with MSE. Test-time flip averaging and box NMS cleanup follow. Cost scales with person count: people means head passes plus detection.
Worked example
Box . Expanded about its centre : new size , crop . Resized to , the wrist at crop-relative maps to heatmap at resolution. A -pixel detector shift moves the peak by one heatmap cell, which flip averaging halves.
Code
# Crop expansion and heatmap mapping for one box.x, y, w, h, k = 100, 50, 60, 120, 1.25cx, cy = x + w / 2, y + h / 2nw, nh = w * k, h * kprint((round(cx - nw / 2), round(cy - nh / 2), nw, nh))# -> (92, 35, 75.0, 150.0)print((round(0.7 * 48, 1), round(0.5 * 64, 1)))# -> (33.6, 32.0)Watch Out For
Detector misses capping pose recall
No box means no skeleton, however good the head is. Symptom: high joint precision but missing small or occluded people. Fix: lower detector thresholds for pose pipelines and audit person recall before tuning heatmaps.
Tight crops clipping limbs
Snug boxes cut hands and feet. Symptom: systematic wrist and ankle misses at frame edges. Fix: keep the expansion and aspect handling; never crop to raw detector output.
The Quick Version
- Detect people, expand boxes, estimate one skeleton per crop.
- Normalized single-person inputs give the best joint accuracy.
- Compute grows linearly with crowd size.
- Detector recall caps the whole pipeline.
- Choose top-down when precision beats frame budget.