Skip to content
AI360Xpert
Beta

Top-Down Pose Estimation

Top-down pose detects each person first, crops them large, and estimates one clean skeleton per crop for maximum accuracy.

Top-down pose estimation frames each person first then solves one skeleton inside each frame.
Top-down pose estimation frames each person first then solves one skeleton inside each frame.

Why Does This Exist?

Bottom-up speed comes with grouping errors, and OpenPose swaps limbs in crowds. Top-down methods (SimpleBaseline, Xiao et al., 2018; HRNet, Sun et al., 2019) buy accuracy with compute: run a person detector, enlarge each box, and feed normalized crops to a single-person head. Each head sees one centred, scale-normalized body, so heatmaps stay sharp even for small figures.

Think of It Like This

Portrait booths at a busy fair

Instead of sketching the crowd from a tower, the fair builds one booth per visitor: frame them, seat them centred, good lighting, full canvas. Each artist draws one sitter well. The cost is booths and artists per visitor, and a missed framing ruins that portrait.

Where it stops: fifty visitors need fifty sittings, so rush hour queues grow linearly.

How It Actually Works

A detector (Faster R-CNN or similar) outputs person boxes, expanded by about 1.25×1.25\times and cropped with aspect ratio 4:34:3. The pose head (ResNet-deconv in SimpleBaseline, HRNet in the accuracy leader) predicts 1717 heatmaps per crop with MSE. Test-time flip averaging and box NMS cleanup follow. Cost scales with person count: 5050 people means 5050 head passes plus detection.

Worked example

Box (x,y,w,h)=(100,50,60,120)(x, y, w, h) = (100, 50, 60, 120). Expanded 1.25×1.25\times about its centre (130,110)(130, 110): new size 75×15075 \times 150, crop (92,35,75,150)(92, 35, 75, 150). Resized to 256×192256 \times 192, the wrist at crop-relative (0.7,0.5)(0.7, 0.5) maps to heatmap (0.7×48,0.5×64)=(33.6,32)(0.7 \times 48, 0.5 \times 64) = (33.6, 32) at 1/41/4 resolution. A 44-pixel detector shift moves the peak by one heatmap cell, which flip averaging halves.

Code

# Crop expansion and heatmap mapping for one box.x, y, w, h, k = 100, 50, 60, 120, 1.25cx, cy = x + w / 2, y + h / 2nw, nh = w * k, h * kprint((round(cx - nw / 2), round(cy - nh / 2), nw, nh))# -> (92, 35, 75.0, 150.0)print((round(0.7 * 48, 1), round(0.5 * 64, 1)))# -> (33.6, 32.0)

Watch Out For

Detector misses capping pose recall

No box means no skeleton, however good the head is. Symptom: high joint precision but missing small or occluded people. Fix: lower detector thresholds for pose pipelines and audit person recall before tuning heatmaps.

Tight crops clipping limbs

Snug boxes cut hands and feet. Symptom: systematic wrist and ankle misses at frame edges. Fix: keep the 1.25×1.25\times expansion and 4:34:3 aspect handling; never crop to raw detector output.

The Quick Version

  • Detect people, expand boxes, estimate one skeleton per crop.
  • Normalized single-person inputs give the best joint accuracy.
  • Compute grows linearly with crowd size.
  • Detector recall caps the whole pipeline.
  • Choose top-down when precision beats frame budget.