2D Human Pose Basics
2D human pose finds body joints as image points and links them into a skeleton, turning pixels into a measurable stick figure.
Why Does This Exist?
Pose estimation promises action understanding, but the base unit is the 2D skeleton: seventeen joints in COCO format (nose, eyes, ears, shoulders, elbows, wrists, hips, knees, ankles) that downstream tasks count, track, and classify. Direct coordinate regression proved brittle, so the field predicts one heatmap per joint and reads the peak. This page is the 2D parent: representation, heatmaps, and metrics. Children cover OpenPose, top-down, bottom-up, and 3D lifting.
Think of It Like This
Pushpins and string on a corkboard photo
Pin a photo to corkboard, push one pin per joint through the print, and tie string along limbs. The pins are heatmap peaks; the string is the skeleton topology. Measuring stride length means measuring pin distances, not re-reading the photo. Occluded ankles are pins you place by guess with low confidence.
Where it stops: pins flatten depth, so a punch toward the camera looks like a short jab.
How It Actually Works
Networks output heatmaps at input resolution trained with MSE against Gaussian blobs ( pixels) centred on each joint. The predicted joint is the argmax, refined by a quarter-pixel shift toward the second-highest neighbour. Object Keypoint Similarity (OKS) scores each joint as with per-joint falloff and scale , averaged into AP over thresholds. PCK counts joints within a fraction of torso size as correct.
Worked example
Wrist ground truth , prediction , scale , wrist . Distance . OKS . A hip with larger forgives more: same distance scores . Threshold keeps both as correct.
Code
import math
# OKS for one joint.d, s, k = 5.0, 200.0, 0.089print(round(math.exp(-(d ** 2) / (2 * s ** 2 * k ** 2)), 3))# -> 0.961Watch Out For
Argmax quantization on small heatmaps
-resolution argmax snaps joints to a -pixel grid. Symptom: jittery tracking on video. Fix: apply the quarter-pixel refinement and consider DARK-style distribution-aware decoding.
Left-right swaps on symmetric poses
Mirror-symmetric joints confuse appearance-only heads. Symptom: crossed wrists on push-ups. Fix: train with flip augmentation plus visibility flags, and check swap rate separately from OKS.
The Quick Version
- 2D pose outputs seventeen heatmaps, one per COCO joint, and reads peaks.
- Gaussian targets with MSE train far better than direct coordinate regression.
- OKS weights errors by joint type and person scale into AP scores.
- Quarter-pixel refinement removes heatmap grid snap.
- This page parents OpenPose, top-down, bottom-up, and 3D children.