3D Human Pose Lifting
3D human pose lifts flat 2D skeletons into depth-aware joint positions, resolving the missing camera-axis coordinate with priors or multiple views.
Why Does This Exist?
2D skeletons flatten the punch moving toward the camera into a short jab: the depth axis is gone. 3D pose recovers per joint in camera or root-relative space for animation, biomechanics, and robotics. Monocular lifting is ambiguous (many 3D poses project to one 2D skeleton), so methods add temporal motion (VideoPose3D, Pavllo et al., 2019), body-shape priors (SMPL mesh recovery, Kanazawa et al., 2018), or multi-view geometry.
Think of It Like This
Rebuilding a statue from its shadow
A lamp casts a dancer's shadow on a wall. Many arm positions throw the same silhouette. Watching the shadow move over seconds rules out statues that would need teleporting elbows, and knowing human bone lengths rules out rubber arms. 3D lifting is shadow plus motion plus anatomy.
Where it stops: a single still shadow keeps infinite statue candidates, so monocular single-frame depth stays a guess.
How It Actually Works
Two-stage lifting takes trusted 2D joints and regresses root-relative 3D with temporal convolutions over frames, which disambiguates depth through motion parallax. Mesh methods regress SMPL pose and shape parameters and project joints through the body model. Direct volumetric heads classify depth bins per joint. Evaluation uses MPJPE (mean per-joint position error in millimetres) and PA-MPJPE after Procrustes alignment; mm is a strong monocular score on Human3.6M.
Worked example
Wrist 2D at with root at and estimated depth mm toward the camera. Lifting outputs root-relative in scaled units. A mm error on each axis gives joint error mm. Averaged over joints, that is the MPJPE contribution of one wrist.
Code
import math
# MPJPE contribution of one joint with 10mm error per axis.print(round(math.sqrt(10 ** 2 + 10 ** 2 + 10 ** 2), 1))# -> 17.3Watch Out For
Trusting monocular depth on single frames
Single-image is underdetermined. Symptom: plausible skeletons with wildly wrong reach. Fix: use video lifting or multi-view fusion for measurement tasks; treat single-frame depth as visualization only.
Comparing MPJPE across alignment protocols
PA-MPJPE after alignment reads much lower than raw MPJPE. Symptom: method A looks better only because of friendlier alignment. Fix: compare identical protocols and report both numbers.
The Quick Version
- 3D pose adds the depth axis that 2D projection destroys.
- Temporal lifting, SMPL priors, and multi-view geometry tame the ambiguity.
- MPJPE in millimetres is the standard error; PA-MPJPE aligns first.
- Single-frame monocular depth is a guess, video is a measurement.
- Lift from trusted 2D joints rather than raw pixels when data is scarce.