Skip to content
AI360Xpert
Beta

3D Human Pose Lifting

3D human pose lifts flat 2D skeletons into depth-aware joint positions, resolving the missing camera-axis coordinate with priors or multiple views.

3D pose estimation adds the missing depth axis to a flat skeleton using motion and body-shape priors.
3D pose estimation adds the missing depth axis to a flat skeleton using motion and body-shape priors.

Why Does This Exist?

2D skeletons flatten the punch moving toward the camera into a short jab: the depth axis is gone. 3D pose recovers (x,y,z)(x, y, z) per joint in camera or root-relative space for animation, biomechanics, and robotics. Monocular lifting is ambiguous (many 3D poses project to one 2D skeleton), so methods add temporal motion (VideoPose3D, Pavllo et al., 2019), body-shape priors (SMPL mesh recovery, Kanazawa et al., 2018), or multi-view geometry.

Think of It Like This

Rebuilding a statue from its shadow

A lamp casts a dancer's shadow on a wall. Many arm positions throw the same silhouette. Watching the shadow move over seconds rules out statues that would need teleporting elbows, and knowing human bone lengths rules out rubber arms. 3D lifting is shadow plus motion plus anatomy.

Where it stops: a single still shadow keeps infinite statue candidates, so monocular single-frame depth stays a guess.

How It Actually Works

Two-stage lifting takes trusted 2D joints and regresses root-relative 3D with temporal convolutions over ≈243\approx 243 frames, which disambiguates depth through motion parallax. Mesh methods regress SMPL pose and shape parameters and project joints through the body model. Direct volumetric heads classify depth bins per joint. Evaluation uses MPJPE (mean per-joint position error in millimetres) and PA-MPJPE after Procrustes alignment; 5050 mm is a strong monocular score on Human3.6M.

Worked example

Wrist 2D at (100,200)(100, 200) with root at (120,210)(120, 210) and estimated depth +300+300 mm toward the camera. Lifting outputs root-relative (x,y,z)=(−20,−10,300)(x, y, z) = (-20, -10, 300) in scaled units. A 1010 mm error on each axis gives joint error 100+100+100≈17.3\sqrt{100 + 100 + 100} \approx 17.3 mm. Averaged over 1717 joints, that is the MPJPE contribution of one wrist.

Code

import math
# MPJPE contribution of one joint with 10mm error per axis.print(round(math.sqrt(10 ** 2 + 10 ** 2 + 10 ** 2), 1))# -> 17.3

Watch Out For

Trusting monocular depth on single frames

Single-image zz is underdetermined. Symptom: plausible skeletons with wildly wrong reach. Fix: use video lifting or multi-view fusion for measurement tasks; treat single-frame depth as visualization only.

Comparing MPJPE across alignment protocols

PA-MPJPE after alignment reads much lower than raw MPJPE. Symptom: method A looks better only because of friendlier alignment. Fix: compare identical protocols and report both numbers.

The Quick Version

  • 3D pose adds the depth axis that 2D projection destroys.
  • Temporal lifting, SMPL priors, and multi-view geometry tame the ambiguity.
  • MPJPE in millimetres is the standard error; PA-MPJPE aligns first.
  • Single-frame monocular depth is a guess, video is a measurement.
  • Lift from trusted 2D joints rather than raw pixels when data is scarce.