Skip to content
AI360Xpert
Beta

Bottom-Up Pose Estimation

Bottom-up pose finds every joint in the image once, then groups joints into people with tags or fields, keeping runtime flat in crowds.

Bottom-up pose estimation spots all joints first then clusters matching tags into individual skeletons.
Bottom-up pose estimation spots all joints first then clusters matching tags into individual skeletons.

Why Does This Exist?

Top-down accuracy costs one network pass per person, which breaks live crowd analytics at fifty people. Bottom-up methods (Associative Embedding, Newell et al., 2017; HigherHRNet, Cheng et al., 2020) run the heavy network once: detect all joints of each type image-wide, predict a grouping tag per joint, and cluster tags into skeletons. Runtime stays nearly flat as crowds grow, at some accuracy cost on overlap.

Think of It Like This

Name tags at a reunion mixer

Instead of photographing guests one by one, the host hands every left wrist a colour-coded sticker at the door in a single pass. Afterwards anyone can cluster matching stickers into families without re-photographing. Twins with swapped stickers merge families, which is exactly the grouping failure.

Where it stops: sticker printing (tag learning) is subtle work, and near-identical stickers split or merge wrongly.

How It Actually Works

The network outputs 1717 joint heatmaps plus 1717 tag maps (one scalar per joint) or 2D grouping vectors. Detections come from heatmap NMS; joints whose tags fall within a learned radius join one person. Training pulls same-person tags together and pushes different-person tags apart with a push-pull loss alongside heatmap MSE. HigherHRNet adds a transposed-convolution branch so small people get higher-resolution heatmaps without a detector.

Worked example

Three left-wrist detections with tags 1.01.0, 1.11.1, 5.25.2 and radius 0.50.5. First two cluster (gap 0.10.1), third starts a new person (gap 4.14.1). A fourth detection at tag 1.61.6 sits 0.50.5 from the cluster edge: borderline, and the grouping threshold decides whether the crowd counts two people or three.

Code

# Tag clustering with radius 0.5.tags, radius, clusters = [1.0, 1.1, 5.2, 1.6], 0.5, []for t in tags:    placed = False    for c in clusters:        if abs(t - sum(c) / len(c)) <= radius:            c.append(t)            placed = True            break    if not placed:        clusters.append([t])print(clusters)# -> [[1.0, 1.1], [5.2], [1.6]]

Watch Out For

Tag collapse on overlapping bodies

Wrestling and hugging drive tags together. Symptom: merged skeletons exactly where top-down would separate them. Fix: raise input resolution, verify crowd-slice AP separately, and fall back to top-down for contact sports.

Tuning grouping radius on empty rooms

A radius tuned on sparse scenes shatters crowds. Symptom: one person split into two skeletons at events. Fix: tune the radius on crowded validation data, not on studio captures.

The Quick Version

  • One network pass finds all joints plus grouping tags or fields.
  • Tags cluster into skeletons with near-flat runtime in crowds.
  • HigherHRNet adds resolution for small people without a detector.
  • Overlap and tag noise are the accuracy price.
  • Choose bottom-up for live crowds, top-down for precision work.