Skip to content
AI360Xpert
Beta

PointRend Mask Refinement

PointRend renders coarse masks first, then spends extra compute only on uncertain edge points instead of the whole high-resolution grid.

PointRend finds the most uncertain pixels on a coarse mask and re-examines just those points with fine-grained features.
PointRend finds the most uncertain pixels on a coarse mask and re-examines just those points with fine-grained features.

Why Does This Exist?

Uniform high-resolution masks waste compute: interiors are already decided, only edges are uncertain, yet dense heads reconvolve millions of settled pixels. PointRend (Kirillov et al., 2020) treats segmentation like adaptive rendering. Start from a coarse 7×77 \times 7 or 28×2828 \times 28 prediction, pick the NN points with probabilities nearest 0.50.5, and reclassify only those with a small MLP fed by fine FPN features plus the coarse score. Repeat subdivision until edges are crisp.

Think of It Like This

A mapmaker who inks only the coastline

A mapmaker sketches continents roughly, then ignores the settled interiors and walks only the blurry coastline with a fine pen, checking the terrain at each uncertain step. The ocean and the deep inland never get revisited. Compute goes to the shoreline where the map is actually wrong.

Where it stops: if the coarse sketch places the continent wrongly, no coastline walk fixes it.

How It Actually Works

Each subdivision upsamples the current mask 2×2\times with bilinear interpolation, scores uncertainty as ∣p−0.5∣|p - 0.5| closeness, and samples the top N≈784N \approx 784 points per box. A 33-layer MLP reads each point's fine feature vector plus its coarse prediction and outputs a refined logit. Training samples a mix of uniform and uncertain points so the MLP sees both. At inference five subdivisions lift 7×77 \times 7 to 224×224224 \times 224 while touching a fraction of the pixels.

Worked example

A 28×2828 \times 28 coarse mask has 784784 cells. Uncertainty picks 200200 edge points. The MLP reclassifies those 200200 instead of all 784784, a 4×4\times saving per step. A bicycle spoke pixel at coarse score 0.520.52 receives fine stride-44 features showing metal texture and flips to 0.910.91 foreground, while a settled road pixel at 0.030.03 is never revisited.

Code

# Uncertainty sampling: pick points nearest 0.5.scores = [0.03, 0.52, 0.91, 0.48, 0.60]ranked = sorted(scores, key=lambda p: abs(p - 0.5))print(ranked[:3])# -> [0.52, 0.48, 0.6]

Watch Out For

Sampling too few points on lace structures

Thin spokes and wires need dense coverage; NN tuned for people misses them. Symptom: broken spokes despite sharp person edges. Fix: raise NN per subdivision and verify small-object boundary recall separately.

Feeding the MLP coarse-only features

Without fine FPN vectors the MLP just repeats the coarse vote. Symptom: subdivisions change nothing. Fix: concatenate stride-44 features with the coarse score at every sampled point.

The Quick Version

  • PointRend subdivides coarse masks and reclassifies only uncertain points.
  • A tiny MLP fuses fine features with the coarse score per point.
  • Five steps can lift 7×77 \times 7 to 224×224224 \times 224 at a fraction of dense cost.
  • interiors are never recomputed; edges get the budget.
  • Coarse errors and starved point budgets are the two failure modes.