PointRend Mask Refinement
PointRend renders coarse masks first, then spends extra compute only on uncertain edge points instead of the whole high-resolution grid.
Why Does This Exist?
Uniform high-resolution masks waste compute: interiors are already decided, only edges are uncertain, yet dense heads reconvolve millions of settled pixels. PointRend (Kirillov et al., 2020) treats segmentation like adaptive rendering. Start from a coarse or prediction, pick the points with probabilities nearest , and reclassify only those with a small MLP fed by fine FPN features plus the coarse score. Repeat subdivision until edges are crisp.
Think of It Like This
A mapmaker who inks only the coastline
A mapmaker sketches continents roughly, then ignores the settled interiors and walks only the blurry coastline with a fine pen, checking the terrain at each uncertain step. The ocean and the deep inland never get revisited. Compute goes to the shoreline where the map is actually wrong.
Where it stops: if the coarse sketch places the continent wrongly, no coastline walk fixes it.
How It Actually Works
Each subdivision upsamples the current mask with bilinear interpolation, scores uncertainty as closeness, and samples the top points per box. A -layer MLP reads each point's fine feature vector plus its coarse prediction and outputs a refined logit. Training samples a mix of uniform and uncertain points so the MLP sees both. At inference five subdivisions lift to while touching a fraction of the pixels.
Worked example
A coarse mask has cells. Uncertainty picks edge points. The MLP reclassifies those instead of all , a saving per step. A bicycle spoke pixel at coarse score receives fine stride- features showing metal texture and flips to foreground, while a settled road pixel at is never revisited.
Code
# Uncertainty sampling: pick points nearest 0.5.scores = [0.03, 0.52, 0.91, 0.48, 0.60]ranked = sorted(scores, key=lambda p: abs(p - 0.5))print(ranked[:3])# -> [0.52, 0.48, 0.6]Watch Out For
Sampling too few points on lace structures
Thin spokes and wires need dense coverage; tuned for people misses them. Symptom: broken spokes despite sharp person edges. Fix: raise per subdivision and verify small-object boundary recall separately.
Feeding the MLP coarse-only features
Without fine FPN vectors the MLP just repeats the coarse vote. Symptom: subdivisions change nothing. Fix: concatenate stride- features with the coarse score at every sampled point.
The Quick Version
- PointRend subdivides coarse masks and reclassifies only uncertain points.
- A tiny MLP fuses fine features with the coarse score per point.
- Five steps can lift to at a fraction of dense cost.
- interiors are never recomputed; edges get the budget.
- Coarse errors and starved point budgets are the two failure modes.