Mask2Former Universal Masks
Mask2Former sharpens MaskFormer with masked attention, multi-scale decoding, and stronger matching, topping semantic, instance, and panoptic boards at once.
Why Does This Exist?
MaskFormer proved mask classification unifies the three segmentation tasks but lagged on small objects: global cross-attention wastes query capacity on irrelevant pixels and single-scale decoding blurs detail. Mask2Former (Cheng et al., 2022) fixes both with masked cross-attention (queries attend only inside their current mask forecast), a multi-scale decoder that cycles to features, and improved matching with stronger mask losses. One model then led semantic, instance, and panoptic leaderboards together.
Think of It Like This
Detectives who search only their own beat
MaskFormer detectives each re-interview the whole city per round. Mask2Former hands each detective a beat map from the last round and says search only inside your outline, then rotates beats from precinct overview down to street level. Attention stops wandering, small alleys get visited, and rounds converge faster.
Where it stops: a wrong first-round outline traps its detective on the wrong beat until the outline updates.
How It Actually Works
The pixel decoder builds a to feature pyramid. Each of decoder layers takes the previous mask forecast, binarizes it, and masks cross-attention to those pixels only. Queries then self-attend, update masks, and move to the next scale, coarse to fine. Training keeps Hungarian matching with classification plus binary-mask point-sampled Dice and focal losses (a PointRend-style sampling trick keeps memory sane). Deformable attention in the pixel decoder sharpens small features.
Worked example
Query forecasts a blob on a map: of pixels, so attention touches percent of the map instead of all of it. At the scale a -pixel traffic light spans cells instead of at , and the query finally resolves it. Three rounds move its mask IoU from to as the beat outline tightens.
Code
# Attention saving from masked versus global cross-attention.total, kept = 14400, 900print(f"{100 * kept / total:.1f}% of pixels attended")# -> 6.2% of pixels attendedWatch Out For
Masked attention locking onto early errors
A bad round-one mask blinds later rounds to the true region. Symptom: confident wrong blobs that never recover. Fix: keep learnable query warm-up, deep supervision on intermediate masks, and enough training iterations for outlines to migrate.
Under-provisioning for video and depth extension
Ports to video instance and 3D occupancy raise memory fast with masked attention over frames. Symptom: OOM when copying image settings. Fix: shorten clip length or query count first; do not copy image configs blindly.
The Quick Version
- Masked cross-attention limits each query to its current mask forecast.
- Multi-scale decoding cycles up to for small-object detail.
- Point-sampled mask losses keep high-resolution training affordable.
- One trained model serves semantic, instance, and panoptic inference.
- Early mask errors can trap attention, so deep supervision matters.