Skip to content
AI360Xpert
Beta

Mask2Former Universal Masks

Mask2Former sharpens MaskFormer with masked attention, multi-scale decoding, and stronger matching, topping semantic, instance, and panoptic boards at once.

Mask2Former restricts each query to its current mask region while decoding from coarse to fine feature scales.
Mask2Former restricts each query to its current mask region while decoding from coarse to fine feature scales.

Why Does This Exist?

MaskFormer proved mask classification unifies the three segmentation tasks but lagged on small objects: global cross-attention wastes query capacity on irrelevant pixels and single-scale decoding blurs detail. Mask2Former (Cheng et al., 2022) fixes both with masked cross-attention (queries attend only inside their current mask forecast), a multi-scale decoder that cycles 1/321/32 to 1/81/8 features, and improved matching with stronger mask losses. One model then led semantic, instance, and panoptic leaderboards together.

Think of It Like This

Detectives who search only their own beat

MaskFormer detectives each re-interview the whole city per round. Mask2Former hands each detective a beat map from the last round and says search only inside your outline, then rotates beats from precinct overview down to street level. Attention stops wandering, small alleys get visited, and rounds converge faster.

Where it stops: a wrong first-round outline traps its detective on the wrong beat until the outline updates.

How It Actually Works

The pixel decoder builds a 1/81/8 to 1/321/32 feature pyramid. Each of 99 decoder layers takes the previous mask forecast, binarizes it, and masks cross-attention to those pixels only. Queries then self-attend, update masks, and move to the next scale, coarse to fine. Training keeps Hungarian matching with classification plus binary-mask point-sampled Dice and focal losses (a PointRend-style sampling trick keeps memory sane). Deformable attention in the pixel decoder sharpens small features.

Worked example

Query 77 forecasts a 30×3030 \times 30 blob on a 120×120120 \times 120 map: 900900 of 14,40014{,}400 pixels, so attention touches 66 percent of the map instead of all of it. At the 1/81/8 scale a 66-pixel traffic light spans 66 cells instead of 22 at 1/321/32, and the query finally resolves it. Three rounds move its mask IoU from 0.40.4 to 0.750.75 as the beat outline tightens.

Code

# Attention saving from masked versus global cross-attention.total, kept = 14400, 900print(f"{100 * kept / total:.1f}% of pixels attended")# -> 6.2% of pixels attended

Watch Out For

Masked attention locking onto early errors

A bad round-one mask blinds later rounds to the true region. Symptom: confident wrong blobs that never recover. Fix: keep learnable query warm-up, deep supervision on intermediate masks, and enough training iterations for outlines to migrate.

Under-provisioning for video and depth extension

Ports to video instance and 3D occupancy raise memory fast with masked attention over frames. Symptom: OOM when copying image settings. Fix: shorten clip length or query count first; do not copy image configs blindly.

The Quick Version

  • Masked cross-attention limits each query to its current mask forecast.
  • Multi-scale decoding cycles 1/321/32 up to 1/81/8 for small-object detail.
  • Point-sampled mask losses keep high-resolution training affordable.
  • One trained model serves semantic, instance, and panoptic inference.
  • Early mask errors can trap attention, so deep supervision matters.