Skip to content
AI360Xpert
Beta

MaskFormer Mask Classification

MaskFormer stops labelling pixels and starts proposing masks: a transformer predicts one mask plus one class per query, then pixels join the best mask.

MaskFormer lets each transformer query claim one region and one class, then assigns pixels by mask membership.
MaskFormer lets each transformer query claim one region and one class, then assigns pixels by mask membership.

Why Does This Exist?

Per-pixel classifiers decide each pixel alone and stitch instances afterwards with boxes or centres, which is why Panoptic FPN and Panoptic-DeepLab need merger logic. MaskFormer (Cheng et al., 2021) flips the task: predict N≈100N \approx 100 binary masks with one class label each, match them to ground truth with bipartite matching, and let pixels belong to masks. One architecture serves semantic, instance, and panoptic tasks with only the inference rule changing.

Think of It Like This

Casting a play instead of auditioning extras one by one

Old casting interviews every extra separately and assembles crowds later. MaskFormer holds 100100 role auditions: each role (query) claims a costume region and a character name together. Extras join the role whose costume fits. Recasting for a tragedy or a comedy uses the same stage with different final assignments.

Where it stops: 100100 roles cap crowded scenes, and unused roles must learn to claim cleanly nothing.

How It Actually Works

A backbone plus pixel decoder builds high-resolution per-pixel embeddings. A transformer decoder refines NN queries with masked cross-attention into pixel features. Each query outputs a class distribution (plus no-object) and a mask embedding; the binary mask is the dot product of that embedding with every pixel embedding, thresholded. Training matches predictions to truth with Hungarian matching on class plus mask cost, then applies cross-entropy plus Dice plus focal mask loss. Semantic inference marginalizes query masks per class; panoptic inference assigns each pixel to its highest-confidence query.

Worked example

100100 queries on an image with 77 true segments. Hungarian matching pairs 77 queries to truth; 9393 match no-object. Query 1212 predicts person with confidence 0.90.9 and a mask covering 1,2001{,}200 pixels at IoU 0.80.8. A pixel inside with mask score 2.02.0 (sigmoid 0.880.88) joins query 1212 as person. A pixel claimed by two queries goes to the higher class×maskclass \times mask product, no rulebook needed.

Code

# Pixel assignment between two competing queries.cands = [("q12 person", 0.9 * 0.88), ("q31 car", 0.6 * 0.75)]print(max(cands, key=lambda c: c[1])[0])# -> q12 person

Watch Out For

Too few queries for dense scenes

100100 queries undercount datasets with 150150 instances. Symptom: merged small objects. Fix: raise NN for instance-heavy data and confirm query count exceeds the max instances per image.

Judging masks by pixel accuracy

Mask-classification models trade per-pixel calibration for region coherence. Symptom: good PQ, odd log-loss. Fix: evaluate with PQ and mask AP, not cross-entropy alone.

The Quick Version

  • MaskFormer predicts 100100 masks with one class each instead of labelling pixels.
  • Hungarian matching pairs predictions to truth during training.
  • Same model serves semantic, instance, and panoptic via different inference rules.
  • Masked cross-attention keeps queries focused on their regions.
  • Query count must exceed the densest scene's instance count.