MaskFormer Mask Classification
MaskFormer stops labelling pixels and starts proposing masks: a transformer predicts one mask plus one class per query, then pixels join the best mask.
Why Does This Exist?
Per-pixel classifiers decide each pixel alone and stitch instances afterwards with boxes or centres, which is why Panoptic FPN and Panoptic-DeepLab need merger logic. MaskFormer (Cheng et al., 2021) flips the task: predict binary masks with one class label each, match them to ground truth with bipartite matching, and let pixels belong to masks. One architecture serves semantic, instance, and panoptic tasks with only the inference rule changing.
Think of It Like This
Casting a play instead of auditioning extras one by one
Old casting interviews every extra separately and assembles crowds later. MaskFormer holds role auditions: each role (query) claims a costume region and a character name together. Extras join the role whose costume fits. Recasting for a tragedy or a comedy uses the same stage with different final assignments.
Where it stops: roles cap crowded scenes, and unused roles must learn to claim cleanly nothing.
How It Actually Works
A backbone plus pixel decoder builds high-resolution per-pixel embeddings. A transformer decoder refines queries with masked cross-attention into pixel features. Each query outputs a class distribution (plus no-object) and a mask embedding; the binary mask is the dot product of that embedding with every pixel embedding, thresholded. Training matches predictions to truth with Hungarian matching on class plus mask cost, then applies cross-entropy plus Dice plus focal mask loss. Semantic inference marginalizes query masks per class; panoptic inference assigns each pixel to its highest-confidence query.
Worked example
queries on an image with true segments. Hungarian matching pairs queries to truth; match no-object. Query predicts person with confidence and a mask covering pixels at IoU . A pixel inside with mask score (sigmoid ) joins query as person. A pixel claimed by two queries goes to the higher product, no rulebook needed.
Code
# Pixel assignment between two competing queries.cands = [("q12 person", 0.9 * 0.88), ("q31 car", 0.6 * 0.75)]print(max(cands, key=lambda c: c[1])[0])# -> q12 personWatch Out For
Too few queries for dense scenes
queries undercount datasets with instances. Symptom: merged small objects. Fix: raise for instance-heavy data and confirm query count exceeds the max instances per image.
Judging masks by pixel accuracy
Mask-classification models trade per-pixel calibration for region coherence. Symptom: good PQ, odd log-loss. Fix: evaluate with PQ and mask AP, not cross-entropy alone.
The Quick Version
- MaskFormer predicts masks with one class each instead of labelling pixels.
- Hungarian matching pairs predictions to truth during training.
- Same model serves semantic, instance, and panoptic via different inference rules.
- Masked cross-attention keeps queries focused on their regions.
- Query count must exceed the densest scene's instance count.