Skip to content
AI360Xpert
Beta

Panoptic FPN Heads

Panoptic FPN runs an instance branch and a semantic branch on one shared pyramid, then merges them with a rule that gives things priority over stuff.

Panoptic FPN predicts thing instances and stuff regions from one backbone then merges them into a single non-overlapping scene map.
Panoptic FPN predicts thing instances and stuff regions from one backbone then merges them into a single non-overlapping scene map.

Why Does This Exist?

Panoptic segmentation demands one label plus one instance ID per pixel, but the two source tasks disagree: the instance branch overlaps boxes while the semantic branch smears classes. Panoptic FPN (Kirillov et al., 2019) made the baseline delightfully boring: one ResNet-FPN backbone, a Mask R-CNN-style instance head for countable things, a small dense head for stuff and things, and a deterministic fusion that resolves every conflict the same way.

Think of It Like This

Two survey teams with one referee book

Two teams map a park: one tags every countable bench and dog, the other colours grass and paths wall to wall. Their maps overlap on bench legs over grass. The referee book says tagged things outrank coloured stuff, higher-confidence things outrank lower ones, and leftover pixels take the stuff colour. No negotiation, same ruling every park.

Where it stops: rigid rules cannot learn that a low-confidence child behind a high-confidence pole still owns those pixels.

How It Actually Works

The shared FPN feeds both heads. The semantic head upsamples 1/321/32 through 1/41/4 features with grouped convolutions to a dense KK-class map. The instance head produces boxes, classes, and 28×2828 \times 28 masks. Fusion sorts thing instances by score, pastes them over the semantic map, deletes stuff pixels claimed by things, resolves thing-on-thing overlap by score order, and fills the rest from stuff predictions. Panoptic Quality PQ=SQ×RQPQ = SQ \times RQ scores the result: segmentation quality times recognition F1 over matched segments.

Worked example

Pixel claimed by person (score 0.90.9), car (score 0.70.7), and stuff road. Fusion keeps person. A second pixel claimed only by stuff sky stays sky. With 33 matched segments at IoUs 0.80.8, 0.70.7, 0.90.9: SQ=0.8SQ = 0.8. With 33 true positives, 11 false positive, 00 misses: RQ=3/(3+0.5)≈0.857RQ = 3 / (3 + 0.5) \approx 0.857. PQ≈0.686PQ \approx 0.686.

Code

# Panoptic Quality from segmentation and recognition quality.sq, tp, fp, fn = 0.8, 3, 1, 0rq = tp / (tp + 0.5 * fp + 0.5 * fn)print(round(sq * rq, 3))# -> 0.686

Watch Out For

Overlapping training masks breaking fusion

Panoptic training expects non-overlapping ground truth. Symptom: the model learns to double-claim pixels and PQ stalls. Fix: preprocess annotations to strict panoptic format with one label per pixel before training.

Stuff-first fusion ordering

Pasting stuff over things erases small instances under road paint. Symptom: vanishing pedestrians on crossings. Fix: keep the canonical order: high-score things first, stuff only for unclaimed pixels.

The Quick Version

  • One FPN backbone feeds an instance head and a dense semantic head.
  • Fusion pastes things by score over stuff and resolves overlaps deterministically.
  • Every pixel ends with exactly one class and one instance ID.
  • PQ multiplies segmentation quality by recognition F1.
  • Simple and debuggable, but hand rules cap accuracy versus learned mergers.