Skip to content
AI360Xpert
Beta

SOLO and SOLOv2 Masks

SOLO turns instance segmentation into grid-cell classification: each cell predicts the mask of the object whose centre falls inside it.

SOLO divides the image into a grid and lets the cell containing each object centre output that whole instance mask.
SOLO divides the image into a grid and lets the cell containing each object centre output that whole instance mask.

Why Does This Exist?

Boxes are an awkward middleman for masks: they need anchors, NMS tuning, and RoI alignment, all inherited from detection. SOLO (Wang et al., 2020) drops them. Divide the image into an S×SS \times S grid; cell (i,j)(i, j) predicts the class of whatever object centres there plus its full-image mask. Two people side by side fall in different cells and separate naturally. SOLOv2 replaces static per-cell channels with dynamic kernels predicted per cell, sharpening masks and adding Matrix NMS.

Think of It Like This

Assigned seats with portrait duties

A classroom grid assigns every student a seat. The rule: whoever sits in a chair paints the full portrait of the classmate whose centre of mass is over that chair. Neighbours never fight over one canvas because centres fall in exactly one seat. SOLOv2 upgrades each painter from a fixed stencil to a custom brush mixed for that sitter.

Where it stops: twins sharing one chair (two centres in one cell) still collide, so grids must be fine enough.

How It Actually Works

FPN levels use grids like 1212 to 4040 per side matched to object scale. Each cell outputs CC class scores and an H×WH \times W mask (SOLO) or a 1×11 \times 1 kernel convolved over mask features (SOLOv2). Centre sampling assigns positives; Dice plus focal losses train masks. Matrix NMS decays duplicate scores in one parallel step using mask IoU instead of sequential box suppression.

Worked example

20×2020 \times 20 grid on a 640×640640 \times 640 image: each cell spans 3232 pixels. Person A centres at (100,200)(100, 200), cell (3,6)(3, 6). Person B centres at (140,200)(140, 200), cell (4,6)(4, 6). Different cells, so each predicts its own mask channel with no box overlap logic. SOLOv2 instead predicts a 99-weight kernel per active cell and convolves it over a 160×160160 \times 160 feature map to render the mask on demand.

Code

# Grid cell assignment for two centres on a 640px image with S=20.S, size = 20, 640for name, x, y in [("A", 100, 200), ("B", 140, 200)]:    print(name, (int(x // (size / S)), int(y // (size / S))))# -> A (3, 6)# -> B (4, 6)

Watch Out For

Coarse grids merging neighbours

Small SS puts two centres in one cell and one mask wins. Symptom: merged twins in crowds. Fix: use FPN-matched fine grids for small objects and confirm centre separation on validation crops.

Porting box NMS thresholds to Matrix NMS

Box-tuned IoU thresholds oversuppress soft mask duplicates. Symptom: missing overlapping instances. Fix: retune the Matrix NMS decay for mask IoU; do not copy detector settings.

The Quick Version

  • SOLO predicts one mask per grid cell keyed by object centre, with no boxes or anchors.
  • FPN levels carry different grid sizes matched to object scale.
  • SOLOv2 predicts dynamic kernels per cell for sharper, cheaper masks.
  • Matrix NMS suppresses duplicates in one parallel mask-IoU step.
  • Grids must be fine enough that neighbour centres separate.