YOLACT Instance Masks
YOLACT predicts image-wide mask prototypes once and per-object coefficients, then blends them linearly for real-time instance masks.
Why Does This Exist?
Mask R-CNN-style two-stage models repool and reconvolve every proposal, which caps speed near FPS. YOLACT (Bolya et al., 2019) splits the work: a Protonet branch predicts full-image prototype masks once, while the detection head predicts one coefficient vector per box. Mask is cropped to its box. Assembly is a matrix multiply, so the system runs above FPS.
Think of It Like This
Paint-by-numbers with shared base coats
A workshop pre-mixes base-coat washes covering the whole canvas: leftness, roundness, vertical edges. Each order ticket lists blending ratios instead of a custom painting session. One ticket says leftness minus roundness, cropped to its frame. Tickets are cheap; the washes are shared.
Where it stops: washes cannot render lace. Fine holes and whiskers fall between prototypes.
How It Actually Works
Prototypes come from a fully convolutional branch at resolution with no per-RoI alignment. Coefficients pass through so subtraction is possible. Training combines classification, box, and mask BCE with a mask weight. Fast NMS replaces sequential suppression. YOLACT++ adds deformable convolutions and a better anchor design.
Worked example
Prototypes at one pixel hold and . Detection predicts coefficients . Combination , sigmoid , so the pixel is foreground. A neighbour with prototypes gives , sigmoid , background. Cropping to the predicted box removes leakage outside it.
Code
import math
# Linear assembly of two prototypes at two pixels.p1, p2 = [0.8, 0.1], [0.2, 0.9]coef = [1.5, -2.0]out = [1 / (1 + math.exp(-(coef[0] * a + coef[1] * b))) for a, b in zip(p1, p2)]print([round(v, 2) for v in out])# -> [0.69, 0.16]Watch Out For
Expecting Mask R-CNN edges at 30 FPS
Shared low-resolution prototypes smooth whiskers and wires. Symptom: AP on small objects trails two-stage models. Fix: choose YOLACT for speed budgets, PointRend or Mask R-CNN for edge fidelity.
Skipping box cropping at inference
Raw assemblies span the image and bleed across neighbours. Symptom: ghost masks between adjacent people. Fix: always crop to the predicted box and threshold at before scoring.
The Quick Version
- One prototype branch plus one coefficient vector per detection replaces per-RoI reconvolution.
- Assembly is a linear blend plus sigmoid, cropped to the box.
- Real-time speeds above FPS on a single GPU generation of its era.
- Small-object and fine-edge accuracy trails two-stage models.
- YOLACT++ adds deformable features and stronger anchors.