Skip to content
AI360Xpert
Beta

YOLACT Instance Masks

YOLACT predicts image-wide mask prototypes once and per-object coefficients, then blends them linearly for real-time instance masks.

YOLACT blends a small set of shared full-image prototypes with per-box coefficients to assemble each instance mask.
YOLACT blends a small set of shared full-image prototypes with per-box coefficients to assemble each instance mask.

Why Does This Exist?

Mask R-CNN-style two-stage models repool and reconvolve every proposal, which caps speed near 1010 FPS. YOLACT (Bolya et al., 2019) splits the work: a Protonet branch predicts k≈32k \approx 32 full-image prototype masks once, while the detection head predicts one coefficient vector per box. Mask jj is σ(∑cajcPc)\sigma(\sum_c a_{jc} P_c) cropped to its box. Assembly is a matrix multiply, so the system runs above 3030 FPS.

Think of It Like This

Paint-by-numbers with shared base coats

A workshop pre-mixes 3232 base-coat washes covering the whole canvas: leftness, roundness, vertical edges. Each order ticket lists blending ratios instead of a custom painting session. One ticket says 0.90.9 leftness minus 0.40.4 roundness, cropped to its frame. Tickets are cheap; the washes are shared.

Where it stops: 3232 washes cannot render lace. Fine holes and whiskers fall between prototypes.

How It Actually Works

Prototypes come from a fully convolutional branch at 1/41/4 resolution with no per-RoI alignment. Coefficients pass through tanh⁡\tanh so subtraction is possible. Training combines classification, box, and mask BCE with a 6.125×6.125\times mask weight. Fast NMS replaces sequential suppression. YOLACT++ adds deformable convolutions and a better anchor design.

Worked example

Prototypes P1,P2P_1, P_2 at one pixel hold 0.80.8 and 0.20.2. Detection jj predicts coefficients [1.5,−2.0][1.5, -2.0]. Combination 1.5(0.8)−2.0(0.2)=1.2−0.4=0.81.5(0.8) - 2.0(0.2) = 1.2 - 0.4 = 0.8, sigmoid ≈0.69\approx 0.69, so the pixel is foreground. A neighbour with prototypes [0.1,0.9][0.1, 0.9] gives 0.15−1.8=−1.650.15 - 1.8 = -1.65, sigmoid ≈0.16\approx 0.16, background. Cropping to the predicted box removes leakage outside it.

Code

import math
# Linear assembly of two prototypes at two pixels.p1, p2 = [0.8, 0.1], [0.2, 0.9]coef = [1.5, -2.0]out = [1 / (1 + math.exp(-(coef[0] * a + coef[1] * b))) for a, b in zip(p1, p2)]print([round(v, 2) for v in out])# -> [0.69, 0.16]

Watch Out For

Expecting Mask R-CNN edges at 30 FPS

Shared low-resolution prototypes smooth whiskers and wires. Symptom: AP on small objects trails two-stage models. Fix: choose YOLACT for speed budgets, PointRend or Mask R-CNN for edge fidelity.

Skipping box cropping at inference

Raw assemblies span the image and bleed across neighbours. Symptom: ghost masks between adjacent people. Fix: always crop to the predicted box and threshold at 0.50.5 before scoring.

The Quick Version

  • One prototype branch plus one coefficient vector per detection replaces per-RoI reconvolution.
  • Assembly is a linear blend plus sigmoid, cropped to the box.
  • Real-time speeds above 3030 FPS on a single GPU generation of its era.
  • Small-object and fine-edge accuracy trails two-stage models.
  • YOLACT++ adds deformable features and stronger anchors.