Skip to content
AI360Xpert
Beta

CondInst Conditional Masks

CondInst predicts a tiny custom filter per detected instance and convolves it over shared features, so neighbours get different masks from the same map.

CondInst mixes a custom filter per instance and stamps it over shared features so each object renders its own mask.
CondInst mixes a custom filter per instance and stamps it over shared features so each object renders its own mask.

Why Does This Exist?

Proposal-free heads like SOLO separate by grid cell and YOLACT blends shared prototypes, but overlapping same-cell neighbours still confuse fixed decoders. CondInst (Tian et al., 2020) gives each instance its own decoder: the controller head predicts a compact filter vector (≈169\approx 169 weights) per location, and that filter convolves the shared 1/81/8 mask features to render exactly one object. Same features in, different filter per instance out.

Think of It Like This

Cookie cutters forged per order

A bakery keeps one sheet of dough (shared features) and forges a custom cutter per order from the ticket photo. Twin orders get twin cutters with shifted teeth. One press each, and identical dough yields distinct cookies. Fixed stencils would stamp the same shape twice.

Where it stops: forging a cutter per pixel of a crowd costs controller compute, so dense scenes need score filtering first.

How It Actually Works

The FCOS-style head outputs classes, box offsets, centreness, and controller parameters per FPN location. Mask features append two relative-coordinate channels measured from the instance centre, giving the dynamic filter a sense of place. The generated 3×33 \times 3 filters (two layers plus a 1×11 \times 1 prediction layer) convolve a small patch around each positive location. Dice plus focal loss trains masks; only top-scoring instances render at inference.

Worked example

Two pedestrians overlap with centres 1212 pixels apart on the 1/81/8 map. Location A generates filter wAw_A, location B generates wBw_B. Applied to identical shared features, wAw_A fires 0.90.9 on the left silhouette and 0.10.1 on the right, while wBw_B fires the reverse, because the relative-coordinate channels shift the response. Fixed shared weights would output one merged blob.

Code

# Relative coordinates separate two instances sharing features.def rel(px: int, cx: int, stride: int = 8) -> float:    return (px - cx) / stride
print((rel(100, 100), rel(112, 100)))# -> (0.0, 1.5)

Watch Out For

Dropping relative coordinates

Without the centre-relative channels the dynamic filter cannot tell twins apart. Symptom: merged neighbours despite per-instance filters. Fix: always concatenate the two coordinate maps before the dynamic convolution.

Rendering every low-score location

Dense locations each forge filters, most for background. Symptom: slow inference and noisy specks. Fix: filter by classification score first and render only survivors.

The Quick Version

  • CondInst predicts custom convolution filters per instance location.
  • Dynamic filters convolve shared mask features plus relative coordinates.
  • Overlapping neighbours separate because their filters and coordinates differ.
  • No RoI alignment or boxes are needed for the mask step.
  • Score filtering keeps per-instance forging affordable.