CondInst Conditional Masks
CondInst predicts a tiny custom filter per detected instance and convolves it over shared features, so neighbours get different masks from the same map.
Why Does This Exist?
Proposal-free heads like SOLO separate by grid cell and YOLACT blends shared prototypes, but overlapping same-cell neighbours still confuse fixed decoders. CondInst (Tian et al., 2020) gives each instance its own decoder: the controller head predicts a compact filter vector ( weights) per location, and that filter convolves the shared mask features to render exactly one object. Same features in, different filter per instance out.
Think of It Like This
Cookie cutters forged per order
A bakery keeps one sheet of dough (shared features) and forges a custom cutter per order from the ticket photo. Twin orders get twin cutters with shifted teeth. One press each, and identical dough yields distinct cookies. Fixed stencils would stamp the same shape twice.
Where it stops: forging a cutter per pixel of a crowd costs controller compute, so dense scenes need score filtering first.
How It Actually Works
The FCOS-style head outputs classes, box offsets, centreness, and controller parameters per FPN location. Mask features append two relative-coordinate channels measured from the instance centre, giving the dynamic filter a sense of place. The generated filters (two layers plus a prediction layer) convolve a small patch around each positive location. Dice plus focal loss trains masks; only top-scoring instances render at inference.
Worked example
Two pedestrians overlap with centres pixels apart on the map. Location A generates filter , location B generates . Applied to identical shared features, fires on the left silhouette and on the right, while fires the reverse, because the relative-coordinate channels shift the response. Fixed shared weights would output one merged blob.
Code
# Relative coordinates separate two instances sharing features.def rel(px: int, cx: int, stride: int = 8) -> float: return (px - cx) / stride
print((rel(100, 100), rel(112, 100)))# -> (0.0, 1.5)Watch Out For
Dropping relative coordinates
Without the centre-relative channels the dynamic filter cannot tell twins apart. Symptom: merged neighbours despite per-instance filters. Fix: always concatenate the two coordinate maps before the dynamic convolution.
Rendering every low-score location
Dense locations each forge filters, most for background. Symptom: slow inference and noisy specks. Fix: filter by classification score first and render only survivors.
The Quick Version
- CondInst predicts custom convolution filters per instance location.
- Dynamic filters convolve shared mask features plus relative coordinates.
- Overlapping neighbours separate because their filters and coordinates differ.
- No RoI alignment or boxes are needed for the mask step.
- Score filtering keeps per-instance forging affordable.