OCRNet Object-Contextual Parsing
OCRNet groups pixels into soft object regions first, then relabels each pixel by its similarity to those regions instead of by its lonely local patch.
Why Does This Exist?
PSPNet and DeepLab ASPP gather context from fixed squares and dilation rings, which mix unrelated objects that happen to sit nearby. A pixel on a bus window pulls in sky because the square says so. OCRNet (Yuan et al., 2020) gathers context from the object's own kind: first estimate rough class regions, pool one descriptor per region, then ask each pixel how much it resembles each descriptor.
This page covers region-guided context. It sits after a backbone like HRNet, which pairs well because its sharp maps make cleaner regions.
Think of It Like This
Sorting mail by pigeonhole summaries
A mailroom has rough pigeonholes for letters, parcels, and magazines. Instead of reading each envelope in isolation, the clerk compares it against the summary card of each hole: average size, weight, stamps. A padded envelope that looks like a magazine matches the parcel card best and goes there. Pixels are envelopes; region descriptors are the summary cards.
Where it stops: if the first rough sort is badly wrong, the summary cards describe the wrong piles and the second pass inherits the error.
How It Actually Works
A backbone predicts a coarse -class soft map. For each class , OCR pools a region descriptor as the softmax-weighted average of pixel features belonging to . Then each pixel feature computes similarity scores against all descriptors, and the object-contextual feature is the weighted sum . The final head fuses with this context vector. Pixels on the same object reinforce each other even when far apart.
Worked example
Three region descriptors: bus, road, sky. A window pixel has dot-product similarities , , with them. Softmax gives weights , , . Its context vector is percent bus prototype, so the fused prediction stays bus despite the local glass texture looking sky-like. A fixed pool over the same spot would have mixed percent sky and flipped the vote.
Code
import math
# Softmax over region similarities [bus, road, sky].sims = [4.0, 0.5, 2.0]exps = [math.exp(s) for s in sims]total = sum(exps)print([round(e / total, 2) for e in exps])# -> [0.86, 0.03, 0.12]Watch Out For
Trusting OCR when the coarse map collapses
If the backbone never predicts some class coarsely, its descriptor is noise and those pixels get no help. Symptom: rare classes improve less than reported. Fix: check coarse-map recall per class first, and train with class-balanced sampling before judging OCR.
Stacking OCR on tiny feature maps
Region pooling on a map makes mushy regions. Symptom: no gain over ASPP. Fix: run OCR on stride- maps from HRNet or a dilated backbone.
The Quick Version
- OCRNet first estimates soft object regions, then pools one descriptor per class.
- Each pixel is relabelled using similarity to those descriptors, not fixed neighbour squares.
- Distant pixels of the same object reinforce each other across the image.
- Quality depends on a decent coarse map and a high-resolution feature grid.
- Pairs naturally with HRNet backbones and replaces ASPP-style context.