Skip to content
AI360Xpert
Beta

OCRNet Object-Contextual Parsing

OCRNet groups pixels into soft object regions first, then relabels each pixel by its similarity to those regions instead of by its lonely local patch.

OCRNet forms a prototype for each object region then boosts every pixel that resembles its own region prototype.
OCRNet forms a prototype for each object region then boosts every pixel that resembles its own region prototype.

Why Does This Exist?

PSPNet and DeepLab ASPP gather context from fixed squares and dilation rings, which mix unrelated objects that happen to sit nearby. A pixel on a bus window pulls in sky because the square says so. OCRNet (Yuan et al., 2020) gathers context from the object's own kind: first estimate rough class regions, pool one descriptor per region, then ask each pixel how much it resembles each descriptor.

This page covers region-guided context. It sits after a backbone like HRNet, which pairs well because its sharp maps make cleaner regions.

Think of It Like This

Sorting mail by pigeonhole summaries

A mailroom has rough pigeonholes for letters, parcels, and magazines. Instead of reading each envelope in isolation, the clerk compares it against the summary card of each hole: average size, weight, stamps. A padded envelope that looks like a magazine matches the parcel card best and goes there. Pixels are envelopes; region descriptors are the summary cards.

Where it stops: if the first rough sort is badly wrong, the summary cards describe the wrong piles and the second pass inherits the error.

How It Actually Works

A backbone predicts a coarse KK-class soft map. For each class kk, OCR pools a region descriptor fkf_k as the softmax-weighted average of pixel features belonging to kk. Then each pixel feature xix_i computes similarity scores siks_{ik} against all KK descriptors, and the object-contextual feature is the weighted sum ∑ksikfk\sum_k s_{ik} f_k. The final head fuses xix_i with this context vector. Pixels on the same object reinforce each other even when far apart.

Worked example

Three region descriptors: bus, road, sky. A window pixel xx has dot-product similarities 4.04.0, 0.50.5, 2.02.0 with them. Softmax gives weights 0.860.86, 0.030.03, 0.120.12. Its context vector is 8686 percent bus prototype, so the fused prediction stays bus despite the local glass texture looking sky-like. A fixed 6×66 \times 6 pool over the same spot would have mixed 4040 percent sky and flipped the vote.

Code

import math
# Softmax over region similarities [bus, road, sky].sims = [4.0, 0.5, 2.0]exps = [math.exp(s) for s in sims]total = sum(exps)print([round(e / total, 2) for e in exps])# -> [0.86, 0.03, 0.12]

Watch Out For

Trusting OCR when the coarse map collapses

If the backbone never predicts some class coarsely, its descriptor is noise and those pixels get no help. Symptom: rare classes improve less than reported. Fix: check coarse-map recall per class first, and train with class-balanced sampling before judging OCR.

Stacking OCR on tiny feature maps

Region pooling on a 1/321/32 map makes mushy regions. Symptom: no gain over ASPP. Fix: run OCR on stride-88 maps from HRNet or a dilated backbone.

The Quick Version

  • OCRNet first estimates soft object regions, then pools one descriptor per class.
  • Each pixel is relabelled using similarity to those descriptors, not fixed neighbour squares.
  • Distant pixels of the same object reinforce each other across the image.
  • Quality depends on a decent coarse map and a high-resolution feature grid.
  • Pairs naturally with HRNet backbones and replaces ASPP-style context.