Skip to content
AI360Xpert
Beta

HRNet High-Resolution Networks

HRNet keeps a full-resolution stream alive from input to output and fuses it with lower streams, so position accuracy never has to be recovered.

HRNet processes high medium and low resolution streams side by side and exchanges information between them at every stage.
HRNet processes high medium and low resolution streams side by side and exchanges information between them at every stage.

Why Does This Exist?

Most backbones downsample first and ask the head to recover position later, which is why encoder-decoder designs need skips and pose heads need heatmap upsampling. HRNet (Sun et al., 2019) refuses the bottleneck: a 1/41/4-resolution stream starts early and stays alive to the output, joined by 1/81/8, 1/161/16, and 1/321/32 streams that supply context. Repeated fusion means the high stream always knows the semantics and the low streams always know the position.

This page is the shared backbone reference for both segmentation and pose. Task heads differ; the representation trick is one idea, so it lives here once.

Think of It Like This

A string quartet that never stops listening

Most orchestras rehearse sections separately and merge at the concert. HRNet is a quartet playing together from bar one: the first violin (high resolution) never leaves the room, and the cello (low resolution, deep context) hums underneath. Every eight bars they glance at each other and retune. No part is reconstructed from memory at the end.

Where it stops: four live streams cost memory, so the quartet needs a bigger stage than a single downsampled trumpet.

How It Actually Works

Stage 11 is a high-resolution convolution stream. Stage 22 adds a 1/21/2-resolution branch; stage 33 adds 1/41/4; stage 44 adds 1/81/8, all relative to that stream. Each multi-resolution block exchanges information: every stream upsamples or downsamples into every other stream and sums. A segmentation head reads the fused high stream directly. A pose head regresses heatmaps from it with little extra upsampling.

Worked example

Input 256×192256 \times 192 for pose. The high stream holds 64×4864 \times 48 throughout. After a fusion, a wrist keypoint at (200,150)(200, 150) in the input sits at (50,37.5)(50, 37.5) in the high stream, inside one cell of ground truth. The 1/321/32 stream sees the whole arm posture and tells the high stream which blob is the wrist versus the elbow, cutting the left-right swap rate that plagues bottleneck backbones.

Code

# Where an input keypoint lands in the HRNet high stream (stride 4).x, y, stride = 200, 150, 4print((x / stride, y / stride))# -> (50.0, 37.5)
# Streams alive in stage 4: full, half, quarter, eighth of the high stream.print([1 / (2 ** i) for i in range(4)])# -> [1.0, 0.5, 0.25, 0.125]

Watch Out For

Running all four streams on small GPUs

Four parallel streams with repeated fusion use notably more memory than a ResNet of similar depth. Symptom: out-of-memory at batch sizes that worked for ResNet-50. Fix: use HRNet-W18 for prototyping, enable gradient checkpointing, or freeze early stages.

Downsampling the HRNet output out of habit

Some ports add a stride-22 stem before the head, throwing away the resolution HRNet preserved. Symptom: pose and mask accuracy drops to ResNet levels. Fix: read the highest-resolution stream directly and upsample at most 2×2\times in the head.

The Quick Version

  • HRNet maintains a high-resolution stream end to end instead of downsampling then recovering.
  • Lower-resolution streams join in stages and all streams fuse repeatedly.
  • Segmentation heads read sharp maps directly; pose heads need little extra upsampling.
  • Accuracy on boundaries and keypoints rises; memory cost rises too.
  • Use the small W18 width when hardware is tight.