HRNet High-Resolution Networks
HRNet keeps a full-resolution stream alive from input to output and fuses it with lower streams, so position accuracy never has to be recovered.
Why Does This Exist?
Most backbones downsample first and ask the head to recover position later, which is why encoder-decoder designs need skips and pose heads need heatmap upsampling. HRNet (Sun et al., 2019) refuses the bottleneck: a -resolution stream starts early and stays alive to the output, joined by , , and streams that supply context. Repeated fusion means the high stream always knows the semantics and the low streams always know the position.
This page is the shared backbone reference for both segmentation and pose. Task heads differ; the representation trick is one idea, so it lives here once.
Think of It Like This
A string quartet that never stops listening
Most orchestras rehearse sections separately and merge at the concert. HRNet is a quartet playing together from bar one: the first violin (high resolution) never leaves the room, and the cello (low resolution, deep context) hums underneath. Every eight bars they glance at each other and retune. No part is reconstructed from memory at the end.
Where it stops: four live streams cost memory, so the quartet needs a bigger stage than a single downsampled trumpet.
How It Actually Works
Stage is a high-resolution convolution stream. Stage adds a -resolution branch; stage adds ; stage adds , all relative to that stream. Each multi-resolution block exchanges information: every stream upsamples or downsamples into every other stream and sums. A segmentation head reads the fused high stream directly. A pose head regresses heatmaps from it with little extra upsampling.
Worked example
Input for pose. The high stream holds throughout. After a fusion, a wrist keypoint at in the input sits at in the high stream, inside one cell of ground truth. The stream sees the whole arm posture and tells the high stream which blob is the wrist versus the elbow, cutting the left-right swap rate that plagues bottleneck backbones.
Code
# Where an input keypoint lands in the HRNet high stream (stride 4).x, y, stride = 200, 150, 4print((x / stride, y / stride))# -> (50.0, 37.5)
# Streams alive in stage 4: full, half, quarter, eighth of the high stream.print([1 / (2 ** i) for i in range(4)])# -> [1.0, 0.5, 0.25, 0.125]Watch Out For
Running all four streams on small GPUs
Four parallel streams with repeated fusion use notably more memory than a ResNet of similar depth. Symptom: out-of-memory at batch sizes that worked for ResNet-50. Fix: use HRNet-W18 for prototyping, enable gradient checkpointing, or freeze early stages.
Downsampling the HRNet output out of habit
Some ports add a stride- stem before the head, throwing away the resolution HRNet preserved. Symptom: pose and mask accuracy drops to ResNet levels. Fix: read the highest-resolution stream directly and upsample at most in the head.
The Quick Version
- HRNet maintains a high-resolution stream end to end instead of downsampling then recovering.
- Lower-resolution streams join in stages and all streams fuse repeatedly.
- Segmentation heads read sharp maps directly; pose heads need little extra upsampling.
- Accuracy on boundaries and keypoints rises; memory cost rises too.
- Use the small W18 width when hardware is tight.