Skip to content
AI360Xpert
Beta

RegNet Design Spaces

RegNet stops hand-tuning single networks and searches whole design spaces instead, finding simple width and depth rules that beat years of manual tweaks.

Sampled design spaces converge on quantized linear widths 48, 104, 208 and 440, simpler stages that beat hand tuned depth.
Sampled design spaces converge on quantized linear widths 48, 104, 208 and 440, simpler stages that beat hand tuned depth.

Why Does This Exist?

Hand-designed networks accumulate quirks: one stage has extra blocks because it helped once, widths double on schedule because that was the habit. Each new model inherits unexplained choices and nobody knows which ones matter.

RegNet (Radosavovic et al., 2020) searches populations instead of instances. Sample thousands of networks from a parameterized space, plot error against FLOPs for the whole population, and keep the space whose distribution dominates. The surviving rule is simple: stage widths follow a quantized linear function wj=w0⋅wajw_j = w_0 \cdot w_a^{j} rounded to stages. This page covers that space-level method. The residual blocks inside live in residual networks and the scaling comparison in EfficientNet.

Think of It Like This

Breeding seed varieties instead of one plant

A farmer testing one plant learns about that plant. Testing whole seed varieties across hundreds of plots learns which variety wins in most soils, then the best plots reveal the simple trait that mattered, such as stalk thickness growing steadily with height.

Single-network tuning is the one plant. Design spaces are the varieties. The quantized linear width rule is the stalk trait.

How It Actually Works

The AnyNetX space parameterizes depth per stage, width per stage, bottleneck ratio, group width and stride. Sampling 500 networks at 400 MFLOPs and plotting mean error shows which spaces dominate; the best spaces share depth near 20 blocks, bottleneck ratio 1 (no bottleneck) and group width that grows with width.

Widths quantize cleanly. With w0=48w_0 = 48 and slope wa=2.1w_a = 2.1, block widths 48×2.1j48 \times 2.1^{j} for j=0,1,2,3j = 0, 1, 2, 3 give 48, 101, 212 and 445, quantized into four stages near 48, 104, 208 and 440. RegNetX keeps this with grouped 3x3 convolutions; RegNetY adds squeeze-excitation per block. RegNetY-400MF reaches about 75% top-1 where a matched ResNet-50 costs ten times the FLOPs for one point more.

Code

w0, wa = 48, 2.1print([round(w0 * wa ** j) for j in range(4)])# -> [48, 101, 212, 445]

Watch Out For

Copying RegNet widths at a different FLOP budget

The quantized widths are fit per regime; 400 MFLOP widths starve a 4 GFLOP model. The symptom is a large RegNet that underperforms a plain ResNet. Pick the RegNet variant whose MFLOP tag matches the budget instead of widening a small one by hand.

Searching one seed and declaring a space winner

Single-network comparisons swing several tenths of a point with seeds. The symptom is a published tweak that never reproduces. Compare population means and best-fit curves across dozens of samples, the way the RegNet study does, before adopting a space change.

The Quick Version

  • RegNet compares whole design spaces by sampling hundreds of networks, not by tuning one instance.
  • Winning spaces use simple regular stages with widths on a quantized linear curve such as 48, 104, 208 and 440.
  • RegNetX uses grouped convolutions and RegNetY adds squeeze-excitation per block.
  • RegNetY at 400 MFLOPs nears 75% top-1 where ResNet-50 needs ten times the compute for one point more.
  • Match the variant tag to the FLOP budget instead of hand-stretching a small RegNet.