SqueezeNet Fire Modules
SqueezeNet matches AlexNet accuracy with 50 times fewer weights by squeezing channels through 1x1 layers before every 3x3, so the model file fits in half a megabyte.
Why Does This Exist?
AlexNet holds 60 million parameters, about 240 MB of float32 weights. That cannot ship inside a car camera or update over a slow connection. But much of AlexNet is waste: wide 3x3 layers processing hundreds of channels that a narrow projection could compress first.
SqueezeNet (Iandola et al., 2016) applies three rules: replace most 3x3 filters with 1x1, cut input channels to remaining 3x3 layers with squeeze layers, and downsample late so early maps stay large. The fire module is the unit. This page covers that compression pattern. 1x1 bottlenecks explain the squeeze step and MobileNet continues the efficiency story.
Think of It Like This
Funneling a crowd through a narrow gate
A stadium empties 1,000 fans through a wide concourse with expensive flooring. Funneling them through a narrow gate into a small lobby, then opening back out, paves far less floor while everyone still exits.
The squeeze 1x1 is the gate. The small lobby is the narrow channel count. The expand 1x1 plus 3x3 pair is the reopening. Where the analogy stops: the gate learns which fans to group rather than passing everyone through.
How It Actually Works
A fire module squeezes channels to with a 1x1, then expands to channels of 1x1 plus channels of 3x3, concatenated. With , , : squeeze costs , expand-1x1 costs , expand-3x3 costs . Total against a direct 3x3 at , a 12-fold cut per module.
Full SqueezeNet holds about 1.2 million parameters against AlexNet at 60 million, a 50-fold cut at matched accuracy near 57% top-1. Late downsampling keeps 55x55 maps deep into the net where AlexNet pools early, trading activation memory for parameter savings.
Code
direct = 3 * 3 * 128 * 128fire = 128 * 16 + 16 * 64 + 3 * 3 * 16 * 64print(direct, fire)# -> (147456, 12288)
print(round(direct / fire))# -> 12
print(round(60_000_000 / 1_200_000))# -> 50Watch Out For
Reading parameter savings as latency savings
SqueezeNet saves weights, not activations: late downsampling holds large maps that cost memory bandwidth. The symptom is a tiny file that still runs slowly. Profile activation traffic, not just file size, when latency is the goal.
Squeezing past the point of recovery
Squeeze ratios near 8 to 1 starve the expand 3x3 of inputs. The symptom is accuracy that plateaus far below baseline. Keep squeeze near one eighth of expand total and widen before adding depth.
The Quick Version
- Fire modules squeeze channels with a 1x1, then expand with parallel 1x1 and 3x3 branches.
- One 128-channel module drops from 147,456 weights to 12,288, a 12-fold cut.
- Full SqueezeNet matches AlexNet near 57% top-1 with 1.2 million parameters against 60 million.
- Late downsampling preserves accuracy but keeps activation memory high despite the tiny file.
- Parameter count and inference speed are different budgets; profile both.