MobileNet Efficient Convolutions
MobileNet splits each convolution into a per-channel spatial pass and a 1x1 mix, cutting phone inference cost nearly ninefold while keeping accuracy close to full networks.
Why Does This Exist?
VGG-16 needs about 15 billion mult-adds per image and 138 million weights. No phone runs that at 30 frames per second on a battery. But phones are where cameras live, so vision had to shrink by an order of magnitude without collapsing accuracy.
MobileNet (Howard et al., 2017) is that shrink. It replaces nearly every standard convolution with a depthwise separable pair and adds two knobs, width multiplier and resolution multiplier, that trade accuracy for latency in one line. MobileNetV2 (2018) added inverted residuals with linear bottlenecks. This page covers the efficiency mechanism. The separable math detail lives in depthwise separable convolutions and scaling strategy in EfficientNet.
Think of It Like This
Sorting mail by street then by name
A mailroom sorts 256 bins of letters first by street (each carrier handles one street, never mixing) and then one clerk recombines streets into names at each address. Two simple passes replace one giant sort where everyone touches everything.
Depthwise is the street pass. The 1x1 pointwise step is the clerk. Where the analogy stops: the clerk step learns its mix during training rather than following fixed rules.
How It Actually Works
A standard 3x3 from 128 to 256 channels costs weights. The separable version costs , about 8.7 times fewer. Mult-adds fall by the same ratio, which is how MobileNet reaches roughly 569 million mult-adds against VGG-level baselines in the tens of billions.
Width multiplier scales every channel count: quarters the compute of most layers. Resolution multiplier scales input size: 160 instead of 224 cuts cost by . V2 inverts the bottleneck: expand 6x with a 1x1, filter with depthwise 3x3, project back with a linear 1x1, and shortcut across when shapes match. Linear matters because ReLU on a narrow 16-channel bottleneck destroys information the shortcut must carry.
Code
standard = 3 * 3 * 128 * 256separable = 3 * 3 * 128 + 128 * 256print(standard, separable)# -> (294912, 33920)
print(round(standard / separable, 1))# -> 8.7
print(round((160 / 224) ** 2, 2))# -> 0.51Watch Out For
Deploying full width on a latency budget
Full MobileNetV2 at 224 still misses 30 FPS on older phones. The symptom is framedrops blamed on the app. Profile with or input 160 first; accuracy drops a few points while latency halves.
Adding ReLU after the narrow projection in V2 blocks
V2 projections stay linear on purpose. Forcing a ReLU onto a 16-channel bottleneck clips the signal the residual must carry. The symptom is accuracy that never recovers no matter the schedule. Keep projections linear and put nonlinearities only in the expanded width.
The Quick Version
- MobileNet replaces standard convolutions with depthwise plus pointwise pairs, cutting cost about 8.7 times per layer.
- Width and resolution multipliers trade accuracy for latency in predictable halves and quarters.
- V2 uses inverted residuals with linear bottlenecks so narrow shortcuts carry full information.
- A full MobileNet runs near 569 million mult-adds where VGG-class nets need tens of billions.
- Keep V2 projections linear and profile reduced widths before blaming the app for framedrops.