SlowFast Networks for Video
SlowFast watches video at two speeds: a slow path studies sharp appearances while a fast path catches quick motion, then they compare notes.
Why Does This Exist?
Uniform frame sampling wastes resolution on still backgrounds and starves fast motion of temporal detail: one rate cannot serve both. Two-stream networks split appearance from motion but pay for optical flow. SlowFast exists as the RGB-only answer with two frame rates inside one network. A slow path samples sparsely at full detail for semantics; a fast path samples densely at reduced channels for tempo. Lateral connections fuse them at several depths. The task context is action recognition.
Think of It Like This
A painter and a drummer watching a dance
A painter sketches one pose per bar in full detail while a drummer taps every beat with eyes half closed. The painter knows who dances; the drummer knows the rhythm. Between bars the drummer's taps annotate the painter's canvas with tempo marks. SlowFast pairs the same two observers: the slow path paints semantics, the fast path drums motion, and lateral links annotate across. The analogy stops at the channels: the drummer uses fewer neurons per tap (about one-eighth the channels) because rhythm needs less detail than identity.
How It Actually Works
The slow path takes roughly 8 frames across the clip through a full 3D ResNet; the fast path takes roughly 32 frames (4x the rate) through a thin network with about one-eighth the channels. Time-strided convolutions in the fast path preserve tempo while lateral concatenations inject motion features into the slow path at each stage. One classifier head reads the fused representation.
A worked budget
A 64-frame window feeds 8 frames to the slow path and 32 to the fast path. If a slow frame costs 1.0 unit of compute, a fast frame at one-eighth channels costs about 0.125, so the fast path adds 32 x 0.125 = 4.0 units against the slow path's 8.0: tempo costs half as much as semantics while quadrupling the frame rate. That ratio is the design's whole economy.
Watch Out For
Slow-only ablations that hide tempo failures
The slow path alone scores well on scene-biased benchmarks, tempting teams to delete the fast path for speed. The symptom is a lean model that cannot tell fast claps from slow waves. Fix it by benchmarking on motion-sensitive splits (something-something style) where tempo decides, before amputating.
The Quick Version
- Slow path: sparse frames, full channels, semantics. Fast path: dense frames, thin channels, tempo.
- Lateral connections fuse motion into semantics at multiple depths.
- No optical flow needed: tempo comes from raw RGB at high frame rate.
- Fast path costs a fraction per frame by design, keeping the budget sane.
- Verify on motion-sensitive data or scene bias will hide tempo blindness.