Video Classification Basics
Video classification labels a whole clip with one tag, sampling frames smartly and pooling them so the majority verdict survives boring stretches.
Why Does This Exist?
Tagging millions of uploads, flagging highlight reels and routing content all need one label per clip, not per frame or per second. Frame classifiers drown in redundancy while missing the point. Video classification exists as the cheapest useful video task: sample, embed, pool, decide. It is the parent of action recognition (human verbs) and the sibling of localization tasks that also ask when. The parent picture is video understanding.
Think of It Like This
Judging a parade from snapshots
You cannot watch a three-hour parade, so you photograph one float per block and judge the parade's theme from the spread. Ten marching bands and two clowns means a music parade even if one photo caught a nap. Clip classifiers judge parades the same way: sparse snapshots, pooled verdicts, robust to the odd dull frame. The analogy stops at the pooling: you weigh the finale more, while mean pooling weighs every sample equally unless attention says otherwise.
How It Actually Works
Uniform or segment-based sampling picks K frames or short snippets across the clip. A 2D or 3D backbone embeds each, and temporal pooling (mean, max, attention-weighted, or a small transformer) fuses them into one clip vector for the classifier. Training uses video-level cross-entropy; inference averages several multi-crop views. Sparse segment sampling (as in TSN-style designs) covers minutes of footage with a dozen frames.
A worked vote
Twelve sampled snippets from a soccer compilation score goal-celebration at 0.9 on four snippets, midfield play at 0.7 on six, and crowd shots near uniform on two. Mean pooling gives celebration (4 x 0.9 + 6 x 0.2 + noise) / 12 = 0.40 against play (4 x 0.1 + 6 x 0.7 + noise) / 12 = 0.38: celebration edges it on peak strength despite fewer snippets. Attention pooling would widen the gap by down-weighting the crowd frames explicitly.
Watch Out For
Single-label verdicts on multi-event clips
One tag per clip fails compilations and long untrimmed videos where several events share the runtime. The symptom is oscillating predictions across seeds as pooling ties break randomly. Fix it by switching to multi-label heads or to temporal localization when clips hold more than one event.
The Quick Version
- One label per clip from sampled snippets, pooled and classified.
- Segment sampling stretches a dozen frames across minutes affordably.
- Pooling choice (mean, max, attention) decides how dull stretches count.
- Multi-view inference averages crops and clips for stability.
- Multi-event videos need multi-label or localization, not single tags.