Background Subtraction
Background subtraction keeps a running portrait of the empty scene and flags any pixel that stops matching it, turning a fixed camera into a motion sensor.
Why Does This Exist?
A security camera watches an empty hallway 99 percent of the night. Running a full detector on every frame wastes power to rediscover emptiness, and optical flow measures motion without saying what moved. What you actually want is a binary answer per pixel: did something new arrive?
Background subtraction maintains a model of the scene at rest and compares each frame against it. Pixels that disagree enough become foreground, and connected foreground blobs become detections to hand to classical object tracking. It is the cheapest detector that exists when the camera does not move, which is why it still runs inside traffic counters and wildlife cameras.
Think of It Like This
A museum guard's mental photo
A guard memorizes the gallery at opening: painting here, bench there, empty floor. All day she compares the live room against that memory. A visitor standing before a painting pops out instantly, not because she analyzed the visitor but because the visitor matches nothing in her memory.
The running average is her memory, slowly updating as afternoon light shifts through the windows. The threshold is her judgment about what counts as "different enough." And a moved bench fools her exactly once, until her memory absorbs the new arrangement.
How It Actually Works
1. Model the background per pixel
The simplest model is a running average: , where is the background estimate, the current frame, and (say ) the learning rate. Worked: background pixel , new frame pixel . Update: . Difference , far above a threshold like , so the pixel is foreground. The model barely moved, which is the point: background adapts slowly, intruders spike instantly.
2. Threshold into a mask
marks foreground. One global is fragile under shadows and auto-exposure, so real systems use per-pixel variance (Mixture of Gaussians keeps 3 to 5 Gaussian modes per pixel to survive waving trees and flickering screens) plus shadow suppression that discounts brightness-only changes.
3. Clean the mask
Raw masks are speckled. Morphological operations open the mask to delete dots, close it to fill holes inside objects, then connected components turn blobs into boxes. A minimum-area filter drops the remaining flicker.
Code
import numpy as np
def update_foreground(bg: np.ndarray, frame: np.ndarray, alpha: float = 0.05, thresh: float = 25.0): bg = (1 - alpha) * bg + alpha * frame mask = np.abs(frame - bg) > thresh return bg, mask
bg = np.full((2, 2), 100.0)bg, mask = update_foreground(bg, np.array([[100.0, 180.0], [100.0, 102.0]]))print(bg[0, 1], mask[0, 1]) # -> 104.0 TrueWatch Out For
Stationary objects melt into the background
A suitcase left on the platform is foreground today and furniture next week, because the running average absorbs it. The absorption speed is : fast learning forgets stopped objects quickly. If your task includes abandoned-object alarms, you need a second slow model or an explicit stationary-object test, not just a tuned .
Sudden light changes are fake motion
Someone flips a light switch and every pixel disagrees at once: full-frame false foreground. Per-pixel Gaussians recover over seconds, but the alarm already fired. Gate global-change frames by the fraction of pixels that flipped, and normalize illumination before differencing where you can.
The Quick Version
- Background subtraction models the empty scene and thresholds per-pixel disagreement into a foreground mask.
- The running-average example moves to on a intruder pixel and flags it at threshold .
- Mixture of Gaussians per pixel survives waving trees and flicker that kill single-value models.
- Morphology plus connected components turns speckled masks into detection boxes.
- Fixed cameras only; stopped objects absorb and light switches cause full-frame false alarms.