Skip to content
AI360Xpert
Beta

Background Subtraction

Background subtraction keeps a running portrait of the empty scene and flags any pixel that stops matching it, turning a fixed camera into a motion sensor.

Each frame is differenced against the running background model, and pixels past the threshold become the foreground mask.
Each frame is differenced against the running background model, and pixels past the threshold become the foreground mask.

Why Does This Exist?

A security camera watches an empty hallway 99 percent of the night. Running a full detector on every frame wastes power to rediscover emptiness, and optical flow measures motion without saying what moved. What you actually want is a binary answer per pixel: did something new arrive?

Background subtraction maintains a model of the scene at rest and compares each frame against it. Pixels that disagree enough become foreground, and connected foreground blobs become detections to hand to classical object tracking. It is the cheapest detector that exists when the camera does not move, which is why it still runs inside traffic counters and wildlife cameras.

Think of It Like This

A museum guard's mental photo

A guard memorizes the gallery at opening: painting here, bench there, empty floor. All day she compares the live room against that memory. A visitor standing before a painting pops out instantly, not because she analyzed the visitor but because the visitor matches nothing in her memory.

The running average is her memory, slowly updating as afternoon light shifts through the windows. The threshold is her judgment about what counts as "different enough." And a moved bench fools her exactly once, until her memory absorbs the new arrangement.

How It Actually Works

1. Model the background per pixel

The simplest model is a running average: B←(1−α)B+αIB \leftarrow (1 - \alpha) B + \alpha I, where BB is the background estimate, II the current frame, and α\alpha (say 0.050.05) the learning rate. Worked: background pixel 100100, new frame pixel 180180. Update: B=0.95⋅100+0.05⋅180=104B = 0.95 \cdot 100 + 0.05 \cdot 180 = 104. Difference ∣180−104∣=76|180 - 104| = 76, far above a threshold like T=25T = 25, so the pixel is foreground. The model barely moved, which is the point: background adapts slowly, intruders spike instantly.

2. Threshold into a mask

∣I−B∣>T|I - B| > T marks foreground. One global TT is fragile under shadows and auto-exposure, so real systems use per-pixel variance (Mixture of Gaussians keeps 3 to 5 Gaussian modes per pixel to survive waving trees and flickering screens) plus shadow suppression that discounts brightness-only changes.

3. Clean the mask

Raw masks are speckled. Morphological operations open the mask to delete dots, close it to fill holes inside objects, then connected components turn blobs into boxes. A minimum-area filter drops the remaining flicker.

Code

import numpy as np
def update_foreground(bg: np.ndarray, frame: np.ndarray, alpha: float = 0.05, thresh: float = 25.0):    bg = (1 - alpha) * bg + alpha * frame    mask = np.abs(frame - bg) > thresh    return bg, mask
bg = np.full((2, 2), 100.0)bg, mask = update_foreground(bg, np.array([[100.0, 180.0], [100.0, 102.0]]))print(bg[0, 1], mask[0, 1])  # -> 104.0 True

Watch Out For

Stationary objects melt into the background

A suitcase left on the platform is foreground today and furniture next week, because the running average absorbs it. The absorption speed is α\alpha: fast learning forgets stopped objects quickly. If your task includes abandoned-object alarms, you need a second slow model or an explicit stationary-object test, not just a tuned α\alpha.

Sudden light changes are fake motion

Someone flips a light switch and every pixel disagrees at once: full-frame false foreground. Per-pixel Gaussians recover over seconds, but the alarm already fired. Gate global-change frames by the fraction of pixels that flipped, and normalize illumination before differencing where you can.

The Quick Version

  • Background subtraction models the empty scene and thresholds per-pixel disagreement into a foreground mask.
  • The running-average example moves 100100 to 104104 on a 180180 intruder pixel and flags it at threshold 2525.
  • Mixture of Gaussians per pixel survives waving trees and flicker that kill single-value models.
  • Morphology plus connected components turns speckled masks into detection boxes.
  • Fixed cameras only; stopped objects absorb and light switches cause full-frame false alarms.