Kalman Filter Tracking
The Kalman filter tracks by predicting where the object should be from its motion, then splitting the difference between that prediction and the new measurement by trust.
Why Does This Exist?
Detections are noisy and intermittent. A detector jitters a box by a few pixels each frame, misses frames entirely during occlusion, and gives no velocity to aim with. Naive tracking that copies the last detection inherits all of that: jittery boxes, lost IDs behind every pillar.
The Kalman filter (1960) is the optimal estimator for linear motion with Gaussian noise, and the workhorse motion model inside classical object tracking pipelines like SORT. It keeps a running belief (position plus velocity) with an uncertainty, predicts it forward each frame, and corrects it with each measurement, weighting prediction against measurement by their relative uncertainties. During occlusion there is no measurement, so it coasts on prediction until the object reappears.
Think of It Like This
Catching a ball with your eyes half closed
You watch a ball arc toward you, blink, and your brain keeps simulating the arc during the blink. When your eyes open, you see the ball slightly off from where you imagined and instantly revise. The longer the blink, the less you trust your simulation; the blurrier your vision, the less you trust your eyes.
Prediction is the simulation during the blink. The measurement is the bleary glimpse. The Kalman gain is the instant revision rule: trust the glimpse more when your simulation is stale, trust the simulation more when your vision is blurry.
How It Actually Works
1. State and prediction
The state holds position and velocity, . Each step predicts forward with the motion model: , unchanged (constant velocity), while uncertainty grows by process noise . A growing means "I am getting less sure while coasting."
2. Update by the Kalman gain
A new measurement arrives with noise . The gain in the scalar case splits the difference: state prediction , uncertainty . Worked: predicted position , measurement , gain . Updated position , exactly halfway because both sides were trusted equally. High chases measurements; low rides the model.
3. Coast through occlusion
With no detection, skip the update and predict again. Uncertainty compounds each missed frame, so the gate for matching new detections widens honestly instead of pretending confidence. Particle filters generalize this to non-Gaussian messes where one bell curve cannot hold the belief.
Code
def kalman_update(pred: float, measure: float, gain: float) -> float: return pred + gain * (measure - pred)
print(kalman_update(12.0, 12.5, 0.5)) # -> 12.25print(kalman_update(12.0, 12.5, 0.9)) # -> 12.45Watch Out For
Constant velocity is a lie during maneuvers
The model assumes velocity barely changes, so a sharp turn or sudden stop leaves predictions flying straight while the object is gone. The filter recovers once measurements return, but IDs may already have switched. If tracks swap at every turn, raise process noise or move to an interacting-multiple-model setup.
Tuning Q and R by feel
(motion surprise) and (sensor slop) decide everything, and hand-tuned values from one camera fail on another. Measure from your detector's real jitter on static scenes, then set from the fastest accelerations you must follow. Numbers from data beat numbers from vibes.
The Quick Version
- The Kalman filter predicts position from velocity, then corrects with the measurement weighted by the gain.
- Prediction with measurement at gain lands exactly at .
- Uncertainty grows while coasting, so occlusion handling degrades honestly instead of snapping.
- Optimal only for linear motion with Gaussian noise; maneuvers need more process noise.
- The motion core of SORT-style trackers; nonlinear cases graduate to particle filters.