Scale-Invariant Feature Transform
SIFT finds the same points at any zoom and rotation, then fingerprints each with a 128 number vector that survives viewpoint change.
Why Does This Exist?
Harris breaks under zoom, and raw patches break under rotation. Real photo pairs mix both: a tourist photo versus a wide plaza shot. SIFT was built to survive scale, rotation, and moderate lighting change in one pipeline.
It chains four ideas: Difference of Gaussians extrema for scale, subpixel refinement, orientation voting for rotation, and pooled gradient histograms for identity. That chain made panorama stitching and object retrieval practical before deep features.
Think of It Like This
A lighthouse with its own compass
A lighthouse must be found from near or far, day or fog. First you spot its flash among rocks. Then you read its compass bearing. Then you record its flash pattern.
SIFT does the same. DoG spots the flash across scales. Orientation voting reads the bearing. The 128 value vector records the pattern. The analogy stops at weather: heavy perspective and day night change still defeat it.
How It Actually Works
Detection finds -neighbor DoG extrema, refines each to subpixel accuracy with a Taylor fit, drops points with , and drops edge responses when Hessian ratio . Orientation assignment builds a -bin gradient histogram around each survivor and creates one keypoint per peak above percent of the max.
The descriptor
Each keypoint samples a rotated window, splits it into cells of , and votes gradients into orientation bins per cell. That gives floats, normalized to unit length, clamped at , and renormalized to resist nonlinear lighting.
Worked numbers
A keypoint with dominant orientation degrees rotates its sampling grid by degrees before pooling. Suppose one cell holds gradient mass across its bins. The in bin three dominates, so that direction shapes of the final numbers. Two true views might land at Euclidean distance after normalization, while a false pair lands near , which is why the ratio test separates them.
Watch Out For
Running default octaves on tiny images
SIFT doubles the input image by default to catch fine detail. On small thumbnails that just magnifies blur and multiplies junk points. Disable the doubling or raise contrast threshold for thumbnails and video frames.
Expecting real time on CPU
A float descriptor per point plus DoG pyramids costs real milliseconds. SIFT suits offline stitching and retrieval. For live tracking use ORB or learned binary features.
The Quick Version
- DoG extrema plus refinement give repeatable points across scale.
- Orientation voting gives rotation invariance.
- Sixteen cells times eight bins give the 128 float descriptor.
- Normalization plus clamping handles lighting shifts.
- Accurate and slow, ideal for stitching and retrieval rather than live tracking.