Skip to content
AI360Xpert
Beta

Scale-Invariant Feature Transform

SIFT finds the same points at any zoom and rotation, then fingerprints each with a 128 number vector that survives viewpoint change.

SIFT detects DoG extrema, assigns orientation, and pools gradients into 128 values.
SIFT detects DoG extrema, assigns orientation, and pools gradients into 128 values.

Why Does This Exist?

Harris breaks under zoom, and raw patches break under rotation. Real photo pairs mix both: a tourist photo versus a wide plaza shot. SIFT was built to survive scale, rotation, and moderate lighting change in one pipeline.

It chains four ideas: Difference of Gaussians extrema for scale, subpixel refinement, orientation voting for rotation, and pooled gradient histograms for identity. That chain made panorama stitching and object retrieval practical before deep features.

Think of It Like This

A lighthouse with its own compass

A lighthouse must be found from near or far, day or fog. First you spot its flash among rocks. Then you read its compass bearing. Then you record its flash pattern.

SIFT does the same. DoG spots the flash across scales. Orientation voting reads the bearing. The 128 value vector records the pattern. The analogy stops at weather: heavy perspective and day night change still defeat it.

How It Actually Works

Detection finds 2626-neighbor DoG extrema, refines each to subpixel accuracy with a Taylor fit, drops points with ∣D∣<0.03|D| < 0.03, and drops edge responses when Hessian ratio r>10r > 10. Orientation assignment builds a 3636-bin gradient histogram around each survivor and creates one keypoint per peak above 8080 percent of the max.

The descriptor

Each keypoint samples a 16×1616 \times 16 rotated window, splits it into 1616 cells of 4×44 \times 4, and votes gradients into 88 orientation bins per cell. That gives 16×8=12816 \times 8 = 128 floats, normalized to unit length, clamped at 0.20.2, and renormalized to resist nonlinear lighting.

Worked numbers

A keypoint with dominant orientation 3535 degrees rotates its sampling grid by −35-35 degrees before pooling. Suppose one cell holds gradient mass [0,2,9,4,1,0,0,0][0, 2, 9, 4, 1, 0, 0, 0] across its 88 bins. The 99 in bin three dominates, so that direction shapes 88 of the final 128128 numbers. Two true views might land at Euclidean distance 0.30.3 after normalization, while a false pair lands near 1.11.1, which is why the ratio test separates them.

Watch Out For

Running default octaves on tiny images

SIFT doubles the input image by default to catch fine detail. On small thumbnails that just magnifies blur and multiplies junk points. Disable the doubling or raise contrast threshold for thumbnails and video frames.

Expecting real time on CPU

A 128128 float descriptor per point plus DoG pyramids costs real milliseconds. SIFT suits offline stitching and retrieval. For live tracking use ORB or learned binary features.

The Quick Version

  • DoG extrema plus refinement give repeatable points across scale.
  • Orientation voting gives rotation invariance.
  • Sixteen cells times eight bins give the 128 float descriptor.
  • Normalization plus clamping handles lighting shifts.
  • Accurate and slow, ideal for stitching and retrieval rather than live tracking.