Skip to content
AI360Xpert
Beta

Difference of Gaussians

DoG subtracts two nearby blurs to spotlight structures at one size, which gives SIFT a fast way to hunt keypoints across scales.

DoG subtracts adjacent Gaussian blurs and finds extrema across space and scale.
DoG subtracts adjacent Gaussian blurs and finds extrema across space and scale.

Why Does This Exist?

A Laplacian pyramid finds scale tuned structure but costs a lot of separable second derivatives. SIFT needed something cheaper that still peaks on blobs and corners at each size. Subtracting two Gaussian blurs approximates the Laplacian with only blurs and subtraction.

That trick powers the SIFT detector. Build a scale space, subtract neighbors, then look for extrema in space and scale. For the blob view first, read blob detection.

Think of It Like This

Two frosted panes of glass

Look through one lightly frosted pane, then a slightly more frosted pane. Each view alone looks blurry. Hold the difference between the views and only structures at one size pop out.

Nearby Gaussian widths play the frosted panes. Their difference cancels both finer grain and larger shading, leaving the middle size. The analogy stops at math: DoG approximates σ2∇2G\sigma^2 \nabla^2 G, it does not equal it.

How It Actually Works

Blur image II with Gaussians G(x,y,kσ)G(x,y,k\sigma) and G(x,y,σ)G(x,y,\sigma), then subtract:

D(x,y,σ)=L(x,y,kσ)−L(x,y,σ)D(x,y,\sigma) = L(x,y,k\sigma) - L(x,y,\sigma)

where LL is the blurred image. SIFT stacks several DD images per octave, downsamples, and repeats. Each pixel is compared to 2626 neighbors: 88 in its own scale plus 99 above and 99 below. A pixel larger or smaller than all 2626 is a candidate keypoint.

Worked numbers

Take k=2≈1.414k = \sqrt{2} \approx 1.414 and base σ=1.6\sigma = 1.6, the SIFT defaults. Adjacent blurs use σ=1.6\sigma = 1.6 and kσ≈2.26k\sigma \approx 2.26. A disk of radius 55 peaks when the DoG scale matches it, while a pixel with center value 4040 against all 2626 neighbors below 3030 becomes a bright extremum. Later SIFT stages reject low contrast points with ∣D∣<0.03|D| < 0.03 and edge-like points by Hessian ratio, but DoG supplies the candidates.

Watch Out For

DoG extrema are not finished keypoints

DoG fires on noise, edges, and low contrast texture. Without contrast filtering, edge suppression, and subpixel refinement, you match junk. Never feed raw DoG extrema into a matcher.

Octave count copied blindly

Too few octaves miss large objects, too many waste time on tiny thumbnails. Match octave count and scales per octave to your smallest and largest target sizes.

The Quick Version

  • DoG equals the difference of two nearby Gaussian blurs.
  • It approximates scale normalized LoG at a fraction of the cost.
  • SIFT finds 2626-neighbor extrema over space and scale.
  • Contrast and edge tests prune the raw candidates.
  • Octaves extend detection from fine detail to large structure.