Harris Corner Detector
Harris turns every pixel into a corner score from local gradients, so you can keep points that stay put when the camera shifts.
Why Does This Exist?
Tracking and stitching need points you can find again. Flat patches all look alike, and edges slide along their own direction. Corners lock in two directions at once, so they survive small viewpoint changes.
Harris gave this a cheap test. It uses only first derivatives in a small window, no explicit eigenvalue solver, and one response value per pixel. Read the overview in corner and keypoint detection first, then come back for the exact score.
Think of It Like This
A tent corner in the wind
Picture a tent stretched on a lawn. Push the middle of a flat sheet and you cannot tell where you pushed. Push along a seam and the fabric slides. Pull the corner stake and the whole shape answers at once.
Harris looks for tent stakes in the image. Flat means no pull in any direction. Edge means pull in one direction. Corner means pull in every direction. The analogy stops at physics: pixels do not stretch, only gradient energy counts.
How It Actually Works
Start with horizontal and vertical derivatives and at each pixel, usually from Sobel filters. Around each pixel, form the structure tensor over a Gaussian window :
summarizes gradient energy in the neighborhood. Let its eigenvalues be and . Both small means flat. One large and one small means edge. Both large means corner.
The response shortcut
Eigenvalues cost too much per pixel, so Harris uses determinant and trace:
Here and , with near . You threshold and keep local maxima.
Worked numbers
Take a window where , , and . Then and trace . With , . That large positive marks a corner. An edge with sums , , and gives , trace , and , so it is rejected.
Harris is rotation invariant because rotating the image rotates both eigenvectors together. It is not scale invariant: zoom in and the window sees a curve instead of a corner.
Watch Out For
Threshold picked on one image
A threshold that looks clean on a textured poster fires everywhere on gravel and nowhere on soft faces. Tune and the threshold on your own contrast range, then add non-maximum suppression. A raw threshold without local maxima gives clusters, not points.
Expecting scale invariance
Harris has no pyramid, so matching a close-up against a wide shot fails. If scale changes, move to scale invariant features or a pyramid detector instead of raising the threshold.
The Quick Version
- Harris builds a structure tensor from and in each window.
- Response avoids an eigenvalue solver.
- Large positive means corner, negative means edge, small means flat.
- Rotation invariant but scale sensitive, so it suits tracking at fixed distance.
- Threshold plus non-maximum suppression turns scores into usable points.