Skip to content
AI360Xpert
Beta

Shi-Tomasi Corner Detector

Shi-Tomasi keeps corners whose weaker gradient direction is still strong, which picks points a tracker can actually follow.

Shi-Tomasi scores a window by its smaller eigenvalue and keeps points above threshold.
Shi-Tomasi scores a window by its smaller eigenvalue and keeps points above threshold.

Why Does This Exist?

Harris finds corners, but trackers need corners that survive from frame to frame. A point with one strong gradient direction and one weak direction looks like a corner at one threshold yet slides under noise. Optical flow then drifts.

Shi and Tomasi asked which points are good to track and answered with the weaker eigenvalue. If the smaller eigenvalue is large, both directions carry signal, so Lucas-Kanade can solve for motion. Start with Harris for the shared tensor, then read this as the stricter rule.

Think of It Like This

Two good grips on a box

You must carry a box with both hands. One strong grip and one loose grip means the box twists. Two firm grips mean you control it.

Harris asks whether the combined grip energy is high. Shi-Tomasi checks your weaker hand alone. The analogy stops at tracking: eigenvalues do not get tired, but image noise plays the role of a slippery glove.

How It Actually Works

Build the same structure tensor MM from IxI_x and IyI_y under a Gaussian window. Let its eigenvalues be λ1≥λ2\lambda_1 \ge \lambda_2. The Shi-Tomasi score is simply:

RST=λ2=min⁡(λ1,λ2)R_{ST} = \lambda_2 = \min(\lambda_1, \lambda_2)

Keep the pixel when RSTR_{ST} exceeds a threshold and survives non-maximum suppression. OpenCV's goodFeaturesToTrack adds two practical controls: a quality level relative to the best score, and a minimum distance between kept points.

Worked numbers

Suppose a window gives M=[4.01.01.03.0]M = \begin{bmatrix} 4.0 & 1.0 \\ 1.0 & 3.0 \end{bmatrix}. Trace is 7.07.0 and determinant is 11.011.0. Eigenvalues solve λ2−7λ+11=0\lambda^2 - 7\lambda + 11 = 0, so λ1≈4.79\lambda_1 \approx 4.79 and λ2≈2.21\lambda_2 \approx 2.21. Score RST=2.21R_{ST} = 2.21. With threshold 1.51.5, you keep it. With threshold 3.03.0, you drop it even though Harris with k=0.04k = 0.04 gives R=11.0−1.96=9.04R = 11.0 - 1.96 = 9.04 and would keep it. That gap is the point: Shi-Tomasi rejects lopsided windows.

Like Harris, this is rotation invariant and scale sensitive. It pairs with pyramidal Lucas-Kanade for video because the kept points constrain motion in both axes.

Watch Out For

Quality level copied across scenes

A quality level of 0.010.01 keeps hundreds of points indoors and almost none on a white wall. Set it per scene contrast, and enforce minimum distance, or points cluster on the one textured poster while the rest of the frame has nothing to track.

Using it for wide baseline matching

Shi-Tomasi has no descriptor and no scale handling. It suits consecutive video frames, not day versus night photo matching. For that jump to feature descriptors and a real matcher.

The Quick Version

  • Shi-Tomasi scores each window by min⁡(λ1,λ2)\min(\lambda_1, \lambda_2) of the structure tensor.
  • A large smaller eigenvalue means gradients support motion solving in both axes.
  • Quality level and minimum distance turn scores into spread out trackable points.
  • Rotation invariant but not scale invariant.
  • Best partner is pyramidal Lucas-Kanade on video, not wide baseline photo pairs.