Skip to content
AI360Xpert
Beta

Correlation in Image Processing

Slide a small template over a photo and score every spot by how well they line up. High scores mark matches, which is how template matching finds objects.

Correlation centers the template on each pixel and sums direct overlaps, so the brightest response marks the best match.
Correlation centers the template on each pixel and sums direct overlaps, so the brightest response marks the best match.

Why Does This Exist?

Say you need to find a logo in a scanned page or track a marker across video frames. You know exactly what the target looks like, so learning a detector is overkill. You want to slide the known pattern over the image and ask, at every spot, how well they agree.

Correlation is that direct comparison. It lives inside the broader family of spatial filtering: same sliding machinery, but the template is used exactly as drawn, with no flipping. Deep learning libraries call their sliding operation convolution, yet they skip the flip too, so they are really doing correlation.

Think of It Like This

Tracing a stencil over a map

Cut a stencil in the shape of a lake and slide it over a paper map. At most positions the cutout shows forest or roads, a poor fit. At one position the lake on the map fills the cutout exactly, and you stop. The template is your stencil, the correlation score is how much of the cutout matches, and the peak is where you stop sliding. The analogy ends where brightness does: a stencil ignores ink darkness, while raw correlation rewards bright regions even when shapes disagree.

How It Actually Works

Correlation centers a template TT of size m×nm \times n on each image pixel (x,y)(x, y) and sums the direct products. T(i,j)T(i, j) is the template value at row ii and column jj, and II is the image:

R(x,y)=∑i∑jT(i,j)⋅I(x+i,y+j)R(x, y) = \sum_{i} \sum_{j} T(i, j) \cdot I(x + i, y + j)

R(x,y)R(x, y) is the response map. Bright peaks in RR mark positions where the image looks like the template. True mathematical convolution flips TT both ways before this sum; when the kernel is symmetric the two agree, and when it is not, they differ. That flip is the only distinction.

Worked example

Take a 3×33 \times 3 patch and an asymmetric 2×22 \times 2 template:

I=[123456789],T=[1234]I = \begin{bmatrix} 1 & 2 & 3 \\ 4 & 5 & 6 \\ 7 & 8 & 9 \end{bmatrix}, \quad T = \begin{bmatrix} 1 & 2 \\ 3 & 4 \end{bmatrix}

Correlation at the top-left overlap multiplies directly: 1⋅1+2⋅2+4⋅3+5⋅4=1+4+12+20=371 \cdot 1 + 2 \cdot 2 + 4 \cdot 3 + 5 \cdot 4 = 1 + 4 + 12 + 20 = 37. Convolution first flips TT to [4321]\begin{bmatrix} 4 & 3 \\ 2 & 1 \end{bmatrix}, giving 1⋅4+2⋅3+4⋅2+5⋅1=4+6+8+5=231 \cdot 4 + 2 \cdot 3 + 4 \cdot 2 + 5 \cdot 1 = 4 + 6 + 8 + 5 = 23. Same patch, same numbers, different answer, and the flip explains all of it.

Normalized cross-correlation fixes brightness bias by subtracting local means and dividing by local contrast before scoring, so a bright but shapeless region can no longer outscore a dim but exact match.

Watch Out For

Bright regions win without normalization

Raw correlation multiplies intensities directly, so a bright patch of sky can outscore a dim exact match of your template. The symptom is peaks glued to the brightest image areas no matter the pattern. Switch to normalized cross-correlation, which centers and scales each window before comparing.

OpenCV filter2D does not flip

OpenCV's filter2D performs correlation even though many call it convolution, while signal-processing convolve flips. With a symmetric blur kernel you will never notice. With an asymmetric edge or motion kernel the response lands mirrored, so check whether your function flips before reasoning about direction.

The Quick Version

  • Correlation scores every image position by direct overlap with an unflipped template.
  • Peaks in the response map mark the best matches, which powers template matching.
  • Convolution differs only by flipping the kernel both ways first.
  • Normalize by local mean and contrast when brightness varies across the image.
  • Learned CNN kernels are trained in place, so the flip question never arises there.