Skip to content
AI360Xpert
Beta

Pixels

A pixel is one cell of the image grid holding a single light measurement. Millions of them tile into rows and columns, and their spacing sets the finest detail any image can carry.

A pixel is one addressed cell of the image grid, and its row, column and channel values are the smallest unit any vision algorithm can see.
A pixel is one addressed cell of the image grid, and its row, column and channel values are the smallest unit any vision algorithm can see.

Why Does This Exist?

Everything downstream reads pixels: filters slide over them, detectors draw boxes in their coordinates, and networks ingest them as tensor entries. When a bounding box lands one cell off or a crop coordinate flips, the bug is always in pixel addressing, never in the model.

A pixel is the smallest addressable unit of a digital image: one grid position (j,i)(j, i) holding one measurement per channel. This page covers addressing, size, and what a pixel is not. The grid as a whole, with its resolution and channel promises, lives in image fundamentals.

Think of It Like This

Squares on graph paper

Take a sheet of graph paper and fill each square with a single crayon shade. From across the room the squares merge into a drawing. Up close there is nothing but flat squares with hard edges.

That is a pixel: one square, one shade, addressed by column and row. Zooming an image on screen is just drawing each square bigger.

The analogy stops here: sensor pixels have gaps between them and uneven sensitivity, so a real pixel is a small light bucket with dead borders, not a perfect square of color.

How It Actually Works

A pixel at column jj and row ii stores a vector with one entry per channel. In an 8-bit RGB image that vector is (R,G,B)(R, G, B) with each entry an integer from 00 to 255255. In a grayscale image it is a single brightness value. The pixel has no width in world units: it is a sample, and its footprint on the scene depends on the optics and distance (see image formation).

Coordinates start top-left and y grows down

Image coordinates put the origin (0,0)(0, 0) at the top-left corner, with xx running right along columns and yy running down along rows. So pixel (x,y)(x, y) is row yy, column xx. This is upside down relative to math plots, and every flipped overlay you have ever debugged came from forgetting it.

Memory lays rows end to end. In a row-major array of width WW, pixel (x,y)(x, y) sits at linear index y×W+xy \times W + x.

Worked example: finding one pixel in memory

Take a grayscale image with W=640W = 640. The pixel at x=100x = 100, y=50y = 50 is the 100th entry of the 50th row, at index 50×640+100=32,10050 \times 640 + 100 = 32{,}100. Change that one entry and exactly one dot on screen changes.

Pixels are not dots of paint

On a display, one image pixel typically drives three colored subpixels (red, green, blue stripes) whose light blends in your eye. On a sensor, one pixel is a photosite that counts photons through a color filter. The stored number is a measurement with noise, not a fact about the world.

Code

width = 640  # row length in pixelsx, y = 100, 50index = y * width + x  # row-major address of pixel (x, y)print(index)# -> 32100

Watch Out For

Y-down coordinates flip your overlays

Math libraries plot yy growing upward; images store yy growing downward. Drawing a detection box with math convention on an image puts it mirrored about the horizontal middle. Convert once at the boundary: yimage=H−1−ymathy_{image} = H - 1 - y_{math}.

Off-by-one crops from confusing corners with sizes

A box stored as (x,y,w,h)(x, y, w, h) needs slice img[y:y+h, x:x+w], not img[y:y+h-1, x:x+w-1]: Python slices already exclude the end. The minus-one habit silently shaves a row and column off every training crop.

The Quick Version

  • A pixel is one grid cell holding one measurement per channel: a triple in RGB, a single value in grayscale.
  • Coordinates start at the top-left with yy growing downward, opposite to math plots.
  • In row-major memory, pixel (x,y)(x, y) of a width-WW image sits at index y×W+xy \times W + x.
  • Display subpixels and sensor photosites are physical hardware; the stored pixel is just their number.
  • Pixel (100, 50) of a 640-wide image lives at index 32,100.