Watershed Segmentation
Read brightness as hills and valleys, flood the valleys with water, and draw borders where floods from different valleys meet. Touching objects finally split apart.
Why Does This Exist?
Thresholding merges every touching coin into one blob, and region growing happily floods across the touching point too, because neither knows two objects are present. Cell counting, grain sizing, and coin counting all need the opposite: find the touching point and cut there. Watershed does exactly that by reading the image as terrain and splitting where separate floods collide.
It leans on one prerequisite: the distance transform, which turns a binary mask into a height map peaking at object centers. That operation belongs to morphological methods. This page covers the flooding idea, markers, and the over-segmentation trap.
Think of It Like This
Rain flooding mountain valleys
Rain falls on mountain terrain and each valley collects its own lake. The lakes rise until neighbouring lakes threaten to merge, and at every threatened meeting point engineers build a dam. The dams are the watershed lines, and they sit exactly on the ridges between valleys.
It stops holding on noisy terrain. Every pothole becomes its own tiny valley with its own tiny lake, so unmarked watershed builds thousands of dams across what is really one hillside. Real halves need surveyed lake sites, which is what markers supply.
How It Actually Works
Valleys, floods, and dams
Treat intensity (or gradient magnitude) as elevation. Puncture each regional minimum and flood from below at a uniform rate. When two floods from different minima would merge, build a dam at the meeting pixels. When the water reaches the peaks, the dams form closed one-pixel borders around every catchment basin.
The distance transform trick
Floods need valleys inside each object, but a flat binary mask has none. The distance transform fixes that: replace every foreground pixel with its distance to the nearest background pixel. Take the 1D mask , two touching runs of three. Distances read , with peaks of at index and index . Two peaks, two objects, and the zero valley between them is where the dam goes. In 2D the same peaks sit at cell centers, far from the touching points.
Markers tame the flood
Raw watershed floods from every local minimum, and noise supplies thousands, which is the famous over-segmentation: a single cell shatters into mosaic. Marker-based watershed floods only from supplied minima: sure-foreground markers from thresholded distance peaks, sure-background from dilated background, and everything between marked unknown. Fewer minima, fewer basins, and each basin corresponds to a real object.
Code
One-dimensional distance transform and marker peaks on the fixture mask:
mask = [0, 1, 1, 1, 0, 1, 1, 1, 0]
def distance_1d(m): d = [0] * len(m) for i, v in enumerate(m): d[i] = 0 if v == 0 else (d[i - 1] + 1 if i else 1) for i in range(len(m) - 2, -1, -1): if m[i]: d[i] = min(d[i], d[i + 1] + 1) return d
d = distance_1d(mask)peaks = [i for i in range(1, len(d) - 1) if d[i] > d[i - 1] and d[i] >= d[i + 1]]print(f"distances {d}, markers at {peaks}")# -> distances [0, 1, 2, 1, 0, 1, 2, 1, 0], markers at [2, 6]Two markers at indices and mark two objects, matching the hand trace, and the dam falls at the zero valley between them.
Watch Out For
Running watershed with no markers
Symptom: one cell comes back tiled into dozens of tiny basins, each with its own border. Every noise minimum became a flood source. Never run the raw transform on real images. Always supply markers, typically sure-foreground from the distance transform plus sure-background from dilation, and treat the unmarked gap as unknown.
Merging objects with lazy background markers
Symptom: two close but separate cells share one basin because the background marker never reached the gap between them. The sure-background region must genuinely separate the objects. Dilate the background aggressively and check the unknown zone forms a continuous ring around each foreground marker before flooding.
The Quick Version
- Watershed reads the image as terrain, floods from minima, and dams the meeting lines between separate floods.
- The distance transform creates one peak per object center, turning touching blobs into separate valleys.
- Markers restrict flooding to real minima: sure foreground, sure background, and an unknown ring between them.
- Without markers, noise minima cause severe over-segmentation into hundreds of meaningless basins.
- It splits touching objects that thresholding and region growing merge, but it cannot invent boundaries in flat, textureless gaps.