Skip to content
AI360Xpert
Beta

GrabCut Interactive Segmentation

Draw one loose box around the object and GrabCut learns what foreground and background look like, then refines the border itself by repeating graph cut with color models.

One loose box seeds colour models that retrain and re-cut in a loop until the border settles.
One loose box seeds colour models that retrain and re-cut in a loop until the border settles.

Why Does This Exist?

Graph cut finds the optimal binary labeling but demands per-pixel evidence: somebody must say which pixels are foreground and which are background. Painting exact seeds is slow, expert work. GrabCut (Rother, Kolmogorov, and Blake, 2004) asks for one loose rectangle instead and manufactures the rest itself, which made it the default photo cutout tool in editors for a decade.

The prerequisite is the graph-cut machinery: t-links, n-links, and min-cut optimality. This page covers the color models, the tri-map, and the iteration loop. Fully automatic multi-class labeling belongs to semantic segmentation instead.

Think of It Like This

A bouncer with a improving description

A bouncer gets one rough instruction, "everyone inside the rope line might belong, everyone outside definitely does not," plus a description of the guest of honor that sharpens every round. Each round the bouncer re-sorts the crowd, studies the two groups afresh, and redraws the line. After a few rounds the line sits exactly on the guest list.

It stops holding when the description cannot discriminate. If the guest wears the same uniform as the crowd, camouflage in vision terms, no amount of re-sorting separates them and the bouncer needs extra information, which is why GrabCut accepts corrective strokes.

How It Actually Works

The tri-map from one box

Everything outside the user box is marked sure background. Everything inside starts as unknown, tentatively foreground. That three-way labeling (sure background, probable foreground, probable background after the first cut) is the tri-map, and it is the only human input the algorithm ever sees unless corrections arrive later.

Color models both sides learn

Foreground and background colors are each modeled as a Gaussian mixture with five components in RGB space. Unknown pixels get assigned to their most likely component, the mixtures refit to their assigned pixels, and a graph cut re-labels every pixel using the fresh mixtures as its data term. Assign, learn, cut, repeat, usually about five rounds, until labels stop changing. The smoothness term keeps borders clean while the learned colors pull interiors, and a final border-matting pass estimates partial alpha along the edge for hair-soft transitions.

Take a 100×100100 \times 100 photo with a 60×6060 \times 60 box: 6,4006{,}400 pixels start as sure background and 3,6003{,}600 as unknown. A typical run might settle with roughly 2,1002{,}100 of the unknown labeled foreground, the mixtures having learned, say, skin and shirt tones on one side and grass tones on the other. The exact split depends on the image; the mechanism does not.

Code

The standard OpenCV call sequence. The mask values below are the API contract, not measured output:

import cv2import numpy as np
image = cv2.imread("person.jpg")mask = np.zeros(image.shape[:2], np.uint8)rect = (10, 10, 200, 300)  # one loose box around the subjectbgd = np.zeros((1, 65), np.float64)fgd = np.zeros((1, 65), np.float64)
cv2.grabCut(image, mask, rect, bgd, fgd, 5, cv2.GC_INIT_WITH_RECT)# mask now holds GC_BGD (0), GC_FGD (1), GC_PR_BGD (2), GC_PR_FGD (3)foreground = np.where((mask == 1) | (mask == 3), 255, 0).astype("uint8")

Five iterations from a rectangle init; extra user strokes would switch the mode to mask init for corrections.

Watch Out For

Camouflage bleeding across the border

Symptom: the cut leaks into background regions whose colors match the subject, like a green shirt against foliage, and more iterations make it worse as the mixtures absorb the leak. Color is the only evidence GrabCut learns. Fix it with corrective strokes (sure-foreground inside the subject, sure-background in the leak) and rerun in mask mode.

Losing hair, wires, and thin straps

Symptom: fine structures vanish because the smoothness term charges per boundary pixel and thin shapes are nearly all boundary. The energy genuinely prices them out. Fix it by over-protecting thin parts with foreground strokes, or accept the limit and finish sub-pixel strands with border matting or manual alpha work.

The Quick Version

  • GrabCut needs one bounding box: outside is sure background, inside starts unknown.
  • Five-component Gaussian mixtures learn both sides' colors, then graph cut re-labels with those models as evidence.
  • Assign, refit, and cut iterate about five times until the border stops moving.
  • Border matting finishes the edge with partial alpha for soft transitions like hair.
  • Camouflaged colors and hair-thin structures are the two classic failures; corrective strokes fix the first.