Skip to content
AI360Xpert
Glossary
Definition

K-Anonymity

A dataset property where every combination of identifying-ish columns is shared by at least k people, so no row stands alone.

Reached by coarsening quasi-identifiers — columns that don't name anyone alone but combine into an identifier. A postcode district becomes a region, a date of birth becomes a five-year band, and you generalise until every group holds at least kk people.

The arithmetic explains why it's needed. Three thousand postcode districts, 36,500 birth dates and two values for sex give 219 million combinations for 67 million people, so about 85% of the occupied combinations hold exactly one individual. Dropping the name column achieves nothing.

Two known holes. Homogeneity: if all kk people in a group share the same diagnosis, you've learned it without identifying anyone — which is what l-diversity patches. And composition: two separately k-anonymous releases can be intersected to break both, which nothing in the framework prevents.