Skip to content
AI360Xpert
Core ML

Visual explainer

Tensor Processing Units

How TPUs trade general-purpose flexibility for massive, hardware-accelerated matrix multiplication using a systolic array.

CPU fetches memory per step, while TPU flows data through a matrix unit.
CPU fetches memory per step, while TPU flows data through a matrix unit.

Every time a CPU or GPU performs a calculation, it fetches data from memory or a register, does the math, and writes it back. This memory traffic limits speed. A TPU bypasses this bottleneck by passing data directly from one arithmetic unit to the next.

The Systolic Array

Activations flow right, partial sums flow down through the systolic array.
Activations flow right, partial sums flow down through the systolic array.

The heart of a TPU is the systolic array—a vast grid of Multiply-Accumulate (MAC) units. Neural network weights are loaded into the grid and stay put. Input data (activations) flows continuously from left to right, while partial sums cascade downward. No intermediate memory fetches are required.

Where It Breaks

Branching and sparse operations stall the array; it needs predictable, dense math.
Branching and sparse operations stall the array; it needs predictable, dense math.

TPUs are rigid. If your model relies heavily on dynamic control flow (like complex if/else branching) or sparse matrices with mostly zeros, the systolic array stalls. The hardware is designed exclusively to keep dense matrix multiplications running at full capacity.

The Quick Version

  • Memory wall: Traditional chips waste time fetching data for every calculation.
  • Direct flow: TPUs pass results directly between ALUs without returning to memory.
  • Heart of the chip: A systolic array performs massive, synchronized matrix math.
  • The tradeoff: They are terrible at branching logic and sparse data structures.