Visual explainer
Tensor Processing Units
How TPUs trade general-purpose flexibility for massive, hardware-accelerated matrix multiplication using a systolic array.
Every time a CPU or GPU performs a calculation, it fetches data from memory or a register, does the math, and writes it back. This memory traffic limits speed. A TPU bypasses this bottleneck by passing data directly from one arithmetic unit to the next.
The Systolic Array
The heart of a TPU is the systolic array—a vast grid of Multiply-Accumulate (MAC) units. Neural network weights are loaded into the grid and stay put. Input data (activations) flows continuously from left to right, while partial sums cascade downward. No intermediate memory fetches are required.
Where It Breaks
TPUs are rigid. If your model relies heavily on dynamic control flow (like complex if/else branching) or sparse matrices with mostly zeros, the systolic array stalls. The hardware is designed exclusively to keep dense matrix multiplications running at full capacity.
The Quick Version
- Memory wall: Traditional chips waste time fetching data for every calculation.
- Direct flow: TPUs pass results directly between ALUs without returning to memory.
- Heart of the chip: A systolic array performs massive, synchronized matrix math.
- The tradeoff: They are terrible at branching logic and sparse data structures.
What to Read Next
- TransformersA visual walkthrough of the Transformer architecture, from self-attention and positional encodings to the full encoder-decoder stack.
- Attention MechanismInstead of compressing an entire sequence into one vector, attention lets the decoder dynamically weigh and blend the full input for each output step.