Visual explainer
Transformers
A visual walkthrough of the Transformer architecture, from self-attention and positional encodings to the full encoder-decoder stack.
Traditional networks process text sequentially, creating a bottleneck and struggling with long-term dependencies. Transformers break this limit by processing everything at once. Their secret is Self-Attention — a mechanism where every token looks at every other token to figure out its contextual meaning. In the example above, the word "bank" heavily attends to "river", clarifying that it's a river bank, not a financial institution.
Knowing the Order
Because self-attention reads all tokens in parallel, the model has no inherent sense of sequence. If you scramble a sentence, a naive parallel model would output the exact same final representations. To fix this, Transformers inject Positional Encodings — unique mathematical signals — directly into the word embeddings. This addition creates position-aware vectors, letting the network know exactly where each word sits in the text.
The End-to-End Architecture
The complete Transformer relies on a tightly coupled Encoder-Decoder structure:
- The Encoder takes the input sequence and passes it through layers of Self-Attention and Feed-Forward networks. This builds a deep, contextual representation of the entire text.
- The Decoder uses that learned representation via Cross-Attention to generate the final output, one token at a time. It uses Masked Attention to ensure it can only look at tokens it has already generated, preventing it from "cheating" by looking into the future.
What to Read Next
- Attention MechanismInstead of compressing an entire sequence into one vector, attention lets the decoder dynamically weigh and blend the full input for each output step.
- BERT vs GPTTwo fundamentally different ways to read text: understanding the whole context vs generating the next word.
- How A Decision Head WorksVisualizing how a System 1 router replaces autoregressive text generation with a single-pass encoder and typed decision heads.