Visual explainer
Tokenization
How raw text is split into discrete chunks before being processed by neural networks, and why this mechanism creates strange failure modes.
Raw strings are continuous text, but neural networks require discrete numerical inputs. A tokenizer bridges this gap by slicing the text into chunks—words, subwords, or characters—and mapping each piece to a unique integer ID from a fixed vocabulary.
Subword Merging (BPE)
Modern language models use subword tokenization, like Byte-Pair Encoding (BPE). It starts with individual characters and iteratively merges the most frequently adjacent pairs. Common words become a single token, while rare or complex words remain split across multiple smaller tokens, balancing sequence length against vocabulary size.
Where It Breaks
Because models only "see" the integer IDs, they are entirely blind to the characters inside those chunks. If a word like "strawberry" is split irregularly, the model cannot easily count its letters. This is the root cause of many classic LLM failures in spelling, character-level math, and rhyme generation.
The Quick Version
- Discrete inputs: Networks need numbers, so text must be chunked.
- Subword merges: Frequent character pairs combine into efficient single tokens.
- Blind spots: LLMs cannot natively see inside their own tokens.
- Failure modes: Spelling and arithmetic often break due to unexpected token boundaries.
What to Read Next
- Attention MechanismInstead of compressing an entire sequence into one vector, attention lets the decoder dynamically weigh and blend the full input for each output step.
- TransformersA visual walkthrough of the Transformer architecture, from self-attention and positional encodings to the full encoder-decoder stack.