Skip to content
AI360Xpert
Core ML

Visual explainer

Tokenization

How raw text is split into discrete chunks before being processed by neural networks, and why this mechanism creates strange failure modes.

Models compute on numbers, not text.
Models compute on numbers, not text.

Raw strings are continuous text, but neural networks require discrete numerical inputs. A tokenizer bridges this gap by slicing the text into chunks—words, subwords, or characters—and mapping each piece to a unique integer ID from a fixed vocabulary.

Subword Merging (BPE)

Subword tokenization merges frequent characters into chunks.
Subword tokenization merges frequent characters into chunks.

Modern language models use subword tokenization, like Byte-Pair Encoding (BPE). It starts with individual characters and iteratively merges the most frequently adjacent pairs. Common words become a single token, while rare or complex words remain split across multiple smaller tokens, balancing sequence length against vocabulary size.

Where It Breaks

Weird token splits make LLMs bad at math and spelling.
Weird token splits make LLMs bad at math and spelling.

Because models only "see" the integer IDs, they are entirely blind to the characters inside those chunks. If a word like "strawberry" is split irregularly, the model cannot easily count its letters. This is the root cause of many classic LLM failures in spelling, character-level math, and rhyme generation.

The Quick Version

  • Discrete inputs: Networks need numbers, so text must be chunked.
  • Subword merges: Frequent character pairs combine into efficient single tokens.
  • Blind spots: LLMs cannot natively see inside their own tokens.
  • Failure modes: Spelling and arithmetic often break due to unexpected token boundaries.