Visual explainer
Word Embeddings
How computers represent words as dense vectors, allowing semantic meaning to be treated as math.
Computers require numbers, not letters, to process text. Early machine learning approaches assigned a unique, completely isolated ID to every word, creating sparse vectors with a single '1' and thousands of zeros. This one-hot encoding perfectly separates words, but it entirely destroys semantic meaning. Because every word vector is orthogonal to every other word vector, "cat" is exactly as far from "dog" as it is from "car". The model has no way to learn that two words share a relationship.
The Embedding Space
Word embeddings solve this fundamental limitation by mapping words to continuous coordinates instead of discrete bins. Instead of being isolated points in a sparse grid, words are placed in a dense, multidimensional space where similar concepts naturally cluster together. The geometry of the space physically mirrors the semantic relationships of the language itself. Words that appear in similar contexts in training data end up with remarkably similar coordinate values.
Semantic Math
Because meaning is now represented as raw coordinates, complex linguistic relationships become predictable, simple distances. The mathematical direction and distance that turns the concept of "Man" into "Woman" is the exact same vector that turns "King" into "Queen". You can subtract and add concepts directly, allowing a neural network to reason about analogies and semantic composition purely through vector arithmetic.
Where It Breaks
Traditional static embeddings assign exactly one permanent vector to each vocabulary word, completely regardless of the surrounding context. Polysemous words like "bank" are pulled in entirely opposite directions by their distinct meanings during training. The vector ultimately ends up stranded in a dead zone halfway between the "finance" cluster and the "nature" cluster, meaning the model effectively loses both definitions.
The Quick Version
- One-hot fails because it treats all words as equidistant, ignoring meaning.
- Continuous embeddings group semantically similar words into spatial clusters.
- Vector math allows complex linguistic relationships to be calculated predictably.
- Static vectors collapse polysemous words into an average, losing specific context.