Embeddings and Word Representations
Dense numerical vectors that map discrete vocabulary tokens into a continuous geometric space. Words with similar contextual usage cluster together and preserve linear relational analogies.
Why Does This Exist?
In classical natural language processing, words were treated as atomic, indivisible symbols. Under a one-hot encoding across a vocabulary of size , every distinct word is represented as a 50,000-dimensional sparse unit vector containing a single 1 and 49,999 zeros.
This discrete representation suffers from two catastrophic mathematical limitations. First, memory scale: maintaining and computing transformations across high-dimensional sparse vectors creates massive matrix multiplications. Second, metric blindness: because every one-hot vector is perpendicular to every other vector, the inner product for all , and the Euclidean distance . The vector for "doctor" is just as distant from "physician" as it is from "sandpaper". A statistical model trained on sentences about "physicians" cannot generalize to "doctors" without explicit co-occurrence data.
Dense word embeddings solve this by mapping discrete tokens into a continuous low-dimensional vector space (, typically 100 to 768). In this continuous space, geometric proximity mirrors semantic relatedness, and linear spatial translations correspond to semantic transformations, providing foundational inputs for multi-layer perceptrons and transformers.
Think of It Like This
A spatial town map versus an unorganized filing cabinet
Consider an alphabetical filing cabinet containing 50,000 unindexed folders. Folder #12,410 is labeled "Doctor", folder #34,812 is labeled "Physician", and folder #41,209 is labeled "Sandpaper". If an assistant pulls folder #12,410, nothing about its physical shelf position indicates that folder #34,812 contains closely related medical protocols. Every folder sits in its own isolated drawer slot, equidistant from every other file.
Now replace the cabinet with a three-dimensional map of a sprawling city. Rather than assigning arbitrary folder numbers, every profession is positioned according to geographic coordinates: latitude represents technical specialization, longitude represents social interaction, and altitude represents institutional authority.
"Doctor" and "Physician" receive almost identical coordinates on the medical district hill. "Nurse" sits directly adjacent. If you start at "Nurse", walk the exact vector difference that separates a "Paralegal" from an "Attorney" (the legal authority offset), you arrive directly at "Doctor". Geometric distance encodes functional meaning, and spatial directions encode relational roles.
How It Actually Works
The Distributional Hypothesis and Vector Geometry
Dense word representations ground themselves in the Distributional Hypothesis articulated by linguists Zellig Harris (1954) and J.R. Firth (1957): "You shall know a word by the company it keeps." Words occurring in similar linguistic contexts share similar semantic meanings.
In a continuous vector space, each vocabulary token is associated with a dense embedding vector . Semantic similarity is quantified using the cosine similarity metric:
Unlike one-hot encodings, continuous embeddings exhibit linear compositional structure. Because directional vectors capture linguistic relations (such as gender, verb tense, or capital cities), semantic analogies can be computed using simple vector addition and subtraction:
Word2Vec Optimization Objectives
Modern word embeddings are typically learned through predictive self-supervised objectives, popularized by Word2Vec:
-
Continuous Bag-of-Words (CBOW): Predicts a center target word given its surrounding context window:
-
Skip-Gram with Negative Sampling (SGNS): Predicts surrounding context words given the center word . Because computing the full softmax denominator over vocabulary is computationally prohibitive, SGNS formulates a binary logistic regression objective separating true context pairs from noise words :
Global matrix factorization approaches such as GloVe achieve similar representations by directly fitting the log-ratios of global word co-occurrence probabilities.
Worked Example
Consider a simplified 3-dimensional embedding space () where the dimensions roughly represent [royalty, masculinity, animacy].
Let the learned vectors for four words be:
Let an unrelated non-living token be:
Step 1: Compute the Vector Analogy
Calculate the predicted target vector for the analogy "king is to man as ? is to woman":
Evaluating dimension by dimension:
Thus, .
Step 2: Calculate Cosine Similarity to Candidate Words
Compute the dot product between and :
Compute the norms:
Evaluate the cosine similarity:
The target vector matches with directional alignment. In contrast, evaluating cosine similarity against yields a low alignment of , confirming that the arithmetic preserves relational structure.
Code
import numpy as np
def cosine_similarity(u: np.ndarray, v: np.ndarray) -> float: """Compute cosine similarity between two dense vectors.""" dot_product = float(np.dot(u, v)) norm_u = float(np.linalg.norm(u)) norm_v = float(np.linalg.norm(v)) if norm_u == 0.0 or norm_v == 0.0: return 0.0 return round(dot_product / (norm_u * norm_v), 4)
def solve_analogy( a: np.ndarray, b: np.ndarray, c: np.ndarray, candidates: dict[str, np.ndarray]) -> tuple[str, float]: """Solve analogy: a is to b as ? is to c -> target = a - b + c.""" target = a - b + c best_word = "" highest_sim = -1.0 for word, vec in candidates.items(): sim = cosine_similarity(target, vec) if sim > highest_sim: highest_sim = sim best_word = word return best_word, highest_sim
# Semantic 3D space: [royalty, masculinity, animacy]embeddings = { "king": np.array([0.90, 0.80, 0.95]), "man": np.array([0.10, 0.85, 0.92]), "woman": np.array([0.12, -0.80, 0.94]), "queen": np.array([0.92, -0.75, 0.96]), "apple": np.array([0.05, 0.02, 0.10]),}
# Analogy query: king - man + woman ?predicted_word, score = solve_analogy( a=embeddings["king"], b=embeddings["man"], c=embeddings["woman"], candidates={"queen": embeddings["queen"], "apple": embeddings["apple"]})
print(f"Nearest word: {predicted_word}")# -> Nearest word: queen
print(f"Cosine score: {score}")# -> Cosine score: 0.9986Watch Out For
Static embeddings collapse polysemous words into an average vector centroid
Static word representations assign exactly one immutable vector to each string in the vocabulary.
For polysemous words—words that carry multiple distinct definitions such as "bank" (a financial depository versus the muddy bank of a river) or "apple" (the fruit versus the technology enterprise)—the training objective forces the embedding vector to sit at the mathematical weighted average of both usages. Consequently, the vector aligns poorly with either actual meaning in specific sentences. This fundamental limitation prompted the development of contextualized representation architectures such as ELMo and transformer-based self-attention.
The Quick Version
- Word embeddings compress high-dimensional sparse one-hot tokens into compact, dense continuous vectors where proximity represents semantic similarity.
- Grounded in the distributional hypothesis, models learn word vectors from contextual co-occurrence patterns using objectives like Skip-Gram negative sampling.
- Linear vector arithmetic enables geometric analogy solving ().