Visual explainer
Word2Vec
Embed words into continuous vector space to capture semantic relationships and enable arithmetic like king - man + woman = queen.
Before Word2Vec, natural language processing primarily treated words as unique, isolated IDs (one-hot encoding). Because every word is a separate orthogonal axis, "apple" is mathematically no closer to "orange" than it is to "car". Word2Vec introduced a fundamental shift: mapping discrete words into a dense, continuous vector space where physical proximity represents semantic similarity.
Learning from Context
Word2Vec learns these relationships without manual labeling. It slides a fixed-size window over massive text corpora. By continuously tweaking its vectors to predict a target word from its surrounding context (Continuous Bag-of-Words) or predicting the context from a target word (Skip-gram), the model forces words that appear in similar environments to occupy the same spatial neighborhood.
Semantic Geometry
Because the space is entirely structured by context, the directions and distances between words acquire distinct semantic meaning. Traversing the space algebraically allows for vector arithmetic. Starting at the vector for "king", subtracting the concept of "man", and adding "woman" lands near the vector for "queen".
While extremely powerful, standard Word2Vec assigns exactly one fixed vector per word. This forces polysemous words (like "bank" as a financial institution vs. a river edge) into compromised middle grounds, a limitation solved by later contextual models like BERT.
What to Read Next
- Word EmbeddingsHow computers represent words as dense vectors, allowing semantic meaning to be treated as math.
- Embedding SpaceHow models convert fuzzy semantic meaning into precise mathematical distances.
- Camera Projection and Pinhole ModelHow a 3D world maps to a 2D sensor using the pinhole model and similar triangles.