Skip to content
AI360Xpert
Core ML

Visual explainer

Word2Vec

Embed words into continuous vector space to capture semantic relationships and enable arithmetic like king - man + woman = queen.

From one-hot encoding to dense embeddings.
From one-hot encoding to dense embeddings.

Before Word2Vec, natural language processing primarily treated words as unique, isolated IDs (one-hot encoding). Because every word is a separate orthogonal axis, "apple" is mathematically no closer to "orange" than it is to "car". Word2Vec introduced a fundamental shift: mapping discrete words into a dense, continuous vector space where physical proximity represents semantic similarity.

Learning from Context

Training via sliding context window.
Training via sliding context window.

Word2Vec learns these relationships without manual labeling. It slides a fixed-size window over massive text corpora. By continuously tweaking its vectors to predict a target word from its surrounding context (Continuous Bag-of-Words) or predicting the context from a target word (Skip-gram), the model forces words that appear in similar environments to occupy the same spatial neighborhood.

Semantic Geometry

Directions encode semantic relationships.
Directions encode semantic relationships.

Because the space is entirely structured by context, the directions and distances between words acquire distinct semantic meaning. Traversing the space algebraically allows for vector arithmetic. Starting at the vector for "king", subtracting the concept of "man", and adding "woman" lands near the vector for "queen".

While extremely powerful, standard Word2Vec assigns exactly one fixed vector per word. This forces polysemous words (like "bank" as a financial institution vs. a river edge) into compromised middle grounds, a limitation solved by later contextual models like BERT.