ELMo
Contextualized word vectors computed from stacked bidirectional language models. Downstream tasks learn a weighted combination of character, syntactic, and semantic layers.
Why Does This Exist?
Prior to 2018, the standard foundation of natural language systems relied on static word embeddings such as Word2Vec and GloVe. These models assigned a single immutable vector to each string in the vocabulary.
This static assumption broke down on polysemy. In the sentence "The river bank overflowed after heavy rain" and the sentence "The investment bank approved the commercial acquisition", the word "bank" carries entirely different syntactic roles and semantic definitions. Static embeddings collapsed both senses into a single compromise centroid vector, degrading downstream performance on question answering, named entity recognition, and sentiment classification.
Introduced by Peters et al. in 2018, ELMo (Embeddings from Language Models) broke this bottleneck by computing deep contextualized representations. Instead of assigning a fixed vector lookup, ELMo processes the complete sentence through a multi-layer bidirectional Language Model (biLM). The resulting representation for each word is a function of the entire sentence context. Furthermore, ELMo exposed a key mechanistic insight: lower recurrent layers encode syntax and part-of-speech roles, while higher layers capture complex semantic senses.
Think of It Like This
A specialist contractor interviewed across hierarchical management tiers
Imagine an independent contractor entering a large corporation to solve an operational problem. Depending on who you ask about the contractor's role, you receive different descriptions:
At the front desk (Layer 0, Subword/Character CNN), security checks their physical badge and appearance—verifying spelling, morphological prefixes, and base credentials regardless of context.
On the engineering floor (Layer 1, Syntactic biLSTM), the immediate team lead describes their operational grammar: how their role interfaces with adjacent team members, whether they act as an operative or a reviewer, and their procedural placement in daily standups.
In the executive boardroom (Layer 2, Semantic biLSTM), the vice president describes their high-level business objective: what market problem they solve and how their presence changes quarterly strategy.
If the company assigns the contractor to write code (a syntactic downstream task like POS tagging), the team heavily relies on the engineering lead's evaluation ( receives high weight). If they assign the contractor to renegotiate a corporate merger (a semantic task like sentiment or coreference), the executive's assessment dominates ( receives high weight). ELMo lets each downstream task tune its own mixture of these management perspectives.
How It Actually Works
Bidirectional Language Models and Layer-Wise Pooling
ELMo generates representations by pre-training a multi-layer bidirectional Language Model (biLM) on large unlabeled text corpora, then freezing the model and extracting linear layer combinations for downstream architectures.
1. The Bidirectional Language Model (biLM)
Given a sequence of tokens , ELMo first converts each token into a context-independent vector using a character-level Convolutional Neural Network (CNN). This character foundation prevents out-of-vocabulary failures by extracting robust subword morphological features.
Next, the token passes through stacked layers of bidirectional LSTMs:
-
Forward LSTM: Computes directional hidden states at step and layer , optimizing the probability of predicting token given past context :
-
Backward LSTM: Computes directional hidden states at step and layer , optimizing the probability of predicting given future context :
At each step , layer outputs the concatenated representation:
Defining , the biLM produces a set of representations for each token:
2. Downstream Task-Specific Aggregation
Rather than selecting only the topmost LSTM layer, ELMo allows downstream task models (such as an NER classifier or a relation extractor) to learn a task-specific linear combination across all internal layers:
where:
-
are softmax-normalized layer weights:
-
is a learnable scalar scale parameter allowing the downstream model to scale ELMo's overall vector norm to match its internal representations.
Empirical probes demonstrate that receives highest weight for syntactic tasks like part-of-speech tagging, whereas dominates for semantic tasks like word sense disambiguation.
Worked Example
Consider an ELMo model with biLSTM layers (total representations: layer 0, layer 1, and layer 2) with feature dimension .
Let the extracted vectors for the polysemous word "bank" in the sentence "The river bank overflowed" be:
Let a downstream word sense disambiguation classifier learn unnormalized softmax logits .
Step 1: Compute Softmax Layer Weights
Compute exponentials:
Sum of exponentials:
Normalized weights:
Notice that the semantic layer () receives of the total attention.
Step 2: Compute Weighted Representation Sum
Multiply each vector by its scalar weight:
Sum across dimensions:
Step 3: Apply Downstream Scale
Code
import numpy as np
def compute_elmo_representation( layer_vectors: list[np.ndarray], task_logits: np.ndarray, gamma: float = 1.0) -> tuple[np.ndarray, np.ndarray]: """Compute task-specific ELMo contextualized representation.""" # Numerically stable softmax layer weights shift = task_logits - np.max(task_logits) exp_weights = np.exp(shift) s_weights = exp_weights / np.sum(exp_weights)
# Weighted linear combination: sum_j s_j * h_j stacked = np.stack(layer_vectors, axis=0) # Shape: (L+1, d) weighted_sum = np.sum(s_weights[:, np.newaxis] * stacked, axis=0)
# Apply task scalar scale parameter gamma elmo_out = gamma * weighted_sum return np.round(elmo_out, 4), np.round(s_weights, 4)
# Three layer vectors for token 'bank' (Char CNN, Layer 1 biLSTM, Layer 2 biLSTM)h_0 = np.array([1.00, 0.50, 0.20, -0.40], dtype=np.float64)h_1 = np.array([0.80, 1.20, -0.60, 0.50], dtype=np.float64)h_2 = np.array([1.50, 2.00, -1.00, 1.20], dtype=np.float64)
# Learned downstream logits favoring semantic layerlogits = np.array([0.20, 0.80, 1.80], dtype=np.float64)scale_gamma = 1.50
elmo_vector, weights = compute_elmo_representation( [h_0, h_1, h_2], logits, gamma=scale_gamma)
print(f"Layer Softmax Weights: {weights.tolist()}")# -> Layer Softmax Weights: [0.1286, 0.2343, 0.6371]
print(f"Contextualized Vector: {elmo_vector.tolist()}")# -> Contextualized Vector: [1.9076, 2.4296, -1.128, 1.2455]Watch Out For
ELMo concatenates independent forward and backward passes rather than true joint cross-conditioning
A frequent misconception is assuming ELMo's bidirectional LSTMs attend jointly to left and right context simultaneously.
In reality, ELMo trains two strictly independent unidirectional models: the forward LSTM sees only preceding tokens, while the backward LSTM sees only succeeding tokens. Their hidden vectors are concatenated post-hoc: . True bidirectional interaction—where each token's representation attends to all other tokens jointly across intermediate layers—was only realized later with masked language modeling in BERT and self-attention transformers.
The Quick Version
- ELMo replaces fixed word embeddings with dynamic, sentence-dependent vectors extracted from a deep bidirectional Language Model.
- A character CNN provides out-of-vocabulary resilience, while stacked biLSTMs capture directional linguistic context.
- Downstream tasks learn a softmax-weighted sum across all model layers, prioritizing syntax or semantics depending on the application.