Skip to content
AI360Xpert
Beta

ELMo

Contextualized word vectors computed from stacked bidirectional language models. Downstream tasks learn a weighted combination of character, syntactic, and semantic layers.

ELMo constructs contextual representations by passing tokens through stacked forward and backward LSTMs, then computing a learned task-specific weighted sum of all layer vectors.
ELMo constructs contextual representations by passing tokens through stacked forward and backward LSTMs, then computing a learned task-specific weighted sum of all layer vectors.

Why Does This Exist?

Prior to 2018, the standard foundation of natural language systems relied on static word embeddings such as Word2Vec and GloVe. These models assigned a single immutable vector to each string in the vocabulary.

This static assumption broke down on polysemy. In the sentence "The river bank overflowed after heavy rain" and the sentence "The investment bank approved the commercial acquisition", the word "bank" carries entirely different syntactic roles and semantic definitions. Static embeddings collapsed both senses into a single compromise centroid vector, degrading downstream performance on question answering, named entity recognition, and sentiment classification.

Introduced by Peters et al. in 2018, ELMo (Embeddings from Language Models) broke this bottleneck by computing deep contextualized representations. Instead of assigning a fixed vector lookup, ELMo processes the complete sentence through a multi-layer bidirectional Language Model (biLM). The resulting representation for each word is a function of the entire sentence context. Furthermore, ELMo exposed a key mechanistic insight: lower recurrent layers encode syntax and part-of-speech roles, while higher layers capture complex semantic senses.

Think of It Like This

A specialist contractor interviewed across hierarchical management tiers

Imagine an independent contractor entering a large corporation to solve an operational problem. Depending on who you ask about the contractor's role, you receive different descriptions:

At the front desk (Layer 0, Subword/Character CNN), security checks their physical badge and appearance—verifying spelling, morphological prefixes, and base credentials regardless of context.

On the engineering floor (Layer 1, Syntactic biLSTM), the immediate team lead describes their operational grammar: how their role interfaces with adjacent team members, whether they act as an operative or a reviewer, and their procedural placement in daily standups.

In the executive boardroom (Layer 2, Semantic biLSTM), the vice president describes their high-level business objective: what market problem they solve and how their presence changes quarterly strategy.

If the company assigns the contractor to write code (a syntactic downstream task like POS tagging), the team heavily relies on the engineering lead's evaluation (s1s_1 receives high weight). If they assign the contractor to renegotiate a corporate merger (a semantic task like sentiment or coreference), the executive's assessment dominates (s2s_2 receives high weight). ELMo lets each downstream task tune its own mixture of these management perspectives.

How It Actually Works

Bidirectional Language Models and Layer-Wise Pooling

ELMo generates representations by pre-training a multi-layer bidirectional Language Model (biLM) on large unlabeled text corpora, then freezing the model and extracting linear layer combinations for downstream architectures.

1. The Bidirectional Language Model (biLM)

Given a sequence of NN tokens (t1,t2,…,tN)(t_1, t_2, \dots, t_N), ELMo first converts each token into a context-independent vector xk∈Rdembedx_k \in \mathbb{R}^{d_{\text{embed}}} using a character-level Convolutional Neural Network (CNN). This character foundation prevents out-of-vocabulary failures by extracting robust subword morphological features.

Next, the token passes through LL stacked layers of bidirectional LSTMs:

  • Forward LSTM: Computes directional hidden states h⃗k,jLM\vec{h}_{k, j}^{\text{LM}} at step kk and layer j∈{1,…,L}j \in \{1, \dots, L\}, optimizing the probability of predicting token tkt_k given past context (t1,…,tk−1)(t_1, \dots, t_{k-1}):

    p(t1,t2,…,tN)=∏k=1Np(tk∣t1,…,tk−1)p(t_1, t_2, \dots, t_N) = \prod_{k=1}^N p(t_k \mid t_1, \dots, t_{k-1})
  • Backward LSTM: Computes directional hidden states h←k,jLM\overleftarrow{h}_{k, j}^{\text{LM}} at step kk and layer jj, optimizing the probability of predicting tkt_k given future context (tk+1,…,tN)(t_{k+1}, \dots, t_N):

    p(t1,t2,…,tN)=∏k=1Np(tk∣tk+1,…,tN)p(t_1, t_2, \dots, t_N) = \prod_{k=1}^N p(t_k \mid t_{k+1}, \dots, t_N)

At each step kk, layer jj outputs the concatenated representation:

hk,jLM=[h⃗k,jLM ; h←k,jLM]∈R2dlstmh_{k, j}^{\text{LM}} = \left[ \vec{h}_{k, j}^{\text{LM}} \,;\, \overleftarrow{h}_{k, j}^{\text{LM}} \right] \in \mathbb{R}^{2 d_{\text{lstm}}}

Defining hk,0LM=[xk ; xk]h_{k, 0}^{\text{LM}} = [x_k \,;\, x_k], the biLM produces a set of 2L+12L + 1 representations for each token:

Rk={hk,jLM∣j=0,…,L}R_k = \left\{ h_{k, j}^{\text{LM}} \mid j = 0, \dots, L \right\}

2. Downstream Task-Specific Aggregation

Rather than selecting only the topmost LSTM layer, ELMo allows downstream task models (such as an NER classifier or a relation extractor) to learn a task-specific linear combination across all internal layers:

ELMoktask=γtask∑j=0Lsjtaskhk,jLM\text{ELMo}_k^{\text{task}} = \gamma^{\text{task}} \sum_{j=0}^L s_j^{\text{task}} h_{k, j}^{\text{LM}}

where:

  • stask=[s0task,…,sLtask]⊤s^{\text{task}} = [s_0^{\text{task}}, \dots, s_L^{\text{task}}]^\top are softmax-normalized layer weights:

    sjtask=exp⁡(wjtask)∑i=0Lexp⁡(witask)s_j^{\text{task}} = \frac{\exp(w_j^{\text{task}})}{\sum_{i=0}^L \exp(w_i^{\text{task}})}
  • γtask\gamma^{\text{task}} is a learnable scalar scale parameter allowing the downstream model to scale ELMo's overall vector norm to match its internal representations.

Empirical probes demonstrate that s1s_1 receives highest weight for syntactic tasks like part-of-speech tagging, whereas s2s_2 dominates for semantic tasks like word sense disambiguation.

Worked Example

Consider an ELMo model with L=2L = 2 biLSTM layers (total L+1=3L+1 = 3 representations: layer 0, layer 1, and layer 2) with feature dimension d=4d = 4.

Let the extracted vectors for the polysemous word "bank" in the sentence "The river bank overflowed" be:

hk,0=[1.000.500.20−0.40](Character CNN)h_{k, 0} = \begin{bmatrix} 1.00 \\ 0.50 \\ 0.20 \\ -0.40 \end{bmatrix} \quad (\text{Character CNN}) hk,1=[0.801.20−0.600.50](Layer 1: Syntactic Noun)h_{k, 1} = \begin{bmatrix} 0.80 \\ 1.20 \\ -0.60 \\ 0.50 \end{bmatrix} \quad (\text{Layer 1: Syntactic Noun}) hk,2=[1.502.00−1.001.20](Layer 2: Semantic River/Geology)h_{k, 2} = \begin{bmatrix} 1.50 \\ 2.00 \\ -1.00 \\ 1.20 \end{bmatrix} \quad (\text{Layer 2: Semantic River/Geology})

Let a downstream word sense disambiguation classifier learn unnormalized softmax logits wtask=[0.20,0.80,1.80]⊤w^{\text{task}} = [0.20, 0.80, 1.80]^\top.

Step 1: Compute Softmax Layer Weights

Compute exponentials:

e0.20≈1.2214,e0.80≈2.2255,e1.80≈6.0496e^{0.20} \approx 1.2214, \quad e^{0.80} \approx 2.2255, \quad e^{1.80} \approx 6.0496

Sum of exponentials:

∑j=02ewj=1.2214+2.2255+6.0496=9.4965\sum_{j=0}^2 e^{w_j} = 1.2214 + 2.2255 + 6.0496 = 9.4965

Normalized weights:

s0=1.22149.4965≈0.1286,s1=2.22559.4965≈0.2343,s2=6.04969.4965≈0.6371s_0 = \frac{1.2214}{9.4965} \approx 0.1286, \quad s_1 = \frac{2.2255}{9.4965} \approx 0.2343, \quad s_2 = \frac{6.0496}{9.4965} \approx 0.6371

Notice that the semantic layer (s2s_2) receives 63.71%63.71\% of the total attention.

Step 2: Compute Weighted Representation Sum

Multiply each vector by its scalar weight:

s0hk,0=[0.1286, 0.0643, 0.0257, −0.0514]⊤s_0 h_{k, 0} = [0.1286, \, 0.0643, \, 0.0257, \, -0.0514]^\top s1hk,1=[0.1874, 0.2812, −0.1406, 0.1172]⊤s_1 h_{k, 1} = [0.1874, \, 0.2812, \, -0.1406, \, 0.1172]^\top s2hk,2=[0.9557, 1.2742, −0.6371, 0.7645]⊤s_2 h_{k, 2} = [0.9557, \, 1.2742, \, -0.6371, \, 0.7645]^\top

Sum across dimensions:

hcombo=[0.1286+0.1874+0.95570.0643+0.2812+1.27420.0257−0.1406−0.6371−0.0514+0.1172+0.7645]=[1.27171.6197−0.75200.8303]h_{\text{combo}} = \begin{bmatrix} 0.1286 + 0.1874 + 0.9557 \\ 0.0643 + 0.2812 + 1.2742 \\ 0.0257 - 0.1406 - 0.6371 \\ -0.0514 + 0.1172 + 0.7645 \end{bmatrix} = \begin{bmatrix} 1.2717 \\ 1.6197 \\ -0.7520 \\ 0.8303 \end{bmatrix}

Step 3: Apply Downstream Scale γ=1.50\gamma = 1.50

ELMoktask=1.50×[1.27171.6197−0.75200.8303]=[1.90762.4296−1.12801.2455]\text{ELMo}_k^{\text{task}} = 1.50 \times \begin{bmatrix} 1.2717 \\ 1.6197 \\ -0.7520 \\ 0.8303 \end{bmatrix} = \begin{bmatrix} 1.9076 \\ 2.4296 \\ -1.1280 \\ 1.2455 \end{bmatrix}

Code

import numpy as np
def compute_elmo_representation(    layer_vectors: list[np.ndarray],    task_logits: np.ndarray,    gamma: float = 1.0) -> tuple[np.ndarray, np.ndarray]:    """Compute task-specific ELMo contextualized representation."""    # Numerically stable softmax layer weights    shift = task_logits - np.max(task_logits)    exp_weights = np.exp(shift)    s_weights = exp_weights / np.sum(exp_weights)
    # Weighted linear combination: sum_j s_j * h_j    stacked = np.stack(layer_vectors, axis=0)  # Shape: (L+1, d)    weighted_sum = np.sum(s_weights[:, np.newaxis] * stacked, axis=0)
    # Apply task scalar scale parameter gamma    elmo_out = gamma * weighted_sum    return np.round(elmo_out, 4), np.round(s_weights, 4)
# Three layer vectors for token 'bank' (Char CNN, Layer 1 biLSTM, Layer 2 biLSTM)h_0 = np.array([1.00, 0.50, 0.20, -0.40], dtype=np.float64)h_1 = np.array([0.80, 1.20, -0.60, 0.50], dtype=np.float64)h_2 = np.array([1.50, 2.00, -1.00, 1.20], dtype=np.float64)
# Learned downstream logits favoring semantic layerlogits = np.array([0.20, 0.80, 1.80], dtype=np.float64)scale_gamma = 1.50
elmo_vector, weights = compute_elmo_representation(    [h_0, h_1, h_2],    logits,    gamma=scale_gamma)
print(f"Layer Softmax Weights: {weights.tolist()}")# -> Layer Softmax Weights: [0.1286, 0.2343, 0.6371]
print(f"Contextualized Vector: {elmo_vector.tolist()}")# -> Contextualized Vector: [1.9076, 2.4296, -1.128, 1.2455]

Watch Out For

ELMo concatenates independent forward and backward passes rather than true joint cross-conditioning

A frequent misconception is assuming ELMo's bidirectional LSTMs attend jointly to left and right context simultaneously.

In reality, ELMo trains two strictly independent unidirectional models: the forward LSTM sees only preceding tokens, while the backward LSTM sees only succeeding tokens. Their hidden vectors are concatenated post-hoc: [h⃗ ; h←][\vec{h} \,;\, \overleftarrow{h}]. True bidirectional interaction—where each token's representation attends to all other tokens jointly across intermediate layers—was only realized later with masked language modeling in BERT and self-attention transformers.

The Quick Version

  • ELMo replaces fixed word embeddings with dynamic, sentence-dependent vectors extracted from a deep bidirectional Language Model.
  • A character CNN provides out-of-vocabulary resilience, while stacked biLSTMs capture directional linguistic context.
  • Downstream tasks learn a softmax-weighted sum across all model layers, prioritizing syntax or semantics depending on the application.