Seq2Seq with Attention Mechanism
Dynamic alignment that frees encoder-decoder models from fixed-size context vectors. At each decoding step, the model computes a weighted average over all source representations.
Why Does This Exist?
The original sequence-to-sequence (Seq2Seq) architecture introduced by Sutskever et al. (2014) and Cho et al. (2014) revolutionized neural machine translation and speech recognition. A recurrent encoder processed an input sequence of variable length , updating its hidden state iteratively until the final token was reached. The final hidden state served as the sole context vector , handed over to initialize the decoder.
This design created a crippling mathematical flaw: the fixed-length vector bottleneck. Compressing a complex 40-word legal document or a multi-clause compound sentence into a single 512-dimensional vector forced the network to overwrite early details with later tokens. In practice, translation quality measured by BLEU score collapsed precipitously as sentence lengths exceeded 20 words.
Attention mechanisms, pioneered by Bahdanau et al. in 2014 and refined by Luong et al. in 2015, dissolved this bottleneck entirely. Instead of discarding intermediate encoder representations, the decoder retains direct random access to all hidden states . At every step of target generation, the decoder constructs a bespoke, dynamic context vector by computing a softmax alignment distribution across the entire source sequence. This breakthrough served as the direct intellectual ancestor of modern self-attention in transformers.
Think of It Like This
A simultaneous interpreter consulting a full written transcript
Imagine a conference interpreter translating a technical keynote speech from German into English.
In the classic fixed-vector Seq2Seq setup, the interpreter is forced to listen to a full three-minute paragraph in complete silence, close their eyes, memorize the entire passage as a single mental impression, and then begin speaking English from raw memory. Early German verbs and noun clauses get garbled or forgotten, because human working memory cannot store forty words of nuanced syntax in a single mental snapshot.
With an attention mechanism, the interpreter works with a live, real-time written stenograph transcript of the German speech resting on their desk. As they speak each English sentence, their eyes actively glance back to the exact German phrase currently being rendered: when translating the German compound adjective, their gaze zooms in on words four and five; when translating the sentence-final separable German verb prefix, their gaze darts to the very end of the paragraph. They generate an output word while dynamically referencing the specific source words that justify it.
How It Actually Works
Additive Versus Multiplicative Alignment Formulations
An attention-equipped sequence-to-sequence model consists of an encoder producing a sequence of hidden state vectors , and a decoder generating target tokens using decoder hidden states .
At each target decoding step , the model computes a dynamic context vector as the weighted sum of all encoder states:
where the scalar attention weights form a probability distribution derived from a softmax over scalar alignment scores :
The alignment score measures how well source position matches the decoding query state. Two major formulations govern this scoring:
1. Bahdanau Attention (Additive Attention, 2014)
Bahdanau et al. formulate alignment using a single-layer feed-forward network with a non-linear activation. Scoring evaluates the previous decoder state against encoder state :
where , , and are trainable projection matrices. The context vector is computed before the current decoder hidden state is updated:
2. Luong Attention (Multiplicative Attention, 2015)
Luong et al. simplify the alignment calculation by utilizing the current decoder hidden state , and propose three distinct scoring functions:
- Dot: (requires )
- General (Multiplicative):
- Concat:
Once is assembled, Luong combines the context vector and current decoder state through an attentional hidden state layer :
The predictive probability for the next token is then projected directly: .
Worked Example
Suppose an encoder processes a 3-word source sentence () with 2-dimensional hidden states :
The decoder is currently generating target token "noire" (French feminine adjective for "black"). Its query state is:
Using Luong Dot scoring ():
Step 1: Compute Raw Alignment Scores
Step 2: Evaluate Softmax Probabilities
Compute exponentials:
Sum of exponentials:
Calculate normalized weights:
Step 3: Compute the Dynamic Context Vector
Because of the attention mass centers on , carries the semantic information of "black", providing the exact features required to output the feminine adjective form "noire".
Code
import numpy as np
def compute_seq2seq_attention( encoder_states: np.ndarray, decoder_query: np.ndarray, method: str = "dot", W_general: np.ndarray | None = None) -> tuple[np.ndarray, np.ndarray]: """Compute attention alignment distribution and dynamic context vector.""" # encoder_states: (T_x, d), decoder_query: (d,) if method == "dot": scores = np.dot(encoder_states, decoder_query) elif method == "general": if W_general is None: raise ValueError("W_general matrix required for general attention") # s_t^T * W * h_i transformed = np.dot(encoder_states, W_general.T) scores = np.dot(transformed, decoder_query) else: raise ValueError(f"Unsupported scoring method: {method}")
# Numerically stable softmax shift_scores = scores - np.max(scores) exp_scores = np.exp(shift_scores) alpha = exp_scores / np.sum(exp_scores)
# Dynamic context vector: c_t = sum_i alpha_i * h_i context_vector = np.sum(alpha[:, np.newaxis] * encoder_states, axis=0)
return np.round(context_vector, 4), np.round(alpha, 4)
# 3 encoder hidden states (dimension d = 2)h_encoder = np.array([ [1.00, 0.00], # "The" [0.00, 2.00], # "black" [1.00, 1.00], # "cat"], dtype=np.float64)
# Current decoder state querying for "noire"s_decoder = np.array([0.20, 1.50], dtype=np.float64)
c_t, alpha_weights = compute_seq2seq_attention(h_encoder, s_decoder, method="dot")
print(f"Alignment Weights (alpha): {alpha_weights.tolist()}")# -> Alignment Weights (alpha): [0.0456, 0.75, 0.2044]
print(f"Context Vector (c_t): {c_t.tolist()}")# -> Context Vector (c_t): [0.25, 1.7044]Watch Out For
Computing quadratic cross-attention matrices without key-query dimension scaling causes gradient saturation
In dot-product attention scoring , when hidden state dimension grows large (e.g., ), the variance of the dot product scales linearly with : .
This causes raw alignment scores to explode in magnitude, pushing the subsequent softmax function into saturated regions where local gradients approach zero (). In deep architectures, this prevents backpropagation from updating alignment projections. Vaswani et al. solved this by scaling scores by , a critical practice whenever deploying dot-product attention in high dimensions.
The Quick Version
- Attention replaces the fixed single-vector encoder bottleneck with a dynamically calculated, step-specific context vector .
- Alignment scores evaluate compatibility between decoder query states and all source encoder hidden states.
- Bahdanau computes additive attention via non-linear projection, while Luong utilizes fast multiplicative dot products.