Skip to content
AI360Xpert
Beta

Seq2Seq with Attention Mechanism

Dynamic alignment that frees encoder-decoder models from fixed-size context vectors. At each decoding step, the model computes a weighted average over all source representations.

Seq2Seq attention enables the decoder to query all encoder states dynamically, constructing a customized context vector for each generated output token.
Seq2Seq attention enables the decoder to query all encoder states dynamically, constructing a customized context vector for each generated output token.

Why Does This Exist?

The original sequence-to-sequence (Seq2Seq) architecture introduced by Sutskever et al. (2014) and Cho et al. (2014) revolutionized neural machine translation and speech recognition. A recurrent encoder processed an input sequence of variable length TxT_x, updating its hidden state iteratively until the final token was reached. The final hidden state hTxh_{T_x} served as the sole context vector cc, handed over to initialize the decoder.

This design created a crippling mathematical flaw: the fixed-length vector bottleneck. Compressing a complex 40-word legal document or a multi-clause compound sentence into a single 512-dimensional vector forced the network to overwrite early details with later tokens. In practice, translation quality measured by BLEU score collapsed precipitously as sentence lengths exceeded 20 words.

Attention mechanisms, pioneered by Bahdanau et al. in 2014 and refined by Luong et al. in 2015, dissolved this bottleneck entirely. Instead of discarding intermediate encoder representations, the decoder retains direct random access to all hidden states h1,…,hTxh_1, \dots, h_{T_x}. At every step tt of target generation, the decoder constructs a bespoke, dynamic context vector ctc_t by computing a softmax alignment distribution across the entire source sequence. This breakthrough served as the direct intellectual ancestor of modern self-attention in transformers.

Think of It Like This

A simultaneous interpreter consulting a full written transcript

Imagine a conference interpreter translating a technical keynote speech from German into English.

In the classic fixed-vector Seq2Seq setup, the interpreter is forced to listen to a full three-minute paragraph in complete silence, close their eyes, memorize the entire passage as a single mental impression, and then begin speaking English from raw memory. Early German verbs and noun clauses get garbled or forgotten, because human working memory cannot store forty words of nuanced syntax in a single mental snapshot.

With an attention mechanism, the interpreter works with a live, real-time written stenograph transcript of the German speech resting on their desk. As they speak each English sentence, their eyes actively glance back to the exact German phrase currently being rendered: when translating the German compound adjective, their gaze zooms in on words four and five; when translating the sentence-final separable German verb prefix, their gaze darts to the very end of the paragraph. They generate an output word while dynamically referencing the specific source words that justify it.

How It Actually Works

Additive Versus Multiplicative Alignment Formulations

An attention-equipped sequence-to-sequence model consists of an encoder producing a sequence of hidden state vectors h=(h1,h2,…,hTx)h = (h_1, h_2, \dots, h_{T_x}), and a decoder generating target tokens y1,y2,…,yTyy_1, y_2, \dots, y_{T_y} using decoder hidden states s1,s2,…,sTys_1, s_2, \dots, s_{T_y}.

At each target decoding step tt, the model computes a dynamic context vector ctc_t as the weighted sum of all encoder states:

ct=∑i=1Txαt,ihic_t = \sum_{i=1}^{T_x} \alpha_{t, i} h_i

where the scalar attention weights αt,i\alpha_{t, i} form a probability distribution derived from a softmax over scalar alignment scores et,ie_{t, i}:

αt,i=exp⁡(et,i)∑k=1Txexp⁡(et,k)\alpha_{t, i} = \frac{\exp(e_{t, i})}{\sum_{k=1}^{T_x} \exp(e_{t, k})}

The alignment score et,ie_{t, i} measures how well source position ii matches the decoding query state. Two major formulations govern this scoring:

1. Bahdanau Attention (Additive Attention, 2014)

Bahdanau et al. formulate alignment using a single-layer feed-forward network with a non-linear tanh⁡\tanh activation. Scoring evaluates the previous decoder state st−1s_{t-1} against encoder state hih_i:

et,i=va⊤tanh⁡(Wast−1+Uahi)e_{t, i} = v_a^\top \tanh(W_a s_{t-1} + U_a h_i)

where Wa∈Rda×ddecW_a \in \mathbb{R}^{d_a \times d_{\text{dec}}}, Ua∈Rda×dencU_a \in \mathbb{R}^{d_a \times d_{\text{enc}}}, and va∈Rdav_a \in \mathbb{R}^{d_a} are trainable projection matrices. The context vector ctc_t is computed before the current decoder hidden state sts_t is updated:

st=f(st−1,yt−1,ct)s_t = f(s_{t-1}, y_{t-1}, c_t)

2. Luong Attention (Multiplicative Attention, 2015)

Luong et al. simplify the alignment calculation by utilizing the current decoder hidden state sts_t, and propose three distinct scoring functions:

  • Dot: et,i=st⊤hie_{t, i} = s_t^\top h_i (requires ddec=dencd_{\text{dec}} = d_{\text{enc}})
  • General (Multiplicative): et,i=st⊤Wahie_{t, i} = s_t^\top W_a h_i
  • Concat: et,i=va⊤tanh⁡(Wa[st ; hi])e_{t, i} = v_a^\top \tanh(W_a [s_t \,;\, h_i])

Once ctc_t is assembled, Luong combines the context vector and current decoder state through an attentional hidden state layer s~t\tilde{s}_t:

s~t=tanh⁡(Wc[ct ; st])\tilde{s}_t = \tanh(W_c [c_t \,;\, s_t])

The predictive probability for the next token is then projected directly: p(yt∣y<t,x)=softmax(Wss~t)p(y_t \mid y_{<t}, x) = \text{softmax}(W_s \tilde{s}_t).

Worked Example

Suppose an encoder processes a 3-word source sentence (Tx=3T_x = 3) with 2-dimensional hidden states h1,h2,h3∈R2h_1, h_2, h_3 \in \mathbb{R}^2:

h1=[1.000.00]("The"),h2=[0.002.00]("black"),h3=[1.001.00]("cat")h_1 = \begin{bmatrix} 1.00 \\ 0.00 \end{bmatrix} \quad (\text{"The"}), \quad h_2 = \begin{bmatrix} 0.00 \\ 2.00 \end{bmatrix} \quad (\text{"black"}), \quad h_3 = \begin{bmatrix} 1.00 \\ 1.00 \end{bmatrix} \quad (\text{"cat"})

The decoder is currently generating target token yt=y_t = "noire" (French feminine adjective for "black"). Its query state is:

st=[0.201.50]s_t = \begin{bmatrix} 0.20 \\ 1.50 \end{bmatrix}

Using Luong Dot scoring (et,i=st⊤hie_{t, i} = s_t^\top h_i):

Step 1: Compute Raw Alignment Scores

et,1=st⊤h1=(0.20×1.00)+(1.50×0.00)=0.20e_{t, 1} = s_t^\top h_1 = (0.20 \times 1.00) + (1.50 \times 0.00) = 0.20 et,2=st⊤h2=(0.20×0.00)+(1.50×2.00)=3.00e_{t, 2} = s_t^\top h_2 = (0.20 \times 0.00) + (1.50 \times 2.00) = 3.00 et,3=st⊤h3=(0.20×1.00)+(1.50×1.00)=0.20+1.50=1.70e_{t, 3} = s_t^\top h_3 = (0.20 \times 1.00) + (1.50 \times 1.00) = 0.20 + 1.50 = 1.70

Step 2: Evaluate Softmax Probabilities αt,i\alpha_{t, i}

Compute exponentials:

e0.20≈1.2214,e3.00≈20.0855,e1.70≈5.4739e^{0.20} \approx 1.2214, \quad e^{3.00} \approx 20.0855, \quad e^{1.70} \approx 5.4739

Sum of exponentials:

∑k=13eet,k=1.2214+20.0855+5.4739=26.7808\sum_{k=1}^3 e^{e_{t, k}} = 1.2214 + 20.0855 + 5.4739 = 26.7808

Calculate normalized weights:

αt,1=1.221426.7808≈0.0456(4.56%)\alpha_{t, 1} = \frac{1.2214}{26.7808} \approx 0.0456 \quad (4.56\%) αt,2=20.085526.7808≈0.7500(75.00%)\alpha_{t, 2} = \frac{20.0855}{26.7808} \approx 0.7500 \quad (75.00\%) αt,3=5.473926.7808≈0.2044(20.44%)\alpha_{t, 3} = \frac{5.4739}{26.7808} \approx 0.2044 \quad (20.44\%)

Step 3: Compute the Dynamic Context Vector ctc_t

ct=0.0456[1.000.00]+0.7500[0.002.00]+0.2044[1.001.00]c_t = 0.0456 \begin{bmatrix} 1.00 \\ 0.00 \end{bmatrix} + 0.7500 \begin{bmatrix} 0.00 \\ 2.00 \end{bmatrix} + 0.2044 \begin{bmatrix} 1.00 \\ 1.00 \end{bmatrix} ct,1=(0.0456×1.00)+(0.7500×0.00)+(0.2044×1.00)=0.0456+0.2044=0.2500c_{t, 1} = (0.0456 \times 1.00) + (0.7500 \times 0.00) + (0.2044 \times 1.00) = 0.0456 + 0.2044 = 0.2500 ct,2=(0.0456×0.00)+(0.7500×2.00)+(0.2044×1.00)=1.5000+0.2044=1.7044c_{t, 2} = (0.0456 \times 0.00) + (0.7500 \times 2.00) + (0.2044 \times 1.00) = 1.5000 + 0.2044 = 1.7044 ct=[0.25001.7044]c_t = \begin{bmatrix} 0.2500 \\ 1.7044 \end{bmatrix}

Because 75%75\% of the attention mass centers on h2h_2, ctc_t carries the semantic information of "black", providing the exact features required to output the feminine adjective form "noire".

Code

import numpy as np
def compute_seq2seq_attention(    encoder_states: np.ndarray,    decoder_query: np.ndarray,    method: str = "dot",    W_general: np.ndarray | None = None) -> tuple[np.ndarray, np.ndarray]:    """Compute attention alignment distribution and dynamic context vector."""    # encoder_states: (T_x, d), decoder_query: (d,)    if method == "dot":        scores = np.dot(encoder_states, decoder_query)    elif method == "general":        if W_general is None:            raise ValueError("W_general matrix required for general attention")        # s_t^T * W * h_i        transformed = np.dot(encoder_states, W_general.T)        scores = np.dot(transformed, decoder_query)    else:        raise ValueError(f"Unsupported scoring method: {method}")
    # Numerically stable softmax    shift_scores = scores - np.max(scores)    exp_scores = np.exp(shift_scores)    alpha = exp_scores / np.sum(exp_scores)
    # Dynamic context vector: c_t = sum_i alpha_i * h_i    context_vector = np.sum(alpha[:, np.newaxis] * encoder_states, axis=0)
    return np.round(context_vector, 4), np.round(alpha, 4)
# 3 encoder hidden states (dimension d = 2)h_encoder = np.array([    [1.00, 0.00],  # "The"    [0.00, 2.00],  # "black"    [1.00, 1.00],  # "cat"], dtype=np.float64)
# Current decoder state querying for "noire"s_decoder = np.array([0.20, 1.50], dtype=np.float64)
c_t, alpha_weights = compute_seq2seq_attention(h_encoder, s_decoder, method="dot")
print(f"Alignment Weights (alpha): {alpha_weights.tolist()}")# -> Alignment Weights (alpha): [0.0456, 0.75, 0.2044]
print(f"Context Vector (c_t): {c_t.tolist()}")# -> Context Vector (c_t): [0.25, 1.7044]

Watch Out For

Computing quadratic cross-attention matrices without key-query dimension scaling causes gradient saturation

In dot-product attention scoring et,i=st⊤hie_{t, i} = s_t^\top h_i, when hidden state dimension dd grows large (e.g., d≥512d \ge 512), the variance of the dot product scales linearly with dd: Var⁡(st⊤hi)=d\operatorname{Var}(s_t^\top h_i) = d.

This causes raw alignment scores to explode in magnitude, pushing the subsequent softmax function into saturated regions where local gradients approach zero (αi(1−αi)≈0\alpha_i(1 - \alpha_i) \approx 0). In deep architectures, this prevents backpropagation from updating alignment projections. Vaswani et al. solved this by scaling scores by 1dk\frac{1}{\sqrt{d_k}}, a critical practice whenever deploying dot-product attention in high dimensions.

The Quick Version

  • Attention replaces the fixed single-vector encoder bottleneck with a dynamically calculated, step-specific context vector ctc_t.
  • Alignment scores evaluate compatibility between decoder query states and all source encoder hidden states.
  • Bahdanau computes additive attention via non-linear projection, while Luong utilizes fast multiplicative dot products.