Paper breakdown
Neural Machine Translation by Jointly Learning to Align and Translate
Introduces the attention mechanism to sequence-to-sequence models, allowing networks to dynamically align inputs and outputs, removing the single vector bottleneck.
Paper: Neural Machine Translation by Jointly Learning to Align and Translate
Authors: Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio · 2014
Read the paperThe Problem
Before this paper, the state-of-the-art for neural machine translation was the standard encoder-decoder architecture (Seq2Seq). The encoder would process a source sentence of any length and compress all its semantic information into a single, fixed-length context vector. The decoder would then generate the translation from this single vector.
This created a severe information bottleneck. A single vector cannot adequately capture the full meaning and nuance of a long, complex sentence. As a result, the performance of these models dropped precipitously as the length of the input sentence increased. They simply couldn't remember everything they needed to translate.
The Idea
Instead of forcing the entire input sentence into one vector, what if the model could keep all the intermediate representations of the input words and "look back" at them dynamically during translation?
The authors proposed a novel architecture where the decoder learns to "attend" to different parts of the source sentence at each step of generating an output word. By jointly learning to align (finding which source words are relevant right now) and translate, the network naturally solves the bottleneck of the fixed-length vector.
How It Works
The architecture introduces a dynamic mechanism between a bidirectional RNN encoder and an RNN decoder:
The Encoder
Instead of just taking the final hidden state, the model collects a sequence of hidden states, or annotations, one for each word in the source sentence. Because it uses a bidirectional RNN, each annotation contains information about the preceding and following words, centering around a specific word.
The Alignment Model
At each decoding step, a small feedforward neural network scores how well the current state of the decoder aligns with each of the encoder's annotations. These scores are passed through a softmax function to create a weight distribution across the source words.
The Context Vector
The model calculates a weighted sum of the encoder annotations using these alignment weights. This creates a bespoke context vector for just that specific step of the translation.
The Decoder
The decoder uses this dynamic context vector, along with its previous hidden state and the previously generated word, to predict the next word. It repeats this process, shifting its "attention" across the source sentence as it translates.
Why It Mattered
This paper fundamentally altered the trajectory of deep learning by introducing the attention mechanism. It completely eliminated the rapid degradation in performance for long sentences, allowing neural machine translation models to match and soon surpass traditional statistical methods.
More broadly, it introduced the concept of soft alignment—letting a neural network learn what parts of an input are important on the fly. This made the models inherently more interpretable, as one could visualize the alignment weights to see exactly which source words the model focused on when generating each target word.
What Came After
Attention quickly became the standard for nearly all sequential tasks in deep learning, moving far beyond just translation. The success of this mechanism directly paved the way for models that relied entirely on attention without recurrence, culminating in the seminal 2017 paper Attention Is All You Need, which introduced the Transformer architecture that powers modern Large Language Models today.