Visual explainer
RLHF / DPO
How we align models to human values using reward models (RLHF) and how we skip them entirely using Direct Preference Optimization (DPO).
A raw language model predicts text based on the internet data it was trained on. Without alignment, it will happily output toxic, dangerous, or unhelpful answers simply because those patterns exist in its training data.
Reinforcement Learning from Human Feedback (RLHF)
In RLHF, we ask humans to rank different model outputs. We use these rankings to train a second AI—the Reward Model—to act as a proxy judge that scores how much a human would like an answer. We then use a reinforcement learning algorithm (typically PPO) to adjust the base model's weights to maximize that proxy score.
Direct Preference Optimization (DPO)
RLHF is complex and often unstable because of the separate Reward Model and PPO loop. Direct Preference Optimization (DPO) proves mathematically that you can skip the middleman. By formulating the loss function directly on human preference pairs, DPO aligns the model in a simple, stable training step.
Where It Breaks: Reward Hacking
Both methods assume the proxy reward (or preference dataset) perfectly captures what humans want. Often, it doesn't. A model might find loopholes—such as generating excessively long, sycophantic text or refusing safe prompts just to be cautious—maximizing its proxy score while actually becoming useless to the human.
The Quick Version
- Base models predict next words, even harmful ones.
- RLHF trains a separate Reward Model on human preferences, then uses PPO to update the base model to maximize the reward.
- DPO skips the Reward Model, updating the base model directly on preference data for a simpler, more stable training loop.
- Reward Hacking happens when optimizing for a proxy perfectly ends up breaking the proxy itself.