Visual explainer
Mixture Of Experts
How large language models scale capacity without increasing compute by routing tokens to specialised sub-networks.
Traditional dense models are monolithic—every single parameter is activated for every token. As models scale to hundreds of billions of parameters to increase capacity, the compute required becomes a severe bottleneck.
A Mixture of Experts (MoE) architecture solves this by replacing dense layers with a collection of smaller, independent neural networks called experts. Instead of activating everything, a learned router directs each token to only the most relevant sub-networks, creating a sparse model that scales capacity massively while keeping compute per token relatively low.
The Routing Mechanism
The router evaluates an incoming token and assigns a probability (or gate value) to each expert. Instead of passing the token through all available experts, it selects only the Top-K (often just the top 1 or 2).
The token is processed exclusively by these chosen sub-networks, and their outputs are multiplied by their respective gate probabilities before being summed. The unselected experts remain completely idle for that token, saving massive amounts of compute.
Load Balancing and Expert Collapse
MoE models are highly susceptible to a failure mode called expert collapse. If the router discovers early on that one expert is slightly better at a task, it may start sending the vast majority of tokens to that single expert.
This creates a vicious cycle: the active expert receives all the gradients and gets even better, while the other experts are starved of updates and become useless. To prevent this, models apply an auxiliary balancing loss during training, which forces the router to distribute tokens evenly across all experts, ensuring the entire model's capacity is utilized effectively.