Visual explainer
LoRA (Low-Rank Adaptation)
How LoRA trains massive models by freezing the original weights and updating only a tiny, low-rank delta matrix instead.
To teach a large language model a new skill, the traditional approach is full fine-tuning. This requires calculating and applying updates to every single weight in the network—a process so memory-intensive it often requires clusters of GPUs for modern LLMs.
LoRA fundamentally changes this math by exploiting the inherent low rank of model updates.
The Low-Rank Primitive
Instead of modifying the original weights, LoRA freezes the pre-trained model entirely and learns a separate, smaller adjustment. A full delta matrix would still be huge. LoRA's key insight is that model updates have a low intrinsic rank. It breaks the giant delta matrix into two small matrices ( and ) connected by a narrow bottleneck rank (). Multiplying them reconstructs the delta, but uses a tiny fraction of the parameters.
The Forward Pass
In practice, the model processes data through both paths in parallel. The input is multiplied by the frozen pre-trained weights, and simultaneously passed through the learned and matrices. The outputs are then summed together. During inference, these weights can be pre-merged, meaning LoRA adds exactly zero latency to the final model.
The Rank Bottleneck
Because the update is forced through a tight bottleneck, LoRA excels at style transfer or format instruction—superficial changes that require very little information. However, it struggles with deep knowledge injection. If a task requires complex structural changes to the model's fundamental understanding, a low rank cannot carry enough information to express it.