Language Models2021intermediate11 min read
LoRA: Low-Rank Adaptation of Large Language Models
LoRA: التكيّف مُنخَفِض الرُّتبة للنماذج اللغوية الكبيرة
Hu, E. J. · Shen, Y. · Wallis, P. · Allen-Zhu, Z. · Li, Y. · Wang, S. · Wang, L. · Chen, W. — ICLR
The problem
Fine-tuning a like GPT-3 (175B parameters) means retraining every single weight — and storing a separate 350GB copy for each task. With dozens of tasks, storage alone reaches petabytes, and the GPU memory needed for states makes training infeasible for most teams. Existing -efficient methods (adapters, ) either add latency or reduce usable sequence length, and often fall short of quality.
The contribution
LoRA freezes all pre-trained weights and injects pairs of small trainable low- matrices (B·A) alongside each . Because the change ΔW = BA has very low rank (often r = 1–4), the drop by 10,000× while matching or exceeding full fine-tuning quality. At deployment, BA merges back into W₀, so there is zero additional inference latency — and switching tasks means swapping a 35MB file instead of a 350GB model.
The impact
LoRA made fine-tuning large language models accessible to anyone with a single GPU. It is the de facto standard for customizing LLMs — from Stable Diffusion image generation to enterprise chatbots. Its idea that weight updates live on a low-dimensional became a foundational insight, spawning , LoRA+, DoRA, and hundreds of variants. With over 9,000 citations, LoRA democratized the adaptation of billion-parameter models.
Imagine a grand piano with 175 billion strings, each tuned over years of concert performances. To play jazz instead of classical, you don't re-tune every string — you clip a small effects pedal to the output. The pedal has only a few knobs (low rank), yet it reshapes the piano's entire sound.
LoRA is that pedal: it leaves the original instrument untouched and learns a tiny adjustment that, when plugged in, transforms the model's behavior. Unplug it, and the piano is exactly as it was. Plug in a different pedal, and you play blues.
The problem: fine-tuning is expensive and wasteful
The standard recipe for adapting a pre-trained language model is full fine-tuning: take the pre-trained weights and update every parameter by following gradients on your task's data. The objective is to maximize the conditional log-:
For a 175-billion-parameter model like GPT-3, this means:
- Storage: each fine-tuned copy is ~350 GB. Ten tasks → 3.5 TB just in checkpoints.
- Memory: the Adam optimizer stores two extra state vectors per parameter, so training requires ~1.2 TB of GPU VRAM.
- No sharing: the 175B parameters are duplicated in full for every task, even though most of the learned knowledge is shared.
Existing parameter-efficient alternatives tried to fix this but had trade-offs. insert small modules between Transformer layers, but they add sequential depth — increasing inference latency by up to 30% in online scenarios. Prefix tuning prepends trainable tokens to the input, but these steal sequence length from the actual task and perform poorly in low-data regimes.
The key insight: weight updates have low intrinsic rank
Why does LoRA work at all? The answer lies in an observation from prior research: although large language models have billions of parameters, the space they actually use when adapting to a new task is surprisingly small. Aghajanyan et al. (2020) showed that pre-trained models have a low — you can project the parameter space to a much smaller subspace and still learn effectively.
LoRA extends this observation: if the full model lives on a low-dimensional surface, then the change in weights during adaptation () should also be low-rank. Think of it this way: the already knows almost everything it needs — adaptation only needs to nudge a few specific directions in weight space. Those few directions form a low-rank subspace.
The authors confirmed this empirically: even with a rank as low as or , LoRA matched full fine-tuning on GPT-3 175B across multiple benchmarks. The top singular vectors of adaptation matrices trained with different ranks overlap heavily, showing that the useful information concentrates in very few dimensions.
How LoRA works: the reparametrization
The idea is beautifully simple. For a pre-trained weight matrix , instead of learning a full update (which would have parameters), we constrain to be a low-rank product:
where and , with . The total trainable parameters drop from to .
During , the output becomes:
The and the low-rank update operate on the same input in parallel, and their outputs are summed. This is the key design: no extra sequential depth like adapters.
Initialization: is initialized with random Gaussian values and is initialized to zero, so at the start of training. The model begins exactly as the pre-trained one. A is applied to to stabilize training across different values of .
Where to inject LoRA in a Transformer
A Transformer has four weight matrices in each module (, , , ) and two in each MLP block. The authors experimented with applying LoRA to different subsets of these matrices and found that adapting both and gives the best results for a given parameter budget.
This is a revealing finding: even a rank of just spread across both query and value matrices captures enough information — it is better to adapt more matrices at lower rank than fewer matrices at higher rank. In practical terms, with applied to and across all 96 layers of GPT-3 175B, the total trainable parameters are only 18 million — roughly 0.01% of the full model.
The MLP layers are left frozen in the paper's experiments. While adapting them too would likely help, the attention matrices alone proved sufficient for matching full fine-tuning quality.
Zero-latency deployment: merge and swap
Here is LoRA's most elegant trick. During training, and are kept separate so we can compute gradients for and alone. But at deployment, we simply compute:
and replace the original weight matrix with this merged version. The resulting model has exactly the same architecture as the original — no adapter layers, no extra modules, no added computation. Inference runs at the same speed as a model that was fully fine-tuned.
To switch tasks, subtract from to recover , then add a different for the new task. The LoRA matrices for a GPT-3 175B task are only ~35 MB — meaning 100 task-specific adapters plus the base model fit in ~354 GB, versus ~35 TB if each task required a full copy.
What ΔW actually learns
The authors investigated the relationship between and the original , and their findings offer a window into why adaptation works. Three key results emerged:
- is correlated with , but not redundant. The adaptation matrix amplifies certain directions in , but specifically those directions that are not already emphasized. It finds the features the pre-trained model learned but did not prioritize.
- The is huge. For , the adaptation matrix amplifies its chosen directions by a factor of ~21×. The pre-trained model already encoded the task-relevant knowledge — it just needed these specific directions turned up.
- Different tasks amplify different directions. This explains why a generic pre-trained model can be adapted to many different tasks: each task needs a different, small set of weight-space directions amplified.
This analysis reveals a beautiful picture: a large pre-trained model is like a Swiss Army knife with every tool folded in. LoRA doesn't add new tools — it simply unfolds the right ones for each task.
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def lora_forward(x, W0, A, B, alpha, r):
"""
x: input vector/matrix (batch, d_in)
W0: frozen pre-trained weight (d_out, d_in)
A: low-rank down-projection (r, d_in) ← Gaussian init
B: low-rank up-projection (d_out, r) ← zero init
alpha: scaling constant
r: rank of the low-rank update
"""
# Original path: frozen, no gradient needed
h_frozen = x @ W0.T
# LoRA path: only A and B receive gradients
h_lora = x @ A.T @ B.T # x→(r)→(d_out), bottleneck!
# Combine with scaling
return h_frozen + (alpha / r) * h_lora
# At deployment time, merge and forget:
# W_merged = W0 + (alpha / r) * B @ A
# Now inference is just: h = x @ W_merged.T
# No extra compute. No extra latency. That's the trick.
# To switch tasks:
# W0 = W_merged - (alpha / r) * B @ A (recover original)
# W_new = W0 + (alpha / r) * B_task2 @ A_task2Results across model scales
The authors evaluated LoRA across four model families — RoBERTa (125M/355M), DeBERTa (1.5B), GPT-2 (345M/774M), and GPT-3 (175B) — on tasks ranging from natural language understanding () to natural language generation (E2E, WebNLG, DART).
On RoBERTa base with just 0.3M trainable parameters, LoRA achieved 87.2% average on GLUE — surpassing full fine-tuning (86.4%) and adapter methods (84.4–85.4%). On GPT-3 175B, LoRA with 4.7M parameters matched or exceeded full fine-tuning (175B parameters) on WikiSQL, MNLI, and SAMSum. The most striking result: LoRA with on GPT-3 (only 4.7M parameters — 0.0027% of the model) achieved 73.4% on WikiSQL, just 0.4% below full fine-tuning.
Unlike prefix tuning, LoRA's performance does not degrade when you add more trainable parameters. And unlike adapters, it introduces zero additional inference latency, since the low-rank matrices merge into the base weights at deployment time.
Why it mattered
2019
Adapter Tuning (Houlsby et al.)
Inserted small bottleneck modules between Transformer layers. Reduced trainable parameters, but added sequential depth and inference latency.
2021
Prefix Tuning (Li & Liang)
Prepended trainable continuous tokens to the input. No inference latency, but stole sequence length and was difficult to optimize.
2021
LoRA (this paper)
Froze all weights and injected parallel low-rank updates. No inference latency, no sequence length reduction, matched or exceeded full fine-tuning quality.
2023
QLoRA (Dettmers et al.)
Combined LoRA with 4-bit quantization of the frozen weights. Fine-tuned a 65B model on a single 48GB GPU — pushing accessibility even further.
2024
LoRA+ and DoRA
LoRA+ showed that using different learning rates for A and B improves performance. DoRA decomposed weights into magnitude and direction, achieving closer parity with full fine-tuning.
LoRA's impact extends far beyond language models. It became the standard method for customizing Stable Diffusion image generators, where users train small LoRA adapters to generate specific artistic styles or characters. In enterprise settings, companies deploy a single base LLM with hundreds of task-specific LoRA adapters — one for legal documents, another for customer support, another for code generation — all sharing the same GPU memory. The idea that adaptation is low-rank opened a research direction that continues to produce new variants every month.
CitationHu, Shen, Wallis, Allen-Zhu, Li, Wang, Wang, Chen. LoRA: Low-Rank Adaptation of Large Language Models. ICLR, 2022.
Terms in this paper
- Low-Rank Adaptationالتكيّف مُنخَفِض الرُّتبة
- Low-Rank Decompositionالتفكيك مُنخَفِض الرُّتبة
- Intrinsic Dimensionالبُعد الجوهري
- Intrinsic Rankالرُّتبة الجوهرية
- Amplification Factorمُعامِل التضخيم
- Scaling Factorمعامل القياس
- Query Projectionإسقاط الاستعلام
- Value Projectionإسقاط القيمة
- Task Switchingتبديل المهام
- QLoRAQLoRA