Language Models2021intermediate11 min read

LoRA: Low-Rank Adaptation of Large Language Models

LoRA: التكيّف مُنخَفِض الرُّتبة للنماذج اللغوية الكبيرة

Hu, E. J. · Shen, Y. · Wallis, P. · Allen-Zhu, Z. · Li, Y. · Wang, S. · Wang, L. · Chen, W. — ICLR

The problem

Fine-tuning a like GPT-3 (175B parameters) means retraining every single weight — and storing a separate 350GB copy for each task. With dozens of tasks, storage alone reaches petabytes, and the GPU memory needed for states makes training infeasible for most teams. Existing -efficient methods (adapters, ) either add latency or reduce usable sequence length, and often fall short of quality.

The contribution

LoRA freezes all pre-trained weights and injects pairs of small trainable low- matrices (B·A) alongside each . Because the change ΔW = BA has very low rank (often r = 1–4), the drop by 10,000× while matching or exceeding full fine-tuning quality. At deployment, BA merges back into W₀, so there is zero additional inference latency — and switching tasks means swapping a 35MB file instead of a 350GB model.

The impact

LoRA made fine-tuning large language models accessible to anyone with a single GPU. It is the de facto standard for customizing LLMs — from Stable Diffusion image generation to enterprise chatbots. Its idea that weight updates live on a low-dimensional became a foundational insight, spawning , LoRA+, DoRA, and hundreds of variants. With over 9,000 citations, LoRA democratized the adaptation of billion-parameter models.

Imagine a grand piano with 175 billion strings, each tuned over years of concert performances. To play jazz instead of classical, you don't re-tune every string — you clip a small effects pedal to the output. The pedal has only a few knobs (low rank), yet it reshapes the piano's entire sound.

LoRA is that pedal: it leaves the original instrument untouched and learns a tiny adjustment that, when plugged in, transforms the model's behavior. Unplug it, and the piano is exactly as it was. Plug in a different pedal, and you play blues.

The problem: fine-tuning is expensive and wasteful

The standard recipe for adapting a pre-trained language model is full fine-tuning: take the pre-trained weights W0W_0 and update every parameter by following gradients on your task's data. The objective is to maximize the conditional log-:

max⁡Φ∑(x,y)∈Z∑t=1∣y∣log⁡PΦ(yt∣x,y<t)\max_{\Phi} \sum_{(x,y)\in Z}\sum_{t=1}^{|y|} \log P_{\Phi}(y_t \mid x, y_{<t})

For a 175-billion-parameter model like GPT-3, this means:

  • Storage: each fine-tuned copy is ~350 GB. Ten tasks → 3.5 TB just in checkpoints.
  • Memory: the Adam optimizer stores two extra state vectors per parameter, so training requires ~1.2 TB of GPU VRAM.
  • No sharing: the 175B parameters are duplicated in full for every task, even though most of the learned knowledge is shared.

Existing parameter-efficient alternatives tried to fix this but had trade-offs. insert small modules between Transformer layers, but they add sequential depth — increasing inference latency by up to 30% in online scenarios. Prefix tuning prepends trainable tokens to the input, but these steal sequence length from the actual task and perform poorly in low-data regimes.

Open in Lab
Compare the storage and memory costs of full fine-tuning vs. LoRA across model scales.
The demo wakes as you arrive…

The key insight: weight updates have low intrinsic rank

Why does LoRA work at all? The answer lies in an observation from prior research: although large language models have billions of parameters, the space they actually use when adapting to a new task is surprisingly small. Aghajanyan et al. (2020) showed that pre-trained models have a low — you can project the parameter space to a much smaller subspace and still learn effectively.

LoRA extends this observation: if the full model lives on a low-dimensional surface, then the change in weights during adaptation (ΔW\Delta W) should also be low-rank. Think of it this way: the already knows almost everything it needs — adaptation only needs to nudge a few specific directions in weight space. Those few directions form a low-rank subspace.

The authors confirmed this empirically: even with a rank as low as r=1r = 1 or r=2r = 2, LoRA matched full fine-tuning on GPT-3 175B across multiple benchmarks. The top singular vectors of adaptation matrices trained with different ranks overlap heavily, showing that the useful information concentrates in very few dimensions.

Open in Lab
Drag the rank slider to see how even r = 1 or 2 captures most of the useful adaptation signal. The overlap bars show subspace similarity between different ranks.
The demo wakes as you arrive…

How LoRA works: the reparametrization

The idea is beautifully simple. For a pre-trained weight matrix W0∈Rd×kW_0 \in \mathbb{R}^{d \times k}, instead of learning a full update ΔW∈Rd×k\Delta W \in \mathbb{R}^{d \times k} (which would have d×kd \times k parameters), we constrain ΔW\Delta W to be a low-rank product:

ΔW=BA\Delta W = BA

where B∈Rd×rB \in \mathbb{R}^{d \times r} and A∈Rr×kA \in \mathbb{R}^{r \times k}, with r≪min⁡(d,k)r \ll \min(d, k). The total trainable parameters drop from d×kd \times k to r×(d+k)r \times (d + k).

During , the output becomes:

h=W0x+ΔWx=W0x+BAxh = W_0 x + \Delta W x = W_0 x + B A x

The W0W_0 and the low-rank update BABA operate on the same input xx in parallel, and their outputs are summed. This is the key design: no extra sequential depth like adapters.

Initialization: AA is initialized with random Gaussian values and BB is initialized to zero, so ΔW=BA=0\Delta W = BA = 0 at the start of training. The model begins exactly as the pre-trained one. A α/r\alpha / r is applied to ΔWx\Delta W x to stabilize training across different values of rr.

h=W0x+αrBAxh = W_0 x + \frac{\alpha}{r} B A x
LoRA forward pass — the complete equation — W₀ = frozen pre-trained weights · B·A = trainable low-rank update · α/r = scaling constant · The two paths run in parallel and their outputs are summed
Open in Lab
Click each component to see how the frozen weights W₀ and the trainable low-rank matrices B and A combine during the forward pass.
The demo wakes as you arrive…

Where to inject LoRA in a Transformer

A Transformer has four weight matrices in each module (WqW_q, WkW_k, WvW_v, WoW_o) and two in each MLP block. The authors experimented with applying LoRA to different subsets of these matrices and found that adapting both WqW_q and WvW_v gives the best results for a given parameter budget.

This is a revealing finding: even a rank of just r=4r = 4 spread across both query and value matrices captures enough information — it is better to adapt more matrices at lower rank than fewer matrices at higher rank. In practical terms, with r=4r = 4 applied to WqW_q and WvW_v across all 96 layers of GPT-3 175B, the total trainable parameters are only 18 million — roughly 0.01% of the full model.

The MLP layers are left frozen in the paper's experiments. While adapting them too would likely help, the attention matrices alone proved sufficient for matching full fine-tuning quality.

Open in Lab
Click on different weight matrices to see where LoRA adapters are placed. Toggle adapting different combinations and watch the parameter count change.
The demo wakes as you arrive…

Zero-latency deployment: merge and swap

Here is LoRA's most elegant trick. During training, W0W_0 and BABA are kept separate so we can compute gradients for AA and BB alone. But at deployment, we simply compute:

W=W0+BAW = W_0 + BA

and replace the original weight matrix with this merged version. The resulting model has exactly the same architecture as the original — no adapter layers, no extra modules, no added computation. Inference runs at the same speed as a model that was fully fine-tuned.

To switch tasks, subtract BABA from WW to recover W0W_0, then add a different B′A′B'A' for the new task. The LoRA matrices for a GPT-3 175B task are only ~35 MB — meaning 100 task-specific adapters plus the base model fit in ~354 GB, versus ~35 TB if each task required a full copy.

Open in Lab
Watch the merge process: during training, the paths are separate. Press "Deploy" to merge BA into W₀. Press "Switch Task" to swap adapters.
The demo wakes as you arrive…

What ΔW actually learns

The authors investigated the relationship between ΔW\Delta W and the original WW, and their findings offer a window into why adaptation works. Three key results emerged:

  • ΔW\Delta W is correlated with WW, but not redundant. The adaptation matrix amplifies certain directions in WW, but specifically those directions that are not already emphasized. It finds the features the pre-trained model learned but did not prioritize.
  • The is huge. For r=4r = 4, the adaptation matrix amplifies its chosen directions by a factor of ~21×. The pre-trained model already encoded the task-relevant knowledge — it just needed these specific directions turned up.
  • Different tasks amplify different directions. This explains why a generic pre-trained model can be adapted to many different tasks: each task needs a different, small set of weight-space directions amplified.

This analysis reveals a beautiful picture: a large pre-trained model is like a Swiss Army knife with every tool folded in. LoRA doesn't add new tools — it simply unfolds the right ones for each task.

The same idea in code

LoRA forward pass in pure NumPypython

Simplified to show the idea — not the real implementation.

import numpy as np

def lora_forward(x, W0, A, B, alpha, r):
    """
    x:  input vector/matrix      (batch, d_in)
    W0: frozen pre-trained weight (d_out, d_in)
    A:  low-rank down-projection  (r, d_in)   ← Gaussian init
    B:  low-rank up-projection    (d_out, r)   ← zero init
    alpha: scaling constant
    r: rank of the low-rank update
    """
    # Original path: frozen, no gradient needed
    h_frozen = x @ W0.T

    # LoRA path: only A and B receive gradients
    h_lora = x @ A.T @ B.T   # x→(r)→(d_out), bottleneck!

    # Combine with scaling
    return h_frozen + (alpha / r) * h_lora

# At deployment time, merge and forget:
# W_merged = W0 + (alpha / r) * B @ A
# Now inference is just: h = x @ W_merged.T
# No extra compute. No extra latency. That's the trick.

# To switch tasks:
# W0 = W_merged - (alpha / r) * B @ A   (recover original)
# W_new = W0 + (alpha / r) * B_task2 @ A_task2

Results across model scales

The authors evaluated LoRA across four model families — RoBERTa (125M/355M), DeBERTa (1.5B), GPT-2 (345M/774M), and GPT-3 (175B) — on tasks ranging from natural language understanding () to natural language generation (E2E, WebNLG, DART).

On RoBERTa base with just 0.3M trainable parameters, LoRA achieved 87.2% average on GLUE — surpassing full fine-tuning (86.4%) and adapter methods (84.4–85.4%). On GPT-3 175B, LoRA with 4.7M parameters matched or exceeded full fine-tuning (175B parameters) on WikiSQL, MNLI, and SAMSum. The most striking result: LoRA with rv=2r_v = 2 on GPT-3 (only 4.7M parameters — 0.0027% of the model) achieved 73.4% on WikiSQL, just 0.4% below full fine-tuning.

Unlike prefix tuning, LoRA's performance does not degrade when you add more trainable parameters. And unlike adapters, it introduces zero additional inference latency, since the low-rank matrices merge into the base weights at deployment time.

Open in Lab
Adjust the rank r and select which weight matrices to adapt. Watch how parameter count and accuracy change — notice that r = 1 or 2 already performs remarkably well.
The demo wakes as you arrive…

Why it mattered

  1. 2019

    Adapter Tuning (Houlsby et al.)

    Inserted small bottleneck modules between Transformer layers. Reduced trainable parameters, but added sequential depth and inference latency.

  2. 2021

    Prefix Tuning (Li & Liang)

    Prepended trainable continuous tokens to the input. No inference latency, but stole sequence length and was difficult to optimize.

  3. 2021

    LoRA (this paper)

    Froze all weights and injected parallel low-rank updates. No inference latency, no sequence length reduction, matched or exceeded full fine-tuning quality.

  4. 2023

    QLoRA (Dettmers et al.)

    Combined LoRA with 4-bit quantization of the frozen weights. Fine-tuned a 65B model on a single 48GB GPU — pushing accessibility even further.

  5. 2024

    LoRA+ and DoRA

    LoRA+ showed that using different learning rates for A and B improves performance. DoRA decomposed weights into magnitude and direction, achieving closer parity with full fine-tuning.

LoRA's impact extends far beyond language models. It became the standard method for customizing Stable Diffusion image generators, where users train small LoRA adapters to generate specific artistic styles or characters. In enterprise settings, companies deploy a single base LLM with hundreds of task-specific LoRA adapters — one for legal documents, another for customer support, another for code generation — all sharing the same GPU memory. The idea that adaptation is low-rank opened a research direction that continues to produce new variants every month.

CitationHu, Shen, Wallis, Allen-Zhu, Li, Wang, Wang, Chen. LoRA: Low-Rank Adaptation of Large Language Models. ICLR, 2022.

Terms in this paper