Model Compression2022intermediate10 min read

GPTQ: Accurate Post-Training Quantization for Generative Pre-Trained Transformers

GPTQ: تكميم دقيق بعد التدريب للمحوِّلات التوليدية المُدرَّبة مسبقاً

Frantar, E. · Ashkboos, S. · Hoefler, T. · Alistarh, D. — ICLR

The problem

By late 2022, the largest language models like GPT-3 (175B parameters) and OPT-175B required hundreds of gigabytes just to store their weights, needing multiple high-end GPUs even for alone. techniques existed, but the accurate ones required expensive retraining, and the fast methods — like simple — only worked at 8 bits and collapsed at 4 or 3 bits. No one had successfully quantized a 175-billion- model to 3–4 bits without catastrophic accuracy loss.

The contribution

GPTQ: a one-shot post-training method based on approximate second-order () information. Three key innovations make it scalable: (1) quantizing all rows in a fixed column order so the Hessian inverse is computed only once, (2) lazy batch updates that process 128 columns at a time for efficiency, and (3) of the inverse Hessian for numerical stability. GPTQ quantizes OPT-175B to 3–4 bits in approximately four GPU hours with negligible increase, enabling a 175B model to run on a single GPU for the first time. It also achieves 3.25× inference speedup on A100 and 4.5× on A6000 GPUs.

The impact

GPTQ proved that aggressive post-training quantization to 3–4 bits is feasible for the largest language models, democratizing access to models that previously required expensive multi-GPU setups. It became the de facto standard for quantization in the open-source LLM ecosystem, powering tools like AutoGPTQ and the Hugging Face Transformers integration. Its -wise Hessian-based framework directly influenced successor methods like AWQ, SqueezeLLM, and QuIP, and its insights on fixed-order quantization and lazy batching became foundational techniques in the quantization literature. GPTQ also paved the way for QLoRA, which combines 4-bit quantization with parameter-efficient fine-tuning.

Imagine a master sculptor working with a marble statue that weighs ten tons. She needs to move it into a small gallery, but the doorway is too narrow. She cannot just chop pieces off randomly — the statue would lose its beauty. Instead, she carefully shaves each curve, checking after every cut whether the statue still looks right. If a cut on the left arm changes the balance, she compensates by adjusting the right shoulder.

GPTQ does this with numbers. Every weight in a is a chisel mark on the statue. GPTQ rounds each weight to a cruder value (fewer bits), but after every rounding it adjusts the neighboring weights so the layer's output barely changes. The key tool is the Hessian matrix — a mathematical map that tells you which weights are sensitive and which can absorb error.

The memory wall: why large models need quantization

GPT-3 has 175 billion parameters. Stored in 16-bit floating point (FP16), that is 326 GB of memory — more than any single GPU can hold. Running inference requires splitting the model across multiple expensive GPUs, making it inaccessible to most researchers and developers.

The fundamental bottleneck is not computation but . During , each token requires loading the entire model from memory. The speed of generation is limited by how fast you can read weights, not how fast you can multiply them. If you could store each weight in 4 bits instead of 16, you would need 4× less memory and could load weights 4× faster — a direct speedup for generation.

This is where quantization comes in: replacing high-precision numbers (16-bit floats) with low-precision ones (3–4 bit integers). The challenge is doing this without destroying the model's accuracy.

Open in Lab
Adjust the model size and bit-width to see how quantization reduces memory. Notice how 4-bit quantization can fit a 175B model on a single GPU.
The demo wakes as you arrive…

The core idea: quantize one layer at a time, minimize reconstruction error

Rather than quantizing the entire model at once, GPTQ works layer by layer. For each linear layer with weight matrix W\mathbf{W} and a small set of inputs X\mathbf{X} (just 128 samples), the goal is to find quantized weights W^\hat{\mathbf{W}} that keep the layer's output as close as possible to the original. Formally:

argminW^  ∥WX−W^X∥22\text{argmin}_{\hat{\mathbf{W}}} \; \|\mathbf{W}\mathbf{X} - \hat{\mathbf{W}}\mathbf{X}\|_2^2
Layer-wise reconstruction objective — Find quantized weights W^\hat{\mathbf{W}} that minimize the squared difference between the original and quantized layer outputs, measured on the calibration data X\mathbf{X}. This is not about the final model loss — it is about preserving each layer's input-output behavior. If every layer barely changes, the whole model barely changes.

Think of this like a relay race: if every runner keeps almost the same pace as the original team, the total time stays almost the same. One runner slowing down slightly does not ruin the race — but every runner collapsing would. The layer-wise approach ensures each "runner" (layer) maintains its pace.

The predecessor: Optimal Brain Quantization (OBQ)

GPTQ builds on Optimal Brain Quantization (OBQ), which itself extends the classic Optimal Brain Surgeon framework from to quantization. OBQ's idea: quantize weights one at a time, and after each quantized weight, adjust all remaining full-precision weights to compensate for the rounding error. The adjustment direction is guided by the inverse Hessian H−1\mathbf{H}^{-1}.

Concretely, when quantizing weight wqw_q at position qq, the quantization error is wq−quant(wq)w_q - \text{quant}(w_q), and the optimal adjustment to all remaining weights is:

δF=−wq−quant(wq)[HF−1]qq⋅(HF−1):,q\boldsymbol{\delta}_F = -\frac{w_q - \text{quant}(w_q)}{[\mathbf{H}_F^{-1}]_{qq}} \cdot (\mathbf{H}_F^{-1})_{:,q}
OBQ weight update — distributing quantization error via the Hessian — The quantization error of wqw_q is divided by the diagonal Hessian entry (which measures that weight's importance) and then spread along the qq-th column of the inverse Hessian to all remaining weights. Weights that are correlated with wqw_q — according to the Hessian — receive larger compensating adjustments.

The problem: OBQ processes each row of W\mathbf{W} independently and picks a different greedy order for each row. This means the Hessian inverse must be updated separately for every row of every layer. For a weight matrix of size drow×dcold_{\text{row}} \times d_{\text{col}}, the runtime is O(drow⋅dcol3)O(d_{\text{row}} \cdot d_{\text{col}}^3) — cubic in the column dimension. On ResNet-50 (25M parameters), this takes about an hour. On OPT-175B, it would take years.

GPTQ's three innovations: from years to hours

GPTQ makes three modifications to OBQ that together reduce the runtime by over three orders of magnitude — enough to quantize OPT-175B in about four hours on one GPU.

Open in Lab
Walk through the three key innovations that make GPTQ fast enough for 175B-parameter models. Toggle each step to see its effect on runtime.
The demo wakes as you arrive…

Innovation 1 — Fixed column order. The authors discovered that on large, heavily-parameterized layers, the greedy ordering of OBQ gives only marginal improvement over a fixed arbitrary order. The reason is intuitive: in a layer with thousands of weights, the handful of "bad" rounding decisions caused by a suboptimal order are diluted among thousands of "good" compensations. By quantizing all rows in the same fixed column order, the Hessian inverse HF−1\mathbf{H}_F^{-1} is shared across all rows. The inverse needs to be updated only dcold_{\text{col}} times instead of drow×dcold_{\text{row}} \times d_{\text{col}} times, reducing runtime by a factor of drowd_{\text{row}}.

Innovation 2 — Lazy batch updates. Even with a fixed order, updating the entire remaining weight matrix after each single weight is memory-bandwidth-bound — too many memory reads and writes for too few floating-point operations. GPTQ batches the updates: it processes blocks of B=128B = 128 columns at a time, keeping intermediate updates local to the block. Only after finishing a block does it apply a single large matrix update to the remaining weights. This transforms the workload from many small memory-bound operations into a few large compute-bound matrix multiplications that GPUs excel at.

Innovation 3 — Cholesky-based Hessian inverse. Repeatedly removing rows and columns from the inverse Hessian via Gaussian elimination accumulates numerical errors. GPTQ pre-computes the Cholesky decomposition of H−1\mathbf{H}^{-1} once, then reads off the needed inverse entries row by row. This is both faster and more numerically stable, adding a small dampening term λI\lambda \mathbf{I} to handle near-singular Hessians.

The algorithm: column by column, block by block

Here is the full GPTQ procedure for one layer. The input is the weight matrix W\mathbf{W} and the Hessian H=2XX⊤\mathbf{H} = 2\mathbf{X}\mathbf{X}^\top computed from calibration data. The output is the quantized matrix W^\hat{\mathbf{W}}.

GPTQ Algorithm — simplified pseudocodepython

Simplified to show the idea — not the real implementation.

# Input: W (weight matrix), H (Hessian), B (block size = 128) # Output: Q (quantized weight matrix)
H_inv = cholesky(inverse(H + λI))   # Step 3: stable inverse Q = copy(W)
for block_start in range(0, num_cols, B):
    block = columns[block_start : block_start + B]
    error_accumulator = zeros(num_rows, B)

    for j in block:                   # within-block loop
        q_j = quantize(Q[:, j])       # round to nearest grid point
        error = Q[:, j] - q_j         # quantization error
        Q[:, j] = q_j                 # commit quantized column

        # Compensate: adjust remaining block columns
        # using Hessian inverse (Step 1 insight: same for all rows)
        Q[:, j+1:block_end] -= (error / H_inv[j,j]) * H_inv[j, j+1:block_end]

        error_accumulator[:, j - block_start] = error / H_inv[j,j]

    # Step 2: lazy batch update — one big matrix multiply
    Q[:, block_end:] -= error_accumulator @ H_inv[block, block_end:]
Open in Lab
Watch GPTQ quantize a weight matrix column by column. See how each rounding error is compensated by adjusting the remaining weights. The Hessian guides where the error flows.
The demo wakes as you arrive…

Results: 4-bit weights with near-lossless accuracy

The headline result: GPTQ quantizes OPT-175B and BLOOM-176B to 4 bits with a perplexity increase of less than 0.05 points on WikiText-2 — compared to round-to-nearest (RTN) which often increases perplexity by 1–2 points or more. At 3 bits, GPTQ still produces usable models where RTN completely fails with perplexity exploding by orders of magnitude.

Crucially, the quality gap between GPTQ and RTN grows with model size. On small models (125M parameters), RTN is competitive. But on OPT-175B, RTN at 4 bits gives a perplexity of 10.54 versus GPTQ's 8.37 (FP16 baseline: 8.34). The larger the model, the more GPTQ matters — precisely the regime where compression is most needed.

Open in Lab
Compare GPTQ vs RTN vs FP16 baseline across model sizes. Notice how GPTQ stays close to the baseline while RTN diverges on larger models and lower bit-widths.
The demo wakes as you arrive…

Practical impact: 175B on a single GPU

Beyond accuracy, GPTQ delivers real inference speedups. The authors developed custom GPU kernels that dequantize weights on-the-fly during matrix multiplication. Since autoregressive generation is memory-bandwidth-bound (not compute-bound), loading 4× fewer bytes per weight translates directly to faster generation.

The results: approximately 3.25× speedup on NVIDIA A100 and 4.5× on NVIDIA A6000 GPUs for end-to-end generative inference. More importantly, a 4-bit OPT-175B fits entirely in the 80 GB of a single A100 — the first time a 175-billion-parameter model ran on one GPU for generative tasks.

Limitations and open questions

GPTQ quantizes weights only — activations remain in FP16. For most generative scenarios this is fine because the bottleneck is weight loading, not activation computation. But for batch inference with many simultaneous requests, activation quantization becomes important too, and GPTQ does not address this.

The method also does not accelerate the actual multiply operations because mainstream hardware lacks native support for mixed-precision formats like FP16 × INT4. The speedup comes entirely from reduced memory traffic. Finally, at 2-bit and ternary quantization, accuracy degrades noticeably — GPTQ pushes boundaries but does not eliminate the accuracy-compression tradeoff.

Legacy: the quantization ecosystem

  1. 1989

    Optimal Brain Damage (LeCun et al.)

    Introduced using second-order (Hessian) information to decide which weights to remove from a neural network. The diagonal Hessian approximation made pruning decisions principled rather than arbitrary.

  2. 1993

    Optimal Brain Surgeon (Hassibi & Stork)

    Extended OBD to use the full (non-diagonal) inverse Hessian for weight removal, allowing remaining weights to compensate. The direct ancestor of GPTQ's error compensation mechanism.

  3. 2022

    LLM.int8() (Dettmers et al.)

    Showed that 8-bit quantization of large models requires handling activation outliers specially. Introduced mixed-precision decomposition keeping outlier dimensions in FP16.

  4. 2022

    OBQ / Optimal Brain Compression (Frantar & Alistarh)

    Generalized Optimal Brain Surgeon to post-training quantization. Achieved state-of-the-art accuracy but with cubic runtime that limited it to models under 100M parameters.

  5. 2022

    GPTQ (this paper)

    Made Hessian-based quantization practical for 175B-parameter models. Fixed column order, lazy batching, and Cholesky decomposition cut runtime by 1000×. First to run OPT-175B on a single GPU.

  6. 2023

    QLoRA (Dettmers et al.)

    Combined 4-bit GPTQ-style quantization with Low-Rank Adaptation for fine-tuning. Enabled fine-tuning a 65B model on a single 48-GB GPU by keeping the base model in 4-bit and training only low-rank adapters in higher precision.

  7. 2023

    AWQ, SqueezeLLM, QuIP

    Successor methods that build on GPTQ's framework. AWQ protects salient weights via activation-aware scaling. SqueezeLLM handles outliers with sparse storage. QuIP adds incoherence processing for better error bounds.

CitationFrantar, Ashkboos, Hoefler, Alistarh. GPTQ: Accurate Post-Training Quantization for Generative Pre-Trained Transformers. ICLR, 2023.

Terms in this paper