Model Compression2022intermediate11 min read

LLM.int8(): 8-Bit Matrix Multiplication for Transformers at Scale

LLM.int8(): ضرب المصفوفات بدقة 8 بت للمحوِّلات على نطاق واسع

Dettmers, T. · Lewis, M. · Belkada, Y. · Zettlemoyer, L. — NeurIPS

The problem

Large language models like GPT-3 (175B parameters) require enormous memory just to load their weights for — often hundreds of gigabytes. An obvious solution is to quantize weights from 16-bit to 8-bit integers, halving memory. But prior 8-bit methods had only been tested on models up to 350M parameters. At larger scales, they cause severe performance degradation, and nobody understood why.

The contribution

A two-part quantization method called LLM.int8() that enables degradation-free 8-bit inference for transformers up to 175B parameters. First, vector-wise quantization uses per-row and per-column scaling constants instead of one per , improving precision. Second, mixed-precision decomposition isolates rare but critical feature dimensions (~0.1%) into 16-bit multiplication while keeping the remaining 99.9% in 8-bit. The paper also provides the first systematic analysis of emergent outlier features in large transformers: they appear in only ~6 feature dimensions yet dominate model performance, and their sudden emergence at 6.7B parameters explains why prior quantization methods failed at scale.

The impact

LLM.int8() made large model inference accessible on consumer hardware for the first time — OPT-175B and BLOOM could run on a single server with consumer GPUs. It was integrated into Hugging Face Transformers via the bitsandbytes library, becoming the default way to load large models in reduced precision. The outlier analysis directly informed GPTQ and QLoRA, two of the most influential quantization methods that followed. The paper established that understanding emergent features is essential for any compression technique applied to large models.

Imagine a warehouse of identical-looking filing cabinets — millions of them, each holding a single number written to 16 decimal places. You need to move the entire warehouse to a building with half the floor space. Most numbers are small and well-behaved: rounding them to fewer digits loses nothing important.

But scattered across a handful of specific cabinets — the same six cabinet positions on every floor — are giant numbers, 20 times larger than anything else. If you round those, the entire filing system collapses. The information they carry is disproportionately critical.

LLM.int8() is the moving strategy: round 99.9% of the cabinets to save space, but hand-carry those six critical positions at full precision. Nothing is lost, and the warehouse fits.

The memory wall: why large models need compression

A 175-billion model stored in 16-bit floating point requires about 350 GB of GPU memory — just to hold the weights, before any activations or batch data. That exceeds the capacity of even the most expensive enterprise GPUs. In 2022, running inference on such a model required a multi-GPU cluster costing tens of thousands of dollars.

The feed-forward and projection layers account for 95% of all parameters and 65–85% of all computation. If we could quantize these layers from 16-bit to 8-bit integers, we would halve the memory footprint: 175 GB instead of 350 GB. This puts models like OPT-175B and BLOOM within reach of a single server with consumer GPUs.

But quantization is not free. Reducing precision means representing continuous values with fewer discrete bins. If the distribution of values is well-behaved, the rounding error is small. If it is not — if there are extreme outliers — quantization can destroy the information the model needs to function.

Open in Lab
Compare GPU memory requirements for different model sizes in 16-bit vs 8-bit. Models that were previously inaccessible become runnable on consumer hardware.
The demo wakes as you arrive…

How quantization works: from 16 bits to 8

Quantization maps high-precision floating-point numbers to low-precision integers. The simplest method is : divide every value in a tensor by the largest absolute value, multiply by 127, and round to the nearest integer. This maps the tensor into the Int8 range [−127,127][-127, 127].

The idea is like rescaling a thermometer. If temperatures range from −30°-30° to +30°+30°, you rescale so that −30°-30° maps to −127-127 and +30°+30° maps to +127+127. Every degree gets about 4 bins — plenty of resolution. But if one day hits 300°300°, then 300°300° maps to 127127 and all normal temperatures are squeezed into a tiny range around zero. One outlier destroys the precision for everything else.

Xi8=⌊127∥Xf16∥∞⋅Xf16⌉=⌊sx⋅Xf16⌉X_{\mathrm{i8}} = \left\lfloor \frac{127}{\|X_{\mathrm{f16}}\|_\infty} \cdot X_{\mathrm{f16}} \right\rceil = \left\lfloor s_x \cdot X_{\mathrm{f16}} \right\rceil
Absmax quantization — scaling by the infinity norm — The scaling factor sx=127/∥X∥∞s_x = 127 / \|X\|_\infty maps the entire tensor into [−127,127][-127, 127]. If the largest value is 10, then sx=12.7s_x = 12.7 and the tensor has good resolution. If the largest value is 100, then sx=1.27s_x = 1.27 and small values lose precision. The ⌊⋅⌉\lfloor \cdot \rceil denotes rounding to the nearest integer.
Open in Lab
Drag the outlier slider to see how a single large value compresses the quantization bins for all other values. Watch how precision degrades.
The demo wakes as you arrive…

A more sophisticated method is , which shifts the distribution so that asymmetric ranges use the full [−127,127][-127, 127] interval. This helps with activations that are all positive (like ReLU outputs), since absmax would waste half its range on negative values that never appear.

However, zeropoint quantization requires special instructions that GPUs do not natively support efficiently, making it slower in practice. The paper focuses on absmax as the practical choice and enhances it with two innovations.

Innovation 1: vector-wise quantization

Standard quantization uses a single for the entire tensor — one number to rule all values. If any single value is an outlier, it sets the scale for millions of other values. This is called tensor-wise quantization.

The key insight is that matrix multiplication can be viewed as a collection of independent inner products. Each inner product involves one row of the hidden states and one column of the . Since they are independent, each can have its own scaling constant.

Think of it like grading exams in different classrooms. Tensor-wise quantization is like grading all students on the same curve — one genius student skews the scale for everyone. Vector-wise quantization is like giving each classroom its own curve. Each inner product gets an independent scale that fits its own range of values, so no single outlier can dominate the entire computation.

Cf16≈1cf16x⊗cf16w⋅Ci32=S⋅Q(Af16) Q(Bf16)C_{\mathrm{f16}} \approx \frac{1}{c^x_{\mathrm{f16}} \otimes c^w_{\mathrm{f16}}} \cdot C_{\mathrm{i32}} = S \cdot Q(A_{\mathrm{f16}}) \, Q(B_{\mathrm{f16}})
Vector-wise quantization — independent scaling per row and column — Each row of the hidden states gets its own scaling constant cxc^x, and each column of the weight matrix gets its own cwc^w. To recover the output, we denormalize by the outer product cx⊗cwc^x \otimes c^w. This increases precision without requiring custom hardware — it uses the same Int8 multiply instructions, just with finer normalization.
Open in Lab
Compare tensor-wise vs vector-wise quantization. Toggle between them to see how independent scaling constants improve precision for each inner product.
The demo wakes as you arrive…

The discovery: emergent outlier features at scale

The paper's most important contribution may be its analysis of emergent outlier features. As transformers grow, something dramatic happens to the hidden states during inference: certain feature dimensions develop values with magnitudes 20 times larger than any other dimension.

These outliers are not random noise. They are highly systematic: in a 6.7B parameter processing a 2048- sequence, about 150,000 outlier values appear across all layers — but they concentrate in only 6 feature dimensions out of thousands. The same 6 dimensions, across every , every token, every sequence.

The emergence follows a . Below 6B parameters, outliers appear sporadically in about 25% of layers. Between 6B and 6.7B, a sharp shift occurs: outliers suddenly appear in 100% of layers and 75% of all sequence positions. This is the exact point where standard quantization methods collapse.

Open in Lab
Slide the model size to watch outlier features emerge. Notice the sharp phase transition at 6.7B parameters where outliers take over every layer.
The demo wakes as you arrive…

Why do these outliers matter so much? The paper measured what happens when you set outlier dimensions to zero before computing attention. The result is devastating: the top-1 probability drops by more than 20%, and validation increases by 600–1000%. Yet these outliers represent only about 0.1% of all features.

By contrast, removing the same number of random (non-outlier) dimensions reduces the top-1 probability by at most 0.3% and increases perplexity by just 0.1%.

These few dimensions carry information that the entire model depends on. They act as a kind of information highway — critical channels through which the model routes its most important signals. Quantizing them carelessly is like cutting the main power cable to a building: a tiny physical change with catastrophic consequences.

Innovation 2: mixed-precision decomposition

Since outliers are systematic and sparse — the same few feature dimensions across all layers — the solution is elegant: separate them out. Before performing the matrix multiplication, identify the outlier columns (those with any value exceeding a threshold α=6.0\alpha = 6.0), extract them into their own sub-matrices, and multiply those in 16-bit precision. Everything else — the 99.9% — stays in 8-bit.

The process works in four steps. First, scan the hidden state to find which feature dimensions contain outliers above the threshold. Second, decompose the input matrix and weight matrix: pull out the outlier columns into separate matrices. Third, perform 16-bit matrix multiplication on the outlier sub-matrices and 8-bit vector-wise quantized multiplication on the rest. Fourth, add the two results together in 16-bit.

This is computationally inexpensive because the outlier dimensions are so few — at most 7 out of thousands. The extra memory from 16-bit storage of these columns is about 0.1%.

Cf16≈∑h∈OXf16hWf16h  +  Sf16⋅∑h∉OXi8hWi8hC_{\mathrm{f16}} \approx \sum_{h \in O} X^h_{\mathrm{f16}} W^h_{\mathrm{f16}} \;+\; S_{\mathrm{f16}} \cdot \sum_{h \notin O} X^h_{\mathrm{i8}} W^h_{\mathrm{i8}}
Mixed-precision decomposition — 16-bit outliers + 8-bit regular values — The set OO contains the outlier feature dimensions — those with values exceeding threshold α\alpha. The first sum handles outliers in full 16-bit precision. The second sum handles everything else in 8-bit with vector-wise quantization (denormalized by SS). The final output is accumulated in 16-bit. Since ∣O∣≤7|O| \leq 7, the overhead is negligible.
Open in Lab
Watch the decomposition in action: outlier columns are extracted for 16-bit multiplication while everything else goes through 8-bit quantization.
The demo wakes as you arrive…

Putting it together: the LLM.int8() method

LLM.int8() is the combination of vector-wise quantization and mixed-precision decomposition. Given 16-bit inputs and weights, the method proceeds as follows:

Step 1: Identify outlier feature dimensions — any column where at least one value has magnitude ≥6.0\geq 6.0.

Step 2: Decompose both the input and weight matrices into outlier and non-outlier sub-matrices along those feature dimensions.

Step 3: Multiply the outlier sub-matrices in 16-bit floating point. Independently, quantize the non-outlier sub-matrices using absmax vector-wise quantization (row-wise for inputs, column-wise for weights), perform Int8 matrix multiplication, and dequantize the Int32 result using the outer product of the scaling constants.

Step 4: Add the 16-bit outlier result and the dequantized 8-bit result to produce the final 16-bit output.

The method requires no retraining, no data, and no changes to the model architecture. A pretrained 16-bit can be loaded and converted on-the-fly.

Using LLM.int8() with Hugging Face Transformerspython

Simplified to show the idea — not the real implementation.

from transformers import AutoModelForCausalLM, AutoTokenizer
# Load a 175B model in 8-bit — fits on consumer GPUs model = AutoModelForCausalLM.from_pretrained(
    "facebook/opt-175b",
    load_in_8bit=True,           # <-- enables LLM.int8()
    device_map="auto"            # distributes across available GPUs
) tokenizer = AutoTokenizer.from_pretrained("facebook/opt-175b")
# Use the model — no performance degradation inputs = tokenizer("The future of AI is", return_tensors="pt") output = model.generate(**inputs, max_new_tokens=50) print(tokenizer.decode(output[0]))

Results: zero degradation up to 175B parameters

The paper evaluates quantization methods across model sizes from 125M to 175B parameters using two metrics: C4 perplexity (language modeling quality) and zeroshot accuracy on tasks.

The results are striking. Standard absmax quantization begins degrading at 2.7B parameters and collapses by 13B — an 8-bit 13B model performs worse than an 8-bit 6.7B model, inverting the expected scaling trend. Zeropoint quantization holds longer but also fails at 13B. Vector-wise quantization alone delays the problem but cannot solve it.

LLM.int8() is the only method that preserves full 16-bit performance at every scale tested. At 175B parameters (OPT-175B), the 8-bit model matches the 16-bit model exactly on all benchmarks, including WinoGrande, HellaSwag, PIQA, and LAMBADA. The memory footprint drops by nearly 2x — from about 350 GB to about 175 GB — making it possible to run the model on hardware that previously could not support it.

Open in Lab
Track quantization performance as models scale. Only LLM.int8() maintains full 16-bit accuracy at every scale tested.
The demo wakes as you arrive…

Why this matters: democratizing large models

Before LLM.int8(), running a 175B parameter model required 8 enterprise A100 GPUs with 80 GB each — hardware costing over $100,000. After LLM.int8(), the same model runs on 8 consumer RTX 3090 GPUs with 24 GB each. A model previously restricted to major corporations became accessible to academic labs.

The method was integrated into Hugging Face Transformers through the bitsandbytes library, reaching millions of practitioners with a single flag: load_in_8bit=True. This was not just a memory optimization — it was an accessibility revolution.

Perhaps more importantly, the outlier analysis changed how the field thinks about quantization. The discovery that emergent features concentrate in a handful of dimensions and are critical for model performance informed every major quantization method that followed, including GPTQ (4-bit post-training quantization) and QLoRA (quantized ), both of which were co-authored by Tim Dettmers.

Timeline: the quantization revolution

  1. 2015

    Deep Compression (Han et al.)

    Introduced pruning, quantization, and Huffman coding to compress neural networks by 35–49x, establishing model compression as a research field.

  2. 2020

    GPT-3 and the memory wall

    GPT-3 with 175B parameters demonstrated remarkable few-shot learning but required hundreds of GB of GPU memory, making it inaccessible to most researchers.

  3. 2022

    LLM.int8() (this paper)

    First degradation-free 8-bit quantization for transformers up to 175B parameters. Discovered emergent outlier features and solved them with mixed-precision decomposition.

  4. 2022

    GPTQ (Frantar et al.)

    Pushed quantization to 4 bits using second-order information to minimize layer-wise error. Made 175B models fit in ~45 GB.

  5. 2023

    QLoRA (Dettmers et al.)

    Combined 4-bit quantization with Low-Rank Adaptation, enabling fine-tuning of 65B parameter models on a single 48 GB GPU. Built directly on LLM.int8() insights.

  6. 2024

    Quantization becomes the default

    4-bit and 8-bit quantization became standard practice for deploying large models. GGUF, AWQ, and bitsandbytes made quantized inference accessible to everyone.

CitationDettmers, Lewis, Belkada, Zettlemoyer. LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale. NeurIPS, 2022.

Terms in this paper