Model Compression2022intermediate11 min read
LLM.int8(): 8-Bit Matrix Multiplication for Transformers at Scale
LLM.int8(): ضرب المصفوفات بدقة 8 بت للمحوِّلات على نطاق واسع
Dettmers, T. · Lewis, M. · Belkada, Y. · Zettlemoyer, L. — NeurIPS
The problem
Large language models like GPT-3 (175B parameters) require enormous memory just to load their weights for — often hundreds of gigabytes. An obvious solution is to quantize weights from 16-bit to 8-bit integers, halving memory. But prior 8-bit methods had only been tested on models up to 350M parameters. At larger scales, they cause severe performance degradation, and nobody understood why.
The contribution
A two-part quantization method called LLM.int8() that enables degradation-free 8-bit inference for transformers up to 175B parameters. First, vector-wise quantization uses per-row and per-column scaling constants instead of one per , improving precision. Second, mixed-precision decomposition isolates rare but critical feature dimensions (~0.1%) into 16-bit multiplication while keeping the remaining 99.9% in 8-bit. The paper also provides the first systematic analysis of emergent outlier features in large transformers: they appear in only ~6 feature dimensions yet dominate model performance, and their sudden emergence at 6.7B parameters explains why prior quantization methods failed at scale.
The impact
LLM.int8() made large model inference accessible on consumer hardware for the first time — OPT-175B and BLOOM could run on a single server with consumer GPUs. It was integrated into Hugging Face Transformers via the bitsandbytes library, becoming the default way to load large models in reduced precision. The outlier analysis directly informed GPTQ and QLoRA, two of the most influential quantization methods that followed. The paper established that understanding emergent features is essential for any compression technique applied to large models.
Imagine a warehouse of identical-looking filing cabinets — millions of them, each holding a single number written to 16 decimal places. You need to move the entire warehouse to a building with half the floor space. Most numbers are small and well-behaved: rounding them to fewer digits loses nothing important.
But scattered across a handful of specific cabinets — the same six cabinet positions on every floor — are giant numbers, 20 times larger than anything else. If you round those, the entire filing system collapses. The information they carry is disproportionately critical.
LLM.int8() is the moving strategy: round 99.9% of the cabinets to save space, but hand-carry those six critical positions at full precision. Nothing is lost, and the warehouse fits.
The memory wall: why large models need compression
A 175-billion model stored in 16-bit floating point requires about 350 GB of GPU memory — just to hold the weights, before any activations or batch data. That exceeds the capacity of even the most expensive enterprise GPUs. In 2022, running inference on such a model required a multi-GPU cluster costing tens of thousands of dollars.
The feed-forward and projection layers account for 95% of all parameters and 65–85% of all computation. If we could quantize these layers from 16-bit to 8-bit integers, we would halve the memory footprint: 175 GB instead of 350 GB. This puts models like OPT-175B and BLOOM within reach of a single server with consumer GPUs.
But quantization is not free. Reducing precision means representing continuous values with fewer discrete bins. If the distribution of values is well-behaved, the rounding error is small. If it is not — if there are extreme outliers — quantization can destroy the information the model needs to function.
How quantization works: from 16 bits to 8
Quantization maps high-precision floating-point numbers to low-precision integers. The simplest method is : divide every value in a tensor by the largest absolute value, multiply by 127, and round to the nearest integer. This maps the tensor into the Int8 range .
The idea is like rescaling a thermometer. If temperatures range from to , you rescale so that maps to and maps to . Every degree gets about 4 bins — plenty of resolution. But if one day hits , then maps to and all normal temperatures are squeezed into a tiny range around zero. One outlier destroys the precision for everything else.
A more sophisticated method is , which shifts the distribution so that asymmetric ranges use the full interval. This helps with activations that are all positive (like ReLU outputs), since absmax would waste half its range on negative values that never appear.
However, zeropoint quantization requires special instructions that GPUs do not natively support efficiently, making it slower in practice. The paper focuses on absmax as the practical choice and enhances it with two innovations.
Innovation 1: vector-wise quantization
Standard quantization uses a single for the entire tensor — one number to rule all values. If any single value is an outlier, it sets the scale for millions of other values. This is called tensor-wise quantization.
The key insight is that matrix multiplication can be viewed as a collection of independent inner products. Each inner product involves one row of the hidden states and one column of the . Since they are independent, each can have its own scaling constant.
Think of it like grading exams in different classrooms. Tensor-wise quantization is like grading all students on the same curve — one genius student skews the scale for everyone. Vector-wise quantization is like giving each classroom its own curve. Each inner product gets an independent scale that fits its own range of values, so no single outlier can dominate the entire computation.
The discovery: emergent outlier features at scale
The paper's most important contribution may be its analysis of emergent outlier features. As transformers grow, something dramatic happens to the hidden states during inference: certain feature dimensions develop values with magnitudes 20 times larger than any other dimension.
These outliers are not random noise. They are highly systematic: in a 6.7B parameter processing a 2048- sequence, about 150,000 outlier values appear across all layers — but they concentrate in only 6 feature dimensions out of thousands. The same 6 dimensions, across every , every token, every sequence.
The emergence follows a . Below 6B parameters, outliers appear sporadically in about 25% of layers. Between 6B and 6.7B, a sharp shift occurs: outliers suddenly appear in 100% of layers and 75% of all sequence positions. This is the exact point where standard quantization methods collapse.
Why do these outliers matter so much? The paper measured what happens when you set outlier dimensions to zero before computing attention. The result is devastating: the top-1 probability drops by more than 20%, and validation increases by 600–1000%. Yet these outliers represent only about 0.1% of all features.
By contrast, removing the same number of random (non-outlier) dimensions reduces the top-1 probability by at most 0.3% and increases perplexity by just 0.1%.
These few dimensions carry information that the entire model depends on. They act as a kind of information highway — critical channels through which the model routes its most important signals. Quantizing them carelessly is like cutting the main power cable to a building: a tiny physical change with catastrophic consequences.
Innovation 2: mixed-precision decomposition
Since outliers are systematic and sparse — the same few feature dimensions across all layers — the solution is elegant: separate them out. Before performing the matrix multiplication, identify the outlier columns (those with any value exceeding a threshold ), extract them into their own sub-matrices, and multiply those in 16-bit precision. Everything else — the 99.9% — stays in 8-bit.
The process works in four steps. First, scan the hidden state to find which feature dimensions contain outliers above the threshold. Second, decompose the input matrix and weight matrix: pull out the outlier columns into separate matrices. Third, perform 16-bit matrix multiplication on the outlier sub-matrices and 8-bit vector-wise quantized multiplication on the rest. Fourth, add the two results together in 16-bit.
This is computationally inexpensive because the outlier dimensions are so few — at most 7 out of thousands. The extra memory from 16-bit storage of these columns is about 0.1%.
Putting it together: the LLM.int8() method
LLM.int8() is the combination of vector-wise quantization and mixed-precision decomposition. Given 16-bit inputs and weights, the method proceeds as follows:
Step 1: Identify outlier feature dimensions — any column where at least one value has magnitude .
Step 2: Decompose both the input and weight matrices into outlier and non-outlier sub-matrices along those feature dimensions.
Step 3: Multiply the outlier sub-matrices in 16-bit floating point. Independently, quantize the non-outlier sub-matrices using absmax vector-wise quantization (row-wise for inputs, column-wise for weights), perform Int8 matrix multiplication, and dequantize the Int32 result using the outer product of the scaling constants.
Step 4: Add the 16-bit outlier result and the dequantized 8-bit result to produce the final 16-bit output.
The method requires no retraining, no data, and no changes to the model architecture. A pretrained 16-bit can be loaded and converted on-the-fly.
Simplified to show the idea — not the real implementation.
from transformers import AutoModelForCausalLM, AutoTokenizer
# Load a 175B model in 8-bit — fits on consumer GPUs model = AutoModelForCausalLM.from_pretrained(
"facebook/opt-175b",
load_in_8bit=True, # <-- enables LLM.int8()
device_map="auto" # distributes across available GPUs
) tokenizer = AutoTokenizer.from_pretrained("facebook/opt-175b")
# Use the model — no performance degradation inputs = tokenizer("The future of AI is", return_tensors="pt") output = model.generate(**inputs, max_new_tokens=50) print(tokenizer.decode(output[0]))Results: zero degradation up to 175B parameters
The paper evaluates quantization methods across model sizes from 125M to 175B parameters using two metrics: C4 perplexity (language modeling quality) and zeroshot accuracy on tasks.
The results are striking. Standard absmax quantization begins degrading at 2.7B parameters and collapses by 13B — an 8-bit 13B model performs worse than an 8-bit 6.7B model, inverting the expected scaling trend. Zeropoint quantization holds longer but also fails at 13B. Vector-wise quantization alone delays the problem but cannot solve it.
LLM.int8() is the only method that preserves full 16-bit performance at every scale tested. At 175B parameters (OPT-175B), the 8-bit model matches the 16-bit model exactly on all benchmarks, including WinoGrande, HellaSwag, PIQA, and LAMBADA. The memory footprint drops by nearly 2x — from about 350 GB to about 175 GB — making it possible to run the model on hardware that previously could not support it.
Why this matters: democratizing large models
Before LLM.int8(), running a 175B parameter model required 8 enterprise A100 GPUs with 80 GB each — hardware costing over $100,000. After LLM.int8(), the same model runs on 8 consumer RTX 3090 GPUs with 24 GB each. A model previously restricted to major corporations became accessible to academic labs.
The method was integrated into Hugging Face Transformers through the bitsandbytes library, reaching millions of practitioners with a single flag: load_in_8bit=True. This was not just a memory optimization — it was an accessibility revolution.
Perhaps more importantly, the outlier analysis changed how the field thinks about quantization. The discovery that emergent features concentrate in a handful of dimensions and are critical for model performance informed every major quantization method that followed, including GPTQ (4-bit post-training quantization) and QLoRA (quantized ), both of which were co-authored by Tim Dettmers.
Timeline: the quantization revolution
2015
Deep Compression (Han et al.)
Introduced pruning, quantization, and Huffman coding to compress neural networks by 35–49x, establishing model compression as a research field.
2020
GPT-3 and the memory wall
GPT-3 with 175B parameters demonstrated remarkable few-shot learning but required hundreds of GB of GPU memory, making it inaccessible to most researchers.
2022
LLM.int8() (this paper)
First degradation-free 8-bit quantization for transformers up to 175B parameters. Discovered emergent outlier features and solved them with mixed-precision decomposition.
2022
GPTQ (Frantar et al.)
Pushed quantization to 4 bits using second-order information to minimize layer-wise error. Made 175B models fit in ~45 GB.
2023
QLoRA (Dettmers et al.)
Combined 4-bit quantization with Low-Rank Adaptation, enabling fine-tuning of 65B parameter models on a single 48 GB GPU. Built directly on LLM.int8() insights.
2024
Quantization becomes the default
4-bit and 8-bit quantization became standard practice for deploying large models. GGUF, AWQ, and bitsandbytes made quantized inference accessible to everyone.
CitationDettmers, Lewis, Belkada, Zettlemoyer. LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale. NeurIPS, 2022.
Terms in this paper
- Quantizationالتكميم
- Outlierالقيمة الشاذة
- Mixed Precision Trainingالتدريب بالدقة المختلطة
- Half-Precisionنصف الدقة
- Dequantizationفك التكميم
- Model Compressionضغط النماذج
- Dynamic Rangeالنطاق الديناميكي
- Emergent Abilitiesالقدرات المعرفية الناشئة فجأة
- Inferenceالاستدلال
- GPUوحدة معالجة الرسوميات
- Transformerالمحوِّل
- Embeddingالتضمين
- Activation Functionدالة التنشيط
- Layerالطبقة الحسابية
- Weight Matrixمصفوفة الأوزان