Systems & Optimization2019advanced12 min read
ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
ZeRO: تحسينات الذاكرة نحو تدريب نماذج بتريليون معامل
Rajbhandari, S. · Rasley, J. · Ruwase, O. · He, Y. — SC
The problem
By 2019, large models with billions of parameters offered significant accuracy gains, but them was blocked by limits. replicates the entire model on every GPU — including , gradients, and parameters — creating massive . A 1.5B parameter model with optimizer needs at least 24 GB of memory for model states alone, far beyond what a single GPU can spare after accounting for activations. can split the model, but it requires extensive code refactoring, suffers from low compute efficiency due to sequential dependencies, and scales poorly. There was no solution that combined the memory efficiency of model parallelism with the compute efficiency and ease of use of data parallelism.
The contribution
ZeRO (Zero Redundancy Optimizer) eliminates memory redundancy in data-parallel training through three progressive stages. Stage 1 partitions optimizer states across GPUs (4× memory reduction). Stage 2 also partitions gradients (8× reduction). Stage 3 partitions parameters too (linear reduction with number of GPUs). Additionally, ZeRO-R optimizes residual memory: with CPU offloading, constant-size buffers, and a memory defragmentation system. The result trains models up to 100B parameters on 400 GPUs with , achieving 15 petaflops — an 8× increase in trainable model size and 10× improvement in performance over prior state of the art.
The impact
ZeRO became the backbone of DeepSpeed and transformed how the industry trains large language models. Its staged memory partitioning strategy was adopted in PyTorch FSDP and is now standard practice for training models from tens of billions to hundreds of billions of parameters. ZeRO enabled the training of Turing-NLG (17B) at the time of publication and directly influenced the scaling trajectory that led to GPT-3, Megatron-Turing, and beyond. The ZeRO family grew to include ZeRO-Offload, ZeRO-Infinity, and ZeRO++, extending its reach from GPU clusters down to single-GPU setups.
Imagine a team of chefs preparing the same recipe in a large kitchen. Normally, every chef keeps their own copy of the recipe book, their own full spice rack, and their own complete set of measuring cups — even though they all follow the same recipe. The kitchen overflows with duplicate supplies while the actual cooking space shrinks.
ZeRO is the head chef who says: "Tear the recipe book into sections — each chef keeps only their pages. Split the spice rack — each chef holds only their spices. When you need a spice you don't have, call out and the chef who has it passes it over." Now the same kitchen can handle a recipe ten times more complex, because all the space freed from duplicates is available for actual cooking.
The memory wall: why GPUs run out of space
To understand ZeRO, we first need to understand what actually fills GPU memory during training. When training with mixed precision and the Adam optimizer, the memory consumed by a model with Ψ parameters breaks down into two categories: model states and residual states.
Model states include three components. First, the parameters themselves — 2Ψ bytes. Second, the fp16 gradients — another 2Ψ bytes. Third, the optimizer states: Adam maintains an fp32 copy of the parameters (4Ψ bytes), plus fp32 momentum (4Ψ bytes) and fp32 variance (4Ψ bytes) — totaling 12Ψ bytes. Combined, model states alone consume 16Ψ bytes. For a 1.5B parameter model like GPT-2, that is 24 GB — and we have not even counted activations yet.
Residual states include activations (which can consume 60 GB for GPT-2 at 32), temporary buffers (around 6 GB for large models), and unusable fragmented memory. These compete with model states for the same GPU memory.
The parallelism dilemma: memory vs efficiency
Before ZeRO, practitioners faced a painful trade-off between two parallelism strategies.
Data parallelism replicates the full model on every GPU. Each GPU processes a different data batch, computes gradients, and synchronizes them via . This is efficient — every GPU does useful compute simultaneously — and easy to implement. But it wastes memory: every GPU holds a full copy of all model states (16Ψ bytes), even though they are all identical.
Model parallelism splits the model across GPUs, so each device holds only a fraction of the parameters. This solves the memory problem but creates new ones: splitting requires manual code refactoring per model, dependencies between layers create sequential bottlenecks that leave GPUs idle, and communication of activations between pipeline stages adds . In practice, model parallelism rarely achieves more than 50% compute efficiency.
ZeRO asks: what if we could partition the model states like model parallelism does, but keep the execution pattern of data parallelism?
ZeRO-DP: three stages of memory freedom
ZeRO's core idea is elegantly simple: in standard data parallelism, every GPU maintains an identical copy of all model states. This is pure redundancy. ZeRO partitions these states across GPUs instead of replicating them, then uses collective communication to gather pieces when needed.
The insight is that not every GPU needs every piece of state at every moment. Optimizer states are only needed during the parameter update step. Gradients are only needed after the . And parameters, while needed during forward and backward passes, can be temporarily gathered from other GPUs.
ZeRO-DP introduces this partitioning in three cumulative stages, each one partitioning an additional type of model state. The stages are designed so that earlier stages give the biggest memory savings with the least communication overhead.
Stage 1 — Optimizer State Partitioning (P_os): Each GPU stores only 1/N of the optimizer states (where N is the number of GPUs). After each backward pass, gradients are reduced as usual with all-reduce. Then each GPU updates only its own partition of parameters using its partition of optimizer states. Finally, an broadcasts the updated parameters to all GPUs. Memory for optimizer states drops from 12Ψ to 12Ψ/N — a 4× savings on 64 GPUs with no additional communication.
Stage 2 — Partitioning (P_os+g): Building on Stage 1, each GPU also keeps only the gradients it needs — those corresponding to its optimizer partition. Instead of all-reduce (which gives every GPU all gradients), ZeRO uses (which gives each GPU only its slice). After the update, all-gather synchronizes parameters as before. Memory drops from 14Ψ to (2 + 14/N)Ψ — an 8× savings on 64 GPUs, still with no additional communication.
Stage 3 — Parameter Partitioning (P_os+g+p): Now each GPU stores only 1/N of the parameters too. During forward and backward passes, each layer's full parameters are reconstructed on the fly via all-gather from the GPUs that own the relevant slices. After use, the gathered parameters are discarded. Memory drops to 16Ψ/N — a linear reduction. This stage adds 1.5× communication overhead (two extra all-gathers per step), but enables truly massive models.
ZeRO-R: taming residual memory
Even after ZeRO-DP frees model state memory, residual memory — activations, temporary buffers, and fragmentation — can still be the bottleneck. ZeRO-R addresses each one.
Activation partitioning and CPU offload: Activation checkpointing already reduces activation memory by recomputing activations during the backward pass instead of storing them all. ZeRO-R goes further. It partitions checkpointed activations across GPUs, and for very large models, offloads them to CPU memory. Since activations are only needed during the backward pass, the offloaded data can be prefetched back to GPU just before it is needed.
Constant-size buffers: Large all-reduce operations are most efficient when fusing many small tensors into one large buffer. But naively allocating a buffer proportional to model size creates its own memory pressure. ZeRO-R uses a fixed-size buffer that is large enough for efficiency but does not grow with model size.
Memory defragmentation: During training, memory is allocated and freed in varying sizes, creating gaps (fragmentation) that cannot be used even though total free memory seems sufficient. ZeRO-R pre-allocates contiguous memory blocks for activations and gradients, avoiding fragmentation and ensuring memory is usable.
Communication analysis: the cost of partitioning
Memory partitioning is only useful if the communication cost remains manageable. ZeRO achieves this through careful choice of collective communication primitives.
In standard data parallelism, after the backward pass, an all-reduce operation synchronizes gradients across all GPUs. The all-reduce sends 2Ψ data elements per GPU — this is the baseline communication volume.
ZeRO Stage 1 replaces the all-reduce with a reduce-scatter (which gives each GPU only its partition of the reduced gradients) followed by an all-gather (which distributes updated parameters). The combined volume equals the all-reduce baseline: zero overhead.
ZeRO Stage 2 uses reduce-scatter for gradients, which is actually more efficient than all-reduce. Combined with the all-gather for updated parameters, total volume remains equal to the baseline.
ZeRO Stage 3 adds two additional all-gather operations — one before the and one before the backward pass — to reconstruct full parameters from their partitions. This increases total communication to approximately 1.5× the baseline. For high- interconnects like NVLink, this overhead is acceptable for the dramatic memory savings it enables.
Scaling to 100 billion and beyond
The ZeRO paper demonstrated remarkable scaling results. Using ZeRO-100B (a combination of ZeRO Stage 1 with model parallelism), the authors trained models with over 100 billion parameters on 400 NVIDIA V100 GPUs, achieving super-linear speedup and 15 petaflops of sustained .
Critically, ZeRO also made large-scale training accessible without model parallelism. Using ZeRO-DP alone (no model parallelism), models up to 13 billion parameters could be trained — an order of magnitude larger than what standard data parallelism supports. This is especially significant because model parallelism requires specialized code for each model architecture, while ZeRO works as a drop-in replacement for standard data parallelism.
The paper's theoretical analysis showed that with 1024 GPUs, ZeRO Stage 3 could support training of models with over 1 trillion parameters — roughly 6× beyond what model parallelism alone could handle on the same hardware.
ZeRO in practice: DeepSpeed configuration
Simplified to show the idea — not the real implementation.
import deepspeed
# DeepSpeed configuration with ZeRO Stage 2
ds_config = {
"train_batch_size": 32,
"fp16": {"enabled": True}, # Mixed precision training
"zero_optimization": {
"stage": 2, # Partition optimizer states + gradients
"allgather_partitions": True, # Gather updated params after optimizer step
"reduce_scatter": True, # Use reduce-scatter instead of all-reduce
"overlap_comm": True, # Overlap communication with computation
"contiguous_gradients": True, # Reduce memory fragmentation
},
}
# Wrap model and optimizer — ZeRO is transparent to the training loop
model_engine, optimizer, _, _ = deepspeed.initialize(
model=model,
model_parameters=model.parameters(),
config=ds_config,
)
# Training loop stays identical to standard PyTorch
for batch in dataloader:
loss = model_engine(batch)
model_engine.backward(loss)
model_engine.step()
The ZeRO family and what it enabled
2019
ZeRO (original paper)
Three-stage optimizer state, gradient, and parameter partitioning. Trained 100B+ parameters on 400 GPUs with super-linear speedup. Published at SC 2020.
2020
DeepSpeed & Turing-NLG
ZeRO integrated into DeepSpeed framework. Used to train Turing-NLG (17.2B parameters), the largest language model at the time, with record-breaking accuracy.
2021
ZeRO-Offload
Extended ZeRO to offload optimizer states and gradients to CPU memory, enabling billion-scale training on a single GPU. Democratized large model training.
2021
ZeRO-Infinity
Offloaded all model states to both CPU and NVMe storage. Enabled training of models with trillions of parameters by leveraging the full memory hierarchy.
2022
PyTorch FSDP
PyTorch adopted ZeRO's core ideas in Fully Sharded Data Parallel (FSDP), making ZeRO-style training a first-class citizen in the PyTorch ecosystem.
2023
ZeRO++
Added quantized communication and hierarchical partitioning to reduce cross-node communication by up to 4×, extending ZeRO's efficiency to low-bandwidth clusters.
ZeRO's legacy extends far beyond any single paper or system. It demonstrated that the memory wall in deep learning training is not fundamental — it is an artifact of redundant data management. By treating model states as shared resources that can be partitioned, gathered, and released on demand, ZeRO turned the memory scaling problem from multiplicative (every GPU holds everything) to additive (each GPU holds a fraction). This insight underpins every modern large-scale training system, from Megatron-LM's to PyTorch FSDP, and remains the foundation upon which trillion-parameter training is built.
CitationRajbhandari, Rasley, Ruwase, He. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models. SC (International Conference for High Performance Computing, Networking, Storage and Analysis), 2020.
Terms in this paper
- Data Parallelismتوازي البيانات
- Model Parallelismتوازي النموذج
- Optimizer Statesحالات المُحسِّن
- Gradientالتدرج التفاضلي
- Mixed Precision Trainingالتدريب بالدقة المختلطة
- All-Reduceالاختزال الجماعي
- Redundancyالتكرار اللغوي
- Throughputمعدل التدفق والإنتاجية
- Memoryالذاكرة الحوسبية
- Shardingالتجزئة
- Activation Checkpointingنقاط فحص التنشيطات
- FP16الدقة العائمة بنظام 16 بت
- Gradient Accumulationتجميع التدرجات الحسابية
- GPUوحدة معالجة الرسوميات
- Memory Bandwidthنطاق الذاكرة