Model Efficiency & Scaling2022advanced12 min read
Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer
برامج المُوتِّرات V: ضبط الشبكات العصبية الكبيرة عبر النقل الفوري للمعاملات الفائقة
Yang, G. · Hu, E. J. · Babuschkin, I. · Sidor, S. · Liu, X. · Farhi, D. · Ryder, N. · Pachocki, J. · Chen, W. · Gao, J. — NeurIPS
The problem
large neural networks requires choosing hyperparameters — , initialization scale, optimizer settings — that critically affect performance. For billion-parameter models, each tuning experiment costs thousands of GPU-hours. Practitioners often guess or copy settings from smaller models, but under standard parameterization (SP), optimal hyperparameters shift unpredictably as width changes. There was no principled way to tune cheaply on a small model and transfer those settings to a large one.
The contribution
The paper introduces μTransfer — a tuning paradigm built on Maximal Update Parameterization (μP). In μP, initialization variances, learning rates, and output multipliers are scaled with network width so that optimal hyperparameters remain stable across model sizes. This means you can tune on a small (e.g. 13M parameters) and transfer the optimal learning rate, , and initialization directly to a large target model (e.g. 350M or 6.7B parameters), achieving state-of-the-art performance without any tuning on the large model. The authors verify this on both and ResNet architectures.
The impact
μP and μTransfer have become foundational tools for training large language models efficiently. Cerebras-GPT, and multiple industry labs adopted μP to reduce tuning budgets by orders of magnitude. The framework provided theoretical grounding for the empirical observation that some hyperparameters transfer across scales, and formalized the distinction between the (lazy) regime and the feature-learning regime. It influenced how the community thinks about scaling, initialization, and optimization at scale.
Imagine a pilot learning to fly in a small Cessna before stepping into a Boeing 747. In normal flight schools, nothing transfers: the throttle sensitivity, the turn radius, the landing speed — all change drastically with the aircraft's size. But what if someone redesigned the cockpit controls so that "push the throttle halfway" always means "cruise at the optimal speed," regardless of whether the plane weighs 1 ton or 400 tons?
That is exactly what μP does for neural networks. It redesigns how weights, learning rates, and initialization scale with model size — so that the "flight manual" (hyperparameters) you write for a tiny model works perfectly for a massive one. The paper calls this zero-shot hyperparameter transfer: tune once on a small model, deploy on the giant, never touch the knobs again.
The core problem: hyperparameter tuning at scale is unaffordable
Every needs hyperparameters: the learning rate controls how big each update step is, the initialization scale sets the starting point for weights, and optimizer settings like momentum determine how past gradients influence the current step. Choosing these well is the difference between a model that converges to strong performance and one that diverges or stalls.
For small models, you can afford to try hundreds of combinations — a or random search over learning rates and initialization scales. But for a model with billions of parameters, a single training run might cost $100,000 in compute. Running 200 experiments is simply not feasible.
Practitioners have long observed that hyperparameters tuned on a small model sometimes work on a larger model — but sometimes they fail catastrophically. Under standard parameterization, the optimal learning rate can shift by orders of magnitude when you double the width. There is no guarantee, and no theory explaining when transfer works and when it doesn't.
Standard vs. Maximal Update Parameterization
To understand μP, we first need to see what goes wrong with standard parameterization (SP) — the default in frameworks like PyTorch. In SP (including Xavier and He initialization), each weight in a is drawn from a distribution with variance proportional to . The learning rate is the same for all parameters.
This setup works fine at a fixed width. But as you increase the hidden dimension , trouble appears. The activations grow with width, the gradients shift in scale, and the learning rate that was perfect for width 256 becomes disastrous at width 4096. Worse, to keep the network stable at larger widths, you must shrink the learning rate — which pushes the network into the regime (also called the NTK regime), where hidden features barely change from their initialization. The network functions like a fixed kernel machine rather than learning rich features.
μP solves this by prescribing width-dependent scaling rules. The key changes for a layer with fan-in (compared to a base width ) are:
Hidden-layer initialization: variance scales as (same as SP).
Hidden-layer learning rate: scales as , i.e., the learning rate shrinks with width — but the per-coordinate update to activations stays .
Output layer multiplier: the output logits are multiplied by , preventing them from growing with width.
scaling (for Transformers): the temperature scales as instead of .
The result is that every activation, every , and every parameter update has a width-independent coordinate size. The network behaves the same dynamically at width 64 as at width 8192.
The three rules that define μP
μP is not an arbitrary recipe — it emerges from three precise mathematical requirements. Think of them as engineering constraints: the network must satisfy all three to guarantee that hyperparameters transfer across widths.
Rule 1 — Stable initialization: The pre-activations at every layer must have -sized coordinates at initialization. This means that if you look at any single coordinate of the hidden state, its typical magnitude does not depend on width .
Rule 2 — Stable output: The network output must remain throughout training. If the output logits grow with width, the loss landscape changes shape and the optimal learning rate shifts.
Rule 3 — Maximal feature updates: The change in hidden features at each training step must also be . This is the maximal part — updates are as large as possible without causing instability. If updates shrink to zero (as in the NTK regime), the network stops learning features.
The abc-parameterization framework
The mathematical machinery behind μP is the . For each in a layer with fan-in , three exponents control its scaling:
The layer's output is multiplied by (the output multiplier). The initialization variance is (controlling the starting scale). The learning rate is (controlling how fast the layer updates).
Different choices of give different parameterizations. Standard parameterization corresponds to one set of values; NTK parameterization to another. μP corresponds to the unique set of values that satisfy all three desiderata above.
Two regimes: kernel (lazy) vs. feature learning
This is perhaps the deepest conceptual insight in the paper. Neural networks can operate in two fundamentally different regimes as width grows:
In the kernel (NTK / lazy) regime, the hidden features barely change from their random initialization during training. The network effectively becomes a linear model around its initialization — its predictions are governed by a fixed kernel (the ). This is mathematically elegant but fundamentally limits what the network can learn, because the internal representations stay frozen.
In the feature learning regime, hidden features do evolve during training. The network discovers meaningful patterns in the data and builds increasingly useful internal representations. This is what makes deep learning powerful — and it is what μP preserves.
The choice of parameterization determines which regime the network lands in. Standard parameterization pushes wide networks toward the kernel regime. μP is the unique parameterization that keeps networks in the feature-learning regime at any width.
μTransfer: the practical recipe
The paper turns the theory into a simple three-step recipe called μTransfer:
Step 1 — Parametrize in μP: Set up the target model (the large one you actually want to train) using μP scaling rules. This means adjusting initialization variances, per-layer learning rates, and output multipliers according to the abc-exponents.
Step 2 — Tune on a proxy: Build a small "proxy" model — same architecture and depth, but much narrower (e.g. 13M parameters instead of 350M). Run a large hyperparameter search on this proxy. Because of μP, the optimal learning rate, momentum, and other non- hyperparameters are the same at any width.
Step 3 — Transfer: Take the best hyperparameters from the proxy and plug them directly into the large model. Train the large model once — no further tuning needed.
The authors demonstrate this concretely: a 200-sample random search on a 13M-parameter proxy yields hyperparameters that, when transferred to BERT-large (350M), outperform published BERT-large results. The total tuning cost equals one BERT-large pretraining run. Similarly, tuning on a 40M proxy produces results exceeding published GPT-3 6.7B numbers, at only 7% of pretraining cost.
Verifying μP: the coordinate check
How do you know your μP implementation is correct? The paper introduces a simple diagnostic called the coordinate check. The idea: train the model for a few steps at several different widths and plot the average coordinate size of activations (or logits) over training steps. If μP is implemented correctly, the curves for different widths should overlap — the coordinate size is independent of width. If they fan out (grow with width), something is wrong.
This test is cheap to run and catches implementation bugs immediately. It has become the standard verification tool for μP implementations. The open-source mup library (installable via pip install mup) automates both the parameterization and the coordinate check.
Simplified to show the idea — not the real implementation.
from mup import set_base_shapes, MuAdam, MuReadout
# 1. Define base (proxy) model and target model
base_model = MyTransformer(width=256)
target_model = MyTransformer(width=4096)
# 2. Register base shapes — mup calculates width ratios
set_base_shapes(target_model, base_model)
# 3. Use MuAdam — automatically scales per-layer learning rates
optimizer = MuAdam(target_model.parameters(), lr=0.001)
# 4. Coordinate check: train a few steps, verify activations
# are width-independent (curves overlap for different widths)
Key results: BERT, GPT-3, and beyond
The paper validates μTransfer on two major architectures:
BERT-large (350M): Hyperparameters tuned on a 13M-parameter proxy are transferred to BERT-large. The resulting model outperforms the published BERT-large baseline, and the total cost of tuning plus one training run equals a single BERT-large pretraining.
GPT-3 6.7B: Hyperparameters tuned on a 40M proxy transfer directly to the 6.7B model. The result exceeds published GPT-3 6.7B performance, with tuning cost only 7% of pretraining. This is the first demonstration that hyperparameter transfer works across a 150× size gap.
The authors also verify that μTransfer works for ResNets on image classification, showing the approach is not specific to Transformers. Critically, they show that regularization-related hyperparameters (like and ) do not transfer — only optimization hyperparameters (learning rate, momentum, initialization) do.
Impact and adoption
μP and μTransfer changed the economics of training large models. Before this work, hyperparameter tuning at the billion-parameter scale was essentially guesswork combined with expensive trial-and-error. After it, teams could allocate a fixed compute budget to tuning a tiny proxy, then apply those settings with mathematical confidence.
Cerebras adopted μP for their GPT family, demonstrating that a 200-sample random search on a 40M proxy could set the learning rate for models up to 13B parameters. Multiple research labs have since integrated the mup library into their training pipelines. The framework also deepened the community's understanding of the relationship between parameterization, feature learning, and scaling laws — connecting practical recipes to deep theoretical insights from the Tensor Programs series.
The Tensor Programs journey
2019
Tensor Programs I: GP behavior
Yang showed that wide networks of any architecture converge to Gaussian Processes, laying the mathematical foundation.
2020
Tensor Programs II: NTK universality
Extended the framework to the Neural Tangent Kernel for arbitrary architectures, unifying the kernel view of deep learning.
2021
Tensor Programs IV: Feature learning in infinite width
Discovered the Maximal Update Parameterization — the unique scaling that enables feature learning as width → ∞. This is the theoretical foundation of μP.
2022
Tensor Programs V: μTransfer (this paper)
Turned theory into practice: showed that μP enables zero-shot hyperparameter transfer, verified on BERT-large and GPT-3 6.7B, and released the mup library.
2023
Industry adoption: Cerebras-GPT and beyond
Cerebras released GPT models trained with μP, validating μTransfer at production scale. Other labs followed with their own implementations.
CitationYang, Hu, Babuschkin, Sidor, Liu, Farhi, Ryder, Pachocki, Chen, Gao. Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer. NeurIPS, 2022.
Terms in this paper
- Hyperparameterالمعلمة الفائقة
- Learning Rateمعدل التعلم
- Feature Learningتعلُّم السمات
- Infinite-Width Limitنهاية العرض اللانهائي
- Neural Tangent Kernelنواة المماس العصبي
- Scaling Lawقانون التحجيم
- Weight initializationتهيئة الأوزان
- Adamخوارزمية آدام
- Convergenceالتقارب الحسابي
- Batch Normalizationتسوية الدفعات الحسابية
- Over-Parameterizationالإفراط في المعايرة
- Lazy Trainingالتدريب الكسول
- Activation Functionدالة التنشيط
- Gradientالتدرج التفاضلي
- Loss functionدالة الخسارة