Model Efficiency & Scaling2022advanced12 min read

Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

برامج المُوتِّرات V: ضبط الشبكات العصبية الكبيرة عبر النقل الفوري للمعاملات الفائقة

Yang, G. · Hu, E. J. · Babuschkin, I. · Sidor, S. · Liu, X. · Farhi, D. · Ryder, N. · Pachocki, J. · Chen, W. · Gao, J. — NeurIPS

The problem

large neural networks requires choosing hyperparameters — , initialization scale, optimizer settings — that critically affect performance. For billion-parameter models, each tuning experiment costs thousands of GPU-hours. Practitioners often guess or copy settings from smaller models, but under standard parameterization (SP), optimal hyperparameters shift unpredictably as width changes. There was no principled way to tune cheaply on a small model and transfer those settings to a large one.

The contribution

The paper introduces μTransfer — a tuning paradigm built on Maximal Update Parameterization (μP). In μP, initialization variances, learning rates, and output multipliers are scaled with network width so that optimal hyperparameters remain stable across model sizes. This means you can tune on a small (e.g. 13M parameters) and transfer the optimal learning rate, , and initialization directly to a large target model (e.g. 350M or 6.7B parameters), achieving state-of-the-art performance without any tuning on the large model. The authors verify this on both and ResNet architectures.

The impact

μP and μTransfer have become foundational tools for training large language models efficiently. Cerebras-GPT, and multiple industry labs adopted μP to reduce tuning budgets by orders of magnitude. The framework provided theoretical grounding for the empirical observation that some hyperparameters transfer across scales, and formalized the distinction between the (lazy) regime and the feature-learning regime. It influenced how the community thinks about scaling, initialization, and optimization at scale.

Imagine a pilot learning to fly in a small Cessna before stepping into a Boeing 747. In normal flight schools, nothing transfers: the throttle sensitivity, the turn radius, the landing speed — all change drastically with the aircraft's size. But what if someone redesigned the cockpit controls so that "push the throttle halfway" always means "cruise at the optimal speed," regardless of whether the plane weighs 1 ton or 400 tons?

That is exactly what μP does for neural networks. It redesigns how weights, learning rates, and initialization scale with model size — so that the "flight manual" (hyperparameters) you write for a tiny model works perfectly for a massive one. The paper calls this zero-shot hyperparameter transfer: tune once on a small model, deploy on the giant, never touch the knobs again.

The core problem: hyperparameter tuning at scale is unaffordable

Every needs hyperparameters: the learning rate controls how big each update step is, the initialization scale sets the starting point for weights, and optimizer settings like momentum determine how past gradients influence the current step. Choosing these well is the difference between a model that converges to strong performance and one that diverges or stalls.

For small models, you can afford to try hundreds of combinations — a or random search over learning rates and initialization scales. But for a model with billions of parameters, a single training run might cost $100,000 in compute. Running 200 experiments is simply not feasible.

Practitioners have long observed that hyperparameters tuned on a small model sometimes work on a larger model — but sometimes they fail catastrophically. Under standard parameterization, the optimal learning rate can shift by orders of magnitude when you double the width. There is no guarantee, and no theory explaining when transfer works and when it doesn't.

Open in Lab
Compare the cost of hyperparameter tuning: brute-force search on the large model vs. μTransfer from a small proxy. Drag the model size slider to see how costs diverge.
The demo wakes as you arrive…

Standard vs. Maximal Update Parameterization

To understand μP, we first need to see what goes wrong with standard parameterization (SP) — the default in frameworks like PyTorch. In SP (including Xavier and He initialization), each weight in a is drawn from a distribution with variance proportional to 1/fan_in1/\text{fan\_in}. The learning rate η\eta is the same for all parameters.

This setup works fine at a fixed width. But as you increase the hidden dimension nn, trouble appears. The activations grow with width, the gradients shift in scale, and the learning rate that was perfect for width 256 becomes disastrous at width 4096. Worse, to keep the network stable at larger widths, you must shrink the learning rate — which pushes the network into the regime (also called the NTK regime), where hidden features barely change from their initialization. The network functions like a fixed kernel machine rather than learning rich features.

μP solves this by prescribing width-dependent scaling rules. The key changes for a layer with fan-in nn (compared to a base width n0n_0) are:

Hidden-layer initialization: variance scales as 1/n1/n (same as SP).

Hidden-layer learning rate: scales as η0⋅(n0/n)\eta_0 \cdot (n_0 / n), i.e., the learning rate shrinks with width — but the per-coordinate update to activations stays Θ(1)\Theta(1).

Output layer multiplier: the output logits are multiplied by n0/nn_0 / n, preventing them from growing with width.

scaling (for Transformers): the temperature scales as 1/d1/d instead of 1/d1/\sqrt{d}.

The result is that every activation, every , and every parameter update has a width-independent coordinate size. The network behaves the same dynamically at width 64 as at width 8192.

Open in Lab
Watch how activation magnitudes change with width under SP vs. μP. Under SP they grow; under μP they stay constant.
The demo wakes as you arrive…

The three rules that define μP

μP is not an arbitrary recipe — it emerges from three precise mathematical requirements. Think of them as engineering constraints: the network must satisfy all three to guarantee that hyperparameters transfer across widths.

Rule 1 — Stable initialization: The pre-activations at every layer must have Θ(1)\Theta(1)-sized coordinates at initialization. This means that if you look at any single coordinate of the hidden state, its typical magnitude does not depend on width nn.

Rule 2 — Stable output: The network output must remain Θ(1)\Theta(1) throughout training. If the output logits grow with width, the loss landscape changes shape and the optimal learning rate shifts.

Rule 3 — Maximal feature updates: The change in hidden features Δh\Delta h at each training step must also be Θ(1)\Theta(1). This is the maximal part — updates are as large as possible without causing instability. If updates shrink to zero (as in the NTK regime), the network stops learning features.

Coordinate size of hi(ℓ)=Θ(1),Coordinate size of Δhi(ℓ)=Θ(1),Output f(x)=Θ(1)\text{Coordinate size of } h^{(\ell)}_i = \Theta(1), \quad \text{Coordinate size of } \Delta h^{(\ell)}_i = \Theta(1), \quad \text{Output } f(x) = \Theta(1)
The three μP desiderata — all must hold as width n → ∞ — hi(ℓ)h^{(\ell)}_i is a single coordinate of the hidden representation at layer ℓ\ell. The Θ(1)\Theta(1) notation means the quantity stays bounded and non-vanishing as width grows. When all three hold simultaneously, the loss landscape geometry is preserved across scales.

The abc-parameterization framework

The mathematical machinery behind μP is the . For each WW in a layer with fan-in nn, three exponents (a,b,c)(a, b, c) control its scaling:

The layer's output is multiplied by nan^a (the output multiplier). The initialization variance is n−2bn^{-2b} (controlling the starting scale). The learning rate is η⋅n−c\eta \cdot n^{-c} (controlling how fast the layer updates).

Different choices of (a,b,c)(a, b, c) give different parameterizations. Standard parameterization corresponds to one set of values; NTK parameterization to another. μP corresponds to the unique set of (a,b,c)(a, b, c) values that satisfy all three desiderata above.

W(ℓ)=na⋅W~(ℓ),W~ij(ℓ)∼N(0,n−2b),ηℓ=η⋅n−cW^{(\ell)} = n^a \cdot \tilde{W}^{(\ell)}, \quad \tilde{W}^{(\ell)}_{ij} \sim \mathcal{N}(0, n^{-2b}), \quad \eta_\ell = \eta \cdot n^{-c}
The abc-parameterization — the general template for scaling neural network layers — W~\tilde{W} is the raw weight matrix whose entries are drawn from a Gaussian with width-dependent variance. The multiplier nan^a rescales the layer output, and n−cn^{-c} adjusts the per-layer learning rate. Standard parameterization sets a=0,b=1/2,c=0a=0, b=1/2, c=0. μP for hidden layers sets a=0,b=1/2,c=1a=0, b=1/2, c=1 with an output multiplier n−1n^{-1} on the last layer.
Open in Lab
Adjust the abc exponents and watch how activations, gradients, and updates scale with width. Only the μP setting keeps all three stable.
The demo wakes as you arrive…

Two regimes: kernel (lazy) vs. feature learning

This is perhaps the deepest conceptual insight in the paper. Neural networks can operate in two fundamentally different regimes as width grows:

In the kernel (NTK / lazy) regime, the hidden features h(ℓ)h^{(\ell)} barely change from their random initialization during training. The network effectively becomes a linear model around its initialization — its predictions are governed by a fixed kernel (the ). This is mathematically elegant but fundamentally limits what the network can learn, because the internal representations stay frozen.

In the feature learning regime, hidden features do evolve during training. The network discovers meaningful patterns in the data and builds increasingly useful internal representations. This is what makes deep learning powerful — and it is what μP preserves.

The choice of parameterization determines which regime the network lands in. Standard parameterization pushes wide networks toward the kernel regime. μP is the unique parameterization that keeps networks in the feature-learning regime at any width.

Open in Lab
Compare how hidden features evolve during training in the kernel regime vs. the feature-learning regime. In the kernel regime, representations are frozen; in the feature-learning regime, they adapt.
The demo wakes as you arrive…

μTransfer: the practical recipe

The paper turns the theory into a simple three-step recipe called μTransfer:

Step 1 — Parametrize in μP: Set up the target model (the large one you actually want to train) using μP scaling rules. This means adjusting initialization variances, per-layer learning rates, and output multipliers according to the abc-exponents.

Step 2 — Tune on a proxy: Build a small "proxy" model — same architecture and depth, but much narrower (e.g. 13M parameters instead of 350M). Run a large hyperparameter search on this proxy. Because of μP, the optimal learning rate, momentum, and other non- hyperparameters are the same at any width.

Step 3 — Transfer: Take the best hyperparameters from the proxy and plug them directly into the large model. Train the large model once — no further tuning needed.

The authors demonstrate this concretely: a 200-sample random search on a 13M-parameter proxy yields hyperparameters that, when transferred to BERT-large (350M), outperform published BERT-large results. The total tuning cost equals one BERT-large pretraining run. Similarly, tuning on a 40M proxy produces results exceeding published GPT-3 6.7B numbers, at only 7% of pretraining cost.

Open in Lab
The μTransfer workflow: tune hyperparameters on a small proxy model, then transfer them directly to the large target model.
The demo wakes as you arrive…

Verifying μP: the coordinate check

How do you know your μP implementation is correct? The paper introduces a simple diagnostic called the coordinate check. The idea: train the model for a few steps at several different widths and plot the average coordinate size of activations (or logits) over training steps. If μP is implemented correctly, the curves for different widths should overlap — the coordinate size is independent of width. If they fan out (grow with width), something is wrong.

This test is cheap to run and catches implementation bugs immediately. It has become the standard verification tool for μP implementations. The open-source mup library (installable via pip install mup) automates both the parameterization and the coordinate check.

Coordinate check with the mup librarypython

Simplified to show the idea — not the real implementation.

from mup import set_base_shapes, MuAdam, MuReadout

# 1. Define base (proxy) model and target model
base_model = MyTransformer(width=256)
target_model = MyTransformer(width=4096)

# 2. Register base shapes — mup calculates width ratios
set_base_shapes(target_model, base_model)

# 3. Use MuAdam — automatically scales per-layer learning rates
optimizer = MuAdam(target_model.parameters(), lr=0.001)

# 4. Coordinate check: train a few steps, verify activations
#    are width-independent (curves overlap for different widths)
Open in Lab
Simulated coordinate check: under SP, activation magnitude grows with width (curves fan out); under μP, curves overlap perfectly.
The demo wakes as you arrive…

Key results: BERT, GPT-3, and beyond

The paper validates μTransfer on two major architectures:

BERT-large (350M): Hyperparameters tuned on a 13M-parameter proxy are transferred to BERT-large. The resulting model outperforms the published BERT-large baseline, and the total cost of tuning plus one training run equals a single BERT-large pretraining.

GPT-3 6.7B: Hyperparameters tuned on a 40M proxy transfer directly to the 6.7B model. The result exceeds published GPT-3 6.7B performance, with tuning cost only 7% of pretraining. This is the first demonstration that hyperparameter transfer works across a 150× size gap.

The authors also verify that μTransfer works for ResNets on image classification, showing the approach is not specific to Transformers. Critically, they show that regularization-related hyperparameters (like and ) do not transfer — only optimization hyperparameters (learning rate, momentum, initialization) do.

Impact and adoption

μP and μTransfer changed the economics of training large models. Before this work, hyperparameter tuning at the billion-parameter scale was essentially guesswork combined with expensive trial-and-error. After it, teams could allocate a fixed compute budget to tuning a tiny proxy, then apply those settings with mathematical confidence.

Cerebras adopted μP for their GPT family, demonstrating that a 200-sample random search on a 40M proxy could set the learning rate for models up to 13B parameters. Multiple research labs have since integrated the mup library into their training pipelines. The framework also deepened the community's understanding of the relationship between parameterization, feature learning, and scaling laws — connecting practical recipes to deep theoretical insights from the Tensor Programs series.

The Tensor Programs journey

  1. 2019

    Tensor Programs I: GP behavior

    Yang showed that wide networks of any architecture converge to Gaussian Processes, laying the mathematical foundation.

  2. 2020

    Tensor Programs II: NTK universality

    Extended the framework to the Neural Tangent Kernel for arbitrary architectures, unifying the kernel view of deep learning.

  3. 2021

    Tensor Programs IV: Feature learning in infinite width

    Discovered the Maximal Update Parameterization — the unique scaling that enables feature learning as width → ∞. This is the theoretical foundation of μP.

  4. 2022

    Tensor Programs V: μTransfer (this paper)

    Turned theory into practice: showed that μP enables zero-shot hyperparameter transfer, verified on BERT-large and GPT-3 6.7B, and released the mup library.

  5. 2023

    Industry adoption: Cerebras-GPT and beyond

    Cerebras released GPT models trained with μP, validating μTransfer at production scale. Other labs followed with their own implementations.

CitationYang, Hu, Babuschkin, Sidor, Liu, Farhi, Ryder, Pachocki, Chen, Gao. Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer. NeurIPS, 2022.

Terms in this paper