Optimization2023intermediate12 min read

Symbolic Discovery of Optimization Algorithms

الاكتشاف الرمزي لخوارزميات الأمثَلة

Chen, X. · Liang, C. · Huang, D. · Real, E. · Wang, K. · Liu, Y. · Pham, H. · Dong, X. · Luong, T. · Hsieh, C.-J. · Lu, Y. · Le, Q.V. — NeurIPS

The problem

By 2023, and had been the de facto standard optimizers for training deep neural networks for nearly a decade. Despite hundreds of handcrafted alternatives, none had consistently displaced Adam across vision, language, and multimodal tasks. Meanwhile, attempts to automatically discover better optimizers via learning-to-optimize (L2O) or reinforcement learning produced black-box algorithms that failed to generalize from small proxy tasks to real-world training at scale. The community was stuck: human intuition seemed exhausted, and machine search was too brittle.

The contribution

A program-search framework that formulates discovery as symbolic evolution over an infinite program space, using evolutionary search with warm-start, for pruning, and for . The method discovers Lion (EvoLved ), a remarkably simple optimizer that tracks only (no second moment), uses the for uniform-magnitude updates, and achieves 88.3% zero-shot and 91.1% accuracy on ImageNet — surpassing previous best results — while saving up to 5× compute on JFT. Lion is deployed in production at Google.

The impact

Lion demonstrated that machine-discovered optimizers can genuinely surpass a decade of human-designed ones at scale. It proved that adaptive per-parameter scaling (Adam's hallmark) is not necessary — uniform sign updates with proper momentum tracking can match or beat it. Lion's memory savings (one state tensor vs. two) matter at the billion- parameter frontier. The program-search methodology opened a new path: instead of designing optimizers by intuition, search for them symbolically and let simplicity emerge from selection pressure.

Adam is like a seasoned hiker who measures both the slope and the roughness of the terrain at every step, carefully adjusting stride length for each direction. It works, but the hiker needs two notebooks — one for average slope, one for average roughness — and does a lot of arithmetic at each step.

Lion is a different kind of hiker: it throws away the roughness notebook entirely, keeps only one notebook for slope history, and at each step simply asks "uphill or downhill?" then takes a fixed-size step in that direction. Fewer notebooks, simpler decisions — yet this hiker reaches better viewpoints, faster.

Discovering optimizers: algorithms as programs

The key insight behind this paper is that an optimization algorithm is just a short program: it takes a , some historical state, and a , and outputs an update to the weights. If you can represent optimizers as programs, you can search the space of all possible programs to find ones that train neural networks better than anything humans have designed.

The authors start from AdamW as a warm-start seed and apply : each generation picks a parent program, mutates it (insert, delete, or modify a statement), evaluates the child on a small proxy task, and keeps the best. The search space includes 45 math functions — from basic arithmetic to trigonometric and interpolation operations — with no limit on program length or local variables.

This creates an infinite and sparse search space: over 2 million random programs were tested and none matched AdamW. Efficient techniques make the search tractable: abstract execution prunes invalid, duplicate, and redundant programs (giving ~10× speedup via caching and ~3× shorter programs); warm-start and restart balance exploration vs. exploitation; and the total cost is ~3,000 TPU V2 days across multiple search runs.

Open in Lab
Follow the program search pipeline: AdamW seeds the population, evolution mutates programs, abstract execution prunes the space, and funnel selection picks generalizable winners.
The demo wakes as you arrive…

Bridging the gap: from proxy to production

A program that wins on a tiny proxy task (20 minutes on one TPU chip) might fail spectacularly on a real target task (days on 512 TPU chips). This generalization gap is the central challenge of automated optimizer discovery, and previous methods like L2O and neural optimizer search failed precisely here.

The authors address this with two strategies. First, funnel selection: promising programs from the search are evaluated on progressively larger tasks — 10×, then 100× the proxy. Only programs that beat the baseline at each scale advance to the next. Second, program simplification: redundant statements are removed (abstract execution identifies that ~70% of statements in evolved programs are dead code), non-essential operations are ablated, and the surviving program is manually cleaned into its simplest mathematical form.

This process also reveals meta-overfitting: search fitness keeps rising but performance on larger tasks declines. Runs that meta-overfit later tend to find algorithms that generalize better — a useful heuristic the authors exploit by running multiple restarts.

Lion: the evolved sign momentum optimizer

After search, funnel selection, and simplification, the winning program reduces to a strikingly simple algorithm: Lion (EvoLved Sign Momentum). Compared to Adam, which tracks both first and second moments of the gradient, Lion tracks only the first moment (momentum) and computes the update by taking the sign of an interpolation between the current gradient and the momentum.

The intuition is: instead of asking "how steep is this slope and how noisy has it been?" (Adam), Lion asks "which direction does the combined evidence point?" and takes a uniform step. Every parameter gets updated by exactly +1 or −1 (times the learning rate), regardless of gradient magnitude. This creates uniform update magnitudes across all dimensions — a property no handcrafted optimizer had exploited successfully at scale.

Two β values control the behavior: β₁ = 0.9 governs the interpolation for computing the update (more weight on the current gradient), while β₂ = 0.99 governs the for tracking momentum (remembering ~10× longer history). This separation — using different blending factors for the update and the state — is something human designers had not tried, and it emerged naturally from the search.

Open in Lab
Compare how Adam and Lion compute one update step. Notice Lion's uniform ±1 updates vs Adam's scaled updates.
The demo wakes as you arrive…

The Lion update rule

The goal is to update the model weights θ at each training step. Lion does this in three stages: (1) compute an interpolated direction from the current gradient and the stored momentum, (2) apply the sign function to get a uniform ±1 update, (3) then separately update the momentum for the next step with a different blending factor. is applied as decoupled , just like in AdamW.

ct=β1 mt−1+(1−β1) gtc_t = \beta_1 \, m_{t-1} + (1 - \beta_1) \, g_t
Step 1: Interpolate gradient and momentum — Blend the current gradient gₜ with the stored momentum mₜ₋₁ using β₁ = 0.9. This puts 10% weight on the current gradient and 90% on history, creating a combined direction signal.
θt=θt−1−ηt(sign(ct)+λ θt−1)\theta_t = \theta_{t-1} - \eta_t \bigl(\text{sign}(c_t) + \lambda \, \theta_{t-1}\bigr)
Step 2: Update weights with sign and weight decay — Apply sign(cₜ) — each parameter moves by exactly +1 or −1 — then add decoupled weight decay λθ. Multiply the whole update by the learning rate ηₜ. Because sign always produces ±1, Lion needs a smaller learning rate (3–10× smaller than Adam).
mt=β2 mt−1+(1−β2) gtm_t = \beta_2 \, m_{t-1} + (1 - \beta_2) \, g_t
Step 3: Update momentum for the next step — After the weights are updated, refresh the momentum with β₂ = 0.99. This is a slower EMA than β₁, so momentum remembers ~10× longer gradient history. The crucial insight: the blending factor for the update (β₁) differs from the one for state tracking (β₂).

The idea in code

Lion optimizer — the complete algorithm in ~15 linespython

Simplified to show the idea — not the real implementation.

import numpy as np

def lion_step(weight, gradient, momentum, lr, beta1=0.9, beta2=0.99, wd=0.1):
    """One step of the Lion optimizer."""
    # Step 1: Interpolate gradient and momentum for the update direction
    c = beta1 * momentum + (1 - beta1) * gradient

    # Step 2: Sign of c → uniform ±1 update + decoupled weight decay
    update = np.sign(c) + wd * weight

    # Apply the update
    weight = weight - lr * update

    # Step 3: Update momentum for the NEXT step (separate β₂)
    momentum = beta2 * momentum + (1 - beta2) * gradient

    return weight, momentum

# Key differences from Adam:
# • No second moment (v) → saves ~50% optimizer memory
# • sign() produces uniform ±1 → use 3-10x smaller learning rate
# • β₁ ≠ β₂ → decouples update direction from state tracking
# • No epsilon, no bias correction → fewer hyperparameters

Why does sign work? Regularization through noise

The sign operation discards gradient magnitude and keeps only direction. This might seem like throwing away useful information, but it introduces a form of implicit regularization: by ignoring how steep a dimension is, the optimizer adds noise to the updates. This noise prevents the model from converging to sharp minima and pushes it toward flatter regions of the — which are known to generalize better.

Empirical evidence confirms this: -B/16 trained with Lion has a higher training error than with AdamW, yet achieves 2% higher validation accuracy. The authors also measure landscape flatness by perturbing converged weights with Gaussian noise. Lion's solution retains much lower error under perturbation, confirming convergence to a flatter region. This behavior resembles (SAM), but Lion achieves it implicitly through the sign operation rather than an explicit min-max objective.

Open in Lab
Compare the loss landscape around the converged solution: Adam tends to find sharper minima, Lion finds flatter ones.
The demo wakes as you arrive…

Memory and efficiency: half the state, faster steps

Adam tracks two state tensors per parameter: the first moment (mean of gradients) and the second moment (mean of squared gradients). Lion tracks only one: the momentum. For a model with P parameters, this saves P floats of optimizer state — roughly a 50% reduction in optimizer memory. When training billion-parameter models, this directly translates to fewer accelerator chips. For example, training ViT-B/16 with 4,096 requires at least 16 TPU V4 chips with AdamW but only 8 with Lion (both using bfloat16 momentum).

Lion is also 2–15% faster in wall-clock time per step due to its simplicity: no square root, no epsilon, no bias correction — just one interpolation, one sign, and one momentum update.

Open in Lab
Drag the model size slider to see how optimizer memory scales: Adam always needs 2× the state of Lion.
The demo wakes as you arrive…

Practical tuning: learning rate, weight decay, batch size

Because sign(·) always produces ±1, Lion's updates have a larger norm than Adam's scaled updates. This means Lion needs a smaller learning rate — typically 3–10× smaller than what you would use for Adam. To maintain the same effective weight decay strength (which is lr × λ), the weight decay coefficient λ must be increased by the same factor.

For example, if you use lr = 1e−3 and λ = 1.0 for AdamW, a good starting point for Lion is lr = 1e−4 and λ = 10.0. The initial, peak, and end values of the learning rate schedule should all be scaled by the same ratio.

Lion also prefers larger batch sizes. Its performance advantage over AdamW grows with batch size: at batch size 32K, Lion achieves a 2.5% accuracy gain. The authors hypothesize that larger batches give a more reliable gradient direction, which matters more when the optimizer only uses the sign of that direction. Even at small batch sizes (64), Lion remains competitive — it just shines brighter at scale.

Lion is also more robust to choices: heatmaps of accuracy across learning rate and weight decay values show a broader plateau for Lion than for AdamW.

Results across vision, language, and generation

Lion was evaluated across an unusually broad range of architectures (, MLP, ResNet, U-Net, Hybrid) and tasks (image classification, , diffusion models, language modeling, and fine-tuning). Key highlights:

Image classification: On ImageNet trained from scratch, Lion boosted ViT-B/16 accuracy by 1.96% over AdamW (77.44% vs 75.48%). On JFT pre-training, Lion enabled ViT-L/16 to match ViT-H/14 trained by AdamW while being 2× smaller — or equivalently, saved up to 5× compute to reach the same performance.

Vision-language contrastive learning: Using BASIC-L, Lion achieved 88.3% zero-shot ImageNet accuracy — a 2.6% gain over the Adafactor baseline and 2% above the prior state-of-the-art. Fine-tuning accuracy reached 91.1%.

Diffusion models: On 256×256 ImageNet generation, Lion matched AdamW's final FID in 2.3× fewer iterations (FID 4.1 vs 4.7).

Language modeling: On PG-19, Lion achieved up to 2× speedup over AdamW. On 7.5B parameter models trained on 300B tokens, Lion matched training perplexity but improved in-context learning scores across NLG and NLU benchmarks.

Fine-tuning: On GLUE with T5 models (Base to 11B), Lion won 10–12 out of 12 scores at every scale.

Open in Lab
Compare Lion vs AdamW across five task categories. Each axis shows Lion's relative improvement.
The demo wakes as you arrive…

Limitations and when Lion may not help

The authors are transparent about where Lion does not shine. On ResNets, the difference from AdamW and SGD is small — convolutional networks may be easier to optimize than Transformers. With strong (RandAug + Mixup), the regularization benefit of sign shrinks because the augmentation already regularizes. On very large, high-quality datasets (the Imagen base model, large-scale autoregressive pretraining), Lion matches AdamW in perplexity but does not clearly surpass it — the data quality itself reduces the gap between optimizers.

Lion also prefers batch sizes ≥ 64. With very small batches, the gradient direction becomes too noisy, and the sign operation amplifies that noise rather than filtering it. Finally, while Lion saves one state tensor, it still requires momentum storage in bfloat16, which can be expensive at the multi-billion parameter scale. Factorizing the momentum (as Adafactor does for second moments) is suggested as future work.

Lion in context: the optimizer timeline

  1. 2014

    Adam

    Adaptive moment estimation. Tracks first and second moments of gradients. Became the default optimizer for deep learning for nearly a decade.

  2. 2017

    Decoupled Weight Decay (AdamW)

    Loshchilov and Hutter showed that proper weight decay in Adam should be decoupled from the gradient update. AdamW became the new standard.

  3. 2020

    Sharpness-Aware Minimization (SAM)

    Showed that seeking flat minima explicitly improves generalization. Lion achieves a similar effect implicitly through the sign operation.

  4. 2020

    AutoML-Zero

    Ambitious attempt to evolve full ML pipelines from scratch, but only on toy tasks. Lion's search framework builds on this foundation but targets real-world scale.

  5. 2023

    Lion (this paper)

    First machine-discovered optimizer to consistently outperform Adam at state-of-the-art scale. Deployed in Google production systems.

CitationChen, Liang, Huang, Real, Wang, Liu, Pham, Dong, Luong, Hsieh, Lu, Le. Symbolic Discovery of Optimization Algorithms. NeurIPS, 2023.

Terms in this paper