Optimization2023intermediate12 min read
Symbolic Discovery of Optimization Algorithms
الاكتشاف الرمزي لخوارزميات الأمثَلة
Chen, X. · Liang, C. · Huang, D. · Real, E. · Wang, K. · Liu, Y. · Pham, H. · Dong, X. · Luong, T. · Hsieh, C.-J. · Lu, Y. · Le, Q.V. — NeurIPS
The problem
By 2023, and had been the de facto standard optimizers for training deep neural networks for nearly a decade. Despite hundreds of handcrafted alternatives, none had consistently displaced Adam across vision, language, and multimodal tasks. Meanwhile, attempts to automatically discover better optimizers via learning-to-optimize (L2O) or reinforcement learning produced black-box algorithms that failed to generalize from small proxy tasks to real-world training at scale. The community was stuck: human intuition seemed exhausted, and machine search was too brittle.
The contribution
A program-search framework that formulates discovery as symbolic evolution over an infinite program space, using evolutionary search with warm-start, for pruning, and for . The method discovers Lion (EvoLved ), a remarkably simple optimizer that tracks only (no second moment), uses the for uniform-magnitude updates, and achieves 88.3% zero-shot and 91.1% accuracy on ImageNet — surpassing previous best results — while saving up to 5× compute on JFT. Lion is deployed in production at Google.
The impact
Lion demonstrated that machine-discovered optimizers can genuinely surpass a decade of human-designed ones at scale. It proved that adaptive per-parameter scaling (Adam's hallmark) is not necessary — uniform sign updates with proper momentum tracking can match or beat it. Lion's memory savings (one state tensor vs. two) matter at the billion- parameter frontier. The program-search methodology opened a new path: instead of designing optimizers by intuition, search for them symbolically and let simplicity emerge from selection pressure.
Adam is like a seasoned hiker who measures both the slope and the roughness of the terrain at every step, carefully adjusting stride length for each direction. It works, but the hiker needs two notebooks — one for average slope, one for average roughness — and does a lot of arithmetic at each step.
Lion is a different kind of hiker: it throws away the roughness notebook entirely, keeps only one notebook for slope history, and at each step simply asks "uphill or downhill?" then takes a fixed-size step in that direction. Fewer notebooks, simpler decisions — yet this hiker reaches better viewpoints, faster.
Discovering optimizers: algorithms as programs
The key insight behind this paper is that an optimization algorithm is just a short program: it takes a , some historical state, and a , and outputs an update to the weights. If you can represent optimizers as programs, you can search the space of all possible programs to find ones that train neural networks better than anything humans have designed.
The authors start from AdamW as a warm-start seed and apply : each generation picks a parent program, mutates it (insert, delete, or modify a statement), evaluates the child on a small proxy task, and keeps the best. The search space includes 45 math functions — from basic arithmetic to trigonometric and interpolation operations — with no limit on program length or local variables.
This creates an infinite and sparse search space: over 2 million random programs were tested and none matched AdamW. Efficient techniques make the search tractable: abstract execution prunes invalid, duplicate, and redundant programs (giving ~10× speedup via caching and ~3× shorter programs); warm-start and restart balance exploration vs. exploitation; and the total cost is ~3,000 TPU V2 days across multiple search runs.
Bridging the gap: from proxy to production
A program that wins on a tiny proxy task (20 minutes on one TPU chip) might fail spectacularly on a real target task (days on 512 TPU chips). This generalization gap is the central challenge of automated optimizer discovery, and previous methods like L2O and neural optimizer search failed precisely here.
The authors address this with two strategies. First, funnel selection: promising programs from the search are evaluated on progressively larger tasks — 10×, then 100× the proxy. Only programs that beat the baseline at each scale advance to the next. Second, program simplification: redundant statements are removed (abstract execution identifies that ~70% of statements in evolved programs are dead code), non-essential operations are ablated, and the surviving program is manually cleaned into its simplest mathematical form.
This process also reveals meta-overfitting: search fitness keeps rising but performance on larger tasks declines. Runs that meta-overfit later tend to find algorithms that generalize better — a useful heuristic the authors exploit by running multiple restarts.
Lion: the evolved sign momentum optimizer
After search, funnel selection, and simplification, the winning program reduces to a strikingly simple algorithm: Lion (EvoLved Sign Momentum). Compared to Adam, which tracks both first and second moments of the gradient, Lion tracks only the first moment (momentum) and computes the update by taking the sign of an interpolation between the current gradient and the momentum.
The intuition is: instead of asking "how steep is this slope and how noisy has it been?" (Adam), Lion asks "which direction does the combined evidence point?" and takes a uniform step. Every parameter gets updated by exactly +1 or −1 (times the learning rate), regardless of gradient magnitude. This creates uniform update magnitudes across all dimensions — a property no handcrafted optimizer had exploited successfully at scale.
Two β values control the behavior: β₁ = 0.9 governs the interpolation for computing the update (more weight on the current gradient), while β₂ = 0.99 governs the for tracking momentum (remembering ~10× longer history). This separation — using different blending factors for the update and the state — is something human designers had not tried, and it emerged naturally from the search.
The Lion update rule
The goal is to update the model weights θ at each training step. Lion does this in three stages: (1) compute an interpolated direction from the current gradient and the stored momentum, (2) apply the sign function to get a uniform ±1 update, (3) then separately update the momentum for the next step with a different blending factor. is applied as decoupled , just like in AdamW.
The idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def lion_step(weight, gradient, momentum, lr, beta1=0.9, beta2=0.99, wd=0.1):
"""One step of the Lion optimizer."""
# Step 1: Interpolate gradient and momentum for the update direction
c = beta1 * momentum + (1 - beta1) * gradient
# Step 2: Sign of c → uniform ±1 update + decoupled weight decay
update = np.sign(c) + wd * weight
# Apply the update
weight = weight - lr * update
# Step 3: Update momentum for the NEXT step (separate β₂)
momentum = beta2 * momentum + (1 - beta2) * gradient
return weight, momentum
# Key differences from Adam:
# • No second moment (v) → saves ~50% optimizer memory
# • sign() produces uniform ±1 → use 3-10x smaller learning rate
# • β₁ ≠ β₂ → decouples update direction from state tracking
# • No epsilon, no bias correction → fewer hyperparametersWhy does sign work? Regularization through noise
The sign operation discards gradient magnitude and keeps only direction. This might seem like throwing away useful information, but it introduces a form of implicit regularization: by ignoring how steep a dimension is, the optimizer adds noise to the updates. This noise prevents the model from converging to sharp minima and pushes it toward flatter regions of the — which are known to generalize better.
Empirical evidence confirms this: -B/16 trained with Lion has a higher training error than with AdamW, yet achieves 2% higher validation accuracy. The authors also measure landscape flatness by perturbing converged weights with Gaussian noise. Lion's solution retains much lower error under perturbation, confirming convergence to a flatter region. This behavior resembles (SAM), but Lion achieves it implicitly through the sign operation rather than an explicit min-max objective.
Memory and efficiency: half the state, faster steps
Adam tracks two state tensors per parameter: the first moment (mean of gradients) and the second moment (mean of squared gradients). Lion tracks only one: the momentum. For a model with P parameters, this saves P floats of optimizer state — roughly a 50% reduction in optimizer memory. When training billion-parameter models, this directly translates to fewer accelerator chips. For example, training ViT-B/16 with 4,096 requires at least 16 TPU V4 chips with AdamW but only 8 with Lion (both using bfloat16 momentum).
Lion is also 2–15% faster in wall-clock time per step due to its simplicity: no square root, no epsilon, no bias correction — just one interpolation, one sign, and one momentum update.
Practical tuning: learning rate, weight decay, batch size
Because sign(·) always produces ±1, Lion's updates have a larger norm than Adam's scaled updates. This means Lion needs a smaller learning rate — typically 3–10× smaller than what you would use for Adam. To maintain the same effective weight decay strength (which is lr × λ), the weight decay coefficient λ must be increased by the same factor.
For example, if you use lr = 1e−3 and λ = 1.0 for AdamW, a good starting point for Lion is lr = 1e−4 and λ = 10.0. The initial, peak, and end values of the learning rate schedule should all be scaled by the same ratio.
Lion also prefers larger batch sizes. Its performance advantage over AdamW grows with batch size: at batch size 32K, Lion achieves a 2.5% accuracy gain. The authors hypothesize that larger batches give a more reliable gradient direction, which matters more when the optimizer only uses the sign of that direction. Even at small batch sizes (64), Lion remains competitive — it just shines brighter at scale.
Lion is also more robust to choices: heatmaps of accuracy across learning rate and weight decay values show a broader plateau for Lion than for AdamW.
Results across vision, language, and generation
Lion was evaluated across an unusually broad range of architectures (, MLP, ResNet, U-Net, Hybrid) and tasks (image classification, , diffusion models, language modeling, and fine-tuning). Key highlights:
Image classification: On ImageNet trained from scratch, Lion boosted ViT-B/16 accuracy by 1.96% over AdamW (77.44% vs 75.48%). On JFT pre-training, Lion enabled ViT-L/16 to match ViT-H/14 trained by AdamW while being 2× smaller — or equivalently, saved up to 5× compute to reach the same performance.
Vision-language contrastive learning: Using BASIC-L, Lion achieved 88.3% zero-shot ImageNet accuracy — a 2.6% gain over the Adafactor baseline and 2% above the prior state-of-the-art. Fine-tuning accuracy reached 91.1%.
Diffusion models: On 256×256 ImageNet generation, Lion matched AdamW's final FID in 2.3× fewer iterations (FID 4.1 vs 4.7).
Language modeling: On PG-19, Lion achieved up to 2× speedup over AdamW. On 7.5B parameter models trained on 300B tokens, Lion matched training perplexity but improved in-context learning scores across NLG and NLU benchmarks.
Fine-tuning: On GLUE with T5 models (Base to 11B), Lion won 10–12 out of 12 scores at every scale.
Limitations and when Lion may not help
The authors are transparent about where Lion does not shine. On ResNets, the difference from AdamW and SGD is small — convolutional networks may be easier to optimize than Transformers. With strong (RandAug + Mixup), the regularization benefit of sign shrinks because the augmentation already regularizes. On very large, high-quality datasets (the Imagen base model, large-scale autoregressive pretraining), Lion matches AdamW in perplexity but does not clearly surpass it — the data quality itself reduces the gap between optimizers.
Lion also prefers batch sizes ≥ 64. With very small batches, the gradient direction becomes too noisy, and the sign operation amplifies that noise rather than filtering it. Finally, while Lion saves one state tensor, it still requires momentum storage in bfloat16, which can be expensive at the multi-billion parameter scale. Factorizing the momentum (as Adafactor does for second moments) is suggested as future work.
Lion in context: the optimizer timeline
2014
Adam
Adaptive moment estimation. Tracks first and second moments of gradients. Became the default optimizer for deep learning for nearly a decade.
2017
Decoupled Weight Decay (AdamW)
Loshchilov and Hutter showed that proper weight decay in Adam should be decoupled from the gradient update. AdamW became the new standard.
2020
Sharpness-Aware Minimization (SAM)
Showed that seeking flat minima explicitly improves generalization. Lion achieves a similar effect implicitly through the sign operation.
2020
AutoML-Zero
Ambitious attempt to evolve full ML pipelines from scratch, but only on toy tasks. Lion's search framework builds on this foundation but targets real-world scale.
2023
Lion (this paper)
First machine-discovered optimizer to consistently outperform Adam at state-of-the-art scale. Deployed in Google production systems.
CitationChen, Liang, Huang, Real, Wang, Liu, Pham, Dong, Luong, Hsieh, Lu, Le. Symbolic Discovery of Optimization Algorithms. NeurIPS, 2023.
Terms in this paper
- Optimizerالـمُحسِّن
- Momentumالزخم
- Sign Functionدالة الإشارة
- Weight Decayاضمحلال الأوزان
- Learning Rateمعدل التعلم
- Program Searchالبحث البرمجي
- Adamخوارزمية آدام
- AdamWآدَم مع اضمحلال أوزان مفصول
- Sharpness-Aware Minimizationالأمثَلة الواعية بالحدّة
- Exponential Moving Average (EMA)المتوسط المتحرك الأُسِّي
- Gradient Descentالانحدار التدريجي
- Batch Sizeحجم الدفعة الحسابية
- Fine-Tuningالضبط الدقيق
- Vision Transformer (ViT)محوِّل الرؤية (ViT)
- Diffusion Modelنموذج الانتشار