Adam
خوارزمية آدام
محسن تكيفي متطور يدمج بين مزايا العزم التراكمي ومعدل التعلم المتغير لكل معلمة بشكل مستقل.
Adam
Also translated asالـمُحسِّن التكيفي العزمي، محسن آدام
First appears in this corpus in: Méthode Générale pour la Résolution des Systèmes d'Équations Simultanées (1847)
Appears in these papers
- Adaptive Subgradient Methods for Online Learning and Stochastic Optimization2011in the sky ✦
- Adam: A Method for Stochastic Optimization2014in the sky ✦
- Adam: A Method for Stochastic Optimization2014in the sky ✦
- Cyclical Learning Rates for Training Neural Networks2017in the sky ✦
- Cyclical Learning Rates for Training Neural Networks2017in the sky ✦
- Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks2015in the sky ✦
- Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World2017in the sky ✦
- The Falcon Series of Open Language Models2023in the sky ✦
- Google's Neural Machine Translation System: Bridging the Gap Between Human and Machine Translation2016in the sky ✦
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher2022in the sky ✦
- Méthode Générale pour la Résolution des Systèmes d'Équations Simultanées1847in the sky ✦
- Large Batch Optimization for Deep Learning: Training BERT in 76 Minutes2019in the sky ✦
- Symbolic Discovery of Optimization Algorithms2023in the sky ✦
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism2019in the sky ✦
- Some Methods of Speeding Up the Convergence of Iteration Methods1964in the sky ✦
- Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer2022in the sky ✦
- Natural Gradient Works Efficiently in Learning1998in the sky ✦
- A Method for Solving the Convex Programming Problem with Convergence Rate O(1/k²)1983in the sky ✦
- A Method for Solving the Convex Programming Problem with Convergence Rate O(1/k²)1983in the sky ✦
- OPT: Open Pre-Trained Transformer Language Models2022in the sky ✦
- QLoRA: Efficient Finetuning of Quantized LLMs2023in the sky ✦
- RMSProp: Divide the Gradient by a Running Average of Its Recent Magnitude2012in the sky ✦
- SGDR: Stochastic Gradient Descent with Warm Restarts2017in the sky ✦
- Shampoo: Preconditioned Stochastic Tensor Optimization2018in the sky ✦
- Auto-Encoding Variational Bayes2013in the sky ✦
- Robust Speech Recognition via Large-Scale Weak Supervision2022in the sky ✦
- ZeRO: Memory Optimizations Toward Training Trillion Parameter Models2019in the sky ✦