Learning Rate
معدل التعلم
مُعامِل فائق يتحكم في حجم خطوة تحديث المعاملات أثناء النزول التدريجي. قيمة كبيرة جداً تسبب عدم استقرار، وصغيرة جداً تسبب بطء التقارب.
A hyperparameter that controls the step size of each parameter update during gradient descent. Too large causes instability; too small causes slow convergence.
Also translated asحجم خطوة التحديث، خطوة التحديث، خطوة التحديث الحسابي، سرعة التعلم، سرعة التعلُّم، معدل تعلُّم، وتيرة التكيف الحسابي
First appears in this corpus in: Méthode Générale pour la Résolution des Systèmes d'Équations Simultanées (1847)
Appears in these papers
- Asynchronous Methods for Deep Reinforcement Learning2016in the sky ✦
- Adaptive Subgradient Methods for Online Learning and Stochastic Optimization2011in the sky ✦
- Adaptive Subgradient Methods for Online Learning and Stochastic Optimization2011in the sky ✦
- Adam: A Method for Stochastic Optimization2014in the sky ✦
- Decoupled Weight Decay Regularization2019in the sky ✦
- Parameter-Efficient Transfer Learning for NLP2019in the sky ✦
- ImageNet Classification with Deep Convolutional Neural Networks2012in the sky ✦
- ALIGN: Scaling Up Visual and Vision-Language Representation Learning with Noisy Text Supervision2021in the sky ✦
- Alpaca: A Strong, Replicable Instruction-Following Model2023in the sky ✦
- Backpropagation Through Time: What It Does and How to Do It1990in the sky ✦
- Learning Representations by Back-Propagating Errors1986in the sky ✦
- Learning Representations by Back-Propagating Errors1986in the sky ✦
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift2015in the sky ✦
- Practical Bayesian Optimization of Machine Learning Algorithms2012in the sky ✦
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding2018in the sky ✦
- Bootstrap Your Own Latent: A New Approach to Self-Supervised Learning2020in the sky ✦
- Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks2017in the sky ✦
- Cyclical Learning Rates for Training Neural Networks2017in the sky ✦
- Cyclical Learning Rates for Training Neural Networks2017in the sky ✦
- Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks2015in the sky ✦
- Continuous Control with Deep Reinforcement Learning2015in the sky ✦
- DeepSeek-V3 Technical Report2024in the sky ✦
- Training Data-Efficient Image Transformers & Distillation Through Attention2021in the sky ✦
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data2024in the sky ✦
- Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World2017in the sky ✦
- Dropout: A Simple Way to Prevent Neural Networks from Overfitting2014in the sky ✦
- The Falcon Series of Open Language Models2023in the sky ✦
- The Falcon Series of Open Language Models2023in the sky ✦
- Fully Convolutional Networks for Semantic Segmentation2015in the sky ✦
- Fusion-in-Decoder: Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering2021in the sky ✦
- Finetuned Language Models Are Zero-Shot Learners2022in the sky ✦
- Gaussian Error Linear Units (GELUs)2016in the sky ✦
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher2022in the sky ✦
- Greedy Function Approximation: A Gradient Boosting Machine2001in the sky ✦
- On the Difficulty of Training Recurrent Neural Networks2013in the sky ✦
- Méthode Générale pour la Résolution des Systèmes d'Équations Simultanées1847in the sky ✦
- Méthode Générale pour la Résolution des Systèmes d'Équations Simultanées1847in the sky ✦
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling2014in the sky ✦
- The Organization of Behavior: A Neuropsychological Theory1949in the sky ✦
- Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization2018in the sky ✦
- Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization2018in the sky ✦
- Optimizing Neural Networks with Kronecker-Factored Approximate Curvature2015in the sky ✦
- When Does Label Smoothing Help?2019in the sky ✦
- Large Batch Optimization for Deep Learning: Training BERT in 76 Minutes2019in the sky ✦
- Large Batch Optimization for Deep Learning: Training BERT in 76 Minutes2019in the sky ✦
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour2017in the sky ✦
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour2017in the sky ✦
- LightGBM: A Highly Efficient Gradient Boosting Decision Tree2017in the sky ✦
- LightGBM: A Highly Efficient Gradient Boosting Decision Tree2017in the sky ✦
- LIMA: Less Is More for Alignment2023in the sky ✦
- Symbolic Discovery of Optimization Algorithms2023in the sky ✦
- Symbolic Discovery of Optimization Algorithms2023in the sky ✦
- Visualizing the Loss Landscape of Neural Nets2018in the sky ✦
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks2019in the sky ✦
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks2019in the sky ✦
- Mixed Precision Training2018in the sky ✦
- Some Methods of Speeding Up the Convergence of Iteration Methods1964in the sky ✦
- Some Methods of Speeding Up the Convergence of Iteration Methods1964in the sky ✦
- Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer2022in the sky ✦
- Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer2022in the sky ✦
- Natural Gradient Works Efficiently in Learning1998in the sky ✦
- A Method for Solving the Convex Programming Problem with Convergence Rate O(1/k²)1983in the sky ✦
- A Method for Solving the Convex Programming Problem with Convergence Rate O(1/k²)1983in the sky ✦
- Neural Tangent Kernel: Convergence and Generalization in Neural Networks2018in the sky ✦
- OPT: Open Pre-Trained Transformer Language Models2022in the sky ✦
- OPT: Open Pre-Trained Transformer Language Models2022in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Proximal Policy Optimization Algorithms2017in the sky ✦
- Prioritized Experience Replay2015in the sky ✦
- Q-Learning1992in the sky ✦
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation2018in the sky ✦
- Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation2014in the sky ✦
- RMSProp: Divide the Gradient by a Running Average of Its Recent Magnitude2012in the sky ✦
- RMSProp: Divide the Gradient by a Running Average of Its Recent Magnitude2012in the sky ✦
- RoBERTa: A Robustly Optimized BERT Pretraining Approach2019in the sky ✦
- Scaling Laws for Neural Language Models2020in the sky ✦
- Sequence to Sequence Learning with Neural Networks2014in the sky ✦
- SGDR: Stochastic Gradient Descent with Warm Restarts2017in the sky ✦
- SGDR: Stochastic Gradient Descent with Warm Restarts2017in the sky ✦
- Shampoo: Preconditioned Stochastic Tensor Optimization2018in the sky ✦
- Sharpness-Aware Minimization for Efficiently Improving Generalization2021in the sky ✦
- Exploring Simple Siamese Representation Learning2021in the sky ✦
- Support-Vector Networks1995in the sky ✦
- Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows2021in the sky ✦
- Learning to Predict by the Methods of Temporal Differences1988in the sky ✦
- Toolformer: Language Models Can Teach Themselves to Use Tools2023in the sky ✦
- Trust Region Policy Optimization2015in the sky ✦
- Universal Language Model Fine-Tuning for Text Classification2018in the sky ✦
- Robust Speech Recognition via Large-Scale Weak Supervision2022in the sky ✦
- XGBoost: A Scalable Tree Boosting System2016in the sky ✦