Weight Decay
اضمحلال الأوزان
تقنية تمنع الإفراط عبر فرض غرامة عددية تقلص الأوزان الضخمة وتجعلها قريبة من الصفر.
Weight Decay
Also translated asكبح الأوزان الزائدة، تقليص القيم البنيوية، اضمحلال أوزان
First appears in this corpus in: Ridge Regression: Biased Estimation for Nonorthogonal Problems (1970)
Appears in these papers
- Adaptive Subgradient Methods for Online Learning and Stochastic Optimization2011in the sky ✦
- Decoupled Weight Decay Regularization2019in the sky ✦
- ImageNet Classification with Deep Convolutional Neural Networks2012in the sky ✦
- ALIGN: Scaling Up Visual and Vision-Language Representation Learning with Noisy Text Supervision2021in the sky ✦
- Alpaca: A Strong, Replicable Instruction-Following Model2023in the sky ✦
- Backpropagation Through Time: What It Does and How to Do It1990in the sky ✦
- Neural Networks and the Bias/Variance Dilemma1992in the sky ✦
- Training Data-Efficient Image Transformers & Distillation Through Attention2021in the sky ✦
- Extracting and Composing Robust Features with Denoising Autoencoders2008in the sky ✦
- Dropout: A Simple Way to Prevent Neural Networks from Overfitting2014in the sky ✦
- The Falcon Series of Open Language Models2023in the sky ✦
- The Falcon Series of Open Language Models2023in the sky ✦
- Explaining and Harnessing Adversarial Examples2015in the sky ✦
- On the Difficulty of Training Recurrent Neural Networks2013in the sky ✦
- Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets2022in the sky ✦
- Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets2022in the sky ✦
- Large Batch Optimization for Deep Learning: Training BERT in 76 Minutes2019in the sky ✦
- Large Batch Optimization for Deep Learning: Training BERT in 76 Minutes2019in the sky ✦
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour2017in the sky ✦
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour2017in the sky ✦
- Regression Shrinkage and Selection via the Lasso1996in the sky ✦
- LIMA: Less Is More for Alignment2023in the sky ✦
- Symbolic Discovery of Optimization Algorithms2023in the sky ✦
- Symbolic Discovery of Optimization Algorithms2023in the sky ✦
- Visualizing the Loss Landscape of Neural Nets2018in the sky ✦
- mixup: Beyond Empirical Risk Minimization2018in the sky ✦
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications2017in the sky ✦
- Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer2022in the sky ✦
- A Neural Probabilistic Language Model2003in the sky ✦
- Neural Tangent Kernel: Convergence and Generalization in Neural Networks2018in the sky ✦
- OPT: Open Pre-Trained Transformer Language Models2022in the sky ✦
- Optimal Brain Damage1989in the sky ✦
- Ridge Regression: Biased Estimation for Nonorthogonal Problems1970in the sky ✦
- RoBERTa: A Robustly Optimized BERT Pretraining Approach2019in the sky ✦
- SGDR: Stochastic Gradient Descent with Warm Restarts2017in the sky ✦
- SGDR: Stochastic Gradient Descent with Warm Restarts2017in the sky ✦
- Sharpness-Aware Minimization for Efficiently Improving Generalization2021in the sky ✦
- Sharpness-Aware Minimization for Efficiently Improving Generalization2021in the sky ✦