Dropout
الإسقاط العشوائي للعصبونات
تقنية كبح للإفراط تعطل عشوائياً نسبة من العصبونات في كل خطوة تدريب لإجبار الشبكة على عدم الاعتماد على مسارات محددة.
Dropout
Also translated asالتعطيل المؤقت لمسارات المعالجة، آلية كبح الإفراط بالتحييد الدوري، الإسقاط الترشيحي
First appears in this corpus in: The Monte Carlo Method (1949)
Appears in these papers
- ImageNet Classification with Deep Convolutional Neural Networks2012in the sky ✦
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift2015in the sky ✦
- BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer2019in the sky ✦
- Neural Networks and the Bias/Variance Dilemma1992in the sky ✦
- Reducing the Dimensionality of Data with Neural Networks2006in the sky ✦
- Deep Learning2015in the sky ✦
- Deep Learning2015in the sky ✦
- Training Data-Efficient Image Transformers & Distillation Through Attention2021in the sky ✦
- Deep Interest Network for Click-Through Rate Prediction2018in the sky ✦
- Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World2017in the sky ✦
- Dropout: A Simple Way to Prevent Neural Networks from Overfitting2014in the sky ✦
- Dropout: A Simple Way to Prevent Neural Networks from Overfitting2014in the sky ✦
- EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks2019in the sky ✦
- Fast R-CNN2015in the sky ✦
- Explaining and Harnessing Adversarial Examples2015in the sky ✦
- Fusion-in-Decoder: Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering2021in the sky ✦
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness2022in the sky ✦
- Generative Adversarial Networks2014in the sky ✦
- Graph Attention Networks2018in the sky ✦
- Gaussian Error Linear Units (GELUs)2016in the sky ✦
- Gaussian Error Linear Units (GELUs)2016in the sky ✦
- Going Deeper with Convolutions2014in the sky ✦
- Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets2022in the sky ✦
- LIMA: Less Is More for Alignment2023in the sky ✦
- LIMA: Less Is More for Alignment2023in the sky ✦
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks2019in the sky ✦
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism2019in the sky ✦
- mixup: Beyond Empirical Risk Minimization2018in the sky ✦
- mixup: Beyond Empirical Risk Minimization2018in the sky ✦
- The Monte Carlo Method1949in the sky ✦
- Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer2022in the sky ✦
- N-BEATS: Neural Basis Expansion Analysis for Interpretable Time Series Forecasting2019in the sky ✦
- No Free Lunch Theorems for Optimization1997in the sky ✦
- OPT: Open Pre-Trained Transformer Language Models2022in the sky ✦
- A Time Series Is Worth 64 Words: Long-Term Forecasting with Transformers2023in the sky ✦
- A Time Series Is Worth 64 Words: Long-Term Forecasting with Transformers2023in the sky ✦
- Image-to-Image Translation with Conditional Adversarial Networks2017in the sky ✦
- PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation2017in the sky ✦
- Proximal Policy Optimization Algorithms2017in the sky ✦
- Ridge Regression: Biased Estimation for Nonorthogonal Problems1970in the sky ✦
- RoBERTa: A Robustly Optimized BERT Pretraining Approach2019in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦
- Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions2018in the sky ✦
- Universal Language Model Fine-Tuning for Text Classification2018in the sky ✦
- VQA: Visual Question Answering2015in the sky ✦