Batch Size
حجم الدفعة الحسابية
إجمالي عدد عينات البيانات التي يتم معالجتها وحساب تدرجاتها معاً قبل كل عملية تحديث للأوزان البنيوية.
Batch Size
Also translated asسعة المجموعة الجزئية للمعالجة، النطاق العددي لحزمة التدريب
First appears in this corpus in: ImageNet Classification with Deep Convolutional Neural Networks (2012)
Appears in these papers
- Decoupled Weight Decay Regularization2019in the sky ✦
- Parameter-Efficient Transfer Learning for NLP2019in the sky ✦
- ImageNet Classification with Deep Convolutional Neural Networks2012in the sky ✦
- ALIGN: Scaling Up Visual and Vision-Language Representation Learning with Noisy Text Supervision2021in the sky ✦
- ALIGN: Scaling Up Visual and Vision-Language Representation Learning with Noisy Text Supervision2021in the sky ✦
- Alpaca: A Strong, Replicable Instruction-Following Model2023in the sky ✦
- Barlow Twins: Self-Supervised Learning via Redundancy Reduction2021in the sky ✦
- BART: Denoising Sequence-to-Sequence Pre-Training for Natural Language Generation, Translation, and Comprehension2019in the sky ✦
- Practical Bayesian Optimization of Machine Learning Algorithms2012in the sky ✦
- Bootstrap Your Own Latent: A New Approach to Self-Supervised Learning2020in the sky ✦
- Contriever: Unsupervised Dense Information Retrieval with Contrastive Learning2022in the sky ✦
- A ConvNet for the 2020s2022in the sky ✦
- Cyclical Learning Rates for Training Neural Networks2017in the sky ✦
- Cyclical Learning Rates for Training Neural Networks2017in the sky ✦
- DistilBERT, a Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter2019in the sky ✦
- The Falcon Series of Open Language Models2023in the sky ✦
- Fusion-in-Decoder: Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering2021in the sky ✦
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning2023in the sky ✦
- A Generalist Agent2022in the sky ✦
- Glow: Generative Flow with Invertible 1×1 Convolutions2018in the sky ✦
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher2022in the sky ✦
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher2022in the sky ✦
- Group Normalization2018in the sky ✦
- Group Normalization2018in the sky ✦
- Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization2018in the sky ✦
- Large Batch Optimization for Deep Learning: Training BERT in 76 Minutes2019in the sky ✦
- Large Batch Optimization for Deep Learning: Training BERT in 76 Minutes2019in the sky ✦
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour2017in the sky ✦
- LIMA: Less Is More for Alignment2023in the sky ✦
- Symbolic Discovery of Optimization Algorithms2023in the sky ✦
- Symbolic Discovery of Optimization Algorithms2023in the sky ✦
- LLaMA: Open and Efficient Foundation Language Models2023in the sky ✦
- Visualizing the Loss Landscape of Neural Nets2018in the sky ✦
- Mixed Precision Training2018in the sky ✦
- Momentum Contrast for Unsupervised Visual Representation Learning2020in the sky ✦
- OPT: Open Pre-Trained Transformer Language Models2022in the sky ✦
- OPT: Open Pre-Trained Transformer Language Models2022in the sky ✦
- Efficient Memory Management for Large Language Model Serving with PagedAttention2023in the sky ✦
- Efficient Memory Management for Large Language Model Serving with PagedAttention2023in the sky ✦
- RoBERTa: A Robustly Optimized BERT Pretraining Approach2019in the sky ✦
- RoBERTa: A Robustly Optimized BERT Pretraining Approach2019in the sky ✦
- Scaling Laws for Neural Language Models2020in the sky ✦
- Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks2019in the sky ✦
- A Simple Framework for Contrastive Learning of Visual Representations2020in the sky ✦
- A Simple Framework for Contrastive Learning of Visual Representations2020in the sky ✦
- Exploring Simple Siamese Representation Learning2021in the sky ✦
- Fast Inference from Transformers via Speculative Decoding2023in the sky ✦
- SwAV: Unsupervised Learning of Visual Features by Contrasting Cluster Assignments2020in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦
- Toolformer: Language Models Can Teach Themselves to Use Tools2023in the sky ✦
- Robust Speech Recognition via Large-Scale Weak Supervision2022in the sky ✦
- ZeRO: Memory Optimizations Toward Training Trillion Parameter Models2019in the sky ✦