Layer Normalization
التسوية الطبقية
تقنية تسوية اقترحها با وكيروس وهينتون، تُطبِّع الأنشطة داخل كل طبقة بحساب المتوسط والتباين عبر أبعاد السمات لكل عيّنة على حدة، فتعمل جيداً مع الشبكات العودية والدفعات الصغيرة، وأصبحت التسوية الافتراضية في المحوِّلات.
A normalization technique proposed by Ba, Kiros, and Hinton that normalizes activations within each layer by computing the mean and variance across feature dimensions for each sample independently, working well with recurrent networks and small batches, and now the default normalization in Transformers.
Also translated asLayer Norm، LayerNorm، تسوية الطبقة، التطبيع الطبقي
First appears in this corpus in: Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift (2015)
Appears in these papers
- Parameter-Efficient Transfer Learning for NLP2019in the sky ✦
- Parameter-Efficient Transfer Learning for NLP2019in the sky ✦
- BART: Denoising Sequence-to-Sequence Pre-Training for Natural Language Generation, Translation, and Comprehension2019in the sky ✦
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift2015in the sky ✦
- BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer2019in the sky ✦
- BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer2019in the sky ✦
- BLOOM: A 176B-Parameter Open-Access Multilingual Language Model2022in the sky ✦
- BLOOM: A 176B-Parameter Open-Access Multilingual Language Model2022in the sky ✦
- A ConvNet for the 2020s2022in the sky ✦
- A ConvNet for the 2020s2022in the sky ✦
- DeBERTa: Decoding-Enhanced BERT with Disentangled Attention2020in the sky ✦
- DeBERTa: Decoding-Enhanced BERT with Disentangled Attention2020in the sky ✦
- Decision Transformer: Reinforcement Learning via Sequence Modeling2021in the sky ✦
- Mastering Diverse Domains Through World Models2023in the sky ✦
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness2022in the sky ✦
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher2022in the sky ✦
- Improving Language Understanding by Generative Pre-Training2018in the sky ✦
- Language Models Are Unsupervised Multitask Learners2019in the sky ✦
- Do Transformers Really Perform Bad for Graph Representation?2021in the sky ✦
- Do Transformers Really Perform Bad for Graph Representation?2021in the sky ✦
- Group Normalization2018in the sky ✦
- Group Normalization2018in the sky ✦
- Informer: Beyond Efficient Transformer for Long Sequence Time Series Forecasting2021in the sky ✦
- Layer Normalization2016in the sky ✦
- Layer Normalization2016in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- LLaMA: Open and Efficient Foundation Language Models2023in the sky ✦
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism2019in the sky ✦
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism2019in the sky ✦
- PaLM: Scaling Language Modeling with Pathways2022in the sky ✦
- SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers2021in the sky ✦
- Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows2021in the sky ✦
- Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows2021in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer2019in the sky ✦
- A Decoder-Only Foundation Model for Time Series Forecasting2024in the sky ✦
- A Decoder-Only Foundation Model for Time Series Forecasting2024in the sky ✦
- Attention Is All You Need2017in the sky ✦
- An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale2020in the sky ✦
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations2020in the sky ✦
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations2020in the sky ✦