Feed Forward Network (FFN)
شبكة التغذية الأمامية
شبكة عصبية خطية متتالية تطبق تحولات مستقلة على كل وحدة لغوية بعد حساب الانتباه.
Feed Forward Network (FFN)
Also translated asشبكة التمرير الأمامي، الشبكة الخطية المتتالية
First appears in this corpus in: Neural Machine Translation by Jointly Learning to Align and Translate (2014)
Appears in these papers
- Parameter-Efficient Transfer Learning for NLP2019in the sky ✦
- Neural Machine Translation by Jointly Learning to Align and Translate2014in the sky ✦
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting2021in the sky ✦
- BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer2019in the sky ✦
- BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer2019in the sky ✦
- DeepSeek-V3 Technical Report2024in the sky ✦
- End-to-End Object Detection with Transformers2020in the sky ✦
- Deep Interest Network for Click-Through Rate Prediction2018in the sky ✦
- DistilBERT, a Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter2019in the sky ✦
- Are Transformers Effective for Time Series Forecasting?2023in the sky ✦
- Graph Attention Networks2018in the sky ✦
- Gaussian Error Linear Units (GELUs)2016in the sky ✦
- Improving Language Understanding by Generative Pre-Training2018in the sky ✦
- Language Models Are Unsupervised Multitask Learners2019in the sky ✦
- Highway Networks2015in the sky ✦
- Informer: Beyond Efficient Transformer for Long Sequence Time Series Forecasting2021in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- LLaMA: Open and Efficient Foundation Language Models2023in the sky ✦
- LoRA: Low-Rank Adaptation of Large Language Models2021in the sky ✦
- Mistral 7B2023in the sky ✦
- Mistral 7B2023in the sky ✦
- Mixtral of Experts2024in the sky ✦
- PaLM: Scaling Language Modeling with Pathways2022in the sky ✦
- A Time Series Is Worth 64 Words: Long-Term Forecasting with Transformers2023in the sky ✦
- A Time Series Is Worth 64 Words: Long-Term Forecasting with Transformers2023in the sky ✦
- Scaling Laws for Neural Language Models2020in the sky ✦
- Segment Anything2023in the sky ✦
- Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows2021in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer2019in the sky ✦
- Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions2018in the sky ✦
- A Decoder-Only Foundation Model for Time Series Forecasting2024in the sky ✦
- Attention Is All You Need2017in the sky ✦