Transformer
المحوِّل
بنية شبكة عصبية قدّمها فاسواني وزملاؤه في جوجل عام 2017 في ورقة «Attention Is All You Need»، تعتمد كلياً على الانتباه الذاتي بدل الشبكات العودية أو الالتفافية، وأصبحت الأساس الحاسوبي لكل النماذج اللغوية والبصرية والمتعدّدة الوسائط الكبرى الحديثة.
A neural network architecture introduced by Vaswani and colleagues at Google in the 2017 paper 'Attention Is All You Need' that relies entirely on self-attention instead of recurrence or convolution, becoming the computational foundation of essentially every major modern language, vision, and multimodal model.
Also translated asTransformer، Attention Is All You Need، بنية المحوِّل، محوِّل فاسواني
First appears in this corpus in: A Logical Calculus of the Ideas Immanent in Nervous Activity (1943)
Appears in these papers
- Decoupled Weight Decay Regularization2019in the sky ✦
- Parameter-Efficient Transfer Learning for NLP2019in the sky ✦
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware2023in the sky ✦
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware2023in the sky ✦
- Highly Accurate Protein Structure Prediction with AlphaFold2021in the sky ✦
- Grandmaster Level in StarCraft II Using Multi-Agent Reinforcement Learning2019in the sky ✦
- Discovering Faster Matrix Multiplication Algorithms with Reinforcement Learning2022in the sky ✦
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting2021in the sky ✦
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting2021in the sky ✦
- Learning Representations by Back-Propagating Errors1986in the sky ✦
- BART: Denoising Sequence-to-Sequence Pre-Training for Natural Language Generation, Translation, and Comprehension2019in the sky ✦
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift2015in the sky ✦
- BEiT: BERT Pre-Training of Image Transformers2021in the sky ✦
- Learning Long-Term Dependencies with Gradient Descent is Difficult1994in the sky ✦
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding2018in the sky ✦
- BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer2019in the sky ✦
- BLEU: A Method for Automatic Evaluation of Machine Translation2002in the sky ✦
- BLIP-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models2023in the sky ✦
- BLOOM: A 176B-Parameter Open-Access Multilingual Language Model2022in the sky ✦
- Chronos: Learning the Language of Time Series2024in the sky ✦
- Learning Transferable Visual Models from Natural Language Supervision2021in the sky ✦
- Contriever: Unsupervised Dense Information Retrieval with Contrastive Learning2022in the sky ✦
- A ConvNet for the 2020s2022in the sky ✦
- Gradient-Based Learning Applied to Document Recognition1998in the sky ✦
- DALL·E: Zero-Shot Text-to-Image Generation2021in the sky ✦
- DALL·E: Zero-Shot Text-to-Image Generation2021in the sky ✦
- Denoising Diffusion Probabilistic Models2020in the sky ✦
- DeBERTa: Decoding-Enhanced BERT with Disentangled Attention2020in the sky ✦
- DeBERTa: Decoding-Enhanced BERT with Disentangled Attention2020in the sky ✦
- Decision Transformer: Reinforcement Learning via Sequence Modeling2021in the sky ✦
- Decision Transformer: Reinforcement Learning via Sequence Modeling2021in the sky ✦
- Deep Learning2015in the sky ✦
- Deep Speech 2: End-to-End Speech Recognition in English and Mandarin2015in the sky ✦
- DeepSeek-V3 Technical Report2024in the sky ✦
- Deformable Convolutional Networks2017in the sky ✦
- Training Data-Efficient Image Transformers & Distillation Through Attention2021in the sky ✦
- End-to-End Object Detection with Transformers2020in the sky ✦
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion2023in the sky ✦
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion2023in the sky ✦
- Deep Interest Network for Click-Through Rate Prediction2018in the sky ✦
- Emerging Properties in Self-Supervised Vision Transformers2021in the sky ✦
- DistilBERT, a Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter2019in the sky ✦
- Learning Distributed Representations of Concepts1986in the sky ✦
- Are Transformers Effective for Time Series Forecasting?2023in the sky ✦
- Dropout: A Simple Way to Prevent Neural Networks from Overfitting2014in the sky ✦
- ELECTRA: Pre-Training Text Encoders as Discriminators Rather Than Generators2020in the sky ✦
- Deep Contextualized Word Representations2018in the sky ✦
- Are Emergent Abilities of Large Language Models a Mirage?2023in the sky ✦
- The Falcon Series of Open Language Models2023in the sky ✦
- Flamingo: a Visual Language Model for Few-Shot Learning2022in the sky ✦
- Finetuned Language Models Are Zero-Shot Learners2022in the sky ✦
- Graph Attention Networks2018in the sky ✦
- A Generalist Agent2022in the sky ✦
- A Generalist Agent2022in the sky ✦
- Semi-Supervised Classification with Graph Convolutional Networks2017in the sky ✦
- Gaussian Error Linear Units (GELUs)2016in the sky ✦
- Gemini: A Family of Highly Capable Multimodal Models2023in the sky ✦
- Gemini: A Family of Highly Capable Multimodal Models2023in the sky ✦
- Gemma: Open Models Based on Gemini Research and Technology2024in the sky ✦
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models2021in the sky ✦
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models2021in the sky ✦
- GloVe: Global Vectors for Word Representation2014in the sky ✦
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher2022in the sky ✦
- Improving Language Understanding by Generative Pre-Training2018in the sky ✦
- Language Models Are Unsupervised Multitask Learners2019in the sky ✦
- Language Models Are Few-Shot Learners2020in the sky ✦
- GPT-4 Technical Report2023in the sky ✦
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints2023in the sky ✦
- On the Difficulty of Training Recurrent Neural Networks2013in the sky ✦
- Do Transformers Really Perform Bad for Graph Representation?2021in the sky ✦
- Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets2022in the sky ✦
- Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection2023in the sky ✦
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling2014in the sky ✦
- Highway Networks2015in the sky ✦
- ImageBind: One Embedding Space to Bind Them All2023in the sky ✦
- Informer: Beyond Efficient Transformer for Long Sequence Time Series Forecasting2021in the sky ✦
- Jukebox: A Generative Model for Music2020in the sky ✦
- Jukebox: A Generative Model for Music2020in the sky ✦
- When Does Label Smoothing Help?2019in the sky ✦
- Large Batch Optimization for Deep Learning: Training BERT in 76 Minutes2019in the sky ✦
- Layer Normalization2016in the sky ✦
- Learning to Summarize from Human Feedback2020in the sky ✦
- Symbolic Discovery of Optimization Algorithms2023in the sky ✦
- Llama 2: Open Foundation and Fine-Tuned Chat Models2023in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- LLaMA: Open and Efficient Foundation Language Models2023in the sky ✦
- Visual Instruction Tuning2023in the sky ✦
- LLM.int8(): 8-Bit Matrix Multiplication for Transformers at Scale2022in the sky ✦
- LLM.int8(): 8-Bit Matrix Multiplication for Transformers at Scale2022in the sky ✦
- LoRA: Low-Rank Adaptation of Large Language Models2021in the sky ✦
- Long Short-Term Memory1997in the sky ✦
- Masked Autoencoders Are Scalable Vision Learners2022in the sky ✦
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces2023in the sky ✦
- A Logical Calculus of the Ideas Immanent in Nervous Activity1943in the sky ✦
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism2019in the sky ✦
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism2019in the sky ✦
- Mistral 7B2023in the sky ✦
- Mixtral of Experts2024in the sky ✦
- Neural Message Passing for Quantum Chemistry2017in the sky ✦
- Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer2022in the sky ✦
- MusicLM: Generating Music From Text2023in the sky ✦
- A Neural Probabilistic Language Model2003in the sky ✦
- No Free Lunch Theorems for Optimization1997in the sky ✦
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models2024in the sky ✦
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models2024in the sky ✦
- OPT: Open Pre-Trained Transformer Language Models2022in the sky ✦
- Efficient Memory Management for Large Language Model Serving with PagedAttention2023in the sky ✦
- PaLM: Scaling Language Modeling with Pathways2022in the sky ✦
- A Time Series Is Worth 64 Words: Long-Term Forecasting with Transformers2023in the sky ✦
- A Time Series Is Worth 64 Words: Long-Term Forecasting with Transformers2023in the sky ✦
- Perceiver: General Perception with Iterative Attention2021in the sky ✦
- Perceiver: General Perception with Iterative Attention2021in the sky ✦
- The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain1958in the sky ✦
- Textbooks Are All You Need2023in the sky ✦
- π₀: A Vision-Language-Action Flow Model for General Robot Control2024in the sky ✦
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks2020in the sky ✦
- REALM: Retrieval-Augmented Language Model Pre-Training2020in the sky ✦
- Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning1992in the sky ✦
- Deep Residual Learning for Image Recognition2015in the sky ✦
- Learning Phrase Representations Using RNN Encoder-Decoder for Statistical Machine Translation2014in the sky ✦
- Recurrent Neural Networks (RNNs): A Gentle Introduction and Overview2019in the sky ✦
- Recurrent Neural Networks (RNNs): A Gentle Introduction and Overview2019in the sky ✦
- RoBERTa: A Robustly Optimized BERT Pretraining Approach2019in the sky ✦
- RT-1: Robotics Transformer for Real-World Control at Scale2022in the sky ✦
- RT-1: Robotics Transformer for Real-World Control at Scale2022in the sky ✦
- Efficiently Modeling Long Sequences with Structured State Spaces2022in the sky ✦
- Scaling Laws for Neural Language Models2020in the sky ✦
- SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers2021in the sky ✦
- Segment Anything2023in the sky ✦
- Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks2019in the sky ✦
- Sequence to Sequence Learning with Neural Networks2014in the sky ✦
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention2015in the sky ✦
- Video Generation Models as World Simulators2024in the sky ✦
- Fast Inference from Transformers via Speculative Decoding2023in the sky ✦
- Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows2021in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer2019in the sky ✦
- A Decoder-Only Foundation Model for Time Series Forecasting2024in the sky ✦
- Attention Is All You Need2017in the sky ✦
- Attention Is All You Need2017in the sky ✦
- Universal Language Model Fine-Tuning for Text Classification2018in the sky ✦
- Approximation by Superpositions of a Sigmoidal Function1989in the sky ✦
- Very Deep Convolutional Networks for Large-Scale Image Recognition2014in the sky ✦
- Rapid Object Detection Using a Boosted Cascade of Simple Features2001in the sky ✦
- An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale2020in the sky ✦
- VQA: Visual Question Answering2015in the sky ✦
- Neural Discrete Representation Learning2017in the sky ✦
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations2020in the sky ✦
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations2020in the sky ✦
- WaveNet: A Generative Model for Raw Audio2016in the sky ✦
- Robust Speech Recognition via Large-Scale Weak Supervision2022in the sky ✦
- XLNet: Generalized Autoregressive Pretraining for Language Understanding2019in the sky ✦