Attention
آلية الانتباه
ميكانيكية رياضية تمكن النموذج من التركيز الموجه على أجزاء محددة وهامة من السياق وتجاهل الفائض.
Attention
Also translated asالتركيز الموجه على السياق، معيار الترجيح السياقي
First appears in this corpus in: ELIZA — A Computer Program for the Study of Natural Language Communication (1966)
Appears in these papers
- Highly Accurate Protein Structure Prediction with AlphaFold2021in the sky ✦
- Neural Machine Translation by Jointly Learning to Align and Translate2014in the sky ✦
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting2021in the sky ✦
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting2021in the sky ✦
- BEiT: BERT Pre-Training of Image Transformers2021in the sky ✦
- BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer2019in the sky ✦
- Neural Networks and the Bias/Variance Dilemma1992in the sky ✦
- BLEU: A Method for Automatic Evaluation of Machine Translation2002in the sky ✦
- BLOOM: A 176B-Parameter Open-Access Multilingual Language Model2022in the sky ✦
- BLOOM: A 176B-Parameter Open-Access Multilingual Language Model2022in the sky ✦
- Dynamic Routing Between Capsules2017in the sky ✦
- Chronos: Learning the Language of Time Series2024in the sky ✦
- Learning Transferable Visual Models from Natural Language Supervision2021in the sky ✦
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT2020in the sky ✦
- A ConvNet for the 2020s2022in the sky ✦
- DeBERTa: Decoding-Enhanced BERT with Disentangled Attention2020in the sky ✦
- Training Data-Efficient Image Transformers & Distillation Through Attention2021in the sky ✦
- End-to-End Object Detection with Transformers2020in the sky ✦
- Diffusion Models Beat GANs on Image Synthesis2021in the sky ✦
- Diffusion Models Beat GANs on Image Synthesis2021in the sky ✦
- Deep Interest Network for Click-Through Rate Prediction2018in the sky ✦
- Deep Interest Network for Click-Through Rate Prediction2018in the sky ✦
- Emerging Properties in Self-Supervised Vision Transformers2021in the sky ✦
- DINOv2: Learning Robust Visual Features Without Supervision2023in the sky ✦
- Dueling Network Architectures for Deep Reinforcement Learning2016in the sky ✦
- ELECTRA: Pre-Training Text Encoders as Discriminators Rather Than Generators2020in the sky ✦
- ELIZA — A Computer Program for the Study of Natural Language Communication1966in the sky ✦
- Deep Contextualized Word Representations2018in the sky ✦
- The Falcon Series of Open Language Models2023in the sky ✦
- Enriching Word Vectors with Subword Information2017in the sky ✦
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning2023in the sky ✦
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning2023in the sky ✦
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness2022in the sky ✦
- Graph Attention Networks2018in the sky ✦
- Graph Attention Networks2018in the sky ✦
- A Generalist Agent2022in the sky ✦
- Gaussian Processes for Machine Learning2006in the sky ✦
- Semi-Supervised Classification with Graph Convolutional Networks2017in the sky ✦
- Gemini: A Family of Highly Capable Multimodal Models2023in the sky ✦
- Gemini: A Family of Highly Capable Multimodal Models2023in the sky ✦
- Gemma: Open Models Based on Gemini Research and Technology2024in the sky ✦
- How Powerful Are Graph Neural Networks?2019in the sky ✦
- Google's Neural Machine Translation System: Bridging the Gap Between Human and Machine Translation2016in the sky ✦
- Google's Neural Machine Translation System: Bridging the Gap Between Human and Machine Translation2016in the sky ✦
- Going Deeper with Convolutions2014in the sky ✦
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher2022in the sky ✦
- Improving Language Understanding by Generative Pre-Training2018in the sky ✦
- Language Models Are Unsupervised Multitask Learners2019in the sky ✦
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints2023in the sky ✦
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints2023in the sky ✦
- Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization2017in the sky ✦
- Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets2022in the sky ✦
- Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection2023in the sky ✦
- A Tutorial on Hidden Markov Models and Selected Applications in Speech Recognition1989in the sky ✦
- The Mathematics of Statistical Machine Translation: Parameter Estimation1993in the sky ✦
- Informer: Beyond Efficient Transformer for Long Sequence Time Series Forecasting2021in the sky ✦
- Jukebox: A Generative Model for Music2020in the sky ✦
- Large Batch Optimization for Deep Learning: Training BERT in 76 Minutes2019in the sky ✦
- Layer Normalization2016in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- LLaMA: Open and Efficient Foundation Language Models2023in the sky ✦
- LLM.int8(): 8-Bit Matrix Multiplication for Transformers at Scale2022in the sky ✦
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism2019in the sky ✦
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism2019in the sky ✦
- Mistral 7B2023in the sky ✦
- Mixtral of Experts2024in the sky ✦
- Neural Message Passing for Quantum Chemistry2017in the sky ✦
- Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer2022in the sky ✦
- MusicLM: Generating Music From Text2023in the sky ✦
- A Neural Probabilistic Language Model2003in the sky ✦
- Efficient Memory Management for Large Language Model Serving with PagedAttention2023in the sky ✦
- Efficient Memory Management for Large Language Model Serving with PagedAttention2023in the sky ✦
- PaLM: Scaling Language Modeling with Pathways2022in the sky ✦
- A Time Series Is Worth 64 Words: Long-Term Forecasting with Transformers2023in the sky ✦
- A Time Series Is Worth 64 Words: Long-Term Forecasting with Transformers2023in the sky ✦
- Perceiver: General Perception with Iterative Attention2021in the sky ✦
- Perceiver: General Perception with Iterative Attention2021in the sky ✦
- π₀: A Vision-Language-Action Flow Model for General Robot Control2024in the sky ✦
- QLoRA: Efficient Finetuning of Quantized LLMs2023in the sky ✦
- REALM: Retrieval-Augmented Language Model Pre-Training2020in the sky ✦
- Learning Phrase Representations Using RNN Encoder-Decoder for Statistical Machine Translation2014in the sky ✦
- Recurrent Neural Networks (RNNs): A Gentle Introduction and Overview2019in the sky ✦
- Recurrent Neural Networks (RNNs): A Gentle Introduction and Overview2019in the sky ✦
- RoBERTa: A Robustly Optimized BERT Pretraining Approach2019in the sky ✦
- Unsupervised Representation Learning by Predicting Image Rotations2018in the sky ✦
- Unsupervised Representation Learning by Predicting Image Rotations2018in the sky ✦
- RT-1: Robotics Transformer for Real-World Control at Scale2022in the sky ✦
- Efficiently Modeling Long Sequences with Structured State Spaces2022in the sky ✦
- Scaling Laws for Neural Language Models2020in the sky ✦
- SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers2021in the sky ✦
- Sequence to Sequence Learning with Neural Networks2014in the sky ✦
- Shampoo: Preconditioned Stochastic Tensor Optimization2018in the sky ✦
- Show and Tell: A Neural Image Caption Generator2015in the sky ✦
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention2015in the sky ✦
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention2015in the sky ✦
- Spatial Transformer Networks2015in the sky ✦
- Fast Inference from Transformers via Speculative Decoding2023in the sky ✦
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?2024in the sky ✦
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer2019in the sky ✦
- Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions2018in the sky ✦
- A Decoder-Only Foundation Model for Time Series Forecasting2024in the sky ✦
- Attention Is All You Need2017in the sky ✦
- Universal Language Model Fine-Tuning for Text Classification2018in the sky ✦
- An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale2020in the sky ✦
- VQA: Visual Question Answering2015in the sky ✦
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations2020in the sky ✦
- Wide & Deep Learning for Recommender Systems2016in the sky ✦
- XLNet: Generalized Autoregressive Pretraining for Language Understanding2019in the sky ✦