Self-Attention
الانتباه الذاتي
آلية انتباه تحسب فيها كل عنصر في التسلسل درجات الصلة مع جميع العناصر الأخرى في التسلسل ذاته، مما يتيح تفاعلات مباشرة بين كل زوج.
An attention mechanism where each element in a sequence computes relevance scores against all other elements in the same sequence, enabling direct pairwise interactions.
Also translated asآلية الانتباه الذاتي، آلية رصد العلاقات البينية، الانتباه الداخلي، مقياس الارتباط الداخلي للرموز
First appears in this corpus in: ELIZA — A Computer Program for the Study of Natural Language Communication (1966)
Appears in these papers
- Parameter-Efficient Transfer Learning for NLP2019in the sky ✦
- Highly Accurate Protein Structure Prediction with AlphaFold2021in the sky ✦
- Atlas: Few-shot Learning with Retrieval Augmented Language Models2023in the sky ✦
- Neural Machine Translation by Jointly Learning to Align and Translate2014in the sky ✦
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting2021in the sky ✦
- Learning Long-Term Dependencies with Gradient Descent is Difficult1994in the sky ✦
- BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer2019in the sky ✦
- BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer2019in the sky ✦
- BLIP-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models2023in the sky ✦
- BLIP-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models2023in the sky ✦
- Chronos: Learning the Language of Time Series2024in the sky ✦
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT2020in the sky ✦
- A ConvNet for the 2020s2022in the sky ✦
- DALL·E: Zero-Shot Text-to-Image Generation2021in the sky ✦
- DALL·E: Zero-Shot Text-to-Image Generation2021in the sky ✦
- Denoising Diffusion Probabilistic Models2020in the sky ✦
- DeBERTa: Decoding-Enhanced BERT with Disentangled Attention2020in the sky ✦
- DeBERTa: Decoding-Enhanced BERT with Disentangled Attention2020in the sky ✦
- Decision Transformer: Reinforcement Learning via Sequence Modeling2021in the sky ✦
- Decision Transformer: Reinforcement Learning via Sequence Modeling2021in the sky ✦
- Training Data-Efficient Image Transformers & Distillation Through Attention2021in the sky ✦
- Training Data-Efficient Image Transformers & Distillation Through Attention2021in the sky ✦
- End-to-End Object Detection with Transformers2020in the sky ✦
- Emerging Properties in Self-Supervised Vision Transformers2021in the sky ✦
- DINOv2: Learning Robust Visual Features Without Supervision2023in the sky ✦
- DistilBERT, a Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter2019in the sky ✦
- Are Transformers Effective for Time Series Forecasting?2023in the sky ✦
- Are Transformers Effective for Time Series Forecasting?2023in the sky ✦
- ELIZA — A Computer Program for the Study of Natural Language Communication1966in the sky ✦
- Fusion-in-Decoder: Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering2021in the sky ✦
- Flamingo: a Visual Language Model for Few-Shot Learning2022in the sky ✦
- Semi-Supervised Classification with Graph Convolutional Networks2017in the sky ✦
- Improving Language Understanding by Generative Pre-Training2018in the sky ✦
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints2023in the sky ✦
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints2023in the sky ✦
- Do Transformers Really Perform Bad for Graph Representation?2021in the sky ✦
- Do Transformers Really Perform Bad for Graph Representation?2021in the sky ✦
- Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection2023in the sky ✦
- Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection2023in the sky ✦
- Informer: Beyond Efficient Transformer for Long Sequence Time Series Forecasting2021in the sky ✦
- Informer: Beyond Efficient Transformer for Long Sequence Time Series Forecasting2021in the sky ✦
- High-Resolution Image Synthesis with Latent Diffusion Models2022in the sky ✦
- Visual Instruction Tuning2023in the sky ✦
- LoRA: Low-Rank Adaptation of Large Language Models2021in the sky ✦
- Long Short-Term Memory1997in the sky ✦
- Masked Autoencoders Are Scalable Vision Learners2022in the sky ✦
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces2023in the sky ✦
- Mistral 7B2023in the sky ✦
- Mistral 7B2023in the sky ✦
- A Time Series Is Worth 64 Words: Long-Term Forecasting with Transformers2023in the sky ✦
- Perceiver: General Perception with Iterative Attention2021in the sky ✦
- Perceiver: General Perception with Iterative Attention2021in the sky ✦
- Recurrent Neural Networks (RNNs): A Gentle Introduction and Overview2019in the sky ✦
- RoFormer: Enhanced Transformer with Rotary Position Embedding2021in the sky ✦
- RT-1: Robotics Transformer for Real-World Control at Scale2022in the sky ✦
- Scaling Laws for Neural Language Models2020in the sky ✦
- SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers2021in the sky ✦
- Segment Anything2023in the sky ✦
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention2015in the sky ✦
- Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows2021in the sky ✦
- Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows2021in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦
- A Decoder-Only Foundation Model for Time Series Forecasting2024in the sky ✦
- A Decoder-Only Foundation Model for Time Series Forecasting2024in the sky ✦
- Attention Is All You Need2017in the sky ✦
- Attention Is All You Need2017in the sky ✦
- An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale2020in the sky ✦
- XLNet: Generalized Autoregressive Pretraining for Language Understanding2019in the sky ✦