Cross-Attention
الانتباه التبادلي
آلية انتباه تربط بين نمطين مختلفين من البيانات: الاستعلامات تأتي من نمط (مثل سمات الصورة) والمفاتيح والقيم من نمط آخر (مثل تضمينات النص)، مما يُتيح لكل موقع مكاني أن يتعلّم أي رموز نصية تهمّه.
An attention mechanism connecting two different data modalities: queries come from one modality (e.g. image features) and keys/values from another (e.g. text embeddings), allowing each spatial location to learn which text tokens are relevant to it.
Also translated asآلية الربط بين سياقين، الانتباه العابر للنمط، الانتباه العابر للوسائط، الانتباه المتبادل، الانتباه المتقاطع، الانتباه عبر المصادر، التركيز البيني للمصفوفات، انتباه المرمِّز-فاكّ الترميز
First appears in this corpus in: Improving Language Understanding by Generative Pre-Training (2018)
Appears in these papers
- ALIGN: Scaling Up Visual and Vision-Language Representation Learning with Noisy Text Supervision2021in the sky ✦
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware2023in the sky ✦
- Atlas: Few-shot Learning with Retrieval Augmented Language Models2023in the sky ✦
- Atlas: Few-shot Learning with Retrieval Augmented Language Models2023in the sky ✦
- BART: Denoising Sequence-to-Sequence Pre-Training for Natural Language Generation, Translation, and Comprehension2019in the sky ✦
- BLIP-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models2023in the sky ✦
- BLIP-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models2023in the sky ✦
- DALL·E: Zero-Shot Text-to-Image Generation2021in the sky ✦
- End-to-End Object Detection with Transformers2020in the sky ✦
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion2023in the sky ✦
- Fusion-in-Decoder: Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering2021in the sky ✦
- Fusion-in-Decoder: Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering2021in the sky ✦
- Flamingo: a Visual Language Model for Few-Shot Learning2022in the sky ✦
- Flamingo: a Visual Language Model for Few-Shot Learning2022in the sky ✦
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models2021in the sky ✦
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models2021in the sky ✦
- Improving Language Understanding by Generative Pre-Training2018in the sky ✦
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints2023in the sky ✦
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints2023in the sky ✦
- Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection2023in the sky ✦
- ImageBind: One Embedding Space to Bind Them All2023in the sky ✦
- Imagen: Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding2022in the sky ✦
- Imagen: Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding2022in the sky ✦
- Informer: Beyond Efficient Transformer for Long Sequence Time Series Forecasting2021in the sky ✦
- Jukebox: A Generative Model for Music2020in the sky ✦
- High-Resolution Image Synthesis with Latent Diffusion Models2022in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- Visual Instruction Tuning2023in the sky ✦
- Perceiver: General Perception with Iterative Attention2021in the sky ✦
- Perceiver: General Perception with Iterative Attention2021in the sky ✦
- RETRO: Improving Language Models by Retrieving from Trillions of Tokens2022in the sky ✦
- Segment Anything2023in the sky ✦
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer2019in the sky ✦
- Robust Speech Recognition via Large-Scale Weak Supervision2022in the sky ✦