Grouped Query Attention
انتباه الاستعلام المُجمَّع
آلية انتباه وسيطة بين الانتباه متعدد الرؤوس (حيث لكل رأس استعلام زوج مفتاح-قيمة خاص) والانتباه بالاستعلام الأحادي (حيث يتشارك الجميع زوجاً واحداً). تُقسَم رؤوس الاستعلام إلى مجموعات، كل مجموعة تتشارك زوج مفتاح-قيمة واحداً، مما يوازن بين كفاءة الذاكرة والتعبيرية.
An attention mechanism between multihead (each query head has its own KV pair) and multiquery (all share one KV pair). Query heads are split into groups, each group sharing one KV pair, balancing memory efficiency with expressiveness.
Also translated asالانتباه بالاستعلام المُجمَّع، GQA
First appears in this corpus in: The Falcon Series of Open Language Models (2023)
Appears in these papers
- DeepSeek-V3 Technical Report2024in the sky ✦
- The Falcon Series of Open Language Models2023in the sky ✦
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning2023in the sky ✦
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints2023in the sky ✦
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints2023in the sky ✦
- Llama 2: Open Foundation and Fine-Tuned Chat Models2023in the sky ✦
- Llama 2: Open Foundation and Fine-Tuned Chat Models2023in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- Mistral 7B2023in the sky ✦
- Mistral 7B2023in the sky ✦
- Mixtral of Experts2024in the sky ✦
- Mixtral of Experts2024in the sky ✦