Multi-Query Attention
انتباه الاستعلام المتعدّد
تعديل على الانتباه متعدّد الرؤوس حيث تتشارك جميع الرؤوس إسقاط مفتاح وقيمة واحداً بينما يحتفظ كل رأس بإسقاط استعلام خاصّ به. يُقلّل الذاكرة والحوسبة أثناء التوليد الانحداري بشكل كبير.
A modification of multi-head attention where all heads share one key and one value projection while each head retains its own query projection. Dramatically reduces memory and compute during autoregressive generation.
Also translated asالانتباه متعدّد الاستعلامات، MQA
First appears in this corpus in: PaLM: Scaling Language Modeling with Pathways (2022)
Appears in these papers
- Gemini: A Family of Highly Capable Multimodal Models2023in the sky ✦
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints2023in the sky ✦
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints2023in the sky ✦
- PaLM: Scaling Language Modeling with Pathways2022in the sky ✦
- PaLM: Scaling Language Modeling with Pathways2022in the sky ✦