FlashAttention
FlashAttention
خوارزمية واعية بنمط الوصول للذاكرة تحسب الانتباه الدقيق دون تجسيد مصفوفة الانتباه الكاملة n×n في ذاكرة المعالج الرسومي. تُخفّض استهلاك الذاكرة من O(n²) إلى O(n) وتُسرّع زمن التنفيذ. استُخدمت في تدريب LLaMA.
An IO-aware algorithm that computes exact attention without materializing the full n×n attention matrix in GPU memory. Reduces memory usage from O(n²) to O(n) and speeds up wall-clock time. Used in LLaMA training.
Also translated asالانتباه السريع، انتباه فلاش
First appears in this corpus in: FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning (2023)
Appears in these papers
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning2023in the sky ✦
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints2023in the sky ✦
- LLaMA: Open and Efficient Foundation Language Models2023in the sky ✦
- LLaMA: Open and Efficient Foundation Language Models2023in the sky ✦