Efficient Attention
الانتباه عالي الكفاءة
تطبيقات لآلية الانتباه تتجنب تخزين مصفوفة الانتباه الكاملة n×n في الذاكرة، مما يُخفِّض استهلاك الذاكرة من O(n²) إلى O(n). يستخدم LLaMA كلاً من الانتباه الموفّر للذاكرة وFlashAttention.
Attention implementations that avoid storing the full n×n attention matrix in memory, reducing memory usage from O(n²) to O(n). LLaMA uses both memory-efficient attention and FlashAttention.
Also translated asالانتباه الفعّال، الانتباه الموفّر للذاكرة
First appears in this corpus in: LLaMA: Open and Efficient Foundation Language Models (2023)
Appears in these papers