Glossary

Efficient Attention

الانتباه عالي الكفاءة

تطبيقات لآلية الانتباه تتجنب تخزين مصفوفة الانتباه الكاملة n×n في الذاكرة، مما يُخفِّض استهلاك الذاكرة من O(n²) إلى O(n). يستخدم LLaMA كلاً من الانتباه الموفّر للذاكرة وFlashAttention.

Attention implementations that avoid storing the full n×n attention matrix in memory, reducing memory usage from O(n²) to O(n). LLaMA uses both memory-efficient attention and FlashAttention.

Also translated asالانتباه الفعّال، الانتباه الموفّر للذاكرة

First appears in this corpus in: LLaMA: Open and Efficient Foundation Language Models (2023)

Appears in these papers