Glossary

Multi-Head Latent Attention

الانتباه الكامن متعدد الرؤوس

آلية انتباه تضغط جميع متجهات المفتاح والقيمة في متجه كامن مضغوط وحيد لكل رمز، ثم تُعيد بناء المفاتيح والقيم الكاملة لحظياً عبر مصفوفات إسقاط عكسي. تُقلّص ذاكرة KV التخزينية بعشرات المرات مقارنة بالانتباه متعدد الرؤوس المعياري.

An attention mechanism that compresses all key and value vectors into a single compact latent vector per token, then reconstructs full keys and values on the fly via up-projection matrices. Reduces KV cache memory by tens of times compared to standard Multi-Head Attention.

Also translated asالانتباه الكامن، MLA