المعجم

الانتباه الكامن متعدد الرؤوس

Multi-Head Latent Attention

آلية انتباه تضغط جميع متجهات المفتاح والقيمة في متجه كامن مضغوط وحيد لكل رمز، ثم تُعيد بناء المفاتيح والقيم الكاملة لحظياً عبر مصفوفات إسقاط عكسي. تُقلّص ذاكرة KV التخزينية بعشرات المرات مقارنة بالانتباه متعدد الرؤوس المعياري.

An attention mechanism that compresses all key and value vectors into a single compact latent vector per token, then reconstructs full keys and values on the fly via up-projection matrices. Reduces KV cache memory by tens of times compared to standard Multi-Head Attention.

تُرجم أيضاًالانتباه الكامن، MLA