الانتباه المُقسَّم إلى صفحات
PagedAttention
آلية إدارة ذاكرة في محرِّك vLLM مستوحاة من الذاكرة الوهمية في نظم التشغيل، تُخزِّن ذاكرة المفاتيح والقيم للنماذج اللغوية الكبيرة في صفحات صغيرة غير متجاورة، فتُقلِّل التجزئة وتُحسِّن استغلال الذاكرة وتُتيح إنتاجية خدمية أعلى بكثير من الانتباه التقليدي.
A memory-management mechanism at the heart of the vLLM engine, inspired by virtual memory in operating systems, that stores the key–value cache of large language models in small non-contiguous pages, reducing fragmentation, improving memory utilization, and enabling far higher serving throughput than conventional attention.
تُرجم أيضاًPagedAttention، Paged Attention، انتباه ذاكرة الصفحات، الانتباه المُصفَّح