KV Cache
ذاكرة المفاتيح والقيم
في المحوِّلات، تُخزَّن مصفوفات المفاتيح والقيم لكل الرموز السابقة أثناء الاستدلال الارتجاعي. هذه الذاكرة تنمو خطياً مع طول التسلسل، مستهلكة ذاكرة المعالج الرسومي ومبطئة عملية التوليد. مامبا تتجنّبها كلياً بحالة مخفية ثابتة الحجم.
In Transformers, the Key and Value matrices from all previous tokens are stored during autoregressive inference. This cache grows linearly with sequence length, consuming GPU memory and slowing generation. Mamba avoids this entirely with a fixed-size hidden state.
Also translated asالكاش الحوسبي لتسريع التوليد النصي، ذاكرة KV المؤقتة، ذاكرة استرجاع الحسابات السياقية السابقة، مخبأة المفاتيح والقيم
First appears in this corpus in: GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints (2023)
Appears in these papers
- DeepSeek-V3 Technical Report2024in the sky ✦
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints2023in the sky ✦
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints2023in the sky ✦
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces2023in the sky ✦
- Mistral 7B2023in the sky ✦
- Mistral 7B2023in the sky ✦
- Efficient Memory Management for Large Language Model Serving with PagedAttention2023in the sky ✦
- Efficient Memory Management for Large Language Model Serving with PagedAttention2023in the sky ✦
- Fast Inference from Transformers via Speculative Decoding2023in the sky ✦