Kernel Fusion
دمج النواة
تقنية تحسين للمعالجات الرسومية تجمع عدة عمليات متتالية في نواة GPU واحدة لتقليل عبء القراءة والكتابة من الذاكرة. مامبا تدمج التمييز والمسح وضرب المُخرَج في نواة واحدة.
A GPU optimization technique that combines multiple sequential operations into a single GPU kernel to reduce memory read/write overhead. Mamba fuses discretization, scan, and output multiplication into one kernel.
Also translated asتوحيد نوى المعالج، دمج عمليات النواة، دمج نواة المعالج، صهر النوى
First appears in this corpus in: FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (2022)
Appears in these papers
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning2023in the sky ✦
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness2022in the sky ✦
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness2022in the sky ✦
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces2023in the sky ✦