Glossary

Double Quantization

التكميم المزدوج

تقنية تُكمِّم ثوابت تقييس التكميم ذاتها (من FP32 إلى 8-بت) لتقليص العبء الذاكري الإضافي، مما يوفّر نحو 0.37 بت لكل مُعامِل أي قرابة 3 غيغابايت لنموذج بـ 65 مليار مُعامِل.

A technique that quantizes the quantization scaling constants themselves (from FP32 to 8-bit) to reduce memory overhead, saving roughly 0.37 bits per parameter — about 3 GB for a 65B model.

Also translated asالتكميم المضاعف، تكميم من مستويين

First appears in this corpus in: QLoRA: Efficient Finetuning of Quantized LLMs (2023)

Appears in these papers