Model Compression
ضغط النماذج
مجموعة تقنيات لتقليل حجم النموذج و/أو تكلفته الحوسبية مع الحفاظ على أكبر قدر من الأداء، تشمل التقطير والتشذيب والتكميم.
A family of techniques to reduce model size and/or computational cost while preserving as much performance as possible, including distillation, pruning, and quantization.
Also translated asتصغير النماذج، اختزال النماذج
First appears in this corpus in: Optimal Brain Damage (1989)
Appears in these papers
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding2016in the sky ✦
- DistilBERT, a Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter2019in the sky ✦
- DistilBERT, a Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter2019in the sky ✦
- GPTQ: Accurate Post-Training Quantization for Generative Pre-Trained Transformers2022in the sky ✦
- GPTQ: Accurate Post-Training Quantization for Generative Pre-Trained Transformers2022in the sky ✦
- Distilling the Knowledge in a Neural Network2015in the sky ✦
- LLM.int8(): 8-Bit Matrix Multiplication for Transformers at Scale2022in the sky ✦
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks2019in the sky ✦
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications2017in the sky ✦
- Optimal Brain Damage1989in the sky ✦
- Optimal Brain Damage1989in the sky ✦