Knowledge Distillation
تقطير المعرفة
تقنية ضغط نماذج قدّمها هينتون وفينيالز ودين، يُدرَّب فيها نموذج «طالب» صغير على محاكاة التوزيعات الاحتمالية اللينة لنموذج «معلِّم» أكبر بدل التسميات الصلبة، فينقل جزءاً كبيراً من دقة المعلم بحجم أصغر بكثير.
A model compression technique introduced by Hinton, Vinyals, and Dean in which a small 'student' network is trained to mimic the soft probability distributions of a larger 'teacher' network rather than the hard labels, transferring much of the teacher's accuracy at a fraction of the size.
Also translated asKnowledge Distillation، التقطير المعرفي، تقطير النماذج، تعليم الطالب من المعلِّم
First appears in this corpus in: Distilling the Knowledge in a Neural Network (2015)
Appears in these papers
- Alpaca: A Strong, Replicable Instruction-Following Model2023in the sky ✦
- Alpaca: A Strong, Replicable Instruction-Following Model2023in the sky ✦
- Atlas: Few-shot Learning with Retrieval Augmented Language Models2023in the sky ✦
- Bootstrap Your Own Latent: A New Approach to Self-Supervised Learning2020in the sky ✦
- Consistency Models2023in the sky ✦
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding2016in the sky ✦
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning2025in the sky ✦
- DeepSeek-V3 Technical Report2024in the sky ✦
- Training Data-Efficient Image Transformers & Distillation Through Attention2021in the sky ✦
- Training Data-Efficient Image Transformers & Distillation Through Attention2021in the sky ✦
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data2024in the sky ✦
- Emerging Properties in Self-Supervised Vision Transformers2021in the sky ✦
- Emerging Properties in Self-Supervised Vision Transformers2021in the sky ✦
- DINOv2: Learning Robust Visual Features Without Supervision2023in the sky ✦
- DINOv2: Learning Robust Visual Features Without Supervision2023in the sky ✦
- DistilBERT, a Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter2019in the sky ✦
- DistilBERT, a Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter2019in the sky ✦
- Gemini: A Family of Highly Capable Multimodal Models2023in the sky ✦
- Gemini: A Family of Highly Capable Multimodal Models2023in the sky ✦
- Gemma: Open Models Based on Gemini Research and Technology2024in the sky ✦
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher2022in the sky ✦
- HyDE: Precise Zero-Shot Dense Retrieval without Relevance Labels2022in the sky ✦
- Distilling the Knowledge in a Neural Network2015in the sky ✦
- Distilling the Knowledge in a Neural Network2015in the sky ✦
- When Does Label Smoothing Help?2019in the sky ✦
- When Does Label Smoothing Help?2019in the sky ✦
- Llama 2: Open Foundation and Fine-Tuned Chat Models2023in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- LLM.int8(): 8-Bit Matrix Multiplication for Transformers at Scale2022in the sky ✦
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications2017in the sky ✦
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications2017in the sky ✦
- MusicLM: Generating Music From Text2023in the sky ✦
- PaLM: Scaling Language Modeling with Pathways2022in the sky ✦
- Textbooks Are All You Need2023in the sky ✦
- RETRO: Improving Language Models by Retrieving from Trillions of Tokens2022in the sky ✦
- RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback2023in the sky ✦
- RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback2023in the sky ✦
- Self-Instruct: Aligning Language Models with Self-Generated Instructions2022in the sky ✦
- Self-Instruct: Aligning Language Models with Self-Generated Instructions2022in the sky ✦
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection2023in the sky ✦
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection2023in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision2023in the sky ✦