Distillation
التقطير
عملية نقل قدرات نموذج كبير (المعلم) إلى نموذج أصغر (الطالب) بتدريب الأصغر على مخرجات الأكبر، مما يحافظ على معظم الأداء بموارد أقل.
The process of transferring a large model's (teacher) capabilities to a smaller model (student) by training the smaller one on the larger one's outputs, preserving most performance at lower cost.
Also translated asتقطير المعرفة، تكثيف واختزال القدرات الحوسبية، نقل المعرفة، نقل المعرفة من نموذج ضخم لأصغر
First appears in this corpus in: DistilBERT, a Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter (2019)
Appears in these papers
- Consistency Models2023in the sky ✦
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning2025in the sky ✦
- Training Data-Efficient Image Transformers & Distillation Through Attention2021in the sky ✦
- DistilBERT, a Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter2019in the sky ✦
- Informer: Beyond Efficient Transformer for Long Sequence Time Series Forecasting2021in the sky ✦
- Informer: Beyond Efficient Transformer for Long Sequence Time Series Forecasting2021in the sky ✦
- Fast Inference from Transformers via Speculative Decoding2023in the sky ✦