Supervised Fine-Tuning
الضبط الدقيق الخاضع للإشراف
مرحلة تدريب يُعاد فيها ضبط نموذج مُدرَّب مسبقاً على مجموعة بيانات مُعلَّمة بشرياً من أزواج مدخلات-مخرجات، لتوجيه سلوكه نحو مهام محددة.
A training stage where a pre-trained model is re-tuned on a human-labeled dataset of input-output pairs to steer its behavior toward specific tasks.
Also translated asSFT، التدريب الإشرافي الدقيق، التدريب الدقيق بالإشراف، التنقيح الإشرافي، الضبط الدقيق الإشرافي، الضبط الموجَّه
First appears in this corpus in: Learning to Summarize from Human Feedback (2020)
Appears in these papers
- Alpaca: A Strong, Replicable Instruction-Following Model2023in the sky ✦
- Alpaca: A Strong, Replicable Instruction-Following Model2023in the sky ✦
- Evaluating Large Language Models Trained on Code2021in the sky ✦
- DeepSeek-V3 Technical Report2024in the sky ✦
- Direct Preference Optimization: Your Language Model Is Secretly a Reward Model2023in the sky ✦
- The Falcon Series of Open Language Models2023in the sky ✦
- A Generalist Agent2022in the sky ✦
- Gemini: A Family of Highly Capable Multimodal Models2023in the sky ✦
- Gemma: Open Models Based on Gemini Research and Technology2024in the sky ✦
- Gemma: Open Models Based on Gemini Research and Technology2024in the sky ✦
- GPT-4 Technical Report2023in the sky ✦
- Training Language Models to Follow Instructions with Human Feedback2022in the sky ✦
- Training Language Models to Follow Instructions with Human Feedback2022in the sky ✦
- Learning to Summarize from Human Feedback2020in the sky ✦
- Llama 2: Open Foundation and Fine-Tuned Chat Models2023in the sky ✦
- Llama 2: Open Foundation and Fine-Tuned Chat Models2023in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- Mixtral of Experts2024in the sky ✦
- RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback2023in the sky ✦
- RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback2023in the sky ✦