Pre-training
التدريب المسبق
مرحلة التدريب الأولية واسعة النطاق حيث يتعلم النموذج من مجموعات بيانات ضخمة للتنبؤ بالرمز التالي، مكتسباً معرفة واسعة باللغة والاستدلال.
The initial large-scale training phase where a model learns from massive datasets to predict the next token, acquiring broad knowledge of language and reasoning.
Also translated asالتدريب الأولي، مرحلة ما قبل التدريب
First appears in this corpus in: Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation (2014)
Appears in these papers
- Parameter-Efficient Transfer Learning for NLP2019in the sky ✦
- BART: Denoising Sequence-to-Sequence Pre-Training for Natural Language Generation, Translation, and Comprehension2019in the sky ✦
- BART: Denoising Sequence-to-Sequence Pre-Training for Natural Language Generation, Translation, and Comprehension2019in the sky ✦
- BLIP-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models2023in the sky ✦
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT2020in the sky ✦
- Contriever: Unsupervised Dense Information Retrieval with Contrastive Learning2022in the sky ✦
- A ConvNet for the 2020s2022in the sky ✦
- DeBERTa: Decoding-Enhanced BERT with Disentangled Attention2020in the sky ✦
- DeepSeek-V3 Technical Report2024in the sky ✦
- DistilBERT, a Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter2019in the sky ✦
- ELECTRA: Pre-Training Text Encoders as Discriminators Rather Than Generators2020in the sky ✦
- Emergent Abilities of Large Language Models2022in the sky ✦
- Gemma: Open Models Based on Gemini Research and Technology2024in the sky ✦
- Gemma: Open Models Based on Gemini Research and Technology2024in the sky ✦
- GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding2018in the sky ✦
- GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding2018in the sky ✦
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints2023in the sky ✦
- Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection2023in the sky ✦
- Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection2023in the sky ✦
- Learning to Summarize from Human Feedback2020in the sky ✦
- Symbolic Discovery of Optimization Algorithms2023in the sky ✦
- Llama 2: Open Foundation and Fine-Tuned Chat Models2023in the sky ✦
- Llama 2: Open Foundation and Fine-Tuned Chat Models2023in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- Mixtral of Experts2024in the sky ✦
- Neural Collaborative Filtering2017in the sky ✦
- π₀: A Vision-Language-Action Flow Model for General Robot Control2024in the sky ✦
- π₀: A Vision-Language-Action Flow Model for General Robot Control2024in the sky ✦
- Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation2014in the sky ✦
- Self-Instruct: Aligning Language Models with Self-Generated Instructions2022in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦
- End-to-End Training of Deep Visuomotor Policies2016in the sky ✦
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations2020in the sky ✦
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations2020in the sky ✦
- XLNet: Generalized Autoregressive Pretraining for Language Understanding2019in the sky ✦