Tokenizer
مـُجزئ النصوص
المكون البرمجي المكلف بتقسيم وتحويل النص اللفظي البشري إلى سلسلة من الرموز والأرقام الفهرسية المفهومة للنموذج.
Tokenizer
Also translated asوحدة ترميز الكلمات، خوارزمية فك العبارات إلى رموز رقمية
First appears in this corpus in: Enriching Word Vectors with Subword Information (2017)
Appears in these papers
- BLOOM: A 176B-Parameter Open-Access Multilingual Language Model2022in the sky ✦
- BLOOM: A 176B-Parameter Open-Access Multilingual Language Model2022in the sky ✦
- Evaluating Large Language Models Trained on Code2021in the sky ✦
- The Falcon Series of Open Language Models2023in the sky ✦
- The Falcon Series of Open Language Models2023in the sky ✦
- Enriching Word Vectors with Subword Information2017in the sky ✦
- Finetuned Language Models Are Zero-Shot Learners2022in the sky ✦
- Gemma: Open Models Based on Gemini Research and Technology2024in the sky ✦
- Language Models Are Unsupervised Multitask Learners2019in the sky ✦
- Language Models Are Few-Shot Learners2020in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- LLaMA: Open and Efficient Foundation Language Models2023in the sky ✦
- RoBERTa: A Robustly Optimized BERT Pretraining Approach2019in the sky ✦
- Robust Speech Recognition via Large-Scale Weak Supervision2022in the sky ✦