CommonCrawl
CommonCrawl
مشروع مفتوح يجمع بيانات الويب بشكل دوري عبر زاحف ويب، وينتج أرشيفات ضخمة تحتوي على مليارات صفحات الإنترنت. المصدر الخام الأساسي لمعظم مجموعات بيانات التدريب اللغوي الكبيرة.
An open project that periodically crawls the web, producing massive archives of billions of web pages. The primary raw source for most large language model training datasets.