ديب سيك آر1
DeepSeek-R1
نموذج لغوي كبير من شركة ديب سيك متخصّص بالاستدلال، دُرِّب أساساً بالتعلم المعزَّز الخالص (بخوارزمية GRPO) بدل الضبط الدقيق المُشرَف، فبرزت لديه سلوكيات مثل التأمل الذاتي والتحقق تلقائياً.
A reasoning-focused large language model from DeepSeek trained primarily with pure reinforcement learning (using the GRPO algorithm) rather than supervised fine-tuning, causing behaviors such as self-reflection and verification to emerge on their own.
تُرجم أيضاًDeepSeek-R1، نموذج ديب سيك للتفكير المنطقي