Proximal Policy Optimization (PPO)
خوارزمية تحسين السياسة القريبة
خوارزمية تعزيز متطورة تضمن تحديث استراتيجية الأفعال بنسب آمنة وتدريجية لمنع انهيار الأداء التدريبي.
Proximal Policy Optimization (PPO)
Also translated asمحسن استراتيجيات الأفعال المحدود التغير، معيار التعديل الآمن والتدريجي للقرارات
First appears in this corpus in: Asynchronous Methods for Deep Reinforcement Learning (2016)
Appears in these papers
- Asynchronous Methods for Deep Reinforcement Learning2016in the sky ✦
- DeepSeek-V3 Technical Report2024in the sky ✦
- High-Dimensional Continuous Control Using Generalized Advantage Estimation2016in the sky ✦
- Learning Dexterous In-Hand Manipulation2019in the sky ✦
- Learning to Summarize from Human Feedback2020in the sky ✦
- Llama 2: Open Foundation and Fine-Tuned Chat Models2023in the sky ✦
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection2023in the sky ✦