Proximal Policy Optimization
التحسين التقريبي للسياسة
خوارزمية تعلم معزز واسعة الاستخدام تُحدّث السياسة بخطوات مستقرة عبر قصّ نسبة الاحتمال بين السياسة الجديدة والقديمة، وتتطلب نموذج قيمة منفصلاً.
A widely used RL algorithm that updates the policy in stable steps by clipping the probability ratio between new and old policies, requiring a separate value model.
Also translated asPPO، التحسين الاقترابي للسياسة، التحسين القريب للسياسة
First appears in this corpus in: Policy Gradient Methods for Reinforcement Learning with Function Approximation (1999)
Appears in these papers
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning2025in the sky ✦
- Direct Preference Optimization: Your Language Model Is Secretly a Reward Model2023in the sky ✦
- Training Language Models to Follow Instructions with Human Feedback2022in the sky ✦
- Training Language Models to Follow Instructions with Human Feedback2022in the sky ✦
- Learning Dexterous In-Hand Manipulation2019in the sky ✦
- Learning to Summarize from Human Feedback2020in the sky ✦
- Llama 2: Open Foundation and Fine-Tuned Chat Models2023in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Proximal Policy Optimization Algorithms2017in the sky ✦
- WebGPT: Browser-Assisted Question-Answering with Human Feedback2021in the sky ✦