Glossary

Proximal Policy Optimization

التحسين التقريبي للسياسة

خوارزمية تعلم معزز واسعة الاستخدام تُحدّث السياسة بخطوات مستقرة عبر قصّ نسبة الاحتمال بين السياسة الجديدة والقديمة، وتتطلب نموذج قيمة منفصلاً.

A widely used RL algorithm that updates the policy in stable steps by clipping the probability ratio between new and old policies, requiring a separate value model.

Also translated asPPO، التحسين الاقترابي للسياسة، التحسين القريب للسياسة

First appears in this corpus in: Policy Gradient Methods for Reinforcement Learning with Function Approximation (1999)

Appears in these papers