On-Policy
خوارزمية التعلم من السياسة الحالية
خوارزميات تعزيز تتعلم وتحدث استراتيجيتها بالاعتماد الحصري على الأفعال المباشرة التي يتخذها العميل حالياً في البيئة.
On-Policy
Also translated asالتعلم التدعيمي الملتزم باستراتيجية العمل، مواءمة الأفعال المباشرة الحية
First appears in this corpus in: Policy Gradient Methods for Reinforcement Learning with Function Approximation (1999)
Appears in these papers
- Asynchronous Methods for Deep Reinforcement Learning2016in the sky ✦
- Asynchronous Methods for Deep Reinforcement Learning2016in the sky ✦
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures2018in the sky ✦
- Learning Dexterous In-Hand Manipulation2019in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦