Deterministic Policy Gradient
تدرّج السياسة الحتمية
نظرية تُثبت أنه يمكن حساب تدرج دالة الهدف لسياسة حتمية بالتكامل على فضاء الحالات فقط (بدلاً من فضاء الحالات والأفعال معاً كما في السياسات العشوائية)، مما يُقلّل التباين ويُسرّع التقارب.
A theorem proving that the objective gradient for a deterministic policy can be computed by integrating over the state space alone (instead of both state and action spaces as in stochastic policies), reducing variance and speeding convergence.
Also translated asDPG
First appears in this corpus in: Continuous Control with Deep Reinforcement Learning (2015)
Appears in these papers