تدرّج السياسة الحتمية
Deterministic Policy Gradient
نظرية تُثبت أنه يمكن حساب تدرج دالة الهدف لسياسة حتمية بالتكامل على فضاء الحالات فقط (بدلاً من فضاء الحالات والأفعال معاً كما في السياسات العشوائية)، مما يُقلّل التباين ويُسرّع التقارب.
A theorem proving that the objective gradient for a deterministic policy can be computed by integrating over the state space alone (instead of both state and action spaces as in stochastic policies), reducing variance and speeding convergence.
تُرجم أيضاًDPG
أول ظهور في هذه المجموعة: التحكم المستمر بالتعلُّم المعزِّز العميق (2015)
يظهر في هذه الأوراق