Deterministic Policy
السياسة الحتمية
سياسة تُخرج فعلاً واحداً محدداً لكل حالة بدلاً من توزيع احتمالي على الأفعال. الممثل يُعيد مباشرةً القيمة الدقيقة للفعل الأفضل، مما يبسّط حساب التدرج لكنه يتطلب آلية استكشاف خارجية.
A policy that outputs a single specific action for each state rather than a probability distribution over actions. The actor directly returns the exact best action value, simplifying gradient computation but requiring an external exploration mechanism.
Also translated asالسياسة القطعية، السياسة المحدَّدة
First appears in this corpus in: Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor (2018)
Appears in these papers