Entropy Bonus
مكافأة العشوائية الدلالية
حدّ يُضاف لدالة الهدف يكافئ السياسة على الاحتفاظ بتوزيع احتمالي منتشر بدل التركيز على فعل واحد، مما يشجع الاستكشاف ويمنع الانهيار المبكر.
A term added to the objective that rewards the policy for maintaining a spread-out probability distribution rather than concentrating on one action, encouraging exploration and preventing premature convergence.
Also translated asحافز الإنتروبيا، مكافأة التنوع
First appears in this corpus in: Asynchronous Methods for Deep Reinforcement Learning (2016)
Appears in these papers
- Asynchronous Methods for Deep Reinforcement Learning2016in the sky ✦
- Asynchronous Methods for Deep Reinforcement Learning2016in the sky ✦
- First Return, Then Explore2021in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦