Clipped Double Q-Learning
تعلُّم Q المزدوج المقصوص
تمديد لتعلُّم Q المزدوج يستخدم شبكتي هدف ويأخذ الأدنى من تقديراتهما لقيمة الحالة التالية، مما يكبح انحياز المبالغة في التقدير الذي يُصيب تعلُّم Q العادي.
An extension of Double Q-learning that uses two target networks and takes the minimum of their value estimates for the next state, suppressing the overestimation bias that plagues standard Q-learning.
Also translated asQ-Learning المزدوج بالقصّ
First appears in this corpus in: QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation (2018)
Appears in these papers