Dueling Network
الشبكة الثنائية
بنية شبكة عصبية تفصل تقدير قيمة الحالة V(s) عن تقدير ميزة الفعل A(s,a) في تيارين مستقلين، ثم تدمجهما لحساب Q(s,a). هذا يسمح بتعلُّم قيم الحالات حتى من تجارب لا يؤثر فيها اختيار الفعل.
A neural network architecture that separates state-value estimation V(s) from action-advantage estimation A(s,a) into two independent streams, then combines them to compute Q(s,a). This allows learning state values even from experiences where action choice did not matter.
Also translated asشبكة ثنائية التيار، البنية الثنائية
First appears in this corpus in: Rainbow: Combining Improvements in Deep Reinforcement Learning (2018)
Appears in these papers