Double Q-Learning
Q-Learning المزدوج
تعديل على Q-Learning يحتفظ بجدولي Q مستقلين: أحدهما يختار الفعل الأفضل والآخر يُقيّم قيمته، مما يقضي على انحياز التعظيم الناتج عن استخدام نفس الجدول للاختيار والتقييم.
A modification of Q-Learning that maintains two independent Q-tables: one selects the best action, the other evaluates its value, eliminating the maximization bias from using the same table for both selection and evaluation.
Also translated asالتعلم المزدوج لقيم الجودة
First appears in this corpus in: Q-Learning (1992)
Appears in these papers
- Deep Reinforcement Learning with Double Q-Learning2016in the sky ✦
- Deep Reinforcement Learning with Double Q-Learning2016in the sky ✦
- Dueling Network Architectures for Deep Reinforcement Learning2016in the sky ✦
- Prioritized Experience Replay2015in the sky ✦
- Q-Learning1992in the sky ✦
- Rainbow: Combining Improvements in Deep Reinforcement Learning2018in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- TD3: Addressing Function Approximation Error in Actor-Critic Methods2018in the sky ✦