Overestimation Bias
انحياز المبالغة في التقدير
ظاهرة منهجية في Q-learning حيث يؤدي استخدام عامل الحد الأقصى (max) على تقديرات مُشوَّشة إلى المبالغة المتكررة في تقدير قيم الأفعال، مما يُضعف جودة السياسة المُتعلَّمة.
A systematic phenomenon in Q-learning where the max operator over noisy estimates consistently overestimates action values, degrading learned policy quality.
Also translated asتحيُّز الإفراط في التقدير
First appears in this corpus in: Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor (2018)
Appears in these papers
- Conservative Q-Learning for Offline Reinforcement Learning2020in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- TD3: Addressing Function Approximation Error in Actor-Critic Methods2018in the sky ✦
- TD3: Addressing Function Approximation Error in Actor-Critic Methods2018in the sky ✦