Glossary

Overestimation Bias

انحياز المبالغة في التقدير

ظاهرة منهجية في Q-learning حيث يؤدي استخدام عامل الحد الأقصى (max) على تقديرات مُشوَّشة إلى المبالغة المتكررة في تقدير قيم الأفعال، مما يُضعف جودة السياسة المُتعلَّمة.

A systematic phenomenon in Q-learning where the max operator over noisy estimates consistently overestimates action values, degrading learned policy quality.

Also translated asتحيُّز الإفراط في التقدير

First appears in this corpus in: Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor (2018)

Appears in these papers