TD Error
خطأ الفارق الزمني
الفرق بين التنبؤ الحالي والهدف المُستمَد من الخطوة التالية: δₜ = rₜ₊₁ + γV(sₜ₊₁) − V(sₜ). هذه الإشارة الأساسية تقود التعلم في جميع خوارزميات التعلم المعزز القائمة على القيمة.
The difference between the current prediction and the target derived from the next step: δₜ = rₜ₊₁ + γV(sₜ₊₁) − V(sₜ). This fundamental signal drives learning in all value-based reinforcement learning algorithms.
Also translated asإشارة الخطأ الزمني، فارق التنبؤ المتعاقب
First appears in this corpus in: Learning to Predict by the Methods of Temporal Differences (1988)
Appears in these papers
- Prioritized Experience Replay2015in the sky ✦
- Prioritized Experience Replay2015in the sky ✦
- Rainbow: Combining Improvements in Deep Reinforcement Learning2018in the sky ✦
- Rainbow: Combining Improvements in Deep Reinforcement Learning2018in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- Temporal Difference Learning and TD-Gammon1995in the sky ✦
- Temporal Difference Learning and TD-Gammon1995in the sky ✦
- Learning to Predict by the Methods of Temporal Differences1988in the sky ✦
- Learning to Predict by the Methods of Temporal Differences1988in the sky ✦