Glossary

TD Error

خطأ الفارق الزمني

الفرق بين التنبؤ الحالي والهدف المُستمَد من الخطوة التالية: δₜ = rₜ₊₁ + γV(sₜ₊₁) − V(sₜ). هذه الإشارة الأساسية تقود التعلم في جميع خوارزميات التعلم المعزز القائمة على القيمة.

The difference between the current prediction and the target derived from the next step: δₜ = rₜ₊₁ + γV(sₜ₊₁) − V(sₜ). This fundamental signal drives learning in all value-based reinforcement learning algorithms.

Also translated asإشارة الخطأ الزمني، فارق التنبؤ المتعاقب

First appears in this corpus in: Learning to Predict by the Methods of Temporal Differences (1988)

Appears in these papers