تعلُّم الفروق الزمنية
TD Learning
عائلة خوارزميات تعلُّم معزَّز طوّرها ريتشارد ساتون، تُحدِّث تقدير القيمة بفارق بين تنبؤ لحظي وتنبؤ لاحق (خطأ TD) بدل انتظار المكافأة النهائية للحلقة، مما جمع بين مزايا برمجة دينامية ومونتي كارلو.
A family of reinforcement learning algorithms developed by Richard Sutton that updates the value estimate from the difference between a current prediction and a later prediction (the TD error) rather than waiting for the episode's final return, combining the strengths of dynamic programming and Monte Carlo methods.
تُرجم أيضاًTD Learning، Temporal Difference Learning، التعلم بالفارق الزمني، خوارزميات TD