Glossary

Temporal Difference Learning

التعلُّم بالفرق الزمني

عائلة خوارزميات في التعلّم المعزّز تُحدِّث تقديرات دالة القيمة بناءً على الفرق بين تنبؤات متعاقبة بدل انتظار النتيجة النهائية. تجمع بين فكرة التمهيد من البرمجة الديناميكية وأخذ العيّنات من أساليب مونت كارلو.

A family of reinforcement learning algorithms that update value-function estimates based on the difference between successive predictions rather than waiting for the final outcome. Combines the bootstrapping idea from dynamic programming with the sampling approach of Monte Carlo methods.

Also translated asتعلّم الفرق الزمني، أسلوب الفرق الزمني

First appears in this corpus in: TD3: Addressing Function Approximation Error in Actor-Critic Methods (2018)

Appears in these papers