Temporal Difference Learning
التعلُّم بالفرق الزمني
عائلة خوارزميات في التعلّم المعزّز تُحدِّث تقديرات دالة القيمة بناءً على الفرق بين تنبؤات متعاقبة بدل انتظار النتيجة النهائية. تجمع بين فكرة التمهيد من البرمجة الديناميكية وأخذ العيّنات من أساليب مونت كارلو.
A family of reinforcement learning algorithms that update value-function estimates based on the difference between successive predictions rather than waiting for the final outcome. Combines the bootstrapping idea from dynamic programming with the sampling approach of Monte Carlo methods.
Also translated asتعلّم الفرق الزمني، أسلوب الفرق الزمني
First appears in this corpus in: TD3: Addressing Function Approximation Error in Actor-Critic Methods (2018)
Appears in these papers
- Decision Transformer: Reinforcement Learning via Sequence Modeling2021in the sky ✦
- Decision Transformer: Reinforcement Learning via Sequence Modeling2021in the sky ✦
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances2022in the sky ✦
- TD3: Addressing Function Approximation Error in Actor-Critic Methods2018in the sky ✦