Q-Learning
تعلم دالة الجودة (Q)
خوارزمية تعلم تعزيز مباشرة تحدّث قيم جودة الأفعال (Q) مستقلة عن السياسة السلوكية المتبعة لحظياً.
Q-Learning
Also translated asخوارزمية جدول الأفعال والمنفعة، التعلم التدعيمي غير المقيد بالسياسة المسبقة
First appears in this corpus in: Learning to Predict by the Methods of Temporal Differences (1988)
Appears in these papers
- Conservative Q-Learning for Offline Reinforcement Learning2020in the sky ✦
- Conservative Q-Learning for Offline Reinforcement Learning2020in the sky ✦
- Deep Reinforcement Learning with Double Q-Learning2016in the sky ✦
- Human-Level Control Through Deep Reinforcement Learning2015in the sky ✦
- Between MDPs and Semi-MDPs — A Framework for Temporal Abstraction in Reinforcement Learning1999in the sky ✦
- Between MDPs and Semi-MDPs — A Framework for Temporal Abstraction in Reinforcement Learning1999in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Q-Learning1992in the sky ✦
- Learning to Predict by the Methods of Temporal Differences1988in the sky ✦