Bellman Equation
معادلة بيلمان الرياضية
المعادلة المحورية التي تفكك قيمة الحالة الحالية لتربطها بالمكافأة الفورية زائد قيمة الحالة التالية.
Bellman Equation
Also translated asمعادلة التحديث العودي لجدوى الأفعال، معيار الربط الزمني للعوائد المتوقعة
First appears in this corpus in: Dynamic Programming (1957)
Appears in these papers
- Conservative Q-Learning for Offline Reinforcement Learning2020in the sky ✦
- Conservative Q-Learning for Offline Reinforcement Learning2020in the sky ✦
- Continuous Control with Deep Reinforcement Learning2015in the sky ✦
- Decision Transformer: Reinforcement Learning via Sequence Modeling2021in the sky ✦
- Deep Reinforcement Learning with Double Q-Learning2016in the sky ✦
- Human-Level Control Through Deep Reinforcement Learning2015in the sky ✦
- Dueling Network Architectures for Deep Reinforcement Learning2016in the sky ✦
- Dynamic Programming1957in the sky ✦
- Dynamic Programming1957in the sky ✦
- Between MDPs and Semi-MDPs — A Framework for Temporal Abstraction in Reinforcement Learning1999in the sky ✦
- Between MDPs and Semi-MDPs — A Framework for Temporal Abstraction in Reinforcement Learning1999in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Q-Learning1992in the sky ✦
- Q-Learning1992in the sky ✦
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation2018in the sky ✦
- Rainbow: Combining Improvements in Deep Reinforcement Learning2018in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- TD3: Addressing Function Approximation Error in Actor-Critic Methods2018in the sky ✦