Discount Factor
مُعامل الخصم
عدد بين 0 و1 يحدد مقدار أهمية المكافآت المستقبلية مقارنة بالمكافآت الفورية في معادلة بيلمان.
A number between 0 and 1 determining how much future rewards matter relative to immediate rewards in the Bellman Equation.
Also translated asعامل التخفيض الزمني، عامل التخميد الحسابي للمكافآت المؤجلة، نسبة ترجيح الحاضر على المستقبل
First appears in this corpus in: Learning to Predict by the Methods of Temporal Differences (1988)
Appears in these papers
- Conservative Q-Learning for Offline Reinforcement Learning2020in the sky ✦
- Continuous Control with Deep Reinforcement Learning2015in the sky ✦
- Decision Transformer: Reinforcement Learning via Sequence Modeling2021in the sky ✦
- Mastering Diverse Domains Through World Models2023in the sky ✦
- Dueling Network Architectures for Deep Reinforcement Learning2016in the sky ✦
- High-Dimensional Continuous Control Using Generalized Advantage Estimation2016in the sky ✦
- High-Dimensional Continuous Control Using Generalized Advantage Estimation2016in the sky ✦
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures2018in the sky ✦
- Learning Dexterous In-Hand Manipulation2019in the sky ✦
- MuZero: Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model2020in the sky ✦
- Between MDPs and Semi-MDPs — A Framework for Temporal Abstraction in Reinforcement Learning1999in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Q-Learning1992in the sky ✦
- Q-Learning1992in the sky ✦
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation2018in the sky ✦
- Rainbow: Combining Improvements in Deep Reinforcement Learning2018in the sky ✦
- Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning1992in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- Temporal Difference Learning and TD-Gammon1995in the sky ✦
- Learning to Predict by the Methods of Temporal Differences1988in the sky ✦
- TD3: Addressing Function Approximation Error in Actor-Critic Methods2018in the sky ✦