Reward
المكافأة
القيمة العددية التحفيزية التي يتلقاها العميل من البيئة لتقييم مدى جودة الفعل المتخذ.
Reward
Also translated asالقيمة التحفيزية، دالة التعزيز الإيجابي
First appears in this corpus in: Dynamic Programming (1957)
Appears in these papers
- Asynchronous Methods for Deep Reinforcement Learning2016in the sky ✦
- AI Safety via Debate2018in the sky ✦
- Discovering Faster Matrix Multiplication Algorithms with Reinforcement Learning2022in the sky ✦
- A General Reinforcement Learning Algorithm That Masters Chess, Shogi, and Go Through Self-Play2018in the sky ✦
- Concrete Problems in AI Safety2016in the sky ✦
- Conservative Q-Learning for Offline Reinforcement Learning2020in the sky ✦
- Continuous Control with Deep Reinforcement Learning2015in the sky ✦
- Continuous Control with Deep Reinforcement Learning2015in the sky ✦
- Deep Reinforcement Learning with Double Q-Learning2016in the sky ✦
- Human-Level Control Through Deep Reinforcement Learning2015in the sky ✦
- Mastering Diverse Domains Through World Models2023in the sky ✦
- Dueling Network Architectures for Deep Reinforcement Learning2016in the sky ✦
- Dynamic Programming1957in the sky ✦
- High-Dimensional Continuous Control Using Generalized Advantage Estimation2016in the sky ✦
- High-Dimensional Continuous Control Using Generalized Advantage Estimation2016in the sky ✦
- Google's Neural Machine Translation System: Bridging the Gap Between Human and Machine Translation2016in the sky ✦
- First Return, Then Explore2021in the sky ✦
- Hindsight Experience Replay2017in the sky ✦
- Curiosity-Driven Exploration by Self-Supervised Prediction2017in the sky ✦
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures2018in the sky ✦
- Training Language Models to Follow Instructions with Human Feedback2022in the sky ✦
- Between MDPs and Semi-MDPs — A Framework for Temporal Abstraction in Reinforcement Learning1999in the sky ✦
- Between MDPs and Semi-MDPs — A Framework for Temporal Abstraction in Reinforcement Learning1999in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Proximal Policy Optimization Algorithms2017in the sky ✦
- Prioritized Experience Replay2015in the sky ✦
- Q-Learning1992in the sky ✦
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation2018in the sky ✦
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation2018in the sky ✦
- Reflexion: Language Agents with Verbal Reinforcement Learning2023in the sky ✦
- Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning1992in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- Temporal Difference Learning and TD-Gammon1995in the sky ✦
- Temporal Difference Learning and TD-Gammon1995in the sky ✦
- Learning to Predict by the Methods of Temporal Differences1988in the sky ✦
- Trust Region Policy Optimization2015in the sky ✦
- World Models2018in the sky ✦
- World Models2018in the sky ✦