Value Function
دالة تقييم العوائد
دالة رياضية تحسب إجمالي العوائد والمكافآت المتوقعة مستقبلاً عند الانطلاق من وضعية أو حالة معينة.
Value Function
Also translated asدالة حساب القيمة المتوقعة للحالة، معيار الجدوى الاستراتيجية للموقع
First appears in this corpus in: The Monte Carlo Method (1949)
Appears in these papers
- Asynchronous Methods for Deep Reinforcement Learning2016in the sky ✦
- Grandmaster Level in StarCraft II Using Multi-Agent Reinforcement Learning2019in the sky ✦
- Discovering Faster Matrix Multiplication Algorithms with Reinforcement Learning2022in the sky ✦
- A General Reinforcement Learning Algorithm That Masters Chess, Shogi, and Go Through Self-Play2018in the sky ✦
- A General Reinforcement Learning Algorithm That Masters Chess, Shogi, and Go Through Self-Play2018in the sky ✦
- Decision Transformer: Reinforcement Learning via Sequence Modeling2021in the sky ✦
- Decision Transformer: Reinforcement Learning via Sequence Modeling2021in the sky ✦
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning2025in the sky ✦
- DeepSeek-V3 Technical Report2024in the sky ✦
- Direct Preference Optimization: Your Language Model Is Secretly a Reward Model2023in the sky ✦
- Mastering Diverse Domains Through World Models2023in the sky ✦
- Mastering Diverse Domains Through World Models2023in the sky ✦
- Dynamic Programming1957in the sky ✦
- Dynamic Programming1957in the sky ✦
- High-Dimensional Continuous Control Using Generalized Advantage Estimation2016in the sky ✦
- High-Dimensional Continuous Control Using Generalized Advantage Estimation2016in the sky ✦
- Hindsight Experience Replay2017in the sky ✦
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures2018in the sky ✦
- Learning Dexterous In-Hand Manipulation2019in the sky ✦
- Learning Dexterous In-Hand Manipulation2019in the sky ✦
- Learning to Summarize from Human Feedback2020in the sky ✦
- The Monte Carlo Method1949in the sky ✦
- MuZero: Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model2020in the sky ✦
- MuZero: Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model2020in the sky ✦
- Between MDPs and Semi-MDPs — A Framework for Temporal Abstraction in Reinforcement Learning1999in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Proximal Policy Optimization Algorithms2017in the sky ✦
- Prioritized Experience Replay2015in the sky ✦
- Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning1992in the sky ✦
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances2022in the sky ✦
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances2022in the sky ✦
- Temporal Difference Learning and TD-Gammon1995in the sky ✦
- Temporal Difference Learning and TD-Gammon1995in the sky ✦
- TD3: Addressing Function Approximation Error in Actor-Critic Methods2018in the sky ✦