Advantage
الميزة
مقدار تفوّق ناتج معين على متوسط أداء المجموعة، يُحدّد اتجاه وقوة تحديث السياسة في التعلم المعزز.
The degree to which a specific output exceeds the group's average performance, determining the direction and strength of the policy update in RL.
Also translated asالأفضلية، فرق الأداء، مقياس تميز الفعل مقارنة بالمتوسط، ميزة الخطوة البديلة
First appears in this corpus in: Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning (1992)
Appears in these papers
- Asynchronous Methods for Deep Reinforcement Learning2016in the sky ✦
- Asynchronous Methods for Deep Reinforcement Learning2016in the sky ✦
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning2025in the sky ✦
- Deep Reinforcement Learning with Double Q-Learning2016in the sky ✦
- Dueling Network Architectures for Deep Reinforcement Learning2016in the sky ✦
- Dueling Network Architectures for Deep Reinforcement Learning2016in the sky ✦
- High-Dimensional Continuous Control Using Generalized Advantage Estimation2016in the sky ✦
- High-Dimensional Continuous Control Using Generalized Advantage Estimation2016in the sky ✦
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures2018in the sky ✦
- Learning Dexterous In-Hand Manipulation2019in the sky ✦
- Learning Dexterous In-Hand Manipulation2019in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Proximal Policy Optimization Algorithms2017in the sky ✦
- Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning1992in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- Trust Region Policy Optimization2015in the sky ✦
- Trust Region Policy Optimization2015in the sky ✦