Policy Gradient
تدرج السياسة التشغيلية
أساليب تحسين تحدث معالم استراتيجية الأفعال مباشرة عبر تتبع تدرج العوائد لزيادة احتمال الخطوات الناجحة.
Policy Gradient
Also translated asالتحسين المباشر لميل استراتيجية الأفعال، توجيه مسارات الاختيار آلياً
First appears in this corpus in: Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning (1992)
Appears in these papers
- Asynchronous Methods for Deep Reinforcement Learning2016in the sky ✦
- Asynchronous Methods for Deep Reinforcement Learning2016in the sky ✦
- Mastering the Game of Go with Deep Neural Networks and Tree Search2016in the sky ✦
- Grandmaster Level in StarCraft II Using Multi-Agent Reinforcement Learning2019in the sky ✦
- Conservative Q-Learning for Offline Reinforcement Learning2020in the sky ✦
- Continuous Control with Deep Reinforcement Learning2015in the sky ✦
- Decision Transformer: Reinforcement Learning via Sequence Modeling2021in the sky ✦
- Decision Transformer: Reinforcement Learning via Sequence Modeling2021in the sky ✦
- Mastering Diverse Domains Through World Models2023in the sky ✦
- Mastering Diverse Domains Through World Models2023in the sky ✦
- High-Dimensional Continuous Control Using Generalized Advantage Estimation2016in the sky ✦
- High-Dimensional Continuous Control Using Generalized Advantage Estimation2016in the sky ✦
- Google's Neural Machine Translation System: Bridging the Gap Between Human and Machine Translation2016in the sky ✦
- First Return, Then Explore2021in the sky ✦
- Curiosity-Driven Exploration by Self-Supervised Prediction2017in the sky ✦
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures2018in the sky ✦
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures2018in the sky ✦
- Learning Dexterous In-Hand Manipulation2019in the sky ✦
- Learning to Summarize from Human Feedback2020in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Proximal Policy Optimization Algorithms2017in the sky ✦
- Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning1992in the sky ✦
- Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning1992in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention2015in the sky ✦
- STaR: Bootstrapping Reasoning With Reasoning2022in the sky ✦
- TD3: Addressing Function Approximation Error in Actor-Critic Methods2018in the sky ✦
- Trust Region Policy Optimization2015in the sky ✦
- Trust Region Policy Optimization2015in the sky ✦