Target Network
شبكة الهدف
نسخة منفصلة من شبكة Q تُجمَّد أوزانها لآلاف الخطوات ثم تُحدَّث دورياً من الشبكة الرئيسية، مما يُثبّت هدف الأمثَلَة ويمنع التشعّب الناجم عن ملاحقة هدف متحرك.
A separate copy of the Q-network whose weights are frozen for thousands of steps then periodically updated from the main network, stabilizing the optimization target and preventing divergence from chasing a moving target.
Also translated asالشبكة المُجمَّدة، شبكة الهدف المُثبّتة
First appears in this corpus in: Continuous Control with Deep Reinforcement Learning (2015)
Appears in these papers
- Bootstrap Your Own Latent: A New Approach to Self-Supervised Learning2020in the sky ✦
- Bootstrap Your Own Latent: A New Approach to Self-Supervised Learning2020in the sky ✦
- Conservative Q-Learning for Offline Reinforcement Learning2020in the sky ✦
- Continuous Control with Deep Reinforcement Learning2015in the sky ✦
- Continuous Control with Deep Reinforcement Learning2015in the sky ✦
- Deep Reinforcement Learning with Double Q-Learning2016in the sky ✦
- Deep Reinforcement Learning with Double Q-Learning2016in the sky ✦
- Human-Level Control Through Deep Reinforcement Learning2015in the sky ✦
- Human-Level Control Through Deep Reinforcement Learning2015in the sky ✦
- Mastering Diverse Domains Through World Models2023in the sky ✦
- Dueling Network Architectures for Deep Reinforcement Learning2016in the sky ✦
- Dueling Network Architectures for Deep Reinforcement Learning2016in the sky ✦
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation2018in the sky ✦
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation2018in the sky ✦
- Rainbow: Combining Improvements in Deep Reinforcement Learning2018in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- TD3: Addressing Function Approximation Error in Actor-Critic Methods2018in the sky ✦
- TD3: Addressing Function Approximation Error in Actor-Critic Methods2018in the sky ✦