Replay Buffer
ذاكرة التجارب
بنية بيانات دائرية تخزّن الانتقالات الأخيرة (عادةً مليون انتقال) التي يسحب منها الوكيل دفعات مصغرة عشوائية لتدريب شبكته، مما يكسر الترابط الزمني ويُحسّن كفاءة البيانات.
A circular data structure that stores the most recent transitions (typically one million) from which the agent samples random mini-batches to train its network, breaking temporal correlation and improving data efficiency.
Also translated asالمخزن المؤقت لإعادة التشغيل، ذاكرة الإعادة
First appears in this corpus in: Continuous Control with Deep Reinforcement Learning (2015)
Appears in these papers
- Asynchronous Methods for Deep Reinforcement Learning2016in the sky ✦
- Conservative Q-Learning for Offline Reinforcement Learning2020in the sky ✦
- Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks2017in the sky ✦
- Continuous Control with Deep Reinforcement Learning2015in the sky ✦
- Deep Reinforcement Learning with Double Q-Learning2016in the sky ✦
- Human-Level Control Through Deep Reinforcement Learning2015in the sky ✦
- Human-Level Control Through Deep Reinforcement Learning2015in the sky ✦
- Mastering Diverse Domains Through World Models2023in the sky ✦
- Mastering Diverse Domains Through World Models2023in the sky ✦
- Hindsight Experience Replay2017in the sky ✦
- Hindsight Experience Replay2017in the sky ✦
- MuZero: Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model2020in the sky ✦
- MuZero: Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model2020in the sky ✦
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models2024in the sky ✦
- Prioritized Experience Replay2015in the sky ✦
- Prioritized Experience Replay2015in the sky ✦
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation2018in the sky ✦
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation2018in the sky ✦
- Rainbow: Combining Improvements in Deep Reinforcement Learning2018in the sky ✦
- Rainbow: Combining Improvements in Deep Reinforcement Learning2018in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- TD3: Addressing Function Approximation Error in Actor-Critic Methods2018in the sky ✦