Episode
جولة تفاعلية كاملة
الدورة أو الجولة التفاعلية الكاملة للعميل وتبدأ من وضعية الانطلاق وتستمر حتى بلوغ النهاية أو الفشل.
Episode
Also translated asمحاولة تشغيلية من البداية للنهاية، الحلقة التدريبية المغلقة للعميل
First appears in this corpus in: Q-Learning (1992)
Appears in these papers
- Human-Level Control Through Deep Reinforcement Learning2015in the sky ✦
- High-Dimensional Continuous Control Using Generalized Advantage Estimation2016in the sky ✦
- First Return, Then Explore2021in the sky ✦
- Hindsight Experience Replay2017in the sky ✦
- Learning Dexterous In-Hand Manipulation2019in the sky ✦
- Learning Dexterous In-Hand Manipulation2019in the sky ✦
- Between MDPs and Semi-MDPs — A Framework for Temporal Abstraction in Reinforcement Learning1999in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Q-Learning1992in the sky ✦
- Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning1992in the sky ✦
- RT-1: Robotics Transformer for Real-World Control at Scale2022in the sky ✦
- RT-1: Robotics Transformer for Real-World Control at Scale2022in the sky ✦