Off-Policy
خوارزمية التعلم خارج السياسة الحالية
خوارزميات تعزيز تمتلك مرونة التعلم وتحديث استراتيجيتها بالاعتماد على سجلات تفاعلية سابقة أو سياسات سلوكية أخرى.
Off-Policy
Also translated asالتعلم التدعيمي المستقل عن استراتيجية التنفيذ، مواءمة الأفعال عبر السجلات التاريخية
First appears in this corpus in: Q-Learning (1992)
Appears in these papers
- Asynchronous Methods for Deep Reinforcement Learning2016in the sky ✦
- Conservative Q-Learning for Offline Reinforcement Learning2020in the sky ✦
- Human-Level Control Through Deep Reinforcement Learning2015in the sky ✦
- Hindsight Experience Replay2017in the sky ✦
- Hindsight Experience Replay2017in the sky ✦
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures2018in the sky ✦
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures2018in the sky ✦
- Prioritized Experience Replay2015in the sky ✦
- Q-Learning1992in the sky ✦
- Q-Learning1992in the sky ✦
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation2018in the sky ✦
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation2018in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦