Exploration
الاستكشاف (تجربة أفعال جديدة)
رغبة العميل الذكي في تجربة أفعال ومسارات جديدة غير مجربة لاستكشاف فرص وعوائد أفضل في البيئة.
Exploration
Also translated asالبحث العشوائي عن فرص أفضل، التقصّي البيئي المفتوح
First appears in this corpus in: Equation of State Calculations by Fast Computing Machines (1953)
Appears in these papers
- Asynchronous Methods for Deep Reinforcement Learning2016in the sky ✦
- Asynchronous Methods for Deep Reinforcement Learning2016in the sky ✦
- Mastering the Game of Go Without Human Knowledge2017in the sky ✦
- Mastering the Game of Go with Deep Neural Networks and Tree Search2016in the sky ✦
- Discovering Faster Matrix Multiplication Algorithms with Reinforcement Learning2022in the sky ✦
- A General Reinforcement Learning Algorithm That Masters Chess, Shogi, and Go Through Self-Play2018in the sky ✦
- A General Reinforcement Learning Algorithm That Masters Chess, Shogi, and Go Through Self-Play2018in the sky ✦
- Practical Bayesian Optimization of Machine Learning Algorithms2012in the sky ✦
- Practical Bayesian Optimization of Machine Learning Algorithms2012in the sky ✦
- Concrete Problems in AI Safety2016in the sky ✦
- Constitutional AI: Harmlessness from AI Feedback2022in the sky ✦
- Continuous Control with Deep Reinforcement Learning2015in the sky ✦
- Continuous Control with Deep Reinforcement Learning2015in the sky ✦
- Deep Reinforcement Learning with Double Q-Learning2016in the sky ✦
- Human-Level Control Through Deep Reinforcement Learning2015in the sky ✦
- Mastering Diverse Domains Through World Models2023in the sky ✦
- Mastering Diverse Domains Through World Models2023in the sky ✦
- Stochastic Relaxation, Gibbs Distributions, and the Bayesian Restoration of Images1984in the sky ✦
- First Return, Then Explore2021in the sky ✦
- First Return, Then Explore2021in the sky ✦
- Hindsight Experience Replay2017in the sky ✦
- Hindsight Experience Replay2017in the sky ✦
- Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization2018in the sky ✦
- Curiosity-Driven Exploration by Self-Supervised Prediction2017in the sky ✦
- Curiosity-Driven Exploration by Self-Supervised Prediction2017in the sky ✦
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures2018in the sky ✦
- Learning Dexterous In-Hand Manipulation2019in the sky ✦
- Equation of State Calculations by Fast Computing Machines1953in the sky ✦
- MuZero: Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model2020in the sky ✦
- Between MDPs and Semi-MDPs — A Framework for Temporal Abstraction in Reinforcement Learning1999in the sky ✦
- Between MDPs and Semi-MDPs — A Framework for Temporal Abstraction in Reinforcement Learning1999in the sky ✦
- Proximal Policy Optimization Algorithms2017in the sky ✦
- Q-Learning1992in the sky ✦
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation2018in the sky ✦
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation2018in the sky ✦
- Rainbow: Combining Improvements in Deep Reinforcement Learning2018in the sky ✦
- Rainbow: Combining Improvements in Deep Reinforcement Learning2018in the sky ✦
- Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning1992in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- Optimization by Simulated Annealing1983in the sky ✦
- Temporal Difference Learning and TD-Gammon1995in the sky ✦
- Temporal Difference Learning and TD-Gammon1995in the sky ✦
- TD3: Addressing Function Approximation Error in Actor-Critic Methods2018in the sky ✦
- Voyager: An Open-Ended Embodied Agent with Large Language Models2023in the sky ✦
- WebArena: A Realistic Web Environment for Building Autonomous Agents2023in the sky ✦