Reinforcement Learning
التعلم المعزز
نموذج تعلم آلي يتعلم فيه الوكيل اتخاذ قرارات متتالية عبر التفاعل مع بيئة تُعيد له إشارات مكافأة أو عقوبة، بهدف تعظيم المكافأة التراكمية.
Also translated asالتعلم التعزيزي، التعلم بالمكافأة، التعلم القائم على المكافأة والعقاب، التعلم التدعيمي التفاعلي، تدريب بالتعلُّم المعزز
First appears in this corpus in: The Monte Carlo Method (1949)
Appears in these papers
- Asynchronous Methods for Deep Reinforcement Learning2016in the sky ✦
- AI Safety via Debate2018in the sky ✦
- AI Safety via Debate2018in the sky ✦
- Alpaca: A Strong, Replicable Instruction-Following Model2023in the sky ✦
- Mastering the Game of Go Without Human Knowledge2017in the sky ✦
- Mastering the Game of Go Without Human Knowledge2017in the sky ✦
- Mastering the Game of Go with Deep Neural Networks and Tree Search2016in the sky ✦
- Discovering Faster Matrix Multiplication Algorithms with Reinforcement Learning2022in the sky ✦
- A General Reinforcement Learning Algorithm That Masters Chess, Shogi, and Go Through Self-Play2018in the sky ✦
- A General Reinforcement Learning Algorithm That Masters Chess, Shogi, and Go Through Self-Play2018in the sky ✦
- Concrete Problems in AI Safety2016in the sky ✦
- Concrete Problems in AI Safety2016in the sky ✦
- Constitutional AI: Harmlessness from AI Feedback2022in the sky ✦
- Representation Learning with Contrastive Predictive Coding2018in the sky ✦
- Conservative Q-Learning for Offline Reinforcement Learning2020in the sky ✦
- Continuous Control with Deep Reinforcement Learning2015in the sky ✦
- Decision Transformer: Reinforcement Learning via Sequence Modeling2021in the sky ✦
- Deep Learning2015in the sky ✦
- Deep Reinforcement Learning from Human Preferences2017in the sky ✦
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning2025in the sky ✦
- DeepSeek-V3 Technical Report2024in the sky ✦
- Human-Level Control Through Deep Reinforcement Learning2015in the sky ✦
- Mastering Diverse Domains Through World Models2023in the sky ✦
- Mastering Diverse Domains Through World Models2023in the sky ✦
- Dueling Network Architectures for Deep Reinforcement Learning2016in the sky ✦
- Dynamic Programming1957in the sky ✦
- ELECTRA: Pre-Training Text Encoders as Discriminators Rather Than Generators2020in the sky ✦
- High-Dimensional Continuous Control Using Generalized Advantage Estimation2016in the sky ✦
- Google's Neural Machine Translation System: Bridging the Gap Between Human and Machine Translation2016in the sky ✦
- Google's Neural Machine Translation System: Bridging the Gap Between Human and Machine Translation2016in the sky ✦
- First Return, Then Explore2021in the sky ✦
- First Return, Then Explore2021in the sky ✦
- Hindsight Experience Replay2017in the sky ✦
- Curiosity-Driven Exploration by Self-Supervised Prediction2017in the sky ✦
- Curiosity-Driven Exploration by Self-Supervised Prediction2017in the sky ✦
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures2018in the sky ✦
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures2018in the sky ✦
- Training Language Models to Follow Instructions with Human Feedback2022in the sky ✦
- On Information and Sufficiency1951in the sky ✦
- Learning Dexterous In-Hand Manipulation2019in the sky ✦
- Learning to Summarize from Human Feedback2020in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- The Monte Carlo Method1949in the sky ✦
- MuZero: Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model2020in the sky ✦
- Natural Gradient Works Efficiently in Learning1998in the sky ✦
- Learning to Reason with LLMs2024in the sky ✦
- Between MDPs and Semi-MDPs — A Framework for Temporal Abstraction in Reinforcement Learning1999in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Proximal Policy Optimization Algorithms2017in the sky ✦
- Prioritized Experience Replay2015in the sky ✦
- Q-Learning1992in the sky ✦
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation2018in the sky ✦
- Rainbow: Combining Improvements in Deep Reinforcement Learning2018in the sky ✦
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned2022in the sky ✦
- Reflexion: Language Agents with Verbal Reinforcement Learning2023in the sky ✦
- Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning1992in the sky ✦
- Deep Residual Learning for Image Recognition2015in the sky ✦
- RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback2023in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- Some Studies in Machine Learning Using the Game of Checkers1959in the sky ✦
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances2022in the sky ✦
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances2022in the sky ✦
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention2015in the sky ✦
- Optimization by Simulated Annealing1983in the sky ✦
- Temporal Difference Learning and TD-Gammon1995in the sky ✦
- Temporal Difference Learning and TD-Gammon1995in the sky ✦
- Learning to Predict by the Methods of Temporal Differences1988in the sky ✦
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models2023in the sky ✦
- Trust Region Policy Optimization2015in the sky ✦
- Computing Machinery and Intelligence1950in the sky ✦
- Computing Machinery and Intelligence1950in the sky ✦
- End-to-End Training of Deep Visuomotor Policies2016in the sky ✦
- End-to-End Training of Deep Visuomotor Policies2016in the sky ✦
- Voyager: An Open-Ended Embodied Agent with Large Language Models2023in the sky ✦
- WebGPT: Browser-Assisted Question-Answering with Human Feedback2021in the sky ✦
- WebGPT: Browser-Assisted Question-Answering with Human Feedback2021in the sky ✦
- World Models2018in the sky ✦