Policy
السياسة
في التعلم المعزز، الاستراتيجية التي تربط الحالات بالأفعال. في RLHF، النموذج اللغوي ذاته — يربط الأوامر بالإجابات.
In RL, the strategy that maps states to actions. In RLHF, the language model itself — mapping prompts to responses.
Also translated asاستراتيجية العمل، سياسة التوليد، سياسة النموذج، منهجية الاختيار والترجيح، نموذج السياسة
First appears in this corpus in: The Monte Carlo Method (1949)
Appears in these papers
- Asynchronous Methods for Deep Reinforcement Learning2016in the sky ✦
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware2023in the sky ✦
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware2023in the sky ✦
- Mastering the Game of Go Without Human Knowledge2017in the sky ✦
- Discovering Faster Matrix Multiplication Algorithms with Reinforcement Learning2022in the sky ✦
- A General Reinforcement Learning Algorithm That Masters Chess, Shogi, and Go Through Self-Play2018in the sky ✦
- A General Reinforcement Learning Algorithm That Masters Chess, Shogi, and Go Through Self-Play2018in the sky ✦
- Concrete Problems in AI Safety2016in the sky ✦
- Conservative Q-Learning for Offline Reinforcement Learning2020in the sky ✦
- Continuous Control with Deep Reinforcement Learning2015in the sky ✦
- Deep Reinforcement Learning from Human Preferences2017in the sky ✦
- Deep Reinforcement Learning with Double Q-Learning2016in the sky ✦
- Direct Preference Optimization: Your Language Model Is Secretly a Reward Model2023in the sky ✦
- Human-Level Control Through Deep Reinforcement Learning2015in the sky ✦
- Dueling Network Architectures for Deep Reinforcement Learning2016in the sky ✦
- Dynamic Programming1957in the sky ✦
- High-Dimensional Continuous Control Using Generalized Advantage Estimation2016in the sky ✦
- First Return, Then Explore2021in the sky ✦
- Hindsight Experience Replay2017in the sky ✦
- Curiosity-Driven Exploration by Self-Supervised Prediction2017in the sky ✦
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures2018in the sky ✦
- Training Language Models to Follow Instructions with Human Feedback2022in the sky ✦
- On Information and Sufficiency1951in the sky ✦
- Learning Dexterous In-Hand Manipulation2019in the sky ✦
- Learning Dexterous In-Hand Manipulation2019in the sky ✦
- Learning to Summarize from Human Feedback2020in the sky ✦
- The Monte Carlo Method1949in the sky ✦
- Natural Gradient Works Efficiently in Learning1998in the sky ✦
- Between MDPs and Semi-MDPs — A Framework for Temporal Abstraction in Reinforcement Learning1999in the sky ✦
- Between MDPs and Semi-MDPs — A Framework for Temporal Abstraction in Reinforcement Learning1999in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Proximal Policy Optimization Algorithms2017in the sky ✦
- Prioritized Experience Replay2015in the sky ✦
- Q-Learning1992in the sky ✦
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation2018in the sky ✦
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation2018in the sky ✦
- Reflexion: Language Agents with Verbal Reinforcement Learning2023in the sky ✦
- Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning1992in the sky ✦
- RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback2023in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- Optimization by Simulated Annealing1983in the sky ✦
- Temporal Difference Learning and TD-Gammon1995in the sky ✦
- Temporal Difference Learning and TD-Gammon1995in the sky ✦
- End-to-End Training of Deep Visuomotor Policies2016in the sky ✦
- End-to-End Training of Deep Visuomotor Policies2016in the sky ✦
- WebGPT: Browser-Assisted Question-Answering with Human Feedback2021in the sky ✦
- World Models2018in the sky ✦
- World Models2018in the sky ✦