Markov Decision Process (MDP)
عملية ماركوف لاتخاذ القرار
الإطار الرياضي الرسمي المستخدم لنمذجة وصياغة مسائل التعلم بالتعزيز وبيئاتها التفاعلية.
Markov Decision Process (MDP)
Also translated asالصياغة الرياضية لبيئات التعزيز، هيكلية القرارات الاحتمالية المتعاقبة
First appears in this corpus in: Dynamic Programming (1957)
Appears in these papers
- Decision Transformer: Reinforcement Learning via Sequence Modeling2021in the sky ✦
- Dynamic Programming1957in the sky ✦
- Between MDPs and Semi-MDPs — A Framework for Temporal Abstraction in Reinforcement Learning1999in the sky ✦
- Between MDPs and Semi-MDPs — A Framework for Temporal Abstraction in Reinforcement Learning1999in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation2018in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦