Offline Reinforcement Learning
التعلُّم المعزَّز دون اتصال
نموذج في التعلُّم المعزَّز يتعلّم فيه الوكيل سياسة من مجموعة بيانات ثابتة جُمعت مسبقاً، دون أي تفاعل إضافي مع البيئة. يُستخدم حين يكون الاستكشاف المباشر مكلفاً أو خطيراً كما في الرعاية الصحية والروبوتات.
A reinforcement learning paradigm where the agent learns a policy from a fixed, previously collected dataset without any further interaction with the environment. Used when online exploration is expensive or dangerous, such as in healthcare and robotics.
Also translated asالتعلُّم المعزَّز المُنفصل، التعلُّم المعزَّز الدُّفعي
First appears in this corpus in: Conservative Q-Learning for Offline Reinforcement Learning (2020)
Appears in these papers