Behavior Policy
سياسة السلوك (الاستكشافية)
السياسة التي يتبعها الوكيل فعلياً أثناء التفاعل مع البيئة لجمع التجارب. في Q-Learning تكون عادةً سياسة ε-جشعة تمزج بين الاستكشاف والاستغلال.
The policy the agent actually follows while interacting with the environment to collect experience. In Q-Learning this is typically an ε-greedy policy mixing exploration and exploitation.
Also translated asسياسة التفاعل، سياسة جمع البيانات
First appears in this corpus in: Q-Learning (1992)
Appears in these papers
- Conservative Q-Learning for Offline Reinforcement Learning2020in the sky ✦
- Conservative Q-Learning for Offline Reinforcement Learning2020in the sky ✦
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures2018in the sky ✦
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures2018in the sky ✦
- Q-Learning1992in the sky ✦