Target Policy
سياسة الهدف (المُتعلَّمة)
السياسة التي يسعى الوكيل لتعلّمها وتحسينها. في Q-Learning هي دائماً السياسة الجشعة (اختيار الفعل الأعلى قيمة Q) بصرف النظر عن سياسة السلوك المُتّبعة.
The policy the agent is trying to learn and improve. In Q-Learning this is always the greedy policy (pick the highest Q-value action) regardless of the behavior policy being followed.
Also translated asالسياسة المستهدفة، سياسة التقييم المثلى
First appears in this corpus in: Q-Learning (1992)
Appears in these papers
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures2018in the sky ✦
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures2018in the sky ✦
- Q-Learning1992in the sky ✦
- TD3: Addressing Function Approximation Error in Actor-Critic Methods2018in the sky ✦