Epsilon-Greedy
ε-الجشع
استراتيجية اختيار أفعال تختار فعلاً عشوائياً باحتمال ε (استكشاف) وأفضل فعل معروف باحتمال 1−ε (استغلال)، مع تقليل ε تدريجياً أثناء التدريب للانتقال من الاستكشاف إلى الاستغلال.
An action selection strategy that picks a random action with probability ε (exploration) and the best known action with probability 1−ε (exploitation), gradually decreasing ε during training to shift from exploration to exploitation.
Also translated asإبسيلون الجشع، الاستكشاف بإبسيلون
First appears in this corpus in: Human-Level Control Through Deep Reinforcement Learning (2015)
Appears in these papers
- Deep Reinforcement Learning with Double Q-Learning2016in the sky ✦
- Human-Level Control Through Deep Reinforcement Learning2015in the sky ✦
- Human-Level Control Through Deep Reinforcement Learning2015in the sky ✦
- Dueling Network Architectures for Deep Reinforcement Learning2016in the sky ✦
- Dueling Network Architectures for Deep Reinforcement Learning2016in the sky ✦