Rollout
التوليد التجريبي
مرحلة في التعلم المعزز يُولّد فيها نموذج السياسة الحالي مخرجات متعددة لمجموعة أسئلة، تُستخدم لاحقاً لحساب المكافآت وتحديث السياسة.
A phase in RL where the current policy model generates multiple outputs for a set of questions, later used to compute rewards and update the policy.
Also translated asالتشغيل التجريبي، توليد العيّنات
First appears in this corpus in: Mastering the Game of Go with Deep Neural Networks and Tree Search (2016)
Appears in these papers
- Mastering the Game of Go with Deep Neural Networks and Tree Search2016in the sky ✦
- Learning Dexterous In-Hand Manipulation2019in the sky ✦
- Learning Dexterous In-Hand Manipulation2019in the sky ✦
- MuZero: Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model2020in the sky ✦
- World Models2018in the sky ✦