Semi-Supervised Reinforcement Learning
التعلّم المعزّز شبه المُشرَف
إعداد تعلّم معزّز يستطيع فيه الوكيل رؤية مكافأته الحقيقية في نسبة صغيرة فقط من الحلقات، لكنه يُقيَّم على أساس جميع الحلقات.
A reinforcement learning setting where the agent can only see its true reward on a small fraction of episodes, but its performance is evaluated based on all episodes.
Also translated asالتعلّم المعزّز بإشراف جزئي