Glossary

SARSA

خوارزمية SARSA

خوارزمية تعلم معزز داخل السياسة تُحدِّث قيم Q باستخدام الفعل الذي اتُّخِذ فعلاً في الحالة التالية (بدلاً من أفضل فعل كما في Q-Learning)، مما يجعلها تُقيّم السياسة التي تتبعها فعلاً.

An on-policy reinforcement learning algorithm that updates Q-values using the action actually taken in the next state (rather than the best action as in Q-Learning), making it evaluate the policy it actually follows.

Also translated asحالة-فعل-مكافأة-حالة-فعل

First appears in this corpus in: Policy Gradient Methods for Reinforcement Learning with Function Approximation (1999)

Appears in these papers