Function Approximation
تقريب الدوال
استبدال الجدول بدالة ذات معاملات (كشبكة عصبية) تُقدّر قيم Q لأي حالة، حتى تلك التي لم تُزَر من قبل. هذا ما مكّن DQN من التعامل مع فضاءات حالات هائلة كالصور.
Replacing the table with a parameterized function (like a neural network) that estimates Q-values for any state, even unvisited ones. This is what enabled DQN to handle massive state spaces like images.
Also translated asالتقريب الدالّي، تعميم القيم عبر شبكة عصبية
First appears in this corpus in: Learning to Predict by the Methods of Temporal Differences (1988)
Appears in these papers
- Deep Reinforcement Learning with Double Q-Learning2016in the sky ✦
- Human-Level Control Through Deep Reinforcement Learning2015in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Temporal Difference Learning and TD-Gammon1995in the sky ✦
- Temporal Difference Learning and TD-Gammon1995in the sky ✦
- Learning to Predict by the Methods of Temporal Differences1988in the sky ✦
- TD3: Addressing Function Approximation Error in Actor-Critic Methods2018in the sky ✦
- TD3: Addressing Function Approximation Error in Actor-Critic Methods2018in the sky ✦