REINFORCE
خوارزمية REINFORCE
أول خوارزمية عملية تُمكّن الشبكات العصبية من التعلّم عبر المكافآت فقط دون إجابات صحيحة. الفكرة: جرّب أفعالًا عشوائية، وإن حصلت على مكافأة جيدة، عدّل الأوزان لجعل تسلسل الأفعال الذي أدّى لهذه المكافأة أكثر احتمالًا في المستقبل. رياضيًّا: اضرب المكافأة في تدرّج لوغاريتم احتمال الفعل المتّخذ.
The first practical policy gradient algorithm, updating weights by multiplying the reward by the gradient of the log-probability of the action taken, without requiring an environment model.
Also translated asخوارزمية التعزيز الإحصائي، مُقدِّر دالّة الرصيد السياسي، تعزيز ويليامز
First appears in this corpus in: Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning (1992)
Appears in these papers
- Mastering the Game of Go with Deep Neural Networks and Tree Search2016in the sky ✦
- Policy Gradient Methods for Reinforcement Learning with Function Approximation1999in the sky ✦
- Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning1992in the sky ✦
- RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback2023in the sky ✦