Soft Update
التحديث الناعم
أسلوب لتحديث الشبكات الهدفية حيث تُنقل الأوزان الجديدة بنسبة صغيرة τ فقط في كل خطوة: θ' ← τθ + (1−τ)θ'. يمنع القفزات الحادة في أهداف التعلّم ويُبقي التدريب مستقراً.
A method for updating target networks where new weights are blended in by a small fraction τ each step: θ' ← τθ + (1−τ)θ'. Prevents sharp jumps in learning targets and keeps training stable.
Also translated asالتحديث التدريجي، التحديث السلس
First appears in this corpus in: Continuous Control with Deep Reinforcement Learning (2015)
Appears in these papers
- Continuous Control with Deep Reinforcement Learning2015in the sky ✦
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor2018in the sky ✦
- TD3: Addressing Function Approximation Error in Actor-Critic Methods2018in the sky ✦
- TD3: Addressing Function Approximation Error in Actor-Critic Methods2018in the sky ✦