AdaGrad
خوارزمية أداغراد
خوارزمية أمثلة تكيُّفية تقسم معدل التعلم على الجذر التربيعي لمجموع مربعات التدرجات السابقة لكل معلمة، مما يفيد البيانات المتناثرة لكنه قد يوقف التعلم.
An adaptive optimizer that divides the learning rate by the square root of the sum of all past squared gradients per parameter, beneficial for sparse data but may halt learning.
Also translated asالتدرج التكيُّفي، AdaGrad
First appears in this corpus in: RMSProp: Divide the Gradient by a Running Average of Its Recent Magnitude (2012)
Appears in these papers