Layerwise Adaptation
التكييف الطبقي
استراتيجية أمثَلة تُعدّل معدل التعلم أو حجم الخطوة لكل طبقة في الشبكة العصبية بشكل مستقل، بدلاً من استخدام معدل تعلم واحد عام. تستفيد من اختلاف المقاييس والإحصائيات بين الطبقات.
An optimization strategy that adjusts the learning rate or step size for each layer of a neural network independently, rather than using a single global learning rate. Leverages the fact that different layers have different scales and gradient statistics.
Also translated asالتكيُّف حسب الطبقة
First appears in this corpus in: Large Batch Optimization for Deep Learning: Training BERT in 76 Minutes (2019)
Appears in these papers