Pre-normalization
التسوية المسبقة
تطبيق التسوية على مدخلات كل طبقة فرعية بدلاً من مخرجاتها. يحسّن استقرار التدريب للشبكات العميقة جداً. اعتُمد في GPT-3 وLLaMA.
Applying normalization to the input of each sub-layer rather than the output. Improves training stability for very deep networks. Adopted in GPT-3 and LLaMA.
Also translated asالتطبيع المسبق، التسوية قبل الطبقة
First appears in this corpus in: Layer Normalization (2016)
Appears in these papers