التسوية المُسبَقة
Pre-LN
ترتيب يُطبَّق فيه تسوية الطبقة قَبل الطبقة الفرعية (الانتباه أو التغذية الأمامية) وليس بعدها. يجعل مسار الاتصال التخطّي نظيفاً ويُحسّن استقرار التدريب في الأعماق الكبيرة. اعتمده GPT-2 وأغلب النماذج اللاحقة.
An arrangement where layer normalization is applied before the sub-layer (attention or FFN) rather than after. Keeps the residual path clean and improves training stability at extreme depths. Adopted by GPT-2 and most subsequent models.
تُرجم أيضاًPre-LayerNorm، تسوية الطبقة قبل الطبقة الفرعية