المعجم

التسوية المُسبَقة

Pre-LN

ترتيب يُطبَّق فيه تسوية الطبقة قَبل الطبقة الفرعية (الانتباه أو التغذية الأمامية) وليس بعدها. يجعل مسار الاتصال التخطّي نظيفاً ويُحسّن استقرار التدريب في الأعماق الكبيرة. اعتمده GPT-2 وأغلب النماذج اللاحقة.

An arrangement where layer normalization is applied before the sub-layer (attention or FFN) rather than after. Keeps the residual path clean and improves training stability at extreme depths. Adopted by GPT-2 and most subsequent models.

تُرجم أيضاًPre-LayerNorm، تسوية الطبقة قبل الطبقة الفرعية