Post-Normalization
التسوية بعد التنشيط
ترتيب يُطبَّق فيه تسوية الطبقات بعد كتلة الانتباه أو الشبكة الأمامية. كان المعيار في GPT-1 والمحوِّل الأصلي لكنه قد يسبب عدم استقرار في التدريب عند النطاقات الكبيرة.
An arrangement where layer normalization is applied after the attention or feed-forward block. Was the standard in GPT-1 and the original Transformer but can cause training instability at large scale.
Also translated asالتسوية اللاحقة، تسوية ما بعد الكتلة
First appears in this corpus in: Layer Normalization (2016)
Appears in these papers