Tied Embeddings
التضمينات المشتركة
تقنية تُشارك فيها أوزان مصفوفة التضمين بين طبقة المدخلات وطبقة المخرجات في النموذج. يوفّر هذا عدداً كبيراً من المعاملات (3.6 مليار في حالة BLOOM) ويُحسّن التعميم.
A technique where the embedding matrix weights are shared between the input embedding layer and the output projection layer. This saves a large number of parameters (3.6B in BLOOM's case) and improves generalization.
Also translated asربط المدخلات والمخرجات
First appears in this corpus in: The Falcon Series of Open Language Models (2023)
Appears in these papers