Joint Embedding
تضمين مشترك
فضاء تمثيل مشترك تُسقَط فيه وسائط مختلفة (مثل الصوت والنص) بحيث تقع العناصر المتطابقة دلالياً قريبة من بعضها. يُدرَّب عادةً بالتعلم التبايُني.
A shared representation space into which different modalities (e.g., audio and text) are projected so that semantically matching items land close together. Typically trained with contrastive learning.
Also translated asتضمين متعدد الوسائط
First appears in this corpus in: MusicLM: Generating Music From Text (2023)
Appears in these papers