Glossary

Multimodal Fusion

الدمج متعدد الوسائط

عملية الجمع بين تمثيلات من وسائط مختلفة (كالصور والنصوص) في تمثيل موحد يستفيد من المعلومات في كلتا الوسيطتين. في VQA، تُدمج سمات الصورة مع سمات السؤال للتنبؤ بالإجابة.

The process of combining representations from different modalities (such as images and text) into a unified representation that leverages information from both. In VQA, image features are fused with question features to predict the answer.

Also translated asدمج الوسائط، الاندماج متعدد الأنماط

First appears in this corpus in: VQA: Visual Question Answering (2015)

Appears in these papers