Multimodal
متعدد الوسائط
نموذج يستطيع معالجة أنواع متعددة من المدخلات — كالنص والصور والصوت — بدلاً من النص وحده، مما يمنحه فهماً أوسع للعالم.
A model that can process multiple types of input — such as text, images, and audio — rather than text alone, giving it a broader understanding of the world.
Also translated asشامل الوسائط، عابر الوسائط، عابر للوسائط، متعدد الأنماط، مُتعدِّد الأنماط
First appears in this corpus in: VQA: Visual Question Answering (2015)
Appears in these papers
- BLIP-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models2023in the sky ✦
- Learning Transferable Visual Models from Natural Language Supervision2021in the sky ✦
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion2023in the sky ✦
- Flamingo: a Visual Language Model for Few-Shot Learning2022in the sky ✦
- Gemini: A Family of Highly Capable Multimodal Models2023in the sky ✦
- GPT-4 Technical Report2023in the sky ✦
- HuggingGPT: Solving AI Tasks with ChatGPT and Its Friends in Hugging Face2023in the sky ✦
- ImageBind: One Embedding Space to Bind Them All2023in the sky ✦
- Visual Instruction Tuning2023in the sky ✦
- Perceiver: General Perception with Iterative Attention2021in the sky ✦
- VQA: Visual Question Answering2015in the sky ✦