Glossary

Large Multimodal Model

نموذج كبير متعدد الوسائط

نموذج ذكاء اصطناعي واسع النطاق قادر على فهم وتوليد المحتوى عبر وسائط متعددة كالنص والصورة، يجمع عادةً بين مرمِّز بصري ونموذج لغوي كبير.

A large-scale AI model capable of understanding and generating content across multiple modalities such as text and images, typically combining a vision encoder with a large language model.

Also translated asLMM، نموذج لغوي-بصري كبير

First appears in this corpus in: Visual Instruction Tuning (2023)

Appears in these papers