Glossary

Visual Language Model

النموذج اللغوي البصري

نموذج يجمع بين فهم الصور أو الفيديو وتوليد النصوص، حيث يستقبل مدخلات بصرية ونصية معاً وينتج ناتجاً نصياً مشروطاً بكلا الوسيطتين.

A model that combines visual understanding (images or video) with text generation, accepting both visual and textual inputs and producing text output conditioned on both modalities.

Also translated asنموذج لغوي بصري، نموذج الرؤية واللغة

First appears in this corpus in: Flamingo: a Visual Language Model for Few-Shot Learning (2022)

Appears in these papers