Glossary

Vision Encoder

مرمِّز بصري

مكوِّن شبكة عصبية يحوِّل الصور الخام إلى تمثيلات سمات مهيكلة، وعادةً يكون محوِّل رؤية (ViT) أو شبكة التفافية.

A neural network component that converts raw images into structured feature representations, typically a Vision Transformer (ViT) or convolutional neural network.

Also translated asالمرمِّز البصري، مرمِّز الرؤية

First appears in this corpus in: Flamingo: a Visual Language Model for Few-Shot Learning (2022)

Appears in these papers