Glossary

Image Encoder

مرمِّز الصور

شبكة عصبية تعالج صورة مُدخلة وتُنتج تمثيلاً كثيفاً للسمات. في SAM يعمل محوِّل رؤية ضخم (ViT-H) مُدرَّب مسبقاً بأسلوب المُرمِّز التلقائي المُقنَّع كمرمِّز صور.

A neural network that processes an input image and produces a dense feature representation (embedding). In SAM, a ViT-H pre-trained with MAE serves as the image encoder.

Also translated asمرمِّز الصورة، وحدة ترميز الصورة

First appears in this corpus in: ALIGN: Scaling Up Visual and Vision-Language Representation Learning with Noisy Text Supervision (2021)

Appears in these papers