Glossary

Vision Transformer (ViT)

محوِّل الرؤية (ViT)

تكييف بنية المحولات اللغوية لمعالجة الصور البصرية عبر تقسيم المشهد إلى مربعات صغيرة والتعامل معها ككلمات.

An architecture that processes images by splitting them into patches treated as tokens and feeding them into a standard Transformer encoder with no convolutions.

Also translated asبنية المحولات المخصصة لمعالجة الصور، محول البيانات البصرية الشبكي، محوِّل الصور، ViT، محوِّل الرؤية (ViT)

First appears in this corpus in: ImageNet: A Large-Scale Hierarchical Image Database (2009)

Appears in these papers