Vision Transformer (ViT)
محوِّل الرؤية (ViT)
تكييف بنية المحولات اللغوية لمعالجة الصور البصرية عبر تقسيم المشهد إلى مربعات صغيرة والتعامل معها ككلمات.
An architecture that processes images by splitting them into patches treated as tokens and feeding them into a standard Transformer encoder with no convolutions.
Also translated asبنية المحولات المخصصة لمعالجة الصور، محول البيانات البصرية الشبكي، محوِّل الصور، ViT، محوِّل الرؤية (ViT)
First appears in this corpus in: ImageNet: A Large-Scale Hierarchical Image Database (2009)
Appears in these papers
- BEiT: BERT Pre-Training of Image Transformers2021in the sky ✦
- BEiT: BERT Pre-Training of Image Transformers2021in the sky ✦
- Dynamic Routing Between Capsules2017in the sky ✦
- A ConvNet for the 2020s2022in the sky ✦
- Training Data-Efficient Image Transformers & Distillation Through Attention2021in the sky ✦
- Training Data-Efficient Image Transformers & Distillation Through Attention2021in the sky ✦
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data2024in the sky ✦
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data2024in the sky ✦
- Emerging Properties in Self-Supervised Vision Transformers2021in the sky ✦
- Emerging Properties in Self-Supervised Vision Transformers2021in the sky ✦
- DINOv2: Learning Robust Visual Features Without Supervision2023in the sky ✦
- DINOv2: Learning Robust Visual Features Without Supervision2023in the sky ✦
- A Generalist Agent2022in the sky ✦
- Gemini: A Family of Highly Capable Multimodal Models2023in the sky ✦
- Group Normalization2018in the sky ✦
- ImageBind: One Embedding Space to Bind Them All2023in the sky ✦
- ImageBind: One Embedding Space to Bind Them All2023in the sky ✦
- ImageNet: A Large-Scale Hierarchical Image Database2009in the sky ✦
- Symbolic Discovery of Optimization Algorithms2023in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- Momentum Contrast for Unsupervised Visual Representation Learning2020in the sky ✦
- SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers2021in the sky ✦
- SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers2021in the sky ✦
- Segment Anything2023in the sky ✦
- SGDR: Stochastic Gradient Descent with Warm Restarts2017in the sky ✦
- Show and Tell: A Neural Image Caption Generator2015in the sky ✦
- Video Generation Models as World Simulators2024in the sky ✦
- Video Generation Models as World Simulators2024in the sky ✦
- Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows2021in the sky ✦
- Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows2021in the sky ✦
- An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale2020in the sky ✦
- An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale2020in the sky ✦