Patch
رُقعة
مربع صغير من البكسلات يُستخرج من الصورة ويُعامَل كرمز مدخل واحد في محوِّل الرؤية.
A small square of pixels extracted from an image, treated as a single input token in a Vision Transformer.
Also translated asجزء مُقتطع، مربع صوري
First appears in this corpus in: Unsupervised Visual Representation Learning by Context Prediction (2015)
Appears in these papers
- Unsupervised Visual Representation Learning by Context Prediction2015in the sky ✦
- A ConvNet for the 2020s2022in the sky ✦
- Masked Autoencoders Are Scalable Vision Learners2022in the sky ✦
- A Time Series Is Worth 64 Words: Long-Term Forecasting with Transformers2023in the sky ✦
- A Time Series Is Worth 64 Words: Long-Term Forecasting with Transformers2023in the sky ✦
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control2023in the sky ✦
- Video Generation Models as World Simulators2024in the sky ✦
- Video Generation Models as World Simulators2024in the sky ✦
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?2024in the sky ✦
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?2024in the sky ✦
- Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows2021in the sky ✦
- Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows2021in the sky ✦
- A Decoder-Only Foundation Model for Time Series Forecasting2024in the sky ✦