Vision Encoder
مرمِّز بصري
مكوِّن شبكة عصبية يحوِّل الصور الخام إلى تمثيلات سمات مهيكلة، وعادةً يكون محوِّل رؤية (ViT) أو شبكة التفافية.
A neural network component that converts raw images into structured feature representations, typically a Vision Transformer (ViT) or convolutional neural network.
Also translated asالمرمِّز البصري، مرمِّز الرؤية
First appears in this corpus in: Flamingo: a Visual Language Model for Few-Shot Learning (2022)
Appears in these papers
- BLIP-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models2023in the sky ✦
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion2023in the sky ✦
- Flamingo: a Visual Language Model for Few-Shot Learning2022in the sky ✦
- Flamingo: a Visual Language Model for Few-Shot Learning2022in the sky ✦
- Gemini: A Family of Highly Capable Multimodal Models2023in the sky ✦
- Gemini: A Family of Highly Capable Multimodal Models2023in the sky ✦
- Visual Instruction Tuning2023in the sky ✦
- RT-1: Robotics Transformer for Real-World Control at Scale2022in the sky ✦
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control2023in the sky ✦