Visual Tokens
رموز بصرية
تمثيلات رُقَع الصورة التي أُسقطت في فضاء التضمين ذاته الذي تعيش فيه الرموز النصية، مما يتيح للنموذج اللغوي معالجة الصور والنصوص في تيار موحّد.
Image patch representations projected into the same embedding space as text tokens, allowing the language model to process images and text in a single unified stream.
Also translated asالتمثيلات البصرية المرمّزة، الرموز المرئية
First appears in this corpus in: BEiT: BERT Pre-Training of Image Transformers (2021)
Appears in these papers