Computer Vision2021intermediate13 min read
Learning Transferable Visual Models from Natural Language Supervision
تعلُّم نماذج بصرية قابلة للنقل عبر الإشراف باللغة الطبيعية
Radford, A. · Kim, J. W. · Hallacy, C. · Ramesh, A. · Goh, G. · Agarwal, S. · Sastry, G. · Askell, A. · Mishkin, P. · Clark, J. · Krueger, G. · Sutskever, I. — ICML
The problem
State-of-the-art vision models were trained on fixed label sets like ImageNet's 1,000 categories. To recognize anything new — a dog breed, a medical scan, a satellite image — you needed to collect new labeled data and fine-tune. This made vision models brittle and expensive to adapt. Meanwhile, NLP had shown that pre-training on raw text yields representations that transfer broadly. Could vision learn the same way — from language itself?
The contribution
CLIP (Contrastive Language-Image Pre-training): a dual- model that jointly trains an image encoder (ResNet or ViT) and a text encoder () on 400 million image-text pairs scraped from the internet. A symmetric contrastive pushes matching pairs together and non-matching pairs apart in a shared . At inference, is as simple as comparing an image against text embeddings of candidate class descriptions — no , no labeled data, no retraining. CLIP matches a fully-supervised ResNet-50 on ImageNet without seeing a single ImageNet training image.
The impact
CLIP is the architectural foundation of the era. It powers the vision tower inside DALL·E, Stable Diffusion, and most text-to-image generators. Open-vocabulary detectors (OWL-ViT, Grounding DINO), segment-anything models (SAM), and the vision encoders of multimodal LLMs all descend from CLIP's idea: align vision and language in one shared space, then let language steer the model.
Traditional image classifiers are like a customs officer with a fixed list of 1,000 items. Show them a passport they know and they'll stamp it fast. Show them anything not on the list — a rare gemstone, a medical instrument — and they're helpless.
CLIP replaces the checklist with fluency in the language of images. Instead of memorizing a fixed list, it learned to read descriptions of what things look like. Now you can ask it anything in plain English — "a photo of a Pembroke Welsh Corgi" — and it understands, even if it has never been explicitly taught that breed.
The problem: vision models that can't see beyond their training labels
By 2020, the dominant recipe in computer vision was: collect a labeled (like ImageNet's 1.28 million images in 1,000 categories), train a CNN or Vision Transformer to predict those labels, then fine-tune it on your target task. This worked — but it had two deep flaws:
-
Fixed vocabulary. The model can only recognize the categories it was trained on. A model trained on ImageNet knows "Labrador retriever" but not "Golden Doodle". Adding a single new class means collecting data and retraining.
-
Expensive supervision. Labeling millions of images costs enormous human effort. ImageNet alone took over 25,000 annotators. Every new domain — medical imaging, satellite photos, art styles — demands a fresh labeling campaign.
Meanwhile, NLP had discovered a shortcut: train on raw text from the internet and let the model learn language structure itself. GPT and BERT needed no manual labels. Could vision follow the same path?
The fuel: 400 million image-text pairs from the web
To learn from language instead of labels, CLIP needs a massive dataset of images paired with natural descriptions. The authors built WebImageText (WIT): 400 million (image, text) pairs harvested from publicly available internet sources. They started with 500,000 seed queries derived from Wikipedia terms, WordNet synsets, and common bigrams, then collected images and their associated text from the web.
Why 400 million? Earlier attempts to learn from captions used small, curated datasets (like MS-COCO's 330,000 image-caption pairs). At that scale, couldn't compete with ImageNet-trained models. CLIP showed that scaling the data by 1,000× — from hundreds of thousands to hundreds of millions — was the key to making language supervision work. The internet itself became the annotator.
The idea: match images to their descriptions
CLIP's training objective is deceptively simple. Take a batch of image-text pairs. Each image has exactly one correct caption, and each caption has exactly one correct image. The model must figure out which image goes with which text — a matching game.
The mechanism: two separate encoders — one for images, one for text — each produce a (an embedding). A matching pair's vectors should point in the same direction; a mismatched pair's vectors should point away. This is : pull correct pairs together, push incorrect pairs apart.
Think of it as a name-tag game at a conference. people walk into a room, each wearing a photo on their badge and holding a description card. The model's job is to match every photo badge to the right description card. In a batch of 32,768 pairs, there are 32,768 correct matches and over a billion wrong ones. The model learns by getting better at this matching game across billions of examples.
Two encoders, one shared space
CLIP has two independent encoders that never share weights but learn to produce embeddings in the same geometric space:
Image encoder — either a ResNet (with pooling replacing global average pooling) or a Vision Transformer (ViT). The image is turned into a single vector. The best variant, ViT-L/14, splits the image into 14×14-pixel patches, processes them through a Transformer encoder, and uses the final as the image embedding.
Text encoder — a standard Transformer (63M parameters, 12 layers, 512-wide, 8 heads) operating on tokenized text with a max context of 76 tokens. The final [EOS] token's activation serves as the text embedding.
Each encoder's output is then linearly projected into a shared embedding space of dimension where becomes the measure of match quality.
The contrastive loss: a symmetric matching game
Given a batch of image-text pairs, CLIP computes an of cosine similarities between every image embedding and every text embedding, scaled by a learnable temperature . The correct matches lie on the diagonal. The loss is the average of two cross-entropy terms:
- Image → text: for each image (row), the correct text is the target among candidates.
- Text → image: for each text (column), the correct image is the target among candidates.
The temperature is a learnable scalar (initialized at 0.07) that controls how sharp or soft the probability distribution is. A smaller makes the model more confident — sharper peaks on the diagonal — while a larger spreads attention more evenly.
Why matching beats captioning
The authors tried three training objectives:
-
Generative captioning — train the model to predict the exact caption word by word (like GPT). This works but is very slow: generating full sentences requires many sequential steps.
-
Bag-of-words prediction — predict which words appear in the caption, ignoring order. Faster but loses nuance.
-
— just decide if an image-text pair is correct or not. This turned out to be 4× more efficient than generative captioning in . The insight: the model doesn't need to generate language to understand vision; it just needs to distinguish correct matches from incorrect ones.
This is like the difference between writing a book review and matching a review to its book — the second task is much simpler but teaches you just as much about what each book is about.
Zero-shot classification: language becomes the interface
After pre-training, CLIP can classify images into any set of categories — without seeing a single labeled example — using a three-step process:
-
Write prompts for each class: "a photo of a
{class name}" — e.g. "a photo of a dog", "a photo of a cat", "a photo of a car". -
Encode each prompt through the text encoder to get text embeddings (one per class).
-
Encode the test image through the image encoder, compute cosine similarity against all text embeddings, and pick the class with the highest similarity.
This is zero-shot transfer: the model performs a task it was never explicitly trained for. The text encoder effectively synthesizes a linear classifier on the fly, using natural language as the specification.
Prompt engineering and ensembling
How you phrase the class description matters. Using just the bare class name "dog" performs worse than "a photo of a dog" because CLIP was trained on captions — full sentences — not single words. The mismatch between training text and inference text hurts performance. This is the same problem that appears across machine learning.
CLIP addresses this with : instead of a single prompt per class, use many — "a photo of a big {label}", "a photo of a small {label}", "a centered satellite photo of {label}", etc. — and average their text embeddings. On ImageNet, prompt ensembling improved accuracy by 3.5 percentage points over the single-prompt baseline. Context-specific prompts help too: for satellite imagery, "a satellite photo of {label}" beats the generic template.
This shows that CLIP doesn't just learn visual concepts — it learns the relationship between visual content and the language used to describe it. The prompt is a steering wheel, and small adjustments in phrasing can meaningfully improve results.
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
# I[n] = image encoder output for the n-th image
# T[n] = text encoder output for the n-th text
# W_i, W_t = learned linear projections
# t = learned temperature parameter
# Step 1: encode and project into shared space
I_e = normalize(I @ W_i, dim=-1) # (N, d) image embeddings
T_e = normalize(T @ W_t, dim=-1) # (N, d) text embeddings
# Step 2: compute cosine similarity matrix, scaled by temperature
logits = (I_e @ T_e.T) * np.exp(t) # (N, N) — element [i,j] = similarity of image i to text j
# Step 3: labels = the diagonal (image i matches text i)
labels = np.arange(N)
# Step 4: symmetric cross-entropy loss
loss_i = cross_entropy(logits, labels, axis=1) # each image picks its text
loss_t = cross_entropy(logits, labels, axis=0) # each text picks its image
loss = (loss_i + loss_t) / 2
# That's the entire training signal.
# No bounding boxes. No segmentation masks. No class labels.
# Just: "this image goes with this text."Results: zero-shot CLIP vs. supervised models
CLIP's zero-shot performance was benchmarked across over 30 datasets spanning image , OCR, action recognition, geolocation, and fine-grained object recognition. The headline result: zero-shot CLIP (ViT-L/14) matched the accuracy of a fully-supervised ResNet-50 on ImageNet — without seeing any ImageNet training image. On some specialized datasets CLIP underperformed, particularly on fine-grained distinctions (car models, flower species) and abstract tasks (counting, distance estimation). But on others — like STL-10, CIFAR-10/100, and several action recognition benchmarks — it surpassed supervised baselines.
The most striking finding was CLIP's robustness: on shifted versions of ImageNet (ImageNet-V2, ImageNet-Sketch, ImageNet-A, ImageNet-R), zero-shot CLIP maintained its accuracy far better than supervised models, closing the gap between standard and shifted accuracy by up to 75%. Supervised models had overfit to ImageNet's specific distribution; CLIP, trained on the diversity of the internet, generalized.
Linear probes: CLIP's features are among the best
Zero-shot isn't the only way to use CLIP. Freeze the image encoder, extract its embeddings, and train a simple linear classifier on top — a . This tests how good the features are, separate from the zero-shot trick. CLIP's linear probe performance rivaled or exceeded the best self-supervised models (SimCLR, BYOL, MoCo) and fully-supervised models, while being more compute-efficient per quality point — particularly with ViT backbones, which the paper found to be roughly 3× more compute-efficient than ResNet variants.
This means CLIP doesn't just build a clever zero-shot interface — it learns genuinely powerful visual representations that transfer well even through traditional linear evaluation.
Limitations and social considerations
CLIP isn't perfect. The paper is transparent about several limitations:
-
Fine-grained tasks. CLIP struggles with distinctions that require specialized knowledge (car models, flower species, tumor grades). Language descriptions may not capture the subtle visual features that experts look for.
-
Abstract and structural tasks. Counting objects, estimating distance, or understanding spatial relationships remain difficult. CLIP learns appearance, not geometry.
-
from web data. WIT inherits the biases of the internet: demographic under-representation, cultural skew, stereotyped associations between images and text. The authors found that CLIP's zero-shot classifiers can reflect societal biases, particularly in race and gender classification.
-
Task specification via prompts. Zero-shot transfer depends entirely on how well you describe the task in language. For truly novel visual concepts — things that have no common name — CLIP has no advantage over supervised baselines.
Why it mattered
2013
Word2Vec
Word embeddings showed that language can be projected into a continuous vector space where geometry encodes meaning. CLIP does the same across modalities.
2020
ViT — pure Transformer for images
Proved Transformers can replace CNNs for vision. CLIP later showed ViT is ~3× more compute-efficient than ResNets as CLIP's image encoder.
2021
CLIP & DALL·E — vision meets language
Released simultaneously by OpenAI. CLIP encodes vision-language alignment; DALL·E decodes it into image generation. Together they launched the multimodal AI revolution.
2022
Stable Diffusion
Uses CLIP's text encoder to guide image generation through diffusion models. CLIP embeddings became the standard conditioning signal for text-to-image systems.
2023
SAM (Segment Anything)
CLIP-style vision encoders power open-vocabulary segmentation: point at anything and describe it in language to get a pixel-perfect mask.
2024
Multimodal LLMs
GPT-4V, Claude, Gemini — all use CLIP-descended vision encoders to see images. CLIP's shared embedding space is the reason language models can understand photos.
CLIP didn't just achieve strong benchmarks — it changed how vision models are built. Before CLIP, vision and language were separate kingdoms. After CLIP, they share a common currency: a shared embedding space where images and words live side by side.
CitationRadford, Kim, Hallacy, Ramesh, Goh, Agarwal, Sastry, Askell, Mishkin, Clark, Krueger, Sutskever. Learning Transferable Visual Models from Natural Language Supervision. ICML, 2021.
Terms in this paper
- Contrastive Learningالتعلم التبايُني
- Zero-Shot Transferنقل بدون تدريب
- Zero-Shot Classificationالتصنيف بدون أمثلة
- Cosine Similarityتشابه جيب التمام
- Dual Encoderالمُرمِّز الثنائي
- Embedding Spaceفضاء التضمين
- Natural Language Supervisionالإشراف بِاللغة الطبيعية
- Prompt Engineeringهندسة التحفيز
- Prompt Ensemblingتجميع المُحفِّزات
- Linear Probeالمسبار الخطي
- Distribution Shiftانزياح التوزيع
- Temperature Parameterمعامل الحرارة
- InfoNCE Lossخسارة InfoNCE
- Multimodalمتعدد الوسائط
- Vision-Language Modelنموذج الرؤية واللغة