Computer Vision2021intermediate13 min read

Learning Transferable Visual Models from Natural Language Supervision

تعلُّم نماذج بصرية قابلة للنقل عبر الإشراف باللغة الطبيعية

Radford, A. · Kim, J. W. · Hallacy, C. · Ramesh, A. · Goh, G. · Agarwal, S. · Sastry, G. · Askell, A. · Mishkin, P. · Clark, J. · Krueger, G. · Sutskever, I. — ICML

The problem

State-of-the-art vision models were trained on fixed label sets like ImageNet's 1,000 categories. To recognize anything new — a dog breed, a medical scan, a satellite image — you needed to collect new labeled data and fine-tune. This made vision models brittle and expensive to adapt. Meanwhile, NLP had shown that pre-training on raw text yields representations that transfer broadly. Could vision learn the same way — from language itself?

The contribution

CLIP (Contrastive Language-Image Pre-training): a dual- model that jointly trains an image encoder (ResNet or ViT) and a text encoder () on 400 million image-text pairs scraped from the internet. A symmetric contrastive pushes matching pairs together and non-matching pairs apart in a shared . At inference, is as simple as comparing an image against text embeddings of candidate class descriptions — no , no labeled data, no retraining. CLIP matches a fully-supervised ResNet-50 on ImageNet without seeing a single ImageNet training image.

The impact

CLIP is the architectural foundation of the era. It powers the vision tower inside DALL·E, Stable Diffusion, and most text-to-image generators. Open-vocabulary detectors (OWL-ViT, Grounding DINO), segment-anything models (SAM), and the vision encoders of multimodal LLMs all descend from CLIP's idea: align vision and language in one shared space, then let language steer the model.

Traditional image classifiers are like a customs officer with a fixed list of 1,000 items. Show them a passport they know and they'll stamp it fast. Show them anything not on the list — a rare gemstone, a medical instrument — and they're helpless.

CLIP replaces the checklist with fluency in the language of images. Instead of memorizing a fixed list, it learned to read descriptions of what things look like. Now you can ask it anything in plain English — "a photo of a Pembroke Welsh Corgi" — and it understands, even if it has never been explicitly taught that breed.

The problem: vision models that can't see beyond their training labels

By 2020, the dominant recipe in computer vision was: collect a labeled (like ImageNet's 1.28 million images in 1,000 categories), train a CNN or Vision Transformer to predict those labels, then fine-tune it on your target task. This worked — but it had two deep flaws:

  • Fixed vocabulary. The model can only recognize the categories it was trained on. A model trained on ImageNet knows "Labrador retriever" but not "Golden Doodle". Adding a single new class means collecting data and retraining.

  • Expensive supervision. Labeling millions of images costs enormous human effort. ImageNet alone took over 25,000 annotators. Every new domain — medical imaging, satellite photos, art styles — demands a fresh labeling campaign.

Meanwhile, NLP had discovered a shortcut: train on raw text from the internet and let the model learn language structure itself. GPT and BERT needed no manual labels. Could vision follow the same path?

Open in Lab
Compare the traditional pipeline (fixed labels, retraining needed) with CLIP's approach (open vocabulary, zero-shot transfer).
The demo wakes as you arrive…

The fuel: 400 million image-text pairs from the web

To learn from language instead of labels, CLIP needs a massive dataset of images paired with natural descriptions. The authors built WebImageText (WIT): 400 million (image, text) pairs harvested from publicly available internet sources. They started with 500,000 seed queries derived from Wikipedia terms, WordNet synsets, and common bigrams, then collected images and their associated text from the web.

Why 400 million? Earlier attempts to learn from captions used small, curated datasets (like MS-COCO's 330,000 image-caption pairs). At that scale, couldn't compete with ImageNet-trained models. CLIP showed that scaling the data by 1,000× — from hundreds of thousands to hundreds of millions — was the key to making language supervision work. The internet itself became the annotator.

The idea: match images to their descriptions

CLIP's training objective is deceptively simple. Take a batch of NN image-text pairs. Each image has exactly one correct caption, and each caption has exactly one correct image. The model must figure out which image goes with which text — a matching game.

The mechanism: two separate encoders — one for images, one for text — each produce a (an embedding). A matching pair's vectors should point in the same direction; a mismatched pair's vectors should point away. This is : pull correct pairs together, push incorrect pairs apart.

Think of it as a name-tag game at a conference. NN people walk into a room, each wearing a photo on their badge and holding a description card. The model's job is to match every photo badge to the right description card. In a batch of 32,768 pairs, there are 32,768 correct matches and over a billion wrong ones. The model learns by getting better at this matching game across billions of examples.

Open in Lab
Drag image cards to their matching text descriptions. This is exactly what CLIP learns — but with 32,768 pairs at a time.
The demo wakes as you arrive…

Two encoders, one shared space

CLIP has two independent encoders that never share weights but learn to produce embeddings in the same geometric space:

Image encoder — either a ResNet (with pooling replacing global average pooling) or a Vision Transformer (ViT). The image is turned into a single vector. The best variant, ViT-L/14, splits the image into 14×14-pixel patches, processes them through a Transformer encoder, and uses the final as the image embedding.

Text encoder — a standard Transformer (63M parameters, 12 layers, 512-wide, 8 heads) operating on tokenized text with a max context of 76 tokens. The final [EOS] token's activation serves as the text embedding.

Each encoder's output is then linearly projected into a shared embedding space of dimension dd where becomes the measure of match quality.

Open in Lab
Click on any component to see how image and text flow through CLIP's dual encoders.
The demo wakes as you arrive…

The contrastive loss: a symmetric matching game

Given a batch of NN image-text pairs, CLIP computes an N×NN \times N of cosine similarities between every image embedding and every text embedding, scaled by a learnable temperature τ\tau. The correct matches lie on the diagonal. The loss is the average of two cross-entropy terms:

  • Image → text: for each image (row), the correct text is the target among NN candidates.
  • Text → image: for each text (column), the correct image is the target among NN candidates.

The temperature τ\tau is a learnable scalar (initialized at 0.07) that controls how sharp or soft the probability distribution is. A smaller τ\tau makes the model more confident — sharper peaks on the diagonal — while a larger τ\tau spreads attention more evenly.

L=12N[∑i=1N−log⁡esim(Ii,Ti)/τ∑j=1Nesim(Ii,Tj)/τ+∑i=1N−log⁡esim(Ti,Ii)/τ∑j=1Nesim(Ti,Ij)/τ]\mathcal{L} = \frac{1}{2N} \left[ \sum_{i=1}^{N} -\log \frac{e^{\text{sim}(I_i, T_i)/\tau}}{\sum_{j=1}^{N} e^{\text{sim}(I_i, T_j)/\tau}} + \sum_{i=1}^{N} -\log \frac{e^{\text{sim}(T_i, I_i)/\tau}}{\sum_{j=1}^{N} e^{\text{sim}(T_i, I_j)/\tau}} \right]
CLIP's symmetric contrastive loss (InfoNCE) — sim(I,T) = cosine similarity between image and text embeddings · τ = learnable temperature · the first sum asks each image to pick its text · the second sum asks each text to pick its image · the diagonal of the similarity matrix holds the correct answers
Open in Lab
Watch the similarity matrix form as CLIP trains. Correct pairs (diagonal) brighten; incorrect pairs (off-diagonal) darken.
The demo wakes as you arrive…

Why matching beats captioning

The authors tried three training objectives:

  1. Generative captioning — train the model to predict the exact caption word by word (like GPT). This works but is very slow: generating full sentences requires many sequential steps.

  2. Bag-of-words prediction — predict which words appear in the caption, ignoring order. Faster but loses nuance.

  3. — just decide if an image-text pair is correct or not. This turned out to be 4× more efficient than generative captioning in . The insight: the model doesn't need to generate language to understand vision; it just needs to distinguish correct matches from incorrect ones.

This is like the difference between writing a book review and matching a review to its book — the second task is much simpler but teaches you just as much about what each book is about.

Zero-shot classification: language becomes the interface

After pre-training, CLIP can classify images into any set of categories — without seeing a single labeled example — using a three-step process:

  1. Write prompts for each class: "a photo of a {class name}" — e.g. "a photo of a dog", "a photo of a cat", "a photo of a car".

  2. Encode each prompt through the text encoder to get NN text embeddings (one per class).

  3. Encode the test image through the image encoder, compute cosine similarity against all NN text embeddings, and pick the class with the highest similarity.

This is zero-shot transfer: the model performs a task it was never explicitly trained for. The text encoder effectively synthesizes a linear classifier on the fly, using natural language as the specification.

Open in Lab
Type any class descriptions, then watch CLIP rank them against the sample image — no retraining needed.
The demo wakes as you arrive…

Prompt engineering and ensembling

How you phrase the class description matters. Using just the bare class name "dog" performs worse than "a photo of a dog" because CLIP was trained on captions — full sentences — not single words. The mismatch between training text and inference text hurts performance. This is the same problem that appears across machine learning.

CLIP addresses this with : instead of a single prompt per class, use many — "a photo of a big {label}", "a photo of a small {label}", "a centered satellite photo of {label}", etc. — and average their text embeddings. On ImageNet, prompt ensembling improved accuracy by 3.5 percentage points over the single-prompt baseline. Context-specific prompts help too: for satellite imagery, "a satellite photo of {label}" beats the generic template.

This shows that CLIP doesn't just learn visual concepts — it learns the relationship between visual content and the language used to describe it. The prompt is a steering wheel, and small adjustments in phrasing can meaningfully improve results.

The same idea in code

CLIP's contrastive training loop (pseudocode from the paper)python

Simplified to show the idea — not the real implementation.

import numpy as np

# I[n] = image encoder output for the n-th image
# T[n] = text encoder output for the n-th text
# W_i, W_t = learned linear projections
# t = learned temperature parameter

# Step 1: encode and project into shared space
I_e = normalize(I @ W_i, dim=-1)   # (N, d) image embeddings
T_e = normalize(T @ W_t, dim=-1)   # (N, d) text embeddings

# Step 2: compute cosine similarity matrix, scaled by temperature
logits = (I_e @ T_e.T) * np.exp(t)  # (N, N) — element [i,j] = similarity of image i to text j

# Step 3: labels = the diagonal (image i matches text i)
labels = np.arange(N)

# Step 4: symmetric cross-entropy loss
loss_i = cross_entropy(logits, labels, axis=1)   # each image picks its text
loss_t = cross_entropy(logits, labels, axis=0)   # each text picks its image
loss = (loss_i + loss_t) / 2

# That's the entire training signal.
# No bounding boxes. No segmentation masks. No class labels.
# Just: "this image goes with this text."

Results: zero-shot CLIP vs. supervised models

CLIP's zero-shot performance was benchmarked across over 30 datasets spanning image , OCR, action recognition, geolocation, and fine-grained object recognition. The headline result: zero-shot CLIP (ViT-L/14) matched the accuracy of a fully-supervised ResNet-50 on ImageNet — without seeing any ImageNet training image. On some specialized datasets CLIP underperformed, particularly on fine-grained distinctions (car models, flower species) and abstract tasks (counting, distance estimation). But on others — like STL-10, CIFAR-10/100, and several action recognition benchmarks — it surpassed supervised baselines.

The most striking finding was CLIP's robustness: on shifted versions of ImageNet (ImageNet-V2, ImageNet-Sketch, ImageNet-A, ImageNet-R), zero-shot CLIP maintained its accuracy far better than supervised models, closing the gap between standard and shifted accuracy by up to 75%. Supervised models had overfit to ImageNet's specific distribution; CLIP, trained on the diversity of the internet, generalized.

Open in Lab
Compare zero-shot CLIP against supervised ResNet-50 across different dataset categories.
The demo wakes as you arrive…

Linear probes: CLIP's features are among the best

Zero-shot isn't the only way to use CLIP. Freeze the image encoder, extract its embeddings, and train a simple linear classifier on top — a . This tests how good the features are, separate from the zero-shot trick. CLIP's linear probe performance rivaled or exceeded the best self-supervised models (SimCLR, BYOL, MoCo) and fully-supervised models, while being more compute-efficient per quality point — particularly with ViT backbones, which the paper found to be roughly 3× more compute-efficient than ResNet variants.

This means CLIP doesn't just build a clever zero-shot interface — it learns genuinely powerful visual representations that transfer well even through traditional linear evaluation.

Limitations and social considerations

CLIP isn't perfect. The paper is transparent about several limitations:

  • Fine-grained tasks. CLIP struggles with distinctions that require specialized knowledge (car models, flower species, tumor grades). Language descriptions may not capture the subtle visual features that experts look for.

  • Abstract and structural tasks. Counting objects, estimating distance, or understanding spatial relationships remain difficult. CLIP learns appearance, not geometry.

  • from web data. WIT inherits the biases of the internet: demographic under-representation, cultural skew, stereotyped associations between images and text. The authors found that CLIP's zero-shot classifiers can reflect societal biases, particularly in race and gender classification.

  • Task specification via prompts. Zero-shot transfer depends entirely on how well you describe the task in language. For truly novel visual concepts — things that have no common name — CLIP has no advantage over supervised baselines.

Why it mattered

  1. 2013

    Word2Vec

    Word embeddings showed that language can be projected into a continuous vector space where geometry encodes meaning. CLIP does the same across modalities.

  2. 2020

    ViT — pure Transformer for images

    Proved Transformers can replace CNNs for vision. CLIP later showed ViT is ~3× more compute-efficient than ResNets as CLIP's image encoder.

  3. 2021

    CLIP & DALL·E — vision meets language

    Released simultaneously by OpenAI. CLIP encodes vision-language alignment; DALL·E decodes it into image generation. Together they launched the multimodal AI revolution.

  4. 2022

    Stable Diffusion

    Uses CLIP's text encoder to guide image generation through diffusion models. CLIP embeddings became the standard conditioning signal for text-to-image systems.

  5. 2023

    SAM (Segment Anything)

    CLIP-style vision encoders power open-vocabulary segmentation: point at anything and describe it in language to get a pixel-perfect mask.

  6. 2024

    Multimodal LLMs

    GPT-4V, Claude, Gemini — all use CLIP-descended vision encoders to see images. CLIP's shared embedding space is the reason language models can understand photos.

CLIP didn't just achieve strong benchmarks — it changed how vision models are built. Before CLIP, vision and language were separate kingdoms. After CLIP, they share a common currency: a shared embedding space where images and words live side by side.

CitationRadford, Kim, Hallacy, Ramesh, Goh, Agarwal, Sastry, Askell, Mishkin, Clark, Krueger, Sutskever. Learning Transferable Visual Models from Natural Language Supervision. ICML, 2021.

Terms in this paper