Multimodal AI2021intermediate12 min read
ALIGN: Scaling Up Visual and Vision-Language Representation Learning with Noisy Text Supervision
ALIGN: توسيع نطاق تعلّم التمثيلات البصرية والبصرية-اللغوية بإشراف نصّي مُشوَّش
Jia, C. · Yang, Y. · Xia, Y. · Chen, Y.-T. · Parekh, Z. · Pham, H. · Le, Q. V. · Sung, Y. · Li, Z. · Duerig, T. — ICML
The problem
By 2021, learning visual representations required expensive, human-curated datasets like ImageNet (14M images, years to label). Vision-language datasets like Conceptual Captions demanded even heavier cleaning — semantic parsing, filtering, balancing — keeping them at roughly 3–10M examples. Meanwhile, NLP had already moved to on raw, uncurated internet text at massive scale. The question was: could vision and vision-language follow the same path, or would noisy data ruin the representations?
The contribution
ALIGN demonstrates that scale beats curation. Using over 1.8 billion noisy image alt-text pairs with only minimal frequency-based filtering — no semantic parsing, no human annotation — a simple dual- architecture trained with achieves state-of-the-art results on Flickr30K and MSCOCO retrieval (outperforming even complex models), 76.4% zero-shot ImageNet accuracy, and 88.64% fine-tuned ImageNet accuracy. The key insight: at sufficient scale, noisy data with only 4× more examples already outperforms carefully cleaned data.
The impact
ALIGN validated the "noisy data at scale" paradigm for vision-language learning, complementing CLIP's concurrent approach of curated data with an allowlist. Together, they launched the modern era of contrastive vision-language pre-training. ALIGN's insight that raw alt-text data suffices directly influenced ImageBind (multi-modal binding across six modalities), Flamingo (few-shot visual language models), and the broader trend of training foundation models on web-scale noisy data without manual curation.
Imagine two approaches to training a museum tour guide. The traditional way: hire art historians to write a precise label for every painting — accurate but painfully slow, and you'll only cover a few thousand works.
ALIGN's way: let the guide wander through every gallery in the world, reading whatever placard is posted next to each painting — some placards are wrong, some are vague, some describe the frame instead of the art. But after seeing a billion paintings with their messy placards, the guide develops an extraordinary intuition for matching art to descriptions. The sheer volume of experience compensates for the imperfections.
That is the core thesis of ALIGN: you don't need perfect labels if you have enough imperfect ones.
The data bottleneck: why curation doesn't scale
By 2021 the recipe for visual was clear: pre-train on a large labeled dataset, then fine-tune on downstream tasks. The problem was that "large labeled dataset" meant either ImageNet (14M images with manually verified class labels) or JFT-300M (Google's internal dataset with 300M labeled images). Both required years of human effort.
For vision-language tasks the situation was worse. Conceptual Captions, the standard dataset, started with billions of raw image-caption pairs from the web — then applied heavy semantic parsing, text cleaning, image-based filtering, and manual quality checks to produce a mere 3.3 million clean pairs. This process threw away 99.7% of the data. The quality was high, but the scale was tiny compared to what NLP models were consuming.
ALIGN asks a radical question: what if we skip almost all the cleaning and just use the noisy data? The hypothesis is that at sufficient scale, the signal in billions of imperfect pairs drowns out the noise.
ALIGN's dataset starts from raw English alt-text data — the text that web developers write in the alt attribute of <img> tags. This text is often noisy: generic strings like "thumbnail", file names like "IMG_2847.jpg", or descriptions of the frame rather than the content. The paper applies only minimal filtering: remove images smaller than 200 pixels, exclude alt-texts shared by more than 10 images, discard texts with rare tokens, and filter texts shorter than 3 words or longer than 20. No semantic parsing, no human review. The result: 1.8 billion image-text pairs — nearly 600× the size of Conceptual Captions.
Architecture: the simplest thing that could work
ALIGN uses a dual-encoder architecture — deliberately simple. There are two independent encoders that never look at each other's inputs:
: (up to the L2 variant) with . The pooled features pass through a linear projection to produce a fixed-size vector. No region features, no object detectors, no attention over patches — just the whole image compressed into a single vector.
: (up to Large) with a 100K wordpiece vocabulary generated from the training data. The [CLS] token embedding passes through a linear layer to match the image embedding dimension.
Both embeddings are L2-normalized and live in a shared space. Matching is done by . This architecture is orders of magnitude faster at retrieval than cross-attention models like UNITER or ViLBERT, because each image and text can be encoded independently and compared via a simple dot product.
The contrastive objective: learning by comparison
The training objective is elegant in its simplicity. Given a batch of N image-text pairs, ALIGN treats each matched pair as a positive example and all other N−1 pairings as negatives. Two symmetric losses are computed:
Image-to-text: for each image, classify which of the N texts is its match. Text-to-image: for each text, classify which of the N images is its match.
Think of it as a conference room with N people, each holding a photo and a caption. Everyone's task is to find whose caption matches their photo, and vice versa. The more people in the room (larger ), the harder the task and the better the learned representations.
A critical training detail is the batch size. ALIGN uses 16,384 pairs per batch, spread across 1,024 TPUv3 cores (16 pairs per core). Embeddings from all cores are concatenated to form the full batch for the contrastive loss. This massive batch means each image is compared against 16,383 negative texts — a much harder classification task that produces better embeddings.
The learned σ converges to approximately 1/64. A smaller σ sharpens the softmax distribution, making the model more discriminative. The paper also applies (0.1) to the softmax losses, which prevents the model from being overconfident and helps with noisy labels.
Scaling: bigger data, bigger models, better results
The reveals clean scaling laws. The paper sweeps image encoders from EfficientNet-B1 to L2 and text encoders from BERT-Mini to BERT-Large. Results improve smoothly with both axes — no diminishing returns, no collapse.
For vision-only tasks like ImageNet KNN, the image encoder size dominates: even with BERT-Mini as the text encoder, EfficientNet-L2 outperforms B7 with BERT-Large. For cross-modal tasks like retrieval, both encoders matter equally. This makes intuitive sense: retrieving the right caption requires understanding both the image content and the linguistic nuance.
The embedding dimension also matters. Going from 80 to 640 dimensions steadily improves performance, so the final model lets the dimension scale with the (L2 uses 1,376-dimensional embeddings). The number of in-batch negatives is equally important — using only 25% of the batch as negatives degrades retrieval R@1 by about 3 points.
Zero-shot transfer and downstream tasks
Once the shared is learned, becomes trivial. To classify an image into one of K classes, feed each class name through the text encoder (e.g. "A photo of a dog"), compute cosine similarity with the image embedding, and pick the highest-scoring class. No training on the target task at all.
ALIGN achieves 76.4% top-1 zero-shot accuracy on ImageNet — comparable to CLIP's 76.2% despite using noisier data. On ImageNet-R (non-natural images like cartoons and sketches), ALIGN reaches 92.2% vs CLIP's 88.9%, showing that noisy training data may actually improve robustness to distribution shifts.
For retrieval, ALIGN sets new records on both Flickr30K and MSCOCO, outperforming even cross-attention models that are orders of magnitude slower. With , ALIGN reaches 95.3% R@1 on Flickr30K image-to-text retrieval — a 7+ point jump over the previous best.
Emergent compositionality: searching with image + text
One of the most striking properties of ALIGN's embedding space is . Because images and text live in the same vector space, you can add or subtract text embeddings from image embeddings to modify the query semantically.
Want to find images of "pandas in Australia"? Take the embedding of a panda image, add the text embedding of "Australia", and retrieve the nearest images. The results show koala-like animals in Australian landscapes — the embedding space has learned to compose visual and linguistic concepts algebraically, much like the famous word2vec analogies ("king − man + woman = queen") but across modalities.
You can also subtract concepts: take a street scene, subtract "cars", and the retrieval returns similar streets but without vehicles. This emergent capability was not explicitly trained — it falls out naturally from the contrastive objective applied at massive scale.
Multilingual extension
Because ALIGN's data pipeline uses no language-specific filters beyond basic text processing, extending it to other languages is straightforward. The paper lifts the English-only constraint, collects 1.8B multilingual image-text pairs (covering 100+ languages), builds a 250K wordpiece vocabulary, and trains ALIGN_mling using the same architecture and hyperparameters.
On Multi30K (multilingual covering English, German, French, and Czech), zero-shot ALIGN_mling outperforms the fine-tuned M3P model on English, German, and French. The largest improvement is +57.8 mean recall on French — a language where specialized multilingual models had struggled. This demonstrates that ALIGN's scale-over-curation philosophy generalizes across languages with minimal engineering effort.
Training recipe
ALIGN is trained from scratch (no pre-trained weights for either encoder) using the LAMB optimizer with a of 1e−5. The warms up linearly to 1e−3 over 10K steps, then decays linearly to zero over 1.2M steps (~12 epochs). Images are resized to 346×346 and randomly cropped to 289×289 during training (center-cropped during evaluation).
The text encoder uses a maximum sequence length of 64 wordpiece tokens, since alt-texts are capped at 20 unigrams. The softmax temperature is initialized at 1.0 and learned during training. Label smoothing of 0.1 is applied to both contrastive losses.
For fine-tuning on retrieval benchmarks, the batch size is reduced from 16,384 to 2,048 to avoid false negatives (when the dataset is small, random pairs may actually be true matches). The learning rate drops to 1e−5, and the model is fine-tuned for 3K–6K steps depending on the dataset.
Simplified to show the idea — not the real implementation.
# ALIGN training loop (simplified pseudocode)
for images, texts in dataloader: # batch of N=16384 pairs
# Encode both modalities
img_emb = l2_normalize(image_encoder(images)) # [N, D]
txt_emb = l2_normalize(text_encoder(texts)) # [N, D]
# Cosine similarity matrix scaled by learned temperature
logits = img_emb @ txt_emb.T / sigma # [N, N]
# Labels: diagonal = matches (index i matches i)
labels = arange(N)
# Symmetric contrastive loss with label smoothing
loss_i2t = cross_entropy(logits, labels, smoothing=0.1)
loss_t2i = cross_entropy(logits.T, labels, smoothing=0.1)
loss = (loss_i2t + loss_t2i) / 2
loss.backward()
optimizer.step() # LAMB optimizer
Legacy and the road ahead
2013
DeViSE — visual-semantic embeddings
Frome et al. learned to map images into word2vec space. An early dual-encoder that could do zero-shot recognition, but trained on ImageNet labels rather than free-form text.
2018
Conceptual Captions — curated image-text data
Sharma et al. built a 3.3M image-caption dataset from web alt-text using heavy filtering and semantic parsing. High quality, but the pipeline discarded 99.7% of raw data.
2021
CLIP — curated contrastive learning
Radford et al. collected 400M image-text pairs using a concept allowlist from Wikipedia. Dual encoder + contrastive loss achieves 76.2% zero-shot ImageNet. The curated-data approach.
2021
ALIGN — noisy data at massive scale
Jia et al. skip curation, use 1.8B noisy alt-text pairs, and match CLIP's zero-shot accuracy (76.4%) while setting new retrieval SOTA. Proof that scale compensates for noise.
2022
Flamingo — few-shot visual language
Alayrac et al. built on CLIP/ALIGN-style encoders to create a model capable of few-shot visual reasoning with interleaved image-text prompts. Extended the contrastive foundation into generative territory.
2023
ImageBind — binding six modalities
Girdhar et al. extended the contrastive alignment idea to six modalities (image, text, audio, depth, thermal, IMU). ALIGN's proof that noisy paired data suffices made this multi-modal generalization tractable.
ALIGN's lasting contribution is not just a model but a principle: you don't need to clean your data if you have enough of it. This philosophy has reshaped how the field collects training data. Before ALIGN, data curation was considered a prerequisite for quality. After ALIGN, researchers increasingly embraced web-scale noisy data as a viable — sometimes superior — alternative. The trend continues today in large multimodal models that train on billions of uncurated image-text pairs from the open web.
CitationJia, Yang, Xia, Chen, Parekh, Pham, Le, Sung, Li, Duerig. Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision. ICML, 2021.
Terms in this paper
- Dual Encoderالمُرمِّز الثنائي
- Contrastive Learningالتعلم التبايُني
- Contrastive Lossالخسارة التبايُنية
- Zero-Shot Classificationالتصنيف بدون أمثلة
- Zero-Shot Transferنقل بدون تدريب
- Image-Text Retrievalاسترجاع الصورة-النص
- Embedding Spaceفضاء التضمين
- Cosine Similarityتشابه جيب التمام
- Batch Sizeحجم الدفعة الحسابية
- Temperatureالحرارة
- EfficientNetإيفيشنت نت
- BERTبيرت
- Vision-Language Modelنموذج الرؤية واللغة
- Transfer Learningنقل التعلم
- Data Augmentationتعزيز البيانات