Multimodal AI2022advanced13 min read
Flamingo: a Visual Language Model for Few-Shot Learning
Flamingo: نموذج لغوي بصري للتعلّم بأمثلة قليلة
Alayrac, J.-B. · Donahue, J. · Luc, P. · Miech, A. · Barr, I. · Hasson, Y. · Lenc, K. · Mensch, A. · Millican, K. · Reynolds, M. · Ring, R. · Rutherford, E. · Cabi, S. · Han, T. · Gong, Z. · Samangooei, S. · Monteiro, M. · Menick, J. · Borgeaud, S. · Brock, A. · Nematzadeh, A. · Sharifzadeh, S. · Binkowski, M. · Barreira, R. · Vinyals, O. · Zisserman, A. · Simonyan, K. — NeurIPS
The problem
By 2022 there were powerful vision models (like CLIP) and powerful language models (like Chinchilla), but combining them was awkward. a language model on visual tasks risked of its language abilities. And every new visual task — captioning, VQA, — still demanded large annotated datasets and task-specific architectures. There was no single model that could look at a few image-text examples and immediately generalize to a new visual task, the way GPT-3 did for language.
The contribution
Flamingo: a family of visual language models (3B, 9B, 80B parameters) that bridge a frozen and a frozen language model using two innovations — a Resampler that compresses variable-length visual features into a fixed set of tokens, and Gated Dense layers that inject visual information into the LM without disrupting its learned representations. Trained on interleaved image-text web data, Flamingo achieves state-of-the-art few-shot performance on 16 benchmarks, surpassing fine-tuned models that used thousands of times more task-specific data.
The impact
Flamingo proved that frozen LMs can acquire visual understanding without catastrophic forgetting, establishing the "bridge and gate" paradigm for multimodal models. Its Perceiver Resampler directly inspired the Q-Former in BLIP-2, and the interleaved vision-language approach influenced Gemini's multimodal architecture. The idea of gated cross-attention for safe visual injection became the standard blueprint for connecting vision encoders to LLMs, visible in LLaVA, PaLM-E, and many others.
Imagine a brilliant lecturer who gives perfect talks but has never seen a photograph in their life. Now hire a visual aide who watches a slideshow and, after each slide, writes a short sticky note summarizing what they saw. The aide slips these notes between the lecturer's script pages. The lecturer never changes their script — they just learn to glance at the sticky notes and weave the visual details into their talk. The aide also carries a magnifying glass that compresses a large, detailed photo into exactly the same number of bullet points every time, no matter the photo's size. That's Flamingo: the lecturer is a frozen language model, the magnifying glass is the Perceiver Resampler, and the sticky notes flow through gated cross-attention.
The problem: vision and language lived in separate worlds
By 2022, the AI community had two kinds of powerful models that rarely talked to each other. On one side, CLIP and similar contrastive models could match images to text — but only via a shared space; they couldn't generate language. On the other side, large language models like Chinchilla could write fluent paragraphs — but were completely blind, processing only text tokens.
The obvious fix — fine-tune the LM on visual data — ran into a brutal obstacle: catastrophic forgetting. When you retrain a 70-billion-parameter language model on captioning data, it can lose its ability to do language tasks. And even if you manage the forgetting, you still need large labeled datasets for every new visual task. Want to do VQA? Train on VQA data. Want captioning? Train on captioning data. There was no visual equivalent of GPT-3's , where you show a few examples in the prompt and the model generalizes on the spot.
Flamingo's solution: bridge, don't merge
Flamingo's key insight is architectural restraint. Instead of merging vision and language into one model, it keeps both the vision encoder and the language model frozen and builds a lightweight bridge between them. This bridge has two components:
- Perceiver Resampler — compresses variable-length visual features from images or video frames into a fixed set of visual tokens.
- Gated Cross-Attention Dense (GATED XATTN-DENSE) layers — new layers inserted between the frozen LM layers that allow the language model to "look at" the visual tokens through cross-attention.
The frozen components (vision encoder + LM) hold 90%+ of the parameters and never change. Only the bridge — the Perceiver Resampler and the gated cross-attention layers — is trained from scratch. This preserves the LM's language capabilities while teaching it to incorporate visual information.
The Perceiver Resampler: one-size-fits-all visual compression
A single image produces hundreds of visual features from the vision encoder (e.g. a 2D spatial grid from an NFNet). A video produces thousands — one grid per frame. Feeding all of these directly into the language model would be prohibitively expensive, especially since the LM's attention is quadratic in sequence length.
The Perceiver Resampler solves this with an elegant idea borrowed from Perceiver: use a fixed set of learned latent queries (64 tokens) that cross-attend to the visual features. No matter whether the input is one image or a 30-frame video, the output is always exactly 64 visual tokens. Think of it as a panel of 64 expert analysts: they each look at the entire visual input and produce one compact summary vector. The same 64 analysts handle any input, so the downstream LM always sees a fixed-size visual context.
Architecturally, the Perceiver Resampler is a small with learned latent vectors as queries and the vision encoder's features (plus learned temporal embeddings for video) as keys and values. It has 194M parameters — tiny compared to the 70B LM.
Gated Cross-Attention: the safe visual injection
The Perceiver Resampler produces 64 visual tokens per image. But how does the language model actually use them? This is where Gated Cross-Attention Dense (GATED XATTN-DENSE) layers come in.
New cross-attention layers are inserted between the existing frozen LM layers. In each GATED XATTN-DENSE layer, the query comes from the language model's hidden states (what the LM is currently "thinking"), while the key and value come from the visual tokens produced by the Perceiver Resampler. This lets each text ask: "Is there anything in the image relevant to what I'm about to generate?"
But there's a critical subtlety. If you simply add cross-attention outputs to the LM's hidden states, you will corrupt the carefully learned language representations at initialization, crashing performance. Flamingo's solution: tanh gating. The output of each new layer is multiplied by , where is a learnable scalar initialized to zero. At initialization, , so the visual signal is completely muted — the model behaves exactly like the original frozen LM. During , gradually increases, letting visual information flow in smoothly. This is the "gate" that opens slowly, preventing the language model from being overwhelmed.
Interleaved input: images and text in one stream
Most vision-language models of the era handled one image and one text at a time. Flamingo breaks this limitation by accepting arbitrarily interleaved sequences of images (or videos) and text, exactly as they appear on a web page: an image, then a paragraph, then another image, then another paragraph.
Special tokens mark the boundaries: <image> precedes each visual input, and <EOC> (End of Chunk) follows each text segment associated with an image. This structure is critical for : you can place several (image, caption) pairs in the prompt as demonstrations, then provide a new image and let the model generate the caption.
To prevent future images from influencing past text (causal consistency), Flamingo uses image-causal masking in the cross-attention: each text token can only attend to visual tokens from the image that precedes it, not from images that come later in the sequence.
Training: four datasets, one objective
Flamingo is trained on four complementary datasets:
- M3W (MultiModal MassiveWeb): text and images extracted from approximately 43 million web pages, preserving their natural interleaved structure. This is the key dataset for in-context few-shot learning.
- ALIGN: 1.8 billion image-text pairs from alt-text on the web. Large but noisy, with short descriptions averaging 12.4 tokens.
- LTIP (Long Text & Image Pairs): 312 million image-text pairs with higher quality and longer descriptions (averaging 20.5 tokens).
- VTP (Video & Text Pairs): 27 million short videos (~22 seconds) paired with descriptive captions.
The training objective is a weighted sum of per-dataset negative log-likelihoods, with weights set to 1.0, 0.2, 0.2, and 0.03 for M3W, ALIGN, LTIP, and VTP respectively. The M3W dataset receives the highest weight because it provides the interleaved structure essential for few-shot capabilities. Ablation studies showed that removing M3W caused a 17% performance drop — confirming that interleaved data is the backbone of few-shot learning.
Few-shot learning: the ultimate payoff
The whole architecture comes together in few-shot evaluation. Given a new visual task — say, answering questions about medical images — you don't fine-tune anything. You simply construct a prompt with a few (image, question, answer) examples, append a new (image, question), and let Flamingo generate the answer. The interleaved architecture, trained on M3W's natural web structure, has learned the meta-skill of in-context learning from visual demonstrations.
Flamingo supports 0-shot (no examples), 4-shot, 8-shot, 16-shot, and 32-shot evaluation. Performance improves log-linearly with the number of shots, and even at 4 shots, Flamingo often matches or exceeds models that were fine-tuned on tens of thousands of labeled examples.
Three sizes: 3B, 9B, and 80B
Flamingo comes in three sizes, all sharing the same 435M-parameter vision encoder and 194M-parameter Perceiver Resampler:
- Flamingo-3B: 1.4B frozen LM + 1.2B trainable GATED XATTN-DENSE (inserted after every LM layer). 3.2B total parameters.
- Flamingo-9B: 7.1B frozen LM + 1.6B trainable GATED XATTN-DENSE (inserted every 4th layer). 9.3B total parameters.
- Flamingo-80B: 70B frozen Chinchilla LM + 10B trainable GATED XATTN-DENSE (inserted every 7th layer). 80B total parameters. Trained on 1536 TPU chips for 15 days.
A key design decision: larger models insert cross-attention less frequently. The 80B model only adds a gated cross-attention layer every 7th frozen LM layer, trading off visual resolution for computational efficiency.
Results: few-shot beats fine-tuned
Flamingo-80B was evaluated on 16 multimodal benchmarks spanning captioning, , and classification. The results were striking:
- With just 32 few-shot examples, Flamingo set new state-of-the-art on 6 out of 16 benchmarks — without any fine-tuning — surpassing models that were fine-tuned on thousands of labeled examples.
- On all benchmarks with published few-shot results, Flamingo set the new state-of-the-art for few-shot performance.
- Performance scaled cleanly with model size: Flamingo-80B > Flamingo-9B > Flamingo-3B across all tasks.
- On open-ended tasks like COCO captioning and VQAv2, Flamingo showed especially strong performance, demonstrating that the generative language model backbone was a major advantage.
When fine-tuned on task-specific data, Flamingo set new state-of-the-art on an additional 5 benchmarks, demonstrating that the few-shot architecture is also an excellent initialization for full fine-tuning.
The idea in code
Simplified to show the idea — not the real implementation.
import torch
import torch.nn as nn
import torch.nn.functional as F
class PerceiverResampler(nn.Module):
"""Compress variable-length visual features into fixed tokens."""
def __init__(self, dim=1024, n_latents=64, n_heads=8):
super().__init__()
self.latents = nn.Parameter(torch.randn(n_latents, dim))
self.cross_attn = nn.MultiheadAttention(dim, n_heads, batch_first=True)
self.norm = nn.LayerNorm(dim)
def forward(self, visual_features):
# visual_features: (batch, n_tokens, dim) — variable n_tokens
latents = self.latents.unsqueeze(0).expand(visual_features.shape[0], -1, -1)
# Cross-attend: latents query the visual features
out, _ = self.cross_attn(query=latents, key=visual_features, value=visual_features)
return self.norm(out) # (batch, 64, dim) — always 64 tokens
class GatedCrossAttentionDense(nn.Module):
"""One GATED XATTN-DENSE layer inserted between frozen LM layers."""
def __init__(self, dim=1024, n_heads=8):
super().__init__()
self.cross_attn = nn.MultiheadAttention(dim, n_heads, batch_first=True)
self.ffn = nn.Sequential(nn.Linear(dim, dim * 4), nn.GELU(), nn.Linear(dim * 4, dim))
self.norm1 = nn.LayerNorm(dim)
self.norm2 = nn.LayerNorm(dim)
# Gate scalars initialized to 0 → tanh(0) = 0 → no visual signal at start
self.alpha_xattn = nn.Parameter(torch.tensor(0.0))
self.alpha_ffn = nn.Parameter(torch.tensor(0.0))
def forward(self, x, visual_tokens):
# x: text hidden states from frozen LM layer
# visual_tokens: output of Perceiver Resampler (batch, 64, dim)
xattn_out, _ = self.cross_attn(
query=self.norm1(x), key=visual_tokens, value=visual_tokens
)
x = x + torch.tanh(self.alpha_xattn) * xattn_out # gated residual
ffn_out = self.ffn(self.norm2(x))
x = x + torch.tanh(self.alpha_ffn) * ffn_out # gated residual
return x
# At initialization: tanh(0) = 0, so both gates are shut.
# The model outputs exactly what the frozen LM would produce.
# During training, alpha values grow, slowly opening the visual gates.What Flamingo unlocked
2021
CLIP — Contrastive Vision-Language Pretraining
Proved that contrastive learning on image-text pairs creates a powerful shared embedding space, but CLIP cannot generate text.
2022
Chinchilla — Compute-Optimal LM
A 70B language model trained with optimal data-to-parameters ratio. Flamingo uses it as the frozen language backbone.
2022
Flamingo — Visual Language Model
Bridges frozen vision and language with Perceiver Resampler and gated cross-attention. State-of-the-art few-shot on 16 benchmarks.
2023
BLIP-2 — Bootstrapped Vision-Language Pretraining
Replaced the Perceiver Resampler with Q-Former, a more lightweight bridge. Directly inspired by Flamingo's freeze-and-bridge approach.
2023
Gemini — Multimodal from the Ground Up
Google DeepMind's natively multimodal model, building on Flamingo's interleaved vision-language paradigm but training vision and language jointly from scratch.
Flamingo's deepest contribution is the "freeze-and-bridge" paradigm. Before Flamingo, multimodal models either trained everything end-to-end (expensive and prone to forgetting) or used simple projection layers (losing rich visual information). Flamingo showed that you can keep powerful pretrained models intact and connect them with carefully gated bridges. This pattern — frozen encoder, bridge module, frozen decoder — is now the standard recipe for building multimodal models, from BLIP-2 to LLaVA to PaLM-E.
CitationAlayrac, Donahue, Luc, Miech, et al.. Flamingo: a Visual Language Model for Few-Shot Learning. NeurIPS, 2022.
Terms in this paper
- Visual Language Modelالنموذج اللغوي البصري
- Few-Shot Learningالتعلّم بأمثلة قليلة
- Cross-Attentionالانتباه التبادلي
- Perceiverبيرسيفر
- In-Context Learningالتعلم في السياق
- Gating Mechanismآلية البوابات
- Vision Encoderمرمِّز بصري
- Multimodalمتعدد الوسائط
- Contrastive Learningالتعلم التبايُني
- Image Captioningوصف الصور
- Visual Question Answeringالإجابة البصرية عن الأسئلة
- Pretrainingالتدريب المسبق
- Fine-Tuningالضبط الدقيق
- Frozen Weightsالأوزان المُجمَّدة