Multimodal AI2022advanced13 min read

Flamingo: a Visual Language Model for Few-Shot Learning

Flamingo: نموذج لغوي بصري للتعلّم بأمثلة قليلة

Alayrac, J.-B. · Donahue, J. · Luc, P. · Miech, A. · Barr, I. · Hasson, Y. · Lenc, K. · Mensch, A. · Millican, K. · Reynolds, M. · Ring, R. · Rutherford, E. · Cabi, S. · Han, T. · Gong, Z. · Samangooei, S. · Monteiro, M. · Menick, J. · Borgeaud, S. · Brock, A. · Nematzadeh, A. · Sharifzadeh, S. · Binkowski, M. · Barreira, R. · Vinyals, O. · Zisserman, A. · Simonyan, K. — NeurIPS

The problem

By 2022 there were powerful vision models (like CLIP) and powerful language models (like Chinchilla), but combining them was awkward. a language model on visual tasks risked of its language abilities. And every new visual task — captioning, VQA, — still demanded large annotated datasets and task-specific architectures. There was no single model that could look at a few image-text examples and immediately generalize to a new visual task, the way GPT-3 did for language.

The contribution

Flamingo: a family of visual language models (3B, 9B, 80B parameters) that bridge a frozen and a frozen language model using two innovations — a Resampler that compresses variable-length visual features into a fixed set of tokens, and Gated Dense layers that inject visual information into the LM without disrupting its learned representations. Trained on interleaved image-text web data, Flamingo achieves state-of-the-art few-shot performance on 16 benchmarks, surpassing fine-tuned models that used thousands of times more task-specific data.

The impact

Flamingo proved that frozen LMs can acquire visual understanding without catastrophic forgetting, establishing the "bridge and gate" paradigm for multimodal models. Its Perceiver Resampler directly inspired the Q-Former in BLIP-2, and the interleaved vision-language approach influenced Gemini's multimodal architecture. The idea of gated cross-attention for safe visual injection became the standard blueprint for connecting vision encoders to LLMs, visible in LLaVA, PaLM-E, and many others.

Imagine a brilliant lecturer who gives perfect talks but has never seen a photograph in their life. Now hire a visual aide who watches a slideshow and, after each slide, writes a short sticky note summarizing what they saw. The aide slips these notes between the lecturer's script pages. The lecturer never changes their script — they just learn to glance at the sticky notes and weave the visual details into their talk. The aide also carries a magnifying glass that compresses a large, detailed photo into exactly the same number of bullet points every time, no matter the photo's size. That's Flamingo: the lecturer is a frozen language model, the magnifying glass is the Perceiver Resampler, and the sticky notes flow through gated cross-attention.

The problem: vision and language lived in separate worlds

By 2022, the AI community had two kinds of powerful models that rarely talked to each other. On one side, CLIP and similar contrastive models could match images to text — but only via a shared space; they couldn't generate language. On the other side, large language models like Chinchilla could write fluent paragraphs — but were completely blind, processing only text tokens.

The obvious fix — fine-tune the LM on visual data — ran into a brutal obstacle: catastrophic forgetting. When you retrain a 70-billion-parameter language model on captioning data, it can lose its ability to do language tasks. And even if you manage the forgetting, you still need large labeled datasets for every new visual task. Want to do VQA? Train on VQA data. Want captioning? Train on captioning data. There was no visual equivalent of GPT-3's , where you show a few examples in the prompt and the model generalizes on the spot.

Flamingo's solution: bridge, don't merge

Flamingo's key insight is architectural restraint. Instead of merging vision and language into one model, it keeps both the vision encoder and the language model frozen and builds a lightweight bridge between them. This bridge has two components:

  1. Perceiver Resampler — compresses variable-length visual features from images or video frames into a fixed set of visual tokens.
  2. Gated Cross-Attention Dense (GATED XATTN-DENSE) layers — new layers inserted between the frozen LM layers that allow the language model to "look at" the visual tokens through cross-attention.

The frozen components (vision encoder + LM) hold 90%+ of the parameters and never change. Only the bridge — the Perceiver Resampler and the gated cross-attention layers — is trained from scratch. This preserves the LM's language capabilities while teaching it to incorporate visual information.

Open in Lab
Click any component to learn its role. Blue = frozen pretrained, purple = trained from scratch.
The demo wakes as you arrive…

The Perceiver Resampler: one-size-fits-all visual compression

A single image produces hundreds of visual features from the vision encoder (e.g. a 2D spatial grid from an NFNet). A video produces thousands — one grid per frame. Feeding all of these directly into the language model would be prohibitively expensive, especially since the LM's attention is quadratic in sequence length.

The Perceiver Resampler solves this with an elegant idea borrowed from Perceiver: use a fixed set of learned latent queries (64 tokens) that cross-attend to the visual features. No matter whether the input is one image or a 30-frame video, the output is always exactly 64 visual tokens. Think of it as a panel of 64 expert analysts: they each look at the entire visual input and produce one compact summary vector. The same 64 analysts handle any input, so the downstream LM always sees a fixed-size visual context.

Architecturally, the Perceiver Resampler is a small with learned latent vectors as queries and the vision encoder's features (plus learned temporal embeddings for video) as keys and values. It has 194M parameters — tiny compared to the 70B LM.

Open in Lab
Watch how variable-length visual features (images or video frames) are compressed into a fixed set of 64 visual tokens.
The demo wakes as you arrive…

Gated Cross-Attention: the safe visual injection

The Perceiver Resampler produces 64 visual tokens per image. But how does the language model actually use them? This is where Gated Cross-Attention Dense (GATED XATTN-DENSE) layers come in.

New cross-attention layers are inserted between the existing frozen LM layers. In each GATED XATTN-DENSE layer, the query comes from the language model's hidden states (what the LM is currently "thinking"), while the key and value come from the visual tokens produced by the Perceiver Resampler. This lets each text ask: "Is there anything in the image relevant to what I'm about to generate?"

But there's a critical subtlety. If you simply add cross-attention outputs to the LM's hidden states, you will corrupt the carefully learned language representations at initialization, crashing performance. Flamingo's solution: tanh gating. The output of each new layer is multiplied by tanh⁡(α)\tanh(\alpha), where α\alpha is a learnable scalar initialized to zero. At initialization, tanh⁡(0)=0\tanh(0) = 0, so the visual signal is completely muted — the model behaves exactly like the original frozen LM. During , α\alpha gradually increases, letting visual information flow in smoothly. This is the "gate" that opens slowly, preventing the language model from being overwhelmed.

Open in Lab
Drag the α slider to see how the tanh gate controls visual information flow into the frozen LM.
The demo wakes as you arrive…
y=x+tanh⁡(αxattn)⋅CrossAttn(x,v)+tanh⁡(αffn)⋅FFN(…)y = x + \tanh(\alpha_{\text{xattn}}) \cdot \text{CrossAttn}(x, v) + \tanh(\alpha_{\text{ffn}}) \cdot \text{FFN}(\ldots)
Gated cross-attention residual connection — x is the hidden state from the frozen LM layer. v is the set of visual tokens from the Perceiver Resampler. Both α_xattn and α_ffn are learnable scalars initialized to 0, so at the start of training tanh(0) = 0 and the output y = x — the model is identical to the original LM. As training progresses, the gates open and the model learns to blend visual context.

Interleaved input: images and text in one stream

Most vision-language models of the era handled one image and one text at a time. Flamingo breaks this limitation by accepting arbitrarily interleaved sequences of images (or videos) and text, exactly as they appear on a web page: an image, then a paragraph, then another image, then another paragraph.

Special tokens mark the boundaries: <image> precedes each visual input, and <EOC> (End of Chunk) follows each text segment associated with an image. This structure is critical for : you can place several (image, caption) pairs in the prompt as demonstrations, then provide a new image and let the model generate the caption.

To prevent future images from influencing past text (causal consistency), Flamingo uses image-causal masking in the cross-attention: each text token can only attend to visual tokens from the image that precedes it, not from images that come later in the sequence.

Open in Lab
See how images and text interleave in a Flamingo prompt, with image-causal masking.
The demo wakes as you arrive…

Training: four datasets, one objective

Flamingo is trained on four complementary datasets:

  • M3W (MultiModal MassiveWeb): text and images extracted from approximately 43 million web pages, preserving their natural interleaved structure. This is the key dataset for in-context few-shot learning.
  • ALIGN: 1.8 billion image-text pairs from alt-text on the web. Large but noisy, with short descriptions averaging 12.4 tokens.
  • LTIP (Long Text & Image Pairs): 312 million image-text pairs with higher quality and longer descriptions (averaging 20.5 tokens).
  • VTP (Video & Text Pairs): 27 million short videos (~22 seconds) paired with descriptive captions.

The training objective is a weighted sum of per-dataset negative log-likelihoods, with weights λ\lambda set to 1.0, 0.2, 0.2, and 0.03 for M3W, ALIGN, LTIP, and VTP respectively. The M3W dataset receives the highest weight because it provides the interleaved structure essential for few-shot capabilities. Ablation studies showed that removing M3W caused a 17% performance drop — confirming that interleaved data is the backbone of few-shot learning.

L=∑mλm E(x,y)∼Dm[−∑llog⁡p(yl∣y<l,x)]\mathcal{L} = \sum_{m} \lambda_m \, \mathbb{E}_{(x,y) \sim \mathcal{D}_m} \left[ -\sum_{l} \log p(y_l \mid y_{<l}, x) \right]
Multi-dataset training objective — The model minimizes a weighted sum of expected negative log-likelihoods across all datasets. Each dataset D_m contributes with weight λ_m. The text y is conditioned on both preceding text tokens y_{<l} and the visual inputs x. The M3W weight (1.0) is 5× the paired datasets (0.2), reflecting the primacy of interleaved data.

Few-shot learning: the ultimate payoff

The whole architecture comes together in few-shot evaluation. Given a new visual task — say, answering questions about medical images — you don't fine-tune anything. You simply construct a prompt with a few (image, question, answer) examples, append a new (image, question), and let Flamingo generate the answer. The interleaved architecture, trained on M3W's natural web structure, has learned the meta-skill of in-context learning from visual demonstrations.

Flamingo supports 0-shot (no examples), 4-shot, 8-shot, 16-shot, and 32-shot evaluation. Performance improves log-linearly with the number of shots, and even at 4 shots, Flamingo often matches or exceeds models that were fine-tuned on tens of thousands of labeled examples.

Open in Lab
Add or remove few-shot examples to see how Flamingo's prompt works.
The demo wakes as you arrive…

Three sizes: 3B, 9B, and 80B

Flamingo comes in three sizes, all sharing the same 435M-parameter vision encoder and 194M-parameter Perceiver Resampler:

  • Flamingo-3B: 1.4B frozen LM + 1.2B trainable GATED XATTN-DENSE (inserted after every LM layer). 3.2B total parameters.
  • Flamingo-9B: 7.1B frozen LM + 1.6B trainable GATED XATTN-DENSE (inserted every 4th layer). 9.3B total parameters.
  • Flamingo-80B: 70B frozen Chinchilla LM + 10B trainable GATED XATTN-DENSE (inserted every 7th layer). 80B total parameters. Trained on 1536 TPU chips for 15 days.

A key design decision: larger models insert cross-attention less frequently. The 80B model only adds a gated cross-attention layer every 7th frozen LM layer, trading off visual resolution for computational efficiency.

Open in Lab
Compare the three model sizes and their parameter distribution.
The demo wakes as you arrive…

Results: few-shot beats fine-tuned

Flamingo-80B was evaluated on 16 multimodal benchmarks spanning captioning, , and classification. The results were striking:

  • With just 32 few-shot examples, Flamingo set new state-of-the-art on 6 out of 16 benchmarks — without any fine-tuning — surpassing models that were fine-tuned on thousands of labeled examples.
  • On all benchmarks with published few-shot results, Flamingo set the new state-of-the-art for few-shot performance.
  • Performance scaled cleanly with model size: Flamingo-80B > Flamingo-9B > Flamingo-3B across all tasks.
  • On open-ended tasks like COCO captioning and VQAv2, Flamingo showed especially strong performance, demonstrating that the generative language model backbone was a major advantage.

When fine-tuned on task-specific data, Flamingo set new state-of-the-art on an additional 5 benchmarks, demonstrating that the few-shot architecture is also an excellent initialization for full fine-tuning.

The idea in code

Flamingo forward pass — Perceiver Resampler + gated cross-attentionpython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn as nn
import torch.nn.functional as F

class PerceiverResampler(nn.Module):
    """Compress variable-length visual features into fixed tokens."""
    def __init__(self, dim=1024, n_latents=64, n_heads=8):
        super().__init__()
        self.latents = nn.Parameter(torch.randn(n_latents, dim))
        self.cross_attn = nn.MultiheadAttention(dim, n_heads, batch_first=True)
        self.norm = nn.LayerNorm(dim)

    def forward(self, visual_features):
        # visual_features: (batch, n_tokens, dim) — variable n_tokens
        latents = self.latents.unsqueeze(0).expand(visual_features.shape[0], -1, -1)
        # Cross-attend: latents query the visual features
        out, _ = self.cross_attn(query=latents, key=visual_features, value=visual_features)
        return self.norm(out)  # (batch, 64, dim) — always 64 tokens

class GatedCrossAttentionDense(nn.Module):
    """One GATED XATTN-DENSE layer inserted between frozen LM layers."""
    def __init__(self, dim=1024, n_heads=8):
        super().__init__()
        self.cross_attn = nn.MultiheadAttention(dim, n_heads, batch_first=True)
        self.ffn = nn.Sequential(nn.Linear(dim, dim * 4), nn.GELU(), nn.Linear(dim * 4, dim))
        self.norm1 = nn.LayerNorm(dim)
        self.norm2 = nn.LayerNorm(dim)
        # Gate scalars initialized to 0 → tanh(0) = 0 → no visual signal at start
        self.alpha_xattn = nn.Parameter(torch.tensor(0.0))
        self.alpha_ffn = nn.Parameter(torch.tensor(0.0))

    def forward(self, x, visual_tokens):
        # x: text hidden states from frozen LM layer
        # visual_tokens: output of Perceiver Resampler (batch, 64, dim)
        xattn_out, _ = self.cross_attn(
            query=self.norm1(x), key=visual_tokens, value=visual_tokens
        )
        x = x + torch.tanh(self.alpha_xattn) * xattn_out  # gated residual
        ffn_out = self.ffn(self.norm2(x))
        x = x + torch.tanh(self.alpha_ffn) * ffn_out      # gated residual
        return x

# At initialization: tanh(0) = 0, so both gates are shut.
# The model outputs exactly what the frozen LM would produce.
# During training, alpha values grow, slowly opening the visual gates.

What Flamingo unlocked

  1. 2021

    CLIP — Contrastive Vision-Language Pretraining

    Proved that contrastive learning on image-text pairs creates a powerful shared embedding space, but CLIP cannot generate text.

  2. 2022

    Chinchilla — Compute-Optimal LM

    A 70B language model trained with optimal data-to-parameters ratio. Flamingo uses it as the frozen language backbone.

  3. 2022

    Flamingo — Visual Language Model

    Bridges frozen vision and language with Perceiver Resampler and gated cross-attention. State-of-the-art few-shot on 16 benchmarks.

  4. 2023

    BLIP-2 — Bootstrapped Vision-Language Pretraining

    Replaced the Perceiver Resampler with Q-Former, a more lightweight bridge. Directly inspired by Flamingo's freeze-and-bridge approach.

  5. 2023

    Gemini — Multimodal from the Ground Up

    Google DeepMind's natively multimodal model, building on Flamingo's interleaved vision-language paradigm but training vision and language jointly from scratch.

Flamingo's deepest contribution is the "freeze-and-bridge" paradigm. Before Flamingo, multimodal models either trained everything end-to-end (expensive and prone to forgetting) or used simple projection layers (losing rich visual information). Flamingo showed that you can keep powerful pretrained models intact and connect them with carefully gated bridges. This pattern — frozen encoder, bridge module, frozen decoder — is now the standard recipe for building multimodal models, from BLIP-2 to LLaVA to PaLM-E.

CitationAlayrac, Donahue, Luc, Miech, et al.. Flamingo: a Visual Language Model for Few-Shot Learning. NeurIPS, 2022.

Terms in this paper