Multimodal AI2023intermediate11 min read
BLIP-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models
BLIP-2: بناء تدريب مسبق للرؤية واللغة انطلاقاً من مرمِّزات صور ونماذج لغوية كبيرة مُجمَّدة
Li, J. · Li, D. · Savarese, S. · Hoi, S. — ICML
The problem
By 2023, vision-language had achieved impressive results, but at prohibitive cost. Models like Flamingo (80B parameters) required end-to-end training of massive architectures, consuming thousands of GPU-days. Meanwhile, powerful unimodal models — ViT image encoders trained on billions of images, and LLMs like OPT and Flan-T5 trained on trillions of tokens — already existed. The challenge was: how do you reuse these frozen experts without the and enormous compute that comes with end-to-end training? Previous approaches like Frozen and Flamingo used image-to-text generation loss alone, which proved insufficient to bridge the deep modality gap between vision and language.
The contribution
BLIP-2: a generic and compute-efficient vision-language pre-training strategy that bootstraps from frozen pre-trained image encoders and frozen LLMs. The key innovation is — a lightweight Querying with only 188M trainable parameters — pre-trained in two stages to bridge the modality gap. Stage 1 learns vision-language alignment using three objectives (ITC, ITM, ITG) with a frozen . Stage 2 connects the Q-Former to a frozen LLM for generative learning. BLIP-2 outperforms Flamingo80B by 8.7% on zero-shot VQAv2 while using 54× fewer trainable parameters, and achieves state-of-the-art on zero-shot (NoCaps 121.6 CIDEr) and .
The impact
BLIP-2 established the dominant paradigm in AI: freeze powerful unimodal models, train a lightweight connector. This "bridge and freeze" pattern was adopted by virtually every major multimodal system that followed. InstructBLIP extended it with . LLaVA simplified the connector to a single . The insight that you can bootstrap from existing foundation models rather than training from scratch made multimodal AI accessible to research labs with limited compute, and permanently changed how the field builds vision-language systems.
Imagine a United Nations conference where a photographer and a novelist sit at the same table but speak entirely different languages. The photographer captures vivid, detailed scenes. The novelist crafts beautiful, reasoned narratives. But neither can understand the other.
Instead of forcing both to learn a new shared language (extremely expensive retraining), you place a skilled interpreter between them — someone who learned to listen to the photographer's visual descriptions and rephrase them as prompts the novelist can work with. That interpreter is BLIP-2's Q-Former: a lightweight bridge that translates vision into language, letting two frozen experts collaborate without either needing to change.
The cost problem: end-to-end VLP is unsustainable
Vision-language pre-training (VLP) had become a victim of its own success. Models kept getting bigger — and so did the bill. Flamingo achieved strong multimodal results, but its 80 billion parameters required enormous compute to train end-to-end. Meanwhile, the ingredients for great vision-language models already existed separately: ViT encoders trained by CLIP understood images superbly, and LLMs like OPT and Flan-T5 understood language brilliantly.
The fundamental question was deceptively simple: can we just connect these existing experts instead of training everything from scratch? The challenge is that an image and an LLM live in completely different representation spaces — raw visual features are meaningless to a language model. Previous attempts like Frozen used a simple image-to-text generation loss to bridge this gap, but this proved too weak. The modality gap between vision and language is deep, and a naive connection leads to catastrophic forgetting in the LLM or poor visual grounding.
The Q-Former: a lightweight bridge between vision and language
The heart of BLIP-2 is the Querying Transformer (Q-Former) — a small transformer that sits between the frozen image encoder and the frozen LLM. Think of it as a translator with 32 "questions" (learnable query embeddings) that it asks the image encoder. These queries attend to the image features through layers and learn to extract exactly the visual information that matters for language tasks.
Architecturally, Q-Former consists of two interleaved sub-modules that share the same layers: (1) an image transformer whose queries interact with frozen image features via cross-attention (inserted every other block), and (2) a text transformer that encodes or decodes text. The queries can also interact with text tokens through the shared self-attention — and different attention masks control this interaction depending on the training objective.
The Q-Former is initialized from BERT-Base weights. The cross-attention layers (which don't exist in BERT) are randomly initialized. This gives the Q-Former a strong language understanding prior from the start.
Two-stage pre-training: learn to see, then learn to speak
BLIP-2's training proceeds in two carefully designed stages. Each stage freezes different parts of the system and teaches Q-Former a different skill.
Stage 1: Vision-language representation learning
In the first stage, Q-Former learns to extract visual features that are most relevant to text — with the image encoder kept completely frozen. The image is passed through the frozen ViT, producing patch embeddings. The 32 learnable queries then interact with these patch embeddings through cross-attention layers, learning which visual regions matter for understanding the accompanying text.
Three objectives train Q-Former jointly in this stage, each with a different self-attention masking strategy that controls how queries and text interact:
Image-Text (ITC): Aligns the query outputs with the text by maximizing similarity of matched image-text pairs and minimizing similarity of unmatched pairs. Uses a unimodal mask — queries and text cannot see each other through self-attention, forcing each modality to produce independent representations.
Image-Grounded Text Generation (ITG): Trains Q-Former to generate text conditioned on the image. Uses a causal mask — queries can attend to each other but text tokens can only see previous tokens plus all queries, mimicking a that uses visual queries as a prefix.
(ITM): A binary classification task that predicts whether an image-text pair is matched. Uses a mask — queries and text fully attend to each other, enabling fine-grained cross-modal alignment.
Stage 2: Vision-to-language generative learning
In the second stage, Q-Former connects to a frozen LLM. The output query embeddings from Q-Former are projected through a fully-connected layer to match the LLM's input dimension, then prepended as soft visual prompts to the LLM's input. The LLM itself remains completely frozen — it never updates a single parameter.
BLIP-2 supports two types of LLMs:
Decoder-based LLMs (e.g., OPT): Trained with standard language modeling loss — the frozen LLM generates text conditioned on the visual prompts from Q-Former.
LLMs (e.g., Flan-T5): Trained with prefix language modeling loss — the text is split into a prefix (concatenated with visual prompts as encoder input) and a suffix (used as the decoder's generation target).
This stage is what gives BLIP-2 its generative capabilities. Because the LLM is frozen, BLIP-2 inherits all the LLM's language abilities — reasoning, instruction following, knowledge — without any risk of catastrophic forgetting.
The idea in code
Simplified to show the idea — not the real implementation.
import torch
import torch.nn as nn
class QFormerBridge(nn.Module):
"""Simplified Q-Former: learnable queries attend to frozen image features,
then project to the LLM's input space."""
def __init__(self, num_queries=32, d_model=768, d_llm=2560):
super().__init__()
# Learnable query embeddings — the 32 "questions" asked to the image
self.queries = nn.Parameter(torch.randn(1, num_queries, d_model))
# Cross-attention: queries attend to image patch features
self.cross_attn = nn.MultiheadAttention(d_model, num_heads=12, batch_first=True)
# Project query outputs to match the LLM's embedding dimension
self.proj = nn.Linear(d_model, d_llm)
def forward(self, image_features):
"""
image_features: (batch, num_patches, d_model) from frozen ViT
Returns: (batch, num_queries, d_llm) — visual prompts for the LLM
"""
B = image_features.size(0)
q = self.queries.expand(B, -1, -1) # (B, 32, 768)
# Queries attend to frozen image features via cross-attention
visual_queries, _ = self.cross_attn(
query=q,
key=image_features,
value=image_features
) # (B, 32, 768)
# Project to LLM space
visual_prompts = self.proj(visual_queries) # (B, 32, 2560)
return visual_prompts
# Usage:
# 1. Frozen ViT encodes the image → patch embeddings
# 2. Q-Former extracts 32 visual queries via cross-attention
# 3. FC layer projects queries to LLM dimension
# 4. Visual prompts are prepended to text tokens → frozen LLM generates outputResults: more with less
BLIP-2's results are remarkable not just for their quality, but for the with which they were achieved:
Zero-shot VQAv2: 65.0% accuracy with OPT-6.7B (65.4% with Flan-T5-XXL) — outperforming Flamingo80B's 56.3% by 8.7 points while using 54× fewer trainable parameters. Fine-tuned, BLIP-2 reaches 82.2%.
Image Captioning: Zero-shot CIDEr of 144.5 on COCO (OPT-6.7B). Fine-tuned, it achieves 145.8 CIDEr. On NoCaps, zero-shot CIDEr reaches 121.6 — a new state-of-the-art that demonstrates strong generalization to novel object categories.
Image-Text Retrieval: On COCO, fine-tuned BLIP-2 achieves 97.6% Recall@1 for image-to-text and 89.7% for text-to-image retrieval.
All this with a training cost that is a fraction of competing approaches — approximately 12 GPU-days compared to the hundreds required by end-to-end methods.
What matters most? Ablation insights
The paper's reveals critical design insights:
is essential. Without Stage 1, both OPT and Flan-T5 give substantially lower performance on zero-shot VQA. OPT in particular suffers from catastrophic forgetting — performance drastically degrades as training proceeds. Stage 1 gives the queries enough visual grounding that the LLM can interpret them without destabilization.
All three Stage 1 losses matter. Removing ITC hurts retrieval. Removing ITG hurts generation. Removing ITM hurts fine-grained understanding. The three losses are complementary, not redundant.
Stronger unimodal models yield better results. Upgrading the image encoder from ViT-L (CLIP) to ViT-g (EVA-CLIP) consistently improves all tasks. Upgrading the LLM from OPT-2.7B to OPT-6.7B or Flan-T5-XXL improves generation quality. BLIP-2's modular design means it automatically benefits from advances in either modality.
What BLIP-2 unlocked
2023
BLIP-2
Freeze image encoder + LLM, train lightweight Q-Former bridge. 54× fewer trainable parameters than Flamingo, better results across VQA, captioning, and retrieval.
2023
InstructBLIP
Extended BLIP-2 by adding instruction tuning to the Q-Former, enabling instruction-following across diverse vision-language tasks without task-specific fine-tuning.
2023
LLaVA — Visual Instruction Tuning
Simplified the connector from Q-Former to a single linear projection layer, proving that even simpler bridges work when the LLM is strong enough and instruction data is rich.
2023
MiniGPT-4
Froze the entire BLIP-2 Q-Former + ViT pipeline and trained only a single linear layer to connect to Vicuna LLM, achieving GPT-4-like visual conversation.
2024
BLIP-3 (xGen-MM)
Replaced Q-Former with a simpler vision token sampler and unified all training under a single next-token prediction loss, scaling to interleaved multi-image input.
BLIP-2's deepest legacy is the paradigm, not the architecture. Before BLIP-2, the prevailing assumption was that multimodal models needed end-to-end training of enormous systems. BLIP-2 proved that a lightweight connector trained in carefully staged phases can achieve superior results by standing on the shoulders of existing foundation models. This "freeze and bridge" pattern became the template for virtually every major multimodal system that followed — from LLaVA to GPT-4V to Gemini. The insight that modularity beats monolithic training when strong unimodal components already exist permanently reshaped how the field builds vision-language AI.
CitationLi, Li, Savarese, Hoi. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. ICML, 2023.
Terms in this paper
- Q-Formerمحوِّل الاستعلام
- Frozen Weightsالأوزان المُجمَّدة
- Vision Encoderمرمِّز بصري
- Large Language Modelالنموذج اللغوي الكبير
- Cross-Attentionالانتباه التبادلي
- Contrastive Learningالتعلم التبايُني
- Image Captioningوصف الصور
- Visual Question Answeringالإجابة البصرية عن الأسئلة
- Image-Text Matchingمطابقة الصورة والنص
- Zero-Shot Transferنقل بدون تدريب
- Bootstrappingالتمهيد الذاتي
- Fine-Tuningالضبط الدقيق
- Embeddingالتضمين
- Self-Attentionالانتباه الذاتي
- Transfer Learningنقل التعلم