Multimodal AI2023intermediate11 min read

Visual Instruction Tuning

الضبط التعليمي البصري

Liu, H. · Li, C. · Wu, Q. · Lee, Y. J. — NeurIPS

The problem

had revolutionized text-only LLMs — models like ChatGPT could follow complex human instructions thanks to on . But in the world, no one had done the same for vision: there was no instruction-following data that paired images with rich conversational questions and answers, and no trained end-to-end to follow visual instructions.

The contribution

LLaVA (Large Language and Vision Assistant): a pipeline that uses language-only GPT-4 to generate 158K multimodal instruction-following examples from existing image-text pairs, then trains an end-to-end model connecting CLIP's to the Vicuna LLM through a simple linear projection. A two-stage training process — first aligning visual features to the language space, then end-to-end on instruction data — produces a model that exhibits impressive visual chat abilities and achieves state-of-the-art on Science QA (92.53%).

The impact

LLaVA proved that connecting a frozen vision encoder to an LLM with a simple projection — then instruction-tuning on GPT-4-generated data — is a surprisingly effective recipe for multimodal AI. It launched a wave of research (LLaVA-1.5, LLaVA-NeXT, and dozens of descendants), established GPT-4-as-judge evaluation, and showed that synthetic instruction data can bootstrap powerful multimodal assistants. The architecture became a blueprint for virtually every open-source multimodal LLM that followed.

A language model without vision is like a brilliant scholar locked in a dark room — they can reason, write, and answer any text question, but cannot see the world. Previous attempts to give them sight were like handing them a photograph and asking "what do you see?" — useful, but limited.

LLaVA does something different: it gives the scholar eyes and a conversation partner. Instead of just captioning images, the model learns to discuss them — answering follow-up questions, reasoning about unusual scenes, and connecting what it sees to what it knows. The trick? A teacher (GPT-4) wrote the conversation scripts by reading image descriptions, so the scholar could practice before ever seeing real images.

The gap: LLMs can talk, but they cannot see

By 2023, instruction-tuned LLMs like ChatGPT had mastered following complex text instructions. Meanwhile, vision-language models like CLIP and BLIP-2 could understand images — but only in rigid ways: generate a caption, answer a multiple-choice question, or classify an object. None of them could hold a free-form conversation about an image the way GPT-4 handles text conversations.

Two walls blocked progress:

  • No multimodal instruction data. Instruction tuning needs examples of "Human asks → Model answers" about images. Such data barely existed, and creating it with human annotators is slow and expensive.

  • No end-to-end multimodal instruction-following model. Existing systems either used separate vision and language modules stitched together with complex (like Flamingo) or couldn't follow open-ended instructions at all.

The first innovation: GPT-4 writes the training data

The key insight is that you don't need a model that can see images to write about them — you just need good text descriptions. The authors took existing image-text pairs from the COCO dataset and converted each image's information into two symbolic forms that a text-only model can understand:

  • Captions — five human-written sentences describing the scene from different angles.

  • Bounding boxes — object labels with their coordinates, like person: [0.68, 0.24, 0.77, 0.69].

They fed these text descriptions to GPT-4 (which cannot see images) and asked it to generate three types of instruction-following responses as if it were seeing the image:

  • Conversations — multi-turn Q&A about the visual content (58K examples).

  • Detailed descriptions — rich, comprehensive image descriptions (23K examples).

  • Complex reasoning — questions requiring step-by-step logical thinking about the scene (77K examples).

This produced 158K instruction-following examples — a full training dataset, generated without any human annotation beyond the original captions.

Open in Lab
Step through how GPT-4 transforms image captions and bounding boxes into conversational instruction-following data.
The demo wakes as you arrive…

The architecture: vision encoder + projection + LLM

LLaVA's architecture is strikingly simple — just three components chained together:

  1. Vision encoder (CLIP ViT-L/14) — the same Vision from CLIP, pre-trained to align images with text. It converts an image into a grid of visual vectors. Think of it as the model's retina: it sees the image and produces a structured representation.

  2. (a single linear matrix W) — this is the "translator" that converts visual features into the same dimensional space as the LLM's word embeddings. One matrix multiplication: Hv=W⋅ZvH_v = W \cdot Z_v. That's it — no cross-attention, no Q-Former, no complex fusion. Just a linear bridge.

  3. Language model (Vicuna) — an instruction-tuned version of LLaMA that generates responses. It receives the projected concatenated with the text instruction tokens, and generates the answer autoregressively.

The visual tokens are simply prepended to the text tokens in the LLM's input sequence. From the LLM's perspective, the image is just more tokens — it processes vision and language in a single unified stream.

Hv=W⋅Zv,Zv=g(Xv)H_v = W \cdot Z_v, \quad Z_v = g(X_v)
Visual feature projection — the bridge between sight and language — g = CLIP vision encoder · X_v = input image · Z_v = visual features · W = trainable projection matrix · H_v = visual tokens in language embedding space
Open in Lab
Click any component to see its role in the LLaVA pipeline.
The demo wakes as you arrive…

Two-stage training: align, then instruct

LLaVA's training happens in two carefully designed stages — each solving a different problem:

Stage 1 — Pre-training for . The goal is to teach the projection layer to translate visual features into the language model's vocabulary. Using 595K filtered image-caption pairs from CC3M, the model learns simple captioning: "describe this image briefly." Only the projection matrix W is trained; both the vision encoder and the LLM stay frozen. Think of it as calibrating a translator — they learn the vocabulary before attempting conversation.

Stage 2 — Fine-tuning end-to-end. Now the model learns to follow instructions about images. Using the 158K GPT-4-generated instruction-following examples, both the projection layer and the LLM weights are updated (the vision encoder stays frozen). This is where the model learns to reason, describe in detail, hold multi-turn conversations, and connect visual content to its pre-trained knowledge.

Open in Lab
Toggle between Stage 1 and Stage 2 to see which components are frozen vs. trained.
The demo wakes as you arrive…
p(Xa∣Xv,Xinstruct)=∏i=1Lpθ(xi∣Xv,Xinstruct,<i,Xa,<i)p(X_a \mid X_v, X_{\text{instruct}}) = \prod_{i=1}^{L} p_\theta(x_i \mid X_v, X_{\text{instruct},<i}, X_{a,<i})
Autoregressive training objective — predict the assistant's answer token by token — X_v = image · X_instruct = instruction tokens · X_a = target answer tokens · θ = trainable parameters · the model maximizes the probability of generating the correct answer given the image and instruction

How the image becomes tokens

The elegance of LLaVA lies in how it unifies vision and language into a single stream. Here is the complete flow:

  1. The CLIP vision encoder processes the image and outputs a grid of feature vectors ZvZ_v (one per image patch, just like in ViT).

  2. The projection matrix W maps each feature into the same dimensional space as the LLM's word embeddings, producing visual tokens HvH_v.

  3. These visual tokens are concatenated with the tokenized text instruction HqH_q into a single input sequence.

  4. The LLM processes this unified sequence — visual tokens first, then text tokens — and generates the response autoregressively.

From the LLM's perspective, there is no difference between "seeing" an image and "reading" a word — both are just embeddings in the same space. The attention mechanism lets every word attend to every image patch and vice versa, enabling rich cross-modal reasoning.

Open in Lab
Watch image patches get projected into the LLM's token space and merged with text tokens.
The demo wakes as you arrive…

The idea in code

LLaVA forward pass — from image to conversationpython

Simplified to show the idea — not the real implementation.

import numpy as np

def clip_vision_encoder(image, patch_size=14):
    """CLIP ViT-L/14: split image into patches, encode each patch."""
    H, W, C = image.shape
    n_patches = (H // patch_size) * (W // patch_size)  # e.g. 256 for 224x224
    # In reality, a full ViT forward pass happens here
    visual_features = np.random.randn(n_patches, 1024)  # Z_v: (256, 1024)
    return visual_features

def project_to_language_space(Z_v, W_proj):
    """The simple linear projection: the entire bridge between vision and language."""
    H_v = Z_v @ W_proj  # (256, 1024) @ (1024, 4096) → (256, 4096)
    return H_v           # visual tokens, same dim as word embeddings

def llava_forward(image, text_tokens, W_proj, llm):
    """
    LLaVA: CLIP encodes the image → project → concatenate with text → LLM generates.
    """
    # Step 1: Extract visual features
    Z_v = clip_vision_encoder(image)           # (256, 1024)

    # Step 2: Project to language embedding space
    H_v = project_to_language_space(Z_v, W_proj)  # (256, 4096)

    # Step 3: Get text embeddings
    H_q = llm.embed(text_tokens)               # (T, 4096)

    # Step 4: Concatenate visual + text tokens into one sequence
    input_embeds = np.concatenate([H_v, H_q], axis=0)  # (256+T, 4096)

    # Step 5: LLM generates the response autoregressively
    response = llm.generate(input_embeds)
    return response

# That's it. The image is just more tokens to the LLM.
# The projection matrix W is all that connects vision to language.

Results: visual chat and Science QA

LLaVA was evaluated in two settings that test complementary capabilities:

Visual chatbot evaluation. Using GPT-4 as a judge, LLaVA scored 85.1% relative to a text-only GPT-4 baseline on a synthetic benchmark of 90 visual questions (conversations, detailed descriptions, and complex reasoning). On the more challenging in-the-wild benchmark with 60 questions from diverse domains, LLaVA scored 67.3% overall — significantly ahead of BLIP-2 (38.1%) and OpenFlamingo (19.1%). On complex reasoning specifically, LLaVA reached an impressive 81.7%.

Science QA. On this benchmark of 21K multimodal science questions, LLaVA achieved 90.92% accuracy. When combined with GPT-4 as a judge (LLaVA proposes, GPT-4 decides when they disagree), the ensemble reached 92.53% — a new state of the art, surpassing specialized methods like Multimodal Chain-of-Thought (91.68%).

Open in Lab
Compare LLaVA's performance against BLIP-2, OpenFlamingo, and the GPT-4 text baseline.
The demo wakes as you arrive…

What matters most: ablation findings

The paper's ablation studies reveal several key insights:

  • Instruction tuning is critical. Without it, the model scores only 21.5% — a 63-point drop. The raw model can barely follow instructions about images.

  • All three data types help. Removing complex reasoning drops performance by 5.7 points; removing detailed descriptions drops it by 2 points. Even conversation data alone is far better than no instruction tuning.

  • Pre-training matters. Skipping Stage 1 (feature alignment) and training directly on Science QA drops accuracy from 90.92% to 85.81% — a 5.1% gap. The alignment stage preserves pre-trained knowledge while building the visual bridge.

  • Visual features before the last layer work better. Using CLIP's second-to-last layer gives 90.92% vs 89.96% for the last layer — likely because the final layer focuses on abstract, global features while earlier layers retain localized spatial details.

  • Scale helps. The 13B model outperforms the 7B model by 1.08 percentage points.

Why it mattered

  1. 2023

    LLaVA (this paper)

    First visual instruction tuning with GPT-4-generated data. Linear projection connects CLIP to Vicuna. 85.1% relative to GPT-4 on visual chat, 92.53% on Science QA.

  2. 2023

    LLaVA-1.5

    Replaced the linear projection with a 2-layer MLP, used CLIP-ViT-L-336px for higher resolution, and added academic VQA data. State-of-the-art on 11 benchmarks with only 1.2M data, training in one day on 8 A100 GPUs.

  3. 2024

    LLaVA-NeXT

    Extended to higher-resolution images with dynamic tiling, stronger LLM backbones (LLaMA-3, Qwen-1.5), and zero-shot video understanding. Pushed the frontier of what open-source multimodal models can achieve.

LLaVA proved that the Transformer-based architecture does not need domain-specific fusion modules to handle multiple modalities. The same that processes text can process images — all you need is a bridge to translate visual features into the language model's token space. This insight is the foundation of every modern multimodal LLM.

CitationLiu, Li, Wu, Lee. Visual Instruction Tuning. NeurIPS, 2023.

Terms in this paper