Multimodal AI2023advanced12 min read

Gemini: A Family of Highly Capable Multimodal Models

Gemini: عائلة نماذج فائقة القدرات تجمع بين وسائط متعددة

Gemini Team · Anil, R. · Borgeaud, S. · Alayrac, J.-B. · Yu, J. · Soricut, R. · Schalkwyk, J. — arXiv

The problem

By late 2023, large language models had shown remarkable text reasoning abilities, but their understanding of images, audio, and video remained limited. Most systems bolted vision or audio modules onto a pre-trained text model as an afterthought — a pipeline that lost cross-modal nuance. The field needed a model trained natively on all modalities from the start, one that could reason across text, images, audio, and video as fluidly as humans do, and that could scale from powerful server deployments down to on-device applications.

The contribution

Gemini introduced a family of natively multimodal models — Ultra, Pro, and Nano — trained jointly on text, image, audio, and video data from the ground up. Rather than stitching separate encoders together, Gemini processes interleaved multimodal sequences through a unified , enabling genuine cross-modal reasoning. Gemini Ultra became the first model to surpass human-expert performance on the (90.04%) and set state-of-the-art results on 30 of 32 benchmarks, including all 20 multimodal benchmarks tested. The Nano variants demonstrated that distillation and 4-bit can bring frontier capabilities to mobile devices.

The impact

Gemini reshaped the competitive landscape of AI. It proved that native multimodal outperforms the pipeline approach of gluing vision onto language models. Its three-tier sizing strategy — Ultra, Pro, Nano — became a template for deploying frontier models across the full compute spectrum. The Nano distillation path directly led to the Gemma open-weight family, democratizing access to Gemini-quality capabilities. The paper's demonstration that a single architecture can excel across text, vision, audio, and video set the direction for all subsequent foundation model development.

Imagine a conductor who grew up playing every instrument in the orchestra — not someone who learned piano first and then tried to conduct. Because she knows how a violin phrase feels under the fingers, how a drum beat sounds from behind the kit, and how a choir lyric reads on the page, she can bring them all together into something no specialist-turned-conductor could.

Previous multimodal AI was like hiring a pianist to conduct: talented at one instrument, but awkwardly translating everything else into piano terms. Gemini is the conductor who speaks every instrument's language natively.

The problem: why bolting modalities together falls short

Before Gemini, the dominant approach to multimodal AI was a two-stage pipeline. First, train a powerful language model on text. Then, attach a separate — often a — and train an adapter to translate image features into the language model's space. Models like Flamingo and PaLI pioneered this approach with impressive results.

But this pipeline has a fundamental limitation: the language model was never designed to think in images. It receives a translated summary of visual information, not the raw signal. Subtle spatial relationships, fine-grained textures, and temporal patterns in video or audio are compressed or lost during translation. The model cannot reason about what it has not truly seen.

This is like learning about music by reading descriptions of songs rather than listening to them. You can memorize facts about tempo and key signatures, but you cannot truly understand how a melody feels — because you have never heard one.

Open in Lab
Compare the pipeline approach (separate encoders stitched to a text model) with Gemini's native approach (all modalities trained together from scratch).
The demo wakes as you arrive…

Architecture: one Transformer, every modality

At its core, Gemini is a Transformer decoder — the same fundamental architecture behind GPT-4 and PaLM. But what makes it different is what goes in. Instead of processing only text tokens, Gemini accepts interleaved sequences of text, image, audio, and video tokens, all feeding into the same decoder stack.

Images are split into patches (like a ViT encoder) and projected into the same embedding dimension as text tokens. Video is treated as a sequence of image frames. Audio signals at 16kHz are captured via Universal Speech Model (USM) features, preserving nuances lost when audio is naively transcribed to text. All these modality tokens flow through the same layers, letting the model learn cross-modal relationships directly.

Gemini uses for efficiency and supports a 32k . The models are optimized for Google's accelerators, with architectural improvements enabling stable training at unprecedented scale.

Open in Lab
Explore how text, image, audio, and video tokens flow through Gemini's unified Transformer decoder. Click each modality to see its tokenization.
The demo wakes as you arrive…

The key innovation is not in any single architectural component — multi-query attention, patch embeddings, and decoder-only Transformers all existed before. The innovation is the decision to train a single model on all modalities jointly from scratch, letting cross-modal understanding emerge naturally rather than being retrofitted.

This is like the difference between learning to cook by following recipes (pipeline approach) and learning by tasting, smelling, and experimenting with ingredients directly (native approach). Both produce meals, but the native cook develops intuitions that no recipe can teach.

Three sizes, one family: Ultra, Pro, and Nano

Gemini 1.0 comes in three sizes, each tailored for a different computational reality. Ultra is the largest and most capable, designed for highly complex reasoning tasks where accuracy matters more than latency. Pro is optimized for cost and latency while retaining strong performance across a wide range of tasks — the workhorse for production deployment at scale. Nano is designed for on-device applications, coming in two sub-variants: Nano-1 (1.8B parameters) for low-memory devices and Nano-2 (3.25B parameters) for higher-memory devices.

The Nano models are particularly interesting because they are distilled from larger Gemini models and then quantized to 4 bits for deployment. This distillation-then-quantization pipeline demonstrates that frontier-level understanding can be compressed into a model small enough to run on a smartphone, without requiring an internet connection.

Open in Lab
Compare the three Gemini sizes: Ultra for the hardest tasks, Pro for balanced deployment, Nano for on-device intelligence.
The demo wakes as you arrive…

Training: data, tokenization, and scale

Gemini models are trained on a dataset that is both multimodal and multilingual, encompassing web documents, books, code, images, audio, and video. The diversity of this corpus is critical: by seeing text alongside the images and sounds it describes, the model learns associations that no text-only corpus could provide.

The tokenizer is , trained on a large sample of the full training corpus. This corpus-aware is key: by training the tokenizer on the same data the model will see, the vocabulary captures patterns specific to that data, including efficient tokenization of non-Latin scripts. This improves both model quality and speed for multilingual users — a single tokenizer handles Arabic, Chinese, code, and English without privileging any one language.

Training scale follows the Chinchilla approach from Hoffmann et al.: the largest models are trained on a compute-optimal number of tokens, while the smaller Nano models are deliberately overtrained — given more tokens than the Chinchilla suggests — to maximize their performance within a fixed inference budget, following the philosophy advocated in LLaMA.

Open in Lab
See how multimodal data flows through Gemini's training pipeline: from raw data sources through tokenization to the unified Transformer.
The demo wakes as you arrive…

Results: surpassing human experts on MMLU

Gemini Ultra achieved 90.04% accuracy on MMLU — the first model to surpass the human-expert threshold of 89.8%. MMLU tests knowledge across 57 diverse subjects from law and biology to history and mathematics. The prior state-of-the-art was 86.4%.

This result used a novel evaluation technique called uncertainty-routed . For questions where the model is confident, it answers directly. For harder questions, it uses chain-of-thought reasoning to work through the problem step by step. The routing is based on the model's own uncertainty estimate — if the top predicted answer has low confidence, the model "thinks out loud" before answering. This combination improved accuracy from 83.7% (5-shot) to 90.04%.

Beyond MMLU, Gemini Ultra also achieved 94.4% on GSM8K (grade-school math), 74.4% on HumanEval (code generation), and 53.2% on MATH (competition-level mathematics). Across all 32 text and reasoning benchmarks, Gemini Ultra advanced the state of the art on 30 of them.

Open in Lab
Compare Gemini Ultra against GPT-4 and prior state-of-the-art across key benchmarks. Click a benchmark name for details.
The demo wakes as you arrive…

Multimodal results: state of the art across vision, audio, and video

The multimodal results are where Gemini truly distinguishes itself. On every one of the 20 multimodal benchmarks tested, Gemini set a new state of the art. This includes image understanding tasks like MMMU (62.4% — college-level multimodal reasoning), TextVQA (visual question answering with text in images), DocVQA (understanding documents), and MathVista (mathematical reasoning from diagrams).

For video, Gemini can process sequences of frames directly within its large context window, enabling temporal reasoning that pipeline models struggle with — like understanding cause-and-effect in a video clip or tracking objects across scenes. For audio, the USM-based encoding preserves pitch, tone, and speaker identity, enabling nuanced understanding beyond simple transcription.

These results confirm the hypothesis behind native multimodal training: a model that has always seen text and images together develops richer cross-modal representations than one that learned text first and images second.

Open in Lab
Explore examples of Gemini's cross-modal reasoning: understanding handwritten math, interpreting charts, and reasoning about images.
The demo wakes as you arrive…

Post-training: alignment, safety, and responsible deployment

After large-scale pretraining, Gemini models undergo post-training to improve quality, align with human preferences, and enforce safety criteria. This involves two stages. First, (SFT) on high-quality demonstration data teaches the model to follow instructions and produce well-structured, helpful responses. Second, () further refines the model to be more helpful, accurate, and harmless.

Safety is treated as a first-class concern, not an afterthought. The team conducted impact assessments, developed model policies covering content safety and representational harms, and performed extensive evaluations including red-teaming. For multimodal outputs specifically, safety datasets and filters were developed to mitigate harmful image-to-text generations. Post-training interventions reduced factuality errors from 6.7% to 3.8% and increased output attribution rates.

This comprehensive approach reflects a maturing understanding in the field: building capable models and deploying them safely are not competing goals — they are complementary requirements.

Cross-modal reasoning: what Gemini can do that others cannot

The paper showcases several qualitative examples that demonstrate genuinely new capabilities. In one example, a student submits a handwritten physics solution. Gemini reads the handwriting, understands the physics problem setup, verifies the mathematical reasoning, identifies an error in the student's work, and generates corrected LaTeX — all in one pass. This requires simultaneously understanding handwriting recognition, physics knowledge, mathematical verification, and LaTeX generation.

In another example, Gemini examines a plot generated by matplotlib code and suggests improvements — understanding both the visual appearance of the chart and the code that produced it. In video understanding, the model can analyze a sequence of soccer game frames and suggest technique improvements for a player.

These examples are not cherry-picked tricks but natural consequences of native multimodal training: when a model has always seen text, images, code, and math together, combining them is not a special capability — it is the default mode of operation.

Code: how multimodal tokens flow through Gemini

Pseudocode: Gemini multimodal input processingpython

Simplified to show the idea — not the real implementation.

# Gemini processes interleaved multimodal sequences
# All modalities become tokens in a single sequence

def process_multimodal_input(text, images, audio, video):
    tokens = []

    # Text: SentencePiece tokenization
    text_tokens = sentencepiece_tokenize(text)
    tokens.extend(text_tokens)

    # Images: split into patches, project to embedding dim
    for img in images:
        patches = split_into_patches(img, patch_size=14)
        img_tokens = visual_encoder(patches)  # ViT-style
        tokens.extend(img_tokens)

    # Audio: extract USM features at 16kHz
    if audio is not None:
        audio_features = usm_encode(audio, sample_rate=16000)
        tokens.extend(audio_features)

    # Video: encode as sequence of frames
    if video is not None:
        frames = extract_frames(video)
        for frame in frames:
            frame_tokens = visual_encoder(split_into_patches(frame))
            tokens.extend(frame_tokens)

    # All tokens flow through the same Transformer decoder
    # Multi-query attention enables efficient processing
    output = transformer_decoder(
        tokens,
        context_length=32768,
        attention_type="multi_query"
    )
    return output

Legacy: from Gemini to the multimodal era

  1. 2022

    Flamingo & PaLI: multimodal pioneers

    Google's Flamingo and PaLI showed that visual encoders could be integrated with language models for multimodal understanding — but as separate modules stitched together after text pretraining.

  2. 2023

    GPT-4V: text-first multimodality

    OpenAI's GPT-4V extended GPT-4 with vision capabilities. Powerful but still a text-first model with vision added — the pipeline approach at its best.

  3. 2023

    Gemini 1.0: native multimodal training

    First model trained natively across text, image, audio, and video. Achieved SOTA on 30/32 benchmarks. Introduced the Ultra/Pro/Nano family structure.

  4. 2024

    Gemini 1.5: million-token context

    Extended the context window from 32K to 1 million tokens with near-perfect recall, enabling understanding of entire books, hours of video, and massive codebases in a single prompt.

  5. 2024

    Gemma: open-weight descendants

    Google released the Gemma family — open-weight models distilled from Gemini technology, making frontier-quality capabilities accessible to the broader research community.

Gemini's core contribution extends beyond its benchmark numbers. It established that native multimodal training produces fundamentally different — and superior — cross-modal understanding compared to the pipeline approach. Every major AI lab has since moved toward native multimodal architectures. The three-tier deployment strategy (server, edge, device) became the standard model for bringing AI capabilities to users everywhere. And the distillation pathway from Ultra through Nano to Gemma demonstrated that open science and frontier capability are not mutually exclusive.

CitationGemini Team, Anil, Borgeaud, Alayrac, Yu, Soricut, Schalkwyk, et al.. Gemini: A Family of Highly Capable Multimodal Models. arXiv, 2023.

Terms in this paper