Generative Models2024intermediate13 min read

Video Generation Models as World Simulators

نماذج توليد الفيديو بوصفها محاكيات للعالم

Brooks, T. · Peebles, B. · Holmes, C. · DePue, W. · Guo, Y. · Jing, L. · Schnurr, D. · Taylor, J. · Luhman, T. · Luhman, E. · Ng, C. · Wang, R. · Ramesh, A. — OpenAI Technical Report

The problem

Before Sora, video generation models were narrow specialists: each handled a fixed resolution, a fixed duration, and a specific category of content. Resizing, cropping, and trimming were standard preprocessing steps that destroyed and compositional information. No single model could generate videos of variable lengths, resolutions, and aspect ratios the way a large language model can handle any length of text. Moreover, there was no clear path to scale video generation models the way transformers had been scaled in language — without a unified "" representation for visual data.

The contribution

Sora: a diffusion that treats videos and images as sequences of spacetime patches in a compressed . A encodes raw pixels into a low-dimensional , which is then decomposed into fixed-size patches — the visual equivalent of text tokens. A transformer-based is trained to denoise these patches, conditioned on text prompts generated by a re-captioning pipeline inspired by DALL·E 3. The result is a single model that generates videos of variable durations (up to one minute), resolutions (up to 1920×1080), and aspect ratios, with emergent capabilities including , , and interactive world simulation.

The impact

Sora demonstrated that the -based transformer paradigm — proven in language (tokens) and images ( patches) — extends naturally to video. It showed that a single architecture can handle visual data at any resolution and duration, removing the need for task-specific preprocessing. Its emergent simulation capabilities — 3D consistency, long-range coherence, object permanence — suggested that scaling video models is a viable path toward general-purpose world simulators. Sora accelerated a wave of open-source video generation research and influenced every subsequent model from Vidu to CogVideoX to Wan.

A language model sees the world as a sequence of word tokens. Give it a prompt, and it assembles tokens one by one into coherent text. Now imagine doing the same thing — but instead of words, the tokens are tiny video puzzle pieces called patches, each capturing a small cube of space and time.

That's Sora's core insight: just as GPT learned that all text is tokens, Sora learned that all video is patches. A 10-second clip and a single photograph are the same kind of data — just different-sized grids of patches. Compress the video, slice it into patches, let a transformer learn how they fit together, and suddenly one model can generate anything from a portrait photo to a minute-long cinematic scene.

The core idea: all visual data is patches

The success of large language models rests on a simple unifying idea: all text — code, math, poetry, legal contracts — is a sequence of tokens. This lets one architecture and one objective handle everything.

Sora applies the same principle to visual data. The key question was: what is the visual equivalent of a text token? The answer: a spacetime patch. Just as Vision Transformer (ViT) slices an image into a grid of 2D patches, Sora extends this to 3D — slicing a video into small cubes that span a few pixels in width, a few in height, and a few frames in time.

But raw video is enormous. A one-minute 1080p video at 30fps contains over 6 billion pixels. Processing patches directly from raw pixels would be computationally impractical. The solution is a two-stage pipeline: first compress, then patch.

Open in Lab
Follow the pipeline: raw video → compression → latent space → spacetime patches → transformer tokens.
The demo wakes as you arrive…

Stage 1: the video compression network

Before creating patches, Sora feeds the raw video through a learned compression network — conceptually similar to the in a . This network reduces both spatial and temporal dimensions, projecting the video into a lower-dimensional latent space. A corresponding maps generated latents back to pixel space.

Think of it as translating a full HD movie into a compact "visual language." The latent representation preserves the essential structure of the video — shapes, motions, textures — while discarding perceptual redundancy. All subsequent training and generation happens in this compressed space, making the problem computationally tractable.

Stage 2: spacetime latent patches

Once the video is in latent space, Sora decomposes it into a sequence of spacetime patches — small 3D cubes carved from the latent representation. Each patch spans a local region in width, height, and time. These patches become the transformer's input tokens, exactly analogous to word tokens in a language model.

The critical benefit: because patches are fixed-size but the grid can be any shape, Sora naturally handles variable resolutions, durations, and aspect ratios. A 16:9 widescreen video simply produces a wider grid of patches; a 5-second clip produces fewer temporal slices than a 1-minute video. At time, you control the output by choosing the shape of the initial noise grid.

Images fit naturally too — they're just videos with a single frame, so the temporal dimension of each patch is one.

Open in Lab
Toggle between different resolutions and durations to see how the patch grid changes shape while individual patches stay the same size.
The demo wakes as you arrive…

The diffusion transformer: denoising patches at scale

Sora is a diffusion transformer (DiT) — a marriage of diffusion models and transformers. During training, clean latent patches are corrupted with Gaussian noise, and the model learns to predict the original clean patches given the noisy input and a text-conditioning signal. At inference time, the model starts from pure noise and iteratively denoises to produce a coherent video.

The text conditioning follows the re-captioning approach from DALL·E 3: instead of relying on short, noisy alt-text, a descriptive captioning model first produces detailed captions for every training video. At inference time, GPT expands short user prompts into rich descriptions before sending them to Sora. This two-stage prompting dramatically improves the faithfulness and quality of generated videos.

The key finding: diffusion transformers scale as effectively for video as they do for images and language. As training compute increases, sample quality improves markedly — not just sharpness, but coherence, physical plausibility, and compositional accuracy.

Open in Lab
Watch the denoising process: from pure noise through intermediate steps to the final clean video frame.
The demo wakes as you arrive…

Training on native sizes: no more cropping

Previous video generation approaches resized all training data to a standard size — for example, 4-second clips at 256×256 pixels. This destroys critical information: the natural composition and framing of the original content.

Sora takes a different approach: it trains on videos at their native resolution and aspect ratio. Because the patch representation is flexible — any grid shape works — the model doesn't need to crop or resize input data. This yields two concrete benefits:

Sampling flexibility: the same model can produce 1920×1080 widescreen videos, 1080×1920 vertical videos for mobile, or anything in between, all without retraining.

Better composition: models trained on square crops often produce awkward framing with subjects partially cut off. Training on native ratios lets the model learn the real-world distribution of scene compositions, leading to more natural and aesthetically pleasing results.

Open in Lab
Compare square-cropped vs. native aspect ratio training on framing quality.
The demo wakes as you arrive…

Language understanding and re-captioning

Text-to-video generation requires videos paired with text captions. The quality of these captions is crucial: vague or inaccurate descriptions lead to a model that can't faithfully follow user prompts.

Sora adopts the re-captioning technique from DALL·E 3. First, a descriptive captioner model is trained to produce highly detailed descriptions of every video in the training set — capturing not just the main subject but background elements, camera motion, lighting, style, and temporal progression. These rich captions replace the original short alt-text.

At inference time, a second trick is applied: GPT transforms the user's short prompt (e.g., "a dog on the beach") into a longer, detailed description (e.g., "a golden retriever running along a sandy beach at golden hour, ocean waves crashing in the background, handheld camera tracking the dog"). This prompt expansion bridges the gap between how users naturally write and what the model was trained on.

Prompting with images and video

Sora isn't limited to text-to-video generation. It can also be prompted with existing images or video, enabling a wide range of editing capabilities:

Image animation: given a static image (for example, from DALL·E) and a text prompt, Sora generates a video that brings the image to life while following the prompt's instructions.

Video extension: Sora can extend a video forward or backward in time. Starting from the same ending, different extensions can produce entirely different beginnings — or the model can extend both directions to create a seamless infinite loop.

Video-to-video editing: using SDEdit-style techniques, Sora can transform the style and environment of an input video zero-shot — for example, placing a scene in a different setting or changing its visual style.

Video interpolation: given two videos with entirely different subjects and compositions, Sora can generate smooth transitions that gradually morph one into the other.

Emergent simulation capabilities

Perhaps the most surprising finding is what happens when the model is scaled up. Without any explicit 3D or physics supervision, Sora develops emergent simulation capabilities — behaviors that no one specifically trained for:

3D consistency: as the virtual camera moves through a scene, people and objects shift correctly in three-dimensional space. The model has implicitly learned something about 3D geometry purely from 2D video data.

Object permanence: characters and objects persist even when occluded (hidden behind something) or when they temporarily leave the frame. The model maintains a coherent mental model of what exists in the scene.

Long-range coherence: across minute-long generations, the same character maintains their appearance — clothing, hair, body shape — across multiple shots and camera angles.

World interaction: the model can simulate simple causal effects. A painter leaves new brush strokes that persist, a person eats food and leaves bite marks — these are actions that change the state of the world over time.

Digital world simulation: Sora can render interactive digital environments like Minecraft, simultaneously controlling a player with a basic policy and rendering the game world in high fidelity.

Open in Lab
Click each emergent capability to see how it manifests and why it matters.
The demo wakes as you arrive…

The idea in code

Sora pipeline: compress → patch → denoise — pseudocodepython

Simplified to show the idea — not the real implementation.

import numpy as np

def sora_pipeline(video, text_prompt, model, encoder, decoder, captioner, gpt):
    """Simplified Sora training + inference pipeline."""

    # ── TRAINING ──
    # Step 1: Re-caption training videos with detailed descriptions
    detailed_caption = captioner.describe(video)     # "golden retriever on wet sand..."

    # Step 2: Compress raw video into latent space
    latent = encoder(video)                          # (T, H, W, C) → (t, h, w, c)
    # Temporal + spatial compression by orders of magnitude

    # Step 3: Extract spacetime patches (the "visual tokens")
    patches = extract_patches(latent, patch_size=(2, 2, 2))  # (N_patches, patch_dim)
    # Variable number of patches = variable resolution/duration

    # Step 4: Add noise and train the DiT to denoise
    noise = np.random.randn_like(patches)
    t = np.random.uniform(0, 1)                      # random timestep
    noisy_patches = (1 - t) * patches + t * noise    # interpolate clean→noise
    predicted_clean = model(noisy_patches, t, detailed_caption)
    loss = mse(predicted_clean, patches)             # predict the original patches

    # ── INFERENCE ──
    # Step 1: Expand user prompt via GPT
    rich_prompt = gpt.expand(text_prompt)            # "a dog" → detailed description

    # Step 2: Choose output shape by setting the patch grid
    grid = init_noise_grid(width=1920, height=1080, duration=60, patch_size=(2,2,2))

    # Step 3: Iteratively denoise
    for step in reversed(range(num_steps)):
        grid = model.denoise_step(grid, step, rich_prompt)

    # Step 4: Decode back to pixels
    output_video = decoder(grid)                     # latent → pixels
    return output_video

Limitations: not a real physics engine (yet)

Despite its impressive emergent capabilities, Sora has clear limitations as a world simulator. It does not accurately model many basic physical interactions: glass shattering doesn't fragment correctly, objects sometimes float or pass through surfaces, and eating food doesn't always produce the expected changes in state.

Long videos can develop temporal incoherencies — objects spontaneously appearing or disappearing, or physical laws being violated in subtle ways. The model has no explicit physics engine or 3D representation; everything it knows about the world was learned implicitly from patterns in video data.

OpenAI positions these not as fundamental barriers but as scaling challenges: the emergent capabilities already observed suggest that more data and compute will push the model further toward genuine physical understanding. Whether this implicit learning can truly replace explicit physical modeling remains one of the most debated questions in the field.

What Sora unlocked

  1. 2022

    Latent Diffusion (Stable Diffusion)

    Moved diffusion to latent space, making high-resolution image generation practical. Sora extends this idea to video.

  2. 2023

    Diffusion Transformers (DiT)

    Replaced U-Net with transformers for image diffusion. Showed that transformers scale better. Sora adopted DiT as its backbone.

  3. 2024

    Sora

    Unified visual generation via spacetime patches. Variable resolution, duration, aspect ratio. Emergent 3D consistency and world simulation.

  4. 2024

    Open-Sora, CogVideoX, Vidu

    Open-source community and industry adopted the spacetime-patch DiT paradigm, validating Sora's architectural choices.

  5. 2025

    Wan, Veo 3, Cosmos

    Next-generation models pushed longer videos, audio sync, and action-controllable world models, building directly on Sora's foundation.

Sora's deepest contribution isn't any one video it can generate — it's the idea that visual data of all kinds can be unified through patches the same way language was unified through tokens. This representation unlocks the same scaling benefits that transformed NLP: one architecture, one training objective, variable-sized inputs and outputs, and emergent capabilities that improve with scale. The question is no longer whether video generation can scale, but how far the emergent capabilities will go.

CitationBrooks, Peebles, Holmes, DePue, Guo, Jing, Schnurr, Taylor, Luhman, Luhman, Ng, Wang, Ramesh. Video Generation Models as World Simulators. OpenAI Technical Report, 2024.

Terms in this paper