Robotics2023advanced12 min read

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

RT-2: نماذج الرؤية-اللغة-الفعل تنقل معرفة الويب إلى التحكم الروبوتي

Brohan, A. · Brown, N. · Carbajal, J. · Chebotar, Y. · Chen, X. · Choromanski, K. · Ding, T. · Driess, D. · Dubey, A. · Finn, C. · Florence, P. · Fu, C. · Gonzalez Arenas, M. · Gopalakrishnan, K. · Han, K. · Hausman, K. · Herzog, A. · Hsu, J. · Ichter, B. · Irpan, A. · Joshi, N. · Julian, R. · Kalashnikov, D. · Kuang, Y. · Leal, I. · Lee, L. · Lee, T. · Levine, S. · Lu, Y. · Michalewski, H. · Mordatch, I. · Pertsch, K. · Rao, K. · Reymann, K. · Ryoo, M. · Salazar, G. · Sanketi, P. · Sermanet, P. · Singh, J. · Singh, A. · Soricut, R. · Tran, H. · Vanhoucke, V. · Vuong, Q. · Wahid, A. · Welker, S. · Wohlhart, P. · Wu, J. · Xia, F. · Xiao, T. · Xu, P. · Xu, S. · Yu, T. · Zitkovich, B. — CoRL

The problem

By 2023, robotic control policies trained on demonstration data could reliably execute tasks they had been shown, but failed when faced with novel objects, unfamiliar instructions, or situations requiring common-sense reasoning. Meanwhile, large vision-language models (VLMs) trained on Internet-scale data possessed rich semantic knowledge — object recognition, spatial reasoning, world knowledge — but could not translate that understanding into physical actions. Prior approaches used VLMs only as high-level planners that issued commands to separate low-level controllers, never bridging the gap between semantic understanding and direct motor control.

The contribution

RT-2 introduces vision-language-action (VLA) models: take a pre-trained vision-language model (PaLI-X or PaLM-E), represent robot actions as text tokens by discretizing each of 8 action dimensions into 256 bins, and co-fine-tune the model on both web-scale vision-language data and demonstrations simultaneously. The robot's actions become just another "language" the model can speak. No architectural changes needed — the same model that answers "what color is this object?" can output "1 128 91 241 5 101 127" to move a robot arm. Over 6,000 evaluation trials showed RT-2 matches RT-1 on seen tasks while dramatically improving generalization: 3× higher success rate on emergent capability tasks, and the ability to perform semantic reasoning like choosing a rock as an improvised hammer.

The impact

RT-2 proved that internet-scale knowledge can transfer directly into physical robot control — not just planning, but the actual motor commands. It established the VLA paradigm that became the foundation for OpenVLA, RT-2-X (Open X-Embodiment), and the wave of models that followed. The paper demonstrated that scaling vision-language models benefits robotics the same way it benefits NLP: bigger models with more web data produce more capable, more generalizable robots, creating a direct bridge between the foundation model revolution and physical-world AI.

Teaching a robot to manipulate objects is like training a chef who has never tasted food. You can show them thousands of recipes and arm movements, and they will copy the exact dishes they have seen — but ask them to "make something refreshing for a hot day" and they are lost, because they have no concept of heat, thirst, or refreshment.

RT-2 is like hiring a chef who has already read every cookbook, watched every cooking show, and browsed every food blog on the Internet — then teaching them knife skills. Now when you say "make something refreshing," they reach for the watermelon, because they already know what refreshing means. The knife skills are new, but the world knowledge is not.

The key trick: the chef writes recipes and arm movements in the same notebook, so the wisdom from cookbooks flows directly into how they move their hands.

The gap: robots that see but don't understand

Before RT-2, robotic control faced a fundamental knowledge barrier. Models like RT-1 were trained on tens of thousands of robot demonstrations and could reliably pick up an apple or move a can to the left. But this knowledge was procedural only — the robot learned motor patterns, not concepts.

Ask RT-1 to "pick up the object you would use as a hammer" and it has no idea what to do. It never saw that instruction in training. It does not know that hammers are heavy and hard, that a rock shares those properties, or that an energy drink is what a tired person needs. All of this is world knowledge — the kind that vision-language models absorb from billions of web pages.

Meanwhile, large vision-language models like PaLI-X and PaLM-E could answer "which object here could serve as a hammer?" with ease — but they could not move a robot arm. They operated in the world of words, not the world of physics. The missing bridge: how do you pour that semantic knowledge into the motor commands themselves?

The idea: robot actions are just another language

RT-2's breakthrough rests on a disarmingly simple idea: treat robot actions as text tokens. A robot arm action has 8 dimensions — positional displacement in x, y, z, rotational displacement in roll, pitch, yaw, gripper extension, and a termination flag. Each continuous dimension is divided into 256 uniform bins, turning it into an integer. The full action becomes a string like "1 128 91 241 5 101 127" — and as far as the language model is concerned, this is just another sentence to generate.

This means you do not need a new architecture. You take an existing vision-language model that already understands images and text, and you teach it one more "language": the language of robot movements. The model's input is a camera image of the robot's workspace plus a natural language instruction like "pick up the blue block." Its output is a sequence of action tokens that the robot executes.

Open in Lab
See how a continuous robot action (arm movement) is discretized into 256 bins per dimension, then expressed as a text string the language model can generate.
The demo wakes as you arrive…

The choice of which tokens represent actions depends on the backbone. PaLI-X already has unique tokens for integers 0–999, so actions simply reuse the number tokens. PaLM-E does not have convenient integer tokens, so the 256 least frequently used tokens in its vocabulary are overwritten with action meanings — a form of where existing symbols acquire new semantics through training.

The training prompt follows a simple Q&A format: "Q: what action should the robot take to [pick up the blue block]? A: 1 128 91 241 5 101 127". From the model's perspective, this is no different from — the same loss function, the same generation.

Architecture: two VLM backbones, one recipe

RT-2 is not a single model — it is a recipe applied to two different vision-language model backbones. Think of it as a cooking technique that works with different base ingredients:

RT-2-PaLI-X uses PaLI-X as its backbone — a ViT-22B visual encoder paired with a 32B text model, totaling up to 55 billion parameters. The visual encoder splits the camera image into patches (just like ViT), projects them into tokens, and the encoder-decoder generates the action string autoregressively.

RT-2-PaLM-E uses PaLM-E — a 12 billion model that weaves visual tokens directly into the text token stream. It processes images and text in a single unified sequence, producing actions as the continuation of that sequence.

Both models take the same input — a camera image plus an instruction — and produce the same output: a string of 8 action tokens. The key finding: the recipe works regardless of backbone architecture, and bigger models generalize better.

Open in Lab
Click any layer to see what it does in the RT-2 pipeline.
The demo wakes as you arrive…

Co-fine-tuning: the secret ingredient

A naïve approach would be to simply fine-tune the VLM on robot data — replace web images and text with robot camera images and action strings. But this causes : the model loses its web knowledge, the very thing we want to transfer.

RT-2's solution is : in every training batch, robot trajectory data and original web data (visual question answering, image captioning, etc.) are mixed together. The robot data is up-sampled to occupy 50–66% of each , ensuring the model gets enough action examples to learn control, while the remaining web data acts as an anchor that prevents the model from forgetting what a hammer looks like or what "tired" means.

This is conceptually similar to how language model mixes instruction data with original pre-training data — but here the stakes are different. If the model forgets "red" it cannot follow "pick up the red cup." The co-training ratio is not just a training detail; it is the mechanism by which web knowledge flows into robotic control.

Open in Lab
Compare fine-tuning (robot data only, forgets web knowledge) with co-fine-tuning (mixed batches, preserves world knowledge for generalization).
The demo wakes as you arrive…

Emergent capabilities: what web knowledge unlocks

The most striking result of RT-2 is not improved performance on known tasks — it is the emergence of entirely new capabilities that were never present in the robot training data. These capabilities come purely from the VLM's web knowledge flowing into robotic control:

  • Symbol understanding: "Move the apple to the number 3" — the robot was never trained on numbers, but the VLM knows what "3" looks like.
  • Semantic reasoning: "Pick up something you could use as a hammer" — the robot grabs a rock, not a sponge, because the VLM knows that hammers are heavy and hard.
  • Human recognition: "Give the snack to the person on the left" — the VLM can distinguish people in the scene.
  • Relative reasoning: "Pick up the smallest object" or "the one closest to the red block" — spatial and comparative reasoning from web .
  • Common-sense inference: "Which drink would help a tired person?" — the robot picks up an energy drink, applying world knowledge no one ever demonstrated to it.
Open in Lab
Explore examples of emergent capabilities — tasks the robot was never trained on but can solve thanks to web knowledge from the VLM backbone.
The demo wakes as you arrive…

Chain of thought: robots that reason before acting

Inspired by chain-of-thought prompting in language models, the authors trained a variant of RT-2 that generates a natural language reasoning plan before outputting action tokens. Instead of input → action, the flow becomes input → plan → action.

For example, given the instruction "pick up something to help clean up a spill," the model first outputs: "I would pick up the sponge because it can absorb liquid." Then it emits the action tokens to grab the sponge.

This is not merely explanatory — the planning step actually improves task success, because the model can decompose multi-step reasoning into explicit intermediate steps. And crucially, the reasoning is inspectable: an operator can read the robot's plan before it acts, catching errors before they become physical mistakes.

Open in Lab
Step through RT-2's chain-of-thought process: the model reasons about the task, outputs a natural-language plan, then generates action tokens.
The demo wakes as you arrive…

Results: 3× better on novel tasks

The evaluation covered more than 6,000 real-robot trials across three categories:

Seen tasks — instructions and objects from the training set. RT-2 matched RT-1's strong performance (both above 90% success), confirming that co-fine-tuning does not sacrifice known-task reliability.

Generalization — same task types but with novel objects, backgrounds, or environments not seen during training. RT-2 approximately doubled RT-1's success rate, showing that VLM knowledge transfers to manipulation even when the exact visual scenario is new.

— entirely new task categories requiring symbol understanding, semantic reasoning, or human recognition. RT-2 achieved over 3× the average success rate of the best baseline (RT-1 and VC-1). The 55B PaLI-X variant consistently outperformed the 12B PaLM-E variant, confirming that larger VLM backbones carry more transferable knowledge.

On the open-source Language Table benchmark, RT-2 reached 90% success versus the prior state-of-the-art of 77%.

Open in Lab
See how increasing model size and adding co-fine-tuning improve generalization performance. The interactive chart lets you compare training recipes.
The demo wakes as you arrive…

What RT-2 unlocked

  1. 2022

    RT-1 — Robotics Transformer

    A 35M parameter Transformer trained on 130K robot demonstrations. High success on seen tasks, but limited generalization — no world knowledge beyond what the robot saw.

  2. 2023

    RT-2 — Vision-Language-Action Models

    Proved internet-scale VLM knowledge transfers directly into motor control. Actions as tokens, co-fine-tuning preserves web knowledge, emergent reasoning appears.

  3. 2023

    RT-2-X (Open X-Embodiment)

    Extended RT-2 to cross-embodiment learning — training on data from 22 different robot types across 21 institutions. A single policy that works on multiple robot bodies.

  4. 2024

    OpenVLA

    Open-source VLA built on the RT-2 recipe with a 7B parameter backbone. Made the VLA paradigm accessible to the broader robotics community without Google-scale compute.

  5. 2025

    Embodied AI wave

    Physical Intelligence (π₀), GR00T, and many others built on the VLA paradigm RT-2 established. Robot foundation models are following the same scaling trajectory as LLMs.

RT-2's deepest contribution is not a benchmark number — it is the proof that the foundation model revolution applies to physical robots, not just text and images. The same recipe that made language models powerful — massive pre-training on web data, then task-specific fine-tuning — works for robotic control. And the connection is literal: actions are tokens, and knowledge flows through the same weights that process language.

The same idea in code

Action tokenization and VLA inference, simplifiedpython

Simplified to show the idea — not the real implementation.

import numpy as np

# Step 1: Discretize a continuous robot action into 256 bins
def tokenize_action(action_vector, n_bins=256):
    """Turn 7 continuous floats + 1 binary into 8 integer tokens."""
    # action_vector: [dx, dy, dz, drx, dry, drz, gripper, terminate]
    tokens = []
    for i, val in enumerate(action_vector):
        if i == 7:  # terminate is already binary
            tokens.append(int(val))
        else:
            # Map from action range [-1, 1] to bin [0, 255]
            bin_idx = int((val + 1) / 2 * (n_bins - 1))
            bin_idx = np.clip(bin_idx, 0, n_bins - 1)
            tokens.append(bin_idx)
    return tokens   # e.g. [1, 128, 91, 241, 5, 101, 127, 0]

# Step 2: Format as a text string for the VLM
def action_to_text(tokens):
    return " ".join(str(t) for t in tokens)
    # "1 128 91 241 5 101 127 0"

# Step 3: The VLM prompt template
def make_prompt(instruction):
    return f"Q: what action should the robot take to {instruction}? A:"

# That's it. The VLM generates "1 128 91 241 5 101 127 0"
# as text, and we parse it back into motor commands.
# The same model that answers "what color is this?" also
# outputs arm movements — language and actions share one brain.

CitationBrohan, Brown, Carbajal, Chebotar, Chen, Choromanski, Ding, Driess, Dubey, Finn, Florence, Fu, Gonzalez Arenas, Gopalakrishnan, Han, Hausman, Herzog, Hsu, Ichter, Irpan, Joshi, Julian, Kalashnikov, Kuang, Leal, Lee, Lee, Levine, Lu, Michalewski, Mordatch, Pertsch, Rao, Reymann, Ryoo, Salazar, Sanketi, Sermanet, Singh, Singh, Soricut, Tran, Vanhoucke, Vuong, Wahid, Welker, Wohlhart, Wu, Xia, Xiao, Xu, Xu, Yu, Zitkovich. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. CoRL, 2023.

Terms in this paper