Robotics2024intermediate12 min read

Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Open X-Embodiment: مجموعات بيانات التعلّم الروبوتي ونماذج RT-X

Open X-Embodiment Collaboration · O'Neill, A. · Rehman, A. · Gupta, A. · Padalkar, A. · Pertsch, K. · Levine, S. · Finn, C. · Karol, H. — ICRA

The problem

In NLP and computer vision, the field converged on large pretrained backbones — one model serves many tasks. Robotics never had that luxury. Every lab trains a separate model for every robot, every task, and every environment. Datasets are small, siloed, and stored in incompatible formats. The " gap" — different robots have different sensors, kinematics, and action spaces — makes it unclear whether data from one robot can help another at all. There was no ImageNet for robotics, and no evidence that cross-robot even works.

The contribution

Two contributions. First, the Open X-Embodiment Dataset: over one million real-robot trajectories from 22 different embodiments, collected across 21 institutions and 60 existing datasets, all unified into the RLDS format. Second, two trained models — RT-1-X (35M parameters) and RT-2-X (55B parameters) — that demonstrate : RT-1-X achieves 50% higher success than robot-specific baselines in small-data domains, and RT-2-X shows 3× improvement on emergent skills the evaluation robot never practiced.

The impact

Open X-Embodiment proved that cross-robot works in robotics — a question the field had debated for years. It created the first large-scale open-source robotics dataset and inspired a wave of generalist robot policies including Octo, RT-H, and π₀. The RLDS-based data standard it established became the de-facto format for multi-robot datasets. By winning the ICRA 2024 Best Paper Award, it signaled that the era of single-robot, single-task learning is giving way to foundation models for robotics.

Imagine 21 hospitals, each with a different type of surgical robot. Every hospital trains its robot from scratch — its own data, its own model, its own format. No hospital can use another's training records because the robots have different arms, different cameras, and different filing systems.

Now imagine a medical consortium that collects every surgery recording from every hospital, translates them into one common format, and trains a single AI surgeon that has watched a million procedures across all robot types. This surgeon doesn't just match the specialists — it outperforms them, because experience on one robot teaches skills that transfer to others.

Open X-Embodiment is that consortium for robotic : 21 institutions, 22 robots, one dataset, one format, and two models (RT-X) that prove generalist beats specialist.

The Problem: every robot is an island

By 2023, the pretrain → fine-tune paradigm had transformed NLP and computer vision. Models like GPT and CLIP learn broad representations from massive internet data and then adapt to specific tasks with minimal . But robotics was stuck in the pre-ImageNet era: every research group trained a separate model for its own robot, its own task set, and its own lab environment.

The reasons were practical. Different robots have different kinematics — a 7-DOF Franka arm moves nothing like a WidowX or a quadruped. Their cameras sit in different positions. Their action spaces range from joint velocities to deltas. And every lab stored data in its own format, making sharing impossible. This is the embodiment gap: the physical differences between robots that make it unclear whether data from one can teach another.

The result was chronic data scarcity. Each lab might collect a few thousand demonstrations, but that's far too few for a deep model to generalize broadly. Meanwhile, in NLP, billion-parameter models were being trained on trillions of tokens. Robotics needed its own data unification moment — and proof that training actually produces positive transfer rather than interference.

Open in Lab
Explore how different robots differ in kinematics, cameras, and action spaces — illustrating the embodiment gap that this paper bridges.
The demo wakes as you arrive…

The Open X-Embodiment Dataset

The first contribution is a dataset of unprecedented scale for robotics. The Open X-Embodiment Dataset pools 60 existing datasets from 34 research labs across 21 institutions worldwide. It contains over 1 million real robot trajectories from 22 different robot embodiments — from single-arm manipulators like the Franka Emika and WidowX, to bi-manual robots and quadrupeds.

The key challenge was format unification. Each lab stored its data differently — different image resolutions, different action representations, different file formats. The authors adopted the RLDS (Reinforcement Learning Dataset) format, which stores data in serialized TFRecord files and accommodates varying action spaces, camera configurations, and sensor modalities. This means a researcher anywhere can load any subset of the dataset with the same data pipeline, regardless of which robot collected it.

To understand what the dataset contains, the authors used the PaLM language model to extract objects and behaviors from the natural language instructions. The dataset spans 527 skills across 160,266 tasks. While pick-and-place dominates, the long tail includes skills like wiping, assembling, cable routing, and drawer manipulation. Objects range from food items to appliances to tools.

Open in Lab
Explore the scale and diversity of the Open X-Embodiment Dataset across robots, institutions, skills, and objects.
The demo wakes as you arrive…

Bridging the embodiment gap: coarse action alignment

Different robots have wildly different action spaces. Some report joint velocities, others report Cartesian end-effector deltas, and gripper conventions vary from binary open/close to continuous force control. To train a single model on all of them, the authors chose a simple but effective compromise: coarse alignment.

Every robot's actions are converted to a 7-dimensional end-effector vector: Δx, Δy, Δz (translation), Δroll, Δpitch, Δyaw (rotation), and gripper opening. For robots that don't use some dimensions (e.g., a 4-DOF arm that cannot roll), the unused dimensions are zeroed out. Each dimension is then discretized into 256 bins, and an eighth "terminate episode" dimension is added — giving a total of 8 tokenized action dimensions.

For observations, one canonical camera view per dataset is selected and resized to a common resolution. Language instructions are included as-is. This coarse alignment intentionally does not try to perfectly harmonize across embodiments — for instance, the same Δx=0.01 might mean different physical displacements on different robots, but the normalization per dataset handles this. The key insight is that even coarse alignment is enough to enable positive transfer.

Open in Lab
See how a continuous robot action is discretized into 256 bins per dimension, creating a unified action language across embodiments.
The demo wakes as you arrive…

Two models, one dataset: RT-1-X and RT-2-X

Rather than designing a new architecture, the authors took two existing models — RT-1 and RT-2 — and trained them on the cross-embodiment data mixture. The results are called RT-1-X and RT-2-X.

RT-1-X is a 35M-parameter designed specifically for robotic control. It processes a history of 15 images through an ImageNet-pretrained , converts the language instruction into a Universal Sentence Encoder (USE) , and weaves the two modalities together via layers. This produces 81 vision-language tokens that feed into a Transformer, which outputs the 8 tokenized action dimensions. RT-1-X is trained only on robotics data from the X-embodiment mixture.

RT-2-X is a 55B-parameter . It takes a pretrained (VLM) and co-fine-tunes it on both its original internet-scale vision-language data and the robotics data in roughly a 1:1 ratio. The clever trick is that RT-2 casts tokenized actions as text tokens — for example, the action might be represented as the string "1 128 91 241 5 101 127". This means the VLM's existing language modeling head can output robot actions without any architectural change, while retaining the semantic knowledge from web .

Open in Lab
Compare the RT-1-X and RT-2-X architectures side by side: how each processes images and language to produce robot actions.
The demo wakes as you arrive…

Training recipe

Both models are trained using standard categorical cross-entropy loss over the tokenized actions. Each of the 8 action dimensions (7 end-effector + 1 termination) is predicted as a problem over 256 bins.

The robotics data mixture used in the experiments includes 9 embodiments (a subset of the full 22 in the dataset — the rest were added over time). The mixing strategy weighs datasets to prevent the largest datasets from dominating training.

RT-1-X is trained solely on the robotics mixture. RT-2-X uses : roughly half of each batch comes from the original VLM data (internet images with text) and half from the robotics mixture. This prevents the large model from forgetting its web-learned knowledge while acquiring robotic skills — a technique similar to how language models use replay buffers to prevent .

The models receive a history of recent images (15 frames for RT-1-X) plus the natural language instruction. One canonical camera view per dataset is selected and resized to a common resolution. Actions are normalized per-dataset before , so the model's output can be de-normalized back to the physical units of whatever robot is being controlled.

L=−∑d=18∑k=1256yd,klog⁡p^d,k\mathcal{L} = -\sum_{d=1}^{8} \sum_{k=1}^{256} y_{d,k} \log \hat{p}_{d,k}
Action prediction loss — Standard categorical cross-entropy summed over all 8 action dimensions. For each dimension d, the ground-truth bin k is one-hot encoded as y, and the model outputs a probability distribution p̂ over 256 bins.

Results: positive transfer across embodiments

The central experimental question is: does training on data from many robots actually help each individual robot? The answer is a decisive yes.

RT-1-X on small-data domains: When evaluated on robots that contribute relatively small datasets to the mixture, RT-1-X outperforms both the original lab-specific methods and RT-1 trained only on that robot's data by an average of 50%. This is the clearest evidence of positive transfer — the model learns general manipulation knowledge from other robots that transfers to robots with limited data.

RT-1-X on large-data domains (Google Robot): On the Google Robot, which contributes the largest single dataset, RT-1-X matches but does not exceed the specialist RT-1. This is expected: when you already have abundant data for your target domain, adding out-of- domain data provides diminishing returns.

RT-2-X on emergent skills: The most striking result is that RT-2-X, the 55B parameter VLA, achieves 3× higher success on skills that are present in the Bridge robot dataset but absent from the Google Robot dataset. These are "emergent" skills that the evaluation robot never practiced — yet the model performs them because it learned them from other embodiments. The 55B model shows significantly higher emergent skill transfer than the 5B variant, confirming that model capacity matters.

Open in Lab
Compare RT-1-X against baselines across different evaluation domains, showing positive transfer in small-data regimes.
The demo wakes as you arrive…

What matters: ablation studies

The paper conducts extensive ablations on the RT-2-X model to understand which design decisions matter most:

Model capacity: The 55B parameter model substantially outperforms the 5B variant, especially on emergent skills. Larger models extract more transferable knowledge from the same data.

Web pretraining is essential: A 55B model trained from scratch without web data performs dramatically worse. The semantic knowledge from internet pretraining (understanding objects, spatial relations, language) is critical for cross-embodiment transfer.

Image history matters, but briefly: Using a short history of images (1–3 frames) helps the model infer motion and dynamics. Longer histories provide diminishing returns.

Data composition: Including all 9 embodiments works better than any subset. Excluding the largest single dataset (Google Robot) or the most diverse subset (Bridge) both hurt performance, suggesting that both volume and diversity contribute.

Open problem — sensors and grippers: Transfer is weakest when source and target robots differ greatly in sensing modality (e.g., adding depth cameras) or actuation type (e.g., suction vs parallel jaw grippers). The coarse action alignment strategy has limits.

Lessons and open problems

Open X-Embodiment taught the robotics community three important lessons. First, cross-embodiment transfer works — training on heterogeneous robots is not a recipe for confusion but for . Second, data diversity beats data volume for any single robot: a small dataset gains more from cross-embodiment training than a large one. Third, model scale unlocks emergent capabilities: the 55B RT-2-X demonstrates spatial reasoning and object relationships that simply don't appear at smaller scales.

Yet several open problems remain. The coarse alignment strategy struggles when robots differ too much in their physical structure (e.g., quadrupeds vs arms). Handling diverse camera viewpoints, gripper types, and action frequencies in a more principled way is an active area of research. And while the paper proves positive transfer, the dataset is still far smaller than the trillions of tokens used to train language models — the field needs orders of magnitude more data to approach true foundation models for robotics.

The open-source release of both the dataset and the data pipeline sparked immediate follow-up work. Projects like Octo (a generalist robot policy), DROID (a larger in-the-wild dataset), and π₀ (a flow-matching based generalist) all build directly on the Open X-Embodiment ecosystem.

  1. 2022

    RT-1 — Robotics Transformer

    Google introduced RT-1, a 35M-parameter Transformer trained on 130k demonstrations from a single robot type, showing that Transformers can learn effective robotic policies.

  2. 2023

    RT-2 — Vision-Language-Action model

    RT-2 co-fine-tuned a large VLM for robotic control, showing that internet-scale knowledge transfers to physical manipulation.

  3. 2023

    Open X-Embodiment (this paper)

    21 institutions pooled 60 datasets into the largest open-source robotics dataset. RT-1-X and RT-2-X demonstrated positive cross-embodiment transfer for the first time at scale.

  4. 2024

    Octo — open-source generalist policy

    Built on the OXE dataset, Octo became the first open-source generalist robot policy designed for efficient fine-tuning to new robots.

  5. 2024

    π₀ — flow-matching generalist

    Physical Intelligence's π₀ used flow matching and large-scale cross-embodiment data to create a dexterous generalist robot policy with emergent capabilities.

CitationOpen X-Embodiment Collaboration. Open X-Embodiment: Robotic Learning Datasets and RT-X Models. ICRA (Best Paper Award), 2024.

Terms in this paper