Robotics2022intermediate11 min read
RT-1: Robotics Transformer for Real-World Control at Scale
RT-1: محوِّل الروبوتات للتحكم الواقعي على نطاق واسع
Brohan, A. · Brown, N. · Carbajal, J. · Chebotar, Y. · Dabis, J. · Finn, C. · Gopalakrishnan, K. · Hausman, K. · Herzog, A. · Hsu, J. · Ibarz, J. · Ichter, B. · Irpan, A. · Jackson, T. · Jesmonth, S. · Joshi, N. J. · Julian, R. · Kalashnikov, D. · Kuang, Y. · Leal, I. · Lee, K.-H. · Levine, S. · Lu, Y. · Malla, U. · Manjunath, D. · Mordatch, I. · Nachum, O. · Parada, C. · Peralta, J. · Perez, E. · Pertsch, K. · Quiambao, J. · Rao, K. · Ryoo, M. · Salazar, G. · Sanketi, P. · Sayed, K. · Singh, J. · Sontakke, S. · Stone, A. · Tan, C. · Tran, H. · Vanhoucke, V. · Vega, S. · Vuong, Q. · Xia, F. · Xiao, T. · Xu, P. · Xu, S. · Yu, T. · Zitkovich, B. — RSS
The problem
By 2022, large models had transformed NLP and computer vision through scale: train once on diverse data, then deploy everywhere. But robotics was stuck in the old paradigm — one small model per task, trained from scratch on narrow, expensive-to-collect datasets. Existing multi-task robot policies like Gato could only handle a handful of real-world tasks, and instruction-following methods generalized poorly to new tasks, objects, or environments. The core question was: can we build a single, high-capacity robot policy that absorbs diverse real-world experience and generalizes the way language models do?
The contribution
RT-1: a -based robot policy that takes camera images and natural-language instructions as input and outputs discretized motor actions at 3 Hz. Its architecture — a FiLM-conditioned EfficientNet for visual-language fusion, TokenLearner for compression, and a Transformer — is compact (35M parameters) yet high-capacity. Trained via on 130k real-robot episodes spanning 700+ tasks collected over 17 months with 13 robots, RT-1 achieves 97% success on seen tasks and generalizes to new tasks (+25%), distractors (+36%), and backgrounds (+18%) better than all baselines. It integrates with SayCan for long-horizon task execution up to 50 stages.
The impact
RT-1 proved that the foundation-model recipe — large diverse data + high-capacity architecture + simple objective — works for real-world robotics, not just language and vision. It directly led to RT-2 (a that transfers web knowledge to robotic control) and the Open X-Embodiment project, which pooled data from 22 robots across 21 institutions. RT-1 established that robotic intelligence can scale with data, marking the birth of robotic foundation models.
Imagine a new chef who can only cook one dish at a time — you have to teach them each recipe from scratch, and they forget the previous one. Now imagine a chef who has watched thousands of cooking videos of hundreds of dishes: they can read a recipe card ("make pasta carbonara"), look at the ingredients in front of them, and figure out the right hand movements — even for a dish they've never made before.
RT-1 is that second chef for robots. Instead of training one robot brain per task, it trains a single brain on 700+ tasks from 130,000 real demonstrations. You hand it a text command and a camera feed, and it figures out the arm and base movements — even for tasks it has never practiced.
The problem: one model per task does not scale
In NLP and computer vision, the recipe for success became clear: collect a massive, diverse dataset, train a high-capacity model once, and then deploy or fine-tune it for many tasks. GPT absorbs text from the entire internet; CLIP learns from 400 million image-text pairs. These models generalize because they've seen enough variety.
Robotics had not yet made this leap. Each task — pick up a sponge, open a drawer, push a bottle — required its own dataset of demonstrations and its own trained policy. Gato (DeepMind, 2022) attempted a generalist , but its real-world robotics capability was limited to a single block-stacking task. Instruction-following methods existed, but their to new objects or environments was weak.
The two key bottlenecks were data and architecture. Robot data is expensive — it requires physical hardware, human operators, and real-world time. And robot policies must run in real time on the robot itself, which limits model size. RT-1 attacks both problems simultaneously.
RT-1 architecture: from pixels and words to motor commands
RT-1's architecture is a pipeline of three stages, each solving a different part of the problem. Think of it as an information highway where the camera image enters as a wide river of pixels, the natural language instruction acts as a traffic sign telling the river where to go, and the output is a compact set of motor commands.
Stage 1 — Visual-Language Fusion (FiLM-EfficientNet): The robot's camera image (300×300 pixels) enters an EfficientNet-B3 backbone, pre-trained on ImageNet. But how does the network know which task to perform? FiLM (Feature-wise Linear Modulation) conditioning injects the language instruction — encoded by a pre-trained Universal Sentence Encoder — directly into the convolutional layers. FiLM modulates each with a task-specific scale and shift, essentially telling the visual network "pay to the green chip bag, not the red can."
Stage 2 — Token Compression (TokenLearner): After FiLM-EfficientNet, each image produces a 9×9 spatial grid of 512-dimensional features — 81 tokens total. For a 6-frame image history, that's 486 tokens — far too many for real-time Transformer . TokenLearner learns to compress these 81 tokens down to just 8 per image by learning spatial attention masks that select the most informative regions. This reduces the Transformer's input from 486 tokens to just 48, making 3 Hz inference feasible on the robot.
Stage 3 — Temporal Reasoning (Transformer): The 48 tokens enter a decoder-only Transformer with 8 layers. The Transformer attends across the 6-frame history, learning temporal patterns — like "the arm is moving toward the object" or "the gripper just closed." Its output is a sequence of action tokens.
Action tokenization: turning continuous motion into tokens
A robot arm moves in continuous space — joint angles are real numbers, gripper openings are fractions. But Transformers excel at discrete sequences: picking one token from a vocabulary, like choosing a word from a dictionary. RT-1 bridges this gap through .
Each action has 11 dimensions: 7 for the arm (x, y, z position, roll, pitch, yaw rotation, and gripper opening), 3 for the mobile base (x, y translation and yaw rotation), and 1 mode-switching command (to terminate the ). Each continuous dimension is discretized into 256 bins, uniformly spanning the range of possible values. So "move the arm 3 cm to the right" becomes token 147 in the x-dimension bin.
This tokenization converts the control problem into a problem: at each timestep, the Transformer predicts 11 categorical distributions (one per dimension), each over 256 classes. This is trained with standard cross-entropy loss — the same objective used to train language models.
Data at scale: 130k episodes, 700+ tasks, 13 robots
The dataset is arguably RT-1's most important contribution. Over 17 months, a fleet of 13 Everyday Robots mobile manipulators collected roughly 130,000 demonstration episodes across more than 700 task instructions. These tasks fall into categories such as picking, placing, opening and closing drawers, rotating objects, and pulling napkins. The instructions are natural-language commands like "pick rxbar chocolate from middle drawer" or "place green jalapeno chip bag to paper bowl."
The paper demonstrates that both data quantity and data diversity matter for generalization. Training on 10× more data improves performance, but adding diverse task types matters even more. A model trained on many simple picking tasks generalizes to novel objects better than one trained on fewer but more complex tasks. The key insight is that structural overlap between tasks — grasping is common to picking, placing, and drawer-opening — enables the model to compose learned skills in new ways.
Generalization: new tasks, new objects, new environments
RT-1 was evaluated on over 3,000 real-world robot trials — not in simulation. The results demonstrate three forms of generalization. First, new task generalization: RT-1 achieves 76% success on completely unseen task instructions, compared to 53% for Gato and 24% for BC-Z (the next-best baseline). Second, robustness to distractors: when novel objects are placed on the table as distractions, RT-1 maintains 83% success vs. 47% for Gato. Third, background generalization: when evaluated in unseen kitchens with different lighting and layouts, RT-1 reaches 59% success vs. 41% for the nearest baseline.
The paper also shows that RT-1 integrates seamlessly with SayCan, a framework where a large language model plans a sequence of high-level steps (e.g., "I spilled my drink" → find sponge → pick up sponge → bring to table → wipe), and RT-1 executes each step. With RT-1 as the low-level controller, SayCan completes long-horizon tasks with up to 50 sequential stages — a dramatic increase over prior work.
Absorbing heterogeneous data: simulation and other robots
Can RT-1 benefit from data that doesn't come from its own robot? The paper tests two scenarios. First, mixing in simulated episodes from a different visual domain. Second, adding real demonstrations from a Kuka IIWA arm — a completely different robot morphology. In both cases, the additional data does not hurt performance on the original tasks and improves generalization to new scenarios. This is a crucial result: it means the path to better robot intelligence is not just collecting more data on one robot, but pooling experience across robots, simulators, and even embodiments — a vision that culminated in the Open X-Embodiment project.
Design choices that matter
The paper ablates several architectural decisions. TokenLearner is essential — without it, the model is too slow for real-time control or must be made so small it can't learn the task diversity. With 8 tokens per image (down from 81), inference runs at 3 Hz while maintaining or improving performance. Image history of 6 frames outperforms single-frame input, as temporal context helps resolve ambiguities (is the arm approaching or retreating?). Action into 256 bins outperforms both continuous regression and coarser discretizations. And FiLM conditioning is superior to late fusion of language and vision.
Simplified to show the idea — not the real implementation.
import numpy as np
NUM_BINS = 256
# Each action dimension has a known [min, max] range
ACTION_RANGES = {
"x": (-0.03, 0.03), # meters
"y": (-0.03, 0.03),
"z": (-0.03, 0.03),
"roll": (-0.25, 0.25), # radians
"pitch": (-0.25, 0.25),
"yaw": (-0.25, 0.25),
"gripper": (0.0, 1.0), # open fraction
}
def tokenize(value: float, lo: float, hi: float) -> int:
"""Map a continuous value to a discrete bin index."""
normalized = (value - lo) / (hi - lo) # → [0, 1]
clipped = np.clip(normalized, 0.0, 1.0)
return int(clipped * (NUM_BINS - 1)) # → {0 .. 255}
def detokenize(token: int, lo: float, hi: float) -> float:
"""Map a bin index back to a continuous value."""
return lo + (token / (NUM_BINS - 1)) * (hi - lo)
# Example: tokenize "move 1.5 cm right in x"
x_value = 0.015
x_lo, x_hi = ACTION_RANGES["x"]
token = tokenize(x_value, x_lo, x_hi) # → 191
reconstructed = detokenize(token, x_lo, x_hi) # → 0.01494…
What RT-1 set in motion
2018
QT-Opt — grasping at scale with RL
Google's Q-learning based grasping policy trained on 580k real grasps. Showed that robot learning benefits from large-scale data, but was limited to a single task.
2022
SayCan — language models ground in robot affordances
Combined LLM planning with robot affordances. The LLM proposes high-level plans and a value function scores which actions the robot can actually do. Required a capable low-level policy — RT-1 became that policy.
2022
Gato — generalist agent
DeepMind's multi-modal generalist that plays games, chats, and does simple robotics. Proved that one model can do many things, but real-world robot capabilities were narrow (one block-stacking task).
2022
RT-1 — this paper
First robot policy to absorb 700+ real-world tasks into one Transformer model. Proved that scale + diversity in robot data enables generalization, not just performance.
2023
RT-2 — vision-language-action model
Built on RT-1's data and recipe. Pre-trained a vision-language model (PaLI-X) on web data, then fine-tuned it to output robot actions. Web knowledge transferred to robotics — e.g., understanding "move the coke can to Taylor Swift" without training on celebrity recognition.
2023
Open X-Embodiment — pooling robot data across institutions
22 robot embodiments, 21 institutions, 1 million+ episodes. Trained RT-X models on the combined dataset. Showed that cross-embodiment transfer works: training on data from many robots improves each robot's individual performance.
RT-1's lasting contribution is not any single architectural trick — it is the demonstration that the foundation-model recipe works for real-world robotics. Large diverse data + high-capacity model + simple training objective = a robot that generalizes. This insight now drives the entire field: from RT-2's web-knowledge transfer to Open X-Embodiment's cross-institution data pooling to the latest vision-language-action models. Every time a research lab trains a general-purpose robot policy on diverse demonstrations, the lineage traces back to RT-1.
CitationBrohan, Brown, Carbajal, Chebotar, et al.. RT-1: Robotics Transformer for Real-World Control at Scale. RSS, 2023.
Terms in this paper
- Action Tokenizationتكميم الأفعال
- Behavioral Cloningالاستنساخ السلوكي
- Imitation Learningالتعلّم بالتقليد
- Transformerالمحوِّل
- Vision Encoderمرمِّز بصري
- Tokenوحدة لغوية (رمز)
- Generalizationالتعميم
- Zero-Shotالنمط الصفري
- Sim-to-Realالنقل من المحاكاة إلى الواقع
- Manipulationالتلاعب
- Foundation Modelنموذج أساسي
- Visuomotor Policyالسياسة البصرية-الحركية
- Embodied AIالذكاء المُجسَّد
- Episodeجولة تفاعلية كاملة
- Backboneالبنية الأساسية