Reinforcement Learning2022advanced13 min read

A Generalist Agent

وكيل واحد لكل المهام

Reed, S. · Żołna, K. · Parisotto, E. · Gómez Colmenarejo, S. · Novikov, A. · Barth-Maron, G. · Giménez, M. · Sulsky, Y. · Kay, J. · Springenberg, J. T. · Eccles, T. · Bruce, J. · Razavi, A. · Edwards, A. · Heess, N. · Chen, Y. · Hadsell, R. · Vinyals, O. · Bordbar, M. · de Freitas, N. — TMLR

The problem

By 2022, AI had produced remarkable specialists: language models that write fluently, vision models that classify accurately, and RL agents that master individual games. But each specialist lives in its own silo — a separate architecture, a separate pipeline, and a separate set of weights for every task. If you want a robot that also chats and captions images, you need three different systems stitched together. There was no single neural network that could act across text, vision, and physical control with the same weights.

The contribution

Gato: a 1.2B-parameter decoder-only that treats every task as sequence modeling. All data — text, images, button presses, joint torques — is serialized into a flat sequence. The model is trained with a masked loss that predicts only actions and text tokens. A single set of weights handles 604 distinct tasks: playing 51 Atari games, captioning images, chatting, stacking blocks with a real Sawyer robot arm, navigating 3D environments, and following language instructions. Gato achieves above-average human scores on 23 Atari games and competitive robotic stacking performance, and can adapt to new tasks with as few as 10 episodes.

The impact

Gato was the first demonstration that a single transformer with one set of weights can operate across hundreds of tasks spanning text, vision, and physical control. It validated the "one model for everything" hypothesis at scale, directly inspiring multi-embodiment policies like RT-1 and RT-2. It showed that the same governing language models extend to multi-modal agents, and reignited debate about the path toward general-purpose AI systems.

Today's AI specialists are like a hospital where every doctor can only treat one organ. The cardiologist is world-class but cannot set a broken bone; the orthopedist is brilliant but cannot read an EKG. If a patient needs both, someone must carry their chart between departments, translate terminology, and hope nothing is lost in transit.

Gato is the general practitioner who trained in every department. The same doctor walks into the Atari arcade and plays Breakout, steps into a radiology room and captions an X-ray, picks up a phone and chats, then walks onto the factory floor and stacks blocks with a robot arm — all drawing on the same medical school education (one set of weights). The GP won't beat every specialist at their own game, but having one brain that understands all these domains unlocks something no specialist can: cross-domain intuition.

The core idea: everything is a sequence

The key insight behind Gato is disarmingly simple: if large language models can master hundreds of text tasks by predicting the next token, why not extend the same idea to every modality? Text, images, joystick inputs, robot joint angles — all of these can be serialized into a flat sequence of discrete tokens. Once everything is tokens, a single transformer can learn to predict the next one, regardless of whether that token represents a word, a pixel patch, a button press, or a motor command.

This is the "bitter lesson" in action: instead of designing specialized architectures with hand-crafted inductive biases for each domain, Gato bets on a single generic architecture and lets scale and data do the work.

Open in Lab
See how Gato serializes different modalities — text, images, continuous values, and discrete actions — into one unified token stream.
The demo wakes as you arrive…

Tokenization: the universal translator

The most critical engineering challenge in Gato is — converting every modality into a common vocabulary of discrete integers. Each modality has its own scheme:

Text is encoded using SentencePiece with a 32,000-subword vocabulary, producing integers in the range [0,32000)[0, 32000). This is standard language model practice.

Images are split into non-overlapping 16×1616 \times 16 patches in raster order (following ), with each pixel normalized to [−1,1][-1, 1] and divided by 16=4\sqrt{16} = 4. The patches are then embedded through a small ResNet block rather than being discretized — they remain continuous vectors until the layer.

Continuous values (like joint torques or proprioceptive readings) are mu-law encoded to [−1,1][-1, 1], then discretized into 1024 uniform bins. These tokens are shifted to the range [32000,33024)[32000, 33024) so they never collide with text tokens.

Discrete values (like Atari button presses) are flattened to integers in [0,1024)[0, 1024).

After tokenization, everything is concatenated in a canonical order: for each timestep, observations come first (text → images → tensors), then a separator token, then actions. Episodes are simply timesteps in chronological order.

F(x)=sgn(x)log⁡(∣x∣μ+1)log⁡(Mμ+1)F(x) = \text{sgn}(x) \frac{\log(\lvert x \rvert \mu + 1)}{\log(M \mu + 1)}
Mu-law companding (μ = 100, M = 256) — Maps continuous values to [−1,1][-1, 1] with logarithmic compression. The sign function preserves polarity. The parameter μ\mu controls how aggressively values near zero are expanded. After companding, values are uniformly binned into 1024 discrete tokens.

Architecture: a decoder-only transformer

Gato's architecture is deliberately simple — a standard decoder-only transformer, much like GPT. The model has 24 layers, 16 heads, an embedding dimension of 2048, and a feedforward hidden size of 8192, totaling 1.18 billion parameters. It uses GEGLU activations, pre-norm layer normalization, and a context window of 1024 tokens.

The interesting part is not the transformer itself but the embedding function that feeds it. Different token types are embedded differently:

Text, discrete, and continuous tokens are embedded through a standard learned lookup table, plus a learnable position encoding based on the token's position within its timestep.

Image patch tokens pass through a single ResNet v2 block (with GroupNorm and GELU) to produce a vector per patch. They also receive a 2D patch position encoding — separate learnable row and column embeddings based on the patch's location in the original image.

The model was deliberately sized at 1.2B parameters — not because bigger wouldn't be better, but because this was the largest model that could run in real-time on the NVIDIA RTX 3090 GPUs inside DeepMind's robot arms. When hardware improves, the model can scale up.

Open in Lab
Explore Gato's architecture: click on each component to see how tokens flow from raw inputs through embedding to the transformer and out to actions.
The demo wakes as you arrive…

Training: masked next-token prediction

Gato is trained as a standard — predict the next token given all previous tokens. But with a crucial twist: the loss is masked. The model only incurs loss on text tokens and action tokens. Observations (images, proprioceptive readings) are input-only — the model processes them but is never asked to predict them.

This makes intuitive sense: predicting the next frame of a video game or the next sensor reading from a robot would be enormously expensive, and the model's job is to act, not to hallucinate sensor data.

L(θ,B)=−∑b=1∣B∣∑l=1Lm(b,l)log⁡pθ ⁣(sl(b)∣s1(b),…,sl−1(b))\mathcal{L}(\theta, B) = -\sum_{b=1}^{|B|} \sum_{l=1}^{L} m(b,l) \log p_\theta\!\left(s_l^{(b)} \mid s_1^{(b)}, \ldots, s_{l-1}^{(b)}\right)
Masked training loss — The masking function m(b,l)=1m(b, l) = 1 only when token ll in sequence bb is a text token or an action token, and 00 otherwise. This means the model receives gradient signal only for the outputs it needs to generate — words and actions — while using observations purely as context. BB is the training batch, LL is the sequence length (1024).

Training ran on a 16×16 TPU v3 slice for 1 million steps with 512 and sequence length 1024, taking approximately 4 days. Each batch mixes subsequences approximately uniformly across domains (Atari, text, robotics, etc.), with some manual upweighting of larger and higher-quality datasets.

An important data curation step: for control tasks, Gato trains on a filtered set of episodes with returns at least 80% of the expert return. This means the model mostly sees good behavior — a form of from near-expert demonstrations.

For conditioning: during training, 25% of sequences are prepended with a prompt from the same task. Half of these prompts come from the end of an episode (acting as goal conditioning), and the other half are sampled uniformly. This teaches the model to infer the task from context.

Open in Lab
See which tokens receive gradient during training: only text and action tokens are targets. Observations are input-only context.
The demo wakes as you arrive…

Deployment: from tokens back to actions

At deployment, Gato works as a standard autoregressive policy. A prompt (typically a successful demonstration) is tokenized and forms the initial sequence. The environment then yields an observation, which is tokenized and appended. Gato samples the action vector one token at a time. Once all action tokens are generated, the action is decoded by inverting the tokenization (un-binning continuous values, un-shifting discrete ones) and sent to the environment. The environment steps, yields a new observation, and the process repeats.

The model always sees all previous observations and actions within its 1024-token context window. For robotics tasks, Transformer-XL memory was used at deployment (though not during training) to extend effective context.

A practical challenge: the 1.2B model barely fits the 20 Hz control rate of the Sawyer robot, overrunning the 50 ms timestep by about 10 ms. This is the fundamental tension in Gato's design — real-time control demands small models, but capability demands large ones.

604 tasks, one set of weights

Gato's training data spans an extraordinary range. Simulated control alone covers 596 tasks across 12 domains: 51 Atari games, 254 DM Lab levels, 45 Meta-World manipulation tasks, 30 DM Control Suite tasks (from both state and pixels), 46 BabyAI language-instruction levels, 16 Procgen games, 38 Modular RL locomotion variants, and more. On top of that, Gato ingests vision-language datasets: MassiveText for pure text, ALIGN and LTIP for image-caption pairs, MS-COCO and Conceptual Captions for captioning, and OKVQA/VQAv2 for visual question answering.

The key result: with a single set of weights, Gato performs above 50% of expert score on 450 out of 604 tasks. It achieves above-average human performance on 23 Atari games and scores over twice the human score on 11 of them. On the robotic RGB Stacking benchmark, it matches the baseline behavior cloning that was trained on that single task alone.

Open in Lab
Explore Gato's 604 tasks across different domains. See how performance compares to expert-level and human baselines.
The demo wakes as you arrive…

Scaling laws: bigger models, better everything

The authors tested three model sizes — 79M, 364M, and 1.18B parameters — and found consistent scaling improvements. For the same number of training tokens, larger models perform significantly better across all domains. This mirrors the scaling laws observed in language models and suggests that Gato's multi-modal, multi-task performance would continue to improve with even larger models.

This is a critical result because it means the architecture is not the bottleneck — data and compute are. As hardware improves and training data grows, the same Gato recipe should yield increasingly capable generalist agents.

Open in Lab
Compare performance across 79M, 364M, and 1.18B parameter models as training progresses.
The demo wakes as you arrive…

Few-shot adaptation: learning new tricks with little data

Can Gato learn entirely new tasks it has never seen? The authors tested this by holding out four tasks from pretraining — cartpole swingup, Meta-World assembly, DM Lab apple foraging, and Atari Boxing — then fine-tuning on tiny datasets of 1 to 1000 episodes.

The results reveal a clear pattern: pretraining on diverse data helps. Gato pretrained on all data consistently outperforms models pretrained only on same-domain data, which in turn outperform models trained from scratch. The one exception is Atari Boxing, where the visual style is so different from other training domains that transfer is difficult.

For robotics, the results are even more dramatic. On the RGB Stacking benchmark, Gato recovers expert performance with just 10 fine-tuning episodes and peaks at 100–1000 episodes. A novel "stack blue on green" task (completely outside training distribution) achieved 60% success rate after fine-tuning, versus 0.5% for a baseline trained from scratch.

What Gato learns inside: embedding analysis

To understand how Gato organizes its internal representations, the authors extracted embeddings from the middle of the network (layer 12 of 24) and visualized them using PCA → .

The result is revealing: embeddings from the same task cluster tightly together, and clusters from the same domain sit near each other. Text and web data overlap (both are language), while Atari games form their own island. Control tasks cluster by their input modality (state vs. pixels). Even a held-out task (cartpole swingup) lands next to its domain neighbors despite never appearing in training.

This means Gato has learned a task-aware internal geometry — different tasks live in different regions of representation space, and the model navigates between them based on context. It is not simply memorizing token patterns; it has developed an internal map of tasks and domains.

Limitations and open questions

Gato is a proof of concept, not a finished product, and it exposes several fundamental limitations:

Context length bottleneck. The 1024-token context window means that for image-heavy environments, the model can only attend to a few timesteps. This severely limits in-context learning — the model simply cannot see enough of a demonstration to learn from it.

Specialist gap. Gato underperforms specialist agents trained on individual tasks. The 1.2B parameter budget is spread across 604 tasks, so each task gets far less effective capacity than a specialist would.

Data collection. Control data requires expensive RL training or human demonstrations. Unlike language data (which can be scraped from the web), there is no "Internet of robot experience" to draw from.

Dialogue quality. Gato's text outputs are often superficial or factually incorrect, far below the quality of dedicated language models like Chinchilla or PaLM at the time.

The path Gato opened

  1. 2021

    Decision Transformer

    Showed that offline RL can be framed as sequence modeling — predict actions conditioned on desired return. Proved transformers work for control, but only single-task.

  2. 2022

    Gato — A Generalist Agent

    Extended Decision Transformer to 604 tasks across text, vision, and control. One model, one set of weights, real robot deployment. Proved multi-modal generalism is feasible.

  3. 2022

    RT-1 — Robotics Transformer

    Built on Gato's insight that robot control is sequence modeling. Focused specifically on real-world manipulation with 130k demonstrations across 700+ tasks.

  4. 2023

    RT-2 — Vision-Language-Action Model

    Merged a vision-language model with robotic control: the same model that understands images and text can also output robot actions. Gato's generalist vision realized with better language understanding.

Gato's deepest legacy is not its performance numbers — specialists still win on individual tasks. Its legacy is the proof of possibility: that a single set of weights can encode knowledge spanning text, vision, and physical control. Before Gato, "one model for everything" was a philosophical position. After Gato, it was an engineering roadmap. The path from Gato to RT-1 to RT-2 shows a clear trajectory: take the generalist hypothesis seriously, add more data and scale, and the specialists start to become unnecessary.

CitationReed, Żołna, Parisotto, Gómez Colmenarejo, Novikov, Barth-Maron, Giménez, Sulsky, Kay, Springenberg, Eccles, Bruce, Razavi, Edwards, Heess, Chen, Hadsell, Vinyals, Bordbar, de Freitas. A Generalist Agent. TMLR, 2022.

Terms in this paper