Reinforcement Learning2021intermediate12 min read

Decision Transformer: Reinforcement Learning via Sequence Modeling

محوِّل القرار: التعلّم المعزَّز عبر نمذجة التسلسلات

Chen, L. · Lu, K. · Rajeswaran, A. · Lee, K. · Grover, A. · Laskin, M. · Abbeel, P. · Srinivas, A. · Mordatch, I. — NeurIPS

The problem

Traditional relies on temporal difference (TD) learning to propagate reward signals backward through time. This approach suffers from the "deadly triad" — combining function approximation, , and off-policy learning leads to instability. Value overestimation compounds errors, and sparse or delayed rewards make extremely difficult. Meanwhile, models had revolutionized language and vision through , but nobody had tried to replace RL algorithms entirely with sequence prediction.

The contribution

Decision Transformer recasts offline RL as autoregressive sequence modeling. Instead of fitting value functions or computing policy gradients, it feeds trajectories of (, state, action) triples into a GPT-style causal Transformer. At test time, specifying a high desired prompts the model to generate expert-level actions. Despite its simplicity, Decision Transformer matches or exceeds state-of-the-art offline RL methods (CQL, BEAR, BRAC) on Atari, OpenAI Gym, and Key-to-Door tasks — and dramatically outperforms TD learning on long-horizon credit assignment.

The impact

Decision Transformer proved that sequence modeling can replace dedicated RL algorithms, opening a new research direction. It directly inspired Gato (DeepMind, 2022), which used the same -as-tokens idea to build a single generalist agent across hundreds of tasks. The paper also catalyzed Trajectory Transformer, Online Decision Transformer, and a wave of works unifying language modeling with decision-making — bridging the gap between LLMs and embodied AI.

Traditional RL is like teaching someone to cook by grading each knife cut and stirring motion, then slowly working backward from a delicious meal to figure out which actions mattered — a process called .

Decision Transformer throws that out. Instead, it hands the learner a recipe book of full cooking sessions — every action from start to finish, annotated with the final meal rating. Want a 5-star dish? The model flips to sessions that scored 5 stars and replays similar moves. No backward credit propagation, no value estimation — just pattern-matching on complete stories.

The problem: why TD learning struggles offline

In standard , an agent interacts with an environment: it observes a state, takes an action, receives a reward, and transitions to a new state. The goal is to learn a policy that maximizes cumulative reward. The dominant approach — temporal difference (TD) learning — estimates a and propagates rewards backward using Bellman backups.

But offline RL removes the ability to explore. The agent has only a fixed dataset of past trajectories. TD learning in this setting faces three compounding problems. First, bootstrapping — estimating values from other estimated values — amplifies errors when the data doesn't cover the full state space. Second, the agent may overestimate the value of actions it has never seen, a phenomenon called value overestimation. Third, when rewards are sparse or delayed, Bellman backups must chain across many timesteps to propagate signal, and each link adds noise.

Methods like Conservative Q-Learning (CQL) address overestimation by penalizing Q-values for unseen actions, but they still rely on the TD learning machinery. The question this paper asks is more radical: can we bypass TD learning entirely?

Open in Lab
Compare the two paradigms — TD learning propagates rewards backward step by step, while Decision Transformer reads the full trajectory forward.
The demo wakes as you arrive…

The insight: reinforcement learning is just sequence completion

GPT predicts the next word given previous words. Decision Transformer predicts the next action given previous returns, states, and actions. The insight is deceptively simple: a trajectory through an environment is just a sequence of tokens.

But there is a crucial twist. If you condition on past rewards, the model learns to imitate whatever behavior produced those rewards — including bad behavior. Instead, Decision Transformer conditions on returns-to-go: the sum of future rewards from each timestep onward. This is the key mechanism that enables control at test time.

At test time, you set the return-to-go to a high desired value — essentially telling the model "generate a trajectory that earns this much reward." The model then produces actions consistent with trajectories in the data that achieved similar returns. This is analogous to prompting a language model with a style : instead of "write formally," you say "achieve return 1000."

Trajectory tokenization: how experiences become tokens

A standard language model processes sequences of word tokens. Decision Transformer processes sequences of experience tokens. Each timestep contributes three tokens, forming a repeating triplet:

  1. Return-to-go (R^t\hat{R}_t): the sum of all future rewards from timestep tt onward.
  2. State (sts_t): the environment observation (e.g., joint angles for a robot, or pixel frames for Atari).
  3. Action (ata_t): the action taken at that timestep.

The full trajectory becomes a long sequence: τ=(R^1,s1,a1,R^2,s2,a2,…,R^T,sT,aT)\tau = (\hat{R}_1, s_1, a_1, \hat{R}_2, s_2, a_2, \ldots, \hat{R}_T, s_T, a_T).

The model sees the last KK timesteps — a of 3K3K tokens — just as GPT uses a fixed context window of text tokens. Each modality has its own linear layer, and a shared timestep embedding (not a per-token ) is added to all three tokens from the same timestep.

Open in Lab
See how a trajectory is sliced into (return-to-go, state, action) triplets and fed as tokens to the Transformer.
The demo wakes as you arrive…
τ=(R^1,  s1,  a1,  R^2,  s2,  a2,  …,  R^T,  sT,  aT)\tau = \left(\hat{R}_1,\; s_1,\; a_1,\; \hat{R}_2,\; s_2,\; a_2,\; \ldots,\; \hat{R}_T,\; s_T,\; a_T\right)
Trajectory Representation — Each timestep contributes a triplet of return-to-go, state, and action. The model autoregressively generates the action token at each step.

Architecture: a GPT that reads experience instead of text

The architecture is deliberately minimal — a GPT model with almost no RL-specific modifications. The three token types (return-to-go, state, action) each pass through their own linear embedding layer to reach a shared embedding dimension. A learned timestep embedding is added (not per-token positional encoding, because one timestep spans three tokens). follows each embedding.

For visual environments like Atari, states pass through a convolutional encoder before the linear embedding. For continuous control tasks like MuJoCo, states are raw vectors.

The embedded tokens are fed into a standard GPT with causal masking — each token can only attend to itself and previous tokens. A linear prediction head maps the hidden state at each state token position to the predicted action. The training loss is simply mean squared error for continuous actions or cross-entropy for discrete actions.

Open in Lab
The Decision Transformer architecture — tokens flow through embeddings, causal Transformer, and action prediction head.
The demo wakes as you arrive…
Decision Transformer — simplified pseudocodepython

Simplified to show the idea — not the real implementation.

def DecisionTransformer(R, s, a, t):
    # Embed each modality separately
    pos = embed_timestep(t)        # shared per-timestep embedding
    R_emb = embed_R(R) + pos       # return-to-go embedding
    s_emb = embed_s(s) + pos       # state embedding
    a_emb = embed_a(a) + pos       # action embedding

    # Interleave as (R1, s1, a1, R2, s2, a2, ...)
    tokens = interleave(R_emb, s_emb, a_emb)

    # Causal GPT — each token sees only past tokens
    hidden = causal_transformer(tokens)

    # Predict actions from state positions
    return action_head(hidden.at_state_positions)

# Training: simple supervised loss loss = MSE(predicted_actions, true_actions)
# Evaluation: set high target return, then generate target_return = desired_performance while not done:
    action = DT(target_return, state, past_actions, timestep)[-1]
    state, reward, done = env.step(action)
    target_return -= reward   # decrement by achieved reward

Return conditioning: steering the agent with a number

The most striking property of Decision Transformer is that it can generate behaviors of different skill levels from a single model. During training, the model sees trajectories of varying quality — from random walks to expert demonstrations. Each trajectory has its own return-to-go values.

At evaluation time, the return-to-go token acts like a prompt. Set it to the maximum return in the training data and the model generates expert-level behavior. Set it to a medium value and you get cautious, moderate behavior. The paper demonstrates a strong correlation between the specified target return and the actual return achieved by the agent across all environments.

Remarkably, on some tasks like Seaquest, the model can even extrapolate beyond the training data — achieving returns higher than any trajectory it was trained on. This suggests the model learns underlying action-reward patterns, not just memorized sequences.

Open in Lab
Adjust the target return and watch how the agent's behavior changes — the same model produces different skill levels.
The demo wakes as you arrive…

Long-term credit assignment via attention

Consider the Key-to-Door environment: a grid world with three phases. In phase 1, the agent is in a room with a key. In phase 2, an empty distractor room. In phase 3, a room with a door. The agent gets a reward of 1 only if it picked up the key in phase 1 and reached the door in phase 3. The actions in phase 2 are irrelevant.

This is a nightmare for TD learning. The reward signal at the door must propagate backward through every single timestep in the distractor room, each Bellman backup adding noise. CQL achieves only 13% success even with 10,000 random trajectories.

Decision Transformer handles this naturally. Self-attention lets the model directly associate the key-pickup event with the final reward, skipping the distractor phase entirely. The paper shows that the Transformer's attention weights peak at exactly two events: picking up the key and reaching the door. With 10,000 random trajectories, Decision Transformer achieves 94.6% success.

This is credit assignment through pattern matching rather than value propagation — the Transformer notices that "key pickup early + door reached late = high return" and directly learns this association.

Open in Lab
Key-to-Door environment — see how Transformer attention skips the distractor room to link key pickup with the door reward.
The demo wakes as you arrive…

Results: matching RL without doing RL

The paper evaluates Decision Transformer across three diverse benchmarks:

Atari (discrete actions, visual inputs, 1% of DQN replay data): Decision Transformer matches CQL on 3 of 4 games and outperforms REM, QR-DQN, and behavior cloning on all 4 games.

OpenAI Gym / D4RL (continuous control — HalfCheetah, Hopper, Walker): On Medium-Expert datasets, DT achieves the highest scores in most tasks, with particularly strong results on Hopper (107.6) and Walker (108.1). On Medium-Replay datasets where data quality varies widely, DT dramatically outperforms CQL on Hopper (82.7 vs 48.6).

Key-to-Door (sparse reward, long-horizon credit assignment): DT achieves 94.6% success vs CQL's 13.3% with 10K random trajectories — the most dramatic gap, showcasing the attention-based credit assignment advantage.

Crucially, Decision Transformer is also competitive with Percentile Behavior Cloning (%BC) — a method that trains only on the best trajectories — but without needing to manually select the right data subset. DT learns which trajectories to imitate based on the return-to-go conditioning.

Robustness to sparse and delayed rewards

A critical experiment tests what happens when rewards are delayed until the end of the episode. In this "delayed return" setting, the agent receives zero reward at every timestep and only gets the total cumulative reward at the final step.

For TD learning, this is catastrophic. CQL drops from 111.0 to 9.0 on Hopper-Medium-Expert with delayed rewards — the chain of Bellman backups collapses entirely when intermediate rewards vanish.

Decision Transformer is barely affected (107.6 → 107.3 on the same task). This makes sense: the model conditions on returns-to-go, which are computed from the total trajectory reward, not step-by-step signals. Whether rewards arrive at each step or only at the end, the return-to-go calculation stays the same.

This robustness to reward sparsity is one of Decision Transformer's most practically important properties, since real-world tasks often have delayed, binary, or extremely sparse reward signals.

What this means: the convergence of language modeling and RL

Decision Transformer is part of a broader trend: the unification of sequence modeling and decision-making. The paper showed that the same GPT architecture that generates coherent essays can generate coherent action sequences — without any RL-specific training machinery.

This has profound implications. First, it means RL problems can benefit from the massive infrastructure built for language models — the optimizers, the scaling laws, the training stability techniques. Second, it suggests that future generalist agents may be trained not with specialized RL pipelines, but with the same pre-train-then-prompt paradigm used for language.

Gato (DeepMind, 2022) took this idea to its logical extreme, training a single Transformer on tokenized trajectories from 604 different tasks — Atari, robotics, captioning, dialogue — achieving reasonable performance across all of them. That result traces directly back to Decision Transformer's proof of concept.

  1. 2019

    Upside-Down RL (UDRL)

    Schmidhuber and Srivastava et al. proposed conditioning policies on desired returns using supervised learning — the conceptual precursor, but limited to single-step context and simple architectures.

  2. 2021

    Decision Transformer

    Replaced RL algorithms with GPT-style sequence modeling on full trajectories. Proved that autoregressive generation conditioned on returns-to-go can match TD learning — and excel at long-horizon tasks.

  3. 2021

    Trajectory Transformer

    Janner et al. (concurrent work) also modeled RL as sequence prediction but additionally predicted states and returns, and discretized continuous values into tokens.

  4. 2022

    Online Decision Transformer

    Zheng et al. extended DT with online fine-tuning, combining offline pretraining with online exploration using entropy regularization.

  5. 2022

    Gato — A Generalist Agent

    DeepMind scaled the trajectory-as-tokens idea to 604 tasks across multiple modalities. A single Transformer played Atari, controlled robots, captioned images, and chatted — all from tokenized sequences.

CitationChen, Lu, Rajeswaran, Lee, Grover, Laskin, Abbeel, Srinivas, Mordatch. Decision Transformer: Reinforcement Learning via Sequence Modeling. NeurIPS, 2021.

Terms in this paper