Robotics2023advanced11 min read
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
سياسة الانتشار: تعلُّم السياسات البصرية-الحركية بانتشار الأفعال
Chi, C. · Xu, Z. · Feng, S. · Cousineau, E. · Du, Y. · Burchfiel, B. · Tedrake, R. · Song, S. — RSS
The problem
Teaching robots to perform tasks from human demonstrations is typically framed as : map observations to actions. But robot action distributions are often — a block can be pushed left or right with equal validity — and existing methods like mixture-of-Gaussians or energy-based models either fail to capture all modes faithfully, suffer from instability, or cannot scale to high-dimensional action sequences. There was no policy representation that combined expressiveness, stability, and the ability to predict coherent action sequences in high dimensions.
The contribution
Diffusion Policy: a conditional diffusion process that generates robot action sequences instead of images. Starting from Gaussian noise, the policy iteratively refines a full action conditioned on visual observations. Key technical innovations include treating observations as conditioning (not part of the joint distribution), receding-horizon that executes a prefix of the predicted sequence before re-planning, and a time-series for tasks requiring high-frequency action changes. Evaluated on 15 tasks across 4 benchmarks, it outperforms prior methods by an average of 46.9%.
The impact
Diffusion Policy became the dominant paradigm for robot . It demonstrated that generative diffusion models — originally designed for images — are a natural fit for action generation, combining expressiveness, stability, and scalability. The approach directly inspired π₀ (Physical Intelligence), ALOHA-ACT, and a wave of diffusion-based robot policies. Within two years, nearly every major robotics lab adopted diffusion-based action generation, making this paper a foundational reference for the post-behavioral-cloning era of robot learning.
Imagine a chef who plans a dish not by writing one step at a time, but by throwing random ingredients onto the counter and then, in a series of quick refinements, rearranging, removing, and adjusting until a perfect plate emerges.
Each refinement makes the dish a little less chaotic and a little more intentional. By the final pass, every element is in its place — and importantly, the same starting mess could lead to entirely different valid dishes depending on the random arrangement.
Diffusion Policy teaches robots this way: start with noise, refine toward action — and let the randomness naturally express the many valid ways to accomplish a task.
The challenge: why robot action prediction is harder than it looks
Teaching a robot to perform a task from human demonstrations seems like a straightforward supervised learning problem: observe what the human does and learn a mapping from observations to actions. In practice, three properties make this far more challenging than typical .
First, human demonstrations are multimodal. Given the same — say, a block sitting in front of a robot arm — there may be equally valid actions: push it left, push it right, or go around it. A policy that averages these modes produces an action that goes straight through the block — a move that belongs to none of the valid strategies.
Second, robot actions are sequential and correlated. A single action in isolation is meaningless; what matters is a coherent trajectory. Predicting one step at a time invites jitter — the policy may switch between valid strategies on consecutive steps, producing incoherent motion.
Third, the can be high-dimensional. A bimanual robot with grippers has 14 or more per timestep, and predicting a sequence of such actions compounds the dimensionality.
The core idea: denoising diffusion for action generation
Diffusion models were originally designed to generate images. The key insight of this paper is that the same mechanism — start with noise and iteratively denoise — works remarkably well for generating robot actions.
In image diffusion (), a neural network learns to predict the noise added to an image, then removes it step by step. In Diffusion Policy, the same process happens in action space: the model starts with a random action sequence sampled from a Gaussian and refines it over denoising steps, conditioned on the robot's current observations.
The noise-prediction network takes the current noisy action sequence , the observation , and the denoising iteration as input, and predicts the noise component. Subtracting this prediction yields a cleaner action sequence. Repeating for steps produces the final clean action trajectory .
The denoising step follows this update rule:
Training: a simple MSE loss
Training Diffusion Policy is elegant in its simplicity. From a dataset of demonstration trajectories, the process samples a clean action sequence , adds noise at a randomly chosen level to produce a corrupted version, and asks the network to predict the noise that was added. The is a standard :
Simplified to show the idea — not the real implementation.
# Sample a batch of demonstration trajectories obs, actions = dataset.sample_batch()
# Pick random noise levels for each sample k = torch.randint(0, K, (batch_size,))
# Sample Gaussian noise matching the action shape noise = torch.randn_like(actions)
# Corrupt the clean actions with noise at level k noisy_actions = noise_schedule.add_noise(actions, noise, k)
# Predict the noise from corrupted actions + observations predicted_noise = noise_net(obs, noisy_actions, k)
# Simple MSE loss — that's the entire objective loss = F.mse_loss(predicted_noise, noise) loss.backward() optimizer.step()Key design decisions
Three design choices unlock the full potential of diffusion for robot control: the visual conditioning strategy, the receding-horizon execution scheme, and the choice between and denoising backbones.
Visual conditioning. A naive approach would model the joint distribution of observations and actions , denoising both together. Diffusion Policy instead conditions on observations: it models . This means the visual encoder runs once per , not once per denoising step — a dramatic speedup that makes real-time control feasible. The visual features are injected into the denoising network through (for the CNN ) or (for the Transformer backbone).
CNN vs Transformer backbone. The CNN backbone (a 1D temporal convolution network) works well out of the box for most tasks. Its inductive bias toward smooth signals, however, can over-smooth actions that need to change rapidly. The Transformer backbone handles sharp action changes better but requires more tuning. The paper recommends starting with the CNN and switching to the Transformer if performance is low on tasks with high-frequency action variation.
Visual encoder. A ResNet-18 with spatial pooling (to preserve spatial information) and GroupNorm instead of BatchNorm (for compatibility with the used in DDPM training) encodes each camera view independently. The encoder is trained end-to-end with the diffusion policy.
Receding horizon: balancing planning and responsiveness
One of the most impactful design choices is how the predicted action sequence is executed. The policy predicts future action steps but only executes the first of them before re-observing and re-planning. This is the receding-horizon control scheme, borrowed from model predictive control.
Think of it like driving: you look far down the road to plan your steering (prediction horizon ), but you only commit to the next few seconds of actual steering wheel movement ( ) before checking the road again. This achieves two things: the long prediction horizon ensures and smooth trajectories, while the short execution horizon keeps the policy responsive to changes.
The observation horizon captures how many past frames the policy sees. Having (the current and previous frame) often suffices, giving the policy implicit velocity information without computational overhead.
Intriguing properties
Diffusion Policy inherits several powerful properties from the diffusion framework, each addressing a known weakness in prior robot learning methods.
Multimodal action distributions. Multimodality arises naturally from two sources: the random initial noise that determines which mode the trajectory converges to, and the stochastic noise added at each denoising step that allows exploration between modes. Unlike mixture-of-Gaussians policies that must specify the number of modes in advance, diffusion policies express arbitrary distributions — including ones with modes that appear and disappear depending on the observation.
Synergy with . Most behavior cloning methods use velocity control because position control amplifies multimodality problems (multiple valid positions exist at each step). Since Diffusion Policy handles multimodality well, it can exploit position control's advantage — less — and consistently outperforms the same architecture with velocity control.
. Energy-based policies (IBC) use InfoNCE loss with negative sampling to estimate the normalization constant. This estimation is noisy, causing training spikes and making selection difficult. Diffusion Policy sidesteps this entirely by learning the score function (gradient of log-probability), which is independent of the normalization constant. The result: smooth training curves and reliable checkpoint selection.
Results: consistent gains across 15 tasks
Diffusion Policy was evaluated across 15 tasks from 4 benchmarks: RoboMimic (Lift, Can, Square, Transport, ToolHang), Push-T, Block Push, and Franka Kitchen. These tasks span single-arm and bimanual robots, 2-DoF to 14-DoF action spaces, rigid and fluid objects, and data from single or multiple human demonstrators.
The results were striking. Diffusion Policy outperformed prior state-of-the-art methods — LSTM-GMM (BC-RNN), IBC, and BET — on every task variant, with an average success-rate improvement of 46.9%. Particularly large gains appeared on tasks requiring precision (ToolHang: from 67% to 100%), multi-stage reasoning (Kitchen p4: from 44% to 99%), and multimodal behavior (Block Push p2: from 71% to 94%).
On real-world tasks — Push-T with a UR5 arm, sauce pouring and spreading with a Franka, mug flipping, and three bimanual tasks (egg beater, mat unrolling, shirt folding) — Diffusion Policy achieved robust performance with as few as 90-250 demonstrations, running at 10 Hz inference using DDIM with 10 denoising steps.
Connections to control theory
The paper provides a reassuring sanity check: in the simplest possible case — a linear dynamical system controlled by a linear feedback policy — Diffusion Policy converges to the correct controller. The optimal denoiser in this case becomes , and DDIM sampling converges to . For trajectory prediction (), the denoiser implicitly learns the dynamics model to predict future actions as .
This means that to clone behavior that depends on state, the diffusion policy must implicitly learn a task-relevant dynamics model — an emergent property that helps explain the method's strong performance on contact-rich tasks where dynamics matter.
Real-time inference with DDIM
A critical requirement for robot control is speed. Full DDPM sampling with 100 denoising steps is too slow for closed-loop control. Diffusion Policy uses DDIM (Denoising Diffusion Implicit Models), which decouples the number of training and inference steps. Training with 100 noise levels but inferring with only 10 steps achieves 0.1 second inference latency on an Nvidia 3080 GPU — fast enough for 10 Hz real-time control.
The paper also uses a square cosine (from iDDPM), which better captures both high-frequency and low-frequency characteristics of action signals compared to the original linear schedule.
Timeline: from image diffusion to robot diffusion
2020
DDPM (Ho et al.)
Denoising Diffusion Probabilistic Models showed that iterative denoising can generate high-quality images, rivaling GANs with stable training.
2021
IBC — Implicit Behavioral Cloning
Florence et al. used energy-based models for robot policies, showing expressiveness for multimodal actions but suffering from training instability.
2022
Diffuser (Janner et al.)
Applied diffusion to trajectory planning, denoising joint state-action sequences. Worked for offline RL but modeled the joint distribution, requiring denoising observations too.
2023
Diffusion Policy (this paper)
Represented visuomotor policies as conditional diffusion on actions only, with receding-horizon control. Outperformed all baselines by 46.9% average on 15 tasks. Opened the door for diffusion-based robot learning.
2024
π₀ and the diffusion policy wave
Physical Intelligence's π₀ and many other systems adopted diffusion-based action generation as the default paradigm for robot imitation learning, scaling to foundation models for robotics.
CitationChi, Xu, Feng, Cousineau, Du, Burchfiel, Tedrake, Song. Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. RSS, 2023.
Terms in this paper
- Diffusion Modelنموذج الانتشار
- DDPMنموذج الانتشار الاحتمالي لإزالة التشويش
- Visuomotor Policyالسياسة البصرية-الحركية
- Behavioral Cloningالاستنساخ السلوكي
- Imitation Learningالتعلّم بالتقليد
- Action Spaceفضاء الأفعال
- Denoisingإزالة الضوضاء
- Noise Scheduleجدول الضوضاء
- Score Functionدالة الرصيد
- Receding Horizonالأفق المتراجع
- Trajectoryمسار تتابع الحالات والأفعال
- Multimodalمتعدد الوسائط
- End-Effectorالذراع الطرفي
- Manipulationالتلاعب
- Convolutional Neural Network (CNN)الشبكة العصبية الالتفافية