Robotics2016advanced12 min read
End-to-End Training of Deep Visuomotor Policies
التدريب الشامل لسياسات بصرية-حركية عميقة
Levine, S. · Finn, C. · Darrell, T. · Abbeel, P. — JMLR
The problem
Practical robot learning systems in 2016 relied on hand-engineered pipelines: a vision module estimated object poses, a state estimator combined sensor readings, and a low-level controller computed motor commands. Each component was designed and tuned separately. Errors in the vision system could not be corrected by the controller, and the controller could not request better visual features. a directly from pixels to torques was considered infeasible — model-free needed far too many real-world samples, and the 92,000-parameter was orders of magnitude larger than policies prior methods could handle.
The contribution
A (GPS) method that decomposes visuomotor learning into two tractable pieces: (1) -centric RL using simple linear-Gaussian controllers that have access to the full state, and (2) that distills those trajectories into a deep CNN policy that only sees camera images. A novel architecture converts convolutional feature maps into 2D coordinates — a compact, spatial representation ideal for control. through GPS lets the vision layers discover task-specific features that outperform features trained only for . The result: the first method to train deep visuomotor policies for complex with direct control, using only minutes of real-world interaction.
The impact
This paper established visuomotor learning as a viable paradigm for real-world robotics. Its spatial softmax layer became a standard component in robotic perception networks. The guided policy search framework influenced a generation of robot learning methods, and the paper's core question — "does joint training beat separate training?" — with its affirmative answer, motivated subsequent work in domain randomization, sim-to-real transfer, diffusion policies, and large-scale robotic grasping (QT-Opt), all of which train perception and control as one unified system.
Imagine learning to catch a ball with two people doing different jobs: one person watches the ball and shouts coordinates, and the other — blindfolded — swings their hand based on those coordinates. If the watcher says "2 meters left" but the catcher's arm can only move in 30 cm increments, the gap is never bridged. Errors in one person's job can never be fixed by the other.
Now imagine one person who can see the ball and move their hand. Their eyes naturally learn to track the features that matter for catching — the spin, the trajectory, the speed — and their arm adjusts in real time. That's end-to-end visuomotor learning: one neural network that sees the camera image and directly outputs motor torques, trained so that the "eyes" and the "hands" improve together.
The problem: modular pipelines break at the seams
Before this paper, a robot that needed to screw a cap onto a bottle would use a pipeline of separate modules:
- A vision system (often hand-crafted) to estimate the bottle's 3D position from the camera image.
- A state estimator to combine joint encoders, pose, and the vision estimate into a single state vector.
- A controller — typically PD or trajectory-following — to compute the torques needed to reach the estimated target.
Each module was designed independently and never updated based on the others' performance. If the vision module was off by 1 cm, the controller could not compensate, and the vision module had no incentive to be more precise since it was never told about task failure. For manipulation tasks requiring millimeter accuracy — like inserting a shape into a sorting cube — this modular approach consistently failed.
The key question of this paper is deceptively simple: does training perception and control jointly end-to-end give better performance than training each separately? Answering "yes" required solving two hard problems: (1) how to train a 92,000-parameter CNN policy with the tiny amount of data a real robot can provide, and (2) how to design a network architecture that preserves the spatial information a controller needs.
The trick: guided policy search (GPS)
Training a deep neural network policy directly with reinforcement learning in the real world is impractical — model-free methods like REINFORCE would need thousands of robot-hours. Guided policy search sidesteps this by splitting the problem into two easier sub-problems that alternate:
Step A — (the "teachers"): For each training condition (e.g. the bottle at position #3), optimize a simple time-varying using trajectory-centric RL with fitted local dynamics. These controllers have access to the full state (joint angles, object positions, velocities) and are cheap to optimize — think of them as expert tutors who can see everything.
Step B — Supervised learning (the "student"): Collect the trajectories from all the teachers and train the deep CNN policy to imitate their actions, but using only camera observations as input. The student must learn to see what the teachers know — it must infer object positions from raw pixels.
The key insight is that after training the student, the teachers adapt. They adjust their trajectories so the student can actually reproduce them. This back-and-forth is formalized as a optimization that converges to a locally optimal solution where the student policy matches the teachers' behavior.
The optimization objective
GPS starts from a constrained optimization: minimize the expected trajectory cost, subject to the constraint that the trajectory distribution matches the policy:
The supervised learning objective for the student policy becomes a weighted regression: minimize the difference between the policy's predicted action and the teacher's mean action, weighted by the precision (inverse covariance) of the teacher's controller — so the student is penalized more where the teacher is confident:
The architecture: spatial softmax for control
Standard CNNs for use to throw away spatial information — a cat is a cat regardless of where it appears. But a robot controller needs spatial information: "the bottle is 3 cm to the left" is exactly what matters. The authors designed a novel architecture with three key innovations:
1. No pooling. The convolutional layers preserve full spatial resolution (109×109 at the final layer).
2. Spatial softmax. Instead of max-pooling, each of the 32 response maps is passed through a softmax over spatial locations, converting each map into a probability distribution over "where is this feature in the image?"
3. Expected position. The expected (x, y) coordinate of each feature is computed as a weighted sum — a soft-argmax. This gives 32 feature points (64 coordinates) that form a compact spatial summary of the scene.
These feature points are concatenated with the robot's joint configuration and fed through two small fully connected layers (40 units each) that output the 7 motor torques. The entire network has only ~92,000 parameters — tiny by vision standards, but huge by 2016 policy search standards.
The full training pipeline
The complete training pipeline has three phases, designed to minimize robot interaction time:
Phase 1 — Pose pretraining: The robot moves the target object through random positions while recording camera images and object poses. These ~1000 image-pose pairs are used to pretrain the convolutional layers for pose regression. The first convolutional layer is initialized from ImageNet weights. This gives the vision layers a head start on learning what objects look like, before any control learning begins.
Phase 2 — Trajectory pretraining: For 15 iterations, guided policy search trains only the simple linear-Gaussian controllers (teachers), without optimizing the full visuomotor policy. A small fully-connected network on the full state is used temporarily to keep the trajectories coordinated. This gives the teachers a head start on learning the motor skills.
Phase 3 — End-to-end training: The full CNN policy is trained with GPS. In each iteration, the motor control layers are first optimized alone (they were not pretrained), then the entire network is fine-tuned end-to-end. This prevents the convolutional layers from forgetting useful features when the motor layers produce large initial errors. Typically 2–4 iterations of GPS suffice.
Real-world experiments: from pixels to torques
The authors evaluated their method on four manipulation tasks on a , each requiring vision-based control with millimeter precision:
- Coat hanger — hang a hanger on a rack, with 2 grasps × 3 rack positions.
- Shape sorting cube — insert a trapezoid into the matching hole, with 9 cube positions.
- Toy hammer — fit the claw under a nail, with 3 grasps × 5 nail positions.
- Bottle cap — screw a cap onto a bottle, with 9 bottle positions.
Three conditions were compared: end-to-end (the full method), pose features (pretrained vision features, only control layers trained with GPS), and pose prediction (vision outputs a 3D pose estimate, control trained separately).
What the network learned to see
Because the spatial softmax forces the network to express its visual understanding as feature points, we can literally see what the policy looks at. After end-to-end training, the learned feature points are qualitatively different from those learned for pose estimation alone:
- Task-relevant features emerge: The policy finds points on both the target object and the robot manipulator — both clearly relevant for guiding contact.
- Distractors are suppressed: The spatial softmax provides : weak activations are suppressed by the normalization, so the policy ignores irrelevant objects even when they were not present during training.
- Goal-driven, not recognition-driven: Pose estimation features concentrate on object edges for 3D localization. End-to-end features instead track the specific contact surfaces that matter for the task — the mouth of the bottle, the tip of the hammer claw.
The spatial softmax in code
Simplified to show the idea — not the real implementation.
import numpy as np
def spatial_softmax(feature_maps):
"""Convert (C, H, W) feature maps to (C, 2) feature point coordinates."""
C, H, W = feature_maps.shape
# 1. Softmax over spatial locations for each channel
flat = feature_maps.reshape(C, -1) # (C, H*W)
weights = np.exp(flat - flat.max(axis=1, keepdims=True))
weights /= weights.sum(axis=1, keepdims=True) # each channel sums to 1
weights = weights.reshape(C, H, W)
# 2. Expected position: weighted mean of (x, y) coordinates
xs = np.linspace(-1, 1, W) # normalized x coords
ys = np.linspace(-1, 1, H) # normalized y coords
grid_x, grid_y = np.meshgrid(xs, ys)
fp_x = (weights * grid_x).sum(axis=(1, 2)) # (C,)
fp_y = (weights * grid_y).sum(axis=(1, 2)) # (C,)
return np.stack([fp_x, fp_y], axis=1) # (C, 2) feature points
# Usage: 32 response maps → 32 feature points → 64 coordinates
# These 64 numbers + joint angles = everything the motor layers see
feature_maps = np.random.randn(32, 109, 109) # simulated conv3 output
points = spatial_softmax(feature_maps)
print(f"Shape: {points.shape}") # (32, 2)
# Each point is a differentiable (x, y) — the whole thing is trainableWhy it mattered
This paper answered a question that had been open since the earliest days of neural network robotics: can end-to-end training from pixels to torques actually work? The answer was not just "yes" but "yes, and it's significantly better than the modular alternative." Three lasting contributions shaped the field:
- The end-to-end principle for robotics. Joint training allows the vision system to learn features that are useful for control, not just for recognition. This principle now underlies virtually all modern robot learning systems.
- Spatial softmax. The soft-argmax operation that converts feature maps to coordinate predictions became a widely adopted building block, appearing in pose estimation, manipulation, and navigation architectures.
- Guided policy search as a bridge. GPS showed that you do not need to choose between the of model-based methods and the expressiveness of deep networks — you can have both by using the former to teach the latter. This teacher-student paradigm appears throughout modern robot learning.
2016
Visuomotor Policies (this paper)
First end-to-end method training deep CNNs from pixels to torques for real-world manipulation, using guided policy search and spatial softmax.
2017
Domain Randomization
Tobin et al. trained visuomotor policies in simulation with randomized textures, lighting, and camera positions, then transferred to the real world — extending the end-to-end principle to sim-to-real transfer.
2018
QT-Opt
Kalashnikov et al. scaled vision-based grasping to 580,000 real grasps across 7 robots, training a Q-function from pixels end-to-end. Proved the paradigm works at industrial scale.
2022
SayCan
Ahn et al. connected language models to visuomotor skills, using end-to-end trained manipulation primitives as the "hands" that a language model "brain" orchestrates.
2023
Diffusion Policy
Chi et al. replaced Gaussian policy outputs with denoising diffusion, enabling multi-modal action distributions while retaining end-to-end visuomotor training.
Every modern robot learning system that maps camera images to actions — from warehouse picking to surgical assistance — carries the DNA of this paper. The question it asked, and decisively answered, remains the design principle: train perception and control together.
CitationLevine, Finn, Darrell, Abbeel. End-to-End Training of Deep Visuomotor Policies. JMLR, 2016.
Terms in this paper
- Policyالسياسة
- Guided Policy Searchالبحث الموجَّه عن السياسة
- Convolutional Neural Network (CNN)الشبكة العصبية الالتفافية
- Spatial SoftmaxSoftmax مكاني
- Trajectoryمسار تتابع الحالات والأفعال
- Reinforcement Learningالتعلم المعزز
- end-to-endمن طرف إلى طرف
- Feature Pointنقطة السمة
- Torqueالعزم