Robotics2016advanced12 min read

End-to-End Training of Deep Visuomotor Policies

التدريب الشامل لسياسات بصرية-حركية عميقة

Levine, S. · Finn, C. · Darrell, T. · Abbeel, P. — JMLR

The problem

Practical robot learning systems in 2016 relied on hand-engineered pipelines: a vision module estimated object poses, a state estimator combined sensor readings, and a low-level controller computed motor commands. Each component was designed and tuned separately. Errors in the vision system could not be corrected by the controller, and the controller could not request better visual features. a directly from pixels to torques was considered infeasible — model-free needed far too many real-world samples, and the 92,000-parameter was orders of magnitude larger than policies prior methods could handle.

The contribution

A (GPS) method that decomposes visuomotor learning into two tractable pieces: (1) -centric RL using simple linear-Gaussian controllers that have access to the full state, and (2) that distills those trajectories into a deep CNN policy that only sees camera images. A novel architecture converts convolutional feature maps into 2D coordinates — a compact, spatial representation ideal for control. through GPS lets the vision layers discover task-specific features that outperform features trained only for . The result: the first method to train deep visuomotor policies for complex with direct control, using only minutes of real-world interaction.

The impact

This paper established visuomotor learning as a viable paradigm for real-world robotics. Its spatial softmax layer became a standard component in robotic perception networks. The guided policy search framework influenced a generation of robot learning methods, and the paper's core question — "does joint training beat separate training?" — with its affirmative answer, motivated subsequent work in domain randomization, sim-to-real transfer, diffusion policies, and large-scale robotic grasping (QT-Opt), all of which train perception and control as one unified system.

Imagine learning to catch a ball with two people doing different jobs: one person watches the ball and shouts coordinates, and the other — blindfolded — swings their hand based on those coordinates. If the watcher says "2 meters left" but the catcher's arm can only move in 30 cm increments, the gap is never bridged. Errors in one person's job can never be fixed by the other.

Now imagine one person who can see the ball and move their hand. Their eyes naturally learn to track the features that matter for catching — the spin, the trajectory, the speed — and their arm adjusts in real time. That's end-to-end visuomotor learning: one neural network that sees the camera image and directly outputs motor torques, trained so that the "eyes" and the "hands" improve together.

The problem: modular pipelines break at the seams

Before this paper, a robot that needed to screw a cap onto a bottle would use a pipeline of separate modules:

  • A vision system (often hand-crafted) to estimate the bottle's 3D position from the camera image.
  • A state estimator to combine joint encoders, pose, and the vision estimate into a single state vector.
  • A controller — typically PD or trajectory-following — to compute the torques needed to reach the estimated target.

Each module was designed independently and never updated based on the others' performance. If the vision module was off by 1 cm, the controller could not compensate, and the vision module had no incentive to be more precise since it was never told about task failure. For manipulation tasks requiring millimeter accuracy — like inserting a shape into a sorting cube — this modular approach consistently failed.

Open in Lab
Toggle between the modular pipeline and the end-to-end approach to see how errors propagate differently.
The demo wakes as you arrive…

The key question of this paper is deceptively simple: does training perception and control jointly end-to-end give better performance than training each separately? Answering "yes" required solving two hard problems: (1) how to train a 92,000-parameter CNN policy with the tiny amount of data a real robot can provide, and (2) how to design a network architecture that preserves the spatial information a controller needs.

The trick: guided policy search (GPS)

Training a deep neural network policy directly with reinforcement learning in the real world is impractical — model-free methods like REINFORCE would need thousands of robot-hours. Guided policy search sidesteps this by splitting the problem into two easier sub-problems that alternate:

Step A — (the "teachers"): For each training condition (e.g. the bottle at position #3), optimize a simple time-varying pi(ut∣xt)=N(Ktixt+kti,  Cti)p_i(\mathbf{u}_t|\mathbf{x}_t) = \mathcal{N}(K_t^i \mathbf{x}_t + k_t^i,\; C_t^i) using trajectory-centric RL with fitted local dynamics. These controllers have access to the full state xt\mathbf{x}_t (joint angles, object positions, velocities) and are cheap to optimize — think of them as expert tutors who can see everything.

Step B — Supervised learning (the "student"): Collect the trajectories from all the teachers and train the deep CNN policy πθ(ut∣ot)\pi_\theta(\mathbf{u}_t|\mathbf{o}_t) to imitate their actions, but using only camera observations ot\mathbf{o}_t as input. The student must learn to see what the teachers know — it must infer object positions from raw pixels.

The key insight is that after training the student, the teachers adapt. They adjust their trajectories so the student can actually reproduce them. This back-and-forth is formalized as a optimization that converges to a locally optimal solution where the student policy matches the teachers' behavior.

Open in Lab
Watch the teacher-student loop: teachers optimize trajectories with full state, the student learns to match them from images alone. Click "Next Iteration" to see them converge.
The demo wakes as you arrive…

The optimization objective

GPS starts from a constrained optimization: minimize the expected trajectory cost, subject to the constraint that the trajectory distribution matches the policy:

min⁡p, πθ  Ep ⁣[ℓ(τ)]s.t.p(ut∣xt)=πθ(ut∣xt)    ∀ xt,ut,t\min_{p,\,\pi_\theta}\; \mathbb{E}_{p}\!\bigl[\ell(\tau)\bigr] \quad \text{s.t.} \quad p(\mathbf{u}_t|\mathbf{x}_t) = \pi_\theta(\mathbf{u}_t|\mathbf{x}_t) \;\;\forall\, \mathbf{x}_t, \mathbf{u}_t, t
Constrained GPS objective — the foundation of the teacher-student loop — Minimize trajectory cost ℓ(τ) under guiding distribution p, subject to the constraint that p matches the policy πθ at every state and time step. BADMM relaxes this into alternating optimization of teachers (p) and student (πθ).

The supervised learning objective for the student policy becomes a weighted regression: minimize the difference between the policy's predicted action and the teacher's mean action, weighted by the precision (inverse covariance) of the teacher's controller — so the student is penalized more where the teacher is confident:

Lθ=12N∑i=1N∑t=1TEpi ⁣[(μπ(ot)−μtip(xt))⊤Cti−1(μπ(ot)−μtip(xt))]\mathcal{L}_\theta = \frac{1}{2N}\sum_{i=1}^{N}\sum_{t=1}^{T} \mathbb{E}_{p_i}\!\Bigl[ \bigl(\mu^\pi(\mathbf{o}_t) - \mu^{p}_{ti}(\mathbf{x}_t)\bigr)^\top C_{ti}^{-1} \bigl(\mu^\pi(\mathbf{o}_t) - \mu^{p}_{ti}(\mathbf{x}_t)\bigr) \Bigr]
Student's supervised loss — weighted regression against teacher trajectories — μᵖ is the teacher's mean action (computed from full state), μᵖⁱ is the policy's output (from observations only), and C⁻¹ is the teacher's precision matrix. The weighting ensures the student focuses on the parts of the trajectory that matter most.

The architecture: spatial softmax for control

Standard CNNs for use to throw away spatial information — a cat is a cat regardless of where it appears. But a robot controller needs spatial information: "the bottle is 3 cm to the left" is exactly what matters. The authors designed a novel architecture with three key innovations:

1. No pooling. The convolutional layers preserve full spatial resolution (109×109 at the final layer).

2. Spatial softmax. Instead of max-pooling, each of the 32 response maps is passed through a softmax over spatial locations, converting each map into a probability distribution over "where is this feature in the image?"

3. Expected position. The expected (x, y) coordinate of each feature is computed as a weighted sum — a soft-argmax. This gives 32 feature points (64 coordinates) that form a compact spatial summary of the scene.

These feature points are concatenated with the robot's joint configuration and fed through two small fully connected layers (40 units each) that output the 7 motor torques. The entire network has only ~92,000 parameters — tiny by vision standards, but huge by 2016 policy search standards.

Open in Lab
Watch how the spatial softmax converts a convolutional response map into a single (x, y) feature point. Toggle between max-pooling and spatial softmax to see the difference.
The demo wakes as you arrive…
sc,i,j=eac,i,j∑i′,j′eac,i′,j′fcx=∑i,jsc,i,j xi,jfcy=∑i,jsc,i,j yi,js_{c,i,j} = \frac{e^{a_{c,i,j}}}{\sum_{i',j'} e^{a_{c,i',j'}}} \qquad f_c^x = \sum_{i,j} s_{c,i,j}\, x_{i,j} \qquad f_c^y = \sum_{i,j} s_{c,i,j}\, y_{i,j}
Spatial softmax + expected position — the soft-argmax — For channel c: normalize activations aₖᵢⱼ into a distribution s over spatial locations, then compute the expected (x, y) position. The result is a differentiable 2D coordinate for each learned feature — no hand-labeled keypoints required.
Open in Lab
Click on each layer to see its role. Notice how the spatial softmax bottleneck converts high-resolution feature maps into compact (x, y) points.
The demo wakes as you arrive…

The full training pipeline

The complete training pipeline has three phases, designed to minimize robot interaction time:

Phase 1 — Pose pretraining: The robot moves the target object through random positions while recording camera images and object poses. These ~1000 image-pose pairs are used to pretrain the convolutional layers for pose regression. The first convolutional layer is initialized from ImageNet weights. This gives the vision layers a head start on learning what objects look like, before any control learning begins.

Phase 2 — Trajectory pretraining: For 15 iterations, guided policy search trains only the simple linear-Gaussian controllers (teachers), without optimizing the full visuomotor policy. A small fully-connected network on the full state is used temporarily to keep the trajectories coordinated. This gives the teachers a head start on learning the motor skills.

Phase 3 — End-to-end training: The full CNN policy is trained with GPS. In each iteration, the motor control layers are first optimized alone (they were not pretrained), then the entire network is fine-tuned end-to-end. This prevents the convolutional layers from forgetting useful features when the motor layers produce large initial errors. Typically 2–4 iterations of GPS suffice.

Open in Lab
Click each phase to see what it does and how much robot time it uses.
The demo wakes as you arrive…

Real-world experiments: from pixels to torques

The authors evaluated their method on four manipulation tasks on a , each requiring vision-based control with millimeter precision:

  • Coat hanger — hang a hanger on a rack, with 2 grasps × 3 rack positions.
  • Shape sorting cube — insert a trapezoid into the matching hole, with 9 cube positions.
  • Toy hammer — fit the claw under a nail, with 3 grasps × 5 nail positions.
  • Bottle cap — screw a cap onto a bottle, with 9 bottle positions.

Three conditions were compared: end-to-end (the full method), pose features (pretrained vision features, only control layers trained with GPS), and pose prediction (vision outputs a 3D pose estimate, control trained separately).

Open in Lab
Compare success rates across all tasks and conditions. Click each task to see the performance breakdown.
The demo wakes as you arrive…

What the network learned to see

Because the spatial softmax forces the network to express its visual understanding as feature points, we can literally see what the policy looks at. After end-to-end training, the learned feature points are qualitatively different from those learned for pose estimation alone:

  • Task-relevant features emerge: The policy finds points on both the target object and the robot manipulator — both clearly relevant for guiding contact.
  • Distractors are suppressed: The spatial softmax provides : weak activations are suppressed by the normalization, so the policy ignores irrelevant objects even when they were not present during training.
  • Goal-driven, not recognition-driven: Pose estimation features concentrate on object edges for 3D localization. End-to-end features instead track the specific contact surfaces that matter for the task — the mouth of the bottle, the tip of the hammer claw.
Open in Lab
Compare feature points from pose estimation vs end-to-end training. Notice how end-to-end features concentrate on task-relevant surfaces.
The demo wakes as you arrive…

The spatial softmax in code

Spatial softmax — from feature maps to 2D coordinatespython

Simplified to show the idea — not the real implementation.

import numpy as np

def spatial_softmax(feature_maps):
    """Convert (C, H, W) feature maps to (C, 2) feature point coordinates."""
    C, H, W = feature_maps.shape

    # 1. Softmax over spatial locations for each channel
    flat = feature_maps.reshape(C, -1)                   # (C, H*W)
    weights = np.exp(flat - flat.max(axis=1, keepdims=True))
    weights /= weights.sum(axis=1, keepdims=True)        # each channel sums to 1
    weights = weights.reshape(C, H, W)

    # 2. Expected position: weighted mean of (x, y) coordinates
    xs = np.linspace(-1, 1, W)  # normalized x coords
    ys = np.linspace(-1, 1, H)  # normalized y coords
    grid_x, grid_y = np.meshgrid(xs, ys)

    fp_x = (weights * grid_x).sum(axis=(1, 2))  # (C,)
    fp_y = (weights * grid_y).sum(axis=(1, 2))  # (C,)

    return np.stack([fp_x, fp_y], axis=1)  # (C, 2) feature points

# Usage: 32 response maps → 32 feature points → 64 coordinates
# These 64 numbers + joint angles = everything the motor layers see
feature_maps = np.random.randn(32, 109, 109)  # simulated conv3 output
points = spatial_softmax(feature_maps)
print(f"Shape: {points.shape}")  # (32, 2)
# Each point is a differentiable (x, y) — the whole thing is trainable

Why it mattered

This paper answered a question that had been open since the earliest days of neural network robotics: can end-to-end training from pixels to torques actually work? The answer was not just "yes" but "yes, and it's significantly better than the modular alternative." Three lasting contributions shaped the field:

  • The end-to-end principle for robotics. Joint training allows the vision system to learn features that are useful for control, not just for recognition. This principle now underlies virtually all modern robot learning systems.
  • Spatial softmax. The soft-argmax operation that converts feature maps to coordinate predictions became a widely adopted building block, appearing in pose estimation, manipulation, and navigation architectures.
  • Guided policy search as a bridge. GPS showed that you do not need to choose between the of model-based methods and the expressiveness of deep networks — you can have both by using the former to teach the latter. This teacher-student paradigm appears throughout modern robot learning.
  1. 2016

    Visuomotor Policies (this paper)

    First end-to-end method training deep CNNs from pixels to torques for real-world manipulation, using guided policy search and spatial softmax.

  2. 2017

    Domain Randomization

    Tobin et al. trained visuomotor policies in simulation with randomized textures, lighting, and camera positions, then transferred to the real world — extending the end-to-end principle to sim-to-real transfer.

  3. 2018

    QT-Opt

    Kalashnikov et al. scaled vision-based grasping to 580,000 real grasps across 7 robots, training a Q-function from pixels end-to-end. Proved the paradigm works at industrial scale.

  4. 2022

    SayCan

    Ahn et al. connected language models to visuomotor skills, using end-to-end trained manipulation primitives as the "hands" that a language model "brain" orchestrates.

  5. 2023

    Diffusion Policy

    Chi et al. replaced Gaussian policy outputs with denoising diffusion, enabling multi-modal action distributions while retaining end-to-end visuomotor training.

Every modern robot learning system that maps camera images to actions — from warehouse picking to surgical assistance — carries the DNA of this paper. The question it asked, and decisively answered, remains the design principle: train perception and control together.

CitationLevine, Finn, Darrell, Abbeel. End-to-End Training of Deep Visuomotor Policies. JMLR, 2016.

Terms in this paper