Robotics2019advanced11 min read

Learning Dexterous In-Hand Manipulation

تعلُّم التلاعب البارع داخل اليد الآلية

OpenAI · Andrychowicz, M. · Baker, B. · Chociej, M. · Józefowicz, R. · McGrew, B. · Pachocki, J. · Petron, A. · Plappert, M. · Powell, G. · Ray, A. · Schneider, J. · Sidor, S. · Tobin, J. · Welinder, P. · Weng, L. · Zaremba, W. — IJRR

The problem

Dexterous — rotating, flipping, and repositioning objects using fingers alone — is one of the most difficult unsolved problems in robotics. The Shadow Dexterous Hand has 24 degrees of freedom and has been commercially available since 2005, yet no one could program it to manipulate objects dexterously. Classical planning requires exact physical models that are impossible to build for such a complex system, and directly on the physical robot is too slow and expensive — deep RL needs millions of trials, but hardware breaks after a few hundred.

The contribution

A complete system for learning entirely in simulation and transferring to a physical robot. Three pillars make transfer possible: (1) extensive of physics, observations, and visual appearance so the cannot overfit to any single simulation; (2) memory-augmented policies that perform implicit on the fly; (3) massive-scale distributed training generating 2 years of simulated experience per hour. The resulting policy naturally discovers human-like grasps — tripod, tip pinch, finger gaiting — without any demonstrations, and achieves up to 50 consecutive block rotations on the physical Shadow Hand.

The impact

This paper proved that can solve robotics problems previously considered impossible for learning-based methods. It established domain randomization as the dominant paradigm for bridging the , directly inspiring subsequent work like Solving Rubik's Cube with a Robot Hand (OpenAI, 2019), DexTreme (2023), and the broader wave of dexterous research. The distributed RL infrastructure (Rapid) became the template for large-scale robotic policy training.

Imagine you are training a surgeon, but you cannot let them touch a real patient until they are already skilled. So you build a thousand simulated operating rooms — some with slippery instruments, some with dim lights, some with organs in unusual positions. The surgeon practices in all of them simultaneously.

On the day of the real surgery, nothing looks exactly like any simulation — but nothing is completely surprising either. The surgeon adapts instantly because every possible surprise was already a Tuesday in one of those training rooms.

This is precisely what OpenAI did with a robotic hand: trained it across thousands of randomized simulations so that reality became just another variation.

The challenge: 24 degrees of freedom, zero demonstrations

The Shadow Dexterous Hand has five fingers driven by 20 pairs of agonist–antagonist tendons, giving it 24 degrees of freedom — comparable to the human hand. The task is in-hand object reorientation: place a block on the palm and rotate it to match a target orientation, over and over, without dropping it.

Classical robotics would require a precise physics model of every tendon, joint, and contact surface. But the Shadow Hand has tendon-based actuation that produces backlash, deformable fingertips, and state-dependent sensor noise — effects that rigid-body simulators cannot faithfully reproduce. This gap between simulation and reality is the reality gap, and it makes direct transfer of policies from simulation to hardware nearly impossible.

Training on the physical robot is equally impractical. requires millions of episodes, each lasting many seconds. Even at 12 Hz control frequency, gathering enough experience would take years — and the hand physically breaks during experimentation, requiring costly repairs.

Open in Lab
Explore the Shadow Dexterous Hand's 24 degrees of freedom. Click joints to see their range of motion and the tendons that control them.
The demo wakes as you arrive…

The key idea: domain randomization

If you cannot make one simulation perfectly match reality, make thousands of simulations that collectively surround reality. This is domain randomization: at the start of every training , randomize the physics, the observations, and the visual appearance so that no two episodes look alike. A policy that succeeds across all of these variations cannot have memorized the quirks of any single simulator — it must have learned something general enough to work in the real world too.

The randomizations fall into three categories:

Physics randomizations. Object dimensions are scaled by a factor in [0.95,1.05][0.95, 1.05], masses by [0.5,1.5][0.5, 1.5], friction coefficients by [0.7,1.3][0.7, 1.3], joint damping by a log-uniform factor in [0.3,3.0][0.3, 3.0], and even gravity gets Gaussian noise with σ=0.4 m/s2\sigma = 0.4\,\text{m/s}^2 per coordinate. These ranges are centered on calibrated values so the distribution straddles reality.

noise. Correlated noise (sampled once per episode) simulates sensor bias, while uncorrelated noise (sampled every step) simulates measurement jitter. Fingertip positions get ± 2 mm\pm\,2\,\text{mm} uncorrelated noise; object orientation gets ± 0.1 rad\pm\,0.1\,\text{rad}.

Unmodeled effects. Action delays, motor backlash, random external forces on the object, and simulated marker occlusion all inject real-world messiness that a pristine simulator would miss.

Open in Lab
Click "Randomize" to see how each episode gets different physics parameters. The green band shows where reality likely falls.
The demo wakes as you arrive…

The brain: LSTM policy trained with PPO

Many randomized parameters — object mass, surface friction, joint stiffness — persist throughout an episode. A smart agent should be able to figure out these hidden properties from a few seconds of interaction and adjust its strategy. This is exactly what system identification does in classical control, but here the policy learns to do it implicitly through memory.

The policy is a : inputs pass through a with ReLU activations, then into an LSTM with 512 hidden units. The LSTM maintains a that accumulates information across time steps. The authors showed that after just 5 seconds of interaction, the LSTM hidden state is predictive of whether the block is bigger or smaller than average — the network has learned to identify physical properties on the fly.

The value network has the same architecture but receives privileged information available only in simulation: true joint angles, joint velocities, and object velocities. This is the approach — the critic can see everything during training, making it easier to estimate returns, while the policy (actor) must cope with noisy, partial observations. At deployment time, only the actor is needed.

Open in Lab
The actor sees only noisy, partial observations. The critic sees everything during training — but is discarded at deployment.
The demo wakes as you arrive…

Training uses (PPO). PPO is an algorithm that updates the policy by maximizing a clipped surrogate objective:

LPPO=E ⁣[min⁡ ⁣(π(at∣st)πold(at∣st)A^t,  clip ⁣(π(at∣st)πold(at∣st), 1−ϵ, 1+ϵ)A^t)]L^{\text{PPO}} = \mathbb{E}\!\left[\min\!\left(\frac{\pi(a_t|s_t)}{\pi_{\text{old}}(a_t|s_t)}\hat{A}_t,\;\text{clip}\!\left(\frac{\pi(a_t|s_t)}{\pi_{\text{old}}(a_t|s_t)},\,1{-}\epsilon,\,1{+}\epsilon\right)\hat{A}_t\right)\right]
PPO clipped surrogate objective — The ratio ππold\frac{\pi}{\pi_{\text{old}}} measures how much the policy has changed. The clipping with ϵ≈0.2\epsilon \approx 0.2 prevents destructively large updates: if the new policy strays too far from the old one, the gradient is zeroed out. The advantage A^t\hat{A}_t (estimated via GAE with λ=0.95\lambda=0.95) tells whether an action was better or worse than average. PPO pushes up the probability of better-than-average actions while keeping updates conservative.

A key design choice is discrete actions. Although joint angles are continuous, the authors discretize each of the 20 action dimensions into 11 bins. They found that discrete action distributions outperform continuous Gaussian policies — possibly because a categorical distribution is more expressive and makes learning a good function simpler.

The reward is straightforward: rt=dt−dt+1r_t = d_t - d_{t+1}, where dtd_t is the rotation angle between the current and target orientation. A bonus of +5+5 is given when a goal is achieved (within 0.4 rad tolerance), and a penalty of −20-20 when the object is dropped. This naturally encourages progress toward the goal orientation while penalizing failure.

Scale: 2 years of experience per hour

Domain randomization massively increases the difficulty of learning. Without randomization, the policy converges in about 3 years of simulated experience (1.5 hours wall-clock). With full randomization, convergence requires around 100 years of simulated experience (50 hours wall-clock). This 30× increase demands serious compute.

The training system, called Rapid, uses 384 worker machines with 16 CPU cores each (6,144 cores total) to generate rollouts in parallel, and a single optimizer machine with 8 NVIDIA V100 GPUs. Workers download the latest policy parameters from Redis, simulate episodes, and push experience back. The optimizer pulls batches from Redis, stages them on GPUs, computes gradients, and averages them across GPUs with MPI. This pipeline generates approximately 2 years of simulated experience per hour.

Scaling experiments showed near-linear speedup up to 16 GPUs and 12,288 CPU cores, with diminishing returns beyond that.

Open in Lab
The Rapid distributed training pipeline: workers generate experience, Redis buffers it, and the optimizer updates the policy on GPUs.
The demo wakes as you arrive…

Eyes for the hand: vision-based pose estimation

For deployment outside a laboratory, the robot cannot rely on motion-capture markers glued to the object. It needs to see. The authors train a to estimate the object's 3D position and orientation from three RGB cameras.

The vision network uses a shared feature extractor per camera (two convolution layers, max-pooling, 4 ResNet blocks, and a spatial softmax layer), then concatenates the three feature vectors and feeds them through a fully connected network to predict 7 values: 3 for position and 4 for orientation (as a quaternion).

The critical trick: the vision model is trained entirely on synthetic images rendered in Unity with randomized appearance — different lighting, textures, materials, camera positions, and noise levels. This visual domain randomization follows the same philosophy as the physics randomization: make the training distribution so broad that real images fall within it. The result is a model with 5°5° rotation error and 9 mm9\,\text{mm} position error on real images — good enough for the control policy to work.

Results: human grasps emerge without human data

The most remarkable finding is what the policy discovers on its own. Without any human demonstrations or hand-crafted reward shaping, the policy naturally develops grasp types catalogued in human manipulation research: the tripod grasp, the tip pinch, the palmar pinch, the quadpod, and even power grasps. It also exhibits dynamic strategies like finger gaiting (walking the object between fingers), finger pivoting, controlled use of gravity, and coordinated translational and torsional forces.

Interestingly, the policy prefers to use the little finger for precision grasps. On the Shadow Hand, the little finger has an extra degree of freedom compared to the index and middle fingers — the opposite of human anatomy. The policy has re-discovered human grasps but adapted them to its own body.

Quantitatively, the best policy achieves a median of 50 consecutive successful block rotations in simulation and 13 on the physical robot (with motion capture) or 11.5 (with vision). With the wrist pitch joint locked — which avoids the most common failure mode — the median on the physical robot rises to 28.5.

Open in Lab
Six grasp types that the policy discovers on its own. Click each to see how it is used during manipulation.
The demo wakes as you arrive…

Ablation: what really matters for transfer

The authors systematically removed groups of randomizations and tested on the physical robot. The results are stark:

With all randomizations, the policy achieves a median of 13 rotations on the physical robot. With no randomizations (vanilla simulator), the median drops to 0 — the policy fails immediately. Removing physics randomizations alone drops performance to a median of 2. Removing unmodeled effects (backlash, delays, random forces) also drops to 2. Observation noise removal has a smaller effect in the motion-capture setup (median 8.5) but a dramatic one in the vision setup (median drops from 11.5 to 3.5).

Training without randomization requires about 3 years of simulated experience. With full randomization, it requires about 100 years — a 30× cost for making the policy robust enough to work in reality.

Memory matters too. Replacing the LSTM policy with a feed-forward network drops the physical robot median from 13 to 3. The LSTM hidden state is predictive of properties (object size) after just 5 seconds, confirming that the memory performs implicit system identification.

Open in Lab
Ablation results: toggle randomization groups and memory to see the impact on physical robot performance.
The demo wakes as you arrive…

Legacy: from blocks to Rubik's cubes and beyond

  1. 2017

    Domain Randomization (Tobin et al.)

    Introduced domain randomization for transferring object pose estimators from simulation to reality using randomized textures and lighting.

  2. 2017

    PPO (Schulman et al.)

    Proximal Policy Optimization became the workhorse algorithm for on-policy reinforcement learning, balancing sample efficiency with training stability.

  3. 2018

    This paper — Learning Dexterous In-Hand Manipulation

    Combined domain randomization, LSTM memory, and massive-scale PPO to achieve unprecedented dexterity on a physical 24-DoF hand, proving sim-to-real transfer viable for complex manipulation.

  4. 2019

    Solving Rubik's Cube with a Robot Hand

    Extended this work's approach to an even harder task — solving a Rubik's cube with in-hand manipulation — using Automatic Domain Randomization (ADR).

  5. 2023

    DexTreme — agile in-hand manipulation

    Transferred agile in-hand manipulation from simulation to reality using lessons from this paper: domain randomization, recurrent policies, and large-scale RL training.

  6. 2023

    ALOHA / ACT — bimanual dexterous manipulation

    Brought dexterous manipulation to low-cost bimanual robots using imitation learning, showing that the dexterity frontier this paper opened continues to expand with new learning paradigms.

CitationOpenAI, Andrychowicz, Baker, Chociej, Józefowicz, McGrew, Pachocki, Petron, Plappert, Powell, Ray, Schneider, Sidor, Tobin, Welinder, Weng, Zaremba. Learning Dexterous In-Hand Manipulation. IJRR, 2019.

Terms in this paper