Robotics2024advanced12 min read

π₀: A Vision-Language-Action Flow Model for General Robot Control

π₀: نموذج تدفُّقي للرؤية واللغة والفعل للتحكم العام بالروبوتات

Black, K. · Brown, N. · Driess, D. · Esmail, A. · Equi, M. · Finn, C. · Fusai, N. · Groom, L. · Hausman, K. · Ichter, B. · Jakubczak, S. · Jones, T. · Ke, L. · Levine, S. · Li-Bell, A. · Mothukuri, M. · Nair, S. · Pertsch, K. · Shi, L. X. · Tanner, J. · Vuong, Q. · Walling, A. · Wang, H. · Zhilinsky, U. — RSS

The problem

By 2024, robot learning had made strides in simple pick-and-place tasks, but building a single policy that generalizes across diverse robots, tasks, and environments remained elusive. Existing vision-language-action models used token prediction for actions, which struggled with the high-frequency, continuous, and multimodal nature of . Meanwhile, diffusion-based policies achieved dexterity but lacked the semantic understanding that comes from internet-scale . No model combined both capabilities at scale.

The contribution

π₀ is a 3.3B-parameter that combines a pre-trained VLM (PaliGemma) with for continuous action generation. The key innovation is an — a separate set of weights that processes robot states and actions via flow matching while sharing with the VLM backbone. Pre-trained on 10,000+ hours of data from 7 robot configurations and 68 tasks plus open-source data (OXE, DROID, Bridge), π₀ can follow language commands zero-shot, and be fine-tuned to complex dexterous tasks like laundry folding and box assembly, significantly outperforming prior VLAs and task-specific methods.

The impact

π₀ established a new paradigm for robot foundation models by demonstrating that the pre-train → fine-tune recipe from language models transfers to robotics when combined with flow matching. It showed that a single model can master tasks ranging from simple pick-and-place to 20-minute laundry folding sequences — the longest dexterous tasks in the end-to-end robot learning literature. π₀ directly inspired the subsequent π₀.5 model with open-world generalization and accelerated the VLA field toward practical deployment.

Imagine an orchestra conductor who has spent years reading every music textbook and watching every concert recording in the world. They understand music deeply — but they have never held a baton. Now hand them a baton and ask them to conduct 7 different orchestras, each with different instruments, in 68 different pieces.

The conductor cannot wave their arms in broad, clumsy sweeps — the musicians need 50 precise micro-gestures per second. So instead of discretizing music into rough beats, the conductor learns to sculpt fluid motion from noise: starting with a random wave and refining it step by step until it becomes the exact gesture the music demands.

π₀ is that conductor. Its musical knowledge comes from a pre-trained on the internet. Its baton technique comes from flow matching. And its repertoire comes from 10,000+ hours of watching robots work.

The challenge: why one robot brain for everything is so hard

By 2024, the field of robot learning was split into two worlds. In one world, vision-language models like RT-2 could understand scenes and follow language instructions, but they treated robot actions as discrete tokens — like words in a sentence. This works for slow, coarse tasks (pick up the cup, move it left), but falls apart for dexterous at 50 Hz: folding fabric, packing eggs, or assembling boxes. You cannot describe 50 precise joint angle adjustments per second as a sequence of vocabulary words.

In the other world, Diffusion Policy and similar methods could produce smooth, high-frequency action trajectories using diffusion models. These excelled at dexterity but were trained from scratch on small, task-specific datasets, with no access to the broad semantic knowledge that internet-scale pre- provides.

The gap between these two worlds is the central problem π₀ addresses: how to build a single model that inherits the semantic understanding of a vision-language model AND produces the precise, continuous, high-frequency actions that dexterous manipulation demands.

Open in Lab
The landscape of robot learning approaches before π₀. Autoregressive VLAs had semantic knowledge but lacked dexterity. Diffusion policies had dexterity but lacked semantics. π₀ occupies the top-right corner where both meet.
The demo wakes as you arrive…

Architecture: the VLM backbone + action expert design

The architecture of π₀ starts with PaliGemma, a 3-billion parameter open-source vision-language model. PaliGemma already understands images and text — it was trained on internet-scale data to do things like image captioning and visual question answering. The key question is: how do you add robot control on top of this without destroying the semantic knowledge?

The answer is a dual-expert design inspired by Transfusion. The model has two sets of transformer weights. The first — the VLM backbone — processes images and language tokens using the original PaliGemma weights. The second — the action expert — is a smaller set of weights (300M parameters) dedicated to processing robot states and actions.

Both experts share the same attention mechanism. This means action tokens can attend to image and language tokens (understanding what to do), while image and language tokens can attend to action tokens (understanding where the robot is). But the feedforward networks — where most of the computation happens — are separate. This separation is crucial: language understanding and precise motor control require fundamentally different computations.

Think of it as a brain with two hemispheres sharing a communication channel. The language hemisphere understands "fold the shirt" and recognizes what a shirt looks like. The motor hemisphere translates that understanding into 50 precise joint commands per second. The shared attention is the corpus callosum connecting them.

Open in Lab
The π₀ architecture: images and language flow through the VLM backbone, while robot states and actions flow through the action expert. Both share attention but use separate feedforward weights. Click on each component to learn more.
The demo wakes as you arrive…

Flow matching: sculpting actions from noise

Flow matching is π₀'s mechanism for generating continuous actions. The intuition is simple: imagine you have a block of marble (random noise) and want to sculpt it into a precise statue (the correct robot action). The model learns a vector field — at every point in the noisy space, it tells you which direction to push the marble to get closer to the statue. Start from noise, follow the vectors for 10 steps, and you arrive at a precise action.

More formally, π₀ models the conditional distribution of action chunks At\mathbf{A}_t given an observation ot\mathbf{o}_t. An action chunk is a sequence of H=50H = 50 future actions — one full second of robot motion at 50 Hz. During training, the model learns to predict the vector field that transforms noise into actions. During , it starts from random Gaussian noise and integrates the vector field in 10 steps.

The critical advantage over standard diffusion is the linear probability path: the transition from noise to signal follows a straight line rather than a curved path. This makes the integration simpler and more efficient — 10 steps suffice where diffusion might need 50 or 100.

Lτ(θ)=Ep(At∣ot), q(Atτ∣At)∥vθ(Atτ,ot)−u(Atτ∣At)∥2L^{\tau}(\theta) = \mathbb{E}_{p(\mathbf{A}_t|\mathbf{o}_t),\,q(\mathbf{A}_t^{\tau}|\mathbf{A}_t)} \left\| \mathbf{v}_\theta(\mathbf{A}_t^{\tau}, \mathbf{o}_t) - \mathbf{u}(\mathbf{A}_t^{\tau}|\mathbf{A}_t) \right\|^2
Flow matching training loss — learning to predict the denoising vector field — The model vθ\mathbf{v}_\theta learns to predict the target vector field u(Atτ∣At)=ϵ−At\mathbf{u}(\mathbf{A}_t^{\tau}|\mathbf{A}_t) = \epsilon - \mathbf{A}_t, where ϵ\epsilon is random noise. The noisy actions are constructed as Atτ=τAt+(1−τ)ϵ\mathbf{A}_t^{\tau} = \tau\mathbf{A}_t + (1-\tau)\epsilon, forming a straight-line interpolation between noise (τ=0\tau{=}0) and the true action (τ=1\tau{=}1). The flow timestep τ\tau is sampled from a beta distribution that emphasizes noisier timesteps.

At inference time, actions are generated by integrating the learned vector field using forward Euler steps:

Atτ+δ=Atτ+δ vθ(Atτ,ot)\mathbf{A}_t^{\tau+\delta} = \mathbf{A}_t^{\tau} + \delta\,\mathbf{v}_\theta(\mathbf{A}_t^{\tau}, \mathbf{o}_t)
Euler integration — 10 steps from noise to action — Starting from random noise At0∼N(0,I)\mathbf{A}_t^0 \sim \mathcal{N}(\mathbf{0}, \mathbf{I}), the model applies 10 denoising steps with step size δ=0.1\delta = 0.1. At each step, it queries the action expert for the predicted vector field, nudging the noisy action toward a clean action. The observation prefix is cached so only the action tokens are recomputed at each step, making inference efficient.
Open in Lab
Watch flow matching sculpt actions from noise. Each step follows the learned vector field along a straight path from random Gaussian noise to precise robot actions.
The demo wakes as you arrive…

Cross-embodiment training: one brain, seven robots

π₀ is not trained on a single robot. It is trained simultaneously on 7 different robot configurations — single-arm robots (UR5e, Franka), dual-arm robots (bimanual UR5e, bimanual Trossen, bimanual ARX), and mobile manipulators (Mobile Trossen, Mobile Fibocom). These robots have different numbers of joints (6 to 18 degrees of freedom), different numbers of cameras (2 or 3), and different action spaces.

The key trick is a universal action representation. All robots are represented as 18-dimensional vectors (the largest in the dataset). Robots with fewer joints are zero-padded. Robots with fewer cameras have their missing image slots masked. A single model processes all of them, learning to share visual and semantic knowledge across embodiments while specializing its motor outputs for each configuration.

In addition to the dexterous data (10,000+ hours), the pre-training mixture includes the Open X-Embodiment dataset from 22 additional robots, plus Bridge v2 and DROID. The total pre-training mixture spans 29+ distinct robot platforms — an unprecedented scale.

Open in Lab
Seven robot configurations feed into a single π₀ model. Click each robot to see its action space dimension and camera setup. The universal 18-D vector accommodates all.
The demo wakes as you arrive…

Training recipe: pre-training then post-training

π₀ follows the same two-phase recipe that made large language models successful. The pre-training phase exposes the model to the full diversity of robot data — all 7 configurations, all 68 tasks, plus OXE data. The goal is to build a base model with broad capabilities. The quality of this data varies: some demonstrations are smooth and efficient, others include mistakes and recoveries. This diversity is valuable because it teaches the model how to handle unexpected situations and recover from errors.

The post-training phase (or ) uses smaller, higher-quality datasets curated for specific downstream tasks. This phase teaches the model to execute the desired task fluently and efficiently, much like how language models are aligned with curated instruction data after pre-training.

The intuition for why both phases matter: training only on high-quality data produces a model that performs perfectly when everything goes right, but breaks when anything goes wrong (because it never saw mistakes). Training only on diverse pre-training data produces a model that can handle anything but does nothing particularly well. The combination gives you a model that attempts the efficient strategy from the post-training data, but falls back on recovery behaviors from the pre-training data when things go sideways.

Open in Lab
Toggle between pre-training and post-training data. Pre-training is diverse but noisy; post-training is clean but narrow. The combination produces robust, fluent behavior.
The demo wakes as you arrive…

Results: from zero-shot to laundry folding

The evaluation spans four experiments. First, zero-shot evaluation after pre-training on 5 tasks (shirt folding, bussing easy/hard, grocery bagging, toast removal). π₀ significantly outperformed all baselines — OpenVLA, Octo, and π₀-small — even when given the same compute budget. OpenVLA particularly struggled because its autoregressive discretization cannot handle 50 Hz action chunks.

Second, language following tests showed that π₀'s VLM backbone enables much better instruction following than π₀-small (which lacks VLM initialization). This translates to better performance with both human-provided and VLM-generated intermediate commands.

Third, fine-tuning to new tasks (bowl stacking, towel folding, tupperware in microwave, paper towel replacement, Franka items in drawer) showed that π₀ outperformed ACT, Diffusion Policy, OpenVLA, and Octo. Pre-training amplified gains especially on tasks closer to the pre-training distribution, and the benefit was largest with small fine-tuning datasets (1 hour).

Fourth, complex multi-stage tasks demonstrated π₀'s ability to master 5–20 minute dexterous sequences: full laundry folding (fetch from dryer, transport, fold), table bussing with novel objects, box assembly from flat cardboard, and egg packing. Pre-training + fine-tuning consistently outperformed both zero-shot and training from scratch.

Open in Lab
Zero-shot evaluation results across 5 tasks. π₀ outperforms all baselines by a large margin. Even π₀ at 160k steps (compute parity) beats all other methods.
The demo wakes as you arrive…

Why this matters: the LLM playbook comes to robotics

π₀ is significant not just for what it achieves, but for the paradigm it validates. Large language models succeeded because of three ingredients: internet-scale pre-training, the transformer architecture, and a pre-train → fine-tune recipe. π₀ shows that this same formula works for robot control when you replace autoregressive text prediction with flow matching for actions.

The implications are profound. Instead of training a separate policy for every task and every robot, you train one generalist foundation model and fine-tune it. This fundamentally changes the economics of robot learning: data from every task and every robot becomes useful to every other task and robot. A laundry-folding demonstration helps the model learn better grocery-bagging, because both involve grasping deformable objects.

The pre-training/post-training split also changes how we think about robot data quality. You no longer need all data to be perfect. Messy, diverse pre-training data teaches the model to handle the real world's chaos. Clean post-training data teaches it to do the specific task well. This mirrors exactly how LLMs are trained: internet text teaches knowledge, instruction tuning teaches behavior.

Context: the road to robot foundation models

  1. 2022

    RT-1 — Robotics Transformer

    Google's first large-scale robot transformer, trained on 130k demonstrations for 700+ tasks. Showed that transformer architectures work for robot control.

  2. 2023

    RT-2 — VLA model

    Showed that VLMs can be fine-tuned into VLAs by treating actions as text tokens. Transferred web knowledge to robotic control but used autoregressive discretization.

  3. 2023

    Diffusion Policy & ACT

    Showed that diffusion models and action chunking enable dexterous manipulation. Task-specific, trained from scratch on small datasets.

  4. 2023

    Open X-Embodiment

    Collaborative effort to create a cross-embodiment robot dataset from 22 robot platforms. Showed that cross-robot training improves generalization.

  5. 2024

    π₀ (this paper)

    Combined VLM pre-training with flow matching in a single cross-embodiment model. Demonstrated the longest dexterous tasks in robot learning literature.

π₀ represents a convergence point in robot learning. It shows that the recipe that worked for language (pre-train broadly, fine-tune narrowly) and the recipe that worked for dexterity (diffusion-based action generation) are not just compatible — they are synergistic. The broad pre-training makes fine-tuning more data-efficient, and the flow matching makes VLM-based control precise enough for real-world dexterity. This convergence suggests that the era of general-purpose robot foundation models has begun.

CitationBlack, Brown, Driess, Esmail, Equi, Finn, Fusai, Groom, Hausman, Ichter, Jakubczak, Jones, Ke, Levine, Li-Bell, Mothukuri, Nair, Pertsch, Shi, Tanner, Vuong, Walling, Wang, Zhilinsky. π₀: A Vision-Language-Action Flow Model for General Robot Control. RSS, 2024.

Terms in this paper