Computer Vision2020intermediate15 min read

RAFT: Recurrent All-Pairs Field Transforms for Optical Flow

RAFT: تحويلات الحقل التكرارية لجميع الأزواج للتدفق البصري

Teed, Z. · Deng, J. — ECCV

The problem

estimation — computing a dense displacement field between two consecutive video frames — is a foundational task in computer vision. By 2020, deep learning methods had achieved strong results, but they relied on pyramids that built separate feature extractors at each resolution. This multi-scale warping pipeline suffered from three problems: errors at coarse levels cascaded to finer levels, small fast-moving objects were missed when downsampled too aggressively, and the architectures were large, slow to train, and hard to generalize across datasets.

The contribution

RAFT introduces a three-stage architecture — , all-pairs , and recurrent iterative updates — that replaces the traditional coarse-to-fine warping pipeline. It builds a single 4D for all pairs of pixels at 1/8 resolution, pools it into a multi-scale pyramid, and then uses a lightweight convolutional to iteratively refine a flow field from zero. RAFT won the Best Paper Award at ECCV 2020, achieving a 16% error reduction on KITTI (F1-all 5.10%) and a 30% error reduction on Sintel final pass (EPE 2.855), with strong cross-dataset and 10× faster convergence.

The impact

RAFT fundamentally reshaped optical flow research by proving that iterative refinement over a single-resolution correlation volume outperforms multi-scale warping. Its architecture became the de facto template for subsequent methods — GMA, FlowFormer, SEA-RAFT — and its correlation lookup idea has been adopted in stereo matching, scene flow, and video depth estimation including Depth Anything. The design principle of "build a rich cost volume once, then refine iteratively" now underpins a wide family of dense correspondence methods.

Imagine you are a detective studying two aerial photographs of the same city taken one second apart. Your job: draw an arrow on every building, every car, every tree, showing exactly where it moved. You could try to guess all the arrows at once, but that's overwhelming. Instead, you start by building a massive reference book — for every spot in photo A, how similar is it to every spot in photo B? Then you make a first rough guess ("nothing moved"), look up each spot in your reference book, and nudge the arrows. You repeat this — look up, nudge, look up, nudge — and each round brings the arrows closer to truth.

That is exactly how RAFT works. The reference book is the correlation volume. The nudging process is the recurrent update operator. And the magic is that this simple loop — look up similarities, adjust the flow — produces more accurate results than any previous method, with a far simpler architecture.

What is optical flow?

Optical flow is the task of estimating a dense displacement field between two consecutive frames of a video. For every pixel (u,v)(u, v) in frame I1I_1, the goal is to find a displacement vector (f1,f2)(f^1, f^2) such that the pixel's corresponding location in frame I2I_2 is (u+f1,v+f2)(u + f^1, v + f^2). The result is a flow field — a map of arrows that describes how every pixel moved.

This is not about tracking a few special points; it is about computing motion for every single pixel in the image. A 1080p frame has over two million pixels, each needing its own displacement vector. Think of it as painting an arrow on every dot in a pointillist painting, showing which way that dot drifted between two moments in time.

Optical flow is used in video stabilization, action recognition, autonomous driving, video editing, and medical imaging. It is a foundational building block for any system that needs to understand motion in video.

The three-stage architecture

RAFT's design can be distilled into three stages, each clean and self-contained:

Stage 1 — Feature Extraction. A shared convolutional network encodes both input images into compact per-pixel feature vectors at 1/8 resolution. A second context processes only the first image and produces features that guide the update operator.

Stage 2 — Correlation Volume. The between every pair of feature vectors from the two images produces a 4D correlation volume of size H×W×H×WH \times W \times H \times W (at 1/8 resolution). This volume is then average-pooled on the last two dimensions to form a four-level correlation pyramid that captures similarity at multiple spatial scales.

Stage 3 — Iterative Updates. A convolutional GRU starts from an initial flow of zero and, at each iteration, looks up a local patch from the correlation pyramid around its current flow estimate, ingests correlation features together with context features, and produces a flow update Δf\Delta f. After many iterations, the flow converges to an accurate estimate.

All three stages are differentiable and trained .

Open in Lab
The three-stage RAFT pipeline: feature extraction, correlation volume construction, and iterative GRU-based flow refinement. Click each stage to learn more.
The demo wakes as you arrive…

Stage 1: Feature extraction — seeing through learned eyes

RAFT uses two encoder networks to convert raw images into compact representations:

Feature Encoder (gθg_\theta) — Applied to both images I1I_1 and I2I_2. It is a convolutional network with residual connections that maps each H×W×3H \times W \times 3 image to a H8×W8×D\frac{H}{8} \times \frac{W}{8} \times D (where D=256D = 256). The encoder uses — this is important because instance normalization operates per-image, making features invariant to global brightness or contrast shifts between the two frames. The same weights are shared for both images. Think of this as giving both photographs to the same analyst who applies the same interpretation rules to each one.

Context Encoder (hθh_\theta) — Applied only to I1I_1. It has the same architecture but uses instead of instance normalization, producing context features and an initial hidden state for the GRU. The context encoder captures information specific to the reference frame — edges, texture boundaries, semantic cues — that helps the update operator decide how to refine the flow at each location. If the feature encoder is like a translator converting images to a common language, the context encoder is like a local guide who knows the terrain of the first image.

Stage 2: The correlation volume — a similarity lookup table

The correlation volume is the heart of RAFT. Given feature maps gθ(I1)g_\theta(I_1) and gθ(I2)g_\theta(I_2), the correlation volume is computed by taking the dot product of every feature vector in the first image with every feature vector in the second image. If both feature maps have spatial dimensions H′×W′H' \times W' (where H′=H/8H' = H/8, W′=W/8W' = W/8), the result is a 4D tensor:

Cijkl=∑dgθ(I1)ij(d)⋅gθ(I2)kl(d)C_{ijkl} = \sum_{d} g_\theta(I_1)_{ij}^{(d)} \cdot g_\theta(I_2)_{kl}^{(d)}
4D Correlation Volume — dot product of all pairs of feature vectors — Entry CijklC_{ijkl} measures how similar the feature at position (i,j)(i,j) in I1I_1 is to the feature at position (k,l)(k,l) in I2I_2. The full volume C∈RH′×W′×H′×W′C \in \mathbb{R}^{H' \times W' \times H' \times W'} contains the answer to the question "for every pixel in frame 1, how well does it match every pixel in frame 2?" — computed once and reused across all iterations.

Think of the correlation volume as a massive phone book. For each pixel in the first image, you can look up how similar it is to any pixel in the second image. High values mean strong matches; low values mean poor matches. Instead of searching the entire phone book every time, RAFT builds a multi-scale index — the correlation pyramid.

Open in Lab
Select a pixel in the first image and see its correlation map across the second image. Brighter regions indicate higher similarity — the pixel's likely correspondence.
The demo wakes as you arrive…

The correlation pyramid: seeing both near and far

A single correlation volume captures fine-grained similarity but has a limited — it only directly shows matches within a local neighborhood. To handle large displacements (fast-moving objects), RAFT constructs a correlation pyramid by repeatedly average-pooling the last two dimensions of CC with kernel sizes 1, 2, 4, and 8:

At level 0, the volume has full H′×W′H' \times W' resolution on the second image's dimensions. At level 3, those dimensions are reduced to H′8×W′8\frac{H'}{8} \times \frac{W'}{8}. Crucially, the first two dimensions (the I1I_1 dimensions) are never pooled — this preserves per-pixel resolution in the reference frame.

The key insight: lower pyramid levels have a wider effective receptive field, capturing large displacements. Higher levels preserve fine spatial detail for small, precise motions. Together, they cover both scenarios. This is fundamentally different from traditional coarse-to-fine approaches because the pyramid is built from a single correlation volume at one feature resolution, not from separate feature extractors at multiple resolutions.

Correlation Lookup. At each iteration, the update operator needs to know "what does the correlation volume say about my current flow estimate?" For each pixel x\mathbf{x} in I1I_1, the current flow maps it to an estimated correspondence x′=x+f\mathbf{x}' = \mathbf{x} + \mathbf{f}. The lookup operator places a grid of radius rr around x′\mathbf{x}' in the correlation volume and bilinearly samples from it. This is done at every level of the pyramid, and the results are concatenated into a single feature vector.

Think of it as using your current guess to check a specific neighborhood in the phone book: "given where I think this pixel went, what do the nearby entries say about that guess?" The multi-level lookup means you check at multiple zoom levels simultaneously — a coarse overview and a fine-grained local view — giving the update operator both global context and local precision.

Open in Lab
Explore the 4-level correlation pyramid. Higher levels cover wider areas with less detail. The lookup operator samples from all levels around the current flow estimate.
The demo wakes as you arrive…

Stage 3: The update operator — iterating toward truth

The update operator is the engine of RAFT. It is inspired by classical algorithms that iteratively refine a solution. Starting from an initial flow f0=0\mathbf{f}_0 = \mathbf{0}, it produces a sequence of improved estimates {f1,f2,…,fN}\{\mathbf{f}_1, \mathbf{f}_2, \ldots, \mathbf{f}_N\} where each update is:

fk+1=fk+Δfk\mathbf{f}_{k+1} = \mathbf{f}_k + \Delta \mathbf{f}_k
Iterative flow update — each step adds a learned correction — The flow update Δfk\Delta \mathbf{f}_k is predicted by the GRU at each iteration. The goal is for the sequence to converge to a fixed point fk→f∗\mathbf{f}_k \rightarrow \mathbf{f}^*, the true optical flow. This is analogous to gradient descent, but instead of hand-crafting the descent direction, the network learns it.

At each iteration kk, the update operator performs four steps:

  1. Correlation lookup — uses fk\mathbf{f}_k to sample a local patch from all pyramid levels.
  2. Input assembly — concatenates the correlation features, the current flow (encoded through two layers), and the context features from the context encoder.
  3. GRU update — passes the assembled input through a convolutional GRU that updates a hidden state hk\mathbf{h}_k. The GRU has separate gates for horizontal and vertical spatial patterns (using 1×51 \times 5 and 5×15 \times 1 convolutions).
  4. Flow prediction — passes hk\mathbf{h}_k through two convolutional layers to produce Δfk\Delta \mathbf{f}_k.

The GRU's hidden state carries memory across iterations — it "remembers" what it has learned in previous refinement steps. This is like a detective who does not start from scratch each round but builds on accumulated evidence. The GRU weights are shared across all iterations, so the total parameter count stays small (5.3M for the full model).

zt=σ(Conv3×3([ht−1,xt],Wz)),rt=σ(Conv3×3([ht−1,xt],Wr))z_t = \sigma(\text{Conv}_{3 \times 3}([h_{t-1}, x_t], W_z)), \quad r_t = \sigma(\text{Conv}_{3 \times 3}([h_{t-1}, x_t], W_r))
GRU gates — controlling information flow — The update gate ztz_t controls how much of the new information to accept vs. how much of the old hidden state to keep. The reset gate rtr_t controls how much of the previous hidden state to forget when computing the candidate state. Together, they let the GRU decide: "should I keep refining my current estimate or make a bigger correction?"
Open in Lab
Watch the flow field converge from zero initialization over 12 iterations. Each step refines the estimate — early iterations make large corrections, later ones fine-tune.
The demo wakes as you arrive…

Convex upsampling: from coarse to full resolution

The update operator produces a flow field at 1/8 resolution. To recover full-resolution flow, RAFT uses learned convex — a technique more principled than .

For each pixel in the full-resolution output, the network predicts a set of 9 weights (a 3×33 \times 3 mask) that form a convex combination of the 9 nearest neighbors in the coarse flow field. The weights are predicted from the final hidden state of the GRU and passed through a softmax to ensure they sum to 1. This means each high-resolution pixel is a weighted average of its coarse neighbors, with the weights learned to respect motion boundaries.

Think of it as filling in a mosaic. Bilinear interpolation would blindly average the four nearest tiles. Convex upsampling asks the network: "these 9 coarse tiles are your neighbors — how should I blend them?" The network learns to give nearly all the weight to tiles on the same side of a motion boundary and almost zero weight to tiles on the other side, preserving sharp edges in the flow.

Open in Lab
Compare bilinear upsampling (blurry edges) vs. learned convex upsampling (sharp motion boundaries). The network predicts per-pixel blending weights that respect edges.
The demo wakes as you arrive…

Training: supervising every iteration

RAFT is trained with a that supervises every iteration of the update operator, not just the final output. For a sequence of NN predicted flows {f1,…,fN}\{\mathbf{f}_1, \ldots, \mathbf{f}_N\} and ground-truth flow fgt\mathbf{f}_{gt}:

L=∑i=1NγN−i∥fgt−fi∥1\mathcal{L} = \sum_{i=1}^{N} \gamma^{N-i} \|\mathbf{f}_{gt} - \mathbf{f}_i\|_1
Exponentially weighted multi-step loss — Each intermediate flow prediction is penalized with L1L_1 distance from the ground truth, weighted by γN−i\gamma^{N-i} where γ=0.8\gamma = 0.8. Later iterations receive higher weight because they are closer to the final answer. This encourages the early iterations to make meaningful progress too — if only the final output were supervised, the early iterations might learn to be lazy.

Why not coarse-to-fine? The paradigm shift

Before RAFT, the dominant approach to optical flow was the coarse-to-fine pyramid. The idea was intuitive: estimate flow at low resolution first (where large motions become small), then warp the second image using that estimate, and refine at higher resolution. Repeat until you reach full resolution.

This approach has three critical weaknesses that RAFT avoids:

Error cascading. Mistakes at coarse levels propagate to finer levels and cannot be corrected. A wrong large-motion estimate at the top level will misalign the warped image for all subsequent levels.

Small object loss. Aggressive downsampling at coarse levels can shrink small objects below a single pixel, making their motion unrecoverable.

Architectural complexity. Each pyramid level needs its own feature extractor, cost volume, and estimator — multiplying parameters and training difficulty.

RAFT sidesteps all three by operating at a single resolution (1/8). The correlation pyramid is not built from multiple feature resolutions — it is built by pooling a single correlation volume. Large motions are captured by the coarse levels of the pyramid, and the recurrent updates can correct early mistakes because each iteration re-examines the correlation evidence from scratch.

Results: new state of the art

RAFT set new records across multiple optical flow benchmarks:

On Sintel (final pass), RAFT achieved an end-point error (EPE) of 2.855 pixels — a 30% reduction from the previous best (4.098 pixels). On KITTI 2015, it achieved an F1-all error of 5.10% — a 16% reduction from the previous best (6.10%).

Beyond raw accuracy, three aspects stood out:

Cross-dataset generalization. When trained only on synthetic data (FlyingChairs + FlyingThings3D) and evaluated on KITTI without fine-tuning, RAFT achieved 5.04 EPE — a 40% improvement over the best prior method. This suggests the architecture learns genuinely transferable motion representations.

Training efficiency. RAFT converges in approximately 100K iterations — roughly 10× fewer than competing methods like FlowNet2 or PWC-Net.

Compact model. The full model has only 5.3M parameters. A smaller variant (RAFT-S) with 1.0M parameters still outperforms all prior methods on Sintel, running at 20 FPS on a 1080Ti GPU.

Open in Lab
RAFT's error reduction across benchmarks compared to prior methods. Lower is better.
The demo wakes as you arrive…

Code: the update loop in pseudocode

RAFT inference loop (simplified)python

Simplified to show the idea — not the real implementation.

# Extract features from both images
fmap1 = feature_encoder(image1)  # [B, 256, H/8, W/8]
fmap2 = feature_encoder(image2)  # shared weights
# Extract context + initial hidden state from image1
cnet, h = context_encoder(image1)
# Build 4D correlation volume (all pairs dot product)
corr_volume = einsum('bchw, bckl -> bhwkl', fmap1, fmap2)
corr_pyramid = build_pyramid(corr_volume)  # 4 levels
# Initialize flow at zero
flow = torch.zeros(B, 2, H//8, W//8)
for k in range(num_iterations):  # typically 12
    # Look up correlation around current flow estimate
    corr_features = corr_lookup(corr_pyramid, flow, radius=4)
    # Encode current flow
    flow_features = flow_encoder(flow)
    # GRU update: correlation + flow + context -> new hidden state
    h = gru(h, corr_features, flow_features, cnet)
    # Predict flow delta from hidden state
    delta_flow = flow_head(h)
    flow = flow + delta_flow

# Convex upsample from 1/8 to full resolution
flow_up = convex_upsample(flow, mask=upsample_head(h))

Legacy: the iterative refinement revolution

  1. 2015

    FlowNet — deep learning enters optical flow

    Dosovitskiy et al. showed that a single CNN can estimate optical flow end-to-end, trained on synthetic data. Opened the door to learned methods.

  2. 2017

    FlowNet2 — stacking networks for refinement

    Ilg et al. stacked multiple FlowNet architectures in a cascade, achieving competitive accuracy but at high computational cost and model size.

  3. 2018

    PWC-Net — pyramids, warping, cost volumes

    Sun et al. combined feature pyramids with iterative warping and cost volumes, becoming the dominant paradigm. Fast and accurate, but suffered from coarse-to-fine error cascading.

  4. 2020

    RAFT — this paper (Best Paper, ECCV 2020)

    Replaced coarse-to-fine warping with all-pairs correlation + recurrent iterative updates. 30% error reduction on Sintel, 16% on KITTI, 10× faster convergence, and strong cross-dataset generalization.

  5. 2021

    GMA — Global Motion Aggregation

    Jiang et al. added self-attention to RAFT's update operator to aggregate motion information globally, improving performance on occluded regions.

  6. 2022

    FlowFormer — Transformers for optical flow

    Huang et al. replaced the GRU with a Transformer-based architecture, building on RAFT's correlation volume but using attention for the update step.

  7. 2024

    SEA-RAFT & Depth Anything

    SEA-RAFT simplified and accelerated RAFT with a mixture-of-Laplace loss. Depth Anything extended RAFT's correlation-based ideas to monocular depth estimation, showing the architecture's influence beyond optical flow.

RAFT proved that you do not need complex multi-scale architectures to estimate dense correspondences. A single correlation volume, a lightweight recurrent updater, and iterative refinement — these three simple ideas, composed well, outperformed years of architectural engineering. The lesson extends beyond optical flow: when in doubt, prefer iterative refinement over single-shot prediction. Let the network correct its own mistakes.

CitationTeed, Deng. RAFT: Recurrent All-Pairs Field Transforms for Optical Flow. ECCV, 2020.

Terms in this paper