Computer Vision2020intermediate15 min read
RAFT: Recurrent All-Pairs Field Transforms for Optical Flow
RAFT: تحويلات الحقل التكرارية لجميع الأزواج للتدفق البصري
Teed, Z. · Deng, J. — ECCV
The problem
estimation — computing a dense displacement field between two consecutive video frames — is a foundational task in computer vision. By 2020, deep learning methods had achieved strong results, but they relied on pyramids that built separate feature extractors at each resolution. This multi-scale warping pipeline suffered from three problems: errors at coarse levels cascaded to finer levels, small fast-moving objects were missed when downsampled too aggressively, and the architectures were large, slow to train, and hard to generalize across datasets.
The contribution
RAFT introduces a three-stage architecture — , all-pairs , and recurrent iterative updates — that replaces the traditional coarse-to-fine warping pipeline. It builds a single 4D for all pairs of pixels at 1/8 resolution, pools it into a multi-scale pyramid, and then uses a lightweight convolutional to iteratively refine a flow field from zero. RAFT won the Best Paper Award at ECCV 2020, achieving a 16% error reduction on KITTI (F1-all 5.10%) and a 30% error reduction on Sintel final pass (EPE 2.855), with strong cross-dataset and 10× faster convergence.
The impact
RAFT fundamentally reshaped optical flow research by proving that iterative refinement over a single-resolution correlation volume outperforms multi-scale warping. Its architecture became the de facto template for subsequent methods — GMA, FlowFormer, SEA-RAFT — and its correlation lookup idea has been adopted in stereo matching, scene flow, and video depth estimation including Depth Anything. The design principle of "build a rich cost volume once, then refine iteratively" now underpins a wide family of dense correspondence methods.
Imagine you are a detective studying two aerial photographs of the same city taken one second apart. Your job: draw an arrow on every building, every car, every tree, showing exactly where it moved. You could try to guess all the arrows at once, but that's overwhelming. Instead, you start by building a massive reference book — for every spot in photo A, how similar is it to every spot in photo B? Then you make a first rough guess ("nothing moved"), look up each spot in your reference book, and nudge the arrows. You repeat this — look up, nudge, look up, nudge — and each round brings the arrows closer to truth.
That is exactly how RAFT works. The reference book is the correlation volume. The nudging process is the recurrent update operator. And the magic is that this simple loop — look up similarities, adjust the flow — produces more accurate results than any previous method, with a far simpler architecture.
What is optical flow?
Optical flow is the task of estimating a dense displacement field between two consecutive frames of a video. For every pixel in frame , the goal is to find a displacement vector such that the pixel's corresponding location in frame is . The result is a flow field — a map of arrows that describes how every pixel moved.
This is not about tracking a few special points; it is about computing motion for every single pixel in the image. A 1080p frame has over two million pixels, each needing its own displacement vector. Think of it as painting an arrow on every dot in a pointillist painting, showing which way that dot drifted between two moments in time.
Optical flow is used in video stabilization, action recognition, autonomous driving, video editing, and medical imaging. It is a foundational building block for any system that needs to understand motion in video.
The three-stage architecture
RAFT's design can be distilled into three stages, each clean and self-contained:
Stage 1 — Feature Extraction. A shared convolutional network encodes both input images into compact per-pixel feature vectors at 1/8 resolution. A second context processes only the first image and produces features that guide the update operator.
Stage 2 — Correlation Volume. The between every pair of feature vectors from the two images produces a 4D correlation volume of size (at 1/8 resolution). This volume is then average-pooled on the last two dimensions to form a four-level correlation pyramid that captures similarity at multiple spatial scales.
Stage 3 — Iterative Updates. A convolutional GRU starts from an initial flow of zero and, at each iteration, looks up a local patch from the correlation pyramid around its current flow estimate, ingests correlation features together with context features, and produces a flow update . After many iterations, the flow converges to an accurate estimate.
All three stages are differentiable and trained .
Stage 1: Feature extraction — seeing through learned eyes
RAFT uses two encoder networks to convert raw images into compact representations:
Feature Encoder () — Applied to both images and . It is a convolutional network with residual connections that maps each image to a (where ). The encoder uses — this is important because instance normalization operates per-image, making features invariant to global brightness or contrast shifts between the two frames. The same weights are shared for both images. Think of this as giving both photographs to the same analyst who applies the same interpretation rules to each one.
Context Encoder () — Applied only to . It has the same architecture but uses instead of instance normalization, producing context features and an initial hidden state for the GRU. The context encoder captures information specific to the reference frame — edges, texture boundaries, semantic cues — that helps the update operator decide how to refine the flow at each location. If the feature encoder is like a translator converting images to a common language, the context encoder is like a local guide who knows the terrain of the first image.
Stage 2: The correlation volume — a similarity lookup table
The correlation volume is the heart of RAFT. Given feature maps and , the correlation volume is computed by taking the dot product of every feature vector in the first image with every feature vector in the second image. If both feature maps have spatial dimensions (where , ), the result is a 4D tensor:
Think of the correlation volume as a massive phone book. For each pixel in the first image, you can look up how similar it is to any pixel in the second image. High values mean strong matches; low values mean poor matches. Instead of searching the entire phone book every time, RAFT builds a multi-scale index — the correlation pyramid.
The correlation pyramid: seeing both near and far
A single correlation volume captures fine-grained similarity but has a limited — it only directly shows matches within a local neighborhood. To handle large displacements (fast-moving objects), RAFT constructs a correlation pyramid by repeatedly average-pooling the last two dimensions of with kernel sizes 1, 2, 4, and 8:
At level 0, the volume has full resolution on the second image's dimensions. At level 3, those dimensions are reduced to . Crucially, the first two dimensions (the dimensions) are never pooled — this preserves per-pixel resolution in the reference frame.
The key insight: lower pyramid levels have a wider effective receptive field, capturing large displacements. Higher levels preserve fine spatial detail for small, precise motions. Together, they cover both scenarios. This is fundamentally different from traditional coarse-to-fine approaches because the pyramid is built from a single correlation volume at one feature resolution, not from separate feature extractors at multiple resolutions.
Correlation Lookup. At each iteration, the update operator needs to know "what does the correlation volume say about my current flow estimate?" For each pixel in , the current flow maps it to an estimated correspondence . The lookup operator places a grid of radius around in the correlation volume and bilinearly samples from it. This is done at every level of the pyramid, and the results are concatenated into a single feature vector.
Think of it as using your current guess to check a specific neighborhood in the phone book: "given where I think this pixel went, what do the nearby entries say about that guess?" The multi-level lookup means you check at multiple zoom levels simultaneously — a coarse overview and a fine-grained local view — giving the update operator both global context and local precision.
Stage 3: The update operator — iterating toward truth
The update operator is the engine of RAFT. It is inspired by classical algorithms that iteratively refine a solution. Starting from an initial flow , it produces a sequence of improved estimates where each update is:
At each iteration , the update operator performs four steps:
- Correlation lookup — uses to sample a local patch from all pyramid levels.
- Input assembly — concatenates the correlation features, the current flow (encoded through two layers), and the context features from the context encoder.
- GRU update — passes the assembled input through a convolutional GRU that updates a hidden state . The GRU has separate gates for horizontal and vertical spatial patterns (using and convolutions).
- Flow prediction — passes through two convolutional layers to produce .
The GRU's hidden state carries memory across iterations — it "remembers" what it has learned in previous refinement steps. This is like a detective who does not start from scratch each round but builds on accumulated evidence. The GRU weights are shared across all iterations, so the total parameter count stays small (5.3M for the full model).
Convex upsampling: from coarse to full resolution
The update operator produces a flow field at 1/8 resolution. To recover full-resolution flow, RAFT uses learned convex — a technique more principled than .
For each pixel in the full-resolution output, the network predicts a set of 9 weights (a mask) that form a convex combination of the 9 nearest neighbors in the coarse flow field. The weights are predicted from the final hidden state of the GRU and passed through a softmax to ensure they sum to 1. This means each high-resolution pixel is a weighted average of its coarse neighbors, with the weights learned to respect motion boundaries.
Think of it as filling in a mosaic. Bilinear interpolation would blindly average the four nearest tiles. Convex upsampling asks the network: "these 9 coarse tiles are your neighbors — how should I blend them?" The network learns to give nearly all the weight to tiles on the same side of a motion boundary and almost zero weight to tiles on the other side, preserving sharp edges in the flow.
Training: supervising every iteration
RAFT is trained with a that supervises every iteration of the update operator, not just the final output. For a sequence of predicted flows and ground-truth flow :
Why not coarse-to-fine? The paradigm shift
Before RAFT, the dominant approach to optical flow was the coarse-to-fine pyramid. The idea was intuitive: estimate flow at low resolution first (where large motions become small), then warp the second image using that estimate, and refine at higher resolution. Repeat until you reach full resolution.
This approach has three critical weaknesses that RAFT avoids:
Error cascading. Mistakes at coarse levels propagate to finer levels and cannot be corrected. A wrong large-motion estimate at the top level will misalign the warped image for all subsequent levels.
Small object loss. Aggressive downsampling at coarse levels can shrink small objects below a single pixel, making their motion unrecoverable.
Architectural complexity. Each pyramid level needs its own feature extractor, cost volume, and estimator — multiplying parameters and training difficulty.
RAFT sidesteps all three by operating at a single resolution (1/8). The correlation pyramid is not built from multiple feature resolutions — it is built by pooling a single correlation volume. Large motions are captured by the coarse levels of the pyramid, and the recurrent updates can correct early mistakes because each iteration re-examines the correlation evidence from scratch.
Results: new state of the art
RAFT set new records across multiple optical flow benchmarks:
On Sintel (final pass), RAFT achieved an end-point error (EPE) of 2.855 pixels — a 30% reduction from the previous best (4.098 pixels). On KITTI 2015, it achieved an F1-all error of 5.10% — a 16% reduction from the previous best (6.10%).
Beyond raw accuracy, three aspects stood out:
Cross-dataset generalization. When trained only on synthetic data (FlyingChairs + FlyingThings3D) and evaluated on KITTI without fine-tuning, RAFT achieved 5.04 EPE — a 40% improvement over the best prior method. This suggests the architecture learns genuinely transferable motion representations.
Training efficiency. RAFT converges in approximately 100K iterations — roughly 10× fewer than competing methods like FlowNet2 or PWC-Net.
Compact model. The full model has only 5.3M parameters. A smaller variant (RAFT-S) with 1.0M parameters still outperforms all prior methods on Sintel, running at 20 FPS on a 1080Ti GPU.
Code: the update loop in pseudocode
Simplified to show the idea — not the real implementation.
# Extract features from both images
fmap1 = feature_encoder(image1) # [B, 256, H/8, W/8]
fmap2 = feature_encoder(image2) # shared weights
# Extract context + initial hidden state from image1
cnet, h = context_encoder(image1)
# Build 4D correlation volume (all pairs dot product)
corr_volume = einsum('bchw, bckl -> bhwkl', fmap1, fmap2)
corr_pyramid = build_pyramid(corr_volume) # 4 levels
# Initialize flow at zero
flow = torch.zeros(B, 2, H//8, W//8)
for k in range(num_iterations): # typically 12
# Look up correlation around current flow estimate
corr_features = corr_lookup(corr_pyramid, flow, radius=4)
# Encode current flow
flow_features = flow_encoder(flow)
# GRU update: correlation + flow + context -> new hidden state
h = gru(h, corr_features, flow_features, cnet)
# Predict flow delta from hidden state
delta_flow = flow_head(h)
flow = flow + delta_flow
# Convex upsample from 1/8 to full resolution
flow_up = convex_upsample(flow, mask=upsample_head(h))Legacy: the iterative refinement revolution
2015
FlowNet — deep learning enters optical flow
Dosovitskiy et al. showed that a single CNN can estimate optical flow end-to-end, trained on synthetic data. Opened the door to learned methods.
2017
FlowNet2 — stacking networks for refinement
Ilg et al. stacked multiple FlowNet architectures in a cascade, achieving competitive accuracy but at high computational cost and model size.
2018
PWC-Net — pyramids, warping, cost volumes
Sun et al. combined feature pyramids with iterative warping and cost volumes, becoming the dominant paradigm. Fast and accurate, but suffered from coarse-to-fine error cascading.
2020
RAFT — this paper (Best Paper, ECCV 2020)
Replaced coarse-to-fine warping with all-pairs correlation + recurrent iterative updates. 30% error reduction on Sintel, 16% on KITTI, 10× faster convergence, and strong cross-dataset generalization.
2021
GMA — Global Motion Aggregation
Jiang et al. added self-attention to RAFT's update operator to aggregate motion information globally, improving performance on occluded regions.
2022
FlowFormer — Transformers for optical flow
Huang et al. replaced the GRU with a Transformer-based architecture, building on RAFT's correlation volume but using attention for the update step.
2024
SEA-RAFT & Depth Anything
SEA-RAFT simplified and accelerated RAFT with a mixture-of-Laplace loss. Depth Anything extended RAFT's correlation-based ideas to monocular depth estimation, showing the architecture's influence beyond optical flow.
RAFT proved that you do not need complex multi-scale architectures to estimate dense correspondences. A single correlation volume, a lightweight recurrent updater, and iterative refinement — these three simple ideas, composed well, outperformed years of architectural engineering. The lesson extends beyond optical flow: when in doubt, prefer iterative refinement over single-shot prediction. Let the network correct its own mistakes.
CitationTeed, Deng. RAFT: Recurrent All-Pairs Field Transforms for Optical Flow. ECCV, 2020.
Terms in this paper
- Optical Flowالانسياب البصري
- Correlationالارتباط
- Iterative Algorithmالخوارزمية التكرارية
- Gated Recurrent Unit (GRU)وحدة البوابات التكرارية (GRU)
- Feature Extractionاستخلاص السمات
- Feature Mapخريطة السمات
- Feature Pyramidهرم السِّمات
- Encoderالمُرمِّز
- Convolutionالالتفاف الرقمي
- Bilinear Interpolationالاستيفاء الثنائي الخطي
- Residual Connectionالوصلة التجاوزية
- Dot Productالضرب النقطي
- Dense Predictionالتنبؤ الكثيف
- Coarse-to-Fineمن الخشن إلى الدقيق
- end-to-endمن طرف إلى طرف