Computer Vision2015intermediate12 min read

Fast R-CNN

Fast R-CNN: اكتشاف أجسام أسرع بمعمارية موحّدة

Girshick, R. — ICCV

The problem

R- (2014) was accurate but impractical: it ran a full CNN forward pass on every one of ~2,000 region proposals, stored features on disk, and trained three separate stages — CNN , SVM classifiers, and bounding-box regressors. took 84 hours and took 47 seconds per image. SPPnet improved speed by sharing convolutions, but its spatial pyramid could not back-propagate through all layers, so the convolutional was impossible. needed an architecture that was both fast and trainable.

The contribution

Fast R-CNN: a single-stage training pipeline for object detection. Pass the full image through a CNN once to get a shared , then for each , use a novel to extract a fixed-size . Two sibling output heads — a classifier and a bounding-box regressor — are trained jointly with a multi-task . All layers, including the convolutional backbone, are fine-tuned end-to-end via . Training is 9× faster than R-CNN, inference is 213× faster, and accuracy improves on PASCAL VOC 2012 from 62.4% to 68.4% mAP.

The impact

Fast R-CNN proved that object detection can be trained end-to-end in a single stage with shared convolutional features — a principle that every subsequent detector adopted. RoI pooling became the standard interface between region proposals and heads, later refined into RoIAlign by Mask R-CNN. The multi-task loss combining classification and localization became the default for two-stage detectors. Most importantly, by isolating as the remaining bottleneck, Fast R-CNN directly motivated Faster R-CNN's Region Proposal Network, which finally made the entire pipeline run on the .

Think of R-CNN as a building inspector who walks to every room, sets up a camera, takes a picture, goes back to the lab, develops the film, and only then decides what's in the room. Two thousand rooms means two thousand round trips.

Fast R-CNN is a building inspector with a drone: fly it once over the entire building, get one high-resolution aerial photo, then zoom into any room on that same photo — crop, classify, measure. One flyover, unlimited zooms.

The problem: R-CNN was accurate but painfully slow

The original R-CNN pipeline had three separate stages, each with its own training procedure:

  1. Feature extraction — run a full CNN forward pass on each of ~2,000 cropped-and-warped region proposals, then cache the features to disk.
  2. Classification — train a linear SVM for each object class on the cached features.
  3. Bounding-box — train a separate regressor to refine each proposal's coordinates.

This multi-stage design was slow (84 hours to train on VOC 2007), wasteful (hundreds of gigabytes of cached features), and fundamentally limited: because each stage was trained independently, errors in one stage could not be corrected by another. The CNN backbone could not be fine-tuned through the SVM or regressor, so detection-specific learning stopped at the feature level.

Open in Lab
Click to compare the R-CNN multi-stage pipeline with Fast R-CNN's unified single-stage approach.
The demo wakes as you arrive…

The key idea: share the expensive computation

The most expensive part of a CNN is the convolutional layers — the early layers that extract edges, textures, and patterns. R-CNN ran these layers 2,000 times per image (once per proposal). But here's the insight: all proposals come from the same image, so they should share the same convolutional features.

Fast R-CNN's solution: run the CNN on the whole image once to produce a single shared feature map. Then, for each region proposal, simply look up the corresponding region on that feature map. This is like building a single detailed blueprint of the entire building, then using it to answer questions about any room — instead of surveying each room from scratch.

SPPnet (He et al., 2014) had the same sharing idea, but its spatial pyramid pooling layer blocked flow to the convolutional layers. Fast R-CNN replaced it with a simpler RoI pooling layer that allows full backpropagation through every layer.

Open in Lab
Toggle between R-CNN (repeated CNN passes) and Fast R-CNN (single pass + feature map lookups) to see the computation difference.
The demo wakes as you arrive…

RoI Pooling: one size fits all

Region proposals come in all shapes and sizes — a person is tall and thin, a car is wide and low — but the fully connected layers at the end of the network demand a fixed-size input. RoI pooling solves this mismatch.

Here's how it works. Given a region proposal (say 21 × 14 on the feature map) and a target output size (say 7 × 7), RoI pooling divides the region into a 7 × 7 grid of sub-windows. Each sub-window is approximately 3 × 2 cells. Within each sub-window, selects the largest value. The result: a fixed 7 × 7 × C feature regardless of the original region's shape.

Think of it as a cookie cutter on the feature map: no matter the region's shape, RoI pooling stamps out the same fixed-size summary. This is a simplified single-level version of spatial pyramid pooling from SPPnet — but critically, it allows gradients to flow backward through the pooling into the convolutional layers, enabling end-to-end training.

Open in Lab
Drag the RoI box over the feature map. Watch how regions of different sizes always produce a fixed 7 × 7 output through grid-based max pooling.
The demo wakes as you arrive…

The Fast R-CNN architecture

The full pipeline flows as follows:

  1. The entire input image passes through a pre-trained CNN backbone (e.g., VGG16) up to the last , producing a shared feature map.

  2. An external (Selective Search) proposes ~2,000 candidate regions on the original image.

  3. Each proposal is projected onto the shared feature map, and RoI pooling extracts a fixed-size (e.g., 7 × 7 × 512) feature vector.

  4. The feature vector passes through two fully connected layers (fc6 and fc7, each 4096 units with and ).

  5. The output forks into two sibling heads:

    • A (K+1)-way softmax producing class probabilities (K object classes + background).
    • A bounding-box regressor producing 4 refined coordinates per class.
  6. Both heads are trained jointly with a multi-task loss that combines classification cross-entropy and bounding-box .

The genius is that steps 1–6 form a single differentiable computation graph. Gradients flow from both output heads, through the fully connected layers, through RoI pooling, and into the convolutional backbone — updating every simultaneously.

Open in Lab
Click any stage in the architecture to learn what it does.
The demo wakes as you arrive…

Multi-task loss: classify and locate in one shot

Instead of training classification and localization as separate stages (as R-CNN did), Fast R-CNN trains both tasks simultaneously using a single combined loss function. This is one of the paper's core innovations: the multi-task loss.

For each RoI, the network outputs two things: a probability distribution p=(p0,…,pK)p = (p_0, \ldots, p_K) over K+1 classes (including background), and bounding-box regression offsets tu=(txu,tyu,twu,thu)t^u = (t^u_x, t^u_y, t^u_w, t^u_h) for each class uu. The true class is uu and the true bounding-box target is vv.

The key design choice is the indicator [u≥1][u \geq 1]: background proposals (class 0) have no ground-truth box, so their regression loss is zeroed out. Only foreground objects contribute to the localization loss. The λ\lambda controls the balance between the two tasks — the paper uses λ=1\lambda = 1, giving equal weight to classification and localization.

L(p,u,tu,v)=Lcls(p,u)+λ[u≥1]⋅Lloc(tu,v)L(p, u, t^u, v) = L_{cls}(p, u) + \lambda [u \geq 1] \cdot L_{loc}(t^u, v)
Fast R-CNN multi-task loss — L_cls = cross-entropy for classification · L_loc = smooth L1 for box regression · [u ≥ 1] zeros out localization loss for background proposals · λ balances the two tasks (set to 1)

Smooth L1 loss: robustness without explosions

For bounding-box regression, Fast R-CNN introduced the smooth L1 loss, which blends the best properties of L1 and L2 losses:

  • When the error is small (|x| < 1), it behaves like L2 (squared error): smooth, differentiable, and converges precisely.
  • When the error is large (|x| ≥ 1), it behaves like L1 (absolute error): grows linearly rather than quadratically, preventing outliers from producing explosive gradients.

This is critical for training stability. Bounding-box targets can include large offsets (when a proposal is far from the ground truth), and L2 loss would turn those into enormous gradients that destabilize training. Smooth L1 caps the gradient magnitude at 1 for large errors, making training robust to outliers without requiring aggressive .

smoothL1(x)={0.5x2if ∣x∣<1∣x∣−0.5otherwisesmooth_{L1}(x) = \begin{cases} 0.5x^2 & \text{if } |x| < 1 \\ |x| - 0.5 & \text{otherwise} \end{cases}
Smooth L1 loss — the best of both worlds — Small errors get L2's smooth gradients for precise convergence · Large errors get L1's linear penalty to avoid gradient explosions · The transition at |x| = 1 is continuous and differentiable
Open in Lab
Drag the slider to see how smooth L1 transitions between L2 (small errors) and L1 (large errors). Compare with pure L1 and L2 losses.
The demo wakes as you arrive…

Training: hierarchical sampling and shared features

A subtle but important training innovation is hierarchical sampling of mini-batches. Instead of randomly sampling RoIs from across many images (which would require computing a separate feature map for each image in the batch), Fast R-CNN samples 2 images per and 64 RoIs from each image, for a total of 128 RoIs.

Of these 128 RoIs, 25% are positive ( ≥ 0.5 with a ground-truth box) and 75% are negative (IoU in [0.1, 0.5)). This 1:3 ratio ensures the network sees enough negative examples (background) without being overwhelmed by them.

Because all 64 RoIs from the same image share the same convolutional feature map, the expensive forward pass happens only twice per mini-batch (once per image) rather than 128 times. This makes training both faster and memory-efficient.

The same idea in code

Fast R-CNN forward pass — simplifiedpython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn as nn
from torchvision.ops import roi_pool

class FastRCNN(nn.Module):
    def __init__(self, backbone, num_classes=21):
        """
        backbone:    pretrained CNN (e.g. VGG16 up to conv5)
        num_classes: K object classes + 1 background
        """
        super().__init__()
        self.backbone = backbone            # shared conv layers
        self.roi_pool = lambda feat, rois: roi_pool(
            feat, rois, output_size=(7, 7), spatial_scale=1/16
        )
        self.fc6 = nn.Linear(512 * 7 * 7, 4096)
        self.fc7 = nn.Linear(4096, 4096)
        self.cls_head = nn.Linear(4096, num_classes)      # class scores
        self.box_head = nn.Linear(4096, num_classes * 4)   # box offsets per class

    def forward(self, image, rois):
        # Step 1: one CNN pass on the whole image → shared feature map
        feat_map = self.backbone(image)

        # Step 2: RoI pooling → fixed-size features for each proposal
        pooled = self.roi_pool(feat_map, rois)   # (N_rois, 512, 7, 7)

        # Step 3: flatten → FC layers
        x = pooled.flatten(1)                    # (N_rois, 512*7*7)
        x = torch.relu(self.fc6(x))
        x = torch.relu(self.fc7(x))

        # Step 4: two sibling heads
        cls_scores = self.cls_head(x)            # (N_rois, num_classes)
        box_offsets = self.box_head(x)            # (N_rois, num_classes * 4)
        return cls_scores, box_offsets

Results and speed

On PASCAL VOC 2007, Fast R-CNN with VGG16 achieves 70.0% mAP — surpassing R-CNN's 66.0%. On VOC 2012 it reaches 68.4% mAP vs. R-CNN's 62.4%.

But the real story is speed. Training time dropped from 84 hours (R-CNN) to 9.5 hours — a 9× speedup. Test-time inference dropped from 47 seconds per image to 0.32 seconds (excluding proposal generation) — a 146× speedup. Including Selective Search (~2 seconds), total inference is ~2.3 seconds, which is still a 213× speedup over R-CNN's full pipeline.

The paper also compares against SPPnet, showing that Fast R-CNN trains 3× faster and tests 10× faster while achieving higher accuracy, thanks to the ability to fine-tune all layers.

Open in Lab
Training and testing speed comparison across R-CNN, SPPnet, and Fast R-CNN.
The demo wakes as you arrive…

Fast R-CNN solved the training and feature-sharing problems, but one piece of the pipeline remained outside the : Selective Search. This CPU-based algorithm takes about 2 seconds per image — longer than Fast R-CNN's entire neural network inference (0.32s). It accounts for 86% of total test time.

This observation pointed directly to the next breakthrough: what if we replace Selective Search with a small neural network that proposes regions on the GPU? That idea became the Region Proposal Network (RPN) in Faster R-CNN, published just one month later by Ren, He, Sun, and Girshick himself.

Why it mattered

Fast R-CNN established three principles that shaped every subsequent object detector:

Shared convolutional features. Never repeat the same computation. Run the CNN once, reuse the feature map for all proposals. This idea lives on in every modern detector — from Faster R-CNN to YOLO to DETR.

. Classification and localization benefit from being trained together. Joint losses became the standard for all detection architectures.

End-to-end trainability. Every component should be differentiable so gradients can optimize the entire system. When something blocks gradient flow (like SPPnet's pooling or Selective Search), it becomes the target for the next innovation.

  1. 2014

    R-CNN

    First to use CNNs for object detection — extracted features from each region proposal independently. Accurate but extremely slow (47s per image).

  2. 2014

    SPPnet

    Shared convolutional computation across proposals using spatial pyramid pooling, but could not back-propagate through the pooling layer.

  3. 2015

    Fast R-CNN

    Unified training with RoI pooling and multi-task loss. All layers trainable end-to-end. 9× faster training, 213× faster inference than R-CNN.

  4. 2015

    Faster R-CNN

    Replaced Selective Search with a Region Proposal Network (RPN), making the entire pipeline run on the GPU. Near real-time detection at 5 fps.

  5. 2017

    Mask R-CNN

    Extended Faster R-CNN with a mask branch for instance segmentation. Replaced RoI pooling with RoIAlign for sub-pixel accuracy.

CitationGirshick, R.. Fast R-CNN. ICCV, 2015.

Terms in this paper