Computer Vision2015intermediate13 min read

Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks

Faster R-CNN: نحو رصد الأجسام آنيًّا بشبكات المقترحات الإقليمية

Ren, S. · He, K. · Girshick, R. · Sun, J. — NeurIPS

The problem

By 2015, Fast R- had made the detection network itself fast, but the region proposal step — typically — ran on the CPU and took nearly 2 seconds per image, dwarfing the 0.2-second detection. Region proposals had become the computational bottleneck of the entire pipeline.

The contribution

(RPN): a fully convolutional network that slides over the shared and, at each position, predicts whether an object is present () and refines the coordinates of k anchor boxes of different scales and aspect ratios. The RPN shares convolutional features with the downstream Fast R-CNN detector, making proposals nearly cost-free (~10 ms per image). An strategy unifies both networks into one, achieving 73.2% mAP on PASCAL VOC 2007 at 5 fps with VGG-16 — and 78.8% mAP on VOC 2012.

The impact

Faster R-CNN established the modern two-stage detection paradigm: propose then classify. Its anchor-box mechanism became the default for nearly every detector — one-stage (YOLO, SSD, RetinaNet) and two-stage (Mask R-CNN, Cascade R-CNN) alike. Feature Pyramid Networks extended its shared- idea, and DETR later replaced anchors with learned queries. Faster R-CNN remains one of the most cited papers in computer vision history.

Imagine you're an airport security officer scanning luggage on a screen. Before Faster R-CNN, one colleague would slowly circle every suspicious shape by hand (Selective Search), then hand you each circled region to classify. You were fast, but your colleague was painfully slow.

Faster R-CNN gives you both the same set of X-ray glasses (shared convolutional features). Now your colleague, looking at the same X-ray image you see, instantly points to candidate regions with pre-set frames of different sizes (anchor boxes). You then zoom into each frame and decide: "this is a laptop," "this is a water bottle," "nothing here."

Because you share the same X-ray view, your colleague's pointing costs almost nothing — detection becomes a single, unified system.

The bottleneck: proposals were slower than detection

The evolution from R-CNN to Fast R-CNN solved the repeated-computation problem: instead of running a CNN on every proposed region independently, Fast R-CNN runs the CNN once on the whole image and pools features from the shared feature map for each region. Detection became fast — but the proposals still came from Selective Search, a handcrafted CPU algorithm that tested ~2000 region candidates per image in roughly 2 seconds.

The GPU detection network sat idle, waiting for the CPU to finish generating proposals. It was as if you had built a high-speed highway but were stuck at a manual toll booth at the entrance.

Open in Lab
Compare the old pipeline (Selective Search on CPU) with the new one (RPN on GPU sharing features). Watch the bottleneck disappear.
The demo wakes as you arrive…

The idea: let the network propose its own regions

The key observation is beautifully simple: the convolutional feature map that Fast R-CNN already computes for detection contains rich spatial information about where objects are. Why not use that same feature map to also propose regions?

The Region Proposal Network (RPN) is a small network that slides over this shared feature map. At each spatial location, it does two things simultaneously:

  • Predicts an objectness score — is there any object here, regardless of class?
  • Refines the coordinates of k anchor boxes — predefined reference rectangles of different sizes and shapes centered at that location.

Because the RPN operates on features the backbone has already computed, generating proposals costs almost nothing extra — roughly 10 milliseconds per image.

Anchor boxes: pre-set frames for multi-scale detection

How does the RPN handle objects of vastly different sizes — a car versus a person versus a traffic light — without using image pyramids or pyramids?

The answer is anchor boxes. At each sliding-window position on the feature map, the network places k reference rectangles (anchors) of different scales and aspect ratios. The original paper uses 3 scales (128², 256², 512²) × 3 aspect ratios (1:1, 1:2, 2:1) = 9 anchors per position.

Think of it as placing 9 transparent frames of different sizes and shapes at every point on the feature map — like a photographer trying different viewfinders. The network then learns to say "yes, there's an object in this frame" or "no, skip this one," and for the positive frames, it learns how to nudge the frame boundaries to better fit the actual object.

This is far more efficient than scanning the image at multiple resolutions. The anchors act as built-in multi-scale priors, and the network only needs to learn refinements from these starting positions rather than raw coordinates from scratch.

Open in Lab
Hover over the feature map grid to see the 9 anchor boxes at each position. Toggle scales and aspect ratios to understand coverage.
The demo wakes as you arrive…

Inside the RPN: a small network on top of the backbone

The RPN architecture is surprisingly compact. On top of the shared convolutional feature map (e.g. from the last of VGG-16 or ResNet), two things happen:

Step 1: Sliding window. A 3×3 convolutional layer slides over the feature map, producing a 256-dimensional (or 512-dimensional) at each spatial position. This is the "looking through every viewfinder" step.

Step 2: Twin heads. Two sibling 1×1 convolutional layers branch from this intermediate layer:

  • The classification head outputs 2k scores (object vs. not-object for each of k anchors).
  • The head outputs 4k coordinates (dx, dy, dw, dh refinements for each anchor).

For k = 9 anchors, each position produces 18 classification scores and 36 bounding-box refinements. On a feature map of size W×H, the total number of anchors is W × H × 9 — typically around 20,000 for a 1000×600 input image.

Open in Lab
Click each layer of the RPN to understand its role — from shared backbone to twin output heads.
The demo wakes as you arrive…

Training the RPN: positive, negative, and the IoU threshold

With ~20,000 anchors per image, most will be background. To train the RPN, each anchor is labeled as either positive (contains an object) or negative (background):

  • An anchor is positive if it has the highest with any ground-truth box, or if its IoU with any ground-truth box exceeds 0.7.
  • An anchor is negative if its IoU with all ground-truth boxes is below 0.3.
  • Anchors between 0.3 and 0.7 IoU are ignored during training — they neither help nor hurt.

From these labeled anchors, a of 256 is sampled per image (up to 128 positive, the rest negative). This balanced sampling prevents the overwhelming majority of negatives from dominating the learning signal.

Open in Lab
Drag the predicted box to change overlap with the ground truth. Watch the IoU score and anchor label update.
The demo wakes as you arrive…

The multi-task loss: classification + regression

The RPN is trained with a multi-task that combines two objectives. First, understand the intuition: we want the network to learn both where objects are (by labeling anchors as object/background) and how to adjust anchor boundaries to fit objects tightly. These are a classification problem and a regression problem respectively, combined into one loss.

L({pi},{ti})=1Ncls∑iLcls(pi,pi∗)+λ1Nreg∑ipi∗Lreg(ti,ti∗)L(\{p_i\}, \{t_i\}) = \frac{1}{N_{cls}} \sum_i L_{cls}(p_i, p_i^*) + \lambda \frac{1}{N_{reg}} \sum_i p_i^* L_{reg}(t_i, t_i^*)
RPN multi-task loss — The Region Proposal Network is trained to perform two tasks at the same time. First, it learns to determine whether an anchor is likely to contain an object or belong to the background. Second, it learns how to refine the anchor's position and size so that it better matches the object. The total loss combines these classification and localization objectives, with a balancing factor ensuring that both tasks contribute appropriately during training. Localization errors are computed only for anchors that correspond to real objects.

The term pi∗p_i^* in the regression loss acts as a gate: regression loss is computed only for positive anchors. There is no point in refining a box if there is no object inside it. The two normalizations (NclsN_{cls} for mini-batch size and NregN_{reg} for the number of anchor positions) ensure the two tasks contribute equally to learning.

The full Faster R-CNN pipeline

Here is the complete detection flow, from raw image to final bounding boxes:

Stage 1 — Shared backbone. The input image passes through a deep CNN (VGG-16 or ResNet). The output is a feature map — a compressed, semantically rich representation of the image.

Stage 2 — RPN proposals. The RPN slides over this feature map, producing ~20,000 anchors. After filtering by objectness score and applying (keeping only the highest-scoring, non-overlapping proposals), roughly 300 top proposals remain.

Stage 3 — . Each proposal is mapped back to the shared feature map and pooled into a fixed-size feature vector (e.g. 7×7) using RoI Pooling. This allows the downstream classifier to accept proposals of any size.

Stage 4 — Fast R-CNN head. The pooled features pass through fully connected layers that produce two outputs per proposal: a class probability distribution ( over C+1 classes including background) and 4 refined bounding-box coordinates per class.

Stage 5 — Final NMS. Non-maximum suppression removes duplicate detections, yielding the final set of bounding boxes with class labels and confidence scores.

Open in Lab
Step through the full Faster R-CNN pipeline — from raw pixels to final detections.
The demo wakes as you arrive…

Four-step alternating training

Training a unified network where proposal and detection share features is not trivial. The paper proposes a four-step alternating training strategy:

  • Step 1: Train the RPN alone, initialized from an ImageNet-pretrained backbone.
  • Step 2: Train a separate Fast R-CNN detector using the proposals from Step 1's RPN. This detector is also initialized from ImageNet but does not yet share features with the RPN.
  • Step 3: Re-initialize the RPN with the detector's shared convolutional layers (from Step 2), but freeze these shared layers and fine-tune only the RPN-specific layers. Now both networks share the same backbone.
  • Step 4: Keeping the shared layers frozen, fine-tune the Fast R-CNN-specific layers.

After these four steps, both networks share the same convolutional features, forming a single unified network. In practice, later work showed that approximate joint training (backpropagating through both networks simultaneously) also works well.

Translation-invariant anchors

A critical property of the anchor mechanism is . The same set of anchors and the same prediction function are applied at every position on the feature map. If an object shifts in the image, the same anchor at a different position will detect it.

This is analogous to how a convolutional filter detects the same edge everywhere — the RPN extends this principle to object proposals. Translation invariance means the size does not grow with image size, and the network generalizes to objects at any location.

In contrast, earlier methods like MultiBox used position-specific predictions — they needed separate parameters for each spatial location, couldn't generalize across positions, and required far more parameters.

The RPN in code

Region Proposal Network — core logicpython

Simplified to show the idea — not the real implementation.

import numpy as np

def generate_anchors(feat_h, feat_w, stride=16):
    """Place 9 anchors (3 scales × 3 ratios) at each feature map cell."""
    scales = [128, 256, 512]
    ratios = [0.5, 1.0, 2.0]  # h/w ratios
    anchors = []
    for y in range(feat_h):
        for x in range(feat_w):
            cx, cy = x * stride + stride // 2, y * stride + stride // 2
            for s in scales:
                for r in ratios:
                    w = s * np.sqrt(r)
                    h = s / np.sqrt(r)
                    anchors.append([cx - w/2, cy - h/2, cx + w/2, cy + h/2])
    return np.array(anchors)  # (feat_h * feat_w * 9, 4)

def rpn_forward(feature_map, W_conv, W_cls, W_reg, k=9):
    """Simplified RPN forward pass."""
    H, W, C = feature_map.shape
    # Step 1: 3×3 sliding window → intermediate features
    intermediate = conv3x3(feature_map, W_conv)  # (H, W, 256)
    # Step 2a: classification head → objectness scores
    cls_scores = conv1x1(intermediate, W_cls)     # (H, W, 2*k)
    # Step 2b: regression head → box refinements
    reg_deltas = conv1x1(intermediate, W_reg)     # (H, W, 4*k)
    return cls_scores, reg_deltas

# After RPN: filter by score, apply deltas to anchors, then NMS
# → ~300 proposals fed into the Fast R-CNN detection head.

Results: speed and accuracy

On PASCAL VOC 2007, Faster R-CNN with VGG-16 achieved 73.2% mAP at 5 fps — a significant speed-up over the Selective Search baseline while matching or exceeding its accuracy. With the deeper ResNet-101 backbone, accuracy reached 78.8% mAP on VOC 2012.

The RPN itself generates proposals with higher at fewer proposals than Selective Search: 300 RPN proposals achieve the same recall as 2,000 Selective Search regions. This is because the RPN proposals are learned and tuned for the specific task, while Selective Search is a generic, hand-engineered algorithm.

On the MS COCO benchmark, Faster R-CNN set new state-of-the-art results, demonstrating that the framework scales to larger, more challenging datasets with more object categories.

Open in Lab
Compare recall at different proposal counts — 300 RPN proposals match 2000 Selective Search regions.
The demo wakes as you arrive…

Non-maximum suppression: removing duplicate detections

Both the RPN and the final detection stage produce overlapping predictions — many neighboring anchors may fire for the same object. Non-maximum suppression (NMS) cleans this up:

  • Sort all proposals by their objectness score (or class confidence).
  • Take the highest-scoring proposal.
  • Remove any remaining proposal whose IoU with the selected one exceeds a threshold (typically 0.7 for RPN, 0.3 for final detection).
  • Repeat until no proposals remain.

Think of NMS as a "winner-takes-all" filter: when multiple overlapping frames claim the same object, only the most confident one survives. This is essential because without it, the same car might be detected dozens of times.

Why it mattered

  1. 2014

    R-CNN — region-based CNN detection

    Girshick et al. ran a CNN on each of ~2000 Selective Search proposals independently. Accurate but very slow — 47 seconds per image on a GPU.

  2. 2015

    Fast R-CNN — shared features, one-pass detection

    Girshick ran the CNN once on the full image and pooled features per region from the shared feature map. Detection sped up 25×, but Selective Search remained the bottleneck.

  3. 2015

    Faster R-CNN — learned proposals via RPN

    Ren et al. replaced Selective Search with a Region Proposal Network sharing features with the detector. Proposals became nearly free. 73.2% mAP on VOC 2007 at 5 fps.

  4. 2017

    Feature Pyramid Networks (FPN)

    Lin et al. built a multi-scale feature pyramid on top of the Faster R-CNN backbone, dramatically improving detection of small objects.

  5. 2017

    Mask R-CNN — extending to instance segmentation

    He et al. added a mask prediction branch to Faster R-CNN, enabling pixel-level segmentation alongside detection. Used RoI Align instead of RoI Pooling.

  6. 2020

    DETR — anchors replaced by learned queries

    Carion et al. applied Transformers to detection, using learned object queries instead of anchors and Hungarian matching instead of NMS. A fundamentally new paradigm built in response to the anchor-based tradition Faster R-CNN established.

From R-CNN's 47 seconds per image to Faster R-CNN's 0.2 seconds — a 200× speedup in 18 months. The key was moving proposals inside the network and sharing computation. This principle — unify, share, learn — is the thread connecting every advance in this timeline.

CitationRen, He, Girshick, Sun. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. NeurIPS, 2015.

Terms in this paper