Computer Vision2017intermediate10 min read

Feature Pyramid Networks for Object Detection

شبكات هرم السِّمات لكشف الأجسام

Lin, T.-Y. · Dollár, P. · Girshick, R. · He, K. · Hariharan, B. · Belongie, S. — CVPR

The problem

Object detectors like Faster R- run on a single from the last convolutional layer. That layer has rich semantic meaning — it knows "this is a cat" — but its spatial is 32× smaller than the input. Small objects vanish. The classic fix, an image pyramid (run the detector on resized copies of the image), works but multiplies compute and memory. Deep learning detectors abandoned pyramids for speed, and paid for it with poor small-object performance.

The contribution

FPN reuses the pyramid that any ConvNet already builds internally. A bottom-up pass (the normal forward pass) produces feature maps at 1/4, 1/8, 1/16, and 1/32 of the input resolution — each richer in semantics but coarser in space. A top-down pathway upsamples the semantic-rich top and merges it with the detail-rich bottom via lateral connections (1×1 conv + element-wise addition). The result: every pyramid level is both semantically strong and spatially precise. Plugged into Faster R-CNN, FPN achieved state-of-the-art on COCO without bells and whistles.

The impact

FPN became the default neck in virtually every modern detector: Mask R-CNN, RetinaNet, YOLOv3+, FCOS, and the Swin Transformer detector all build on it. It showed that feature fusion is cheap, effective, and -agnostic — an idea so successful it's now taken for granted. Every COCO leaderboard entry since 2017 uses some form of feature pyramid.

Imagine a security guard watching a parking lot through a single camera mounted on a tall pole. Cars look fine, but license plates are unreadable blurs — too small to resolve from that height.

The naive fix: install cameras at multiple heights. Now you can read plates and see cars, but you need 4× the hardware and 4× the video processing.

FPN's trick: keep the one tall camera but run a cable down the pole. At each height, a small screen shows the zoomed-out overview from above, annotated with the fine details visible at that height. Every screen gets the big picture and local detail — one camera, one pass, all scales covered.

The problem: one scale does not fit all

A street scene may contain a pedestrian 300 pixels tall and a traffic sign 20 pixels wide — a 15× difference in scale. A detector that operates on a single feature faces an impossible trade-off:

  • High-resolution maps (early layers) see small objects but lack semantic understanding — they can detect edges but not identify them as "stop sign."
  • Low-resolution maps (deep layers) understand categories but smear spatial detail — they know "person" but can't localize one who is 20 pixels tall.

Before FPN, the field had two flawed solutions. Image pyramids (running the backbone on resized copies of the image) give multi-scale features but at huge compute cost. Single-shot detectors like SSD predicted from multiple layers but never enriched the early layers with deep semantics.

Open in Lab
Toggle between strategies to see how each handles multi-scale detection. FPN gives every level both detail and semantics.
The demo wakes as you arrive…

Bottom-up pathway: the pyramid you already have

Every ConvNet naturally builds a pyramid during its forward pass. In ResNet, for example, the output of each residual stage (conv2, conv3, conv4, conv5) has half the spatial resolution of the previous one but double the channels. FPN calls these outputs (C₂, C₃, C₄, C₅), with spatial strides (4, 8, 16, 32) relative to the input.

This is the bottom-up pathway. Think of it as climbing a mountain: at each rest stop you see more of the landscape (more semantic context), but the details below your feet get smaller (lower spatial resolution). C₅ is the summit — the deepest semantics, the lowest resolution.

Top-down pathway & lateral connections: the FPN recipe

The core innovation is a three-step merge at each pyramid level, working from top to bottom:

Step 1 — Upsample. Take the feature map one level above (coarser, richer semantics) and upsample it 2× using nearest-neighbor interpolation, so it matches the spatial size of the level below.

Step 2 — Reduce channels. Apply a 1×1 to the corresponding bottom-up map (Cₙ) to reduce its channels to a fixed width d = 256. This is the lateral connection — it acts as a translator that makes the bottom-up features speak the same "language" ( dimension) as the top-down flow.

Step 3 — Merge. Element-wise add the upsampled top-down map and the channel-reduced bottom-up map. Then apply a 3×3 convolution to smooth aliasing artifacts from upsampling. The result is Pₙ — a pyramid level that inherits deep semantics from above and spatial precision from its own resolution.

This process starts at P₅ (which is just C₅ after a 1×1 conv) and cascades down to P₂. The final pyramid (P₂, P₃, P₄, P₅) has the same spatial sizes as (C₂, C₃, C₄, C₅) but every level now carries strong semantic features.

Open in Lab
Step through the three operations that build each FPN level. The lateral connection is the key innovation.
The demo wakes as you arrive…
Pn=Conv3×3 ⁣(  Upsample(Pn+1)+Conv1×1(Cn)  )P_n = \text{Conv}_{3\times3}\!\Big(\;\text{Upsample}(P_{n+1}) + \text{Conv}_{1\times1}(C_n)\;\Big)
The FPN merge equation — one line that builds every pyramid level — Upsample the richer level above · add the matching bottom-up map (after a 1×1 channel reducer) · smooth with a 3×3 conv. Repeat from top to bottom.

The full FPN architecture

Putting it all together, FPN is a neck — it sits between the backbone (ResNet, VGG, etc.) and the detection head (RPN, Fast R-CNN classifiers). It takes the raw backbone features and transforms them into a multi-scale feature pyramid where every level is semantically enriched.

The architecture is backbone-agnostic: swap ResNet-50 for ResNet-101 or any other ConvNet, and FPN still works. This modularity is why it became the standard building block.

Open in Lab
Explore the full FPN architecture. Click any level to see its role and resolution. Toggle backbone stages on and off.
The demo wakes as you arrive…

FPN with RPN: proposals at every scale

In the original Faster R-CNN, the (RPN) slides over a single feature map and uses anchor boxes of multiple sizes and aspect ratios to propose regions. But detecting tiny objects from a coarse 1/32-scale map is like reading fine print from across the room.

FPN attaches the same RPN head (a small 3×3 conv followed by classification and regression branches) to each pyramid level independently. The key simplification: since each level already handles a different spatial scale, anchors only need to vary in (1:1, 1:2, 2:1), not in size. P₂ handles the smallest objects (32² anchor area), P₃ handles 64², P₄ handles 128², and P₅ handles 256². Scale variation is encoded in the pyramid itself, not in the anchors.

This single change raised Average by 8 points over the baseline Faster R-CNN on COCO — an enormous improvement from an architectural insight rather than a training trick.

k=⌊k0+log⁡2(wh/224)⌋k = \lfloor k_0 + \log_2(\sqrt{wh}/224) \rfloor
RoI-to-level assignment for Fast R-CNN on FPN — Each region of interest (RoI) of width w and height h is assigned to pyramid level Pₖ. Small RoIs go to fine levels (P₂), large RoIs to coarse levels (P₅). k₀ = 4 means a 224×224 RoI maps to P₄ — the canonical ImageNet scale.

Anchor assignment across pyramid levels

One of FPN's elegant simplifications is how it distributes detection responsibility across pyramid levels. Instead of overloading one feature map with anchors of every size (as Faster R-CNN did), FPN assigns each scale to the pyramid level whose naturally matches it. Small objects live on the high-resolution P₂, and large objects live on the low-resolution P₅. The detector head is shared across all levels — the same weights, applied to scale-appropriate features.

Open in Lab
Click objects of different sizes to see which pyramid level detects them. Small → P₂, large → P₅.
The demo wakes as you arrive…

FPN in code

Minimal FPN top-down pathwaypython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn as nn
import torch.nn.functional as F

class FPN(nn.Module):
    """Minimal Feature Pyramid Network on top of a ResNet backbone."""

    def __init__(self, in_channels_list, out_channels=256):
        super().__init__()
        # Lateral connections: 1x1 conv to unify channel dims
        self.laterals = nn.ModuleList([
            nn.Conv2d(in_ch, out_channels, 1)
            for in_ch in in_channels_list        # e.g. [256, 512, 1024, 2048]
        ])
        # Smoothing convolutions: reduce aliasing after addition
        self.smooths = nn.ModuleList([
            nn.Conv2d(out_channels, out_channels, 3, padding=1)
            for _ in in_channels_list
        ])

    def forward(self, features):
        """features: [C2, C3, C4, C5] from ResNet."""
        # 1. Apply lateral 1x1 convs
        laterals = [l(f) for l, f in zip(self.laterals, features)]

        # 2. Top-down pathway: start from deepest, add to each level
        for i in range(len(laterals) - 1, 0, -1):
            upsampled = F.interpolate(
                laterals[i], scale_factor=2, mode='nearest'
            )
            laterals[i - 1] = laterals[i - 1] + upsampled

        # 3. Smooth each merged map
        pyramid = [s(l) for s, l in zip(self.smooths, laterals)]
        return pyramid  # [P2, P3, P4, P5]

Experiments: what each component contributes

The authors performed a meticulous ablation study on COCO to isolate the value of each FPN component:

  • Lateral connections only (no top-down): AR₁₀₀ = 44.9. The channel unification helps but there's no semantic enrichment flowing downward.
  • Top-down only (no laterals): AR₁₀₀ = 44.1. Semantic info flows down but without the spatial anchoring from the bottom-up maps.
  • Full FPN (top-down + laterals): AR₁₀₀ = 56.3 — an 8-point jump over the baseline single-scale RPN (AR₁₀₀ = 48.3 on C₅ alone).

For with Fast R-CNN on top, FPN improved COCO AP by 2.3 points and PASCAL-style AP by 3.8 points over a strong single-scale baseline. The final system achieved 59.1% AP₅₀ on COCO test-dev, surpassing all single-model entries including COCO 2016 challenge winners.

Open in Lab
Compare how removing each FPN component affects Average Recall.
The demo wakes as you arrive…

Beyond boxes: FPN for segmentation

The authors also showed FPN can generate proposals. By attaching a small fully convolutional head (a 5×5 MLP predicting 14×14 masks) to each pyramid level, FPN produced mask proposals that outperformed DeepMask and SharpMask. This capability directly paved the way for Mask R-CNN, which replaced the mask proposal head with a proper mask prediction branch and achieved the first practical instance segmentation system.

Design principles that made FPN last

Impact: FPN's descendants

  1. 2017

    FPN published (CVPR)

    Lin et al. introduce FPN. State-of-the-art on COCO detection with a simple top-down + lateral architecture on top of ResNet.

  2. 2017

    RetinaNet

    Lin et al. use FPN as the backbone neck in a one-stage detector, paired with focal loss to handle class imbalance. Closed the gap between one-stage and two-stage detectors.

  3. 2017

    Mask R-CNN

    He et al. build on FPN + Faster R-CNN, adding a mask prediction branch. Unified object detection and instance segmentation in one model.

  4. 2018

    YOLOv3 adopts FPN-like neck

    Redmon & Farhadi integrate multi-scale prediction with FPN-style feature fusion into YOLO, dramatically improving small object detection.

  5. 2019

    PANet, NAS-FPN, BiFPN

    Researchers began stacking, searching, and bidirectionalizing feature pyramids — all descendants of FPN's core insight.

  6. 2021

    Swin Transformer detector

    Liu et al. show that even with a Transformer backbone replacing CNNs, the FPN neck remains essential for multi-scale detection.

FPN's legacy is not a single detector but an architectural pattern: enrich multi-scale features by flowing information both up and down. Every time you see a modern detector with a "neck" — PANet, BiFPN, NAS-FPN — you are looking at FPN's intellectual offspring.

CitationLin, Dollár, Girshick, He, Hariharan, Belongie. Feature Pyramid Networks for Object Detection. CVPR, 2017.

Terms in this paper