Computer Vision2015intermediate10 min read

Fully Convolutional Networks for Semantic Segmentation

الشبكات الالتفافية الكاملة للتجزئة الدلالية

Long, J. · Shelhamer, E. · Darrell, T. — CVPR

The problem

networks like VGG and AlexNet condense an entire image into a single label — "cat" or "dog." But many vision tasks need a label for every pixel: self-driving cars must know which pixels are road, pedestrians, or sky. Before this paper, pixel-wise labeling relied on patchwise classification or hand-crafted pipelines that were slow and could not be trained end-to-end.

The contribution

: adapt any classification network (AlexNet, VGG, GoogLeNet) into a that outputs a map the same size as the input. Replace fully-connected layers with 1×1 convolutions so the network accepts any input size, then upsample the coarse output using learned transposed convolutions. Skip connections fuse fine-grained early-layer features with deep semantic features, refining boundaries from 32-pixel (FCN-32s) to 8-pixel stride (FCN-8s).

The impact

FCN established the paradigm for dense prediction that every modern segmentation model follows. U-Net, DeepLab, Mask R-, and SegNet are all direct descendants. It proved that classification backbones can be repurposed for pixel-level tasks, making the standard starting point for segmentation. The paper's 62.2% mean on PASCAL VOC 2012 exceeded the prior state of the art by a wide margin and ran 286× faster than patchwise approaches.

A classification network is like a funnel: it pours an image in at the top and squeezes out a single drop — the class label — at the bottom. All spatial information is crushed.

FCN turns the funnel into a megaphone: after squeezing, it expands the signal back to the original image size, so every pixel gets its own label. And by connecting the wide mouth of the funnel to the expanding bell of the megaphone — skip connections — it recovers the fine details that the squeezing lost.

The problem: classifiers throw away location

By 2014, deep convolutional networks like AlexNet and VGG had conquered image classification. But classification answers what is in the image, not where. requires both: every pixel must be assigned a class label.

The standard approach at the time was patchwise classification: crop a small patch around each pixel, classify it with a CNN, repeat for every pixel. This was painfully slow and threw away context — a pixel's patch didn't know what the neighboring patches saw. Other methods used hand-crafted superpixels or CRF post-processing bolted onto CNN features, creating fragmented pipelines that could not be trained end-to-end.

Open in Lab
Classification gives one label for the whole image. Segmentation gives a label for every pixel.
The demo wakes as you arrive…

The key insight: fully convolutional = any size in, same size out

Classification networks end with fully-connected layers that demand a fixed input size (e.g. 224×224) and output a fixed-length vector. These layers destroy spatial layout — they treat every neuron as equally connected to every output, regardless of position.

FCN's insight: replace every fully-connected layer with a 1×1 . A 1×1 convolution applies a learned linear combination at each spatial position independently. The output is no longer a vector but a spatial map — a heat map of class scores at every location. Because convolutions don't care about input size, the network now accepts images of any dimension and produces correspondingly-sized output.

Think of it this way: a fully-connected layer is like reading an entire essay and writing one summary sentence. A 1×1 convolution is like a proofreader who walks through the essay paragraph by paragraph, writing a margin note at each one. The same weights (the same critical eye), applied at every position.

Open in Lab
Left: fully-connected layer flattens spatial dimensions. Right: 1×1 convolution preserves the spatial grid.
The demo wakes as you arrive…

The decoder: expanding back to full resolution

After the convolutional layers of VGG or AlexNet, the is 32× smaller than the input (e.g. 7×7 from a 224×224 image) due to repeated . We need to expand this coarse map back to the original resolution — pixel by pixel.

FCN uses transposed convolution (also called or upconvolution). Think of a normal convolution as many-to-one: many input pixels contribute to one output pixel. A transposed convolution reverses this: one input pixel contributes to many output pixels. It inserts zeros between input values, then applies a learned filter — effectively spreading and blending each coarse prediction across a larger spatial area.

The key advantage: these weights are learned, not hand-designed. The network discovers how to interpolate between coarse predictions in a way that best recovers the original spatial detail.

y=WTxy = W^T x
Transposed convolution — upsampling by learned expansion — Where a normal convolution multiplies input patches by W to produce smaller output, the transposed version multiplies by Wᵀ to expand back. The same weights, reversed direction.
Open in Lab
Watch how a 2×2 coarse map expands to 4×4 through transposed convolution. Each input cell "spreads" its value through the learned filter.
The demo wakes as you arrive…

Skip connections: recovering fine detail

Upsampling the final layer alone (FCN-32s) produces segmentations that are blobby and imprecise — edges are smeared, small objects are lost. The reason: 32× downsampling discards too much spatial detail. But earlier layers in the network still have that detail — pool3 has 8× stride, pool4 has 16× stride.

The paper's second key idea: skip connections that fuse predictions from the deep, coarse layer with features from shallower, finer layers:

  • FCN-32s — upsample the final layer 32× directly. Fast but coarse.

  • FCN-16s — combine the 2× upsampled final prediction with pool4 features, then upsample 16×. Edges become noticeably sharper.

  • FCN-8s — add pool3 features too. Three streams fused. Boundaries become crisp enough for practical use.

Think of it as a conference call: the deep layers contribute what objects are present (a zoomed-out view), while the shallow layers contribute where exactly the boundaries fall (a zoomed-in view). Fusing both gives segmentations that are simultaneously accurate in identity and precise in location.

Open in Lab
Toggle between FCN-32s, FCN-16s, and FCN-8s to see how skip connections progressively sharpen segmentation boundaries.
The demo wakes as you arrive…

Putting it all together: the FCN architecture

The full FCN architecture has two paths:

The () takes a pretrained classification network — VGG-16 is the default — and removes its fully-connected layers. Five groups of convolution + pooling shrink the spatial resolution from H×W down to H/32 × W/32 while building increasingly abstract representations. This encoder is initialized with ImageNet-pretrained weights — transfer learning that gives the network a massive head start.

The () maps the coarse, deep features back to pixel-level predictions. First, two 1×1 convolution layers replace VGG's fc6 and fc7, outputting a map with C channels (one per class). Then transposed convolutions upsample back to the input resolution, with skip connections optionally fusing pool3 and pool4 features along the way.

The entire network is trained end-to-end with per-pixel cross-entropy : every pixel's predicted class distribution is compared to its ground-truth label, and flows gradients through both the decoder and the encoder.

Open in Lab
Click any stage to see its role. The encoder shrinks spatial dimensions while the decoder expands them back. Skip connections bridge the two paths.
The demo wakes as you arrive…

The same idea in code

FCN-8s — from classifier to dense predictorpython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn as nn

class FCN8s(nn.Module):
    """Turn VGG-16 into a fully convolutional segmenter."""

    def __init__(self, n_classes=21):
        super().__init__()
        # --- Encoder: VGG-16 conv blocks (pretrained) ---
        # pool1-pool5 shrink spatial dims by 2× each → total 32×
        self.block1 = vgg_block(3,   64,  2)    # /2
        self.block2 = vgg_block(64,  128, 2)    # /4
        self.block3 = vgg_block(128, 256, 3)    # /8   ← skip source
        self.block4 = vgg_block(256, 512, 3)    # /16  ← skip source
        self.block5 = vgg_block(512, 512, 3)    # /32

        # --- Replace FC layers with 1×1 convolutions ---
        self.fc6 = nn.Conv2d(512, 4096, 1)      # was: nn.Linear(25088, 4096)
        self.fc7 = nn.Conv2d(4096, 4096, 1)     # was: nn.Linear(4096, 4096)
        self.score = nn.Conv2d(4096, n_classes, 1)

        # --- Skip connection projections ---
        self.score_pool4 = nn.Conv2d(512, n_classes, 1)
        self.score_pool3 = nn.Conv2d(256, n_classes, 1)

        # --- Decoder: learned transposed convolutions ---
        self.up2x  = nn.ConvTranspose2d(n_classes, n_classes, 4,  stride=2,  padding=1)
        self.up4x  = nn.ConvTranspose2d(n_classes, n_classes, 4,  stride=2,  padding=1)
        self.up8x  = nn.ConvTranspose2d(n_classes, n_classes, 16, stride=8,  padding=4)

    def forward(self, x):
        p1 = self.block1(x)       # H/2
        p2 = self.block2(p1)      # H/4
        p3 = self.block3(p2)      # H/8   — fine features
        p4 = self.block4(p3)      # H/16  — medium features
        p5 = self.block5(p4)      # H/32  — coarse, semantic-rich

        s = self.score(self.fc7(self.fc6(p5)))   # coarse scores

        # Skip fusions: deep + medium + fine
        s = self.up2x(s) + self.score_pool4(p4)  # FCN-16s level
        s = self.up4x(s) + self.score_pool3(p3)  # FCN-8s level
        s = self.up8x(s)                          # back to full res
        return s                                  # (B, n_classes, H, W)

Training: transfer learning meets dense supervision

FCN's strategy leverages two powerful ideas:

Transfer learning from classification. The encoder starts from VGG-16 pretrained on ImageNet's 1.2 million images. These weights already encode rich hierarchical features — edges, textures, parts, objects — that transfer beautifully to segmentation. Only the decoder and skip projections are learned from scratch.

Per-pixel cross-entropy loss. Every pixel independently contributes to the loss — the network receives dense supervision from tens of thousands of labeled pixels per image. This is far richer than classification's single label per image.

The authors trained in stages: first FCN-32s (no skips), then FCN-16s (adding pool4 skip), then FCN-8s (adding pool3 skip). Each stage initializes from the previous one. This staged approach stabilizes training — the coarse prediction is learned first, then progressively refined. The is 10⁻⁴ with and high (0.9), and the new layers use 20× higher learning rate since they start from scratch.

Results and the stride refinement story

The progression from FCN-32s to FCN-8s tells a clear story of refinement:

  • FCN-32s: 59.4% mean IoU on PASCAL VOC 2011. The segmentation captures object identity but boundaries are rough — edges wander by up to 16 pixels.

  • FCN-16s: 62.4% mean IoU. Adding pool4 sharpens edges noticeably. Small objects start appearing.

  • FCN-8s: 62.7% mean IoU. Adding pool3 further refines boundaries. The improvement from 32s to 8s is most visible at object edges and thin structures.

On PASCAL VOC 2012, FCN-8s achieved 62.2% mean IoU — a large improvement over the prior state of the art. ran at ~0.2 seconds per image, compared to ~1 minute for patchwise methods — a 286× speedup.

Open in Lab
Compare segmentation quality across FCN-32s, FCN-16s, and FCN-8s. Notice how edges sharpen with each skip connection.
The demo wakes as you arrive…

Why it mattered

  1. 2012

    AlexNet — CNNs conquer classification

    Deep CNNs trained on ImageNet proved that learned features crush hand-crafted ones. Classification solved — but segmentation still relied on patchwise methods.

  2. 2014

    VGG — deeper is better

    VGG-16/19 showed that simply stacking 3×3 convolutions deeper improves classification. Its clean, uniform architecture made it the ideal backbone for FCN to adapt.

  3. 2015

    FCN — pixels in, pixels out

    This paper. Proved classifiers can become dense predictors with three changes: 1×1 convolutions, transposed convolutions, and skip connections.

  4. 2015

    U-Net — the symmetric encoder-decoder

    Ronneberger et al. extended FCN's skip connections into a symmetric architecture with concatenation instead of addition — dominating biomedical segmentation and becoming the most cited segmentation paper.

  5. 2017

    DeepLab v2 — dilated convolutions

    Instead of downsampling then upsampling, dilated convolutions expand the receptive field without losing resolution — an alternative path inspired by FCN's insight.

  6. 2017

    Mask R-CNN — instance segmentation

    Added an FCN branch to Faster R-CNN: for each detected object, predict a pixel mask. FCN's dense prediction idea extended from scene-level to object-level segmentation.

  7. 2021

    SegFormer — Transformers enter segmentation

    Replaced the CNN encoder with a Transformer backbone while keeping FCN's decoder philosophy: hierarchical features, multi-scale fusion, dense output.

FCN's legacy is not a specific network — it is a design pattern. The idea that classifiers contain spatial information worth recovering, and that skip connections can bridge the gap between semantics and spatial precision, reshaped the entire field of dense prediction. Every time a model outputs a per-pixel map — whether for segmentation, depth estimation, optical flow, or super-resolution — it owes something to this paper.

CitationLong, Shelhamer, Darrell. Fully Convolutional Networks for Semantic Segmentation. CVPR, 2015.

Terms in this paper