Computer Vision2015intermediate10 min read
Fully Convolutional Networks for Semantic Segmentation
الشبكات الالتفافية الكاملة للتجزئة الدلالية
Long, J. · Shelhamer, E. · Darrell, T. — CVPR
The problem
networks like VGG and AlexNet condense an entire image into a single label — "cat" or "dog." But many vision tasks need a label for every pixel: self-driving cars must know which pixels are road, pedestrians, or sky. Before this paper, pixel-wise labeling relied on patchwise classification or hand-crafted pipelines that were slow and could not be trained end-to-end.
The contribution
: adapt any classification network (AlexNet, VGG, GoogLeNet) into a that outputs a map the same size as the input. Replace fully-connected layers with 1×1 convolutions so the network accepts any input size, then upsample the coarse output using learned transposed convolutions. Skip connections fuse fine-grained early-layer features with deep semantic features, refining boundaries from 32-pixel (FCN-32s) to 8-pixel stride (FCN-8s).
The impact
FCN established the paradigm for dense prediction that every modern segmentation model follows. U-Net, DeepLab, Mask R-, and SegNet are all direct descendants. It proved that classification backbones can be repurposed for pixel-level tasks, making the standard starting point for segmentation. The paper's 62.2% mean on PASCAL VOC 2012 exceeded the prior state of the art by a wide margin and ran 286× faster than patchwise approaches.
A classification network is like a funnel: it pours an image in at the top and squeezes out a single drop — the class label — at the bottom. All spatial information is crushed.
FCN turns the funnel into a megaphone: after squeezing, it expands the signal back to the original image size, so every pixel gets its own label. And by connecting the wide mouth of the funnel to the expanding bell of the megaphone — skip connections — it recovers the fine details that the squeezing lost.
The problem: classifiers throw away location
By 2014, deep convolutional networks like AlexNet and VGG had conquered image classification. But classification answers what is in the image, not where. requires both: every pixel must be assigned a class label.
The standard approach at the time was patchwise classification: crop a small patch around each pixel, classify it with a CNN, repeat for every pixel. This was painfully slow and threw away context — a pixel's patch didn't know what the neighboring patches saw. Other methods used hand-crafted superpixels or CRF post-processing bolted onto CNN features, creating fragmented pipelines that could not be trained end-to-end.
The key insight: fully convolutional = any size in, same size out
Classification networks end with fully-connected layers that demand a fixed input size (e.g. 224×224) and output a fixed-length vector. These layers destroy spatial layout — they treat every neuron as equally connected to every output, regardless of position.
FCN's insight: replace every fully-connected layer with a 1×1 . A 1×1 convolution applies a learned linear combination at each spatial position independently. The output is no longer a vector but a spatial map — a heat map of class scores at every location. Because convolutions don't care about input size, the network now accepts images of any dimension and produces correspondingly-sized output.
Think of it this way: a fully-connected layer is like reading an entire essay and writing one summary sentence. A 1×1 convolution is like a proofreader who walks through the essay paragraph by paragraph, writing a margin note at each one. The same weights (the same critical eye), applied at every position.
The decoder: expanding back to full resolution
After the convolutional layers of VGG or AlexNet, the is 32× smaller than the input (e.g. 7×7 from a 224×224 image) due to repeated . We need to expand this coarse map back to the original resolution — pixel by pixel.
FCN uses transposed convolution (also called or upconvolution). Think of a normal convolution as many-to-one: many input pixels contribute to one output pixel. A transposed convolution reverses this: one input pixel contributes to many output pixels. It inserts zeros between input values, then applies a learned filter — effectively spreading and blending each coarse prediction across a larger spatial area.
The key advantage: these weights are learned, not hand-designed. The network discovers how to interpolate between coarse predictions in a way that best recovers the original spatial detail.
Skip connections: recovering fine detail
Upsampling the final layer alone (FCN-32s) produces segmentations that are blobby and imprecise — edges are smeared, small objects are lost. The reason: 32× downsampling discards too much spatial detail. But earlier layers in the network still have that detail — pool3 has 8× stride, pool4 has 16× stride.
The paper's second key idea: skip connections that fuse predictions from the deep, coarse layer with features from shallower, finer layers:
-
FCN-32s — upsample the final layer 32× directly. Fast but coarse.
-
FCN-16s — combine the 2× upsampled final prediction with pool4 features, then upsample 16×. Edges become noticeably sharper.
-
FCN-8s — add pool3 features too. Three streams fused. Boundaries become crisp enough for practical use.
Think of it as a conference call: the deep layers contribute what objects are present (a zoomed-out view), while the shallow layers contribute where exactly the boundaries fall (a zoomed-in view). Fusing both gives segmentations that are simultaneously accurate in identity and precise in location.
Putting it all together: the FCN architecture
The full FCN architecture has two paths:
The () takes a pretrained classification network — VGG-16 is the default — and removes its fully-connected layers. Five groups of convolution + pooling shrink the spatial resolution from H×W down to H/32 × W/32 while building increasingly abstract representations. This encoder is initialized with ImageNet-pretrained weights — transfer learning that gives the network a massive head start.
The () maps the coarse, deep features back to pixel-level predictions. First, two 1×1 convolution layers replace VGG's fc6 and fc7, outputting a map with C channels (one per class). Then transposed convolutions upsample back to the input resolution, with skip connections optionally fusing pool3 and pool4 features along the way.
The entire network is trained end-to-end with per-pixel cross-entropy : every pixel's predicted class distribution is compared to its ground-truth label, and flows gradients through both the decoder and the encoder.
The same idea in code
Simplified to show the idea — not the real implementation.
import torch
import torch.nn as nn
class FCN8s(nn.Module):
"""Turn VGG-16 into a fully convolutional segmenter."""
def __init__(self, n_classes=21):
super().__init__()
# --- Encoder: VGG-16 conv blocks (pretrained) ---
# pool1-pool5 shrink spatial dims by 2× each → total 32×
self.block1 = vgg_block(3, 64, 2) # /2
self.block2 = vgg_block(64, 128, 2) # /4
self.block3 = vgg_block(128, 256, 3) # /8 ← skip source
self.block4 = vgg_block(256, 512, 3) # /16 ← skip source
self.block5 = vgg_block(512, 512, 3) # /32
# --- Replace FC layers with 1×1 convolutions ---
self.fc6 = nn.Conv2d(512, 4096, 1) # was: nn.Linear(25088, 4096)
self.fc7 = nn.Conv2d(4096, 4096, 1) # was: nn.Linear(4096, 4096)
self.score = nn.Conv2d(4096, n_classes, 1)
# --- Skip connection projections ---
self.score_pool4 = nn.Conv2d(512, n_classes, 1)
self.score_pool3 = nn.Conv2d(256, n_classes, 1)
# --- Decoder: learned transposed convolutions ---
self.up2x = nn.ConvTranspose2d(n_classes, n_classes, 4, stride=2, padding=1)
self.up4x = nn.ConvTranspose2d(n_classes, n_classes, 4, stride=2, padding=1)
self.up8x = nn.ConvTranspose2d(n_classes, n_classes, 16, stride=8, padding=4)
def forward(self, x):
p1 = self.block1(x) # H/2
p2 = self.block2(p1) # H/4
p3 = self.block3(p2) # H/8 — fine features
p4 = self.block4(p3) # H/16 — medium features
p5 = self.block5(p4) # H/32 — coarse, semantic-rich
s = self.score(self.fc7(self.fc6(p5))) # coarse scores
# Skip fusions: deep + medium + fine
s = self.up2x(s) + self.score_pool4(p4) # FCN-16s level
s = self.up4x(s) + self.score_pool3(p3) # FCN-8s level
s = self.up8x(s) # back to full res
return s # (B, n_classes, H, W)Training: transfer learning meets dense supervision
FCN's strategy leverages two powerful ideas:
Transfer learning from classification. The encoder starts from VGG-16 pretrained on ImageNet's 1.2 million images. These weights already encode rich hierarchical features — edges, textures, parts, objects — that transfer beautifully to segmentation. Only the decoder and skip projections are learned from scratch.
Per-pixel cross-entropy loss. Every pixel independently contributes to the loss — the network receives dense supervision from tens of thousands of labeled pixels per image. This is far richer than classification's single label per image.
The authors trained in stages: first FCN-32s (no skips), then FCN-16s (adding pool4 skip), then FCN-8s (adding pool3 skip). Each stage initializes from the previous one. This staged approach stabilizes training — the coarse prediction is learned first, then progressively refined. The is 10⁻⁴ with and high (0.9), and the new layers use 20× higher learning rate since they start from scratch.
Results and the stride refinement story
The progression from FCN-32s to FCN-8s tells a clear story of refinement:
-
FCN-32s: 59.4% mean IoU on PASCAL VOC 2011. The segmentation captures object identity but boundaries are rough — edges wander by up to 16 pixels.
-
FCN-16s: 62.4% mean IoU. Adding pool4 sharpens edges noticeably. Small objects start appearing.
-
FCN-8s: 62.7% mean IoU. Adding pool3 further refines boundaries. The improvement from 32s to 8s is most visible at object edges and thin structures.
On PASCAL VOC 2012, FCN-8s achieved 62.2% mean IoU — a large improvement over the prior state of the art. ran at ~0.2 seconds per image, compared to ~1 minute for patchwise methods — a 286× speedup.
Why it mattered
2012
AlexNet — CNNs conquer classification
Deep CNNs trained on ImageNet proved that learned features crush hand-crafted ones. Classification solved — but segmentation still relied on patchwise methods.
2014
VGG — deeper is better
VGG-16/19 showed that simply stacking 3×3 convolutions deeper improves classification. Its clean, uniform architecture made it the ideal backbone for FCN to adapt.
2015
FCN — pixels in, pixels out
This paper. Proved classifiers can become dense predictors with three changes: 1×1 convolutions, transposed convolutions, and skip connections.
2015
U-Net — the symmetric encoder-decoder
Ronneberger et al. extended FCN's skip connections into a symmetric architecture with concatenation instead of addition — dominating biomedical segmentation and becoming the most cited segmentation paper.
2017
DeepLab v2 — dilated convolutions
Instead of downsampling then upsampling, dilated convolutions expand the receptive field without losing resolution — an alternative path inspired by FCN's insight.
2017
Mask R-CNN — instance segmentation
Added an FCN branch to Faster R-CNN: for each detected object, predict a pixel mask. FCN's dense prediction idea extended from scene-level to object-level segmentation.
2021
SegFormer — Transformers enter segmentation
Replaced the CNN encoder with a Transformer backbone while keeping FCN's decoder philosophy: hierarchical features, multi-scale fusion, dense output.
FCN's legacy is not a specific network — it is a design pattern. The idea that classifiers contain spatial information worth recovering, and that skip connections can bridge the gap between semantics and spatial precision, reshaped the entire field of dense prediction. Every time a model outputs a per-pixel map — whether for segmentation, depth estimation, optical flow, or super-resolution — it owes something to this paper.
CitationLong, Shelhamer, Darrell. Fully Convolutional Networks for Semantic Segmentation. CVPR, 2015.
Terms in this paper
- Semantic Segmentationالتجزئة الدلالية للصورة
- Fully Convolutional Network (FCN)الشبكة الالتفافية الكاملة
- Dense Predictionالتنبؤ الكثيف
- Deconvolutionالتفاف معكوس
- Upsamplingرفع الدقة
- Skip Connectionالاتصال التجاوزي
- Feature Mapخريطة السمات
- Poolingالتجميع المكاني
- Strideخطوة التخطّي الحسابية
- Receptive Fieldالحقل الاستقبالي للعصبون
- Transfer Learningنقل التعلم
- Fine-Tuningالضبط الدقيق