Computer Vision2017intermediate8 min read
DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
DeepLab: التجزئة الدلالية للصور بالشبكات الالتفافية العميقة والالتفاف المُوسَّع وحقول مارکوف الشرطية كاملة الاتصال
Chen, L.-C. · Papandreou, G. · Kokkinos, I. · Murphy, K. · Yuille, A. L. — IEEE TPAMI
The problem
Deep convolutional networks (DCNNs) trained for image shrink spatial resolution through repeated and striding — great for recognizing "what" is in the image, catastrophic for knowing "where." Applying them to faces three walls: (1) output feature maps are 32× smaller than the input, losing fine detail; (2) objects appear at many scales, but fixed-size filters see only one; (3) pooling-driven spatial invariance blurs object boundaries precisely where segmentation must be pixel-sharp.
The contribution
DeepLab introduces three interlocking solutions: (1) () — filters with learned gaps that widen the without shrinking resolution or adding parameters; (2) (ASPP) — parallel atrous convolutions at multiple dilation rates that capture objects and context at several scales simultaneously; (3) Fully connected CRF post-processing that couples DCNN predictions with pixel-level color and position similarity to recover sharp object boundaries. Together they reach 79.7% mIOU on PASCAL VOC 2012.
The impact
DeepLab established atrous as the standard tool for , replacing the need for layers. ASPP became a building block in segmentation architectures for years. The paper spawned DeepLabv3 and DeepLabv3+, which dominate autonomous driving, medical imaging, and satellite analysis to this day. Its core ideas — controlling resolution via dilation and refining boundaries with graphical models — influenced SegFormer, Mask R-CNN, and virtually every modern segmentation system.
Imagine you are a cartographer tracing every building, road, and park on an aerial photo. A normal magnifying glass (standard convolution) shows detail but only a tiny patch — you must zoom out to see context, which blurs the lines you need to trace.
DeepLab gives you a perforated magnifying glass: holes in the lens let light from a wider area through, so you see both the fine boundary and the surrounding context, without ever zooming out.
After tracing, a meticulous proofreader (the CRF) walks along every boundary, checking: does the edge I drew match the color change in the actual photo? Where it doesn't, the proofreader snaps the line to the true edge.
The problem: classification networks destroy spatial detail
Deep convolutional networks like VGG-16 and ResNet-101 are exceptional at answering "what is in this image?" But semantic segmentation needs a harder answer: "what is each pixel?"
These classification networks shrink images through repeated pooling and striding. A 512×512 input becomes a 16×16 — 32 times smaller. That is fine for classification, where a "cat" label does not need a location, but devastating for segmentation, where each of the 262,144 input pixels needs its own label.
Three specific challenges emerge:
-
Resolution collapse. Repeated max-pooling and striding reduce feature maps to of the input, destroying the fine spatial information that boundaries depend on.
-
Scale variation. A person in the foreground and a car in the background occupy vastly different numbers of pixels. Fixed-size filters cannot capture both scales.
-
Boundary blur. The spatial invariance that makes classification robust also makes predicted boundaries mushy. Pooling blurs precisely where segmentation must be sharp.
Solution 1: atrous convolution — see wider without shrinking
The word "atrous" comes from the French à trous — "with holes." The idea is simple but powerful: instead of packing weights side by side, spread them apart by inserting zeros (holes) between them.
A standard 3×3 convolution sees a 3×3 patch. An atrous convolution with dilation rate spreads those same 9 weights over a 5×5 area; at , they cover a 9×9 area. The filter "reaches" a wider receptive field without learning extra weights or shrinking the feature map.
Think of it like a hand with spread fingers pressing into sand: the same five fingers, but the spread lets them sense a larger area. No extra fingers (parameters) are needed, and the sand (feature map) stays at full resolution.
In practice, DeepLab takes a classification network (VGG-16 or ResNet-101), removes the last few pooling/striding layers, and replaces subsequent convolutions with atrous convolutions at rate or higher. The network now produces feature maps at instead of of the input resolution — 4× denser, with the same number of parameters.
Solution 2: ASPP — capturing objects at every scale
A single dilation rate sees one scale. But a street scene has pedestrians, cars, buildings, and sky — objects spanning vastly different pixel areas. Atrous Spatial Pyramid Pooling (ASPP) solves this by running several atrous convolutions in parallel, each with a different dilation rate (e.g., ).
Imagine four scouts looking at the same landscape through binoculars set to different zoom levels: one sees a close-up of a leaf, another a tree, the third a grove, the fourth the whole forest. ASPP merges their reports into a single, scale-aware feature map.
Each branch captures context at its own scale. Small dilation rates catch fine textures and small objects; large rates capture global context like "this region is sky." The outputs are concatenated, giving every pixel a description of its neighborhood.
Solution 3: CRF — sharpening boundaries with pixel-level reasoning
Even with atrous convolution, DCNN outputs are spatially smooth — edges are blurry because convolution inherently averages over local neighborhoods. The , however, changes sharply at object boundaries: one pixel is "person," the next is "background."
DeepLab uses a Fully Connected (CRF) as a post-processing step to fix this. Think of it as an image spell-checker that asks two questions for every pixel pair in the image:
-
Color similarity: Do these two pixels look alike in the original image (similar RGB values)? If yes, they probably belong to the same class.
-
Spatial proximity: Are they close together? Nearby pixels with similar colors should share a label.
Unlike short-range CRFs that only check immediate neighbors, DeepLab's fully connected CRF checks every pixel against every other pixel. This lets it recover thin structures and long boundaries that local models miss.
Putting it all together — the DeepLab pipeline
The complete DeepLab system is a three-stage pipeline:
Stage 1 — with atrous convolution. Take a pre-trained classification network (VGG-16 or ResNet-101). Remove the last pooling layers and replace convolutions with atrous convolutions. The output is a dense feature map at resolution.
Stage 2 — ASPP. Feed the dense features into parallel atrous convolutions at rates . Concatenate their outputs. Each pixel now has multi-scale context.
Stage 3 — Score map + bilinear + CRF. A convolution maps ASPP features to per-class scores. upsamples to full resolution. The fully connected CRF refines boundaries.
Click any stage below to see it in action:
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def atrous_conv2d(image, kernel, rate=1):
"""2D atrous convolution: same kernel, wider reach."""
H, W = image.shape
k = kernel.shape[0]
# Effective kernel size with dilation
eff_k = k + (k - 1) * (rate - 1)
out_h, out_w = H - eff_k + 1, W - eff_k + 1
out = np.zeros((out_h, out_w))
for i in range(out_h):
for j in range(out_w):
# Sample with gaps of `rate` between kernel elements
patch = image[i:i+eff_k:rate, j:j+eff_k:rate]
out[i, j] = np.sum(patch * kernel)
return out
def aspp(feature_map, kernel, rates=[6, 12, 18, 24]):
"""Atrous Spatial Pyramid Pooling: multi-scale context."""
branches = []
for r in rates:
branch = atrous_conv2d(feature_map, kernel, rate=r)
branches.append(branch)
# In practice: concatenate along channel axis
# Here we average for simplicity
min_h = min(b.shape[0] for b in branches)
min_w = min(b.shape[1] for b in branches)
cropped = [b[:min_h, :min_w] for b in branches]
return np.mean(cropped, axis=0)
# rate=1 → standard 3×3 conv, sees 3×3 area
# rate=2 → same 9 weights, sees 5×5 area
# rate=4 → same 9 weights, sees 9×9 area
# No extra parameters. Just wider vision.Why it mattered
2015
FCN — Fully Convolutional Networks
Long et al. showed that classification CNNs can be converted to dense predictors by replacing FC layers with convolutions and using deconvolution for upsampling. The first end-to-end segmentation model.
2015
DeepLabv1 — Atrous convolution + CRF
First use of atrous convolution in segmentation and CRF post-processing. Proved the combination works better than deconvolution for boundary recovery.
2017
DeepLabv2 — ASPP added
This paper. Added multi-scale ASPP, switched to ResNet-101 backbone, and reached 79.7% mIOU on PASCAL VOC 2012.
2017
DeepLabv3 — Improved ASPP, no CRF
Refined ASPP with batch normalization and image-level features. Performance improved enough to drop CRF post-processing entirely.
2018
DeepLabv3+ — Encoder-decoder with atrous separable convolution
Added a decoder module for sharper boundaries and used depthwise separable atrous convolution for efficiency. The current standard for many segmentation tasks.
2021
SegFormer — Transformers for segmentation
Replaced convolution with hierarchical vision transformers and a lightweight MLP decoder. No CRF needed — attention naturally captures long-range context that atrous convolution approximated.
FCN proved segmentation could be end-to-end. DeepLab proved it could be precise. SegFormer proved it could be done without convolution at all — but the DNA of ASPP's multi-scale philosophy runs through every modern segmentation model.
CitationChen, Papandreou, Kokkinos, Murphy, Yuille. DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs. IEEE TPAMI, 2017.
Terms in this paper
- Semantic Segmentationالتجزئة الدلالية للصورة
- Atrous Convolutionالالتفاف المُوسَّع
- Dilated Convolutionالالتفاف المتوسِّع
- Receptive Fieldالحقل الاستقبالي للعصبون
- Feature Mapخريطة السمات
- Conditional Random Fieldالحقل العشوائي الشرطي
- Atrous Spatial Pyramid Poolingالتجميع الهرمي المكاني بالالتفاف المُوسَّع
- Fully Convolutional Networkشبكة التفافية كاملة
- Dense Predictionالتنبؤ الكثيف
- Upsamplingرفع الدقة
- Poolingالتجميع المكاني
- Strideخطوة التخطّي الحسابية
- Multi-Scaleمتعدد المقاييس
- Bilinear Interpolationالاستيفاء الثنائي الخطي