Computer Vision2017intermediate8 min read

DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs

DeepLab: التجزئة الدلالية للصور بالشبكات الالتفافية العميقة والالتفاف المُوسَّع وحقول مارکوف الشرطية كاملة الاتصال

Chen, L.-C. · Papandreou, G. · Kokkinos, I. · Murphy, K. · Yuille, A. L. — IEEE TPAMI

The problem

Deep convolutional networks (DCNNs) trained for image shrink spatial resolution through repeated and striding — great for recognizing "what" is in the image, catastrophic for knowing "where." Applying them to faces three walls: (1) output feature maps are 32× smaller than the input, losing fine detail; (2) objects appear at many scales, but fixed-size filters see only one; (3) pooling-driven spatial invariance blurs object boundaries precisely where segmentation must be pixel-sharp.

The contribution

DeepLab introduces three interlocking solutions: (1) () — filters with learned gaps that widen the without shrinking resolution or adding parameters; (2) (ASPP) — parallel atrous convolutions at multiple dilation rates that capture objects and context at several scales simultaneously; (3) Fully connected CRF post-processing that couples DCNN predictions with pixel-level color and position similarity to recover sharp object boundaries. Together they reach 79.7% mIOU on PASCAL VOC 2012.

The impact

DeepLab established atrous as the standard tool for , replacing the need for layers. ASPP became a building block in segmentation architectures for years. The paper spawned DeepLabv3 and DeepLabv3+, which dominate autonomous driving, medical imaging, and satellite analysis to this day. Its core ideas — controlling resolution via dilation and refining boundaries with graphical models — influenced SegFormer, Mask R-CNN, and virtually every modern segmentation system.

Imagine you are a cartographer tracing every building, road, and park on an aerial photo. A normal magnifying glass (standard convolution) shows detail but only a tiny patch — you must zoom out to see context, which blurs the lines you need to trace.

DeepLab gives you a perforated magnifying glass: holes in the lens let light from a wider area through, so you see both the fine boundary and the surrounding context, without ever zooming out.

After tracing, a meticulous proofreader (the CRF) walks along every boundary, checking: does the edge I drew match the color change in the actual photo? Where it doesn't, the proofreader snaps the line to the true edge.

The problem: classification networks destroy spatial detail

Deep convolutional networks like VGG-16 and ResNet-101 are exceptional at answering "what is in this image?" But semantic segmentation needs a harder answer: "what is each pixel?"

These classification networks shrink images through repeated pooling and striding. A 512×512 input becomes a 16×16 — 32 times smaller. That is fine for classification, where a "cat" label does not need a location, but devastating for segmentation, where each of the 262,144 input pixels needs its own label.

Three specific challenges emerge:

  • Resolution collapse. Repeated max-pooling and striding reduce feature maps to 132\frac{1}{32} of the input, destroying the fine spatial information that boundaries depend on.

  • Scale variation. A person in the foreground and a car in the background occupy vastly different numbers of pixels. Fixed-size filters cannot capture both scales.

  • Boundary blur. The spatial invariance that makes classification robust also makes predicted boundaries mushy. Pooling blurs precisely where segmentation must be sharp.

Open in Lab
Click "Classify" vs "Segment" to see how repeated pooling destroys spatial detail needed for per-pixel labeling.
The demo wakes as you arrive…

Solution 1: atrous convolution — see wider without shrinking

The word "atrous" comes from the French à trous — "with holes." The idea is simple but powerful: instead of packing weights side by side, spread them apart by inserting zeros (holes) between them.

A standard 3×3 convolution sees a 3×3 patch. An atrous convolution with dilation rate r=2r = 2 spreads those same 9 weights over a 5×5 area; at r=4r = 4, they cover a 9×9 area. The filter "reaches" a wider receptive field without learning extra weights or shrinking the feature map.

Think of it like a hand with spread fingers pressing into sand: the same five fingers, but the spread lets them sense a larger area. No extra fingers (parameters) are needed, and the sand (feature map) stays at full resolution.

y[i]=∑kx[i+r⋅k] w[k]y[i] = \sum_{k} x[i + r \cdot k] \, w[k]
Atrous (dilated) convolution — x = input feature map · w = filter weights · r = dilation rate · When r = 1 this is standard convolution. Larger r widens the receptive field without adding parameters or reducing resolution.
Open in Lab
Drag the dilation rate slider. Watch the same 3×3 kernel reach wider without growing in parameters.
The demo wakes as you arrive…

In practice, DeepLab takes a classification network (VGG-16 or ResNet-101), removes the last few pooling/striding layers, and replaces subsequent convolutions with atrous convolutions at rate r=2r = 2 or higher. The network now produces feature maps at 18\frac{1}{8} instead of 132\frac{1}{32} of the input resolution — 4× denser, with the same number of parameters.

Solution 2: ASPP — capturing objects at every scale

A single dilation rate sees one scale. But a street scene has pedestrians, cars, buildings, and sky — objects spanning vastly different pixel areas. Atrous Spatial Pyramid Pooling (ASPP) solves this by running several atrous convolutions in parallel, each with a different dilation rate (e.g., r=6,12,18,24r = 6, 12, 18, 24).

Imagine four scouts looking at the same landscape through binoculars set to different zoom levels: one sees a close-up of a leaf, another a tree, the third a grove, the fourth the whole forest. ASPP merges their reports into a single, scale-aware feature map.

Each branch captures context at its own scale. Small dilation rates catch fine textures and small objects; large rates capture global context like "this region is sky." The outputs are concatenated, giving every pixel a description of its neighborhood.

Open in Lab
Toggle each ASPP branch to see what scale of context it captures. Combine them to see the full multi-scale feature map.
The demo wakes as you arrive…

Solution 3: CRF — sharpening boundaries with pixel-level reasoning

Even with atrous convolution, DCNN outputs are spatially smooth — edges are blurry because convolution inherently averages over local neighborhoods. The , however, changes sharply at object boundaries: one pixel is "person," the next is "background."

DeepLab uses a Fully Connected (CRF) as a post-processing step to fix this. Think of it as an image spell-checker that asks two questions for every pixel pair in the image:

  • Color similarity: Do these two pixels look alike in the original image (similar RGB values)? If yes, they probably belong to the same class.

  • Spatial proximity: Are they close together? Nearby pixels with similar colors should share a label.

Unlike short-range CRFs that only check immediate neighbors, DeepLab's fully connected CRF checks every pixel against every other pixel. This lets it recover thin structures and long boundaries that local models miss.

E(x)=∑iψu(xi)+∑i<jψp(xi,xj)E(\mathbf{x}) = \sum_i \psi_u(x_i) + \sum_{i < j} \psi_p(x_i, x_j)
CRF energy function — ψ_u = unary potential from the DCNN (how confident is the network about this pixel's class?) · ψ_p = pairwise potential (do pixels i and j look similar in color and position? if so, penalize different labels). The CRF finds the labeling x that minimizes total energy.
Open in Lab
Watch how CRF iterations progressively sharpen blurry DCNN boundaries. Click to step through iterations.
The demo wakes as you arrive…

Putting it all together — the DeepLab pipeline

The complete DeepLab system is a three-stage pipeline:

Stage 1 — with atrous convolution. Take a pre-trained classification network (VGG-16 or ResNet-101). Remove the last pooling layers and replace convolutions with atrous convolutions. The output is a dense feature map at 18\frac{1}{8} resolution.

Stage 2 — ASPP. Feed the dense features into parallel atrous convolutions at rates r=6,12,18,24r = 6, 12, 18, 24. Concatenate their outputs. Each pixel now has multi-scale context.

Stage 3 — Score map + bilinear + CRF. A 1×11 \times 1 convolution maps ASPP features to per-class scores. upsamples to full resolution. The fully connected CRF refines boundaries.

Click any stage below to see it in action:

Open in Lab
Click each stage to explore how the image flows through the pipeline.
The demo wakes as you arrive…

The same idea in code

Atrous convolution and ASPP in NumPypython

Simplified to show the idea — not the real implementation.

import numpy as np

def atrous_conv2d(image, kernel, rate=1):
    """2D atrous convolution: same kernel, wider reach."""
    H, W = image.shape
    k = kernel.shape[0]
    # Effective kernel size with dilation
    eff_k = k + (k - 1) * (rate - 1)
    out_h, out_w = H - eff_k + 1, W - eff_k + 1
    out = np.zeros((out_h, out_w))
    for i in range(out_h):
        for j in range(out_w):
            # Sample with gaps of `rate` between kernel elements
            patch = image[i:i+eff_k:rate, j:j+eff_k:rate]
            out[i, j] = np.sum(patch * kernel)
    return out

def aspp(feature_map, kernel, rates=[6, 12, 18, 24]):
    """Atrous Spatial Pyramid Pooling: multi-scale context."""
    branches = []
    for r in rates:
        branch = atrous_conv2d(feature_map, kernel, rate=r)
        branches.append(branch)
    # In practice: concatenate along channel axis
    # Here we average for simplicity
    min_h = min(b.shape[0] for b in branches)
    min_w = min(b.shape[1] for b in branches)
    cropped = [b[:min_h, :min_w] for b in branches]
    return np.mean(cropped, axis=0)

# rate=1 → standard 3×3 conv, sees 3×3 area
# rate=2 → same 9 weights, sees 5×5 area
# rate=4 → same 9 weights, sees 9×9 area
# No extra parameters. Just wider vision.

Why it mattered

  1. 2015

    FCN — Fully Convolutional Networks

    Long et al. showed that classification CNNs can be converted to dense predictors by replacing FC layers with convolutions and using deconvolution for upsampling. The first end-to-end segmentation model.

  2. 2015

    DeepLabv1 — Atrous convolution + CRF

    First use of atrous convolution in segmentation and CRF post-processing. Proved the combination works better than deconvolution for boundary recovery.

  3. 2017

    DeepLabv2 — ASPP added

    This paper. Added multi-scale ASPP, switched to ResNet-101 backbone, and reached 79.7% mIOU on PASCAL VOC 2012.

  4. 2017

    DeepLabv3 — Improved ASPP, no CRF

    Refined ASPP with batch normalization and image-level features. Performance improved enough to drop CRF post-processing entirely.

  5. 2018

    DeepLabv3+ — Encoder-decoder with atrous separable convolution

    Added a decoder module for sharper boundaries and used depthwise separable atrous convolution for efficiency. The current standard for many segmentation tasks.

  6. 2021

    SegFormer — Transformers for segmentation

    Replaced convolution with hierarchical vision transformers and a lightweight MLP decoder. No CRF needed — attention naturally captures long-range context that atrous convolution approximated.

FCN proved segmentation could be end-to-end. DeepLab proved it could be precise. SegFormer proved it could be done without convolution at all — but the DNA of ASPP's multi-scale philosophy runs through every modern segmentation model.

CitationChen, Papandreou, Kokkinos, Murphy, Yuille. DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs. IEEE TPAMI, 2017.

Terms in this paper