Computer Vision2016intermediate10 min read

SSD: Single Shot MultiBox Detector

SSD: كاشف متعدد الصناديق بتمريرة واحدة

Liu, W. · Anguelov, D. · Erhan, D. · Szegedy, C. · Reed, S. · Fu, C.-Y. · Berg, A. C. — ECCV

The problem

By 2015, accurate required two stages: first generate region proposals (like Faster R-), then classify each one. This "propose-then-classify" pipeline was accurate but slow — too slow for real-time applications like autonomous driving and video surveillance. Meanwhile, single-shot detectors like YOLO v1 were fast but sacrificed accuracy, especially on small objects, because they predicted from a single, coarse .

The contribution

SSD: a that predicts from multiple feature maps at different resolutions simultaneously. A VGG-16 extracts features, then six progressively smaller feature maps (38×38 down to 1×1) each predict bounding boxes with default (anchor) boxes of varying aspect ratios. The network predicts class scores and box offsets for every in a single — no step needed. and extensive stabilize . SSD300 achieved 74.3% mAP on VOC2007 at 59 FPS — faster than Faster R-CNN and more accurate than YOLO.

The impact

SSD proved that single-stage detectors could match two-stage accuracy while running in real time. Its prediction idea — detect small objects from early, high-resolution layers and large objects from deep, low-resolution layers — became the foundational principle behind Feature Pyramid Networks (FPN) and RetinaNet. Every modern detector, from YOLOv3 onward, uses multi-scale features that trace back to SSD's design.

Imagine a security guard monitoring a building through six screens, each showing the same scene at a different zoom level. On the close-up screen, he spots a dropped wallet. On the wide-angle screen, he catches a delivery truck pulling in. He doesn't check one screen, then the next — he scans all six at once and calls out everything he sees in a single breath.

That's SSD. Previous detectors were like a guard who first circled the building to list "suspicious areas" (region proposals), then went back to investigate each one. SSD skips the patrol entirely — one look through all six screens, and every object is detected.

The problem: speed vs accuracy — pick one?

Before SSD, the object detection landscape was split into two camps:

  • Two-stage detectors like Faster R-CNN first generated candidate regions, then classified each one. Accurate (73.2% mAP on VOC2007) but slow — only 7 FPS. Every potential object required its own forward pass through the classifier.

  • Single-stage detectors like YOLO v1 predicted everything in one pass. Fast (45 FPS) but inaccurate on small objects, because the entire image was divided into a coarse 7×7 grid — objects smaller than a grid cell were effectively invisible.

The community needed a detector that was both fast and accurate. The key insight SSD brought: you don't need to choose between speed and accuracy if you predict at multiple scales simultaneously.

Open in Lab
Compare the two-stage pipeline (propose → classify) with SSD's single-shot approach. Notice how SSD eliminates the proposal step entirely.
The demo wakes as you arrive…

The core idea: detect at every scale

SSD's architecture has two parts: a backbone network (VGG-16, pre-trained on ImageNet) that extracts features, and a set of auxiliary convolutional layers that progressively shrink the feature maps. Think of it as a pyramid of six observation decks at different heights — each deck surveys the scene at a different zoom level:

  • Conv4_3 (38×38) — high resolution, small → catches small objects like cups and remote controls.
  • Conv7 (19×19) — medium resolution → spots medium objects like chairs and monitors.
  • Conv8_2 (10×10), Conv9_2 (5×5) — lower resolution, wider view → detects large objects like tables and people.
  • Conv10_2 (3×3), Conv11_2 (1×1) — lowest resolution, broadest view → handles very large objects and full-scene context.

At every position on every feature map, the network places a set of default boxes (also called anchor boxes) with different aspect ratios. For each default box, it predicts two things: (1) class probabilities for every category plus "background," and (2) four offsets (Δcx, Δcy, Δw, Δh) to adjust the box to fit the actual object.

Open in Lab
Click on different feature map layers to see how each scale detects different-sized objects. Larger feature maps detect smaller objects.
The demo wakes as you arrive…

Default boxes: pre-shaped guesses at every position

Instead of predicting boxes from scratch (which is hard to learn), SSD starts from a set of pre-defined default boxes — think of them as templates. At each position on a feature map, SSD tiles a small collection of boxes with different aspect ratios (1:1, 2:1, 1:2, 3:1, 1:3) and two square boxes at slightly different scales. The network's job is not to create boxes from nothing but to nudge these templates to better fit the objects.

Each feature map uses a different base scale for its default boxes. The scales grow linearly from 0.2 (for the finest map) to 0.9 (for the coarsest map), so the six maps together tile the full range of object sizes. With 4–6 default boxes per position across six maps, SSD produces a total of 8,732 predictions — all in a single forward pass.

sk=smin⁡+smax⁡−smin⁡m−1(k−1),k∈[1,m]s_k = s_{\min} + \frac{s_{\max} - s_{\min}}{m - 1}(k - 1), \quad k \in [1, m]
Default box scale formula — how box sizes grow across feature maps — s_min = 0.2, s_max = 0.9, m = number of feature maps. Each feature map k gets a base scale s_k. Aspect ratios {1, 2, 3, 1/2, 1/3} define width and height: w = s_k√a_r, h = s_k/√a_r.
Open in Lab
Drag the scale slider and toggle aspect ratios to see how default boxes cover an image cell.
The demo wakes as you arrive…

Architecture walkthrough: from pixels to 8,732 predictions

The SSD pipeline flows through three stages:

Stage 1 — Feature extraction (VGG-16 backbone). The input image (300×300) passes through VGG-16's convolutional layers up to Conv5_3. The fully-connected layers FC6 and FC7 are converted to convolutional layers (using subsampling), producing the first two prediction feature maps. This conversion is what makes SSD fully convolutional — no fixed-size input required in principle.

Stage 2 — Auxiliary convolution layers. Four additional pairs of convolutional layers are added after the VGG backbone. Each pair halves the spatial dimensions, creating the pyramid of feature maps: 10×10, 5×5, 3×3, 1×1. The progressively shrinking maps capture progressively larger objects, because each position's receptive field covers a bigger chunk of the original image — like zooming out step by step.

Stage 3 — Prediction layers. At each of the six feature maps, a 3×3 convolutional filter is applied for each default box, producing (C+4) outputs per box: C class scores plus 4 offsets. There are no fully-connected layers, no region proposals, and no feature resampling. The entire pipeline is a single feedforward network.

Open in Lab
Click on any layer to see its role, dimensions, and number of predictions.
The demo wakes as you arrive…

Training: matching, mining, and the loss function

Training SSD involves three key mechanisms that work together:

Matching strategy. Each ground-truth box is first matched to the default box with the highest . Then, any remaining default box with IoU > 0.5 to any is also marked as positive. This ensures every ground-truth object has at least one matched predictor, while also allowing nearby default boxes to learn from the same object.

Hard negative mining. With 8,732 default boxes and typically only a handful of objects, the vast majority of boxes are "background." Blindly training on all negatives would overwhelm the positives. SSD sorts negative boxes by their confidence loss (how badly they misclassified background) and keeps only the top ones, maintaining a 3:1 ratio of negatives to positives. This focuses training on the hardest mistakes.

Data augmentation. Each training image is randomly transformed: horizontal flips, random crops with minimum IoU constraints (0.1, 0.3, 0.5, 0.7, 0.9), color distortions, and random expansion (zoom out). This aggressive augmentation is critical for small object detection — a random crop that makes an object fill most of the frame teaches the detector to see it at a larger scale.

L(x,c,l,g)=1N[Lconf(x,c)+α Lloc(x,l,g)]L(x, c, l, g) = \frac{1}{N}\Big[L_{\text{conf}}(x, c) + \alpha \, L_{\text{loc}}(x, l, g)\Big]
SSD training loss — classification + localization — N = number of matched default boxes · L_conf = softmax cross-entropy over class confidences · L_loc = Smooth L1 loss over box offsets (Δcx, Δcy, Δw, Δh) · α = 1 (balances the two terms) · When N = 0, loss is set to 0.

After prediction: Non-Maximum Suppression

With 8,732 predictions per image, many boxes will overlap on the same object. (NMS) cleans this up: it first filters out all boxes with confidence below 0.01, then for each class, sorts the remaining boxes by confidence and iteratively removes any box that overlaps (IoU > 0.45) with a higher-confidence box. The surviving boxes are the final detections. Think of it as a committee vote — when multiple boxes claim the same object, only the most confident one gets to speak.

Open in Lab
Click "Detect" to see raw predictions, then "Apply NMS" to watch overlapping boxes vanish.
The demo wakes as you arrive…

Results: real-time accuracy

SSD delivered on both fronts — speed and accuracy — closing the gap that defined the field:

SSD300 (300×300 input): 74.3% mAP on VOC2007 at 59 FPS on a Titan X GPU. This matched Faster R-CNN's accuracy while being 8× faster.

SSD512 (512×512 input): 76.8% mAP on VOC2007 at 22 FPS. Trading some speed for higher resolution improved small-object detection significantly.

On COCO, SSD512 achieved 26.8% mAP — competitive with Faster R-CNN (21.9%) and ION (23.6%) while maintaining real-time performance. The ablation studies confirmed that multi-scale prediction is the single most important factor: using six feature maps instead of one improved mAP from 62.4% to 74.3%.

Open in Lab
Compare SSD variants against Faster R-CNN and YOLO on the speed-accuracy tradeoff.
The demo wakes as you arrive…

The same idea in code

SSD prediction head — predicting boxes and classes from a feature mappython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn as nn

class SSDPredictionHead(nn.Module):
    """For each feature map, predict class scores + box offsets
    for every default box at every spatial position."""

    def __init__(self, in_channels, num_classes, num_boxes):
        super().__init__()
        # 3x3 conv: predict class scores for each default box
        self.cls_conv = nn.Conv2d(
            in_channels,
            num_boxes * num_classes,   # e.g. 6 boxes × 21 classes
            kernel_size=3, padding=1
        )
        # 3x3 conv: predict 4 offsets (dx, dy, dw, dh) per box
        self.loc_conv = nn.Conv2d(
            in_channels,
            num_boxes * 4,             # e.g. 6 boxes × 4 offsets
            kernel_size=3, padding=1
        )

    def forward(self, feature_map):
        # feature_map: (batch, channels, H, W)
        cls = self.cls_conv(feature_map)  # (batch, boxes*classes, H, W)
        loc = self.loc_conv(feature_map)  # (batch, boxes*4, H, W)
        return cls, loc

# SSD applies this head to EACH of its 6 feature maps:
# Conv4_3 (38×38, 4 boxes) + Conv7 (19×19, 6 boxes) + ...
# Total: 38²×4 + 19²×6 + 10²×6 + 5²×6 + 3²×4 + 1²×4 = 8732 boxes

Why multi-scale detection changed everything

  1. 2014

    R-CNN

    The first deep learning object detector. Used selective search to propose ~2,000 regions, then classified each with a CNN. Accurate but extremely slow — 47 seconds per image.

  2. 2015

    Faster R-CNN

    Replaced selective search with a Region Proposal Network (RPN) sharing features with the detector. Reached 73.2% mAP on VOC2007 but still only 7 FPS.

  3. 2015

    YOLO v1

    First single-stage detector — one forward pass, no proposals. 45 FPS but only 63.4% mAP, struggling with small objects due to its coarse 7×7 grid.

  4. 2016

    SSD

    Bridged the gap: multi-scale prediction from 6 feature maps delivered both speed (59 FPS) and accuracy (74.3% mAP). Proved single-stage can match two-stage quality.

  5. 2017

    Feature Pyramid Network (FPN)

    Formalized SSD's multi-scale idea by adding top-down pathways with lateral connections, enriching low-level features with high-level semantics.

  6. 2017

    RetinaNet

    Combined FPN with Focal Loss — a new loss function that down-weights easy negatives automatically, replacing SSD's hard negative mining with a more elegant solution.

SSD's multi-scale design became the standard playbook. FPN refined the idea by feeding semantic information back down to fine-grained layers, and RetinaNet solved the problem more elegantly with focal loss. But the core principle — predicting at multiple resolutions — was SSD's contribution to the field.

CitationLiu, Anguelov, Erhan, Szegedy, Reed, Fu, Berg. SSD: Single Shot MultiBox Detector. ECCV, 2016.

Terms in this paper