Computer Vision2018intermediate11 min read

YOLOv3: An Incremental Improvement

YOLOv3: تحسين تدريجي

Redmon, J. · Farhadi, A. — arXiv

The problem

Two-stage detectors like Faster R-CNN first propose candidate regions, then classify each one — accurate but far too slow for real-time video. YOLOv2 was fast but struggled with small objects and could not assign multiple labels to the same object. The field needed a that was both fast enough for real-time use and accurate enough across all object sizes.

The contribution

YOLOv3 introduced three key upgrades: (1) -53, a deeper with residual connections that matches ResNet-152 accuracy at twice the speed; (2) prediction at three sizes (13×13, 26×26, 52×52), inspired by Networks, dramatically improving small-; (3) independent logistic classifiers replacing , enabling multi-label for overlapping categories. At 320×320 it runs in 22 ms at 28.2 , matching SSD accuracy but three times faster.

The impact

YOLOv3 became the go-to detector for real-time applications — autonomous driving, surveillance, robotics, and mobile vision. Its multi-scale design and Darknet backbone set the template that YOLOv4, v5, v7, and v8 all built upon. The paper's casual tone and honest reporting of failures made it one of the most cited and beloved papers in computer vision. Joseph Redmon's decision to stop YOLO research over ethical concerns also sparked a wider conversation about responsible AI.

Imagine you're in a helicopter surveying a city. A is like stopping, hovering over each suspicious spot, pulling out binoculars, and carefully identifying what's there. Accurate, but painfully slow when you need to cover the whole city.

YOLO is the pilot who photographs the entire city in one click, then instantly reads the photo to find every car, building, and person. YOLOv3 upgrades the camera with three zoom lenses — wide-angle for big buildings, medium for cars, and telephoto for pedestrians — so nothing escapes regardless of size.

The problem: fast or accurate, pick one

By 2018, object detection had split into two camps:

  • Two-stage detectors (Faster R-CNN, R-FCN): first propose regions, then classify each one. High accuracy but too slow for real-time video — hundreds of milliseconds per frame.

  • Single-stage detectors (YOLOv2, SSD): process the whole image at once and predict boxes directly. Fast enough for video, but weaker on small objects and unable to handle overlapping categories.

YOLOv2 in particular had a glaring weakness: it used a relatively shallow backbone, detected at a single scale, and relied on softmax — which forced every object into exactly one class. A "woman" wearing a "dress" could only be labeled as one or the other.

Open in Lab
Toggle between single-scale (YOLOv2) and multi-scale (YOLOv3) detection. Notice how small objects are missed at the coarse 13×13 grid.
The demo wakes as you arrive…

Darknet-53: a deeper, faster backbone

The first upgrade is the feature extractor. YOLOv2 used Darknet-19 (19 convolutional layers). YOLOv3 replaces it with Darknet-53 — 53 convolutional layers organized into residual blocks, borrowing the skip-connection idea from ResNet.

Each pairs a 1×1 convolution (to compress channels) with a 3×3 convolution (to extract spatial features), then adds the input back. This is the same "shortcut highway" pattern from ResNet: gradients flow through the without decaying, so the network can be much deeper without suffering from the .

The result: Darknet-53 achieves the same classification accuracy as ResNet-152 on , but runs nearly twice as fast because it uses strided convolutions instead of max-pooling and has no fully connected layers — every parameter is a convolutional filter.

Open in Lab
Click on any residual block to see its internal structure. Notice how the 1×1 convolution compresses channels before the 3×3 expands them.
The demo wakes as you arrive…

Bounding box prediction: anchors meet sigmoids

YOLOv3 doesn't predict box coordinates from scratch. Instead, it starts from anchor boxes — pre-computed box shapes found by running on the data's ground-truth boxes. Think of anchors as a set of "starting guesses" for common object shapes: tall and thin (like a person standing), wide and short (like a car), or roughly square (like a face).

The network then predicts offsets from these anchors. For each , the network outputs four numbers — tx,ty,tw,tht_x, t_y, t_w, t_h — plus an tot_o.

bx=σ(tx)+cxby=σ(ty)+cybw=pwetwbh=phethb_x = \sigma(t_x) + c_x \qquad b_y = \sigma(t_y) + c_y \qquad b_w = p_w e^{t_w} \qquad b_h = p_h e^{t_h}
Bounding box prediction from anchor offsets — σ(tₓ), σ(tᵧ) = sigmoid-squeezed center offsets within the grid cell · cₓ, cᵧ = grid cell top-left corner · pʷ, pʰ = anchor width and height · eᵗʷ, eᵗʰ = exponential scaling of anchor dimensions

Why the sigmoid on center coordinates? Without it, the predicted center could land anywhere in the image — making training unstable. The sigmoid squeezes the output between 0 and 1, so the center stays within its own . It's like saying: "you can nudge the anchor's center anywhere inside this one cell, but you can't teleport it across the image."

The width and height use exponential scaling: the network predicts how much to stretch or shrink the anchor. Predicting tw=0t_w = 0 means "keep the anchor's width"; tw=1t_w = 1 means "multiply it by e≈2.7e \approx 2.7."

The objectness score σ(to)\sigma(t_o) tells us: "is there actually an object here, or is this box just background noise?" Each ground-truth object is assigned to exactly one anchor — the one with the highest . If an anchor overlaps a ground-truth object by more than 0.5 but isn't the best match, it is simply ignored during training.

Open in Lab
Drag the sliders to change tₓ, tᵧ, tᵤ, tₕ and watch how the predicted box (blue) transforms from the anchor (dashed).
The demo wakes as you arrive…

Multi-scale detection: three zoom levels

The biggest leap in YOLOv3 is predicting objects at three different scales, inspired by Feature Pyramid Networks (FPN). Here's the intuition:

Imagine you're looking at a city photo. To find skyscrapers, you need a wide-angle view — a coarse grid where each cell covers a large area. To find street signs, you need to zoom in — a fine grid where each cell covers a small area. YOLOv3 does both, plus one in between.

The network produces feature maps at three resolutions. For a 416×416 input:

  • 13×13 ( 32) — each cell sees a 32×32 pixel region. Large objects detected here using the 3 largest anchors.

  • 26×26 (stride 16) — each cell sees a 16×16 region. Medium objects with mid-sized anchors.

  • 52×52 (stride 8) — each cell sees an 8×8 region. Small objects with the 3 smallest anchors.

To build the finer scales, YOLOv3 upsamples the deeper (13×13) feature map and concatenates it with the earlier (26×26) feature map from the backbone. This is the FPN idea: the deep layers carry rich semantic meaning ("this is a dog") while the shallow layers carry fine spatial detail ("the edge is at pixel 47"). Merging them gives both.

Open in Lab
Select a detection scale to see which objects it catches. The feature pyramid fuses deep semantics with shallow spatial detail.
The demo wakes as you arrive…

Multi-label classification: beyond softmax

YOLOv2 used softmax for class prediction, which forces all class probabilities to sum to 1. This means each detected object gets exactly one label. But real-world datasets often have overlapping categories: "woman" and "person", "car" and "vehicle".

YOLOv3 replaces softmax with independent logistic classifiers — one sigmoid per class. Each class gets its own binary yes/no decision. A bounding box can now be simultaneously "woman" (0.95) and "person" (0.98), because the scores don't compete.

The training loss for class prediction switches from cross-entropy (which assumes mutual exclusion) to — one loss term per class, independent of the others. This seemingly small change unlocked YOLO for datasets with hierarchical or overlapping labels.

Putting it all together

The complete YOLOv3 architecture works as an information pipeline with three stages:

Stage 1 — (Darknet-53): The input image passes through 53 convolutional layers with residual blocks. Each Darknet block uses and activation. Strided convolutions (stride=2) replace max-pooling for downsampling, reducing information loss.

Stage 2 — Feature pyramid (multi-scale ): Features from three depths of the backbone are tapped off. The deepest features are upsampled and concatenated with earlier features, creating a pyramid where each level combines semantic richness with spatial precision.

Stage 3 — Detection heads: At each of the three scales, a small convolutional sub-network predicts a 3D tensor of shape S×S×[3×(4+1+C)]S \times S \times [3 \times (4 + 1 + C)] — for 3 anchors, each with 4 box coordinates, 1 objectness score, and C class scores. For COCO (80 classes) that's S×S×255S \times S \times 255.

In total, for a 416×416 input, YOLOv3 predicts (132+262+522)×3=10,647(13^2 + 26^2 + 52^2) \times 3 = 10{,}647 bounding boxes. Most are suppressed by low objectness and .

Open in Lab
Click on any stage to expand its internal structure and data flow.
The demo wakes as you arrive…

The core idea in code

YOLOv3 bounding box decoding — from network output to pixel coordinatespython

Simplified to show the idea — not the real implementation.

import numpy as np

def sigmoid(x):
    return 1 / (1 + np.exp(-x))

def decode_yolo_boxes(raw_output, anchors, grid_size, img_size=416):
    """Decode raw YOLO output into (x, y, w, h) in pixel coordinates.

    raw_output: (S, S, 3, 5+C) — 3 anchors, 4 coords + 1 obj + C classes
    anchors:    (3, 2) — width, height of each anchor in pixels
    grid_size:  S — e.g. 13, 26, or 52
    """
    S = grid_size
    stride = img_size / S

    # Build the grid of cell offsets: cx, cy for every cell
    cx = np.arange(S).reshape(1, S, 1)   # (1, S, 1)
    cy = np.arange(S).reshape(S, 1, 1)   # (S, 1, 1)

    # Extract raw predictions
    tx = raw_output[..., 0]   # center x offset
    ty = raw_output[..., 1]   # center y offset
    tw = raw_output[..., 2]   # width scale
    th = raw_output[..., 3]   # height scale
    to = raw_output[..., 4]   # objectness

    # Decode — this is the whole bounding-box formula
    bx = (sigmoid(tx) + cx) * stride        # pixel x of center
    by = (sigmoid(ty) + cy) * stride        # pixel y of center
    bw = anchors[:, 0] * np.exp(tw)         # pixel width
    bh = anchors[:, 1] * np.exp(th)         # pixel height
    obj = sigmoid(to)                        # objectness probability

    return np.stack([bx, by, bw, bh, obj], axis=-1)

# Example: decode the 13×13 scale with the 3 largest anchors
# anchors_large = np.array([[116,90], [156,198], [373,326]])
# boxes = decode_yolo_boxes(raw_13x13, anchors_large, 13)

Performance: speed vs. accuracy

YOLOv3 hits a distinctive sweet spot on the speed-accuracy curve:

  • At 320×320, it runs in 22 ms (45 FPS) at 28.2 mAP, matching SSD but 3× faster.

  • At 416×416, it achieves 31.0 mAP in 29 ms.

  • At 608×608, it reaches 33.0 mAP in 51 ms — comparable to RetinaNet's 57.5 mAP@50 but 3.8× faster.

On the older mAP@50 metric (which YOLOv3 excels at), it scores 57.9 mAP@50 at 608×608, nearly matching RetinaNet. However, on the stricter mAP@[.5:.95] metric, YOLOv3 falls behind — the paper honestly notes it struggles with precise box localization at higher IoU thresholds.

The honest reporting is part of what makes this paper special. Redmon openly lists what didn't work: dropped mAP by 2 points, linear x/y predictions destabilized training, and dual IoU thresholds "produced similar results." This transparency is rare and valuable.

Open in Lab
Hover over each detector to compare inference time vs. mAP. YOLOv3 dominates the top-left (fast + accurate) region.
The demo wakes as you arrive…

Things we tried that didn't work

Why YOLOv3 matters

YOLOv3 was not a single breakthrough moment — it was a masterclass in engineering refinement. By combining three proven ideas (deeper residual backbones, feature pyramids, and independent classifiers) into one cohesive system, Redmon showed that incremental improvements, done right, can transform a detector's capabilities.

More importantly, YOLOv3 set the architectural template that the entire YOLO family follows to this day: a deep convolutional backbone → multi-scale feature fusion → detection heads at each scale. YOLOv4, v5, v7, v8, and v11 all iterate on this same skeleton.

  1. 2016

    YOLOv1

    The original "You Only Look Once" — framed detection as a single regression problem. Revolutionary speed but coarse localization and struggles with small objects.

  2. 2017

    YOLOv2 / YOLO9000

    Added anchor boxes, batch normalization, and multi-resolution training. YOLO9000 could detect 9000+ categories by jointly training on detection and classification data.

  3. 2018

    YOLOv3

    Darknet-53, multi-scale FPN-style detection, independent logistic classifiers. The architectural template for all future YOLO versions.

  4. 2020

    YOLOv4

    Bochkovskiy et al. added CSPDarknet backbone, PANet neck, and a bag of training tricks (Mosaic augmentation, CIoU loss). Pushed accuracy without sacrificing speed.

  5. 2020

    DETR

    Facebook's DEtection TRansformer replaced anchors and NMS with a Transformer encoder-decoder and bipartite matching — a completely different paradigm.

  6. 2023

    YOLOv8

    Ultralytics unified detection, segmentation, and pose estimation in one framework. Anchor-free detection head, C2f modules, decoupled classification and localization.

CitationRedmon, Farhadi. YOLOv3: An Incremental Improvement. arXiv, 2018.

Terms in this paper