Computer Vision2015intermediate13 min read

You Only Look Once: Unified, Real-Time Object Detection

نظرة واحدة تكفي: نظام موحَّد لرصد الكائنات في الزمن الحقيقي

Redmon, J. · Divvala, S. · Girshick, R. · Farhadi, A. — CVPR

The problem

By 2015, systems like R- worked by generating thousands of region proposals, then classifying each one separately. This multi-stage pipeline was accurate but painfully slow — R-CNN took ~47 seconds per image, and even Fast R-CNN needed a separate step that became the speed bottleneck. These systems couldn't run in real-time, making them unusable for applications like autonomous driving, robotics, or live video analysis.

The contribution

YOLO reframes object detection as a single problem: one neural network predicts bounding boxes and class probabilities directly from full images in one evaluation. The image is divided into an S×S grid; each cell predicts B bounding boxes with confidence scores and C class probabilities. This unified architecture runs at 45 FPS (155 FPS for Fast YOLO), enabling while maintaining competitive accuracy. Because YOLO reasons globally about the image, it makes far fewer background mistakes than region-based methods.

The impact

YOLO launched the revolution, proving that speed and accuracy need not be traded off. It inspired SSD, YOLOv2–v8, RetinaNet, and every modern real-time detector. The idea that detection can be a single, end-to-end differentiable network changed the field's default assumption — from "propose then classify" to "predict everything at once." Today, YOLO descendants power autonomous vehicles, surveillance systems, medical imaging, and smartphone cameras worldwide.

Traditional object detectors work like a security guard checking every room in a building one by one, asking "is there something here?" at each door. By the time they finish, the intruder has moved.

YOLO is a security camera on the ceiling: one wide glance captures the entire floor plan — every person, every object, every location — in a single snapshot. It doesn't need to open doors one at a time; it sees everything at once.

That's the core trade: instead of careful, sequential search, YOLO bets that one smart look at the whole scene is faster and sufficient.

The problem: detection pipelines are slow and fragmented

Before YOLO, object detection was a multi-stage affair. The dominant approach — R-CNN and its descendants — followed a "propose then classify" pipeline:

  1. Generate region proposals: algorithms like would scan the image and suggest ~2000 candidate rectangles that might contain objects.
  2. Classify each region: a CNN would extract features from each proposal, then a classifier would decide what (if anything) was there.
  3. Refine boxes: post-processing would adjust coordinates and eliminate duplicates.

Each stage was trained separately with different objectives. The pipeline was accurate but suffered from two fatal flaws: it was slow (R-CNN: ~47 seconds/image, even Fast R-CNN couldn't reach real-time), and each component was optimized in isolation — the region proposer didn't know what the classifier needed, and vice versa.

Open in Lab
Compare the multi-stage R-CNN pipeline with YOLO's single-pass approach. Press Play to watch the difference in processing flow.
The demo wakes as you arrive…

The idea: detection as regression

YOLO's insight is radical in its simplicity: treat detection as a single regression problem. Instead of proposing and classifying regions, predict all bounding boxes and class probabilities for the entire image simultaneously, from raw pixels to final detections in one neural network evaluation.

Here's how it works. The input image is divided into an S × S grid (the paper uses 7 × 7 = 49 cells). Each is responsible for detecting objects whose center falls within it. Every cell predicts:

  • B bounding boxes (B = 2 in the paper), each with 5 values: center coordinates (x, y), width and height (w, h), and a
  • C class probabilities (C = 20 for PASCAL VOC)

The confidence score reflects both how likely the box contains an object and how accurate the predicted box is. Formally, it encodes Pr⁡(Object)×IoU\Pr(\text{Object}) \times \text{IoU}, where measures how well the predicted box overlaps the . If no object is present, confidence should be zero; if an object is present, confidence equals the IoU between predicted and actual boxes.

The final output is a single of size S × S × (B × 5 + C). For PASCAL VOC with S=7, B=2, C=20, that's 7 × 7 × 30 = 1470 values — the entire detection compressed into a single, dense prediction.

Open in Lab
Hover over grid cells to see what each one predicts. Each cell outputs 2 bounding boxes (with confidence) + 20 class probabilities.
The demo wakes as you arrive…

Measuring overlap: Intersection over Union

How do you measure whether a predicted bounding box is "correct"? You need a number that is 1.0 when the predicted box perfectly overlaps the ground truth and 0.0 when they don't overlap at all. That number is IoU ().

Think of it as a Venn diagram: the intersection is the area where both boxes overlap; the union is the total area covered by either box. IoU = intersection ÷ union. An IoU of 0.5 or above is typically considered a "correct" detection.

In YOLO, IoU plays a dual role. During , it is the ground truth target for the confidence score: the network learns to predict how well its own boxes will overlap the real objects. During evaluation, IoU determines whether a predicted box counts as a or a .

IoU=Area of IntersectionArea of Union\text{IoU} = \frac{\text{Area of Intersection}}{\text{Area of Union}}
IoU — the universal overlap metric for bounding boxes — Ranges from 0 (no overlap) to 1 (perfect overlap). A detection is typically "correct" at IoU ≥ 0.5. YOLO trains the confidence score to predict this value.
Open in Lab
Drag the predicted box to see how IoU changes. Notice how even small misalignments reduce IoU significantly.
The demo wakes as you arrive…

The network architecture

YOLO's network is inspired by GoogLeNet but replaces inception modules with simpler 1×1 reduction layers followed by 3×3 convolutional layers. The architecture has:

  • 24 convolutional layers that extract features at increasing levels of abstraction — from edges in the early layers to object parts and complete shapes in the deeper layers
  • 2 fully connected layers that take the final and produce the S × S × (B × 5 + C) output tensor

The network takes a 448 × 448 pixel image as input (larger than the 224 × 224 used by most classifiers, because detection needs finer spatial detail). The convolutional layers progressively shrink the spatial dimensions while growing the channel depth, building a rich, hierarchical representation.

There's also a fast version — Fast YOLO — with only 9 convolutional layers and fewer filters. It sacrifices some accuracy for extreme speed: 155 FPS, making it one of the fastest detectors ever published.

Open in Lab
Click any layer group to see its role in the feature extraction pipeline.
The demo wakes as you arrive…

The loss function: one formula to train them all

Since YOLO is a single network, it needs a single that simultaneously teaches it three things: where the objects are (localization), how confident each detection is (objectness), and what each object is ().

The loss uses sum-squared error for speed, but with careful weighting to handle a key imbalance: most grid cells contain no object. Without correction, the overwhelming "no object" signal would drown out the real detections. YOLO addresses this with two balancing weights:

  • λcoord=5\lambda_{\text{coord}} = 5 — amplifies the localization loss so the network pays more attention to getting boxes right
  • λnoobj=0.5\lambda_{\text{noobj}} = 0.5 — dampens the confidence loss for cells without objects so they don't dominate training

A second subtlety: the loss uses square root of width and height rather than raw values. Why? A 2-pixel error matters much more in a 10-pixel box than in a 200-pixel box. Taking the square root compresses large values, making the loss more sensitive to small-object errors — a clever normalization trick.

L=  λcoord∑i∑j1ijobj[(xi−x^i)2+(yi−y^i)2+(wi−w^i)2+(hi−h^i)2]+∑i∑j1ijobj(Ci−C^i)2+λnoobj∑i∑j1ijnoobj(Ci−C^i)2+∑i1iobj∑c∈classes(pi(c)−p^i(c))2\begin{aligned} \mathcal{L} =\;& \lambda_{\text{coord}} \sum_i \sum_j \mathbb{1}_{ij}^{\text{obj}} \Big[ (x_i-\hat{x}_i)^2 + (y_i-\hat{y}_i)^2 \\ &\qquad + (\sqrt{w_i}-\sqrt{\hat{w}_i})^2 + (\sqrt{h_i}-\sqrt{\hat{h}_i})^2 \Big] \\ &+ \sum_i \sum_j \mathbb{1}_{ij}^{\text{obj}} (C_i-\hat{C}_i)^2 \\ &+ \lambda_{\text{noobj}} \sum_i \sum_j \mathbb{1}_{ij}^{\text{noobj}} (C_i-\hat{C}_i)^2 \\ &+ \sum_i \mathbb{1}_i^{\text{obj}} \sum_{c\in\text{classes}} (p_i(c)-\hat{p}_i(c))^2 \end{aligned}
YOLO's multi-part loss — localization + objectness + classification — Three parts summed: (1) box coordinate loss ×5 for cells with objects, using √w,√h for scale invariance · (2) confidence loss for boxes with objects · (3) dampened confidence loss ×0.5 for boxes without objects · (4) class probability loss for cells with objects
Open in Lab
Toggle each loss component on/off to see how it contributes. Notice how λ_coord and λ_noobj rebalance the training signal.
The demo wakes as you arrive…

Cleaning up: non-maximum suppression

YOLO predicts 98 bounding boxes per image (7 × 7 cells × 2 boxes each). Many of these overlap heavily for the same object — several neighboring cells may detect the same car, producing redundant boxes.

(NMS) cleans this up in three steps:

  1. Discard all boxes with confidence below a threshold (e.g., 0.25)
  2. Select the box with the highest confidence
  3. Remove any remaining box that overlaps with the selected box beyond an IoU threshold (e.g., 0.5) — it's probably detecting the same object

Repeat steps 2–3 until no boxes remain. The result is a clean set of detections, usually just one box per object.

Think of it as a classroom where multiple students raise their hands to answer the same question. NMS picks the most confident student and tells the others with similar answers to put their hands down.

Open in Lab
Step through NMS — watch redundant boxes get suppressed one by one.
The demo wakes as you arrive…

Training strategy and design choices

Several training details make YOLO work in practice:

Pre-training on ImageNet: the first 20 convolutional layers are pre-trained on ImageNet classification at 224 × 224. This gives the network strong general-purpose extractors before it ever sees a detection task. The resolution is then doubled to 448 × 448 for detection , because detecting objects requires more spatial detail than classification.

activations: all layers use Leaky (slope 0.1 for negative values) instead of standard ReLU. This prevents "dead neurons" — units that stop learning because their gradient is zero for all negative inputs.

Aggressive : random scaling, translations up to 20% of the image size, exposure and saturation adjustments in HSV color space. These augmentations force the network to be robust to varying object sizes, positions, and lighting conditions.

assignment: when multiple bounding boxes are predicted for the same cell, only the one with the highest IoU to the ground truth is designated the "responsible" predictor. This specialization encourages each predictor to get better at certain aspect ratios and object sizes over time.

Strengths and limitations

Speed vs. accuracy: the real-time frontier

YOLO's contribution is best understood through the speed-accuracy landscape of its era. At one extreme, DPM (Deformable Parts Model) achieved 33.7 mAP at less than 1 FPS. At the other extreme, Fast YOLO achieved 52.7 mAP at 155 FPS.

The key comparison is with Fast R-CNN: it scored 70.0 mAP but at only 0.5 FPS. YOLO scored 63.4 mAP at 45 FPS — a modest accuracy drop for a 90× speed increase. And when YOLO's detections were combined with Fast R-CNN (using YOLO to eliminate background false positives), the combination hit 75.0 mAP — the best result on VOC 2007 at the time.

This showed something profound: YOLO and R-CNN make complementary errors. R-CNN is precise but fooled by backgrounds; YOLO is fast and globally aware. Together they're stronger than either alone.

Open in Lab
Each dot is a detection system. The ideal is the top-right corner — high mAP *and* high FPS. Notice YOLO's unique position in the speed frontier.
The demo wakes as you arrive…

The idea in code

YOLO's grid-based prediction — the core conceptpython

Simplified to show the idea — not the real implementation.

import numpy as np

def yolo_predict(image, model, S=7, B=2, C=20):
    """
    Run YOLO on one image.
    Returns: S×S×(B*5 + C) tensor of predictions.
    Each cell: [x1,y1,w1,h1,conf1, x2,y2,w2,h2,conf2, p1,...,p20]
    """
    # The entire image → one forward pass → one tensor
    prediction = model(image)  # shape: (S, S, B*5 + C) = (7, 7, 30)
    return prediction

def decode_predictions(pred, S=7, B=2, C=20, conf_thresh=0.25):
    """Decode YOLO output into usable bounding boxes."""
    boxes = []
    for row in range(S):
        for col in range(S):
            cell = pred[row, col]
            class_probs = cell[B*5:]  # last C values = class probabilities

            for b in range(B):
                offset = b * 5
                x, y = cell[offset], cell[offset+1]     # center (relative to cell)
                w, h = cell[offset+2], cell[offset+3]   # size (relative to image)
                conf = cell[offset+4]                     # P(object) × IoU

                # Final score = confidence × class probability
                scores = conf * class_probs
                best_class = np.argmax(scores)
                best_score = scores[best_class]

                if best_score > conf_thresh:
                    # Convert cell-relative coords to image-relative
                    abs_x = (col + x) / S
                    abs_y = (row + y) / S
                    boxes.append((abs_x, abs_y, w, h, best_score, best_class))
    return boxes

def nms(boxes, iou_thresh=0.5):
    """Non-maximum suppression: keep only the best box per object."""
    boxes = sorted(boxes, key=lambda b: b[4], reverse=True)  # sort by score
    keep = []
    while boxes:
        best = boxes.pop(0)
        keep.append(best)
        boxes = [b for b in boxes if iou(best, b) < iou_thresh]
    return keep

# That's the whole pipeline:
# 1. One forward pass → S×S×30 tensor
# 2. Decode into boxes with scores
# 3. NMS to remove duplicates
# No region proposals. No multi-stage pipeline. Just look once.

Why it mattered

  1. 2014

    R-CNN — region-based detection begins

    Girshick et al. proposed extracting ~2000 region proposals, running CNN features on each, then classifying with SVMs. Accurate but 47 seconds per image.

  2. 2015

    Faster R-CNN — learnable proposals

    Ren et al. replaced Selective Search with a Region Proposal Network (RPN), making proposals part of the network. Faster, but still two-stage.

  3. 2015

    YOLO — the single-stage revolution

    Redmon et al. reframed detection as regression. One look, one network, real-time speed. Changed the default assumption of the field.

  4. 2016

    SSD — multi-scale single-stage detection

    Liu et al. predicted from multiple feature maps at different resolutions, improving small object detection while keeping single-stage speed.

  5. 2017

    YOLOv2/YOLO9000 — better, faster, stronger

    Added batch normalization, anchor boxes, multi-scale training, and passthrough layers. YOLO9000 could detect 9000+ categories using hierarchical classification.

  6. 2017

    RetinaNet — focal loss solves class imbalance

    Lin et al. showed single-stage detectors lagged in accuracy because of extreme foreground-background imbalance. Focal loss down-weighted easy negatives, closing the gap with two-stage detectors.

  7. 2018

    YOLOv3 — multi-scale predictions

    Redmon added predictions at three scales using a feature pyramid, greatly improving small object detection while maintaining real-time speed.

  8. 2023

    YOLOv8 — modern unified framework

    Ultralytics released YOLOv8 with an anchor-free design, decoupled detection heads, and a clean API. YOLO became both a model family and an industry standard.

YOLO didn't just make detection faster — it changed what detection could be. Before YOLO, real-time object detection was considered impractical. After YOLO, it became the baseline expectation. Every single-stage detector, from SSD to YOLOv3 to RetinaNet, traces its lineage to this paper's central bet: one look is enough.

CitationRedmon, Divvala, Girshick, Farhadi. You Only Look Once: Unified, Real-Time Object Detection. CVPR, 2016.

Terms in this paper