Computer Vision2015intermediate13 min read
You Only Look Once: Unified, Real-Time Object Detection
نظرة واحدة تكفي: نظام موحَّد لرصد الكائنات في الزمن الحقيقي
Redmon, J. · Divvala, S. · Girshick, R. · Farhadi, A. — CVPR
The problem
By 2015, systems like R- worked by generating thousands of region proposals, then classifying each one separately. This multi-stage pipeline was accurate but painfully slow — R-CNN took ~47 seconds per image, and even Fast R-CNN needed a separate step that became the speed bottleneck. These systems couldn't run in real-time, making them unusable for applications like autonomous driving, robotics, or live video analysis.
The contribution
YOLO reframes object detection as a single problem: one neural network predicts bounding boxes and class probabilities directly from full images in one evaluation. The image is divided into an S×S grid; each cell predicts B bounding boxes with confidence scores and C class probabilities. This unified architecture runs at 45 FPS (155 FPS for Fast YOLO), enabling while maintaining competitive accuracy. Because YOLO reasons globally about the image, it makes far fewer background mistakes than region-based methods.
The impact
YOLO launched the revolution, proving that speed and accuracy need not be traded off. It inspired SSD, YOLOv2–v8, RetinaNet, and every modern real-time detector. The idea that detection can be a single, end-to-end differentiable network changed the field's default assumption — from "propose then classify" to "predict everything at once." Today, YOLO descendants power autonomous vehicles, surveillance systems, medical imaging, and smartphone cameras worldwide.
Traditional object detectors work like a security guard checking every room in a building one by one, asking "is there something here?" at each door. By the time they finish, the intruder has moved.
YOLO is a security camera on the ceiling: one wide glance captures the entire floor plan — every person, every object, every location — in a single snapshot. It doesn't need to open doors one at a time; it sees everything at once.
That's the core trade: instead of careful, sequential search, YOLO bets that one smart look at the whole scene is faster and sufficient.
The problem: detection pipelines are slow and fragmented
Before YOLO, object detection was a multi-stage affair. The dominant approach — R-CNN and its descendants — followed a "propose then classify" pipeline:
- Generate region proposals: algorithms like would scan the image and suggest ~2000 candidate rectangles that might contain objects.
- Classify each region: a CNN would extract features from each proposal, then a classifier would decide what (if anything) was there.
- Refine boxes: post-processing would adjust coordinates and eliminate duplicates.
Each stage was trained separately with different objectives. The pipeline was accurate but suffered from two fatal flaws: it was slow (R-CNN: ~47 seconds/image, even Fast R-CNN couldn't reach real-time), and each component was optimized in isolation — the region proposer didn't know what the classifier needed, and vice versa.
The idea: detection as regression
YOLO's insight is radical in its simplicity: treat detection as a single regression problem. Instead of proposing and classifying regions, predict all bounding boxes and class probabilities for the entire image simultaneously, from raw pixels to final detections in one neural network evaluation.
Here's how it works. The input image is divided into an S × S grid (the paper uses 7 × 7 = 49 cells). Each is responsible for detecting objects whose center falls within it. Every cell predicts:
- B bounding boxes (B = 2 in the paper), each with 5 values: center coordinates (x, y), width and height (w, h), and a
- C class probabilities (C = 20 for PASCAL VOC)
The confidence score reflects both how likely the box contains an object and how accurate the predicted box is. Formally, it encodes , where measures how well the predicted box overlaps the . If no object is present, confidence should be zero; if an object is present, confidence equals the IoU between predicted and actual boxes.
The final output is a single of size S × S × (B × 5 + C). For PASCAL VOC with S=7, B=2, C=20, that's 7 × 7 × 30 = 1470 values — the entire detection compressed into a single, dense prediction.
Measuring overlap: Intersection over Union
How do you measure whether a predicted bounding box is "correct"? You need a number that is 1.0 when the predicted box perfectly overlaps the ground truth and 0.0 when they don't overlap at all. That number is IoU ().
Think of it as a Venn diagram: the intersection is the area where both boxes overlap; the union is the total area covered by either box. IoU = intersection ÷ union. An IoU of 0.5 or above is typically considered a "correct" detection.
In YOLO, IoU plays a dual role. During , it is the ground truth target for the confidence score: the network learns to predict how well its own boxes will overlap the real objects. During evaluation, IoU determines whether a predicted box counts as a or a .
The network architecture
YOLO's network is inspired by GoogLeNet but replaces inception modules with simpler 1×1 reduction layers followed by 3×3 convolutional layers. The architecture has:
- 24 convolutional layers that extract features at increasing levels of abstraction — from edges in the early layers to object parts and complete shapes in the deeper layers
- 2 fully connected layers that take the final and produce the S × S × (B × 5 + C) output tensor
The network takes a 448 × 448 pixel image as input (larger than the 224 × 224 used by most classifiers, because detection needs finer spatial detail). The convolutional layers progressively shrink the spatial dimensions while growing the channel depth, building a rich, hierarchical representation.
There's also a fast version — Fast YOLO — with only 9 convolutional layers and fewer filters. It sacrifices some accuracy for extreme speed: 155 FPS, making it one of the fastest detectors ever published.
The loss function: one formula to train them all
Since YOLO is a single network, it needs a single that simultaneously teaches it three things: where the objects are (localization), how confident each detection is (objectness), and what each object is ().
The loss uses sum-squared error for speed, but with careful weighting to handle a key imbalance: most grid cells contain no object. Without correction, the overwhelming "no object" signal would drown out the real detections. YOLO addresses this with two balancing weights:
- — amplifies the localization loss so the network pays more attention to getting boxes right
- — dampens the confidence loss for cells without objects so they don't dominate training
A second subtlety: the loss uses square root of width and height rather than raw values. Why? A 2-pixel error matters much more in a 10-pixel box than in a 200-pixel box. Taking the square root compresses large values, making the loss more sensitive to small-object errors — a clever normalization trick.
Cleaning up: non-maximum suppression
YOLO predicts 98 bounding boxes per image (7 × 7 cells × 2 boxes each). Many of these overlap heavily for the same object — several neighboring cells may detect the same car, producing redundant boxes.
(NMS) cleans this up in three steps:
- Discard all boxes with confidence below a threshold (e.g., 0.25)
- Select the box with the highest confidence
- Remove any remaining box that overlaps with the selected box beyond an IoU threshold (e.g., 0.5) — it's probably detecting the same object
Repeat steps 2–3 until no boxes remain. The result is a clean set of detections, usually just one box per object.
Think of it as a classroom where multiple students raise their hands to answer the same question. NMS picks the most confident student and tells the others with similar answers to put their hands down.
Training strategy and design choices
Several training details make YOLO work in practice:
Pre-training on ImageNet: the first 20 convolutional layers are pre-trained on ImageNet classification at 224 × 224. This gives the network strong general-purpose extractors before it ever sees a detection task. The resolution is then doubled to 448 × 448 for detection , because detecting objects requires more spatial detail than classification.
activations: all layers use Leaky (slope 0.1 for negative values) instead of standard ReLU. This prevents "dead neurons" — units that stop learning because their gradient is zero for all negative inputs.
Aggressive : random scaling, translations up to 20% of the image size, exposure and saturation adjustments in HSV color space. These augmentations force the network to be robust to varying object sizes, positions, and lighting conditions.
assignment: when multiple bounding boxes are predicted for the same cell, only the one with the highest IoU to the ground truth is designated the "responsible" predictor. This specialization encourages each predictor to get better at certain aspect ratios and object sizes over time.
Strengths and limitations
Speed vs. accuracy: the real-time frontier
YOLO's contribution is best understood through the speed-accuracy landscape of its era. At one extreme, DPM (Deformable Parts Model) achieved 33.7 mAP at less than 1 FPS. At the other extreme, Fast YOLO achieved 52.7 mAP at 155 FPS.
The key comparison is with Fast R-CNN: it scored 70.0 mAP but at only 0.5 FPS. YOLO scored 63.4 mAP at 45 FPS — a modest accuracy drop for a 90× speed increase. And when YOLO's detections were combined with Fast R-CNN (using YOLO to eliminate background false positives), the combination hit 75.0 mAP — the best result on VOC 2007 at the time.
This showed something profound: YOLO and R-CNN make complementary errors. R-CNN is precise but fooled by backgrounds; YOLO is fast and globally aware. Together they're stronger than either alone.
The idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def yolo_predict(image, model, S=7, B=2, C=20):
"""
Run YOLO on one image.
Returns: S×S×(B*5 + C) tensor of predictions.
Each cell: [x1,y1,w1,h1,conf1, x2,y2,w2,h2,conf2, p1,...,p20]
"""
# The entire image → one forward pass → one tensor
prediction = model(image) # shape: (S, S, B*5 + C) = (7, 7, 30)
return prediction
def decode_predictions(pred, S=7, B=2, C=20, conf_thresh=0.25):
"""Decode YOLO output into usable bounding boxes."""
boxes = []
for row in range(S):
for col in range(S):
cell = pred[row, col]
class_probs = cell[B*5:] # last C values = class probabilities
for b in range(B):
offset = b * 5
x, y = cell[offset], cell[offset+1] # center (relative to cell)
w, h = cell[offset+2], cell[offset+3] # size (relative to image)
conf = cell[offset+4] # P(object) × IoU
# Final score = confidence × class probability
scores = conf * class_probs
best_class = np.argmax(scores)
best_score = scores[best_class]
if best_score > conf_thresh:
# Convert cell-relative coords to image-relative
abs_x = (col + x) / S
abs_y = (row + y) / S
boxes.append((abs_x, abs_y, w, h, best_score, best_class))
return boxes
def nms(boxes, iou_thresh=0.5):
"""Non-maximum suppression: keep only the best box per object."""
boxes = sorted(boxes, key=lambda b: b[4], reverse=True) # sort by score
keep = []
while boxes:
best = boxes.pop(0)
keep.append(best)
boxes = [b for b in boxes if iou(best, b) < iou_thresh]
return keep
# That's the whole pipeline:
# 1. One forward pass → S×S×30 tensor
# 2. Decode into boxes with scores
# 3. NMS to remove duplicates
# No region proposals. No multi-stage pipeline. Just look once.Why it mattered
2014
R-CNN — region-based detection begins
Girshick et al. proposed extracting ~2000 region proposals, running CNN features on each, then classifying with SVMs. Accurate but 47 seconds per image.
2015
Faster R-CNN — learnable proposals
Ren et al. replaced Selective Search with a Region Proposal Network (RPN), making proposals part of the network. Faster, but still two-stage.
2015
YOLO — the single-stage revolution
Redmon et al. reframed detection as regression. One look, one network, real-time speed. Changed the default assumption of the field.
2016
SSD — multi-scale single-stage detection
Liu et al. predicted from multiple feature maps at different resolutions, improving small object detection while keeping single-stage speed.
2017
YOLOv2/YOLO9000 — better, faster, stronger
Added batch normalization, anchor boxes, multi-scale training, and passthrough layers. YOLO9000 could detect 9000+ categories using hierarchical classification.
2017
RetinaNet — focal loss solves class imbalance
Lin et al. showed single-stage detectors lagged in accuracy because of extreme foreground-background imbalance. Focal loss down-weighted easy negatives, closing the gap with two-stage detectors.
2018
YOLOv3 — multi-scale predictions
Redmon added predictions at three scales using a feature pyramid, greatly improving small object detection while maintaining real-time speed.
2023
YOLOv8 — modern unified framework
Ultralytics released YOLOv8 with an anchor-free design, decoupled detection heads, and a clean API. YOLO became both a model family and an industry standard.
YOLO didn't just make detection faster — it changed what detection could be. Before YOLO, real-time object detection was considered impractical. After YOLO, it became the baseline expectation. Every single-stage detector, from SSD to YOLOv3 to RetinaNet, traces its lineage to this paper's central bet: one look is enough.
CitationRedmon, Divvala, Girshick, Farhadi. You Only Look Once: Unified, Real-Time Object Detection. CVPR, 2016.
Terms in this paper
- Object Detectionرصد وتحديد الكائنات
- Bounding Boxمربع الإحاطة
- Grid Cellخلية الشبكة
- Confidence Scoreدرجة الثقة
- Non-Maximum Suppressionكبت غير أعظمي
- Intersection over Unionالتقاطع على الاتحاد
- Single-Stage Detectorالكاشف أحادي المرحلة
- Real-Time Detectionالرصد في الزمن الحقيقي
- Responsible Predictorالمتنبئ المسؤول
- End-to-End Detectionالرصد من طرف إلى طرف