Computer Vision2017intermediate11 min read
Focal Loss for Dense Object Detection
دالة الخسارة البؤرية للكشف الكثيف عن الأجسام
Lin, T.-Y. · Goyal, P. · Girshick, R. · He, K. · Dollár, P. — ICCV
The problem
One-stage object detectors like SSD and YOLO are fast because they evaluate a dense grid of candidate locations in a single pass — no proposal stage needed. But they consistently trailed the accuracy of slower two-stage detectors like Faster R-CNN. Why? Because a typical image produces ~100,000 candidate locations, of which only a handful contain actual objects. This extreme foreground-background imbalance (roughly 1:1000) means the loss signal is drowned by tens of thousands of easy negatives — background patches the model already classifies correctly but whose small individual losses, summed together, overwhelm the meaningful gradients from the rare hard examples.
The contribution
: a simple reshaping of cross-entropy that adds a (1 − p_t)^γ. When the model is already confident (p_t is high), the factor shrinks the loss toward zero; when the model struggles (p_t is low), the loss is barely touched. This automatically down-weights the flood of easy negatives and focuses on the hard, informative examples. Combined with a clean one-stage detector called RetinaNet — built on a with two simple subnetworks for and box regression — focal loss closes the accuracy gap: RetinaNet matches two-stage speed while surpassing two-stage accuracy for the first time.
The impact
Focal loss proved that the accuracy gap between one-stage and two-stage detectors was not an architectural problem — it was a training problem. The idea spread far beyond : focal loss is now standard in medical imaging, NLP with imbalanced labels, dense prediction tasks, and any setting where easy examples dominate training. RetinaNet's architecture also popularized using FPN as a universal backbone for dense prediction.
Imagine a search-and-rescue team scanning a city after an earthquake. 99% of the buildings are clearly intact — a glance is enough. But the team keeps filing detailed "all clear" reports for every safe building, and by the time they reach the few collapsed ones, they're too exhausted to investigate properly.
Focal loss is like a new protocol: "If a building is obviously fine, spend zero time on the report. Save all your energy for the ones that are hard to assess."
The result: the same team, same equipment, but now they find every survivor — because they stopped wasting effort on what was already obvious.
The problem: easy negatives drown the signal
Object detection works by examining thousands of candidate boxes across the image. A like Faster R-CNN handles this in two phases: first a Region Proposal Network filters ~100k raw anchors down to ~2,000 promising ones, then a classifier examines only those 2,000. This filtering elegantly avoids the imbalance problem — most background is thrown away before it can affect training.
One-stage detectors skip the filtering. They classify all ~100k anchors at once. This is faster, but it means the training loss is computed over all 100k, and roughly 99.9% are easy background. The cross-entropy loss treats every example equally: an the model already classifies correctly at 99% confidence still contributes a small but nonzero loss. Multiply that small loss by 99,000 easy negatives, and it drowns the signal from the ~100 anchors that actually contain objects.
Starting point: cross-entropy and its weakness
The standard cross-entropy loss for binary classification is:
where if the , and otherwise. In other words, is the model's estimated probability for the correct class.
The key weakness: even when is high (say 0.9 — a correctly classified, easy example), is still a nontrivial loss. One such example is harmless; 99,000 of them contribute a collective loss of ~10,000, which overwhelms the few hundred hard positives.
A common patch is α-balancing: weighting the loss by class frequency. Setting α = 0.75 for the rare foreground class and 0.25 for the common background class helps, but it only rebalances how much each class matters — it doesn't distinguish easy from hard within each class.
Focal loss: turn down the volume on easy examples
The insight is elegant: multiply the cross-entropy by a modulating factor that automatically fades to zero as the model becomes more confident. Easy examples suppress themselves; hard examples stay loud.
Let's make this concrete with numbers. Consider an easy negative with and γ = 2:
- Cross-entropy:
- Focal loss factor:
- Focal loss: — 100× smaller
Now a hard example with :
- Cross-entropy:
- Focal loss factor:
- Focal loss: — barely reduced
The easy example's contribution drops by two orders of magnitude. The hard example's loss stays almost intact. This is the whole mechanism: a single multiplicative term that converts cross-entropy from a "count every vote equally" function into a "listen to the struggling students" function.
Where does the loss come from?
The paper includes a revealing analysis (Figure 4): after training, plot the cumulative distribution of the loss across all foreground and background examples.
With standard cross-entropy (γ = 0), loss is spread nearly uniformly — the flood of easy backgrounds contributes roughly as much total loss as the hard cases. With focal loss (γ = 2), the picture changes dramatically: for background examples, nearly all the loss is concentrated in the hardest few percent. The easy 90%+ contribute almost nothing.
For foreground examples, the distribution barely changes — focal loss doesn't penalize hard foreground examples, it only silences the background noise.
RetinaNet: a clean one-stage detector
To prove that focal loss — not architecture — was the missing ingredient, the authors deliberately kept the detector simple. RetinaNet has three parts:
1. Backbone: Feature Pyramid Network (FPN) — built on top of ResNet, FPN constructs a multi-scale feature pyramid. Lower levels (P3) have high resolution for small objects; upper levels (P7) have low resolution but large receptive fields for big objects. Lateral connections blend fine spatial detail from bottom-up features with semantic richness from top-down features.
2. Classification subnet — a small fully-convolutional network (4 conv layers + sigmoid) attached to each FPN level. For every spatial position and every anchor (A = 9), it predicts K class probabilities. Parameters are shared across all pyramid levels.
3. Box regression subnet — an identical structure but predicting 4 offset values (dx, dy, dw, dh) per anchor, specifying how to shift and resize each to fit the detected object.
Anchors: tiling the search space
At every spatial position on every FPN level, RetinaNet places A = 9 anchor boxes: 3 aspect ratios (1:2, 1:1, 2:1) × 3 scales (, , ). These anchors serve as reference templates. The box regression subnet doesn't predict absolute coordinates — it predicts small adjustments to these anchors.
Anchors with IoU ≥ 0.5 against a ground-truth box are assigned as positives; those with IoU < 0.4 are negatives; the rest are ignored during training. This yields ~100k anchors per image, of which typically only ~10–100 are positive.
A critical training trick: the final classification layer is initialized with where π = 0.01, so every anchor starts with a prior probability of only 1% for foreground. Without this, the massive initial loss from 100k anchors all predicting 50% foreground would destabilize training.
One-stage vs. two-stage: what changed
Before this paper, the standard explanation for the one-stage accuracy gap was: "Two-stage detectors have better architecture." The region proposal network provides spatial attention — it tells the classifier where to look. One-stage detectors lack this.
The paper's key insight is different: the gap is not about where you look — it's about how you train. Two-stage detectors achieve class balance implicitly: the proposal stage discards 99% of easy negatives, and the second stage uses a 1:3 positive-to-negative sampling ratio. Focal loss achieves the same balance explicitly, through the itself.
The proof is in the numbers: RetinaNet with focal loss (γ = 2, α = 0.25) achieves 36.0 AP on COCO minival — far above the best OHEM variant at 32.8 AP, and above Faster R-CNN with FPN at 36.2 AP. A cleaner approach to training beat architectural sophistication.
Focal loss vs. hard negative mining
Before focal loss, the main strategy for handling imbalance in one-stage detectors was Online Hard Example Mining (OHEM): sort examples by loss, apply , then keep only the highest-loss ones for each mini-batch.
OHEM has two weaknesses: (1) it completely discards easy examples, losing any remaining learning signal they might offer; (2) it introduces extra hyperparameters — NMS threshold and batch size — that require tuning. The paper tests many OHEM configurations (batch sizes 128, 256, 512; NMS thresholds 0.5, 0.7; with and without 1:3 ratio enforcement) and finds the best reaches 32.8 AP. Focal loss reaches 36.0 with no sampling, no NMS on training data, and only two hyperparameters (γ, α) that are robust across a wide range.
The difference is philosophical: OHEM is a sampling strategy that selects what to train on. Focal loss is a continuous reweighting that trains on everything but listens more carefully to what matters.
The same idea in code
Simplified to show the idea — not the real implementation.
import torch
import torch.nn.functional as F
def focal_loss(logits, targets, gamma=2.0, alpha=0.25):
"""
logits: (N,) raw model outputs before sigmoid
targets: (N,) binary labels {0, 1}
"""
p = torch.sigmoid(logits) # predicted probability
ce = F.binary_cross_entropy_with_logits( # numerically stable CE
logits, targets, reduction='none'
)
p_t = p * targets + (1 - p) * (1 - targets) # prob of correct class
# The magic: this factor → 0 for easy examples, ≈ 1 for hard ones
modulating_factor = (1 - p_t) ** gamma
# α-balancing: slightly higher weight for rare foreground class
alpha_t = alpha * targets + (1 - alpha) * (1 - targets)
loss = alpha_t * modulating_factor * ce # focal loss per example
return loss.sum() # sum (not mean) — paper convention
# That's it. The entire contribution is one multiplicative term: (1 - p_t)^γ
# Standard CE is just focal_loss(..., gamma=0, alpha=0.5)Why it mattered
2016
SSD — fast but inaccurate
Liu et al. introduce the Single Shot MultiBox Detector — a one-stage detector using multi-scale feature maps. Fast inference but 10–20% lower AP than two-stage methods due to class imbalance during training.
2017
FPN — multi-scale feature fusion
Lin et al. propose the Feature Pyramid Network — a top-down architecture with lateral connections that builds rich multi-scale features from any backbone. Becomes the standard backbone for both one-stage and two-stage detectors.
2017
RetinaNet + Focal Loss
This paper. Focal loss solves the class imbalance problem via loss design. RetinaNet, a simple FPN-based one-stage detector, surpasses all existing two-stage detectors for the first time, reaching 39.1 AP on COCO test-dev.
2020
DETR — end-to-end detection with Transformers
Carion et al. replace anchors and NMS with a Transformer encoder-decoder and Hungarian matching. Eliminates hand-designed components entirely, but focal loss remains relevant in many DETR variants to handle residual imbalance.
2024
Focal loss everywhere
Focal loss becomes a standard tool far beyond object detection — medical imaging (tumor vs. healthy tissue), NLP (rare-label classification), dense prediction, and any task with extreme label imbalance.
CitationLin, Goyal, Girshick, He, Dollár. Focal Loss for Dense Object Detection. ICCV, 2017.
Terms in this paper
- Focal Lossخسارة بؤرية
- Class Imbalanceاختلال التوازن بين الفئات
- Single-Stage Detectorالكاشف أحادي المرحلة
- Object Detectionرصد وتحديد الكائنات
- Feature Pyramid Networkشبكة الهرم الاستخلاصي للسمات
- Anchor Boxمربعات الإحاطة المرجعية (المرساة)
- Cross Entropyالعشوائية المتقاطعة
- Hard Negative Miningالتنقيب عن السلبيات الصعبة
- Bounding Boxمربع الإحاطة
- Non-Maximum Suppressionكبت غير أعظمي