Computer Vision2014intermediate12 min read
Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation
هرميّات السمات الغنية للكشف الدقيق عن الأجسام والتجزئة الدلالية
Girshick, R. · Donahue, J. · Darrell, T. · Malik, J. — CVPR
The problem
By 2013, on PASCAL VOC had stagnated. The best systems relied on complex ensembles of hand-crafted features like HOG and SIFT, combined with deformable part models. Meanwhile, CNNs had just conquered image (AlexNet, 2012), but nobody had shown how to use those powerful learned features for detection — a harder task that requires not just recognizing what's in the image but also where each object is.
The contribution
R-: a simple, scalable pipeline that improved mAP on PASCAL VOC 2012 by over 30% relative (from ~40% to 53.3%). The recipe: (1) generate ~2000 class-agnostic region proposals via , (2) warp each to 227×227 and extract a 4096-d from a CNN pre-trained on ImageNet then fine-tuned for detection, (3) classify each region with per-class linear SVMs, (4) refine bounding boxes with learned regressors. The key insight: supervised on a large auxiliary dataset (ImageNet), followed by domain-specific , yields massive gains when labeled detection data is scarce.
The impact
R-CNN proved that CNN features trained for classification transfer powerfully to detection, igniting the modern era of deep-learning-based object detection. It spawned Fast R-CNN, Faster R-CNN, and Mask R-CNN, and its "pre-train then fine-tune" paradigm became the standard recipe across all of computer vision. Single-stage detectors like SSD and YOLO emerged partly as responses to R-CNN's speed bottleneck, but they all stand on the foundation R-CNN built: learned features beat hand-crafted ones for localization too.
Before R-CNN, detecting objects in an image was like searching for a friend in a stadium by checking every seat one by one with a magnifying glass — exhausting and slow.
R-CNN's insight is what you'd actually do: scan the crowd from afar to spot a few likely areas (region proposals), then walk over and look closely at each one with your best lens (a CNN).
Even smarter: R-CNN's lens was first trained to recognize thousands of objects in an encyclopedia (ImageNet). When transferred to the stadium task, it already knew what people look like — it just needed a quick refresher on this specific stadium's seating chart (fine-tuning).
The problem: detection stalled while classification soared
By 2013, object detection on PASCAL VOC had been stuck for years. The dominant approach — the (DPM) — built detectors from hand-crafted HOG features arranged in a star-shaped grammar of parts. DPM was a decade-old recipe: compute HOG descriptors, build part models, run sliding windows at multiple scales. It was slow, complex, and its accuracy had plateaued.
Meanwhile, AlexNet had just won ImageNet 2012 by a stunning margin, proving that deep CNNs learn far richer features than any hand-designed descriptor. The burning question was: can these powerful classification features also solve detection?
Detection is fundamentally harder than classification — you must predict what and where. A classifier only needs one label per image; a detector needs a and label for every object instance. This demands both recognition power and spatial precision.
Stage 1: Spotting candidate regions (Selective Search)
Instead of sliding a window over every possible location and scale — which would mean millions of evaluations — R-CNN uses Selective Search to propose roughly 2,000 candidate regions per image. Think of it as a fast pre-filter: before deploying the expensive CNN, a lightweight algorithm identifies bounding boxes that are likely to contain something.
Selective Search works bottom-up: it starts from a fine-grained over-segmentation of the image (many tiny regions), then iteratively merges similar neighboring regions based on color, texture, size, and containment. At every merge step, it records the bounding box of the merged region as a proposal. This hierarchical grouping naturally produces proposals at multiple scales — from small objects like cups to large ones like cars — without any .
The result is a set of roughly 2,000 class-agnostic bounding boxes. "Class-agnostic" means the algorithm doesn't know what the object is — it just flags regions that look like they contain something interesting. The what question comes later, when the CNN examines each region.
Stage 2: Seeing with borrowed eyes (CNN feature extraction)
Each of the ~2,000 proposed regions is warped (resized) to 227×227 pixels — the fixed input size that AlexNet expects — regardless of the region's original aspect ratio. This brute-force resizing distorts some objects, but R-CNN shows it works well enough in practice.
The warped region is then fed through a CNN (originally AlexNet, later VGGNet) that was pre-trained on ImageNet for the 1000-class classification task. But instead of taking the final classification output, R-CNN extracts the 4096-dimensional activation vector from the layer just before the final classifier (fc7). This vector is the region's feature representation — a compact summary of everything the CNN "sees" in that patch.
Here is the crucial insight about : the CNN was never trained on PASCAL VOC or on detection at all, yet its features are incredibly useful because ImageNet taught it to recognize edges, textures, parts, and objects. The lower layers detect universal visual patterns (edges, corners, color blobs), while higher layers detect object-specific features (eyes, wheels, fur). These features transfer to new tasks with remarkable effectiveness.
To bridge the gap between ImageNet classification and VOC detection, R-CNN fine-tunes the pre-trained CNN on the detection data. It replaces AlexNet's 1000-way classifier with a (K+1)-way classifier (K object classes + background) and continues with warped region proposals, using an IoU threshold of 0.5 to decide positive vs. negative examples.
The key insight: transfer learning bridges the data gap
The deepest lesson of R-CNN is not the pipeline itself but the transfer learning recipe it validated. ImageNet has 1.2 million labeled images for classification; PASCAL VOC has only a few thousand for detection. Training a deep CNN from scratch on VOC would overfit immediately.
R-CNN's solution: pre-train the CNN on ImageNet's abundant data, where it learns rich visual features — a universal "visual vocabulary". Then fine-tune on the small detection dataset, where the network adapts those general features to the specific task. This two-step recipe — pre-training then fine-tuning — turned out to be one of the most important ideas in , eventually extending to NLP (BERT, GPT) and beyond.
The ablation studies in the paper tell a striking story: without fine-tuning, R-CNN's CNN features already outperformed HOG-based methods. With fine-tuning, mAP jumped by another 8 points. The fine-tuned fc7 features were the most discriminative — proof that the highest layers adapt most to the new domain.
Stage 3: Classifying each region (linear SVMs)
After extracting a 4096-d feature vector from each , R-CNN feeds it into one linear SVM per object class. For a 20-class dataset like PASCAL VOC, that means 20 separate binary classifiers, each asking: "Does this region contain class X?"
Why SVMs instead of the CNN's own classifier? The paper explored this question carefully. The softmax classifier trained during fine-tuning actually performs slightly worse than SVMs. The reason: fine-tuning uses loosely labeled data (any region with IoU ≥ 0.5 is "positive"), while SVM training uses stricter labels (only ground-truth boxes are positive, and IoU < 0.3 is negative). This stricter positive/negative definition gives SVMs a sharper decision boundary.
Once all SVMs have scored every proposal for every class, (NMS) eliminates redundant overlapping detections. For each class independently, NMS ranks proposals by SVM score and iteratively discards any proposal whose IoU with a higher-scoring proposal exceeds a threshold (typically 0.3). The result: one clean bounding box per detected object.
Stage 4: Refining the box (bounding box regression)
Selective Search proposes rough bounding boxes, but they rarely align perfectly with the actual object. R-CNN adds a bounding box regressor — a simple linear model trained to predict four offsets () that shift and resize the proposed box to better fit the ground-truth object.
Think of it as fine-tuning the crop: the initial proposal says "there's probably a dog around here", and the regressor says "shift left by 5 pixels, shrink the height by 10%". This refinement typically boosts mAP by 3–4 points.
Measuring overlap: Intersection over Union (IoU)
R-CNN uses IoU everywhere: to decide if a proposal is positive or negative during training, to set the threshold for NMS, and to evaluate final detection quality. IoU measures how well a predicted bounding box aligns with the ground-truth box:
Training: a multi-stage recipe
R-CNN's training is a three-stage sequential process — one of its main drawbacks compared to later methods:
Stage A — Fine-tune the CNN. Take AlexNet pre-trained on ImageNet. Replace the 1000-way classifier with a randomly initialized (K+1)-way classifier. Train on warped region proposals: proposals with IoU ≥ 0.5 with any ground-truth box are positive; the rest are negative. Use with a of 0.001 (1/10 of the original ImageNet rate).
Stage B — Train SVMs. Freeze the fine-tuned CNN and extract features for all proposals. Train one linear SVM per class, using ground-truth boxes as positives and proposals with IoU < 0.3 as . Note the stricter labels compared to Stage A.
Stage C — Train bounding box regressors. Using the same frozen features, train per-class linear regressors to predict box offsets. Only proposals with IoU ≥ 0.6 with a ground-truth box are used as training examples.
This sequential pipeline is complex and disk-space-hungry (features must be cached for all proposals across all images). Fast R-CNN would later unify all three stages into a single end-to-end trainable network.
Results: a 30% leap
R-CNN's impact was immediate and dramatic:
On PASCAL VOC 2012, R-CNN achieved 53.3% mAP — a relative improvement of more than 30% over the previous best (DPM at ~40%).
On ILSVRC 2013 detection, R-CNN achieved 31.4% mAP, compared to the winner OverFeat's 24.3%.
The ablation studies revealed a clear pattern: (1) CNN features alone (pool5) already outperform HOG. (2) Higher layers (fc6, fc7) give progressively better features. (3) Fine-tuning adds ~8 mAP points. (4) Bounding box regression adds another 3–4 points. (5) Replacing AlexNet with the deeper VGGNet-16 pushed mAP to 66.0% on VOC 2007.
Beyond boxes: semantic segmentation
R-CNN also tackled — classifying every pixel in the image, not just drawing boxes. The approach: for each region proposal, compute CNN features on (a) the full rectangular region, (b) only the foreground pixels (with background masked out), and (c) the foreground of a slightly larger region. Concatenate these feature vectors and train an SVM per class.
This achieved competitive results on VOC 2011 segmentation, demonstrating that the same CNN features useful for detection also capture the information needed for pixel-level classification. Later work like Fully Convolutional Networks (FCN) would surpass this approach, but R-CNN showed the path.
Limitations and what came next
R-CNN had clear limitations that motivated its successors:
Speed. Running the CNN on ~2,000 proposals independently took ~47 seconds per image — far too slow for real-time use. Most computation was redundant since overlapping proposals share pixels.
Multi-stage training. Three separate training stages (CNN fine-tuning, SVM training, bbox regressor training) made the pipeline complex and hard to optimize jointly.
Disk space. Caching features for all proposals across all images required hundreds of gigabytes.
Fixed warping. Forcing every region to 227×227 destroys aspect ratio information that could help recognition.
These problems inspired a rapid succession of improvements:
2014
R-CNN — This Paper
Region proposals + CNN features + SVM classification. Proved CNN features transfer to detection. 53.3% mAP on VOC 2012.
2015
Fast R-CNN
Runs the CNN once on the whole image, then pools features per region (RoI Pooling). Unified training, 9× faster than R-CNN.
2015
Faster R-CNN
Replaces Selective Search with a learned Region Proposal Network (RPN). Near real-time detection. End-to-end trainable.
2016
SSD & YOLO — Single-Stage Detectors
Skip proposals entirely. Predict boxes and classes directly from feature maps. Real-time detection at the cost of some accuracy.
2017
Mask R-CNN
Extends Faster R-CNN with a mask branch for instance segmentation. Detects, classifies, and segments each object in one pass.
The complete R-CNN pipeline in code
Simplified to show the idea — not the real implementation.
import numpy as np
def selective_search(image):
"""Generate ~2000 region proposals (bounding boxes)."""
# Bottom-up segmentation → iterative merging by color,
# texture, size, containment → collect merged bounding boxes
return proposals # list of (x, y, w, h)
def extract_features(cnn, image, proposals):
"""Warp each proposal to 227×227 and extract fc7 features."""
features = []
for (x, y, w, h) in proposals:
region = image[y:y+h, x:x+w]
warped = resize(region, (227, 227)) # brute-force warp
feat = cnn.forward(warped, layer='fc7') # 4096-d vector
features.append(feat)
return np.array(features) # (2000, 4096)
def classify_and_detect(features, svms, regressors, proposals):
"""Score each proposal with per-class SVMs, then refine boxes."""
detections = []
for cls in range(num_classes):
scores = svms[cls].predict(features) # (2000,)
# Non-Maximum Suppression: keep best, discard overlaps
kept = nms(proposals, scores, iou_threshold=0.3)
# Bounding box regression: refine box coordinates
for idx in kept:
offsets = regressors[cls].predict(features[idx])
refined_box = apply_offsets(proposals[idx], offsets)
detections.append((cls, refined_box, scores[idx]))
return detectionsWhy R-CNN changed everything
CitationGirshick, Donahue, Darrell, Malik. Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation. CVPR, 2014.
Terms in this paper
- Object Detectionرصد وتحديد الكائنات
- Region Proposalالمنطقة المرشحة
- Selective Searchالبحث الانتقائي
- Transfer Learningنقل التعلم
- Fine-Tuningالضبط الدقيق
- Bounding Boxمربع الإحاطة
- Non-Maximum Suppressionكبت غير أعظمي
- Support Vector Machineآلة ناقلات الدعم (SVM)
- Feature Extractionاستخلاص السمات
- Mean Average Precisionمتوسط الدقة المتوسطة
- Intersection over Unionالتقاطع على الاتحاد
- Convolutional Neural Network (CNN)الشبكة العصبية الالتفافية