Computer Vision2005intermediate12 min read
Histograms of Oriented Gradients for Human Detection
مُدرَّجات الاتجاهات التكرارية لاكتشاف البشر
Dalal, N. · Triggs, B. — CVPR
The problem
Before 2005, detecting people in photographs was unreliable. Existing descriptors like Haar wavelets (used in Viola-Jones face detection) struggled with the high variability of human poses, clothing, and backgrounds. Edge-based features existed but were too sensitive to lighting changes, and no single hand-crafted feature set could robustly capture the human silhouette across different conditions.
The contribution
HOG: a dense feature descriptor that divides the image into small cells, computes a histogram of orientations in each cell, groups cells into overlapping blocks for contrast , and concatenates everything into a single high-dimensional feature . Paired with a linear SVM, this pipeline achieved near-perfect on the MIT dataset and set a new standard. The key insight: local gradient orientation distributions, with proper normalization, capture human body shape robustly across lighting and pose variations.
The impact
HOG+SVM became the dominant pedestrian detection pipeline for nearly a decade. It powered real-world systems in autonomous driving, surveillance, and robotics. The descriptor's design principles — dense overlapping grids, local orientation histograms, block normalization — directly influenced later in architectures. R- (2014) used HOG-era sliding-window thinking as a stepping stone to its region-proposal approach, marking the bridge between hand-crafted and learned features.
Imagine you're a police sketch artist. A witness describes a suspect: "tall, broad shoulders, arms at the sides." You don't need to know the color of their shirt or the exact lighting — you need outlines and proportions. You sketch the silhouette as a pattern of edges: strong vertical lines for the legs, a horizontal bar at the shoulders, a rounded contour at the head.
HOG does exactly this. It throws away color and brightness, keeps only the direction and strength of edges at every point, and bundles these into a compact description. A classifier then learns: "this pattern of edges looks like a person."
Why edges? The case for gradient-based features
Raw values are a terrible feature for person detection. Move the person into shade and every pixel changes. Change their shirt from white to black and half the pixels flip. But the edges — where intensity changes sharply — remain remarkably stable. The boundary between arm and background, the line of the shoulders, the contour of the head: these persist across lighting, clothing, and skin tone.
Earlier methods recognized this. Canny edge detection finds edges but loses their orientation and strength. SIFT computes oriented gradients around sparse keypoints. HOG's breakthrough was to compute dense, overlapping gradient histograms across the entire detection window, capturing the global shape of the human body rather than isolated keypoints.
The HOG pipeline: from pixels to person
The HOG descriptor is built in five stages, each refining the previous one. Think of it as an assembly line where raw pixels enter one end and a compact shape description emerges from the other. Understanding each stage, and why it matters, is the key to understanding how hand-crafted features once powered state-of-the-art detection.
Stage 1: Computing gradients
The first step is finding edges. At every pixel, we compute how fast the intensity changes horizontally and vertically. Dalal and Triggs found that the simplest possible works best: subtract the pixel on the left from the pixel on the right for the horizontal gradient, and the pixel above from the pixel below for the vertical one. No Gaussian smoothing, no Sobel operators — just the raw centered difference .
From the two components and , we compute the magnitude (how strong the edge is) and the orientation (which direction it points). The magnitude tells us how much of an edge is here; the orientation tells us which way it runs.
Stage 2: Building cell histograms
Next, we divide the image into small spatial regions called cells — typically 8×8 pixels. Within each cell, we build a histogram: we spread the full 0°–180° range of unsigned orientations into 9 bins (each covering 20°), and each pixel votes for the bin matching its gradient direction, with its vote weighted by the gradient magnitude.
Why unsigned (0°–180°) instead of signed (0°–360°)? Because for silhouette detection, a dark-to-light edge is the same shape information as a light-to-dark edge. A person's arm creates a vertical edge regardless of whether the arm is darker or lighter than the background. Dalal and Triggs confirmed experimentally that unsigned gradients give better results for pedestrian detection.
Think of each cell histogram as a compass rose: it tells you which edge directions dominate this small patch of the image. A cell on the shoulder will have strong horizontal bins; a cell on the leg will have strong vertical bins.
Stage 3: Block normalization — defeating illumination
Cell histograms capture local edge patterns, but their magnitudes depend on lighting. A person in bright sunlight has higher gradient magnitudes than the same person in shadow — the pattern of edge directions is the same, but the histogram values are scaled differently.
To fix this, HOG groups cells into larger blocks — typically 2×2 cells (16×16 pixels) — and normalizes the concatenated histograms within each block. The blocks overlap: shifting one cell at a time, so each cell appears in multiple blocks with different normalization contexts. This redundancy is deliberate — it makes the descriptor more robust.
Dalal and Triggs tested four normalization schemes and found that L2-norm followed by clipping at 0.2 and renormalizing (called L2-Hys) works best:
Stage 4: The final feature vector
After normalization, the histograms from all blocks are concatenated into one long vector. For a typical 64×128 pixel detection window with 8×8 cells, 2×2 cell blocks, and 9 orientation bins:
- The window has 8×16 = 128 cells
- Overlapping blocks: 7×15 = 105 blocks
- Each block: 2×2 cells × 9 bins = 36 values
- Final descriptor: 105 × 36 = 3,780 dimensions
This 3,780-element vector is a complete description of the gradient structure within the detection window. It encodes where every edge is, how strong it is, and which direction it points — all normalized against local illumination.
Stage 5: The linear SVM classifier
The 3,780-dimensional HOG vector is fed into a linear SVM — a classifier that learns a single separating "person" from "not person" in this high-dimensional space. The SVM learns which gradient patterns correspond to human body shapes during : strong horizontal gradients at shoulder height, strong vertical gradients at leg positions, a rounded gradient pattern at head level.
Why a linear SVM? Because the HOG features are so well-engineered that a simple linear boundary suffices — nonlinear kernels gave only marginal improvement at much higher computational cost. This is a hallmark of good : when the features are right, the classifier can be simple.
Detection: the sliding window approach
To find people in a full image, the HOG+SVM pipeline uses a : a 64×128 detection window moves across the image at multiple scales, computing the HOG descriptor at every position and feeding it to the SVM. The SVM says "person" or "not person" for each window. Overlapping positive detections are merged using to produce the final bounding boxes.
This is computationally expensive — thousands of windows per image at multiple scales — but it was practical for offline processing and near-real-time with optimizations. The sliding window paradigm would later be replaced by region proposals in R-CNN, but HOG proved that dense scanning with a strong descriptor works.
Design choices that made the difference
Dalal and Triggs ran extensive ablation studies — systematically changing one component at a time — and identified the choices that matter most:
- Fine gradient scale — the filter outperformed Sobel, Gaussian derivatives, and larger masks. Smoothing hurts because it blurs the very edges we're trying to capture.
- Unsigned gradients with fine orientation binning — 9 bins over 0°–180° outperformed signed gradients (0°–360°) and coarser binning. For shape detection, you need direction, not polarity.
- Overlapping block normalization — this was the single most impactful design choice. Without it, performance dropped dramatically. Local contrast normalization is what makes HOG robust to illumination.
- Linear SVM — a simple classifier on top of well-engineered features. This validates the principle that feature quality matters more than classifier complexity.
HOG vs SIFT: dense vs sparse
SIFT and HOG both use gradient histograms, but they answer different questions. SIFT asks: "What distinctive keypoints exist and how can I describe the region around each one?" HOG asks: "What is the overall gradient structure of the entire detection window?"
SIFT is sparse — it finds special points and ignores the rest. This makes it excellent for matching and recognition of specific objects. HOG is dense — it covers every cell in the window. This makes it excellent for detecting categories of objects (like "any person") where the shape varies but follows a pattern.
Think of it this way: SIFT is like describing a face by its moles and freckles (unique landmarks). HOG is like describing a face by the overall arrangement of edges — eyes here, nose there, chin below — which works for any face, not just one specific person.
HOG in code
Simplified to show the idea — not the real implementation.
import numpy as np
def compute_gradients(image):
"""Compute gradient magnitude and orientation at every pixel."""
gx = np.zeros_like(image, dtype=float)
gy = np.zeros_like(image, dtype=float)
gx[:, 1:-1] = image[:, 2:] - image[:, :-2] # [-1, 0, 1] filter
gy[1:-1, :] = image[2:, :] - image[:-2, :]
magnitude = np.sqrt(gx**2 + gy**2)
orientation = np.degrees(np.arctan2(gy, gx)) % 180 # unsigned: 0-180
return magnitude, orientation
def cell_histogram(magnitude, orientation, bins=9):
"""Build a 9-bin orientation histogram for one 8x8 cell."""
hist = np.zeros(bins)
bin_width = 180.0 / bins
for i in range(magnitude.shape[0]):
for j in range(magnitude.shape[1]):
bin_idx = int(orientation[i, j] / bin_width) % bins
hist[bin_idx] += magnitude[i, j] # vote weighted by magnitude
return hist
def hog_descriptor(image, cell_size=8, block_size=2, bins=9):
"""Full HOG: cells → blocks → normalize → concatenate."""
mag, ori = compute_gradients(image)
h, w = image.shape
cells_y, cells_x = h // cell_size, w // cell_size
# Build cell histograms
cells = np.zeros((cells_y, cells_x, bins))
for cy in range(cells_y):
for cx in range(cells_x):
y0, x0 = cy * cell_size, cx * cell_size
cells[cy, cx] = cell_histogram(
mag[y0:y0+cell_size, x0:x0+cell_size],
ori[y0:y0+cell_size, x0:x0+cell_size], bins)
# Block normalization (L2-Hys)
descriptor = []
for by in range(cells_y - block_size + 1):
for bx in range(cells_x - block_size + 1):
block = cells[by:by+block_size, bx:bx+block_size].flatten()
norm = np.sqrt(np.sum(block**2) + 1e-6)
block = block / norm
block = np.minimum(block, 0.2) # clip
norm = np.sqrt(np.sum(block**2) + 1e-6)
block = block / norm # renormalize
descriptor.append(block)
return np.concatenate(descriptor) # 3780-D for 64x128 windowR-HOG and C-HOG: rectangular vs circular blocks
Dalal and Triggs explored two block geometries. R-HOG (Rectangular HOG) uses square grids of cells — the version described above and the one that became standard. C-HOG (Circular HOG) uses cells arranged in log-polar sectors radiating from a central point, similar to SIFT's descriptor layout.
R-HOG performed slightly better for pedestrian detection, was simpler to implement, and became the default. C-HOG has advantages for rotationally symmetric objects but adds complexity without significant gains for upright human detection.
Impact: from hand-crafted to learned features
1999
SIFT (Lowe)
Scale-invariant sparse keypoint descriptors using gradient histograms. Proved that oriented gradients are powerful feature primitives.
2001
Viola-Jones (Haar features)
Real-time face detection using Haar wavelets and boosted cascades. Fast but struggled with full-body detection and pose variation.
2005
HOG (Dalal & Triggs)
Dense gradient histograms with block normalization — near-perfect pedestrian detection. Set the standard for a decade.
2008
DPM (Deformable Part Models)
Extended HOG with deformable part models — a root filter plus movable part filters, capturing pose variation. Won PASCAL VOC three years running.
2012
AlexNet
Deep CNNs shattered ImageNet records, learning features automatically instead of hand-crafting them. Beginning of the end for hand-designed descriptors.
2014
R-CNN (Girshick et al.)
Replaced HOG+SVM with CNN features + region proposals. Used HOG-era thinking (sliding window → region proposal) as the bridge to modern detection.
Limitations
CitationDalal, Triggs. Histograms of Oriented Gradients for Human Detection. CVPR, 2005.
Terms in this paper
- Histogram of Oriented Gradientsمُدرَّج الاتجاهات التكرارية
- Gradientالتدرج التفاضلي
- Feature Extractionاستخلاص السمات
- Object Detectionرصد وتحديد الكائنات
- Support Vector Machineآلة ناقلات الدعم (SVM)
- Sliding Windowالنافذة المنزلقة
- Normalizationالمعايرة القياسية للبيانات
- Kernelالنواة الحسابية
- Convolutionالالتفاف الرقمي
- Bounding Boxمربع الإحاطة
- Feature Mapخريطة السمات
- Linear Classifierمُصنِّف خطي
- Pedestrian Detectionاكتشاف المُشاة