Computer Vision2005intermediate12 min read

Histograms of Oriented Gradients for Human Detection

مُدرَّجات الاتجاهات التكرارية لاكتشاف البشر

Dalal, N. · Triggs, B. — CVPR

The problem

Before 2005, detecting people in photographs was unreliable. Existing descriptors like Haar wavelets (used in Viola-Jones face detection) struggled with the high variability of human poses, clothing, and backgrounds. Edge-based features existed but were too sensitive to lighting changes, and no single hand-crafted feature set could robustly capture the human silhouette across different conditions.

The contribution

HOG: a dense feature descriptor that divides the image into small cells, computes a histogram of orientations in each cell, groups cells into overlapping blocks for contrast , and concatenates everything into a single high-dimensional feature . Paired with a linear SVM, this pipeline achieved near-perfect on the MIT dataset and set a new standard. The key insight: local gradient orientation distributions, with proper normalization, capture human body shape robustly across lighting and pose variations.

The impact

HOG+SVM became the dominant pedestrian detection pipeline for nearly a decade. It powered real-world systems in autonomous driving, surveillance, and robotics. The descriptor's design principles — dense overlapping grids, local orientation histograms, block normalization — directly influenced later in architectures. R- (2014) used HOG-era sliding-window thinking as a stepping stone to its region-proposal approach, marking the bridge between hand-crafted and learned features.

Imagine you're a police sketch artist. A witness describes a suspect: "tall, broad shoulders, arms at the sides." You don't need to know the color of their shirt or the exact lighting — you need outlines and proportions. You sketch the silhouette as a pattern of edges: strong vertical lines for the legs, a horizontal bar at the shoulders, a rounded contour at the head.

HOG does exactly this. It throws away color and brightness, keeps only the direction and strength of edges at every point, and bundles these into a compact description. A classifier then learns: "this pattern of edges looks like a person."

Why edges? The case for gradient-based features

Raw values are a terrible feature for person detection. Move the person into shade and every pixel changes. Change their shirt from white to black and half the pixels flip. But the edges — where intensity changes sharply — remain remarkably stable. The boundary between arm and background, the line of the shoulders, the contour of the head: these persist across lighting, clothing, and skin tone.

Earlier methods recognized this. Canny edge detection finds edges but loses their orientation and strength. SIFT computes oriented gradients around sparse keypoints. HOG's breakthrough was to compute dense, overlapping gradient histograms across the entire detection window, capturing the global shape of the human body rather than isolated keypoints.

The HOG pipeline: from pixels to person

The HOG descriptor is built in five stages, each refining the previous one. Think of it as an assembly line where raw pixels enter one end and a compact shape description emerges from the other. Understanding each stage, and why it matters, is the key to understanding how hand-crafted features once powered state-of-the-art detection.

Open in Lab
Step through the five stages of the HOG pipeline. Click each stage to see what it does and why.
The demo wakes as you arrive…

Stage 1: Computing gradients

The first step is finding edges. At every pixel, we compute how fast the intensity changes horizontally and vertically. Dalal and Triggs found that the simplest possible works best: subtract the pixel on the left from the pixel on the right for the horizontal gradient, and the pixel above from the pixel below for the vertical one. No Gaussian smoothing, no Sobel operators — just the raw centered difference [−1,0,1][-1, 0, 1].

From the two components gxg_x and gyg_y, we compute the magnitude (how strong the edge is) and the orientation (which direction it points). The magnitude tells us how much of an edge is here; the orientation tells us which way it runs.

gx=I(x+1,y)−I(x−1,y),gy=I(x,y+1)−I(x,y−1)g_x = I(x+1,y) - I(x-1,y), \quad g_y = I(x,y+1) - I(x,y-1)
Horizontal and vertical gradient components — The simplest centered-difference filter. No smoothing needed — Dalal & Triggs showed it outperforms Sobel and Gaussian-derivative filters for this task.
m=gx2+gy2,θ=arctan⁡(gy/gx)m = \sqrt{g_x^2 + g_y^2}, \quad \theta = \arctan(g_y / g_x)
Gradient magnitude and orientation — m = edge strength · θ = edge direction. These two values at each pixel are the raw material that HOG organizes into histograms.
Open in Lab
Hover over pixels to see how the gradient is computed from neighbors. Arrows show direction, brightness shows magnitude.
The demo wakes as you arrive…

Stage 2: Building cell histograms

Next, we divide the image into small spatial regions called cells — typically 8×8 pixels. Within each cell, we build a histogram: we spread the full 0°–180° range of unsigned orientations into 9 bins (each covering 20°), and each pixel votes for the bin matching its gradient direction, with its vote weighted by the gradient magnitude.

Why unsigned (0°–180°) instead of signed (0°–360°)? Because for silhouette detection, a dark-to-light edge is the same shape information as a light-to-dark edge. A person's arm creates a vertical edge regardless of whether the arm is darker or lighter than the background. Dalal and Triggs confirmed experimentally that unsigned gradients give better results for pedestrian detection.

Think of each cell histogram as a compass rose: it tells you which edge directions dominate this small patch of the image. A cell on the shoulder will have strong horizontal bins; a cell on the leg will have strong vertical bins.

Open in Lab
Click on different cells to see how gradient orientations are binned into a 9-bin histogram. Taller bars mean more edge energy in that direction.
The demo wakes as you arrive…

Stage 3: Block normalization — defeating illumination

Cell histograms capture local edge patterns, but their magnitudes depend on lighting. A person in bright sunlight has higher gradient magnitudes than the same person in shadow — the pattern of edge directions is the same, but the histogram values are scaled differently.

To fix this, HOG groups cells into larger blocks — typically 2×2 cells (16×16 pixels) — and normalizes the concatenated histograms within each block. The blocks overlap: shifting one cell at a time, so each cell appears in multiple blocks with different normalization contexts. This redundancy is deliberate — it makes the descriptor more robust.

Dalal and Triggs tested four normalization schemes and found that L2-norm followed by clipping at 0.2 and renormalizing (called L2-Hys) works best:

v=f∥f∥22+ϵ2,vi←min⁡(vi,0.2),v←v∥v∥22+ϵ2v = \frac{f}{\sqrt{\|f\|_2^2 + \epsilon^2}}, \quad v_i \leftarrow \min(v_i, 0.2), \quad v \leftarrow \frac{v}{\sqrt{\|v\|_2^2 + \epsilon^2}}
L2-Hys block normalization — f = raw histogram vector of a block · normalize by L2 · clip values above 0.2 (caps dominant gradients) · renormalize. This suppresses illumination effects while keeping shape information intact.
Open in Lab
See how overlapping blocks cover the detection window. Toggle illumination to watch normalization stabilize the histograms.
The demo wakes as you arrive…

Stage 4: The final feature vector

After normalization, the histograms from all blocks are concatenated into one long vector. For a typical 64×128 pixel detection window with 8×8 cells, 2×2 cell blocks, and 9 orientation bins:

  • The window has 8×16 = 128 cells
  • Overlapping blocks: 7×15 = 105 blocks
  • Each block: 2×2 cells × 9 bins = 36 values
  • Final descriptor: 105 × 36 = 3,780 dimensions

This 3,780-element vector is a complete description of the gradient structure within the detection window. It encodes where every edge is, how strong it is, and which direction it points — all normalized against local illumination.

Stage 5: The linear SVM classifier

The 3,780-dimensional HOG vector is fed into a linear SVM — a classifier that learns a single separating "person" from "not person" in this high-dimensional space. The SVM learns which gradient patterns correspond to human body shapes during : strong horizontal gradients at shoulder height, strong vertical gradients at leg positions, a rounded gradient pattern at head level.

Why a linear SVM? Because the HOG features are so well-engineered that a simple linear boundary suffices — nonlinear kernels gave only marginal improvement at much higher computational cost. This is a hallmark of good : when the features are right, the classifier can be simple.

Detection: the sliding window approach

To find people in a full image, the HOG+SVM pipeline uses a : a 64×128 detection window moves across the image at multiple scales, computing the HOG descriptor at every position and feeding it to the SVM. The SVM says "person" or "not person" for each window. Overlapping positive detections are merged using to produce the final bounding boxes.

This is computationally expensive — thousands of windows per image at multiple scales — but it was practical for offline processing and near-real-time with optimizations. The sliding window paradigm would later be replaced by region proposals in R-CNN, but HOG proved that dense scanning with a strong descriptor works.

Open in Lab
Watch the sliding window scan across the image. Green boxes are positive detections; the SVM confidence is shown for each.
The demo wakes as you arrive…

Design choices that made the difference

Dalal and Triggs ran extensive ablation studies — systematically changing one component at a time — and identified the choices that matter most:

  • Fine gradient scale — the [−1,0,1][-1, 0, 1] filter outperformed Sobel, Gaussian derivatives, and larger masks. Smoothing hurts because it blurs the very edges we're trying to capture.
  • Unsigned gradients with fine orientation binning — 9 bins over 0°–180° outperformed signed gradients (0°–360°) and coarser binning. For shape detection, you need direction, not polarity.
  • Overlapping block normalization — this was the single most impactful design choice. Without it, performance dropped dramatically. Local contrast normalization is what makes HOG robust to illumination.
  • Linear SVM — a simple classifier on top of well-engineered features. This validates the principle that feature quality matters more than classifier complexity.

HOG vs SIFT: dense vs sparse

SIFT and HOG both use gradient histograms, but they answer different questions. SIFT asks: "What distinctive keypoints exist and how can I describe the region around each one?" HOG asks: "What is the overall gradient structure of the entire detection window?"

SIFT is sparse — it finds special points and ignores the rest. This makes it excellent for matching and recognition of specific objects. HOG is dense — it covers every cell in the window. This makes it excellent for detecting categories of objects (like "any person") where the shape varies but follows a pattern.

Think of it this way: SIFT is like describing a face by its moles and freckles (unique landmarks). HOG is like describing a face by the overall arrangement of edges — eyes here, nose there, chin below — which works for any face, not just one specific person.

Open in Lab
Left: SIFT keypoints (sparse, each with its own descriptor). Right: HOG grid (dense, covering the entire window). Both use gradient histograms differently.
The demo wakes as you arrive…

HOG in code

Simplified HOG descriptor computationpython

Simplified to show the idea — not the real implementation.

import numpy as np

def compute_gradients(image):
    """Compute gradient magnitude and orientation at every pixel."""
    gx = np.zeros_like(image, dtype=float)
    gy = np.zeros_like(image, dtype=float)
    gx[:, 1:-1] = image[:, 2:] - image[:, :-2]   # [-1, 0, 1] filter
    gy[1:-1, :] = image[2:, :] - image[:-2, :]
    magnitude = np.sqrt(gx**2 + gy**2)
    orientation = np.degrees(np.arctan2(gy, gx)) % 180  # unsigned: 0-180
    return magnitude, orientation

def cell_histogram(magnitude, orientation, bins=9):
    """Build a 9-bin orientation histogram for one 8x8 cell."""
    hist = np.zeros(bins)
    bin_width = 180.0 / bins
    for i in range(magnitude.shape[0]):
        for j in range(magnitude.shape[1]):
            bin_idx = int(orientation[i, j] / bin_width) % bins
            hist[bin_idx] += magnitude[i, j]    # vote weighted by magnitude
    return hist

def hog_descriptor(image, cell_size=8, block_size=2, bins=9):
    """Full HOG: cells → blocks → normalize → concatenate."""
    mag, ori = compute_gradients(image)
    h, w = image.shape
    cells_y, cells_x = h // cell_size, w // cell_size

    # Build cell histograms
    cells = np.zeros((cells_y, cells_x, bins))
    for cy in range(cells_y):
        for cx in range(cells_x):
            y0, x0 = cy * cell_size, cx * cell_size
            cells[cy, cx] = cell_histogram(
                mag[y0:y0+cell_size, x0:x0+cell_size],
                ori[y0:y0+cell_size, x0:x0+cell_size], bins)

    # Block normalization (L2-Hys)
    descriptor = []
    for by in range(cells_y - block_size + 1):
        for bx in range(cells_x - block_size + 1):
            block = cells[by:by+block_size, bx:bx+block_size].flatten()
            norm = np.sqrt(np.sum(block**2) + 1e-6)
            block = block / norm
            block = np.minimum(block, 0.2)       # clip
            norm = np.sqrt(np.sum(block**2) + 1e-6)
            block = block / norm                  # renormalize
            descriptor.append(block)

    return np.concatenate(descriptor)  # 3780-D for 64x128 window

R-HOG and C-HOG: rectangular vs circular blocks

Dalal and Triggs explored two block geometries. R-HOG (Rectangular HOG) uses square grids of cells — the version described above and the one that became standard. C-HOG (Circular HOG) uses cells arranged in log-polar sectors radiating from a central point, similar to SIFT's descriptor layout.

R-HOG performed slightly better for pedestrian detection, was simpler to implement, and became the default. C-HOG has advantages for rotationally symmetric objects but adds complexity without significant gains for upright human detection.

Impact: from hand-crafted to learned features

  1. 1999

    SIFT (Lowe)

    Scale-invariant sparse keypoint descriptors using gradient histograms. Proved that oriented gradients are powerful feature primitives.

  2. 2001

    Viola-Jones (Haar features)

    Real-time face detection using Haar wavelets and boosted cascades. Fast but struggled with full-body detection and pose variation.

  3. 2005

    HOG (Dalal & Triggs)

    Dense gradient histograms with block normalization — near-perfect pedestrian detection. Set the standard for a decade.

  4. 2008

    DPM (Deformable Part Models)

    Extended HOG with deformable part models — a root filter plus movable part filters, capturing pose variation. Won PASCAL VOC three years running.

  5. 2012

    AlexNet

    Deep CNNs shattered ImageNet records, learning features automatically instead of hand-crafting them. Beginning of the end for hand-designed descriptors.

  6. 2014

    R-CNN (Girshick et al.)

    Replaced HOG+SVM with CNN features + region proposals. Used HOG-era thinking (sliding window → region proposal) as the bridge to modern detection.

Limitations

CitationDalal, Triggs. Histograms of Oriented Gradients for Human Detection. CVPR, 2005.

Terms in this paper