Computer Vision2004intermediate9 min read

Distinctive Image Features from Scale-Invariant Keypoints

سمات صور مميّزة من نقاط مفتاحية ثابتة القياس

Lowe, D. G. — International Journal of Computer Vision (IJCV)

The problem

Recognizing an object in a new image — rotated, scaled, partially occluded, under different lighting — required hand-engineered features that were brittle and task-specific. Global features like color histograms collapsed under viewpoint changes; edge-based templates broke with rotation. No general-purpose method existed to detect the same physical point across images and describe it in a way that survived real-world transformations.

The contribution

(Scale-Invariant Transform): a four-stage pipeline that detects keypoints stable across scale and rotation, then builds a 128-dimensional for each one. Stage 1 builds a Gaussian and finds extrema in the Difference-of-Gaussians (DoG). Stage 2 refines locations with a Taylor expansion and discards low-contrast or edge-like points. Stage 3 assigns a dominant orientation from local gradients. Stage 4 encodes a 16×16 patch around the as 4×4 histograms with 8 bins each — a compact, distinctive fingerprint invariant to scale, rotation, and moderate illumination change.

The impact

SIFT defined the feature-matching paradigm that dominated computer vision for a decade. Panorama stitching, 3D reconstruction, object recognition, and augmented reality all ran on SIFT or its direct descendants (SURF, ORB, RootSIFT). Its design principles — scale-space detection, gradient-based description, ratio-test matching — remain deeply embedded in modern pipelines, even those powered by .

Imagine you're a detective with a magnifying glass, looking at a crime-scene photo. You notice a distinctive scratch on a doorknob — a mark that looks the same whether you zoom in, step back, or tilt the photo.

You jot down its location (where on the door), its orientation (which way the scratch runs), and a sketch of the surrounding pattern (the wood grain and paint chips nearby).

Later, when someone hands you a different photo of the same door — taken from across the room, slightly rotated — your notes still match. That's SIFT: find distinctive spots, describe their neighborhoods, and match them across views.

The problem: objects change appearance across views

Before SIFT, recognizing an object across different images was a patchwork of brittle tricks. Change the camera distance and a template shrinks or grows. Rotate the camera and edges rearrange. Change the lighting and intensities shift.

The core challenge: find the same physical point in two images and describe it so that the description survives scale changes, rotation, and illumination shifts — all without any learning or data. The description must be local (so partial occlusion doesn't destroy it) and distinctive (so each point's description differs from every other).

Open in Lab
Drag the sliders to change scale, rotation, and brightness. Watch how the raw pixel patch changes — but SIFT's descriptor (right) stays nearly identical.
The demo wakes as you arrive…

Stage 1: building the scale space

Objects exist at a particular scale: a sugar cube on a table is clearly visible, but from an airplane the cube vanishes. To detect features at every possible scale, SIFT constructs a scale space — a stack of progressively blurred copies of the image.

Each blur level uses a Gaussian with a larger standard deviation σ\sigma. The image is organized into octaves: each halves the image , and within each octave the blur increases by a constant factor kk. This means a feature that appears small in one octave may appear large in another — SIFT checks them all.

L(x,y,σ)=G(x,y,σ)∗I(x,y)L(x, y, \sigma) = G(x, y, \sigma) * I(x, y)
Scale-space representation — The image I convolved with a Gaussian G of width σ produces a blurred version L. Stacking these for increasing σ gives the scale space — a "tower" of blurs that lets SIFT see the image at every zoom level.
Open in Lab
Click any image in the pyramid to see the corresponding blur level. Each octave halves the resolution.
The demo wakes as you arrive…

Next, SIFT approximates the Laplacian of Gaussian (LoG) — a blob detector — with a cheaper operation: the (DoG). Each DoG image is simply one blur level minus the next. The result highlights regions where intensity changes abruptly — exactly the places that make good keypoints.

D(x,y,σ)=L(x,y,kσ)−L(x,y,σ)D(x, y, \sigma) = L(x, y, k\sigma) - L(x, y, \sigma)
Difference of Gaussians (DoG) — Subtracting two adjacent Gaussian blurs approximates the Laplacian — a "blob detector" that lights up at locations and scales where the image has distinctive structure.
Open in Lab
Watch how subtracting adjacent blurs produces the DoG images, then see which points survive as extrema across 26 neighbors (8 same-level + 9 above + 9 below).
The demo wakes as you arrive…

Stage 2: refining keypoint locations

The raw extrema from Stage 1 are approximate — they sit on the discrete grid. SIFT refines each candidate using a Taylor expansion of the DoG function, fitting a 3D quadratic to find the sub-pixel, sub-scale true extremum. Think of it as focusing a microscope: the pixel grid gives a blurry location, the Taylor fit sharpens it.

Two filters then discard bad candidates:

  • Low-contrast filter: if the DoG value at the refined extremum is below a threshold (|D| < 0.03), the point is too faint to be reliable — discard it.

  • Edge filter: a point along an edge (like a wall) is poorly localized in one direction. SIFT checks the ratio of the principal curvatures (eigenvalues of the ). If the ratio exceeds a threshold (typically 10), the point is edge-like — discard it.

x^=−∂2D∂x2−1∂D∂x\hat{x} = -\frac{\partial^2 D}{\partial x^2}^{-1} \frac{\partial D}{\partial x}
Sub-pixel keypoint refinement — A Taylor expansion around the candidate keypoint yields a quadratic. Setting the derivative to zero gives the offset x-hat — the sub-pixel shift to the true extremum.

Stage 3: assigning orientation

To achieve rotation invariance, SIFT measures the dominant gradient direction around each keypoint. It builds an orientation histogram with 36 bins (each covering 10°), weighted by gradient magnitude and a Gaussian window centered on the keypoint.

The bin with the highest peak becomes the keypoint's canonical orientation. Any bin within 80% of the peak spawns an additional keypoint at the same location with that orientation — so a corner viewed from two angles can generate two entries. From this point on, all measurements are made relative to this orientation, which is why rotating the image does not change the final descriptor.

Open in Lab
See gradient arrows around a keypoint and how they form a 36-bin histogram. The tallest peak sets the canonical orientation (red arrow).
The demo wakes as you arrive…

Stage 4: building the 128-D descriptor

Now each keypoint has a location, a scale, and an orientation. The final step is to describe what the neighborhood looks like in a way that is compact, distinctive, and invariant.

SIFT takes a 16×16 pixel patch around the keypoint (aligned to the canonical orientation and scaled appropriately), divides it into a 4×4 grid of sub-regions (each 4×4 pixels), and in each sub-region computes a histogram of 8 gradient orientations. Stacking 4×4×8 = 128 values into a single produces the SIFT descriptor.

The descriptor is then normalized to unit length for illumination invariance: if all pixel intensities double, gradients double, but after the vector is unchanged. A clipping step (capping values at 0.2 and re-normalizing) further reduces the effect of large gradients from non-linear illumination changes.

Open in Lab
Hover over sub-regions to see their 8-bin gradient histograms. Together they form the 128-D descriptor shown on the right.
The demo wakes as you arrive…

Matching: finding correspondences

Given two images, SIFT extracts keypoints from each, then matches descriptors using nearest-neighbor search in the 128-D space. But a raw nearest neighbor is often wrong — a descriptor might simply match the closest one available, even if the true match is absent.

Lowe's elegant solution is the ratio test: compute the distance to the closest and second-closest neighbors. If the ratio d1/d2<0.8d_1/d_2 < 0.8, the match is accepted; otherwise, it's rejected as ambiguous. This single test eliminates ~90% of false matches while keeping ~95% of correct ones.

Open in Lab
Keypoints extracted from two views of the same object. Lines connect matched descriptors. Toggle the ratio test to see how it filters false matches.
The demo wakes as you arrive…

The complete SIFT pipeline

Let's now walk through all four stages together. From a raw image, SIFT produces a set of keypoints, each with a location (x, y), a scale (σ), an orientation (θ), and a 128-D descriptor — a complete "ID card" for that spot in the image. Two images of the same scene may produce hundreds of these ID cards, and matching them reveals which points are the same physical location.

Open in Lab
Step through the four SIFT stages on a sample image.
The demo wakes as you arrive…

The idea in code

SIFT descriptor construction (simplified)python

Simplified to show the idea — not the real implementation.

import numpy as np

def compute_gradients(patch):
    """Compute gradient magnitude and orientation for a 16x16 patch."""
    dy = patch[2:, 1:-1] - patch[:-2, 1:-1]  # vertical gradient
    dx = patch[1:-1, 2:] - patch[1:-1, :-2]   # horizontal gradient
    magnitude = np.sqrt(dx**2 + dy**2)
    orientation = np.arctan2(dy, dx)           # in radians
    return magnitude, orientation

def build_descriptor(magnitude, orientation, n_bins=8, grid=4):
    """Build 128-D SIFT descriptor from 16x16 gradient patch."""
    h, w = magnitude.shape
    cell_h, cell_w = h // grid, w // grid
    descriptor = []
    for i in range(grid):
        for j in range(grid):
            # Extract the 4x4 sub-region
            mag = magnitude[i*cell_h:(i+1)*cell_h, j*cell_w:(j+1)*cell_w]
            ori = orientation[i*cell_h:(i+1)*cell_h, j*cell_w:(j+1)*cell_w]
            # Build 8-bin histogram weighted by magnitude
            hist, _ = np.histogram(ori, bins=n_bins,
                                   range=(-np.pi, np.pi),
                                   weights=mag)
            descriptor.extend(hist)
    descriptor = np.array(descriptor)
    # Normalize → clip → renormalize (illumination invariance)
    descriptor /= (np.linalg.norm(descriptor) + 1e-7)
    descriptor = np.clip(descriptor, 0, 0.2)
    descriptor /= (np.linalg.norm(descriptor) + 1e-7)
    return descriptor  # 128-D vector: the keypoint's fingerprint

# Each keypoint in the image gets one of these 128-D descriptors.
# Matching = finding the nearest descriptor in the other image.

Why it changed everything

SIFT's design principles — build a scale space, detect stable extrema, describe with local gradients, match with a ratio test — became the blueprint for a generation of feature methods. Even deep-learned features like SuperPoint and SuperGlue use SIFT-era evaluation protocols and often compare against SIFT as a baseline.

  1. 1999

    Lowe's initial SIFT paper

    Object Recognition from Local Scale-Invariant Features, presented at ICCV. Introduced the scale-space extrema approach and basic descriptor.

  2. 2004

    Full SIFT paper (this paper)

    Published in IJCV with the complete 4-stage pipeline, ratio test, and extensive evaluation. Became the standard reference for local feature matching.

  3. 2006

    SURF — speed over accuracy

    Bay et al. approximated SIFT using box filters and integral images, trading some accuracy for 3–5× faster computation. Made real-time matching feasible.

  4. 2008

    Photo Tourism / Bundler

    Snavely et al. used SIFT to reconstruct 3D scenes from thousands of internet photos. Proved SIFT scales to massive unstructured photo collections.

  5. 2011

    ORB — SIFT's free alternative

    Rublee et al. combined FAST keypoints with BRIEF descriptors and added rotation invariance. Patent-free, real-time, and nearly as accurate as SIFT for many tasks.

  6. 2020

    SuperGlue — learned matching

    Sarlin et al. replaced hand-crafted matching with a graph neural network that learns to match keypoints. Used SIFT's pipeline as the baseline to beat.

  7. 2020

    SIFT patent expires

    Lowe's US patent expired, making SIFT freely available for commercial use. OpenCV moved SIFT from the non-free module to the main library.

CitationLowe, D. G.. Distinctive Image Features from Scale-Invariant Keypoints. International Journal of Computer Vision (IJCV), 2004.

Terms in this paper