Computer Vision1998foundational6 min read

Gradient-Based Learning Applied to Document Recognition

التعلُّم بالتدرُّج لتمييز الوثائق

LeCun, Y. · Bottou, L. · Bengio, Y. · Haffner, P. — Proceedings of the IEEE

The problem

Recognizing handwritten characters with classic pattern recognition required hand-designed feature extractors — brittle, labor-intensive, and different for every task. Fully-connected networks on raw pixels ignore image structure and explode in parameters: a 32×32 image into 100 hidden units already needs 100,000 weights, and a digit shifted by one pixel looks like a completely different input.

The contribution

-5: a trained by . Three architectural priors — local receptive fields, shared weights, and spatial — encode what we know about images: patterns are local, the same pattern can appear anywhere, and exact position matters less than presence. The entire pipeline, from raw pixels to class label, is differentiable and learned.

The impact

LeNet read millions of bank checks at AT&T in the 1990s — 's first industrial deployment. It defined the CNN template — conv, pool, repeat, classify — that AlexNet scaled up in 2012 to ignite the modern deep learning era. Every vision model since carries its DNA.

Imagine checking a page for the letter "L" with a rubber stamp carved with an L-shape.

You press it at every position; where the ink underneath matches the carving, the stamp "lights up."

That's a : one small pattern detector, reused everywhere.

LeNet's insight is that the network can carve its own stamps — the filters are learned by gradient descent, not designed by an engineer.

The problem: hand-crafted features don't scale

Before this paper, a digit recognizer was two systems glued together: a hand-designed feature extractor (edge detectors, stroke counters — years of expert tuning) and a small trainable classifier on top. Change the task from digits to letters? Start over.

The naive fix — a fully-connected network on raw pixels — fails for two reasons:

  • A 32×32 image into 100 hidden units already costs 100,000 weights, and scaling to real images (224×224) makes this 5 million per layer. is inevitable.

  • Pixels are not independent: nearby pixels form strokes, strokes form parts, parts form digits. Full connectivity throws that structure away and must re-learn it from scratch.

Open in Lab
Drag the image size and watch fully-connected parameters explode while convolution stays flat.
The demo wakes as you arrive…

The idea: three priors about images

LeNet bakes what we know about images into the architecture itself:

  1. Local receptive fields — each neuron looks at a small patch (5×5), not the whole image. Strokes are local; a pixel and its neighbors are the meaningful unit.

  2. — the same filter scans every position. A "7" corner is a "7" corner whether it appears top-left or bottom-right. This is the rubber stamp: one set of weights, applied everywhere.

  3. () — after detecting a feature, shrink the map by keeping local summaries. A digit shifted by one pixel still looks similar after pooling. This buys .

Convolution: the sliding dot product

A (filter) is a small matrix — say 5×5 — of learned weights. It slides across the input image one position at a time. At each position, we compute the dot product of the with the image patch beneath it: multiply element-wise, sum the products. The result is a single number — high if the patch "matches" the kernel, low if it doesn't.

Doing this at every position produces a : a new image where each pixel tells you "how strongly this pattern was detected here." Six different kernels produce six feature maps — six different "stamps" pressed everywhere.

y[i,j]=∑u,vw[u,v]⋅x[i+u, j+v]+by[i,j] = \sum_{u,v} w[u,v] \cdot x[i+u,\, j+v] + b
2D Convolution — One output pixel = dot product of filter w with image patch. Same w at every position — that is weight sharing.
Open in Lab
Click cells and switch kernels. A vertical-edge kernel responds strongly at the boundary between dark and light regions.
The demo wakes as you arrive…

Pooling: tolerance to small shifts

After convolution, pooling reduces spatial dimensions while keeping the strongest signals. takes the maximum value in each small block (e.g. 2×2); (used in LeNet) takes the mean.

Pooling serves two purposes: it halves the spatial size (reducing computation downstream), and it makes the representation tolerant to small translations — a digit shifted by one pixel produces nearly the same pooled output.

Open in Lab
Hover a 2×2 block. In max mode, only the winner survives. In avg mode, all four contribute.
The demo wakes as you arrive…

Stack and repeat: the feature hierarchy

LeNet stacks convolution and pooling multiple times. With each layer, the features grow more abstract: the first layer detects edges, the second combines edges into corners and curves, and deeper layers assemble these into digit parts and whole digits.

This hierarchy emerges automatically — nobody tells the network what features to learn. Backpropagation, the algorithm from the 1986 paper, adjusts every filter so that the final classification improves. The features are discovered, not designed.

Open in Lab
Click each level to see what the network learns to detect — from simple edges to whole digits.
The demo wakes as you arrive…

LeNet-5: the full architecture

LeNet-5 chains these ideas into a complete pipeline. A 32×32 grayscale input passes through two rounds of convolution + pooling, then a third convolution that covers the entire remaining spatial extent (effectively becoming fully-connected), followed by two dense layers that map to 10 output classes (digits 0–9).

The entire network has about 60,000 parameters — compared to the millions a fully-connected network would need. And it was trained end-to-end: raw pixels in, class label out, gradient descent adjusts everything.

Open in Lab
Click any layer to see its shape, parameter count, and what it does.
The demo wakes as you arrive…

The same idea in code

Convolution = a sliding dot productpython

Simplified to show the idea — not the real implementation.

import numpy as np

def convolve2d(image, kernel):
    """Slide `kernel` over `image`; each output pixel is a dot product."""
    H, W = image.shape
    k = kernel.shape[0]                      # e.g. 5 for a 5×5 filter
    out = np.zeros((H - k + 1, W - k + 1))
    for i in range(out.shape[0]):
        for j in range(out.shape[1]):
            patch = image[i:i+k, j:j+k]      # the paper under the stamp
            out[i, j] = np.sum(patch * kernel)  # press the stamp
    return out

# A vertical-edge "stamp". LeNet's point: don't write this by hand —
# initialize randomly and let backpropagation carve it.
vertical_edge = np.array([[ 1, 0, -1],
                          [ 1, 0, -1],
                          [ 1, 0, -1]])

# For a 32×32 image with a 5×5 kernel:
# FC layer: 32*32 * 100 = 102,400 weights
# Conv layer: 5*5 + 1 bias = 26 weights (per filter)
# That's 3,938× fewer parameters. Per filter.

Why it mattered

  1. 2006

    Hinton — Deep belief nets revive depth

    Hinton, Osindero & Teh showed deep networks could be pre-trained layer-by-layer, then fine-tuned end-to-end — reigniting the field's confidence that depth was worth pursuing.

  2. 2010

    Xavier init & ReLU — Vanishing gradients solved

    Glorot & Bengio's Xavier initialization kept activations in range at the start of training, while Nair & Hinton's ReLU (derivative = 1 when active) eliminated the vanishing gradient problem that had blocked deep CNNs for years.

  3. 2012

    AlexNet — LeNet's recipe, scaled to GPUs

    AlexNet won ImageNet by a 10-point margin — a deep CNN trained end-to-end with backpropagation, exactly LeNet's conv–pool–classify recipe scaled up with more layers, more data, and GPU compute. The moment deep learning became the dominant paradigm.

  4. 2014

    GoogLeNet & VGG — Depth at scale

    GoogLeNet introduced inception modules (parallel filters of different sizes) to go wider efficiently, while VGG showed that simple 3×3 convolutions stacked very deep outperform complex hand-designed filters.

  5. 2015

    ResNet — 150+ layers with residual connections

    He et al. introduced skip connections that add the input directly to the output of each block — the gradient can always flow backward through the identity path. CNNs scaled to 152 layers and won every major vision benchmark.

  6. 2020

    Vision Transformer — attention replaces convolution

    Dosovitskiy et al. showed that treating image patches as tokens and applying pure self-attention matches and then exceeds CNN performance at scale — challenging two decades of convolutional dominance while preserving the end-to-end training principle.

  7. 2026

    Every image model alive

    Every image classifier, object detector, and medical imaging model uses convolution or its descendants. The 1998 paper's three priors — locality, weight sharing, and pooling — remain the foundation of visual AI.

CitationLeCun, Bottou, Bengio, Haffner. Gradient-Based Learning Applied to Document Recognition. Proceedings of the IEEE, 1998.

Terms in this paper