Computer Vision2015intermediate11 min read

Unsupervised Visual Representation Learning by Context Prediction

تعلُّم تمثيلات بصرية بدون إشراف عبر التنبؤ بالسياق

Doersch, C. · Gupta, A. · Efros, A. A. — ICCV

The problem

By 2015, the best visual features came from convolutional networks trained on millions of labeled images (ImageNet). But labeling is expensive: only a tiny fraction of the world's images have labels. Unsupervised methods — autoencoders, generative models — had failed to learn features competitive with supervised ones on real, full-resolution images. The field needed a way to extract rich visual representations from unlabeled data alone.

The contribution

A self-supervised : extract random pairs of patches from an image in a 3×3 grid and train a ConvNet to predict which of 8 possible relative positions the second occupies. The hypothesis is that solving this spatial puzzle forces the network to recognize objects and parts. Two AlexNet-style branches with shared weights process each patch separately, then fuse for . Careful design avoids trivial shortcuts — gaps, jitter, and color-channel manipulation block texture-continuation and chromatic-aberration cheats. The learned features achieve state-of-the-art among unsupervised methods on PASCAL VOC 2007 detection (46.3% mAP), and enable unsupervised visual discovery of object categories.

The impact

This paper launched the era of visual pretext tasks — designing clever puzzles whose solutions demand semantic understanding. It proved that spatial context within a single image is a rich supervisory signal. RotNet, Jigsaw, and Colorization all followed this template. The "pretext task → transfer" pipeline it established became the foundation that contrastive methods like MoCo and SimCLR later refined, ultimately leading to self-supervised features that match or exceed supervised ImageNet features.

Imagine you're assembling a jigsaw puzzle but the box lid is missing — you've never seen the full picture. You pick up two pieces and try to figure out how they fit together. If one piece shows a cat's ear and the other shows whiskers, you know the whiskers go below the ear — but only because you know what a cat looks like.

That's exactly this paper's trick: force a to play the same game with image patches. It never sees labels. It never sees the full picture. It just tries to predict spatial relationships — and to do that well, it must learn what objects are.

The problem: labels are expensive, images are free

By 2015, the recipe for good visual features was clear: train a deep convolutional network on ImageNet's 1.3 million labeled images. The features from the hidden layers — especially the fully-connected layers — transferred beautifully to other tasks like on PASCAL VOC.

But this recipe has a bottleneck: human labels. ImageNet took years of crowd-sourced annotation. The internet has billions of unlabeled images, but can't touch them. Generative models (autoencoders, Boltzmann machines) struggled with full-resolution natural images — they got lost in low-level pixel details like texture instead of learning high-level concepts like "cat" or "car."

The question was: can we design a task that requires semantic understanding, yet needs zero labels?

The idea: predict where a patch belongs

The insight comes from word embeddings in NLP. Skip-gram models learn rich word representations by predicting context words — without any labeled data. The authors ask: can we do the same with image patches?

Here is the setup. Take an image and divide it conceptually into a 3×3 grid. Pick the center patch. Now randomly pick one of the 8 surrounding patches. Give the network only these two patches — with no information about where they came from — and ask: where is the second patch relative to the first? This is an 8-way classification problem.

Why does this work? Consider two patches from a face — one showing an eye, the other showing a mouth. To predict that the mouth is below the eye, the network must recognize both parts and understand facial geometry. Random texture patches, by contrast, give no spatial signal. The task naturally pushes the network toward learning about objects and their structure.

Open in Lab
Click any of the 8 surrounding positions to see which pair the network would receive. The network must classify one of 8 spatial relationships.
The demo wakes as you arrive…

Architecture: a Siamese AlexNet

The network follows a late-fusion design. Each patch passes through its own AlexNet-style branch (conv1 through an fc6-equivalent) — but the two branches share weights. This is critical: shared weights mean the same -extraction function is applied to both patches, producing a single reusable per patch.

Only at the top — two fully connected layers — do the representations from both patches merge for joint reasoning. Because joint capacity is limited, the network is forced to do most of its semantic work within each branch. The final outputs a probability distribution over the 8 possible spatial configurations.

Think of it as two identical librarians who each read a page independently, then meet briefly to decide how the pages relate. Each librarian must extract as much meaning as possible on their own, because the meeting is short.

Open in Lab
Click any layer to see its role. Dotted lines indicate shared weights between the two branches.
The demo wakes as you arrive…

Avoiding trivial shortcuts

A pretext task is useless if the network can solve it without learning semantics. The authors discovered three shortcuts the network could exploit:

1. Boundary patterns. If patches touch, textures or edges might continue across the boundary, giving away the answer. Fix: leave a gap of ~48 pixels (half the patch width) between patches.

2. Exact alignment. Even with a gap, long straight lines could span patches. Fix: jitter each patch location randomly by up to ±7 pixels.

3. — the most surprising shortcut. Camera lenses focus different wavelengths slightly differently, causing the green channel to shrink toward the image center relative to red and blue. The network learned to detect the green–magenta separation in each patch, infer its absolute position on the lens, and then compute the relative position trivially.

Fix: either project out the green–magenta axis from every pixel, or randomly drop 2 of 3 color channels and replace them with Gaussian noise. Both force the network to rely on content, not color artifacts.

Open in Lab
Toggle between original and color-projected patches. Notice how the green–magenta gradient disappears after projection, forcing the network to use content cues.
The demo wakes as you arrive…

Training details

The network was trained on the ImageNet 2012 (~1.3M images), using only the images — all labels discarded. Images were resized to 150K–450K pixels preserving aspect ratio. Patches were sampled at 96×96 resolution from a grid pattern, with each patch participating in up to 8 pairings.

Preprocessing for each patch included mean subtraction, color projection or channel dropping, and random downsampling to as few as 100 pixels (then upsampling) to build robustness to pixelation.

A critical challenge was saddle-point collapse: with standard , the network would predict a uniform distribution over all 8 classes, with fc6 and fc7 activations collapsing to zero. The optimization got stuck in a where the upper layers ignored lower-layer features. The fix was (without the learnable scale and shift parameters), which forced activations to vary across examples and broke the degeneracy. High (0.999) further accelerated learning. took approximately four weeks on a single K40 GPU.

The idea in code

Context prediction — patch sampling and 8-way classificationpython

Simplified to show the idea — not the real implementation.

import numpy as np

def sample_patch_pair(image, patch_size=96, gap=48, jitter=7):
    """Sample a center patch and one of 8 neighbors from an image."""
    h, w = image.shape[:2]
    stride = patch_size + gap  # distance between patch centers

    # Random center patch location
    cy = np.random.randint(stride, h - stride)
    cx = np.random.randint(stride, w - stride)

    # 8 relative positions: top-left, top, top-right, left, right, etc.
    offsets = [(-1,-1), (-1,0), (-1,1),
               ( 0,-1),         ( 0,1),
               ( 1,-1), ( 1,0), ( 1,1)]
    label = np.random.randint(8)
    dy, dx = offsets[label]

    # Add jitter to break exact alignment
    jy = np.random.randint(-jitter, jitter + 1)
    jx = np.random.randint(-jitter, jitter + 1)

    center = crop(image, cy, cx, patch_size)
    neighbor = crop(image, cy + dy*stride + jy,
                           cx + dx*stride + jx, patch_size)
    return center, neighbor, label   # label ∈ {0, ..., 7}

def drop_color_channels(patch):
    """Randomly keep only 1 of 3 color channels, noise the rest."""
    keep = np.random.randint(3)
    noisy = patch.copy().astype(float)
    sigma = patch[:, :, keep].std() / 100
    for c in range(3):
        if c != keep:
            noisy[:, :, c] = np.random.normal(0, sigma, patch.shape[:2])
    return noisy

# The network: two AlexNet branches with SHARED weights → fc6 each,
# concatenate → fc7 → fc8 → softmax over 8 classes.
# Batch normalization (no γ, β) prevents saddle-point collapse.

Downstream results: does the pretext task help real tasks?

The learned features were evaluated via two transfer-learning experiments:

Object Detection (PASCAL VOC 2007). The pre-trained conv layers were plugged into the R- pipeline and fine-tuned on VOC. The self-supervised model achieved 46.3% mAP with color dropping — a 6% boost over training from scratch, and the best unsupervised result at the time. With a VGG-style backbone and the rescaling trick from Krähenbühl et al., performance reached 61.7% mAP — remarkably close to the fully-supervised ImageNet-pretrained VGG at 68.6% mAP.

Surface Normal Estimation (NYUv2). The self-supervised features performed almost identically to fully-supervised ImageNet features on the geometry task — suggesting that the spatial reasoning learned from context transfers to 3D understanding, something ImageNet classification doesn't directly encourage.

Visual Data Mining. Using the learned features for nearest-neighbor matching and geometric verification, the system discovered object categories (cats, people, birds, monitors) from unlabeled Pascal VOC 2011 images — including deformable objects like birds and torsos that hand-crafted features couldn't handle.

Open in Lab
Click each stage to see how the pretext task flows into downstream applications.
The demo wakes as you arrive…

How hard is the pretext task?

The network achieved 38.4% on the 8-way classification task (chance is 12.5%), suggesting the task is genuinely hard — human performance is comparable. There was negligible : accuracy on ImageNet's training set was 39.5% vs. 40.3% on the .

An interesting experiment tested whether the task is easier on objects: patches sampled from within PASCAL ground-truth bounding boxes gave 39.2% accuracy — almost the same. Patches from cars specifically reached 45.6%. The network learns about both objects and scene layout.

Legacy: the pretext task era

This paper established a powerful recipe: invent a task that is free to label, hard to solve without understanding, and whose features transfer. The community responded with a flood of pretext tasks: predicting image rotations, solving jigsaw puzzles, colorizing grayscale images, predicting video frame order, and counting visual primitives.

But the deeper contribution is philosophical. Before this paper, unsupervised visual learning was largely about reconstruction — autoencoders, RBMs, sparse coding. This work showed that discrimination tasks (classification, prediction) can learn richer features than generative tasks, because they don't waste capacity on pixel-level details.

The pretext task paradigm eventually gave way to — SimCLR and MoCo showed that the specific pretext task matters less than the learning objective (pulling similar representations together, pushing dissimilar ones apart). But the idea that spatial structure within a single image contains enough signal to learn about the visual world? That started here.

  1. 2015

    Context Prediction (this paper)

    Predicting relative patch positions as a self-supervised pretext task. First demonstration that spatial context within a single image can train useful features.

  2. 2016

    Context Encoders (Inpainting)

    Predicting missing image regions (inpainting) as a pretext task, extending context prediction from relative positions to full region reconstruction.

  3. 2016

    Jigsaw Puzzles

    Solving permuted jigsaw puzzles as a pretext task — a multi-patch generalization of context prediction.

  4. 2018

    RotNet

    Predicting image rotation (0°, 90°, 180°, 270°). Simpler pretext task, comparable features — proved that spatial reasoning generalizes beyond patches.

  5. 2019

    MoCo

    Momentum Contrast shifted focus from pretext tasks to contrastive objectives. A momentum encoder and dynamic queue replaced hand-crafted puzzles with instance discrimination.

  6. 2020

    SimCLR

    Simple contrastive learning with large batches and strong augmentations. No pretext task at all — just pull augmented views of the same image together. The pretext era's spiritual successor.

CitationDoersch, Gupta, Efros. Unsupervised Visual Representation Learning by Context Prediction. ICCV, 2015.

Terms in this paper