Computer Vision2017intermediate9 min read

Image-to-Image Translation with Conditional Adversarial Networks

ترجمة الصور بالشبكات التنافسية التوليدية المشروطة

Isola, P. · Zhu, J.-Y. · Zhou, T. · Efros, A. A. — CVPR

The problem

Image-to-image translation — turning segmentation maps into photos, edges into objects, day scenes into night — is a recurring challenge across . Before pix2pix, every translation task required its own hand-crafted : L1 for reconstruction, for style, domain-specific metrics for each application. This meant reinventing the objective from scratch every time, and the resulting L1/L2 losses produced blurry outputs because they average over all plausible outputs.

The contribution

A general-purpose framework for paired image-to-image translation using conditional GANs. Three key innovations: (1) the on an input image so the learns a mapping rather than sampling from noise, (2) a generator with skip connections that preserve fine spatial details across the , and (3) a PatchGAN that classifies overlapping image patches rather than the whole image, focusing the adversarial signal on local texture realism. The combined L1 + produces sharp, realistic outputs without task-specific engineering.

The impact

pix2pix established conditional GANs as a practical tool and became the standard baseline for paired image translation. It inspired CycleGAN (unpaired translation), pix2pixHD (high-resolution synthesis), SPADE (semantic synthesis), and an entire ecosystem of GAN-based image manipulation tools. Artists and researchers worldwide adopted it for applications from architectural rendering to medical imaging, making it one of the most influential papers in generative computer vision.

Imagine a translation booth at the United Nations. On one side sits a speaker delivering a speech in one language (an edge map, a segmentation mask, a sketch). On the other side sits a translator (the Generator) who must produce the same speech in a completely different language (a photorealistic image).

But there's a twist: a fact-checker (the Discriminator) listens to both the original and the translation and decides whether the translation sounds native or obviously machine-generated.

The translator keeps improving until the fact-checker can't tell real translations from generated ones. And critically, this setup works for any pair of languages — you don't need to redesign the booth for each one.

The problem: every pixel-to-pixel task needs its own loss

Many problems in computer vision are really the same underlying task: take a 2D image as input and produce a different 2D image as output. turns photos into label maps. Colorization turns grayscale into color. turns photos into line drawings. turns thumbnails into high-res images.

Before pix2pix, researchers treated each of these as a separate problem, hand-designing a specific loss function for each one. And the simplest pixel-level losses like L1 or L2 had a fatal flaw: they average over all plausible outputs. If there are two equally valid colors for a pixel, L2 picks the mean — which is neither. The result: universally blurry outputs.

Open in Lab
Toggle between L1 loss (blurry average) and adversarial loss (sharp, realistic). Notice how L1 hedges its bets while the GAN commits to one plausible output.
The demo wakes as you arrive…

The idea: let the network learn the loss function

The core insight of pix2pix is deceptively simple: instead of hand-designing a loss function, learn it. A uses two networks in opposition. The Generator takes an input image and produces an output image. The Discriminator sees both the input and the output and decides whether the output is a real photograph or a generated fake. The Generator improves by fooling the Discriminator; the Discriminator improves by catching fakes.

The word conditional is key. Unlike a standard GAN that generates images from random noise, a conditional GAN is conditioned on an input image — the Generator doesn't invent from nothing, it translates. The Discriminator doesn't just ask "is this image real?", it asks "is this a real translation of that specific input?".

Open in Lab
Watch the full pix2pix pipeline: input image → Generator → output. The Discriminator checks both the input-output pair against real pairs.
The demo wakes as you arrive…

The objective function: adversarial loss + L1

The training objective combines two complementary forces. The adversarial loss pushes the Generator to produce outputs that look real — it handles high-frequency details and textures. The L1 loss pushes the Generator to be accurate — it captures the overall structure and low-frequency content. Neither alone is sufficient: adversarial loss without L1 produces vivid but structurally incorrect images; L1 without adversarial loss produces correct but blurry images. Together, they get both right.

LcGAN(G,D)=Ex,y[log⁡D(x,y)]+Ex,z[log⁡(1−D(x,G(x,z)))]\mathcal{L}_{\text{cGAN}}(G, D) = \mathbb{E}_{x,y}[\log D(x, y)] + \mathbb{E}_{x,z}[\log(1 - D(x, G(x, z)))]
Conditional adversarial loss — x = input image · y = real target · z = noise · G(x,z) = generated output · D judges pairs: D(x,y) should → 1 for real pairs, D(x,G(x,z)) should → 0 for fake pairs. G tries to maximize the chance D makes a mistake.
LL1(G)=Ex,y,z[∥y−G(x,z)∥1]\mathcal{L}_{L1}(G) = \mathbb{E}_{x,y,z}\left[\|y - G(x, z)\|_1\right]
L1 reconstruction loss — The mean absolute difference between the generated image and the ground truth. Captures the global structure — colors, shapes, layout — while the adversarial loss handles the fine texture.
G∗=arg⁡min⁡Gmax⁡D  LcGAN(G,D)+λ LL1(G)G^* = \arg\min_G \max_D\; \mathcal{L}_{\text{cGAN}}(G, D) + \lambda\,\mathcal{L}_{L1}(G)
The full pix2pix objective — λ = 100 in all experiments. The Generator minimizes, the Discriminator maximizes. The large λ ensures structural accuracy while the adversarial term adds realism.

The Generator: U-Net with skip connections

A standard - compresses the input into a small bottleneck vector, then decompresses it back to an image. The problem? Fine details — edges, textures, exact positions — get lost in the bottleneck. Every pixel must squeeze through a narrow information highway.

The U-Net solves this by adding skip connections: each encoder layer is directly connected to its mirror decoder layer. Think of it as building a shortcut bridge across the U-shape. The encoder's early layers capture low-level details (edges, colors); these details bypass the bottleneck entirely and flow directly to the decoder. The bottleneck carries only the high-level semantic understanding, while the skip connections carry the spatial precision.

The result: the Generator can produce outputs that are both semantically correct (right objects in right places) and spatially precise (sharp edges, aligned textures).

Open in Lab
Click layers to see how information flows: through the bottleneck (semantic path) and via skip connections (detail path). Removing skip connections loses sharp details.
The demo wakes as you arrive…

Concretely, the U-Net encoder has 8 downsampling layers: each applies a with stride 2 (halving spatial dimensions), , and activation. The decoder mirrors this with 8 upsampling layers using transposed convolutions. At each level, the decoder concatenates the skip-connected encoder features with its own upsampled features before the next convolution — doubling the channel count and preserving both semantic and spatial information.

The Discriminator: PatchGAN — think locally

A traditional discriminator looks at the entire image and outputs a single real/fake score. But L1 already handles the global, low-frequency structure well enough. What's missing is high-frequency realism: sharp edges, realistic textures, plausible micro-patterns.

PatchGAN flips the approach: instead of one score for the whole image, it outputs a grid of scores, one for each overlapping N×N patch. Each score says "is this local region real or fake?" The discriminator becomes a texture critic that examines local neighborhoods independently.

The default 70×70 patch size hit the sweet spot in the paper's experiments: large enough to capture meaningful texture patterns, small enough to apply across different image sizes. Compared to a full-image discriminator, PatchGAN has far fewer parameters and runs faster, while producing outputs that are sharper and more textured.

Open in Lab
Hover over the image to see the PatchGAN receptive field. Each patch gets its own real/fake score. Compare patch sizes: 1×1 (PixelGAN), 70×70, 286×286 (full image).
The demo wakes as you arrive…

Putting it all together: the full architecture

The full pix2pix system combines three ideas into one coherent pipeline:

  • Conditioning: the Generator and Discriminator both receive the input image, so the system learns to translate rather than generate from scratch.
  • U-Net skip connections: the Generator preserves spatial details that would be lost in a standard bottleneck architecture.
  • PatchGAN: the Discriminator focuses on local texture realism, complementing the L1 loss that handles global structure.

Training alternates between updating the Discriminator (to better distinguish real from fake pairs) and updating the Generator (to better fool the Discriminator while staying close to the ground truth via L1). The Generator uses as its only source of stochasticity — traditional noise vectors z had no observable effect in the authors' experiments.

Open in Lab
Click on any component to see its role. Toggle skip connections on/off to see their effect.
The demo wakes as you arrive…

The same idea in code

Simplified pix2pix training looppython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn as nn

# --- Core pix2pix training step ---

def train_step(G, D, real_input, real_target, opt_G, opt_D, lambda_L1=100):
    """One training iteration for pix2pix."""

    # 1. Generate a fake image from the input
    fake_target = G(real_input)

    # 2. Train the Discriminator
    #    It sees (input, real_target) → should output 1 (real)
    #    It sees (input, fake_target) → should output 0 (fake)
    pred_real = D(real_input, real_target)       # patch grid of scores
    pred_fake = D(real_input, fake_target.detach())
    loss_D = 0.5 * (bce(pred_real, ones) + bce(pred_fake, zeros))

    opt_D.zero_grad()
    loss_D.backward()
    opt_D.step()

    # 3. Train the Generator
    #    Adversarial: fool D into thinking fake is real
    #    L1: stay close to the ground truth
    pred_fake_for_G = D(real_input, fake_target)
    loss_G_adv = bce(pred_fake_for_G, ones)     # fool D
    loss_G_L1  = nn.L1Loss()(fake_target, real_target)
    loss_G = loss_G_adv + lambda_L1 * loss_G_L1

    opt_G.zero_grad()
    loss_G.backward()
    opt_G.step()

    return loss_D.item(), loss_G.item()

# bce = nn.BCEWithLogitsLoss()
# ones/zeros = torch.ones/zeros matching the PatchGAN output grid

Results: one architecture, many tasks

The authors demonstrated pix2pix on a striking variety of tasks — all with the same architecture and hyperparameters, only swapping the training data:

  • Labels → Street scenes (Cityscapes)
  • Architectural labels → Building facades
  • Grayscale → Color (colorization)
  • Edges → Photos (handbags, shoes)
  • Aerial photos → Maps (and maps → aerial)
  • Day → Night
  • BW → Color

In Amazon Mechanical Turk evaluations, generated images from several tasks fooled human judges about 20% of the time — meaning one in five generated images was indistinguishable from a real photograph. The PatchGAN discriminator with 70×70 patches consistently produced the sharpest and most realistic results across all tasks.

Open in Lab
Browse through different translation tasks. The same network architecture, different training data.
The demo wakes as you arrive…

Why it mattered

pix2pix opened two major research directions: (1) unpaired translation with CycleGAN, which removed the need for paired training data entirely, and (2) high-resolution and semantically-guided synthesis, leading to pix2pixHD, SPADE, and eventually the generative AI tools that power modern image editing applications.

  1. 2014

    GAN (Goodfellow et al.)

    The foundational framework: a generator and discriminator trained adversarially. Generated images from noise, but no control over what is generated.

  2. 2014

    cGAN (Mirza & Osindero)

    Added conditional information (class labels) to both generator and discriminator, enabling controlled generation. The conceptual predecessor of pix2pix.

  3. 2015

    U-Net (Ronneberger et al.)

    Encoder-decoder with skip connections for biomedical segmentation. pix2pix adopted this architecture as its generator backbone.

  4. 2017

    pix2pix (this paper)

    Unified conditional GAN framework for general-purpose paired image translation. Introduced PatchGAN and the L1 + adversarial objective.

  5. 2017

    CycleGAN (Zhu et al.)

    Extended image translation to unpaired data using cycle-consistency loss. No need for matched input-output pairs.

  6. 2018

    pix2pixHD (Wang et al.)

    Scaled pix2pix to 2048×1024 resolution with multi-scale generators and discriminators, enabling photorealistic image synthesis.

  7. 2019

    SPADE / GauGAN (Park et al.)

    Semantic-layout-driven synthesis using spatially-adaptive normalization. Users draw semantic maps and get photorealistic landscapes.

CitationIsola, Zhu, Zhou, Efros. Image-to-Image Translation with Conditional Adversarial Networks. CVPR, 2017.

Terms in this paper