Computer Vision2016intermediate11 min read

Image Style Transfer Using Convolutional Neural Networks

نقل الأسلوب الفني للصور باستخدام الشبكات العصبية الالتفافية

Gatys, L. A. · Ecker, A. S. · Bethge, M. — CVPR

The problem

Rendering a photograph in the style of a famous painting was a long-standing challenge. Previous texture-transfer algorithms used low-level statistics and could not separate semantic content from visual style. A method that extracts paint strokes from Van Gogh and applies them to a cityscape photo had no way to preserve recognizable buildings and sky while changing only the texture and color palette.

The contribution

The paper shows that a pre-trained for object recognition (VGG-19) already separates content from style internally. Content is captured by feature maps in higher layers; style is captured by the correlations between feature maps — encoded as Gram matrices — across multiple layers. By starting from and optimizing pixel values to simultaneously match the content features of one image and the Gram matrices of another, the algorithm produces stunning artistic renderings. The ratio α/β controls whether the result leans more toward the photograph or the painting.

The impact

This paper ignited the field of Neural Style Transfer and demonstrated that deep networks learn much more than classification — they learn rich, separable visual representations. It inspired real-time feed-forward style transfer, arbitrary style transfer, and eventually contributed to the vision behind modern text-to-image models like DALL·E. Apps like Prisma brought neural art to millions of phones within months.

Imagine you're an art student. Your teacher hands you a photograph of a city and says: "Paint this scene, but in Van Gogh's style."

You look at the photo and note what is in it — buildings, a river, the sky. Then you study Van Gogh's Starry Night and note how he paints — thick, swirling brushstrokes, vivid blues and yellows, turbulent texture.

You paint a new canvas that keeps the city's layout but uses Van Gogh's brushwork everywhere. That is exactly what this algorithm does — except the "eye" that sees content and style is a pre-trained neural network, and the "hand" that paints is .

The problem: content and style are tangled

When we look at a photograph of a building painted by Monet, we effortlessly separate two things: the content (a building by a river) and the style (soft pastel colors, visible brushwork). But for a computer, an image is just a grid of numbers. Traditional image-processing tools had no principled way to tease these apart.

Previous texture-transfer methods used hand-designed statistics — histograms of oriented gradients, wavelet coefficients — that captured low-level patterns well but couldn't understand "there is a building here." So they either destroyed the objects in the photo or failed to transfer the painting's character.

The key insight: convolutional neural networks trained for object recognition already build internal representations that separate what is in an image from how it looks. We just need to find where each lives inside the network and how to extract it.

VGG-19: the eye that sees content and style

The authors use VGG-19, a 19-layer convolutional network trained on ImageNet for object recognition. They do not retrain it or add any layers — they simply pass images through it and read off the internal activations.

Recall that a CNN builds a hierarchy of feature maps: early layers detect edges and colors, middle layers detect textures and patterns, and deep layers detect entire objects and their arrangements. This hierarchy is exactly the lever we need:

  • Content lives in the deep layers — where the network knows "this is a building" but has forgotten the exact pixel colors.
  • Style lives in the correlations across feature maps at every level — the texture fingerprint that says "thick swirling strokes in blue and yellow" without caring where the strokes are.
Open in Lab
Click through VGG layers to see how content becomes abstract and style patterns emerge at different depths.
The demo wakes as you arrive…

Content representation: what is in the image

Pass a photograph through VGG-19 and record the feature maps at a chosen deep layer (the paper uses conv4_2). These feature maps form the content : a set of activation patterns that encode the spatial arrangement of objects — buildings, sky, river — but have shed exact pixel values and textures.

Think of it like a detailed pencil sketch: the shapes and layout are all there, but the colors and brushwork are gone. To measure how close a generated image's content is to the original photo, we compare their feature maps at this layer.

Lcontent(p⃗,x⃗,l)=12∑i,j(Fijl−Pijl)2\mathcal{L}_{\text{content}}(\vec{p}, \vec{x}, l) = \frac{1}{2} \sum_{i,j} \left( F_{ij}^{l} - P_{ij}^{l} \right)^2
Content loss — squared difference of feature maps — F = feature maps of the generated image · P = feature maps of the content photo · both at layer l · the loss is small when the generated image has the same high-level structure as the photo

Read it as a difference report: for every filter in the chosen layer, how different is the generated image's response from the photo's response? If the generated image has a building in the same place as the photo, those filter responses will match closely, even if the colors and textures are completely different.

Style representation: the Gram matrix fingerprint

Style is harder to pin down than content. A Van Gogh painting has a distinctive texture — swirling strokes, thick impasto, vivid complementary colors — that is present everywhere, not tied to any particular object.

The key idea: style is captured by which feature maps fire together. If, across the entire image, a "blue" filter and a "curved edge" filter tend to activate at the same locations, that correlation IS the style signal. The mathematical tool for capturing all such pairwise correlations is the .

To build the Gram matrix for a given layer: take all NN feature maps (each of size MM pixels), flatten each into a row vector, and compute the N×NN \times N matrix of dot products. Entry (i,j)(i,j) measures how much feature ii and feature jj co-activate across the image. The Gram matrix is essentially a correlation fingerprint of how different visual patterns relate to each other — the "DNA" of the artist's style.

Gijl=∑kFikl FjklG_{ij}^{l} = \sum_{k} F_{ik}^{l} \, F_{jk}^{l}
Gram matrix — correlations between feature maps — F = feature maps at layer l · each entry G(i,j) is the dot product of (flattened) feature map i with feature map j · this discards spatial information while keeping texture statistics

Think of each as a specialist reporter covering the image: one reports "edges here," another reports "blue here," a third reports "curves here." The Gram matrix records which reporters tend to file their stories from the same locations. In a Van Gogh painting, the "curves" reporter and the "blue" reporter frequently agree — that pattern of agreement is what makes the painting look like a Van Gogh.

Crucially, the Gram matrix sums over all spatial positions, so it captures what patterns co-occur without caring where. This is exactly what we want: style is location-independent.

Open in Lab
Watch how two feature maps produce a Gram matrix entry. High co-activation means the two patterns tend to appear together.
The demo wakes as you arrive…

Style loss: matching the texture fingerprint

The compares the Gram matrices of the generated image against the Gram matrices of the style painting at multiple VGG layers (typically conv1_1 through conv5_1). Using multiple layers is critical: low layers capture fine textures (brush hair marks), middle layers capture medium patterns (swirl shapes), and deep layers capture large-scale style (color palette arrangements).

El=14Nl2Ml2∑i,j(Gijl−Aijl)2E_{l} = \frac{1}{4 N_{l}^{2} M_{l}^{2}} \sum_{i,j} \left( G_{ij}^{l} - A_{ij}^{l} \right)^2
Style loss per layer — G = Gram matrix of the generated image · A = Gram matrix of the style image · N = number of feature maps · M = feature map size · the normalization ensures layers with different sizes contribute equally
Lstyle(a⃗,x⃗)=∑l=0Lwl El\mathcal{L}_{\text{style}}(\vec{a}, \vec{x}) = \sum_{l=0}^{L} w_{l} \, E_{l}
Total style loss — weighted sum across layers — w_l = weight for layer l · summing across layers captures style at every scale from fine brush marks to broad composition

Putting it together: the α/β balancing act

The total loss is a weighted sum of and style loss. The algorithm starts from a white noise image and iteratively adjusts its pixels using gradient descent (specifically L-BFGS, a quasi-Newton optimizer) to minimize this combined loss.

The weights α and β control the trade-off: high α/β preserves the photograph's structure with subtle style hints; low α/β lets the painting's texture dominate, turning the photo into an almost abstract artwork.

This is not training a network — the VGG weights are frozen. We are optimizing the image itself. Each pixel is a variable, and gradient descent sculpts the pixel values until the image simultaneously looks like the photo (in terms of feature maps) and feels like the painting (in terms of Gram matrices).

Ltotal(p⃗,a⃗,x⃗)=α Lcontent(p⃗,x⃗)+β Lstyle(a⃗,x⃗)\mathcal{L}_{\text{total}}(\vec{p}, \vec{a}, \vec{x}) = \alpha \, \mathcal{L}_{\text{content}}(\vec{p}, \vec{x}) + \beta \, \mathcal{L}_{\text{style}}(\vec{a}, \vec{x})
Total loss — the single objective — p = content photo · a = style artwork · x = generated image (starts as noise) · α = content weight · β = style weight · gradient descent adjusts every pixel of x to minimize this
Open in Lab
Drag α/β to see how the balance shifts between preserving the photo and adopting the painting's style.
The demo wakes as you arrive…

The full pipeline

Here is the complete algorithm step by step:

  1. Pass the content photo through VGG-19 and record feature maps at conv4_2 → this is the content target.
  2. Pass the style painting through VGG-19 and compute Gram matrices at conv1_1, conv2_1, conv3_1, conv4_1, conv5_1 → these are the style targets.
  3. Initialize the output image from white noise (or from the content photo for faster convergence).
  4. In each iteration: pass the output image through VGG-19, compute content loss against the content target and style loss against the style targets, sum them with weights α and β, backpropagate to compute gradients with respect to the pixels of the output image, and update the pixels using L-BFGS.
  5. Repeat until convergence (typically a few hundred iterations).
Open in Lab
Step through the full pipeline. Watch how the noise image gradually transforms into a styled result.
The demo wakes as you arrive…

The same idea in code

Neural style transfer — content loss, style loss, and optimizationpython

Simplified to show the idea — not the real implementation.

import numpy as np

def content_loss(F, P):
    """Squared difference between feature maps of generated (F) and content (P)."""
    return 0.5 * np.sum((F - P) ** 2)

def gram_matrix(F):
    """Gram matrix: correlations between feature maps.
    F shape: (N_filters, M_pixels) — each row is a flattened feature map."""
    return F @ F.T    # (N, N) — entry (i,j) = how much filter i and j co-activate

def style_loss_layer(F, A):
    """Style loss for one layer. F = generated, A = style target."""
    G = gram_matrix(F)       # Gram matrix of generated image
    N, M = F.shape
    return np.sum((G - A) ** 2) / (4 * N**2 * M**2)

def total_loss(content_F, content_P, style_Fs, style_As, alpha, beta):
    """The full objective: alpha * content + beta * style."""
    L_content = content_loss(content_F, content_P)
    L_style = sum(style_loss_layer(F, A) for F, A in zip(style_Fs, style_As))
    return alpha * L_content + beta * L_style

# The "magic": we don't train the network. We freeze VGG and optimize
# the IMAGE PIXELS themselves. Each pixel is a variable, and L-BFGS
# updates every pixel to minimize total_loss.
# Start from white noise → iterate → out comes art.

Why do Gram matrices capture style?

This is a deep question the paper answers empirically but later work has formalized. The Gram matrix can be understood as matching the distribution of features across the image. Later research showed that minimizing the Gram matrix difference is equivalent to minimizing a specific form of Maximum Mean Discrepancy — a statistical distance between distributions.

Intuitively: two images have the same style if the "recipe" of visual ingredients is the same — the same proportions of edges, colors, and textures co-occurring, regardless of where they appear. The Gram matrix is exactly that recipe card.

The paper also shows that reconstructing an image from Gram matrices alone (without any content constraint) produces textures that look like the original painting — wavy blue and yellow strokes from Starry Night, for instance — confirming that Gram matrices truly encode style.

Open in Lab
See how reconstructing from Gram matrices alone produces texture that matches the original painting's style.
The demo wakes as you arrive…

Key design choices and their effects

Several decisions shape the final result:

  • Which content layer? Deeper layers preserve abstract structure; shallower layers preserve more exact pixel details. The paper uses conv4_2 — deep enough to capture objects but abstract enough to allow style freedom.
  • Which style layers? Using only one layer captures style at one scale. Using all five (conv1_1 through conv5_1) captures everything from fine grain to broad composition.
  • Initialization. Starting from white noise produces artistic results but converges slowly. Starting from the content photo converges faster and preserves more structure.
  • vs. . The authors found that replacing max with average pooling slightly improves gradient flow and produces smoother results.

Why it mattered

  1. 2015

    Gatys et al. — Neural Style Transfer

    The original arXiv paper showed that VGG feature maps and Gram matrices can separate content from style, enabling artistic rendering of photographs.

  2. 2016

    Johnson et al. — Feed-forward Style Transfer

    Trained a feed-forward network to apply a fixed style in a single forward pass instead of iterative optimization — real-time style transfer became possible.

  3. 2016

    Prisma app — Neural art goes mainstream

    Prisma brought neural style transfer to millions of smartphones, becoming one of the most downloaded apps. It demonstrated consumer appetite for AI-powered creative tools.

  4. 2017

    Pix2Pix — Paired image-to-image translation

    Isola et al. used conditional GANs for paired image translation. The perceptual loss framework from style transfer influenced its feature-matching approach.

  5. 2017

    CycleGAN — Unpaired image translation

    Zhu et al. enabled style transfer between domains without paired examples — horses ↔ zebras, summer ↔ winter — extending the content-style separation idea.

  6. 2021

    DALL·E — Text-to-image generation

    OpenAI's DALL·E showed that the content-style separation insight scales to text-conditioned generation. Users could request "a painting of a cat in the style of Monet" — neural style understanding made general.

CitationGatys, Ecker, Bethge. Image Style Transfer Using Convolutional Neural Networks. CVPR, 2016.

Terms in this paper