AI Safety2015intermediate12 min read

Explaining and Harnessing Adversarial Examples

فهم الأمثلة الخصومية واستثمارها

Goodfellow, I. J. · Shlens, J. · Szegedy, C. — ICLR

The problem

Neural networks achieve impressive accuracy, yet they are shockingly fragile: adding an imperceptibly small perturbation to an input image can flip the 's prediction with high confidence. Szegedy et al. (2014) discovered this vulnerability but attributed it to the extreme nonlinearity and of deep networks. No fast method existed to generate these adversarial examples, making impractically slow.

The contribution

Goodfellow et al. overturn the nonlinearity hypothesis: adversarial vulnerability comes from linearity, not nonlinearity. In high dimensions, a 's with a small perturbation can grow linearly with the number of dimensions even when each perturbation element is tiny. This insight yields the Fast Sign Method (FGSM) — a single step that generates adversarial examples instantly. Using FGSM for adversarial reduces MNIST error from 0.94% to 0.78% and drops adversarial error from 89.4% to 17.9%.

The impact

FGSM became the standard baseline for adversarial attacks and a practical tool for adversarial training. The linear explanation unified a scattered field: it explains why adversarial examples transfer across models, why ensembles don't help, and why RBF networks are naturally immune. The paper launched the modern adversarial research program and directly inspired PGD, Mixup, certified defenses, and adversarial training at scale.

Imagine you're a teacher grading exams. You draw a straight line on the answer sheet to separate pass from fail. A student discovers that by smudging every answer just a tiny bit — each smudge too small to notice — the total shift pushes their score across the line. The smudges are invisible individually, but hundreds of tiny pushes in the same direction add up to a huge shove.

That's exactly what adversarial examples do to neural networks. The network's decision boundary is roughly a flat plane in a space with thousands of dimensions. A perturbation that's imperceptible -by-pixel can, across all those dimensions, shove the input across the boundary.

The mystery: invisible noise fools state-of-the-art models

In 2014, Szegedy et al. made a startling discovery: take any image correctly classified by a deep , add a carefully crafted perturbation so small that the two images look identical to the human eye, and the network confidently predicts a completely wrong class. A panda becomes a gibbon. A stop sign becomes a speed limit.

Three facts made the phenomenon especially puzzling:

  • The perturbations are imperceptible — each pixel changes by less than 1/255 of the dynamic range.
  • The same fools different architectures trained on different data. The trick is not model-specific.
  • Standard defenses — , , model averaging — offered no protection.
Open in Lab
The classic FGSM demonstration: a panda + tiny noise = "gibbon" at 99.3% confidence. Drag ε to see how perturbation strength affects the attack.
The demo wakes as you arrive…

The key insight: linearity is the cause, not nonlinearity

The prevailing theory blamed adversarial examples on deep networks being too nonlinear — highly curved decision boundaries with tiny pockets of misclassification. Goodfellow et al. argued the opposite: the real culprit is that modern networks are too linear.

Why linear? Because we designed them that way. activations are piecewise linear. gates operate in their linear regime. Maxout units select among linear functions. Even networks are tuned to avoid saturation — staying in the near-linear middle. This linearity makes networks easy to optimize with , but it also makes them easy to fool.

Here is the core mathematical argument. Consider a weight vector w\mathbf{w} and an input x\mathbf{x}. Now add a perturbation η\boldsymbol{\eta} to get the adversarial input x~=x+η\tilde{\mathbf{x}} = \mathbf{x} + \boldsymbol{\eta}. The dot product becomes:

w⊤x~=w⊤x+w⊤η\mathbf{w}^\top \tilde{\mathbf{x}} = \mathbf{w}^\top \mathbf{x} + \mathbf{w}^\top \boldsymbol{\eta}

The adversarial perturbation changes the activation by w⊤η\mathbf{w}^\top \boldsymbol{\eta}. To maximize this change while keeping each element of η\boldsymbol{\eta} tiny (bounded by ϵ\epsilon), we set η=ϵ⋅sign(w)\boldsymbol{\eta} = \epsilon \cdot \text{sign}(\mathbf{w}). The resulting activation shift is ϵ⋅∥w∥1\epsilon \cdot \|\mathbf{w}\|_1.

If w\mathbf{w} has nn dimensions with average element magnitude mm, the shift is ϵ⋅m⋅n\epsilon \cdot m \cdot n. Each element of η\boldsymbol{\eta} is at most ϵ\epsilon — invisible — but the total effect grows linearly with nn. In a space with thousands of dimensions, that tiny per-pixel nudge accumulates into a massive output change.

Open in Lab
Each dimension contributes a tiny ε push. Watch the total effect grow linearly with n.
The demo wakes as you arrive…

The Fast Gradient Sign Method (FGSM)

The linear explanation doesn't just explain adversarial examples — it immediately gives us a recipe to generate them. If linearity is the problem, then linearizing the around the current input gives the optimal perturbation direction.

The idea: compute the gradient of the with respect to the input, then take the sign of each element. This gives the direction that maximally increases the loss at each pixel while keeping the perturbation within the ℓ∞\ell_\infty ball of radius ϵ\epsilon.

One forward pass, one backward pass, one element-wise sign — and the adversarial example is ready. Compare this to the previous method by Szegedy et al., which required expensive L-BFGS over many iterations.

η=ϵ⋅sign ⁣(∇xJ(θ,x,y))\boldsymbol{\eta} = \epsilon \cdot \text{sign}\!\left(\nabla_{\mathbf{x}} J(\boldsymbol{\theta}, \mathbf{x}, y)\right)
The Fast Gradient Sign Method — adversarial perturbation in one step — θ = model parameters · x = clean input · y = true label · J = loss function · ∇ₓJ = gradient of the loss w.r.t. the input · sign = element-wise sign (−1, 0, or +1) · ε = perturbation budget (e.g. 0.007 for ImageNet, 0.25 for MNIST)

Think of FGSM as climbing the steepest hill on the loss landscape, but taking exactly one step of fixed size ϵ\epsilon in every dimension. The sign function ensures every pixel pushes the loss upward by the same amount — it's the most efficient use of a limited perturbation budget.

Open in Lab
Step through FGSM: forward pass → compute gradient → take sign → scale by ε → add to input.
The demo wakes as you arrive…

How effective is FGSM?

Devastatingly effective. A simple classifier on MNIST, attacked with ϵ=0.25\epsilon = 0.25, jumps from 1.6% error to 99.9% error — nearly every single example is misclassified. A maxout network, a much stronger model, still reaches 89.4% error under the same attack with 97.6% average confidence on incorrect predictions.

On CIFAR-10 with ϵ=0.1\epsilon = 0.1, a convolutional maxout network reaches 87.15% error with 96.6% confidence on wrong labels. The models are not just wrong — they are confidently wrong.

Open in Lab
Drag ε from 0 to 0.3 to see how error rate and confidence change. Even tiny ε values cause dramatic accuracy drops.
The demo wakes as you arrive…

Why adversarial examples transfer across models

One of the most puzzling facts about adversarial examples is their transferability: an adversarial image crafted to fool Model A often fools Model B too — even if B has a different architecture and was trained on different data. Under the nonlinearity hypothesis, this makes no sense: why would two very different nonlinear functions share the same blind spots?

The linear explanation makes this natural. Different models trained on the same task learn similar weight vectors, because generalizes — that's its whole point. If Model A and Model B have similar weight directions, then the adversarial perturbation sign(∇xJ)\text{sign}(\nabla_\mathbf{x} J) for Model A will have a large positive dot product with Model B's weights too. The perturbation moves in a broad direction in input space, not a fragile pinpoint, so it transfers readily.

Open in Lab
Tracing ε values from −15 to +15: adversarial misclassification is stable across a wide range, not a thin sliver.
The demo wakes as you arrive…

Harnessing adversarial examples: adversarial training

If FGSM can generate adversarial examples cheaply, why not train on them? The idea is simple but powerful: at each training step, generate adversarial versions of the current batch and train on both the clean and adversarial examples. The model learns to resist the strongest perturbation it can find, not just memorize the training set.

The adversarial training objective mixes the clean and adversarial losses:

J~(θ,x,y)=α J(θ,x,y)+(1−α) J ⁣(θ,x+ϵ⋅sign(∇xJ),y)\tilde{J}(\boldsymbol{\theta}, \mathbf{x}, y) = \alpha \, J(\boldsymbol{\theta}, \mathbf{x}, y) + (1 - \alpha) \, J\!\left(\boldsymbol{\theta}, \mathbf{x} + \epsilon \cdot \text{sign}(\nabla_{\mathbf{x}} J), y\right)
Adversarial training objective — learn from worst-case inputs — α = mixing coefficient (0.5 in the paper) · First term: loss on clean input · Second term: loss on FGSM-perturbed input · The model is updated to minimize both simultaneously.

The results are striking. On MNIST with a larger maxout network (1600 units per ):

  • Clean test error: improved from 0.94% to 0.78% — adversarial training acts as a regularizer, beating dropout alone.
  • Adversarial error: dropped from 89.4% to 17.9% — the model is far more robust.
  • The learned weights became more localized and interpretable, focusing on semantically meaningful features rather than diffuse -like patterns.

Adversarial training can be understood in multiple ways: as playing a game with an adversary, as minimizing an upper bound on the expected cost over noisy inputs, or as a form of where the model requests labels for the hardest examples.

Open in Lab
Compare weight visualizations: naively trained (noisy, diffuse) vs adversarially trained (localized, interpretable).
The demo wakes as you arrive…

What doesn't work — and why

The paper systematically eliminates competing defenses:

  • Dropout and standard : don't change the fundamental linear relationship between input perturbation and output. The model is still easy to fool.
  • Generative models: an MP-DBM (a with 0.88% MNIST error) reaches 97.5% error on adversarial examples. Being generative is not protective.
  • Ensembles: an ensemble of 12 maxout networks reaches 91.1% error on adversarial examples designed for the full ensemble. Averaging doesn't wash out adversarial perturbations.
  • Random noise: training with random ±ϵ\pm\epsilon noise yields 86.2% adversarial error vs 17.9% for FGSM adversarial training. Random noise is "easy" — FGSM finds the "hard" examples.

The one model family that is naturally resistant is RBF networks. They achieve only 55.4% error on adversarial examples because they respond with low confidence (1.2%) when fooled. But RBF networks lack the capacity to generalize well, creating a fundamental tension: easy-to-optimize linear models are easy to attack, while attack-resistant nonlinear models are hard to train.

Adversarial training vs weight decay: a subtle distinction

For logistic , the FGSM perturbation is exact (not an approximation). The adversarial training objective becomes minimizing:

Ex,y[ζ ⁣(y(∣∣w∣∣1−w⊤x−b))]\mathbb{E}_{x,y} \left[ \zeta\!\left(y(||\mathbf{w}||_1 - \mathbf{w}^\top \mathbf{x} - b)\right) \right]

This resembles L1L_1 regularization, but with a crucial difference: the L1L_1 penalty is subtracted from the activation, not added to the cost. When the model makes confident correct predictions, the penalty can saturate and deactivate. L1L_1 is more pessimistic — it never turns off, even when the model has sufficient margin. In practice, adversarial training with ϵ=0.25\epsilon = 0.25 works well on MNIST, while L1L_1 with coefficient 0.0025 already causes .

FGSM in code

FGSM attack and adversarial training — complete implementationpython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn.functional as F

def fgsm_attack(model, x, y, epsilon=0.25):
    """Generate adversarial examples using FGSM.
    x: input images (batch)
    y: true labels
    epsilon: perturbation budget
    """
    x.requires_grad_(True)
    loss = F.cross_entropy(model(x), y)
    loss.backward()

    # The core of FGSM: sign of the gradient, scaled by epsilon
    perturbation = epsilon * x.grad.sign()
    x_adv = (x + perturbation).clamp(0, 1)

    return x_adv

def adversarial_train_step(model, optimizer, x, y, epsilon=0.25, alpha=0.5):
    """One step of adversarial training: mix clean and FGSM losses."""
    # Clean loss
    clean_loss = F.cross_entropy(model(x), y)

    # Generate adversarial examples (detach model from FGSM graph)
    x_adv = fgsm_attack(model, x.clone(), y, epsilon)

    # Adversarial loss
    adv_loss = F.cross_entropy(model(x_adv.detach()), y)

    # Combined objective: α * clean + (1-α) * adversarial
    total_loss = alpha * clean_loss + (1 - alpha) * adv_loss

    optimizer.zero_grad()
    total_loss.backward()
    optimizer.step()

    return total_loss.item()

The capacity question: which models can resist?

Not all models are equally vulnerable. The paper draws a clear line between models that can learn robustness and those that cannot:

  • Linear models (softmax regression) cannot resist adversarial perturbation. They respond confidently in every direction of input space and have no mechanism to "say I don't know."
  • RBF networks naturally resist adversarial examples because they respond confidently only near their learned centers. Far from the data, they default to low confidence. But they sacrifice — they can't transfer knowledge to unseen regions.
  • Deep networks with hidden layers can, in principle, learn to resist adversarial perturbation — the universal approximator theorem guarantees this. But they need to be explicitly trained to do so. Standard training doesn't ask for robustness, so the model doesn't learn it.

This creates a spectrum: linear models are high- but low- (they respond to everything but are easily fooled), while RBF networks are high-precision but low-recall (they respond only to familiar inputs but miss novel ones).

Open in Lab
The precision-recall tradeoff: linear models respond to everything (easy to fool), RBF networks respond only to familiar inputs (hard to fool, but limited).
The demo wakes as you arrive…

The legacy of FGSM

This paper is a turning point. Before it, adversarial examples were a curiosity. After it, they became a research program. The linear explanation provided a unifying framework that made the phenomenon understandable and actionable.

  1. 2014

    Intriguing Properties (Szegedy et al.)

    Discovery of adversarial examples via L-BFGS optimization. Slow and expensive but proved the vulnerability exists.

  2. 2015

    FGSM (this paper)

    The linear explanation + a one-step attack + adversarial training. Made the field practical and unified.

  3. 2018

    PGD (Madry et al.)

    Projected Gradient Descent: multi-step FGSM that finds stronger adversarial examples. The gold standard for adversarial training.

  4. 2018

    Mixup (Zhang et al.)

    Training on convex combinations of examples and labels. A softer form of adversarial regularization that smooths the decision boundary.

  5. 2019

    Certified Defenses

    Randomized smoothing and other methods that provide mathematical guarantees of robustness within a certified radius.

  6. 2020

    Adversarial Training at Scale

    Adversarial training applied to ImageNet-scale models, bridging the gap between research benchmarks and real-world deployment.

CitationGoodfellow, Shlens, Szegedy. Explaining and Harnessing Adversarial Examples. ICLR, 2015.

Terms in this paper