AI Safety2015intermediate12 min read
Explaining and Harnessing Adversarial Examples
فهم الأمثلة الخصومية واستثمارها
Goodfellow, I. J. · Shlens, J. · Szegedy, C. — ICLR
The problem
Neural networks achieve impressive accuracy, yet they are shockingly fragile: adding an imperceptibly small perturbation to an input image can flip the 's prediction with high confidence. Szegedy et al. (2014) discovered this vulnerability but attributed it to the extreme nonlinearity and of deep networks. No fast method existed to generate these adversarial examples, making impractically slow.
The contribution
Goodfellow et al. overturn the nonlinearity hypothesis: adversarial vulnerability comes from linearity, not nonlinearity. In high dimensions, a 's with a small perturbation can grow linearly with the number of dimensions even when each perturbation element is tiny. This insight yields the Fast Sign Method (FGSM) — a single step that generates adversarial examples instantly. Using FGSM for adversarial reduces MNIST error from 0.94% to 0.78% and drops adversarial error from 89.4% to 17.9%.
The impact
FGSM became the standard baseline for adversarial attacks and a practical tool for adversarial training. The linear explanation unified a scattered field: it explains why adversarial examples transfer across models, why ensembles don't help, and why RBF networks are naturally immune. The paper launched the modern adversarial research program and directly inspired PGD, Mixup, certified defenses, and adversarial training at scale.
Imagine you're a teacher grading exams. You draw a straight line on the answer sheet to separate pass from fail. A student discovers that by smudging every answer just a tiny bit — each smudge too small to notice — the total shift pushes their score across the line. The smudges are invisible individually, but hundreds of tiny pushes in the same direction add up to a huge shove.
That's exactly what adversarial examples do to neural networks. The network's decision boundary is roughly a flat plane in a space with thousands of dimensions. A perturbation that's imperceptible -by-pixel can, across all those dimensions, shove the input across the boundary.
The mystery: invisible noise fools state-of-the-art models
In 2014, Szegedy et al. made a startling discovery: take any image correctly classified by a deep , add a carefully crafted perturbation so small that the two images look identical to the human eye, and the network confidently predicts a completely wrong class. A panda becomes a gibbon. A stop sign becomes a speed limit.
Three facts made the phenomenon especially puzzling:
The key insight: linearity is the cause, not nonlinearity
The prevailing theory blamed adversarial examples on deep networks being too nonlinear — highly curved decision boundaries with tiny pockets of misclassification. Goodfellow et al. argued the opposite: the real culprit is that modern networks are too linear.
Why linear? Because we designed them that way. activations are piecewise linear. gates operate in their linear regime. Maxout units select among linear functions. Even networks are tuned to avoid saturation — staying in the near-linear middle. This linearity makes networks easy to optimize with , but it also makes them easy to fool.
Here is the core mathematical argument. Consider a weight vector and an input . Now add a perturbation to get the adversarial input . The dot product becomes:
The adversarial perturbation changes the activation by . To maximize this change while keeping each element of tiny (bounded by ), we set . The resulting activation shift is .
If has dimensions with average element magnitude , the shift is . Each element of is at most — invisible — but the total effect grows linearly with . In a space with thousands of dimensions, that tiny per-pixel nudge accumulates into a massive output change.
The Fast Gradient Sign Method (FGSM)
The linear explanation doesn't just explain adversarial examples — it immediately gives us a recipe to generate them. If linearity is the problem, then linearizing the around the current input gives the optimal perturbation direction.
The idea: compute the gradient of the with respect to the input, then take the sign of each element. This gives the direction that maximally increases the loss at each pixel while keeping the perturbation within the ball of radius .
One forward pass, one backward pass, one element-wise sign — and the adversarial example is ready. Compare this to the previous method by Szegedy et al., which required expensive L-BFGS over many iterations.
Think of FGSM as climbing the steepest hill on the loss landscape, but taking exactly one step of fixed size in every dimension. The sign function ensures every pixel pushes the loss upward by the same amount — it's the most efficient use of a limited perturbation budget.
How effective is FGSM?
Devastatingly effective. A simple classifier on MNIST, attacked with , jumps from 1.6% error to 99.9% error — nearly every single example is misclassified. A maxout network, a much stronger model, still reaches 89.4% error under the same attack with 97.6% average confidence on incorrect predictions.
On CIFAR-10 with , a convolutional maxout network reaches 87.15% error with 96.6% confidence on wrong labels. The models are not just wrong — they are confidently wrong.
Why adversarial examples transfer across models
One of the most puzzling facts about adversarial examples is their transferability: an adversarial image crafted to fool Model A often fools Model B too — even if B has a different architecture and was trained on different data. Under the nonlinearity hypothesis, this makes no sense: why would two very different nonlinear functions share the same blind spots?
The linear explanation makes this natural. Different models trained on the same task learn similar weight vectors, because generalizes — that's its whole point. If Model A and Model B have similar weight directions, then the adversarial perturbation for Model A will have a large positive dot product with Model B's weights too. The perturbation moves in a broad direction in input space, not a fragile pinpoint, so it transfers readily.
Harnessing adversarial examples: adversarial training
If FGSM can generate adversarial examples cheaply, why not train on them? The idea is simple but powerful: at each training step, generate adversarial versions of the current batch and train on both the clean and adversarial examples. The model learns to resist the strongest perturbation it can find, not just memorize the training set.
The adversarial training objective mixes the clean and adversarial losses:
The results are striking. On MNIST with a larger maxout network (1600 units per ):
- Clean test error: improved from 0.94% to 0.78% — adversarial training acts as a regularizer, beating dropout alone.
- Adversarial error: dropped from 89.4% to 17.9% — the model is far more robust.
- The learned weights became more localized and interpretable, focusing on semantically meaningful features rather than diffuse -like patterns.
Adversarial training can be understood in multiple ways: as playing a game with an adversary, as minimizing an upper bound on the expected cost over noisy inputs, or as a form of where the model requests labels for the hardest examples.
What doesn't work — and why
The paper systematically eliminates competing defenses:
- Dropout and standard : don't change the fundamental linear relationship between input perturbation and output. The model is still easy to fool.
- Generative models: an MP-DBM (a with 0.88% MNIST error) reaches 97.5% error on adversarial examples. Being generative is not protective.
- Ensembles: an ensemble of 12 maxout networks reaches 91.1% error on adversarial examples designed for the full ensemble. Averaging doesn't wash out adversarial perturbations.
- Random noise: training with random noise yields 86.2% adversarial error vs 17.9% for FGSM adversarial training. Random noise is "easy" — FGSM finds the "hard" examples.
The one model family that is naturally resistant is RBF networks. They achieve only 55.4% error on adversarial examples because they respond with low confidence (1.2%) when fooled. But RBF networks lack the capacity to generalize well, creating a fundamental tension: easy-to-optimize linear models are easy to attack, while attack-resistant nonlinear models are hard to train.
Adversarial training vs weight decay: a subtle distinction
For logistic , the FGSM perturbation is exact (not an approximation). The adversarial training objective becomes minimizing:
This resembles regularization, but with a crucial difference: the penalty is subtracted from the activation, not added to the cost. When the model makes confident correct predictions, the penalty can saturate and deactivate. is more pessimistic — it never turns off, even when the model has sufficient margin. In practice, adversarial training with works well on MNIST, while with coefficient 0.0025 already causes .
FGSM in code
Simplified to show the idea — not the real implementation.
import torch
import torch.nn.functional as F
def fgsm_attack(model, x, y, epsilon=0.25):
"""Generate adversarial examples using FGSM.
x: input images (batch)
y: true labels
epsilon: perturbation budget
"""
x.requires_grad_(True)
loss = F.cross_entropy(model(x), y)
loss.backward()
# The core of FGSM: sign of the gradient, scaled by epsilon
perturbation = epsilon * x.grad.sign()
x_adv = (x + perturbation).clamp(0, 1)
return x_adv
def adversarial_train_step(model, optimizer, x, y, epsilon=0.25, alpha=0.5):
"""One step of adversarial training: mix clean and FGSM losses."""
# Clean loss
clean_loss = F.cross_entropy(model(x), y)
# Generate adversarial examples (detach model from FGSM graph)
x_adv = fgsm_attack(model, x.clone(), y, epsilon)
# Adversarial loss
adv_loss = F.cross_entropy(model(x_adv.detach()), y)
# Combined objective: α * clean + (1-α) * adversarial
total_loss = alpha * clean_loss + (1 - alpha) * adv_loss
optimizer.zero_grad()
total_loss.backward()
optimizer.step()
return total_loss.item()The capacity question: which models can resist?
Not all models are equally vulnerable. The paper draws a clear line between models that can learn robustness and those that cannot:
- Linear models (softmax regression) cannot resist adversarial perturbation. They respond confidently in every direction of input space and have no mechanism to "say I don't know."
- RBF networks naturally resist adversarial examples because they respond confidently only near their learned centers. Far from the data, they default to low confidence. But they sacrifice — they can't transfer knowledge to unseen regions.
- Deep networks with hidden layers can, in principle, learn to resist adversarial perturbation — the universal approximator theorem guarantees this. But they need to be explicitly trained to do so. Standard training doesn't ask for robustness, so the model doesn't learn it.
This creates a spectrum: linear models are high- but low- (they respond to everything but are easily fooled), while RBF networks are high-precision but low-recall (they respond only to familiar inputs but miss novel ones).
The legacy of FGSM
This paper is a turning point. Before it, adversarial examples were a curiosity. After it, they became a research program. The linear explanation provided a unifying framework that made the phenomenon understandable and actionable.
2014
Intriguing Properties (Szegedy et al.)
Discovery of adversarial examples via L-BFGS optimization. Slow and expensive but proved the vulnerability exists.
2015
FGSM (this paper)
The linear explanation + a one-step attack + adversarial training. Made the field practical and unified.
2018
PGD (Madry et al.)
Projected Gradient Descent: multi-step FGSM that finds stronger adversarial examples. The gold standard for adversarial training.
2018
Mixup (Zhang et al.)
Training on convex combinations of examples and labels. A softer form of adversarial regularization that smooths the decision boundary.
2019
Certified Defenses
Randomized smoothing and other methods that provide mathematical guarantees of robustness within a certified radius.
2020
Adversarial Training at Scale
Adversarial training applied to ImageNet-scale models, bridging the gap between research benchmarks and real-world deployment.
CitationGoodfellow, Shlens, Szegedy. Explaining and Harnessing Adversarial Examples. ICLR, 2015.
Terms in this paper
- Adversarial Exampleالعينات العدائية المضللة
- Adversarial Attackالهجوم العدائي الموجه
- Adversarial Trainingالتدريب التنافسي
- Robustnessالمتانة
- Gradientالتدرج التفاضلي
- Normالمعيار التفاضلي
- Regularizationالضبط الهيكلي
- ReLUدالة الوحدة الخطية المصححة
- Softmaxسوفت ماكس
- Generalizationالتعميم