Computer Vision2017intermediate12 min read
Dynamic Routing Between Capsules
التوجيه الديناميكي بين الكبسولات
Sabour, S. · Frosst, N. · Hinton, G. E. — NeurIPS
The problem
CNNs use to achieve translation , but this throws away precise spatial information — the exact position, orientation, and scale of detected features. A knows a face has two eyes and a mouth, but doesn't check whether they are in the right spatial arrangement. It can be fooled by a scrambled face. Worse, CNNs need exponentially more data or detectors to handle viewpoint changes beyond translation.
The contribution
Capsule Networks: instead of scalar feature detectors, use groups of neurons (capsules) whose output vectors encode both the presence AND the of an entity. A -by-agreement iteratively sends each lower-level capsule's output to the higher-level capsule whose own output best matches the prediction — like puzzle pieces finding their correct assembly. The system achieves state-of-the-art on MNIST and excels at segmenting highly overlapping digits.
The impact
Capsule Networks introduced a fundamentally different way to think about visual : entities as vectors with pose, not scalars with presence. Though they haven't replaced CNNs at scale, they inspired research on equivariant representations, part-whole hierarchies, and alternatives to that preserve geometric information — ideas that echo in Vision Transformers and geometric deep learning.
A CNN detecting a face is like a checklist inspector: ✓ two eyes, ✓ nose, ✓ mouth — "it's a face!" even if the eyes are below the mouth and the nose is sideways. It checks what's there but not how the pieces fit together.
A is more like an architect reading a blueprint: each part reports not just "I exist" but "here's my exact position, size, and angle." The architect then checks whether all the parts agree on where the whole should be — only then does she sign off on the building.
The problem: max pooling destroys spatial relationships
Convolutional Neural Networks detect features using learned filters — edges, textures, parts — and that works brilliantly (as we saw with LeNet). But there is a critical weakness in how CNNs handle the arrangement of those features:
-
Max pooling discards position. After detecting "eye" and "mouth," max pooling keeps only that they were found somewhere in a local region. The network cannot verify whether the eye is above the mouth — it just needs both activations to be high.
-
Viewpoint changes require exponential resources. Translation invariance is built in, but rotation, scaling, and 3D viewpoint changes are not. To handle them, a CNN must either replicate filters at many orientations (exponential growth) or see every viewpoint in data (exponential data).
-
No part-whole reasoning. A CNN has no mechanism to ask: "do these detected parts agree on where the whole object should be?" It treats features as independent votes, not as a geometric puzzle.
The idea: from scalar neurons to vector capsules
The core insight is deceptively simple: replace individual neurons that output a single number with small groups of neurons (capsules) that output a .
In a CNN, a says: "edge detected here, with activation 0.9." That scalar tells you how strongly a feature was detected, but nothing about its pose — its position, orientation, scale, thickness, or deformation.
A capsule says something richer:
-
The length of its output vector (between 0 and 1) represents the probability that the entity exists — like the old scalar activation.
-
The orientation (direction) of the vector represents the instantiation parameters — the pose: where the entity is, how big it is, how it is oriented, how it is deformed.
This is a fundamentally different kind of representation. A single capsule for the concept "eye" doesn't just fire strongly when an eye is present — its vector encodes which eye, in what position, at what angle. Think of it as upgrading from a light switch (on/off) to a compass (direction and magnitude).
The squashing function: keeping vectors in [0, 1]
Since the length of a capsule's output vector represents a probability, it must stay between 0 and 1. But we also need to preserve the vector's direction (which encodes the pose). The solves both problems: it shrinks short vectors toward zero and long vectors toward length 1, while keeping the orientation unchanged.
Think of it like a tube that compresses water flow: a trickle stays a trickle, a torrent gets capped just below maximum, but the direction of flow never changes.
Dynamic routing: parts vote for wholes
This is the paper's most important contribution. How does a lower-level capsule (say, "eye") decide which higher-level capsule ("face" vs "car") to send its output to?
The answer is routing by agreement — an iterative process that works like a committee of experts finding consensus:
-
Predict. Each lower capsule makes a for every possible parent capsule by multiplying its output by a learned transformation matrix : . This is the capsule saying: "if I am part of entity , then should have this pose."
-
Weight and sum. Each parent sums these predictions, weighted by coupling coefficients (which start equal), then squashes the sum to get its output .
-
Measure agreement. If a prediction closely matches the parent's output (high ), the increases — the part "agrees" with this whole, so send more signal there.
-
Repeat for a few iterations (typically 3). Parts that agree with a whole reinforce each other; parts that disagree get routed elsewhere.
This is like an eye capsule saying "the face should be here at this angle," a nose capsule saying "the face should be here at this angle," and if they agree — the face capsule activates strongly.
The routing algorithm step by step
The routing algorithm runs between every pair of adjacent capsule layers. Here is the complete procedure:
- Initialize all routing logits , so every lower capsule sends equal weight to all parents.
- For iterations (typically 3):
- Compute coupling coefficients
- Compute each parent's weighted input:
- Apply squashing:
- Update routing logits:
The dot product is the agreement — it measures how well capsule 's prediction for parent matched what actually computed. High agreement increases the coupling; low agreement decreases it.
CapsNet architecture
The paper presents a simple three- architecture for MNIST digit recognition:
Layer 1 — Conv1: A standard with 256 kernels of size 9×9 and activation. This converts pixel intensities into local feature detectors — nothing capsule-specific yet.
Layer 2 — PrimaryCapsules: A convolutional capsule layer with 32 channels of 8D capsules, using 9×9 kernels with 2. This produces 32 × 6 × 6 = 1,152 capsule outputs, each an 8-dimensional vector. These are the lowest-level entities: oriented edges, simple shapes. The squashing function replaces ReLU.
Layer 3 — DigitCaps: 10 capsules (one per digit class), each 16-dimensional. Every DigitCaps capsule receives input from all 1,152 PrimaryCapsules via learned transformation matrices. Dynamic routing runs between these two capsule layers.
The length of each DigitCaps vector gives the probability. With only 8.2M parameters (vs 35.4M for a comparable CNN baseline), CapsNet achieves 0.25% test error on MNIST.
Margin loss: one loss per capsule
Since each DigitCaps capsule independently represents whether a digit class is present, the paper uses a separate for each capsule. This is crucial for allowing the network to detect multiple digits simultaneously (as in the overlapping digits task).
The intuition: if digit is present (), we want its capsule's length to be at least . If digit is absent (), we want the length to be at most . The penalizes violations of these margins.
The down-weighting factor on the absent-class term prevents the network from shrinking all capsule vectors to zero at the start of training — a critical practical detail.
Reconstruction: proving the capsule encodes pose
To verify that DigitCaps vectors actually encode meaningful pose information — not just a discriminative signal — the paper adds a reconstruction : three fully-connected layers (512 → 1024 → 784) that reconstruct the input image from the capsule vector of the correct digit.
During training, all vectors except the correct digit's are masked to zero. The decoder then tries to reconstruct the original 28×28 image from just the 16D vector. The reconstruction loss (mean squared error, scaled by 0.0005) is added to the margin loss.
This serves as a powerful regularizer: it forces the capsule to encode the specific visual details of this digit — its thickness, slant, width — not just "it's a 7." When you perturb individual dimensions of the DigitCaps vector and feed them to the decoder, you can see each dimension controlling a specific visual property: stroke thickness, width, rotation, localized deformations.
Overlapping digits: where capsules truly shine
The most compelling result in the paper is on MultiMNIST: two digits from different classes overlaid on the same image, with bounding boxes overlapping by about 80%.
A standard CNN struggles because max pooling merges features from both digits into an indistinguishable soup. CapsNet, using routing-by-agreement, can segment the two digits without pixel-level supervision — each DigitCaps capsule "claims" the parts that agree with its predicted pose, and the other capsule claims the rest.
The reconstruction decoder proves this works: when you reconstruct each digit separately from its capsule vector, you get clean, separated images of each digit — even though they were physically overlapping in the input. CapsNet achieves 5.2% error on MultiMNIST with just 11.36M parameters, matching a sequential model that was tested on a much easier (less overlapping) version of the task.
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def squash(s):
"""Non-linear squashing: short → near 0, long → near 1, direction preserved."""
sq_norm = np.sum(s ** 2, axis=-1, keepdims=True)
scale = sq_norm / (1 + sq_norm)
return scale * s / (np.sqrt(sq_norm) + 1e-8)
def routing(u_hat, r=3):
"""
u_hat: (n_lower, n_upper, dim_upper) — prediction vectors.
Returns: (n_upper, dim_upper) — output of upper capsules.
"""
n_lower, n_upper, _ = u_hat.shape
b = np.zeros((n_lower, n_upper)) # routing logits, start at 0
for iteration in range(r):
c = np.exp(b) / np.exp(b).sum(axis=1, keepdims=True) # coupling coeffs (softmax)
s = np.einsum('ij,ijd->jd', c, u_hat) # weighted sum of predictions
v = squash(s) # squash to get output
if iteration < r - 1:
# Update routing logits by agreement (dot product)
agreement = np.einsum('ijd,jd->ij', u_hat, v)
b += agreement
return v # (n_upper, dim_upper)
# Example: 1152 PrimaryCapsules → 10 DigitCaps
# Each primary capsule predicts a 16D vector for each digit capsule
# u_hat[i,j] = W_ij @ u_i (learned transformation)
# After 3 routing iterations, each digit capsule's vector length
# gives the probability of that digit being present.Equivariance vs invariance
A key philosophical difference between CNNs and capsules lies in how they handle viewpoint changes:
CNNs aim for invariance: max pooling makes the output the same regardless of where a feature appears. A "7" activates the same neuron whether it's in the top-left or bottom-right. But this throws away where it is — information that matters for understanding spatial arrangements.
Capsules aim for : when the input changes (rotates, scales, shifts), the capsule's output vector changes correspondingly — in a predictable, structured way. The capsule tracks how the entity's pose changed, rather than ignoring the change.
This is why capsules can generalize to novel viewpoints that CNNs cannot. When trained only on translated MNIST digits and tested on affinely transformed digits (affNIST), an under-trained CapsNet achieved 79% accuracy compared to 66% for a CNN with similar parameters — without ever seeing rotations or shearing during training.
Why it mattered
2011
Transforming Autoencoders
Hinton, Krizhevsky, and Wang introduced the capsule concept and proposed using transformation matrices to learn part-whole relationships, but required external supervision for the transformations.
2017
Dynamic Routing Between Capsules
Sabour, Frosst, and Hinton presented a complete capsule system with learned routing. Achieved state-of-the-art on MNIST and demonstrated capsules' segmentation power on overlapping digits.
2018
Matrix Capsules with EM Routing
Hinton, Sabour, and Frosst replaced dynamic routing with EM (Expectation-Maximization) routing, using pose matrices instead of pose vectors for a richer representation.
2019
Stacked Capsule Autoencoders
Kosiorek, Sabour, Teh, and Hinton combined capsules with autoencoders for unsupervised object discovery and part decomposition — removing the need for labeled data.
2020
Vision Transformer (ViT)
Though not a capsule network, ViT echoes capsule ideas: reasoning about spatial relationships between patches, with self-attention serving a role analogous to routing-by-agreement.
CitationSabour, Frosst, Hinton. Dynamic Routing Between Capsules. NeurIPS, 2017.
Terms in this paper
- Capsule Networkالشبكة الكبسولية
- Dynamic Routingالتوجيه الديناميكي
- Squashing Functionدالة السحق
- Coupling Coefficientمعامل الارتباط
- Prediction Vectorمتجه التنبؤ
- Margin Lossخسارة الهامش
- Part-Whole Relationshipعلاقة الأجزاء بالكلّ
- Poseهيئة
- Equivarianceالتساوي التغايُري