Computer Vision2017intermediate12 min read

Dynamic Routing Between Capsules

التوجيه الديناميكي بين الكبسولات

Sabour, S. · Frosst, N. · Hinton, G. E. — NeurIPS

The problem

CNNs use to achieve translation , but this throws away precise spatial information — the exact position, orientation, and scale of detected features. A knows a face has two eyes and a mouth, but doesn't check whether they are in the right spatial arrangement. It can be fooled by a scrambled face. Worse, CNNs need exponentially more data or detectors to handle viewpoint changes beyond translation.

The contribution

Capsule Networks: instead of scalar feature detectors, use groups of neurons (capsules) whose output vectors encode both the presence AND the of an entity. A -by-agreement iteratively sends each lower-level capsule's output to the higher-level capsule whose own output best matches the prediction — like puzzle pieces finding their correct assembly. The system achieves state-of-the-art on MNIST and excels at segmenting highly overlapping digits.

The impact

Capsule Networks introduced a fundamentally different way to think about visual : entities as vectors with pose, not scalars with presence. Though they haven't replaced CNNs at scale, they inspired research on equivariant representations, part-whole hierarchies, and alternatives to that preserve geometric information — ideas that echo in Vision Transformers and geometric deep learning.

A CNN detecting a face is like a checklist inspector: ✓ two eyes, ✓ nose, ✓ mouth — "it's a face!" even if the eyes are below the mouth and the nose is sideways. It checks what's there but not how the pieces fit together.

A is more like an architect reading a blueprint: each part reports not just "I exist" but "here's my exact position, size, and angle." The architect then checks whether all the parts agree on where the whole should be — only then does she sign off on the building.

The problem: max pooling destroys spatial relationships

Convolutional Neural Networks detect features using learned filters — edges, textures, parts — and that works brilliantly (as we saw with LeNet). But there is a critical weakness in how CNNs handle the arrangement of those features:

  • Max pooling discards position. After detecting "eye" and "mouth," max pooling keeps only that they were found somewhere in a local region. The network cannot verify whether the eye is above the mouth — it just needs both activations to be high.

  • Viewpoint changes require exponential resources. Translation invariance is built in, but rotation, scaling, and 3D viewpoint changes are not. To handle them, a CNN must either replicate filters at many orientations (exponential growth) or see every viewpoint in data (exponential data).

  • No part-whole reasoning. A CNN has no mechanism to ask: "do these detected parts agree on where the whole object should be?" It treats features as independent votes, not as a geometric puzzle.

Open in Lab
Drag the face parts around. The CNN always says "face!" as long as parts are present. The capsule network checks spatial agreement.
The demo wakes as you arrive…

The idea: from scalar neurons to vector capsules

The core insight is deceptively simple: replace individual neurons that output a single number with small groups of neurons (capsules) that output a .

In a CNN, a says: "edge detected here, with activation 0.9." That scalar tells you how strongly a feature was detected, but nothing about its pose — its position, orientation, scale, thickness, or deformation.

A capsule says something richer:

  • The length of its output vector (between 0 and 1) represents the probability that the entity exists — like the old scalar activation.

  • The orientation (direction) of the vector represents the instantiation parameters — the pose: where the entity is, how big it is, how it is oriented, how it is deformed.

This is a fundamentally different kind of representation. A single capsule for the concept "eye" doesn't just fire strongly when an eye is present — its vector encodes which eye, in what position, at what angle. Think of it as upgrading from a light switch (on/off) to a compass (direction and magnitude).

Open in Lab
Compare a scalar neuron (left) with a capsule (right). The scalar only has magnitude. The capsule vector encodes pose information in its direction.
The demo wakes as you arrive…

The squashing function: keeping vectors in [0, 1]

Since the length of a capsule's output vector represents a probability, it must stay between 0 and 1. But we also need to preserve the vector's direction (which encodes the pose). The solves both problems: it shrinks short vectors toward zero and long vectors toward length 1, while keeping the orientation unchanged.

Think of it like a tube that compresses water flow: a trickle stays a trickle, a torrent gets capped just below maximum, but the direction of flow never changes.

vj=∥sj∥21+∥sj∥2⋅sj∥sj∥\mathbf{v}_j = \frac{\|\mathbf{s}_j\|^2}{1 + \|\mathbf{s}_j\|^2} \cdot \frac{\mathbf{s}_j}{\|\mathbf{s}_j\|}
The squashing function — ||s_j||² / (1 + ||s_j||²) = scaling factor that maps any length to [0, 1) · s_j / ||s_j|| = unit vector preserving direction. Short inputs → near-zero output, long inputs → near-one output. Direction is always preserved.
Open in Lab
Drag the input vector length and watch how squashing maps it to [0, 1) while preserving direction. Compare with sigmoid on the right.
The demo wakes as you arrive…

Dynamic routing: parts vote for wholes

This is the paper's most important contribution. How does a lower-level capsule (say, "eye") decide which higher-level capsule ("face" vs "car") to send its output to?

The answer is routing by agreement — an iterative process that works like a committee of experts finding consensus:

  1. Predict. Each lower capsule ii makes a for every possible parent capsule jj by multiplying its output ui\mathbf{u}_i by a learned transformation matrix WijW_{ij}: u^j∣i=Wijui\hat{\mathbf{u}}_{j|i} = W_{ij} \mathbf{u}_i. This is the capsule saying: "if I am part of entity jj, then jj should have this pose."

  2. Weight and sum. Each parent sums these predictions, weighted by coupling coefficients cijc_{ij} (which start equal), then squashes the sum to get its output vj\mathbf{v}_j.

  3. Measure agreement. If a prediction u^j∣i\hat{\mathbf{u}}_{j|i} closely matches the parent's output vj\mathbf{v}_j (high ), the cijc_{ij} increases — the part "agrees" with this whole, so send more signal there.

  4. Repeat for a few iterations (typically 3). Parts that agree with a whole reinforce each other; parts that disagree get routed elsewhere.

This is like an eye capsule saying "the face should be here at this angle," a nose capsule saying "the face should be here at this angle," and if they agree — the face capsule activates strongly.

cij=exp⁡(bij)∑kexp⁡(bik)c_{ij} = \frac{\exp(b_{ij})}{\sum_k \exp(b_{ik})}
Coupling coefficients via routing softmax — b_ij = log-prior (starts at 0) + accumulated agreement scores. Softmax ensures all coupling coefficients from capsule i sum to 1 — the part distributes its vote across possible wholes.
Open in Lab
Watch 3 iterations of routing. Lower capsules send predictions upward; coupling coefficients shift toward the capsule with highest agreement.
The demo wakes as you arrive…

The routing algorithm step by step

The routing algorithm runs between every pair of adjacent capsule layers. Here is the complete procedure:

  • Initialize all routing logits bij=0b_{ij} = 0, so every lower capsule sends equal weight to all parents.
  • For rr iterations (typically 3):
  • Compute coupling coefficients ci=softmax(bi)\mathbf{c}_i = \text{softmax}(\mathbf{b}_i)
  • Compute each parent's weighted input: sj=∑iciju^j∣i\mathbf{s}_j = \sum_i c_{ij} \hat{\mathbf{u}}_{j|i}
  • Apply squashing: vj=squash(sj)\mathbf{v}_j = \text{squash}(\mathbf{s}_j)
  • Update routing logits: bij←bij+u^j∣i⋅vjb_{ij} \leftarrow b_{ij} + \hat{\mathbf{u}}_{j|i} \cdot \mathbf{v}_j

The dot product u^j∣i⋅vj\hat{\mathbf{u}}_{j|i} \cdot \mathbf{v}_j is the agreement — it measures how well capsule ii's prediction for parent jj matched what jj actually computed. High agreement increases the coupling; low agreement decreases it.

CapsNet architecture

The paper presents a simple three- architecture for MNIST digit recognition:

Layer 1 — Conv1: A standard with 256 kernels of size 9×9 and activation. This converts pixel intensities into local feature detectors — nothing capsule-specific yet.

Layer 2 — PrimaryCapsules: A convolutional capsule layer with 32 channels of 8D capsules, using 9×9 kernels with 2. This produces 32 × 6 × 6 = 1,152 capsule outputs, each an 8-dimensional vector. These are the lowest-level entities: oriented edges, simple shapes. The squashing function replaces ReLU.

Layer 3 — DigitCaps: 10 capsules (one per digit class), each 16-dimensional. Every DigitCaps capsule receives input from all 1,152 PrimaryCapsules via learned transformation matrices. Dynamic routing runs between these two capsule layers.

The length of each DigitCaps vector gives the probability. With only 8.2M parameters (vs 35.4M for a comparable CNN baseline), CapsNet achieves 0.25% test error on MNIST.

Open in Lab
Click any layer to see dimensions, parameter count, and role in the pipeline.
The demo wakes as you arrive…

Margin loss: one loss per capsule

Since each DigitCaps capsule independently represents whether a digit class is present, the paper uses a separate for each capsule. This is crucial for allowing the network to detect multiple digits simultaneously (as in the overlapping digits task).

The intuition: if digit kk is present (Tk=1T_k = 1), we want its capsule's length ∥vk∥\|\mathbf{v}_k\| to be at least m+=0.9m^+ = 0.9. If digit kk is absent (Tk=0T_k = 0), we want the length to be at most m−=0.1m^- = 0.1. The penalizes violations of these margins.

The down-weighting factor λ=0.5\lambda = 0.5 on the absent-class term prevents the network from shrinking all capsule vectors to zero at the start of training — a critical practical detail.

Lk=Tkmax⁡(0, m+−∥vk∥)2+λ (1−Tk) max⁡(0, ∥vk∥−m−)2L_k = T_k \max(0,\, m^+ - \|\mathbf{v}_k\|)^2 + \lambda\,(1 - T_k)\,\max(0,\, \|\mathbf{v}_k\| - m^-)^2
Margin loss per digit capsule — T_k = 1 if digit k is present · m⁺ = 0.9 (upper target) · m⁻ = 0.1 (lower target) · λ = 0.5 (down-weights absent classes). Total loss = sum over all 10 digit capsules.

Reconstruction: proving the capsule encodes pose

To verify that DigitCaps vectors actually encode meaningful pose information — not just a discriminative signal — the paper adds a reconstruction : three fully-connected layers (512 → 1024 → 784) that reconstruct the input image from the capsule vector of the correct digit.

During training, all vectors except the correct digit's are masked to zero. The decoder then tries to reconstruct the original 28×28 image from just the 16D vector. The reconstruction loss (mean squared error, scaled by 0.0005) is added to the margin loss.

This serves as a powerful regularizer: it forces the capsule to encode the specific visual details of this digit — its thickness, slant, width — not just "it's a 7." When you perturb individual dimensions of the DigitCaps vector and feed them to the decoder, you can see each dimension controlling a specific visual property: stroke thickness, width, rotation, localized deformations.

Open in Lab
Drag the sliders to perturb individual dimensions of the capsule vector and watch how the reconstructed digit changes.
The demo wakes as you arrive…

Overlapping digits: where capsules truly shine

The most compelling result in the paper is on MultiMNIST: two digits from different classes overlaid on the same image, with bounding boxes overlapping by about 80%.

A standard CNN struggles because max pooling merges features from both digits into an indistinguishable soup. CapsNet, using routing-by-agreement, can segment the two digits without pixel-level supervision — each DigitCaps capsule "claims" the parts that agree with its predicted pose, and the other capsule claims the rest.

The reconstruction decoder proves this works: when you reconstruct each digit separately from its capsule vector, you get clean, separated images of each digit — even though they were physically overlapping in the input. CapsNet achieves 5.2% error on MultiMNIST with just 11.36M parameters, matching a sequential model that was tested on a much easier (less overlapping) version of the task.

Open in Lab
Two digits overlap. The capsule network segments them via routing. Toggle to see each digit's reconstruction.
The demo wakes as you arrive…

The same idea in code

Dynamic routing between capsules, completepython

Simplified to show the idea — not the real implementation.

import numpy as np

def squash(s):
    """Non-linear squashing: short → near 0, long → near 1, direction preserved."""
    sq_norm = np.sum(s ** 2, axis=-1, keepdims=True)
    scale = sq_norm / (1 + sq_norm)
    return scale * s / (np.sqrt(sq_norm) + 1e-8)

def routing(u_hat, r=3):
    """
    u_hat: (n_lower, n_upper, dim_upper) — prediction vectors.
    Returns: (n_upper, dim_upper) — output of upper capsules.
    """
    n_lower, n_upper, _ = u_hat.shape
    b = np.zeros((n_lower, n_upper))       # routing logits, start at 0

    for iteration in range(r):
        c = np.exp(b) / np.exp(b).sum(axis=1, keepdims=True)  # coupling coeffs (softmax)
        s = np.einsum('ij,ijd->jd', c, u_hat)                 # weighted sum of predictions
        v = squash(s)                                          # squash to get output

        if iteration < r - 1:
            # Update routing logits by agreement (dot product)
            agreement = np.einsum('ijd,jd->ij', u_hat, v)
            b += agreement

    return v  # (n_upper, dim_upper)

# Example: 1152 PrimaryCapsules → 10 DigitCaps
# Each primary capsule predicts a 16D vector for each digit capsule
# u_hat[i,j] = W_ij @ u_i  (learned transformation)
# After 3 routing iterations, each digit capsule's vector length
# gives the probability of that digit being present.

Equivariance vs invariance

A key philosophical difference between CNNs and capsules lies in how they handle viewpoint changes:

CNNs aim for invariance: max pooling makes the output the same regardless of where a feature appears. A "7" activates the same neuron whether it's in the top-left or bottom-right. But this throws away where it is — information that matters for understanding spatial arrangements.

Capsules aim for : when the input changes (rotates, scales, shifts), the capsule's output vector changes correspondingly — in a predictable, structured way. The capsule tracks how the entity's pose changed, rather than ignoring the change.

This is why capsules can generalize to novel viewpoints that CNNs cannot. When trained only on translated MNIST digits and tested on affinely transformed digits (affNIST), an under-trained CapsNet achieved 79% accuracy compared to 66% for a CNN with similar parameters — without ever seeing rotations or shearing during training.

Open in Lab
Rotate the input digit. The CNN neuron's activation stays flat (invariant). The capsule vector rotates correspondingly (equivariant).
The demo wakes as you arrive…

Why it mattered

  1. 2011

    Transforming Autoencoders

    Hinton, Krizhevsky, and Wang introduced the capsule concept and proposed using transformation matrices to learn part-whole relationships, but required external supervision for the transformations.

  2. 2017

    Dynamic Routing Between Capsules

    Sabour, Frosst, and Hinton presented a complete capsule system with learned routing. Achieved state-of-the-art on MNIST and demonstrated capsules' segmentation power on overlapping digits.

  3. 2018

    Matrix Capsules with EM Routing

    Hinton, Sabour, and Frosst replaced dynamic routing with EM (Expectation-Maximization) routing, using pose matrices instead of pose vectors for a richer representation.

  4. 2019

    Stacked Capsule Autoencoders

    Kosiorek, Sabour, Teh, and Hinton combined capsules with autoencoders for unsupervised object discovery and part decomposition — removing the need for labeled data.

  5. 2020

    Vision Transformer (ViT)

    Though not a capsule network, ViT echoes capsule ideas: reasoning about spatial relationships between patches, with self-attention serving a role analogous to routing-by-agreement.

CitationSabour, Frosst, Hinton. Dynamic Routing Between Capsules. NeurIPS, 2017.

Terms in this paper