Multimodal AI2021intermediate13 min read

Perceiver: General Perception with Iterative Attention

Perceiver: إدراك عام بالانتباه التكراري

Jaegle, A. · Gimeno, F. · Brock, A. · Zisserman, A. · Vinyals, O. · Carreira, J. — ICML

The problem

By 2021, deep learning models for perception were locked into specific modalities. ConvNets exploited 2D grid structure but couldn't handle audio or point clouds. Transformers were flexible but scaled quadratically with input size — applying to 50,000 pixels was computationally infeasible. Every new modality demanded a new architecture, and fusion required bespoke engineering decisions about when and how to combine signals.

The contribution

The Perceiver: a single -based architecture that handles arbitrary modalities without domain-specific assumptions. The key innovation is an asymmetric that projects high-dimensional inputs (M ~ 50,000) into a compact latent array (N ~ 512), reducing complexity from O(M²) to O(MN). Deep Transformer processing then operates only in the low-dimensional at O(LN²). Iterative cross-attention allows the model to revisit the input multiple times. makes the architecture parameter-efficient and can be viewed as a recurrent network unrolled in depth. The Perceiver achieves 78.0% top-1 on ImageNet (matching ResNet-50 and ViT), competitive results on AudioSet across audio, video, and multimodal settings, and strong performance on ModelNet40 point clouds — all with essentially the same architecture.

The impact

The Perceiver demonstrated that a single, modality-agnostic architecture could compete with specialized models across vision, audio, video, and 3D. It directly influenced Perceiver IO (structured outputs), Flamingo (visual language models), and Gato (generalist agents). The cross-attention bottleneck pattern became a standard building block in multimodal architectures, and the paper advanced the broader trend toward general-purpose foundation models that process any input through a unified interface.

Imagine a stadium of 50,000 fans all shouting information at once. A standard Transformer tries to let every fan hear every other fan — that's 2.5 billion conversations. Impossible.

The Perceiver places a small panel of 512 journalists in the press box. Instead of fans talking to each other, the journalists interview the crowd (cross-attention), then debate among themselves (self-attention), then go back to the crowd for follow-up questions. After a few rounds, the panel knows everything important.

Best of all, the same panel works whether the crowd is shouting pixel values, audio samples, 3D coordinates, or all of them at once.

The problem: one architecture per modality

Deep learning in 2021 had a modality problem. For images, the field relied on convolutional neural networks — architectures that bake in the 2D grid structure of pixels, use local receptive fields, and share weights across spatial dimensions. These inductive biases are powerful for images but meaningless for audio waveforms, and actively harmful for irregular data like point clouds.

Want to process stereo images? You need to decide between early or late fusion. Moving to audio? Replace 2D convolutions with 1D. Adding video? Build a 3D architecture. Point clouds from a Lidar sensor? Use PointNet. Each modality demanded its own architecture, its own design choices, its own engineering effort.

Transformers offered an alternative: they make few assumptions about input structure and can model relationships between any pair of elements. But standard self-attention has a fatal flaw — its complexity is quadratic in the number of inputs. An ImageNet image has 50,176 pixels. Self-attention on that many elements requires billions of operations per layer. This is why the Vision Transformer (ViT) first reduces the input to ~200 patches using a 2D convolution — smuggling domain-specific structure back in.

Open in Lab
Each modality demands a different architecture. The Perceiver uses one architecture for all of them.
The demo wakes as you arrive…

The key idea: an asymmetric attention bottleneck

The Perceiver's solution is elegant. Instead of letting all inputs attend to each other (quadratic cost), it introduces a small set of latent units — a learned array of N vectors (typically 512) that acts as an information bottleneck. The inputs don't talk to each other directly. They talk to the latent array, and the latent array talks to itself.

This is done via cross-attention: the (Q) comes from the latent array, while the key (K) and value (V) come from the input byte array. Because Q has N indices and K/V have M indices, the attention matrix is N × M instead of M × M. For ImageNet, this means 512 × 50,176 ≈ 25 million operations instead of 50,176² ≈ 2.5 billion. The cost drops from quadratic to linear in M.

Think of it as a funnel: the wide input is compressed through a narrow neck into a compact . All the expensive processing happens after the funnel, in the small latent space.

Open in Lab
The input byte array (M elements) is projected through the latent bottleneck (N elements) via cross-attention. Adjust N to see how it affects complexity.
The demo wakes as you arrive…

In standard self-attention, Q, K, and V all come from the same array of M elements. The attention matrix QKᵀ is therefore M × M, and the whole operation is O(M²). The Perceiver breaks this symmetry:

CrossAttn(X,Z)=softmax ⁣(QZ KX⊤D)VX  ,QZ∈RN×D,  KX,VX∈RM×D\text{CrossAttn}(X, Z) = \text{softmax}\!\left(\frac{Q_Z \, K_X^\top}{\sqrt{D}}\right) V_X \;,\quad Q_Z \in \mathbb{R}^{N \times D},\; K_X, V_X \in \mathbb{R}^{M \times D}
Cross-attention — Q from latent array Z, K and V from input array X — The query Q is projected from the latent array Z (N indices), while the key K and value V are projected from the input X (M indices). The attention matrix is N × M, so the overall cost is O(MN) instead of O(M²). Since N ≪ M, this is a dramatic reduction. The output has shape N × D — the same size as the latent array.

The full architecture: cross-attend, self-attend, repeat

The Perceiver alternates between two modules:

Cross-attention module: projects the high-dimensional input byte array into the low-dimensional latent array. This is the information bottleneck — the latent array absorbs the most relevant parts of the input.

Latent Transformer: a deep stack of self-attention blocks operating entirely in the latent space. Because the latent array is small (N = 512), these layers are cheap. The architecture uses GPT-2-style Transformer blocks, and the best ImageNet model uses 6 blocks per Transformer and 8 cross-attention iterations — totaling 48 latent Transformer blocks.

The total complexity is O(MN + LN²), where L is the number of latent Transformer layers. The first term is the cross-attention cost (linear in M), and the second is the latent processing cost (independent of M). This decomposition is what makes the architecture scalable: you can go deep without the input size being a factor.

O(MN⏟cross-attn+LN2⏟latent transformer)vs.O(LM2)  (standard Transformer)\mathcal{O}\bigl(\underbrace{MN}_{\text{cross-attn}} + \underbrace{LN^2}_{\text{latent transformer}}\bigr) \quad \text{vs.} \quad \mathcal{O}(LM^2) \;\text{(standard Transformer)}
Perceiver complexity vs standard Transformer complexity — With M = 50,176 pixels, N = 512 latents, and L = 48 layers: the standard Transformer needs O(48 × 50,176²) ≈ 1.2 × 10¹¹ operations. The Perceiver needs O(50,176 × 512 + 48 × 512²) ≈ 3.8 × 10⁷ — more than 3,000× cheaper.
Open in Lab
The Perceiver architecture: cross-attention compresses the input, the latent Transformer processes it, and the cycle repeats iteratively.
The demo wakes as you arrive…

Iterative refinement and weight sharing

The bottleneck is tight — 512 latent vectors must summarize 50,000+ inputs. A single pass might miss details. The Perceiver addresses this with iterative cross-attention: after the latent Transformer processes the initial summary, the model goes back to the full input and cross-attends again, refining its understanding.

The best ImageNet model uses 8 iterations. Each iteration gives the latent array another chance to extract information it missed before. Think of a journalist reviewing interview notes, returning for follow-up questions, and building a richer story with each pass.

A crucial design choice is weight sharing: all cross-attention modules after the first share weights, and all latent Transformer blocks share weights. This reduces the parameter count by approximately 10×, turning the architecture into something functionally equivalent to a — but unrolled in depth, not in time. Weight sharing also acts as a regularizer, reducing and actually improving validation performance on ImageNet.

Open in Lab
Watch how the latent array refines its representation with each cross-attention iteration. More iterations capture finer details.
The demo wakes as you arrive…

Position encodings: telling the model where things are

Attention is permutation-invariant — it returns the same output regardless of input order. For language this is solved with sequence position encodings. But for images, audio, video, and point clouds, the Perceiver needs something richer.

The paper uses Fourier feature position encodings. For each input element, the position along each dimension d is encoded as bands of sinusoids:

pos(xd)=[sin⁡(fkπxd),  cos⁡(fkπxd)]k=1K,fk∈{1,…,μ2}\text{pos}(x_d) = \bigl[\sin(f_k \pi x_d),\; \cos(f_k \pi x_d)\bigr]_{k=1}^{K}, \quad f_k \in \{1, \ldots, \tfrac{\mu}{2}\}
Fourier feature position encoding — Each spatial dimension d is encoded using K frequency bands linearly spaced from 1 to μ/2 (the Nyquist frequency for resolution μ). This produces a position vector of size d(2K + 1) for each input element. For images, d = 2 (x, y); for video, d = 3 (x, y, t). The encoding is concatenated with input features, not added — unlike standard Transformer position encodings.

The beauty of this design is its generality. For images, you use 2D positions. For audio, 1D. For video, 3D. For point clouds, 3D spatial coordinates. You don't need to redesign the architecture — just adapt the position encoding. And the encoding is optional: a learned position encoding with no knowledge of spatial structure still achieves 70.9% on ImageNet (Table 2 in the paper), proving the architecture isn't secretly depending on spatial priors.

For multimodal inputs (audio + video), each modality uses its own position encoding dimensionality, plus a learned modality-specific to distinguish which signal is which.

Permutation test: architecture without spatial assumptions

To prove that the Perceiver doesn't secretly rely on spatial structure, the authors ran a striking experiment: they randomly permuted the pixel order of all ImageNet images (using the same permutation for all images, applied after position features are computed).

Models with 2D convolutions collapsed: ResNet-50 dropped from 73.5% to 39.4%, and ViT dropped from 76.7% to 61.7%. Their architectures assume spatial locality — when pixels are scrambled, local filters become meaningless.

The Perceiver was completely unaffected: 78.0% before and after permutation. Its architecture treats every input as a bag of features — position information comes from features, not from spatial assumptions. This makes the Perceiver a natural fit for modalities with no grid structure at all: point clouds, sensor arrays, molecular data.

Open in Lab
Permute the pixels — ConvNets collapse, the Perceiver doesn't flinch. Click to scramble and compare.
The demo wakes as you arrive…

Results: one model, many modalities

The Perceiver was evaluated across four modalities with essentially the same architecture:

ImageNet (images): 78.0% top-1 accuracy, matching ResNet-50 (77.6%) and ViT (77.9%) — without any 2D convolutions. The model directly attends to all 50,176 pixels.

AudioSet (audio): 38.4 mAP using raw audio waveforms (61,440 samples), matching -14 and outperforming most specialized audio models. Mel spectrograms gave similar performance (38.4 mAP).

AudioSet (video): 25.8 mAP from video alone, competitive with state-of-the-art.

AudioSet (audio + video): 43.4 mAP with raw audio, demonstrating effective multimodal fusion. A simple trick — randomly dropping the entire video stream during with 30% probability — boosted performance by 3% by preventing the model from over-relying on the higher-bandwidth video signal.

ModelNet40 (point clouds): 85.7% accuracy, outperforming both ViT and ResNet baselines adapted for point clouds, though below the specialized PointNet++ (91.9%) which uses hand-crafted geometric features.

Open in Lab
Perceiver performance across modalities compared to specialized baselines.
The demo wakes as you arrive…

Attention maps: how the Perceiver sees

The paper visualizes cross-attention maps at different iterations, revealing fascinating behavior. The first cross-attention module — which has unique weights — produces attention maps that clearly show the structure of the input image. Some latent indices specialize in detecting the main subject (a dog, for example), visible directly in the attention pattern.

Later iterations, which share weights, produce high-frequency plaid-like patterns that sweep across the image at different spatial frequencies. These tartan-like patterns appear to reflect the Fourier structure of the position encodings. While later modules produce similar patterns overall, the specific pixel sets they attend to differ — the model isn't redundantly re-reading the same information, but progressively refining its understanding.

Why this matters: toward general perception

The Perceiver represents a philosophical shift in deep learning architecture design. Instead of encoding domain knowledge into the architecture itself — 2D convolutions for images, 1D for audio, specialized models for 3D — it moves as much as possible into the data layer: position encodings, modality tags, and raw features. The architecture itself remains generic.

This aligns with what Richard Sutton called "The Bitter Lesson" (2019): methods that leverage general computation scale better than methods that leverage human knowledge. The Perceiver doesn't outperform the best specialized models on every benchmark, but it matches them with one architecture. As data and compute grow, the generalist approach tends to win.

The cross-attention bottleneck pattern has since appeared in Perceiver IO (which adds structured outputs using decoder cross-attention), Flamingo (which uses it to fuse vision and language for visual question answering), and Gato (DeepMind's generalist agent that plays games, chats, and controls robots — all with one set of weights).

Perceiver forward pass (PyTorch-style pseudocode)python

Simplified to show the idea — not the real implementation.

# x: input byte array [B, M, C]  (e.g. 50176 pixels × channels)
# z: learned latent array [1, N, D] (e.g. 512 × 1024), broadcast over batch

z = self.latent_array.expand(B, -1, -1)         # [B, N, D]

for i in range(num_iterations):                  # e.g. 8 iterations
    # Cross-attention: latent queries the input
    z = cross_attention(q=z, kv=x) + z           # [B, N, D]

    # Latent Transformer: self-attention in small space
    for block in self.transformer_blocks:         # e.g. 6 blocks
        z = self_attention(z) + z                 # [B, N, D]  O(N²)
        z = feed_forward(z) + z

logits = self.classifier(z.mean(dim=1))           # global average pool → class

Timeline: from Transformers to general perception

  1. 2017

    Transformer (Vaswani et al.)

    Introduced self-attention for sequence modeling. Powerful but quadratic in sequence length, and designed primarily for language.

  2. 2019

    Set Transformer (Lee et al.)

    Used cross-attention to project large input sets to smaller arrays, reducing attention complexity. A direct precursor to the Perceiver's bottleneck design.

  3. 2021

    ViT — Vision Transformer

    Applied Transformers to images by splitting them into 16×16 patches. Impressive results but still requires a 2D convolutional preprocessing step.

  4. 2021

    Perceiver (this paper)

    Removed modality-specific assumptions entirely. One architecture handles images, audio, video, point clouds, and multimodal combinations.

  5. 2022

    Perceiver IO

    Extended the Perceiver with decoder cross-attention to produce structured outputs (not just classification), enabling tasks like optical flow and language modeling.

  6. 2022

    Flamingo & Gato

    Flamingo used cross-attention bottlenecks to fuse vision and language for few-shot visual QA. Gato used a single Transformer as a generalist agent across 600+ tasks. Both inherit the Perceiver's modality-agnostic philosophy.

The Perceiver's deepest contribution isn't a performance number — it's a design principle. Domain-specific priors are powerful when you know the domain, but they lock you in. By pushing domain knowledge from architecture into data (position encodings, modality tags), the Perceiver showed that a single, general architecture can be competitive across modalities. As AI moves toward foundation models that see, hear, and act in the world, this principle becomes foundational.

CitationJaegle, Gimeno, Brock, Zisserman, Vinyals, Carreira. Perceiver: General Perception with Iterative Attention. ICML, 2021.

Terms in this paper