Time Series2021intermediate10 min read

Informer: Beyond Efficient Transformer for Long Sequence Time Series Forecasting

Informer: ما وراء المحوِّل الفعّال للتنبؤ بالسلاسل الزمنية الطويلة

Zhou, H. · Zhang, S. · Peng, J. · Zhang, S. · Li, J. · Xiong, H. · Zhang, W. — AAAI

The problem

Long sequence Forecasting (LSTF) — predicting hundreds of future steps from long historical windows — is critical for energy planning, weather, and finance. Transformers showed promise, but three problems blocked their use: (1) the canonical has O(L²) time and memory cost, making long inputs prohibitive; (2) stacking layers compounds this cost; (3) the standard predicts one step at a time, which is both slow and accumulates errors over long horizons.

The contribution

Informer introduces three innovations: (1) ProbSparse self-, which uses KL-divergence to identify the few "active" queries and reduces complexity from O(L²) to O(L log L); (2) a self-attention distilling mechanism that halves the sequence length between encoder layers using Conv1d + MaxPool, shrinking total memory to O((2−ε)L log L); (3) a generative-style decoder that predicts the entire output sequence in one instead of step-by-step. Informer won the AAAI 2021 Outstanding Paper Award.

The impact

Informer proved that Transformers could handle long-horizon Time Series Forecasting efficiently. It directly inspired Autoformer (auto-correlation for seasonal decomposition), FEDformer (frequency-domain attention), PatchTST (patching time series like vision tokens), and DLinear (a linear baseline that challenged complexity). The ProbSparse attention idea influenced efficient attention research beyond time series. Informer remains a milestone in the Transformer-for-time-series lineage.

Imagine a teacher grading 1,000 exam papers. The standard approach — reading every word of every paper — takes forever. A smarter teacher first skims each paper to spot which ones need careful reading (the interesting answers) and which are nearly identical boilerplate. She reads only the interesting ones in full, gives the rest a default score, and finishes in a fraction of the time.

Then she summarizes her notes after each round, keeping only the highlights, so her stack of papers shrinks by half each time.

Finally, instead of writing feedback one sentence at a time — where early mistakes compound — she writes the entire feedback letter in one go, seeing the full picture.

Informer does exactly this with time-series data: skip the boring attention, compress the sequence layer by layer, and predict everything at once.

The challenge: why standard Transformers fail at long time series

Consider predicting hourly electricity consumption for the next 20 days — that is 480 output steps. To capture repeating patterns (daily cycles, weekly trends), the model needs a long input window too, perhaps 720 steps or more. In a standard Transformer, self-attention compares every input position to every other, producing an L×LL \times L attention matrix. With L=720L = 720, that is over half a million entries — per head, per layer. Stack 6 layers with 8 heads, and the cost becomes enormous.

But the problem runs deeper than raw compute. The authors observed that in time-series self-attention, most of the attention scores are nearly uniform — they carry almost no useful information. Only a small minority of -key pairs produce sharp, concentrated attention patterns. The attention distribution follows a long-tail pattern: a few "active" queries dominate while the rest are "lazy," spreading attention uniformly across all keys like background noise.

This observation is the key insight behind Informer: if most queries are lazy, why compute them at all?

Open in Lab
Most queries produce near-uniform attention (the "lazy" tail). Only a few produce sharp, informative patterns (the "active" head). Informer focuses on the active ones.
The demo wakes as you arrive…

Innovation 1: ProbSparse self-attention — attend only to what matters

The core question is: how do you identify the "active" queries before computing the full attention matrix? Informer answers this with a sparsity measurement based on .

The intuition: if a query qiq_i attends uniformly to all keys, its attention distribution is uniform — it carries no information. If qiq_i concentrates its attention on a few keys, its distribution is far from uniform — it is informative. We can measure this gap using the Kullback-Leibler divergence between the query's actual attention distribution and the uniform distribution. Queries with high divergence are the important ones.

After dropping a constant term, the authors derive a practical sparsity score:

Mˉ(qi,K)=max⁡j{qikj⊤d}−1LK∑j=1LKqikj⊤d\bar{M}(q_i, K) = \max_j \left\{ \frac{q_i k_j^\top}{\sqrt{d}} \right\} - \frac{1}{L_K} \sum_{j=1}^{L_K} \frac{q_i k_j^\top}{\sqrt{d}}
Query sparsity measurement — max score minus mean score — For each query qiq_i, compute its maximum attention score with any key minus its average score across all keys. A large gap means qiq_i has a dominant key it strongly focuses on — it is "active." A gap near zero means qiq_i spreads attention uniformly — it is "lazy." Only the top-uu queries by this score are kept, where u=c⋅ln⁡LQu = c \cdot \ln L_Q and cc is a sampling factor (set to 5 in practice).

Once we select the top-uu active queries into a reduced matrix Qˉ\bar{Q}, the ProbSparse attention is simply standard attention applied only to those queries:

A(Q,K,V)=Softmax ⁣(QˉK⊤d)V\mathcal{A}(Q, K, V) = \text{Softmax}\!\left(\frac{\bar{Q} K^\top}{\sqrt{d}}\right) V
ProbSparse self-attention — only active queries attend — Qˉ\bar{Q} contains only c⋅ln⁡LQc \cdot \ln L_Q rows instead of LQL_Q. The "lazy" queries receive a default value: the mean of all values VV. This reduces the attention computation from O(L2)O(L^2) to O(Llog⁡L)O(L \log L) while preserving the important interactions.
Open in Lab
Toggle between canonical and ProbSparse attention. Watch how ProbSparse selects only the top queries (red) and assigns a default value to the rest (gray).
The demo wakes as you arrive…

Innovation 2: self-attention distilling — compress as you go

Even with ProbSparse attention reducing cost per layer, stacking multiple encoder layers still builds up memory. The dominant features from one layer are a small fraction of the input — so why pass the full sequence to the next layer?

Informer's self-attention distilling operation answers this by halving the sequence length between layers. After each attention layer, the output passes through a 1D convolution (to blend adjacent time steps) followed by max-pooling with stride 2. This keeps the most prominent features and discards the rest — like distilling a long document into bullet points, then distilling the bullet points further.

Picture a pyramid: the bottom layer (closest to the input) sees the full sequence of length LL, the next sees L/2L/2, then L/4L/4, and so on. Each layer operates on a smaller sequence, so the total memory across all JJ layers is a geometric series:

Memory=O ⁣(∑j=0J−1L2jlog⁡L2j)=O ⁣((2−ϵ)Llog⁡L)\text{Memory} = O\!\left(\sum_{j=0}^{J-1} \frac{L}{2^j} \log \frac{L}{2^j}\right) = O\!\left((2 - \epsilon) L \log L\right)
Distilling memory — geometric shrinkage across layers — Without distilling, JJ stacked layers cost O(J⋅L2)O(J \cdot L^2) for standard attention or O(J⋅Llog⁡L)O(J \cdot L \log L) for ProbSparse alone. With distilling, the geometric sum converges to roughly 2Llog⁡L2 L \log L — nearly the cost of a single layer regardless of depth.
Open in Lab
Watch the encoder pyramid: each layer halves the sequence. The total memory (shown on the right) converges instead of growing with depth.
The demo wakes as you arrive…

Innovation 3: generative decoder — predict everything at once

Standard Transformer decoders use dynamic decoding: predict step 1, feed it back as input, predict step 2, feed it back, and so on. For 480 output steps, that means 480 sequential forward passes. Besides being slow, each prediction error feeds into the next step, causing error accumulation over long horizons.

Informer replaces this with a generative-style decoder that produces all output steps in a single forward pass. The trick: instead of starting from scratch, the decoder receives a "start token" consisting of the last segment of the known input sequence (providing local context), followed by placeholder zeros where the predictions should go. The decoder fills in all the zeros simultaneously using to the encoder's feature map.

Think of it like a fill-in-the-blank exam: the student sees the beginning of the story (the start token) and a blank section to complete (the zeros), and writes the entire answer at once rather than word by word.

Open in Lab
Compare step-by-step decoding (errors compound) vs generative decoding (all at once). Click "Run" to see how errors accumulate differently.
The demo wakes as you arrive…

Putting it together: the full Informer architecture

The complete architecture combines all three innovations into an framework designed for long sequences:

Input : Each time step is embedded by combining a scalar projection of its value, a learned , and temporal features (hour-of-day, day-of-week, etc.) that inject calendar knowledge.

Encoder: Multiple ProbSparse self-attention layers, each followed by a Conv1d + MaxPool distilling operation that halves the sequence. Parallel encoder stacks at different resolutions provide robustness. The final feature map is a compressed, attention-distilled representation of the input.

Decoder: Receives the start token (known tail of input) concatenated with zero placeholders. Uses masked ProbSparse self-attention (so predictions cannot look ahead at future zeros) plus cross-attention to the encoder feature map. Produces all predictions in one forward pass through a final linear projection.

Open in Lab
Explore the full Informer architecture. Click on each component to see how it processes the time-series data from input to output.
The demo wakes as you arrive…

In code: ProbSparse attention step by step

ProbSparse Self-Attention — core logicpython

Simplified to show the idea — not the real implementation.

import torch
import math

def prob_sparse_attention(Q, K, V, sample_factor=5):
    """ProbSparse self-attention: O(L log L) instead of O(L²)."""
    L_Q, d = Q.shape
    L_K = K.shape[0]
    
    # Step 1: Sample u = c * ln(L_Q) keys for sparsity measurement
    u = max(1, int(sample_factor * math.log(L_Q)))
    
    # Step 2: Compute sparsity score M for each query
    #   M(q_i) = max_j(q_i · k_j / √d) - mean_j(q_i · k_j / √d)
    scores = Q @ K.T / math.sqrt(d)         # [L_Q, L_K]
    M = scores.max(dim=-1).values - scores.mean(dim=-1)  # [L_Q]
    
    # Step 3: Select top-u queries with highest sparsity
    top_idx = M.topk(u).indices               # [u]
    Q_bar = Q[top_idx]                         # [u, d]
    
    # Step 4: Compute attention only for selected queries
    attn_scores = Q_bar @ K.T / math.sqrt(d)  # [u, L_K]
    attn_weights = torch.softmax(attn_scores, dim=-1)
    
    # Step 5: Active queries get attention output;
    #         lazy queries get mean(V) as default
    out = V.mean(dim=0).unsqueeze(0).expand(L_Q, -1).clone()
    out[top_idx] = attn_weights @ V            # [u, d]
    
    return out

Experiments: faster and more accurate on long horizons

Informer was evaluated on four real-world datasets: ETTh1/ETTh2 (hourly electricity transformer temperature), ETTm1 (15-minute electricity), and ECL (electricity consumption of 321 clients). All experiments used prediction horizons from 24 to 720 steps — far longer than previous work.

The results showed three clear advantages. First, on long horizons (168–720 steps), Informer consistently outperformed all baselines — including LSTMa, DeepAR, and the vanilla Transformer — in both MSE and MAE. The improvement was especially large on the longest horizons, where standard Transformers degraded rapidly. Second, Informer's memory usage scaled as O(Llog⁡L)O(L \log L) compared to O(L2)O(L^2), allowing it to handle sequences 5–10× longer than what vanilla Transformers could fit in memory. Third, the generative decoder made 3–4× faster by avoiding step-by-step prediction.

Open in Lab
Compare MSE across prediction horizons. Informer maintains accuracy where other models break down.
The demo wakes as you arrive…

Legacy: from Informer to the modern time-series Transformer zoo

  1. 2017

    Transformer (Vaswani et al.)

    Introduced the self-attention mechanism with O(L²) complexity. Revolutionized NLP but was not designed for long time-series inputs.

  2. 2019

    LogTrans & Sparse Transformer

    Early attempts at efficient attention using fixed sparse patterns or log-sparse masks. Reduced complexity but used hand-crafted patterns, not data-driven selection.

  3. 2021

    Informer (this paper)

    Data-driven sparse attention via ProbSparse, self-attention distilling for memory efficiency, and generative decoder for fast inference. AAAI Best Paper.

  4. 2021

    Autoformer (Wu et al.)

    Replaced dot-product attention with auto-correlation for capturing seasonal periodicity. Added explicit trend-seasonal decomposition inside the network.

  5. 2022

    DLinear (Zeng et al.)

    A provocative baseline: a simple linear layer outperformed many Transformer models on standard benchmarks, questioning whether attention is needed at all for Time Series Forecasting.

  6. 2023

    PatchTST (Nie et al.)

    Segmented time series into patches (like vision tokens), dramatically reducing sequence length and enabling channel-independent modeling. Showed that treating time series like images unlocks Transformer power.

Informer opened the door, but the field quickly moved beyond it. Autoformer showed that domain-specific inductive biases (seasonal decomposition) can outperform generic . DLinear challenged whether Transformers add value at all for simple forecasting. PatchTST showed that the input representation matters as much as the attention mechanism. Yet all of these models exist because Informer first proved that Transformers could work for long-horizon forecasting — the question shifted from "can they?" to "how best?"

CitationZhou, Zhang, Peng, Zhang, Li, Xiong, Zhang. Informer: Beyond Efficient Transformer for Long Sequence Time Series Forecasting. AAAI, 2021.

Terms in this paper