Time Series2021intermediate10 min read
Informer: Beyond Efficient Transformer for Long Sequence Time Series Forecasting
Informer: ما وراء المحوِّل الفعّال للتنبؤ بالسلاسل الزمنية الطويلة
Zhou, H. · Zhang, S. · Peng, J. · Zhang, S. · Li, J. · Xiong, H. · Zhang, W. — AAAI
The problem
Long sequence Forecasting (LSTF) — predicting hundreds of future steps from long historical windows — is critical for energy planning, weather, and finance. Transformers showed promise, but three problems blocked their use: (1) the canonical has O(L²) time and memory cost, making long inputs prohibitive; (2) stacking layers compounds this cost; (3) the standard predicts one step at a time, which is both slow and accumulates errors over long horizons.
The contribution
Informer introduces three innovations: (1) ProbSparse self-, which uses KL-divergence to identify the few "active" queries and reduces complexity from O(L²) to O(L log L); (2) a self-attention distilling mechanism that halves the sequence length between encoder layers using Conv1d + MaxPool, shrinking total memory to O((2−ε)L log L); (3) a generative-style decoder that predicts the entire output sequence in one instead of step-by-step. Informer won the AAAI 2021 Outstanding Paper Award.
The impact
Informer proved that Transformers could handle long-horizon Time Series Forecasting efficiently. It directly inspired Autoformer (auto-correlation for seasonal decomposition), FEDformer (frequency-domain attention), PatchTST (patching time series like vision tokens), and DLinear (a linear baseline that challenged complexity). The ProbSparse attention idea influenced efficient attention research beyond time series. Informer remains a milestone in the Transformer-for-time-series lineage.
Imagine a teacher grading 1,000 exam papers. The standard approach — reading every word of every paper — takes forever. A smarter teacher first skims each paper to spot which ones need careful reading (the interesting answers) and which are nearly identical boilerplate. She reads only the interesting ones in full, gives the rest a default score, and finishes in a fraction of the time.
Then she summarizes her notes after each round, keeping only the highlights, so her stack of papers shrinks by half each time.
Finally, instead of writing feedback one sentence at a time — where early mistakes compound — she writes the entire feedback letter in one go, seeing the full picture.
Informer does exactly this with time-series data: skip the boring attention, compress the sequence layer by layer, and predict everything at once.
The challenge: why standard Transformers fail at long time series
Consider predicting hourly electricity consumption for the next 20 days — that is 480 output steps. To capture repeating patterns (daily cycles, weekly trends), the model needs a long input window too, perhaps 720 steps or more. In a standard Transformer, self-attention compares every input position to every other, producing an attention matrix. With , that is over half a million entries — per head, per layer. Stack 6 layers with 8 heads, and the cost becomes enormous.
But the problem runs deeper than raw compute. The authors observed that in time-series self-attention, most of the attention scores are nearly uniform — they carry almost no useful information. Only a small minority of -key pairs produce sharp, concentrated attention patterns. The attention distribution follows a long-tail pattern: a few "active" queries dominate while the rest are "lazy," spreading attention uniformly across all keys like background noise.
This observation is the key insight behind Informer: if most queries are lazy, why compute them at all?
Innovation 1: ProbSparse self-attention — attend only to what matters
The core question is: how do you identify the "active" queries before computing the full attention matrix? Informer answers this with a sparsity measurement based on .
The intuition: if a query attends uniformly to all keys, its attention distribution is uniform — it carries no information. If concentrates its attention on a few keys, its distribution is far from uniform — it is informative. We can measure this gap using the Kullback-Leibler divergence between the query's actual attention distribution and the uniform distribution. Queries with high divergence are the important ones.
After dropping a constant term, the authors derive a practical sparsity score:
Once we select the top- active queries into a reduced matrix , the ProbSparse attention is simply standard attention applied only to those queries:
Innovation 2: self-attention distilling — compress as you go
Even with ProbSparse attention reducing cost per layer, stacking multiple encoder layers still builds up memory. The dominant features from one layer are a small fraction of the input — so why pass the full sequence to the next layer?
Informer's self-attention distilling operation answers this by halving the sequence length between layers. After each attention layer, the output passes through a 1D convolution (to blend adjacent time steps) followed by max-pooling with stride 2. This keeps the most prominent features and discards the rest — like distilling a long document into bullet points, then distilling the bullet points further.
Picture a pyramid: the bottom layer (closest to the input) sees the full sequence of length , the next sees , then , and so on. Each layer operates on a smaller sequence, so the total memory across all layers is a geometric series:
Innovation 3: generative decoder — predict everything at once
Standard Transformer decoders use dynamic decoding: predict step 1, feed it back as input, predict step 2, feed it back, and so on. For 480 output steps, that means 480 sequential forward passes. Besides being slow, each prediction error feeds into the next step, causing error accumulation over long horizons.
Informer replaces this with a generative-style decoder that produces all output steps in a single forward pass. The trick: instead of starting from scratch, the decoder receives a "start token" consisting of the last segment of the known input sequence (providing local context), followed by placeholder zeros where the predictions should go. The decoder fills in all the zeros simultaneously using to the encoder's feature map.
Think of it like a fill-in-the-blank exam: the student sees the beginning of the story (the start token) and a blank section to complete (the zeros), and writes the entire answer at once rather than word by word.
Putting it together: the full Informer architecture
The complete architecture combines all three innovations into an framework designed for long sequences:
Input : Each time step is embedded by combining a scalar projection of its value, a learned , and temporal features (hour-of-day, day-of-week, etc.) that inject calendar knowledge.
Encoder: Multiple ProbSparse self-attention layers, each followed by a Conv1d + MaxPool distilling operation that halves the sequence. Parallel encoder stacks at different resolutions provide robustness. The final feature map is a compressed, attention-distilled representation of the input.
Decoder: Receives the start token (known tail of input) concatenated with zero placeholders. Uses masked ProbSparse self-attention (so predictions cannot look ahead at future zeros) plus cross-attention to the encoder feature map. Produces all predictions in one forward pass through a final linear projection.
In code: ProbSparse attention step by step
Simplified to show the idea — not the real implementation.
import torch
import math
def prob_sparse_attention(Q, K, V, sample_factor=5):
"""ProbSparse self-attention: O(L log L) instead of O(L²)."""
L_Q, d = Q.shape
L_K = K.shape[0]
# Step 1: Sample u = c * ln(L_Q) keys for sparsity measurement
u = max(1, int(sample_factor * math.log(L_Q)))
# Step 2: Compute sparsity score M for each query
# M(q_i) = max_j(q_i · k_j / √d) - mean_j(q_i · k_j / √d)
scores = Q @ K.T / math.sqrt(d) # [L_Q, L_K]
M = scores.max(dim=-1).values - scores.mean(dim=-1) # [L_Q]
# Step 3: Select top-u queries with highest sparsity
top_idx = M.topk(u).indices # [u]
Q_bar = Q[top_idx] # [u, d]
# Step 4: Compute attention only for selected queries
attn_scores = Q_bar @ K.T / math.sqrt(d) # [u, L_K]
attn_weights = torch.softmax(attn_scores, dim=-1)
# Step 5: Active queries get attention output;
# lazy queries get mean(V) as default
out = V.mean(dim=0).unsqueeze(0).expand(L_Q, -1).clone()
out[top_idx] = attn_weights @ V # [u, d]
return outExperiments: faster and more accurate on long horizons
Informer was evaluated on four real-world datasets: ETTh1/ETTh2 (hourly electricity transformer temperature), ETTm1 (15-minute electricity), and ECL (electricity consumption of 321 clients). All experiments used prediction horizons from 24 to 720 steps — far longer than previous work.
The results showed three clear advantages. First, on long horizons (168–720 steps), Informer consistently outperformed all baselines — including LSTMa, DeepAR, and the vanilla Transformer — in both MSE and MAE. The improvement was especially large on the longest horizons, where standard Transformers degraded rapidly. Second, Informer's memory usage scaled as compared to , allowing it to handle sequences 5–10× longer than what vanilla Transformers could fit in memory. Third, the generative decoder made 3–4× faster by avoiding step-by-step prediction.
Legacy: from Informer to the modern time-series Transformer zoo
2017
Transformer (Vaswani et al.)
Introduced the self-attention mechanism with O(L²) complexity. Revolutionized NLP but was not designed for long time-series inputs.
2019
LogTrans & Sparse Transformer
Early attempts at efficient attention using fixed sparse patterns or log-sparse masks. Reduced complexity but used hand-crafted patterns, not data-driven selection.
2021
Informer (this paper)
Data-driven sparse attention via ProbSparse, self-attention distilling for memory efficiency, and generative decoder for fast inference. AAAI Best Paper.
2021
Autoformer (Wu et al.)
Replaced dot-product attention with auto-correlation for capturing seasonal periodicity. Added explicit trend-seasonal decomposition inside the network.
2022
DLinear (Zeng et al.)
A provocative baseline: a simple linear layer outperformed many Transformer models on standard benchmarks, questioning whether attention is needed at all for Time Series Forecasting.
2023
PatchTST (Nie et al.)
Segmented time series into patches (like vision tokens), dramatically reducing sequence length and enabling channel-independent modeling. Showed that treating time series like images unlocks Transformer power.
Informer opened the door, but the field quickly moved beyond it. Autoformer showed that domain-specific inductive biases (seasonal decomposition) can outperform generic . DLinear challenged whether Transformers add value at all for simple forecasting. PatchTST showed that the input representation matters as much as the attention mechanism. Yet all of these models exist because Informer first proved that Transformers could work for long-horizon forecasting — the question shifted from "can they?" to "how best?"
CitationZhou, Zhang, Peng, Zhang, Li, Xiong, Zhang. Informer: Beyond Efficient Transformer for Long Sequence Time Series Forecasting. AAAI, 2021.
Terms in this paper
- Self-Attentionالانتباه الذاتي
- Sparse Attentionالانتباه المتناثر
- KL Divergenceتباعد KL
- Encoder-Decoderمرمِّز-فاكّ ترميز
- Distillationالتقطير
- Positional Encodingالترميز الموضعي
- Time Seriesسلسلة زمنية
- Forecastالتنبؤ
- Multi-Head Attentionالانتباه المتعدد المسارات
- Softmaxسوفت ماكس
- Queryمصفوفة الطلب (الاستعلام)
- Keyمصفوفة المفتاح (الاستدلال)
- Feed Forward Network (FFN)شبكة التغذية الأمامية
- Residual Connectionالوصلة التجاوزية
- Layer Normalizationالتسوية الطبقية