Time Series2021intermediate10 min read
Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting
Autoformer: بنية تفكيك تقدّمية مع ارتباط ذاتي للتنبؤ بعيد المدى بالسلاسل الزمنية
Wu, H. · Xu, J. · Wang, J. · Long, M. — NeurIPS
The problem
Long-term is critical for applications like energy planning, weather warning, and disease tracking, but existing -based models face two fundamental problems. First, the intricate temporal patterns of long-term series — where trends, seasons, and intertwine — make it hard to discover reliable dependencies. Second, standard has quadratic complexity, and the sparse approximations (ProbSparse, LogSparse, LSH) that fix this are all point-wise, creating an information utilization bottleneck: they connect individual time steps rather than coherent sub-sequences.
The contribution
Autoformer introduces two core innovations. First, a progressive architecture that embeds moving-average-based -seasonal separation as an inner block inside every and layer — not as preprocessing. This lets the model iteratively refine the trend and seasonal components throughout the forecasting process. Second, an Auto-Correlation mechanism that replaces self-: it uses FFT-based to discover periodic dependencies and aggregates similar sub-series via time-delay rolling, achieving O(L log L) complexity while connecting entire sub-sequences rather than scattered points. On six benchmarks covering energy, traffic, economics, weather, and disease, Autoformer achieves a 38% relative MSE reduction over the previous state of the art.
The impact
Autoformer demonstrated that the Transformer's self-attention is not the only — or best — way to model temporal dependencies in time series. By proving that series-level periodicity connections outperform point-wise attention, it opened the door to a new family of forecasting architectures. It directly influenced DLinear (which questioned whether even simple linear models suffice) and the broader debate on what inductive biases time series models need. The inner decomposition idea has been adopted by subsequent models, and the paper remains a key reference in the time series forecasting literature.
Imagine trying to predict the tide a month from now. A weather forecaster wouldn't stare at the raw water-level graph — it's a tangled mess of rising ocean, daily tides, and random waves. Instead, she'd peel the signal apart: first extract the slow seasonal rise, then isolate the 12-hour tidal cycle, then handle the noise separately.
Standard Transformers try to read the raw mess directly, connecting individual water-level readings with attention. Autoformer works like the forecaster: it continuously peels apart the trend from the cycles inside every layer, and instead of linking individual readings, it matches entire tidal cycles — recognizing that today's tide looks like yesterday's tide shifted by 12 hours.
The problem: why Transformers struggle with long-term series
Informer and related Transformer-based forecasters improved efficiency with sparse attention variants, but they all share a fundamental limitation: point-wise connections. Self-attention computes a score between individual time steps — "how relevant is step 42 to step 100?" — then aggregates values from selected points. This works for language, where a single word can carry critical meaning, but time series are different.
In a time series, meaning lives in sub-sequences, not individual points. A single temperature reading at 3 PM means little on its own; it's the pattern across the full day — the diurnal cycle — that carries predictive power. Point-wise attention cannot naturally capture this: it sees trees, not the forest.
The second problem is entangled patterns. Real-world series mix long-term trends (gradual warming), seasonal cycles (summer/winter), and noise. When these are tangled together, it is hard for any model to find reliable dependencies. Classical time series analysis solved this decades ago with decomposition — but only as a preprocessing step applied once to the raw data.
Innovation 1: progressive decomposition inside the network
Classical methods like ARIMA decompose the time series before feeding it to the model: extract trend, extract , forecast each separately. But this is a one-shot operation on the known past — it cannot refine the decomposition as the model builds its prediction of the unknown future.
Autoformer makes decomposition an inner block of the network. After every sub-layer (attention, feed-forward), the intermediate representation is passed through a Series Decomposition Block that splits it into trend and seasonal components using a simple . The trend is the smoothed version (capturing slow changes), and the seasonal part is the residual (capturing oscillations).
Think of it like peeling an onion: each layer removes another shell of trend, leaving a purer seasonal signal for the next layer to work with. The decoder accumulates the trend pieces from every layer, progressively building the final trend prediction.
Innovation 2: Auto-Correlation replaces self-attention
The key insight is that time series have periodicity: the same phase position across different periods tends to exhibit similar patterns. Monday morning traffic looks like last Monday morning traffic. The temperature at noon today resembles noon yesterday.
Auto-Correlation exploits this directly. Instead of asking "which individual time steps are similar?" (as self-attention does), it asks "which time delays produce the highest correlation?" — a question that naturally identifies periodic structure. The mechanism has two stages:
Stage 1 — Period discovery: Compute the autocorrelation R(τ) between the query and key series using FFT. Peaks in R(τ) indicate likely period lengths. Select the top-k delays with the highest autocorrelation as the most probable periods.
Stage 2 — Time delay aggregation: For each selected delay τᵢ, roll the value series by τᵢ positions. This aligns sub-series that sit at the same phase of the estimated period. Aggregate these aligned sub-series using -weighted autocorrelation confidences. The result captures entire sub-sequence patterns rather than individual points.
Putting it together: the Autoformer architecture
Autoformer retains the skeleton of the Transformer but replaces every self-attention with Auto-Correlation and wraps every sub-layer output in a Series Decomposition Block.
Encoder — receives the past I time steps. Each of its N layers applies: (1) Auto-Correlation + residual → decomposition (discard trend, keep seasonal), (2) feed-forward + residual → decomposition (discard trend, keep seasonal). The encoder progressively strips away the trend, outputting a pure seasonal representation.
Decoder — is initialized with two streams: (a) the seasonal component of the last I/2 steps concatenated with zeros of length O (the ), and (b) the trend-cyclical component concatenated with the mean of the input. Each of its M layers has three decomposition steps: self-Auto-Correlation, cross-Auto-Correlation with the encoder output, and feed-forward. Each step extracts trend which is accumulated via learnable projections. The final prediction is the sum of the refined seasonal component and the accumulated trend.
Simplified to show the idea — not the real implementation.
import torch
import torch.nn as nn
class SeriesDecomp(nn.Module):
"""Separate trend from seasonal using moving average."""
def __init__(self, kernel_size: int = 25):
super().__init__()
self.kernel_size = kernel_size
# Padding to keep length unchanged
self.avg = nn.AvgPool1d(
kernel_size=kernel_size, stride=1,
padding=0 # manual padding below
)
def forward(self, x):
# x shape: (batch, length, channels)
# Pad both sides to preserve length
pad = self.kernel_size // 2
front = x[:, :1, :].repeat(1, pad, 1)
end = x[:, -1:, :].repeat(1, pad, 1)
x_padded = torch.cat([front, x, end], dim=1)
# Moving average → trend
# AvgPool1d expects (batch, channels, length)
trend = self.avg(x_padded.permute(0, 2, 1))
trend = trend.permute(0, 2, 1)
# Seasonal = original − trend
seasonal = x - trend
return seasonal, trendEfficiency: O(L log L) with richer connections
Auto-Correlation achieves O(L log L) complexity through two properties. First, the autocorrelation of all L lags is computed at once via FFT (O(L log L) by the Wiener–Khinchin theorem). Second, only k = ⌊c·log L⌋ delays are selected, and each aggregation step processes a length-L series, giving O(L log L) total.
But complexity alone doesn't tell the full story. What matters is information density per connection. Self-attention connects individual points (each connection carries one scalar value). Auto-Correlation connects entire sub-series at each delay (each connection carries L values). This means Auto-Correlation achieves both lower complexity and richer information flow — breaking the efficiency-vs-accuracy tradeoff that plagued sparse attention methods.
Results: 38% improvement across six benchmarks
Autoformer was evaluated on six real-world datasets covering energy (ETT, Electricity), traffic, exchange rates, weather, and influenza-like illness (ILI). The input length was fixed at 96 (36 for ILI), and prediction horizons ranged from 96 to 720 steps.
Across all benchmarks and all prediction lengths, Autoformer achieved consistent state-of-the-art performance. Highlights include a 74% MSE reduction on ETT (from 1.334 to 0.339 at predict-336), 61% on Exchange, and 43% on ILI. The averaged improvement across settings was 38%.
Notably, Autoformer's performance degraded gracefully as the prediction horizon increased. While baselines often showed sharp jumps in error at longer horizons, Autoformer maintained stable accuracy — a crucial property for real-world long-term planning applications.
What Autoformer learns: interpretable periodicity
One striking property of Auto-Correlation is interpretability. By examining the learned delay values from the decoder, researchers found that the model automatically discovers meaningful periodicities from the data:
For the hourly Traffic dataset, the top delays clustered at 24 (daily cycle) and 168 (weekly cycle). For the daily Exchange dataset, delays indicated monthly, quarterly, and yearly periods. The model didn't need to be told about these cycles — it discovered them from the raw data through the autocorrelation computation.
This is a significant advantage over self-attention, which offers no interpretable structure: attention weights between individual time steps don't reveal what periodic patterns the model is using.
Context in the time series forecasting landscape
1970
ARIMA
Box and Jenkins formalize autoregressive integrated moving average models. The gold standard for statistical time series forecasting for decades.
2017
Transformer
Vaswani et al. introduce the self-attention mechanism. Initially for NLP, but its ability to model long-range dependencies attracts time series researchers.
2019
LogTrans
LogSparse attention selects time steps at exponentially increasing intervals. Reduces complexity to O(L(log L)²) but still uses point-wise connections.
2021
Informer
ProbSparse attention with KL-divergence selection. O(L log L) complexity. First Transformer to seriously target long-term series forecasting.
2021
Autoformer
Breaks away from point-wise attention entirely. Progressive decomposition + Auto-Correlation achieve 38% improvement over Informer and predecessors.
2023
DLinear
Zeng et al. show that a single linear layer can match or beat Transformers on many benchmarks — sparking debate about whether complex architectures are needed at all.
CitationWu, Xu, Wang, Long. Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting. NeurIPS, 2021.
Terms in this paper
- Autocorrelationالارتباط الذاتي
- Time Seriesسلسلة زمنية
- Forecastingالتنبؤ
- Transformerالمحوِّل
- Encoderالمُرمِّز
- Decoderمفكّ الترميز
- Attentionآلية الانتباه
- Moving Averageالمتوسط المتحرك
- Stochastic Processعملية عشوائية
- Fast Fourier Transformتحويل فورييه السريع
- Softmaxسوفت ماكس
- Backpropagationالتحديث التراجعي
- Mean Squared Errorمتوسط مربع الخطأ
- Mean Absolute Errorمتوسط الخطأ المطلق