Time Series2021intermediate10 min read

Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting

Autoformer: بنية تفكيك تقدّمية مع ارتباط ذاتي للتنبؤ بعيد المدى بالسلاسل الزمنية

Wu, H. · Xu, J. · Wang, J. · Long, M. — NeurIPS

The problem

Long-term is critical for applications like energy planning, weather warning, and disease tracking, but existing -based models face two fundamental problems. First, the intricate temporal patterns of long-term series — where trends, seasons, and intertwine — make it hard to discover reliable dependencies. Second, standard has quadratic complexity, and the sparse approximations (ProbSparse, LogSparse, LSH) that fix this are all point-wise, creating an information utilization bottleneck: they connect individual time steps rather than coherent sub-sequences.

The contribution

Autoformer introduces two core innovations. First, a progressive architecture that embeds moving-average-based -seasonal separation as an inner block inside every and layer — not as preprocessing. This lets the model iteratively refine the trend and seasonal components throughout the forecasting process. Second, an Auto-Correlation mechanism that replaces self-: it uses FFT-based to discover periodic dependencies and aggregates similar sub-series via time-delay rolling, achieving O(L log L) complexity while connecting entire sub-sequences rather than scattered points. On six benchmarks covering energy, traffic, economics, weather, and disease, Autoformer achieves a 38% relative MSE reduction over the previous state of the art.

The impact

Autoformer demonstrated that the Transformer's self-attention is not the only — or best — way to model temporal dependencies in time series. By proving that series-level periodicity connections outperform point-wise attention, it opened the door to a new family of forecasting architectures. It directly influenced DLinear (which questioned whether even simple linear models suffice) and the broader debate on what inductive biases time series models need. The inner decomposition idea has been adopted by subsequent models, and the paper remains a key reference in the time series forecasting literature.

Imagine trying to predict the tide a month from now. A weather forecaster wouldn't stare at the raw water-level graph — it's a tangled mess of rising ocean, daily tides, and random waves. Instead, she'd peel the signal apart: first extract the slow seasonal rise, then isolate the 12-hour tidal cycle, then handle the noise separately.

Standard Transformers try to read the raw mess directly, connecting individual water-level readings with attention. Autoformer works like the forecaster: it continuously peels apart the trend from the cycles inside every layer, and instead of linking individual readings, it matches entire tidal cycles — recognizing that today's tide looks like yesterday's tide shifted by 12 hours.

The problem: why Transformers struggle with long-term series

Informer and related Transformer-based forecasters improved efficiency with sparse attention variants, but they all share a fundamental limitation: point-wise connections. Self-attention computes a score between individual time steps — "how relevant is step 42 to step 100?" — then aggregates values from selected points. This works for language, where a single word can carry critical meaning, but time series are different.

In a time series, meaning lives in sub-sequences, not individual points. A single temperature reading at 3 PM means little on its own; it's the pattern across the full day — the diurnal cycle — that carries predictive power. Point-wise attention cannot naturally capture this: it sees trees, not the forest.

The second problem is entangled patterns. Real-world series mix long-term trends (gradual warming), seasonal cycles (summer/winter), and noise. When these are tangled together, it is hard for any model to find reliable dependencies. Classical time series analysis solved this decades ago with decomposition — but only as a preprocessing step applied once to the raw data.

Open in Lab
Left: point-wise self-attention connects individual time steps. Right: Auto-Correlation connects entire sub-series across periods. Click points or periods to compare.
The demo wakes as you arrive…

Innovation 1: progressive decomposition inside the network

Classical methods like ARIMA decompose the time series before feeding it to the model: extract trend, extract , forecast each separately. But this is a one-shot operation on the known past — it cannot refine the decomposition as the model builds its prediction of the unknown future.

Autoformer makes decomposition an inner block of the network. After every sub-layer (attention, feed-forward), the intermediate representation is passed through a Series Decomposition Block that splits it into trend and seasonal components using a simple . The trend is the smoothed version (capturing slow changes), and the seasonal part is the residual (capturing oscillations).

Think of it like peeling an onion: each layer removes another shell of trend, leaving a purer seasonal signal for the next layer to work with. The decoder accumulates the trend pieces from every layer, progressively building the final trend prediction.

Xt=AvgPool ⁣(Padding(X)),Xs=X−Xt\mathcal{X}_t = \text{AvgPool}\!\bigl(\text{Padding}(\mathcal{X})\bigr), \qquad \mathcal{X}_s = \mathcal{X} - \mathcal{X}_t
Series Decomposition Block — moving average separation — Given an input series 𝒳 ∈ ℝ^{L×d}, a moving average (AvgPool with padding) extracts the trend-cyclical component 𝒳_t by smoothing periodic fluctuations. The seasonal component 𝒳_s is simply the residual. This block is applied after every attention and feed-forward sub-layer throughout the entire architecture.
Open in Lab
Drag the window size slider to see how the moving average decomposes a signal into trend and seasonal parts. Larger windows produce smoother trends.
The demo wakes as you arrive…

Innovation 2: Auto-Correlation replaces self-attention

The key insight is that time series have periodicity: the same phase position across different periods tends to exhibit similar patterns. Monday morning traffic looks like last Monday morning traffic. The temperature at noon today resembles noon yesterday.

Auto-Correlation exploits this directly. Instead of asking "which individual time steps are similar?" (as self-attention does), it asks "which time delays produce the highest correlation?" — a question that naturally identifies periodic structure. The mechanism has two stages:

Stage 1 — Period discovery: Compute the autocorrelation R(τ) between the query and key series using FFT. Peaks in R(τ) indicate likely period lengths. Select the top-k delays with the highest autocorrelation as the most probable periods.

Stage 2 — Time delay aggregation: For each selected delay τᵢ, roll the value series by τᵢ positions. This aligns sub-series that sit at the same phase of the estimated period. Aggregate these aligned sub-series using -weighted autocorrelation confidences. The result captures entire sub-sequence patterns rather than individual points.

RQ,K(τ)=F−1 ⁣(F(Q)⋅F∗(K))R_{Q,K}(\tau) = \mathcal{F}^{-1}\!\bigl(\mathcal{F}(Q) \cdot \mathcal{F}^*(K)\bigr)
Autocorrelation via FFT (Wiener–Khinchin theorem) — The autocorrelation R(τ) for all lags τ ∈ {1, …, L} is computed in one shot using FFT. ℱ is the Fast Fourier Transform, ℱ* is its complex conjugate, and ℱ⁻¹ is the inverse. This achieves O(L log L) complexity instead of the O(L²) of naïve computation.
Auto-Correlation(Q,K,V)=∑i=1kRoll(V,τi)⋅R^Q,K(τi)\text{Auto-Correlation}(Q,K,V) = \sum_{i=1}^{k} \text{Roll}(V,\tau_i) \cdot \hat{R}_{Q,K}(\tau_i)
Time Delay Aggregation — The top-k delays τ₁…τₖ (where k = ⌊c·log L⌋) are selected from the autocorrelation peaks. For each delay, the value series V is rolled by τᵢ to align same-phase sub-series. The softmax-normalized autocorrelation R̂(τᵢ) weights the aggregation. Roll(V, τ) shifts elements cyclically — those pushed past the end reappear at the start.
Open in Lab
Watch how Auto-Correlation discovers periodicity via FFT, then rolls and aggregates sub-series from matching phases. Toggle between standard and FFT views.
The demo wakes as you arrive…

Putting it together: the Autoformer architecture

Autoformer retains the skeleton of the Transformer but replaces every self-attention with Auto-Correlation and wraps every sub-layer output in a Series Decomposition Block.

Encoder — receives the past I time steps. Each of its N layers applies: (1) Auto-Correlation + residual → decomposition (discard trend, keep seasonal), (2) feed-forward + residual → decomposition (discard trend, keep seasonal). The encoder progressively strips away the trend, outputting a pure seasonal representation.

Decoder — is initialized with two streams: (a) the seasonal component of the last I/2 steps concatenated with zeros of length O (the ), and (b) the trend-cyclical component concatenated with the mean of the input. Each of its M layers has three decomposition steps: self-Auto-Correlation, cross-Auto-Correlation with the encoder output, and feed-forward. Each step extracts trend which is accumulated via learnable projections. The final prediction is the sum of the refined seasonal component and the accumulated trend.

Open in Lab
Click on any block to see its role. The blue decomposition blocks peel away trend; the green Auto-Correlation blocks discover periodic dependencies.
The demo wakes as you arrive…
Series Decomposition Block — core PyTorch implementationpython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn as nn

class SeriesDecomp(nn.Module):
    """Separate trend from seasonal using moving average."""
    def __init__(self, kernel_size: int = 25):
        super().__init__()
        self.kernel_size = kernel_size
        # Padding to keep length unchanged
        self.avg = nn.AvgPool1d(
            kernel_size=kernel_size, stride=1,
            padding=0  # manual padding below
        )

    def forward(self, x):
        # x shape: (batch, length, channels)
        # Pad both sides to preserve length
        pad = self.kernel_size // 2
        front = x[:, :1, :].repeat(1, pad, 1)
        end = x[:, -1:, :].repeat(1, pad, 1)
        x_padded = torch.cat([front, x, end], dim=1)

        # Moving average → trend
        # AvgPool1d expects (batch, channels, length)
        trend = self.avg(x_padded.permute(0, 2, 1))
        trend = trend.permute(0, 2, 1)

        # Seasonal = original − trend
        seasonal = x - trend
        return seasonal, trend

Efficiency: O(L log L) with richer connections

Auto-Correlation achieves O(L log L) complexity through two properties. First, the autocorrelation of all L lags is computed at once via FFT (O(L log L) by the Wiener–Khinchin theorem). Second, only k = ⌊c·log L⌋ delays are selected, and each aggregation step processes a length-L series, giving O(L log L) total.

But complexity alone doesn't tell the full story. What matters is information density per connection. Self-attention connects individual points (each connection carries one scalar value). Auto-Correlation connects entire sub-series at each delay (each connection carries L values). This means Auto-Correlation achieves both lower complexity and richer information flow — breaking the efficiency-vs-accuracy tradeoff that plagued sparse attention methods.

Results: 38% improvement across six benchmarks

Autoformer was evaluated on six real-world datasets covering energy (ETT, Electricity), traffic, exchange rates, weather, and influenza-like illness (ILI). The input length was fixed at 96 (36 for ILI), and prediction horizons ranged from 96 to 720 steps.

Across all benchmarks and all prediction lengths, Autoformer achieved consistent state-of-the-art performance. Highlights include a 74% MSE reduction on ETT (from 1.334 to 0.339 at predict-336), 61% on Exchange, and 43% on ILI. The averaged improvement across settings was 38%.

Notably, Autoformer's performance degraded gracefully as the prediction horizon increased. While baselines often showed sharp jumps in error at longer horizons, Autoformer maintained stable accuracy — a crucial property for real-world long-term planning applications.

Open in Lab
MSE comparison across prediction horizons. Lower is better. Notice how Autoformer's curve stays flat while baselines diverge.
The demo wakes as you arrive…

What Autoformer learns: interpretable periodicity

One striking property of Auto-Correlation is interpretability. By examining the learned delay values from the decoder, researchers found that the model automatically discovers meaningful periodicities from the data:

For the hourly Traffic dataset, the top delays clustered at 24 (daily cycle) and 168 (weekly cycle). For the daily Exchange dataset, delays indicated monthly, quarterly, and yearly periods. The model didn't need to be told about these cycles — it discovered them from the raw data through the autocorrelation computation.

This is a significant advantage over self-attention, which offers no interpretable structure: attention weights between individual time steps don't reveal what periodic patterns the model is using.

Context in the time series forecasting landscape

  1. 1970

    ARIMA

    Box and Jenkins formalize autoregressive integrated moving average models. The gold standard for statistical time series forecasting for decades.

  2. 2017

    Transformer

    Vaswani et al. introduce the self-attention mechanism. Initially for NLP, but its ability to model long-range dependencies attracts time series researchers.

  3. 2019

    LogTrans

    LogSparse attention selects time steps at exponentially increasing intervals. Reduces complexity to O(L(log L)²) but still uses point-wise connections.

  4. 2021

    Informer

    ProbSparse attention with KL-divergence selection. O(L log L) complexity. First Transformer to seriously target long-term series forecasting.

  5. 2021

    Autoformer

    Breaks away from point-wise attention entirely. Progressive decomposition + Auto-Correlation achieve 38% improvement over Informer and predecessors.

  6. 2023

    DLinear

    Zeng et al. show that a single linear layer can match or beat Transformers on many benchmarks — sparking debate about whether complex architectures are needed at all.

CitationWu, Xu, Wang, Long. Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting. NeurIPS, 2021.

Terms in this paper