Time Series2023intermediate12 min read

Are Transformers Effective for Time Series Forecasting?

هل المحوِّلات فعّالة في التنبؤ بالسلاسل الزمنية؟

Zeng, A. · Chen, M. · Zhang, L. · Xu, Q. — AAAI

The problem

By 2022, -based models had become the dominant approach for long-term , with Informer winning AAAI 2021 Best Paper and Autoformer appearing at NeurIPS 2021. These models introduced increasingly complex variants to capture temporal dependencies. However, self-attention is fundamentally — it extracts semantic correlations between paired elements regardless of order. , by contrast, are inherently ordered sequences where temporal position is the most critical information. Moreover, these Transformer models were only compared against baselines that suffer from error accumulation, making the comparisons unfair.

The contribution

The authors proposed DLinear, an embarrassingly simple model that decomposes a time series into and remainder using a , then applies two single- linear networks to forecast each component. Using (not autoregressive), DLinear outperformed all existing Transformer models — Informer, Autoformer, FEDformer, Pyraformer, LogTrans — on 9 benchmarks in most cases, often by 25–50%. The paper also showed that Transformer models fail to improve with larger look-back windows, while DLinear improves consistently — suggesting Transformers do not truly extract temporal relations.

The impact

This paper sent shockwaves through the time series community. It forced researchers to rethink whether architectural complexity was the right path for temporal modeling. It directly inspired PatchTST, which responded by showing that Transformers can be effective when properly designed (using patching instead of point-wise tokens). The paper established DLinear as a mandatory — any new forecasting model must now beat it to be taken seriously. It also highlighted the critical but overlooked distinction between iterated and direct multi-step forecasting strategies.

Imagine a city that keeps building taller and taller skyscrapers to see further into the distance. Each new tower — Informer, Autoformer, FEDformer — adds more floors, more glass, more elevators. But one day, someone climbs a simple hill on the edge of town and discovers they can see further than anyone in the skyscrapers. Why? Because the skyscrapers were designed for a city skyline — spotting landmarks and buildings (semantic tokens) — not for watching a continuous horizon (time series). The hill is DLinear: no floors, no elevators, just a direct line of sight.

The question: does self-attention actually help with time series?

Transformers conquered NLP and by finding semantic relationships between tokens — which word relates to which, which image patch connects to which. Self-attention excels at this because it treats every pair of elements equally, regardless of position. But this very property — — is a fundamental mismatch for time series.

In a time series, the order of data points is the signal. Stock prices, temperature readings, electricity consumption — they are continuous flows where position carries meaning. Self-attention, by contrast, is "anti-ordering": it sees every pair but ignores their sequential relationship. While positional encodings partially compensate, the fundamental architecture was designed for semantic correlation, not temporal dynamics.

The authors posed a sharp question: if we strip away all the architectural complexity and use the simplest possible model with direct multi-step forecasting, will the Transformer's supposed advantages survive?

Open in Lab
Shuffle the time series — self-attention gives the same output. A linear layer uses position directly.
The demo wakes as you arrive…

The hidden variable: how you forecast matters more than what you forecast with

Before comparing architectures, there is a critical distinction that previous papers glossed over: how multi-step predictions are generated.

Iterated Multi-Step (IMS) forecasting trains a single-step predictor and applies it recursively — predict step 1, feed it back, predict step 2, and so on. This is the autoregressive approach used by RNNs and classical methods. It works well for short horizons, but errors compound with each step. By the time you reach step 720, the accumulated noise can be devastating.

Direct Multi-Step (DMS) forecasting predicts all future steps at once. The model takes the and outputs the entire forecast horizon in a single pass. No error accumulation, because there is no feedback loop.

The key insight: all existing Transformer-based models (Informer, Autoformer, FEDformer) use DMS forecasting with non-autoregressive decoders. But the baselines they compared against — RNNs, DeepAR, classical methods — all use IMS. This is not a fair comparison. The Transformers' supposed advantage might come entirely from the forecasting strategy, not the architecture.

Open in Lab
Watch how errors accumulate in IMS forecasting vs the single-pass DMS approach.
The demo wakes as you arrive…

DLinear: decompose, then project

The DLinear architecture is built on two observations. First, a single linear layer is the simplest possible way to aggregate historical information for future prediction — it directly maps L input time steps to T output time steps through a learned . Second, decomposing a time series into trend and remainder (as Autoformer showed) is a model-agnostic technique that can boost any architecture.

DLinear combines both: it applies a moving average with kernel size 25 to extract the trend component, subtracts it to get the remainder, then feeds each through its own single-layer linear network. The two outputs are summed to produce the final forecast.

That is the entire model. No attention heads, no multi-layer encoders, no positional embeddings, no feed-forward networks. Just plus two matrix multiplications.

Open in Lab
Trace the data flow through DLinear: decomposition splits, two linear layers process, and outputs merge.
The demo wakes as you arrive…
X^=Ws⋅Xs+Wt⋅Xt\hat{X} = W_s \cdot X_s + W_t \cdot X_t
DLinear forecast — sum of two linear projections — The look-back window X is decomposed into trend Xₜ and remainder Xₛ via moving average. Two weight matrices Wₛ ∈ ℝᵀˣᴸ and Wₜ ∈ ℝᵀˣᴸ project each component from L input steps to T output steps. The final forecast is their sum.

Why decomposition matters: separating the slow from the fast

A raw time series is a superposition of multiple dynamics operating at different scales. Think of it like ocean waves: there is the tide (the slow, long-term trend) and the ripples (the fast, seasonal fluctuations). Predicting both at once with a single model forces it to simultaneously learn slow dynamics and fast oscillations — a harder optimization problem.

The decomposition strategy borrowed from Autoformer applies a moving average to smooth out the ripples and extract the tide. The difference between the raw signal and the smoothed tide is the remainder — the seasonal or cyclical component.

The confirmed that decomposition helps most when there is a clear trend in the data (Exchange-Rate, ILI, ETT datasets) and provides less benefit for trendless data (Traffic, Weather). This makes intuitive sense: if there is no tide to separate, the decomposition step is redundant.

Open in Lab
See how the moving average separates trend from remainder on different dataset types.
The demo wakes as you arrive…

The idea in code

DLinear — the complete model in ~30 linespython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn as nn

class MovingAvg(nn.Module):
    """Extracts the trend from a time series via moving average."""
    def __init__(self, kernel_size=25):
        super().__init__()
        self.kernel_size = kernel_size
        self.avg = nn.AvgPool1d(kernel_size, stride=1, padding=0)

    def forward(self, x):
        # Pad the edges to keep the same length
        front = x[:, :1, :].repeat(1, (self.kernel_size - 1) // 2, 1)
        end = x[:, -1:, :].repeat(1, (self.kernel_size - 1) // 2, 1)
        x_padded = torch.cat([front, x, end], dim=1)
        trend = self.avg(x_padded.permute(0, 2, 1)).permute(0, 2, 1)
        return trend

class DLinear(nn.Module):
    """Decomposition + two linear layers. That's it."""
    def __init__(self, look_back=96, horizon=96, channels=1):
        super().__init__()
        self.decomp = MovingAvg(kernel_size=25)
        self.linear_trend = nn.Linear(look_back, horizon)    # Wt
        self.linear_remainder = nn.Linear(look_back, horizon) # Ws

    def forward(self, x):
        # x: (batch, look_back, channels)
        trend = self.decomp(x)                  # smooth trend
        remainder = x - trend                   # what's left
        # Project each component from L → T time steps
        trend_out = self.linear_trend(trend.permute(0, 2, 1))
        remainder_out = self.linear_remainder(remainder.permute(0, 2, 1))
        forecast = (trend_out + remainder_out).permute(0, 2, 1)
        return forecast  # (batch, horizon, channels)

# Total parameters for L=96, T=96, C=1:
# Wt: 96×96 = 9,216   Ws: 96×96 = 9,216   Total: ~18K
# Compare: Informer ≈ 14M, Autoformer ≈ 15M, FEDformer ≈ 21M

Results: simplicity wins across 9 benchmarks

The results were striking. DLinear outperformed all Transformer-based methods in most settings across 9 datasets: ETTh1, ETTh2, ETTm1, ETTm2, Traffic, Electricity, Weather, Exchange-Rate, and ILI.

The margins were not small. On Exchange-Rate, DLinear beat the state-of-the-art FEDformer by roughly 50% in . On Weather, the improvement exceeded 25%. Even the naive "Repeat-C" baseline (just repeat the last value) outperformed all Transformers on Exchange-Rate by over 30%.

Efficiency was equally dramatic. DLinear requires only 0.04G MACs compared to Informer's 3.93G, Autoformer's 4.41G, and FEDformer's 4.41G. runs in 0.4 ms versus 40–164 ms for the Transformer variants. Parameters: 140K vs 14–21 million.

Open in Lab
Compare MSE across models and datasets — DLinear consistently matches or beats Transformers.
The demo wakes as you arrive…

The look-back window test: can Transformers use more history?

A model that truly captures temporal relations should improve when given more historical data. If you can see 30 days instead of 4, you should forecast better — period.

The authors tested all models with look-back windows from 24 to 720 steps. The results revealed a damning pattern: Transformer-based models either stagnated or got worse with larger windows. Informer, Autoformer, and Pyraformer showed flat or increasing MSE as the look-back window grew. FEDformer was slightly better but still plateaued.

DLinear, by contrast, consistently improved with longer look-back windows. On Traffic, its MSE dropped from 0.65 (L=24) to around 0.41 (L=720). On Electricity, from 0.23 to 0.13. The improvement was dramatic and consistent.

This is perhaps the paper's most devastating finding. It suggests that Transformers are not actually learning temporal dependencies from the data — they are merely mapping inputs to outputs without exploiting the temporal structure that additional history provides.

Open in Lab
Drag the slider to increase the look-back window and watch each model's MSE respond.
The demo wakes as you arrive…

Reading the weights: what DLinear actually learns

Because DLinear is a linear model, its learned weights W ∈ ℝᵀˣᴸ directly reveal which input time steps influence which output steps. This is a level of that Transformers, with their multi-headed attention layers, cannot match.

For the Electricity dataset (hourly granularity), the remainder weights showed clear diagonal bands spaced 24 steps apart — the model learned daily periodicity. The value at output step 1 depends most on input steps 1, 25, 49, 73 — i.e., the same hour on previous days. The Traffic weights revealed both daily (24-step) and weekly (168-step) periodicity.

For Exchange-Rate (no periodicity), the trend weights showed higher values near the most recent time steps — the model learned that recent values matter most for predicting aperiodic financial data. This interpretability is not just a nice feature; it provides validation that the model is learning sensible temporal patterns.

Open in Lab
Explore DLinear's learned weights — periodic patterns emerge as diagonal bands.
The demo wakes as you arrive…

Dissecting the Transformers: what helps, what hurts

The authors systematically removed components from existing Transformer models to understand what actually contributes to performance. When self-attention layers and auxiliary designs were stripped from Informer step by step — gradually converting it into a linear model — performance improved at each step. The closer the model got to being a plain linear layer, the better it forecasted.

strategies also told a revealing story. Informer collapsed without positional embeddings (it uses single time steps as tokens), while FEDformer — which relies more on Fourier transforms than self-attention — was more robust to removing embeddings. This suggests that FEDformer's strength comes from its frequency analysis, not from the Transformer architecture itself.

Even training data size did not save the Transformers. On the Traffic dataset, reducing training data from 17,544 hours to 8,760 hours (one year) actually improved predictions for all models — contradicting the argument that Transformers would shine with more data.

Limitations: what DLinear cannot do

The authors were transparent about DLinear's limitations. As a single-layer linear model, it has limited capacity — it cannot capture nonlinear dynamics, change points, or complex multi-scale interactions. It was designed as a baseline, not as the final answer to time series forecasting.

Later work, notably PatchTST, showed that Transformers can work for time series when the tokenization is redesigned. Instead of treating each time step as a (which fragments the temporal signal), PatchTST groups consecutive time steps into patches — similar to how Vision Transformers handle images. This preserves local temporal structure before self-attention operates, and PatchTST consistently outperformed DLinear.

The debate DLinear started is still ongoing: simplicity as a baseline forced the community to be more rigorous about what Transformers actually contribute, leading to better Transformer designs rather than abandoning them entirely.

Timeline: the great time series debate

  1. 2021

    Informer — AAAI Best Paper

    ProbSparse self-attention for long-term time series forecasting. Introduced DMS decoding to Transformers for time series, achieving state-of-the-art on ETT benchmarks.

  2. 2021

    Autoformer — NeurIPS

    Introduced seasonal-trend decomposition inside the Transformer and replaced self-attention with auto-correlation. The decomposition idea would later inspire DLinear.

  3. 2022

    FEDformer — ICML

    Frequency-enhanced decomposed Transformer. Used Fourier and wavelet transforms for O(L) complexity. Closest Transformer competitor to DLinear.

  4. 2023

    DLinear — AAAI (this paper)

    A single-layer linear model that outperformed all existing Transformers. Questioned the fundamental effectiveness of self-attention for temporal modeling.

  5. 2023

    PatchTST — ICLR

    Responded to DLinear by showing Transformers work when time steps are grouped into patches. Restored Transformer credibility for time series with proper tokenization.

CitationZeng, Chen, Zhang, Xu. Are Transformers Effective for Time Series Forecasting?. AAAI, 2023.

Terms in this paper