Time Series2023intermediate12 min read
Are Transformers Effective for Time Series Forecasting?
هل المحوِّلات فعّالة في التنبؤ بالسلاسل الزمنية؟
Zeng, A. · Chen, M. · Zhang, L. · Xu, Q. — AAAI
The problem
By 2022, -based models had become the dominant approach for long-term , with Informer winning AAAI 2021 Best Paper and Autoformer appearing at NeurIPS 2021. These models introduced increasingly complex variants to capture temporal dependencies. However, self-attention is fundamentally — it extracts semantic correlations between paired elements regardless of order. , by contrast, are inherently ordered sequences where temporal position is the most critical information. Moreover, these Transformer models were only compared against baselines that suffer from error accumulation, making the comparisons unfair.
The contribution
The authors proposed DLinear, an embarrassingly simple model that decomposes a time series into and remainder using a , then applies two single- linear networks to forecast each component. Using (not autoregressive), DLinear outperformed all existing Transformer models — Informer, Autoformer, FEDformer, Pyraformer, LogTrans — on 9 benchmarks in most cases, often by 25–50%. The paper also showed that Transformer models fail to improve with larger look-back windows, while DLinear improves consistently — suggesting Transformers do not truly extract temporal relations.
The impact
This paper sent shockwaves through the time series community. It forced researchers to rethink whether architectural complexity was the right path for temporal modeling. It directly inspired PatchTST, which responded by showing that Transformers can be effective when properly designed (using patching instead of point-wise tokens). The paper established DLinear as a mandatory — any new forecasting model must now beat it to be taken seriously. It also highlighted the critical but overlooked distinction between iterated and direct multi-step forecasting strategies.
Imagine a city that keeps building taller and taller skyscrapers to see further into the distance. Each new tower — Informer, Autoformer, FEDformer — adds more floors, more glass, more elevators. But one day, someone climbs a simple hill on the edge of town and discovers they can see further than anyone in the skyscrapers. Why? Because the skyscrapers were designed for a city skyline — spotting landmarks and buildings (semantic tokens) — not for watching a continuous horizon (time series). The hill is DLinear: no floors, no elevators, just a direct line of sight.
The question: does self-attention actually help with time series?
Transformers conquered NLP and by finding semantic relationships between tokens — which word relates to which, which image patch connects to which. Self-attention excels at this because it treats every pair of elements equally, regardless of position. But this very property — — is a fundamental mismatch for time series.
In a time series, the order of data points is the signal. Stock prices, temperature readings, electricity consumption — they are continuous flows where position carries meaning. Self-attention, by contrast, is "anti-ordering": it sees every pair but ignores their sequential relationship. While positional encodings partially compensate, the fundamental architecture was designed for semantic correlation, not temporal dynamics.
The authors posed a sharp question: if we strip away all the architectural complexity and use the simplest possible model with direct multi-step forecasting, will the Transformer's supposed advantages survive?
The hidden variable: how you forecast matters more than what you forecast with
Before comparing architectures, there is a critical distinction that previous papers glossed over: how multi-step predictions are generated.
Iterated Multi-Step (IMS) forecasting trains a single-step predictor and applies it recursively — predict step 1, feed it back, predict step 2, and so on. This is the autoregressive approach used by RNNs and classical methods. It works well for short horizons, but errors compound with each step. By the time you reach step 720, the accumulated noise can be devastating.
Direct Multi-Step (DMS) forecasting predicts all future steps at once. The model takes the and outputs the entire forecast horizon in a single pass. No error accumulation, because there is no feedback loop.
The key insight: all existing Transformer-based models (Informer, Autoformer, FEDformer) use DMS forecasting with non-autoregressive decoders. But the baselines they compared against — RNNs, DeepAR, classical methods — all use IMS. This is not a fair comparison. The Transformers' supposed advantage might come entirely from the forecasting strategy, not the architecture.
DLinear: decompose, then project
The DLinear architecture is built on two observations. First, a single linear layer is the simplest possible way to aggregate historical information for future prediction — it directly maps L input time steps to T output time steps through a learned . Second, decomposing a time series into trend and remainder (as Autoformer showed) is a model-agnostic technique that can boost any architecture.
DLinear combines both: it applies a moving average with kernel size 25 to extract the trend component, subtracts it to get the remainder, then feeds each through its own single-layer linear network. The two outputs are summed to produce the final forecast.
That is the entire model. No attention heads, no multi-layer encoders, no positional embeddings, no feed-forward networks. Just plus two matrix multiplications.
Why decomposition matters: separating the slow from the fast
A raw time series is a superposition of multiple dynamics operating at different scales. Think of it like ocean waves: there is the tide (the slow, long-term trend) and the ripples (the fast, seasonal fluctuations). Predicting both at once with a single model forces it to simultaneously learn slow dynamics and fast oscillations — a harder optimization problem.
The decomposition strategy borrowed from Autoformer applies a moving average to smooth out the ripples and extract the tide. The difference between the raw signal and the smoothed tide is the remainder — the seasonal or cyclical component.
The confirmed that decomposition helps most when there is a clear trend in the data (Exchange-Rate, ILI, ETT datasets) and provides less benefit for trendless data (Traffic, Weather). This makes intuitive sense: if there is no tide to separate, the decomposition step is redundant.
The idea in code
Simplified to show the idea — not the real implementation.
import torch
import torch.nn as nn
class MovingAvg(nn.Module):
"""Extracts the trend from a time series via moving average."""
def __init__(self, kernel_size=25):
super().__init__()
self.kernel_size = kernel_size
self.avg = nn.AvgPool1d(kernel_size, stride=1, padding=0)
def forward(self, x):
# Pad the edges to keep the same length
front = x[:, :1, :].repeat(1, (self.kernel_size - 1) // 2, 1)
end = x[:, -1:, :].repeat(1, (self.kernel_size - 1) // 2, 1)
x_padded = torch.cat([front, x, end], dim=1)
trend = self.avg(x_padded.permute(0, 2, 1)).permute(0, 2, 1)
return trend
class DLinear(nn.Module):
"""Decomposition + two linear layers. That's it."""
def __init__(self, look_back=96, horizon=96, channels=1):
super().__init__()
self.decomp = MovingAvg(kernel_size=25)
self.linear_trend = nn.Linear(look_back, horizon) # Wt
self.linear_remainder = nn.Linear(look_back, horizon) # Ws
def forward(self, x):
# x: (batch, look_back, channels)
trend = self.decomp(x) # smooth trend
remainder = x - trend # what's left
# Project each component from L → T time steps
trend_out = self.linear_trend(trend.permute(0, 2, 1))
remainder_out = self.linear_remainder(remainder.permute(0, 2, 1))
forecast = (trend_out + remainder_out).permute(0, 2, 1)
return forecast # (batch, horizon, channels)
# Total parameters for L=96, T=96, C=1:
# Wt: 96×96 = 9,216 Ws: 96×96 = 9,216 Total: ~18K
# Compare: Informer ≈ 14M, Autoformer ≈ 15M, FEDformer ≈ 21MResults: simplicity wins across 9 benchmarks
The results were striking. DLinear outperformed all Transformer-based methods in most settings across 9 datasets: ETTh1, ETTh2, ETTm1, ETTm2, Traffic, Electricity, Weather, Exchange-Rate, and ILI.
The margins were not small. On Exchange-Rate, DLinear beat the state-of-the-art FEDformer by roughly 50% in . On Weather, the improvement exceeded 25%. Even the naive "Repeat-C" baseline (just repeat the last value) outperformed all Transformers on Exchange-Rate by over 30%.
Efficiency was equally dramatic. DLinear requires only 0.04G MACs compared to Informer's 3.93G, Autoformer's 4.41G, and FEDformer's 4.41G. runs in 0.4 ms versus 40–164 ms for the Transformer variants. Parameters: 140K vs 14–21 million.
The look-back window test: can Transformers use more history?
A model that truly captures temporal relations should improve when given more historical data. If you can see 30 days instead of 4, you should forecast better — period.
The authors tested all models with look-back windows from 24 to 720 steps. The results revealed a damning pattern: Transformer-based models either stagnated or got worse with larger windows. Informer, Autoformer, and Pyraformer showed flat or increasing MSE as the look-back window grew. FEDformer was slightly better but still plateaued.
DLinear, by contrast, consistently improved with longer look-back windows. On Traffic, its MSE dropped from 0.65 (L=24) to around 0.41 (L=720). On Electricity, from 0.23 to 0.13. The improvement was dramatic and consistent.
This is perhaps the paper's most devastating finding. It suggests that Transformers are not actually learning temporal dependencies from the data — they are merely mapping inputs to outputs without exploiting the temporal structure that additional history provides.
Reading the weights: what DLinear actually learns
Because DLinear is a linear model, its learned weights W ∈ ℝᵀˣᴸ directly reveal which input time steps influence which output steps. This is a level of that Transformers, with their multi-headed attention layers, cannot match.
For the Electricity dataset (hourly granularity), the remainder weights showed clear diagonal bands spaced 24 steps apart — the model learned daily periodicity. The value at output step 1 depends most on input steps 1, 25, 49, 73 — i.e., the same hour on previous days. The Traffic weights revealed both daily (24-step) and weekly (168-step) periodicity.
For Exchange-Rate (no periodicity), the trend weights showed higher values near the most recent time steps — the model learned that recent values matter most for predicting aperiodic financial data. This interpretability is not just a nice feature; it provides validation that the model is learning sensible temporal patterns.
Dissecting the Transformers: what helps, what hurts
The authors systematically removed components from existing Transformer models to understand what actually contributes to performance. When self-attention layers and auxiliary designs were stripped from Informer step by step — gradually converting it into a linear model — performance improved at each step. The closer the model got to being a plain linear layer, the better it forecasted.
strategies also told a revealing story. Informer collapsed without positional embeddings (it uses single time steps as tokens), while FEDformer — which relies more on Fourier transforms than self-attention — was more robust to removing embeddings. This suggests that FEDformer's strength comes from its frequency analysis, not from the Transformer architecture itself.
Even training data size did not save the Transformers. On the Traffic dataset, reducing training data from 17,544 hours to 8,760 hours (one year) actually improved predictions for all models — contradicting the argument that Transformers would shine with more data.
Limitations: what DLinear cannot do
The authors were transparent about DLinear's limitations. As a single-layer linear model, it has limited capacity — it cannot capture nonlinear dynamics, change points, or complex multi-scale interactions. It was designed as a baseline, not as the final answer to time series forecasting.
Later work, notably PatchTST, showed that Transformers can work for time series when the tokenization is redesigned. Instead of treating each time step as a (which fragments the temporal signal), PatchTST groups consecutive time steps into patches — similar to how Vision Transformers handle images. This preserves local temporal structure before self-attention operates, and PatchTST consistently outperformed DLinear.
The debate DLinear started is still ongoing: simplicity as a baseline forced the community to be more rigorous about what Transformers actually contribute, leading to better Transformer designs rather than abandoning them entirely.
Timeline: the great time series debate
2021
Informer — AAAI Best Paper
ProbSparse self-attention for long-term time series forecasting. Introduced DMS decoding to Transformers for time series, achieving state-of-the-art on ETT benchmarks.
2021
Autoformer — NeurIPS
Introduced seasonal-trend decomposition inside the Transformer and replaced self-attention with auto-correlation. The decomposition idea would later inspire DLinear.
2022
FEDformer — ICML
Frequency-enhanced decomposed Transformer. Used Fourier and wavelet transforms for O(L) complexity. Closest Transformer competitor to DLinear.
2023
DLinear — AAAI (this paper)
A single-layer linear model that outperformed all existing Transformers. Questioned the fundamental effectiveness of self-attention for temporal modeling.
2023
PatchTST — ICLR
Responded to DLinear by showing Transformers work when time steps are grouped into patches. Restored Transformer credibility for time series with proper tokenization.
CitationZeng, Chen, Zhang, Xu. Are Transformers Effective for Time Series Forecasting?. AAAI, 2023.
Terms in this paper
- Time Series Forecastingالتنبؤ بالسلاسل الزمنية
- Decompositionالتفكيك
- Direct Multi-Step Forecastingالتنبؤ المباشر متعدد الخطوات
- Iterated Multi-Step Forecastingالتنبؤ التكراري متعدد الخطوات
- Look-Back Windowنافذة المراجعة
- Trendاتجاه
- Moving Averageالمتوسط المتحرك
- Self-Attentionالانتباه الذاتي
- Mean Squared Error (MSE)متوسط الخطأ المربّع
- Mean Absolute Error (MAE)متوسط الخطأ المطلق
- Autoregressiveذاتي الانحدار
- Permutation-Invariantثابت أمام التبديل
- Multivariate Time Seriesسلسلة زمنية متعددة المتغيّرات