Time Series2023intermediate12 min read
A Time Series Is Worth 64 Words: Long-Term Forecasting with Transformers
سلسلة زمنية تساوي 64 كلمة: التنبؤ طويل المدى باستخدام المحوِّلات
Nie, Y. · Nguyen, N. H. · Sinthong, P. · Kalagnanam, J. — ICLR
The problem
-based models for time series forecasting in 2022 faced two critical failures. First, they tokenized each time step as a single point, losing local temporal patterns and creating excessively long input sequences with quadratic cost. Second, they mixed all channels (variables) into a single at each time step, which caused on cross- noise rather than learning useful patterns. A simple one-layer linear model (DLinear) outperformed every Transformer variant — Informer, Autoformer, FEDformer — on standard benchmarks, raising the question: are Transformers even useful for time series?
The contribution
PatchTST: a channel-independent time series Transformer built on two design choices. (1) Patching — segmenting each univariate time series into subseries-level patches that serve as input tokens, retaining local semantics and quadratically reducing attention cost. (2) — processing each variable through a shared Transformer backbone separately rather than mixing all channels into one token. Together, these achieve a 21% reduction over the best Transformers, outperform DLinear on large datasets, and enable masked self-supervised pre- that further improves over supervised training.
The impact
PatchTST rescued Transformers for time series, proving that the right and channel design — not complex attention mechanisms — was the missing piece. Patching became the default input strategy for virtually all subsequent time series Transformers and foundation models, including TimesFM. The channel-independence insight spread to other architectures too. PatchTST positioned itself as a foundational building block for time series foundation models.
Imagine you're a speed reader trying to understand a long paragraph. Reading it letter by letter is agonizingly slow and the characters carry no meaning individually. Reading word by word is natural — each word is a semantic unit, and you can cover the whole paragraph faster. Now imagine you're a weather forecaster staring at a year of hourly temperature readings: 8,760 individual points. Feeding each point as a separate token to a Transformer is the letter-by-letter approach. PatchTST groups those points into patches — like words — so the model reads faster, sees more context, and actually understands local patterns like "this morning's gradual warmup."
The problem: point-wise tokens throw away local meaning
Before PatchTST, most Transformer-based time series models — Informer, Autoformer, FEDformer — treated each time step as a separate input token. This is equivalent to tokenizing English text one character at a time: it destroys local meaning.
A single temperature reading of 23.5°C carries almost no semantic information by itself. But a group of readings [22.1, 22.8, 23.2, 23.5] tells a story: the temperature is steadily rising. That local pattern — the trend, the shape — is exactly what point-wise tokenization loses.
Worse, the computational cost is devastating. The Transformer's has O(N²) time and memory complexity, where N is the number of input tokens. For a of 336 time steps, that's 336² = 112,896 attention entries. Training on the Traffic dataset (862 channels) with point-wise tokens took over 10,000 seconds — impractical for real-world deployment.
The coup de grâce came from Zeng et al. (2022): a single-layer linear model called DLinear outperformed every Transformer variant on common benchmarks. This wasn't a minor win — it was a fundamental challenge to the usefulness of Transformers for time series.
Key idea 1: Patching — from letters to words
The patching idea comes directly from . Vision Transformer (ViT) splits an image into 16×16 pixel patches and treats each patch as a token. PatchTST applies the same principle to time series: instead of one point per token, group P consecutive time steps into a single patch token.
Imagine a long ribbon of temperature readings. You cut this ribbon into segments of length P = 16, with a stride S = 8 (so patches overlap by 8 steps). Each segment becomes one input token to the Transformer. This has three immediate benefits:
1. Local semantics are preserved. Each patch captures a local shape — a rising trend, a periodic oscillation, a plateau. This is the "word" that carries meaning, versus the meaningless "letter" of a single data point.
2. Quadratic attention cost drops dramatically. With stride S = 8 and look-back window L = 336, the number of tokens drops from 336 to roughly 336/8 ≈ 42. The attention cost drops from 336² to 42² — a 64× reduction. On the Traffic dataset, training time dropped from 10,040 seconds to 464 seconds — a 22× speedup.
3. The model can see longer history. Because each token now summarizes P time steps, the same 42 tokens can cover 336 steps instead of just 42. You can even extend to 512 steps (64 patches) within the same computational budget, giving the model more context for long-term forecasting.
Key idea 2: Channel-independence — one instrument at a time
A multivariate time series has M channels (variables) — for example, 862 road sensors in the Traffic dataset, or 21 weather measurements (temperature, humidity, pressure...). Previous Transformer models used : at each time step, they stacked all M values into a single vector and projected it into the space. This forces the model to learn cross-channel correlations at every layer.
Think of it like an orchestra: channel-mixing forces all instruments to read from the same sheet of music, blending their sounds before understanding each one individually.
PatchTST takes the opposite approach: channel-independence. Each univariate series is fed through the Transformer backbone separately. The Transformer weights are shared across all channels (like a shared teacher), but each series generates its own attention maps. The final predictions from all channels are concatenated at the output.
Why does this work better? Three reasons:
- Adaptability: Different channels may have very different temporal patterns. In the Electricity dataset, some series are smooth and periodic while others are noisy. With channel-independence, each series learns its own attention pattern.
- Less overfitting: Channel-mixing models need to learn M×M cross-channel interactions, which requires much more data. Channel-independent models converge faster with less training data.
- Noise isolation: If one channel is noisy, channel-mixing projects that noise into all other channels via the embedding. Channel-independence keeps noise contained.
The architecture: vanilla Transformer, smart design
PatchTST's architecture is deliberately simple — it uses the vanilla Transformer with no custom attention mechanisms. The innovation is entirely in how the input is prepared (patching + channel-independence), not in the Transformer itself.
The forward process for each univariate series goes through these stages:
1. . Each series is normalized to zero mean and unit standard deviation. This handles between training and testing data. The mean and standard deviation are saved and added back to the output predictions.
2. Patching. The normalized series is divided into N overlapping patches of length P with stride S, creating a patch matrix of shape P × N.
3. Projection + Position Embedding. Each patch is linearly projected into the Transformer's latent space of dimension D via a trainable projection matrix. Learnable position embeddings are added to preserve temporal order.
4. Transformer Encoder. The projected patches pass through standard encoder layers: , , feed-forward network with GELU activation, and residual connections.
5. Flatten + Linear Head. The encoder output is flattened and a linear layer maps it to the prediction horizon T.
Default hyperparameters: 3 encoder layers, 16 attention heads, D = 128, dimension = 256, = 0.2, patch length P = 16, stride S = 8.
Self-supervised learning: mask patches, not points
PatchTST isn't just a supervised forecasting model — it also supports self-supervised pre-training via masked patch reconstruction, inspired by masked autoencoders in computer vision and masked language modeling in NLP.
The idea is straightforward: randomly select 40% of the non-overlapping patches and replace them with zeros. The model is trained to reconstruct the original values of those masked patches using MSE loss. Because entire patches are masked (not individual points), the model can't simply interpolate from neighboring time steps — it must learn higher-level temporal representations.
Why is this important? Think of it as reading a paragraph with entire words blacked out rather than individual letters. You need to understand sentence structure and context to fill in the missing words — mere letter-level guessing won't work.
After pre-training, the model can be fine-tuned for forecasting in two ways:
- : Freeze the encoder, train only the prediction head for 20 epochs. Already competitive with supervised training from scratch.
- End-to-end : First linear probing for 10 epochs, then unfreeze everything for 20 more epochs. This achieves the best results — even beating supervised training on large datasets like Weather, Traffic, and Electricity.
Crucially, channel-independence makes natural: because each channel is processed independently with shared weights, you can pre-train on a dataset with 321 channels (Electricity) and fine-tune on one with 21 channels (Weather) — the number of variables doesn't need to match.
The idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def create_patches(series, patch_len=16, stride=8):
"""Segment a univariate time series into overlapping patches."""
L = len(series)
# Pad the end by repeating the last value
padded = np.concatenate([series, np.full(stride, series[-1])])
patches = []
for start in range(0, L - patch_len + stride + 1, stride):
patches.append(padded[start : start + patch_len])
return np.array(patches) # shape: (N, P)
def patchtst_forward(multivariate_series, transformer, proj, head):
"""Channel-independent PatchTST: each channel goes through separately."""
M, L = multivariate_series.shape # M channels, L time steps
predictions = []
for i in range(M):
x = multivariate_series[i] # one univariate series
mean, std = x.mean(), x.std()
x = (x - mean) / (std + 1e-8) # instance normalization
patches = create_patches(x) # (N, P) patch tokens
tokens = proj(patches) # (N, D) projected to latent space
tokens += pos_embedding # add learned positional encoding
z = transformer.encode(tokens) # (N, D) — vanilla encoder
pred = head(z.flatten()) # (T,) — forecast T steps ahead
pred = pred * std + mean # undo normalization
predictions.append(pred)
return np.stack(predictions) # (M, T) — all channelsResults: Transformers work when you tokenize right
PatchTST was evaluated on 8 popular benchmarks: Weather, Traffic, Electricity, ILI, and the four ETT datasets. Two variants were tested: PatchTST/42 (look-back L = 336, 42 patches — fair comparison to DLinear) and PatchTST/64 (L = 512, 64 patches — pushing performance further).
Against Transformer baselines: PatchTST/64 achieved an overall 21.0% reduction in MSE and 16.7% reduction in MAE compared to the best Transformer-based models (FEDformer, Autoformer, Informer). This is an enormous margin.
Against DLinear: PatchTST outperformed DLinear on the large datasets — Weather, Traffic, and Electricity — which are the most reliable benchmarks because they have enough data to resist overfitting. On Weather with prediction horizon 96, PatchTST/64 achieved MSE 0.149 vs DLinear's 0.176. On Traffic at horizon 96: 0.360 vs 0.410.
Self-supervised boost: On large datasets, masked pre-training followed by fine-tuning beat supervised training from scratch. On Electricity at horizon 96: self-supervised MSE 0.126 vs supervised 0.130.
Transfer learning: Pre-training on Electricity and fine-tuning on Weather achieved MSE 0.145 at horizon 96 — better than supervised training from scratch (0.152) despite the domains being completely different.
comparison: Against dedicated self-supervised methods (TS2Vec, TNC, TS-TCC, BTSF), PatchTST achieved 34–49% improvement in MSE on ETTh1.
Ablation: every piece matters
The paper systematically tested every design choice. Removing patching while keeping channel-independence raised MSE on Weather from 0.152 to 0.164. Removing channel-independence while keeping patching raised it to 0.168. Removing both (original TST model) raised it to 0.177. Both pieces matter, and their combination is better than either alone.
Varying patch length: MSE remained stable across patch lengths from 4 to 40, showing robustness. Lengths between 8 and 16 were generally optimal.
Varying look-back window: PatchTST consistently improved with longer windows — unlike FEDformer, Autoformer, and Informer, which often degraded or plateaued when the look-back window exceeded 96. This is direct evidence that patching allows the model to absorb more historical information.
Instance normalization: Removing it degraded performance, especially on datasets with distribution shift (ILI, Traffic). But even without it, PatchTST still outperformed other Transformers — confirming that the gains come primarily from patching and channel-independence.
What PatchTST influenced
2017
Transformer (Attention Is All You Need)
The original Transformer architecture. Used for machine translation, later adapted for every modality.
2021
ViT — Vision Transformer
Introduced 16×16 image patches as tokens. Proved patching works for Transformers beyond NLP, inspiring PatchTST.
2021
Informer
ProbSparse attention for efficient long-term forecasting. Still used point-wise tokens and channel-mixing.
2022
DLinear — Are Transformers Effective?
A simple linear model that outperformed all Transformers. Challenged the community to rethink Transformer design for time series.
2023
PatchTST (this paper)
Patching + channel-independence rescued Transformers for time series. 21% MSE reduction over the best Transformers. Enabled self-supervised pre-training.
2024
TimesFM
Google's time series foundation model. Built on patching as default tokenization, directly influenced by PatchTST's design.
CitationNie, Nguyen, Sinthong, Kalagnanam. A Time Series Is Worth 64 Words: Long-term Forecasting with Transformers. ICLR, 2023.
Terms in this paper
- Patchرُقعة
- Patch Embeddingتضمين الرُّقع
- Channelالقناة البنيوية
- Instance Normalizationتسوية النسخة
- Transfer Learningنقل التعلم
- Masked Autoencoderمُرمِّز تلقائي مُقنَّع
- Multi-Head Attentionالانتباه المتعدد المسارات
- Transformerالمحوِّل
- Mean Squared Error (MSE)متوسط الخطأ المربّع
- Self-Supervised Learningالتعلم ذاتي الإشراف
- Representation Learningتعلم التمثيلات الرقمية
- Positional Embeddingالتضمين الموضعي
- Linear Probingالاختبار الخطي
- Fine-Tuningالضبط الدقيق
- Attentionآلية الانتباه