Time Series2024intermediate11 min read
A Decoder-Only Foundation Model for Time Series Forecasting
نموذج أساسي بفاكّ ترميز فقط للتنبؤ بالسلاسل الزمنية
Das, A. · Kong, W. · Sen, R. · Zhou, Y. — ICML
The problem
By 2023, large language models had demonstrated that a single pre-trained model can handle countless NLP tasks . But was still fragmented: every dataset needed its own custom-trained model. There was no vocabulary, no grammar, and not enough public data to train a . Worse, a forecasting model must handle variable context lengths, prediction horizons, and temporal granularities — from minutes to years — all at time.
The contribution
TimesFM: a 200M-parameter , pre-trained on a corpus of ~100 billion time-points mixing real-world data (Google Trends, Wikipedia page views) and synthetic patterns (ARMA, seasonal, trend). The model treats time-series segments as patches (analogous to tokens), uses causal with input length 32 and output patch length 128, and can forecast any horizon via auto-regressive decoding with fewer steps. Zero-shot performance on Monash, Darts, and ETT benchmarks matches or beats state-of-the-art supervised models trained specifically on each dataset — at a tiny fraction of LLM costs.
The impact
TimesFM demonstrated that the foundation-model paradigm scales to time-series, not just language and vision. It proved that a relatively small model (200M params), trained from scratch on time-series data alone, can vastly outperform repurposed LLMs like GPT-3 for forecasting. This opened the floodgates for time-series foundation models — Chronos, Moirai, and others followed — and established patching + -only as the standard recipe.
Think of a seasoned meteorologist who has studied thousands of weather charts from every continent: tropical monsoons, arctic cold fronts, desert heat waves. When handed a chart from a city they have never visited, they can still forecast the next few days — because they recognize the patterns, not the city.
Traditional forecasting models are like local weather stations: each one only knows its own city. TimesFM is the global meteorologist. It reads time-series the way GPT reads text — breaking them into chunks (patches), processing them through a causal Transformer, and predicting what comes next — except the "language" is the universal grammar of trends, seasons, and noise.
The problem: every dataset needs its own model
In NLP, GPT showed that a single model can handle translation, summarization, and question answering — all zero-shot. But in Forecasting, the picture was fragmented. You want to forecast retail demand? Train a DeepAR model on that retail data. Energy prices? Train a different model. Traffic flow? Another model entirely.
This creates three painful bottlenecks. First, training cost: every new forecasting task demands GPU time and engineering effort. Second, data hunger: many real-world time-series are short — quarterly or yearly data might have only a few dozen points — far too little for . Third, no transfer: a model trained on electricity data knows nothing about hospital admissions, even though both might share seasonal patterns.
The natural question is: can we build a foundation model for time-series the way GPT is a foundation model for text? The challenge is that time-series has no vocabulary, no grammar, and wildly varying granularities — from 10-minute sensor readings to annual GDP figures.
Key idea 1: patches are the tokens of time-series
In language models, text is split into tokens — subwords that the model processes one at a time. TimesFM does the same with time-series: it splits the input into non-overlapping patches of fixed length (32 time-points by default). Each patch is the time-series equivalent of a word.
Why patching? Three reasons. First, it compresses the sequence: a 512-point time-series becomes just 16 tokens, making attention much cheaper. Second, each patch captures local structure — a daily cycle, a weekly rhythm — that a single point cannot. Third, patching was already proven by PatchTST to improve forecasting accuracy.
Each patch is fed through an Input — a small MLP with a skip connection — to produce a vector of size model_dim (1280 in the 200M model). Positional encodings are added so the model knows the order of patches. The result is a sequence of vectors ready for the Transformer.
Key idea 2: decoder-only training for flexible forecasting
PatchTST uses an -only architecture: it sees all patches at once and outputs the entire forecast in one shot. This works when the prediction horizon is fixed, but a foundation model needs to handle any horizon at inference time.
TimesFM solves this by using a decoder-only architecture — the same paradigm as GPT-2. Each output token can only attend to input tokens that came before it (causal attention). This means the model learns to predict the next patch given all previous patches, and at inference time it can auto-regressively generate as many future patches as needed.
The critical difference from GPT is the output patch length. In language models, each step generates one token. TimesFM generates an output patch of 128 time-points per step — 4× the input patch length. This means forecasting 512 future points requires only 4 auto-regressive steps instead of 16. Fewer steps means less error accumulation and faster inference.
The architecture: patched decoder in detail
The full TimesFM architecture has three stages, like an assembly line:
Stage 1 — Input Processing: The raw time-series is split into patches of size 32. Each patch passes through an Input Residual Block (MLP with skip connection) to become a model_dim vector. Positional encodings are added. A padding mask handles variable-length inputs.
Stage 2 — Stacked Transformer: 20 transformer layers with causal multi-head (16 heads) and feed-forward networks. The model dimension is 1280. Each layer uses . This stack contains the bulk of the 200M parameters.
Stage 3 — Output Mapping: Each transformer output passes through an Output Residual Block that maps it from model_dim to an output patch of 128 time-points.
The training loss is simply averaged over all output patches in a mini-batch. No special tricks — just MSE on the predicted vs actual future values.
The fuel: 100 billion time-points
A foundation model is only as good as its pretraining data. TimesFM draws from three major sources:
Google Trends (~0.5B points): Search interest over time for ~22k head queries across hourly, daily, weekly, and monthly granularities from 2007 to 2022.
Wikipedia Pageviews (~300B points): Hourly views of all Wikimedia pages from 2012 to 2023, aggregated into hourly, daily, weekly, and monthly series. This is by far the largest source, providing the sheer volume needed for a foundation model.
(~6B points): 3 million generated time-series, each 2048 points long, combining ARMA processes, seasonal patterns (sines and cosines), trends (linear, exponential), and step functions. Synthetic data ensures the model sees granularities and patterns underrepresented in the real data.
Additional real-world datasets (M4, Electricity, Traffic, Weather) round out the mix. The training loader samples 80% real data and 20% synthetic, with equal weight given to different granularity groups.
The idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def patchify(series, patch_len=32):
"""Break a time-series into non-overlapping patches."""
n = len(series) // patch_len
patches = series[:n * patch_len].reshape(n, patch_len)
return patches
def input_residual_block(patch, W1, W2, b1, b2):
"""MLP with skip connection — maps patch → model_dim vector."""
h = np.maximum(0, patch @ W1 + b1) # ReLU hidden layer
out = h @ W2 + b2 # linear output
return out + patch @ W_skip # skip connection
def timesfm_forward(series, model, patch_len=32, output_len=128):
"""One forward pass of TimesFM."""
patches = patchify(series, patch_len)
# Stage 1: patches → input tokens
tokens = []
for j, patch in enumerate(patches):
token = input_residual_block(patch, *model.input_params)
token += model.pos_encoding[j] # add positional encoding
tokens.append(token)
# Stage 2: causal transformer — each token sees only past tokens
outputs = model.causal_transformer(tokens) # (N, model_dim)
# Stage 3: each output token → predict next output_len points
forecasts = []
for o in outputs:
pred = output_residual_block(o, *model.output_params)
forecasts.append(pred) # shape: (output_len,)
return forecasts # last forecast = the actual future prediction
# KEY INSIGHT: output_len (128) > patch_len (32)
# → fewer auto-regressive steps to cover the full horizon
# → less error accumulation than token-by-token generationResults: one model rivals the best specialists
TimesFM was evaluated zero-shot on three groups — datasets it had never seen during training:
Monash Archive (18 datasets): Covering domains from finance to weather to traffic, with granularities from minutes to years. TimesFM achieved the best geometric mean of scaled MAE, outperforming supervised baselines like DeepAR and N-BEATS, and beating llmtime (GPT-3 based) by more than 25%.
Darts (8 datasets): Univariate datasets with interesting seasonal patterns. TimesFM performed within statistical significance of the best models (seasonal ARIMA and llmtime), despite being fully zero-shot.
ETT (4 datasets, horizons 96 and 192): Electricity transformer temperature data. TimesFM matched or beat PatchTST — a state-of-the-art supervised long-horizon forecaster trained specifically on these datasets.
All results came from the same 200M model with no task-specific tuning. The model trained in just 2 days on 16 TPUv5e cores.
Ablation: what each design choice contributes
The paper includes careful ablation studies that justify the key design decisions:
Scaling (17M → 70M → 200M): Performance improves monotonically with compute (FLOPs), following a power-law pattern similar to LLM scaling laws. This suggests even larger models could be better.
Output patch length (8 → 128): Longer output patches consistently reduce error on long-horizon tasks. Going from output length 8 to 128 on ETT datasets shows a clear monotonic improvement, confirming that fewer auto-regressive steps help.
Input patch length (8 → 128): Sweet spot at 16–32. Too short (8) loses local structure; too long (128) shifts the model toward encoder-decoder territory, losing the benefits of decoder-only training.
Synthetic data: Removing synthetic data hurts performance on underrepresented granularities (quarterly, yearly, 10-minute). On well-represented hourly data (ETTh), there is almost no difference — but on 15-minute data (ETTm), synthetic data makes a significant difference.
What TimesFM unlocked
2022
PatchTST — Patching for time-series
Showed that treating time-series segments as patches and using a Transformer encoder achieves state-of-the-art long-horizon forecasting. Established patching as a key design principle.
2023
llmtime — LLMs as forecasters
Showed that GPT-3 can do zero-shot forecasting by encoding time-series values as text. Promising but expensive and often outperformed by smaller domain-specific models.
2024
TimesFM — First practical TS foundation model
200M parameter decoder-only model, pre-trained on 100B time-points. Zero-shot performance matches supervised SOTA across Monash, Darts, and ETT benchmarks.
2024
Chronos — Tokenizing time-series values
Amazon's approach: quantize time-series values into bins and train a T5-based model on the resulting tokens. Complementary to TimesFM's patching approach.
2024
Moirai & others — The foundation model wave
Multiple time-series foundation models emerge. Moirai (Salesforce) handles multivariate forecasting. The paradigm TimesFM helped establish is now standard.
TimesFM's legacy is the proof of concept: you can build a foundation model for time-series that generalizes across domains, granularities, and horizons. The recipe — patch the input, use a decoder-only Transformer, train on massive mixed data — is now the starting point for every new entry in this space. Just as BERT proved that pre-training works for NLP, TimesFM proved it works for forecasting.
CitationDas, Kong, Sen, Zhou. A Decoder-Only Foundation Model for Time Series Forecasting. ICML, 2024.
Terms in this paper
- Foundation Modelنموذج أساسي
- Zero-Shotالنمط الصفري
- Decoder-Onlyوحدة تفكيك ترميز فقط
- Patchرُقعة
- Patch Embeddingتضمين الرُّقع
- Autoregressiveذاتي الانحدار
- Time Series Forecastingالتنبؤ بالسلاسل الزمنية
- Causal Maskقناع سببي
- Positional Encodingالترميز الموضعي
- Residual Blockالكتلة المتبقّية
- Mean Absolute Errorمتوسط الخطأ المطلق
- Mean Squared Errorمتوسط مربع الخطأ
- Self-Attentionالانتباه الذاتي
- Feed Forward Network (FFN)شبكة التغذية الأمامية
- Multi-Head Attentionالانتباه المتعدد المسارات