Time Series2024intermediate11 min read

A Decoder-Only Foundation Model for Time Series Forecasting

نموذج أساسي بفاكّ ترميز فقط للتنبؤ بالسلاسل الزمنية

Das, A. · Kong, W. · Sen, R. · Zhou, Y. — ICML

The problem

By 2023, large language models had demonstrated that a single pre-trained model can handle countless NLP tasks . But was still fragmented: every dataset needed its own custom-trained model. There was no vocabulary, no grammar, and not enough public data to train a . Worse, a forecasting model must handle variable context lengths, prediction horizons, and temporal granularities — from minutes to years — all at time.

The contribution

TimesFM: a 200M-parameter , pre-trained on a corpus of ~100 billion time-points mixing real-world data (Google Trends, Wikipedia page views) and synthetic patterns (ARMA, seasonal, trend). The model treats time-series segments as patches (analogous to tokens), uses causal with input length 32 and output patch length 128, and can forecast any horizon via auto-regressive decoding with fewer steps. Zero-shot performance on Monash, Darts, and ETT benchmarks matches or beats state-of-the-art supervised models trained specifically on each dataset — at a tiny fraction of LLM costs.

The impact

TimesFM demonstrated that the foundation-model paradigm scales to time-series, not just language and vision. It proved that a relatively small model (200M params), trained from scratch on time-series data alone, can vastly outperform repurposed LLMs like GPT-3 for forecasting. This opened the floodgates for time-series foundation models — Chronos, Moirai, and others followed — and established patching + -only as the standard recipe.

Think of a seasoned meteorologist who has studied thousands of weather charts from every continent: tropical monsoons, arctic cold fronts, desert heat waves. When handed a chart from a city they have never visited, they can still forecast the next few days — because they recognize the patterns, not the city.

Traditional forecasting models are like local weather stations: each one only knows its own city. TimesFM is the global meteorologist. It reads time-series the way GPT reads text — breaking them into chunks (patches), processing them through a causal Transformer, and predicting what comes next — except the "language" is the universal grammar of trends, seasons, and noise.

The problem: every dataset needs its own model

In NLP, GPT showed that a single model can handle translation, summarization, and question answering — all zero-shot. But in Forecasting, the picture was fragmented. You want to forecast retail demand? Train a DeepAR model on that retail data. Energy prices? Train a different model. Traffic flow? Another model entirely.

This creates three painful bottlenecks. First, training cost: every new forecasting task demands GPU time and engineering effort. Second, data hunger: many real-world time-series are short — quarterly or yearly data might have only a few dozen points — far too little for . Third, no transfer: a model trained on electricity data knows nothing about hospital admissions, even though both might share seasonal patterns.

The natural question is: can we build a foundation model for time-series the way GPT is a foundation model for text? The challenge is that time-series has no vocabulary, no grammar, and wildly varying granularities — from 10-minute sensor readings to annual GDP figures.

Open in Lab
Compare how text and time-series are broken into processable units.
The demo wakes as you arrive…

Key idea 1: patches are the tokens of time-series

In language models, text is split into tokens — subwords that the model processes one at a time. TimesFM does the same with time-series: it splits the input into non-overlapping patches of fixed length (32 time-points by default). Each patch is the time-series equivalent of a word.

Why patching? Three reasons. First, it compresses the sequence: a 512-point time-series becomes just 16 tokens, making attention much cheaper. Second, each patch captures local structure — a daily cycle, a weekly rhythm — that a single point cannot. Third, patching was already proven by PatchTST to improve forecasting accuracy.

Each patch is fed through an Input — a small MLP with a skip connection — to produce a vector of size model_dim (1280 in the 200M model). Positional encodings are added so the model knows the order of patches. The result is a sequence of vectors ready for the Transformer.

tj=InputResidualBlock(y~j⊙(1−m~j))+PEjt_j = \text{InputResidualBlock}(\tilde{y}_j \odot (1 - \tilde{m}_j)) + PE_j
Input token construction — patch + masking + position — Each patch ỹⱼ is element-wise multiplied with the inverted padding mask (to zero out padded positions), passed through a Residual Block to produce a model_dim vector, then summed with the positional encoding PEⱼ.
Open in Lab
Drag the slider to change patch length and see how the time-series is segmented.
The demo wakes as you arrive…

Key idea 2: decoder-only training for flexible forecasting

PatchTST uses an -only architecture: it sees all patches at once and outputs the entire forecast in one shot. This works when the prediction horizon is fixed, but a foundation model needs to handle any horizon at inference time.

TimesFM solves this by using a decoder-only architecture — the same paradigm as GPT-2. Each output token can only attend to input tokens that came before it (causal attention). This means the model learns to predict the next patch given all previous patches, and at inference time it can auto-regressively generate as many future patches as needed.

The critical difference from GPT is the output patch length. In language models, each step generates one token. TimesFM generates an output patch of 128 time-points per step — 4× the input patch length. This means forecasting 512 future points requires only 4 auto-regressive steps instead of 16. Fewer steps means less error accumulation and faster inference.

Open in Lab
Compare auto-regressive steps needed for different output patch lengths.
The demo wakes as you arrive…
oj=StackedTransformer((t1,m˙1),…,(tj,m˙j))o_j = \text{StackedTransformer}\bigl((t_1, \dot{m}_1), \ldots, (t_j, \dot{m}_j)\bigr)
Causal transformer output — each patch sees only past patches — The stacked transformer processes all input tokens using causal self-attention. Output token oⱼ encodes all information from patches 1 through j, respecting the causal mask.
y^pj+1:pj+h=OutputResidualBlock(oj)\hat{y}_{pj+1:pj+h} = \text{OutputResidualBlock}(o_j)
Output mapping — each transformer output predicts an output patch — Output token oⱼ is passed through another Residual Block to produce h = 128 future time-points. This is the key difference from LLMs where output = 1 token.

The architecture: patched decoder in detail

The full TimesFM architecture has three stages, like an assembly line:

Stage 1 — Input Processing: The raw time-series is split into patches of size 32. Each patch passes through an Input Residual Block (MLP with skip connection) to become a model_dim vector. Positional encodings are added. A padding mask handles variable-length inputs.

Stage 2 — Stacked Transformer: 20 transformer layers with causal multi-head (16 heads) and feed-forward networks. The model dimension is 1280. Each layer uses . This stack contains the bulk of the 200M parameters.

Stage 3 — Output Mapping: Each transformer output passes through an Output Residual Block that maps it from model_dim to an output patch of 128 time-points.

The training loss is simply averaged over all output patches in a mini-batch. No special tricks — just MSE on the predicted vs actual future values.

Open in Lab
Explore the three stages of TimesFM architecture interactively.
The demo wakes as you arrive…
L=1N∑j=1NMSE(y^pj+1:pj+h,  ypj+1:pj+h)\mathcal{L} = \frac{1}{N} \sum_{j=1}^{N} \text{MSE}\bigl(\hat{y}_{pj+1:pj+h},\; y_{pj+1:pj+h}\bigr)
Training loss — MSE averaged over all output patches — For each of the N input patches, the model predicts h = 128 future points. The loss is the mean of MSE across all patches in the batch. For probabilistic forecasting, quantile heads can replace MSE — but the paper focuses on point forecasts.

The fuel: 100 billion time-points

A foundation model is only as good as its pretraining data. TimesFM draws from three major sources:

Google Trends (~0.5B points): Search interest over time for ~22k head queries across hourly, daily, weekly, and monthly granularities from 2007 to 2022.

Wikipedia Pageviews (~300B points): Hourly views of all Wikimedia pages from 2012 to 2023, aggregated into hourly, daily, weekly, and monthly series. This is by far the largest source, providing the sheer volume needed for a foundation model.

(~6B points): 3 million generated time-series, each 2048 points long, combining ARMA processes, seasonal patterns (sines and cosines), trends (linear, exponential), and step functions. Synthetic data ensures the model sees granularities and patterns underrepresented in the real data.

Additional real-world datasets (M4, Electricity, Traffic, Weather) round out the mix. The training loader samples 80% real data and 20% synthetic, with equal weight given to different granularity groups.

Open in Lab
Explore the composition and scale of the TimesFM pretraining corpus.
The demo wakes as you arrive…

The idea in code

TimesFM core logic — patching, causal transformer, output predictionpython

Simplified to show the idea — not the real implementation.

import numpy as np

def patchify(series, patch_len=32):
    """Break a time-series into non-overlapping patches."""
    n = len(series) // patch_len
    patches = series[:n * patch_len].reshape(n, patch_len)
    return patches

def input_residual_block(patch, W1, W2, b1, b2):
    """MLP with skip connection — maps patch → model_dim vector."""
    h = np.maximum(0, patch @ W1 + b1)  # ReLU hidden layer
    out = h @ W2 + b2                     # linear output
    return out + patch @ W_skip            # skip connection

def timesfm_forward(series, model, patch_len=32, output_len=128):
    """One forward pass of TimesFM."""
    patches = patchify(series, patch_len)

    # Stage 1: patches → input tokens
    tokens = []
    for j, patch in enumerate(patches):
        token = input_residual_block(patch, *model.input_params)
        token += model.pos_encoding[j]     # add positional encoding
        tokens.append(token)

    # Stage 2: causal transformer — each token sees only past tokens
    outputs = model.causal_transformer(tokens)  # (N, model_dim)

    # Stage 3: each output token → predict next output_len points
    forecasts = []
    for o in outputs:
        pred = output_residual_block(o, *model.output_params)
        forecasts.append(pred)              # shape: (output_len,)

    return forecasts  # last forecast = the actual future prediction

# KEY INSIGHT: output_len (128) > patch_len (32)
# → fewer auto-regressive steps to cover the full horizon
# → less error accumulation than token-by-token generation

Results: one model rivals the best specialists

TimesFM was evaluated zero-shot on three groups — datasets it had never seen during training:

Monash Archive (18 datasets): Covering domains from finance to weather to traffic, with granularities from minutes to years. TimesFM achieved the best geometric mean of scaled MAE, outperforming supervised baselines like DeepAR and N-BEATS, and beating llmtime (GPT-3 based) by more than 25%.

Darts (8 datasets): Univariate datasets with interesting seasonal patterns. TimesFM performed within statistical significance of the best models (seasonal ARIMA and llmtime), despite being fully zero-shot.

ETT (4 datasets, horizons 96 and 192): Electricity transformer temperature data. TimesFM matched or beat PatchTST — a state-of-the-art supervised long-horizon forecaster trained specifically on these datasets.

All results came from the same 200M model with no task-specific tuning. The model trained in just 2 days on 16 TPUv5e cores.

Open in Lab
Compare TimesFM zero-shot performance against supervised baselines.
The demo wakes as you arrive…

Ablation: what each design choice contributes

The paper includes careful ablation studies that justify the key design decisions:

Scaling (17M → 70M → 200M): Performance improves monotonically with compute (FLOPs), following a power-law pattern similar to LLM scaling laws. This suggests even larger models could be better.

Output patch length (8 → 128): Longer output patches consistently reduce error on long-horizon tasks. Going from output length 8 to 128 on ETT datasets shows a clear monotonic improvement, confirming that fewer auto-regressive steps help.

Input patch length (8 → 128): Sweet spot at 16–32. Too short (8) loses local structure; too long (128) shifts the model toward encoder-decoder territory, losing the benefits of decoder-only training.

Synthetic data: Removing synthetic data hurts performance on underrepresented granularities (quarterly, yearly, 10-minute). On well-represented hourly data (ETTh), there is almost no difference — but on 15-minute data (ETTm), synthetic data makes a significant difference.

What TimesFM unlocked

  1. 2022

    PatchTST — Patching for time-series

    Showed that treating time-series segments as patches and using a Transformer encoder achieves state-of-the-art long-horizon forecasting. Established patching as a key design principle.

  2. 2023

    llmtime — LLMs as forecasters

    Showed that GPT-3 can do zero-shot forecasting by encoding time-series values as text. Promising but expensive and often outperformed by smaller domain-specific models.

  3. 2024

    TimesFM — First practical TS foundation model

    200M parameter decoder-only model, pre-trained on 100B time-points. Zero-shot performance matches supervised SOTA across Monash, Darts, and ETT benchmarks.

  4. 2024

    Chronos — Tokenizing time-series values

    Amazon's approach: quantize time-series values into bins and train a T5-based model on the resulting tokens. Complementary to TimesFM's patching approach.

  5. 2024

    Moirai & others — The foundation model wave

    Multiple time-series foundation models emerge. Moirai (Salesforce) handles multivariate forecasting. The paradigm TimesFM helped establish is now standard.

TimesFM's legacy is the proof of concept: you can build a foundation model for time-series that generalizes across domains, granularities, and horizons. The recipe — patch the input, use a decoder-only Transformer, train on massive mixed data — is now the starting point for every new entry in this space. Just as BERT proved that pre-training works for NLP, TimesFM proved it works for forecasting.

CitationDas, Kong, Sen, Zhou. A Decoder-Only Foundation Model for Time Series Forecasting. ICML, 2024.

Terms in this paper