Time Series2024intermediate10 min read

Chronos: Learning the Language of Time Series

Chronos: تعلُّم لغة السلاسل الزمنية

Ansari, A.F. · Stella, L. · Turkmen, C. · Zhang, X. · Mercado, P. · Shen, H. · Shchur, O. · Rangapuram, S.S. · Pineda Arango, S. · Kapoor, S. · Zschiegner, J. · Maddix, D.C. · Wang, H. · Mahoney, M.W. · Torkkola, K. · Wilson, A.G. · Bohlke-Schneider, M. · Wang, Y. — TMLR

The problem

forecasting traditionally requires a separate model for each dataset. forecasters improved accuracy by learning across many series in one dataset, but they still could not generalize to unseen datasets. Meanwhile, large language models showed remarkable abilities on text, but adapting them to time series either required heavy prompt engineering, per-task , or billion-parameter models too expensive for practical use. The field lacked a simple, scalable framework that could pretrain on diverse time series and forecast new ones out of the box.

The contribution

Chronos: a framework that tokenizes real-valued time series via scaling and into a fixed , then trains standard T5 language models (20M–710M parameters) on these tokens using the . No time-series-specific architecture changes are needed. The authors also introduce KernelSynth, a method that generates synthetic training data from Gaussian processes, and TSMixup for . On a 42-dataset benchmark, Chronos significantly outperforms baselines in-domain and achieves comparable or superior zero-shot performance on unseen datasets versus models trained specifically on them.

The impact

Chronos demonstrated that time series forecasting can be recast as a language modeling problem with minimal adaptation. It validated that pretrained language model architectures — without any time-series-specific design — can achieve state-of-the-art forecasting when trained on tokenized time series at scale. This opened the door for the forecasting community to leverage the entire LLM toolchain — scaling laws, efficient , and continual — and inspired a wave of foundation models for time series.

Imagine a translator who has read thousands of books in every language. Give them a book in a new language they've never formally studied, and they can still guess the next sentence — because patterns of storytelling are universal. Chronos does the same for time series: it translates numbers into "words," reads millions of different time series "stories" — sales, weather, energy — and learns that patterns of change are universal too. Hand it a brand-new series it has never seen, and it continues the story.

The core idea: treat numbers as words

Language models predict the next from a finite vocabulary. Time series values, on the other hand, are continuous real numbers — there is no natural "dictionary" for them. Chronos bridges this gap with a beautifully simple pipeline.

First, each time series is scaled by dividing every value by the mean of the absolute values in the historical context. This normalization — called — maps series of vastly different magnitudes into a comparable range. A retail series with sales in the thousands and a temperature series around 20°C are brought to the same neighborhood.

Second, the scaled values are quantized into a fixed number of bins. Think of a number line divided into equally-spaced buckets: each real value falls into one bucket, and that bucket's index becomes the token. The original T5 vocabulary is replaced entirely by these bin indices plus a few special tokens for padding and missing values.

The result: a time series becomes a sequence of integer tokens — exactly the format a language model expects. No patching, no heads, no custom architecture. Just tokens fed into a standard T5 -.

Open in Lab
Drag the slider to see how real-valued time series data is scaled and quantized into discrete tokens.
The demo wakes as you arrive…

Scaling: speaking the same dialect

Different time series live at wildly different scales. A series of stock prices might range from 100 to 300, while a sensor reading might hover between 0.001 and 0.01. If we quantize them directly, the stock prices would saturate the upper bins while the sensor readings would all land in a single bin — the model would see entirely different "dialects."

Mean scaling solves this by dividing each value by the mean absolute value of the historical context. Formally, given a context window of length C, the scale factor is the average of the absolute values. This centers the typical magnitude around 1.0, and critically, it preserves zeros — if sales are zero on a holiday, the tokenized value remains zero. This matters because zeros in time series often carry semantic meaning, unlike in natural language.

x~i=xis,s=1C∑j=1C∣xj∣\tilde{x}_i = \frac{x_i}{s}, \quad s = \frac{1}{C}\sum_{j=1}^{C}|x_j|
Mean Scaling — Each value is divided by the mean of absolute values in the context window, normalizing the series to a comparable range while preserving zeros.

Quantization: building a vocabulary for numbers

After scaling, each value is still a real number. The quantization step maps it to one of B bin indices. Chronos places B bin centers along the real line with B−1 edges separating them. Each scaled value falls into the bin whose boundaries contain it, and the bin index becomes the token.

Think of a piano keyboard: continuous sound frequencies are discretized into 88 keys. You lose some precision — a pitch between C and C♯ maps to whichever key is closer — but you gain the ability to write sheet music: a symbolic, discrete representation that a can natively process.

The default vocabulary size in Chronos is 4096 bins. The authors found this sweet spot through ablation: too few bins lose precision, too many waste capacity on rarely-used tokens. The bin edges follow the quantiles of a standard normal distribution, which concentrates resolution where most scaled values lie (near zero) and spaces bins more widely in the tails.

q(x~)=kif bk−1≤x~<bk,d(k)=ckq(\tilde{x}) = k \quad \text{if } b_{k-1} \leq \tilde{x} < b_k, \qquad d(k) = c_k
Quantization and Dequantization — The quantization function q maps a scaled value to a bin index; the dequantization function d maps a bin index back to its center value.
Open in Lab
Click on bins to see how continuous values map to discrete tokens and back.
The demo wakes as you arrive…

Training and inference: the language of change

Once tokenized, a time series looks exactly like a sentence. Chronos uses the T5 encoder-decoder architecture and trains it with the standard cross-entropy loss — the same objective used for machine translation. The encoder reads the tokenized context (the historical values), and the decoder predicts the tokenized future one token at a time.

At inference, Chronos generates forecasts autoregressively: it samples a token from the predicted categorical distribution, maps it back to a real value via dequantization and rescaling, then feeds it back as input for the next step. By sampling multiple trajectories (typically 20), Chronos produces a full predictive distribution — not just a point forecast. This gives users quantiles, confidence intervals, and a sense of uncertainty, which is critical for decision-making.

This design choice — modeling real values via a categorical distribution over bins — is sometimes called "regression via ." The key advantage is flexibility: the model can represent arbitrary predictive distributions, including multimodal ones, without any parametric assumptions. A Gaussian head can only produce bell curves; Chronos can produce any shape.

L=−∑t=C+1C+Hlog⁡pθ(zt∣z1:C)\mathcal{L} = -\sum_{t=C+1}^{C+H} \log p_\theta(z_t \mid \mathbf{z}_{1:C})
Cross-Entropy Training Loss — The model is trained to maximize the probability of the correct token at each future time step, conditioned on the tokenized context.
Open in Lab
Follow the full Chronos pipeline from raw time series through tokenization, model processing, and probabilistic output.
The demo wakes as you arrive…

Data augmentation: inventing new stories

A language model needs billions of sentences to learn grammar. Time series datasets are far smaller — even a comprehensive collection of public datasets provides only tens of thousands of individual series. Chronos addresses this data scarcity with two augmentation strategies.

TSMixup randomly selects a small set of real time series from different datasets and creates new series as convex combinations of them. Imagine blending the sales curve of a bookshop with the electricity load of a factory — the result is a synthetic series that inherits patterns from both, pushing the model to learn more general features rather than memorizing specific datasets.

KernelSynth takes a more principled approach: it generates entirely synthetic time series from Gaussian processes. A is defined by a kernel function that controls the shape of the patterns it produces. KernelSynth randomly composes basic kernels — linear, periodic, radial basis function — using addition and multiplication, then samples time series from the resulting Gaussian process. This produces an unlimited supply of diverse, realistic-looking series with known statistical properties.

In experiments, adding KernelSynth data significantly improved zero-shot performance on unseen datasets, validating that synthetic diversity generalizes better than more real data of limited variety.

Open in Lab
Explore how composing different kernel functions generates diverse synthetic time series.
The demo wakes as you arrive…

Results: what Chronos achieves

The evaluation covered 42 datasets spanning energy, finance, healthcare, nature, retail, transport, and web traffic. The authors split these into two benchmarks: Benchmark I used 15 datasets from the training corpus (in-domain), and Benchmark II used 27 held-out datasets (zero-shot).

On in-domain data, the largest Chronos model (710M parameters, based on T5-Large) significantly outperformed all baselines, including both classical statistical models (ARIMA, ETS, Theta) and deep learning methods (DeepAR, PatchTST, TFT). The gains were especially large on probabilistic metrics like weighted quantile loss.

On zero-shot data — datasets the model had never seen during training — Chronos performed comparably or better than models that were specifically trained on those datasets. This is the key finding: a model trained only on tokenized time series, with no task-specific heads or features, generalizes to new domains.

Scaling analysis showed clear improvements from 20M to 710M parameters, suggesting that the language modeling scaling laws apply to tokenized time series as well.

Open in Lab
Compare Chronos model sizes against baselines across in-domain and zero-shot benchmarks.
The demo wakes as you arrive…

Key design decisions and ablations

The paper includes a thorough exploring several design axes. Model size matters: performance improves consistently from Chronos-Mini (20M) to Chronos-Large (710M). Initialization from pretrained T5 weights helps compared to random initialization, even though the original T5 vocabulary is completely replaced — the patterns and positional encodings transfer useful priors.

Vocabulary size of 4096 bins hits the sweet spot: smaller vocabularies (e.g., 64) lack resolution for fine-grained patterns, while much larger ones (e.g., 16384) offer diminishing returns. The proportion of also matters: a mix of roughly 50% real and 50% KernelSynth data maximizes zero-shot performance.

The authors note several limitations. Chronos processes each time step as one token, which is less efficient than patching approaches that group multiple steps. It currently handles only univariate forecasting. And because quantization has finite resolution, very small differences between values can be lost — a form of information inherent to the approach.

Historical context

  1. 2017

    Transformer architecture

    Vaswani et al. introduced the Transformer for machine translation, establishing self-attention as the dominant sequence modeling mechanism.

  2. 2020

    T5 — Text-to-Text Transfer Transformer

    Raffel et al. unified NLP tasks into a text-to-text framework. T5's encoder-decoder became the architectural backbone for Chronos.

  3. 2020

    DeepAR — Probabilistic forecasting with RNNs

    Salinas et al. showed that autoregressive RNN models trained across many time series can produce calibrated probabilistic forecasts.

  4. 2023

    PatchTST — Patching for efficient time series transformers

    Nie et al. proposed patching time series into sub-sequences, treating each patch as a token, dramatically reducing sequence length and improving transformer efficiency.

  5. 2023

    LLMTime — Zero-shot forecasting with GPT-3

    Gruver et al. showed that encoding time series as digit strings lets pretrained LLMs forecast without any training, but required billion-parameter models.

  6. 2024

    Chronos — This paper

    Ansari et al. demonstrated that small-to-medium language models (20M–710M), trained from scratch on tokenized time series, achieve state-of-the-art zero-shot forecasting without any time-series-specific architecture changes.

Chronos represents a philosophical shift in time series forecasting. Rather than engineering specialized architectures with domain-specific features — lag indices, seasonal components, trend extractors — it asks: can a general-purpose sequence model learn these patterns from data if we just give it the right representation? The answer turns out to be yes, and this aligns the forecasting field with the broader trend in AI: scale and generality win over specialization.

CitationAnsari, Stella, Turkmen, Zhang, Mercado, Shen, Shchur, Rangapuram, Pineda Arango, Kapoor, Zschiegner, Maddix, Wang, Mahoney, Torkkola, Wilson, Bohlke-Schneider, Wang. Chronos: Learning the Language of Time Series. TMLR, 2024.

Terms in this paper