Time Series2020intermediate12 min read
DeepAR: Probabilistic Forecasting with Autoregressive Recurrent Networks
DeepAR: التنبؤ الاحتمالي باستخدام الشبكات التكرارية الذاتية الانحدار
Salinas, D. · Flunkert, V. · Gasthaus, J. · Januschowski, T. — International Journal of Forecasting
The problem
Classical methods like ARIMA and Exponential Smoothing fit one model per time series. When you have thousands or millions of related series — product demands, server loads, energy readings — fitting individual models is slow, misses cross-series patterns, and fails entirely for new items with no history (the cold-start problem). Most classical methods also produce only point forecasts, not the full probability distribution needed for risk-sensitive decisions like inventory planning or capacity allocation.
The contribution
DeepAR: a global that learns from many related time series simultaneously. It uses an to map historical observations and covariates to the parameters of a chosen function (Gaussian for real-valued data, negative binomial for count data). During it generates sample paths to produce coherent probabilistic forecasts — full predictive distributions, not just point estimates. The model handles widely varying scales via per-series normalization and addresses cold-start through learned item embeddings. Empirical results show ~15% accuracy improvement over state-of-the-art methods across multiple real-world datasets.
The impact
DeepAR pioneered the "global model" paradigm for time series: one neural network trained across thousands of series, replacing the one-model-per-series tradition. It became the backbone of Amazon's demand forecasting system (Amazon Forecast) and inspired a generation of neural forecasting methods. Its autoregressive sampling strategy for coherent probabilistic forecasts influenced successors like Chronos, Informer, and the broader GluonTS ecosystem. DeepAR showed that deep learning could be practical for production forecasting — not just an academic curiosity.
Imagine a weather forecaster who has studied every city's weather for 20 years. When a brand-new town asks for tomorrow's forecast, the forecaster doesn't panic — they recognize patterns: "this looks like coastal weather in spring." They don't give a single temperature; they say "expect 18–24°C, most likely 21°C."
Classical forecasting is like hiring a separate forecaster per city, each studying only their own records. DeepAR is the seasoned forecaster who learns from all cities at once, spots shared patterns, handles newcomers gracefully, and always gives a range — not just a number.
The problem: one model per series does not scale
A large retailer may track demand for 500,000 products across 1,000 stores. Classical methods like ARIMA or Exponential Smoothing fit a separate model to each individual time series. This creates three serious problems:
Scale: fitting and maintaining half a million models is operationally expensive.
Sparsity: many products have short or intermittent sales histories, giving individual models too little data to learn meaningful patterns.
: a brand-new product has zero history. Individual models literally cannot start.
And even when classical models work, they typically output a single predicted value — a point forecast. But business decisions need uncertainty: should I order 100 or 200 units? The answer depends on how confident the forecast is. What's needed is a full probability distribution over future values.
The idea: learn one model across all series, output distributions
DeepAR's two key insights are:
1. Global model: Instead of fitting one model per time series, train a single recurrent neural network on all related time series simultaneously. The network learns shared patterns — seasonality, trend shapes, effects — that transfer across series. Series-specific behavior is captured through learned item embeddings and per-series scale normalization.
2. Probabilistic output: Instead of predicting a single number, the network outputs the parameters of a probability distribution at each time step. For real-valued data, it outputs the mean μ and standard deviation σ of a Gaussian. For count data, it outputs the parameters of a negative binomial distribution. This means DeepAR produces a full predictive distribution, not just a point estimate.
The architecture is autoregressive: at each time step, the model conditions on the previous observation and covariates, feeds them through an LSTM, and outputs distribution parameters. Think of it as a chain: each link (time step) uses the previous link's output to predict the next.
The architecture: LSTM meets likelihood
At each time step t, DeepAR performs four operations:
Step 1 — Gather inputs: Collect the previous observation z(t−1), the covariates x(t) (time features like day-of-week, month, holidays, plus any item-specific covariates), and the previous h(t−1).
Step 2 — LSTM : Feed these inputs into a multi-layer LSTM, which produces a new hidden state h(t). The hidden state is the network's compressed memory — like a summary of everything it has seen so far for this series.
Step 3 — Distribution parameters: Pass h(t) through two fully connected layers to produce distribution parameters. For a Gaussian likelihood: one layer outputs μ(t) (the predicted mean) and another outputs σ(t) (the predicted standard deviation, passed through a softplus to ensure positivity).
Step 4 — Likelihood: The likelihood ℓ(z(t) | θ(t)) measures how probable the actual observed value z(t) is under the predicted distribution. During , we maximize this likelihood (equivalently, minimize the ).
Training: maximize likelihood with teacher forcing
DeepAR is trained by maximizing the log-likelihood of the observed data. The loss function is:
During training, the model uses : at each time step, it receives the actual previous observation z(t−1) as input, not its own prediction. This is like a student practicing with an answer key — they always see the correct previous answer, which stabilizes and accelerates learning.
However, during inference there are no future observations to feed. The model must use its own sampled predictions as inputs — creating a fundamentally different regime. This train/test gap is a known challenge of autoregressive models, and DeepAR mitigates it through careful data augmentation (random window selection during training).
Inference: Monte Carlo sample paths
At inference time, DeepAR generates forecasts through Monte Carlo sampling. The process is:
-
Run the LSTM through the known history (conditioning range) to build up the hidden state — exactly like during training.
-
At the first prediction step, sample a value from the predicted distribution: z̃(t₀) ~ Gaussian(μ(t₀), σ(t₀)).
-
Feed that sampled value back as input to the next step (instead of the true observation, which we don't have).
-
Repeat for each step in the prediction horizon.
-
Do the entire process K times (e.g., K=100) to get K sample paths.
Each sample path is a plausible future trajectory. Together, the K paths form an empirical approximation of the full predictive distribution. From these paths you can compute any (P10, P50, P90), prediction intervals, or expected values — whatever your downstream decision system needs.
Practical tricks: scale handling and item embeddings
Real-world time series datasets have two practical challenges that DeepAR addresses with elegant solutions:
Scale normalization: Product A might sell 10,000 units/day while Product B sells 2. If the network sees raw values, it will be dominated by high-volume series. DeepAR divides each series by a scale factor v_i (typically the mean of the absolute values in the conditioning range). The network learns patterns in the normalized space, and predictions are scaled back. This lets a single network handle series spanning many orders of magnitude.
Item embeddings: Each time series gets a learnable vector — a compact fingerprint that captures series-specific characteristics the network can't infer from covariates alone. These embeddings are fed as additional input to the LSTM. For new items (cold start), the embedding starts at a default value and quickly adapts as data arrives. Think of it as giving each product a name tag that the network learns to read.
Choosing the right likelihood
A powerful feature of DeepAR is its modularity around the likelihood function. The network architecture stays the same — only the output layer and loss function change:
Gaussian likelihood — for real-valued data (temperature, prices, energy consumption). The network outputs μ and σ, and the loss is the negative log of the Gaussian density evaluated at the true observation.
Negative binomial likelihood — for count data (product demand, website clicks, event counts). This distribution naturally handles the over-dispersion and zero-inflation common in count data, unlike a Gaussian which can predict negative values.
Other options are equally plug-and-play: Beta for data in [0, 1], Bernoulli for binary outcomes, or mixtures for complex distributions. The only requirement is that you can evaluate the log-likelihood and take gradients — which any standard parametric distribution satisfies.
The idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def deepar_forward(lstm, z_prev, x_t, h_prev, W_mu, W_sigma):
"""One time step of DeepAR: inputs → LSTM → distribution params."""
# Concatenate previous value, covariates, and optional embedding
inp = np.concatenate([z_prev, x_t])
# LSTM produces new hidden state (compressed memory)
h_t = lstm.forward(h_prev, inp) # shape: (hidden_dim,)
# Project hidden state → distribution parameters
mu = W_mu @ h_t # predicted mean
sigma = softplus(W_sigma @ h_t) # predicted std (always > 0)
return h_t, mu, sigma
def deepar_train_loss(model, series, covariates):
"""Negative log-likelihood loss over one time series."""
loss = 0.0
h = np.zeros(model.hidden_dim) # initial hidden state
for t in range(1, len(series)):
h, mu, sigma = deepar_forward(
model.lstm,
z_prev=series[t-1:t], # teacher forcing: real value
x_t=covariates[t],
h_prev=h,
W_mu=model.W_mu,
W_sigma=model.W_sigma,
)
# Gaussian negative log-likelihood
loss += 0.5 * np.log(2 * np.pi * sigma**2) + (series[t] - mu)**2 / (2 * sigma**2)
return loss / (len(series) - 1)
# During INFERENCE: replace series[t-1] with a SAMPLE from Gaussian(mu, sigma).
# Repeat K times → K sample paths → empirical predictive distribution.Results: accuracy across real-world datasets
DeepAR was evaluated on several large-scale forecasting datasets spanning retail demand, automotive parts, and electricity consumption:
~15% improvement in weighted quantile loss over state-of-the-art baselines including ARIMA, ETS (Exponential Smoothing), and Seasonal Naïve across most datasets.
Electricity dataset (370 hourly series): DeepAR significantly outperformed classical methods, especially at longer horizons where capturing cross-series patterns matters most.
Parts dataset (intermittent demand): The negative binomial likelihood excelled here, correctly modeling the many-zeros-then-a-spike pattern that Gaussian methods handle poorly.
Traffic dataset (862 road occupancy series): DeepAR captured daily and weekly seasonality patterns shared across sensors, producing well-calibrated uncertainty bands.
The key finding: a single global model consistently outperformed a basket of individually-fitted expert models — while requiring far less manual configuration.
The encoder-decoder perspective
DeepAR can be understood through the lens of the framework used in sequence-to-sequence models:
The conditioning range (historical observations) acts as the : the LSTM reads through the known history, accumulating context into its hidden state. At the end of this phase, the hidden state is a compressed summary of everything the model knows about this series.
The prediction range acts as the : starting from the encoder's final hidden state, the LSTM generates future values step by step, feeding each sampled prediction back as input. The key difference from machine translation is that both encoder and decoder share the same LSTM — there is no separate decoder network. This weight sharing is what makes DeepAR efficient and what enables it to seamlessly transition from reading history to generating forecasts.
What DeepAR enabled
2017
DeepAR
Global autoregressive RNN for probabilistic forecasting. One model across thousands of series, Monte Carlo sample paths, pluggable likelihoods.
2019
GluonTS
Open-source toolkit for probabilistic time series modeling built on MXNet. Made DeepAR and similar models accessible to practitioners.
2019
Amazon Forecast
AWS managed service for time series forecasting, with DeepAR as a core algorithm. Brought neural forecasting to production at scale.
2021
Informer
Transformer-based long-horizon forecasting with ProbSparse attention. Extended the probabilistic forecasting agenda to attention-based architectures.
2023
PatchTST
Patching + Transformer for time series. Showed that splitting series into patches and applying self-attention rivals RNN-based methods.
2024
Chronos
Pre-trained language model for time series via tokenization. Inherits DeepAR's autoregressive and probabilistic DNA, applied at foundation-model scale.
CitationSalinas, Flunkert, Gasthaus, Januschowski. DeepAR: Probabilistic forecasting with autoregressive recurrent networks. International Journal of Forecasting, 2020.
Terms in this paper
- Autoregressive Modelالنموذج التوليدي التراجعي
- Probabilistic Forecastingالتنبؤ الاحتمالي
- Recurrent Neural Network (RNN)الشبكة العصبية التكرارية
- LSTMشبكة الذاكرة الطويلة قصيرة المدى
- Likelihoodالأرجحية
- Negative Log-Likelihoodسالب لوغاريتم الأرجحية
- Gaussian Distributionالتوزيع الغاوسي
- Teacher Forcingالتوجيه بالمرجع
- Monte Carloأساليب محاكاة مونت كارلو
- Covariateالمتغير المرافق
- Hidden Stateالحالة المخفية
- Cold Startالبداية الباردة
- Embeddingالتضمين
- Quantileالمئين
- Time Seriesسلسلة زمنية