Time Series1970foundational12 min read

Time Series Analysis: Forecasting and Control

تحليل السلاسل الزمنية: التنبؤ والتحكّم

Box, G. E. P. · Jenkins, G. M. — Holden-Day

The problem

Before 1970 most forecasting relied on ad-hoc rules: moving averages, exponential smoothing, or manually fitting trend lines. There was no unified, principled framework for deciding how much of the past to use, how to handle trends and seasonality, or how to check whether a model actually captured the data's structure. Practitioners had to guess whether a series followed an autoregressive pattern, a moving-average pattern, or some combination, and there was no systematic procedure for choosing, estimating, diagnosing, and refining a model.

The contribution

Box and Jenkins unified three components into a single framework called ARIMA (AutoRegressive Integrated ). The AR part captures how the current value depends on past values. The I (Integrated) part uses to remove non-. The MA part models the influence of past errors. They proposed a systematic cycle: identify the model orders from plots, estimate parameters by maximum likelihood, check residuals for remaining structure, and iterate until the model is adequate. This became known as the Box-Jenkins methodology.

The impact

The Box-Jenkins methodology became the gold standard for time series forecasting in economics, engineering, and the natural sciences. It provided the theoretical foundation for seasonal models (SARIMA), transfer function models, and intervention analysis. Its influence extends to modern deep learning forecasters like DeepAR, N-BEATS, Prophet, and Autoformer, which all inherit ARIMA's decomposition intuition even as they replace its linear assumptions with neural networks. Every statistical forecasting package today implements Box-Jenkins at its core.

Imagine you're watching the surface of a lake. Some days it's calm, some days there are waves. A trend is a slow tide raising the whole lake — you need to subtract it before you can study the waves. An AR process is a wave that remembers its own height: if it was high a moment ago, it's probably still high. An MA process is a stone tossed into the lake: the splash creates ripples that linger for a few moments, then fade. ARIMA says: remove the tide (differencing), then model the remaining surface as a combination of self-remembering waves (AR) and fading ripples (MA).

The problem: forecasting without a framework

By the late 1960s, time series forecasting was a patchwork of disconnected tools. Engineers used exponential smoothing, economists fitted trend-cycle decompositions by hand, and statisticians studied autoregressive processes in theory but had no practical workflow for applying them to real data.

The core questions had no unified answer:

  • How much of the past should influence the forecast?
  • How do you remove trends and seasonality systematically?
  • Once you fit a model, how do you know it's actually capturing the data's structure rather than missing something important?

What was needed was an end-to-end methodology: a single framework that guides the analyst from raw data to a validated, interpretable forecast.

Foundation: stationarity — the language ARIMA speaks

Before fitting any model, the data must be stationary: its mean and variance should not change over time. Think of a stationary series as a conversation that stays on topic — it fluctuates around a constant level. A non-stationary series is like a conversation that drifts — the level itself keeps moving.

Most real-world series are non-stationary. Stock prices trend upward, sales have seasonal peaks, temperatures have annual cycles. ARIMA handles non-stationarity by differencing: subtracting each observation from the previous one. If the original series yty_t has a trend, the differenced series yt−yt−1y_t - y_{t-1} often does not. The number of times you difference is called the order of integration dd — this is the "I" in ARIMA.

Open in Lab
Toggle differencing to watch a trending series become stationary. Notice how the mean stops drifting after one round of differencing.
The demo wakes as you arrive…

The AR component: the series remembers itself

The Autoregressive component models the idea that the current value is a weighted combination of past values. An AR(1) process says: "today's value equals a fraction of yesterday's, plus random noise." An AR(2) extends this to: "today depends on both yesterday and the day before."

The mental model is inertia: a hot day is more likely to be followed by another hot day than a sudden cold snap. The coefficient ϕ\phi controls the strength of this memory. When ∣ϕ∣<1|\phi| < 1 the memory fades — the series is stable and always returns to its mean. When ∣ϕ∣=1|\phi| = 1 you get a random walk — the series wanders without ever returning, which is non-stationary.

yt=c+ϕ1yt−1+ϕ2yt−2+⋯+ϕpyt−p+εty_t = c + \phi_1 y_{t-1} + \phi_2 y_{t-2} + \cdots + \phi_p y_{t-p} + \varepsilon_t
AR(p) — autoregressive model of order p — The current value is a linear combination of the p most recent past values plus white noise ε. The coefficients φ₁…φₚ are learned from data.

The MA component: the echoes of surprise

While AR models memory of past values, the Moving Average component models the memory of past surprises — the forecast errors. An MA(1) process says: "today's value is the mean plus a fraction of yesterday's surprise."

The mental model is a stone dropped in a pond: the splash (the shock, εt−1\varepsilon_{t-1}) creates ripples that affect the surface for a few time steps, then fade away. The crucial difference from AR is that MA shocks have finite memory — an MA(q) process forgets everything more than q steps ago, no matter how large the shock was.

yt=μ+εt+θ1εt−1+θ2εt−2+⋯+θqεt−qy_t = \mu + \varepsilon_t + \theta_1 \varepsilon_{t-1} + \theta_2 \varepsilon_{t-2} + \cdots + \theta_q \varepsilon_{t-q}
MA(q) — moving average model of order q — The current value is the mean plus a weighted sum of the q most recent forecast errors. Each θ controls how much a past shock still echoes today.

Putting it together: ARIMA(p, d, q)

ARIMA combines all three ideas into a single model. First, difference the series dd times to make it stationary. Then model the differenced series as an ARMA(p, q) process — a mixture of autoregressive memory and moving-average shocks.

The notation ARIMA(p, d, q) tells you everything:

  • p = number of AR terms (how far back the memory reaches)
  • d = number of differences (how many times you subtract to remove trends)
  • q = number of MA terms (how many past shocks still echo)

For example, ARIMA(1,1,1) means: difference once to remove the trend, then model the result with one AR term and one MA term. This simple specification turns out to be surprisingly powerful for many real-world series.

ϕ(B)(1−B)d yt=θ(B) εt\phi(B)(1 - B)^d \, y_t = \theta(B) \, \varepsilon_t
ARIMA(p, d, q) in backshift operator notation — B is the backshift operator: Byₜ = yₜ₋₁. The polynomial φ(B) captures AR, (1−B)ᵈ captures differencing, and θ(B) captures MA. This compact notation makes the algebraic structure transparent.
Open in Lab
Adjust p, d, q sliders to see how each component shapes the generated time series. Watch how AR adds persistence, differencing removes trend, and MA adds short-lived shocks.
The demo wakes as you arrive…

The Box-Jenkins methodology: a systematic cycle

The true innovation of Box and Jenkins was not just the ARIMA model itself, but a methodology — a repeatable, disciplined cycle for building time series models. Think of it as a scientific method for forecasting:

Step 1 — Identification: Examine autocorrelation (ACF) and (PACF) plots to determine the orders p, d, q. The ACF measures total correlation at each lag; the PACF isolates the direct effect of each lag after removing intermediate effects. Together they form a fingerprint that reveals the model class.

Step 2 — Estimation: Fit the model parameters (φ, θ) using , which finds the parameter values that make the observed data most probable under the model.

Step 3 — Diagnostic checking: Examine the residuals. If the model is adequate, residuals should be — no remaining autocorrelation, no patterns. If patterns remain, return to Step 1 with a revised model.

This iterate-until-clean cycle is what made Box-Jenkins a methodology rather than just a formula.

Open in Lab
Walk through the three stages of the Box-Jenkins cycle. Click each stage to see what happens inside.
The demo wakes as you arrive…

Reading the fingerprints: ACF and PACF

The autocorrelation function (ACF) and partial autocorrelation function (PACF) are the primary diagnostic tools for identifying ARIMA model orders. They reveal the temporal DNA of a time series.

For a pure AR(p) process, the PACF shows sharp spikes at lags 1 through p, then cuts off to zero — like a set of dominoes that fall exactly p deep. The ACF, meanwhile, decays gradually (exponentially or with damped oscillations).

For a pure MA(q) process, the pattern reverses: the ACF shows sharp spikes at lags 1 through q then cuts off, while the PACF decays gradually.

For mixed ARMA(p,q) processes, both functions decay gradually — which makes identification harder and is where experience and information criteria like AIC become essential.

Open in Lab
Select a process type (AR, MA, or ARMA) and adjust the order. The ACF and PACF plots update live, showing the characteristic fingerprint of each model.
The demo wakes as you arrive…

Estimation and diagnostics: fitting and checking

Once you've identified candidate orders (p, d, q), the next step is estimation. Box and Jenkins advocated maximum likelihood estimation (MLE), which finds the parameter values ϕ^\hat\phi, θ^\hat\theta that maximize the probability of observing the actual data under the model. For Gaussian ARIMA, this is equivalent to minimizing the sum of squared one-step-ahead prediction errors.

After estimation comes the critical diagnostic step: examining the residuals ε^t=yt−y^t\hat\varepsilon_t = y_t - \hat{y}_t. If the model is adequate, these residuals should look like white noise — independent random values with no remaining pattern. Tools for checking include the residual ACF plot (no significant spikes), the Ljung-Box test (a formal test for residual autocorrelation), and QQ-plots (checking normality).

If diagnostics reveal problems — significant residual autocorrelation, or non-constant variance — you iterate: revise the model orders, re-estimate, and re-check. The cycle continues until the residuals are clean.

Forecasting: predicting with honest uncertainty

Once a model passes diagnostics, it can forecast future values. ARIMA forecasts are optimal in a precise sense: they minimize mean squared error among all linear predictors when the true process is ARIMA.

A key feature is that ARIMA provides prediction intervals — not just point forecasts. As the forecast horizon grows, the intervals widen, reflecting increasing uncertainty. This honest uncertainty quantification was a major advance over earlier methods that gave only point predictions with no indication of reliability.

For an AR process, the forecast converges to the mean exponentially fast. For an integrated process (d ≥ 1), the forecast follows the last observed trend and the prediction intervals grow without bound — the model candidly admits it cannot predict far ahead when the series has no fixed level to return to.

Open in Lab
Generate a time series, fit an ARIMA model, and watch it forecast ahead. Notice how prediction intervals fan out with increasing horizon.
The demo wakes as you arrive…

Seasonal ARIMA: capturing repeating patterns

Many real-world series have seasonality — monthly retail sales peak in December, electricity demand surges every summer, airline passengers follow annual cycles. Box and Jenkins extended ARIMA to handle this with , written as ARIMA(p,d,q)(P,D,Q)ₘ, where m is the seasonal period (12 for monthly, 4 for quarterly).

The idea is elegant: apply the same AR-I-MA logic at two scales simultaneously. The non-seasonal part (p,d,q) captures short-term dependencies within a season. The seasonal part (P,D,Q) captures dependencies across seasons — how this January relates to last January. Seasonal differencing (subtracting the value from m steps ago) removes the seasonal pattern the same way ordinary differencing removes a trend.

The famous "Airline model," ARIMA(0,1,1)(0,1,1)₁₂, was Box and Jenkins' showcase: one non-seasonal MA term, one seasonal MA term, one ordinary difference, one seasonal difference. Despite its simplicity, it models international airline passenger data remarkably well and remains a benchmark to this day.

Open in Lab
Explore the Airline model. Toggle seasonal differencing to see the seasonal pattern removed. Adjust seasonal MA to model the remaining structure.
The demo wakes as you arrive…

The same idea in code

ARIMA forecasting with statsmodelspython

Simplified to show the idea — not the real implementation.

import numpy as np
from statsmodels.tsa.arima.model import ARIMA
from statsmodels.graphics.tsaplots import plot_acf, plot_pacf

# Generate sample data: random walk + noise
np.random.seed(42)
y = np.cumsum(np.random.randn(200)) + 0.5 * np.random.randn(200)

# Step 1: Identify — look at ACF/PACF of differenced series
diff_y = np.diff(y)  # first difference removes the random walk trend
# plot_acf(diff_y)   # would show the fingerprint
# plot_pacf(diff_y)  # would show the fingerprint

# Step 2: Estimate — fit ARIMA(1, 1, 1)
model = ARIMA(y, order=(1, 1, 1))
fitted = model.fit()
print(fitted.summary())

# Step 3: Diagnose — check residuals
residuals = fitted.resid
# If residuals look like white noise → model is adequate
# If not → try different (p, d, q) and repeat

# Forecast 20 steps ahead with 95% prediction intervals
forecast = fitted.get_forecast(steps=20)
mean_forecast = forecast.predicted_mean
conf_int = forecast.conf_int(alpha=0.05)  # 95% intervals

# The intervals widen with horizon — honest uncertainty!

Why it mattered: the legacy of Box-Jenkins

  1. 1927

    Yule's autoregressive model

    G. Udny Yule introduced autoregressive models to study sunspot cycles, establishing the idea that a time series can regress on its own past.

  2. 1937

    Slutsky's moving average insight

    Eugen Slutsky showed that summing random shocks (a moving average process) can generate realistic economic cycles, establishing MA as a fundamental building block.

  3. 1938

    Wold decomposition theorem

    Herman Wold proved any stationary process can be decomposed into a deterministic part and an MA(∞) part — the theoretical warrant for ARMA modeling.

  4. 1960

    Kalman filter

    Rudolf Kálmán introduced the Kalman filter for state-space estimation. Later work showed ARIMA can be cast as a state-space model and estimated via the Kalman filter, unifying the two traditions.

  5. 1970

    Box & Jenkins — Time Series Analysis

    The landmark textbook unifies AR, I, and MA into ARIMA, introduces the identify-estimate-diagnose cycle, and provides the seasonal extension. It defines the field for decades.

  6. 2017

    DeepAR — neural probabilistic forecasting

    Amazon's DeepAR uses autoregressive RNNs to produce probabilistic forecasts. The autoregressive idea descends directly from ARIMA, but replaces linear coefficients with learned neural networks.

  7. 2017

    Prophet — decomposition for practitioners

    Facebook's Prophet decomposes series into trend, seasonality, and holidays — a modern echo of ARIMA's decomposition philosophy, designed for analysts who are not time series specialists.

  8. 2020

    N-BEATS — pure neural forecasting

    N-BEATS uses stacks of fully-connected networks with backward and forward residual links, beating ARIMA and statistical methods on many benchmarks without any time series-specific assumptions.

  9. 2021

    Autoformer — decomposition meets attention

    Autoformer embeds seasonal-trend decomposition inside a Transformer architecture with auto-correlation attention. The decomposition idea traces back to ARIMA's separation of trend from stationary structure.

Every time you see a forecasting model that decomposes a series into trend and seasonality, checks residuals for remaining patterns, or expresses uncertainty through prediction intervals — you are seeing Box and Jenkins' intellectual DNA. ARIMA was the seed from which the entire field of systematic time series forecasting grew.

CitationBox, G. E. P., Jenkins, G. M.. Time Series Analysis: Forecasting and Control. Holden-Day, 1970.

Terms in this paper