Probability & Statistics1958foundational9 min read

The Regression Analysis of Binary Sequences

تحليل الانحدار للمتتاليات الثنائية

Cox, D. R. — Journal of the Royal Statistical Society: Series B

The problem

Classical linear predicts a continuous number — but many real problems have binary outcomes (0 or 1). Fitting a straight line to binary data can yield "probabilities" above 1 or below 0, which are nonsensical. Researchers needed a principled statistical framework for modeling how the of a binary event changes with one or more explanatory variables, complete with rigorous tests and estimation.

The contribution

Cox formalized : the of a binary outcome as a linear function of the predictors. The logistic () function maps this linear score to a probability in [0, 1]. He derived for the coefficients, identified sufficient statistics (the key summaries that capture all the data's information about each parameter), and built exact — eliminating nuisance parameters without approximation. This gave binary data a rigorous regression theory analogous to normal-theory linear regression.

The impact

Logistic regression became the most widely used model in statistics and machine learning. It is the default first model for any binary task — from medical diagnosis to spam filtering. The sigmoid function Cox employed became the of early neural networks, and the cross-entropy loss used to train modern classifiers is the negative log-likelihood Cox maximized. Logistic regression's ideas grew into generalized linear models and the Cox proportional hazards model for survival analysis.

Imagine a dimmer switch that controls a light. Linear regression is like a dial with no stops: you can twist it past "fully on" or past "fully off" into territory that doesn't exist — the bulb can't be brighter than 100% or darker than 0%.

Logistic regression replaces that dial with an S-shaped cam: no matter how far you turn, the brightness smoothly saturates at 0% and 100%. The position of your hand (the input) maps to a brightness (the probability) that always makes physical sense.

Cox's paper is the engineering blueprint for that cam.

The problem: straight lines can't predict yes or no

In classical linear regression, you predict a continuous output yy from inputs xx: y=α+βxy = \alpha + \beta x. But what if yy can only be 0 or 1 — a patient lives or dies, a part is defective or not?

If you fit a straight line through 0's and 1's, the line inevitably goes above 1 and below 0 for extreme values of xx. Those predictions are meaningless as probabilities. You need a model whose output is always a valid probability — between 0 and 1.

Open in Lab
Toggle between a straight line and the logistic curve to see why linear regression fails for binary data.
The demo wakes as you arrive…

The solution: the logistic function

Cox adopted the logistic function (also called the sigmoid) to map any real number to a probability between 0 and 1. The core idea has three layers — think of it as a pipeline of transformations:

1 — Linear score. Compute a weighted sum z=α+βxz = \alpha + \beta x exactly as in ordinary regression. This score can range from −∞-\infty to +∞+\infty.

Layer 2 — Sigmoid squeeze. Feed zz into the S-shaped sigmoid function σ(z)=1/(1+e−z)\sigma(z) = 1 / (1 + e^{-z}). This «squeezes» any real number into the interval (0,1)(0, 1).

Layer 3 — Interpret as probability. The output σ(z)\sigma(z) is the model's estimate of P(Y=1∣x)P(Y = 1 \mid x) — the probability that the outcome is 1 given the input xx.

This three-step pipeline is the beating heart of logistic regression, and its influence extends far beyond statistics: the sigmoid function later became the canonical activation function in neural networks, and the entire pipeline is what every binary classifier in deep learning still performs at its .

P(Y=1∣x)=σ(α+βx)=11+e−(α+βx)P(Y=1 \mid x) = \sigma(\alpha + \beta x) = \frac{1}{1 + e^{-(\alpha + \beta x)}}
The logistic regression model — probability as a sigmoid of a linear score — α = intercept (baseline log-odds) · β = how much a one-unit increase in x shifts the log-odds · the sigmoid σ wraps the linear score into [0, 1]
Open in Lab
Drag α and β to see how the sigmoid curve shifts and steepens.
The demo wakes as you arrive…

The logit: turning probability inside out

The sigmoid maps a linear score to a probability. Its inverse — the — maps a probability back to a linear score. This inverse view is what makes logistic regression elegant and interpretable.

Start with the odds of the event: if the probability is pp, the odds are p/(1−p)p / (1 - p). A probability of 0.75 means odds of 3:1 — "three times more likely to happen than not." Now take the natural logarithm of the odds to get the log-odds: logit(p)=ln⁡(p/(1−p))\text{logit}(p) = \ln(p / (1 - p)).

Cox's key insight: the logit of the probability is a linear function of the predictors. This is the simplest possible relationship between the inputs and the transformed outcome — and it is exactly why the model is called log-istic regression.

Interpreting coefficients becomes natural: β\beta is the change in log-odds per unit change in xx. A positive β\beta means higher xx makes the event more likely; a negative β\beta means higher xx makes it less likely.

logit(p)=ln⁡ ⁣(p1−p)=α+βx\text{logit}(p) = \ln\!\left(\frac{p}{1-p}\right) = \alpha + \beta x
The logit link — log-odds are linear in the predictors — p/(1−p) = the odds · ln converts odds to log-odds · α + βx is a straight line — the same form as linear regression, but in log-odds space instead of probability space
Open in Lab
Move a probability on the left and watch it map through odds → log-odds on the right.
The demo wakes as you arrive…

Finding the best curve: maximum likelihood estimation

In ordinary regression you minimize squared errors. But for binary outcomes, squared error is not the right ruler — a prediction of 0.99 for an actual 1 should be rewarded more than a prediction of 0.51.

Cox used maximum likelihood estimation: find the values of α\alpha and β\beta that make the observed data most probable under the model. For each data point: if yi=1y_i = 1, the model's contribution is pip_i; if yi=0y_i = 0, it is 1−pi1 - p_i. The likelihood is the product of all these terms, and we maximize it.

In practice we maximize the log-likelihood — a sum instead of a product — because sums are numerically stable and easier to differentiate. This log-likelihood is exactly the negative of what modern deep learning calls the binary cross-entropy loss. Every you have ever trained for classification is optimizing the same objective Cox wrote down in 1958.

ℓ(α,β)=∑i=1n[yiln⁡pi+(1−yi)ln⁡(1−pi)]\ell(\alpha, \beta) = \sum_{i=1}^{n} \left[ y_i \ln p_i + (1 - y_i) \ln(1 - p_i) \right]
Log-likelihood — the objective function of logistic regression — Each term rewards the model for assigning high probability to the outcome that actually happened · maximizing this sum ≡ minimizing binary cross-entropy loss
Open in Lab
Drag the curve to fit the data — watch the log-likelihood rise as the fit improves.
The demo wakes as you arrive…

Sufficient statistics and conditional inference

One of Cox's most elegant contributions was identifying the sufficient statistics for logistic regression. A is a summary of the data that captures all the information the data contains about a parameter — you can throw away the raw data and lose nothing.

Cox showed that the logistic model has the same sufficient statistics as a normal-theory linear model: Y=∑yiY = \sum y_i (the total count of 1's) and X=∑yixiX = \sum y_i x_i (a weighted sum of the predictor values). YY tells you everything about the intercept α\alpha, and XX tells you everything about the slope β\beta.

This led to a powerful technique: conditional inference. To estimate β\beta without worrying about the nuisance parameter α\alpha, simply condition on the observed value of YY. This eliminates α\alpha from the analysis entirely — no approximation needed. It is the binary-data analogue of Fisher's exact test for contingency tables.

Multiple predictors: scaling to higher dimensions

The model extends naturally to multiple predictors. With pp predictors x1,…,xpx_1, \ldots, x_p, the linear score becomes:

z=α+β1x1+β2x2+⋯+βpxpz = \alpha + \beta_1 x_1 + \beta_2 x_2 + \cdots + \beta_p x_p

Each βj\beta_j measures the effect of predictor jj on the log-odds, holding all other predictors constant. This is the same logic as multiple linear regression, but operating in log-odds space.

In modern notation, this is z=w⊤x+bz = \mathbf{w}^\top \mathbf{x} + b — the same dot-product-plus- that is the fundamental building block of every neural network layer. The entire output layer of a binary classifier in deep learning is this equation followed by a sigmoid — literally the logistic regression model.

P(Y=1∣x)=11+e−(w⊤x+b)P(Y=1 \mid \mathbf{x}) = \frac{1}{1 + e^{-(\mathbf{w}^\top \mathbf{x} + b)}}
Multi-predictor logistic regression in modern vector notation — w = weight vector · x = feature vector · b = bias · this is the output neuron of every binary classifier
Open in Lab
Move the weight vector and bias to see how the decision boundary shifts in 2D feature space.
The demo wakes as you arrive…

The same idea in code

Logistic regression from scratchpython

Simplified to show the idea — not the real implementation.

import numpy as np

def sigmoid(z):
    """The logistic function — maps any real number to (0, 1)."""
    return 1 / (1 + np.exp(-z))

def log_likelihood(X, y, w):
    """Log-likelihood: how well do the weights explain the data?"""
    p = sigmoid(X @ w)
    return np.sum(y * np.log(p + 1e-12) + (1 - y) * np.log(1 - p + 1e-12))

def fit_logistic(X, y, lr=0.01, steps=1000):
    """Gradient ascent on the log-likelihood."""
    w = np.zeros(X.shape[1])
    for _ in range(steps):
        p = sigmoid(X @ w)
        gradient = X.T @ (y - p)      # direction of steepest ascent
        w += lr * gradient             # nudge weights to increase likelihood
    return w

# That's it — the sigmoid + log-likelihood + gradient ascent is the
# entire algorithm. Modern libraries add regularization and smarter
# optimizers, but the core is unchanged since Cox (1958).

Historical context: from bioassay to deep learning

The logistic function itself dates to Verhulst (1845), who used it to model population growth. Berkson (1944) applied it to bioassay — estimating the dose of a drug at which 50% of subjects respond. But it was Cox's 1958 paper that unified these threads into a general regression framework with rigorous inference, making the logistic model a tool for any binary-outcome problem.

  1. 1845

    Verhulst's Logistic Curve

    Pierre-François Verhulst introduced the logistic function to model population growth with a carrying capacity — the first appearance of the S-curve.

  2. 1944

    Berkson's Logit Model

    Joseph Berkson applied the logistic function to bioassay dose-response curves, coining the term "logit" and advocating it over the probit model.

  3. 1958

    Cox: The Regression Analysis of Binary Sequences

    Formalized logistic regression as a general-purpose statistical framework with maximum likelihood estimation, sufficient statistics, and conditional inference.

  4. 1972

    Generalized Linear Models

    Nelder and Wedderburn unified logistic regression, Poisson regression, and others into a single framework — generalized linear models — with iteratively reweighted least squares as the fitting algorithm.

  5. 1986

    Backpropagation & the Sigmoid Neuron

    Rumelhart, Hinton, and Williams used the sigmoid (logistic) function as the activation in neural networks trained by backpropagation — directly inheriting Cox's function.

  6. 2012

    Deep Learning Revolution

    AlexNet won ImageNet using sigmoid/softmax output layers with cross-entropy loss — the multi-class generalization of Cox's binary log-likelihood.

  7. 2026

    Still the Default Baseline

    Logistic regression remains the standard first-attempt classifier. Studies consistently show it matches or rivals complex models in many real-world settings.

Why it still matters

CitationCox, D. R.. The Regression Analysis of Binary Sequences. Journal of the Royal Statistical Society: Series B, 1958.

Terms in this paper