Probability & Statistics1958foundational9 min read
The Regression Analysis of Binary Sequences
تحليل الانحدار للمتتاليات الثنائية
Cox, D. R. — Journal of the Royal Statistical Society: Series B
The problem
Classical linear predicts a continuous number — but many real problems have binary outcomes (0 or 1). Fitting a straight line to binary data can yield "probabilities" above 1 or below 0, which are nonsensical. Researchers needed a principled statistical framework for modeling how the of a binary event changes with one or more explanatory variables, complete with rigorous tests and estimation.
The contribution
Cox formalized : the of a binary outcome as a linear function of the predictors. The logistic () function maps this linear score to a probability in [0, 1]. He derived for the coefficients, identified sufficient statistics (the key summaries that capture all the data's information about each parameter), and built exact — eliminating nuisance parameters without approximation. This gave binary data a rigorous regression theory analogous to normal-theory linear regression.
The impact
Logistic regression became the most widely used model in statistics and machine learning. It is the default first model for any binary task — from medical diagnosis to spam filtering. The sigmoid function Cox employed became the of early neural networks, and the cross-entropy loss used to train modern classifiers is the negative log-likelihood Cox maximized. Logistic regression's ideas grew into generalized linear models and the Cox proportional hazards model for survival analysis.
Imagine a dimmer switch that controls a light. Linear regression is like a dial with no stops: you can twist it past "fully on" or past "fully off" into territory that doesn't exist — the bulb can't be brighter than 100% or darker than 0%.
Logistic regression replaces that dial with an S-shaped cam: no matter how far you turn, the brightness smoothly saturates at 0% and 100%. The position of your hand (the input) maps to a brightness (the probability) that always makes physical sense.
Cox's paper is the engineering blueprint for that cam.
The problem: straight lines can't predict yes or no
In classical linear regression, you predict a continuous output from inputs : . But what if can only be 0 or 1 — a patient lives or dies, a part is defective or not?
If you fit a straight line through 0's and 1's, the line inevitably goes above 1 and below 0 for extreme values of . Those predictions are meaningless as probabilities. You need a model whose output is always a valid probability — between 0 and 1.
The solution: the logistic function
Cox adopted the logistic function (also called the sigmoid) to map any real number to a probability between 0 and 1. The core idea has three layers — think of it as a pipeline of transformations:
1 — Linear score. Compute a weighted sum exactly as in ordinary regression. This score can range from to .
Layer 2 — Sigmoid squeeze. Feed into the S-shaped sigmoid function . This «squeezes» any real number into the interval .
Layer 3 — Interpret as probability. The output is the model's estimate of — the probability that the outcome is 1 given the input .
This three-step pipeline is the beating heart of logistic regression, and its influence extends far beyond statistics: the sigmoid function later became the canonical activation function in neural networks, and the entire pipeline is what every binary classifier in deep learning still performs at its .
The logit: turning probability inside out
The sigmoid maps a linear score to a probability. Its inverse — the — maps a probability back to a linear score. This inverse view is what makes logistic regression elegant and interpretable.
Start with the odds of the event: if the probability is , the odds are . A probability of 0.75 means odds of 3:1 — "three times more likely to happen than not." Now take the natural logarithm of the odds to get the log-odds: .
Cox's key insight: the logit of the probability is a linear function of the predictors. This is the simplest possible relationship between the inputs and the transformed outcome — and it is exactly why the model is called log-istic regression.
Interpreting coefficients becomes natural: is the change in log-odds per unit change in . A positive means higher makes the event more likely; a negative means higher makes it less likely.
Finding the best curve: maximum likelihood estimation
In ordinary regression you minimize squared errors. But for binary outcomes, squared error is not the right ruler — a prediction of 0.99 for an actual 1 should be rewarded more than a prediction of 0.51.
Cox used maximum likelihood estimation: find the values of and that make the observed data most probable under the model. For each data point: if , the model's contribution is ; if , it is . The likelihood is the product of all these terms, and we maximize it.
In practice we maximize the log-likelihood — a sum instead of a product — because sums are numerically stable and easier to differentiate. This log-likelihood is exactly the negative of what modern deep learning calls the binary cross-entropy loss. Every you have ever trained for classification is optimizing the same objective Cox wrote down in 1958.
Sufficient statistics and conditional inference
One of Cox's most elegant contributions was identifying the sufficient statistics for logistic regression. A is a summary of the data that captures all the information the data contains about a parameter — you can throw away the raw data and lose nothing.
Cox showed that the logistic model has the same sufficient statistics as a normal-theory linear model: (the total count of 1's) and (a weighted sum of the predictor values). tells you everything about the intercept , and tells you everything about the slope .
This led to a powerful technique: conditional inference. To estimate without worrying about the nuisance parameter , simply condition on the observed value of . This eliminates from the analysis entirely — no approximation needed. It is the binary-data analogue of Fisher's exact test for contingency tables.
Multiple predictors: scaling to higher dimensions
The model extends naturally to multiple predictors. With predictors , the linear score becomes:
Each measures the effect of predictor on the log-odds, holding all other predictors constant. This is the same logic as multiple linear regression, but operating in log-odds space.
In modern notation, this is — the same dot-product-plus- that is the fundamental building block of every neural network layer. The entire output layer of a binary classifier in deep learning is this equation followed by a sigmoid — literally the logistic regression model.
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def sigmoid(z):
"""The logistic function — maps any real number to (0, 1)."""
return 1 / (1 + np.exp(-z))
def log_likelihood(X, y, w):
"""Log-likelihood: how well do the weights explain the data?"""
p = sigmoid(X @ w)
return np.sum(y * np.log(p + 1e-12) + (1 - y) * np.log(1 - p + 1e-12))
def fit_logistic(X, y, lr=0.01, steps=1000):
"""Gradient ascent on the log-likelihood."""
w = np.zeros(X.shape[1])
for _ in range(steps):
p = sigmoid(X @ w)
gradient = X.T @ (y - p) # direction of steepest ascent
w += lr * gradient # nudge weights to increase likelihood
return w
# That's it — the sigmoid + log-likelihood + gradient ascent is the
# entire algorithm. Modern libraries add regularization and smarter
# optimizers, but the core is unchanged since Cox (1958).Historical context: from bioassay to deep learning
The logistic function itself dates to Verhulst (1845), who used it to model population growth. Berkson (1944) applied it to bioassay — estimating the dose of a drug at which 50% of subjects respond. But it was Cox's 1958 paper that unified these threads into a general regression framework with rigorous inference, making the logistic model a tool for any binary-outcome problem.
1845
Verhulst's Logistic Curve
Pierre-François Verhulst introduced the logistic function to model population growth with a carrying capacity — the first appearance of the S-curve.
1944
Berkson's Logit Model
Joseph Berkson applied the logistic function to bioassay dose-response curves, coining the term "logit" and advocating it over the probit model.
1958
Cox: The Regression Analysis of Binary Sequences
Formalized logistic regression as a general-purpose statistical framework with maximum likelihood estimation, sufficient statistics, and conditional inference.
1972
Generalized Linear Models
Nelder and Wedderburn unified logistic regression, Poisson regression, and others into a single framework — generalized linear models — with iteratively reweighted least squares as the fitting algorithm.
1986
Backpropagation & the Sigmoid Neuron
Rumelhart, Hinton, and Williams used the sigmoid (logistic) function as the activation in neural networks trained by backpropagation — directly inheriting Cox's function.
2012
Deep Learning Revolution
AlexNet won ImageNet using sigmoid/softmax output layers with cross-entropy loss — the multi-class generalization of Cox's binary log-likelihood.
2026
Still the Default Baseline
Logistic regression remains the standard first-attempt classifier. Studies consistently show it matches or rivals complex models in many real-world settings.
Why it still matters
CitationCox, D. R.. The Regression Analysis of Binary Sequences. Journal of the Royal Statistical Society: Series B, 1958.
Terms in this paper
- Logistic Regressionالانحدار اللوجستي الاحتمالي
- Sigmoidدالة سيجمويد
- Maximum Likelihood Estimationتقدير الأرجحية القصوى
- Cross Entropyالعشوائية المتقاطعة
- Logitالدرجة الخام
- Log-Oddsلوغاريتم الأرجحية
- Decision Boundaryحدّ القرار
- Sufficient Statisticالإحصاء الكافي
- Conditional Inferenceالاستدلال الشرطي