Signal Processing1995intermediate11 min read
An Information-Maximization Approach to Blind Separation and Blind Deconvolution
نهج تعظيم المعلومات للفصل الأعمى وفك الالتفاف الأعمى
Bell, A. J. · Sejnowski, T. J. — Neural Computation
The problem
In many real-world settings — from audio recordings to brain scans to seismic data — sensors observe mixtures of underlying source signals, not the sources themselves. Classical methods like PCA can decorrelate the signals (remove second-order correlations), but is not separation: two decorrelated signals can still share higher-order statistical structure. True separation requires statistical independence, which means removing all dependencies, not just correlations. Before this paper, no practical neural-network existed to achieve this from the observed mixtures alone — "blindly", with no knowledge of the mixing process or the source distributions.
The contribution
Bell and Sejnowski showed that maximizing the joint of the outputs of a with nonlinear () activation functions is equivalent to minimizing the between those outputs — which is exactly the condition for statistical independence. The resulting "Infomax" learning rule adjusts an W using only the observed mixtures, requiring no knowledge of the mixing process or source distributions. The algorithm successfully separated up to 10 mixed speakers and also performed blind (removing unknown echoes from a speech signal). This established (ICA) as a practical, general-purpose tool for blind signal processing.
The impact
Infomax ICA became one of the most cited algorithms in signal processing and neuroscience. It is the standard tool for artifact removal in EEG/fMRI brain imaging, audio source separation, and telecommunications. The paper bridged and , showing that Shannon's entropy could drive a neural network to discover hidden structure. It directly inspired sparse coding (Olshausen & Field 1996), which showed that Infomax-like principles applied to natural images produce receptive fields resembling those in the visual cortex — linking ICA to computational neuroscience. Amari's (1996) and the extended Infomax for sub-Gaussian sources (Lee et al. 1999) made it scalable and universal.
Imagine a soundboard operator who accidentally recorded a concert with every microphone channel fused together — guitar, drums, and vocals all tangled into each wire. She has no recording of any instrument alone, no diagram of the stage layout, and no knowledge of how the sounds were mixed. All she has are the tangled wires.
Infomax hands her a set of unmixing knobs: one knob per channel. She turns them — blindly — following a single rule: maximize the information flowing through each channel independently. As the knobs converge, each wire begins to carry a single instrument again. The tangled concert unravels itself.
The cocktail party problem: untangling mixed signals
The classic setup: independent source signals (voices at a party) are mixed by an unknown . What we observe is the mixture . Each sensor (microphone) gets a different linear combination of all sources.
The goal of is to find an unmixing such that — recovering the original sources without knowing or the distributions of the sources. The word "blind" means we work from the mixtures alone.
PCA can decorrelate the mixtures (remove linear correlations), but decorrelation only uses second-order statistics (covariances). Two signals can be uncorrelated yet statistically dependent — think of a signal and its square. True separation requires statistical independence, which involves all higher-order moments.
Why PCA falls short: decorrelation ≠ independence
Consider two uniform source signals scattered inside a square. Mixing them rotates and scales the square into a parallelogram. PCA finds the axes of maximum — it "un-rotates" the parallelogram — but the result is still a rotated square, not the original one. PCA captures the spread of the data but is blind to the edges, which encode higher-order structure.
ICA goes further: it finds axes along which the data is not just uncorrelated but statistically independent. For uniform sources, this means finding the axes aligned with the edges of the parallelogram — which are exactly the directions of the original sources. The key difference: PCA uses only the (second-order), while ICA exploits the full shape of the (all orders).
The Infomax principle: maximize information, get independence for free
The brilliant insight of Bell and Sejnowski is an unexpected connection between two seemingly different goals:
Goal 1 — : Pass the mixed signals through a network , where is a bounded nonlinear function (the sigmoid), and maximize the joint entropy of the outputs. This is Linsker's "Infomax" principle: push as much information as possible through the network.
Goal 2 — Independence: Make the outputs statistically independent, i.e., minimize their mutual information.
The key theorem: because is fixed and invertible, maximizing is equivalent to minimizing the mutual information among the components of (the linear outputs before the nonlinearity). In other words, the network automatically finds independent components by trying to be an efficient information channel. The nonlinearity is what gives the network access to higher-order statistics — without it, we'd be back to PCA.
The math: from entropy to a learning rule
The network computes two stages. First, the linear unmixing: . Then, the element-wise nonlinearity: (the logistic sigmoid).
The joint entropy of the output can be written using the change-of-variables formula for probability densities. After simplification, the of with respect to yields a remarkably clean learning rule.
The term has a geometric meaning: it prevents the determinant of from shrinking to zero (which would collapse all outputs into a single line). Think of it as an "anti-collapse" force that keeps the output representation spread out.
The term is where the magic happens: measures how far each output is from being uniformly distributed after the sigmoid. It pushes each to match the distribution that would make uniform — which happens only when contains exactly one independent source. This is the term that goes beyond PCA: it uses the nonlinearity to access higher-order moments.
The sigmoid as a density matcher
Why specifically a sigmoid? The logistic sigmoid is the cumulative distribution function (CDF) of the logistic distribution. When the network successfully separates a source, it pushes to match that source. Passing through the sigmoid then maps it to a uniform distribution on — because applying any random variable's CDF to itself yields a uniform. Uniform outputs have maximum entropy for a bounded variable, which is exactly what the algorithm is trying to achieve.
This means the sigmoid implicitly assumes that the sources have a distribution shaped like the logistic (super-Gaussian: peaked center, heavy tails). For sub-Gaussian sources (flatter than Gaussian, like uniform), the original algorithm struggles — which is why Lee, Girolami and Sejnowski (1999) later extended it with switchable nonlinearities.
The deep connection: entropy ↔ mutual information ↔ independence
The mutual information between the outputs can be decomposed as:
This is the difference between the sum of individual entropies and the joint entropy. Independence means , which happens when the joint entropy equals the sum of marginals — the outputs carry no redundant information.
Because the sigmoid is a fixed invertible function, and differ only by a term that depends on through the log-. Maximizing with respect to simultaneously maximizes — which, since the marginals are bounded by the sigmoid's range, forces the mutual information toward zero. The algorithm achieves independence as a byproduct of efficient information transmission.
Beyond separation: blind deconvolution
Bell and Sejnowski showed that the same principle extends to blind deconvolution: removing unknown echoes and reverberation from a signal. Instead of an instantaneous mixing matrix , the signal is convolved with an unknown (impulse response). The unmixing becomes a filter too — a set of delay-weighted connections — and the same entropy-maximization rule adapts the filter weights.
Think of it this way: in blind separation, each sensor receives a spatial mixture (superposition of sources at one instant). In blind deconvolution, a single sensor receives a temporal mixture (superposition of a source with its own delayed copies, i.e., echoes). Both are linear mixing — one across channels, the other across time — and information maximization handles both.
The full algorithm: step by step
Here is the complete Infomax ICA pipeline, from raw mixtures to recovered sources. The algorithm is surprisingly simple given what it achieves.
Simplified to show the idea — not the real implementation.
import numpy as np
def sigmoid(u):
return 1.0 / (1.0 + np.exp(-u))
def infomax_ica(X, n_iter=200, lr=0.01):
"""
X: (n_sources, n_samples) — mixed signals, zero-mean
Returns W: unmixing matrix, S: estimated sources
"""
n, T = X.shape
W = np.eye(n) # start with identity
for _ in range(n_iter):
u = W @ X # linear unmixing
y = sigmoid(u) # nonlinear squashing
phi = 1 - 2 * y # score function φ(u)
# Natural gradient update (Amari 1996):
# avoids computing W⁻¹, converges faster
dW = (np.eye(n) + phi @ u.T / T) @ W
W += lr * dW
S = W @ X # final source estimates
return W, S
# That's it. The sigmoid + entropy maximization automatically
# finds statistically independent components — no knowledge
# of the mixing matrix or source distributions needed.Why it mattered: from cocktail parties to brains
Infomax ICA did not just solve a signal processing problem — it created a bridge between information theory, neural networks, and neuroscience. The paper showed that a neural network can discover hidden structure in data by following a purely information-theoretic objective, with no labels, no teacher, and no knowledge of the data-generating process.
This unsupervised, information-driven approach directly inspired Olshausen and Field's sparse coding (1996), which applied similar principles to natural images and found that the learned features resemble the orientation-selective receptive fields of neurons in the primary visual cortex — suggesting that the brain itself might perform something like ICA.
1989
Linsker's Infomax Principle
Linsker proposed that neural networks should maximize mutual information between inputs and outputs. But in linear networks, this only yields PCA — no higher-order structure.
1994
Comon formalizes ICA
Comon gave Independent Component Analysis its name and formal framework using higher-order cumulants. But a practical neural-network algorithm was still missing.
1995
Bell & Sejnowski — Infomax ICA
The paper that connected information maximization to ICA through the sigmoid nonlinearity. Separated up to 10 mixed speakers and performed blind deconvolution.
1996
Amari's natural gradient
Amari replaced the costly matrix inverse in the learning rule with a multiplicative natural gradient, making Infomax fast enough for real-world applications.
1996
Sparse coding — ICA meets the visual cortex
Olshausen and Field showed that Infomax-like principles applied to natural images produce receptive fields resembling those in the visual cortex, bridging signal processing and neuroscience.
1999
Extended Infomax
Lee, Girolami and Sejnowski extended Infomax with switchable nonlinearities to handle both sub-Gaussian and super-Gaussian sources, making it universal.
1999
FastICA
Hyvärinen proposed FastICA — a fixed-point algorithm maximizing non-Gaussianity, faster than gradient-based methods for many problems. Together with Infomax, these became the two standard ICA algorithms.
Today, Infomax ICA is embedded in every major neuroscience toolbox (EEGLAB, MNE-Python, FieldTrip) for separating brain signals from artifacts. It is a standard preprocessing step in telecommunications for separating superimposed radio signals. The paper's deeper legacy is the demonstration that information theory provides a principled objective for unsupervised learning — an idea that echoes through variational autoencoders, contrastive learning, and modern representation learning.
CitationBell, A. J. and Sejnowski, T. J.. An Information-Maximization Approach to Blind Separation and Blind Deconvolution. Neural Computation, 1995.
Terms in this paper
- Independent Component Analysisتحليل المكوّنات المستقلة
- Mutual Informationالمعلومات المتبادلة
- Entropyالعشوائية الدلالية
- Blind Source Separationالفصل الأعمى للمصادر
- Sigmoidدالة سيجمويد
- Mixing Matrixمصفوفة المَزج
- Unmixing Matrixمصفوفة فكّ المَزج
- Cocktail Party Problemمسألة حفلة الكوكتيل
- Information Maximizationتعظيم المعلومات
- Natural Gradientالمُتدرِّج الطبيعي