Signal Processing1995intermediate11 min read

An Information-Maximization Approach to Blind Separation and Blind Deconvolution

نهج تعظيم المعلومات للفصل الأعمى وفك الالتفاف الأعمى

Bell, A. J. · Sejnowski, T. J. — Neural Computation

The problem

In many real-world settings — from audio recordings to brain scans to seismic data — sensors observe mixtures of underlying source signals, not the sources themselves. Classical methods like PCA can decorrelate the signals (remove second-order correlations), but is not separation: two decorrelated signals can still share higher-order statistical structure. True separation requires statistical independence, which means removing all dependencies, not just correlations. Before this paper, no practical neural-network existed to achieve this from the observed mixtures alone — "blindly", with no knowledge of the mixing process or the source distributions.

The contribution

Bell and Sejnowski showed that maximizing the joint of the outputs of a with nonlinear () activation functions is equivalent to minimizing the between those outputs — which is exactly the condition for statistical independence. The resulting "Infomax" learning rule adjusts an W using only the observed mixtures, requiring no knowledge of the mixing process or source distributions. The algorithm successfully separated up to 10 mixed speakers and also performed blind (removing unknown echoes from a speech signal). This established (ICA) as a practical, general-purpose tool for blind signal processing.

The impact

Infomax ICA became one of the most cited algorithms in signal processing and neuroscience. It is the standard tool for artifact removal in EEG/fMRI brain imaging, audio source separation, and telecommunications. The paper bridged and , showing that Shannon's entropy could drive a neural network to discover hidden structure. It directly inspired sparse coding (Olshausen & Field 1996), which showed that Infomax-like principles applied to natural images produce receptive fields resembling those in the visual cortex — linking ICA to computational neuroscience. Amari's (1996) and the extended Infomax for sub-Gaussian sources (Lee et al. 1999) made it scalable and universal.

Imagine a soundboard operator who accidentally recorded a concert with every microphone channel fused together — guitar, drums, and vocals all tangled into each wire. She has no recording of any instrument alone, no diagram of the stage layout, and no knowledge of how the sounds were mixed. All she has are the tangled wires.

Infomax hands her a set of unmixing knobs: one knob per channel. She turns them — blindly — following a single rule: maximize the information flowing through each channel independently. As the knobs converge, each wire begins to carry a single instrument again. The tangled concert unravels itself.

The cocktail party problem: untangling mixed signals

The classic setup: nn independent source signals s1,s2,…,sns_1, s_2, \ldots, s_n (voices at a party) are mixed by an unknown n×nn \times n AA. What we observe is the mixture x=As\mathbf{x} = A\mathbf{s}. Each sensor (microphone) gets a different linear combination of all sources.

The goal of is to find an unmixing WW such that u=Wx≈s\mathbf{u} = W\mathbf{x} \approx \mathbf{s} — recovering the original sources without knowing AA or the distributions of the sources. The word "blind" means we work from the mixtures alone.

PCA can decorrelate the mixtures (remove linear correlations), but decorrelation only uses second-order statistics (covariances). Two signals can be uncorrelated yet statistically dependent — think of a signal and its square. True separation requires statistical independence, which involves all higher-order moments.

Open in Lab
Watch how two independent source signals get mixed by matrix A into tangled observations, then recovered by the unmixing matrix W. Press "Mix" then "Unmix".
The demo wakes as you arrive…

Why PCA falls short: decorrelation ≠ independence

Consider two uniform source signals scattered inside a square. Mixing them rotates and scales the square into a parallelogram. PCA finds the axes of maximum — it "un-rotates" the parallelogram — but the result is still a rotated square, not the original one. PCA captures the spread of the data but is blind to the edges, which encode higher-order structure.

ICA goes further: it finds axes along which the data is not just uncorrelated but statistically independent. For uniform sources, this means finding the axes aligned with the edges of the parallelogram — which are exactly the directions of the original sources. The key difference: PCA uses only the (second-order), while ICA exploits the full shape of the (all orders).

Open in Lab
Left: PCA axes align with maximum variance — they decorrelate but don't recover the original sources. Right: ICA axes align with the edges of the data — they find the independent sources.
The demo wakes as you arrive…

The Infomax principle: maximize information, get independence for free

The brilliant insight of Bell and Sejnowski is an unexpected connection between two seemingly different goals:

Goal 1 — : Pass the mixed signals through a network y=g(Wx)\mathbf{y} = g(W\mathbf{x}), where gg is a bounded nonlinear function (the sigmoid), and maximize the joint entropy H(y)H(\mathbf{y}) of the outputs. This is Linsker's "Infomax" principle: push as much information as possible through the network.

Goal 2 — Independence: Make the outputs statistically independent, i.e., minimize their mutual information.

The key theorem: because gg is fixed and invertible, maximizing H(y)H(\mathbf{y}) is equivalent to minimizing the mutual information among the components of u=Wx\mathbf{u} = W\mathbf{x} (the linear outputs before the nonlinearity). In other words, the network automatically finds independent components by trying to be an efficient information channel. The nonlinearity is what gives the network access to higher-order statistics — without it, we'd be back to PCA.

Open in Lab
The Infomax architecture: mixed signals x pass through unmixing matrix W, then through sigmoid g. Maximizing the output entropy H(y) drives W toward the inverse of the mixing matrix.
The demo wakes as you arrive…

The math: from entropy to a learning rule

The network computes two stages. First, the linear unmixing: u=Wx\mathbf{u} = W\mathbf{x}. Then, the element-wise nonlinearity: yi=g(ui)=11+e−uiy_i = g(u_i) = \frac{1}{1 + e^{-u_i}} (the logistic sigmoid).

The joint entropy of the output y\mathbf{y} can be written using the change-of-variables formula for probability densities. After simplification, the of H(y)H(\mathbf{y}) with respect to WW yields a remarkably clean learning rule.

ΔW∝[W−T+φ(u) xT]\Delta W \propto [W^{-T} + \varphi(\mathbf{u})\,\mathbf{x}^T]
The Infomax learning rule — The goal of Infomax is to learn representations whose outputs carry as much information as possible about the inputs while becoming statistically independent from one another. The learning rule therefore balances two objectives: preventing different outputs from collapsing into the same signal and encouraging each output to capture distinct structure in the data. When training converges, the outputs tend to correspond to the underlying independent factors that generated the observations.

The term W−TW^{-T} has a geometric meaning: it prevents the determinant of WW from shrinking to zero (which would collapse all outputs into a single line). Think of it as an "anti-collapse" force that keeps the output representation spread out.

The term φ(u)xT\varphi(\mathbf{u})\mathbf{x}^T is where the magic happens: φ(ui)=1−2g(ui)\varphi(u_i) = 1 - 2g(u_i) measures how far each output is from being uniformly distributed after the sigmoid. It pushes each uiu_i to match the distribution that would make yiy_i uniform — which happens only when uiu_i contains exactly one independent source. This is the term that goes beyond PCA: it uses the nonlinearity to access higher-order moments.

The sigmoid as a density matcher

Why specifically a sigmoid? The logistic sigmoid g(u)=1/(1+e−u)g(u) = 1/(1+e^{-u}) is the cumulative distribution function (CDF) of the logistic distribution. When the network successfully separates a source, it pushes uiu_i to match that source. Passing uiu_i through the sigmoid then maps it to a uniform distribution on [0,1][0,1] — because applying any random variable's CDF to itself yields a uniform. Uniform outputs have maximum entropy for a bounded variable, which is exactly what the algorithm is trying to achieve.

This means the sigmoid implicitly assumes that the sources have a distribution shaped like the logistic (super-Gaussian: peaked center, heavy tails). For sub-Gaussian sources (flatter than Gaussian, like uniform), the original algorithm struggles — which is why Lee, Girolami and Sejnowski (1999) later extended it with switchable nonlinearities.

Open in Lab
Slide the source distribution from super-Gaussian to sub-Gaussian. Watch how the sigmoid maps it: super-Gaussian → near-uniform output (high entropy), sub-Gaussian → U-shaped output (low entropy). The algorithm works best when the sigmoid matches the source distribution.
The demo wakes as you arrive…

The deep connection: entropy ↔ mutual information ↔ independence

The mutual information between the outputs can be decomposed as:

I(u1;u2;…;un)=∑iH(ui)−H(u)I(u_1; u_2; \ldots; u_n) = \sum_i H(u_i) - H(\mathbf{u})

This is the difference between the sum of individual entropies and the joint entropy. Independence means I=0I = 0, which happens when the joint entropy equals the sum of marginals — the outputs carry no redundant information.

Because the sigmoid is a fixed invertible function, H(y)H(\mathbf{y}) and H(u)H(\mathbf{u}) differ only by a term that depends on WW through the log-. Maximizing H(y)H(\mathbf{y}) with respect to WW simultaneously maximizes H(u)H(\mathbf{u}) — which, since the marginals H(ui)H(u_i) are bounded by the sigmoid's range, forces the mutual information II toward zero. The algorithm achieves independence as a byproduct of efficient information transmission.

max⁡WH(y)  ⟺  min⁡W∑iH(ui)−H(u)  ⟺  min⁡W  I(u1;…;un)\max_W H(\mathbf{y}) \;\Longleftrightarrow\; \min_W \sum_i H(u_i) - H(\mathbf{u}) \;\Longleftrightarrow\; \min_W\; I(u_1;\ldots;u_n)
The Infomax–Independence equivalence — Maximizing the joint output entropy = minimizing mutual information among the linear outputs = driving them toward statistical independence. One objective, three equivalent readings.
Open in Lab
Watch how the entropy decomposes as the unmixing matrix converges. The gap between marginal sum and joint entropy (= mutual information) shrinks toward zero as the algorithm finds the independent sources.
The demo wakes as you arrive…

Beyond separation: blind deconvolution

Bell and Sejnowski showed that the same principle extends to blind deconvolution: removing unknown echoes and reverberation from a signal. Instead of an instantaneous mixing matrix AA, the signal is convolved with an unknown (impulse response). The unmixing becomes a filter too — a set of delay-weighted connections — and the same entropy-maximization rule adapts the filter weights.

Think of it this way: in blind separation, each sensor receives a spatial mixture (superposition of sources at one instant). In blind deconvolution, a single sensor receives a temporal mixture (superposition of a source with its own delayed copies, i.e., echoes). Both are linear mixing — one across channels, the other across time — and information maximization handles both.

The full algorithm: step by step

Here is the complete Infomax ICA pipeline, from raw mixtures to recovered sources. The algorithm is surprisingly simple given what it achieves.

Infomax ICA — the complete algorithmpython

Simplified to show the idea — not the real implementation.

import numpy as np

def sigmoid(u):
    return 1.0 / (1.0 + np.exp(-u))

def infomax_ica(X, n_iter=200, lr=0.01):
    """
    X: (n_sources, n_samples) — mixed signals, zero-mean
    Returns W: unmixing matrix, S: estimated sources
    """
    n, T = X.shape
    W = np.eye(n)                       # start with identity

    for _ in range(n_iter):
        u = W @ X                       # linear unmixing
        y = sigmoid(u)                  # nonlinear squashing
        phi = 1 - 2 * y                 # score function φ(u)

        # Natural gradient update (Amari 1996):
        # avoids computing W⁻¹, converges faster
        dW = (np.eye(n) + phi @ u.T / T) @ W
        W += lr * dW

    S = W @ X                           # final source estimates
    return W, S

# That's it. The sigmoid + entropy maximization automatically
# finds statistically independent components — no knowledge
# of the mixing matrix or source distributions needed.

Why it mattered: from cocktail parties to brains

Infomax ICA did not just solve a signal processing problem — it created a bridge between information theory, neural networks, and neuroscience. The paper showed that a neural network can discover hidden structure in data by following a purely information-theoretic objective, with no labels, no teacher, and no knowledge of the data-generating process.

This unsupervised, information-driven approach directly inspired Olshausen and Field's sparse coding (1996), which applied similar principles to natural images and found that the learned features resemble the orientation-selective receptive fields of neurons in the primary visual cortex — suggesting that the brain itself might perform something like ICA.

  1. 1989

    Linsker's Infomax Principle

    Linsker proposed that neural networks should maximize mutual information between inputs and outputs. But in linear networks, this only yields PCA — no higher-order structure.

  2. 1994

    Comon formalizes ICA

    Comon gave Independent Component Analysis its name and formal framework using higher-order cumulants. But a practical neural-network algorithm was still missing.

  3. 1995

    Bell & Sejnowski — Infomax ICA

    The paper that connected information maximization to ICA through the sigmoid nonlinearity. Separated up to 10 mixed speakers and performed blind deconvolution.

  4. 1996

    Amari's natural gradient

    Amari replaced the costly matrix inverse in the learning rule with a multiplicative natural gradient, making Infomax fast enough for real-world applications.

  5. 1996

    Sparse coding — ICA meets the visual cortex

    Olshausen and Field showed that Infomax-like principles applied to natural images produce receptive fields resembling those in the visual cortex, bridging signal processing and neuroscience.

  6. 1999

    Extended Infomax

    Lee, Girolami and Sejnowski extended Infomax with switchable nonlinearities to handle both sub-Gaussian and super-Gaussian sources, making it universal.

  7. 1999

    FastICA

    Hyvärinen proposed FastICA — a fixed-point algorithm maximizing non-Gaussianity, faster than gradient-based methods for many problems. Together with Infomax, these became the two standard ICA algorithms.

Today, Infomax ICA is embedded in every major neuroscience toolbox (EEGLAB, MNE-Python, FieldTrip) for separating brain signals from artifacts. It is a standard preprocessing step in telecommunications for separating superimposed radio signals. The paper's deeper legacy is the demonstration that information theory provides a principled objective for unsupervised learning — an idea that echoes through variational autoencoders, contrastive learning, and modern representation learning.

CitationBell, A. J. and Sejnowski, T. J.. An Information-Maximization Approach to Blind Separation and Blind Deconvolution. Neural Computation, 1995.

Terms in this paper