Neuroscience2020intermediate19 min read

Growing Neural Cellular Automata

كائنٌ ينمو من خليّة واحدة ويرمّم نفسه

Mordvintsev, A. · Randazzo, E. · Niklasson, E. · Levin, M. — Distill

The problem

Biology builds bodies without a blueprint reader. Cells hold identical genomes, sense only their immediate surroundings, and still arrive at a specific anatomy — then defend it. We know many genes that are required for regeneration, but not an algorithm that is sufficient for it. Cellular automata have modelled local interaction since von Neumann, and hand-designed or evolved rules have produced beautiful patterns, but nobody could take an arbitrary target shape and obtain a local rule that grows it and then holds it.

The contribution

A cellular automaton whose is a small differentiable network — about eight thousand parameters — shared by every cell and applied to a sixteen-channel state vector. Perception is a fixed Sobel , the rule emits a residual increment, cells update on independent coin flips, and cells without living neighbours are zeroed. Training is against a pixel loss. Two training choices do the real work: a pool of past states that turns the target into an , and damage applied to pool samples that widens the basin around it.

The impact

It showed that a completely decentralised system with no global state can be programmed by gradient descent to reach and defend a specific configuration, and to recover from failures it never saw. The line of work continued through the Differentiable Self-organizing Systems thread, and it sits close to graph neural networks, swarm robotics, and any setting where identical agents must agree on a global outcome using local messages alone.

A salamander loses a leg and grows it back — bone, muscle, nerve, in the right order, at the right length, and then it stops. No cell in that leg has ever seen a salamander. Each one knows the chemistry at its own surface and which neighbours are pressing against it.

The puzzle is not how a cell divides. It is how a crowd of cells, each blind past its own membrane, agrees on where the leg ends.

The body that builds itself

Almost every multicellular organism starts as one cell, and that cell's descendants arrange themselves into the same anatomy every time. This is morphogenesis, and it is the clearest case of self-organisation we have: the arrangement is not imposed from outside, it falls out of local decisions made in parallel by cells that share a genome and differ only in what they have sensed and stored.

What makes it more than impressive is that it is defended. Cut an early mammalian embryo in half and each half makes a whole individual. Damage a salamander's eye, limb, or parts of its brain and it is rebuilt. Something in the cell collective represents the finished shape well enough to notice a deviation and correct it — and it does so without any single cell holding that shape.

The open question is not which genes are required; many are known. It is which algorithm would be sufficient — what a cell would have to compute, from its own surroundings alone, for a specific anatomy to follow.

Cellular automata have been the natural way to ask that question since von Neumann used them for self-replication. A grid of cells, the same rule applied to every cell at every step, each new state depending only on a small neighbourhood. Turing's reaction-diffusion patterns, Conway's Game of Life, the Gray-Scott model, and more recent continuous relatives like SmoothLife and Lenia all show how much behaviour a handful of local constraints can produce.

The trouble is the direction of travel. Those rules were written, or searched for, and then run to see what appeared. Going the other way — naming a shape first and obtaining a local rule that produces it — has been attempted with evolutionary search, notably Miller's self-repairing French flag, but it does not reach an arbitrary target.

This paper changes one thing about the setup, and that one thing opens the door: it makes the cell state continuous, and therefore the update rule differentiable. A differentiable rule can be fitted with the same machinery as any other network, which means the target shape can be stated as a loss instead of being designed into the rule by hand.

A cell is sixteen numbers

Each grid cell holds a vector of sixteen real values. The first three are the colour a reader sees. The fourth is alpha, and it carries a definition rather than an appearance: a cell with alpha above 0.1 is mature, its immediate neighbours are growing, and anything further out is empty and has every channel forced to zero on every step. So "which cells exist" is not tracked anywhere. It is re-derived from the state itself, continuously.

The remaining twelve channels are given no meaning at all. Nothing in the loss mentions them. They are the cell's private memory and its only means of signalling, and the rule is free to use them as chemical concentrations, as a position estimate, as a clock, or as anything else that helps. What the cells eventually agree to put there is learned, and it is where the whole coordination lives.

This is where the biological analogy is tightest. Every cell runs an identical rule, exactly as every cell in an organism carries the same genome. Two cells behave differently only because their state vectors differ — because of what they have received, emitted and stored.

One rule, four phases

The update must apply the same operation to every cell and depend only on a 3×3 neighbourhood, which is the definition of a convolution. The paper splits it into four phases, and the first is the one with an argument behind it.

Perception is a fixed 3×3 convolution — classical Sobel filters estimating the partial derivatives of every channel in x and y. The kernel is not learned, and it could have been. The reason is biological: real cells mostly read chemical gradients, so each cell is given two gradient estimates per channel plus its own state, concatenated into a forty-eight-dimensional percept. Fixing the kernel also fixes what "local" means for this model, which later turns out to be a lever rather than a constraint.

Open in Lab
Walk the four phases on a live neighbourhood, then press "strip the neighbours" and walk them again. Phases one to three still compute a perfectly reasonable increment for the centre cell; phase four throws it away and zeroes the cell, because a cell with no mature neighbour is not a cell at all.
The demo wakes as you arrive…
Perception, and the update it feedspython

Simplified to show the idea — not the real implementation.

def perceive(state_grid):
    # A fixed kernel, not a learned one: cells read gradients,
    # the way real cells read chemical concentration gradients.
    sobel_x = [[-1, 0, +1],
               [-2, 0, +2],
               [-1, 0, +1]]
    sobel_y = transpose(sobel_x)
    grad_x = conv2d(sobel_x, state_grid)
    grad_y = conv2d(sobel_y, state_grid)
    # 16 own channels + 16 x-gradients + 16 y-gradients = 48.
    return concat(state_grid, grad_x, grad_y, axis=2)

def update(perception_vector):
    # Runs per cell. The same weights in every cell: ~8k in total.
    x = dense(perception_vector, output_len=128)
    x = relu(x)
    # Zero-initialised, so an untrained rule does nothing at all
    # instead of destroying the seed on the very first step.
    ds = dense(x, output_len=16, weights_init=0.0)
    return ds  # an increment, never a replacement state

The percept passes through a 1×1 convolution, a ReLU, and a second 1×1 convolution — about eight thousand parameters in total, and because the convolutions are 1×1 they are really a small network applied independently at every position. Two details matter more than the architecture. The output is an increment added to the existing state, in the manner of a residual network, so a cell's default is continuity. And the last layer is initialised to zero, so before training the rule does nothing at all: the seed survives step one instead of being destroyed by noise. There is no ReLU on the output, because an increment has to be able to subtract.

The third phase removes the clock. A conventional automaton updates every cell at once, which assumes something is synchronising them, and a self-organising system should not need that. Instead each cell updates with probability one half on each step, independently of the rest. Mechanically this is dropout applied to the update vector rather than to weights.

The fourth phase is the alive mask already described. Note the order: the rule computes, the coin decides whether the computation is applied, and only then is existence re-checked.

st+1(x)=m(x) [ st(x)+b(x)⋅fθ(Pst,x) ]s_{t+1}(x) = m(x)\,\big[\, s_t(x) + b(x)\cdot f_\theta(P s_t, x) \,\big]
One step, with both masks made explicit — The cell's next state is its current state plus an increment, and the two masks sit on either side of that sum. The Bernoulli variable is the coin the cell flips to decide whether it acts this step; it is drawn afresh per cell and per step, so no two cells are synchronised. The other mask is existence — one where some mature cell sits in the neighbourhood, zero otherwise — and it multiplies the whole state rather than the increment. The perception operator is the fixed convolution, and the learned rule is identical everywhere; the only thing that distinguishes two cells is the state that rule is reading.

Experiment 1 — a rule that reaches the shape

The first experiment is deliberately naive. Initialise the grid to zeros except one seed cell at the centre, with every channel except RGB set to one. Run the update rule for a number of steps sampled uniformly from 64 to 96 — random, so the pattern has to be present across a window rather than at one instant. Apply a pixel-wise L2 loss between the RGBA channels and the target image. Then differentiate through the whole unrolled sequence, which is backpropagation through time, exactly as for a recurrent network.

It works, and it works on the first honest attempt: a single cell becomes a lizard, a smiley, an eye, within a few thousand training iterations. One practical caveat is flagged by the authors — training destabilises in the later stages, with the loss jumping suddenly — and they fix it by normalising each parameter's gradient separately before the update.

So the paper now has a growth function that reaches the target. The rest of the paper is about how far that is from having one that means it.

Open in Lab
Let it grow from the single seed, then tap somewhere else on the grid to plant a second one. A whole second organism appears, and the two settle a border between them, though nothing in the rule counts bodies. Then drag the update probability down: the same shape arrives, just later, because no cell was ever waiting on a clock.
The demo wakes as you arrive…

Now run the trained rule past the horizon it was trained on. The loss was applied somewhere between step 64 and step 96 and never afterwards, so anything the rule does at step 200 is unconstrained — and it shows. Across training runs with identical settings, some patterns decay to nothing, some keep growing and lose the shape entirely, and some happen to sit almost still.

This is not a bug to be patched. It is the difference between two requirements that a loss measured inside one window cannot tell apart: be at the target then, and be at the target and stay there. Every rule that satisfies the first is scored identically, whether or not it satisfies the second, so which of them training lands on is left to chance.

Experiment 2 — what persists, exists

Reframe the grid as a dynamical system. Every cell is a system with the same dynamics, coupled to its neighbours, and training adjusts those shared dynamics. Experiment 1 obtained a trajectory: starting from the seed, the system passes through the target. What is wanted is stronger — that the target be an attractor, a state the system returns to after being pushed off it.

The direct way to train for that is to unroll much longer and apply the loss repeatedly along the way, so the rule is shaped to come back to the target from wherever it has wandered. The cost is prohibitive: every intermediate activation in the episode must be kept for the backward pass, so memory grows with the length of the unroll.

The paper's alternative gets the same effect at the price of a short unroll. Keep a pool of states, initially filled with copies of the seed. Sample a batch from the pool, train the usual short episode on it, then write the resulting final states back into the pool. The next time those states are drawn they are the starting point, so the effective unroll stretches across training steps instead of within one.

The sample pool, and the one line that keeps it honestpython

Simplified to show the idea — not the real implementation.

def pool_training():
    seed = zeros(64, 64, 16)
    seed[64 // 2, 64 // 2, 3:] = 1.0   # alpha and hidden channels on
    pool = [seed] * 1024

    for i in range(training_iterations):
        idxs, batch = pool.sample(32)
        batch = sort_desc(batch, loss(batch))
        # Reseed the WORST sample, not a random one: this both
        # preserves growth-from-seed and evicts junk from the pool.
        batch[0] = seed
        outputs, loss = train(batch, target)
        # Today's endings are tomorrow's beginnings.
        pool[idxs] = outputs
Open in Lab
Scrub to around the step where the loss was applied and the two rules are hard to separate — which is exactly what a loss measured at that moment can see. Then keep scrubbing. One of them has no reason to stop and does not stop; the other has settled into a state it comes back to.
The demo wakes as you arrive…

Two details in that loop repay attention. Replacing one sample with a fresh seed on every batch prevents the equivalent of catastrophic forgetting: without it the pool drifts towards late-stage states, and the rule slowly loses the ability to grow from a single cell at all. And replacing the highest-loss sample specifically, rather than a random one, makes early training much more stable, because the pool is continuously cleaned of the junk states that a half-trained rule produces.

The pool also changes what the rule is being asked to do, and it changes it over time. Early on the pool is full of broken, half-formed, incorrect states, and the rule is being trained to recover from them. Later, as the rule improves, the pool fills with states that are nearly the target, and training shifts to polishing. Nobody designed that curriculum; it is a side effect of feeding the system its own output.

Experiment 3 — regeneration, mostly unasked for

Before training anything new, the authors damage the models from Experiment 2 and watch. Each model develops for a hundred steps, then the pattern is cut: each half removed in turn, and a square taken out of the middle. Nothing in their training ever mentioned damage.

Several of them repair themselves anyway. The lizard in particular grows back convincingly. The reason is not mysterious once the attractor framing is in place: training for persistence selects rules that grow, stabilise and do not self-destruct, and a damaged state sits near enough to states those rules already know how to handle that the same dynamics carry it home.

But it is a tendency, not a guarantee, and the failures are worth naming because they are recognisably biological. Some models answer a wound with uncontrolled growth. Some are so over-stabilised that they ignore it entirely. Some destroy themselves — especially under heavier damage.

To get regeneration reliably, widen the on purpose. The change is three lines in the training loop: sample from the pool as before, then zero a randomly placed circular region in a few of the lowest-loss samples before the training step runs. The rule now has to get from damaged states back to the target — and those damaged states are drawn from the pool, which means they are damaged versions of states the rule actually reaches.

The recovery that results is both stronger and broader than the damage seen in training. The models are damaged with erased circles, and they regenerate from rectangular cuts they never met. That generalisation is the paper's clearest evidence that what was learned is a basin, and not a lookup table of remembered repairs.

Open in Lab
Cut a small disc from any of the three regimes and most of them close it. Now widen the wound and take out the centre. The rule trained only to hold its shape settles for a ragged approximation and stops there, while the one trained on wounds rebuilds cleanly — the gap the paper reports is widest exactly here, in the middle.
The demo wakes as you arrive…

Experiment 4 — turn the sensors and the body turns

Recall that perception was fixed rather than learned: two Sobel filters estimating gradients along two perpendicular axes. Read that as biology and each cell has two sensors pointed along those axes, each reporting how a concentration changes along its own direction. So what happens if the sensors are pointed somewhere else?

Rotating them is a rotation applied to the pair of Sobel kernels, and it is done after training, to a finished model, with every weight untouched. The pattern then grows rotated by the chosen angle. Nothing was retrained, and no cell was ever told what orientation means.

It is worth being precise about why this is more than a change of reference frame. On a continuous plane it would be nearly trivial. On a pixel lattice it is not: rotating a grid of pixels is not a bijection, and normally requires interpolating between cells that no longer line up. That the pattern survives the operation says the tolerate conditions well outside the ones they were fitted under.

[KxKy]=[cos⁡θ−sin⁡θsin⁡θcos⁡θ][SobelxSobely]\begin{bmatrix} K_x \\ K_y \end{bmatrix} = \begin{bmatrix} \cos\theta & -\sin\theta \\ \sin\theta & \cos\theta \end{bmatrix} \begin{bmatrix} Sobel_x \\ Sobel_y \end{bmatrix}
Rotating the axes a cell measures along — The two perception kernels are treated as a pair of vectors and turned through the angle as a unit. At zero these are the ordinary horizontal and vertical Sobel filters; at any other angle each kernel becomes a blend of both. Nothing downstream changes — the same forty-eight numbers arrive at the same learned rule, and the rule cannot tell that the frame it is reading in has moved.
Open in Lab
Turn the dial and watch both overlap figures. The match against the turned body barely moves while the match against the upright body collapses, which is how you know the organism rotated rather than degraded. The kernel underneath shows what actually changed: nine numbers, and not one weight.
The demo wakes as you arrive…

Why no cell is allowed a clock

The stochastic update is easy to read as a regularisation trick, and it is not one. Updating every cell simultaneously requires a global clock, and a global clock is exactly the kind of centralised structure the paper is trying to do without. If the result depended on perfect synchrony it would be a much weaker result, because no physical collective of cells or robots has that.

So each cell waits an independent random interval, modelled as a per-cell mask that zeroes the update with probability one half. Go back to the growth grid above and push the update probability down: the same body arrives, only after more steps. Push it all the way up and growth becomes a clean synchronised wavefront. The shape does not care.

This is also what makes the closing speculation coherent. The authors sketch a physical version — a grid of tiny independent computers, roughly ten kilobytes of read-only memory each to hold the rule, a few hundred bytes of working memory, a coloured diode, and the ability to pass sixteen numbers to their neighbours. That device is only buildable because nothing in the model needs the whole grid to step at the same moment.

What it does not solve

What it opened

The relationship to convolutional networks is not an analogy but an identity: in the authors' own phrasing, the model could equally be described as a recurrent residual convolutional network with per-pixel dropout. The fixed 3×3 perception is the same commitment to a that the neocognitron made, arrived at from the opposite direction — here is not an efficiency argument, it is the thing being modelled.

Two further connections are worth holding onto. The target-as-attractor framing is the same move that Hopfield networks make, with the basin trained rather than constructed, and with the stored state spread across a grid rather than across weights. And training on deliberately corrupted inputs to obtain robust recovery is the strategy behind denoising autoencoders, applied here to a spatial system in which the corruption is a wound.

The twelve hidden channels are a with no assigned meaning, invented by the cells because coordination demanded it. And the learned update is a dynamics function of exactly the kind that world models fit, here shared by thousands of agents at once. Downstream, the thread continued into self-classifying grids and into the wider family of message-passing systems — and the swarm robotics the discussion points at, from Kilobots to mergeable nervous systems, still has its collective behaviour designed by hand.

CitationMordvintsev, Randazzo, Niklasson, Levin. Growing Neural Cellular Automata. Distill, 2020.

Terms in this paper