Deep Learning2018advanced9 min read

Neural Ordinary Differential Equations

المعادلات التفاضلية العادية العصبية

Chen, T. Q. · Rubanova, Y. · Bettencourt, J. · Duvenaud, D. — NeurIPS

The problem

Deep residual networks (ResNets) stack dozens or hundreds of discrete layers, each applying a small transformation. Memory scales linearly with depth because must store every intermediate activation. The number of layers is fixed at design time, so the model cannot adapt its computation to easy versus hard inputs. And discrete normalizing flows require restricted architectures to keep the Jacobian determinant tractable.

The contribution

Instead of stacking discrete layers, parameterize the derivative of the with a : dh/dt = f(h(t), t, θ). The output is computed by a black-box . The computes exact gradients in constant memory — no need to store intermediate activations. The model adapts its depth per input (the solver takes more steps for harder inputs). Continuous normalizing flows follow naturally, allowing unrestricted architectures and exact log- training via the instantaneous .

The impact

(Best Paper, NeurIPS 2018) bridged differential equations and , spawning continuous normalizing flows (FFJORD), latent ODE models for irregular time series, and the theoretical framework that led to and score-based generative models. It showed that neural networks are not just stacks of layers but can be viewed as dynamical systems — an insight that reshaped how we think about depth, memory, and generative modeling.

A ResNet is like climbing a staircase: at each step you rise by a fixed height, and the architect decided exactly how many steps there are. You must remember every single step to retrace your path.

A Neural ODE replaces the staircase with a smooth ramp: instead of counting steps, you describe the slope at every point, and a solver rolls a ball along the ramp to reach the top. The ball can speed up on easy stretches and slow down on tricky curves — adapting its effort automatically. And to retrace the path, you just roll the ball backward from the top; no need to memorize every point along the way.

From ResNet to ODE: the continuous limit

The story begins with a residual network. In a ResNet, each layer computes:

ht+1=ht+f(ht,θt)h_{t+1} = h_t + f(h_t, \theta_t)

This looks like Euler's method for solving an ODE — a technique from 1768 that approximates a smooth curve by taking small discrete steps. Each ResNet layer is one Euler step. Stack enough of them with small enough step sizes, and you approach a continuous transformation.

Chen et al. asked the natural question: what if we go all the way? Instead of prescribing a fixed number of layers, define the continuous dynamics directly:

dh(t)dt=f(h(t),t,θ)\frac{dh(t)}{dt} = f(h(t), t, \theta)

Now a single neural network ff describes the velocity of the hidden state at every moment in time. To compute the output, hand this equation to an ODE solver that traces the trajectory from t0t_0 to t1t_1 — choosing its own step sizes, taking more steps when the landscape is complex and fewer when it's smooth.

Open in Lab
Left: ResNet takes fixed discrete steps. Right: Neural ODE traces a smooth continuous curve. Toggle the number of Euler steps to see how discrete approaches continuous.
The demo wakes as you arrive…

The formulation: hidden state as a flowing river

Think of each data point's hidden state as a particle dropped into a river. The neural network ff defines the river's current — its — at every location and time. The particle follows the flow, and where it ends up is the network's output.

The key equation is the initial value problem:

h(t1)=h(t0)+∫t0t1f(h(t),t,θ) dth(t_1) = h(t_0) + \int_{t_0}^{t_1} f(h(t), t, \theta)\, dt
The Neural ODE forward pass — h(t₀) is the input · f is a neural network defining the velocity · the integral is evaluated by an ODE solver (e.g. Dormand–Prince) · h(t₁) is the output

The ODE solver is a black box: any standard numerical solver works — Euler, Runge-Kutta, Dormand-Prince. Modern adaptive solvers monitor their own error and choose step sizes automatically. This means the model's effective depth is not a you set — it emerges from the problem's complexity. Easy inputs need few solver steps; hard inputs get more.

Open in Lab
Watch how different ODE solvers (Euler, RK4, Adaptive) trace the same trajectory. Adaptive solvers take fewer steps on smooth regions and more on sharp turns.
The demo wakes as you arrive…

The adjoint method: constant-memory backpropagation

Standard backpropagation through a deep network stores every layer's activations — memory grows with depth. For a continuous model with potentially thousands of solver steps, this is prohibitive.

The adjoint sensitivity method (from optimal control theory, Pontryagin 1962) offers a brilliant alternative. Define the adjoint state as the of the loss with respect to the hidden state at time tt:

a(t)=∂L∂h(t)a(t) = \frac{\partial L}{\partial h(t)}

This adjoint itself satisfies an ODE that runs backward in time:

da(t)dt=−a(t)⊤∂f(h(t),t,θ)∂h\frac{da(t)}{dt} = -a(t)^{\top} \frac{\partial f(h(t), t, \theta)}{\partial h}
The adjoint ODE — gradients flowing backward — Run the dynamics backward from t₁ to t₀. No stored activations needed — reconstruct h(t) on the fly. Memory cost: O(1) regardless of solver steps.

The gradient with respect to the parameters θ\theta is also computed by an integral running backward:

dLdθ=−∫t1t0a(t)⊤∂f(h(t),t,θ)∂θ dt\frac{dL}{d\theta} = -\int_{t_1}^{t_0} a(t)^{\top} \frac{\partial f(h(t), t, \theta)}{\partial \theta}\, dt

In practice, the forward state h(t)h(t), the adjoint a(t)a(t), and the parameter gradient are all integrated together as one augmented system — a single backward ODE solve replaces backpropagation through all layers.

Open in Lab
Click "Forward" to see the ODE solver trace the trajectory. Then click "Backward" to watch the adjoint method compute gradients — note how no intermediate states are stored.
The demo wakes as you arrive…

Continuous normalizing flows: free-form generative modeling

A normalizing flow transforms a simple distribution (like a Gaussian) into a complex one by applying a chain of invertible transformations. The challenge: to compute the probability of a data point under the model, you need the determinant of each transformation's Jacobian — and for discrete flows, this constrains the architecture severely.

Neural ODEs unlock a continuous version. As a distribution evolves through the ODE, its log-density changes according to the instantaneous change of variables formula:

∂log⁡p(h(t))∂t=−tr ⁣(∂f∂h(t))\frac{\partial \log p(h(t))}{\partial t} = -\text{tr}\!\left(\frac{\partial f}{\partial h(t)}\right)
Instantaneous change of variables — The trace (sum of diagonal elements) of the Jacobian replaces the full determinant. This is much cheaper to compute and imposes no architectural constraints on f.

This is revolutionary: discrete normalizing flows required carefully designed invertible layers (coupling layers, autoregressive transforms). Continuous normalizing flows can use any neural network architecture for ff — no invertibility constraint, no special structure. The continuous dynamics guarantees invertibility automatically because smooth ODE flows never cross their own trajectories.

Grathwohl et al. (2018) later extended this idea into FFJORD, using the Hutchinson trace estimator to make the computation even more efficient.

Open in Lab
Watch a simple Gaussian distribution transform into a complex shape through continuous flow. The trajectories never cross — invertibility is guaranteed.
The demo wakes as you arrive…

Three advantages at a glance

Neural ODEs offer three properties that discrete networks cannot match:

  • Constant memory: the adjoint method computes exact gradients without storing intermediate activations. Memory does not grow with the number of solver steps.
  • Adaptive computation: the ODE solver automatically allocates more computation to harder inputs and less to easier ones. There is no fixed "depth" — the model decides on the fly.
  • Free-form generative models: continuous normalizing flows impose no architectural constraints, only requiring the trace of the Jacobian instead of the full determinant.
Open in Lab
Compare memory usage of ResNet (grows with depth) versus Neural ODE (constant). Drag the depth slider to see the difference.
The demo wakes as you arrive…

The idea in code

Neural ODE forward pass with torchdiffeqpython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn as nn
from torchdiffeq import odeint

class ODEFunc(nn.Module):
    """The neural network f that defines dh/dt."""
    def __init__(self, dim):
        super().__init__()
        self.net = nn.Sequential(
            nn.Linear(dim, 64),
            nn.Tanh(),
            nn.Linear(64, dim),
        )

    def forward(self, t, h):
        # t is the current time, h is the hidden state
        return self.net(h)

class NeuralODE(nn.Module):
    def __init__(self, func, t_span):
        super().__init__()
        self.func = func        # the velocity field f
        self.t_span = t_span    # [t0, t1]

    def forward(self, h0):
        # The ODE solver is the "forward pass"
        # It traces h from t0 to t1 using f
        solution = odeint(
            self.func, h0, self.t_span,
            method='dopri5'  # Dormand-Prince (adaptive)
        )
        return solution[-1]  # h(t1)

# Usage: one network replaces an entire ResNet
func = ODEFunc(dim=64)
model = NeuralODE(func, t_span=torch.tensor([0.0, 1.0]))
h0 = torch.randn(32, 64)  # batch of 32, dim 64
h1 = model(h0)            # output — solver chose its own depth

Latent ODEs: modeling irregular time series

A powerful application: combine a Neural ODE with a to model time series that arrive at irregular intervals — patient vitals recorded at random hospital visits, sensor readings at uneven timestamps.

The idea: encode the observations into a latent state z(t0)z(t_0), then let the Neural ODE evolve zz forward in continuous time. Because the ODE naturally handles any time interval, there is no need to discretize time into fixed bins or pad missing values. The reads off predictions at whatever times you need.

This Latent ODE model became a foundational tool for healthcare, climate science, and any domain where data arrives unpredictably.

Open in Lab
Irregular observations (dots) are encoded into a latent state. The ODE evolves it continuously, producing predictions at any desired time.
The demo wakes as you arrive…

Why it changed everything

  1. 2015

    ResNet

    He et al. introduce residual connections: h_{t+1} = h_t + f(h_t). Several researchers notice the link to Euler discretization of ODEs.

  2. 2018

    Neural ODE (this paper)

    Chen et al. go all the way: replace discrete layers with a continuous ODE, train with the adjoint method, and introduce continuous normalizing flows. Best Paper at NeurIPS.

  3. 2019

    FFJORD

    Grathwohl et al. extend continuous normalizing flows with the Hutchinson trace estimator, making them scalable and practical for density estimation.

  4. 2019

    Latent ODEs

    Rubanova et al. combine Neural ODEs with VAEs for irregularly-sampled time series, creating a natural framework for healthcare and sensor data.

  5. 2020

    Score-based generative models

    Song et al. unify score matching and diffusion with SDEs — a direct intellectual descendant of the continuous dynamics perspective Neural ODE introduced.

  6. 2022

    Flow matching

    Lipman et al. propose flow matching — a simpler training objective for continuous normalizing flows that avoids simulation during training entirely.

The thread from Neural ODE to flow matching is direct: continuous normalizing flows → score-based generative models → flow matching. Each builds on the fundamental insight that transformations between distributions can be modeled as continuous flows governed by differential equations. Modern image and video generation inherits this DNA.

CitationChen, Rubanova, Bettencourt, Duvenaud. Neural Ordinary Differential Equations. NeurIPS, 2018.

Terms in this paper