Deep Learning2018advanced9 min read
Neural Ordinary Differential Equations
المعادلات التفاضلية العادية العصبية
Chen, T. Q. · Rubanova, Y. · Bettencourt, J. · Duvenaud, D. — NeurIPS
The problem
Deep residual networks (ResNets) stack dozens or hundreds of discrete layers, each applying a small transformation. Memory scales linearly with depth because must store every intermediate activation. The number of layers is fixed at design time, so the model cannot adapt its computation to easy versus hard inputs. And discrete normalizing flows require restricted architectures to keep the Jacobian determinant tractable.
The contribution
Instead of stacking discrete layers, parameterize the derivative of the with a : dh/dt = f(h(t), t, θ). The output is computed by a black-box . The computes exact gradients in constant memory — no need to store intermediate activations. The model adapts its depth per input (the solver takes more steps for harder inputs). Continuous normalizing flows follow naturally, allowing unrestricted architectures and exact log- training via the instantaneous .
The impact
(Best Paper, NeurIPS 2018) bridged differential equations and , spawning continuous normalizing flows (FFJORD), latent ODE models for irregular time series, and the theoretical framework that led to and score-based generative models. It showed that neural networks are not just stacks of layers but can be viewed as dynamical systems — an insight that reshaped how we think about depth, memory, and generative modeling.
A ResNet is like climbing a staircase: at each step you rise by a fixed height, and the architect decided exactly how many steps there are. You must remember every single step to retrace your path.
A Neural ODE replaces the staircase with a smooth ramp: instead of counting steps, you describe the slope at every point, and a solver rolls a ball along the ramp to reach the top. The ball can speed up on easy stretches and slow down on tricky curves — adapting its effort automatically. And to retrace the path, you just roll the ball backward from the top; no need to memorize every point along the way.
From ResNet to ODE: the continuous limit
The story begins with a residual network. In a ResNet, each layer computes:
This looks like Euler's method for solving an ODE — a technique from 1768 that approximates a smooth curve by taking small discrete steps. Each ResNet layer is one Euler step. Stack enough of them with small enough step sizes, and you approach a continuous transformation.
Chen et al. asked the natural question: what if we go all the way? Instead of prescribing a fixed number of layers, define the continuous dynamics directly:
Now a single neural network describes the velocity of the hidden state at every moment in time. To compute the output, hand this equation to an ODE solver that traces the trajectory from to — choosing its own step sizes, taking more steps when the landscape is complex and fewer when it's smooth.
The formulation: hidden state as a flowing river
Think of each data point's hidden state as a particle dropped into a river. The neural network defines the river's current — its — at every location and time. The particle follows the flow, and where it ends up is the network's output.
The key equation is the initial value problem:
The ODE solver is a black box: any standard numerical solver works — Euler, Runge-Kutta, Dormand-Prince. Modern adaptive solvers monitor their own error and choose step sizes automatically. This means the model's effective depth is not a you set — it emerges from the problem's complexity. Easy inputs need few solver steps; hard inputs get more.
The adjoint method: constant-memory backpropagation
Standard backpropagation through a deep network stores every layer's activations — memory grows with depth. For a continuous model with potentially thousands of solver steps, this is prohibitive.
The adjoint sensitivity method (from optimal control theory, Pontryagin 1962) offers a brilliant alternative. Define the adjoint state as the of the loss with respect to the hidden state at time :
This adjoint itself satisfies an ODE that runs backward in time:
The gradient with respect to the parameters is also computed by an integral running backward:
In practice, the forward state , the adjoint , and the parameter gradient are all integrated together as one augmented system — a single backward ODE solve replaces backpropagation through all layers.
Continuous normalizing flows: free-form generative modeling
A normalizing flow transforms a simple distribution (like a Gaussian) into a complex one by applying a chain of invertible transformations. The challenge: to compute the probability of a data point under the model, you need the determinant of each transformation's Jacobian — and for discrete flows, this constrains the architecture severely.
Neural ODEs unlock a continuous version. As a distribution evolves through the ODE, its log-density changes according to the instantaneous change of variables formula:
This is revolutionary: discrete normalizing flows required carefully designed invertible layers (coupling layers, autoregressive transforms). Continuous normalizing flows can use any neural network architecture for — no invertibility constraint, no special structure. The continuous dynamics guarantees invertibility automatically because smooth ODE flows never cross their own trajectories.
Grathwohl et al. (2018) later extended this idea into FFJORD, using the Hutchinson trace estimator to make the computation even more efficient.
Three advantages at a glance
Neural ODEs offer three properties that discrete networks cannot match:
- Constant memory: the adjoint method computes exact gradients without storing intermediate activations. Memory does not grow with the number of solver steps.
- Adaptive computation: the ODE solver automatically allocates more computation to harder inputs and less to easier ones. There is no fixed "depth" — the model decides on the fly.
- Free-form generative models: continuous normalizing flows impose no architectural constraints, only requiring the trace of the Jacobian instead of the full determinant.
The idea in code
Simplified to show the idea — not the real implementation.
import torch
import torch.nn as nn
from torchdiffeq import odeint
class ODEFunc(nn.Module):
"""The neural network f that defines dh/dt."""
def __init__(self, dim):
super().__init__()
self.net = nn.Sequential(
nn.Linear(dim, 64),
nn.Tanh(),
nn.Linear(64, dim),
)
def forward(self, t, h):
# t is the current time, h is the hidden state
return self.net(h)
class NeuralODE(nn.Module):
def __init__(self, func, t_span):
super().__init__()
self.func = func # the velocity field f
self.t_span = t_span # [t0, t1]
def forward(self, h0):
# The ODE solver is the "forward pass"
# It traces h from t0 to t1 using f
solution = odeint(
self.func, h0, self.t_span,
method='dopri5' # Dormand-Prince (adaptive)
)
return solution[-1] # h(t1)
# Usage: one network replaces an entire ResNet
func = ODEFunc(dim=64)
model = NeuralODE(func, t_span=torch.tensor([0.0, 1.0]))
h0 = torch.randn(32, 64) # batch of 32, dim 64
h1 = model(h0) # output — solver chose its own depthLatent ODEs: modeling irregular time series
A powerful application: combine a Neural ODE with a to model time series that arrive at irregular intervals — patient vitals recorded at random hospital visits, sensor readings at uneven timestamps.
The idea: encode the observations into a latent state , then let the Neural ODE evolve forward in continuous time. Because the ODE naturally handles any time interval, there is no need to discretize time into fixed bins or pad missing values. The reads off predictions at whatever times you need.
This Latent ODE model became a foundational tool for healthcare, climate science, and any domain where data arrives unpredictably.
Why it changed everything
2015
ResNet
He et al. introduce residual connections: h_{t+1} = h_t + f(h_t). Several researchers notice the link to Euler discretization of ODEs.
2018
Neural ODE (this paper)
Chen et al. go all the way: replace discrete layers with a continuous ODE, train with the adjoint method, and introduce continuous normalizing flows. Best Paper at NeurIPS.
2019
FFJORD
Grathwohl et al. extend continuous normalizing flows with the Hutchinson trace estimator, making them scalable and practical for density estimation.
2019
Latent ODEs
Rubanova et al. combine Neural ODEs with VAEs for irregularly-sampled time series, creating a natural framework for healthcare and sensor data.
2020
Score-based generative models
Song et al. unify score matching and diffusion with SDEs — a direct intellectual descendant of the continuous dynamics perspective Neural ODE introduced.
2022
Flow matching
Lipman et al. propose flow matching — a simpler training objective for continuous normalizing flows that avoids simulation during training entirely.
The thread from Neural ODE to flow matching is direct: continuous normalizing flows → score-based generative models → flow matching. Each builds on the fundamental insight that transformations between distributions can be modeled as continuous flows governed by differential equations. Modern image and video generation inherits this DNA.
CitationChen, Rubanova, Bettencourt, Duvenaud. Neural Ordinary Differential Equations. NeurIPS, 2018.
Terms in this paper
- Neural ODEالمعادلة التفاضلية العصبية
- Ordinary Differential Equationمعادلة تفاضلية عادية
- Adjoint Sensitivity Methodطريقة الحساسية المرافقة
- Continuous Normalizing Flowتدفق التسوية المستمر
- Residual Connectionالوصلة التجاوزية
- ODE Solverحالّ المعادلات التفاضلية
- Continuous-Depth Networkشبكة مستمرة العمق
- Euler Methodطريقة أويلر
- Backpropagationالتحديث التراجعي
- Latent Variableالمتغير الكامن