Continuous Normalizing Flows
Instead of stacking discrete invertible neural network layers, continuous normalizing flows let data points drift smoothly along velocity vector fields defined by ordinary differential equations. This continuous flow turns arbitrary data distributions into simple Gaussians while exact probabilities are tracked along the path.
Why Does This Exist?
Discrete normalizing flows like RealNVP and Glow obtain tractable likelihoods by constraining layer architectures. To keep Jacobian matrices triangular and readily invertible, models split features into partitions or restrict channel transformations to invertible matrices. These structural handcuffs limit expressive power per layer, forcing practitioners to stack dozens of coupling blocks and store every intermediate activation tensor in GPU memory during backpropagation.
Continuous Normalizing Flows (CNFs), built on Neural Ordinary Differential Equations (Neural ODEs), remove these architectural restrictions. Rather than defining transformations through a discrete chain of layers , a CNF defines the transformation through a continuous-time differential equation:
Because any smooth, Lipschitz-continuous vector field produces a bijective, unique trajectory that never crosses itself, the network requires zero architectural constraints. It can be an arbitrary multi-layer perceptron or residual convolutional network. Furthermore, the instantaneous change-of-variables theorem replaces the difficult matrix determinant with a simple matrix trace.
Think of It Like This
Floating downstream through a mountain river
Imagine placing a fleet of toy boats across a wide lake and opening floodgates into a twisting canyon river. Each boat's position changes continuously over time, steered purely by the local water velocity at its exact location.
Because two boats cannot occupy the same water molecule at the same instant, their paths never intersect. As the river flows from time to , the scattered boats gradually gather into a predictable, compact basin. If you want to know where a boat started, simply reverse the river's flow and watch it float backward along its exact trajectory. You never need separate rules for moving forward versus backward—the water's continuous velocity field handles both.
How It Actually Works
Instantaneous Change of Variables and Adjoint Sensitivity
In discrete flows, the density transform requires the full Jacobian determinant:
In a continuous flow governed by , the continuous counterpart is the instantaneous change of variables formula:
Integrating this differential equation from data space to prior space yields the total change in log-density:
Computing the exact trace of the Jacobian matrix normally takes operations using vector-Jacobian products. FFJORD (Free-form Jacobian of Reversible Dynamics) scales this to high dimensions using the Hutchinson trace estimator:
where is a random noise vector drawn from a standard normal or Rademacher distribution ( with equal probability). By computing , automatic differentiation evaluates the trace in operations using a single backward pass.
To optimize parameters without storing all intermediate states along the integration path, the adjoint sensitivity method defines an adjoint state:
The adjoint satisfies its own reverse-time ordinary differential equation:
The loss gradient with respect to model parameters is obtained by integrating backwards from to :
This guarantees memory consumption with respect to integration depth, eliminating the activation caching bottleneck of discrete deep networks.
Worked Example
Consider a 2-dimensional continuous flow from to . Let the observed data vector be .
The parameterized vector field is:
The Jacobian is diagonal:
Let us numerically integrate using an explicit Euler solver with step size (2 steps):
-
First step ():
- Velocity at :
- Position update:
- Instantaneous trace:
- Density update:
-
Second step ():
- Velocity at :
- Position update:
- Density update:
-
Total Log-Likelihood:
- Base prior log-density for standard 2D Gaussian :
- Combined data log-likelihood:
Code
from typing import Callable, Tupleimport math
Vector = list[float]
def vector_field(z: Vector, t: float) -> Tuple[Vector, float]: """Evaluates dz/dt = f(z, t) and the exact divergence (trace of Jacobian).""" dz1 = -0.5 * z[0] dz2 = -0.8 * z[1] + 0.1 * t trace = -0.5 + (-0.8) return [dz1, dz2], trace
def euler_cnf_step( z: Vector, t: float, dt: float, vf: Callable[[Vector, float], Tuple[Vector, float]]) -> Tuple[Vector, float]: """Single forward integration step tracking state and accumulated log-density.""" dz, trace = vf(z, t) z_next = [z[i] + dt * dz[i] for i in range(len(z))] d_log_p = -trace * dt return z_next, d_log_p
# Simulate trajectory from t=0.0 to t=1.0 with dt=0.5z_state = [1.0, 2.0]total_d_log_p = 0.0t_current = 0.0dt = 0.5
for step in range(2): z_state, delta_log_p = euler_cnf_step(z_state, t_current, dt, vector_field) total_d_log_p += delta_log_p t_current += dt
# Compute prior log density under 2D standard Gaussiannorm_sq = z_state[0] ** 2 + z_state[1] ** 2log_p_prior = -math.log(2.0 * math.pi) - 0.5 * norm_sqlog_p_x = log_p_prior + total_d_log_p
print("Terminal state z(1):", [round(v, 4) for v in z_state])# -> Terminal state z(1): [0.5625, 0.745]print("Log prior p(z):", round(log_p_prior, 4))# -> Log prior p(z): -2.2736print("Accumulated divergence integral:", round(total_d_log_p, 4))# -> Accumulated divergence integral: 1.3print("Total data log-likelihood log p(x):", round(log_p_x, 4))# -> Total data log-likelihood log p(x): -0.9736Watch Out For
Stiff ODE trajectories causing runaway function evaluations (NFE)
Symptom: Training time per epoch multiplies tenfold over several hundred steps, with adaptive ODE solvers (like Dormand-Prince / dopri5) hanging indefinitely.
As a continuous flow neural network trains without regularization, its learned vector field can develop extreme velocity gradients and turbulent swirls. Adaptive solvers must shrink their integration step sizes toward zero to satisfy local error tolerances, causing the Number of Function Evaluations (NFE) to explode from 20 to over 1,000 per sample.
The solution is trajectory regularization. Penalize both the Frobenius norm of the Jacobian and the kinetic energy of the path: . This straightens the flow paths, enforces smooth dynamics, and keeps NFE small throughout training.
The Quick Version
- Continuous Normalizing Flows replace discrete layer stacks with neural vector fields integrated over continuous time.
- Any standard feedforward neural network can define the dynamics; invertibility is guaranteed by ODE trajectory uniqueness.
- The instantaneous change of variables simplifies log-determinants into the trace of the Jacobian.
- Hutchinson's trace estimator approximates the trace in operations using random projection vectors.
- The adjoint sensitivity method enables exact backpropagation backwards in time with constant memory.