RealNVP and Glow
Transform complex data into simple bell curves by passing inputs through reversible blocks that scale and shift one half using the other half. Because each step is invertible with a triangular Jacobian, sampling and exact likelihood computation take only a single pass.
Why Does This Exist?
Generative models like GANs and VAEs force painful trade-offs. Generative Adversarial Networks optimize an adversarial game that lacks an explicit density function, making them notorious for mode collapse and impossible to evaluate for exact data log-likelihood. Variational Autoencoders optimize an evidence lower bound (), but the bound is loose, and their approximate posteriors blur high-frequency details. Autoregressive models provide exact likelihoods, but generating an image of resolution requires sequential forward passes—taking minutes per single sample.
RealNVP (Real-valued Non-Volume Preserving flows) and Glow solve this trilemma by constructing bijective, invertible neural transformations . By enforcing that is invertible and has a tractable Jacobian determinant, the change-of-variables theorem computes the exact log-likelihood in a single forward pass:
Sampling is equally fast: draw a random latent vector and evaluate the inverse mapping in a single forward step.
Think of It Like This
A two-chamber origami fold
Imagine folding a stamped sheet of paper into a neat envelope. In the first step, you hold the left half perfectly stationary against the desk while folding, stretching, and stamping new patterns onto the right half using instructions read directly from the left half.
Because the left half never changed, anyone opening the envelope can immediately inspect the left half, re-read the exact same origami instructions, and un-stretch the right half back into its pristine original state. Repeating this process while alternating which half stays still lets you fold arbitrary paper shapes into a standardized square without ever losing a single fiber of information.
How It Actually Works
Affine Coupling and Invertible 1x1 Convolutions
The core computational bottleneck in normalizing flows is the Jacobian determinant . For an input of dimension , computing the determinant of an unconstrained matrix costs arithmetic operations—prohibitive for high-dimensional images where .
RealNVP circumvents this with affine coupling layers. The input vector is partitioned along feature dimensions into two parts: and , where . The forward transformation is defined as:
Here, is a scale network and is a translation network. Crucially, and can be arbitrarily complex neural networks (such as deep residual convolutional networks) with no invertibility constraints of their own.
Because depends only on , and depends on and , the Jacobian matrix is lower triangular:
The determinant of a triangular matrix is the product of its diagonal elements. Its log-determinant reduces to a simple sum that evaluates in linear time:
The exact inverse requires no matrix inversion:
Glow enhances RealNVP by organizing transformations into a repeating Step of Flow containing three components:
- Actnorm (Activation Normalization): Replaces batch normalization with an affine transformation with per-channel scale and bias initialized to give zero mean and unit variance on the first batch. The log-determinant is .
- Invertible Convolution: Generalizes channel permutation. For a tensor of shape , a weight matrix mixes channels: . Parameterizing via LU decomposition () allows evaluating in time.
- Affine Coupling Layer: Applies the RealNVP split-scale-shift mechanism.
Worked Example
Let us trace a single affine coupling block on a 4-dimensional vector where and .
-
Split input into two halves:
-
Evaluate arbitrary neural networks and . Suppose the network outputs:
-
Compute scaling factors via exponentiation:
-
Compute output :
-
Compute the contribution to log-likelihood (the Jacobian log-determinant):
-
Invert from back to : The reconstruction is numerically exact.
Code
from typing import Tupleimport math
class AffineCouplingLayer: """Affine coupling layer operating on vectors with exact inversion."""
def __init__(self, dim: int) -> None: self.dim = dim self.split_dim = dim // 2 # Weights for simple linear mappings modeling s and t self.w_s = [0.2, -0.1] self.w_t = [-0.5, 0.3] self.b_t = [1.0, 0.0]
def _net(self, x1: list[float]) -> Tuple[list[float], list[float]]: # Scale s and translation t conditioned on x1 s = [self.w_s[0] * x1[0], self.w_s[1] * x1[1]] t = [self.w_t[0] * x1[0] + self.b_t[0], self.w_t[1] * x1[1] + self.b_t[1]] return s, t
def forward(self, x: list[float]) -> Tuple[list[float], float]: x1 = x[:self.split_dim] x2 = x[self.split_dim:] s, t = self._net(x1)
y1 = list(x1) y2 = [x2[i] * math.exp(s[i]) + t[i] for i in range(len(x2))] log_det = sum(s) return y1 + y2, log_det
def inverse(self, y: list[float]) -> list[float]: y1 = y[:self.split_dim] y2 = y[self.split_dim:] s, t = self._net(y1)
x1 = list(y1) x2 = [(y2[i] - t[i]) * math.exp(-s[i]) for i in range(len(y2))] return x1 + x2
layer = AffineCouplingLayer(dim=4)x_in = [2.0, -1.0, 3.0, 4.0]y_out, log_det_val = layer.forward(x_in)x_recovered = layer.inverse(y_out)
print("Forward y:", [round(v, 4) for v in y_out])# -> Forward y: [2.0, -1.0, 4.4755, 4.1207]print("Log det:", round(log_det_val, 4))# -> Log det: 0.5print("Reconstructed x:", [round(v, 4) for v in x_recovered])# -> Reconstructed x: [2.0, -1.0, 3.0, 4.0]Watch Out For
Numerical explosion from unconstrained exponential scaling
Symptom: During training, loss becomes NaN or sudden gradient explosions occur within the coupling layers after a few thousand steps.
In an affine coupling layer, . If the neural network outputs unbounded activations like , . The inverse operation multiplies by , resulting in severe numerical underflow and vanishing gradients.
The remedy is to pass the scale network outputs through a bounding activation: where , or clamp to . This guarantees stable conditioning throughout both forward and reverse paths.
The Quick Version
- Normalizing flows chain bijective layers to map complex empirical data distributions into simple isotropic Gaussians.
- RealNVP splits feature dimensions in half, applying identity to the first half and affine scale-and-shift to the second half.
- Triangular Jacobians reduce determinant computation from an intractable to linear time.
- Glow extends RealNVP with actnorm and invertible convolutions to mix channels without fixed permutations.
- Both exact log-likelihood evaluation and sample synthesis execute in a single parallelizable forward pass.