Skip to content
AI360Xpert
Beta

RealNVP and Glow

Transform complex data into simple bell curves by passing inputs through reversible blocks that scale and shift one half using the other half. Because each step is invertible with a triangular Jacobian, sampling and exact likelihood computation take only a single pass.

RealNVP affine coupling block splitting features for triangular Jacobian log-determinants alongside Glow 1x1 convolutions and actnorm steps
RealNVP affine coupling block splitting features for triangular Jacobian log-determinants alongside Glow 1x1 convolutions and actnorm steps

Why Does This Exist?

Generative models like GANs and VAEs force painful trade-offs. Generative Adversarial Networks optimize an adversarial game that lacks an explicit density function, making them notorious for mode collapse and impossible to evaluate for exact data log-likelihood. Variational Autoencoders optimize an evidence lower bound (ELBO\text{ELBO}), but the bound is loose, and their approximate posteriors blur high-frequency details. Autoregressive models provide exact likelihoods, but generating an image of resolution 256×256×3256 \times 256 \times 3 requires 196,608196{,}608 sequential forward passes—taking minutes per single sample.

RealNVP (Real-valued Non-Volume Preserving flows) and Glow solve this trilemma by constructing bijective, invertible neural transformations fθ:X→Zf_\theta: \mathcal{X} \to \mathcal{Z}. By enforcing that fθf_\theta is invertible and has a tractable Jacobian determinant, the change-of-variables theorem computes the exact log-likelihood in a single forward pass:

log⁡pX(x)=log⁡pZ(fθ(x))+log⁡∣det⁡(∂fθ(x)∂x)∣\log p_X(x) = \log p_Z(f_\theta(x)) + \log \left| \det \left( \frac{\partial f_\theta(x)}{\partial x} \right) \right|

Sampling is equally fast: draw a random latent vector z∼N(0,I)z \sim \mathcal{N}(0, I) and evaluate the inverse mapping x=fθ−1(z)x = f_\theta^{-1}(z) in a single forward step.

Think of It Like This

A two-chamber origami fold

Imagine folding a stamped sheet of paper into a neat envelope. In the first step, you hold the left half perfectly stationary against the desk while folding, stretching, and stamping new patterns onto the right half using instructions read directly from the left half.

Because the left half never changed, anyone opening the envelope can immediately inspect the left half, re-read the exact same origami instructions, and un-stretch the right half back into its pristine original state. Repeating this process while alternating which half stays still lets you fold arbitrary paper shapes into a standardized square without ever losing a single fiber of information.

How It Actually Works

Affine Coupling and Invertible 1x1 Convolutions

The core computational bottleneck in normalizing flows is the Jacobian determinant det⁡(J)\det(J). For an input of dimension DD, computing the determinant of an unconstrained D×DD \times D matrix costs O(D3)\mathcal{O}(D^3) arithmetic operations—prohibitive for high-dimensional images where D≈105D \approx 10^5.

RealNVP circumvents this with affine coupling layers. The input vector x∈RDx \in \mathbb{R}^D is partitioned along feature dimensions into two parts: x1:dx_{1:d} and xd+1:Dx_{d+1:D}, where d<Dd < D. The forward transformation y=f(x)y = f(x) is defined as:

y1:d=x1:dy_{1:d} = x_{1:d} yd+1:D=xd+1:D⊙exp⁡(s(x1:d))+t(x1:d)y_{d+1:D} = x_{d+1:D} \odot \exp(s(x_{1:d})) + t(x_{1:d})

Here, s:Rd→RD−ds: \mathbb{R}^d \to \mathbb{R}^{D-d} is a scale network and t:Rd→RD−dt: \mathbb{R}^d \to \mathbb{R}^{D-d} is a translation network. Crucially, ss and tt can be arbitrarily complex neural networks (such as deep residual convolutional networks) with no invertibility constraints of their own.

Because y1:dy_{1:d} depends only on x1:dx_{1:d}, and yd+1:Dy_{d+1:D} depends on x1:dx_{1:d} and xd+1:Dx_{d+1:D}, the Jacobian matrix is lower triangular:

J=∂y∂x=[Id0∂yd+1:D∂x1:ddiag⁡(exp⁡(s(x1:d)))]J = \frac{\partial y}{\partial x} = \begin{bmatrix} I_d & 0 \\ \frac{\partial y_{d+1:D}}{\partial x_{1:d}} & \operatorname{diag}(\exp(s(x_{1:d}))) \end{bmatrix}

The determinant of a triangular matrix is the product of its diagonal elements. Its log-determinant reduces to a simple sum that evaluates in O(D)\mathcal{O}(D) linear time:

log⁡∣det⁡J∣=∑i=1D−ds(x1:d)i\log |\det J| = \sum_{i=1}^{D-d} s(x_{1:d})_i

The exact inverse requires no matrix inversion:

x1:d=y1:dx_{1:d} = y_{1:d} xd+1:D=(yd+1:D−t(y1:d))⊙exp⁡(−s(y1:d))x_{d+1:D} = (y_{d+1:D} - t(y_{1:d})) \odot \exp(-s(y_{1:d}))

Glow enhances RealNVP by organizing transformations into a repeating Step of Flow containing three components:

  1. Actnorm (Activation Normalization): Replaces batch normalization with an affine transformation yi,j=s⊙xi,j+by_{i,j} = s \odot x_{i,j} + b with per-channel scale ss and bias bb initialized to give zero mean and unit variance on the first batch. The log-determinant is h⋅w⋅∑log⁡∣s∣h \cdot w \cdot \sum \log |s|.
  2. Invertible 1×11 \times 1 Convolution: Generalizes channel permutation. For a tensor of shape h×w×ch \times w \times c, a weight matrix W∈Rc×cW \in \mathbb{R}^{c \times c} mixes channels: y=xWy = x W. Parameterizing WW via LU decomposition (W=PL(U+diag⁡(s))W = P L (U + \operatorname{diag}(s))) allows evaluating log⁡∣det⁡W∣=h⋅w⋅∑log⁡∣s∣\log |\det W| = h \cdot w \cdot \sum \log |s| in O(c)\mathcal{O}(c) time.
  3. Affine Coupling Layer: Applies the RealNVP split-scale-shift mechanism.

Worked Example

Let us trace a single affine coupling block on a 4-dimensional vector x=[2.0,−1.0,3.0,4.0]x = [2.0, -1.0, 3.0, 4.0] where D=4D=4 and d=2d=2.

  1. Split input into two halves: x1:2=[2.0,−1.0],x3:4=[3.0,4.0]x_{1:2} = [2.0, -1.0], \quad x_{3:4} = [3.0, 4.0]

  2. Evaluate arbitrary neural networks s(x1:2)s(x_{1:2}) and t(x1:2)t(x_{1:2}). Suppose the network outputs: s(x1:2)=[0.4,0.1],t(x1:2)=[0.0,−0.3]s(x_{1:2}) = [0.4, 0.1], \quad t(x_{1:2}) = [0.0, -0.3]

  3. Compute scaling factors via exponentiation: exp⁡(s)=[exp⁡(0.4),exp⁡(0.1)]=[1.49182,1.10517]\exp(s) = [\exp(0.4), \exp(0.1)] = [1.49182, 1.10517]

  4. Compute output yy: y1:2=x1:2=[2.0,−1.0]y_{1:2} = x_{1:2} = [2.0, -1.0] y3=3.0×1.49182+0.0=4.47546y_3 = 3.0 \times 1.49182 + 0.0 = 4.47546 y4=4.0×1.10517+(−0.3)=4.42068−0.3=4.12068y_4 = 4.0 \times 1.10517 + (-0.3) = 4.42068 - 0.3 = 4.12068 y=[2.0,−1.0,4.47546,4.12068]y = [2.0, -1.0, 4.47546, 4.12068]

  5. Compute the contribution to log-likelihood (the Jacobian log-determinant): log⁡∣det⁡J∣=s1+s2=0.4+0.1=0.5000\log |\det J| = s_1 + s_2 = 0.4 + 0.1 = 0.5000

  6. Invert from y=[2.0,−1.0,4.47546,4.12068]y = [2.0, -1.0, 4.47546, 4.12068] back to xx: x1:2=y1:2=[2.0,−1.0]x_{1:2} = y_{1:2} = [2.0, -1.0] s(y1:2)=[0.4,0.1],t(y1:2)=[0.0,−0.3]s(y_{1:2}) = [0.4, 0.1], \quad t(y_{1:2}) = [0.0, -0.3] x3=(4.47546−0.0)×exp⁡(−0.4)=4.47546×0.67032=3.0000x_3 = (4.47546 - 0.0) \times \exp(-0.4) = 4.47546 \times 0.67032 = 3.0000 x4=(4.12068−(−0.3))×exp⁡(−0.1)=4.42068×0.90484=4.0000x_4 = (4.12068 - (-0.3)) \times \exp(-0.1) = 4.42068 \times 0.90484 = 4.0000 The reconstruction is numerically exact.

Code

from typing import Tupleimport math
class AffineCouplingLayer:    """Affine coupling layer operating on vectors with exact inversion."""
    def __init__(self, dim: int) -> None:        self.dim = dim        self.split_dim = dim // 2        # Weights for simple linear mappings modeling s and t        self.w_s = [0.2, -0.1]        self.w_t = [-0.5, 0.3]        self.b_t = [1.0, 0.0]
    def _net(self, x1: list[float]) -> Tuple[list[float], list[float]]:        # Scale s and translation t conditioned on x1        s = [self.w_s[0] * x1[0], self.w_s[1] * x1[1]]        t = [self.w_t[0] * x1[0] + self.b_t[0], self.w_t[1] * x1[1] + self.b_t[1]]        return s, t
    def forward(self, x: list[float]) -> Tuple[list[float], float]:        x1 = x[:self.split_dim]        x2 = x[self.split_dim:]        s, t = self._net(x1)
        y1 = list(x1)        y2 = [x2[i] * math.exp(s[i]) + t[i] for i in range(len(x2))]        log_det = sum(s)        return y1 + y2, log_det
    def inverse(self, y: list[float]) -> list[float]:        y1 = y[:self.split_dim]        y2 = y[self.split_dim:]        s, t = self._net(y1)
        x1 = list(y1)        x2 = [(y2[i] - t[i]) * math.exp(-s[i]) for i in range(len(y2))]        return x1 + x2

layer = AffineCouplingLayer(dim=4)x_in = [2.0, -1.0, 3.0, 4.0]y_out, log_det_val = layer.forward(x_in)x_recovered = layer.inverse(y_out)
print("Forward y:", [round(v, 4) for v in y_out])# -> Forward y: [2.0, -1.0, 4.4755, 4.1207]print("Log det:", round(log_det_val, 4))# -> Log det: 0.5print("Reconstructed x:", [round(v, 4) for v in x_recovered])# -> Reconstructed x: [2.0, -1.0, 3.0, 4.0]

Watch Out For

Numerical explosion from unconstrained exponential scaling

Symptom: During training, loss becomes NaN or sudden gradient explosions occur within the coupling layers after a few thousand steps.

In an affine coupling layer, y2=x2⊙exp⁡(s(x1))+t(x1)y_2 = x_2 \odot \exp(s(x_1)) + t(x_1). If the neural network s(x1)s(x_1) outputs unbounded activations like si=15.0s_i = 15.0, exp⁡(si)≈3.26×106\exp(s_i) \approx 3.26 \times 10^6. The inverse operation multiplies by exp⁡(−si)\exp(-s_i), resulting in severe numerical underflow and vanishing gradients.

The remedy is to pass the scale network outputs through a bounding activation: s=α⋅tanh⁡(s~)s = \alpha \cdot \tanh(\tilde{s}) where α≈2.0\alpha \approx 2.0, or clamp exp⁡(s)\exp(s) to [ϵ,1/ϵ][\epsilon, 1/\epsilon]. This guarantees stable conditioning throughout both forward and reverse paths.

The Quick Version

  • Normalizing flows chain bijective layers to map complex empirical data distributions into simple isotropic Gaussians.
  • RealNVP splits feature dimensions in half, applying identity to the first half and affine scale-and-shift to the second half.
  • Triangular Jacobians reduce determinant computation from an intractable O(D3)\mathcal{O}(D^3) to linear O(D)\mathcal{O}(D) time.
  • Glow extends RealNVP with actnorm and invertible 1×11 \times 1 convolutions to mix channels without fixed permutations.
  • Both exact log-likelihood evaluation and sample synthesis execute in a single parallelizable forward pass.