Skip to content
AI360Xpert
Beta

Neural Rendering & 3D

Neural rendering synthesizes photorealistic novel views of 3D objects by optimizing neural networks via differentiable volumetric ray marching.

Neural rendering samples 3D points along camera rays, queries an MLP for volume density and view-dependent color, and integrates pixel values via differentiable volume rendering
Neural rendering samples 3D points along camera rays, queries an MLP for volume density and view-dependent color, and integrates pixel values via differentiable volume rendering

Why Does This Exist?

In traditional computer graphics, 3D scenes are modeled using discrete polygonal meshes, point clouds, or volumetric voxel grids. While highly optimized for rasterization hardware, these discrete representations present severe limitations when reconstructing real-world environments from photographs:

  1. Discretization and Memory Limits: High-resolution voxel grids scale cubically as O(V3)\mathcal{O}(V^3); storing a fine 102431024^3 voxel grid requires gigabytes of memory, yet mostly contains empty space.
  2. Non-Differentiable Rendering: Traditional triangle rasterization involves discrete rasterization boundaries (an edge is either on a pixel or off); taking derivatives of pixel colors with respect to mesh vertex coordinates produces discontinuous gradients, making end-to-end gradient descent impossible.

Neural rendering bridges computer vision and computer graphics by parameterizing 3D scenes as continuous implicit neural functions. Models like Neural Radiance Fields (NeRF; Mildenhall et al., 2020) and 3D Gaussian Splatting (Kerbl et al., 2023) use differentiable image formation models: by casting rays through virtual camera lenses and integrating light absorption mathematically, every parameter of the 3D scene can be optimized directly against raw 2D photographs using standard Adam gradient descent.

Think of It Like This

A colored fog sculpture lit by a flashlight

Imagine an artist sculpting a transparent glass cube filled with multicolored, swirling smoke. In dense sections of smoke, the fog is thick and opaque; in empty sections, the fog is clear air. Furthermore, the smoke contains tiny pearlescent flakes that shine gold when viewed from the front, but turn deep crimson when viewed from the side.

To determine what an observer sees through a specific peephole:

  1. You trace a line of sight (a ray of light) from the observer's eye through the glass cube.
  2. As the line penetrates the cube, you record how thick the smoke is (σ\sigma) and what color it shines (c\mathbf{c}) at every millimeter.
  3. If the ray first hits thick red smoke, that red color reflects into the eye and occludes whatever lies behind it. If it passes through thin mist, deeper layers remain partially visible.

Neural rendering does not store physical glass or smoke: a neural network acts as a continuous mathematical inquiry box that tells the ray marcher the exact smoke density and color at any coordinate in the universe.

How It Actually Works

Continuous Radiance Fields and Differentiable Volume Rendering

The flagship mathematical formulation of neural rendering is the Neural Radiance Field (NeRF). A 3D scene is represented as a continuous function Fθ:(x,d)→(c,σ)F_\theta: (\mathbf{x}, \mathbf{d}) \to (\mathbf{c}, \sigma), parameterized by a multi-layer perceptron (MLP) with weights θ\theta:

  • x=(x,y,z)∈R3\mathbf{x} = (x, y, z) \in \mathbb{R}^3 is the 3D spatial location.
  • d=(θ,ϕ)∈S2\mathbf{d} = (\theta, \phi) \in \mathbb{S}^2 is the 2D viewing direction vector.
  • σ∈R+\sigma \in \mathbb{R}^+ is the differential volume density (how much light is absorbed/blocked at x\mathbf{x}).
  • c=(r,g,b)∈[0,1]3\mathbf{c} = (r, g, b) \in [0, 1]^3 is the directional emitted radiance (color).

To enforce physical consistency, volume density σ(x)\sigma(\mathbf{x}) depends strictly on spatial location x\mathbf{x} (geometry is view-invariant), whereas color c(x,d)\mathbf{c}(\mathbf{x}, \mathbf{d}) depends on both location and viewing angle to capture non-Lambertian effects like glossy specular reflections.

1. High-Frequency Positional Encoding

Deep neural networks exhibit a spectral bias toward learning low-frequency functions. To capture fine 3D surface textures and crisp edges, NeRF applies a sinusoidal positional encoding γ:R→R2L\gamma: \mathbb{R} \to \mathbb{R}^{2L}:

γ(p)=[sin⁡(20πp),cos⁡(20πp),…,sin⁡(2L−1πp),cos⁡(2L−1πp)]\gamma(p) = \left[ \sin(2^0 \pi p), \cos(2^0 \pi p), \dots, \sin(2^{L-1} \pi p), \cos(2^{L-1} \pi p) \right]

Setting L=10L = 10 for coordinates x\mathbf{x} projects each spatial coordinate into a 60-dimensional frequency space.

2. Differentiable Volume Rendering (Numerical Quadrature)

For a camera ray r(t)=o+td\mathbf{r}(t) = \mathbf{o} + t\mathbf{d} spanning near bound tnt_n and far bound tft_f, the expected color C(r)C(\mathbf{r}) is defined by the volume rendering integral:

C(r)=∫tntfT(t)σ(r(t))c(r(t),d) dtC(\mathbf{r}) = \int_{t_n}^{t_f} T(t) \sigma(\mathbf{r}(t)) \mathbf{c}(\mathbf{r}(t), \mathbf{d}) \, dt

where T(t)=exp⁡(−∫tntσ(r(s)) ds)T(t) = \exp\left(-\int_{t_n}^t \sigma(\mathbf{r}(s)) \, ds\right) is the accumulated transmittance (the probability that the ray traverses from tnt_n to tt without hitting an occluding particle).

To compute this on a GPU, the integral is discretized via stratified sampling into KK intervals [ti,ti+1][t_i, t_{i+1}] of width δi=ti+1−ti\delta_i = t_{i+1} - t_i:

C^(r)=∑i=1KTi(1−exp⁡(−σiδi))ci\hat{C}(\mathbf{r}) = \sum_{i=1}^K T_i \left( 1 - \exp(-\sigma_i \delta_i) \right) \mathbf{c}_i

Ti=exp⁡(−∑j=1i−1σjδj)T_i = \exp\left( -\sum_{j=1}^{i-1} \sigma_j \delta_j \right)

Because every operation in this summation is smooth and differentiable with respect to σi\sigma_i and ci\mathbf{c}_i, the photometric reconstruction loss:

L=∑r∈R∥C^(r)−Cgroundtruth(r)∥22\mathcal{L} = \sum_{\mathbf{r} \in \mathcal{R}} \left\| \hat{C}(\mathbf{r}) - C_{\text{groundtruth}}(\mathbf{r}) \right\|_2^2

propagates exact analytical gradients back to the weights θ\theta of the MLP without needing 3D supervision, depth sensors, or surface meshes.

Worked Example

Consider a single camera ray sampled at K=2K = 2 intervals with uniform step size δ1=δ2=1.0\delta_1 = \delta_2 = 1.0:

  • Sample 1: Volume density σ1=0.5\sigma_1 = 0.5, Color c1=[1.0,0.0,0.0]\mathbf{c}_1 = [1.0, 0.0, 0.0] (pure red).
  • Sample 2: Volume density σ2=2.0\sigma_2 = 2.0, Color c2=[0.0,1.0,0.0]\mathbf{c}_2 = [0.0, 1.0, 0.0] (pure green).
  1. Calculate Sample 1 Transmittance and Alpha:

    • Transmittance T1=exp⁡(0)=1.0T_1 = \exp(0) = 1.0 (no obstruction prior to sample 1).
    • Absorption probability α1=1−exp⁡(−σ1δ1)=1−exp⁡(−0.5)≈1−0.6065=0.3935\alpha_1 = 1 - \exp(-\sigma_1 \delta_1) = 1 - \exp(-0.5) \approx 1 - 0.6065 = 0.3935.
    • Contribution of sample 1: C1=T1α1c1=1.0×0.3935×[1.0,0.0,0.0]=[0.3935,0.0,0.0]\mathbf{C}_1 = T_1 \alpha_1 \mathbf{c}_1 = 1.0 \times 0.3935 \times [1.0, 0.0, 0.0] = [0.3935, 0.0, 0.0]
  2. Calculate Sample 2 Transmittance and Alpha:

    • Transmittance T2=exp⁡(−σ1δ1)=exp⁡(−0.5)≈0.6065T_2 = \exp(-\sigma_1 \delta_1) = \exp(-0.5) \approx 0.6065 (only 60.6% of light penetrates past sample 1).
    • Absorption probability α2=1−exp⁡(−σ2δ2)=1−exp⁡(−2.0)≈1−0.1353=0.8647\alpha_2 = 1 - \exp(-\sigma_2 \delta_2) = 1 - \exp(-2.0) \approx 1 - 0.1353 = 0.8647.
    • Contribution of sample 2: C2=T2α2c2=0.6065×0.8647×[0.0,1.0,0.0]≈0.5244×[0.0,1.0,0.0]=[0.0,0.5244,0.0]\mathbf{C}_2 = T_2 \alpha_2 \mathbf{c}_2 = 0.6065 \times 0.8647 \times [0.0, 1.0, 0.0] \approx 0.5244 \times [0.0, 1.0, 0.0] = [0.0, 0.5244, 0.0]
  3. Total Pixel Color: C^(r)=C1+C2=[0.3935,0.5244,0.0]\hat{C}(\mathbf{r}) = \mathbf{C}_1 + \mathbf{C}_2 = [0.3935, 0.5244, 0.0] The pixel reflects a blend of both layers, with the denser background layer partially occluded by the semi-transparent foreground fog.

Code

import numpy as np

def volume_render_ray(    densities: np.ndarray,  # Shape: (K,) volume densities sigma >= 0    colors: np.ndarray,     # Shape: (K, 3) RGB colors in [0, 1]    step_sizes: np.ndarray, # Shape: (K,) delta distances between sample points) -> np.ndarray:    """Discretized differentiable volume rendering via numerical quadrature."""    k = densities.shape[0]
    # 1. Opacity for each sample interval: alpha_i = 1 - exp(-sigma_i * delta_i)    sigmas_deltas = densities * step_sizes    alphas = 1.0 - np.exp(-sigmas_deltas)
    # 2. Accumulated transmittance: T_i = exp(-sum_{j=1}^{i-1} sigma_j * delta_j)    # Cumulative sum with 0 prepended for T_1 = 1.0    accumulated_optical_depth = np.cumsum(sigmas_deltas)    transmittance = np.ones(k, dtype=np.float64)    transmittance[1:] = np.exp(-accumulated_optical_depth[:-1])
    # 3. Composite sample weights: w_i = T_i * alpha_i    weights = transmittance * alphas
    # 4. Integrate pixel color: C = sum(weights * colors)    pixel_color = np.sum(weights[:, np.newaxis] * colors, axis=0)    return pixel_color

# Test with 2 sample pointsdens = np.array([0.5, 2.0])cols = np.array([    [1.0, 0.0, 0.0],  # Red foreground    [0.0, 1.0, 0.0],  # Green background])deltas = np.array([1.0, 1.0])
rendered_pixel = volume_render_ray(dens, cols, deltas)print("Rendered RGB Pixel:", np.round(rendered_pixel, 4))# -> Rendered RGB Pixel: [0.3935 0.5245 0.    ]

Watch Out For

Ray marching computational latency during inference

Evaluating a NeRF model for a single 800×800800 \times 800 image requires casting 640,000640{,}000 camera rays. Sampling 192 points per ray forces over 120 million forward passes through a 256-width MLP for just one image frame, causing rendering to crawl at 0.1 FPS.

If your production pipeline requires real-time 60+ FPS novel view synthesis (such as VR/AR headsets or interactive web visualizers), do not evaluate raw coordinate MLPs per ray at runtime. Bake the trained NeRF into a multiresolution hash grid (Instant-NGP) or transition to explicit rasterized primitives like 3D Gaussian Splatting, which replaces volumetric ray marching with tile-based GPU alpha-blending rasterization.

The Quick Version

  • Neural rendering represents continuous 3D geometry and view-dependent color using neural networks optimized end-to-end via 2D photometric loss.
  • High-frequency sinusoidal positional encodings overcome the spectral bias of MLPs, preserving crisp textures and sharp surface boundaries.
  • Volume rendering uses numerical quadrature to integrate transmittance and opacity along camera rays in a fully differentiable pipeline.