Neural Rendering & 3D
Neural rendering synthesizes photorealistic novel views of 3D objects by optimizing neural networks via differentiable volumetric ray marching.
Why Does This Exist?
In traditional computer graphics, 3D scenes are modeled using discrete polygonal meshes, point clouds, or volumetric voxel grids. While highly optimized for rasterization hardware, these discrete representations present severe limitations when reconstructing real-world environments from photographs:
- Discretization and Memory Limits: High-resolution voxel grids scale cubically as ; storing a fine voxel grid requires gigabytes of memory, yet mostly contains empty space.
- Non-Differentiable Rendering: Traditional triangle rasterization involves discrete rasterization boundaries (an edge is either on a pixel or off); taking derivatives of pixel colors with respect to mesh vertex coordinates produces discontinuous gradients, making end-to-end gradient descent impossible.
Neural rendering bridges computer vision and computer graphics by parameterizing 3D scenes as continuous implicit neural functions. Models like Neural Radiance Fields (NeRF; Mildenhall et al., 2020) and 3D Gaussian Splatting (Kerbl et al., 2023) use differentiable image formation models: by casting rays through virtual camera lenses and integrating light absorption mathematically, every parameter of the 3D scene can be optimized directly against raw 2D photographs using standard Adam gradient descent.
Think of It Like This
A colored fog sculpture lit by a flashlight
Imagine an artist sculpting a transparent glass cube filled with multicolored, swirling smoke. In dense sections of smoke, the fog is thick and opaque; in empty sections, the fog is clear air. Furthermore, the smoke contains tiny pearlescent flakes that shine gold when viewed from the front, but turn deep crimson when viewed from the side.
To determine what an observer sees through a specific peephole:
- You trace a line of sight (a ray of light) from the observer's eye through the glass cube.
- As the line penetrates the cube, you record how thick the smoke is () and what color it shines () at every millimeter.
- If the ray first hits thick red smoke, that red color reflects into the eye and occludes whatever lies behind it. If it passes through thin mist, deeper layers remain partially visible.
Neural rendering does not store physical glass or smoke: a neural network acts as a continuous mathematical inquiry box that tells the ray marcher the exact smoke density and color at any coordinate in the universe.
How It Actually Works
Continuous Radiance Fields and Differentiable Volume Rendering
The flagship mathematical formulation of neural rendering is the Neural Radiance Field (NeRF). A 3D scene is represented as a continuous function , parameterized by a multi-layer perceptron (MLP) with weights :
- is the 3D spatial location.
- is the 2D viewing direction vector.
- is the differential volume density (how much light is absorbed/blocked at ).
- is the directional emitted radiance (color).
To enforce physical consistency, volume density depends strictly on spatial location (geometry is view-invariant), whereas color depends on both location and viewing angle to capture non-Lambertian effects like glossy specular reflections.
1. High-Frequency Positional Encoding
Deep neural networks exhibit a spectral bias toward learning low-frequency functions. To capture fine 3D surface textures and crisp edges, NeRF applies a sinusoidal positional encoding :
Setting for coordinates projects each spatial coordinate into a 60-dimensional frequency space.
2. Differentiable Volume Rendering (Numerical Quadrature)
For a camera ray spanning near bound and far bound , the expected color is defined by the volume rendering integral:
where is the accumulated transmittance (the probability that the ray traverses from to without hitting an occluding particle).
To compute this on a GPU, the integral is discretized via stratified sampling into intervals of width :
Because every operation in this summation is smooth and differentiable with respect to and , the photometric reconstruction loss:
propagates exact analytical gradients back to the weights of the MLP without needing 3D supervision, depth sensors, or surface meshes.
Worked Example
Consider a single camera ray sampled at intervals with uniform step size :
- Sample 1: Volume density , Color (pure red).
- Sample 2: Volume density , Color (pure green).
-
Calculate Sample 1 Transmittance and Alpha:
- Transmittance (no obstruction prior to sample 1).
- Absorption probability .
- Contribution of sample 1:
-
Calculate Sample 2 Transmittance and Alpha:
- Transmittance (only 60.6% of light penetrates past sample 1).
- Absorption probability .
- Contribution of sample 2:
-
Total Pixel Color: The pixel reflects a blend of both layers, with the denser background layer partially occluded by the semi-transparent foreground fog.
Code
import numpy as np
def volume_render_ray( densities: np.ndarray, # Shape: (K,) volume densities sigma >= 0 colors: np.ndarray, # Shape: (K, 3) RGB colors in [0, 1] step_sizes: np.ndarray, # Shape: (K,) delta distances between sample points) -> np.ndarray: """Discretized differentiable volume rendering via numerical quadrature.""" k = densities.shape[0]
# 1. Opacity for each sample interval: alpha_i = 1 - exp(-sigma_i * delta_i) sigmas_deltas = densities * step_sizes alphas = 1.0 - np.exp(-sigmas_deltas)
# 2. Accumulated transmittance: T_i = exp(-sum_{j=1}^{i-1} sigma_j * delta_j) # Cumulative sum with 0 prepended for T_1 = 1.0 accumulated_optical_depth = np.cumsum(sigmas_deltas) transmittance = np.ones(k, dtype=np.float64) transmittance[1:] = np.exp(-accumulated_optical_depth[:-1])
# 3. Composite sample weights: w_i = T_i * alpha_i weights = transmittance * alphas
# 4. Integrate pixel color: C = sum(weights * colors) pixel_color = np.sum(weights[:, np.newaxis] * colors, axis=0) return pixel_color
# Test with 2 sample pointsdens = np.array([0.5, 2.0])cols = np.array([ [1.0, 0.0, 0.0], # Red foreground [0.0, 1.0, 0.0], # Green background])deltas = np.array([1.0, 1.0])
rendered_pixel = volume_render_ray(dens, cols, deltas)print("Rendered RGB Pixel:", np.round(rendered_pixel, 4))# -> Rendered RGB Pixel: [0.3935 0.5245 0. ]Watch Out For
Ray marching computational latency during inference
Evaluating a NeRF model for a single image requires casting camera rays. Sampling 192 points per ray forces over 120 million forward passes through a 256-width MLP for just one image frame, causing rendering to crawl at 0.1 FPS.
If your production pipeline requires real-time 60+ FPS novel view synthesis (such as VR/AR headsets or interactive web visualizers), do not evaluate raw coordinate MLPs per ray at runtime. Bake the trained NeRF into a multiresolution hash grid (Instant-NGP) or transition to explicit rasterized primitives like 3D Gaussian Splatting, which replaces volumetric ray marching with tile-based GPU alpha-blending rasterization.
The Quick Version
- Neural rendering represents continuous 3D geometry and view-dependent color using neural networks optimized end-to-end via 2D photometric loss.
- High-frequency sinusoidal positional encodings overcome the spectral bias of MLPs, preserving crisp textures and sharp surface boundaries.
- Volume rendering uses numerical quadrature to integrate transmittance and opacity along camera rays in a fully differentiable pipeline.