3D Gaussian Splatting
Instead of querying a dense neural network for every light ray, represent a 3D scene as millions of fuzzy colored ellipsoids and splat them directly onto the camera lens.
Why Does This Exist?
Neural Radiance Fields (NeRFs) proved that neural networks could synthesize photorealistic novel views of complex real-world scenes. However, NeRFs require evaluating a multi-layer perceptron dozens or hundreds of times along every individual camera ray. At resolution, rendering a single image requires hundreds of millions of forward passes through a neural network, limiting frame rates to fractions of a frame per second and requiring hours or days of compute to optimize.
Traditional point clouds render fast, but raw points leave holes and cannot model view-dependent specularity or semi-transparent boundaries. Triangle meshes struggle to capture complex geometries like hair, foliage, and diffuse reflections without heavy manual cleanup.
3D Gaussian Splatting (3DGS) replaces implicit neural ray marching with explicit, differentiable volumetric primitives. By defining a scene as millions of parameterized 3D Gaussians and projecting them directly onto screen tiles using a custom GPU rasterizer, 3DGS achieves equal or superior visual fidelity to NeRFs while training in minutes and rendering at over 100 frames per second.
Think of It Like This
A cloud of colored spray-paint droplets against glass
Imagine recreating a sculpture by suspending millions of tiny, translucent droplets of colored spray paint in mid-air. Each droplet has a specific position, a stretched oval shape, an orientation, an opacity, and a color that slightly shifts depending on where you stand.
To see what the sculpture looks like from any angle, you hold a clear glass pane in front of it and press the droplets flat against the glass surface. Droplets in front cover up droplets behind them according to their transparency. Rather than examining the entire depth of the room point by point, you simply stamp the overlapping colored patches onto the glass in order from front to back.
How It Actually Works
Differentiable Projection and Covariance Formulation
A 3D Gaussian is centered at mean and defined by a 3D covariance matrix :
Because a valid covariance matrix must be positive semi-definite, direct gradient descent on is prone to invalid configurations. 3DGS decomposes into a scaling vector and a rotation quaternion :
where is a diagonal scaling matrix and is the orthonormal rotation matrix derived from normalized quaternion .
To render these 3D Gaussians from a camera view defined by viewing transform and projective transformation, the 3D covariance is projected into a 2D screen-space covariance using the local affine Jacobian of the projective transformation:
Each Gaussian carries an opacity and view-dependent color coefficients represented via spherical harmonics (SH) up to degree 3.
During rendering, the screen is subdivided into pixel tiles. The rasterizer discards Gaussians outside the viewing frustum (or beyond a confidence radius of ). The remaining Gaussians are assigned 64-bit keys containing their tile ID and view depth, sorted rapidly via GPU radix sort, and alpha-blended front-to-back:
Rasterization terminates early when accumulated opacity reaches (typically transmittance ), avoiding needless compute for occluded splats.
Worked Example
Consider a single 3D Gaussian positioned at camera coordinates with focal length pixels and optical center at the origin .
-
Affine Jacobian at depth : The perspective projection maps . The Jacobian evaluated at is:
-
Covariance in camera space : Suppose the Gaussian is an axis-aligned ellipsoid with radii meters (no rotation, ):
-
Screen-space 2D Covariance :
In screen space, the standard deviations are pixels and pixels.
-
Pixel evaluation at offset pixels:
With base opacity , the effective alpha at this pixel is .
Code
import torchimport torch.nn as nn
def compute_2d_covariance( scale: torch.Tensor, rotation_q: torch.Tensor, view_matrix: torch.Tensor, fov_x: float, fov_y: float, width: float, height: float, mean_cam: torch.Tensor) -> torch.Tensor: """Project 3D Gaussian scale and rotation to 2D screen covariance.""" # 1. Normalize quaternion and construct 3D rotation matrix R q = rotation_q / torch.norm(rotation_q, dim=-1, keepdim=True) r, i, j, k = q[0], q[1], q[2], q[3] R = torch.stack([ torch.stack([1 - 2*(j**2 + k**2), 2*(i*j - k*r), 2*(i*k + j*r)]), torch.stack([2*(i*j + k*r), 1 - 2*(i**2 + k**2), 2*(j*k - i*r)]), torch.stack([2*(i*k - j*r), 2*(j*k + i*r), 1 - 2*(i**2 + j**2)]) ]) # 2. 3D covariance Sigma = R @ S @ S.T @ R.T S = torch.diag(scale) M = R @ S Sigma = M @ M.T
# 3. Perspective projection Jacobian J fx = width / (2.0 * torch.tan(torch.tensor(fov_x / 2.0))) fy = height / (2.0 * torch.tan(torch.tensor(fov_y / 2.0))) x, y, z = mean_cam[0], mean_cam[1], mean_cam[2] J = torch.tensor([ [fx / z, 0.0, -(fx * x) / (z * z)], [0.0, fy / z, -(fy * y) / (z * z)] ], dtype=torch.float32)
# 4. Transform to screen space: Sigma_2d = J @ W @ Sigma @ W.T @ J.T W = view_matrix[:3, :3] Sigma_cam = W @ Sigma @ W.T Sigma_2d = J @ Sigma_cam @ J.T # Add small low-pass filter to avoid aliasing artifacts Sigma_2d[0, 0] += 0.3 Sigma_2d[1, 1] += 0.3 return Sigma_2d
# Example with a unit Gaussian at z=2.0 metersscale = torch.tensor([0.02, 0.01, 0.02])quat = torch.tensor([1.0, 0.0, 0.0, 0.0])view_mat = torch.eye(4)mean_cam = torch.tensor([0.0, 0.0, 2.0])
cov2d = compute_2d_covariance( scale, quat, view_mat, fov_x=1.0472, fov_y=1.0472, width=800.0, height=800.0, mean_cam=mean_cam)
print(f"cov_xx: {cov2d[0, 0].item():.2f}, cov_yy: {cov2d[1, 1].item():.2f}")# -> cov_xx: 27.63, cov_yy: 7.13Watch Out For
Floating splat artifacts and uncontrolled densification
During training, large gradient errors in under-reconstructed empty spaces cause the optimization to spawn huge transparent Gaussians that float in front of the camera, appearing as blurry halos or needle-like artifacts from alternate angles.
To eliminate floaters, track average positional gradient magnitude . When gradients exceed a threshold , split Gaussians with large scales into two smaller ones, or clone Gaussians with small scales. Crucially, enforce an opacity reset step every 3,000 iterations that zeros out opacities , pruning away any splats that cannot earn back opacity during subsequent training steps.
The Quick Version
- Replaces slow ray-marching neural networks with millions of explicit 3D ellipsoids parameterized by position, scale, rotation, opacity, and spherical harmonic colors.
- Differentiable camera projection maps 3D covariance directly to 2D screen covariance using local projective Jacobians.
- Tile-based GPU radix sorting and front-to-back alpha blending enable photorealistic novel-view synthesis at over 100 frames per second.