Skip to content
AI360Xpert
Beta

3D Gaussian Splatting

Instead of querying a dense neural network for every light ray, represent a 3D scene as millions of fuzzy colored ellipsoids and splat them directly onto the camera lens.

Anisotropic 3D Gaussians are projected via viewing Jacobians to 2D screen splats and composited using tile-based alpha blending.
Anisotropic 3D Gaussians are projected via viewing Jacobians to 2D screen splats and composited using tile-based alpha blending.

Why Does This Exist?

Neural Radiance Fields (NeRFs) proved that neural networks could synthesize photorealistic novel views of complex real-world scenes. However, NeRFs require evaluating a multi-layer perceptron dozens or hundreds of times along every individual camera ray. At 1920×10801920 \times 1080 resolution, rendering a single image requires hundreds of millions of forward passes through a neural network, limiting frame rates to fractions of a frame per second and requiring hours or days of compute to optimize.

Traditional point clouds render fast, but raw points leave holes and cannot model view-dependent specularity or semi-transparent boundaries. Triangle meshes struggle to capture complex geometries like hair, foliage, and diffuse reflections without heavy manual cleanup.

3D Gaussian Splatting (3DGS) replaces implicit neural ray marching with explicit, differentiable volumetric primitives. By defining a scene as millions of parameterized 3D Gaussians and projecting them directly onto screen tiles using a custom GPU rasterizer, 3DGS achieves equal or superior visual fidelity to NeRFs while training in minutes and rendering at over 100 frames per second.

Think of It Like This

A cloud of colored spray-paint droplets against glass

Imagine recreating a sculpture by suspending millions of tiny, translucent droplets of colored spray paint in mid-air. Each droplet has a specific position, a stretched oval shape, an orientation, an opacity, and a color that slightly shifts depending on where you stand.

To see what the sculpture looks like from any angle, you hold a clear glass pane in front of it and press the droplets flat against the glass surface. Droplets in front cover up droplets behind them according to their transparency. Rather than examining the entire depth of the room point by point, you simply stamp the overlapping colored patches onto the glass in order from front to back.

How It Actually Works

Differentiable Projection and Covariance Formulation

A 3D Gaussian is centered at mean μ∈R3\mu \in \mathbb{R}^3 and defined by a 3D covariance matrix Σ\Sigma:

G(x)=exp⁡(−12(x−μ)TΣ−1(x−μ))G(x) = \exp\left(-\frac{1}{2} (x - \mu)^T \Sigma^{-1} (x - \mu)\right)

Because a valid covariance matrix must be positive semi-definite, direct gradient descent on Σ\Sigma is prone to invalid configurations. 3DGS decomposes Σ\Sigma into a scaling vector s∈R3s \in \mathbb{R}^3 and a rotation quaternion q∈R4q \in \mathbb{R}^4:

Σ=RSSTRT\Sigma = R S S^T R^T

where S=diag(s)S = \text{diag}(s) is a diagonal scaling matrix and RR is the orthonormal rotation matrix derived from normalized quaternion qq.

To render these 3D Gaussians from a camera view defined by viewing transform WW and projective transformation, the 3D covariance is projected into a 2D screen-space covariance Σ′\Sigma' using the local affine Jacobian JJ of the projective transformation:

Σ′=JWΣWTJT\Sigma' = J W \Sigma W^T J^T

Each Gaussian carries an opacity α∈[0,1]\alpha \in [0, 1] and view-dependent color coefficients cc represented via spherical harmonics (SH) up to degree 3.

During rendering, the screen is subdivided into 16×1616 \times 16 pixel tiles. The rasterizer discards Gaussians outside the viewing frustum (or beyond a 99%99\% confidence radius of 3σ3\sigma). The remaining Gaussians are assigned 64-bit keys containing their tile ID and view depth, sorted rapidly via GPU radix sort, and alpha-blended front-to-back:

C=∑i∈Nciαi∏j=1i−1(1−αj)C = \sum_{i \in \mathcal{N}} c_i \alpha_i \prod_{j=1}^{i-1} (1 - \alpha_j)

Rasterization terminates early when accumulated opacity reaches 1−ϵ1 - \epsilon (typically transmittance T<0.0001T < 0.0001), avoiding needless compute for occluded splats.

Worked Example

Consider a single 3D Gaussian positioned at camera coordinates μ=[0,0,2]T\mu = [0, 0, 2]^T with focal length fx=fy=500f_x = f_y = 500 pixels and optical center at the origin (0,0)(0, 0).

  1. Affine Jacobian JJ at depth z=2z = 2: The perspective projection maps [x,y,z]T→[fxx/z,fyy/z]T[x, y, z]^T \to [f_x x / z, f_y y / z]^T. The Jacobian evaluated at μ=[0,0,2]T\mu = [0, 0, 2]^T is:

    J=[fxz0−fxxz20fyz−fyyz2]=[500200050020]=[2500002500]J = \begin{bmatrix} \frac{f_x}{z} & 0 & -\frac{f_x x}{z^2} \\ 0 & \frac{f_y}{z} & -\frac{f_y y}{z^2} \end{bmatrix} = \begin{bmatrix} \frac{500}{2} & 0 & 0 \\ 0 & \frac{500}{2} & 0 \end{bmatrix} = \begin{bmatrix} 250 & 0 & 0 \\ 0 & 250 & 0 \end{bmatrix}
  2. Covariance in camera space Σ\Sigma: Suppose the Gaussian is an axis-aligned ellipsoid with radii s=[0.02,0.01,0.02]Ts = [0.02, 0.01, 0.02]^T meters (no rotation, R=IR = I):

    Σ=diag(0.022,0.012,0.022)=[0.00040000.00010000.0004]\Sigma = \text{diag}(0.02^2, 0.01^2, 0.02^2) = \begin{bmatrix} 0.0004 & 0 & 0 \\ 0 & 0.0001 & 0 \\ 0 & 0 & 0.0004 \end{bmatrix}
  3. Screen-space 2D Covariance Σ′\Sigma':

    Σ′=JΣJT=[2502×0.0004002502×0.0001]=[62500×0.00040062500×0.0001]=[25.0006.25]\Sigma' = J \Sigma J^T = \begin{bmatrix} 250^2 \times 0.0004 & 0 \\ 0 & 250^2 \times 0.0001 \end{bmatrix} = \begin{bmatrix} 62500 \times 0.0004 & 0 \\ 0 & 62500 \times 0.0001 \end{bmatrix} = \begin{bmatrix} 25.0 & 0 \\ 0 & 6.25 \end{bmatrix}

    In screen space, the standard deviations are σx=25.0=5.0\sigma_x = \sqrt{25.0} = 5.0 pixels and σy=6.25=2.5\sigma_y = \sqrt{6.25} = 2.5 pixels.

  4. Pixel evaluation at offset Δ=[3,0]T\Delta = [3, 0]^T pixels:

    G(Δ)=exp⁡(−12(3225.0))=exp⁡(−0.18)≈0.8353G(\Delta) = \exp\left(-\frac{1}{2} \left(\frac{3^2}{25.0}\right)\right) = \exp(-0.18) \approx 0.8353

    With base opacity α0=0.8\alpha_0 = 0.8, the effective alpha at this pixel is α=α0×G(Δ)=0.8×0.8353=0.6682\alpha = \alpha_0 \times G(\Delta) = 0.8 \times 0.8353 = 0.6682.

Code

import torchimport torch.nn as nn
def compute_2d_covariance(    scale: torch.Tensor,    rotation_q: torch.Tensor,    view_matrix: torch.Tensor,    fov_x: float,    fov_y: float,    width: float,    height: float,    mean_cam: torch.Tensor) -> torch.Tensor:    """Project 3D Gaussian scale and rotation to 2D screen covariance."""    # 1. Normalize quaternion and construct 3D rotation matrix R    q = rotation_q / torch.norm(rotation_q, dim=-1, keepdim=True)    r, i, j, k = q[0], q[1], q[2], q[3]    R = torch.stack([        torch.stack([1 - 2*(j**2 + k**2), 2*(i*j - k*r), 2*(i*k + j*r)]),        torch.stack([2*(i*j + k*r), 1 - 2*(i**2 + k**2), 2*(j*k - i*r)]),        torch.stack([2*(i*k - j*r), 2*(j*k + i*r), 1 - 2*(i**2 + j**2)])    ])        # 2. 3D covariance Sigma = R @ S @ S.T @ R.T    S = torch.diag(scale)    M = R @ S    Sigma = M @ M.T
    # 3. Perspective projection Jacobian J    fx = width / (2.0 * torch.tan(torch.tensor(fov_x / 2.0)))    fy = height / (2.0 * torch.tan(torch.tensor(fov_y / 2.0)))    x, y, z = mean_cam[0], mean_cam[1], mean_cam[2]        J = torch.tensor([        [fx / z, 0.0, -(fx * x) / (z * z)],        [0.0, fy / z, -(fy * y) / (z * z)]    ], dtype=torch.float32)
    # 4. Transform to screen space: Sigma_2d = J @ W @ Sigma @ W.T @ J.T    W = view_matrix[:3, :3]    Sigma_cam = W @ Sigma @ W.T    Sigma_2d = J @ Sigma_cam @ J.T        # Add small low-pass filter to avoid aliasing artifacts    Sigma_2d[0, 0] += 0.3    Sigma_2d[1, 1] += 0.3    return Sigma_2d
# Example with a unit Gaussian at z=2.0 metersscale = torch.tensor([0.02, 0.01, 0.02])quat = torch.tensor([1.0, 0.0, 0.0, 0.0])view_mat = torch.eye(4)mean_cam = torch.tensor([0.0, 0.0, 2.0])
cov2d = compute_2d_covariance(    scale, quat, view_mat,    fov_x=1.0472, fov_y=1.0472,    width=800.0, height=800.0,    mean_cam=mean_cam)
print(f"cov_xx: {cov2d[0, 0].item():.2f}, cov_yy: {cov2d[1, 1].item():.2f}")# -> cov_xx: 27.63, cov_yy: 7.13

Watch Out For

Floating splat artifacts and uncontrolled densification

During training, large gradient errors in under-reconstructed empty spaces cause the optimization to spawn huge transparent Gaussians that float in front of the camera, appearing as blurry halos or needle-like artifacts from alternate angles.

To eliminate floaters, track average positional gradient magnitude ∇μL\nabla_{\mu} L. When gradients exceed a threshold τpos\tau_{\text{pos}}, split Gaussians with large scales into two smaller ones, or clone Gaussians with small scales. Crucially, enforce an opacity reset step every 3,000 iterations that zeros out opacities α→0.01\alpha \to 0.01, pruning away any splats that cannot earn back opacity during subsequent training steps.

The Quick Version

  • Replaces slow ray-marching neural networks with millions of explicit 3D ellipsoids parameterized by position, scale, rotation, opacity, and spherical harmonic colors.
  • Differentiable camera projection maps 3D covariance directly to 2D screen covariance using local projective Jacobians.
  • Tile-based GPU radix sorting and front-to-back alpha blending enable photorealistic novel-view synthesis at over 100 frames per second.