Computer Vision2023intermediate13 min read

3D Gaussian Splatting for Real-Time Radiance Field Rendering

التمثيل بالنقاط الغاوسية ثلاثية الأبعاد لتصيير حقول الإشعاع في الزمن الحقيقي

Kerbl, B. · Kopanas, G. · Leimkühler, T. · Drettakis, G. — ACM Transactions on Graphics (SIGGRAPH)

The problem

By 2023, Neural Fields (NeRF) had achieved remarkable quality for synthesis — generating photorealistic images of a scene from new camera angles. But NeRF and its successors shared a fundamental bottleneck: rendering required marching hundreds of rays through a neural network, making real-time display impossible for full scenes at high resolution. Faster methods like Instant-NGP accelerated but still could not reach 30 fps at 1080p on unbounded, real-world scenes. The field was stuck between quality and speed — no method could deliver both.

The contribution

A complete pipeline for photorealistic novel view synthesis that achieves state-of-the-art quality AND real-time rendering (≥30 fps at 1080p). Three interlocking innovations: (1) an explicit using millions of 3D Gaussians — each with learnable position, covariance, opacity, and spherical harmonic color coefficients; (2) an adaptive density control scheme that clones, splits, and prunes Gaussians during optimization to capture both fine detail and broad structure; (3) a tile-based differentiable GPU rasterizer that projects, sorts, and alpha-composites Gaussians in real time — no ray marching, no neural network at render time.

The impact

3D became the fastest-adopted representation in the history of neural rendering. Within months it spawned hundreds of follow-up papers covering dynamic scenes (4D-GS), SLAM, text-to-3D, avatars, autonomous driving, and more. It proved that explicit primitives with classical rasterization can match or exceed implicit neural representations in quality while being orders of magnitude faster. The work directly influenced video generation systems and became a foundational building block for real-time 3D content creation. GitHub stars exceeded 22,000 — a testament to both research impact and practical utility.

NeRF is like an oracle locked in a dark room. Hand it a camera position and direction through a slot, and it slowly computes what color you'd see — one pixel at a time. The oracle is brilliant but painfully slow: getting a single full image takes seconds.

3D Gaussian Splatting takes a completely different approach. Imagine filling a room with millions of tiny colored glass beads — each one can be stretched into an ellipsoid, tinted with any color, and made more or less transparent. You adjust every bead until, when you look through any window, the scene behind the glass perfectly matches a photograph. No oracle is needed at viewing time — just point a camera and the GPU paints all the beads onto the screen in milliseconds.

The problem: neural rendering is accurate but slow

NeRF represents a scene as a continuous function stored inside a neural network: given any 3D point and viewing direction, the network outputs a color and density. To render a single pixel, you march a ray into the scene, query the network at dozens of sample points along the ray, then composite those samples into a final color via .

This works beautifully — NeRF produces photorealistic images — but it is fundamentally slow. Every pixel requires multiple neural network forward passes. At 1080p, that means over two million pixels × dozens of samples each = hundreds of millions of network evaluations per frame. Even with spatial hash grids (Instant-NGP) or baked representations (Plenoxels), real-time rendering of full unbounded scenes remained elusive.

The deeper issue is architectural: NeRF uses an — the scene lives inside the network's weights and can only be queried, never directly drawn. 3D Gaussian Splatting asks: what if the scene representation were explicit — a collection of simple geometric primitives that the GPU already knows how to paint?

Open in Lab
Compare the NeRF ray-marching pipeline (slow, implicit) with Gaussian Splatting's rasterization pipeline (fast, explicit).
The demo wakes as you arrive…

The core idea: scenes as clouds of 3D Gaussians

A 3D Gaussian is the simplest possible "soft blob" in three-dimensional space. It is defined by a center point (mean μ) and a Σ that controls its size, shape, and orientation. Think of it as a fuzzy ellipsoid: dense in the center, fading smoothly to zero at the edges.

The key insight is that a single 3D Gaussian can represent a wide range of local scene geometry — a flat surface patch (a very thin, flat ellipsoid), a thin edge (a long, narrow ellipsoid), or a soft volume (a round blob). By using anisotropic Gaussians (different scales along different axes), each primitive adapts to the local geometry it needs to represent.

Each Gaussian carries five learnable properties: its position μ ∈ ℝ³, a covariance matrix Σ ∈ ℝ³ˣ³ (parameterized as a rotation quaternion q and a scale vector s for stable optimization), an opacity α ∈ [0,1], and spherical harmonic (SH) coefficients for . The SH coefficients let each Gaussian change color depending on the viewing angle — crucial for capturing specular highlights and other view-dependent effects.

Open in Lab
Drag the sliders to reshape a single 3D Gaussian — see how scale, rotation, opacity, and color create different surface elements.
The demo wakes as you arrive…
G(x)=e−12(x−μ)⊤Σ−1(x−μ)G(\mathbf{x}) = e^{-\frac{1}{2}(\mathbf{x}-\boldsymbol{\mu})^\top \boldsymbol{\Sigma}^{-1}(\mathbf{x}-\boldsymbol{\mu})}
3D Gaussian function — The Gaussian evaluates to 1 at the center μ and falls off smoothly in all directions. The covariance matrix Σ controls how quickly it fades along each axis — a large eigenvalue means the Gaussian extends far in that direction. To ensure Σ stays positive semi-definite during optimization, it is decomposed as Σ = R S Sᵀ Rᵀ, where R is a rotation matrix (from quaternion q) and S is a diagonal scale matrix.

Rendering: from 3D blobs to pixels

Rendering 3D Gaussians into a 2D image follows the same equation used in volume rendering, but computed via rasterization instead of ray marching. The process has three stages:

1. Projection: Each 3D Gaussian is projected onto the image plane using the camera's viewing transformation. The 3D covariance Σ transforms into a 2D covariance Σ′ via a first-order (Jacobian) approximation of the projective mapping. The result is a 2D ellipse on screen — each Gaussian's "splat."

2. Sorting: All Gaussians are sorted by depth. This is essential for correct alpha compositing — nearby Gaussians must be drawn before distant ones.

3. Alpha compositing: For each pixel, the sorted Gaussians that overlap it are blended front-to-back. Each Gaussian contributes its color weighted by its opacity and the 2D Gaussian evaluation at that pixel, multiplied by the (how much light has not yet been absorbed by closer Gaussians). Accumulation stops when transmittance drops below a threshold.

C(p)=∑i∈Nci αi Gi2D(p)∏j=1i−1(1−αj Gj2D(p))C(\mathbf{p}) = \sum_{i \in \mathcal{N}} c_i \, \alpha_i \, G_i^{2D}(\mathbf{p}) \prod_{j=1}^{i-1}\left(1 - \alpha_j \, G_j^{2D}(\mathbf{p})\right)
Alpha compositing equation for splatting — C(p) is the final pixel color. Each Gaussian i contributes color cᵢ, weighted by its opacity αᵢ, its 2D Gaussian value at pixel p, and the accumulated transmittance from all Gaussians closer to the camera. This is mathematically equivalent to NeRF's volume rendering integral but computed as a discrete sum over explicit primitives — making it much faster.
Open in Lab
Watch how overlapping Gaussians blend front-to-back to produce a final pixel color. Drag to reorder.
The demo wakes as you arrive…

The secret to speed: tile-based GPU rasterization

Computing the alpha compositing equation naively — testing every Gaussian against every pixel — would be far too slow. The paper's rendering engine uses a tile-based approach inspired by how modern GPUs render triangles:

Step 1 — : Discard any Gaussian outside the camera's viewing frustum, using a 99% confidence interval to bound each Gaussian's extent.

Step 2 — Tiling: Divide the screen into 16×16 pixel tiles. Assign each surviving Gaussian to every tile its 2D splat overlaps.

Step 3 — Sorting: For each tile, sort its Gaussians by depth using a fast GPU radix sort. The key trick: pack the tile ID and depth into a single integer so one global sort handles all tiles simultaneously.

Step 4 — Compositing: Each tile is processed by one GPU thread block. Threads within the block share a common list of sorted Gaussians via shared memory, compositing in parallel — one thread per pixel. Accumulation stops early when transmittance drops below 1/255.

This design maps perfectly to GPU hardware: tiles correspond to thread blocks, pixels to threads, and the shared Gaussian list exploits shared memory bandwidth. The result: 30–200+ fps at 1080p depending on scene complexity.

Open in Lab
See how the screen is divided into tiles and how Gaussians are sorted and composited within each tile.
The demo wakes as you arrive…

Adaptive density control: growing and pruning the Gaussian cloud

The optimization starts from a sparse produced by Structure-from-Motion (SfM). These initial points are far too few to represent a full scene — imagine trying to paint a detailed landscape with only a hundred dots. The adaptive density control scheme grows and refines the Gaussian set during training:

Cloning — when a Gaussian is too small and its positional gradient is large (the optimizer wants to move it a lot because the region is under-reconstructed), create a copy and move it in the direction of the gradient. This fills in missing geometry. Think of it as mitosis: one cell divides into two to cover more territory.

Splitting — when a Gaussian is too large and its gradient is also large (it's covering too much area and can't represent the fine details within), replace it with two smaller Gaussians. This refines coarse regions into sharper detail. Think of a large paintbrush being traded for two finer ones.

Pruning — periodically remove Gaussians that have become nearly transparent (α close to 0) or excessively large. Also reset all opacities to near-zero periodically, forcing Gaussians to re-earn their existence — those that don't contribute to any view fade away.

This grow-and-prune cycle runs every 100 optimization steps. A typical scene converges to 1–5 million Gaussians — orders of magnitude fewer than the number of MLP evaluations NeRF requires per frame.

Open in Lab
Watch cloning, splitting, and pruning in action as the Gaussian cloud evolves to match the scene.
The demo wakes as you arrive…

Training: the full optimization loop

The full pipeline optimizes all Gaussian parameters jointly using , with the differentiable rasterizer providing gradients. The combines an L1 photometric loss with a structural similarity term:

The L1 loss penalizes per-pixel color differences between the rendered image and the ground-truth photograph. D-SSIM captures structural similarity — it is sensitive to contrast and structure at a patch level, catching errors that pixel-level L1 would miss. The weighting λ = 0.2 was found empirically.

Training typically runs for 30,000 iterations, taking 20–50 minutes on a single GPU. This is competitive with or faster than NeRF methods, and the resulting representation renders in real time — a dramatic improvement over NeRF which required seconds per frame even after training.

L=(1−λ) L1+λ LD-SSIM\mathcal{L} = (1-\lambda)\,\mathcal{L}_1 + \lambda\, \mathcal{L}_{\text{D-SSIM}}
Training loss — L1 plus structural similarity — L1 measures per-pixel absolute difference. D-SSIM measures structural dissimilarity at a patch level (1 - SSIM). Together they balance pixel accuracy and perceptual quality. λ = 0.2 gives more weight to L1 for sharp reconstructions while D-SSIM prevents blurry regions.
3D Gaussian Splatting — simplified training looppython

Simplified to show the idea — not the real implementation.

import torch

def train_step(gaussians, camera, gt_image, iteration):
    """One optimization step for 3D Gaussian Splatting."""

    # 1. Project 3D Gaussians onto 2D image plane
    means_2d, covs_2d, depths = project_gaussians(
        gaussians.means,       # (N, 3) positions
        gaussians.quats,       # (N, 4) rotation quaternions
        gaussians.scales,      # (N, 3) scale vectors
        camera                 # camera intrinsics + extrinsics
    )

    # 2. Sort by depth for correct alpha compositing
    sorted_idx = torch.argsort(depths)

    # 3. Tile-based rasterization: splat onto image
    rendered = tile_rasterize(
        means_2d[sorted_idx],
        covs_2d[sorted_idx],
        gaussians.colors[sorted_idx],    # SH → RGB
        gaussians.opacities[sorted_idx],
        image_size=gt_image.shape[-2:]
    )

    # 4. Compute combined loss
    l1_loss = torch.abs(rendered - gt_image).mean()
    ssim_loss = 1.0 - ssim(rendered, gt_image)
    loss = 0.8 * l1_loss + 0.2 * ssim_loss

    # 5. Backprop through differentiable rasterizer
    loss.backward()
    optimizer.step()

    # 6. Adaptive density control every 100 steps
    if iteration % 100 == 0:
        clone_small_high_grad(gaussians)   # under-reconstruction
        split_large_high_grad(gaussians)   # over-reconstruction
        prune_transparent(gaussians)       # α ≈ 0 → remove

    return loss.item()

# No neural network at render time — just project, sort, composite!

Results: quality meets speed

The paper evaluated on 13 real scenes from three established benchmarks — Mip-NeRF 360, Tanks & Temples, and Deep Blending — plus the synthetic Blender dataset. The results were striking:

  • Quality: On Mip-NeRF 360, Gaussian Splatting matched or exceeded the of Mip-NeRF 360 itself — the reigning state-of-the-art — while also achieving higher SSIM and lower LPIPS scores on most scenes.

  • Speed: Rendering at 1080p ranged from 30 fps on complex outdoor scenes to over 200 fps on smaller indoor scenes — the first method to achieve real-time performance at this quality level on unbounded scenes.

  • Training time: 20–50 minutes on a single NVIDIA A6000 GPU — competitive with the fastest NeRF variants and far faster than the original NeRF (hours to days).

  • Representation size: 1–5 million Gaussians per scene, storing position, covariance, SH coefficients, and opacity. This is a few hundred MB — compact enough for practical deployment.

The quality-speed combination was unprecedented. Previous methods formed a clear trade-off curve; Gaussian Splatting broke through it.

Open in Lab
Each dot is a method. Gaussian Splatting (starred) breaks the quality-speed trade-off frontier.
The demo wakes as you arrive…

View-dependent color via spherical harmonics

Real surfaces rarely look the same from every angle. A metal table has bright specular highlights that shift as you move; a matte wall has subtle shading variations. To capture these view-dependent effects without a neural network, each Gaussian stores spherical harmonic (SH) coefficients instead of a single fixed color.

are a set of basis functions defined on the surface of a sphere. By storing coefficients for each SH band (up to degree 3, giving 16 coefficients per color channel), each Gaussian can represent how its color varies with viewing direction. Low-order terms capture diffuse color; higher-order terms model specular highlights and complex reflectance.

During rendering, the viewing direction from camera to each Gaussian is computed, and the SH coefficients are evaluated to produce the view-specific RGB color. This is much cheaper than querying a neural network and still captures rich appearance variation.

What 3D Gaussian Splatting unlocked

  1. 2020

    NeRF

    Neural Radiance Fields encode scenes as continuous functions inside an MLP. Beautiful quality but seconds-per-frame rendering. Launched the neural rendering revolution.

  2. 2022

    Instant-NGP

    Hash-grid acceleration cut NeRF training from hours to minutes. Near-real-time rendering on small scenes, but full 1080p real-time on unbounded scenes remained out of reach.

  3. 2023

    3D Gaussian Splatting

    Explicit Gaussian primitives + tile-based rasterization. First method to achieve real-time 1080p rendering at state-of-the-art quality. Over 22K GitHub stars.

  4. 2024

    4D Gaussian Splatting & Dynamic Scenes

    Extensions to dynamic scenes, deformable avatars, and SLAM. Gaussians gained a time dimension, enabling real-time novel view synthesis of moving scenes.

  5. 2024

    Impact on Video Generation

    Gaussian-based 3D representations influenced video generation systems and text-to-3D pipelines, bridging the gap between 2D generative models and 3D-consistent output.

The deepest legacy of 3D Gaussian Splatting is the paradigm shift it embodied: moving from implicit neural representations back to explicit geometric primitives, but now equipped with differentiable optimization and modern GPU rasterization. It showed that classical computer graphics ideas — point splatting, alpha compositing, tile-based rendering — gain new power when combined with gradient-based learning. The field did not need a better neural network; it needed a better representation.

CitationKerbl, Kopanas, Leimkühler, Drettakis. 3D Gaussian Splatting for Real-Time Radiance Field Rendering. ACM Transactions on Graphics (SIGGRAPH), 2023.

Terms in this paper