Computer Vision2023intermediate13 min read
3D Gaussian Splatting for Real-Time Radiance Field Rendering
التمثيل بالنقاط الغاوسية ثلاثية الأبعاد لتصيير حقول الإشعاع في الزمن الحقيقي
Kerbl, B. · Kopanas, G. · Leimkühler, T. · Drettakis, G. — ACM Transactions on Graphics (SIGGRAPH)
The problem
By 2023, Neural Fields (NeRF) had achieved remarkable quality for synthesis — generating photorealistic images of a scene from new camera angles. But NeRF and its successors shared a fundamental bottleneck: rendering required marching hundreds of rays through a neural network, making real-time display impossible for full scenes at high resolution. Faster methods like Instant-NGP accelerated but still could not reach 30 fps at 1080p on unbounded, real-world scenes. The field was stuck between quality and speed — no method could deliver both.
The contribution
A complete pipeline for photorealistic novel view synthesis that achieves state-of-the-art quality AND real-time rendering (≥30 fps at 1080p). Three interlocking innovations: (1) an explicit using millions of 3D Gaussians — each with learnable position, covariance, opacity, and spherical harmonic color coefficients; (2) an adaptive density control scheme that clones, splits, and prunes Gaussians during optimization to capture both fine detail and broad structure; (3) a tile-based differentiable GPU rasterizer that projects, sorts, and alpha-composites Gaussians in real time — no ray marching, no neural network at render time.
The impact
3D became the fastest-adopted representation in the history of neural rendering. Within months it spawned hundreds of follow-up papers covering dynamic scenes (4D-GS), SLAM, text-to-3D, avatars, autonomous driving, and more. It proved that explicit primitives with classical rasterization can match or exceed implicit neural representations in quality while being orders of magnitude faster. The work directly influenced video generation systems and became a foundational building block for real-time 3D content creation. GitHub stars exceeded 22,000 — a testament to both research impact and practical utility.
NeRF is like an oracle locked in a dark room. Hand it a camera position and direction through a slot, and it slowly computes what color you'd see — one pixel at a time. The oracle is brilliant but painfully slow: getting a single full image takes seconds.
3D Gaussian Splatting takes a completely different approach. Imagine filling a room with millions of tiny colored glass beads — each one can be stretched into an ellipsoid, tinted with any color, and made more or less transparent. You adjust every bead until, when you look through any window, the scene behind the glass perfectly matches a photograph. No oracle is needed at viewing time — just point a camera and the GPU paints all the beads onto the screen in milliseconds.
The problem: neural rendering is accurate but slow
NeRF represents a scene as a continuous function stored inside a neural network: given any 3D point and viewing direction, the network outputs a color and density. To render a single pixel, you march a ray into the scene, query the network at dozens of sample points along the ray, then composite those samples into a final color via .
This works beautifully — NeRF produces photorealistic images — but it is fundamentally slow. Every pixel requires multiple neural network forward passes. At 1080p, that means over two million pixels × dozens of samples each = hundreds of millions of network evaluations per frame. Even with spatial hash grids (Instant-NGP) or baked representations (Plenoxels), real-time rendering of full unbounded scenes remained elusive.
The deeper issue is architectural: NeRF uses an — the scene lives inside the network's weights and can only be queried, never directly drawn. 3D Gaussian Splatting asks: what if the scene representation were explicit — a collection of simple geometric primitives that the GPU already knows how to paint?
The core idea: scenes as clouds of 3D Gaussians
A 3D Gaussian is the simplest possible "soft blob" in three-dimensional space. It is defined by a center point (mean μ) and a Σ that controls its size, shape, and orientation. Think of it as a fuzzy ellipsoid: dense in the center, fading smoothly to zero at the edges.
The key insight is that a single 3D Gaussian can represent a wide range of local scene geometry — a flat surface patch (a very thin, flat ellipsoid), a thin edge (a long, narrow ellipsoid), or a soft volume (a round blob). By using anisotropic Gaussians (different scales along different axes), each primitive adapts to the local geometry it needs to represent.
Each Gaussian carries five learnable properties: its position μ ∈ ℝ³, a covariance matrix Σ ∈ ℝ³ˣ³ (parameterized as a rotation quaternion q and a scale vector s for stable optimization), an opacity α ∈ [0,1], and spherical harmonic (SH) coefficients for . The SH coefficients let each Gaussian change color depending on the viewing angle — crucial for capturing specular highlights and other view-dependent effects.
Rendering: from 3D blobs to pixels
Rendering 3D Gaussians into a 2D image follows the same equation used in volume rendering, but computed via rasterization instead of ray marching. The process has three stages:
1. Projection: Each 3D Gaussian is projected onto the image plane using the camera's viewing transformation. The 3D covariance Σ transforms into a 2D covariance Σ′ via a first-order (Jacobian) approximation of the projective mapping. The result is a 2D ellipse on screen — each Gaussian's "splat."
2. Sorting: All Gaussians are sorted by depth. This is essential for correct alpha compositing — nearby Gaussians must be drawn before distant ones.
3. Alpha compositing: For each pixel, the sorted Gaussians that overlap it are blended front-to-back. Each Gaussian contributes its color weighted by its opacity and the 2D Gaussian evaluation at that pixel, multiplied by the (how much light has not yet been absorbed by closer Gaussians). Accumulation stops when transmittance drops below a threshold.
The secret to speed: tile-based GPU rasterization
Computing the alpha compositing equation naively — testing every Gaussian against every pixel — would be far too slow. The paper's rendering engine uses a tile-based approach inspired by how modern GPUs render triangles:
Step 1 — : Discard any Gaussian outside the camera's viewing frustum, using a 99% confidence interval to bound each Gaussian's extent.
Step 2 — Tiling: Divide the screen into 16×16 pixel tiles. Assign each surviving Gaussian to every tile its 2D splat overlaps.
Step 3 — Sorting: For each tile, sort its Gaussians by depth using a fast GPU radix sort. The key trick: pack the tile ID and depth into a single integer so one global sort handles all tiles simultaneously.
Step 4 — Compositing: Each tile is processed by one GPU thread block. Threads within the block share a common list of sorted Gaussians via shared memory, compositing in parallel — one thread per pixel. Accumulation stops early when transmittance drops below 1/255.
This design maps perfectly to GPU hardware: tiles correspond to thread blocks, pixels to threads, and the shared Gaussian list exploits shared memory bandwidth. The result: 30–200+ fps at 1080p depending on scene complexity.
Adaptive density control: growing and pruning the Gaussian cloud
The optimization starts from a sparse produced by Structure-from-Motion (SfM). These initial points are far too few to represent a full scene — imagine trying to paint a detailed landscape with only a hundred dots. The adaptive density control scheme grows and refines the Gaussian set during training:
Cloning — when a Gaussian is too small and its positional gradient is large (the optimizer wants to move it a lot because the region is under-reconstructed), create a copy and move it in the direction of the gradient. This fills in missing geometry. Think of it as mitosis: one cell divides into two to cover more territory.
Splitting — when a Gaussian is too large and its gradient is also large (it's covering too much area and can't represent the fine details within), replace it with two smaller Gaussians. This refines coarse regions into sharper detail. Think of a large paintbrush being traded for two finer ones.
Pruning — periodically remove Gaussians that have become nearly transparent (α close to 0) or excessively large. Also reset all opacities to near-zero periodically, forcing Gaussians to re-earn their existence — those that don't contribute to any view fade away.
This grow-and-prune cycle runs every 100 optimization steps. A typical scene converges to 1–5 million Gaussians — orders of magnitude fewer than the number of MLP evaluations NeRF requires per frame.
Training: the full optimization loop
The full pipeline optimizes all Gaussian parameters jointly using , with the differentiable rasterizer providing gradients. The combines an L1 photometric loss with a structural similarity term:
The L1 loss penalizes per-pixel color differences between the rendered image and the ground-truth photograph. D-SSIM captures structural similarity — it is sensitive to contrast and structure at a patch level, catching errors that pixel-level L1 would miss. The weighting λ = 0.2 was found empirically.
Training typically runs for 30,000 iterations, taking 20–50 minutes on a single GPU. This is competitive with or faster than NeRF methods, and the resulting representation renders in real time — a dramatic improvement over NeRF which required seconds per frame even after training.
Simplified to show the idea — not the real implementation.
import torch
def train_step(gaussians, camera, gt_image, iteration):
"""One optimization step for 3D Gaussian Splatting."""
# 1. Project 3D Gaussians onto 2D image plane
means_2d, covs_2d, depths = project_gaussians(
gaussians.means, # (N, 3) positions
gaussians.quats, # (N, 4) rotation quaternions
gaussians.scales, # (N, 3) scale vectors
camera # camera intrinsics + extrinsics
)
# 2. Sort by depth for correct alpha compositing
sorted_idx = torch.argsort(depths)
# 3. Tile-based rasterization: splat onto image
rendered = tile_rasterize(
means_2d[sorted_idx],
covs_2d[sorted_idx],
gaussians.colors[sorted_idx], # SH → RGB
gaussians.opacities[sorted_idx],
image_size=gt_image.shape[-2:]
)
# 4. Compute combined loss
l1_loss = torch.abs(rendered - gt_image).mean()
ssim_loss = 1.0 - ssim(rendered, gt_image)
loss = 0.8 * l1_loss + 0.2 * ssim_loss
# 5. Backprop through differentiable rasterizer
loss.backward()
optimizer.step()
# 6. Adaptive density control every 100 steps
if iteration % 100 == 0:
clone_small_high_grad(gaussians) # under-reconstruction
split_large_high_grad(gaussians) # over-reconstruction
prune_transparent(gaussians) # α ≈ 0 → remove
return loss.item()
# No neural network at render time — just project, sort, composite!Results: quality meets speed
The paper evaluated on 13 real scenes from three established benchmarks — Mip-NeRF 360, Tanks & Temples, and Deep Blending — plus the synthetic Blender dataset. The results were striking:
-
Quality: On Mip-NeRF 360, Gaussian Splatting matched or exceeded the of Mip-NeRF 360 itself — the reigning state-of-the-art — while also achieving higher SSIM and lower LPIPS scores on most scenes.
-
Speed: Rendering at 1080p ranged from 30 fps on complex outdoor scenes to over 200 fps on smaller indoor scenes — the first method to achieve real-time performance at this quality level on unbounded scenes.
-
Training time: 20–50 minutes on a single NVIDIA A6000 GPU — competitive with the fastest NeRF variants and far faster than the original NeRF (hours to days).
-
Representation size: 1–5 million Gaussians per scene, storing position, covariance, SH coefficients, and opacity. This is a few hundred MB — compact enough for practical deployment.
The quality-speed combination was unprecedented. Previous methods formed a clear trade-off curve; Gaussian Splatting broke through it.
View-dependent color via spherical harmonics
Real surfaces rarely look the same from every angle. A metal table has bright specular highlights that shift as you move; a matte wall has subtle shading variations. To capture these view-dependent effects without a neural network, each Gaussian stores spherical harmonic (SH) coefficients instead of a single fixed color.
are a set of basis functions defined on the surface of a sphere. By storing coefficients for each SH band (up to degree 3, giving 16 coefficients per color channel), each Gaussian can represent how its color varies with viewing direction. Low-order terms capture diffuse color; higher-order terms model specular highlights and complex reflectance.
During rendering, the viewing direction from camera to each Gaussian is computed, and the SH coefficients are evaluated to produce the view-specific RGB color. This is much cheaper than querying a neural network and still captures rich appearance variation.
What 3D Gaussian Splatting unlocked
2020
NeRF
Neural Radiance Fields encode scenes as continuous functions inside an MLP. Beautiful quality but seconds-per-frame rendering. Launched the neural rendering revolution.
2022
Instant-NGP
Hash-grid acceleration cut NeRF training from hours to minutes. Near-real-time rendering on small scenes, but full 1080p real-time on unbounded scenes remained out of reach.
2023
3D Gaussian Splatting
Explicit Gaussian primitives + tile-based rasterization. First method to achieve real-time 1080p rendering at state-of-the-art quality. Over 22K GitHub stars.
2024
4D Gaussian Splatting & Dynamic Scenes
Extensions to dynamic scenes, deformable avatars, and SLAM. Gaussians gained a time dimension, enabling real-time novel view synthesis of moving scenes.
2024
Impact on Video Generation
Gaussian-based 3D representations influenced video generation systems and text-to-3D pipelines, bridging the gap between 2D generative models and 3D-consistent output.
The deepest legacy of 3D Gaussian Splatting is the paradigm shift it embodied: moving from implicit neural representations back to explicit geometric primitives, but now equipped with differentiable optimization and modern GPU rasterization. It showed that classical computer graphics ideas — point splatting, alpha compositing, tile-based rendering — gain new power when combined with gradient-based learning. The field did not need a better neural network; it needed a better representation.
CitationKerbl, Kopanas, Leimkühler, Drettakis. 3D Gaussian Splatting for Real-Time Radiance Field Rendering. ACM Transactions on Graphics (SIGGRAPH), 2023.
Terms in this paper
- Gaussian Splattingالتطاير الغاوسي
- Neural Radiance Fieldحقل الإشعاع العصبي
- Novel Viewالمنظر الجديد
- Volume Renderingالتصيير الحجمي
- Alpha Compositingالدمج بالعتامة
- Spherical Harmonicsالتوافقيات الكروية
- Anisotropicلا متناحٍ
- Covariance Matrixمصفوفة التباين المشترك
- Differentiable Renderingالتصيير القابل للاشتقاق
- Point Cloudسحابة النقاط
- Structure from Motionالبنية من الحركة
- View-Dependent Colorاللون المعتمد على زاوية الرؤية
- Radianceالإشعاع
- Camera Poseوضعية الكاميرا
- Scene Representationتمثيل المشهد