Computer Vision2020intermediate10 min read
NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis
NeRF: تمثيل المشاهد كحقول إشعاع عصبية لتوليف المناظر
Mildenhall, B. · Srinivasan, P. P. · Tancik, M. · Barron, J. T. · Ramamoorthi, R. · Ng, R. — ECCV
The problem
Synthesizing photorealistic images of a scene from new camera angles — given only a sparse set of photographs — remained an open problem. Traditional 3D representations like meshes and grids either lacked fine detail or consumed enormous memory. Existing neural approaches could not render complex real-world scenes with view-dependent effects like reflections and translucency at high resolution.
The contribution
Neural Radiance Fields (NeRF): represent a scene as a continuous 5D function — mapping spatial coordinates (x, y, z) and (θ, φ) to and — encoded entirely within an . lifts the low-dimensional inputs into a high-frequency space so the network can capture fine geometric detail. A hierarchical () sampling strategy concentrates samples where matter actually exists, making rendering efficient. Because is naturally differentiable, the network is trained end-to-end from posed 2D images alone.
The impact
NeRF ignited the neural rendering revolution. Within three years it spawned Instant-NGP (real-time ), Mip-NeRF (anti-aliased rendering), and ultimately 3D — which replaced the implicit MLP with explicit Gaussians for real-time rendering. NeRF's core insight — of an implicit scene function — underpins modern 3D content creation, autonomous driving simulation, and AR/VR.
Imagine you own a miniature world locked inside a snow globe. You can photograph it from any angle, but you can never open the globe to touch the objects inside.
Traditional 3D methods try to reconstruct the objects — build a tiny replica from your photos. NeRF takes a radically different approach: it trains a to memorize the light inside the globe. Give it any point in space and a viewing direction, and it tells you exactly what color and how opaque that point is. To take a new photo, you just cast imaginary rays through the globe and ask the network what each ray would see.
The problem: how do you capture a 3D world from flat photos?
You photograph a scene from 20–100 viewpoints. Now you want a photorealistic image from a viewpoint you never photographed — this is novel . Before NeRF, the main approaches each hit a wall:
- Mesh-based methods build explicit triangle surfaces but struggle with fuzzy or transparent objects like smoke, hair, and glass.
- Voxel grids divide space into tiny cubes, but the memory explodes cubically — a 512³ grid already costs 134 million voxels before you store a single color.
- Point clouds are sparse and produce holes when rendered.
NeRF sidesteps all three by storing the scene not as geometry, but as a learned function inside a neural network's weights.
The core idea: a neural network that knows what every point in space looks like
NeRF encodes an entire scene inside a single MLP (multi-layer perceptron). The network is a function that takes a 5D input — a 3D spatial location plus a 2D viewing direction — and outputs two things:
- Volume density — how opaque is this point? Is there solid matter here, or empty air?
- RGB color — what color does this point emit toward the given viewing direction?
The density depends only on the spatial location (a wall is a wall regardless of where you stand), while the color changes with viewing direction — this is how NeRF captures view-dependent effects like specular reflections on a shiny surface.
Positional encoding: teaching the network to see fine detail
MLPs are inherently biased toward learning smooth, low-frequency functions. A raw coordinate like changes too slowly for the network to carve out sharp edges and fine textures. NeRF solves this with positional encoding — the same idea from the Transformer paper, repurposed for 3D.
Each scalar coordinate is lifted into a high-dimensional using sinusoids at exponentially increasing frequencies. Think of it as translating one quiet number into a rich chord of harmonics — the low frequencies capture coarse shape, the high frequencies encode crisp edges and texture.
Volume rendering: from density field to pixels
To render a single pixel, NeRF casts a from the camera's eye through that pixel into the scene. The network is queried at sampled points along the ray. Each point contributes its color weighted by two factors:
- How dense it is — opaque points contribute more color.
- How much light has already been blocked by points closer to the camera — this is the .
The final pixel color is a weighted sum along the entire ray. This is classical volume rendering — the same physics used in medical CT scans and cinematic cloud effects — but here the density field comes from a neural network instead of a physical measurement.
Ray marching: walking through the scene point by point
To render an image, NeRF traces one ray per pixel. Along each ray, it samples points between the near and far bounds of the scene. At each point it queries the MLP for density and color, then composites them using the volume rendering equation.
Imagine walking along a laser beam through fog: at each step you note the fog's thickness (density) and tint (color). A thick cloud blocks what's behind it; clear air lets background colors shine through. The final pixel is everything you saw along the walk, weighted by how much light was left at each step.
Hierarchical sampling: spend samples where they matter
Sampling the ray uniformly is wasteful — most of empty space contributes nothing to the final color. NeRF uses a two-pass strategy:
- Coarse network: sample = 64 points uniformly. Run volume rendering to get a rough density profile along the ray — this tells you where stuff is.
- Fine network: use the coarse weights as a probability distribution to draw = 128 additional samples, biased toward regions with high density. Run a second, larger network on all samples for the final color.
This is like scanning a bookshelf with a flashlight (coarse pass) to find which shelves have books, then focusing your reading light there (fine pass).
The MLP architecture: how the network is wired
The NeRF MLP has a specific structure designed to separate geometry from appearance:
- 8 fully-connected layers (256 channels each, activation) process the positional-encoded spatial coordinates . A feeds the input back into the 5th layer — the same residual idea from ResNet.
- After the 8th layer, the network outputs volume density — this depends only on spatial location.
- A single extra layer takes the 256-dimensional feature vector from the 8th layer, concatenates the positional-encoded viewing direction, and outputs the RGB color.
This asymmetry is deliberate: density is view-independent (geometry doesn't change when you move), but color is view-dependent (a shiny surface looks different from different angles).
Training: learning from photos alone
Training NeRF requires only a set of images with known camera poses (obtained via Structure-from-Motion). For each training step:
- Sample a batch of camera rays from random training images.
- For each ray, sample points and run them through both the coarse and fine networks.
- Render the predicted pixel color using volume rendering.
- Compute the photometric — the between the rendered and actual pixel color.
- Backpropagate the through the rendering equation into the MLP weights.
After ~100–300K iterations (typically 1–2 days on a single ), the network has memorized the scene's geometry and appearance well enough to render photorealistic images from novel viewpoints.
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def positional_encoding(x, L=10):
"""Lift a scalar to 2L sinusoidal features."""
freqs = 2.0 ** np.arange(L) * np.pi
return np.concatenate([np.sin(x * freqs), np.cos(x * freqs)])
def volume_render(colors, densities, deltas):
"""Classic volume rendering along one ray.
colors: (N, 3) — RGB at each sample
densities: (N,) — sigma at each sample
deltas: (N,) — distance between consecutive samples
"""
alpha = 1.0 - np.exp(-densities * deltas) # opacity per sample
transmittance = np.cumprod(1.0 - alpha + 1e-10) # how much light remains
transmittance = np.concatenate([[1.0], transmittance[:-1]]) # shift right
weights = alpha * transmittance # contribution of each sample
pixel_color = (weights[:, None] * colors).sum(axis=0)
return pixel_color # (3,) — one RGB pixel
# The full NeRF pipeline per pixel:
# 1. Cast a ray from the camera through the pixel
# 2. Sample N points along the ray
# 3. Positional-encode each point's (x,y,z) and direction (θ,φ)
# 4. Feed into MLP → get (σ, r, g, b) per point
# 5. Volume-render the (σ, rgb) samples into one pixel color
# 6. MSE loss against ground-truth pixel → backprop into MLPWhy it mattered
2020
NeRF (this paper)
5D implicit scene representation with positional encoding and hierarchical sampling. Photorealistic results but slow — hours to train, seconds to render one image.
2021
Mip-NeRF
Replaced point samples with cone-traced volumes, eliminating aliasing artifacts and improving multi-scale rendering quality.
2022
Instant-NGP
Multi-resolution hash encoding replaced positional encoding, cutting training from hours to seconds. Made NeRF practical for interactive applications.
2023
3D Gaussian Splatting
Replaced the implicit MLP with millions of explicit 3D Gaussians. Achieves real-time (>30 FPS) rendering with quality matching NeRF, enabling AR/VR deployment.
2024
4D Gaussians & Dynamic Scenes
Extensions to 4D (space + time) enable real-time rendering of dynamic scenes — sports replays, telepresence, and virtual production.
NeRF's neural rendering paradigm — train on photos, render from anywhere — has become the foundation for the next generation of 3D content creation, from Google's Immersive View to autonomous driving simulation. Its direct descendant, 3D Gaussian Splatting, solved the speed problem and made real-time neural rendering a reality.
CitationMildenhall, Srinivasan, Tancik, Barron, Ramamoorthi, Ng. NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. ECCV, 2020.
Terms in this paper
- Neural Radiance Fieldحقل الإشعاع العصبي
- Volume Renderingالتصيير الحجمي
- View Synthesisتوليف المناظر
- Implicit Representationالتمثيل الضمني
- Ray Marchingالسير الشعاعي
- Positional Encodingالترميز الموضعي
- Volume Densityالكثافة الحجمية
- Viewing Directionاتجاه النظر
- Hierarchical Samplingاختيار العينات الهرمي
- Novel Viewالمنظر الجديد
- Camera Rayشعاع الكاميرا
- Transmittanceالنفاذية
- Alpha Compositingالدمج بالعتامة
- 5D Coordinateإحداثيات خماسية الأبعاد
- View-Dependent Colorاللون المعتمد على زاوية الرؤية