Computer Vision2020intermediate10 min read

NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis

NeRF: تمثيل المشاهد كحقول إشعاع عصبية لتوليف المناظر

Mildenhall, B. · Srinivasan, P. P. · Tancik, M. · Barron, J. T. · Ramamoorthi, R. · Ng, R. — ECCV

The problem

Synthesizing photorealistic images of a scene from new camera angles — given only a sparse set of photographs — remained an open problem. Traditional 3D representations like meshes and grids either lacked fine detail or consumed enormous memory. Existing neural approaches could not render complex real-world scenes with view-dependent effects like reflections and translucency at high resolution.

The contribution

Neural Radiance Fields (NeRF): represent a scene as a continuous 5D function — mapping spatial coordinates (x, y, z) and (θ, φ) to and — encoded entirely within an . lifts the low-dimensional inputs into a high-frequency space so the network can capture fine geometric detail. A hierarchical () sampling strategy concentrates samples where matter actually exists, making rendering efficient. Because is naturally differentiable, the network is trained end-to-end from posed 2D images alone.

The impact

NeRF ignited the neural rendering revolution. Within three years it spawned Instant-NGP (real-time ), Mip-NeRF (anti-aliased rendering), and ultimately 3D — which replaced the implicit MLP with explicit Gaussians for real-time rendering. NeRF's core insight — of an implicit scene function — underpins modern 3D content creation, autonomous driving simulation, and AR/VR.

Imagine you own a miniature world locked inside a snow globe. You can photograph it from any angle, but you can never open the globe to touch the objects inside.

Traditional 3D methods try to reconstruct the objects — build a tiny replica from your photos. NeRF takes a radically different approach: it trains a to memorize the light inside the globe. Give it any point in space and a viewing direction, and it tells you exactly what color and how opaque that point is. To take a new photo, you just cast imaginary rays through the globe and ask the network what each ray would see.

The problem: how do you capture a 3D world from flat photos?

You photograph a scene from 20–100 viewpoints. Now you want a photorealistic image from a viewpoint you never photographed — this is novel . Before NeRF, the main approaches each hit a wall:

  • Mesh-based methods build explicit triangle surfaces but struggle with fuzzy or transparent objects like smoke, hair, and glass.
  • Voxel grids divide space into tiny cubes, but the memory explodes cubically — a 512³ grid already costs 134 million voxels before you store a single color.
  • Point clouds are sparse and produce holes when rendered.

NeRF sidesteps all three by storing the scene not as geometry, but as a learned function inside a neural network's weights.

Open in Lab
Compare three traditional 3D representations with NeRF's implicit approach. Click each card to see its tradeoffs.
The demo wakes as you arrive…

The core idea: a neural network that knows what every point in space looks like

NeRF encodes an entire scene inside a single MLP (multi-layer perceptron). The network is a function FΘF_\Theta that takes a 5D input — a 3D spatial location (x,y,z)(x, y, z) plus a 2D viewing direction (θ,ϕ)(\theta, \phi) — and outputs two things:

  • Volume density σ\sigma — how opaque is this point? Is there solid matter here, or empty air?
  • RGB color (r,g,b)(r, g, b) — what color does this point emit toward the given viewing direction?

The density depends only on the spatial location (a wall is a wall regardless of where you stand), while the color changes with viewing direction — this is how NeRF captures view-dependent effects like specular reflections on a shiny surface.

FΘ:(x,y,z,θ,ϕ)→(σ,r,g,b)F_\Theta : (x, y, z, \theta, \phi) \rightarrow (\sigma, r, g, b)
The NeRF function — a 5D-to-4D mapping — Input: a 3D point plus a viewing direction. Output: volume density σ (geometry) and RGB color (appearance). The MLP's weights Θ encode the entire scene.
Open in Lab
Click each stage to see how NeRF transforms a 5D coordinate into a color and density.
The demo wakes as you arrive…

Positional encoding: teaching the network to see fine detail

MLPs are inherently biased toward learning smooth, low-frequency functions. A raw coordinate like (0.5,0.3,0.7)(0.5, 0.3, 0.7) changes too slowly for the network to carve out sharp edges and fine textures. NeRF solves this with positional encoding — the same idea from the Transformer paper, repurposed for 3D.

Each scalar coordinate pp is lifted into a high-dimensional using sinusoids at exponentially increasing frequencies. Think of it as translating one quiet number into a rich chord of harmonics — the low frequencies capture coarse shape, the high frequencies encode crisp edges and texture.

γ(p)=[sin⁡(20πp), cos⁡(20πp), …, sin⁡(2L−1πp), cos⁡(2L−1πp)]\gamma(p) = \left[\sin(2^0 \pi p),\, \cos(2^0 \pi p),\, \ldots,\, \sin(2^{L-1} \pi p),\, \cos(2^{L-1} \pi p)\right]
Positional encoding — from one number to 2L features — For spatial coordinates L = 10 (60 features from 3 inputs); for viewing direction L = 4 (24 features from 3 direction components). Without this step, the network produces blurry, over-smooth renderings.
Open in Lab
Drag the coordinate slider and watch how adding higher frequencies reveals finer spatial detail.
The demo wakes as you arrive…

Volume rendering: from density field to pixels

To render a single pixel, NeRF casts a from the camera's eye through that pixel into the scene. The network is queried at NN sampled points along the ray. Each point contributes its color weighted by two factors:

  • How dense it is — opaque points contribute more color.
  • How much light has already been blocked by points closer to the camera — this is the TT.

The final pixel color is a weighted sum along the entire ray. This is classical volume rendering — the same physics used in medical CT scans and cinematic cloud effects — but here the density field comes from a neural network instead of a physical measurement.

C^(r)=∑i=1NTi αi ci,Ti=∏j=1i−1(1−αj),αi=1−e−σiδi\hat{C}(\mathbf{r}) = \sum_{i=1}^{N} T_i \, \alpha_i \, \mathbf{c}_i, \quad T_i = \prod_{j=1}^{i-1}(1 - \alpha_j), \quad \alpha_i = 1 - e^{-\sigma_i \delta_i}
Volume rendering equation — blending colors along a ray — To determine the final color of a pixel, the renderer samples many points along a viewing ray and combines their contributions. Each point contributes according to three factors: its color, its opacity, and how much light remains after passing through all previous points. Dense regions block more light, while transparent regions allow more distant points to remain visible. The final pixel color is the accumulated result of all these contributions.
Open in Lab
Step through a single camera ray. Watch how transmittance drops as the ray hits dense regions.
The demo wakes as you arrive…

Ray marching: walking through the scene point by point

To render an image, NeRF traces one ray per pixel. Along each ray, it samples NN points between the near and far bounds of the scene. At each point it queries the MLP for density and color, then composites them using the volume rendering equation.

Imagine walking along a laser beam through fog: at each step you note the fog's thickness (density) and tint (color). A thick cloud blocks what's behind it; clear air lets background colors shine through. The final pixel is everything you saw along the walk, weighted by how much light was left at each step.

Open in Lab
Cast a ray through the scene. Watch samples accumulate color as the ray passes through objects.
The demo wakes as you arrive…

Hierarchical sampling: spend samples where they matter

Sampling the ray uniformly is wasteful — most of empty space contributes nothing to the final color. NeRF uses a two-pass strategy:

  • Coarse network: sample NcN_c = 64 points uniformly. Run volume rendering to get a rough density profile along the ray — this tells you where stuff is.
  • Fine network: use the coarse weights as a probability distribution to draw NfN_f = 128 additional samples, biased toward regions with high density. Run a second, larger network on all Nc+NfN_c + N_f samples for the final color.

This is like scanning a bookshelf with a flashlight (coarse pass) to find which shelves have books, then focusing your reading light there (fine pass).

Open in Lab
Toggle between uniform and hierarchical sampling. Notice how the fine pass concentrates samples near surfaces.
The demo wakes as you arrive…

The MLP architecture: how the network is wired

The NeRF MLP has a specific structure designed to separate geometry from appearance:

  • 8 fully-connected layers (256 channels each, activation) process the positional-encoded spatial coordinates (x,y,z)(x, y, z). A feeds the input back into the 5th layer — the same residual idea from ResNet.
  • After the 8th layer, the network outputs volume density σ\sigma — this depends only on spatial location.
  • A single extra layer takes the 256-dimensional feature vector from the 8th layer, concatenates the positional-encoded viewing direction, and outputs the RGB color.

This asymmetry is deliberate: density is view-independent (geometry doesn't change when you move), but color is view-dependent (a shiny surface looks different from different angles).

Open in Lab
Click each block to see its role. Notice where density and color branch apart.
The demo wakes as you arrive…

Training: learning from photos alone

Training NeRF requires only a set of images with known camera poses (obtained via Structure-from-Motion). For each training step:

  • Sample a batch of camera rays from random training images.
  • For each ray, sample points and run them through both the coarse and fine networks.
  • Render the predicted pixel color using volume rendering.
  • Compute the photometric — the between the rendered and actual pixel color.
  • Backpropagate the through the rendering equation into the MLP weights.

After ~100–300K iterations (typically 1–2 days on a single ), the network has memorized the scene's geometry and appearance well enough to render photorealistic images from novel viewpoints.

L=∑r∈R[∥C^c(r)−C(r)∥22+∥C^f(r)−C(r)∥22]\mathcal{L} = \sum_{\mathbf{r} \in \mathcal{R}} \left[ \| \hat{C}_c(\mathbf{r}) - C(\mathbf{r}) \|_2^2 + \| \hat{C}_f(\mathbf{r}) - C(\mathbf{r}) \|_2^2 \right]
The training loss — coarse + fine photometric error — The training objective compares the rendered image against the real image captured by the camera. For each sampled viewing ray, the model measures how closely its predicted color matches the ground-truth pixel color. Errors from both the coarse rendering stage and the fine rendering stage are included, allowing the two networks to be trained together and progressively improve reconstruction quality.

The same idea in code

Simplified NeRF rendering in NumPypython

Simplified to show the idea — not the real implementation.

import numpy as np

def positional_encoding(x, L=10):
    """Lift a scalar to 2L sinusoidal features."""
    freqs = 2.0 ** np.arange(L) * np.pi
    return np.concatenate([np.sin(x * freqs), np.cos(x * freqs)])

def volume_render(colors, densities, deltas):
    """Classic volume rendering along one ray.
    colors:    (N, 3)  — RGB at each sample
    densities: (N,)    — sigma at each sample
    deltas:    (N,)    — distance between consecutive samples
    """
    alpha = 1.0 - np.exp(-densities * deltas)          # opacity per sample
    transmittance = np.cumprod(1.0 - alpha + 1e-10)    # how much light remains
    transmittance = np.concatenate([[1.0], transmittance[:-1]])  # shift right
    weights = alpha * transmittance                    # contribution of each sample
    pixel_color = (weights[:, None] * colors).sum(axis=0)
    return pixel_color   # (3,) — one RGB pixel

# The full NeRF pipeline per pixel:
# 1. Cast a ray from the camera through the pixel
# 2. Sample N points along the ray
# 3. Positional-encode each point's (x,y,z) and direction (θ,φ)
# 4. Feed into MLP → get (σ, r, g, b) per point
# 5. Volume-render the (σ, rgb) samples into one pixel color
# 6. MSE loss against ground-truth pixel → backprop into MLP

Why it mattered

  1. 2020

    NeRF (this paper)

    5D implicit scene representation with positional encoding and hierarchical sampling. Photorealistic results but slow — hours to train, seconds to render one image.

  2. 2021

    Mip-NeRF

    Replaced point samples with cone-traced volumes, eliminating aliasing artifacts and improving multi-scale rendering quality.

  3. 2022

    Instant-NGP

    Multi-resolution hash encoding replaced positional encoding, cutting training from hours to seconds. Made NeRF practical for interactive applications.

  4. 2023

    3D Gaussian Splatting

    Replaced the implicit MLP with millions of explicit 3D Gaussians. Achieves real-time (>30 FPS) rendering with quality matching NeRF, enabling AR/VR deployment.

  5. 2024

    4D Gaussians & Dynamic Scenes

    Extensions to 4D (space + time) enable real-time rendering of dynamic scenes — sports replays, telepresence, and virtual production.

NeRF's neural rendering paradigm — train on photos, render from anywhere — has become the foundation for the next generation of 3D content creation, from Google's Immersive View to autonomous driving simulation. Its direct descendant, 3D Gaussian Splatting, solved the speed problem and made real-time neural rendering a reality.

CitationMildenhall, Srinivasan, Tancik, Barron, Ramamoorthi, Ng. NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. ECCV, 2020.

Terms in this paper