Generative Models2016intermediate12 min read

WaveNet: A Generative Model for Raw Audio

WaveNet: نموذج توليدي للصوت الخام

van den Oord, A. · Dieleman, S. · Zen, H. · Simonyan, K. · Vinyals, O. · Graves, A. · Kalchbrenner, N. · Senior, A. · Kavukcuoglu, K. — SSW (ISCA)

The problem

Before WaveNet, speech synthesis relied on two approaches: concatenative systems that stitched together pre-recorded fragments (producing robotic transitions), and parametric systems that generated audio from statistical models of vocal features (producing muffled, unnatural speech). Both operated on hand-engineered audio features, never on the raw itself. Music generation faced similar limits. Meanwhile, raw audio has extreme temporal resolution — at least 16,000 samples per second — and each sample can take one of 65,536 values (16-bit). No had successfully modeled such sequences.

The contribution

WaveNet: a fully probabilistic autoregressive neural network that generates raw audio waveform sample by sample. It uses stacks of dilated causal convolutions to build an exponentially large with few layers, units inspired by LSTMs, and to reduce 65,536 output values to 256 tractable categories. Residual and skip connections enable of very deep networks. When conditioned on text features, it achieves state-of-the-art quality; when trained on music, it generates novel, realistic musical fragments.

The impact

WaveNet proved that neural networks can model raw audio at sample level with unprecedented quality. It was adopted by Google for its Cloud Text-to-Speech service and directly inspired Tacotron 2's neural , Parallel WaveNet for real-time synthesis, and OpenAI's Jukebox for music generation. Its architecture became a standard building block for audio processing and time-series modeling.

Traditional speech synthesizers are like a collage artist: they cut out pre-recorded fragments and paste them together, hoping the seams don't show. Or they're like a paint-by-numbers kit — they fill in predetermined patterns that look like speech but never quite sound human.

WaveNet is more like a pointillist painter: it places one tiny dot at a time, each dot chosen after studying every dot that came before. Up close, it's just dots — 16,000 per second. Step back, and you hear a human voice.

The problem: machines that talk but dont sound real

Before WaveNet, speech synthesis used two families of methods, each with a fundamental flaw:

  • Concatenative synthesis — stitch together short clips of recorded speech. High quality inside each clip, but audible glitches at the boundaries. The system needs a massive database of recordings, and it can't generalize beyond them.

  • Parametric synthesis — a statistical model (typically a ) generates acoustic features like pitch and spectral envelopes, then a vocoder converts those features to audio. Faster and more flexible, but the generated speech sounds muffled and robotic because the vocoder makes simplifying assumptions — for example, that speech signals follow a Gaussian distribution, which they don't.

Both approaches operate on hand-engineered features, never on the raw waveform. They treat audio as something to be assembled from parts, not generated from scratch. And raw audio is terrifyingly high-dimensional: at 16 kHz, a single second is a sequence of 16,000 samples, each taking one of 65,536 possible values. No neural network had successfully generated sequences this long and this detailed.

Open in Lab
Compare concatenative synthesis (glitchy joins) with WaveNet's smooth, sample-level generation.
The demo wakes as you arrive…

The core idea: predict one sample at a time

WaveNet borrows its foundation from PixelCNN: model a complex distribution by breaking it into a chain of simple conditional distributions. For an audio waveform x={x1,x2,…,xT}\mathbf{x} = \{x_1, x_2, \ldots, x_T\}, the joint probability factors as:

p(x)=∏t=1Tp(xt∣x1,…,xt−1)p(\mathbf{x}) = \prod_{t=1}^{T} p(x_t \mid x_1, \ldots, x_{t-1})
Autoregressive factorization — the chain rule of probability applied to audio — Each sample depends on all previous samples. The model learns to output a categorical distribution over 256 possible values for the next sample, given everything before it.

Think of it as a storyteller who never skips ahead: every word (sample) is chosen knowing every word that came before, but without peeking at any word that comes after. This is what makes the model autoregressive — each prediction depends only on the past, maintaining strict temporal causality.

But here's the challenge: the "story" of audio is told at 16,000 words per second. How do you give the model enough memory of the past without making it impossibly deep?

Causal convolutions: no peeking at the future

In a standard , each output can see inputs from both sides — past and future. That's fine for classification, but fatal for generation: you can't predict the next sample if you've already seen it.

A shifts the filter so it only looks at the current and previous time steps. The output at time tt depends only on inputs at times t,t−1,t−2,…t, t-1, t-2, \ldots — never on t+1t+1 or beyond. It's like a one-way mirror: information flows forward in time, never backward.

But causal convolutions have a limitation: the receptive field — the number of past samples the model can consider — grows only linearly with the number of layers. To see 1,024 past samples, you'd need 1,024 layers. That's impractical.

Dilated convolutions: exponential memory with linear layers

The key architectural innovation in WaveNet is the dilated causal convolution. Instead of reading consecutive samples, each layer skips a fixed number of samples between its inputs. The skipping pattern — called the — doubles with each layer: 1, 2, 4, 8, 16, ...

Picture a highway with exits: the first layer checks every exit (dilation 1), the second checks every other exit (dilation 2), the third checks every fourth exit (dilation 4). After just 10 layers with a filter size of 2, the network can see 210=1,0242^{10} = 1,024 past samples — enough to span about 64 milliseconds of audio at 16 kHz. Stack multiple such blocks and you cover hundreds of milliseconds, which is more than enough to capture speech dynamics like syllables, pitch contours, and even rhythm.

This exponential growth is the crucial trick: logarithmic depth for linear receptive field. Ten layers give the reach of a thousand-layer causal network, at a fraction of the cost.

Open in Lab
Watch how dilation doubles per layer, building an exponentially large receptive field. Compare with standard causal convolutions.
The demo wakes as you arrive…
Receptive field=1+B×∑l=0L−1(k−1)×2l=1+B×(k−1)×(2L−1)\text{Receptive field} = 1 + B \times \sum_{l=0}^{L-1} (k-1) \times 2^l = 1 + B \times (k-1) \times (2^L - 1)
Receptive field size — exponential growth from dilation — B = number of dilation blocks · L = layers per block · k = filter size. With B=1, L=10, k=2, the receptive field is 1,024 — from only 10 layers.

μ-law companding: making 65,536 values manageable

Raw audio is stored as 16-bit integers, meaning each sample has 65,536 possible values. Using a output over 65,536 categories would be enormously expensive. WaveNet solves this with μ-law companding — a nonlinear transformation borrowed from telecommunications that compresses the range to just 256 values (8-bit).

The key insight is that human hearing is logarithmic: the difference between volumes 100 and 200 sounds much larger than between 10,000 and 10,100, even though the absolute gap is the same. μ-law encoding respects this by allocating more precision to quiet sounds (where our ears are sensitive) and less to loud sounds (where we can't tell the difference anyway).

f(xt)=sign(xt)ln⁡(1+μ∣xt∣)ln⁡(1+μ)f(x_t) = \text{sign}(x_t) \frac{\ln(1 + \mu |x_t|)}{\ln(1 + \mu)}
μ-law companding transform (μ = 255) — Compresses the input signal nonlinearly. Quiet regions get finer resolution; loud regions are coarsened. After quantizing to 256 bins, the reconstruction quality far exceeds simple linear quantization.
Open in Lab
Compare linear vs μ-law quantization. Notice how μ-law preserves detail in quiet regions.
The demo wakes as you arrive…

Gated activation units: learning what to remember and what to forget

Simple ReLU activations work for many tasks, but audio generation needs something more expressive. WaveNet borrows the gating mechanism from LSTMs: each layer computes two parallel transformations of the input — a filter (tanh) that proposes what content to produce, and a gate (sigmoid) that decides how much of that content to let through. The element-wise product of the two produces the final output.

Think of it as a recording studio: the filter is the musician playing all possible notes, and the gate is the sound engineer at the mixing board, deciding which channels to amplify and which to mute. Together they sculpt the output more precisely than either could alone.

z=tanh⁡(Wf,k∗x)⊙σ(Wg,k∗x)\mathbf{z} = \tanh(W_{f,k} * \mathbf{x}) \odot \sigma(W_{g,k} * \mathbf{x})
Gated activation unit — Wf,kW_{f,k} = learned filter convolution · Wg,kW_{g,k} = learned gate convolution · ∗* = (dilated) convolution · ⊙\odot = element-wise multiplication. The tanh branch proposes content; the sigmoid branch controls flow.

Residual and skip connections: information highways through deep networks

WaveNet uses dozens of layers, and deep networks are notoriously hard to train — gradients vanish or explode as they travel backward through many layers. WaveNet addresses this with two types of shortcut connections:

  • Residual connections add the input of each block directly to its output: output=f(x)+x\text{output} = f(\mathbf{x}) + \mathbf{x}. This creates a "highway" that lets gradients flow past each block unimpeded. The block only needs to learn the residual — what to add to the input — rather than the full transformation.

  • Skip connections send each block's gated output directly to the final output layers, bypassing all subsequent blocks. These create shortcuts from every depth level to the surface, so early layers' signals aren't diluted by passing through the entire stack. All skip signals are summed together and fed through two ReLU + 1×1 convolution layers followed by a softmax to predict the next sample.

Together, they let the network be much deeper than would otherwise be trainable, which directly translates to a larger receptive field and better audio quality.

Open in Lab
Click each component to see its role. Follow the data flow from input through the gated activation, residual addition, and skip path.
The demo wakes as you arrive…

Putting it all together

The full WaveNet architecture stacks everything together:

  1. Input: raw audio samples, μ-law encoded to 256 values, fed as one-hot vectors through a causal 1×1 convolution.

  2. Residual blocks: a stack of dilated causal convolution layers with gated activations. Each block outputs both a residual path (added back to the input for the next block) and a skip path (sent directly to the output). Dilation rates double per layer: 1, 2, 4, ..., 512, then repeat this cycle multiple times.

  3. Output: all skip connections are summed, passed through ReLU → 1×1 conv → ReLU → 1×1 conv → softmax over 256 values. The result is a probability distribution over the next audio sample.

The model is trained to maximize the of the training audio data. At generation time, it samples from the predicted distribution, feeds the sample back as input, and repeats — one sample at a time.

Open in Lab
Explore the full WaveNet architecture. Click any layer to see what it does.
The demo wakes as you arrive…

Conditioning: teaching WaveNet what to say and how to sound

An unconditioned WaveNet generates interesting but uncontrolled audio — babbling voices, ambient textures. To make it useful, the model is conditioned on additional information that guides what it generates.

WaveNet supports two types of :

  • Global conditioning adds a single conditioning vector h\mathbf{h} (like a speaker identity ) to every layer. This is like telling the model "speak in this person's voice" — the same instruction applies everywhere: z=tanh⁡(Wf∗x+VfTh)⊙σ(Wg∗x+VgTh)\mathbf{z} = \tanh(W_f * \mathbf{x} + V_f^T \mathbf{h}) \odot \sigma(W_g * \mathbf{x} + V_g^T \mathbf{h})

  • Local conditioning provides a time-varying signal ht\mathbf{h}_t — such as linguistic features or mel-spectrogram frames from a text-to-speech frontend. This tells the model what to say at each moment: z=tanh⁡(Wf∗x+Vf∗y)⊙σ(Wg∗x+Vg∗y)\mathbf{z} = \tanh(W_f * \mathbf{x} + V_f * \mathbf{y}) \odot \sigma(W_g * \mathbf{x} + V_g * \mathbf{y}) where y\mathbf{y} is the upsampled conditioning signal matching the audio resolution.

With speaker conditioning, a single WaveNet can generate speech in multiple voices — switching seamlessly between speakers by changing the conditioning vector.

Open in Lab
Toggle between unconditioned, globally conditioned, and locally conditioned generation.
The demo wakes as you arrive…

The core in code

WaveNet residual block with gated activationpython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn as nn
import torch.nn.functional as F

class WaveNetResidualBlock(nn.Module):
    """One residual block: dilated causal conv → gated activation → residual + skip."""

    def __init__(self, residual_channels, gate_channels, skip_channels, dilation):
        super().__init__()
        # Dilated causal convolution: kernel_size=2, dilation doubles each layer
        self.conv = nn.Conv1d(
            residual_channels, gate_channels * 2,  # filter + gate outputs
            kernel_size=2, dilation=dilation, padding=0
        )
        self.res_conv = nn.Conv1d(gate_channels, residual_channels, 1)  # 1x1
        self.skip_conv = nn.Conv1d(gate_channels, skip_channels, 1)     # 1x1

    def forward(self, x):
        # Causal padding: pad only the left side
        pad = self.conv.dilation[0]
        x_padded = F.pad(x, (pad, 0))

        h = self.conv(x_padded)
        # Split into filter (tanh) and gate (sigmoid)
        h_filter, h_gate = h.chunk(2, dim=1)
        z = torch.tanh(h_filter) * torch.sigmoid(h_gate)

        skip = self.skip_conv(z)          # → skip output (to final layers)
        residual = self.res_conv(z) + x   # → residual output (to next block)
        return residual, skip

# Stack blocks with exponentially growing dilation:
# dilation = [1, 2, 4, 8, ..., 512, 1, 2, 4, 8, ..., 512, ...]
# Each cycle covers 2^10 = 1024 samples of receptive field.

Results: how WaveNet sounds

WaveNet was evaluated on text-to-speech in both English and Mandarin using Mean Opinion Score (MOS) tests — human listeners rate naturalness on a scale from 1 (completely unnatural) to 5 (completely natural). The results were striking:

For English, WaveNet achieved a MOS of 4.21, far surpassing concatenative synthesis (3.86) and parametric synthesis (2.17). For Mandarin, it scored 4.08 vs. 4.21 for concatenative and 2.98 for parametric. This was the first time a neural model substantially closed the gap with natural speech (rated 4.55 for English).

Beyond speech, when trained on piano music, WaveNet generated novel musical fragments that captured realistic timbre, dynamics, and even rudimentary melodic structure — all without any music-specific engineering.

Open in Lab
Mean Opinion Scores comparing WaveNet with traditional synthesis methods.
The demo wakes as you arrive…

Why it changed everything

  1. 2016

    WaveNet

    Sample-level autoregressive model for raw audio. Dilated causal convolutions give exponential receptive field. First neural TTS to approach human quality.

  2. 2017

    Parallel WaveNet

    Used probability density distillation to train a parallel (non-autoregressive) student network, enabling real-time synthesis — 1000× faster than the original.

  3. 2017

    Tacotron 2

    Combined a sequence-to-sequence model for text-to-spectrogram with a modified WaveNet vocoder. Achieved MOS nearly indistinguishable from human speech (4.53).

  4. 2018

    WaveGlow

    Flow-based model combining WaveNet-style architecture with invertible 1×1 convolutions for parallel, real-time speech synthesis without distillation.

  5. 2020

    Jukebox (OpenAI)

    Extended autoregressive audio generation to full songs — minutes of music with vocals, lyrics, and style conditioning, building on WaveNet's core principles.

  6. 2021

    Google Cloud TTS with WaveNet voices

    WaveNet voices became commercially available across dozens of languages, making neural speech synthesis accessible to millions of developers.

WaveNet's architecture — dilated causal convolutions, gated activations, residual and skip connections — became the template for neural audio synthesis. Its descendants Tacotron 2 and Jukebox pushed the frontier from short clips to full-length speech and music. The core idea — autoregressive generation of raw waveforms — proved that neural networks can master the most fine-grained generative task in audio.

Citationvan den Oord, Dieleman, Zen, Simonyan, Vinyals, Graves, Kalchbrenner, Senior, Kavukcuoglu. WaveNet: A Generative Model for Raw Audio. SSW (ISCA), 2016.

Terms in this paper