Generative Models2018intermediate11 min read
Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions
توليد الكلام الطبيعي بتهيئة WaveNet على تنبؤات طيف ميل
Shen, J. · Pang, R. · Weiss, R. J. · Schuster, M. · Jaitly, N. · Yang, Z. · Chen, Z. · Zhang, Y. · Wang, Y. · Skerry-Ryan, RJ · Saurous, R. A. · Agiomyrgiannakis, Y. · Wu, Y. — ICASSP
The problem
Traditional systems were built from a fragile pipeline of separate components: a text analyzer, a duration model, a pitch (F0) predictor, and a . Each component was engineered by hand with linguistic rules and signal-processing heuristics, so errors cascaded from stage to stage and the resulting speech sounded robotic. WaveNet showed that a neural vocoder could produce natural-sounding audio, but it required hand-engineered linguistic features as input. Tacotron 1 showed that a model could predict spectrograms from text, but it used the Griffin-Lim algorithm for waveform synthesis, producing audible artifacts. No system had combined neural spectrogram prediction with a neural vocoder into a single, fully neural pipeline that achieves human-level naturalness.
The contribution
Tacotron 2 combines a recurrent sequence-to-sequence network with — predicting 80-channel mel spectrograms directly from characters — with a modified WaveNet vocoder that converts those spectrograms into time-domain waveforms. The serves as a compact acoustic intermediate representation that decouples the two networks: the first decides what to say, the second decides how to render it at sample level. This design achieves a (MOS) of 4.53, nearly matching professional recordings (4.58), and dramatically simplifies the WaveNet architecture because the mel spectrogram already captures long-range acoustic structure.
The impact
Tacotron 2 proved that a fully neural text-to-speech pipeline could match human speech quality, setting the benchmark for all subsequent TTS systems. Its mel-spectrogram-as-interface idea became the standard: FastSpeech, VITS, and modern diffusion-based TTS systems all predict mel spectrograms first. The work inspired the decoupling paradigm — separating high-level acoustic modeling from waveform generation — that remains central to speech synthesis today. Whisper, Google's Assistant voice, and virtually every modern TTS system traces its lineage to this architecture.
Imagine you're a choir director reading a script for the first time. You go through two mental stages: first, you mark up the sheet music — deciding the pitch, rhythm, pauses, and emphasis for each phrase. Then you hand that annotated score to the singers, who turn your markings into actual sound waves that fill the concert hall.
Tacotron 2 works the same way. The first network (the spectrogram predictor) reads the text and paints a detailed acoustic blueprint — the mel spectrogram. The second network (WaveNet) takes that blueprint and renders it into the rich, nuanced waveform you actually hear. The blueprint is the contract between "what to say" and "how to sound."
The problem: brittle pipelines and robotic speech
Before Tacotron 2, building a TTS system meant stitching together a long chain of hand-crafted modules:
- A text analyzer that converts raw text into linguistic features (phonemes, stress patterns, part-of-speech tags).
- A duration model that predicts how long each phoneme lasts.
- An F0 predictor that decides the pitch contour.
- A vocoder that synthesizes a waveform from these acoustic features.
Each module was engineered with domain-specific rules and signal-processing tricks. Errors at any stage — a wrong phoneme, an unnatural duration — would cascade downstream, and the speech would sound stilted. Two breakthroughs had appeared but remained incomplete:
WaveNet (2016) showed that a neural vocoder could generate astonishingly natural audio, but it needed hand-crafted linguistic features as input — the same fragile pipeline, just with a better last stage.
Tacotron 1 (2017) showed that a sequence-to-sequence model could predict spectrograms directly from characters, but it used the Griffin-Lim algorithm to convert spectrograms to audio, introducing phase-estimation artifacts that made the speech sound hollow.
The bridge: mel spectrograms as acoustic blueprints
The key insight is choosing the right intermediate representation — the contract between the text-understanding network and the audio-rendering network. Tacotron 2 uses 80-channel mel spectrograms computed from 50 ms frames with a 12.5 ms hop.
A mel spectrogram is a time-frequency picture of sound, but warped to match how the human ear actually perceives pitch. Low frequencies get more resolution (because our ears are most sensitive there), and high frequencies get compressed. Think of it as a heat map of sound: the x-axis is time, the y-axis is frequency (on the mel scale), and the color intensity is energy.
Why mel spectrograms instead of raw linguistic features? Three reasons:
- They are compact. 80 channels at 12.5 ms hops vs. 16,000+ raw audio samples per second — a 200× compression that still captures pitch, timbre, and rhythm.
- They are learnable. No hand-engineered phoneme dictionaries or prosody rules needed.
- They simplify WaveNet. Because the spectrogram already encodes long-range acoustic structure, WaveNet no longer needs an enormous receptive field to "figure out" what word it's in the middle of — it can focus on rendering fine audio detail.
Network 1: from characters to spectrograms
The spectrogram prediction network is a sequence-to-sequence model with three main parts: an , a location-sensitive mechanism, and an .
The encoder takes a sequence of characters, embeds each into a 512-dimensional vector, then passes them through 3 convolutional layers (512 filters, kernel size 5) with and ReLU. These convolutions capture local character context — each filter sees a character and its two neighbors on each side. Finally, a with 256 units per direction produces the encoded representation. Think of the encoder as a reader that understands both the individual letters and how they flow together into words and phrases.
Location-sensitive attention extends the standard mechanism by also looking at where the model was attending in the previous step. This prevents the common failure modes of attention in TTS: repeating a word (attention gets stuck) or skipping a word (attention jumps too far). The attention history acts like a bookmark that says "I've read up to here — move forward, don't repeat."
The decoder is autoregressive: at each step it predicts one frame of the mel spectrogram, then feeds that prediction back as input for the next step. The previous prediction passes through a pre-net (2 fully connected layers of 256 units with ReLU and ) that acts as an information bottleneck — it forces the model to learn a robust attention alignment rather than memorizing exact frame values. The pre-net output and attention context are concatenated and fed through 2 unidirectional LSTM layers (1024 units each). The LSTM output is used to predict the current mel frame and a that signals end of utterance.
The post-net: polishing the spectrogram
The decoder's raw mel predictions are reasonable but lack fine detail. A 5-layer convolutional post-net (512 filters, kernel size 5, with batch normalization and tanh on all but the last layer) predicts a residual — a correction to add to the initial prediction. This is a : instead of asking the post-net to predict the entire spectrogram from scratch, it only needs to learn the difference between "good enough" and "great."
This is exactly like an artist's workflow: the decoder does the rough sketch, and the post-net adds shading, highlights, and fine brushstrokes. The system trains with two loss terms: one on the raw decoder output and one on the post-net-enhanced output, so both the sketch and the refined version learn to be accurate.
Network 2: WaveNet turns the blueprint into sound
The modified WaveNet vocoder takes the mel spectrogram and generates audio samples at 24 kHz — sample by sample, autoregressively. It uses the same core idea as the original WaveNet: stacked dilated causal convolutions that give each sample a large receptive field over past samples.
But there's a crucial simplification. The original WaveNet needed 30 dilated layers to achieve a receptive field large enough to capture long-range speech structure (what word is being said, what the intonation pattern is). Tacotron 2's WaveNet needs far fewer layers because the mel spectrogram already encodes that information. The vocoder just needs enough context to render smooth, natural-sounding audio at the sample level.
The paper's ablation studies confirmed this: reducing the receptive field by a factor of 25 still produced high-quality audio, but eliminating dilated convolutions entirely degraded quality significantly. There's a "sweet spot" — the vocoder needs some local context at the waveform level, but the heavy lifting of linguistic and prosodic planning is already done by the spectrogram network.
Training the two networks
A key design choice: the two networks are trained independently. The spectrogram network is trained first using — at each decoder step, the ground-truth previous mel frame is fed as input (instead of the model's own prediction). This stabilizes and helps the attention learn clean, diagonal alignments.
Once the spectrogram network is trained, it generates mel spectrograms for the training data. These predicted (not ground-truth) spectrograms are used to train the WaveNet vocoder. Why not use ground-truth spectrograms? Because at time, WaveNet will see predicted spectrograms — which may have small imperfections. Training on predicted spectrograms teaches WaveNet to be robust to these imperfections, bridging the train-test gap.
The for the spectrogram network is straightforward (MSE) on both the raw decoder output and the post-net output, plus on the stop token. No adversarial training, no complex perceptual losses — just MSE, and the system achieves near-human quality.
The core idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def encoder(chars, W_embed, conv_filters, lstm):
"""Encode a character sequence into hidden representations."""
# Step 1: character embedding (512-dim)
embedded = W_embed[chars] # (T_text, 512)
# Step 2: 3 conv layers capture local context (2 neighbors each side)
h = embedded
for conv in conv_filters:
h = relu(batch_norm(conv1d(h, conv))) # (T_text, 512)
# Step 3: bidirectional LSTM captures global context
h_enc = bilstm(h, lstm) # (T_text, 512)
return h_enc
def decoder_step(prev_mel, h_enc, attn_state, lstm_state, prenet, lstm):
"""Predict one mel frame, conditioned on encoder output + previous frame."""
# Pre-net: information bottleneck with dropout
h_pre = dropout(relu(prenet[1] @ dropout(relu(prenet[0] @ prev_mel))))
# Location-sensitive attention: "where was I? move forward"
context, attn_state = location_attention(h_enc, lstm_state, attn_state)
# Decoder LSTM: 2 layers, 1024 units each
lstm_input = np.concatenate([h_pre, context])
lstm_out, lstm_state = lstm_forward(lstm_input, lstm_state, lstm)
# Predict mel frame + stop token
mel_frame = linear_project(np.concatenate([lstm_out, context])) # 80 channels
stop_prob = sigmoid(linear_project_stop(np.concatenate([lstm_out, context])))
return mel_frame, stop_prob, attn_state, lstm_stateWhat the ablation studies revealed
The authors systematically removed or modified components to measure their contribution:
- Mel spectrogram vs. linguistic features as WaveNet input. Mel spectrograms won decisively — they produced more natural speech without requiring any text processing pipeline.
- WaveNet receptive field size. Reducing the receptive field by 25× (from 256 ms to about 10 ms) barely affected quality. But removing dilated convolutions entirely (receptive field drops to ~0.5 ms) caused significant degradation.
- Post-net contribution. The 5-layer convolutional post-net consistently improved spectrogram quality by learning to add fine spectral detail that the autoregressive decoder misses.
- The system achieved MOS 4.53 — virtually indistinguishable from professional recordings at MOS 4.58.
Putting it all together
The full Tacotron 2 pipeline flows from left to right: characters → encoder → attention → decoder → post-net → mel spectrogram → WaveNet → audio waveform. The mel spectrogram is the clean interface: the spectrogram network can be improved independently (better attention, better encoder) and the vocoder can be swapped (WaveNet → WaveRNN → HiFi-GAN) without retraining the other half.
This modular design is why the mel-spectrogram interface became the industry standard. Every subsequent TTS system — FastSpeech, VITS, Grad-TTS — predicts mel spectrograms as an intermediate step. The idea of separating "what to say" from "how to render" turned out to be one of the most influential design decisions in speech synthesis.
Why it mattered
2016
WaveNet
Proved that a neural vocoder could generate natural-sounding speech sample by sample, but required hand-crafted linguistic features as input.
2017
Tacotron 1
First end-to-end model mapping characters to spectrograms, but used Griffin-Lim for waveform synthesis, causing audible artifacts.
2018
Tacotron 2
Combined neural spectrogram prediction with neural vocoder, achieving MOS 4.53 — virtually indistinguishable from human speech.
2019
FastSpeech
Replaced the autoregressive decoder with parallel prediction, achieving real-time inference while keeping mel spectrograms as the interface.
2020
HiFi-GAN
A GAN-based vocoder that converts mel spectrograms to audio in real time with quality rivaling WaveNet — showing the mel interface enables vocoder plug-and-play.
2022
Whisper
OpenAI's speech recognition model that processes mel spectrograms as input — the same representation Tacotron 2 established, now used in the reverse direction.
2023
VALL-E & XTTS
Zero-shot voice cloning and multilingual TTS systems that build on the mel spectrogram paradigm Tacotron 2 established.
CitationShen, Pang, Weiss, Schuster, Jaitly, Yang, Chen, Zhang, Wang, Skerry-Ryan, Saurous, Agiomyrgiannakis, Wu. Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions. ICASSP, 2018.
Terms in this paper
- Mel Spectrogramالمخطط الطيفي ميل
- Text-to-Speechتحويل النص إلى كلام
- Vocoderالمشفِّر الصوتي
- Encoder-Decoderمرمِّز-فاكّ ترميز
- Location-Sensitive Attentionالانتباه الحسّاس للموقع
- Autoregressive Modelالنموذج التوليدي التراجعي
- Sequence-to-Sequenceتسلسل إلى تسلسل
- Embeddingالتضمين
- Feed Forward Network (FFN)شبكة التغذية الأمامية
- Batch Normalizationتسوية الدفعات الحسابية
- LSTMشبكة الذاكرة الطويلة قصيرة المدى
- Residual Connectionالوصلة التجاوزية
- Teacher Forcingالتوجيه بالمرجع
- Mean Opinion Scoreمتوسط درجة الرأي