Generative Models2020advanced13 min read
Jukebox: A Generative Model for Music
Jukebox: نموذج توليدي للموسيقى
Dhariwal, P. · Jun, H. · Payne, C. · Kim, J. W. · Radford, A. · Sutskever, I. — arXiv
The problem
Music generation from raw audio is extraordinarily challenging because of scale. A single minute of CD-quality audio contains over 2.6 million samples at 44.1 kHz. Existing symbolic approaches like MIDI capture notes but lose timbre, dynamics, and vocal expression. Direct waveform modeling with networks is computationally intractable at this scale, and previous models could only produce short, low-fidelity clips. No system could generate multi-minute songs with coherent structure, realistic vocals, and stylistic diversity directly in the audio domain.
The contribution
Jukebox introduces a hierarchical VQ-VAE that compresses 44.1 kHz raw audio by factors of 8×, 32×, and 128× into three levels of discrete codes using codebooks of 2048 entries each. Autoregressive Sparse Transformers then model the distribution of these codes at each level: a top-level prior captures long-range musical structure (melody, harmony, rhythm) conditioned on artist, genre, and optionally unaligned lyrics, while two priors progressively recover fine-grained audio detail. The system was trained on 1.2 million songs and can generate diverse, multi-minute songs with coherent structure and rudimentary singing.
The impact
Jukebox demonstrated that raw audio generation at the scale of full songs is feasible, bridging the gap between symbolic music models and waveform synthesis. Its hierarchical VQ-VAE tokenization strategy directly influenced subsequent audio models including MusicLM and AudioLM. The paper showed that -based priors could model long-range musical coherence, paving the way for the current generation of music AI systems. It also revealed that VQ-VAE representations encode rich musical semantics usable for music information retrieval tasks.
Imagine you want to write a detailed biography of someone from a single passport photo. It's impossibly hard — there's not enough information. But what if you had three versions of their life story: a one-sentence summary, a one-page outline, and a full chapter?
You'd write the summary first (capturing the big arc), then expand it into the outline (filling in key events), then flesh out the full chapter (adding dialogue and detail). Each stage adds richness while staying faithful to the level above.
Jukebox does exactly this with music. It compresses a song into three "zoom levels" — a coarse skeleton, a mid-level sketch, and a fine-grained waveform — then learns to generate each level from the one above it. The result: coherent, multi-minute songs generated directly from raw audio.
The scale problem: why raw audio is so hard
To understand why Jukebox is a breakthrough, we need to grasp the scale problem. CD-quality audio is sampled at 44,100 times per second. A 4-minute song contains roughly 10.5 million samples. An that generates one sample at a time — like WaveNet — would need to predict each of those samples sequentially, attending to all previous ones.
Compare this to text: a 4-minute song is equivalent to a "document" of 10.5 million tokens. Even the largest language models process only a few thousand tokens at once. The mismatch is enormous.
Symbolic approaches like MIDI sidestep this by representing music as discrete note events — pitch, duration, velocity. But they discard timbre, vocal expression, recording ambience, and everything that makes a song sound like a particular performance rather than a mechanical playback. The challenge is to work directly in the raw audio domain while making the problem computationally tractable.
The compressor: hierarchical VQ-VAE
Jukebox's first insight is to compress raw audio into a much shorter sequence of discrete codes before modeling it. The tool for this is a Vector Quantized Variational (VQ-VAE).
A VQ-VAE works like a postal system with a fixed set of pre-printed postcards. The reads a chunk of audio and finds the closest matching "postcard" ( vector) from a library of 2048 options. The then reconstructs the audio from just that postcard index. teaches both the encoder and the codebook to make the postcards as expressive as possible.
But a single level of compression isn't enough. Jukebox uses three separate VQ-VAEs at different temporal resolutions. Think of them as three zoom levels on a map — city, street, and building. The bottom level (8× compression) retains fine acoustic detail like consonant textures and breaths. The middle level (32× compression) captures rhythmic patterns and timbral contours. The top level (128× compression) preserves only the broadest musical features — melody, harmony, and structure.
A critical design choice: the three VQ-VAEs are trained independently, not hierarchically. In a traditional hierarchical VQ-VAE (like VQ-VAE-2), upper levels encode what lower levels miss. Here, each level encodes the full audio independently at its own resolution. This means each level's codes are self-contained and can reconstruct audio on their own, though at different fidelity levels.
Why separate encoders? The authors found that a shared hierarchical encoder causes the upper levels to become unused — the bottom level already captures enough detail, so the network learns to ignore the top codes. With separate encoders, each level is forced to independently learn the most useful representation at its resolution.
Training the VQ-VAE: spectral loss and codebook collapse
Each VQ-VAE encoder uses stacked one-dimensional dilated convolutions — the same architecture pioneered by WaveNet — to progressively downsample the audio. The decoder mirrors this with transposed strided convolutions to reconstruct the waveform.
The training loss has three components. First, a reconstruction loss that penalizes the difference between the original and reconstructed audio. Second, a codebook loss that pulls codebook vectors toward the encoder outputs they represent. Third, a that prevents the encoder from jumping erratically between codebook entries.
Jukebox adds a crucial fourth component: . The standard reconstruction loss operates in the time domain (sample-by-sample), but it struggles to preserve high-frequency content. The spectral loss compares the mel spectrograms of the original and reconstructed audio, explicitly penalizing missing harmonics and high-frequency texture. Without it, the middle and top levels lose most frequencies above a few kHz — the music sounds muffled.
A notorious failure mode of VQ-VAEs is : the model learns to use only a handful of codebook entries, wasting the rest. Imagine a library with 2048 postcards, but the postal system only ever picks 50 of them — the rest gather dust.
Jukebox fights this with random restarts: whenever a codebook entry's usage falls below a threshold, it is teleported to a randomly chosen encoder output. This ensures all entries stay "alive" and track the evolving data distribution. The effect is dramatic — codebook entropy stays near the theoretical maximum throughout training.
The composer: autoregressive Transformer priors
Compression alone doesn't generate music — it just makes the sequence short enough to model. The actual generation happens in the prior models: autoregressive Transformers that learn the distribution of VQ-VAE codes and sample new ones.
The system uses three priors matching the three VQ-VAE levels. Generation flows top-down: first, the top-level prior generates the most compressed codes (~345 codes for 24 seconds of audio). These codes encode the broad musical arc — chord progression, melody contour, rhythmic feel. Then the middle upsampler takes those top codes and generates denser middle-level codes, adding rhythmic precision and timbral detail. Finally, the bottom upsampler produces the densest codes, recovering fine acoustic texture.
Each prior is a with mechanisms designed for long sequences. The top-level prior for the 5-billion parameter model uses 79 layers with a mixture of dense, factorized, and stride-based attention patterns to handle the full .
Steering the music: artist, genre, and lyric conditioning
One of Jukebox's most striking features is : the ability to steer generation toward a specific artist style, genre, and even lyrics.
Artist and genre conditioning is straightforward: each artist and genre is mapped to a learned vector, which is concatenated with the positional information and fed into each Transformer layer. This teaches the model to associate certain musical patterns — vocal timbre, instrumentation, tempo preferences — with each artist/genre combination.
Lyric conditioning is more subtle. The lyrics are unaligned — the model receives the full text of the lyrics but no information about which word corresponds to which musical moment. A separate Transformer encoder processes the lyrics, and encoder-decoder attention layers in the music prior learn to attend to relevant parts of the lyrics as it generates each musical . Over training, the model learns a rough alignment between text and audio, enabling rudimentary but recognizable singing.
The lyric encoder uses a key innovation: it processes lyrics at the character level rather than the word level. This lets the model learn phonetic patterns — how letters map to sung sounds — which is essential for generating intelligible singing. The attention pattern between lyrics and music shows a roughly diagonal structure: the model progresses through the lyrics roughly in order, though with repetitions and pauses that match musical phrasing.
Full pipeline: from text to song
Let us trace the full generation pipeline. The user provides a genre (e.g., "rock"), an artist (e.g., "Elvis Presley"), and optionally lyrics. The system then:
Step 1 — Top prior samples: The top-level Sparse Transformer generates ~345 codes for a 24-second chunk, conditioned on artist, genre, and lyrics. Each code corresponds to ~70ms of audio and encodes broad musical features.
Step 2 — Middle upsampler: Conditioned on the top codes, the middle prior generates ~1,380 codes that add rhythmic and timbral detail at 32× compression.
Step 3 — Bottom upsampler: Conditioned on the middle codes, the bottom prior generates ~5,500 codes at 8× compression, recovering consonant textures and fine dynamics.
Step 4 — VQ-VAE decoder: The bottom-level codes are looked up in the codebook and passed through the VQ-VAE decoder to reconstruct the final waveform.
For songs longer than 24 seconds, Jukebox uses : it generates overlapping chunks and cross-fades them, so the model can attend to recent context even when the total song length exceeds the Transformer's context window.
Results and limitations
Jukebox produces multi-minute songs that are remarkably coherent at a local level: melodic phrases sound natural, harmonies resolve logically, and vocal timbre is consistent. The model generates across a wide range of genres — pop, rock, hip-hop, jazz, classical — with recognizable stylistic features. Human evaluators rated Jukebox samples significantly higher than previous raw audio models.
However, significant limitations remain. Long-range structure beyond about 30 seconds remains challenging: songs often drift in key or tempo rather than maintaining a coherent verse-chorus form. Lyrics are only roughly intelligible — listeners can sometimes catch phrases, but the singing is not consistently word-perfect. Generation is extremely slow: producing one minute of audio at the top level takes about 3 hours on a V100 GPU, with upsampling adding more time.
The 5-billion parameter model was trained on 1.2 million songs (roughly 4,000 hours of music per genre) using 256 V100 GPUs. The scale of compute required limits the accessibility of this approach.
Code: sampling from Jukebox
Simplified to show the idea — not the real implementation.
# Set up conditioning: artist, genre, and lyrics
# These embeddings are fed into the Transformer priors
artist = "Elvis Presley"
genre = "Rock"
lyrics = """ Well, since my baby left me I found a new place to dwell """
# Step 1: Top-level prior generates ~345 codes (128x compression) # Each code covers ~70ms of audio — captures melody and harmony top_codes = top_prior.sample(
n_samples=1,
artist=artist,
genre=genre,
lyrics=lyrics,
sample_length_in_seconds=24
)
# Step 2: Middle upsampler adds rhythmic detail (32x compression) # Conditioned on top codes, generates ~1,380 codes mid_codes = middle_upsampler.sample(
top_codes=top_codes
)
# Step 3: Bottom upsampler adds fine texture (8x compression) # Conditioned on middle codes, generates ~5,500 codes bot_codes = bottom_upsampler.sample(
middle_codes=mid_codes
)
# Step 4: VQ-VAE decoder reconstructs the waveform # Looks up codebook vectors and runs through decoder network audio = vqvae.decode(bot_codes, level=0)Timeline: from WaveNet to music AI
2016
WaveNet — autoregressive audio generation
DeepMind introduced WaveNet, proving that autoregressive neural networks can generate raw audio waveforms sample-by-sample with unprecedented quality for speech synthesis.
2017
VQ-VAE — discrete latent representations
Van den Oord et al. introduced VQ-VAE, showing that continuous latent spaces can be replaced with discrete codebooks without losing reconstruction quality, enabling autoregressive modeling of latent codes.
2019
MuseNet — symbolic music with Transformers
OpenAI's MuseNet generated multi-instrument MIDI compositions using a Transformer, demonstrating long-range musical coherence but limited to symbolic (note-level) representation.
2020
Jukebox — raw audio music generation at scale
Combined hierarchical VQ-VAE compression with Sparse Transformer priors to generate multi-minute songs with singing directly in the raw audio domain, conditioned on artist, genre, and lyrics.
2023
MusicLM — text-to-music generation
Google's <NodeLink slug="musiclm">MusicLM</NodeLink> built on Jukebox's hierarchical tokenization strategy, using AudioLM and MuLan to generate music from free-form text descriptions with improved coherence and quality.
Jukebox stands at a pivotal point in the timeline of music AI. It inherited the raw waveform modeling capability of WaveNet and the discrete tokenization idea of VQ-VAE, then demonstrated that these ideas could scale to full songs. Its hierarchical approach — compress, model at multiple scales, upsample — became the template for subsequent systems including AudioLM and MusicLM.
The paper's deeper contribution may be showing that VQ-VAE codes are not just compression artifacts — they encode semantically meaningful musical information that can be extracted and used for music understanding tasks, not just generation.
CitationDhariwal, Jun, Payne, Kim, Radford, Sutskever. Jukebox: A Generative Model for Music. arXiv, 2020.
Terms in this paper
- Codebookجدول الرموز
- Autoregressive Modelالنموذج التوليدي التراجعي
- Hierarchical Samplingاختيار العينات الهرمي
- Mel Spectrogramالمخطط الطيفي ميل
- Commitment Lossخسارة الالتزام
- Latent Spaceالفضاء الكامن
- Upsamplingرفع الدقة
- Conditioningالتوجيه
- Priorالاحتمال القبلي المبدئي
- Embeddingالتضمين
- Transformerالمحوِّل
- Encoderالمُرمِّز
- Decoderمفكّ الترميز
- Reconstruction Errorخطأ إعادة البناء
- Bottleneckعنق الزجاجة