Generative Models2023intermediate13 min read
MusicLM: Generating Music From Text
MusicLM: توليد الموسيقى من النص
Agostinelli, A. · Denk, T. I. · Borsos, Z. · Engel, J. · Verzetti, M. · Caillon, A. · Huang, Q. · Jansen, A. · Roberts, A. · Tagliasacchi, M. · Sharifi, M. · Zeghidour, N. · Frank, C. — arXiv
The problem
By 2022, text-to-image models like DALL·E 2 could generate photorealistic images from descriptions. But music was far behind. Existing audio generators could handle simple sound effects — "whistling with wind blowing" — yet producing rich, multi-instrument music with long-term structure from a single text caption remained an open problem. Two bottlenecks stood in the way: first, generating high-fidelity audio with coherent structure over minutes (not just seconds); and second, the severe scarcity of paired music-text data compared to the billions of image-text pairs available in the vision domain.
The contribution
MusicLM combines three pre-trained models — SoundStream (neural ), w2v-BERT (self-supervised audio representations), and MuLan (joint music-text embeddings) — into a hierarchical pipeline. During training, it uses only audio: MuLan audio embeddings serve as tokens, semantic tokens from w2v-BERT capture long-term structure, and acoustic tokens from SoundStream encode fine audio detail. At , the text from MuLan replaces the audio embedding, enabling text-conditioned generation without any paired data. The model generates 24 kHz music consistent over several minutes, outperforming baselines on MusicCaps, a new 5.5k expert-annotated evaluation dataset released with the paper.
The impact
MusicLM demonstrated that the hierarchical approach from AudioLM could be extended to text-conditioned generation, establishing a new paradigm for music AI. It directly inspired MusicGen (Meta, 2023), which simplified the architecture with a single-stage , and influenced the design of audio generation in multimodal models. MusicCaps became a standard benchmark, and the idea of bridging modalities via a shared (MuLan) without requiring paired training data influenced subsequent work across audio, video, and cross-modal generation including models like Sora.
Imagine an orchestra with no sheet music. Instead, a conductor receives a handwritten note saying: "Play a gentle piano waltz with soft strings in the background."
The conductor doesn't play any instrument. Instead, she works in three stages. First, she sketches a rough outline on a whiteboard — verse, chorus, bridge — capturing the skeleton of the piece. Then a section leader fills in the harmonic details — which chords, which voicings. Finally, each musician adds the fine acoustic texture — reverb, dynamics, bow pressure.
MusicLM works the same way. A text description is converted into a shared language the system understands, then three cascading stages — semantic, coarse acoustic, fine acoustic — build the music from skeleton to polished recording.
The building blocks: three pre-trained models
MusicLM does not train a single giant model from scratch. Instead, it assembles three independently pre-trained and frozen components, each specializing in a different aspect of audio understanding. Think of them as three experts who already know their craft — MusicLM just learns how to coordinate them.
SoundStream is a neural audio codec. It compresses raw audio waveforms into a compact sequence of discrete tokens — called acoustic tokens — using (RVQ). RVQ chains multiple codebooks together: the first captures the coarsest structure, each subsequent codebook refines the residual error, like adding layers of detail to a sketch. For 24 kHz audio, SoundStream produces 600 tokens per second across 12 quantization levels, at a bitrate of just 6 kbps. These tokens are sufficient to reconstruct high-fidelity audio.
w2v-BERT is a self-supervised speech model trained with on audio. From an intermediate layer (layer 7), MusicLM extracts representations that capture semantic content — the kind of information that tells you whether the audio is jazz or classical, whether the tempo is fast or slow, whether there is singing or only instruments. These representations are quantized into 25 semantic tokens per second using with 1024 centroids.
MuLan is a joint music-text embedding model, analogous to CLIP in the image domain. It has two towers — one processes audio, one processes text — both projecting into a shared 128-dimensional space trained with . If a piece of music and a text description refer to the same concept, their embeddings land close together. MuLan embeddings are quantized into just 12 MuLan tokens per audio sequence using RVQ. This is the bridge that connects text to music.
The pipeline: from text to music in three stages
MusicLM generates music through a cascade of three Transformer stages. Each stage takes as input the conditioning tokens plus the output of the previous stage, and predicts the next level of tokens.
Stage 1 — Semantic modeling. This stage takes the MuLan tokens (which encode "what kind of music") and generates the semantic tokens from w2v-BERT. It models the distribution . Think of this as sketching the composition: deciding the melodic contour, rhythmic feel, and temporal structure. This stage is trained on 30-second crops of audio.
Stage 2 — Coarse acoustic modeling. Given the MuLan tokens and the semantic tokens from Stage 1, this stage predicts the first 4 levels of the SoundStream RVQ. It models . These coarse acoustic tokens add harmonic content, rough timbre, and approximate dynamics — like filling in the colors of a painting.
Stage 3 — Fine acoustic modeling. The remaining 8 RVQ levels are generated conditioned on everything above. This adds the final layer of acoustic detail — precise timbre, reverb, stereo texture. The resolution goes from a sketch to a studio-quality recording.
Each stage uses a Transformer with 24 layers, 16 heads, an embedding dimension of 1024, and 430M parameters. They are trained independently on audio-only data.
The MuLan bridge: training without paired data
The central insight of MusicLM is how it sidesteps the paired-data bottleneck. In the image domain, DALL·E 2 has billions of image-text pairs. In the music domain, such data barely exists. MusicLM's solution: train on audio only, and swap the conditioning signal at inference time.
During training, the conditioning signal is the MuLan audio embedding — extracted from the target audio itself. The model learns to generate music given a description of what that music sounds like. No text captions are needed. This allows training on a massive audio-only dataset of 280,000 hours — five million clips.
During inference, the MuLan text embedding from the user's prompt replaces the audio embedding. This works because MuLan was trained with contrastive learning to place matching audio-text pairs close together in embedding space. From the Transformer's perspective, the conditioning tokens look the same whether they came from audio or text.
This design is analogous to DALL·E 2's use of CLIP, but with a crucial simplification: MusicLM skips the "prior" model that DALL·E 2 uses to map text embeddings to image embeddings. The shared embedding space is tight enough that direct substitution works.
Residual vector quantization: digital LEGO for sound
Residual (RVQ) is the backbone of MusicLM's token system. Standard vector quantization (VQ) maps each audio frame to the nearest entry in a single codebook — but a single codebook can only capture so much detail. Increasing codebook size causes an exponential explosion in memory and lookup cost.
RVQ solves this by chaining multiple smaller codebooks. The first codebook quantizes the original signal. The second codebook quantizes the residual — the error left over after the first quantization. The third codebook quantizes the residual of the residual. And so on. The reconstruction is the sum of all codebook outputs.
In SoundStream, 12 codebooks with 1024 entries each give the same representational capacity as a single codebook with entries — an astronomically large number — but with only entries in total. More importantly, the codebooks form a natural hierarchy: the first few levels capture the most important structure, and later levels add refinement. MusicLM exploits this by generating coarse levels (1–4) and fine levels (5–12) in separate stages.
Beyond text: melody conditioning
Some aspects of music are easier to hum than to describe in words. MusicLM addresses this with melody conditioning: the user provides a hummed, whistled, or sung melody alongside a text description. The text controls the style (genre, instrumentation, mood), while the melody controls the melodic contour.
To capture melody invariant to timbre, the authors train a small ViT-based embedding model using on pairs of audio clips with matching melodies but different acoustics (e.g., a song and its cover version). The resulting 192-dimensional melody embeddings are quantized with RVQ (24 quantizers, vocabulary 512) and concatenated with MuLan tokens as conditioning.
The result: you can whistle a tune, type "epic orchestral," and MusicLM will render your whistle as a full orchestral arrangement — same melody, completely different sound.
Results: how good is the music?
MusicLM was evaluated against Mubert (API-based, uses pre-recorded sounds by musicians) and Riffusion (fine-tuned Stable Diffusion on mel spectrograms) on the MusicCaps benchmark. Three metrics capture different quality dimensions.
Fréchet Audio Distance (FAD) measures audio quality without needing a reference. MusicLM achieved FAD = 4.0, compared to 9.6 for Mubert and 13.4 for Riffusion, meaning its generated audio sounds more like real music.
KL Divergence (KLD) measures whether the generated music has similar acoustic characteristics to the reference. MusicLM scored 1.01, vs 1.58 for Mubert and 1.19 for Riffusion — tighter adherence to the text description.
MuLan Cycle Consistency (MCC) measures between the text embedding and the embedding of the generated audio. MusicLM scored 0.51, far above Mubert's 0.32 and Riffusion's 0.34.
In human listening tests, MusicLM was preferred over both baselines in pairwise comparisons, winning 312 out of 600 comparisons, while there remained a gap with the ground truth reference music (472 wins).
Memorization analysis: does it copy the training data?
A critical concern with generative music models is whether they memorize and reproduce copyrighted music from the training set. The authors adapted a memorization methodology from large language models to study this.
They selected random training examples, fed the model a prompt (MuLan tokens + varying lengths of prefix), and compared generated continuations against the actual training data. The results: exact token matches remained below 0.2%, even with a 10-second prompt. Approximate matches (using an optimal transport measure between token histograms) reached about 1% — but closer inspection revealed these were low-diversity sequences (repeating patterns, drones) with an average entropy of just 1.0 bit versus 4.6 bits for normal sequences.
The second acoustic stage introduces further diversity even when semantic tokens match exactly, making literal reproduction of training audio extremely unlikely in practice.
Long generation and story mode
MusicLM generates longer sequences by advancing with a sliding window: it uses the last 15 seconds as a prefix to generate the next 15 seconds, always conditioning on the same text description. This approach produces coherent audio over several minutes — far beyond the 30-second training window.
A variant called story mode changes the text description every 15 seconds. The model produces smooth transitions that maintain tempo consistency while shifting musical context — from "gentle piano" to "energetic drums" — creating a narrative arc through sound.
Limitations and open challenges
Several limitations are inherited from MuLan. The model misunderstands negations — "music without drums" may still produce drums. It also does not adhere to precise temporal ordering described in text, such as "start with piano, then add guitar after 10 seconds."
The MCC metric has a self-referential bias because it uses MuLan itself, which favors MusicLM since MuLan is part of its pipeline. Vocals quality remains limited, and the model does not generate lyrics. High-level song structure (intro, verse, chorus) is also not explicitly modeled.
Timeline: from audio modeling to text-to-music
2020
Jukebox (OpenAI)
First model to generate music with singing voices using hierarchical VQ-VAEs. Showed long-term coherence was possible but quality had noticeable artifacts.
2021
w2v-BERT and MuLan foundations
w2v-BERT combined contrastive learning with masked language modeling for audio. MuLan created a shared embedding space between music audio and natural language text.
2022
AudioLM
Cast audio generation as language modeling with semantic + acoustic tokens. Achieved high fidelity and coherence but had no text conditioning.
2023
MusicLM (this paper)
Combined AudioLM's hierarchical generation with MuLan's text-audio bridge. First model to generate high-quality, text-conditioned music at 24 kHz over minutes. Released MusicCaps benchmark.
2023
MusicGen (Meta)
Simplified MusicLM's multi-stage pipeline into a single autoregressive Transformer with interleaved token patterns. Open-sourced and faster inference.
MusicLM's deepest contribution is not any single component but the architectural recipe: use a pre-trained model to sidestep paired data, decompose generation into hierarchical stages aligned with token granularity, and let each stage specialize. This recipe has since been adapted for video, 3D generation, and multimodal synthesis — confirming that the hierarchical token paradigm extends far beyond music.
CitationAgostinelli, Denk, Borsos, Engel, Verzetti, Caillon, Huang, Jansen, Roberts, Tagliasacchi, Sharifi, Zeghidour, Frank. MusicLM: Generating Music From Text. arXiv, 2023.
Terms in this paper
- Autoregressive Modelالنموذج التوليدي التراجعي
- Semantic Tokenرمز دلالي
- Acoustic Tokenرمز صوتي
- Residual Vector Quantizationالتكميم المتجهي المتبقي
- Contrastive Learningالتعلم التبايُني
- Embedding Spaceفضاء التضمين
- Codebookجدول الرموز
- Sequence-to-Sequenceتسلسل إلى تسلسل
- Hierarchical Samplingاختيار العينات الهرمي
- Conditioningالتوجيه
- Mel Spectrogramالمخطط الطيفي ميل
- Knowledge Distillationتقطير المعرفة
- Self-Supervised Learningالتعلم ذاتي الإشراف
- Encoder-Decoderمرمِّز-فاكّ ترميز
- Tokenizationتجزئة النصوص