Generative Models2023intermediate13 min read

MusicLM: Generating Music From Text

MusicLM: توليد الموسيقى من النص

Agostinelli, A. · Denk, T. I. · Borsos, Z. · Engel, J. · Verzetti, M. · Caillon, A. · Huang, Q. · Jansen, A. · Roberts, A. · Tagliasacchi, M. · Sharifi, M. · Zeghidour, N. · Frank, C. — arXiv

The problem

By 2022, text-to-image models like DALL·E 2 could generate photorealistic images from descriptions. But music was far behind. Existing audio generators could handle simple sound effects — "whistling with wind blowing" — yet producing rich, multi-instrument music with long-term structure from a single text caption remained an open problem. Two bottlenecks stood in the way: first, generating high-fidelity audio with coherent structure over minutes (not just seconds); and second, the severe scarcity of paired music-text data compared to the billions of image-text pairs available in the vision domain.

The contribution

MusicLM combines three pre-trained models — SoundStream (neural ), w2v-BERT (self-supervised audio representations), and MuLan (joint music-text embeddings) — into a hierarchical pipeline. During training, it uses only audio: MuLan audio embeddings serve as tokens, semantic tokens from w2v-BERT capture long-term structure, and acoustic tokens from SoundStream encode fine audio detail. At , the text from MuLan replaces the audio embedding, enabling text-conditioned generation without any paired data. The model generates 24 kHz music consistent over several minutes, outperforming baselines on MusicCaps, a new 5.5k expert-annotated evaluation dataset released with the paper.

The impact

MusicLM demonstrated that the hierarchical approach from AudioLM could be extended to text-conditioned generation, establishing a new paradigm for music AI. It directly inspired MusicGen (Meta, 2023), which simplified the architecture with a single-stage , and influenced the design of audio generation in multimodal models. MusicCaps became a standard benchmark, and the idea of bridging modalities via a shared (MuLan) without requiring paired training data influenced subsequent work across audio, video, and cross-modal generation including models like Sora.

Imagine an orchestra with no sheet music. Instead, a conductor receives a handwritten note saying: "Play a gentle piano waltz with soft strings in the background."

The conductor doesn't play any instrument. Instead, she works in three stages. First, she sketches a rough outline on a whiteboard — verse, chorus, bridge — capturing the skeleton of the piece. Then a section leader fills in the harmonic details — which chords, which voicings. Finally, each musician adds the fine acoustic texture — reverb, dynamics, bow pressure.

MusicLM works the same way. A text description is converted into a shared language the system understands, then three cascading stages — semantic, coarse acoustic, fine acoustic — build the music from skeleton to polished recording.

The building blocks: three pre-trained models

MusicLM does not train a single giant model from scratch. Instead, it assembles three independently pre-trained and frozen components, each specializing in a different aspect of audio understanding. Think of them as three experts who already know their craft — MusicLM just learns how to coordinate them.

SoundStream is a neural audio codec. It compresses raw audio waveforms into a compact sequence of discrete tokens — called acoustic tokens — using (RVQ). RVQ chains multiple codebooks together: the first captures the coarsest structure, each subsequent codebook refines the residual error, like adding layers of detail to a sketch. For 24 kHz audio, SoundStream produces 600 tokens per second across 12 quantization levels, at a bitrate of just 6 kbps. These tokens are sufficient to reconstruct high-fidelity audio.

w2v-BERT is a self-supervised speech model trained with on audio. From an intermediate layer (layer 7), MusicLM extracts representations that capture semantic content — the kind of information that tells you whether the audio is jazz or classical, whether the tempo is fast or slow, whether there is singing or only instruments. These representations are quantized into 25 semantic tokens per second using with 1024 centroids.

MuLan is a joint music-text embedding model, analogous to CLIP in the image domain. It has two towers — one processes audio, one processes text — both projecting into a shared 128-dimensional space trained with . If a piece of music and a text description refer to the same concept, their embeddings land close together. MuLan embeddings are quantized into just 12 MuLan tokens per audio sequence using RVQ. This is the bridge that connects text to music.

Open in Lab
Explore the three token types: MuLan tokens capture "what kind of music," semantic tokens capture the musical structure, and acoustic tokens encode the fine sound detail.
The demo wakes as you arrive…

The pipeline: from text to music in three stages

MusicLM generates music through a cascade of three Transformer stages. Each stage takes as input the conditioning tokens plus the output of the previous stage, and predicts the next level of tokens.

Stage 1 — Semantic modeling. This stage takes the MuLan tokens (which encode "what kind of music") and generates the semantic tokens from w2v-BERT. It models the distribution p(St∣S<t,MA)p(S_t | S_{<t}, M_A). Think of this as sketching the composition: deciding the melodic contour, rhythmic feel, and temporal structure. This stage is trained on 30-second crops of audio.

Stage 2 — Coarse acoustic modeling. Given the MuLan tokens and the semantic tokens from Stage 1, this stage predicts the first 4 levels of the SoundStream RVQ. It models p(At∣A<t,S,MA)p(A_t | A_{<t}, S, M_A). These coarse acoustic tokens add harmonic content, rough timbre, and approximate dynamics — like filling in the colors of a painting.

Stage 3 — Fine acoustic modeling. The remaining 8 RVQ levels are generated conditioned on everything above. This adds the final layer of acoustic detail — precise timbre, reverb, stereo texture. The resolution goes from a sketch to a studio-quality recording.

Each stage uses a Transformer with 24 layers, 16 heads, an embedding dimension of 1024, and 430M parameters. They are trained independently on audio-only data.

Open in Lab
Follow the text prompt through all three stages — from MuLan tokens to semantic tokens to acoustic tokens — and see how each stage adds fidelity.
The demo wakes as you arrive…

The MuLan bridge: training without paired data

The central insight of MusicLM is how it sidesteps the paired-data bottleneck. In the image domain, DALL·E 2 has billions of image-text pairs. In the music domain, such data barely exists. MusicLM's solution: train on audio only, and swap the conditioning signal at inference time.

During training, the conditioning signal is the MuLan audio embedding — extracted from the target audio itself. The model learns to generate music given a description of what that music sounds like. No text captions are needed. This allows training on a massive audio-only dataset of 280,000 hours — five million clips.

During inference, the MuLan text embedding from the user's prompt replaces the audio embedding. This works because MuLan was trained with contrastive learning to place matching audio-text pairs close together in embedding space. From the Transformer's perspective, the conditioning tokens look the same whether they came from audio or text.

This design is analogous to DALL·E 2's use of CLIP, but with a crucial simplification: MusicLM skips the "prior" model that DALL·E 2 uses to map text embeddings to image embeddings. The shared embedding space is tight enough that direct substitution works.

Open in Lab
See how audio and text embeddings converge in MuLan's shared space — and how MusicLM swaps them at inference time.
The demo wakes as you arrive…
Training: p(St∣S<t,MA)⟶Inference: p(St∣S<t,MT)\text{Training: } p(S_t | S_{<t}, M_A) \quad\longrightarrow\quad \text{Inference: } p(S_t | S_{<t}, M_T)
The MuLan swap — audio conditioning in training, text conditioning at inference — During training, the semantic modeling stage conditions on MuLan audio tokens MAM_A extracted from the target audio. At inference, these are replaced by MuLan text tokens MTM_T computed from the user's text prompt. Because both live in the same embedding space, the Transformer generalizes from one to the other.

Residual vector quantization: digital LEGO for sound

Residual (RVQ) is the backbone of MusicLM's token system. Standard vector quantization (VQ) maps each audio frame to the nearest entry in a single codebook — but a single codebook can only capture so much detail. Increasing codebook size causes an exponential explosion in memory and lookup cost.

RVQ solves this by chaining multiple smaller codebooks. The first codebook quantizes the original signal. The second codebook quantizes the residual — the error left over after the first quantization. The third codebook quantizes the residual of the residual. And so on. The reconstruction is the sum of all codebook outputs.

In SoundStream, 12 codebooks with 1024 entries each give the same representational capacity as a single codebook with 1024121024^{12} entries — an astronomically large number — but with only 12×102412 \times 1024 entries in total. More importantly, the codebooks form a natural hierarchy: the first few levels capture the most important structure, and later levels add refinement. MusicLM exploits this by generating coarse levels (1–4) and fine levels (5–12) in separate stages.

Open in Lab
Add RVQ levels one by one and hear how each layer refines the audio — from muddy sketch to crisp recording.
The demo wakes as you arrive…
x^=∑q=1Qeq,eq=arg⁡min⁡c∈Cq∥rq−1−c∥,rq=rq−1−eq\hat{x} = \sum_{q=1}^{Q} e_q, \quad e_q = \arg\min_{c \in \mathcal{C}_q} \left\| r_{q-1} - c \right\|, \quad r_q = r_{q-1} - e_q
Residual vector quantization — each codebook refines the previous residual — The input signal is reconstructed as the sum of selected codewords eqe_q from QQ codebooks. Each codebook Cq\mathcal{C}_q quantizes the residual rq−1r_{q-1} from the previous level, starting with r0=xr_0 = x (the original signal). This progressively captures finer detail without exponential codebook growth.

Beyond text: melody conditioning

Some aspects of music are easier to hum than to describe in words. MusicLM addresses this with melody conditioning: the user provides a hummed, whistled, or sung melody alongside a text description. The text controls the style (genre, instrumentation, mood), while the melody controls the melodic contour.

To capture melody invariant to timbre, the authors train a small ViT-based embedding model using on pairs of audio clips with matching melodies but different acoustics (e.g., a song and its cover version). The resulting 192-dimensional melody embeddings are quantized with RVQ (24 quantizers, vocabulary 512) and concatenated with MuLan tokens as conditioning.

The result: you can whistle a tune, type "epic orchestral," and MusicLM will render your whistle as a full orchestral arrangement — same melody, completely different sound.

Results: how good is the music?

MusicLM was evaluated against Mubert (API-based, uses pre-recorded sounds by musicians) and Riffusion (fine-tuned Stable Diffusion on mel spectrograms) on the MusicCaps benchmark. Three metrics capture different quality dimensions.

Fréchet Audio Distance (FAD) measures audio quality without needing a reference. MusicLM achieved FADVGG_\text{VGG} = 4.0, compared to 9.6 for Mubert and 13.4 for Riffusion, meaning its generated audio sounds more like real music.

KL Divergence (KLD) measures whether the generated music has similar acoustic characteristics to the reference. MusicLM scored 1.01, vs 1.58 for Mubert and 1.19 for Riffusion — tighter adherence to the text description.

MuLan Cycle Consistency (MCC) measures between the text embedding and the embedding of the generated audio. MusicLM scored 0.51, far above Mubert's 0.32 and Riffusion's 0.34.

In human listening tests, MusicLM was preferred over both baselines in pairwise comparisons, winning 312 out of 600 comparisons, while there remained a gap with the ground truth reference music (472 wins).

Open in Lab
Compare MusicLM against baselines across FAD, KLD, MCC, and human preferences.
The demo wakes as you arrive…

Memorization analysis: does it copy the training data?

A critical concern with generative music models is whether they memorize and reproduce copyrighted music from the training set. The authors adapted a memorization methodology from large language models to study this.

They selected random training examples, fed the model a prompt (MuLan tokens + varying lengths of prefix), and compared generated continuations against the actual training data. The results: exact token matches remained below 0.2%, even with a 10-second prompt. Approximate matches (using an optimal transport measure between token histograms) reached about 1% — but closer inspection revealed these were low-diversity sequences (repeating patterns, drones) with an average entropy of just 1.0 bit versus 4.6 bits for normal sequences.

The second acoustic stage introduces further diversity even when semantic tokens match exactly, making literal reproduction of training audio extremely unlikely in practice.

Long generation and story mode

MusicLM generates longer sequences by advancing with a sliding window: it uses the last 15 seconds as a prefix to generate the next 15 seconds, always conditioning on the same text description. This approach produces coherent audio over several minutes — far beyond the 30-second training window.

A variant called story mode changes the text description every 15 seconds. The model produces smooth transitions that maintain tempo consistency while shifting musical context — from "gentle piano" to "energetic drums" — creating a narrative arc through sound.

Limitations and open challenges

Several limitations are inherited from MuLan. The model misunderstands negations — "music without drums" may still produce drums. It also does not adhere to precise temporal ordering described in text, such as "start with piano, then add guitar after 10 seconds."

The MCC metric has a self-referential bias because it uses MuLan itself, which favors MusicLM since MuLan is part of its pipeline. Vocals quality remains limited, and the model does not generate lyrics. High-level song structure (intro, verse, chorus) is also not explicitly modeled.

Timeline: from audio modeling to text-to-music

  1. 2020

    Jukebox (OpenAI)

    First model to generate music with singing voices using hierarchical VQ-VAEs. Showed long-term coherence was possible but quality had noticeable artifacts.

  2. 2021

    w2v-BERT and MuLan foundations

    w2v-BERT combined contrastive learning with masked language modeling for audio. MuLan created a shared embedding space between music audio and natural language text.

  3. 2022

    AudioLM

    Cast audio generation as language modeling with semantic + acoustic tokens. Achieved high fidelity and coherence but had no text conditioning.

  4. 2023

    MusicLM (this paper)

    Combined AudioLM's hierarchical generation with MuLan's text-audio bridge. First model to generate high-quality, text-conditioned music at 24 kHz over minutes. Released MusicCaps benchmark.

  5. 2023

    MusicGen (Meta)

    Simplified MusicLM's multi-stage pipeline into a single autoregressive Transformer with interleaved token patterns. Open-sourced and faster inference.

MusicLM's deepest contribution is not any single component but the architectural recipe: use a pre-trained model to sidestep paired data, decompose generation into hierarchical stages aligned with token granularity, and let each stage specialize. This recipe has since been adapted for video, 3D generation, and multimodal synthesis — confirming that the hierarchical token paradigm extends far beyond music.

CitationAgostinelli, Denk, Borsos, Engel, Verzetti, Caillon, Huang, Jansen, Roberts, Tagliasacchi, Sharifi, Zeghidour, Frank. MusicLM: Generating Music From Text. arXiv, 2023.

Terms in this paper