Datasets › MusicCaps

MusicCaps

Introduced by Andrea Agostinelli et al. in MusicLM: Generating Music From Text26 Jan 2023 archive 2025-07-28

MusicCaps is a dataset composed of 5.5k music-text pairs, with rich text descriptions provided by human experts. For each 10-second music clip, MusicCaps provides:

1) A free-text caption consisting of four sentences on average, describing the music and

2) A list of music aspects, describing genre, mood, tempo, singer voices, instrumentation, dissonances, rhythm, etc.

Source:MusicLM: Generating Music From Text

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Text-to-Music Generation MusicCaps MeLFusion (image-conditioned) FAD 1.12 MeLFusion: Synthesizing Music from Image and Language... schowdhury671/melfusion 21 Compare

Papers archive 2025-07-28

14 shown of 14 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 84. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
ETTA: Elucidating the Design Space of Text-to-Audio Models 1 1 26 Dec 2024 not harvested
FLUX that Plays Music 2 1 1 Sep 2024 ran 6 of 9 samples (3 unverified; 9 pointer-only for licence)
Stable Audio Open 1 1 19 Jul 2024 ran 1 of 2 samples (1 unverified)
Improving Text-To-Audio Models with Synthetic Captions 1 1 18 Jun 2024 ran 3 of 4 samples (1 unverified; 4 pointer-only for licence)
MeLFusion: Synthesizing Music from Image and Language Cues using Diffusion Models 1 1 7 Jun 2024 ran 12 of 12 samples (0 unverified; 12 pointer-only for licence)
Quality-aware Masked Diffusion Transformer for Enhanced Music Generation 2 1 24 May 2024 ran 4 of 7 samples (3 unverified)
Fast Timing-Conditioned Latent Audio Diffusion 2 1 7 Feb 2024 ran 2 of 4 samples (2 unverified)
AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining 2 3 10 Aug 2023 ran 12 of 27 samples (15 unverified; 13 pointer-only for licence)
JEN-1: Text-Guided Universal Music Generation with Omnidirectional Diffusion Models 2 1 9 Aug 2023 not harvested
Simple and Controllable Music Generation 8 3 8 Jun 2023 ran 11 of 18 samples (7 unverified; 7 pointer-only for licence)
Efficient Neural Music Generation 0 1 25 May 2023 not harvested
Noise2Music: Text-conditioned Music Generation with Diffusion Models 0 2 8 Feb 2023 not harvested
MusicLM: Generating Music From Text 5 3 26 Jan 2023 ran 6 of 16 samples (10 unverified)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation 2 1 ran 4 of 4 samples (0 unverified; 4 pointer-only for licence)

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY-SA 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • MusicCaps

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections