Datasets › MusicCaps
MusicCaps
MusicCaps is a dataset composed of 5.5k music-text pairs, with rich text descriptions provided by human experts. For each 10-second music clip, MusicCaps provides:
1) A free-text caption consisting of four sentences on average, describing the music and
2) A list of music aspects, describing genre, mood, tempo, singer voices, instrumentation, dissonances, rhythm, etc.
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Text-to-Music Generation | MusicCaps | MeLFusion (image-conditioned) FAD 1.12 | MeLFusion: Synthesizing Music from Image and Language... | schowdhury671/melfusion | 21 | Compare |
Papers archive 2025-07-28
14 shown of 14 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 84. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| ETTA: Elucidating the Design Space of Text-to-Audio Models | 1 | 1 | 26 Dec 2024 | not harvested |
| FLUX that Plays Music | 2 | 1 | 1 Sep 2024 | ran 6 of 9 samples (3 unverified; 9 pointer-only for licence) |
| Stable Audio Open | 1 | 1 | 19 Jul 2024 | ran 1 of 2 samples (1 unverified) |
| Improving Text-To-Audio Models with Synthetic Captions | 1 | 1 | 18 Jun 2024 | ran 3 of 4 samples (1 unverified; 4 pointer-only for licence) |
| MeLFusion: Synthesizing Music from Image and Language Cues using Diffusion Models | 1 | 1 | 7 Jun 2024 | ran 12 of 12 samples (0 unverified; 12 pointer-only for licence) |
| Quality-aware Masked Diffusion Transformer for Enhanced Music Generation | 2 | 1 | 24 May 2024 | ran 4 of 7 samples (3 unverified) |
| Fast Timing-Conditioned Latent Audio Diffusion | 2 | 1 | 7 Feb 2024 | ran 2 of 4 samples (2 unverified) |
| AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining | 2 | 3 | 10 Aug 2023 | ran 12 of 27 samples (15 unverified; 13 pointer-only for licence) |
| JEN-1: Text-Guided Universal Music Generation with Omnidirectional Diffusion Models | 2 | 1 | 9 Aug 2023 | not harvested |
| Simple and Controllable Music Generation | 8 | 3 | 8 Jun 2023 | ran 11 of 18 samples (7 unverified; 7 pointer-only for licence) |
| Efficient Neural Music Generation | 0 | 1 | 25 May 2023 | not harvested |
| Noise2Music: Text-conditioned Music Generation with Diffusion Models | 0 | 2 | 8 Feb 2023 | not harvested |
| MusicLM: Generating Music From Text | 5 | 3 | 26 Jan 2023 | ran 6 of 16 samples (10 unverified) |
| UniAudio: An Audio Foundation Model Toward Universal Audio Generation | 2 | 1 | ran 4 of 4 samples (0 unverified; 4 pointer-only for licence) |
Dataset loaders archive 2025-07-28
1 loader as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- MusicCaps
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections