Datasets › MagnaTagATune
MagnaTagATune
MagnaTagATune dataset contains 25,863 music clips. Each clip is a 29-seconds-long excerpt belonging to one of the 5223 songs, 445 albums and 230 artists. The clips span a broad range of genres like Classical, New Age, Electronica, Rock, Pop, World, Jazz, Blues, Metal, Punk, and more. Each audio clip is supplied with a vector of binary annotations of 188 tags. These annotations are obtained by humans playing the two-player online TagATune game. In this game, the two players are either presented with the same or a different audio clip. Subsequently, they are asked to come up with tags for their specific audio clip. Afterward, players view each other’s tags and are asked to decide whether they were presented the same audio clip. Tags are only assigned when more than two players agreed. The annotations include tags like ’singer’, ’no singer’, ’violin’, ’drums’, ’classical’, ’jazz’. The top 50 most popular tags are typically used for evaluation to ensure that there is enough training data for each tag. There are 16 parts, and researchers comonnly use parts 1-12 for training, part 13 for validation and parts 14-16 for testing.
Source: Brains on Beats Audio Source: http://mirg.city.ac.uk/codeapps/the-magnatagatune-dataset
Benchmarks archive 2025-07-28
All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Music Auto-Tagging | MagnaTagATune | M2D2 AS+ PR-AUC 41.6 | M2D2: Exploring General-purpose Audio-Language... | nttcslab/m2d +1 | 3 | Compare |
| Music Auto-Tagging | MagnaTagATune (clean) | Short-chunk CNN + Res PR-AUC 46.14 | Evaluation of CNN-based Automatic Music Tagging Models | minzwon/sota-music-tagging-models +6 | 3 | Compare |
Papers archive 2025-07-28
5 shown of 5 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 65. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| M2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP | 2 | 1 | 28 Mar 2025 | not harvested |
| Masked Latent Prediction and Classification for Self-Supervised Audio Representation Learning | 1 | 1 | 17 Feb 2025 | not harvested |
| Audio Embeddings as Teachers for Music Classification | 1 | 2 | 30 Jun 2023 | not harvested |
| Contrastive Learning of Musical Representations | 1 | 1 | 17 Mar 2021 | not harvested |
| Evaluation of CNN-based Automatic Music Tagging Models | 7 | 1 | ran 2 of 9 samples (7 unverified) |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
Unknown
Modalities archive 2025-07-28
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- MagnaTagATune
- MagnaTagATune (clean)
2 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections