{"url":"/dataset/magnatagatune","name":"MagnaTagATune","full_name":null,"description_markdown":"**MagnaTagATune** dataset contains 25,863 music clips. Each clip is a 29-seconds-long excerpt belonging to one of the 5223 songs, 445 albums and 230 artists. The clips span a broad range of genres like Classical, New Age, Electronica, Rock, Pop, World, Jazz, Blues, Metal, Punk, and more. Each audio clip is supplied with a vector of binary annotations of 188 tags. These annotations are obtained by humans playing the two-player online TagATune game. In this game, the two players are either presented with the same or a different audio clip. Subsequently, they are asked to come up with tags for their specific audio clip. Afterward, players view each other’s tags and are asked to decide whether they were presented the same audio clip. Tags are only assigned when more than two players agreed. The annotations include tags like ’singer’, ’no singer’, ’violin’, ’drums’, ’classical’, ’jazz’. The top 50 most popular tags are typically used for evaluation to ensure that there is enough training data for each tag. There are 16 parts, and researchers comonnly use parts 1-12 for training, part 13 for validation and parts 14-16 for testing.\r\n\r\nSource: [Brains on Beats](https://arxiv.org/abs/1606.02627)\r\nAudio Source: [http://mirg.city.ac.uk/codeapps/the-magnatagatune-dataset](http://mirg.city.ac.uk/codeapps/the-magnatagatune-dataset)","description_withheld":null,"homepage":"http://mirg.city.ac.uk/codeapps/the-magnatagatune-dataset","introduced_date":"2009-01-01","introduced_date_note":null,"introduced_by":{"paper":null,"title":"Evaluation of Algorithms Using Games: The Case of Music Tagging","first_author":null,"url":"http://ismir2009.ismir.net/proceedings/OS5-5.pdf"},"license":{"name":"Unknown","url":null},"modalities":[{"name":"Audio","url":"/datasets/modality/audio"}],"tasks":[{"name":"Music Auto-Tagging","url":"/task/music-auto-tagging","datasets_with_task":"/datasets/task/music-auto-tagging"},{"name":"Music Classification","url":"/task/music-classification","datasets_with_task":"/datasets/task/music-classification"},{"name":"Music Tagging","url":"/task/music-tagging","datasets_with_task":"/datasets/task/music-tagging"}],"languages":[],"variants":["MagnaTagATune","MagnaTagATune (clean)"],"data_loaders":[],"num_papers_in_archive":65,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/music-auto-tagging-on-magnatagatune","task":"Music Auto-Tagging","dataset_variant":"MagnaTagATune","rows":3,"metrics":["PR-AUC","ROC AUC"],"first_row_in_archive_order":{"model":"M2D2 AS+","paper":"/paper/m2d2-exploring-general-purpose-audio-language","metrics":{"PR-AUC":"41.6","ROC AUC":"91.8"},"code_links":[{"title":"nttcslab/m2d","url":"https://github.com/nttcslab/m2d"},{"title":"nttcslab/eval-audio-repr","url":"https://github.com/nttcslab/eval-audio-repr"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/music-auto-tagging-on-magnatagatune-clean","task":"Music Auto-Tagging","dataset_variant":"MagnaTagATune (clean)","rows":3,"metrics":["PR-AUC","ROC-AUC"],"first_row_in_archive_order":{"model":"Short-chunk CNN + Res","paper":"/paper/evaluation-of-cnn-based-automatic-music","metrics":{"PR-AUC":"46.14","ROC-AUC":"91.29"},"code_links":[{"title":"minzwon/sota-music-tagging-models","url":"https://github.com/minzwon/sota-music-tagging-models"},{"title":"minzwon/tag-based-music-retrieval","url":"https://github.com/minzwon/tag-based-music-retrieval"},{"title":"pxaris/ccml","url":"https://github.com/pxaris/ccml"},{"title":"cgaroufis/msspt","url":"https://github.com/cgaroufis/msspt"},{"title":"jaehwlee/tf2-music-tagging-models","url":"https://github.com/jaehwlee/tf2-music-tagging-models"},{"title":"Dohppak/Music_DeepEmbedding_Extractor","url":"https://github.com/Dohppak/Music_DeepEmbedding_Extractor"},{"title":"HephaestusProject/pytorch-FCN","url":"https://github.com/HephaestusProject/pytorch-FCN"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/m2d2-exploring-general-purpose-audio-language","title":"M2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP","date":"2025-03-28","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/masked-latent-prediction-and-classification","title":"Masked Latent Prediction and Classification for Self-Supervised Audio Representation Learning","date":"2025-02-17","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/audio-embeddings-as-teachers-for-music","title":"Audio Embeddings as Teachers for Music Classification","date":"2023-06-30","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/contrastive-learning-of-musical","title":"Contrastive Learning of Musical Representations","date":"2021-03-17","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/evaluation-of-cnn-based-automatic-music","title":"Evaluation of CNN-based Automatic Music Tagging Models","date":null,"rows_on_this_dataset":1,"code_links":7,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":9,"samples_ran":2,"samples_unverified":7,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":1,"samples_harvested":9,"samples_ran":2,"samples_unverified":7,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}