{"url":"/dataset/musicbrainz20k","name":"MusicBrainz20K","full_name":null,"description_markdown":"The MusicBrainz20K dataset for entity resolution and entity clustering is based on real records about songs from the MusicBrainz database. Each record is described with the following attributes: artist, title, album, year and length. The records have been modified with the DAPO [1] data generator. The generated dataset consists of five sources and approximately 20K records describing 10K unique song entities. It contains duplicates for 50% of the original records in two to five sources which are generated with a high degree of corruption to stress-test the entity resolution and clustering approaches.\r\n\r\n\r\n[1] Hildebrandt, Kai, et al. \"Large-scale data pollution with Apache Spark.\" IEEE Transactions on Big Data 6.2 (2017): 396-411.","description_withheld":null,"homepage":"https://dbs.uni-leipzig.de/research/projects/object_matching/benchmark_datasets_for_entity_resolution","introduced_date":"2017-09-01","introduced_date_note":null,"introduced_by":null,"license":{"name":"Creative Commons license","url":"https://creativecommons.org/licenses/by/4.0/"},"modalities":[],"tasks":[{"name":"Entity Resolution","url":"/task/entity-resolution","datasets_with_task":"/datasets/task/entity-resolution"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["MusicBrainz20K"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/entity-resolution-on-musicbrainz20k","task":"Entity Resolution","dataset_variant":"MusicBrainz20K","rows":3,"metrics":["F1"],"first_row_in_archive_order":{"model":"ALMSER-GB","paper":"/paper/graph-boosted-active-learning-for-multi","metrics":{"F1":"0.951"},"code_links":[{"title":"wbsg-uni-mannheim/ALMSER-GB","url":"https://github.com/wbsg-uni-mannheim/ALMSER-GB"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/graph-boosted-active-learning-for-multi","title":"Graph-boosted Active Learning for Multi-Source Entity Resolution","date":"2021-09-30","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/scalable-matching-and-clustering-of-entities","title":"Scalable Matching and Clustering of Entities with FAMER","date":"2018-10-01","rows_on_this_dataset":2,"code_links":0,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}