{"url":"/dataset/cmedia","name":"CMedia","full_name":null,"description_markdown":"Cmedia dataset:\r\n\r\nThis dataset consists of 200 Youtube links of pop songs (most of them are Chinese songs), together with their groundtruth files of vocal transcription. We will release 100 of them as the open set for training/validation (training set), and use the other 100 as the hidden set for test.\r\n\r\nThe training set can be downloaded here. We strongly suggest participants to use this dataset as training set (if your algorithm is data-driven), since the property of Cmedia training set is close to Cmedia hidden set.","description_withheld":null,"homepage":"https://www.music-ir.org/mirex/wiki/2020:Singing_Transcription_from_Polyphonic_Music","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["CMedia"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}