{"url":"/dataset/google-speech-commands-musan","name":"Google Speech Commands - Musan","full_name":null,"description_markdown":"This noisy speech test set is created from the Google Speech Commands v2 [1] and the Musan dataset[2]. \r\n\r\nIt could be downloaded here: https://zenodo.org/record/6066174#.Yn7NPJPMLyU\r\n\r\nSpecifically, we created this test set by mixing the speech in the Google Speech Commands v2 test set with random noise in the Musan dataset at different signal to noise ratio -12.5,-10,0,10,20,30 and 40 decibel (dB). \r\n\r\nThe Google Speech Commands v2 dataset is under the Creative Commons BY 4.0 license. It could be downloaded at: http://download.tensorflow.org/data/speech_commands_v0.02.tar.gz\r\n\r\nThe Musan dataset is under Attribution 4.0 International (CC BY 4.0). It could be downlowned at https://www.openslr.org/17/\r\n\r\nCitations:\r\n\r\n[1] Pete Warden, “Speech commands: A dataset for limited-vocabulary speech recognition,” arXiv preprint arXiv:1804.03209, 2018.\r\n\r\n[2] David Snyder, Guoguo Chen, and Daniel Povey, “Musan: A music, speech, and noise corpus,” arXiv preprint arXiv:1510.08484, 2015.","description_withheld":null,"homepage":"https://zenodo.org/record/6066174#.Yn7NPJPMLyU","introduced_date":"2022-04-27","introduced_date_note":null,"introduced_by":{"paper":"/paper/importantaug-a-data-augmentation-agent-for","title":"ImportantAug: a data augmentation agent for speech","first_author":"Viet Anh Trinh","url":null},"license":{"name":"Creative Commons Attribution 4.0 International","url":"https://creativecommons.org/licenses/by/4.0/legalcode"},"modalities":[{"name":"Speech","url":"/datasets/modality/speech"}],"tasks":[{"name":"Speech Recognition","url":"/task/speech-recognition","datasets_with_task":"/datasets/task/speech-recognition"},{"name":"Keyword Spotting","url":"/task/keyword-spotting","datasets_with_task":"/datasets/task/keyword-spotting"},{"name":"Robust Speech Recognition","url":"/task/robust-speech-recognition","datasets_with_task":"/datasets/task/robust-speech-recognition"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Google Speech Commands - Musan"],"data_loaders":[],"num_papers_in_archive":3,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/speech-recognition-on-google-speech-commands","task":"Speech Recognition","dataset_variant":"Google Speech Commands - Musan","rows":1,"metrics":["Error rate - SNR 0dB"],"first_row_in_archive_order":{"model":"ImportantAug","paper":"/paper/importantaug-a-data-augmentation-agent-for","metrics":{"Error rate - SNR 0dB":"13.3"},"code_links":[{"title":"tvanh512/importantAug","url":"https://github.com/tvanh512/importantAug"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/importantaug-a-data-augmentation-agent-for","title":"ImportantAug: a data augmentation agent for speech","date":"2021-12-14","rows_on_this_dataset":1,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}