{"url":"/dataset/wham","name":"WHAM!","full_name":"WSJ0 Hipster Ambient Mixtures","description_markdown":"The **WSJ0 Hipster Ambient Mixtures** (**WHAM!**) dataset pairs each two-speaker mixture in the wsj0-2mix dataset with a unique noise background scene. It has an extension called [WHAMR!](/dataset/whamr) that adds artificial reverberation to the speech signals in addition to the background noise.\r\n\r\nThe noise audio was collected at various urban locations throughout the San Francisco Bay Area in late 2018. The environments primarily consist of restaurants, cafes, bars, and parks. Audio was recorded using an Apogee Sennheiser binaural microphone on a tripod between 1.0 and 1.5 meters off the ground.","description_withheld":null,"homepage":"http://wham.whisper.ai/","introduced_date":"2019-07-02","introduced_date_note":null,"introduced_by":{"paper":"/paper/wham-extending-speech-separation-to-noisy","title":"WHAM!: Extending Speech Separation to Noisy Environments","first_author":"Gordon Wichern","url":null},"license":{"name":"CC BY-NC 4.0","url":"https://creativecommons.org/licenses/by-nc/4.0/"},"modalities":[{"name":"Speech","url":"/datasets/modality/speech"}],"tasks":[{"name":"Speech Enhancement","url":"/task/speech-enhancement","datasets_with_task":"/datasets/task/speech-enhancement"},{"name":"Speech Separation","url":"/task/speech-separation","datasets_with_task":"/datasets/task/speech-separation"},{"name":"Audio Source Separation","url":"/task/audio-source-separation","datasets_with_task":"/datasets/task/audio-source-separation"},{"name":"Speech Dereverberation","url":"/task/speech-dereverberation","datasets_with_task":"/datasets/task/speech-dereverberation"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["WHAMR!","WHAM!"],"data_loaders":[],"num_papers_in_archive":114,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/speech-separation-on-wham","task":"Speech Separation","dataset_variant":"WHAM!","rows":6,"metrics":["SI-SDRi"],"first_row_in_archive_order":{"model":"SepReformer-L + DM","paper":"/paper/separate-and-reconstruct-asymmetric-encoder","metrics":{"SI-SDRi":"18.4"},"code_links":[{"title":"dmlguq456/SepReformer","url":"https://github.com/dmlguq456/SepReformer"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/speech-enhancement-on-wham","task":"Speech Enhancement","dataset_variant":"WHAM!","rows":1,"metrics":["PESQ","SDR","SI-SNR"],"first_row_in_archive_order":{"model":"SepFormer","paper":"/paper/on-using-transformers-for-speech-separation","metrics":{"PESQ":"3.07","SDR":"15.04","SI-SNR":"14.35"},"code_links":[{"title":"speechbrain/speechbrain","url":"https://github.com/speechbrain/speechbrain"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/wanna-hear-your-voice-adaptive-effective-and","title":"Wanna hear your voice? A sample is all we need!","date":"2024-10-01","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/separate-and-reconstruct-asymmetric-encoder","title":"Separate and Reconstruct: Asymmetric Encoder-Decoder for Speech Separation","date":"2024-06-10","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/mossformer-pushing-the-performance-limit-of","title":"MossFormer: Pushing the Performance Limit of Monaural Speech Separation using Gated Single-Head Transformer with Convolution-Augmented Joint Self-Attentions","date":"2023-02-23","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/an-efficient-encoder-decoder-architecture","title":"An efficient encoder-decoder architecture with top-down attention for speech separation","date":"2022-09-30","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":18,"samples_ran":13,"samples_unverified":5,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/on-using-transformers-for-speech-separation","title":"Exploring Self-Attention Mechanisms for Speech Separation","date":"2022-02-06","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/mossformer2-combining-transformer-and-rnn-1","title":"MossFormer2: Combining Transformer and RNN-Free Recurrent Network for Enhanced Time-Domain Monaural Speech Separation","date":null,"rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":4,"samples_ran":4,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":3,"samples_harvested":23,"samples_ran":18,"samples_unverified":5,"pointer_only_for_licence":1,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}