{"url":"/dataset/realman","name":"RealMAN","full_name":"A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization","description_markdown":"The Audio Signal and Information Processing Lab at Westlake University, in collaboration with AISHELL, has released the Real-recorded and annotated Microphone Array speech&Noise (RealMAN) dataset, which provides annotated multi-channel speech and noise recordings for dynamic speech enhancement and localization:\r\n\r\n- Microphone array: A 32-channel microphone array with high-fidelity microphones is used for recording\r\n- Speech source: A loudspeaker is used for playing source speech signals (about 35 hours of Mandarin speech)\r\n- Recording duration and scene: A total of 83.7 hours of speech signals (about 48.3 hours for static speaker and 35.4 hours for moving speaker) are recorded in 32 different scenes, and 144.5 hours of background noise are recorded in 31 different scenes. Both speech and noise recording scenes cover various common indoor, outdoor, semi-outdoor and transportation environments, which enables the training of general-purpose speech enhancement and source localization networks.\r\n- Annotation: To obtain the task-specific annotations, speaker location is annotated with an omni-directional fisheye camera by automatically detecting the loudspeaker. The direct-path signal is set as the target clean speech for speech enhancement, which is obtained by filtering the source speech signal with an estimated direct-path propagation filter.","description_withheld":null,"homepage":"https://github.com/Audio-WestlakeU/RealMAN","introduced_date":"2024-06-28","introduced_date_note":null,"introduced_by":{"paper":"/paper/realman-a-real-recorded-and-annotated","title":"RealMAN: A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization","first_author":null,"url":null},"license":null,"modalities":[{"name":"Audio","url":"/datasets/modality/audio"},{"name":"Speech","url":"/datasets/modality/speech"}],"tasks":[{"name":"Speech Enhancement","url":"/task/speech-enhancement","datasets_with_task":"/datasets/task/speech-enhancement"},{"name":"Automatic Speech Recognition (ASR)","url":"/task/automatic-speech-recognition","datasets_with_task":"/datasets/task/automatic-speech-recognition"}],"languages":[{"name":"English","url":"/datasets/language/english"},{"name":"Chinese","url":"/datasets/language/chinese"}],"variants":["RealMAN"],"data_loaders":[],"num_papers_in_archive":5,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/automatic-speech-recognition-asr-on-realman","task":"Automatic Speech Recognition (ASR)","dataset_variant":"RealMAN","rows":2,"metrics":["CER"],"first_row_in_archive_order":{"model":"CleanMel-L-mask","paper":"/paper/cleanmel-mel-spectrogram-enhancement-for","metrics":{"CER":"14.4"},"code_links":[{"title":"Audio-WestlakeU/CleanMel","url":"https://github.com/Audio-WestlakeU/CleanMel"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/speech-enhancement-on-realman","task":"Speech Enhancement","dataset_variant":"RealMAN","rows":2,"metrics":["DNSMOS","DNSMOS BAK","DNSMOS OVRL","DNSMOS SIG","PESQ-WB"],"first_row_in_archive_order":{"model":"CleanMel-L-map","paper":"/paper/cleanmel-mel-spectrogram-enhancement-for","metrics":{"DNSMOS":"3.82","DNSMOS BAK":"4.03","DNSMOS OVRL":"3.25","DNSMOS SIG":"3.55","PESQ-WB":"2.10"},"code_links":[{"title":"Audio-WestlakeU/CleanMel","url":"https://github.com/Audio-WestlakeU/CleanMel"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/cleanmel-mel-spectrogram-enhancement-for","title":"CleanMel: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR","date":"2025-02-27","rows_on_this_dataset":2,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}