{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/realman-a-real-recorded-and-annotated","title":"RealMAN: A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization","arxiv_id":"2406.19959","date":"2024-06-28","proceeding":null,"authors":["Bing Yang","Changsheng Quan","Yabo Wang","Pengyu Wang","Yujie Yang","Ying Fang","Nian Shao","Hui Bu","Xin Xu","Xiaofei Li"],"abstract":"The training of deep learning-based multichannel speech enhancement and source localization systems relies heavily on the simulation of room impulse response and multichannel diffuse noise, due to the lack of large-scale real-recorded datasets. However, the acoustic mismatch between simulated and real-world data could degrade the model performance when applying in real-world scenarios. To bridge this simulation-to-real gap, this paper presents a new relatively large-scale Real-recorded and annotated Microphone Array speech&Noise (RealMAN) dataset. The proposed dataset is valuable in two aspects: 1) benchmarking speech enhancement and localization algorithms in real scenarios; 2) offering a substantial amount of real-world training data for potentially improving the performance of real-world applications. Specifically, a 32-channel array with high-fidelity microphones is used for recording. A loudspeaker is used for playing source speech signals (about 35 hours of Mandarin speech). A total of 83.7 hours of speech signals (about 48.3 hours for static speaker and 35.4 hours for moving speaker) are recorded in 32 different scenes, and 144.5 hours of background noise are recorded in 31 different scenes. Both speech and noise recording scenes cover various common indoor, outdoor, semi-outdoor and transportation environments, which enables the training of general-purpose speech enhancement and source localization networks. To obtain the task-specific annotations, speaker location is annotated with an omni-directional fisheye camera by automatically detecting the loudspeaker. The direct-path signal is set as the target clean speech for speech enhancement, which is obtained by filtering the source speech signal with an estimated direct-path propagation filter.","url_abs":"https://arxiv.org/abs/2406.19959v2","url_pdf":"https://arxiv.org/pdf/2406.19959v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"links_only","authors_date_abstract":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license), from the Kaggle arXiv metadata snapshot of 2026-09-12"},"code_links":[{"paper_slug":"realman-a-real-recorded-and-annotated","repo_url":"https://github.com/Audio-WestlakeU/RealMAN","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[{"slug":"realman","name":"RealMAN","full_name":"A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2406.19959","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2406.19959"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Audio-WestlakeU/RealMAN","reach":{"status":"ok"}}],"summary":{"ran":11,"unverified":1},"by_repo_kind":{"official":{"samples":12,"ran":11,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":12,"samples":[{"code_sha256_prefix":"58450c67e0f277f5","entry":"MVDR","repo":"Audio-WestlakeU/RealMAN","repo_kind":"official","path":"baselines/SE/models/oracle_beamformer.py","file_url":"https://github.com/Audio-WestlakeU/RealMAN/blob/HEAD/baselines/SE/models/oracle_beamformer.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"58450c67e0f277f5"}},{"code_sha256_prefix":"1ff742028e3c0670","entry":"complex_cart2polar","repo":"Audio-WestlakeU/RealMAN","repo_kind":"official","path":"baselines/SSL/Module.py","file_url":"https://github.com/Audio-WestlakeU/RealMAN/blob/HEAD/baselines/SSL/Module.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"1ff742028e3c0670"}},{"code_sha256_prefix":"8ac536b66083e14b","entry":"complex_conjugate_multiplication","repo":"Audio-WestlakeU/RealMAN","repo_kind":"official","path":"baselines/SSL/Module.py","file_url":"https://github.com/Audio-WestlakeU/RealMAN/blob/HEAD/baselines/SSL/Module.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"8ac536b66083e14b"}},{"code_sha256_prefix":"668f563a71ffe080","entry":"complex_multiplication","repo":"Audio-WestlakeU/RealMAN","repo_kind":"official","path":"baselines/SSL/Module.py","file_url":"https://github.com/Audio-WestlakeU/RealMAN/blob/HEAD/baselines/SSL/Module.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"668f563a71ffe080"}},{"code_sha256_prefix":"549c7a45f439e5f6","entry":"get_free_gpus","repo":"Audio-WestlakeU/RealMAN","repo_kind":"official","path":"baselines/SSL/run_tasks.py","file_url":"https://github.com/Audio-WestlakeU/RealMAN/blob/HEAD/baselines/SSL/run_tasks.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"549c7a45f439e5f6"}},{"code_sha256_prefix":"015a39124594c6cf","entry":"neg_si_snr","repo":"Audio-WestlakeU/RealMAN","repo_kind":"official","path":"baselines/SE/FaSNet_TAC.py","file_url":"https://github.com/Audio-WestlakeU/RealMAN/blob/HEAD/baselines/SE/FaSNet_TAC.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"015a39124594c6cf"}},{"code_sha256_prefix":"5f311dcd67f9e592","entry":"normalize","repo":"Audio-WestlakeU/RealMAN","repo_kind":"official","path":"baselines/SE/data_loaders/realman_enh_dataset.py","file_url":"https://github.com/Audio-WestlakeU/RealMAN/blob/HEAD/baselines/SE/data_loaders/realman_enh_dataset.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5f311dcd67f9e592"}},{"code_sha256_prefix":"982c1de375d26311","entry":"read_single_task","repo":"Audio-WestlakeU/RealMAN","repo_kind":"official","path":"baselines/SSL/run_tasks.py","file_url":"https://github.com/Audio-WestlakeU/RealMAN/blob/HEAD/baselines/SSL/run_tasks.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"982c1de375d26311"}},{"code_sha256_prefix":"29cb92f7668015b1","entry":"read_tasks","repo":"Audio-WestlakeU/RealMAN","repo_kind":"official","path":"baselines/SSL/run_tasks.py","file_url":"https://github.com/Audio-WestlakeU/RealMAN/blob/HEAD/baselines/SSL/run_tasks.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"29cb92f7668015b1"}},{"code_sha256_prefix":"c90498629cda3b7f","entry":"select_microphone_array_for_enh","repo":"Audio-WestlakeU/RealMAN","repo_kind":"official","path":"baselines/SE/data_loaders/realman_enh_dataset.py","file_url":"https://github.com/Audio-WestlakeU/RealMAN/blob/HEAD/baselines/SE/data_loaders/realman_enh_dataset.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c90498629cda3b7f"}},{"code_sha256_prefix":"fd9146f3aa7abf0a","entry":"stft","repo":"Audio-WestlakeU/RealMAN","repo_kind":"official","path":"baselines/SE/models/oracle_beamformer.py","file_url":"https://github.com/Audio-WestlakeU/RealMAN/blob/HEAD/baselines/SE/models/oracle_beamformer.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"fd9146f3aa7abf0a"}},{"code_sha256_prefix":"cd36ae67cf96481f","entry":"istft","repo":"Audio-WestlakeU/RealMAN","repo_kind":"official","path":"baselines/SE/models/oracle_beamformer.py","file_url":"https://github.com/Audio-WestlakeU/RealMAN/blob/HEAD/baselines/SE/models/oracle_beamformer.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"cd36ae67cf96481f"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}