{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/speech-enhancement-based-on-reducing-the","title":"Speech Enhancement Based on Reducing the Detail Portion of Speech Spectrograms in Modulation Domain via Discrete Wavelet Transform","arxiv_id":"1811.03486","date":"2018-11-08","proceeding":null,"authors":["Shih-kuang Lee","Syu-Siang Wang","Yu Tsao","Jeih-weih Hung"],"abstract":"In this paper, we propose a novel speech enhancement (SE) method by\nexploiting the discrete wavelet transform (DWT). This new method reduces the\namount of fast time-varying portion, viz. the DWT-wise detail component, in the\nspectrogram of speech signals so as to highlight the speech-dominant component\nand achieves better speech quality. A particularity of this new method is that\nit is completely unsupervised and requires no prior information about the clean\nspeech and noise in the processed utterance. The presented DWT-based SE method\nwith various scaling factors for the detail part is evaluated with a subset of\nAurora-2 database, and the PESQ metric is used to indicate the quality of\nprocessed speech signals. The preliminary results show that the processed\nspeech signals reveal a higher PESQ score in comparison with the original\ncounterparts. Furthermore, we show that this method can still enhance the\nsignal by totally discarding the detail part (setting the respective scaling\nfactor to zero), revealing that the spectrogram can be down-sampled and thus\ncompressed without the cost of lowered quality. In addition, we integrate this\nnew method with conventional speech enhancement algorithms, including spectral\nsubtraction, Wiener filtering, and spectral MMSE estimation, and show that the\nresulting integration behaves better than the respective component method. As a\nresult, this new method is quite effective in improving the speech quality and\nwell additive to the other SE methods.","url_abs":"http://arxiv.org/abs/1811.03486v1","url_pdf":"http://arxiv.org/pdf/1811.03486v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"speech-enhancement-based-on-reducing-the","repo_url":"https://github.com/SKb10/SE_dwt_ISCSLP_2018","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"speech-enhancement","task_name":"Speech Enhancement"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}