{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/investigating-generative-adversarial-networks","title":"Investigating Generative Adversarial Networks based Speech Dereverberation for Robust Speech Recognition","arxiv_id":"1803.10132","date":"2018-03-27","proceeding":null,"authors":["Ke Wang","Junbo Zhang","Sining Sun","Yujun Wang","Fei Xiang","Lei Xie"],"abstract":"We investigate the use of generative adversarial networks (GANs) in speech\ndereverberation for robust speech recognition. GANs have been recently studied\nfor speech enhancement to remove additive noises, but there still lacks of a\nwork to examine their ability in speech dereverberation and the advantages of\nusing GANs have not been fully established. In this paper, we provide deep\ninvestigations in the use of GAN-based dereverberation front-end in ASR. First,\nwe study the effectiveness of different dereverberation networks (the generator\nin GAN) and find that LSTM leads a significant improvement as compared with\nfeed-forward DNN and CNN in our dataset. Second, further adding residual\nconnections in the deep LSTMs can boost the performance as well. Finally, we\nfind that, for the success of GAN, it is important to update the generator and\nthe discriminator using the same mini-batch data during training. Moreover,\nusing reverberant spectrogram as a condition to discriminator, as suggested in\nprevious studies, may degrade the performance. In summary, our GAN-based\ndereverberation front-end achieves 14%-19% relative CER reduction as compared\nto the baseline DNN dereverberation network when tested on a strong\nmulti-condition training acoustic model.","url_abs":"http://arxiv.org/abs/1803.10132v3","url_pdf":"http://arxiv.org/pdf/1803.10132v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"investigating-generative-adversarial-networks","repo_url":"https://github.com/wangkenpu/rsrgan","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"robust-speech-recognition","task_name":"Robust Speech Recognition"},{"task_slug":"speech-dereverberation","task_name":"Speech Dereverberation"},{"task_slug":"speech-enhancement","task_name":"Speech Enhancement"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}