{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/building-state-of-the-art-distant-speech","title":"Building state-of-the-art distant speech recognition using the CHiME-4 challenge with a setup of speech enhancement baseline","arxiv_id":"1803.10109","date":"2018-03-27","proceeding":null,"authors":["Szu-Jui Chen","Aswin Shanmugam Subramanian","Hainan Xu","Shinji Watanabe"],"abstract":"This paper describes a new baseline system for automatic speech recognition\n(ASR) in the CHiME-4 challenge to promote the development of noisy ASR in\nspeech processing communities by providing 1) state-of-the-art system with a\nsimplified single system comparable to the complicated top systems in the\nchallenge, 2) publicly available and reproducible recipe through the main\nrepository in the Kaldi speech recognition toolkit. The proposed system adopts\ngeneralized eigenvalue beamforming with bidirectional long short-term memory\n(LSTM) mask estimation. We also propose to use a time delay neural network\n(TDNN) based on the lattice-free version of the maximum mutual information\n(LF-MMI) trained with augmented all six microphones plus the enhanced data\nafter beamforming. Finally, we use a LSTM language model for lattice and n-best\nre-scoring. The final system achieved 2.74\\% WER for the real test set in the\n6-channel track, which corresponds to the 2nd place in the challenge. In\naddition, the proposed baseline recipe includes four different speech\nenhancement measures, short-time objective intelligibility measure (STOI),\nextended STOI (eSTOI), perceptual evaluation of speech quality (PESQ) and\nspeech distortion ratio (SDR) for the simulation test set. Thus, the recipe\nalso provides an experimental platform for speech enhancement studies with\nthese performance measures.","url_abs":"http://arxiv.org/abs/1803.10109v1","url_pdf":"http://arxiv.org/pdf/1803.10109v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"distant-speech-recognition","task_name":"Distant Speech Recognition"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"noisy-speech-recognition","task_name":"Noisy Speech Recognition"},{"task_slug":"speech-enhancement","task_name":"Speech Enhancement"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/distant-speech-recognition-on-chime-4-real","task":"Distant Speech Recognition","dataset":"CHiME-4 real 6ch","model":"HMM-TDNN(LFMMI) + LSTMLM + NN-GEV","rank_in_archive_order":2,"of":2,"metrics":{"Word Error Rate (WER)":"2.74"},"uses_additional_data":false},{"leaderboard":"/sota/noisy-speech-recognition-on-chime-real","task":"Noisy Speech Recognition","dataset":"CHiME real","model":"HMM-TDNN(LFMMI) + LSTMLM","rank_in_archive_order":2,"of":5,"metrics":{"Percentage error":"11.4"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}