{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/audio-visual-recognition-of-overlapped-speech","title":"Audio-visual Recognition of Overlapped speech for the LRS2 dataset","arxiv_id":"2001.01656","date":"2020-01-06","proceeding":null,"authors":["Jianwei Yu","Shi-Xiong Zhang","Jian Wu","Shahram Ghorbani","Bo Wu","Shiyin Kang","Shansong Liu","Xunying Liu","Helen Meng","Dong Yu"],"abstract":"Automatic recognition of overlapped speech remains a highly challenging task to date. Motivated by the bimodal nature of human speech perception, this paper investigates the use of audio-visual technologies for overlapped speech recognition. Three issues associated with the construction of audio-visual speech recognition (AVSR) systems are addressed. First, the basic architecture designs i.e. end-to-end and hybrid of AVSR systems are investigated. Second, purposefully designed modality fusion gates are used to robustly integrate the audio and visual features. Third, in contrast to a traditional pipelined architecture containing explicit speech separation and recognition components, a streamlined and integrated AVSR system optimized consistently using the lattice-free MMI (LF-MMI) discriminative criterion is also proposed. The proposed LF-MMI time-delay neural network (TDNN) system establishes the state-of-the-art for the LRS2 dataset. Experiments on overlapped speech simulated from the LRS2 dataset suggest the proposed AVSR system outperformed the audio only baseline LF-MMI DNN system by up to 29.98\\% absolute in word error rate (WER) reduction, and produced recognition performance comparable to a more complex pipelined system. Consistent performance improvements of 4.89\\% absolute in WER reduction over the baseline AVSR system using feature fusion are also obtained.","url_abs":"https://arxiv.org/abs/2001.01656v1","url_pdf":"https://arxiv.org/pdf/2001.01656v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"audio-visual-speech-recognition","task_name":"Audio-Visual Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"lipreading","task_name":"Lipreading"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-separation","task_name":"Speech Separation"},{"task_slug":"visual-speech-recognition","task_name":"Visual Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/audio-visual-speech-recognition-on-lrs2","task":"Audio-Visual Speech Recognition","dataset":"LRS2","model":"LF-MMI TDNN","rank_in_archive_order":5,"of":8,"metrics":{"Test WER":"5.9"},"uses_additional_data":false},{"leaderboard":"/sota/automatic-speech-recognition-on-lrs2","task":"Automatic Speech Recognition (ASR)","dataset":"LRS2","model":"LF-MMI TDNN","rank_in_archive_order":6,"of":9,"metrics":{"Test WER":"6.7"},"uses_additional_data":false},{"leaderboard":"/sota/lipreading-on-lrs2","task":"Lipreading","dataset":"LRS2","model":"LF-MMI TDNN","rank_in_archive_order":20,"of":25,"metrics":{"Word Error Rate (WER)":"48.86"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2001.01656","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}