{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/three-dimensional-lip-motion-network-for-text","title":"Three-Dimensional Lip Motion Network for Text-Independent Speaker Recognition","arxiv_id":"2010.06363","date":"2020-10-13","proceeding":null,"authors":["Jianrong Wang","Tong Wu","Shanyu Wang","Mei Yu","Qiang Fang","Ju Zhang","Li Liu"],"abstract":"Lip motion reflects behavior characteristics of speakers, and thus can be used as a new kind of biometrics in speaker recognition. In the literature, lots of works used two-dimensional (2D) lip images to recognize speaker in a textdependent context. However, 2D lip easily suffers from various face orientations. To this end, in this work, we present a novel end-to-end 3D lip motion Network (3LMNet) by utilizing the sentence-level 3D lip motion (S3DLM) to recognize speakers in both the text-independent and text-dependent contexts. A new regional feedback module (RFM) is proposed to obtain attentions in different lip regions. Besides, prior knowledge of lip motion is investigated to complement RFM, where landmark-level and frame-level features are merged to form a better feature representation. Moreover, we present two methods, i.e., coordinate transformation and face posture correction to pre-process the LSD-AV dataset, which contains 68 speakers and 146 sentences per speaker. The evaluation results on this dataset demonstrate that our proposed 3LMNet is superior to the baseline models, i.e., LSTM, VGG-16 and ResNet-34, and outperforms the state-of-the-art using 2D lip image as well as the 3D face. The code of this work is released at https://github.com/wutong18/Three-Dimensional-Lip- Motion-Network-for-Text-Independent-Speaker-Recognition.","url_abs":"https://arxiv.org/abs/2010.06363v1","url_pdf":"https://arxiv.org/pdf/2010.06363v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"three-dimensional-lip-motion-network-for-text","repo_url":"https://github.com/wutong18/Three-Dimensional-Lip-Motion-Network-for-Text-Independent-Speaker-Recognition","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null},{"paper_slug":"three-dimensional-lip-motion-network-for-text","repo_url":"https://github.com/MindCode-4/code-13/tree/main/Three-Dimensional-Lip-Motion-Network-for-Text-Independent-Speaker-Recognition-master","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"three-dimensional-lip-motion-network-for-text","repo_url":"https://github.com/MindCode-4/code-9/tree/main/Three-Dimensional-Lip-Motion-Network-for-Text-Independent-Speaker-Recognition-master","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"three-dimensional-lip-motion-network-for-text","repo_url":"https://github.com/MindCode-4/code-9/tree/main/Two-Layer-ReLU-Network-Analytically","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"three-dimensional-lip-motion-network-for-text","repo_url":"https://github.com/pwc-1/Paper-9/tree/main/3/Three-Dimensional-Lip-Motion-Network-for-Text-Independent-Speaker-Recognition-master","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"three-dimensional-lip-motion-network-for-text","repo_url":"https://github.com/pwc-1/Paper-9/tree/main/4/Three-Dimensional-Lip-Motion-Network-for-Text-Independent-Speaker-Recognition-master","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"speaker-recognition","task_name":"Speaker Recognition"},{"task_slug":"text-independent-speaker-recognition","task_name":"Text-Independent Speaker Recognition"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}