{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/3d-convolutional-neural-networks-for-cross","title":"3D Convolutional Neural Networks for Cross Audio-Visual Matching Recognition","arxiv_id":"1706.05739","date":"2017-06-18","proceeding":null,"authors":["Amirsina Torfi","Seyed Mehdi Iranmanesh","Nasser M. Nasrabadi","Jeremy Dawson"],"abstract":"Audio-visual recognition (AVR) has been considered as a solution for speech\nrecognition tasks when the audio is corrupted, as well as a visual recognition\nmethod used for speaker verification in multi-speaker scenarios. The approach\nof AVR systems is to leverage the extracted information from one modality to\nimprove the recognition ability of the other modality by complementing the\nmissing information. The essential problem is to find the correspondence\nbetween the audio and visual streams, which is the goal of this work. We\npropose the use of a coupled 3D Convolutional Neural Network (3D-CNN)\narchitecture that can map both modalities into a representation space to\nevaluate the correspondence of audio-visual streams using the learned\nmultimodal features. The proposed architecture will incorporate both spatial\nand temporal information jointly to effectively find the correlation between\ntemporal information for different modalities. By using a relatively small\nnetwork architecture and much smaller dataset for training, our proposed method\nsurpasses the performance of the existing similar methods for audio-visual\nmatching which use 3D CNNs for feature representation. We also demonstrate that\nan effective pair selection method can significantly increase the performance.\nThe proposed method achieves relative improvements over 20% on the Equal Error\nRate (EER) and over 7% on the Average Precision (AP) in comparison to the\nstate-of-the-art method.","url_abs":"http://arxiv.org/abs/1706.05739v5","url_pdf":"http://arxiv.org/pdf/1706.05739v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"3d-convolutional-neural-networks-for-cross","repo_url":"https://github.com/astorfi/lip-reading-deeplearning","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null},{"paper_slug":"3d-convolutional-neural-networks-for-cross","repo_url":"https://github.com/sergeyrachev/neuon","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"speaker-verification","task_name":"Speaker Verification"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1706.05739","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}