{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/text-independent-speaker-verification-using","title":"Text-Independent Speaker Verification Using 3D Convolutional Neural Networks","arxiv_id":"1705.09422","date":"2017-05-26","proceeding":null,"authors":["Amirsina Torfi","Jeremy Dawson","Nasser M. Nasrabadi"],"abstract":"In this paper, a novel method using 3D Convolutional Neural Network (3D-CNN)\narchitecture has been proposed for speaker verification in the text-independent\nsetting. One of the main challenges is the creation of the speaker models. Most\nof the previously-reported approaches create speaker models based on averaging\nthe extracted features from utterances of the speaker, which is known as the\nd-vector system. In our paper, we propose an adaptive feature learning by\nutilizing the 3D-CNNs for direct speaker model creation in which, for both\ndevelopment and enrollment phases, an identical number of spoken utterances per\nspeaker is fed to the network for representing the speakers' utterances and\ncreation of the speaker model. This leads to simultaneously capturing the\nspeaker-related information and building a more robust system to cope with\nwithin-speaker variation. We demonstrate that the proposed method significantly\noutperforms the traditional d-vector verification system. Moreover, the\nproposed system can also be an alternative to the traditional d-vector system\nwhich is a one-shot speaker modeling system by utilizing 3D-CNNs.","url_abs":"http://arxiv.org/abs/1705.09422v7","url_pdf":"http://arxiv.org/pdf/1705.09422v7.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"text-independent-speaker-verification-using","repo_url":"https://github.com/astorfi/3D-convolutional-speaker-recognition","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"text-independent-speaker-verification-using","repo_url":"https://github.com/Dou-Yu-xuan/speaker-recognition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"text-independent-speaker-verification-using","repo_url":"https://github.com/arthurutnehmer/3D-convolutional-speaker-recognition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"text-independent-speaker-verification-using","repo_url":"https://github.com/astorfi/3D-convolutional-speaker-recognition-pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"text-independent-speaker-verification-using","repo_url":"https://github.com/zheyejs/3D-convolutional-speaker-recognition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"speaker-verification","task_name":"Speaker Verification"},{"task_slug":"text-independent-speaker-verification","task_name":"Text-Independent Speaker Verification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}