{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/end-to-end-text-dependent-speaker","title":"End-to-End Text-Dependent Speaker Verification","arxiv_id":"1509.08062","date":"2015-09-27","proceeding":null,"authors":["Georg Heigold","Ignacio Moreno","Samy Bengio","Noam Shazeer"],"abstract":"In this paper we present a data-driven, integrated approach to speaker\nverification, which maps a test utterance and a few reference utterances\ndirectly to a single score for verification and jointly optimizes the system's\ncomponents using the same evaluation protocol and metric as at test time. Such\nan approach will result in simple and efficient systems, requiring little\ndomain-specific knowledge and making few model assumptions. We implement the\nidea by formulating the problem as a single neural network architecture,\nincluding the estimation of a speaker model on only a few utterances, and\nevaluate it on our internal \"Ok Google\" benchmark for text-dependent speaker\nverification. The proposed approach appears to be very effective for big data\napplications like ours that require highly accurate, easy-to-maintain systems\nwith a small footprint.","url_abs":"http://arxiv.org/abs/1509.08062v1","url_pdf":"http://arxiv.org/pdf/1509.08062v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"end-to-end-text-dependent-speaker","repo_url":"https://github.com/Janghyun1230/Speaker_Verification","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"end-to-end-text-dependent-speaker","repo_url":"https://github.com/JanhHyun/Speaker_Verification","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"end-to-end-text-dependent-speaker","repo_url":"https://github.com/dalonlobo/diarization-experiments","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"speaker-verification","task_name":"Speaker Verification"},{"task_slug":"text-dependent-speaker-verification","task_name":"Text-Dependent Speaker Verification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1509.08062","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}