{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/attention-based-models-for-text-dependent","title":"Attention-Based Models for Text-Dependent Speaker Verification","arxiv_id":"1710.10470","date":"2017-10-28","proceeding":null,"authors":["F A Rezaur Rahman Chowdhury","Quan Wang","Ignacio Lopez Moreno","Li Wan"],"abstract":"Attention-based models have recently shown great performance on a range of\ntasks, such as speech recognition, machine translation, and image captioning\ndue to their ability to summarize relevant information that expands through the\nentire length of an input sequence. In this paper, we analyze the usage of\nattention mechanisms to the problem of sequence summarization in our end-to-end\ntext-dependent speaker recognition system. We explore different topologies and\ntheir variants of the attention layer, and compare different pooling methods on\nthe attention weights. Ultimately, we show that attention-based models can\nimproves the Equal Error Rate (EER) of our speaker verification system by\nrelatively 14% compared to our non-attention LSTM baseline model.","url_abs":"http://arxiv.org/abs/1710.10470v3","url_pdf":"http://arxiv.org/pdf/1710.10470v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"attention-based-models-for-text-dependent","repo_url":"https://github.com/1alexandra/speech","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"attention-based-models-for-text-dependent","repo_url":"https://github.com/liyongze/lstm_speaker_verification","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"speaker-recognition","task_name":"Speaker Recognition"},{"task_slug":"speaker-verification","task_name":"Speaker Verification"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"text-dependent-speaker-verification","task_name":"Text-Dependent Speaker Verification"},{"task_slug":"translation","task_name":"Translation"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1710.10470","atlas_url":"https://app.syntology.ai/?focus=1710.10470","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}