{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/frame-level-speaker-embeddings-for-text","title":"Frame-level speaker embeddings for text-independent speaker recognition and analysis of end-to-end model","arxiv_id":"1809.04437","date":"2018-09-12","proceeding":null,"authors":["Suwon Shon","Hao Tang","James Glass"],"abstract":"In this paper, we propose a Convolutional Neural Network (CNN) based speaker\nrecognition model for extracting robust speaker embeddings. The embedding can\nbe extracted efficiently with linear activation in the embedding layer. To\nunderstand how the speaker recognition model operates with text-independent\ninput, we modify the structure to extract frame-level speaker embeddings from\neach hidden layer. We feed utterances from the TIMIT dataset to the trained\nnetwork and use several proxy tasks to study the networks ability to represent\nspeech input and differentiate voice identity. We found that the networks are\nbetter at discriminating broad phonetic classes than individual phonemes. In\nparticular, frame-level embeddings that belong to the same phonetic classes are\nsimilar (based on cosine distance) for the same speaker. The frame level\nrepresentation also allows us to analyze the networks at the frame level, and\nhas the potential for other analyses to improve speaker recognition.","url_abs":"http://arxiv.org/abs/1809.04437v1","url_pdf":"http://arxiv.org/pdf/1809.04437v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"frame-level-speaker-embeddings-for-text","repo_url":"https://github.com/Splinter0/CoughCNN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"speaker-recognition","task_name":"Speaker Recognition"},{"task_slug":"text-independent-speaker-recognition","task_name":"Text-Independent Speaker Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1809.04437","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}