{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unified-hypersphere-embedding-for-speaker","title":"Unified Hypersphere Embedding for Speaker Recognition","arxiv_id":"1807.08312","date":"2018-07-22","proceeding":null,"authors":["Mahdi Hajibabaei","Dengxin Dai"],"abstract":"Incremental improvements in accuracy of Convolutional Neural Networks are\nusually achieved through use of deeper and more complex models trained on\nlarger datasets. However, enlarging dataset and models increases the\ncomputation and storage costs and cannot be done indefinitely. In this work, we\nseek to improve the identification and verification accuracy of a\ntext-independent speaker recognition system without use of extra data or deeper\nand more complex models by augmenting the training and testing data, finding\nthe optimal dimensionality of embedding space and use of more discriminative\nloss functions. Results of experiments on VoxCeleb dataset suggest that: (i)\nSimple repetition and random time-reversion of utterances can reduce prediction\nerrors by up to 18%. (ii) Lower dimensional embeddings are more suitable for\nverification. (iii) Use of proposed logistic margin loss function leads to\nunified embeddings with state-of-the-art identification and competitive\nverification accuracies.","url_abs":"http://arxiv.org/abs/1807.08312v1","url_pdf":"http://arxiv.org/pdf/1807.08312v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unified-hypersphere-embedding-for-speaker","repo_url":"https://github.com/MahdiHajibabaei/unified-embedding","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"caffe2","reach":{"status":"ok"}}],"tasks":[{"task_slug":"speaker-recognition","task_name":"Speaker Recognition"},{"task_slug":"text-independent-speaker-recognition","task_name":"Text-Independent Speaker Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1807.08312","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}