{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/speechnas-towards-better-trade-off-between","title":"SpeechNAS: Towards Better Trade-off between Latency and Accuracy for Large-Scale Speaker Verification","arxiv_id":"2109.08839","date":"2021-09-18","proceeding":null,"authors":["Wentao Zhu","Tianlong Kong","Shun Lu","Jixiang Li","Dawei Zhang","Feng Deng","Xiaorui Wang","Sen yang","Ji Liu"],"abstract":"Recently, x-vector has been a successful and popular approach for speaker verification, which employs a time delay neural network (TDNN) and statistics pooling to extract speaker characterizing embedding from variable-length utterances. Improvement upon the x-vector has been an active research area, and enormous neural networks have been elaborately designed based on the x-vector, eg, extended TDNN (E-TDNN), factorized TDNN (F-TDNN), and densely connected TDNN (D-TDNN). In this work, we try to identify the optimal architectures from a TDNN based search space employing neural architecture search (NAS), named SpeechNAS. Leveraging the recent advances in the speaker recognition, such as high-order statistics pooling, multi-branch mechanism, D-TDNN and angular additive margin softmax (AAM) loss with a minimum hyper-spherical energy (MHE), SpeechNAS automatically discovers five network architectures, from SpeechNAS-1 to SpeechNAS-5, of various numbers of parameters and GFLOPs on the large-scale text-independent speaker recognition dataset VoxCeleb1. Our derived best neural network achieves an equal error rate (EER) of 1.02% on the standard test set of VoxCeleb1, which surpasses previous TDNN based state-of-the-art approaches by a large margin. Code and trained weights are in https://github.com/wentaozhu/speechnas.git","url_abs":"https://arxiv.org/abs/2109.08839v1","url_pdf":"https://arxiv.org/pdf/2109.08839v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"speechnas-towards-better-trade-off-between","repo_url":"https://github.com/wentaozhu/speechnas","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"architecture-search","task_name":"Neural Architecture Search"},{"task_slug":"speaker-recognition","task_name":"Speaker Recognition"},{"task_slug":"speaker-verification","task_name":"Speaker Verification"},{"task_slug":"text-independent-speaker-recognition","task_name":"Text-Independent Speaker Recognition"}],"methods":[{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"test","method_name":"Test"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/speaker-verification-on-voxceleb","task":"Speaker Verification","dataset":"VoxCeleb","model":"SpeechNAS","rank_in_archive_order":17,"of":21,"metrics":{"EER":"1.02"},"uses_additional_data":false},{"leaderboard":"/sota/speaker-verification-on-voxceleb1","task":"Speaker Verification","dataset":"VoxCeleb1","model":"SpeechNAS","rank_in_archive_order":13,"of":16,"metrics":{"EER":"1.02"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}