{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/wav2kws-transfer-learning-from-speech","title":"Wav2KWS: Transfer Learning from Speech Representations for Keyword Spotting","arxiv_id":null,"date":"2021-05-10","proceeding":null,"authors":["D.J SEO","H.S OH","Y.C JUNG"],"abstract":"With the expanding development of on-device artificial intelligence, voice-enabled devices such as smart speakers, wearables,\r\nand other on-device or edge processing systems have been proposed. However, building or obtaining large training datasets that\r\nare essential for robust keyword spotting (KWS) remains cumbersome. To address this problem, we propose a deep neural\r\nnetwork that can rapidly establish a high-performance KWS system from arbitrary keyword instruction sets. We use an encoder\r\npretrained with a large-scale speech corpus as the backbone network and then design an effective transfer network for KWS. To\r\ndemonstrate the feasibility of the proposed network, various experiments were conducted on Google Speech Command Datasets\r\nV1 and V2. In addition, to verify the applicability of the network for different languages, we conducted experiments using three\r\ndifferent Korean speech command datasets. The proposed network outperforms state-of-the-art deep neural networks in both\r\nexperiments. Furthermore, the proposed network can understand real human voice even when trained with synthetic text-to-speech\r\ndata.","url_abs":"https://ieeexplore.ieee.org/document/9427206","url_pdf":"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9427206","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"wav2kws-transfer-learning-from-speech","repo_url":"https://github.com/qute012/Wav2Keyword","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"keyword-spotting","task_name":"Keyword Spotting"},{"task_slug":"text-to-speech","task_name":"Text to Speech"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"},{"task_slug":"text-to-speech-1","task_name":"text-to-speech"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/keyword-spotting-on-google-speech-commands","task":"Keyword Spotting","dataset":"Google Speech Commands","model":"Wav2KWS","rank_in_archive_order":3,"of":42,"metrics":{"Google Speech Commands V1 12":"97.9","Google Speech Commands V2 12":"98.5","Google Speech Commands V2 20":"97.8"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}