{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/towards-on-device-keyword-spotting-using-low","title":"Towards on-Device Keyword Spotting using Low-Footprint Quaternion Neural Models","arxiv_id":null,"date":"2023-09-15","proceeding":"IEEE Workshop on Applications of Signal Processing to Audio and Acoustics 2023 9","authors":["Aryan Chaudhary","Vinayak Abrol"],"abstract":"On-device keyword spotting (KWS) is an essential component for wake-up and user interaction on smart edge devices. Existing low-footprint models are mainly based on 2D and 1D convolutions, where the former is better at capturing invariances while the latter enables faster inference times. In this work, we explore Quaternion neural models as an alternative for effective acoustic modeling for the KWS task. Quaternion models can embed various facets of input features within the multiple dimensions of the quaternion space. This leads to smaller & efficient models as compared to their conventional counterparts. We demonstrate this using quaternion versions of the popular KWS models on the Google Command V2 dataset, where our models achieve comparable performance to existing ones. In addition, we also provide an extensive analysis of the learning behavior in the quaternion network to motivate their use in other speech/audio tasks.","url_abs":"https://ieeexplore.ieee.org/document/10248052","url_pdf":"https://ieeexplore.ieee.org/document/10248052","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"towards-on-device-keyword-spotting-using-low","repo_url":"https://github.com/DataSenseiAryan/GoogleSpeechCommandLowFootprint","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"keyword-spotting","task_name":"Keyword Spotting"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/keyword-spotting-on-google-speech-commands","task":"Keyword Spotting","dataset":"Google Speech Commands","model":"QNN","rank_in_archive_order":29,"of":42,"metrics":{"Google Speech Commands":"98.53","Google Speech Commands V2 35":"98.60"},"uses_additional_data":false},{"leaderboard":"/sota/keyword-spotting-on-google-speech-commands-v2-4","task":"Keyword Spotting","dataset":"Google Speech Commands (v2)","model":"Quaternion Neural Networks","rank_in_archive_order":1,"of":1,"metrics":{"Accuracy(10-fold)":"98.53"},"uses_additional_data":false},{"leaderboard":"/sota/keyword-spotting-on-google-speech-commands-v2-3","task":"Keyword Spotting","dataset":"Google Speech Commands V2 35","model":"QuaternionNeuralNetwork","rank_in_archive_order":1,"of":2,"metrics":{"Accuracy (10-fold)":"98.53"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}