{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/matchboxnet-1d-time-channel-separable-1","title":"MatchboxNet: 1D Time-Channel Separable Convolutional Neural Network Architecture for Speech Commands Recognition","arxiv_id":"2004.08531","date":"2020-04-21","proceeding":null,"authors":[],"abstract":"We present an MatchboxNet - an end-to-end neural network for speech command\nrecognition. MatchboxNet is a deep residual network composed from blocks of 1D\ntime-channel separable convolution, batch-normalization, ReLU and dropout\nlayers. MatchboxNet reaches state-of-the-art accuracy on the Google Speech\nCommands dataset while having significantly fewer parameters than similar\nmodels. The small footprint of MatchboxNet makes it an attractive candidate for\ndevices with limited computational resources. The model is highly scalable, so\nmodel accuracy can be improved with modest additional memory and compute.\nFinally, we show how intensive data augmentation using an auxiliary noise\ndataset improves robustness in the presence of background noise.","url_abs":"http://arxiv.org/abs/2004.08531v2","url_pdf":"http://arxiv.org/pdf/2004.08531v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"matchboxnet-1d-time-channel-separable-1","repo_url":"https://github.com/google-research/google-research/tree/master/kws_streaming","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"time-series","task_name":"Time Series Analysis"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/keyword-spotting-on-google-speech-commands","task":"Keyword Spotting","dataset":"Google Speech Commands","model":"MatchboxNet-3x2x64","rank_in_archive_order":6,"of":42,"metrics":{"Google Speech Commands V1 12":"97.48","Google Speech Commands V2 12":"97.63"},"uses_additional_data":false},{"leaderboard":"/sota/time-series-on-speech-commands","task":"Time Series Analysis","dataset":"Speech Commands","model":"MatchboxNet","rank_in_archive_order":4,"of":6,"metrics":{"% Test Accuracy":"97.40"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2004.08531","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}