Browse State-of-the-Art › Keyword Spotting
Keyword Spotting
113 papers with code · 10 benchmarks · 8 datasets archive 2025-07-28
In speech processing, keyword spotting deals with the identification of keywords in utterances.
( Image credit: Simon Grest )
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
10 leaderboard tables shown for this task, 10 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
8 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
2 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 113 papers with code (407 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
9 Apr 2018 35 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 1 pointer-only (licence)Describes an audio dataset of spoken words designed to help train and evaluate keyword spotting systems.
-
20 Nov 2017 18 repositories listedWe train various neural network architectures for keyword spotting published in literature to compare their accuracy and memory/compute requirements.
-
1 Apr 2021 10 repositories listed Syntology ran 1 of 16 samples · 15 unverifiedThe Transformer architecture has been successful across many domains, including natural language processing, computer vision and speech recognition.
-
12 Jul 2020 7 repositories listed Syntology ran 5 of 14 samples · 9 unverifiedWe present a large-scale comparison of various self-supervised models.
-
5 Apr 2021 5 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedIn the past decade, convolutional neural networks (CNNs) have been widely adopted as the main building block for end-to-end audio classification models, which aim to learn a direct mapping from audio spectrograms to…
-
19 Nov 2018 5 repositories listedWe explore the application of end-to-end stateless temporal modeling to small-footprint keyword spotting as opposed to recurrent networks that model long-term temporal dependencies using internal states.
-
8 Jun 2021 4 repositories listedWe present a broadcasted residual learning method to achieve high accuracy with small model size and computational load.
-
28 Oct 2017 4 repositories listed Syntology ran 1 of 14 samples · 13 unverifiedWe explore the application of deep residual learning and dilated convolutions to the keyword spotting task, using the recently-released Google Speech Commands Dataset as our benchmark.
-
18 Oct 2017 4 repositories listedWe describe Honk, an open-source PyTorch reimplementation of convolutional neural networks for keyword spotting that are included as examples in TensorFlow.
-
9 May 2017 4 repositories listedWell established text line segmentation evaluation schemes such as the Detection Rate or Recognition Accuracy demand for binarized data that is annotated on a pixel level.
-
19 Oct 2021 3 repositories listed Syntology ran 11 of 16 samples · 5 unverifiedHowever, pure Transformer models tend to require more training data compared to CNNs, and the success of the AST relies on supervised pretraining that requires a large amount of labeled data and a complex training…
-
8 Apr 2019 3 repositories listedIn addition, we release the implementation of the proposed and the baseline models including an end-to-end pipeline for training models and evaluating them on mobile devices.
-
29 Mar 2018 3 repositories listedIn this paper, we propose an attention-based end-to-end neural approach for small-footprint keyword spotting (KWS), which aims to simplify the pipelines of building a production-quality KWS system.
-
28 Nov 2016 3 repositories listedWe propose a single neural network architecture for two tasks: on-line keyword spotting and voice activity detection.
-
22 Aug 2024 2 repositories listedBy efficiently exploiting the heterogeneous processing engines of the MCU, the always-on labeling task runs in real-time with an average power cost of up to 8.
-
31 Aug 2023 2 repositories listedThis study presents a novel zero-shot user-defined keyword spotting model that utilizes the audio-phoneme relationship of the keyword to improve performance.
-
4 Oct 2022 2 repositories listedThis paper explores the effectiveness of SSL on small models for KWS and establishes that SSL can enhance the performance of small KWS models when labelled data is scarce.
-
20 May 2022 2 repositories listedPaddleSpeech is an open-source all-in-one speech toolkit.
-
29 Jan 2022 2 repositories listedCatastrophic forgetting is a thorny challenge when updating keyword spotting (KWS) models after deployment.
-
14 Jun 2021 2 repositories listed Syntology ran 1 of 16 samples · 15 unverifiedAdvancements in ultra-low-power tiny machine learning (TinyML) systems promise to unlock an entirely new class of smart applications.
-
3 Apr 2021 2 repositories listedWith just five training examples, we fine-tune the embedding model for keyword spotting and achieve an average F1 score of 0.
-
31 Dec 2020 2 repositories listedTo this end, the football keyword dataset (FKD), as a new keyword spotting dataset in Persian, is collected with crowdsourcing.
-
18 Dec 2020 2 repositories listedThis paper introduces neural architecture search (NAS) for the automatic discovery of small models for keyword spotting (KWS) in limited resource environments.
-
26 Oct 2020 2 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)Testing on the Google Speech Commands Dataset, the proposed QCNN encoder attains a competitive accuracy of 95.
-
21 Aug 2020 2 repositories listedWe describe Howl, an open-source wake word detection toolkit with native support for open speech datasets, like Mozilla Common Voice and Google Speech Commands.
-
9 Oct 2018 2 repositories listedWe propose a practical approach based on federated learning to solve out-of-domain issues with continuously running embedded speech-based models such as wake word detectors.
-
19 Jul 2016 2 repositories listed Syntology ran 0 of 10 samples · 10 unverifiedRobust and far-field speech recognition is critical to enable true hands-free communication.
-
12 Jun 2025 1 repository listed Syntology ran 1 of 10 samples · 9 unverifiedGLAP demonstrates its versatility by achieving competitive performance on standard audio-text retrieval benchmarks like Clotho and AudioCaps, while significantly surpassing existing methods in speech retrieval and…
-
12 Jun 2025 1 repository listedSmall-Footprint Keyword Spotting (SF-KWS) has gained popularity in today's landscape of smart voice-activated devices, smartphones, and Internet of Things (IoT) applications.
-
30 May 2025 1 repository listedFabricated in 40-nm CMOS, Chameleon sets new accuracy records on Omniglot for end-to-end on-chip FSL (96.
Syntology lines on 10 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections