Datasets › Speech Commands

Speech Commands

Introduced by Pete Warden in Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition9 Apr 2018 archive 2025-07-28

Speech Commands is an audio dataset of spoken words designed to help train and evaluate keyword spotting systems .

Benchmarks archive 2025-07-28

All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Keyword Spotting Google Speech Commands TripletLoss-res15 Google Speech Commands V1 12 98.56 Learning Efficient Representations for Keyword Spotting... roman-vygon/triplet_loss_kws 42 Compare
Audio Classification Speech Commands EAT Accuracy 98.3±0.04 EAT: Self-Supervised Pre-Training with Efficient Audio... cwx-worst-one/eat 7 Compare
Time Series Analysis Speech Commands SepTr % Test Accuracy 98.51 SepTr: Separable Transformer for Audio Spectrogram Processing ristea/septr 6 Compare
Speech Recognition Speech Commands Centaurus Accuracy (%) 98.53 Let SSMs be ConvNets: State-space Modeling with Optimal... — 3 Compare

Papers archive 2025-07-28

30 shown of 41 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 392. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Let SSMs be ConvNets: State-space Modeling with Optimal Tensor Contractions 0 1 22 Jan 2025 not harvested
SSAMBA: Self-Supervised Audio Representation Learning with Mamba State Space Model 1 1 20 May 2024 ran 5 of 8 samples (3 unverified)
Work in Progress: Linear Transformers for TinyML 0 1 25 Mar 2024 not harvested
Mixer is more than just a model 0 1 28 Feb 2024 not harvested
EAT: Self-Supervised Pre-Training with Efficient Audio Transformer 1 1 7 Jan 2024 ran 13 of 17 samples (4 unverified)
Towards on-Device Keyword Spotting using Low-Footprint Quaternion Neural Models 1 1 15 Sep 2023 not harvested
Masked Modeling Duo: Learning Representations by Encouraging Both Networks to Model the Input 1 1 26 Oct 2022 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
Liquid Structural State-Space Models 1 1 26 Sep 2022 not harvested
EfficientLEAF: A Faster LEarnable Audio Frontend of Questionable Use 1 3 12 Jul 2022 not harvested
End-to-End Audio Strikes Back: Boosting Augmentations Towards An Efficient Audio Classification Network 1 1 25 Apr 2022 not harvested
SepTr: Separable Transformer for Audio Spectrogram Processing 1 1 17 Mar 2022 not harvested
HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection 1 1 2 Feb 2022 ran 4 of 9 samples (5 unverified)
ImportantAug: a data augmentation agent for speech 1 1 14 Dec 2021 not harvested
Efficiently Modeling Long Sequences with Structured State Spaces 8 1 31 Oct 2021 ran 28 of 55 samples (27 unverified; 3 pointer-only for licence)
FlexConv: Continuous Kernel Convolutions with Differentiable Kernel Sizes 1 2 15 Oct 2021 not harvested
Attention-Free Keyword Spotting 1 1 14 Oct 2021 ran 0 of 10 samples (10 unverified)
Broadcasted Residual Learning for Efficient Keyword Spotting 4 1 8 Jun 2021 not harvested
Wav2KWS: Transfer Learning from Speech Representations for Keyword Spotting 1 1 10 May 2021 not harvested
Augmenting Deep Classifiers with Polynomial Neural Networks 2 1 16 Apr 2021 ran 4 of 6 samples (2 unverified; 6 pointer-only for licence)
End-to-end Keyword Spotting using Neural Architecture Search and Quantization 0 1 14 Apr 2021 not harvested
AST: Audio Spectrogram Transformer 5 3 5 Apr 2021 ran 0 of 3 samples (3 unverified)
PATE-AAE: Incorporating Adversarial Autoencoder into Private Aggregation of Teacher Ensembles for Spoken Command Classification 0 1 2 Apr 2021 not harvested
Keyword Transformer: A Self-Attention Model for Keyword Spotting 10 3 1 Apr 2021 ran 1 of 16 samples (15 unverified)
SubSpectral Normalization for Neural Audio Data Processing 0 3 25 Mar 2021 not harvested
EdgeCRNN: an edgecomputing oriented model of acoustic feature enhancement for keyword spotting 0 1 14 Mar 2021 not harvested
CKConv: Continuous Kernel Convolution For Sequential Data 1 1 4 Feb 2021 not harvested
Learning Efficient Representations for Keyword Spotting with Triplet Loss 1 1 12 Jan 2021 not harvested
Decentralizing Feature Extraction with Quantum Convolutional Neural Network for Automatic Speech Recognition 2 1 26 Oct 2020 ran 1 of 2 samples (1 unverified; 2 pointer-only for licence)
MicroNets: Neural Network Architectures for Deploying TinyML Applications on Commodity Microcontrollers 1 1 21 Oct 2020 not harvested
Neural Architecture Search For Keyword Spotting 0 2 1 Sep 2020 not harvested

The full list of 41 is in the JSON twin.

Dataset loaders archive 2025-07-28

4 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Google Speech Commands
  • Speech Commands

2 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections