Browse State-of-the-Art › Speech Recognition
Speech Recognition
1,373 papers with code · 65 benchmarks · 96 datasets archive 2025-07-28
Speech Recognition is the task of converting spoken language into text. It involves recognizing the words spoken in an audio recording and transcribing them into a written format. The goal is to accurately transcribe the speech in real-time or from recorded audio, taking into account factors such as accents, speaking speed, and background noise.
( Image credit: SpecAugment )
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
235 leaderboard tables shown for this task, 65 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 235 until expanded.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| LibriSpeech test-clean (64 rows) | United Med ASR | High-precision medical speech recognition through synthetic data... | — | — | Compare |
| LibriSpeech test-other (53 rows) | SAMBA ASR | Samba-ASR: State-Of-The-Art Speech Recognition Leveraging... | — | — | Compare |
| Switchboard + Hub500 (30 rows) | IBM (LSTM+Conformer encoder-decoder) | On the limit of English conversational speech recognition | — | — | Compare |
| TIMIT (22 rows) | wav2vec 2.0 | wav2vec 2.0: A Framework for Self-Supervised Learning of Speech... | code | Syntology ran 2 of 9 samples · 7 unverified | Compare |
| AISHELL-1 (18 rows) | FireRedASR-AED | FireRedASR: Open-Source Industrial-Grade Mandarin Speech... | code | — | Compare |
| WSJ eval92 (17 rows) | Speechstew 100M | SpeechStew: Simply Mix All Available Speech Recognition Data to... | — | — | Compare |
| Common Voice German (14 rows) | wav2vec 2.0 XLS-R 1B + TEVR (5-gram) | TEVR: Improving Speech Recognition by Token Entropy Variance Reduction | code | — | Compare |
| swb_hub_500 WER fullSWBCH (12 rows) | IBM (LSTM+Conformer encoder-decoder) | On the limit of English conversational speech recognition | — | — | Compare |
| TUDA (9 rows) | Conformer-Transducer (no LM) | Automatic Speech Recognition in German: A Detailed Error Analysis | — | — | Compare |
| Common Voice French (8 rows) | ConformerCTC-L (5-gram) | Scribosermo: Fast Speech-to-Text models for German and other Languages | code | — | Compare |
| Common Voice Spanish (8 rows) | ConformerCTC-L (4-gram) | NeMo: a toolkit for building AI applications using Neural Modules | code | — | Compare |
| MediaSpeech (8 rows) | Quartznet | MediaSpeech: Multilanguage ASR Benchmark and Dataset | code | — | Compare |
| SLUE (8 rows) | W2V2-L-LL60K (+ TED-LIUM 3 LM) | SLUE: New Benchmark Tasks for Spoken Language Understanding... | code | Syntology ran 0 of 12 samples · 12 unverified | Compare |
| VietMed (8 rows) | XLSR-53-Viet | VietMed: A Dataset and Benchmark for Automatic Speech Recognition... | code | — | Compare |
| WenetSpeech (8 rows) | Paraformer-large | FunASR: A Fundamental End-to-End Speech Recognition Toolkit | code | — | Compare |
| EasyCom (5 rows) | ReVISE (bf) | ReVISE: Self-Supervised Speech Resynthesis with Visual Input for... | — | — | Compare |
| GigaSpeech DEV (5 rows) | SAMBA ASR | Samba-ASR: State-Of-The-Art Speech Recognition Leveraging... | — | — | Compare |
| GigaSpeech TEST (5 rows) | Zipformer+pruned transducer w/ CR-CTC (no external language model) | CR-CTC: Consistency regularization on CTC for improved speech recognition | code | — | Compare |
| Hub5'00 SwitchBoard (5 rows) | LAS + SpecAugment (with LM, Switchboard mild policy) | SpecAugment: A Simple Data Augmentation Method for Automatic... | code | Syntology ran 1 of 18 samples · 17 unverified | Compare |
| Libri-Light test-clean (5 rows) | wav2vec 2.0 Large-10h-LV-60k | wav2vec 2.0: A Framework for Self-Supervised Learning of Speech... | code | Syntology ran 2 of 9 samples · 7 unverified | Compare |
| Libri-Light test-other (5 rows) | wav2vec 2.0 Large-10h-LV-60k | wav2vec 2.0: A Framework for Self-Supervised Learning of Speech... | code | Syntology ran 2 of 9 samples · 7 unverified | Compare |
| CHiME-6 dev_gss12 (4 rows) | ConformerXXL-PS + G-Augment | G-Augment: Searching for the Meta-Structure of Data Augmentation... | — | — | Compare |
| LRS3-TED (4 rows) | Whisper | Whisper-Flamingo: Integrating Visual Features into Whisper for... | code | Syntology ran 5 of 18 samples · 13 unverified | Compare |
| Tedlium (4 rows) | United-MedASR (764M) | High-precision medical speech recognition through synthetic data... | — | — | Compare |
| WSJ dev93 (4 rows) | CTC-CRF ST-NAS | Efficient Neural Architecture Search for End-to-end Speech... | code | — | Compare |
| CHiME-6 eval (3 rows) | ConformerXXL-PS + G-Augment | G-Augment: Searching for the Meta-Structure of Data Augmentation... | — | — | Compare |
| Common Voice vi (3 rows) | khanhld/chunkformer-large-vie | ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription | code | — | Compare |
| Europarl-ASR EN Guest-test (3 rows) | United-MedASR (764M) | High-precision medical speech recognition through synthetic data... | — | — | Compare |
| Fongbe audio (3 rows) | Triphone (39 features) + LDA and MLLT + SGMM | First Automatic Fongbe Continuous Speech Recognition System:... | code | — | Compare |
| Speech Commands (3 rows) | Centaurus | Let SSMs be ConvNets: State-space Modeling with Optimal Tensor Contractions | — | — | Compare |
| SPGISpeech (3 rows) | Icefall - zipformer transducer | — | — | — | Compare |
| VIVOS (3 rows) | khanhld/chunkformer-large-vie | ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription | code | — | Compare |
| WSJ eval93 (3 rows) | Deep Speech 2 | Deep Speech 2: End-to-End Speech Recognition in English and Mandarin | code | Syntology ran 2 of 39 samples · 37 unverified | Compare |
| AISHELL-2 (2 rows) | Paraformer-large | FunASR: A Fundamental End-to-End Speech Recognition Toolkit | code | — | Compare |
| AMI IMH (2 rows) | ConformerXXL-P + Downstream NST | BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised... | — | — | Compare |
| AMI SDM1 (2 rows) | ConformerXXL-P | BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised... | — | — | Compare |
| Common Voice (2 rows) | ConformerXXL-P + Downstream NST | BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised... | — | — | Compare |
| Common Voice English (2 rows) | parakeet-rnnt-1.1b | Fast Conformer with Linearly Scalable Attention for Efficient... | — | — | Compare |
| Common Voice Italian (2 rows) | Whisper (Large v2) | Robust Speech Recognition via Large-Scale Weak Supervision | code | Syntology ran 5 of 59 samples · 54 unverified | Compare |
| Europarl-ASR EN MEP-test (2 rows) | mllp_2021_offline_filt | Europarl-ASR: A Large Corpus of Parliamentary Debates for... | — | — | Compare |
| LibriCSS (2 rows) | TS-SEP | TS-SEP: Joint Diarization and Separation Conditioned on Estimated... | code | — | Compare |
| TED-LIUM (2 rows) | Whisper-LLaMa-7b | HyPoradise: An Open Baseline for Generative Speech Recognition... | code | Syntology ran 3 of 7 samples · 4 unverified | Compare |
| AISHELL-2 Test Android (1 row) | Qwen-Audio | Qwen-Audio: Advancing Universal Audio Understanding via Unified... | code | Syntology ran 5 of 7 samples · 2 unverified | Compare |
| AISHELL-2 Test IOS (1 row) | Qwen-Audio | Qwen-Audio: Advancing Universal Audio Understanding via Unified... | code | Syntology ran 5 of 7 samples · 2 unverified | Compare |
| AISHELL-2 Test Mic (1 row) | Qwen-Audio | Qwen-Audio: Advancing Universal Audio Understanding via Unified... | code | Syntology ran 5 of 7 samples · 2 unverified | Compare |
| CALLHOME En (1 row) | WavLM Large & EEND-vector clustering | WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack... | code | — | Compare |
| CALLHOME Spanish Speech (1 row) | TDT 0-2 | Efficient Sequence Transduction by Jointly Predicting Tokens and Durations | code | Syntology ran 1 of 2 samples · 1 unverified | Compare |
| CAS-VSR-S101 (1 row) | ES³ Base* | ES3: Evolving Self-Supervised Learning of Robust Audio-Visual... | — | — | Compare |
| Common Voice Frisian (1 row) | wav2vec2-large-xls-r-1b-frisian | Improving the previous state-of-the-art Frisian ASR by fine-tuning XLS-R | — | — | Compare |
| Common Voice Japanese (1 row) | Whisper (Large v2) | Robust Speech Recognition via Large-Scale Weak Supervision | code | Syntology ran 5 of 59 samples · 54 unverified | Compare |
| Common Voice Portuguese (1 row) | XLSR53 Wav2Vec2 Portuguese by Orlem Santos | XLSR53 Wav2Vec2 Portuguese by Orlem Santos | code | — | Compare |
| Common Voice Russian (1 row) | Whisper (Large v2) | Robust Speech Recognition via Large-Scale Weak Supervision | code | Syntology ran 5 of 59 samples · 54 unverified | Compare |
| facebook/multilingual_librispeech german (1 row) | TDT 0-4 | Efficient Sequence Transduction by Jointly Predicting Tokens and Durations | code | Syntology ran 1 of 2 samples · 1 unverified | Compare |
| GigaSpeech (1 row) | Conformer/Transformer-AED | GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours... | code | Syntology ran 2 of 10 samples · 8 unverified | Compare |
| Google Speech Commands - Musan (1 row) | ImportantAug | ImportantAug: a data augmentation agent for speech | code | — | Compare |
| Hub5'00 FISHER-SWBD (1 row) | CTC-CRF | CAT: A CTC-CRF based ASR Toolkit Bridging the Hybrid and the... | code | — | Compare |
| Hub5'00 CallHome (1 row) | Espresso | Espresso: A Fast End-to-end Neural Speech Recognition Toolkit | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
| LibriSpeech 100h test-clean (1 row) | Branchformer + GFSA | Graph Convolutions Enrich the Self-Attention in Transformers! | code | Syntology ran 19 of 29 samples · 10 unverified | Compare |
| LibriSpeech 100h test-other (1 row) | Branchformer + GFSA | Graph Convolutions Enrich the Self-Attention in Transformers! | code | Syntology ran 19 of 29 samples · 10 unverified | Compare |
| LibriSpeech train-clean-100 test-clean (1 row) | wav2vec_wav2letter | Self-training and Pre-training are Complementary for Speech Recognition | code | — | Compare |
| LibriSpeech train-clean-100 test-other (1 row) | wav2vec_wav2letter | Self-training and Pre-training are Complementary for Speech Recognition | code | — | Compare |
| LRS2 (1 row) | RAVEn Large | Jointly Learning Visual and Auditory Speech Representations from Raw Data | code | — | Compare |
| Switchboard (300hr) (1 row) | End-to-end LF-MMI | End-to-end speech recognition using lattice-free MMI | — | — | Compare |
| Switchboard CallHome (1 row) | SpeechStew (100M) | SpeechStew: Simply Mix All Available Speech Recognition Data to... | — | — | Compare |
| Switchboard SWBD (1 row) | SpeechStew (100M) | SpeechStew: Simply Mix All Available Speech Recognition Data to... | — | — | Compare |
| AISHELL-2 Android (0 rows) | no rows in the archive | — | — | ||
| AISHELL-2 Mic (0 rows) | no rows in the archive | — | — | ||
| ATCOSIM (0 rows) | no rows in the archive | — | — | ||
| BembaSpeech bem (0 rows) | no rows in the archive | — | — | ||
| collectivat/tv3_parla ca (0 rows) | no rows in the archive | — | — | ||
| Common Voice Interlingua (0 rows) | no rows in the archive | — | — | ||
| Common Voice Kinyarwanda (0 rows) | no rows in the archive | — | — | ||
| Common Voice 1.0 French (0 rows) | no rows in the archive | — | — | ||
| Common Voice 7.0 Swedish (0 rows) | no rows in the archive | — | — | ||
| Common Voice 7.0 Bashkir (0 rows) | no rows in the archive | — | — | ||
| Common Voice 7.0 Italian (0 rows) | no rows in the archive | — | — | ||
| Common Voice 7.0 Assamese (0 rows) | no rows in the archive | — | — | ||
| Common Voice 7.0 Turkish (0 rows) | no rows in the archive | — | — | ||
| Common Voice 7.0 Punjabi (0 rows) | no rows in the archive | — | — | ||
| Common Voice 7.0 Japanese (0 rows) | no rows in the archive | — | — | ||
| Common Voice 7.0 Portuguese (0 rows) | no rows in the archive | — | — | ||
| Common Voice 7.0 Luganda (0 rows) | no rows in the archive | — | — | ||
| Common Voice 7.0 Vietnamese (0 rows) | no rows in the archive | — | — | ||
| Common Voice 7.0 Finnish (0 rows) | no rows in the archive | — | — | ||
| Common Voice 7.0 Abkhaz (0 rows) | no rows in the archive | — | — | ||
| Common Voice 7.0 Arabic (0 rows) | no rows in the archive | — | — | ||
| Common Voice 7.0 Basque (0 rows) | no rows in the archive | — | — | ||
| Common Voice 7.0 French (0 rows) | no rows in the archive | — | — | ||
| Common Voice 7.0 German (0 rows) | no rows in the archive | — | — | ||
| Common Voice 7.0 Hindi (0 rows) | no rows in the archive | — | — | ||
| Common Voice 7.0 Odia (0 rows) | no rows in the archive | — | — | ||
| Common Voice 7.0 Thai (0 rows) | no rows in the archive | — | — | ||
| Common Voice 7.0 Votic (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Swedish (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Kurmanji Kurdish (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Bulgarian (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Marathi (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Latvian (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Georgian (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Indonesian (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Spanish (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Estonian (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Punjabi (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Portuguese (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Ukrainian (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Maltese (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Romanian (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Slovenian (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Mongolian (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Assamese (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Central Kurdish (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Sorbian, Upper (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Interlingua (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Romansh Sursilvan (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Romansh Vallader (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Vietnamese (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Russian (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Santali (Ol Chiki) (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Serbian (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Catalan (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Guarani (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Galician (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Armenian (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Swahili (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Hungarian (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Japanese (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Abkhaz (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Arabic (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Basaa (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Breton (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Czech (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Dutch (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Erzya (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 French (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 German (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Greek (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Hausa (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Hindi (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Kabyle (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Kazakh (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Odia (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Polish (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Sakha (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Slovak (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Tatar (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Urdu (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Uyghur (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Uzbek (0 rows) | no rows in the archive | — | — | ||
| Common Voice 8.0 Votic (0 rows) | no rows in the archive | — | — | ||
| Common Voice Arabic (0 rows) | no rows in the archive | — | — | ||
| Common Voice Assamese (0 rows) | no rows in the archive | — | — | ||
| Common Voice Basque (0 rows) | no rows in the archive | — | — | ||
| Common Voice Breton (0 rows) | no rows in the archive | — | — | ||
| Common Voice Catalan (0 rows) | no rows in the archive | — | — | ||
| Common Voice Chinese (Hong Kong) (0 rows) | no rows in the archive | — | — | ||
| Common Voice Chinese (China) (0 rows) | no rows in the archive | — | — | ||
| Common Voice Chuvash (0 rows) | no rows in the archive | — | — | ||
| Common Voice Czech (0 rows) | no rows in the archive | — | — | ||
| Common Voice Dhivehi (0 rows) | no rows in the archive | — | — | ||
| Common Voice Dutch (0 rows) | no rows in the archive | — | — | ||
| Common Voice Esperanto (0 rows) | no rows in the archive | — | — | ||
| Common Voice Estonian (0 rows) | no rows in the archive | — | — | ||
| Common Voice Finnish (0 rows) | no rows in the archive | — | — | ||
| Common Voice Galician (0 rows) | no rows in the archive | — | — | ||
| Common Voice Georgian (0 rows) | no rows in the archive | — | — | ||
| Common Voice Greek (0 rows) | no rows in the archive | — | — | ||
| Common Voice Hakha Chin (0 rows) | no rows in the archive | — | — | ||
| Common Voice Hindi (0 rows) | no rows in the archive | — | — | ||
| Common Voice Hungarian (0 rows) | no rows in the archive | — | — | ||
| Common Voice Indonesian (0 rows) | no rows in the archive | — | — | ||
| Common Voice Irish (0 rows) | no rows in the archive | — | — | ||
| Common Voice Kyrgyz (0 rows) | no rows in the archive | — | — | ||
| Common Voice Latvian (0 rows) | no rows in the archive | — | — | ||
| Common Voice Lithuanian (0 rows) | no rows in the archive | — | — | ||
| Common Voice Luganda (0 rows) | no rows in the archive | — | — | ||
| Common Voice Maltese (0 rows) | no rows in the archive | — | — | ||
| Common Voice Mongolian (0 rows) | no rows in the archive | — | — | ||
| Common Voice Odia (0 rows) | no rows in the archive | — | — | ||
| Common Voice Persian (0 rows) | no rows in the archive | — | — | ||
| Common Voice Polish (0 rows) | no rows in the archive | — | — | ||
| Common Voice Punjabi (0 rows) | no rows in the archive | — | — | ||
| Common Voice Romanian (0 rows) | no rows in the archive | — | — | ||
| Common Voice Romansh Vallader (0 rows) | no rows in the archive | — | — | ||
| Common Voice Romansh Sursilvan (0 rows) | no rows in the archive | — | — | ||
| Common Voice Sakha (0 rows) | no rows in the archive | — | — | ||
| Common Voice Slovenian (0 rows) | no rows in the archive | — | — | ||
| Common Voice Sorbian, Upper (0 rows) | no rows in the archive | — | — | ||
| Common Voice Swedish (0 rows) | no rows in the archive | — | — | ||
| Common Voice Tamil (0 rows) | no rows in the archive | — | — | ||
| Common Voice Tatar (0 rows) | no rows in the archive | — | — | ||
| Common Voice Thai (0 rows) | no rows in the archive | — | — | ||
| Common Voice Turkish (0 rows) | no rows in the archive | — | — | ||
| Common Voice Ukrainian (0 rows) | no rows in the archive | — | — | ||
| Common Voice Vietnamese (0 rows) | no rows in the archive | — | — | ||
| Common Voice Welsh (0 rows) | no rows in the archive | — | — | ||
| CORAA (0 rows) | no rows in the archive | — | — | ||
| DARPA TIMIT (0 rows) | no rows in the archive | — | — | ||
| FLEURS (0 rows) | no rows in the archive | — | — | ||
| fon (0 rows) | no rows in the archive | — | — | ||
| German ASR Data-Mix (0 rows) | no rows in the archive | — | — | ||
| InterSpeech 2021 ASR mr (0 rows) | no rows in the archive | — | — | ||
| Kazakh Speech Corpus v1.1 (0 rows) | no rows in the archive | — | — | ||
| Librispeech (other) (0 rows) | no rows in the archive | — | — | ||
| MAdel121/arabic-egy-cleaned (validation split) (0 rows) | no rows in the archive | — | — | ||
| Mozilla Common Voice 10.0 (0 rows) | no rows in the archive | — | — | ||
| Mozilla Common Voice 16.1 (0 rows) | no rows in the archive | — | — | ||
| Mozilla Common Voice 7.0 (0 rows) | no rows in the archive | — | — | ||
| Mozilla Common Voice 8.0 (0 rows) | no rows in the archive | — | — | ||
| Mozilla Common Voice 9.0 (0 rows) | no rows in the archive | — | — | ||
| Multilingual LibriSpeech (0 rows) | no rows in the archive | — | — | ||
| Open SLR (0 rows) | no rows in the archive | — | — | ||
| OpenSLR (0 rows) | no rows in the archive | — | — | ||
| OpenSLR gu (0 rows) | no rows in the archive | — | — | ||
| OpenSLR High quality TTS data for Javanese (0 rows) | no rows in the archive | — | — | ||
| OpenSLR High quality TTS data for Sundanese (0 rows) | no rows in the archive | — | — | ||
| OpenSLR km (0 rows) | no rows in the archive | — | — | ||
| OpenSLR kn (0 rows) | no rows in the archive | — | — | ||
| OpenSLR mr (0 rows) | no rows in the archive | — | — | ||
| OpenSLR ne (0 rows) | no rows in the archive | — | — | ||
| OpenSLR te (0 rows) | no rows in the archive | — | — | ||
| Podlodka.io (0 rows) | no rows in the archive | — | — | ||
| projecte-aina/parlament_parla ca (0 rows) | no rows in the archive | — | — | ||
| Robust Speech Event - Catalan Dev Data (0 rows) | no rows in the archive | — | — | ||
| Robust Speech Event - Dev Data (0 rows) | no rows in the archive | — | — | ||
| Sberdevices Golos (farfield) (0 rows) | no rows in the archive | — | — | ||
| Sberdevices Golos (crowd) (0 rows) | no rows in the archive | — | — | ||
| speech-recognition-community-v2/dev_data ca (0 rows) | no rows in the archive | — | — | ||
| SPGI Speech (0 rows) | no rows in the archive | — | — | ||
| tedlium-v3 (0 rows) | no rows in the archive | — | — | ||
| Test split of combined dataset using all datasets mentioned above (0 rows) | no rows in the archive | — | — | ||
| UWB-ATCC dataset (Air Traffic Control Communications) (0 rows) | no rows in the archive | — | — | ||
| VLSP - Task 1 (0 rows) | no rows in the archive | — | — | ||
| Vox Populi (0 rows) | no rows in the archive | — | — | ||
| Wall Street Journal 92 (0 rows) | no rows in the archive | — | — | ||
| Wall Street Journal 93 (0 rows) | no rows in the archive | — | — | ||
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
96 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 96 until expanded.
Subtasks archive 2025-07-28
11 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 1,373 papers with code (6,433 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
17 Feb 2016 41 repositories listed Syntology ran 40 of 77 samples · 37 unverified · 50 pointer-only (licence)Modern mobile devices have access to a wealth of data suitable for learning models, which in turn can greatly improve the user experience on the device.
-
5 Aug 2015 40 repositories listed Syntology ran 12 of 51 samples · 39 unverified · 9 pointer-only (licence)Unlike traditional DNN-HMM models, this model learns all the components of a speech recognizer jointly.
-
9 Apr 2018 35 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 1 pointer-only (licence)Describes an audio dataset of spoken words designed to help train and evaluate keyword spotting systems.
-
8 Dec 2015 35 repositories listed Syntology ran 2 of 39 samples · 37 unverifiedWe show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech--two vastly different languages.
-
18 Apr 2019 30 repositories listed Syntology ran 1 of 18 samples · 17 unverifiedOn LibriSpeech, we achieve 6.
-
20 Jun 2020 25 repositories listed Syntology ran 2 of 9 samples · 7 unverified · 2 pointer-only (licence)We show for the first time that learning powerful representations from speech audio alone followed by fine-tuning on transcribed speech can outperform the best semi-supervised methods while being conceptually simpler.
-
16 May 2020 25 repositories listed Syntology ran 4 of 7 samples · 3 unverified · 2 pointer-only (licence)Recently Transformer and Convolution neural network (CNN) based models have shown promising results in Automatic Speech Recognition (ASR), outperforming Recurrent neural networks (RNNs).
-
17 Dec 2014 24 repositories listed Syntology ran 9 of 9 samples · 0 unverified · 8 pointer-only (licence)We present a state-of-the-art speech recognition system developed using end-to-end deep learning.
-
8 Sep 2014 21 repositories listed Syntology ran 2 of 6 samples · 4 unverified · 6 pointer-only (licence)We present a simple regularization technique for Recurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM) units.
-
8 Mar 2021 19 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Mobile devices such as smartphones and autonomous vehicles increasingly rely on deep neural networks (DNNs) to execute complex inference tasks such as image classification and speech recognition, among others.
-
13 Mar 2015 17 repositories listedSeveral variants of the Long Short-Term Memory (LSTM) architecture for recurrent neural networks have been proposed since its inception in 1995.
-
25 May 2018 16 repositories listed Syntology ran 1 of 15 samples · 14 unverifiedThis paper presents the machine learning architecture of the Snips Voice Platform, a software solution to perform Spoken Language Understanding on microprocessors typical of IoT devices.
-
6 Dec 2022 15 repositories listed Syntology ran 5 of 59 samples · 54 unverified · 18 pointer-only (licence)We study the capabilities of speech processing systems trained simply to predict large amounts of transcripts of audio on the internet.
-
8 May 2018 14 repositories listed Syntology ran 0 of 12 samples · 12 unverifiedSequence-to-sequence attention-based models on subword units allow simple open-vocabulary end-to-end speech recognition.
-
7 Oct 2016 14 repositories listed Syntology ran 7 of 21 samples · 14 unverified · 6 pointer-only (licence)We consider the two related problems of detecting if an example is misclassified or out-of-distribution.
-
24 Jun 2015 14 repositories listedRecurrent sequence generators conditioned on input data through an attention mechanism have recently shown very good performance on a range of tasks in- cluding machine translation, handwriting synthesis and image…
-
7 Feb 2022 12 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedWhile the general idea of self-supervised learning is identical across modalities, the actual algorithms and objectives differ widely because they were developed with a single modality in mind.
-
14 Jun 2021 11 repositories listed Syntology ran 0 of 9 samples · 9 unverifiedSelf-supervised approaches for speech representation learning are challenged by three unique problems: (1) there are multiple sound units in each input utterance, (2) there is no lexicon of input sound units during the…
-
19 Nov 2018 11 repositories listed Syntology ran 5 of 6 samples · 1 unverified · 6 pointer-only (licence)Experiments, that are conducted on several datasets and tasks, show that PyTorch-Kaldi can effectively be used to develop modern state-of-the-art speech recognizers.
-
1 Apr 2021 10 repositories listed Syntology ran 1 of 16 samples · 15 unverifiedThe Transformer architecture has been successful across many domains, including natural language processing, computer vision and speech recognition.
-
5 Apr 2019 10 repositories listedIn this paper, we report state-of-the-art results on LibriSpeech among end-to-end speech recognition models without any external training data.
-
26 Oct 2021 9 repositories listedSelf-supervised learning (SSL) achieves great success in speech recognition, while limited exploration has been attempted for other speech processing tasks.
-
11 Sep 2016 9 repositories listedThis paper presents a simple end-to-end model for speech recognition, combining a convolutional network based acoustic model and a graph decoding.
-
9 Jun 2015 9 repositories listed Syntology ran 0 of 8 samples · 8 unverifiedRecurrent Neural Networks can be trained to produce sequences of tokens given some input, as exemplified by recent results in machine translation and image captioning.
-
31 Oct 2021 8 repositories listed Syntology ran 28 of 55 samples · 27 unverified · 3 pointer-only (licence)A central goal of sequence modeling is designing a single principled model that can address sequence data across a range of modalities and tasks, particularly on long-range dependencies.
-
24 Jun 2020 8 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedThis paper presents XLSR which learns cross-lingual speech representations by pretraining a single model from the raw waveform of speech in multiple languages.
-
18 Dec 2018 8 repositories listedThis paper introduces wav2letter++, the fastest open-source deep learning speech recognition framework.
-
21 Sep 2016 8 repositories listed Syntology ran 0 of 10 samples · 10 unverifiedRecently, there has been an increasing interest in end-to-end speech recognition that directly transcribes speech to text without any predefined alignments.
-
25 Sep 2023 7 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedPre-training speech models on large volumes of data has achieved remarkable success.
-
12 Jul 2020 7 repositories listed Syntology ran 5 of 14 samples · 9 unverifiedWe present a large-scale comparison of various self-supervised models.
Syntology lines on 24 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections