Browse State-of-the-Art › Speech Recognition

Speech Recognition

1,373 papers with code · 65 benchmarks · 96 datasets archive 2025-07-28

AudioSpeech

Speech Recognition is the task of converting spoken language into text. It involves recognizing the words spoken in an audio recording and transcribing them into a written format. The goal is to accurately transcribe the speech in real-time or from recorded audio, taking into account factors such as accents, speaking speed, and background noise.

( Image credit: SpecAugment )

Description from the archive archive 2025-07-28.

Benchmarks archive 2025-07-28

235 leaderboard tables shown for this task, 65 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 235 until expanded.

DatasetBest model (first row in archive order)PaperCodeSyntologyCompare
LibriSpeech test-clean (64 rows) United Med ASR High-precision medical speech recognition through synthetic data... — — Compare
LibriSpeech test-other (53 rows) SAMBA ASR Samba-ASR: State-Of-The-Art Speech Recognition Leveraging... — — Compare
Switchboard + Hub500 (30 rows) IBM (LSTM+Conformer encoder-decoder) On the limit of English conversational speech recognition — — Compare
TIMIT (22 rows) wav2vec 2.0 wav2vec 2.0: A Framework for Self-Supervised Learning of Speech... code Syntology ran 2 of 9 samples · 7 unverified Compare
AISHELL-1 (18 rows) FireRedASR-AED FireRedASR: Open-Source Industrial-Grade Mandarin Speech... code — Compare
WSJ eval92 (17 rows) Speechstew 100M SpeechStew: Simply Mix All Available Speech Recognition Data to... — — Compare
Common Voice German (14 rows) wav2vec 2.0 XLS-R 1B + TEVR (5-gram) TEVR: Improving Speech Recognition by Token Entropy Variance Reduction code — Compare
swb_hub_500 WER fullSWBCH (12 rows) IBM (LSTM+Conformer encoder-decoder) On the limit of English conversational speech recognition — — Compare
TUDA (9 rows) Conformer-Transducer (no LM) Automatic Speech Recognition in German: A Detailed Error Analysis — — Compare
Common Voice French (8 rows) ConformerCTC-L (5-gram) Scribosermo: Fast Speech-to-Text models for German and other Languages code — Compare
Common Voice Spanish (8 rows) ConformerCTC-L (4-gram) NeMo: a toolkit for building AI applications using Neural Modules code — Compare
MediaSpeech (8 rows) Quartznet MediaSpeech: Multilanguage ASR Benchmark and Dataset code — Compare
SLUE (8 rows) W2V2-L-LL60K (+ TED-LIUM 3 LM) SLUE: New Benchmark Tasks for Spoken Language Understanding... code Syntology ran 0 of 12 samples · 12 unverified Compare
VietMed (8 rows) XLSR-53-Viet VietMed: A Dataset and Benchmark for Automatic Speech Recognition... code — Compare
WenetSpeech (8 rows) Paraformer-large FunASR: A Fundamental End-to-End Speech Recognition Toolkit code — Compare
EasyCom (5 rows) ReVISE (bf) ReVISE: Self-Supervised Speech Resynthesis with Visual Input for... — — Compare
GigaSpeech DEV (5 rows) SAMBA ASR Samba-ASR: State-Of-The-Art Speech Recognition Leveraging... — — Compare
GigaSpeech TEST (5 rows) Zipformer+pruned transducer w/ CR-CTC (no external language model) CR-CTC: Consistency regularization on CTC for improved speech recognition code — Compare
Hub5'00 SwitchBoard (5 rows) LAS + SpecAugment (with LM, Switchboard mild policy) SpecAugment: A Simple Data Augmentation Method for Automatic... code Syntology ran 1 of 18 samples · 17 unverified Compare
Libri-Light test-clean (5 rows) wav2vec 2.0 Large-10h-LV-60k wav2vec 2.0: A Framework for Self-Supervised Learning of Speech... code Syntology ran 2 of 9 samples · 7 unverified Compare
Libri-Light test-other (5 rows) wav2vec 2.0 Large-10h-LV-60k wav2vec 2.0: A Framework for Self-Supervised Learning of Speech... code Syntology ran 2 of 9 samples · 7 unverified Compare
CHiME-6 dev_gss12 (4 rows) ConformerXXL-PS + G-Augment G-Augment: Searching for the Meta-Structure of Data Augmentation... — — Compare
LRS3-TED (4 rows) Whisper Whisper-Flamingo: Integrating Visual Features into Whisper for... code Syntology ran 5 of 18 samples · 13 unverified Compare
Tedlium (4 rows) United-MedASR (764M) High-precision medical speech recognition through synthetic data... — — Compare
WSJ dev93 (4 rows) CTC-CRF ST-NAS Efficient Neural Architecture Search for End-to-end Speech... code — Compare
CHiME-6 eval (3 rows) ConformerXXL-PS + G-Augment G-Augment: Searching for the Meta-Structure of Data Augmentation... — — Compare
Common Voice vi (3 rows) khanhld/chunkformer-large-vie ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription code — Compare
Europarl-ASR EN Guest-test (3 rows) United-MedASR (764M) High-precision medical speech recognition through synthetic data... — — Compare
Fongbe audio (3 rows) Triphone (39 features) + LDA and MLLT + SGMM First Automatic Fongbe Continuous Speech Recognition System:... code — Compare
Speech Commands (3 rows) Centaurus Let SSMs be ConvNets: State-space Modeling with Optimal Tensor Contractions — — Compare
SPGISpeech (3 rows) Icefall - zipformer transducer — — — Compare
VIVOS (3 rows) khanhld/chunkformer-large-vie ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription code — Compare
WSJ eval93 (3 rows) Deep Speech 2 Deep Speech 2: End-to-End Speech Recognition in English and Mandarin code Syntology ran 2 of 39 samples · 37 unverified Compare
AISHELL-2 (2 rows) Paraformer-large FunASR: A Fundamental End-to-End Speech Recognition Toolkit code — Compare
AMI IMH (2 rows) ConformerXXL-P + Downstream NST BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised... — — Compare
AMI SDM1 (2 rows) ConformerXXL-P BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised... — — Compare
Common Voice (2 rows) ConformerXXL-P + Downstream NST BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised... — — Compare
Common Voice English (2 rows) parakeet-rnnt-1.1b Fast Conformer with Linearly Scalable Attention for Efficient... — — Compare
Common Voice Italian (2 rows) Whisper (Large v2) Robust Speech Recognition via Large-Scale Weak Supervision code Syntology ran 5 of 59 samples · 54 unverified Compare
Europarl-ASR EN MEP-test (2 rows) mllp_2021_offline_filt Europarl-ASR: A Large Corpus of Parliamentary Debates for... — — Compare
LibriCSS (2 rows) TS-SEP TS-SEP: Joint Diarization and Separation Conditioned on Estimated... code — Compare
TED-LIUM (2 rows) Whisper-LLaMa-7b HyPoradise: An Open Baseline for Generative Speech Recognition... code Syntology ran 3 of 7 samples · 4 unverified Compare
AISHELL-2 Test Android (1 row) Qwen-Audio Qwen-Audio: Advancing Universal Audio Understanding via Unified... code Syntology ran 5 of 7 samples · 2 unverified Compare
AISHELL-2 Test IOS (1 row) Qwen-Audio Qwen-Audio: Advancing Universal Audio Understanding via Unified... code Syntology ran 5 of 7 samples · 2 unverified Compare
AISHELL-2 Test Mic (1 row) Qwen-Audio Qwen-Audio: Advancing Universal Audio Understanding via Unified... code Syntology ran 5 of 7 samples · 2 unverified Compare
CALLHOME En (1 row) WavLM Large & EEND-vector clustering WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack... code — Compare
CALLHOME Spanish Speech (1 row) TDT 0-2 Efficient Sequence Transduction by Jointly Predicting Tokens and Durations code Syntology ran 1 of 2 samples · 1 unverified Compare
CAS-VSR-S101 (1 row) ES³ Base* ES3: Evolving Self-Supervised Learning of Robust Audio-Visual... — — Compare
Common Voice Frisian (1 row) wav2vec2-large-xls-r-1b-frisian Improving the previous state-of-the-art Frisian ASR by fine-tuning XLS-R — — Compare
Common Voice Japanese (1 row) Whisper (Large v2) Robust Speech Recognition via Large-Scale Weak Supervision code Syntology ran 5 of 59 samples · 54 unverified Compare
Common Voice Portuguese (1 row) XLSR53 Wav2Vec2 Portuguese by Orlem Santos XLSR53 Wav2Vec2 Portuguese by Orlem Santos code — Compare
Common Voice Russian (1 row) Whisper (Large v2) Robust Speech Recognition via Large-Scale Weak Supervision code Syntology ran 5 of 59 samples · 54 unverified Compare
facebook/multilingual_librispeech german (1 row) TDT 0-4 Efficient Sequence Transduction by Jointly Predicting Tokens and Durations code Syntology ran 1 of 2 samples · 1 unverified Compare
GigaSpeech (1 row) Conformer/Transformer-AED GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours... code Syntology ran 2 of 10 samples · 8 unverified Compare
Google Speech Commands - Musan (1 row) ImportantAug ImportantAug: a data augmentation agent for speech code — Compare
Hub5'00 FISHER-SWBD (1 row) CTC-CRF CAT: A CTC-CRF based ASR Toolkit Bridging the Hybrid and the... code — Compare
Hub5'00 CallHome (1 row) Espresso Espresso: A Fast End-to-end Neural Speech Recognition Toolkit code Syntology ran 1 of 1 samples · 0 unverified Compare
LibriSpeech 100h test-clean (1 row) Branchformer + GFSA Graph Convolutions Enrich the Self-Attention in Transformers! code Syntology ran 19 of 29 samples · 10 unverified Compare
LibriSpeech 100h test-other (1 row) Branchformer + GFSA Graph Convolutions Enrich the Self-Attention in Transformers! code Syntology ran 19 of 29 samples · 10 unverified Compare
LibriSpeech train-clean-100 test-clean (1 row) wav2vec_wav2letter Self-training and Pre-training are Complementary for Speech Recognition code — Compare
LibriSpeech train-clean-100 test-other (1 row) wav2vec_wav2letter Self-training and Pre-training are Complementary for Speech Recognition code — Compare
LRS2 (1 row) RAVEn Large Jointly Learning Visual and Auditory Speech Representations from Raw Data code — Compare
Switchboard (300hr) (1 row) End-to-end LF-MMI End-to-end speech recognition using lattice-free MMI — — Compare
Switchboard CallHome (1 row) SpeechStew (100M) SpeechStew: Simply Mix All Available Speech Recognition Data to... — — Compare
Switchboard SWBD (1 row) SpeechStew (100M) SpeechStew: Simply Mix All Available Speech Recognition Data to... — — Compare
AISHELL-2 Android (0 rows) no rows in the archive — —
AISHELL-2 Mic (0 rows) no rows in the archive — —
ATCOSIM (0 rows) no rows in the archive — —
BembaSpeech bem (0 rows) no rows in the archive — —
collectivat/tv3_parla ca (0 rows) no rows in the archive — —
Common Voice Interlingua (0 rows) no rows in the archive — —
Common Voice Kinyarwanda (0 rows) no rows in the archive — —
Common Voice 1.0 French (0 rows) no rows in the archive — —
Common Voice 7.0 Swedish (0 rows) no rows in the archive — —
Common Voice 7.0 Bashkir (0 rows) no rows in the archive — —
Common Voice 7.0 Italian (0 rows) no rows in the archive — —
Common Voice 7.0 Assamese (0 rows) no rows in the archive — —
Common Voice 7.0 Turkish (0 rows) no rows in the archive — —
Common Voice 7.0 Punjabi (0 rows) no rows in the archive — —
Common Voice 7.0 Japanese (0 rows) no rows in the archive — —
Common Voice 7.0 Portuguese (0 rows) no rows in the archive — —
Common Voice 7.0 Luganda (0 rows) no rows in the archive — —
Common Voice 7.0 Vietnamese (0 rows) no rows in the archive — —
Common Voice 7.0 Finnish (0 rows) no rows in the archive — —
Common Voice 7.0 Abkhaz (0 rows) no rows in the archive — —
Common Voice 7.0 Arabic (0 rows) no rows in the archive — —
Common Voice 7.0 Basque (0 rows) no rows in the archive — —
Common Voice 7.0 French (0 rows) no rows in the archive — —
Common Voice 7.0 German (0 rows) no rows in the archive — —
Common Voice 7.0 Hindi (0 rows) no rows in the archive — —
Common Voice 7.0 Odia (0 rows) no rows in the archive — —
Common Voice 7.0 Thai (0 rows) no rows in the archive — —
Common Voice 7.0 Votic (0 rows) no rows in the archive — —
Common Voice 8.0 Swedish (0 rows) no rows in the archive — —
Common Voice 8.0 Kurmanji Kurdish (0 rows) no rows in the archive — —
Common Voice 8.0 Bulgarian (0 rows) no rows in the archive — —
Common Voice 8.0 Marathi (0 rows) no rows in the archive — —
Common Voice 8.0 Latvian (0 rows) no rows in the archive — —
Common Voice 8.0 Georgian (0 rows) no rows in the archive — —
Common Voice 8.0 Indonesian (0 rows) no rows in the archive — —
Common Voice 8.0 Spanish (0 rows) no rows in the archive — —
Common Voice 8.0 Estonian (0 rows) no rows in the archive — —
Common Voice 8.0 Punjabi (0 rows) no rows in the archive — —
Common Voice 8.0 Portuguese (0 rows) no rows in the archive — —
Common Voice 8.0 Ukrainian (0 rows) no rows in the archive — —
Common Voice 8.0 Maltese (0 rows) no rows in the archive — —
Common Voice 8.0 Romanian (0 rows) no rows in the archive — —
Common Voice 8.0 Slovenian (0 rows) no rows in the archive — —
Common Voice 8.0 Mongolian (0 rows) no rows in the archive — —
Common Voice 8.0 Assamese (0 rows) no rows in the archive — —
Common Voice 8.0 Central Kurdish (0 rows) no rows in the archive — —
Common Voice 8.0 Sorbian, Upper (0 rows) no rows in the archive — —
Common Voice 8.0 Interlingua (0 rows) no rows in the archive — —
Common Voice 8.0 Romansh Sursilvan (0 rows) no rows in the archive — —
Common Voice 8.0 Romansh Vallader (0 rows) no rows in the archive — —
Common Voice 8.0 Vietnamese (0 rows) no rows in the archive — —
Common Voice 8.0 Russian (0 rows) no rows in the archive — —
Common Voice 8.0 Santali (Ol Chiki) (0 rows) no rows in the archive — —
Common Voice 8.0 Serbian (0 rows) no rows in the archive — —
Common Voice 8.0 Catalan (0 rows) no rows in the archive — —
Common Voice 8.0 Guarani (0 rows) no rows in the archive — —
Common Voice 8.0 Galician (0 rows) no rows in the archive — —
Common Voice 8.0 Armenian (0 rows) no rows in the archive — —
Common Voice 8.0 Swahili (0 rows) no rows in the archive — —
Common Voice 8.0 Hungarian (0 rows) no rows in the archive — —
Common Voice 8.0 Japanese (0 rows) no rows in the archive — —
Common Voice 8.0 Abkhaz (0 rows) no rows in the archive — —
Common Voice 8.0 Arabic (0 rows) no rows in the archive — —
Common Voice 8.0 Basaa (0 rows) no rows in the archive — —
Common Voice 8.0 Breton (0 rows) no rows in the archive — —
Common Voice 8.0 Czech (0 rows) no rows in the archive — —
Common Voice 8.0 Dutch (0 rows) no rows in the archive — —
Common Voice 8.0 Erzya (0 rows) no rows in the archive — —
Common Voice 8.0 French (0 rows) no rows in the archive — —
Common Voice 8.0 German (0 rows) no rows in the archive — —
Common Voice 8.0 Greek (0 rows) no rows in the archive — —
Common Voice 8.0 Hausa (0 rows) no rows in the archive — —
Common Voice 8.0 Hindi (0 rows) no rows in the archive — —
Common Voice 8.0 Kabyle (0 rows) no rows in the archive — —
Common Voice 8.0 Kazakh (0 rows) no rows in the archive — —
Common Voice 8.0 Odia (0 rows) no rows in the archive — —
Common Voice 8.0 Polish (0 rows) no rows in the archive — —
Common Voice 8.0 Sakha (0 rows) no rows in the archive — —
Common Voice 8.0 Slovak (0 rows) no rows in the archive — —
Common Voice 8.0 Tatar (0 rows) no rows in the archive — —
Common Voice 8.0 Urdu (0 rows) no rows in the archive — —
Common Voice 8.0 Uyghur (0 rows) no rows in the archive — —
Common Voice 8.0 Uzbek (0 rows) no rows in the archive — —
Common Voice 8.0 Votic (0 rows) no rows in the archive — —
Common Voice Arabic (0 rows) no rows in the archive — —
Common Voice Assamese (0 rows) no rows in the archive — —
Common Voice Basque (0 rows) no rows in the archive — —
Common Voice Breton (0 rows) no rows in the archive — —
Common Voice Catalan (0 rows) no rows in the archive — —
Common Voice Chinese (Hong Kong) (0 rows) no rows in the archive — —
Common Voice Chinese (China) (0 rows) no rows in the archive — —
Common Voice Chuvash (0 rows) no rows in the archive — —
Common Voice Czech (0 rows) no rows in the archive — —
Common Voice Dhivehi (0 rows) no rows in the archive — —
Common Voice Dutch (0 rows) no rows in the archive — —
Common Voice Esperanto (0 rows) no rows in the archive — —
Common Voice Estonian (0 rows) no rows in the archive — —
Common Voice Finnish (0 rows) no rows in the archive — —
Common Voice Galician (0 rows) no rows in the archive — —
Common Voice Georgian (0 rows) no rows in the archive — —
Common Voice Greek (0 rows) no rows in the archive — —
Common Voice Hakha Chin (0 rows) no rows in the archive — —
Common Voice Hindi (0 rows) no rows in the archive — —
Common Voice Hungarian (0 rows) no rows in the archive — —
Common Voice Indonesian (0 rows) no rows in the archive — —
Common Voice Irish (0 rows) no rows in the archive — —
Common Voice Kyrgyz (0 rows) no rows in the archive — —
Common Voice Latvian (0 rows) no rows in the archive — —
Common Voice Lithuanian (0 rows) no rows in the archive — —
Common Voice Luganda (0 rows) no rows in the archive — —
Common Voice Maltese (0 rows) no rows in the archive — —
Common Voice Mongolian (0 rows) no rows in the archive — —
Common Voice Odia (0 rows) no rows in the archive — —
Common Voice Persian (0 rows) no rows in the archive — —
Common Voice Polish (0 rows) no rows in the archive — —
Common Voice Punjabi (0 rows) no rows in the archive — —
Common Voice Romanian (0 rows) no rows in the archive — —
Common Voice Romansh Vallader (0 rows) no rows in the archive — —
Common Voice Romansh Sursilvan (0 rows) no rows in the archive — —
Common Voice Sakha (0 rows) no rows in the archive — —
Common Voice Slovenian (0 rows) no rows in the archive — —
Common Voice Sorbian, Upper (0 rows) no rows in the archive — —
Common Voice Swedish (0 rows) no rows in the archive — —
Common Voice Tamil (0 rows) no rows in the archive — —
Common Voice Tatar (0 rows) no rows in the archive — —
Common Voice Thai (0 rows) no rows in the archive — —
Common Voice Turkish (0 rows) no rows in the archive — —
Common Voice Ukrainian (0 rows) no rows in the archive — —
Common Voice Vietnamese (0 rows) no rows in the archive — —
Common Voice Welsh (0 rows) no rows in the archive — —
CORAA (0 rows) no rows in the archive — —
DARPA TIMIT (0 rows) no rows in the archive — —
FLEURS (0 rows) no rows in the archive — —
fon (0 rows) no rows in the archive — —
German ASR Data-Mix (0 rows) no rows in the archive — —
InterSpeech 2021 ASR mr (0 rows) no rows in the archive — —
Kazakh Speech Corpus v1.1 (0 rows) no rows in the archive — —
Librispeech (other) (0 rows) no rows in the archive — —
MAdel121/arabic-egy-cleaned (validation split) (0 rows) no rows in the archive — —
Mozilla Common Voice 10.0 (0 rows) no rows in the archive — —
Mozilla Common Voice 16.1 (0 rows) no rows in the archive — —
Mozilla Common Voice 7.0 (0 rows) no rows in the archive — —
Mozilla Common Voice 8.0 (0 rows) no rows in the archive — —
Mozilla Common Voice 9.0 (0 rows) no rows in the archive — —
Multilingual LibriSpeech (0 rows) no rows in the archive — —
Open SLR (0 rows) no rows in the archive — —
OpenSLR (0 rows) no rows in the archive — —
OpenSLR gu (0 rows) no rows in the archive — —
OpenSLR High quality TTS data for Javanese (0 rows) no rows in the archive — —
OpenSLR High quality TTS data for Sundanese (0 rows) no rows in the archive — —
OpenSLR km (0 rows) no rows in the archive — —
OpenSLR kn (0 rows) no rows in the archive — —
OpenSLR mr (0 rows) no rows in the archive — —
OpenSLR ne (0 rows) no rows in the archive — —
OpenSLR te (0 rows) no rows in the archive — —
Podlodka.io (0 rows) no rows in the archive — —
projecte-aina/parlament_parla ca (0 rows) no rows in the archive — —
Robust Speech Event - Catalan Dev Data (0 rows) no rows in the archive — —
Robust Speech Event - Dev Data (0 rows) no rows in the archive — —
Sberdevices Golos (farfield) (0 rows) no rows in the archive — —
Sberdevices Golos (crowd) (0 rows) no rows in the archive — —
speech-recognition-community-v2/dev_data ca (0 rows) no rows in the archive — —
SPGI Speech (0 rows) no rows in the archive — —
tedlium-v3 (0 rows) no rows in the archive — —
Test split of combined dataset using all datasets mentioned above (0 rows) no rows in the archive — —
UWB-ATCC dataset (Air Traffic Control Communications) (0 rows) no rows in the archive — —
VLSP - Task 1 (0 rows) no rows in the archive — —
Vox Populi (0 rows) no rows in the archive — —
Wall Street Journal 92 (0 rows) no rows in the archive — —
Wall Street Journal 93 (0 rows) no rows in the archive — —

Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.

Libraries

Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.

Datasets archive 2025-07-28

96 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 96 until expanded.

Subtasks archive 2025-07-28

11 subtasks in the archive's task tree.

Most implemented papers archive 2025-07-28

30 shown of 1,373 papers with code (6,433 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.

Syntology lines on 24 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections