Browse State-of-the-Art › Automatic Speech Recognition (ASR)
Automatic Speech Recognition (ASR)
622 papers with code · 9 benchmarks · 34 datasets archive 2025-07-28
Automatic Speech Recognition (ASR) involves converting spoken language into written text. It is designed to transcribe spoken words into text in real-time, allowing people to communicate with computers, mobile devices, and other technology using their voice. The goal of Automatic Speech Recognition is to accurately transcribe speech, taking into account variations in accent, pronunciation, and speaking style, as well as background noise and other factors that can affect speech quality.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
9 leaderboard tables shown for this task, 9 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
34 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 34 until expanded.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 622 papers with code (3,012 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
9 Apr 2018 35 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 1 pointer-only (licence)Describes an audio dataset of spoken words designed to help train and evaluate keyword spotting systems.
-
18 Apr 2019 30 repositories listed Syntology ran 1 of 18 samples · 17 unverifiedOn LibriSpeech, we achieve 6.
-
16 May 2020 25 repositories listed Syntology ran 4 of 7 samples · 3 unverified · 2 pointer-only (licence)Recently Transformer and Convolution neural network (CNN) based models have shown promising results in Automatic Speech Recognition (ASR), outperforming Recurrent neural networks (RNNs).
-
25 May 2018 16 repositories listed Syntology ran 1 of 15 samples · 14 unverifiedThis paper presents the machine learning architecture of the Snips Voice Platform, a software solution to perform Spoken Language Understanding on microprocessors typical of IoT devices.
-
7 Oct 2016 14 repositories listed Syntology ran 7 of 21 samples · 14 unverified · 6 pointer-only (licence)We consider the two related problems of detecting if an example is misclassified or out-of-distribution.
-
14 Oct 2021 6 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Motivated by the success of T5 (Text-To-Text Transfer Transformer) in pre-trained natural language processing models, we propose a unified-modal SpeechT5 framework that explores the encoder-decoder pre-training for…
-
7 May 2020 6 repositories listedWe demonstrate that on the widely used LibriSpeech benchmark, ContextNet achieves a word error rate (WER) of 2.
-
8 Jun 2017 6 repositories listedThe CTC network sits on top of the encoder and is jointly trained with the attention-based decoder.
-
8 Nov 2020 5 repositories listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)To the best of our knowledge, we have achieved state-of-the-art end-to-end Transformer based model performance on Switchboard and AMI.
-
4 Nov 2022 4 repositories listedThis paper proposes a modification to RNN-Transducer (RNN-T) models for automatic speech recognition (ASR).
-
2 Jun 2022 4 repositories listed Syntology ran 31 of 49 samples · 18 unverifiedAfter re-examining the design choices for both the macro and micro-architecture of Conformer, we propose Squeezeformer which consistently outperforms the state-of-the-art ASR models under the same training schemes.
-
17 Mar 2020 4 repositories listedWe apply automatic speech recognition (ASR) system to obtain a temporally aligned textual description of the speech (similar to subtitles) and treat it as a separate input alongside video frames and the corresponding…
-
9 Nov 2019 4 repositories listedWhile significant improvements have been made in recent years in terms of end-to-end automatic speech recognition (ASR) performance, such improvements were obtained through the use of very large neural networks, unfit…
-
30 Oct 2018 4 repositories listedThe standard approach to mitigate errors made by an automatic speech recognition system is to use confidence scores associated with each predicted word.
-
6 Sep 2018 4 repositories listedThe goal of this work is to recognise phrases and sentences being spoken by a talking face, with or without the audio.
-
5 Jan 2018 4 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedWe construct targeted audio adversarial examples on automatic speech recognition.
-
5 Dec 2017 4 repositories listedAttention-based encoder-decoder architectures such as Listen, Attend, and Spell (LAS), subsume the acoustic, pronunciation and language model components of a traditional automatic speech recognition (ASR) system into a…
-
29 Jul 2015 4 repositories listedThe performance of automatic speech recognition (ASR) has improved tremendously due to the application of deep neural networks (DNNs).
-
23 Jul 2015 4 repositories listed Syntology ran 0 of 17 samples · 17 unverifiedEnergy disaggregation estimates appliance-by-appliance electricity consumption from a single meter that measures the whole home's electricity demand.
-
13 Feb 2024 3 repositories listedWe found that delicate designs are not necessary, while an embarrassingly simple composition of off-the-shelf speech encoder, LLM, and the only trainable linear projector is competent for the ASR task.
-
7 Nov 2023 3 repositories listedWe demonstrate that finetuning Conformer-transducer models on child speech yields significant improvements in ASR performance on child speech, compared to the non-finetuned models.
-
8 Nov 2022 3 repositories listedIn this paper, we introduce the ATCO2 corpus, a dataset that aims at fostering research on the challenging ATC field, which has lagged behind due to lack of annotated data.
-
28 Sep 2022 3 repositories listed Syntology ran 3 of 4 samples · 1 unverifiedIn this work, we present the Textless Vision-Language Transformer (TVLT), where homogeneous transformer blocks take raw visual and audio inputs for vision-and-language representation learning with minimal…
-
29 Mar 2022 3 repositories listedFirst, a non-deterministic WFST outputs all normalization candidates, and then a neural language model picks the best one -- similar to shallow fusion for automatic speech recognition.
-
4 Nov 2021 3 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedAutomatic Music Transcription (AMT), inferring musical notes from raw audio, is a challenging task at the core of music understanding.
-
7 Oct 2021 3 repositories listedWe present a neural-network-based fast diffuse room impulse response generator (FAST-RIR) for generating room impulse responses (RIRs) for a given acoustic environment.
-
12 Feb 2021 3 repositories listedIn this work, we present a hybrid CTC/Attention model based on a ResNet-18 and Convolution-augmented transformer (Conformer), that can be trained in an end-to-end manner.
-
6 Oct 2020 3 repositories listedThis paper presents the sequence-to-sequence (seq2seq) baseline system for the voice conversion challenge (VCC) 2020.
-
7 Sep 2020 3 repositories listedTherefore, we propose preprocessing methods for KsponSpeech corpus and a baseline model for benchmarks.
-
18 Jun 2020 3 repositories listedWe demonstrate that the cross-accent flaws due to speakers' accents are minimized due to the amount of data, making the system feasible for ATC environments.
Syntology lines on 12 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections