Browse State-of-the-Art › speech-recognition
speech-recognition
1,277 papers with code · 0 benchmarks · 3 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 1,277 papers with code (5,715 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
8 Mar 2021 19 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Mobile devices such as smartphones and autonomous vehicles increasingly rely on deep neural networks (DNNs) to execute complex inference tasks such as image classification and speech recognition, among others.
-
6 Dec 2022 15 repositories listed Syntology ran 5 of 59 samples · 54 unverified · 18 pointer-only (licence)We study the capabilities of speech processing systems trained simply to predict large amounts of transcripts of audio on the internet.
-
26 Oct 2021 9 repositories listedSelf-supervised learning (SSL) achieves great success in speech recognition, while limited exploration has been attempted for other speech processing tasks.
-
4 Sep 2021 9 repositories listedTo address this problem we propose a measure of hardware efficiency of neural architecture search space - matrix efficiency measure (MEM); a search space comprising of hardware-efficient operations; a latency-aware…
-
24 Jun 2020 8 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedThis paper presents XLSR which learns cross-lingual speech representations by pretraining a single model from the raw waveform of speech in multiple languages.
-
14 Oct 2021 6 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Motivated by the success of T5 (Text-To-Text Transfer Transformer) in pre-trained natural language processing models, we propose a unified-modal SpeechT5 framework that explores the encoder-decoder pre-training for…
-
7 May 2020 6 repositories listedWe demonstrate that on the widely used LibriSpeech benchmark, ContextNet achieves a word error rate (WER) of 2.
-
5 Dec 2017 6 repositories listedThe situation gets even worse with distributed training on mobile devices (federated learning), which suffers from higher latency, lower throughput, and intermittent poor connections.
-
3 Apr 2015 6 repositories listedLearning long term dependencies in recurrent networks is difficult due to vanishing and exploding gradients.
-
31 Jul 2024 5 repositories listed Syntology ran 2 of 9 samples · 7 unverifiedThis paper presents a new set of foundation models, called Llama 3.
-
14 Dec 2022 5 repositories listed Syntology ran 5 of 8 samples · 3 unverifiedCurrent self-supervised learning algorithms are often modality-specific and require large amounts of computational resources.
-
12 Oct 2021 5 repositories listedWe integrate the proposed methods into the HuBERT framework.
-
19 Jan 2021 5 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)In this paper, we propose a unified pre-training approach called UniSpeech to learn speech representations with both unlabeled and labeled data, in which supervised phonetic CTC learning and phonetically-aware…
-
10 Dec 2020 5 repositories listedIn this paper, we present a novel two-pass approach to unify streaming and non-streaming end-to-end (E2E) speech recognition in a single model.
-
8 Nov 2020 5 repositories listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)To the best of our knowledge, we have achieved state-of-the-art end-to-end Transformer based model performance on Switchboard and AMI.
-
11 Oct 2020 5 repositories listedWe introduce fairseq S2T, a fairseq extension for speech-to-text (S2T) modeling tasks such as end-to-end speech recognition and speech-to-text translation.
-
7 Feb 2020 5 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 1 pointer-only (licence)We present results on the LibriSpeech dataset showing that limiting the left context for self-attention in the Transformer layers makes decoding computationally tractable for streaming, with only a slight degradation in…
-
11 Oct 2018 5 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)In this paper, we present a novel system that separates the voice of a target speaker from multi-speaker signals, by making use of a reference signal from the target speaker.
-
15 Jun 2017 5 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedMulti-task learning (MTL) has led to successes in many applications of machine learning, from natural language processing and speech recognition to computer vision and drug discovery.
-
12 Aug 2014 5 repositories listedThis approach to decoding enables first-pass speech recognition with a language model, completely unaided by the cumbersome infrastructure of HMM-based systems.
-
13 Jul 2024 4 repositories listedIt is too early to conclude that Mamba is a better alternative to transformers for speech before comparing Mamba with transformers in terms of both performance and efficiency in multiple speech-related tasks.
-
22 May 2023 4 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedExpanding the language coverage of speech technology has the potential to improve access to information for many more people.
-
4 Nov 2022 4 repositories listedThis paper proposes a modification to RNN-Transducer (RNN-T) models for automatic speech recognition (ASR).
-
3 Feb 2022 4 repositories listed Syntology ran 9 of 15 samples · 6 unverifiedIn particular the quantizer projects speech inputs with a randomly initialized matrix, and does a nearest-neighbor lookup in a randomly-initialized codebook.
-
30 Sep 2021 4 repositories listedThis paper explores applying the wav2vec2 framework to speaker recognition instead of speech recognition.
-
24 May 2021 4 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Despite rapid progress in the recent past, current speech recognition systems still require labeled training data which limits this technology to a small fraction of the languages spoken around the globe.
-
13 Mar 2021 4 repositories listedOur linguistic analyses (for Fon and Igbo) provide valuable insights and guidance into the creation of speech recognition models for other African low-resourced languages, as well as guide future NLP research for Fon…
-
2 Feb 2021 4 repositories listedIn this paper, we propose an open source, production first, and production ready speech recognition toolkit called WeNet in which a new two-pass approach is implemented to unify streaming and non-streaming end-to-end…
-
17 Mar 2020 4 repositories listedWe apply automatic speech recognition (ASR) system to obtain a temporally aligned textual description of the speech (similar to subtitles) and treat it as a separate input alongside video frames and the corresponding…
-
9 Nov 2019 4 repositories listedWhile significant improvements have been made in recent years in terms of end-to-end automatic speech recognition (ASR) performance, such improvements were obtained through the use of very large neural networks, unfit…
Syntology lines on 14 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections