Browse State-of-the-Art › Robust Speech Recognition
Robust Speech Recognition
26 papers with code · 0 benchmarks · 4 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
26 shown of 26 papers with code (97 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
6 Dec 2022 15 repositories listed Syntology ran 5 of 59 samples · 54 unverified · 18 pointer-only (licence)We study the capabilities of speech processing systems trained simply to predict large amounts of transcripts of audio on the internet.
-
11 Oct 2021 2 repositories listedSpeech enhancement (SE) aims to suppress the additive noise from a noisy speech signal to improve the speech's perceptual quality and intelligibility.
-
9 Apr 2018 2 repositories listedDeep generative models have achieved great success in unsupervised learning with the ability to capture complex nonlinear relationships between latent generating factors and observations.
-
2 Oct 2016 2 repositories listedOn the Aurora 4 task, the very deep CNN achieves a WER of 8.
-
16 Apr 2025 1 repository listedWe present a geometry-driven method for normalizing dysarthric speech by modeling time, frequency, and amplitude distortions as smooth, local Lie group transformations of spectrograms.
-
MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens14 Mar 2025 1 repository listedAudio-Visual Speech Recognition (AVSR) achieves robust speech recognition in noisy environments by combining auditory and visual information.
-
3 Feb 2025 1 repository listedIn this work, we propose mWhisper-Flamingo for multilingual AVSR which combines the strengths of a pre-trained audio model (Whisper) and video model (AV-HuBERT).
-
19 Sep 2024 1 repository listedTo mitigate this issue, we propose a novel channel-aware data simulation method for robust ASR training.
-
8 Mar 2024 1 repository listed Syntology ran 7 of 9 samples · 2 unverified · 2 pointer-only (licence)As Automatic Speech Recognition (ASR) models become ever more pervasive, it is important to ensure that they make reliable predictions under corruptions present in the physical and digital world.
-
19 Jan 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedTo this end, we propose to extract a language-space noise embedding from the N-best list to represent the noise conditions of source speech, which can promote the denoising process in GER.
-
29 Jun 2023 1 repository listedWe introduce LyricWhiz, a robust, multilingual, and zero-shot automatic lyrics transcription method achieving state-of-the-art performance on various lyrics transcription datasets, even in challenging genres such as…
-
1 Mar 2023 1 repository listedWe introduce MuAViC, a multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation providing 1200 hours of audio-visual speech in 9 languages.
-
22 Feb 2023 1 repository listedIn this paper, we propose a simple yet effective approach called gradient remedy (GR) to solve interference between task gradients in noise-robust speech recognition, from perspectives of both angle and magnitude.
-
4 Jan 2023 1 repository listed Syntology ran 0 of 6 samples · 6 unverifiedWe improve previous lip reading methods using an Efficient Conformer back-end on top of a ResNet-18 visual front-end and by adding intermediate CTC losses between blocks.
-
5 Oct 2022 1 repository listed Syntology ran 3 of 5 samples · 2 unverified · 5 pointer-only (licence)While Self-Supervised Learning has helped reap the benefit of the scale from the available unlabeled data, the learning paradigms are continuously being bettered.
-
1 Aug 2022 1 repository listedMoreover, to validate whether the data simulated by DENT-DDSP are able to replace the scarce in-domain noisy data in the noise-robust ASR tasks, several downstream ASR models with the same architecture are trained using…
-
19 Jul 2022 1 repository listedTo showcase such integration, we performed experiments on carefully designed synthetic datasets for noisy-reverberant multi-channel ST and SLU tasks, which can be used as benchmark corpora for future research.
-
28 Mar 2022 1 repository listedThen, we propose style learning to map the fused feature close to clean feature, in order to learn latent speech information from the latter, i.
-
25 Mar 2022 1 repository listedIn this paper, a noise-aware training framework based on two cascaded neural structures is proposed to jointly optimize speech enhancement and speech recognition.
-
5 Nov 2021 1 repository listedWe apply adaptive versions of state-of-the-art attacks, such as the Imperceptible ASR attack, to our model, and show that our strongest defense is robust to all attacks that use inaudible noise, and can only be broken…
-
11 Feb 2021 1 repository listedA systematic comparison of these two approaches for end-to-end robust ASR has not been attempted before.
-
5 Nov 2020 1 repository listedThen, for each class, probabilities of this class are used to compute a mean vector, which we refer to as mean soft labels.
-
25 Jan 2020 1 repository listedWe then propose a revised encoder that better learns short- and long-term speech dynamics with an efficient combination of recurrent and convolutional networks.
-
23 Jun 2019 1 repository listedWe investigate the potential of stochastic neural networks for learning effective waveform-based acoustic models.
-
12 Apr 2019 1 repository listedThe latent variables allow us to convert the domain of speech according to its context and domain representation.
-
27 Mar 2018 1 repository listedFirst, we study the effectiveness of different dereverberation networks (the generator in GAN) and find that LSTM leads a significant improvement as compared with feed-forward DNN and CNN in our dataset.
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections