Browse State-of-the-Art › Speech-to-Text
Speech-to-Text
129 papers with code · 0 benchmarks · 1 dataset archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 129 papers with code (403 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
21 Sep 2016 8 repositories listed Syntology ran 0 of 10 samples · 10 unverifiedRecently, there has been an increasing interest in end-to-end speech recognition that directly transcribes speech to text without any predefined alignments.
-
21 Oct 2019 7 repositories listed Syntology ran 6 of 21 samples · 15 unverifiedAudio captioning is the novel task of general audio content description using free text.
-
11 Oct 2020 5 repositories listedWe introduce fairseq S2T, a fairseq extension for speech-to-text (S2T) modeling tasks such as end-to-end speech recognition and speech-to-text translation.
-
22 Aug 2023 4 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)What does it take to create the Babel Fish, a tool that can help individuals translate speech between any two languages?
-
15 Feb 2018 4 repositories listedIn this paper we introduce a set of resources and tools aimed at providing support for natural language processing, text-to-speech synthesis and speech recognition for Romanian.
-
13 Feb 2018 4 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedDeep learning models with convolutional and recurrent networks are now ubiquitous and analyze massive amounts of audio, image, video, text and graph data, with applications in automatic translation, speech-to-text,…
-
5 Jan 2018 4 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedWe construct targeted audio adversarial examples on automatic speech recognition.
-
15 Oct 2021 3 repositories listedRecent Speech-to-Text models often require a large amount of hardware resources and are mostly trained in English.
-
23 Aug 2021 3 repositories listedHowever, these alignments tend to be brittle and often fail to generalize to long utterances and out-of-domain text, leading to missing or repeating words.
-
24 May 2018 3 repositories listed Syntology ran 3 of 17 samples · 14 unverifiedIn this survey, we consider seq2seq problems from the RL point of view and provide a formulation combining the power of RL methods in decision-making with sequence-to-sequence models that enable remembering long-term…
-
7 Oct 2023 2 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedPrevious mainstream audio-and-text LLMs use discrete audio tokens to represent both input and output audio; however, they suffer from performance degradation on tasks such as automatic speech recognition, speech-to-text…
-
7 Oct 2022 2 repositories listedThe rapid development of single-modal pre-training has prompted researchers to pay more attention to cross-modal pre-training methods.
-
29 Jun 2022 2 repositories listedIn Spoken Language Understanding (SLU) the task is to extract important information from audio commands, like the intent of what a user wants the system to do and special entities like locations or numbers.
-
20 May 2022 2 repositories listedPaddleSpeech is an open-source all-in-one speech toolkit.
-
18 Mar 2022 2 repositories listed Syntology ran 0 of 7 samples · 7 unverifiedRecently, speech representation learning has improved many speech-related tasks such as speech recognition, speech classification, and speech-to-text translation.
-
7 May 2021 2 repositories listed Syntology ran 0 of 7 samples · 7 unverified · 7 pointer-only (licence)By projecting audio and text features to a common semantic representation, Chimera unifies MT and ST tasks and boosts the performance on ST benchmarks, MuST-C and Augmented Librispeech, to a new state-of-the-art.
-
20 Jul 2020 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)Speech translation has recently become an increasingly popular topic of research, partly due to the development of benchmark datasets.
-
13 Dec 2019 2 repositories listedTo our knowledge this is the largest audio corpus in the public domain for speech recognition, both in terms of number of hours and number of languages.
-
7 Apr 2019 2 repositories listedWhereas conventional spoken language understanding (SLU) systems map speech to text, and then text to intent, end-to-end SLU systems map speech directly to intent through a single trainable model.
-
29 May 2025 1 repository listedTraining large-scale models presents challenges not only in terms of resource requirements but also in terms of their convergence.
-
29 May 2025 1 repository listedIn particular, on the English→German task, the system achieves a BLEU of 24.
-
24 May 2025 1 repository listedRecent advances in Multimodal Large Language Models (MLLMs) have significantly enhanced the naturalness and flexibility of human computer interaction by enabling seamless understanding across text, vision, and audio…
-
27 Apr 2025 1 repository listedIn recent years, end-to-end speech-to-speech (S2S) dialogue systems have garnered increasing research attention due to their advantages over traditional cascaded systems, including achieving lower latency and more…
-
25 Apr 2025 1 repository listedThe model was fine-tuned on MediBeng, a synthetic dataset created to simulate the types of bilingual interactions often found in healthcare environments.
-
1 Apr 2025 1 repository listedThis paper introduces a novel method for automated server provisioning by integrating Transformerbased Named Entity Recognition models with Automated Speech Detection using OpenAI's Whisper.
-
19 Feb 2025 1 repository listedFor instance, we find that task models can tolerate a certain level of noise, and are affected differently by the types of errors in the transcript.
-
16 Feb 2025 1 repository listedIn this paper, we propose DuplexMamba, a Mamba-based end-to-end multimodal duplex model for speech-to-text conversation.
-
13 Feb 2025 1 repository listedWith the growing influence of Large Language Models (LLMs), there is increasing interest in integrating speech representations with them to enable more seamless multi-modal processing and speech understanding.
-
5 Feb 2025 1 repository listedTo do so, we introduce a weakly-supervised method that leverages the perplexity of an off-the-shelf text translation system to identify optimal delays on a per-word basis and create aligned synthetic data.
-
23 Jan 2025 1 repository listedLarge Language Models (LLMs) have made significant progress in various downstream tasks, inspiring the development of Speech Understanding Language Models (SULMs) to enable comprehensive speech-based interactions.
Syntology lines on 10 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections