Browse State-of-the-Art › Speech-to-Text Translation
Speech-to-Text Translation
64 papers with code · 11 benchmarks · 4 datasets archive 2025-07-28
Translate audio signals of speech in one language into text in a foreign language, either in an end-to-end or cascade manner.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
11 leaderboard tables shown for this task, 11 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 11 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 64 papers with code (146 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
11 Oct 2020 5 repositories listedWe introduce fairseq S2T, a fairseq extension for speech-to-text (S2T) modeling tasks such as end-to-end speech recognition and speech-to-text translation.
-
22 Aug 2023 4 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)What does it take to create the Babel Fish, a tool that can help individuals translate speech between any two languages?
-
7 Oct 2023 2 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedPrevious mainstream audio-and-text LLMs use discrete audio tokens to represent both input and output audio; however, they suffer from performance degradation on tasks such as automatic speech recognition, speech-to-text…
-
20 May 2022 2 repositories listedPaddleSpeech is an open-source all-in-one speech toolkit.
-
18 Mar 2022 2 repositories listed Syntology ran 0 of 7 samples · 7 unverifiedRecently, speech representation learning has improved many speech-related tasks such as speech recognition, speech classification, and speech-to-text translation.
-
9 Feb 2022 2 repositories listedSpeech translation datasets provide manual segmentations of the audios, which are not available in real-world scenarios, and existing segmentation methods usually significantly reduce translation quality at inference…
-
2 Jun 2021 2 repositories listedAdapter modules were recently introduced as an efficient alternative to fine-tuning in NLP.
-
7 May 2021 2 repositories listed Syntology ran 0 of 7 samples · 7 unverified · 7 pointer-only (licence)By projecting audio and text features to a common semantic representation, Chimera unifies MT and ST tasks and boosts the performance on ST benchmarks, MuST-C and Augmented Librispeech, to a new state-of-the-art.
-
20 Jul 2020 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)Speech translation has recently become an increasingly popular topic of research, partly due to the development of benchmark datasets.
-
29 May 2025 1 repository listedIn particular, on the English→German task, the system achieves a BLEU of 24.
-
24 May 2025 1 repository listedRecent advances in Multimodal Large Language Models (MLLMs) have significantly enhanced the naturalness and flexibility of human computer interaction by enabling seamless understanding across text, vision, and audio…
-
25 Apr 2025 1 repository listedThe model was fine-tuned on MediBeng, a synthetic dataset created to simulate the types of bilingual interactions often found in healthcare environments.
-
13 Feb 2025 1 repository listedWith the growing influence of Large Language Models (LLMs), there is increasing interest in integrating speech representations with them to enable more seamless multi-modal processing and speech understanding.
-
10 Jan 2025 1 repository listedWhile recent multilingual automatic speech recognition models claim to support thousands of languages, ASR for low-resource languages remains highly unreliable due to limited bimodal speech and text training data.
-
22 Jul 2024 1 repository listedWe introduces LLaST, a framework for building high-performance Large Language model based Speech-to-text Translation systems.
-
19 Jul 2024 1 repository listedWe reveal that the inclusion of code-switching units results in higher translation performance than monolingual settings and that models are better at code-switching translation into English than non-English.
-
9 Jul 2024 1 repository listedSpeech Integrated Large Language Models (SILLMs) combine large language models with speech perception to perform diverse tasks, such as emotion recognition to speaker verification, demonstrating universal audio…
-
27 Jun 2024 1 repository listedRecent efforts to develop NLP technologies for African languages have focused on their standard dialects, resulting in disparities for dialects and varieties for which there are little to no resources or tools.
-
26 Jun 2024 1 repository listedOur models and code are available as open-source resources.
-
20 Jun 2024 1 repository listedThis paper describes the FBK's participation in the Simultaneous Translation Evaluation Campaign at IWSLT 2024.
-
10 Jun 2024 1 repository listedTo fill this gap, we introduce StreamAtt, the first StreamST policy, and propose StreamLAAL, the first StreamST latency metric designed to be comparable with existing metrics for SimulST.
-
5 Jun 2024 1 repository listed Syntology ran 4 of 4 samples · 0 unverifiedSimultaneous speech-to-speech translation (Simul-S2ST, a.
-
18 May 2024 1 repository listed Syntology ran 2 of 4 samples · 2 unverifiedA promising approach to preserving model performance in linearized transformers is to employ position-based re-weighting functions.
-
16 Feb 2024 1 repository listedThe speech encoder seamlessly integrates with the MT model at inference, enabling direct translation from speech to text, across all languages supported by the MT model.
-
30 Dec 2023 1 repository listedThis work evaluated several cutting-edge large-scale foundation models based on self-supervision or weak supervision, including SeamlessM4T, SeamlessM4T v2, and Whisper-large-v3, on three code-switched corpora.
-
1 Nov 2023 1 repository listedConventional speech-to-text translation (ST) systems are trained on single-speaker utterances, and they may not generalize to real-life scenarios where the audio contains conversations by multiple speakers.
-
28 Aug 2023 1 repository listedConsistency regularization methods, such as R-Drop (Liang et al., 2021) and CrossConST (Gao et al., 2023), have achieved impressive supervised and zero-shot performance in the neural machine translation (NMT) field.
-
22 Aug 2023 1 repository listedOur single text encoder, covering 200 languages, substantially outperforms existing sentence embeddings such as LASER3 and LabSE on the xsim and xsim++ multilingual similarity search tasks.
-
24 May 2023 1 repository listedJoint speech-language training is challenging due to the large demand for training data and GPU consumption, as well as the modality gap between speech and language.
-
19 May 2023 1 repository listedThe key point is to bridge the modality gap between speech and text so that useful MT techniques can be applied to ST.
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections