Browse State-of-the-Art › Speech-to-Speech Translation
Speech-to-Speech Translation
39 papers with code · 3 benchmarks · 5 datasets archive 2025-07-28
Speech-to-speech translation (S2ST) consists on translating speech from one language to speech in another language. This can be done with a cascade of automatic speech recognition (ASR), text-to-text machine translation (MT), and text-to-speech (TTS) synthesis sub-systems, which is text-centric. Recently, works on S2ST without relying on intermediate text representation is emerging.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| TAT (8 rows) | Hokkien→En (Two-pass decoding) | Speech-to-speech translation for a real-world unwritten language | code | — | Compare |
| FLEURS X-eng (7 rows) | GenTranslateV2 | GenTranslate: Large Language Models are Generative Multilingual... | code | — | Compare |
| CVSS (2 rows) | SeamlessM4T Large | SeamlessM4T: Massively Multilingual & Multimodal Machine Translation | code | Syntology ran 2 of 2 samples · 0 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
5 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 39 papers with code (117 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
6 Dec 2022 15 repositories listed Syntology ran 5 of 59 samples · 54 unverified · 18 pointer-only (licence)We study the capabilities of speech processing systems trained simply to predict large amounts of transcripts of audio on the internet.
-
7 Sep 2022 6 repositories listed Syntology ran 4 of 15 samples · 11 unverified · 3 pointer-only (licence)We introduce AudioLM, a framework for high-quality audio generation with long-term consistency.
-
22 Aug 2023 4 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)What does it take to create the Babel Fish, a tool that can help individuals translate speech between any two languages?
-
4 Jul 2024 3 repositories listed Syntology ran 10 of 10 samples · 0 unverifiedThis report introduces FunAudioLLM, a model family designed to enhance natural voice interactions between humans and large language models (LLMs).
-
24 May 2023 2 repositories listedWe first pretrain a model on large-scale monolingual speech data, finetune it with a small amount of parallel speech data (20-60 hours), and lastly train with an unsupervised backtranslation objective.
-
21 May 2025 1 repository listedWe propose the unit language to overcome the two modeling challenges.
-
22 Apr 2025 1 repository listedThis paper explores the idea of using phonemes as a textual representation within a conventional multilingual simultaneous speech-to-speech translation pipeline, as opposed to the traditional reliance on text-based…
-
5 Feb 2025 1 repository listedTo do so, we introduce a weakly-supervised method that leverages the perplexity of an off-the-shelf text translation system to identify optimal delays on a per-word basis and create aligned synthetic data.
-
17 Sep 2024 1 repository listedSpeech Emotion Recognition (SER) is a crucial component in developing general-purpose AI agents capable of natural human-computer interaction.
-
17 Jul 2024 1 repository listedThis paper introduces EmoCtrl-TTS, an emotion-controllable zero-shot TTS that can generate highly emotional speech with NVs for any speaker.
-
11 Jun 2024 1 repository listedDirect speech-to-speech translation (S2ST) has achieved impressive translation quality, but it often faces the challenge of slow decoding due to the considerable length of speech sequences.
-
11 Jun 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Simultaneous translation models play a crucial role in facilitating communication.
-
5 Jun 2024 1 repository listed Syntology ran 4 of 4 samples · 0 unverifiedSimultaneous speech-to-speech translation (Simul-S2ST, a.
-
28 May 2024 1 repository listedThere is a rising interest and trend in research towards directly translating speech from one language to another, known as end-to-end speech-to-speech translation.
-
22 May 2024 1 repository listedNon-autoregressive Transformers (NATs) are recently applied in direct speech-to-speech translation systems, which convert speech across different languages without intermediate text data.
-
10 Feb 2024 1 repository listedLeveraging the rich linguistic knowledge and strong reasoning abilities of LLMs, our new paradigm can integrate the rich information in N-best candidates to generate a higher-quality translation result.
-
21 Dec 2023 1 repository listedWe introduce EmphAssess, a prosodic benchmark designed to evaluate the capability of speech-to-speech models to encode and reproduce prosodic emphasis.
-
5 Dec 2023 1 repository listed Syntology ran 4 of 7 samples · 3 unverifiedTo mitigate the problem of the absence of a parallel AV2AV translation dataset, we propose to train our spoken language translation system with the audio-only dataset of A2A.
-
11 Oct 2023 1 repository listed Syntology ran 5 of 7 samples · 2 unverified · 7 pointer-only (licence)However, due to the presence of linguistic and acoustic diversity, the target speech follows a complex multimodal distribution, posing challenges to achieving both high-quality translations and fast decoding speeds for…
-
3 Aug 2023 1 repository listedBy setting both the inputs and outputs of our learning problem as speech units, we propose to train an encoder-decoder model in a many-to-many spoken language translation setting, namely Unit-to-Unit Translation (UTUT).
-
9 Jul 2023 1 repository listedSpeech-to-speech translation systems today do not adequately support use for dialog purposes.
-
1 Jun 2023 1 repository listedRecent work in speech-to-speech translation (S2ST) has focused primarily on offline settings, where the full input utterance is available before any output is given.
-
10 Apr 2023 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 1 pointer-only (licence)ESPnet-ST-v2 is a revamp of the open-source ESPnet-ST toolkit necessitated by the broadening interests of the spoken language translation community.
-
7 Mar 2023 1 repository listed Syntology ran 4 of 6 samples · 2 unverifiedWe propose a cross-lingual neural codec language model, VALL-E X, for cross-lingual speech synthesis.
-
16 Dec 2022 1 repository listedIn this paper, we propose a text-free evaluation metric for end-to-end S2ST, named BLASER, to avoid the dependency on ASR systems.
-
15 Dec 2022 1 repository listedWe enhance the model performance by subword prediction in the first-pass decoder, advanced two-pass decoder architecture design and search strategy, and better training regularization.
-
18 Nov 2022 1 repository listedTo support machine learning of cross-language prosodic mappings and other ways to improve speech-to-speech translation, we present a protocol for collecting closely matched pairs of utterances across languages, a…
-
31 Oct 2022 1 repository listedHowever, direct S2ST suffers from the data scarcity problem because the corpora from speech of the source language to speech of the target language are very rare.
-
21 Oct 2022 1 repository listedIn this paper, we introduce a new and simple method for comparing speech utterances without relying on text transcripts.
-
25 May 2022 1 repository listedSpecifically, a sequence of discrete representations derived in a self-supervised manner are predicted from the model and passed to a vocoder for speech reconstruction, while still facing the following challenges: 1)…
Syntology lines on 10 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections