Browse State-of-the-Art › Speech Representation Learning
Speech Representation Learning
48 papers with code · 0 benchmarks · 0 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
No dataset record in the archive lists this task.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 48 papers with code (131 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
14 Jun 2021 11 repositories listed Syntology ran 0 of 9 samples · 9 unverifiedSelf-supervised approaches for speech representation learning are challenged by three unique problems: (1) there are multiple sound units in each input utterance, (2) there is no lexicon of input sound units during the…
-
Mockingjay: Unsupervised Speech Representation Learning with Deep Bidirectional Transformer Encoders25 Oct 2019 7 repositories listed Syntology ran 3 of 11 samples · 8 unverifiedWe present Mockingjay as a new speech representation learning approach, where bidirectional Transformer encoders are pre-trained on a large amount of unlabeled speech.
-
12 Oct 2021 5 repositories listedWe integrate the proposed methods into the HuBERT framework.
-
19 Jan 2021 5 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)In this paper, we propose a unified pre-training approach called UniSpeech to learn speech representations with both unlabeled and labeled data, in which supervised phonetic CTC learning and phonetically-aware…
-
25 Jan 2019 5 repositories listedWe consider the task of unsupervised extraction of meaningful latent representations of speech by applying autoencoding neural networks to speech waveforms.
-
7 Aug 2021 4 repositories listedIn particular, when compared to published models such as conformer-based wav2vec~2.
-
5 Apr 2019 4 repositories listed Syntology ran 4 of 4 samples · 0 unverified · 3 pointer-only (licence)This paper proposes a novel unsupervised autoregressive neural model for learning generic speech representations.
-
7 Nov 2022 2 repositories listedIn this paper, we extend the pretraining method for cross-lingual multi-speaker speech synthesis tasks, including cross-lingual multi-speaker voice cloning and cross-lingual multi-speaker speech editing.
-
18 Mar 2022 2 repositories listed Syntology ran 0 of 7 samples · 7 unverifiedRecently, speech representation learning has improved many speech-related tasks such as speech recognition, speech classification, and speech-to-text translation.
-
5 Jan 2022 2 repositories listedThe lip-reading WER is further reduced to 26.
-
17 Nov 2021 2 repositories listedOn the CoVoST-2 speech translation benchmark, we improve the previous state of the art by an average of 7.
-
30 Apr 2018 2 repositories listedWe apply these results to pairs of words discovered using an unsupervised algorithm and show an improvement on state-of-the-art in unsupervised representation learning using siamese networks.
-
23 Jan 2025 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedSpecifically, we suggest a unimodal multi-task learning, which distills cross-modal knowledge and aligns the corrupted modalities, by predicting clean audio targets with corrupted videos, and clean video targets with…
-
26 Nov 2024 1 repository listedSelf-supervised learning (SSL) has achieved great success in speech-related tasks.
-
17 Oct 2024 1 repository listed Syntology ran 6 of 10 samples · 4 unverifiedIn this paper, we present EH-MAM (Easy-to-Hard adaptive Masked Acoustic Modeling), a novel self-supervised learning approach for speech representation learning.
-
16 Sep 2024 1 repository listedHowever, we observe that the information aggregated in the CLS token correlates more with speaker identity than with linguistic content.
-
10 Jun 2024 1 repository listedWe present mHuBERT-147, the first general-purpose massively multilingual HuBERT speech representation model trained on 90K hours of clean, open-license data.
-
13 Mar 2024 1 repository listedLastly, we show that the proposed recipe can be applied to other distillation methodologies, such as the recent DPWavLM.
-
21 Feb 2024 1 repository listedWe then show that the quality of the pre-trained model depends mainly on the amount of speech data seen during training, i.
-
18 Oct 2023 1 repository listedUsing a large multilingual audio corpus and self-supervised learning, CLARA develops speech representations enriched with emotions, advancing emotion-aware multilingual speech processing.
-
17 Oct 2023 1 repository listedIn this paper, we present a methodology for linguistic feature extraction, focusing particularly on automatically syllabifying words in multiple languages, with a design to be compatible with a forced-alignment tool,…
-
25 Sep 2023 1 repository listedRecent years have witnessed significant advancements in self-supervised learning (SSL) methods for speech-processing tasks.
-
31 Aug 2023 1 repository listedThis paper proposes a novel semi-supervised TTS framework, QS-TTS, to improve TTS quality with lower supervised data requirements via Vector-Quantized Self-Supervised Speech Representation Learning (VQ-S3RL) utilizing…
-
17 May 2023 1 repository listedIn this paper, we introduce self-distillation and online clustering for self-supervised speech representation learning (DinoSR) which combines masked language modeling, self-distillation, and online clustering.
-
5 May 2023 1 repository listedThe latent space is structured to dissociate the latent dynamical factors that are shared between the modalities from those that are specific to each modality.
-
9 Mar 2023 1 repository listedThis paper presents FaceXHuBERT, a text-less speech-driven 3D facial animation generation method that allows to capture personalized and subtle cues in speech (e.
-
27 Feb 2023 1 repository listedThe SA represents our proposal for an efficient streaming SSRL implementation, while the LLSA solves the latency build-up problem of other streaming attention architectures, such as the masked acausal attention (MAA),…
-
27 Feb 2023 1 repository listedSelf-supervised speech representation learning (SSL) has shown to be effective in various downstream tasks, but SSL models are usually large and slow.
-
14 Nov 2022 1 repository listedIn this paper, we provide a new perspective on self-supervised speech models from how the training targets are obtained.
-
2 Nov 2022 1 repository listedIn this paper, we propose a new Self-Supervised Learning (SSL) algorithm called data2vec-aqc, for speech representation learning from unlabeled speech data.
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections