Papers › MT4SSL: Boosting Self-Supervised Speech Representation Learning by Integrating Multiple Targets

MT4SSL: Boosting Self-Supervised Speech Representation Learning by Integrating Multiple Targets

14 Nov 2022arXiv:2211.07321archive 2025-07-28

Ziyang Ma, Zhisheng Zheng, Changli Tang, Yujin Wang, Xie Chen

In this paper, we provide a new perspective on self-supervised speech models from how the training targets are obtained. We generalize the targets extractor into Offline Targets Extractor (Off-TE) and Online Targets Extractor (On-TE). Based on this, we propose a new multi-tasking learning framework for self-supervised learning, MT4SSL, which stands for Boosting Self-Supervised Speech Representation Learning by Integrating Multiple Targets. MT4SSL uses the K-means algorithm as an Off-TE and a teacher network without gradients as an On-TE, respectively. Our model outperforms previous SSL methods by nontrivial margins on the LibriSpeech benchmark, and is comparable to or even better than the best-performing models with fewer data. Furthermore, we find that using both Off-TE and On-TE results in better convergence in the pre-training phase. With both effectiveness and efficiency, we think doing multi-task learning on self-supervised speech models from our perspective is a promising trend.

PaperPDFCode

Code

ddlbojack/mt4ssl officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Automatic Speech RecognitionMulti-Task LearningRepresentation LearningSelf-Supervised LearningSpeech RecognitionSpeech Representation Learning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Speech Recognition LibriSpeech test-clean MT4SSL Word Error Rate (WER) 3.4 #49 of 64 Archive leaderboard report
Speech Recognition LibriSpeech test-other MT4SSL Word Error Rate (WER) 9.6 #46 of 53 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections