Papers › BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic...

BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognition

27 Sep 2021arXiv:2109.13226archive 2025-07-28

Yu Zhang, Daniel S. Park, Wei Han, James Qin, Anmol Gulati, Joel Shor, Aren Jansen, Yuanzhong Xu, Yanping Huang, Shibo Wang, Zongwei Zhou, Bo Li, Min Ma, William Chan, Jiahui Yu, Yongqiang Wang, Liangliang Cao, Khe Chai Sim, Bhuvana Ramabhadran, Tara N. Sainath, Françoise Beaufays, Zhifeng Chen, Quoc V. Le, Chung-Cheng Chiu, Ruoming Pang, Yonghui Wu

We summarize the results of a host of efforts using giant automatic speech recognition (ASR) models pre-trained using large, diverse unlabeled datasets containing approximately a million hours of audio. We find that the combination of pre-training, self-training and scaling up model size greatly increases data efficiency, even for extremely large tasks with tens of thousands of hours of labeled data. In particular, on an ASR task with 34k hours of labeled data, by fine-tuning an 8 billion parameter pre-trained Conformer model we can match state-of-the-art (SoTA) performance with only 3% of the training data and significantly improve SoTA with the full training set. We also report on the universal benefits gained from using big pre-trained and self-trained models for a large set of downstream tasks that cover a wide range of speech domains and span multiple orders of magnitudes of dataset sizes, including obtaining SoTA performance on many public benchmarks. In addition, we utilize the learned representation of pre-trained networks to achieve SoTA results on non-ASR tasks.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language IdentificationSpeech Emotion RecognitionSpeech Recognitionspeech-recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Language Identification VoxForge ConformerG-P Accuracy 99.8 #1 of 1 Archive leaderboard report
Speech Emotion Recognition CREMA-D ConformerXL-P Accuracy 88.2 #2 of 9 Archive leaderboard report
Speech Recognition AMI IMH ConformerXXL-P + Downstream NST Word Error Rate (WER) 7.8 #1 of 2 Archive leaderboard report
Speech Recognition AMI SDM1 ConformerXXL-P Word Error Rate (WER) 17.7 #1 of 2 Archive leaderboard report
Speech Recognition CHiME-6 dev_gss12 ConformerXXL-PS Word Error Rate (WER) 26.2 #2 of 4 Archive leaderboard report
Speech Recognition CHiME-6 eval ConformerXXL-PS Word Error Rate (WER) 31 #2 of 3 Archive leaderboard report
Speech Recognition Common Voice ConformerXXL-P + Downstream NST Test WER 7.7% #1 of 2 Archive leaderboard report
Speech Recognition TED-LIUM ConformerXXL-PS Word Error Rate (WER) 5 #2 of 2 Archive leaderboard report
Speech Recognition WSJ eval92 ConformerXXL-P Word Error Rate (WER) 1.3 #2 of 17 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections