{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cr-ctc-consistency-regularization-on-ctc-for","title":"CR-CTC: Consistency regularization on CTC for improved speech recognition","arxiv_id":"2410.05101","date":"2024-10-07","proceeding":null,"authors":["Zengwei Yao","Wei Kang","Xiaoyu Yang","Fangjun Kuang","Liyong Guo","Han Zhu","Zengrui Jin","Zhaoqing Li","Long Lin","Daniel Povey"],"abstract":"Connectionist Temporal Classification (CTC) is a widely used method for automatic speech recognition (ASR), renowned for its simplicity and computational efficiency. However, it often falls short in recognition performance. In this work, we propose the Consistency-Regularized CTC (CR-CTC), which enforces consistency between two CTC distributions obtained from different augmented views of the input speech mel-spectrogram. We provide in-depth insights into its essential behaviors from three perspectives: 1) it conducts self-distillation between random pairs of sub-models that process different augmented views; 2) it learns contextual representation through masked prediction for positions within time-masked regions, especially when we increase the amount of time masking; 3) it suppresses the extremely peaky CTC distributions, thereby reducing overfitting and improving the generalization ability. Extensive experiments on LibriSpeech, Aishell-1, and GigaSpeech datasets demonstrate the effectiveness of our CR-CTC. It significantly improves the CTC performance, achieving state-of-the-art results comparable to those attained by transducer or systems combining CTC and attention-based encoder-decoder (CTC/AED). We release our code at https://github.com/k2-fsa/icefall.","url_abs":"https://arxiv.org/abs/2410.05101v4","url_pdf":"https://arxiv.org/pdf/2410.05101v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"cr-ctc-consistency-regularization-on-ctc-for","repo_url":"https://github.com/k2-fsa/icefall","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"computational-efficiency","task_name":"Computational Efficiency"},{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/speech-recognition-on-aishell-1","task":"Speech Recognition","dataset":"AISHELL-1","model":"Zipformer+CR-CTC (no external language model)","rank_in_archive_order":6,"of":18,"metrics":{"Params(M)":"66.2","Word Error Rate (WER)":"4.02"},"uses_additional_data":false},{"leaderboard":"/sota/speech-recognition-on-gigaspeech-dev","task":"Speech Recognition","dataset":"GigaSpeech DEV","model":"Zipformer+pruned transducer w/ CR-CTC\n(no external language model)","rank_in_archive_order":2,"of":5,"metrics":{"Word Error Rate (WER)":"9.95"},"uses_additional_data":false},{"leaderboard":"/sota/speech-recognition-on-gigaspeech-dev","task":"Speech Recognition","dataset":"GigaSpeech DEV","model":"Zipformer+pruned transducer\n(no external language model)","rank_in_archive_order":3,"of":5,"metrics":{"Word Error Rate (WER)":"10.09"},"uses_additional_data":false},{"leaderboard":"/sota/speech-recognition-on-gigaspeech-dev","task":"Speech Recognition","dataset":"GigaSpeech DEV","model":"Zipformer+CR-CTC\n(no external language model)","rank_in_archive_order":4,"of":5,"metrics":{"Word Error Rate (WER)":"10.15"},"uses_additional_data":false},{"leaderboard":"/sota/speech-recognition-on-gigaspeech-test","task":"Speech Recognition","dataset":"GigaSpeech TEST","model":"Zipformer+pruned transducer w/ CR-CTC\n(no external language model)","rank_in_archive_order":1,"of":5,"metrics":{"Word Error Rate (WER)":"10.03"},"uses_additional_data":false},{"leaderboard":"/sota/speech-recognition-on-gigaspeech-test","task":"Speech Recognition","dataset":"GigaSpeech TEST","model":"Zipformer+CR-CTC/AED\n(no external language model)","rank_in_archive_order":2,"of":5,"metrics":{"Word Error Rate (WER)":"10.07"},"uses_additional_data":false},{"leaderboard":"/sota/speech-recognition-on-gigaspeech-test","task":"Speech Recognition","dataset":"GigaSpeech TEST","model":"Zipformer+pruned transducer\n(no external language model)","rank_in_archive_order":3,"of":5,"metrics":{"Word Error Rate (WER)":"10.2"},"uses_additional_data":false},{"leaderboard":"/sota/speech-recognition-on-gigaspeech-test","task":"Speech Recognition","dataset":"GigaSpeech TEST","model":"Zipformer+CR-CTC\n(no external language model)","rank_in_archive_order":4,"of":5,"metrics":{"Word Error Rate (WER)":"10.28"},"uses_additional_data":false},{"leaderboard":"/sota/speech-recognition-on-librispeech-test-clean","task":"Speech Recognition","dataset":"LibriSpeech test-clean","model":"Zipformer+pruned transducer w/ CR-CTC (no  external language model)","rank_in_archive_order":16,"of":64,"metrics":{"Word Error Rate (WER)":"1.88"},"uses_additional_data":false},{"leaderboard":"/sota/speech-recognition-on-librispeech-test-clean","task":"Speech Recognition","dataset":"LibriSpeech test-clean","model":"Zipformer+CR-CTC (no external language model)","rank_in_archive_order":26,"of":64,"metrics":{"Word Error Rate (WER)":"2.02"},"uses_additional_data":false},{"leaderboard":"/sota/speech-recognition-on-librispeech-test-other","task":"Speech Recognition","dataset":"LibriSpeech test-other","model":"Zipformer+pruned transducer w/ CR-CTC\n(no external language model)","rank_in_archive_order":15,"of":53,"metrics":{"Word Error Rate (WER)":"3.95"},"uses_additional_data":false},{"leaderboard":"/sota/speech-recognition-on-librispeech-test-other","task":"Speech Recognition","dataset":"LibriSpeech test-other","model":"Zipformer+CR-CTC\n(no external language model)","rank_in_archive_order":24,"of":53,"metrics":{"Word Error Rate (WER)":"4.35"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2410.05101","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}