{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/pre-training-on-high-resource-speech","title":"Pre-training on high-resource speech recognition improves low-resource speech-to-text translation","arxiv_id":"1809.01431","date":"2018-09-05","proceeding":"NAACL 2019 6","authors":["Sameer Bansal","Herman Kamper","Karen Livescu","Adam Lopez","Sharon Goldwater"],"abstract":"We present a simple approach to improve direct speech-to-text translation\n(ST) when the source language is low-resource: we pre-train the model on a\nhigh-resource automatic speech recognition (ASR) task, and then fine-tune its\nparameters for ST. We demonstrate that our approach is effective by\npre-training on 300 hours of English ASR data to improve Spanish-English ST\nfrom 10.8 to 20.2 BLEU when only 20 hours of Spanish-English ST training data\nare available. Through an ablation study, we find that the pre-trained encoder\n(acoustic model) accounts for most of the improvement, despite the fact that\nthe shared language in these tasks is the target language text, not the source\nlanguage audio. Applying this insight, we show that pre-training on ASR helps\nST even when the ASR language differs from both source and target ST languages:\npre-training on French ASR also improves Spanish-English ST. Finally, we show\nthat the approach improves performance on a true low-resource task:\npre-training on a combination of English ASR and French ASR improves\nMboshi-French ST, where only 4 hours of data are available, from 3.5 to 7.1\nBLEU.","url_abs":"http://arxiv.org/abs/1809.01431v2","url_pdf":"http://arxiv.org/pdf/1809.01431v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"pre-training-on-high-resource-speech","repo_url":"https://github.com/0xSameer/ast","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-to-text","task_name":"Speech-to-Text"},{"task_slug":"speech-to-text-translation","task_name":"Speech-to-Text Translation"},{"task_slug":"translation","task_name":"Translation"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1809.01431","atlas_url":"https://app.syntology.ai/?focus=1809.01431","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}