{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/open-source-automatic-speech-recognition-for","title":"Open Source Automatic Speech Recognition for German","arxiv_id":"1807.10311","date":"2018-07-26","proceeding":null,"authors":["Benjamin Milde","Arne Köhn"],"abstract":"High quality Automatic Speech Recognition (ASR) is a prerequisite for\nspeech-based applications and research. While state-of-the-art ASR software is\nfreely available, the language dependent acoustic models are lacking for\nlanguages other than English, due to the limited amount of freely available\ntraining data. We train acoustic models for German with Kaldi on two datasets,\nwhich are both distributed under a Creative Commons license. The resulting\nmodel is freely redistributable, lowering the cost of entry for German ASR. The\nmodels are trained on a total of 412 hours of German read speech data and we\nachieve a relative word error reduction of 26% by adding data from the Spoken\nWikipedia Corpus to the previously best freely available German acoustic model\nrecipe and dataset. Our best model achieves a word error rate of 14.38 on the\nTuda-De test set. Due to the large amount of speakers and the diversity of\ntopics included in the training data, our model is robust against speaker\nvariation and topic shift.","url_abs":"http://arxiv.org/abs/1807.10311v1","url_pdf":"http://arxiv.org/pdf/1807.10311v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"open-source-automatic-speech-recognition-for","repo_url":"https://github.com/uhh-lt/kaldi-tuda-de","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"open-source-automatic-speech-recognition-for","repo_url":"https://github.com/tudarmstadt-lt/kaldi-tuda-de","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/speech-recognition-on-tuda","task":"Speech Recognition","dataset":"TUDA","model":"Kaldi","rank_in_archive_order":6,"of":9,"metrics":{"Test WER":"14.4%"},"uses_additional_data":true}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}