{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multilingual-bottleneck-features-for-subword","title":"Multilingual bottleneck features for subword modeling in zero-resource languages","arxiv_id":"1803.08863","date":"2018-03-23","proceeding":null,"authors":["Enno Hermann","Sharon Goldwater"],"abstract":"How can we effectively develop speech technology for languages where no\ntranscribed data is available? Many existing approaches use no annotated\nresources at all, yet it makes sense to leverage information from large\nannotated corpora in other languages, for example in the form of multilingual\nbottleneck features (BNFs) obtained from a supervised speech recognition\nsystem. In this work, we evaluate the benefits of BNFs for subword modeling\n(feature extraction) in six unseen languages on a word discrimination task.\nFirst we establish a strong unsupervised baseline by combining two existing\nmethods: vocal tract length normalisation (VTLN) and the correspondence\nautoencoder (cAE). We then show that BNFs trained on a single language already\nbeat this baseline; including up to 10 languages results in additional\nimprovements which cannot be matched by just adding more data from a single\nlanguage. Finally, we show that the cAE can improve further on the BNFs if\nhigh-quality same-word pairs are available.","url_abs":"http://arxiv.org/abs/1803.08863v2","url_pdf":"http://arxiv.org/pdf/1803.08863v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multilingual-bottleneck-features-for-subword","repo_url":"https://github.com/eginhard/cae-utd-utils","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1803.08863","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}