{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/analyzing-hidden-representations-in-end-to","title":"Analyzing Hidden Representations in End-to-End Automatic Speech Recognition Systems","arxiv_id":"1709.04482","date":"2017-09-13","proceeding":"NeurIPS 2017 12","authors":["Yonatan Belinkov","James Glass"],"abstract":"Neural models have become ubiquitous in automatic speech recognition systems.\nWhile neural networks are typically used as acoustic models in more complex\nsystems, recent studies have explored end-to-end speech recognition systems\nbased on neural networks, which can be trained to directly predict text from\ninput acoustic features. Although such systems are conceptually elegant and\nsimpler than traditional systems, it is less obvious how to interpret the\ntrained models. In this work, we analyze the speech representations learned by\na deep end-to-end model that is based on convolutional and recurrent layers,\nand trained with a connectionist temporal classification (CTC) loss. We use a\npre-trained model to generate frame-level features which are given to a\nclassifier that is trained on frame classification into phones. We evaluate\nrepresentations from different layers of the deep model and compare their\nquality for predicting phone labels. Our experiments shed light on important\naspects of the end-to-end model such as layer depth, model complexity, and\nother design choices.","url_abs":"http://arxiv.org/abs/1709.04482v1","url_pdf":"http://arxiv.org/pdf/1709.04482v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"analyzing-hidden-representations-in-end-to","repo_url":"https://github.com/boknilev/asr-repr-analysis","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"torch","reach":null}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1709.04482","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}