{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multimodal-grounding-for-sequence-to-sequence","title":"Multimodal Grounding for Sequence-to-Sequence Speech Recognition","arxiv_id":"1811.03865","date":"2018-11-09","proceeding":null,"authors":["Ozan Caglayan","Ramon Sanabria","Shruti Palaskar","Loïc Barrault","Florian Metze"],"abstract":"Humans are capable of processing speech by making use of multiple sensory\nmodalities. For example, the environment where a conversation takes place\ngenerally provides semantic and/or acoustic context that helps us to resolve\nambiguities or to recall named entities. Motivated by this, there have been\nmany works studying the integration of visual information into the speech\nrecognition pipeline. Specifically, in our previous work, we propose a\nmultistep visual adaptive training approach which improves the accuracy of an\naudio-based Automatic Speech Recognition (ASR) system. This approach, however,\nis not end-to-end as it requires fine-tuning the whole model with an adaptation\nlayer. In this paper, we propose novel end-to-end multimodal ASR systems and\ncompare them to the adaptive approach by using a range of visual\nrepresentations obtained from state-of-the-art convolutional neural networks.\nWe show that adaptive training is effective for S2S models leading to an\nabsolute improvement of 1.4% in word error rate. As for the end-to-end systems,\nalthough they perform better than baseline, the improvements are slightly less\nthan adaptive training, 0.8 absolute WER reduction in single-best models. Using\nensemble decoding, end-to-end models reach a WER of 15% which is the lowest\nscore among all systems.","url_abs":"http://arxiv.org/abs/1811.03865v2","url_pdf":"http://arxiv.org/pdf/1811.03865v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multimodal-grounding-for-sequence-to-sequence","repo_url":"https://github.com/srvk/how2-dataset","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"sequence-to-sequence-speech-recognition","task_name":"Sequence-To-Sequence Speech Recognition"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1811.03865","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}