{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-learning-based-multi-modal-addressee","title":"Deep Learning Based Multi-modal Addressee Recognition in Visual Scenes with Utterances","arxiv_id":"1809.04288","date":"2018-09-12","proceeding":null,"authors":["Thao Minh Le","Nobuyuki Shimizu","Takashi Miyazaki","Koichi Shinoda"],"abstract":"With the widespread use of intelligent systems, such as smart speakers,\naddressee recognition has become a concern in human-computer interaction, as\nmore and more people expect such systems to understand complicated social\nscenes, including those outdoors, in cafeterias, and hospitals. Because\nprevious studies typically focused only on pre-specified tasks with limited\nconversational situations such as controlling smart homes, we created a mock\ndataset called Addressee Recognition in Visual Scenes with Utterances (ARVSU)\nthat contains a vast body of image variations in visual scenes with an\nannotated utterance and a corresponding addressee for each scenario. We also\npropose a multi-modal deep-learning-based model that takes different human\ncues, specifically eye gazes and transcripts of an utterance corpus, into\naccount to predict the conversational addressee from a specific speaker's view\nin various real-life conversational scenarios. To the best of our knowledge, we\nare the first to introduce an end-to-end deep learning model that combines\nvision and transcripts of utterance for addressee recognition. As a result, our\nstudy suggests that future addressee recognition can reach the ability to\nunderstand human intention in many social situations previously unexplored, and\nour modality dataset is a first step in promoting research in this field.","url_abs":"http://arxiv.org/abs/1809.04288v1","url_pdf":"http://arxiv.org/pdf/1809.04288v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[{"slug":"arvsu","name":"ARVSU","full_name":"Addressee Recognition in Visual Scenes with Utterances"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}