{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/jointly-learning-to-see-ask-and-guesswhat","title":"Beyond task success: A closer look at jointly learning to see, ask, and GuessWhat","arxiv_id":"1809.03408","date":"2018-09-10","proceeding":"NAACL 2019 6","authors":["Ravi Shekhar","Aashish Venkatesh","Tim Baumgärtner","Elia Bruni","Barbara Plank","Raffaella Bernardi","Raquel Fernández"],"abstract":"We propose a grounded dialogue state encoder which addresses a foundational\nissue on how to integrate visual grounding with dialogue system components. As\na test-bed, we focus on the GuessWhat?! game, a two-player game where the goal\nis to identify an object in a complex visual scene by asking a sequence of\nyes/no questions. Our visually-grounded encoder leverages synergies between\nguessing and asking questions, as it is trained jointly using multi-task\nlearning. We further enrich our model via a cooperative learning regime. We\nshow that the introduction of both the joint architecture and cooperative\nlearning lead to accuracy improvements over the baseline system. We compare our\napproach to an alternative system which extends the baseline with reinforcement\nlearning. Our in-depth analysis shows that the linguistic skills of the two\nmodels differ dramatically, despite approaching comparable performance levels.\nThis points at the importance of analyzing the linguistic output of competing\nsystems beyond numeric comparison solely based on task success.","url_abs":"http://arxiv.org/abs/1809.03408v2","url_pdf":"http://arxiv.org/pdf/1809.03408v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"jointly-learning-to-see-ask-and-guesswhat","repo_url":"https://github.com/anthonysicilia/leather-aacl2022","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"jointly-learning-to-see-ask-and-guesswhat","repo_url":"https://github.com/mmazuecos/reinact2021-impact-of-answers","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"jointly-learning-to-see-ask-and-guesswhat","repo_url":"https://github.com/shekharRavi/Beyond-Task-Success-NAACL2019","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"multi-task-learning","task_name":"Multi-Task Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"visual-grounding","task_name":"Visual Grounding"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1809.03408","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}