{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/end-to-end-optimization-of-goal-driven-and","title":"End-to-end optimization of goal-driven and visually grounded dialogue systems","arxiv_id":"1703.05423","date":"2017-03-15","proceeding":null,"authors":["Florian Strub","Harm de Vries","Jeremie Mary","Bilal Piot","Aaron Courville","Olivier Pietquin"],"abstract":"End-to-end design of dialogue systems has recently become a popular research\ntopic thanks to powerful tools such as encoder-decoder architectures for\nsequence-to-sequence learning. Yet, most current approaches cast human-machine\ndialogue management as a supervised learning problem, aiming at predicting the\nnext utterance of a participant given the full history of the dialogue. This\nvision is too simplistic to render the intrinsic planning problem inherent to\ndialogue as well as its grounded nature, making the context of a dialogue\nlarger than the sole history. This is why only chit-chat and question answering\ntasks have been addressed so far using end-to-end architectures. In this paper,\nwe introduce a Deep Reinforcement Learning method to optimize visually grounded\ntask-oriented dialogues, based on the policy gradient algorithm. This approach\nis tested on a dataset of 120k dialogues collected through Mechanical Turk and\nprovides encouraging results at solving both the problem of generating natural\ndialogues and the task of discovering a specific object in a complex picture.","url_abs":"http://arxiv.org/abs/1703.05423v1","url_pdf":"http://arxiv.org/pdf/1703.05423v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"end-to-end-optimization-of-goal-driven-and","repo_url":"https://github.com/GuessWhatGame/guesswhat","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"end-to-end-optimization-of-goal-driven-and","repo_url":"https://github.com/ibrahimSouleiman/GuessWhat","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"dialogue-management","task_name":"Dialogue Management"},{"task_slug":"management","task_name":"Management"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1703.05423","atlas_url":"https://app.syntology.ai/?focus=1703.05423","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}