{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/dialogue-learning-with-human-in-the-loop","title":"Dialogue Learning With Human-In-The-Loop","arxiv_id":"1611.09823","date":"2016-11-29","proceeding":null,"authors":["Jiwei Li","Alexander H. Miller","Sumit Chopra","Marc'Aurelio Ranzato","Jason Weston"],"abstract":"An important aspect of developing conversational agents is to give a bot the\nability to improve through communicating with humans and to learn from the\nmistakes that it makes. Most research has focused on learning from fixed\ntraining sets of labeled data rather than interacting with a dialogue partner\nin an online fashion. In this paper we explore this direction in a\nreinforcement learning setting where the bot improves its question-answering\nability from feedback a teacher gives following its generated responses. We\nbuild a simulator that tests various aspects of such learning in a synthetic\nenvironment, and introduce models that work in this regime. Finally, real\nexperiments with Mechanical Turk validate the approach.","url_abs":"http://arxiv.org/abs/1611.09823v3","url_pdf":"http://arxiv.org/pdf/1611.09823v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"dialogue-learning-with-human-in-the-loop","repo_url":"https://github.com/facebook/MemNN","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"torch","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"dialogue-learning-with-human-in-the-loop","repo_url":"https://github.com/rohit129/Movie_KnowledgeGraph_QA","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1611.09823","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}