{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/towards-building-large-scale-multimodal","title":"Towards Building Large Scale Multimodal Domain-Aware Conversation Systems","arxiv_id":"1704.00200","date":"2017-04-01","proceeding":null,"authors":["Amrita Saha","Mitesh Khapra","Karthik Sankaranarayanan"],"abstract":"While multimodal conversation agents are gaining importance in several\ndomains such as retail, travel etc., deep learning research in this area has\nbeen limited primarily due to the lack of availability of large-scale, open\nchatlogs. To overcome this bottleneck, in this paper we introduce the task of\nmultimodal, domain-aware conversations, and propose the MMD benchmark dataset.\nThis dataset was gathered by working in close coordination with large number of\ndomain experts in the retail domain. These experts suggested various\nconversations flows and dialog states which are typically seen in multimodal\nconversations in the fashion domain. Keeping these flows and states in mind, we\ncreated a dataset consisting of over 150K conversation sessions between\nshoppers and sales agents, with the help of in-house annotators using a\nsemi-automated manually intense iterative process. With this dataset, we\npropose 5 new sub-tasks for multimodal conversations along with their\nevaluation methodology. We also propose two multimodal neural models in the\nencode-attend-decode paradigm and demonstrate their performance on two of the\nsub-tasks, namely text response generation and best image response selection.\nThese experiments serve to establish baseline performance and open new research\ndirections for each of these sub-tasks. Further, for each of the sub-tasks, we\npresent a `per-state evaluation' of 9 most significant dialog states, which\nwould enable more focused research into understanding the challenges and\ncomplexities involved in each of these states.","url_abs":"http://arxiv.org/abs/1704.00200v3","url_pdf":"http://arxiv.org/pdf/1704.00200v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"towards-building-large-scale-multimodal","repo_url":"https://github.com/kkxkkx/tensorflow_caption","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"response-generation","task_name":"Response Generation"}],"methods":[],"datasets_introduced":[{"slug":"mmd","name":"MMD","full_name":"Multimodal Dialogs"}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}