{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mconv-an-environment-for-multimodal","title":"MConv: An Environment for Multimodal Conversational Search across Multiple Domains","arxiv_id":null,"date":"2021-07-11","proceeding":"SIGIR 2021 7","authors":["Lizi Liao","Le Hong Long","Zheng Zhang","Minlie Huang","Tat-Seng Chua"],"abstract":"Although conversational search has become a hot topic in both dialogue research and IR community, the real breakthrough has been limited by the scale and quality of datasets available. To address this fundamental obstacle, we introduce the Multimodal Multi-domain Conversational dataset (MMConv), a fully annotated collection of human-to-human role-playing dialogues spanning over multiple domains and tasks. The contribution is two-fold. First, beyond the task-oriented multimodal dialogues among user and agent pairs, dialogues are fully annotated with dialogue belief states and dialogue acts. More importantly, we create a relatively comprehensive environment for conducting multimodal conversational search with real user settings, structured venue database, annotated image repository as well as crowd-sourced knowledge database. A detailed description of the data collection procedure along with a summary of data structure and analysis is provided. Second, a set of benchmark results for dialogue state tracking, conversational recommendation, response generation as well as a unified model for multiple tasks are reported. We adopt the state-of-the-art methods for these tasks respectively to demonstrate the usability of the data, discuss limitations of current methods and set baselines for future studies.","url_abs":"https://dl.acm.org/doi/abs/10.1145/3404835.3462970","url_pdf":"https://dl.acm.org/doi/pdf/10.1145/3404835.3462970","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mconv-an-environment-for-multimodal","repo_url":"https://github.com/lizi-git/MMConv","is_official":0,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"conversational-recommendation","task_name":"Conversational Recommendation"},{"task_slug":"conversational-search","task_name":"Conversational Search"},{"task_slug":"dialogue-state-tracking","task_name":"Dialogue State Tracking"},{"task_slug":"response-generation","task_name":"Response Generation"}],"methods":[],"datasets_introduced":[{"slug":"mmconv","name":"MMConv","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/dialogue-state-tracking-on-mmconv","task":"Dialogue State Tracking","dataset":"MMConv","model":"DS-DST","rank_in_archive_order":2,"of":2,"metrics":{"Categorical Accuracy":"91.0","Non-Categorical Accuracy":"23.0","Overall":"18.0"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}