{"url":"/dataset/simmc2-0","name":"SIMMC2.0","full_name":null,"description_markdown":"Next generation task-oriented dialog systems need to understand conversational contexts with their perceived surroundings, to effectively help users in the real-world multimodal environment. Existing task-oriented dialog datasets aimed towards virtual assistance fall short and do not situate the dialog in the user's multimodal context. To overcome, we present a new dataset for Situated and Interactive Multimodal Conversations, SIMMC 2.0, which includes 11K task-oriented user<->assistant dialogs (117K utterances) in the shopping domain, grounded in immersive and photo-realistic scenes.\r\nThe dialogs are collected using a two-phase pipeline: (1) A novel multimodal dialog simulator generates simulated dialog flows, with an emphasis on diversity and richness of interactions, (2) Manual paraphrasing of the generated utterances to collect diverse referring expressions. We provide an in-depth analysis of the collected dataset, and describe in detail the four main benchmark tasks we propose. Our baseline model, powered by the state-of-the-art language model, shows promising results, and highlights new challenges and directions for the community to study.","description_withheld":null,"homepage":"https://github.com/facebookresearch/simmc2/tree/simmc2.0","introduced_date":"2021-04-18","introduced_date_note":null,"introduced_by":{"paper":"/paper/simmc-2-0-a-task-oriented-dialog-dataset-for","title":"SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal Conversations","first_author":"Satwik Kottur","url":null},"license":{"name":"https://github.com/facebookresearch/simmc2/tree/simmc2.0","url":"https://github.com/facebookresearch/simmc2/tree/simmc2.0"},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Dialogue State Tracking","url":"/task/dialogue-state-tracking","datasets_with_task":"/datasets/task/dialogue-state-tracking"},{"name":"Response Generation","url":"/task/response-generation","datasets_with_task":"/datasets/task/response-generation"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["SIMMC2.0"],"data_loaders":[],"num_papers_in_archive":13,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/dialogue-state-tracking-on-simmc2-0","task":"Dialogue State Tracking","dataset_variant":"SIMMC2.0","rows":5,"metrics":["Act F1","Slot F1"],"first_row_in_archive_order":{"model":"PaCE","paper":"/paper/pace-unified-multi-modal-dialogue-pre","metrics":{"Act F1":"97.1","Slot F1":"87.0"},"code_links":[{"title":"AlibabaResearch/DAMO-ConvAI","url":"https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/pace"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/response-generation-on-simmc2-0","task":"Response Generation","dataset_variant":"SIMMC2.0","rows":5,"metrics":["BLEU"],"first_row_in_archive_order":{"model":"PaCE","paper":"/paper/pace-unified-multi-modal-dialogue-pre","metrics":{"BLEU":"34.1"},"code_links":[{"title":"AlibabaResearch/DAMO-ConvAI","url":"https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/pace"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/pace-unified-multi-modal-dialogue-pre","title":"PaCE: Unified Multi-modal Dialogue Pre-training with Progressive and Compositional Experts","date":"2023-05-24","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":0,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/learning-to-embed-multi-modal-contexts-for-1","title":"Learning to Embed Multi-Modal Contexts for Situated Conversational Agents","date":"2022-07-01","rows_on_this_dataset":3,"code_links":0,"syntology":null},{"paper":"/paper/learning-to-embed-multi-modal-contexts-for","title":"Learning to Embed Multi-Modal Contexts for Situated Conversational Agents","date":"2022-01-16","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/multimodal-transformer-networks-for-end-to","title":"Multimodal Transformer Networks for End-to-End Video-Grounded Dialogue Systems","date":"2019-07-02","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":2,"samples_unverified":1,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/language-models-are-unsupervised-multitask","title":"Language Models are Unsupervised Multitask Learners","date":"2019-02-14","rows_on_this_dataset":2,"code_links":21,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":2,"samples_harvested":4,"samples_ran":2,"samples_unverified":2,"pointer_only_for_licence":1,"papers_with_no_sample_that_ran":1,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}