{"url":"/dataset/visdial","name":"VisDial","full_name":"Visual Dialog","description_markdown":"**Visual Dialog** (**VisDial**) dataset contains human annotated questions based on images of MS COCO dataset. This dataset was developed by pairing two subjects on Amazon Mechanical Turk to chat about an image. One person was assigned the job of a ‘questioner’ and the other person acted as an ‘answerer’. The questioner sees only the text description of an image (i.e., an image caption from MS COCO dataset) and the original image remains hidden to the questioner. Their task is to ask questions about this hidden image to “imagine the scene better”. The answerer sees the image, caption and answers the questions asked by the questioner. The two of them can continue the conversation by asking and answering questions for 10 rounds at max.\r\n\r\n**VisDial v1.0** contains 123K dialogues on MS COCO (2017 training set) for training split, 2K dialogues with validation images for validation split and 8K dialogues on test set for test-standard set. The previously released v0.5 and v0.9 versions of VisDial dataset (corresponding to older splits of MS COCO) are considered deprecated.\r\n\r\nSource: [Granular Multimodal Attention Networks for Visual Dialog](https://arxiv.org/abs/1910.05728)\r\nImage Source: [https://arxiv.org/pdf/1611.08669.pdf](https://arxiv.org/pdf/1611.08669.pdf)","description_withheld":null,"homepage":"https://visualdialog.org/","introduced_date":"2017-01-01","introduced_date_note":null,"introduced_by":{"paper":"/paper/visual-dialog","title":"Visual Dialog","first_author":"Abhishek Das","url":null},"license":{"name":"Creative Commons Attribution 4.0 International License","url":"https://creativecommons.org/licenses/by/4.0/"},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"},{"name":"Dialog","url":"/datasets/modality/dialog"}],"tasks":[{"name":"Question Answering","url":"/task/question-answering","datasets_with_task":"/datasets/task/question-answering"},{"name":"Visual Question Answering (VQA)","url":"/task/visual-question-answering","datasets_with_task":"/datasets/task/visual-question-answering"},{"name":"Common Sense Reasoning","url":"/task/common-sense-reasoning","datasets_with_task":"/datasets/task/common-sense-reasoning"},{"name":"Visual Dialog","url":"/task/visual-dialogue","datasets_with_task":"/datasets/task/visual-dialogue"},{"name":"Chat-based Image Retrieval","url":"/task/chat-based-image-retrieval","datasets_with_task":"/datasets/task/chat-based-image-retrieval"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Visual Dialog  v0.9","Visual Dialog v1.0","VisDial v1.0 test-std","VisDial v0.9 val","Visual Dialog v1.0 test-std","Visual Dialog v0.9","VisDial"],"data_loaders":[{"repo":"https://github.com/facebookresearch/ParlAI","url":"https://parl.ai/docs/tasks.html#visdial","frameworks":["pytorch"]}],"num_papers_in_archive":159,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/visual-dialog-on-visual-dialog-v1-0-test-std","task":"Visual Dialog","dataset_variant":"Visual Dialog v1.0 test-std","rows":80,"metrics":["NDCG (x 100)","MRR (x 100)","R@1","R@5","R@10","Mean"],"first_row_in_archive_order":{"model":"Single","paper":null,"metrics":{"MRR (x 100)":"45.75","Mean":"6.54","NDCG (x 100)":"78.7","R@1":"29.5","R@10":"82.45","R@5":"65.7"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/visual-dialog-on-visdial-v09-val","task":"Visual Dialog","dataset_variant":"VisDial v0.9 val","rows":18,"metrics":["MRR","R@1","R@10","R@5","Mean Rank"],"first_row_in_archive_order":{"model":"9xFGA (VGG)","paper":"/paper/factor-graph-attention","metrics":{"MRR":"68.92","Mean Rank":"3.39","R@1":"55.16","R@10":"92.95","R@5":"86.26"},"code_links":[{"title":"idansc/fga","url":"https://github.com/idansc/fga"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/chat-based-image-retrieval-on-visdial","task":"Chat-based Image Retrieval","dataset_variant":"VisDial","rows":3,"metrics":["Recall@10 on 1 rounds","Recall@10 on 2 rounds","Recall@10 on 3 rounds","Hits@10 on 10 Round"],"first_row_in_archive_order":{"model":"ChatGPT & BLIP2","paper":null,"metrics":{"Recall@10 on 1 rounds":"70","Recall@10 on 2 rounds":"73.5","Recall@10 on 3 rounds":"75.75"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/visual-dialog-on-visdial-v10-test-std","task":"Visual Dialog","dataset_variant":"VisDial v1.0 test-std","rows":3,"metrics":["MRR","Mean Rank","NDCG","R@1","R@10","R@5"],"first_row_in_archive_order":{"model":"5xFGA + LS*+","paper":"/paper/ensemble-of-mrr-and-ndcg-models-for-visual","metrics":{"MRR":"0.7124","Mean Rank":"2.96","R@1":"58.28","R@10":"94.45","R@5":"87.55"},"code_links":[{"title":"idansc/mrr-ndcg","url":"https://github.com/idansc/mrr-ndcg"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/common-sense-reasoning-on-visual-dialog-v0-9","task":"Common Sense Reasoning","dataset_variant":"Visual Dialog v0.9","rows":1,"metrics":["1 in 10 R@5"],"first_row_in_archive_order":{"model":"NMN [kottur2018visual]","paper":"/paper/visual-coreference-resolution-in-visual","metrics":{"1 in 10 R@5":"80.1"},"code_links":[{"title":"facebookresearch/corefnmn","url":"https://github.com/facebookresearch/corefnmn"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/common-sense-reasoning-on-visual-dialog-v09","task":"Common Sense Reasoning","dataset_variant":"Visual Dialog  v0.9","rows":1,"metrics":["1 in 10 R@5","Recall@10"],"first_row_in_archive_order":{"model":"PDUN","paper":"/paper/probabilistic-framework-for-solving-visual","metrics":{"1 in 10 R@5":"81.0","Recall@10":"90.5"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/imagescope-unifying-language-guided-image-1","title":"ImageScope: Unifying Language-Guided Image Retrieval via Large Multimodal Model Collective Reasoning","date":"2025-03-13","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/ensemble-of-mrr-and-ndcg-models-for-visual","title":"Ensemble of MRR and NDCG models for Visual Dialog","date":"2021-04-15","rows_on_this_dataset":4,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":2,"samples_ran":2,"samples_unverified":0,"pointer_only_for_licence":2,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/multi-view-attention-networks-for-visual","title":"Multi-View Attention Network for Visual Dialog","date":"2020-04-29","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/iterative-context-aware-graph-inference-for","title":"Iterative Context-Aware Graph Inference for Visual Dialog","date":"2020-04-05","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/efficient-attention-mechanism-for-handling","title":"Efficient Attention Mechanism for Visual Dialog that can Handle All the Interactions between Multiple Inputs","date":"2019-11-26","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/dualvd-an-adaptive-dual-encoding-model-for","title":"DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual Dialogue","date":"2019-11-17","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/probabilistic-framework-for-solving-visual","title":"Probabilistic framework for solving Visual Dialog","date":"2019-09-11","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/reasoning-visual-dialogs-with-structural-and","title":"Reasoning Visual Dialogs with Structural and Partial Observations","date":"2019-04-11","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/factor-graph-attention","title":"Factor Graph Attention","date":"2019-04-11","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/image-question-answer-synergistic-network-for","title":"Image-Question-Answer Synergistic Network for Visual Dialog","date":"2019-02-26","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/making-history-matter-gold-critic-sequence","title":"Making History Matter: History-Advantage Sequence Training for Visual Dialog","date":"2019-02-25","rows_on_this_dataset":2,"code_links":0,"syntology":null},{"paper":"/paper/dual-attention-networks-for-visual-reference","title":"Dual Attention Networks for Visual Reference Resolution in Visual Dialog","date":"2019-02-25","rows_on_this_dataset":2,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":7,"samples_ran":1,"samples_unverified":6,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/recursive-visual-attention-in-visual-dialog","title":"Recursive Visual Attention in Visual Dialog","date":"2018-12-06","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/visual-coreference-resolution-in-visual","title":"Visual Coreference Resolution in Visual Dialog using Neural Module Networks","date":"2018-09-06","rows_on_this_dataset":4,"code_links":1,"syntology":null},{"paper":"/paper/two-can-play-this-game-visual-dialog-with","title":"Two can play this Game: Visual Dialog with Discriminative Question Generation and Answering","date":"2018-03-29","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/are-you-talking-to-me-reasoned-visual-dialog","title":"Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning","date":"2017-11-21","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/visual-reference-resolution-using-attention","title":"Visual Reference Resolution using Attention Memory for Visual Dialog","date":"2017-09-23","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/best-of-both-worlds-transferring-knowledge","title":"Best of Both Worlds: Transferring Knowledge from Discriminative Learning to a Generative Visual Dialog Model","date":"2017-06-05","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/learning-to-reason-end-to-end-module-networks","title":"Learning to Reason: End-to-End Module Networks for Visual Question Answering","date":"2017-04-18","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/visual-dialog","title":"Visual Dialog","date":"2016-11-26","rows_on_this_dataset":6,"code_links":11,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":2,"samples_ran":2,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/hierarchical-question-image-co-attention-for","title":"Hierarchical Question-Image Co-Attention for Visual Question Answering","date":"2016-05-31","rows_on_this_dataset":1,"code_links":9,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":7,"samples_ran":1,"samples_unverified":6,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":4,"samples_harvested":18,"samples_ran":6,"samples_unverified":12,"pointer_only_for_licence":3,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}