Datasets › VisDial

VisDial (Visual Dialog)

Introduced by Abhishek Das et al. in Visual Dialog1 Jan 2017 archive 2025-07-28

Visual Dialog (VisDial) dataset contains human annotated questions based on images of MS COCO dataset. This dataset was developed by pairing two subjects on Amazon Mechanical Turk to chat about an image. One person was assigned the job of a ‘questioner’ and the other person acted as an ‘answerer’. The questioner sees only the text description of an image (i.e., an image caption from MS COCO dataset) and the original image remains hidden to the questioner. Their task is to ask questions about this hidden image to “imagine the scene better”. The answerer sees the image, caption and answers the questions asked by the questioner. The two of them can continue the conversation by asking and answering questions for 10 rounds at max.

VisDial v1.0 contains 123K dialogues on MS COCO (2017 training set) for training split, 2K dialogues with validation images for validation split and 8K dialogues on test set for test-standard set. The previously released v0.5 and v0.9 versions of VisDial dataset (corresponding to older splits of MS COCO) are considered deprecated.

Source: Granular Multimodal Attention Networks for Visual Dialog Image Source: https://arxiv.org/pdf/1611.08669.pdf

Benchmarks archive 2025-07-28

All 6 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Visual Dialog Visual Dialog v1.0 test-std Single NDCG (x 100) 78.7 — — 80 Compare
Visual Dialog VisDial v0.9 val 9xFGA (VGG) MRR 68.92 Factor Graph Attention idansc/fga 18 Compare
Chat-based Image Retrieval VisDial ChatGPT & BLIP2 Recall@10 on 1 rounds 70 — — 3 Compare
Visual Dialog VisDial v1.0 test-std 5xFGA + LS*+ MRR 0.7124 Ensemble of MRR and NDCG models for Visual Dialog idansc/mrr-ndcg 3 Compare
Common Sense Reasoning Visual Dialog v0.9 NMN [kottur2018visual] 1 in 10 R@5 80.1 Visual Coreference Resolution in Visual Dialog using... facebookresearch/corefnmn 1 Compare
Common Sense Reasoning Visual Dialog v0.9 PDUN 1 in 10 R@5 81.0 Probabilistic framework for solving Visual Dialog — 1 Compare

Papers archive 2025-07-28

21 shown of 21 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 159. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
ImageScope: Unifying Language-Guided Image Retrieval via Large Multimodal Model Collective Reasoning 1 1 13 Mar 2025 not harvested
Ensemble of MRR and NDCG models for Visual Dialog 1 4 15 Apr 2021 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
Multi-View Attention Network for Visual Dialog 1 2 29 Apr 2020 not harvested
Iterative Context-Aware Graph Inference for Visual Dialog 1 2 5 Apr 2020 not harvested
Efficient Attention Mechanism for Visual Dialog that can Handle All the Interactions between Multiple Inputs 1 1 26 Nov 2019 not harvested
DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual Dialogue 1 2 17 Nov 2019 not harvested
Probabilistic framework for solving Visual Dialog 0 1 11 Sep 2019 not harvested
Reasoning Visual Dialogs with Structural and Partial Observations 1 2 11 Apr 2019 not harvested
Factor Graph Attention 1 2 11 Apr 2019 not harvested
Image-Question-Answer Synergistic Network for Visual Dialog 0 1 26 Feb 2019 not harvested
Making History Matter: History-Advantage Sequence Training for Visual Dialog 0 2 25 Feb 2019 not harvested
Dual Attention Networks for Visual Reference Resolution in Visual Dialog 2 2 25 Feb 2019 ran 1 of 7 samples (6 unverified)
Recursive Visual Attention in Visual Dialog 1 2 6 Dec 2018 not harvested
Visual Coreference Resolution in Visual Dialog using Neural Module Networks 1 4 6 Sep 2018 not harvested
Two can play this Game: Visual Dialog with Discriminative Question Generation and Answering 0 1 29 Mar 2018 not harvested
Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning 0 1 21 Nov 2017 not harvested
Visual Reference Resolution using Attention Memory for Visual Dialog 0 1 23 Sep 2017 not harvested
Best of Both Worlds: Transferring Knowledge from Discriminative Learning to a Generative Visual Dialog Model 1 1 5 Jun 2017 not harvested
Learning to Reason: End-to-End Module Networks for Visual Question Answering 1 1 18 Apr 2017 not harvested
Visual Dialog 11 6 26 Nov 2016 ran 2 of 2 samples (0 unverified)
Hierarchical Question-Image Co-Attention for Visual Question Answering 9 1 31 May 2016 ran 1 of 7 samples (6 unverified; 1 pointer-only for licence)

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Creative Commons Attribution 4.0 International License

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Visual Dialog v0.9
  • Visual Dialog v1.0
  • VisDial v1.0 test-std
  • VisDial v0.9 val
  • Visual Dialog v1.0 test-std
  • Visual Dialog v0.9
  • VisDial

7 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections