Datasets › VisDial
VisDial (Visual Dialog)
Visual Dialog (VisDial) dataset contains human annotated questions based on images of MS COCO dataset. This dataset was developed by pairing two subjects on Amazon Mechanical Turk to chat about an image. One person was assigned the job of a ‘questioner’ and the other person acted as an ‘answerer’. The questioner sees only the text description of an image (i.e., an image caption from MS COCO dataset) and the original image remains hidden to the questioner. Their task is to ask questions about this hidden image to “imagine the scene better”. The answerer sees the image, caption and answers the questions asked by the questioner. The two of them can continue the conversation by asking and answering questions for 10 rounds at max.
VisDial v1.0 contains 123K dialogues on MS COCO (2017 training set) for training split, 2K dialogues with validation images for validation split and 8K dialogues on test set for test-standard set. The previously released v0.5 and v0.9 versions of VisDial dataset (corresponding to older splits of MS COCO) are considered deprecated.
Source: Granular Multimodal Attention Networks for Visual Dialog Image Source: https://arxiv.org/pdf/1611.08669.pdf
Benchmarks archive 2025-07-28
All 6 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Visual Dialog | Visual Dialog v1.0 test-std | Single NDCG (x 100) 78.7 | — | — | 80 | Compare |
| Visual Dialog | VisDial v0.9 val | 9xFGA (VGG) MRR 68.92 | Factor Graph Attention | idansc/fga | 18 | Compare |
| Chat-based Image Retrieval | VisDial | ChatGPT & BLIP2 Recall@10 on 1 rounds 70 | — | — | 3 | Compare |
| Visual Dialog | VisDial v1.0 test-std | 5xFGA + LS*+ MRR 0.7124 | Ensemble of MRR and NDCG models for Visual Dialog | idansc/mrr-ndcg | 3 | Compare |
| Common Sense Reasoning | Visual Dialog v0.9 | NMN [kottur2018visual] 1 in 10 R@5 80.1 | Visual Coreference Resolution in Visual Dialog using... | facebookresearch/corefnmn | 1 | Compare |
| Common Sense Reasoning | Visual Dialog v0.9 | PDUN 1 in 10 R@5 81.0 | Probabilistic framework for solving Visual Dialog | — | 1 | Compare |
Papers archive 2025-07-28
21 shown of 21 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 159. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
Dataset loaders archive 2025-07-28
1 loader as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
Creative Commons Attribution 4.0 International License
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- Visual Dialog v0.9
- Visual Dialog v1.0
- VisDial v1.0 test-std
- VisDial v0.9 val
- Visual Dialog v1.0 test-std
- Visual Dialog v0.9
- VisDial
7 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections