Datasets › VizWiz

VizWiz (VizWiz-VQA)

Introduced by Danna Gurari et al. in VizWiz Grand Challenge: Answering Visual Questions from Blind People1 Jan 2018 archive 2025-07-28

The VizWiz-VQA dataset originates from a natural visual question answering setting where blind people each took an image and recorded a spoken question about it, together with 10 crowdsourced answers per visual question. The proposed challenge addresses the following two tasks for this dataset: predict the answer to a visual question and (2) predict whether a visual question cannot be answered.

Source: https://vizwiz.org/tasks-and-datasets/vqa/ Image Source: https://vizwiz.org/tasks-and-datasets/vqa/

Benchmarks archive 2025-07-28

All 7 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Image Captioning VizWiz 2020 test-dev IBM Research AI CIDEr 80.67 — — 50 Compare
Visual Question Answering (VQA) VizWiz 2020 VQA PaLI overall 73.3 PaLI: A Jointly-Scaled Multilingual Language-Image Model google-research/big_vision 16 Compare
Image Captioning VizWiz 2020 test IBM Research AI CIDEr 81.04 — — 13 Compare
Visual Question Answering (VQA) VizWiz 2018 LXR955, No Ensemble overall 55.4 LXMERT: Learning Cross-Modality Encoder Representations... huggingface/transformers +8 10 Compare
Visual Question Answering (VQA) VizWiz 2020 Answerability CLIP-Ensemble average_precision 84.13 Less Is More: Linear Layers on CLIP Features as Powerful... — 6 Compare
Visual Question Answering VizWiz Emu-I * Accuracy 38.1 Emu: Generative Pretraining in Multimodality baaivision/emu +1 1 Compare
Visual Question Answering (VQA) VizWiz 2018 Answerability ensemble_two_best average_precision 82.78 — — 1 Compare

Papers archive 2025-07-28

7 shown of 7 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 260. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization 1 1 5 Feb 2024 ran 3 of 5 samples (2 unverified; 5 pointer-only for licence)
Emu: Generative Pretraining in Multimodality 2 1 11 Jul 2023 ran 0 of 2 samples (2 unverified)
PaLI: A Jointly-Scaled Multilingual Language-Image Model 1 1 14 Sep 2022 ran 2 of 4 samples (2 unverified)
Less Is More: Linear Layers on CLIP Features as Powerful VizWiz Model 0 4 10 Jun 2022 not harvested
Decoupled Box Proposal and Featurization with Ultrafine-Grained Semantic Labels Improve Image Captioning and Visual Question Answering 0 1 4 Sep 2019 not harvested
LXMERT: Learning Cross-Modality Encoder Representations from Transformers 9 1 20 Aug 2019 ran 4 of 15 samples (11 unverified; 3 pointer-only for licence)
Towards VQA Models That Can Read 7 1 18 Apr 2019 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • VizWiz Answer Differences 2019
  • VizWiz 2020 VQA
  • VizWiz 2020 test-dev
  • VizWiz 2020 test
  • VizWiz 2020 Answerability
  • VizWiz 2018 Answerability
  • VizWiz 2018
  • VizWiz

8 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections