Datasets › VizWiz
VizWiz (VizWiz-VQA)
The VizWiz-VQA dataset originates from a natural visual question answering setting where blind people each took an image and recorded a spoken question about it, together with 10 crowdsourced answers per visual question. The proposed challenge addresses the following two tasks for this dataset: predict the answer to a visual question and (2) predict whether a visual question cannot be answered.
Source: https://vizwiz.org/tasks-and-datasets/vqa/ Image Source: https://vizwiz.org/tasks-and-datasets/vqa/
Benchmarks archive 2025-07-28
All 7 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Image Captioning | VizWiz 2020 test-dev | IBM Research AI CIDEr 80.67 | — | — | 50 | Compare |
| Visual Question Answering (VQA) | VizWiz 2020 VQA | PaLI overall 73.3 | PaLI: A Jointly-Scaled Multilingual Language-Image Model | google-research/big_vision | 16 | Compare |
| Image Captioning | VizWiz 2020 test | IBM Research AI CIDEr 81.04 | — | — | 13 | Compare |
| Visual Question Answering (VQA) | VizWiz 2018 | LXR955, No Ensemble overall 55.4 | LXMERT: Learning Cross-Modality Encoder Representations... | huggingface/transformers +8 | 10 | Compare |
| Visual Question Answering (VQA) | VizWiz 2020 Answerability | CLIP-Ensemble average_precision 84.13 | Less Is More: Linear Layers on CLIP Features as Powerful... | — | 6 | Compare |
| Visual Question Answering | VizWiz | Emu-I * Accuracy 38.1 | Emu: Generative Pretraining in Multimodality | baaivision/emu +1 | 1 | Compare |
| Visual Question Answering (VQA) | VizWiz 2018 Answerability | ensemble_two_best average_precision 82.78 | — | — | 1 | Compare |
Papers archive 2025-07-28
7 shown of 7 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 260. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization | 1 | 1 | 5 Feb 2024 | ran 3 of 5 samples (2 unverified; 5 pointer-only for licence) |
| Emu: Generative Pretraining in Multimodality | 2 | 1 | 11 Jul 2023 | ran 0 of 2 samples (2 unverified) |
| PaLI: A Jointly-Scaled Multilingual Language-Image Model | 1 | 1 | 14 Sep 2022 | ran 2 of 4 samples (2 unverified) |
| Less Is More: Linear Layers on CLIP Features as Powerful VizWiz Model | 0 | 4 | 10 Jun 2022 | not harvested |
| Decoupled Box Proposal and Featurization with Ultrafine-Grained Semantic Labels Improve Image Captioning and Visual Question Answering | 0 | 1 | 4 Sep 2019 | not harvested |
| LXMERT: Learning Cross-Modality Encoder Representations from Transformers | 9 | 1 | 20 Aug 2019 | ran 4 of 15 samples (11 unverified; 3 pointer-only for licence) |
| Towards VQA Models That Can Read | 7 | 1 | 18 Apr 2019 | not harvested |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- VizWiz Answer Differences 2019
- VizWiz 2020 VQA
- VizWiz 2020 test-dev
- VizWiz 2020 test
- VizWiz 2020 Answerability
- VizWiz 2018 Answerability
- VizWiz 2018
- VizWiz
8 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections