Datasets › RetVQA
RetVQA (Retrieval-Based Visual Question Answering)
The RetVQA dataset is a large-scale dataset designed for Retrieval-Based Visual Question Answering (RetVQA). RetVQA is a more challenging task than traditional VQA, as it requires models to retrieve relevant images from a pool of images before answering a question. The need for RetVQA stems from the fact that information needed to answer a question may be spread across multiple images.
Here is a detailed summary of the RetVQA dataset:
- It is 20 times larger than the closest dataset in this setting, WebQA.
- It was derived from the Visual Genome dataset, utilising its questions and annotations of images.
- It has 418K unique questions and 16,205 unique precise answers.
- The questions are designed to be metadata-independent, meaning they do not rely on information such as captions or tags.
- The questions are divided into five categories:
- color
- shape
- count
- object-attributes
- relation-based.
- The dataset includes both binary (yes/no) questions and open-ended questions that require a generative answer.
- All answers are free-form and fluent, even for binary questions. For example, a binary question may be "Do the rose and sunflower share the same colour?", and a corresponding answer would be "No, the rose and sunflower do not share the same colour".
- Every question in RetVQA requires reasoning over multiple images to arrive at the answer. This contrasts with datasets like WebQA, where a majority of questions can be answered using a single image.
- The dataset has, on average, two relevant images and 24.5 irrelevant images per question. This makes it more challenging than datasets like ISVQA, where images are homogeneous and no explicit retrieval is needed.
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Visual Question Answering (VQA) | RetVQA | MI-BART Accuarcy 76.5 | Answer Mining from a Pool of Images: Towards... | Abhiram4572/mi_bart | 1 | Compare |
Papers archive 2025-07-28
1 shown of 1 paper with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 4. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Answer Mining from a Pool of Images: Towards Retrieval-Based Visual Question Answering | 1 | 1 | 29 Jun 2023 | ran 6 of 7 samples (1 unverified) |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- RetVQA
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections