{"url":"/dataset/retvqa","name":"RetVQA","full_name":"Retrieval-Based Visual Question Answering","description_markdown":"The RetVQA dataset is a large-scale dataset designed for Retrieval-Based Visual Question Answering (RetVQA). RetVQA is a more challenging task than traditional VQA, as it requires models to retrieve relevant images from a pool of images before answering a question. The need for RetVQA stems from the fact that information needed to answer a question may be spread across multiple images. \r\n\r\nHere is a detailed summary of the RetVQA dataset:\r\n\r\n* It is **20 times larger** than the closest dataset in this setting, WebQA.\r\n* It was derived from the Visual Genome dataset, utilising its questions and annotations of images.\r\n* It has **418K unique questions** and **16,205 unique precise answers**.\r\n*  The questions are designed to be **metadata-independent**, meaning they do not rely on information such as captions or tags.\r\n* The questions are divided into five categories: \r\n    * **color** \r\n    * **shape**\r\n    * **count**\r\n    * **object-attributes**\r\n    * **relation-based**.\r\n* The dataset includes both **binary (yes/no)** questions and **open-ended questions** that require a generative answer.\r\n* All answers are free-form and fluent, even for binary questions. For example, a binary question may be \"Do the rose and sunflower share the same colour?\", and a corresponding answer would be \"No, the rose and sunflower do not share the same colour\".\r\n*  Every question in RetVQA requires reasoning over **multiple images** to arrive at the answer. This contrasts with datasets like WebQA, where a majority of questions can be answered using a single image.\r\n* The dataset has, on average, **two relevant images and 24.5 irrelevant images per question**. This makes it more challenging than datasets like ISVQA, where images are homogeneous and no explicit retrieval is needed.","description_withheld":null,"homepage":"https://vl2g.github.io/projects/retvqa/","introduced_date":"2023-06-29","introduced_date_note":null,"introduced_by":{"paper":"/paper/answer-mining-from-a-pool-of-images-towards","title":"Answer Mining from a Pool of Images: Towards Retrieval-Based Visual Question Answering","first_author":"Abhirama Subramanyam Penamakuri","url":null},"license":{"name":"MIT License","url":"https://github.com/Abhiram4572/mi_bart?tab=MIT-1-ov-file"},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Visual Question Answering (VQA)","url":"/task/visual-question-answering","datasets_with_task":"/datasets/task/visual-question-answering"},{"name":"Image Retrieval","url":"/task/image-retrieval","datasets_with_task":"/datasets/task/image-retrieval"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["RetVQA"],"data_loaders":[],"num_papers_in_archive":4,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/visual-question-answering-vqa-on-retvqa","task":"Visual Question Answering (VQA)","dataset_variant":"RetVQA","rows":1,"metrics":["Accuarcy","Accuracy * Fluency"],"first_row_in_archive_order":{"model":"MI-BART","paper":"/paper/answer-mining-from-a-pool-of-images-towards","metrics":{"Accuarcy":"76.5","Accuracy * Fluency":"70.9"},"code_links":[{"title":"Abhiram4572/mi_bart","url":"https://github.com/Abhiram4572/mi_bart"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/answer-mining-from-a-pool-of-images-towards","title":"Answer Mining from a Pool of Images: Towards Retrieval-Based Visual Question Answering","date":"2023-06-29","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":7,"samples_ran":6,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":1,"samples_harvested":7,"samples_ran":6,"samples_unverified":1,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}