{"url":"/dataset/dme-vqa-dataset","name":"DME VQA dataset","full_name":"Diabetic Macular Edema VQA dataset","description_markdown":"Medical VQA dataset built from the [IDRiD](https://ieee-dataport.org/open-access/indian-diabetic-retinopathy-image-dataset-idrid) and [eOphta](https://www.adcis.net/en/third-party/e-ophtha/) datasets. The dataset contains both healthy and unhealthy fundus images.  For each image, a set of pre-defined questions is generated, including questions about regions (e.g. are there hard exudates in this region?), for which an associated mask denotes the location of the region. \r\n\r\nThe motivation for this dataset include the lack of public medical VQA datasets with related questions. In our dataset, questions are related because there is a high-level question about the DME grade of the image, and associated low-level questions that can lead to the answer of the high-level question. This allows to study the consistency of a VQA model i.e. how often the model produces contradictory answers to questions about a given image. Questions about regions are also a novel feature of this dataset.\r\n\r\nThe dataset can be used for general VQA purposes, and also for the more specific purpose of consistency improvement. \r\n\r\nNumber of images :\r\nTrain: 433\r\nVal: 112\r\nTest: 134\r\n\r\nNumber of QA pairs:\r\nTrain: 9779\r\nVal: 2380\r\nTest: 1311\r\n\r\nTo download the dataset, click [here](https://zenodo.org/record/6784358).\r\n\r\nFor more information, check [our paper](https://arxiv.org/abs/2206.13296).","description_withheld":null,"homepage":"https://zenodo.org/record/6784358","introduced_date":"2022-06-27","introduced_date_note":null,"introduced_by":{"paper":"/paper/consistency-preserving-visual-question","title":"Consistency-preserving Visual Question Answering in Medical Imaging","first_author":"Sergio Tascon-Morales","url":null},"license":{"name":"CC BY","url":"https://creativecommons.org/licenses/by/4.0/legalcode"},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Visual Question Answering (VQA)","url":"/task/visual-question-answering","datasets_with_task":"/datasets/task/visual-question-answering"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["DME VQA dataset"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}