Datasets › OCR-VQA
OCR-VQA
The OCR-VQA dataset is a valuable resource for research in the field of Visual Question Answering (VQA). Let me provide you with some details about it:
-
Dataset Overview:
- The OCR-VQA dataset contains a total of 207,572 images along with their associated question-answer pairs.
- These images are related to document content and are accompanied by their corresponding OCR transcriptions¹².
-
Purpose and Significance:
- Visual Question Answering (VQA) tasks require models to reason jointly over visual information (such as images) and natural language inputs (such as questions).
- By using this dataset, researchers can develop and evaluate AI models that can effectively understand and answer questions based on visual content and textual context.
-
Other Related VQA Datasets:
- Apart from OCR-VQA, there are other VQA datasets available for research and benchmarking:
- ScreenQA: Focused on questions related to screen content.
- MP-DocVQA: A dataset for document-based VQA.
- ChartQA: Specifically designed for answering questions about charts.
- InfographicVQA: For handling questions related to infographics.
- Apart from OCR-VQA, there are other VQA datasets available for research and benchmarking:
Source: Conversation with Bing, 3/15/2024 (1) OCR-VQA Dataset | Papers With Code. https://paperswithcode.com/dataset/ocr-vqa. (2) GitHub - anisha2102/docvqa: Document Visual Question Answering. https://github.com/anisha2102/docvqa. (3) VQA: Visual Question Answering. https://visualqa.org/. (4) allenai/aokvqa: Official repository for the A-OKVQA dataset - GitHub. https://github.com/allenai/aokvqa.
Benchmarks archive 2025-07-28
No leaderboard in the archive resolves to this dataset.
Papers archive 2025-07-28
No paper in the archive has a leaderboard row on this dataset; the archive counts 2 papers for it but never published that list.
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
No task tagged in the archive.
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
No modality tagged.
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- OCR-VQA
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections