Home › Datasets › task › Question Answering

Question Answering datasets

archive 2025-07-28

413 datasets carry the task tag "Question Answering" (the task itself: Question Answering), ordered by the archive's paper count. Page 9 of 9: 29 shown of 413. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Question Answering datasets 385–413 of 413

QASiNa (Question Answering Sirah Nabawiyah)
Question Answering Sirah Nabawiyah (QASiNa) Dataset is a reading comprehension dataset consists of QA from Sirah Nabawiyah literature in Indonesian Language
1 paper · 0 benchmarks
QASports (A Question Answering Dataset about Sports)
Sport is one of the most popular and revenue-generating forms of entertainment.
1 paper · 0 benchmarks
RGRS (ResearchGate dataset for Recommending Systems)
RGRS is a dataset for collaboratior recommendation on the ResearchGate academic social network.
1 paper · 0 benchmarks
SCIMAT is a large question-answer dataset for mathematics and science problems; such dataset can have impact on online education, intelligent tutoring and automated grading.
1 paper · 0 benchmarks
Curated QA Benchmark on State of the Union Address 2023.
1 paper · 0 benchmarks
The “Mental Health” forum was used, a forum dedicated to people suffering from schizophrenia and different mental disorders.
1 paper · 1 benchmark
ScienceExamCER is a collection of resources for studying explanation-centered inference, including explanation graphs for 1,680 questions, with 4,950 tablestore rows, and other analyses of the knowledge required to answer elementary and…
1 paper · 0 benchmarks
TQBA++ (Tiny QA Benchmark++)
Ultra-lightweight, multilingual QA eval dataset for rapid testing LLMs.
1 paper · 0 benchmarks
TUMTraffic-VideoQA is a novel dataset designed to understand spatiotemporal video in complex roadside traffic scenarios.
1 paper · 0 benchmarks
Question answering over temporal knowledge graphs (TKGs) is crucial for understanding evolving facts and relationships, yet its development is hindered by limited datasets and difficulties in generating custom QA pairs.
1 paper · 0 benchmarks
TinySocial is a dataset to enable research on Social Visual Question Answering.
1 paper · 0 benchmarks
The data consists of a set of 3 task types and 4 question types, creating 12 total scenarios.
1 paper · 0 benchmarks
The TupleInf Open IE dataset contains Open IE tuples extracted from 263K sentences that were used by the solver in “Answering Complex Questions Using Open Information Extraction” (referred as Tuple KB, T).
1 paper · 0 benchmarks
VCG+112K (Video Instruction Dataset 112K)
Video-ChatGPT introduces the VideoInstruct100K dataset, which employs a semi-automatic annotation pipeline to generate 75K instruction-tuning QA pairs.
1 paper · 0 benchmarks
VDQG (Visual Discriminative Question Generation)
The Visual Discriminative Question Generation (VDQG) dataset contains 11202 ambiguous image pairs collected from Visual Genome.
1 paper · 0 benchmarks
VTQA (Visual Text Question Answering)
VTQA is a dataset containing open-ended questions about image-text pairs.
1 paper · 0 benchmarks
Visual Choice of Plausible Alternatives (VCOPA) is an evaluation dataset containing 380 VCOPA questions and over 1K images with various topics, which is amenable to automatic evaluation, and present the performance of baseline reasoning…
1 paper · 0 benchmarks
Visual Haystacks (VHs) is a "visual-centric" Needle-In-A-Haystack (NIAH) benchmark specifically designed to evaluate the capabilities of Large Multimodal Models (LMMs) in visual retrieval and reasoning over sets of unrelated images.
1 paper · 0 benchmarks
VlogQA (Vietnamese Spoken-Based Machine Reading Comprehension)
The VlogQA consists of 10,076 question-answer pairs based on 1,230 transcript documents sourced from YouTube - an extensive source of user-uploaded content, covering the topics of food and travel in the Vietnamese language.
1 paper · 0 benchmarks
The WangchanX-Legal-ThaiCCL-RAG dataset supports the development Retrieval-Augmented Generation (RAG) for Thai Legal question answering.
1 paper · 0 benchmarks
WikiQAar (English-Arabic Wikipedia Question-Answering)
A publicly available set of question and sentence pairs, collected and annotated for research on open-domain question answering.
1 paper · 0 benchmarks
To collect WikiSuggest, Google Suggest API is used to harvest natural language questions and submit them to Google Search.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Xamarin Q&A consists of two datasets of questions and answers for studying the development of cross-platform mobile applications using the Xamarin framework.
1 paper · 0 benchmarks
catbAbI LM-mode (concatenated-bAbI)
We aim to improve the bAbI benchmark as a means of developing intelligent dialogue agents.
1 paper · 1 benchmark
catbAbI QA-mode (concatenated-bAbI)
We aim to improve the bAbI benchmark as a means of developing intelligent dialogue agents.
1 paper · 1 benchmark
diaforge-utc-r-0725 (DiaFORGE UTC: Unified Tool-Calling Conversations Dataset)
Dataset for our paper Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky which includes 5000 enterprise tools and the corresponding dialogues generated using DiaFORGE UTC data engine.
1 paper · 0 benchmarks
The simply-CLEVR dataset aims to provide a benchmark dataset that can be used for transparent quantitative evaluation of explanation methods (aka heatmaps/XAI methods).
1 paper · 0 benchmarks
VQA NLE synthetic dataset, made with LLaVA-1.5 using features from GQA dataset.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.