Home › Datasets › task › Retrieval
Retrieval datasets
archive 2025-07-28
38 datasets carry the task tag "Retrieval" (the task itself: Retrieval), ordered by the archive's paper count. Page 1 of 1: 38 shown of 38. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Retrieval datasets 1–38 of 38
The Natural Questions corpus is a question answering dataset containing 307,373 training examples, 7,830 development examples, and 7,842 test examples.
1,404 papers · 9 benchmarks
HotpotQA is a question answering dataset collected on the English Wikipedia, containing about 113K crowd-sourced questions that are constructed to require the introduction paragraphs of two Wikipedia articles to answer.
933 papers · 3 benchmarks
OK-VQA (Outside Knowledge Visual Question Answering)
Outside Knowledge Visual Question Answering (OK-VQA) includes more than 14,000 questions that require external knowledge to answer.
368 papers · 2 benchmarks
The task of PubMedQA is to answer research questions with yes/no/maybe (e.g.: Do preoperative statins reduce atrial fibrillation after coronary artery bypass grafting?) using the corresponding abstracts.
276 papers · 3 benchmarks
PopQA is an open-domain QA dataset with 14k QA pairs with fine-grained Wikidata entity ID, Wikipedia page views, and relationship type information.
91 papers · 1 benchmark
This dataset contains 21,889 outfits from polyvore.com, in which 17,316 are for training, 1,497 for validation and 3,076 for testing.
62 papers · 3 benchmarks
Quora Question Pairs (QQP) dataset consists of over 400,000 question pairs, and each question pair is annotated with a binary value indicating whether the two questions are paraphrase of each other.
55 papers · 8 benchmarks
ALCE (Automatic LLMs' Citation Evaluation)
ALCE is a benchmark for Automatic LLMs' Citation Evaluation.
39 papers · 0 benchmarks
In this project, we introduce InfoSeek, a visual question answering dataset tailored for information-seeking questions that cannot be answered with only common sense knowledge.
36 papers · 2 benchmarks
Attribution, Relation, and Order (ARO) benchmark to systematically evaluate the ability of VLMs to understand different types of relationships, attributes, and order information.
35 papers · 0 benchmarks
ProofNet is a benchmark for autoformalization and formal proving of undergraduate-level mathematics.
27 papers · 0 benchmarks
DIOR-RSVG is a large-scale benchmark dataset of remote sensing data (RSVG).
25 papers · 0 benchmarks
QAMPARI is an ODQA benchmark, where question answers are lists of entities, spread across many paragraphs.
19 papers · 0 benchmarks
xCodeEval is one of the largest executable multilingual multitask benchmarks consisting of 17 programming languages with execution-level parallelism.
15 papers · 0 benchmarks
The dataset contains single-shot videos taken from moving cameras in underwater environments.
9 papers · 1 benchmark
BSARD (Belgian Statutory Article Retrieval Dataset)
The Belgian Statutory Article Retrieval Dataset (BSARD) is a French native corpus for studying statutory article retrieval.
6 papers · 1 benchmark
ShapeTalk contains over half a million discriminative utterances produced by contrasting the shapes of common 3D objects for a variety of object classes and degrees of similarity.
6 papers · 0 benchmarks
ComFact is a benchmark for commonsense fact linking, where models are given contexts and trained to identify situationally-relevant commonsense knowledge from KGs.
5 papers · 0 benchmarks
DiSCQ (Discharge Summary Clinical Questions)
DiSCQ is a newly curated question dataset composed of 2,000+ questions paired with the snippets of text (triggers) that prompted each question.
4 papers · 0 benchmarks
COSIAN (a collection of singing voice annotation)
COSIAN is an annotation collection of Japanese popular (J-POP) songs, focusing on singing style and expression of famous solo-singers.
3 papers · 0 benchmarks
FanOutQA is a high quality, multi-hop, multi-document benchmark for large language models using English Wikipedia as its knowledge base.
3 papers · 0 benchmarks
LLeQA (Long-form Legal Question Answering)
LLeQA is a French native dataset for studying information retrieval and long-form question answering in the legal domain.
3 papers · 0 benchmarks
The ToolLens dataset consists of 18,770 concise yet intentionally multifaceted queries, each associated with 1 to 3 verified tools out of a total of 464, designed to better mimic real-world user interactions.
3 papers · 1 benchmark
DTGB (Dynamic Text-attributed Graph Benchmark)
We introduce Dynamic Text-attributed Graph Benchmark (DTGB), a collection of large-scale, time-evolving graphs from diverse domains, with nodes and edges enriched by dynamically changing text attributes and categories.
2 papers · 0 benchmarks
PoseScript is a dataset that pairs a few thousand 3D human poses from AMASS with rich human-annotated descriptions of the body parts and their spatial relationships.
2 papers · 0 benchmarks
RoMQA is a benchmark for robust, multi-evidence, and multi-answer question answering (QA).
2 papers · 0 benchmarks
CRSB (Context Retrieval Supervision Benchmark)
The Official dataset proposed int the paper Context Awareness Gate For Retrieval Augmented Generation
1 paper · 0 benchmarks
- An evaluation test bed for assessing the robustness of sentence embedding models against user-informed misinformation edits.
1 paper · 0 benchmarks
CoreSearch is a dataset for Cross-Document Event Coreference Search.
1 paper · 0 benchmarks
DAPFAM (A Domain‑Aware Patent Retrieval Dataset Aggregated at the Family Level)
Dataset DAPFAM See the accompanying paper: Ayaou et al., 2025 — “DAPFAM: A Domain‑Aware Patent Retrieval Dataset Aggregated at the Family Level” (arXiv:2506.22141).
1 paper · 0 benchmarks
FewDR is a dataset for Few-shot dense retrieval (DR).
1 paper · 0 benchmarks
This is the supporting dataset for the ECCV 2024 paper "MARs: Multi-view Attention Regularizations for Patch-based Feature Recognition of Space Terrain".
1 paper · 0 benchmarks
PTVD is a plot-oriented multimodal dataset in the TV domain.
1 paper · 0 benchmarks
PubMedQA-MetaGen: Metadata-Enriched PubMedQA Corpus Dataset Summary PubMedQA-MetaGen is a metadata-enriched version of the PubMedQA biomedical question-answering dataset, created using the MetaGenBlendedRAG enrichment pipeline.
1 paper · 2 benchmarks
Spiced is a paraphrase dataset of scientific findings annotated for degree of information change.
1 paper · 0 benchmarks
> The StatCan Dialogue Dataset: Retrieving Data Tables through Conversations with Genuine Intents > > Xing Han Lu, Siva Reddy, Harm de Vries > > EACL 2023 | | | | | | | :--: | :--: | :--: | :--: | :--: | | Code | Huggingface | Request on…
1 paper · 1 benchmark
Visual Haystacks (VHs) is a "visual-centric" Needle-In-A-Haystack (NIAH) benchmark specifically designed to evaluate the capabilities of Large Multimodal Models (LMMs) in visual retrieval and reasoning over sets of unrelated images.
1 paper · 0 benchmarks
Multimodal Large Language Models (MLLMs) have shown significant promise in various applications, leading to broad interest from researchers and practitioners alike.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.