Home › Datasets › task › Text Retrieval

Text Retrieval datasets

archive 2025-07-28

37 datasets carry the task tag "Text Retrieval" (the task itself: Text Retrieval), ordered by the archive's paper count. Page 1 of 1: 37 shown of 37. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Text Retrieval datasets 1–37 of 37

The Natural Questions corpus is a question answering dataset containing 307,373 training examples, 7,830 development examples, and 7,842 test examples.
1,404 papers · 9 benchmarks
MS MARCO (Microsoft Machine Reading Comprehension Dataset)
The MS MARCO (Microsoft MAchine Reading Comprehension) is a collection of datasets focused on deep learning in search.
1,036 papers · 7 benchmarks
HotpotQA is a question answering dataset collected on the English Wikipedia, containing about 113K crowd-sourced questions that are constructed to require the introduction paragraphs of two Wikipedia articles to answer.
933 papers · 3 benchmarks
DBpedia (from "DB" for "database") is a project aiming to extract structured content from the information created in the Wikipedia project.
597 papers · 4 benchmarks
FEVER (Fact Extraction and VERification)
FEVER is a publicly available dataset for fact extraction and verification against textual sources.
498 papers · 3 benchmarks
MTEB (Massive Text Embedding Benchmark)
MTEB is a benchmark that spans 8 embedding tasks covering a total of 56 datasets and 112 languages.
155 papers · 6 benchmarks
SciFact is a dataset of 1.4K expert-written claims, paired with evidence-containing abstracts annotated with veracity labels and rationales.
116 papers · 1 benchmark
TREC-COVID is a community evaluation designed to build a test collection that captures the information needs of biomedical researchers using the scientific literature during a pandemic.
73 papers · 1 benchmark
RSICD (Remote Sensing Image Captioning Dataset)
70 papers · 3 benchmarks
SciDocs evaluation framework consists of a suite of evaluation tasks designed for document-level tasks.
57 papers · 2 benchmarks
Quora Question Pairs (QQP) dataset consists of over 400,000 question pairs, and each question pair is annotated with a binary value indicating whether the two questions are paraphrase of each other.
55 papers · 8 benchmarks
A new publicly available dataset for verification of climate change-related claims.
38 papers · 1 benchmark
The IMAGE-CHAT dataset is a large collection of (image, style trait for speaker A, style trait for speaker B, dialogue between A & B) tuples that we collected using crowd-workers, Each dialogue consists of consecutive turns by speaker A…
31 papers · 2 benchmarks
The 20 Newsgroups data set is a collection of approximately 20,000 newsgroup documents, partitioned (nearly) evenly across 20 different newsgroups.
27 papers · 5 benchmarks
DIOR-RSVG is a large-scale benchmark dataset of remote sensing data (RSVG).
25 papers · 0 benchmarks
TripClick is a large-scale dataset of click logs in the health domain, obtained from user interactions of the Trip Database health web search engine.
16 papers · 0 benchmarks
A large-scale video dataset, featuring clips from movies with detailed captions.
15 papers · 1 benchmark
Vidore (Visual Document Retrieval Benchmark)
It is collection regrouping all datasets constituting the ViDoRe benchmark.
13 papers · 0 benchmarks
We present PeerQA, a real-world, scientific, document-level Question Answering (QA) dataset.
12 papers · 3 benchmarks
A large-scale curated dataset of over 152 million tweets, growing daily, related to COVID-19 chatter generated from January 1st to April 4th at the time of writing.
10 papers · 0 benchmarks
BSARD (Belgian Statutory Article Retrieval Dataset)
The Belgian Statutory Article Retrieval Dataset (BSARD) is a French native corpus for studying statutory article retrieval.
6 papers · 1 benchmark
NFCorpus is a full-text English retrieval data set for Medical Information Retrieval.
5 papers · 1 benchmark
LLeQA (Long-form Legal Question Answering)
LLeQA is a French native dataset for studying information retrieval and long-form question answering in the legal domain.
3 papers · 0 benchmarks
LatamXIX (19th Century Latin American Spanish Newspaper Corpus with LLM OCR Correction)
A novel dataset of 19th-century Latin American press texts, which addresses the lack of specialized corpora for historical and linguistic analysis in this region.
2 papers · 0 benchmarks
MAPS-KB is a million-scale probabilistic simile knowledge base, covering 4.3 million triplets over 0.4 million terms from 70 GB corpora.
2 papers · 0 benchmarks
Dataset Card for the ACR Appropriateness Criteria Corpus This dataset contains chunked guidelines and narratives from the ACR Appropriateness Criteria, an set of societal guidelines from the American College of Radiology (ACR) to help…
1 paper · 0 benchmarks
CURE (A dataset for Clinical Understanding & Retrieval Evaluation)
CURE is a retrieval dataset with a monolingual and two cross-lingual conditions, with splits spanning ten medical domains.
1 paper · 0 benchmarks
- An evaluation test bed for assessing the robustness of sentence embedding models against user-informed misinformation edits.
1 paper · 0 benchmarks
CompMix-IR Dataset Overview: Characteristics: CompMix-IR is a heterogeneous knowledge retrieval benchmark dataset, featuring four knowledge types (text, knowledge graphs, tables, and infoboxes), 9,400+ QA pairs, and a corpus of 10 million…
1 paper · 0 benchmarks
DAPFAM (A Domain‑Aware Patent Retrieval Dataset Aggregated at the Family Level)
Dataset DAPFAM See the accompanying paper: Ayaou et al., 2025 — “DAPFAM: A Domain‑Aware Patent Retrieval Dataset Aggregated at the Family Level” (arXiv:2506.22141).
1 paper · 0 benchmarks
DialogCC is a large-scale multi-modal dialogue dataset, which covers diverse real-world topics and various images per dialogue.
1 paper · 0 benchmarks
The Multi-Eup is a new multilingual benchmark dataset, comprising 22K multilingual documents collected from the European Parliament, spanning 24 languages.
1 paper · 0 benchmarks
PolyNews is a multilingual parallel dataset containing news titles 833 language pairs, spanning in 64 languages and 17 scripts.
1 paper · 0 benchmarks
A large-scale dataset for proactive document retrieval that consists of over 2.8 million conversations from Reddit.
1 paper · 0 benchmarks
Spanish Corpus XIX (19th Century Spanish Corpus)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
--- annotationscreators: - no-annotation language: - fr languagecreators: - found license: - cc-by-4.0 multilinguality: - monolingual prettyname: French Legal Cases Dataset sizecategories: - n>1M sourcedatasets: - la-mousse/INCA-17-01-2025…
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.