Home › Datasets › task › Information Retrieval
Information Retrieval datasets
archive 2025-07-28
93 datasets carry the task tag "Information Retrieval" (the task itself: Information Retrieval), ordered by the archive's paper count. Page 1 of 2: 48 shown of 93. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Information Retrieval datasets 1–48 of 93
MS MARCO (Microsoft Machine Reading Comprehension Dataset)
The MS MARCO (Microsoft MAchine Reading Comprehension) is a collection of datasets focused on deep learning in search.
1,036 papers · 7 benchmarks
CoNLL-2003 is a named entity recognition dataset released as a part of CoNLL-2003 shared task: language-independent named entity recognition.
755 papers · 6 benchmarks
BioASQ (Biomedical Semantic Indexing and Question Answering)
BioASQ is a question answering dataset.
192 papers · 1 benchmark
CORD-19 is a free resource of tens of thousands of scholarly articles about COVID-19, SARS-CoV-2, and related coronaviruses for use by the global research community.
163 papers · 1 benchmark
MTEB (Massive Text Embedding Benchmark)
MTEB is a benchmark that spans 8 embedding tasks covering a total of 56 datasets and 112 languages.
155 papers · 6 benchmarks
The Free Music Archive (FMA) is a large-scale dataset for evaluating several tasks in Music Information Retrieval.
128 papers · 2 benchmarks
PubLayNet is a dataset for document layout analysis by automatically matching the XML representations and the content of over 1 million PDF articles that are publicly available on PubMed Central.
123 papers · 1 benchmark
QASC (Question Answering via Sentence Composition)
QASC is a question-answering dataset with a focus on sentence composition.
114 papers · 0 benchmarks
TREC-COVID is a community evaluation designed to build a test collection that captures the information needs of biomedical researchers using the scientific literature during a pandemic.
73 papers · 1 benchmark
MedleyDB, is a dataset of annotated, royalty-free multitrack recordings.
47 papers · 0 benchmarks
MLSUM (MultiLingual SUMmarization)
A large-scale MultiLingual SUMmarization dataset.
45 papers · 4 benchmarks
InsuranceQA is a question answering dataset for the insurance domain, the data stemming from the website Insurance Library.
38 papers · 0 benchmarks
The MSLR-WEB10K dataset consists of 10,000 search queries over the documents from search results.
36 papers · 0 benchmarks
SciTSR is a large-scale table structure recognition dataset, which contains 15,000 tables in PDF format and their corresponding structure labels obtained from LaTeX source files.
36 papers · 0 benchmarks
The MQ2007 dataset consists of queries, corresponding retrieved documents and labels provided by human experts.
32 papers · 0 benchmarks
GuitarSet is a dataset of high-quality guitar recordings and rich annotations.
31 papers · 2 benchmarks
The MQ2008 dataset is a dataset for Learning to Rank.
31 papers · 0 benchmarks
The MSLR-WEB30K dataset consists of 30,000 search queries over the documents from search results.
31 papers · 1 benchmark
WikiReading is a large-scale natural language understanding task and publicly-available dataset with 18 million instances.
26 papers · 0 benchmarks
ASNQ (Answer Sentence Natural Questions)
A large scale dataset to enable the transfer step, exploiting the Natural Questions dataset.
25 papers · 1 benchmark
Composed of 1,395 questions posed by crowdworkers on Wikipedia articles, and a machine translation of the Stanford Question Answering Dataset (Arabic-SQuAD).
24 papers · 0 benchmarks
HeadQA is a multi-choice question answering testbed to encourage research on complex reasoning.
23 papers · 1 benchmark
The George Washington dataset contains 20 pages of letters written by George Washington and his associates in 1755 and thereby categorized into historical collection.
20 papers · 0 benchmarks
The iKala dataset is a singing voice separation dataset that comprises of 252 30-second excerpts sampled from 206 iKala songs (plus 100 hidden excerpts reserved for MIREX data mining contest).
20 papers · 1 benchmark
A dataset on asking Questions for Lack of Clarity in open-domain information-seeking conversations.
19 papers · 0 benchmarks
A dataset with 2,437 dialogues and 10,917 QA pairs.
18 papers · 0 benchmarks
ORCAS is a click-based dataset.
18 papers · 0 benchmarks
ORConvQA (Open-Retrieval Conversational Question Answering)
Enhances QuAC by adapting it to an open-retrieval setting.
17 papers · 0 benchmarks
JuICe is a corpus of 1.5 million examples with a curated test set of 3.7K instances based on online programming assignments.
16 papers · 0 benchmarks
TripClick is a large-scale dataset of click logs in the health domain, obtained from user interactions of the Trip Database health web search engine.
16 papers · 0 benchmarks
The beginnings of a question answering dataset specifically designed for COVID-19, built by hand from knowledge gathered from Kaggle's COVID-19 Open Research Dataset Challenge.
15 papers · 0 benchmarks
The goal of the Robust track is to improve the consistency of retrieval technology by focusing on poorly performing topics.
13 papers · 1 benchmark
ClariQ is an extension of the Qulac dataset with additional new topics, questions, and answers in the training set.
11 papers · 0 benchmarks
Ohsumed includes medical abstracts from the MeSH categories of the year 1991.
11 papers · 2 benchmarks
ReQA (Retrieval Question-Answering)
Retrieval Question-Answering (ReQA) benchmark tests a model’s ability to retrieve relevant answers efficiently from a large set of documents.
10 papers · 0 benchmarks
CLIRMatrix is a large collection of bilingual and multilingual datasets for Cross-Lingual Information Retrieval.
9 papers · 0 benchmarks
ClueWeb22 is the newest iteration of the ClueWeb line of datasets, provides 10 billion web pages affiliated with rich information.
9 papers · 0 benchmarks
GiantMIDI-Piano contains 10,854 unique piano solo pieces composed by 2,786 composers.
9 papers · 0 benchmarks
SciRepEval is a comprehensive benchmark for training and evaluating scientific document representations.
9 papers · 0 benchmarks
The Standardized Project Gutenberg Corpus (SPGC) is an open science approach to a curated version of the complete PG data containing more than 50,000 books and more than 3×109 word-tokens.
9 papers · 0 benchmarks
GermanQuAD is a Question Answering (QA) dataset of 13,722 extractive question/answer pairs in German.
8 papers · 1 benchmark
OpenMIC-2018 is an instrument recognition dataset containing 20,000 examples of Creative Commons-licensed music available on the Free Music Archive.
8 papers · 1 benchmark
QUASAR-S (QUestion Answering by Search And Reading – Stack Overflow)
QUASAR-S is a large-scale dataset aimed at evaluating systems designed to comprehend a natural language query and extract its answer from a large corpus of text.
8 papers · 0 benchmarks
WANDS (Wayfair ANnotation Dataset)
The dataset contains: 42,994 candidate products with data comprising product class, title, description, attributes, category hierarchy, average rating, and number of reviews 480 search query strings with predicted product class 233,448…
8 papers · 0 benchmarks
HiREST (HIerarchical REtrieval and STep-captioning)
HiREST (HIerarchical REtrieval and STep-captioning) dataset is a benchmark that covers hierarchical information retrieval and visual/textual stepwise summarization from an instructional video corpus.
7 papers · 0 benchmarks
MSSD (Music Streaming Sessions Dataset)
The Spotify Music Streaming Sessions Dataset (MSSD) consists of 160 million streaming sessions with associated user interactions, audio features and metadata describing the tracks streamed during the sessions, and snapshots of the…
7 papers · 1 benchmark
ArCOV-19 is an Arabic COVID-19 Twitter dataset that covers the period from 27th of January till 30th of April 2020.
6 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.