Home › Datasets › task › Information Retrieval

Information Retrieval datasets

archive 2025-07-28

93 datasets carry the task tag "Information Retrieval" (the task itself: Information Retrieval), ordered by the archive's paper count. Page 2 of 2: 45 shown of 93. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Information Retrieval datasets 49–93 of 93

BSARD (Belgian Statutory Article Retrieval Dataset)
The Belgian Statutory Article Retrieval Dataset (BSARD) is a French native corpus for studying statutory article retrieval.
6 papers · 1 benchmark
Microsoft Research Social Media Conversation Corpus consists of 127M context-message-response triples from the Twitter FireHose, covering the 3-month period June 2012 through August 2012.
6 papers · 0 benchmarks
TutorialBank is a publicly available dataset which aims to facilitate NLP education and research.
6 papers · 0 benchmarks
CQADupStack is a benchmark dataset for community question-answering research.
5 papers · 1 benchmark
Coached Conversational Preference Elicitation is a dataset consisting of 502 English dialogs with 12,000 annotated utterances between a user and an assistant discussing movie preferences in natural language.
5 papers · 0 benchmarks
Grep-BiasIR (Gender Representation-Bias for Information Retrieval)
Grep-BiasIR is a novel thoroughly-audited dataset which aim to facilitate the studies of gender bias in the retrieved results of IR systems.
5 papers · 0 benchmarks
NFCorpus is a full-text English retrieval data set for Medical Information Retrieval.
5 papers · 1 benchmark
The Bach Doodle Dataset is composed of 21.6 million harmonizations submitted from the Bach Doodle.
4 papers · 0 benchmarks
CCPE-M (Coached Conversational Preference Elicitation dataset for Movies)
A dataset consisting of 502 English dialogs with 12,000 annotated utterances between a user and an assistant discussing movie preferences in natural language.
4 papers · 0 benchmarks
Goal is a novel dataset of football (or 'soccer') highlights videos with transcribed live commentaries in English.
4 papers · 0 benchmarks
GoodSounds dataset contains around 28 hours of recordings of single notes and scales played by 15 different professional musicians, all of them holding a music degree and having some expertise in teaching.
4 papers · 0 benchmarks
MuMu is a new dataset of more than 31k albums classified into 250 genre classes.
4 papers · 0 benchmarks
SV-Ident (Survey Variable Identification)
SV-Ident comprises 4,248 sentences from social science publications in English and German.
4 papers · 2 benchmarks
WikiCLIR is a large-scale (German-English) retrieval data set for Cross-Language Information Retrieval (CLIR).
4 papers · 0 benchmarks
i2b2 De-identification Dataset (Informatics for Integrating Biology and the Bedside (i2b2) Project — De-identification Dataset)
This dataset contains 1304 de-identified longitudinal medical records describing 296 patients.
4 papers · 1 benchmark
COSIAN (a collection of singing voice annotation)
COSIAN is an annotation collection of Japanese popular (J-POP) songs, focusing on singing style and expression of famous solo-singers.
3 papers · 0 benchmarks
DAWT (Densely Annotated Wikipedia Texts)
The DAWT dataset consists of Densely Annotated Wikipedia Texts across multiple languages.
3 papers · 0 benchmarks
DICE: a Dataset of Italian Crime Event news (from Gazzetta di Modena [2011-2021])
The dataset contains the main components of the news articles published online by the newspaper named Gazzetta di Modena: url of the web page, title, sub-title, text, date of publication, crime category assigned to each news article by the…
3 papers · 0 benchmarks
A corpus of 553k news articles from six Persian news websites and agencies with relatively high quality author extracted keyphrases, which is then filtered and cleaned to achieve higher quality keyphrases.
3 papers · 0 benchmarks
A new large-scale retail product dataset for fine-grained image classification.
3 papers · 0 benchmarks
ResQ (Real-world Spatial Question Answering)
ReSQ is a real-world Spatial Question Answering dataset with human-generated questions built on an existing corpus with SpRL annotations.
3 papers · 0 benchmarks
The TREC News Track features modern search tasks in the news domain.
3 papers · 1 benchmark
Nowadays, individuals tend to engage in dialogues with Large Language Models, seeking answers to their questions.
3 papers · 0 benchmarks
BoostCLIR is a bilingual (Japanese-English) corpus of patent abstracts, extracted from the MAREC patent data, and the data from the NTCIR PatentMT workshop collections, accompanied with relevance judgements for the task of patent prior-art…
2 papers · 0 benchmarks
The Large-Scale CLIR Dataset is a retrieval dataset built for Cross-Language Information Retrieval (CLIR).
2 papers · 0 benchmarks
ORCAS-I (Queries Annotated with Intent using Weak Supervision)
A labelled version of the ORCAS click-based dataset of Web queries, which provides 18 million connections to 10 million distinct queries.
2 papers · 1 benchmark
ThreatGram 101 - Extreme Telegram Data (ThreatGram 101 - Extreme Telegram Replies Data with Threat Levels)
Data 1: Raw and Unlabeled; 2 million unlabeled replies from 17 Telegram channels.
2 papers · 1 benchmark
This paper is a condensed report on the second year of the Touché shared task on argument retrieval held at CLEF 2021.
2 papers · 0 benchmarks
CoSQA+ (CoSQA_Plus)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
CompMix-IR Dataset Overview: Characteristics: CompMix-IR is a heterogeneous knowledge retrieval benchmark dataset, featuring four knowledge types (text, knowledge graphs, tables, and infoboxes), 9,400+ QA pairs, and a corpus of 10 million…
1 paper · 0 benchmarks
FZ queries (FindZebra queries)
A set of 248 search queries annotated with the correct diagnosis.
1 paper · 0 benchmarks
GLARE (Guided LexRank for Advanced Retrieval in Legal Analysis)
The Guided Lexrank algorithm is applied to dataset specialappeal.csv to summarize the texts of legal documents.
1 paper · 0 benchmarks
This data adds textual meta-infomation data to two existing corpora for cross language information retrieval: BoostCLIR, and the Large Scale CLIR Dataset (wiki-clir).
1 paper · 0 benchmarks
We introduce OpenDebateEvidence, a comprehensive dataset for argument mining and summarization sourced from the American Competitive Debate community.
1 paper · 0 benchmarks
The Persian Reverse Dictionary Dataset is a collection of 855217 words along with the phrases describing them.
1 paper · 0 benchmarks
Phrase in Context is a curated benchmark for phrase understanding and semantic search, consisting of three tasks of increasing difficulty: Phrase Similarity (PS), Phrase Retrieval (PR) and Phrase Sense Disambiguation (PSD).
1 paper · 0 benchmarks
This dataset is the translation of the MS-marco dataset, marking it the first large-scale urdu IR dataset.
1 paper · 0 benchmarks
Urdu News Headlines Dataset with VOA and BBC An Urdu news headlines dataset is a collection of news headlines in the Urdu language, typically scraped from news websites and social media platforms.
1 paper · 1 benchmark
WMT 2014 Medical (WMT 2014 Medical Translation Task)
The Medical Translation Task of WMT 2014 addresses the problem of domain-specific and genre-specific machine translation.
1 paper · 0 benchmarks
WikiPII, an automatically labeled dataset composed of Wikipedia biography pages, annotated for personal information extraction.
1 paper · 0 benchmarks
diaforge-utc-r-0725 (DiaFORGE UTC: Unified Tool-Calling Conversations Dataset)
Dataset for our paper Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky which includes 5000 enterprise tools and the corresponding dialogues generated using DiaFORGE UTC data engine.
1 paper · 0 benchmarks
MedleyDB 2.0 is a superset of the MedleyDB – a dataset of annotated, royalty-free multitrack recordings.
0 papers · 0 benchmarks
NIAN (Needle in a Needlestack)
The Needle in a Needlestack (NIAN) is a new benchmark designed to measure how well Language Learning Models (LLMs) pay attention to the information in their context window¹.
0 papers · 0 benchmarks
PANACEA (PANACEA dataset - Heterogeneous COVID-19 Claims)
The peer-reviewed publication for this dataset has been presented in the 2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), and can be accessed here:…
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.