Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 4 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 145–192 of 3,998

SciERC dataset is a collection of 500 scientific abstract annotated with scientific entities, their relations, and coreference clusters.
134 papers · 7 benchmarks
BLUE (Biomedical Language Understanding Evaluation)
The BLUE benchmark consists of five different biomedicine text-mining tasks with ten corpora.
133 papers · 0 benchmarks
SearchQA was built using an in-production, commercial search engine.
133 papers · 1 benchmark
The AlpacaEval set contains 805 instructions form self-instruct, open-assistant, vicuna, koala, hh-rlhf.
131 papers · 2 benchmarks
LIAR is a publicly available dataset for fake news detection.
130 papers · 1 benchmark
Europarl (European Parliament Proceedings Parallel Corpus)
A corpus of parallel text in 21 European languages from the proceedings of the European Parliament.
128 papers · 1 benchmark
LogiQA consists of 8,678 QA instances, covering multiple types of deductive reasoning.
127 papers · 1 benchmark
SMAP (Soil Moisture Active Passive)
Soil Moisture Active Passive (SMAP) dataset is a dataset of soil samples and telemetry information using the Mars rover by NASA.
127 papers · 2 benchmarks
WNUT 2017 (WNUT 2017 Emerging and Rare entity recognition)
This shared task focuses on identifying unusual, previously-unseen entities in the context of emerging discussions.
127 papers · 2 benchmarks
WikiHow is a dataset of more than 230,000 article and summary pairs extracted and constructed from an online knowledge base written by different human authors.
127 papers · 2 benchmarks
MSRA-TD500 (MSRA Text Detection 500 Database)
The MSRA-TD500 dataset is a text detection dataset that contains 300 training images and 200 test images.
124 papers · 1 benchmark
The Microsoft Academic Graph is a heterogeneous graph containing scientific publication records, citation relationships between those publications, as well as authors, institutions, journals, conferences, and fields of study.
124 papers · 0 benchmarks
CARER (Contextualized Affect Representations for Emotion Recognition)
CARER is an emotion dataset collected through noisy labels, annotated via distant supervision as in (Go et al., 2009).
123 papers · 2 benchmarks
BBQ (Bias Benchmark for QA)
Bias Benchmark for QA (BBQ) is a dataset consisting of question-sets constructed by the authors that highlight attested social biases against people belonging to protected classes along nine different social dimensions relevant for U.S.
122 papers · 0 benchmarks
MAWPS (MAth Word ProblemS)
MAWPS is an online repository of Math Word Problems, to provide a unified testbed to evaluate different algorithms.
122 papers · 1 benchmark
Multi-News, consists of news articles and human-written summaries of these articles from the site newser.com.
122 papers · 5 benchmarks
The GENIA corpus is the primary collection of biomedical literature compiled and annotated within the scope of the GENIA project.
121 papers · 7 benchmarks
SHAPES (Swarm Heuristics based Adaptive and Penalized Estimation of Splines)
SHAPES is a dataset of synthetic images designed to benchmark systems for understanding of spatial and logical relations among multiple objects.
120 papers · 1 benchmark
SimpleQuestions is a large-scale factoid question answering dataset.
120 papers · 2 benchmarks
VATEX (Video And TEXt)
VATEX is multilingual, large, linguistically complex, and diverse dataset in terms of both video and natural language descriptions.
118 papers · 3 benchmarks
FLoRes-200 doubles the existing language coverage of FLoRes-101.
117 papers · 1 benchmark
VGGFace2 Dataset (Vggface2: A dataset for recognising faces across pose and age)
VGGFace2 is a large-scale face recognition dataset.
117 papers · 0 benchmarks
LLVIP (A Visible-infrared Paired Dataset for Low-light Vision)
Visible-infrared Paired Dataset for Low-light Vision 30976 images (15488 pairs) 24 dark scenes, 2 daytime scenes Support for image-to-image translation (visible to infrared, or infrared to visible), visible and infrared image fusion,…
116 papers · 6 benchmarks
The MRQA (Machine Reading for Question Answering) dataset is a dataset for evaluating the generalization capabilities of reading comprehension systems.
116 papers · 1 benchmark
This corpus comprises of monolingual data for 100+ languages and also includes data for romanized languages.
115 papers · 0 benchmarks
MCTest is a freely available set of stories and associated questions intended for research on the machine comprehension of text.
114 papers · 2 benchmarks
Over a period of three years (2009 - 2011) the daily news and weather forecast airings of the German public tv-station PHOENIX featuring sign language interpretation have been recorded and the weather forecasts of a subset of 386 editions…
114 papers · 2 benchmarks
WHAM! (WSJ0 Hipster Ambient Mixtures)
The WSJ0 Hipster Ambient Mixtures (WHAM!) dataset pairs each two-speaker mixture in the wsj0-2mix dataset with a unique noise background scene.
114 papers · 2 benchmarks
Reading Comprehension with Commonsense Reasoning Dataset (ReCoRD) is a large-scale reading comprehension dataset which requires commonsense reasoning.
111 papers · 1 benchmark
Dataset Summary Mind2Web is a dataset for developing and evaluating generalist agents for the web that can follow language instructions to complete complex tasks on any website.
110 papers · 1 benchmark
MGSM (Multilingual Grade School Math)
Multilingual Grade School Math Benchmark (MGSM) is a benchmark of grade-school math problems.
107 papers · 1 benchmark
Sentiment analysis is increasingly viewed as a vital task both from an academic and a commercial standpoint.
107 papers · 4 benchmarks
VIST (Visual Storytelling)
The Visual Storytelling Dataset (VIST) consists of 210,819 unique photos and 50,000 stories.
107 papers · 2 benchmarks
The Newsela dataset was introduced by Xu et al.
106 papers · 1 benchmark
SLURP (Spoken Language Understanding Resource Package)
A new challenging dataset in English spanning 18 domains, which is substantially bigger and linguistically more diverse than existing datasets.
106 papers · 2 benchmarks
ReDial (Recommendation Dialogues) is an annotated dataset of dialogues, where users recommend movies to each other.
105 papers · 2 benchmarks
ToolBench is an instruction-tuning dataset for tool use, which is created automatically using ChatGPT.
105 papers · 1 benchmark
GAP is a gender-balanced dataset containing 8,908 coreference-labeled pairs of (ambiguous pronoun, antecedent name), sampled from Wikipedia and released by Google AI Language for the evaluation of coreference resolution in practical…
104 papers · 0 benchmarks
ReferIt3D provides two large-scale and complementary visio-linguistic datasets: i) Sr3D, which contains 83.5K template-based utterances leveraging spatial relations among fine-grained object classes to localize a referred object in a…
104 papers · 1 benchmark
CosmosQA is a large-scale dataset of 35.6K problems that require commonsense-based reading comprehension, formulated as multiple-choice questions.
102 papers · 0 benchmarks
QASPER is a dataset for question answering on scientific research papers.
102 papers · 1 benchmark
ConvAI2 (Conversational Intelligence Challenge 2)
The ConvAI2 NeurIPS competition aimed at finding approaches to creating high-quality dialogue agents capable of meaningful open domain conversation.
100 papers · 1 benchmark
CLUE (Chinese Language Understanding Evaluation Benchmark)
CLUE is a Chinese Language Understanding Evaluation benchmark.
99 papers · 8 benchmarks
FIGER (Fine-Grained Entity Recognition)
The FIGER dataset is an entity recognition dataset where entities are labelled using fine-grained system 112 tags, such as person/doctor, art/writtenwork and building/hotel.
96 papers · 2 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.