Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 8 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 337–384 of 3,998

MLSUM (MultiLingual SUMmarization)
A large-scale MultiLingual SUMmarization dataset.
45 papers · 4 benchmarks
RSTPReid (Real Scenario Text-based Person Re-identification)
RSTPReid contains 20505 images of 4,101 persons from 15 cameras.
45 papers · 2 benchmarks
COCO-O(ut-of-distribution) contains 6 domains (sketch, cartoon, painting, weather, handmake, tattoo) of COCO objects which are hard to be detected by most existing detectors.
44 papers · 1 benchmark
EmotionLines contains a total of 29245 labeled utterances from 2000 dialogues.
44 papers · 1 benchmark
Occ3D is a dataset for 3D occupancy prediction, which aims to estimate the detailed occupancy and semantics of objects from multi-view images.
44 papers · 1 benchmark
The SCUT-CTW1500 dataset contains 1,500 images: 1,000 for training and 500 for testing.
44 papers · 3 benchmarks
TurkCorpus, a dataset with 2,359 original sentences from English Wikipedia, each with 8 manual reference simplifications.
44 papers · 1 benchmark
Dataset contains 33,010 molecule-description pairs split into 80\%/10\%/10\% train/val/test splits.
43 papers · 4 benchmarks
Occluded-DukeMTMC contains 15,618 training images, 17,661 gallery images, and 2,210 occluded query images.
43 papers · 1 benchmark
RWC (Real World Computing Music Database)
The RWC (Real World Computing) Music Database is a copyright-cleared music database (DB) that is available to researchers as a common foundation for research.
43 papers · 0 benchmarks
WebQA, is a new benchmark for multimodal multihop reasoning in which systems are presented with the same style of data as humans when searching the web: Snippets and Images.
43 papers · 0 benchmarks
gRefCOCO is the first large-scale Generalized Referring Expression Segmentation dataset that contains multi-target, no-target, and single-target expressions.
43 papers · 2 benchmarks
AVSpeech is a large-scale audio-visual dataset comprising speech clips with no interfering background signals.
42 papers · 0 benchmarks
BUCC (Building and Using Comparable Corpora)
The BUCC mining task is a shared task on parallel sentence extraction from two monolingual corpora with a subset of them assumed to be parallel, and that has been available since 2016.
42 papers · 4 benchmarks
C-GQA (Compositional GQA)
We propose a split built on top of Stanford GQA dataset originally proposed for VQA and name it Compositional GQA (C-GQA) dataset (see supplementary for the details).
42 papers · 0 benchmarks
DuoRC contains 186,089 unique question-answer pairs created from a collection of 7680 pairs of movie plots where each pair in the collection reflects two versions of the same movie.
42 papers · 1 benchmark
We release E-commerce Dialogue Corpus, comprising a training data set, a development set and a test set for retrieval based chatbot.
42 papers · 1 benchmark
EmoContext consists of three-turn English Tweets.
42 papers · 1 benchmark
English Web Treebank is a dataset containing 254,830 word-level tokens and 16,624 sentence-level tokens of webtext in 1174 files annotated for sentence- and word-level tokenization, part-of-speech, and syntactic structure.
42 papers · 0 benchmarks
IconQA (Icon Question Answering)
Current visual question answering (VQA) tasks mainly consider answering human-annotated questions for natural images in the daily-life context.
42 papers · 1 benchmark
PearRead is a dataset of scientific peer reviews.
42 papers · 0 benchmarks
SCROLLS (Standardized CompaRison Over Long Language Sequences)
SCROLLS (Standardized CompaRison Over Long Language Sequences) is an NLP benchmark consisting of a suite of tasks that require reasoning over long texts.
42 papers · 1 benchmark
The UCR Time Series Archive - introduced in 2002, has become an important resource in the time series data mining community, with at least one thousand published papers making use of at least one data set from the archive.
42 papers · 2 benchmarks
This dataset contains product reviews and metadata from Amazon, including 142.8 million reviews spanning May 1996 - July 2014.
41 papers · 5 benchmarks
A creative writing task where the input is 4 random sentences and the output should be a coherent passage with 4 paragraphs that end in the 4 input sentences respectively.
41 papers · 0 benchmarks
NT-VOT211 consists of 211 diverse videos, offering 211,000 well-annotated frames with 8 attributes including camera motion, deformation, fast motion, motion blur, tiny target, distractors, occlusion and out-of-view.
41 papers · 1 benchmark
QVHighlights (Query-based Video Highlights)
The Query-based Video Highlights (QVHighlights) dataset is a dataset for detecting customized moments and highlights from videos given natural language (NL).
41 papers · 4 benchmarks
A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research.
41 papers · 0 benchmarks
AID (Aerial Image Dataset)
AID is a new large-scale aerial image dataset, by collecting sample images from Google Earth imagery.
40 papers · 2 benchmarks
BLURB (Biomedical Language Understanding and Reasoning Benchmark)
BLURB is a collection of resources for biomedical natural language processing.
40 papers · 2 benchmarks
LegalBench is a fascinating project that revolves around legal reasoning and evaluation.
40 papers · 0 benchmarks
SciCite is a dataset of citation intents that addresses multiple scientific domains and is more than five times larger than ACL-ARC.
40 papers · 3 benchmarks
We propose the first question-answering dataset driven by STEM theorems.
40 papers · 1 benchmark
The Argoverse 2 Motion Forecasting Dataset is a curated collection of 250,000 scenarios for training and validation.
39 papers · 0 benchmarks
BookSum is a collection of datasets for long-form narrative summarization.
39 papers · 2 benchmarks
WikiMovies is a dataset for question answering for movies content.
39 papers · 0 benchmarks
CUAD (Contract Understanding Atticus Dataset)
Contract Understanding Atticus Dataset (CUAD) is a dataset for legal contract review.
38 papers · 0 benchmarks
Imagenette is a subset of 10 easily classified classes from Imagenet (bench, English springer, cassette player, chain saw, church, French horn, garbage truck, gas pump, golf ball, parachute).
38 papers · 1 benchmark
InsuranceQA is a question answering dataset for the insurance domain, the data stemming from the website Insurance Library.
38 papers · 0 benchmarks
Tatoeba is a free collection of example sentences with translations geared towards foreign language learners.
38 papers · 2 benchmarks
HOC (Hallmarks of Cancer)
The Hallmarks of Cancer (HOC) corpus consists of 1852 PubMed publication abstracts manually annotated by experts according to the Hallmarks of Cancer taxonomy.
37 papers · 1 benchmark
The SemEval-2018 hypernym discovery evaluation benchmark (Camacho-Collados et al.
37 papers · 3 benchmarks
TextOCR is a dataset to benchmark text recognition on arbitrary shaped scene-text.
37 papers · 0 benchmarks
WMT 2018 is a collection of datasets used in shared tasks of the Third Conference on Machine Translation.
37 papers · 4 benchmarks
The Yelp Reviews Polarity dataset is obtained from the Yelp Dataset Challenge in 2015 (1,569,264 samples that have review text).
37 papers · 0 benchmarks
AdvGLUE (Adversarial GLUE)
Adversarial GLUE (AdvGLUE) is a new multi-task benchmark to quantitatively and thoroughly explore and evaluate the vulnerabilities of modern large-scale language models under various types of adversarial attacks.
36 papers · 1 benchmark
DialogRE is the first human-annotated dialogue-based relation extraction dataset, containing 1,788 dialogues originating from the complete transcripts of a famous American television situation comedy Friends.
36 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.