Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 3 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 97–144 of 3,998

SGD (Schema-Guided Dialogue)
The Schema-Guided Dialogue (SGD) dataset consists of over 20k annotated multi-domain, task-oriented conversations between a human and a virtual assistant.
186 papers · 2 benchmarks
Argoverse 2 (AV2) is a collection of three datasets for perception and forecasting research in the self-driving domain.
185 papers · 3 benchmarks
APPS (Automated Programming Progress Standard)
The APPS dataset consists of problems collected from different open-access coding websites such as Codeforces, Kattis, and more.
180 papers · 1 benchmark
Dataset of hate speech annotated on Internet forum posts in English at sentence-level.
180 papers · 1 benchmark
IFEval (Instruction Following Evaluation Datset)
This dataset evaluates instruction following ability of large language models.
180 papers · 1 benchmark
LRA (Long-Range Arena)
Long-range arena (LRA) is an effort toward systematic evaluation of efficient transformer models.
180 papers · 1 benchmark
ShapeNetCore is a subset of the full ShapeNet dataset with single clean 3D models and manually verified category and alignment annotations.
180 papers · 1 benchmark
VCR (Visual Commonsense Reasoning)
Visual Commonsense Reasoning (VCR) is a large-scale dataset for cognition-level visual understanding.
179 papers · 13 benchmarks
The AI2’s Reasoning Challenge (ARC) dataset is a multiple-choice question-answering dataset, containing questions from science exams from grade 3 to grade 9.
178 papers · 3 benchmarks
QuAC (Question Answering in Context)
Question Answering in Context is a large-scale dataset that consists of around 14K crowdsourced Question Answering dialogs with 98K question-answer pairs in total.
178 papers · 1 benchmark
WMT 2016 is a collection of datasets used in shared tasks of the First Conference on Machine Translation.
178 papers · 16 benchmarks
NELL (Never Ending Language Learning)
NELL is a dataset built from the Web via an intelligent agent called Never-Ending Language Learner.
177 papers · 2 benchmarks
The PAMAP2 Physical Activity Monitoring dataset contains data of 18 different physical activities (such as walking, cycling, playing soccer, etc.), performed by 9 subjects wearing 3 inertial measurement units and a heart rate monitor.
176 papers · 1 benchmark
NExT-QA is a VideoQA benchmark targeting the explanation of video contents.
174 papers · 3 benchmarks
PAWS-X contains 23,659 human translated PAWS evaluation pairs and 296,406 machine translated training pairs in six typologically distinct languages: French, Spanish, German, Chinese, Japanese, and Korean.
172 papers · 0 benchmarks
ATOMIC is an atlas of everyday commonsense reasoning, organized through 877k textual descriptions of inferential knowledge.
170 papers · 0 benchmarks
SentEval is a toolkit for evaluating the quality of universal sentence representations.
168 papers · 1 benchmark
MLQA (MultiLingual Question Answering)
MLQA (MultiLingual Question Answering) is a benchmark dataset for evaluating cross-lingual question answering performance.
167 papers · 1 benchmark
COD10K (Camouflaged/Concealed Object Detection)
Sensory ecologists have found that this s background matching camouflage strategy works by deceiving the visual perceptual system of the observer.
166 papers · 2 benchmarks
SWAG (Situations With Adversarial Generations)
Given a partial description like "she opened the hood of the car," humans can reason about the situation and anticipate what might come next ("then, she examined the engine").
163 papers · 2 benchmarks
This paper introduces the pipeline to scale the largest dataset in egocentric vision EPIC-KITCHENS.
162 papers · 6 benchmarks
MultiRC (Multi-Sentence Reading Comprehension)
MultiRC (Multi-Sentence Reading Comprehension) is a dataset of short paragraphs and multi-sentence questions, i.e., questions that can be answered by combining information from multiple sentences of the paragraph.
162 papers · 1 benchmark
PAWS (Paraphrase Adversaries from Word Scrambling)
Paraphrase Adversaries from Word Scrambling (PAWS) is a dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of…
159 papers · 0 benchmarks
VisDial (Visual Dialog)
Visual Dialog (VisDial) dataset contains human annotated questions based on images of MS COCO dataset.
159 papers · 6 benchmarks
WSJ0-2mix is a speech recognition corpus of speech mixtures using utterances from the Wall Street Journal (WSJ0) corpus.
159 papers · 3 benchmarks
Civil Comments (Jigsaw Unintended Bias in Toxicity Classification)
At the end of 2017 the Civil Comments platform shut down and chose make their ~2m public comments from their platform available in a lasting open archive so that researchers could understand and improve civility in online conversations for…
156 papers · 1 benchmark
A large-scale natural dataset in English to measure stereotypical biases in four domains: gender, profession, race, and religion.
156 papers · 1 benchmark
Total-Text is a text detection dataset that consists of 1,555 images with a variety of text types including horizontal, multi-oriented, and curved text instances.
156 papers · 2 benchmarks
DocRED (Document-Level Relation Extraction Dataset) is a relation extraction dataset constructed from Wikipedia and Wikidata.
155 papers · 4 benchmarks
MIND (MIcrosoft News Dataset)
MIcrosoft News Dataset (MIND) is a large-scale dataset for news recommendation research.
155 papers · 0 benchmarks
SID (See-in-the-Dark)
The See-in-the-Dark (SID) dataset contains 5094 raw short-exposure images, each with a corresponding long-exposure reference image.
155 papers · 3 benchmarks
CC12M (Conceptual 12M)
Conceptual 12M (CC12M) is a dataset with 12 million image-text pairs specifically meant to be used for vision-and-language pre-training.
153 papers · 0 benchmarks
OLID (Offensive Language Identification Dataset)
The OLID is a hierarchical dataset to identify the type and the target of offensive texts in social media.
152 papers · 1 benchmark
ROCStories is a collection of commonsense short stories.
152 papers · 2 benchmarks
FCE (First Certificate in English)
The Cambridge Learner Corpus First Certificate in English (CLC FCE) dataset consists of short texts, written by learners of English as an additional language in response to exam prompts eliciting free-text answers and assessing mastery of…
151 papers · 1 benchmark
SQuAD (Stanford Question Answering Dataset)
The Stanford Question Answering Dataset (SQuAD) is a collection of question-answer pairs derived from Wikipedia articles.
151 papers · 12 benchmarks
ASDiv (Academia Sinica Diverse MWP Dataset)
We present ASDiv (Academia Sinica Diverse MWP Dataset), a diverse (in terms of both language patterns and problem types) English math word problem (MWP) corpus for evaluating the capability of various MWP solvers.
150 papers · 1 benchmark
The WebNLG corpus comprises of sets of triplets describing facts (entities and relations between them) and the corresponding facts in form of natural language text.
149 papers · 17 benchmarks
The ActivityNet-QA dataset contains 58,000 human-annotated QA pairs on 5,800 videos derived from the popular ActivityNet dataset.
146 papers · 2 benchmarks
The SumMe dataset is a video summarization dataset consisting of 25 videos, each annotated with at least 15 human summaries (390 in total).
146 papers · 3 benchmarks
The TVQA dataset is a large-scale video dataset for video question answering.
146 papers · 3 benchmarks
VQA-RAD (Visual Question Answering in Radiology)
VQA-RAD consists of 3,515 question–answer pairs on 315 radiology images.
145 papers · 0 benchmarks
Wizard of Wikipedia is a large dataset with conversations directly grounded with knowledge retrieved from Wikipedia.
145 papers · 1 benchmark
The One Billion Word dataset is a dataset for language modeling.
141 papers · 2 benchmarks
CAMO (Camouflaged Object)
Camouflaged Object (CAMO) dataset specifically designed for the task of camouflaged object segmentation.
139 papers · 2 benchmarks
e-SNLI is used for various goals, such as obtaining full sentence justifications of a model's decisions, improving universal sentence representations and transferring to out-of-domain NLI datasets.
139 papers · 1 benchmark
SEED-Bench consists of 19K multiple choice questions with accurate human annotations (~6 larger than existing benchmarks), which spans 12 evaluation dimensions including the comprehension of both the image and video modality.
137 papers · 0 benchmarks
Cholec80 (Surgical Workflow Dataset)
Cholec80 is an endoscopic video dataset containing 80 videos of cholecystectomy surgeries performed by 13 surgeons.
134 papers · 2 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.