Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 2 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 49–96 of 3,998

Objaverse is a large dataset of objects with 800K+ (and growing) 3D models with descriptive captions, tags, and animations.
393 papers · 2 benchmarks
Speech Commands is an audio dataset of spoken words designed to help train and evaluate keyword spotting systems .
392 papers · 4 benchmarks
DROP (Discrete Reasoning Over Paragraphs)
Discrete Reasoning Over Paragraphs DROP is a crowdsourced, adversarially-created, 96k-question benchmark, in which a system must resolve references in a question, perhaps to multiple input positions, and perform discrete operations over…
382 papers · 3 benchmarks
OK-VQA (Outside Knowledge Visual Question Answering)
Outside Knowledge Visual Question Answering (OK-VQA) includes more than 14,000 questions that require external knowledge to answer.
368 papers · 2 benchmarks
SVAMP (Simple Variations on Arithmetic Math word Problems)
A challenge set for elementary-level Math Word Problems (MWP).
362 papers · 2 benchmarks
WSC (Winograd Schema Challenge)
The Winograd Schema Challenge was introduced both as an alternative to the Turing Test and as a test of a system’s ability to do commonsense reasoning.
361 papers · 2 benchmarks
Automatic image captioning is the task of producing a natural-language utterance (usually a sentence) that correctly reflects the visual content of an image.
352 papers · 2 benchmarks
BIG-bench (Beyond the Imitation Game Benchmark)
The Beyond the Imitation Game Benchmark (BIG-bench) is a collaborative benchmark intended to probe large language models and extrapolate their future capabilities.
349 papers · 121 benchmarks
BookCorpus is a large collection of free novel books written by unpublished authors, which contains 11,038 books (around 74M sentences and 1G words) of 16 different sub-genres (e.g., Romance, Historical, Adventure, etc.).
344 papers · 1 benchmark
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
339 papers · 2 benchmarks
ScienceQA (Science Question Answering)
Science Question Answering (ScienceQA) is a new benchmark that consists of 21,208 multimodal multiple choice questions with diverse science topics and annotations of their answers with corresponding lectures and explanations.
339 papers · 1 benchmark
COPA (Choice of Plausible Alternatives)
The Choice Of Plausible Alternatives (COPA) evaluation provides researchers with a tool for assessing progress in open-domain commonsense causal reasoning.
329 papers · 1 benchmark
MultiWOZ (Multi-domain Wizard-of-Oz)
The Multi-domain Wizard-of-Oz (MultiWOZ) dataset is a large-scale human-human conversational corpus spanning over seven domains, containing 8438 multi-turn dialogues, with each dialogue averaging 14 turns.
328 papers · 8 benchmarks
AffectNet (burak yılmaz)
AffectNet is a large facial expression dataset with around 0.4 million images manually labeled for the presence of eight (neutral, happy, angry, sad, fear, surprise, disgust, contempt) facial expressions along with the intensity of valence…
323 papers · 4 benchmarks
LJSpeech (The LJ Speech Dataset)
This is a public domain speech dataset consisting of 13,100 short audio clips of a single speaker reading passages from 7 non-fiction books.
323 papers · 2 benchmarks
ETT (Electricity Transformer Temperature)
The Electricity Transformer Temperature (ETT) is a crucial indicator in the electric power long-term deployment.
321 papers · 20 benchmarks
MPQA Opinion Corpus (Multi-Perspective Question Answering)
The MPQA Opinion Corpus contains 535 news articles from a wide variety of news sources manually annotated for opinions and other private states (i.e., beliefs, emotions, sentiments, speculations, etc.).
313 papers · 3 benchmarks
BEIR (Benchmarking IR)
BEIR (Benchmarking IR) is a heterogeneous benchmark containing different information retrieval (IR) tasks.
311 papers · 10 benchmarks
DRIVE (Digital Retinal Images for Vessel Extraction)
The Digital Retinal Images for Vessel Extraction (DRIVE) dataset is a dataset for retinal vessel segmentation.
311 papers · 2 benchmarks
The LAMBADA (LAnguage Modeling Broadened to Account for Discourse Aspects) benchmark is an open-ended cloze task which consists of about 10,000 passages from BooksCorpus where a missing target word is predicted in the last sentence of each…
293 papers · 1 benchmark
StrategyQA is a question answering benchmark where the required reasoning steps are implicit in the question, and should be inferred using a strategy.
291 papers · 1 benchmark
MELD (Multimodal EmotionLines Dataset)
Multimodal EmotionLines Dataset (MELD) has been created by enhancing and extending EmotionLines dataset.
289 papers · 3 benchmarks
CoQA (Conversational Question Answering Challenge)
CoQA is a large-scale dataset for building Conversational Question Answering systems.
281 papers · 2 benchmarks
The NewsQA dataset is a crowd-sourced machine reading comprehension dataset of 120,000 question-answer pairs.
272 papers · 1 benchmark
EMNIST (Extended MNIST)
EMNIST (extended MNIST) has 4 times more data than MNIST.
264 papers · 10 benchmarks
The New York Times Annotated Corpus contains over 1.8 million articles written and published by the New York Times between January 1, 1987 and June 19, 2007 with article metadata provided by the New York Times Newsroom, the New York Times…
262 papers · 9 benchmarks
LibriTTS is a multi-speaker English corpus of approximately 585 hours of read English speech at 24kHz sampling rate, prepared by Heiga Zen with the assistance of Google Speech and Google Brain team members.
257 papers · 1 benchmark
OntoNotes 5.0 is a large corpus comprising various genres of text (news, conversational telephone speech, weblogs, usenet newsgroups, broadcast, talk shows) in three languages (English, Chinese, and Arabic) with structural information…
254 papers · 12 benchmarks
The ICDAR 2013 dataset consists of 229 training images and 233 testing images, with word-level annotations provided.
246 papers · 3 benchmarks
MathVista (Mathematical Reasoning of in Visual Contexts)
MathVista is a consolidated Mathematical reasoning benchmark within Visual contexts.
242 papers · 0 benchmarks
The WebQuestions dataset is a question answering dataset using Freebase as the knowledge base and contains 6,642 question-answer pairs.
241 papers · 4 benchmarks
MIMIC-CXR from Massachusetts Institute of Technology presents 371,920 chest X-rays associated with 227,943 imaging studies from 65,079 patients.
240 papers · 3 benchmarks
Omniverse Isaac Gym is a GPU-based physics simulation platform developed by NVIDIA.
240 papers · 2 benchmarks
GOT-10k (Generic Object Tracking Benchmark)
The GOT-10k dataset contains more than 10,000 video segments of real-world moving objects and over 1.5 million manually labelled bounding boxes.
239 papers · 2 benchmarks
CK+ (Extended Cohn-Kanade dataset)
The Extended Cohn-Kanade (CK+) dataset contains 593 video sequences from a total of 123 different subjects, ranging from 18 to 50 years of age with a variety of genders and heritage.
238 papers · 2 benchmarks
Wiki Squirrel (Wikipedia Squirrel)
The data was collected from the English Wikipedia (December 2018).
208 papers · 1 benchmark
The ShanghaiTech Campus dataset has 13 scenes with complex light conditions and camera angles.
207 papers · 4 benchmarks
The NarrativeQA dataset includes a list of documents with Wikipedia summaries, links to full stories, and questions and answers.
206 papers · 1 benchmark
WiC (Words in Context)
WiC is a benchmark for the evaluation of context-sensitive word embeddings.
206 papers · 3 benchmarks
Clotho is an audio captioning dataset, consisting of 4981 audio samples, and each audio sample has five captions (a total of 24 905 captions).
202 papers · 3 benchmarks
HumanML3D is a 3D human motion-language dataset that originates from a combination of HumanAct12 and Amass dataset.
201 papers · 2 benchmarks
Kvasir-SEG is an open-access dataset of gastrointestinal polyp images and corresponding segmentation masks, manually annotated by a medical doctor and then verified by an experienced gastroenterologist.
201 papers · 2 benchmarks
YouCook2 is the largest task-oriented, instructional video dataset in the vision community.
198 papers · 7 benchmarks
WikiQA (Wikipedia open-domain Question Answering)
The WikiQA corpus is a publicly available set of question and sentence pairs, collected and annotated for research on open-domain question answering.
196 papers · 2 benchmarks
Libri-Light is a collection of spoken English audio suitable for training speech recognition systems under limited or no supervision.
194 papers · 2 benchmarks
BC5CDR (BioCreative V CDR corpus)
BC5CDR corpus consists of 1500 PubMed articles with 4409 annotated chemicals, 5818 diseases and 3116 chemical-disease interactions.
191 papers · 4 benchmarks
CMU Multimodal Opinion Sentiment and Emotion Intensity (CMU-MOSEI) is the largest dataset of sentence-level sentiment analysis and emotion recognition in online videos.
190 papers · 3 benchmarks
FewRel (Few-Shot Relation Classification Dataset)
The FewRel (Few-Shot Relation Classification Dataset) contains 100 relations and 70,000 instances from Wikipedia.
189 papers · 3 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.