Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 10 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 433–480 of 3,998

The MIT-BIH Arrhythmia Database contains 48 half-hour excerpts of two-channel ambulatory ECG recordings, obtained from 47 subjects studied by the BIH Arrhythmia Laboratory between 1975 and 1979.
31 papers · 5 benchmarks
MMC4 (Multimodal C4)
Multimodal C4 (MMC4) is an augmentation of the popular text-only c4 corpus with images interleaved.
31 papers · 0 benchmarks
OpinionQA is a dataset for evaluating the alignment of LM opinions with those of 60 US demographic groups over topics ranging from abortion to automation.
31 papers · 0 benchmarks
The SemEval-2013 Task 2 dataset contains data for two subtasks: A, an expression-level subtask, and B, a message-level subtask.
31 papers · 0 benchmarks
TIMIT (TIMIT Acoustic-Phonetic Continuous Speech Corpus)
The TIMIT Acoustic-Phonetic Continuous Speech Corpus is a standard dataset used for evaluation of automatic speech recognition systems.
31 papers · 6 benchmarks
AGQA (Action Genome Question Answering)
Action Genome Question Answering (AGQA) is a benchmark for compositional spatio-temporal reasoning.
30 papers · 0 benchmarks
A testbed for commonsense reasoning about entity knowledge, bridging fact-checking about entities with commonsense inferences.
30 papers · 0 benchmarks
ECtHR (European Court of Human Rights Cases)
ECtHR is a dataset comprising European Court of Human Rights cases, including annotations for paragraph-level rationales.
30 papers · 0 benchmarks
EntityQuestions is a dataset of simple, entity-rich questions based on facts from Wikidata (e.g., "Where was Arve Furset born?
30 papers · 1 benchmark
🤖 Robo3D - The KITTI-C Benchmark KITTI-C is an evaluation benchmark heading toward robust and reliable 3D object detection in autonomous driving.
30 papers · 1 benchmark
VALUE (Video-And-Language Understanding Evaluation)
VALUE is a Video-And-Language Understanding Evaluation benchmark to test models that are generalizable to diverse tasks, domains, and datasets.
30 papers · 0 benchmarks
MassiveText is a collection of large English-language text datasets from multiple sources: web pages, books, news articles, and code.
29 papers · 0 benchmarks
Omni-Realm Benchmark (OmniBenchmark) is a diverse (21 semantic realm-wise datasets) and concise (realm-wise datasets have no concepts overlapping) benchmark for evaluating pre-trained model generalization across semantic…
29 papers · 1 benchmark
WI-LOCNESS (Cambridge English Write & Improve & LOCNESS)
WI-LOCNESS is part of the Building Educational Applications 2019 Shared Task for Grammatical Error Correction.
29 papers · 2 benchmarks
ADNI (Alzheimer's Disease NeuroImaging Initiative)
Alzheimer's Disease Neuroimaging Initiative (ADNI) is a multisite study that aims to improve clinical trials for the prevention and treatment of Alzheimer’s disease (AD).[1] This cooperative study combines expertise and funding from the…
28 papers · 5 benchmarks
Semi-Supervised Object Detection on COCO 10% labeled data
28 papers · 2 benchmarks
EVALution dataset is evenly distributed among the three classes (hypernyms, co-hyponyms and random) and involves three types of parts of speech (noun, verb, adjective).
28 papers · 0 benchmarks
IndicCorp is a large monolingual corpora with around 9 billion tokens covering 12 of the major Indian languages.
28 papers · 0 benchmarks
KAIST (High-quality hyperspectral reconstruction using a spectral prior)
High-quality hyperspectral reconstruction using a spectral prior
28 papers · 1 benchmark
KITTI MOTS (KITTI Multi-Object Tracking and Segmentation (MOTS) Evaluation)
The Multi-Object and Segmentation (MOTS) benchmark [2] consists of 21 training sequences and 29 test sequences.
28 papers · 1 benchmark
KaggleDBQA (KaggleDBQA: Realistic Text-to-SQL dataset)
KaggleDBQA is a challenging cross-domain and complex evaluation dataset of real Web databases, with domain-specific data types, original formatting, and unrestricted questions.
28 papers · 1 benchmark
Dataset is constructed from single intent dataset ATIS.
28 papers · 2 benchmarks
PSG dataset has 48749 images with 133 object classes (80 objects and 53 stuff) and 56 predicate classes.
28 papers · 1 benchmark
Resume contains eight fine-grained entity categories -score from 74.5% to 86.88%.
28 papers · 1 benchmark
THuman2.0 Dataset contains 500 high-quality human scans captured by a dense DLSR rig.
28 papers · 1 benchmark
WikiCoref is an English corpus annotated for anaphoric relations, where all documents are from the English version of Wikipedia.
28 papers · 1 benchmark
The WoZ 2.0 dataset is a newer dialogue state tracking dataset whose evaluation is detached from the noisy output of speech recognition systems.
28 papers · 1 benchmark
XM 3600 (Crossmodal 3600)
Research in massively multilingual image captioning has been severely hampered by a lack of high-quality evaluation datasets.
28 papers · 0 benchmarks
decaNLP (Natural Language Decathlon Benchmark)
Natural Language Decathlon Benchmark (decaNLP) is a challenge that spans ten tasks: question answering, machine translation, summarization, natural language inference, sentiment analysis, semantic role labeling, zero-shot relation…
28 papers · 0 benchmarks
The 20 Newsgroups data set is a collection of approximately 20,000 newsgroup documents, partitioned (nearly) evenly across 20 different newsgroups.
27 papers · 5 benchmarks
AitW (Android in the Wild)
Android in the Wild (AitW) is a dataset for device-control research which is orders of magnitude larger than current datasets.
27 papers · 0 benchmarks
CONAN (COunter NArratives through Nichesourcing)
COunter NArratives through Nichesourcing (CONAN) is a dataset that consists of 4,078 pairs over the 3 languages.
27 papers · 0 benchmarks
CaseHOLD (Case Holdings On Legal Decisions)
CaseHOLD (Case Holdings On Legal Decisions) is a law dataset comprised of over 53,000+ multiple choice questions to identify the relevant holding of a cited case.
27 papers · 2 benchmarks
Chaos NLI is a Natural Language Inference (NLI) dataset with 100 annotations per example (for a total of 464,500 annotations) for some existing data points in the development sets of SNLI, MNLI, and Abductive NLI.
27 papers · 0 benchmarks
Evidence Inference is a corpus for this task comprising 10,000+ prompts coupled with full-text articles describing RCTs.
27 papers · 0 benchmarks
GID (Gaofen Image Dataset)
Gaofen Image Dataset (GID) is a large-scale land-cover dataset constructed with Gaofen-2 (GF-2) satellite images.
27 papers · 0 benchmarks
LDC2017T10 (Abstract Meaning Representation (AMR) Annotation Release 2.0)
Abstract Meaning Representation (AMR) Annotation Release 2.0 was developed by the Linguistic Data Consortium (LDC), SDL/Language Weaver, Inc., the University of Colorado's Computational Language and Educational Research group and the…
27 papers · 1 benchmark
LHQ (Landscapes High-Quality)
A dataset of 90,000 high-resolution nature landscape images, crawled from Unsplash and Flickr and preprocessed with Mask R-CNN and Inception V3.
27 papers · 4 benchmarks
Multi-Modal-CelebA-HQ is a large-scale face image dataset that has 30,000 high-resolution face images selected from the CelebA dataset by following CelebA-HQ.
27 papers · 3 benchmarks
ProofNet is a benchmark for autoformalization and formal proving of undergraduate-level mathematics.
27 papers · 0 benchmarks
RAVDESS (Ryerson Audio-Visual Database of Emotional Speech and Song)
The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) contains 7,356 files (total size: 24.8 GB).
27 papers · 6 benchmarks
Sentiment analysis of codemixed tweets.
27 papers · 0 benchmarks
UVO (Unidentified Video Objects: A Benchmark for Dense, Open-World Segmentation)
UVO is a new benchmark for open-world class-agnostic object segmentation in videos.
27 papers · 2 benchmarks
We have created three new Reading Comprehension datasets constructed using an adversarial model-in-the-loop.
26 papers · 2 benchmarks
Animal Kingdom is a large and diverse dataset that provides multiple annotated tasks to enable a more thorough understanding of natural animal behaviors.
26 papers · 2 benchmarks
CVSS is a massively multilingual-to-English speech to speech translation (S2ST) corpus, covering sentence-level parallel S2ST pairs from 21 languages into English.
26 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.