Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 12 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 529–576 of 3,998

GSCAN (Grounded SCAN)
Grounded SCAN poses a simple task, where an agent must execute action sequences based on a synthetic language instruction.
22 papers · 0 benchmarks
This dataset contains card descriptions of the card game Hearthstone and the code that implements them.
22 papers · 0 benchmarks
100 tasks from LIBERO-100 suite.
22 papers · 1 benchmark
Motion-X is a large-scale 3D expressive whole-body motion dataset, which comprises 15.6M precise 3D whole-body pose annotations (i.e., SMPL-X) covering 81.1K motion sequences from massive scenes, meanwhile providing corresponding semantic…
22 papers · 1 benchmark
PIT (Paraphrase and Semantic Similarity in Twitter)
Paraphrase and Semantic Similarity in Twitter (PIT) presents a constructed Twitter Paraphrase Corpus that contains 18,762 sentence pairs.
22 papers · 1 benchmark
The ECGs in this collection were obtained using a non-commercial, PTB prototype recorder with the following specifications: 16 input channels, (14 for ECGs, 1 for respiration, 1 for line voltage) Input voltage: ±16 mV, compensated offset…
22 papers · 4 benchmarks
Satlas is a remote sensing dataset and benchmark that is large in both breadth, featuring all of the aforementioned applications and more, as well as scale, comprising 290M labels under 137 categories and 7 label modalities.
22 papers · 0 benchmarks
Desc: About of Text8
22 papers · 1 benchmark
TimeQA (Time-Sensitive QA)
This dataset is aimed to study the existing reading comprehension models' capability to perform temporal reasoning, and see whether they are sensitive to the temporal description in the given question.
22 papers · 0 benchmarks
TopiOCQA (pronounced Tapioca) is an open-domain conversational dataset with topic switches on Wikipedia.
22 papers · 0 benchmarks
WebSRC (WebSRC: A Dataset for Web-Based Structural Reading Comprehension)
WebSRC is a novel Web-based Structural Reading Comprehension dataset.
22 papers · 2 benchmarks
XGLUE is an evaluation benchmark XGLUE,which is composed of 11 tasks that span 19 languages.
22 papers · 2 benchmarks
XLCoST (Cross-Lingual Code Snippet)
XLCoST is a benchmark dataset for cross-lingual code intelligence.
22 papers · 0 benchmarks
iSarcasmEval is the first shared task to target intended sarcasm detection: the data for this task was provided and labelled by the authors of the texts themselves.
22 papers · 0 benchmarks
iVQA (Instructional Video Question Answering)
An open-ended VideoQA benchmark that aims to: i) provide a well-defined evaluation by including five correct answer annotations per question and ii) avoid questions which can be answered without the video.
22 papers · 2 benchmarks
CDD-11 (Composite Degradation Dataset 11)
An image restoration dataset
21 papers · 1 benchmark
Consists of 20k English biomedical entity mentions from Reddit expert-annotated with links to SNOMED CT, a widely-used medical knowledge graph.
21 papers · 0 benchmarks
COVID-Fact is a FEVER-like dataset of claims concerning the COVID-19 pandemic.
21 papers · 0 benchmarks
CliCR is a new dataset for domain specific reading comprehension used to construct around 100,000 cloze queries from clinical case reports.
21 papers · 1 benchmark
Event2Mind is a corpus of 25,000 event phrases covering a diverse range of everyday events and situations.
21 papers · 2 benchmarks
ICBHI Respiratory Sound Database (The Respiratory Sound database - ICBHI 2017 Challenge)
The Respiratory Sound database was originally compiled to support the scientific challenge organized at Int.
21 papers · 2 benchmarks
SEP-28k (Stuttering Events in Podcasts)
Stuttering Events in Podcasts (SEP-28k) is a dataset containing over 28k clips labeled with five event types including blocks, prolongations, sound repetitions, word repetitions, and interjections.
21 papers · 0 benchmarks
SMID (Seeing motion in the dark)
This is the low-light image enhancement dataset collected by the CVPR 2018 paper "Seeing Motion in the Dark".
21 papers · 1 benchmark
STREUSLE stands for Supersense-Tagged Repository of English with a Unified Semantics for Lexical Expressions.
21 papers · 1 benchmark
The Terms of Service dataset is a law dataset corresponding to the task of identifying whether contractual terms are potentially unfair.
21 papers · 1 benchmark
The Amazon-Google dataset for entity resolution derives from the online retailers Amazon.com and the product search service of Google accessible through the Google Base Data API.
20 papers · 2 benchmarks
CLOTH (CLOze test by TeacHers)
The Cloze Test by Teachers (CLOTH) benchmark is a collection of nearly 100,000 4-way multiple-choice cloze-style questions from middle- and high school-level English language exams, where the answer fills a blank in a given text.
20 papers · 0 benchmarks
Node classification on Chameleon with the fixed 48%/32%/20% splits provided by Geom-GCN.
20 papers · 2 benchmarks
We introduce an object detection dataset in challenging adverse weather conditions covering 12000 samples in real-world driving scenes and 1500 samples in controlled weather conditions within a fog chamber.
20 papers · 2 benchmarks
ETHOS (multi-labEl haTe speecH detectiOn dataSet)
ETHOS is a hate speech detection dataset.
20 papers · 2 benchmarks
FMB Dataset (Full-time Multi-modality Benchmark Dataset)
FMB contains 1500 well-registered infrared and visible image pairs with 14 annotated pixel-level categories.
20 papers · 1 benchmark
The George Washington dataset contains 20 pages of letters written by George Washington and his associates in 1755 and thereby categorized into historical collection.
20 papers · 0 benchmarks
MINTAKA is a complex, natural, and multilingual dataset designed for experimenting with end-to-end question-answering models.
20 papers · 0 benchmarks
MLQE-PE (Multilingual Quality Estimation and Automatic Post-editing Dataset)
The Multilingual Quality Estimation and Automatic Post-editing (MLQE-PE) Dataset is a dataset for Machine Translation (MT) Quality Estimation (QE) and Automatic Post-Editing (APE).
20 papers · 0 benchmarks
Paralex learns from a collection of 18 million question-paraphrase pairs scraped from WikiAnswers.
20 papers · 1 benchmark
PhotoChat, the first dataset that casts light on the photo sharing behavior in online messaging.
20 papers · 2 benchmarks
Fact-checking (FC) articles which contains pairs (multimodal tweet and a FC-article) from politifact.com.
20 papers · 1 benchmark
A large-scale English dataset for coreference resolution.
20 papers · 1 benchmark
20 real low-resolution images selected from existing datasets or downloaded from internet
20 papers · 0 benchmarks
The Abt-Buy dataset for entity resolution derives from the online retailers Abt.com and Buy.com.
19 papers · 2 benchmarks
BCI (Breast Cancer Immunohistochemical Image Generation)
The evaluation of human epidermal growth factor receptor 2 (HER2) expression is essential to formulate a precise treatment for breast cancer.
19 papers · 1 benchmark
Chaoyang dataset contains 1111 normal, 842 serrated, 1404 adenocarcinoma, 664 adenoma, and 705 normal, 321 serrated, 840 adenocarcinoma, 273 adenoma samples for training and testing, respectively.
19 papers · 2 benchmarks
DDXPlus (DDXPlus: A New Dataset For Automatic Medical Diagnosis)
There has been a rapidly growing interest in Automatic Symptom Detection (ASD) and Automatic Diagnosis (AD) systems in the machine learning research literature, aiming to assist doctors in telemedicine services.
19 papers · 0 benchmarks
DIRHA (Distant-speech Interaction for Robust Home Applications)
DIRHA-English is a multi-microphone database composed of real and simulated sequences of 1-minute.
19 papers · 1 benchmark
Node classification on Deezer Europe with 50%/25%/25% random splits for training/validation/test.
19 papers · 1 benchmark
FNC-1 (Fake News Challenge Stage 1)
FNC-1 was designed as a stance detection dataset and it contains 75,385 labeled headline and article pairs.
19 papers · 2 benchmarks
Node classification on Film with 60%/20%/20% random splits for training/validation/test.
19 papers · 1 benchmark
FoCus (Call for Customized Conversation: Customized Conversation Grounding Persona and Knowledge)
We introduce a new dataset, called FoCus, that supports knowledge-grounded answers that reflect user’s persona.
19 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.